Reliability and validity are the two yardsticks psychologists use to judge whether a measurement or study can be trusted. Reliability asks whether a result would come out the same way again, while validity asks whether it is measuring the right thing in the first place.
Key Takeaways
- Reliability in research refers to the consistency and reproducibility of measurements. It assesses the degree to which a measurement tool produces stable and dependable results when used repeatedly under the same conditions.
- Validity in research refers to the accuracy and meaningfulness of measurements. It examines whether a research instrument or method effectively measures what it claims to measure.
- A reliable instrument may not necessarily be valid, as it might consistently measure something other than the intended concept.
| Reliability | Validity | |
|---|---|---|
| What does it tell you? | The extent to which the results can be reproduced when the research is repeated under the same conditions. | The extent to which the results really measure what they are supposed to measure. |
| How is it assessed? | By checking the consistency of results across time, across different observers, and across parts of the test itself. | By checking how well the results correspond to established theories and other measures of the same concept. |
| How do they relate? | A reliable measurement is not always valid: the results might be reproducible, but they’re not necessarily correct. | A valid measurement is generally reliable: if a test produces accurate results, they should be reproducible. |

While reliability is a prerequisite for validity, it does not guarantee it.
A reliable measure might consistently produce the same result, but that result may not accurately reflect the true value.
For instance, a thermometer could consistently give the same temperature reading, but if it is not calibrated correctly, the measurement would be reliable but not valid.
Assessing Validity
A valid measurement accurately reflects the underlying concept being studied.
For example, a valid intelligence test would accurately assess an individual’s cognitive abilities, while a valid measure of depression would accurately reflect the severity of a person’s depressive symptoms.
Quantitative validity can be assessed through various forms, such as content validity (expert review), criterion validity (comparison with a gold standard), and construct validity (measuring the underlying theoretical construct).
Content Validity
Content validity refers to the extent to which a psychological instrument accurately and fully reflects all the features of the concept being measured.
Content validity is a fundamental consideration in psychometrics, ensuring that a test measures what it purports to measure.
Content validity is not merely about a test appearing valid on the surface, which is face validity. The two ideas are easy to confuse.
Instead, it goes deeper, requiring a systematic and rigorous evaluation of the test content by subject matter experts.
For example, a company using a personality test to screen job applicants needs strong content validity. The items must reflect job-relevant traits.
Content validity is often assessed through expert review, where subject matter experts evaluate the relevance and completeness of the test items.
Criterion Validity
Criterion validity examines how well a measurement tool corresponds to other valid measures of the same concept.
It includes concurrent validity (existing criteria) and predictive validity (future outcomes).
For example, a researcher measuring depression with a self-report inventory can establish criterion validity by checking whether scores correlate with external indicators of depression. These include clinician ratings, missed workdays, or hospital stays.
Criterion validity is important because, without it, tests would not be able to accurately measure in a way consistent with other validated instruments.
Construct Validity
Construct validity assesses how well a particular measurement reflects the theoretical construct (existing theory and knowledge) it is intended to measure.
It goes beyond simply assessing whether a test covers the right material or predicts specific outcomes.
Instead, it focuses on the meaning of the test scores. How do they relate to the underlying theoretical framework?
For instance, a researcher developing a new questionnaire to evaluate aggression must ask whether it truly measures aggression. Or does it just capture assertiveness or dominance instead?
Assessing construct validity involves multiple methods and often relies on the accumulation of evidence over time.
Assessing Reliability
Reliability refers to the consistency and stability of measurement results.
In simpler terms, a reliable tool produces consistent results when applied repeatedly under the same conditions.
Test-Retest Reliability
This method assesses the stability of a measure over time.
Timing is the critical design choice.
The same test is administered to the same group twice, with a reasonable time interval between tests.
The correlation coefficient between the two sets of scores represents the reliability coefficient.
A higher correlation means a more stable measure.
A high correlation indicates that individuals maintain their relative positions within the group despite potential overall shifts in performance.
Rank order is what matters here.
For example, a researcher administers a depression screening test to 100 participants. Two weeks later, they give the exact same test to the same people.
Comparing scores between Time 1 and Time 2 reveals a correlation of 0.85, indicating good test-retest reliability since the scores remained stable over time.
Not every weak correlation is the test’s fault.
The trait itself may simply have changed between sessions rather than the measure being unreliable (Guttman, 1945). Good studies rule out real change with independent evidence that the trait should stay stable over the chosen interval.
Factors influencing test-retest reliability:
- Memory effects: If respondents remember their answers from the first testing, it could artificially inflate the reliability coefficient.
- Time interval between testings: A short interval might lead to inflated reliability due to memory effects, while an excessively long interval increases the chance of genuine changes in the trait being measured.
- Test length and nature of test materials: These can also affect the likelihood of respondents remembering their previous answers.
- Stability of the trait being measured: If the trait itself is unstable and subject to change, test-retest reliability might be low even with a reliable measure.
Interrater Reliability
Interrater reliability assesses the consistency or agreement among judgments made by different raters or observers.
Multiple raters independently assess the same set of targets, and the consistency of their judgments is evaluated.
Training alone will not fix every disagreement.
Adequate training equips raters with the necessary knowledge and skills to apply scoring criteria consistently, reducing systematic errors.
A high interrater reliability indicates that the raters are interchangeable and the rating protocol is reliable.
But how interchangeable, exactly? Raw agreement can be misleading, because two raters can agree simply by chance. A guess would still agree half the time.
Researchers correct for this with Cohen’s kappa (Cohen, 1960), a statistic that adjusts agreement for chance. Zero means chance; one means perfect.
Landis and Koch (1977) proposed interpretive bands for the resulting score. Values of .00 to .20 count as slight agreement, and .21 to .40 as fair. From there, .41 to .60 is moderate, .61 to .80 substantial, and .81 to 1.00 almost perfect.
The bands are only a guide. Even a “substantial” kappa can be too low for a single high-stakes decision, such as a diagnosis.
For example:
- Research: Evaluating the consistency of coding in qualitative research or assessing the agreement among raters evaluating behaviors in observational studies.
- Education: Determining the degree of agreement among teachers grading essays or other subjective assignments.
- Clinical Settings: Evaluating the consistency of diagnoses made by different clinicians based on the same patient information.
Internal Consistency
Internal consistency refers to the consistency of measurement itself. It examines the degree to which different items within a test or scale are measuring the same underlying construct.
For instance, consider a test designed to assess self-esteem.
If the items within the test are internally consistent, individuals with high self-esteem should generally score highly on all or most of the items. Conversely, those with low self-esteem should consistently score lower on those same items.
The pattern should hold across the whole test.
While internal consistency is a necessary condition for validity, it does not guarantee it. A measure can be internally consistent but still not accurately measure the intended construct.
How high should alpha be? Conventionally, an alpha of .70 or above (Cronbach, 1951) is treated as adequate, and anything above .90 counts as high internal consistency.
But a high alpha is not a quality seal. Sijtsma (2009) showed that alpha is best read as a lower-bound estimate of reliability, one that assumes every item measures the trait with equal precision.
A scale with many items can clear .90 even when it is secretly measuring more than one thing. The safer practice is to report alpha alongside a check of the test’s factor structure, rather than treating a large alpha as proof of a good scale.
Methods for estimating internal consistency:
- Split-half reliability divides a test into two parts (such as odd and even number items) and correlates their scores to check consistency.
- Cronbach’s alpha (α) is the most widely used measure of internal consistency. It represents the average of all possible split-half reliability coefficients that could be computed from the test.
Ensuring Validity
- Define concepts clearly: Start with a clear and precise definition of the concepts you want to measure. This clarity will guide the selection or development of appropriate measurement instruments.
- Use established measures: Whenever possible, use well-established and validated measures that are reliable and valid in previous research. If adapting a measure from a different culture or language, carefully translate and validate it for the target population.
- Pilot test instruments: Before conducting the main study, pilot test your measurement instruments with a smaller sample to identify potential issues with wording, clarity, or response options.
- Use multiple measures (mixed methods): Employing multiple methods of data collection (e.g., interviews, observations, surveys) or data analysis can enhance the validity of the findings by providing converging evidence from different sources.
- Address potential biases: Carefully consider factors that could introduce bias into the research, such as sampling methods, data collection procedures, or the researcher’s own preconceptions.
Ensuring Reliability
- Standardize procedures: Establish clear and consistent procedures for data collection, scoring, and analysis. This standardization helps minimize variability due to procedural inconsistencies.
- Train observers or raters: If using multiple raters, provide thorough training to ensure they understand the rating scales, criteria, and procedures. This training enhances interrater reliability by reducing subjective variations in judgments.
- Optimize measurement conditions: Create a controlled and consistent environment for data collection to minimize external factors that could influence responses. For example, ensure participants have adequate privacy, time, and clear instructions.
- Use reliable instruments: Select or develop measurement instruments that have demonstrated good internal consistency reliability, such as a high Cronbach’s alpha coefficient. Address potential issues with reverse-coded items or item heterogeneity that can affect internal consistency.
How should I report validity and reliability in my research?
- Introduction: Discuss previous research on the validity and reliability of the chosen measures, highlighting any limitations or considerations.
- Methodology: Detail the steps taken to ensure validity and reliability, including the measures used, sampling methods, data collection procedures, and steps to address potential biases.
- Results: Report the reliability coefficients obtained (e.g., Cronbach’s alpha, Cohen’s Kappa) and discuss their implications for the study’s findings.
- Discussion: Critically evaluate the validity and reliability of the findings, acknowledging any limitations or areas for improvement.
Validity and Reliability in Qualitative and Quantitative Research
While both qualitative and quantitative research strive to produce credible and trustworthy findings, their approaches to ensuring reliability and validity differ.
Qualitative research emphasizes the richness and depth of understanding, and quantitative research focuses on measurement precision and statistical analysis.
Qualitative Research
While traditional quantitative notions of reliability and validity may not always directly apply, qualitative researchers emphasize trustworthiness and transferability.
The three ideas work together.
Credibility refers to the confidence in the truth and accuracy of the findings, often enhanced through prolonged engagement, persistent observation, and triangulation.
Transferability involves providing rich descriptions of the research context to allow readers to determine the applicability of the findings to other settings.
Confirmability is the degree to which the findings are shaped by the participants’ experiences rather than the researcher’s biases, often addressed through reflexivity and audit trails.
They focus on establishing confidence in the findings by:
- Triangulating data from multiple sources.
- Member checking, allowing participants to verify the interpretations.
- Providing thick, rich descriptions to enhance transferability to other contexts.
Quantitative Research
Quantitative research typically relies more heavily on statistical measures of reliability (e.g., Cronbach’s alpha, test-retest correlations) and validity (e.g., factor analysis, correlations with criterion measures).
The goal is to demonstrate that the measures are consistent, accurate, and meaningfully related to the concepts they are intended to assess.
Critical Evaluation
Reliability and validity are powerful tools for judging research, but neither is a fixed, one-off box to tick. Both come with real strengths and real limitations.
Strengths
- Quantifiable: Reliability and validity can be measured with specific numbers, such as agreement correlations, kappa, or Cronbach’s alpha, rather than simply asserted.
- Actionable checklist: The frameworks give researchers concrete steps to follow, including operational definitions, a standardized procedure, rater training, and adequate sampling of the behavior being studied.
- A precise mathematical link: Spearman (1904) showed that the correlation between two measures can never exceed the square root of the product of their reliabilities, so poor reliability in either measure caps how strongly they can agree.
Limitations
- Reliability is not enough: A perfectly reliable measure can still be worthless, the way a scale that always reads two kilograms too heavy gives identical, wrong readings every time.
- Reliability versus ecological validity: Tightly standardized procedures are easier to make reliable, but standardization can strip away the real-world complexity a task is meant to represent.
- Reliability is not fixed: A measure that is reliable in one sample or setting is not automatically reliable in another; it has to be re-checked whenever the sample, setting, or purpose changes.
- Qualitative research needs different criteria: Interpreting rich, context-dependent material does not fit the classical reliability framework, so qualitative work is judged on credibility, coherence, and researcher reflexivity instead.
Contemporary Research
A single concern runs through the last decade of research on reliability. Psychology has often treated it as a box to tick once in a test manual, rather than a property that must be checked for the sample being tested.
Parsons, Kruijt, and Fox (2019) make the case directly.
Aim: To establish how rarely reliability is actually reported for common cognitive-behavioral tasks, and what that reliability looks like when it is estimated. It also asked what the statistical consequences are of proceeding without that information.
Thousands of dot-probe studies exist.
Method: The authors combined a literature audit of the widely used dot-probe attention-bias task with a worked re-analysis of a public Stroop-task dataset. They computed a split-half internal-consistency estimate and a test-retest intraclass correlation.
The numbers were not reassuring.
Results: Reported dot-probe reliability, where it existed at all, ranged from near zero to about .70. The Stroop re-analysis found split-half correlations of roughly .5-.6 and test-retest correlations of .6-.7, respectable at the group level but modest for studying individual differences.
That gap in power is easy to miss.
Conclusion: Reliability is not a fixed property that can be looked up once. It belongs to a specific measurement, in a specific sample, and should be estimated and reported every time a task is used.
The same theme appears elsewhere too.
This builds on what Hedge, Powell, and Sumner (2018) call the reliability paradox. Classic experimental tasks can produce large, highly replicable group-level effects while still having very poor reliability for measuring differences between individuals.
The very feature that makes a task’s group effect robust, small variation between people, is the same feature that undermines its use for individual-differences research.
Loken and Gelman (2017) add a final piece. When unreliable measurement combines with selective reporting of statistically significant results, published effect sizes can become systematically inflated rather than simply weakened.
A noisy measure occasionally throws up an unusually large effect by chance, and it is precisely those inflated results that clear the bar for publication.

References
Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37-46. https://doi.org/10.1177/001316446002000104
Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297-334. https://doi.org/10.1007/BF02310555
Guttman, L. (1945). A basis for analyzing test-retest reliability. Psychometrika, 10(4), 255-282. https://doi.org/10.1007/BF02288892
Hedge, C., Powell, G., & Sumner, P. (2018). The reliability paradox: Why robust cognitive tasks do not produce reliable individual differences. Behavior Research Methods, 50(3), 1166-1186. https://doi.org/10.3758/s13428-017-0935-1
Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159-174. https://doi.org/10.2307/2529310
Loken, E., & Gelman, A. (2017). Measurement error and the replication crisis. Science, 355(6325), 584-585. https://doi.org/10.1126/science.aal3618
Parsons, S., Kruijt, A.-W., & Fox, E. (2019). Psychological science needs a standard practice of reporting the reliability of cognitive-behavioral measurements. Advances in Methods and Practices in Psychological Science, 2(4), 378-395. https://doi.org/10.1177/2515245919879695
Sijtsma, K. (2009). On the use, the misuse, and the very limited usefulness of Cronbach’s alpha. Psychometrika, 74(1), 107-120. https://doi.org/10.1007/s11336-008-9101-0