Validity In Psychology Research: Types & Examples

In psychology research, validity refers to the extent to which a test or measurement tool accurately measures what it’s intended to measure. It ensures that the research findings are genuine and not due to extraneous factors.

Validity can be categorized into different types based on internal and external validity.

The concept of validity was formulated by Kelley (1927, p. 14), who stated that a test is valid if it measures what it claims to measure.

For example, a test of intelligence should measure intelligence and not something else (such as memory).

Key Takeaways

  • What Validity Means: Kelley (1927) defined a valid test simply as one that measures what it claims to measure, not something else.
  • Three Core Categories: Content, criterion, and construct validity are the three main categories; modern testing standards treat construct validity as the overarching form the other two feed into.
  • Face Validity Is the Weakest Check: A test can “look right” to a naive reader without actually measuring the intended construct, so face validity alone proves little.
  • Convergent and Discriminant Validity Pair Up: A valid measure should correlate with other measures of the same construct, but not with measures of unrelated ones.
  • Internal vs External Validity: Internal validity asks whether the independent variable really caused the result; external validity asks whether that result generalizes beyond the study.
  • Predictive Validity Looks Ahead: A test has predictive validity when it forecasts a future outcome, such as an IQ score predicting later degree attainment.

Internal validity refers to whether the effects observed in a study are due to the independent variable. It excludes other confounding factors.

In other words, there is a causal relationship between the independent and dependent variables.

Internal validity can be improved by controlling extraneous variables, using standardized instructions, counterbalancing, and eliminating demand characteristics and investigator effects.

External validity refers to the extent to which results generalize beyond the study. This includes other settings (ecological validity), other people (population validity), and other times (historical validity).

External validity can be improved by setting experiments more naturally and using random sampling to select participants.

Types of Validity In Psychology

Three main categories of validity are used to assess the validity of the test (i.e., questionnaire, interview, IQ test, etc.): content, criterion, and construct.

  1. Content validity refers to the extent to which a test or measurement represents all aspects of the intended content domain. It assesses whether the test items adequately cover the topic or concept.
  2. Criterion validity assesses the performance of a test based on its correlation with a known external criterion or outcome. It can be further divided into concurrent (measured at the same time) and predictive (measuring future performance) validity.
  3. Construct validity assesses how well a test captures the abstract, unobservable trait it claims to measure. Modern testing standards treat it as the overarching form of validity, with content and criterion evidence feeding into it.

Content validity relies on expert judgement, not a statistical test. For a maths exam, this means checking the item pool covers the syllabus in the right proportions. For a depression inventory, it means checking that all nine DSM symptoms of depression appear and none dominates.

Without a clear content specification to judge against, content validity collapses into simple face validity.

table showing the different types of validity

Face Validity

Face validity is simply whether the test appears (at face value) to measure what it claims to. This is the least sophisticated measure of content-related validity, and is a superficial and subjective assessment based on appearance.

Tests wherein the purpose is clear, even to naïve respondents, are said to have high face validity. Accordingly, tests wherein the purpose is unclear have low face validity (Nevo, 1985).

A direct measurement of face validity is obtained by asking people to rate the validity of a test as it appears to them. This rater could use a Likert scale to assess face validity.

For example:

  1. The test is extremely suitable for a given purpose
  2. The test is very suitable for that purpose;
  3. The test is adequate
  4. The test is inadequate
  5. The test is irrelevant and, therefore, unsuitable

Suitable raters for judging face validity include:

  • Test-takers: People who actually take the test (e.g., a questionnaire, interview, or IQ test) are well placed to judge whether it looks right.
  • Professionals: People who work with the test, such as employers or university administrators, can offer an informed opinion.
  • The general public: People with an interest in the test, such as parents of test-takers, politicians, or teachers, add a further perspective.

The face validity of a test can be considered a robust construct only if a reasonable level of agreement exists among raters.

It should be noted that the term face validity should be avoided when the rating is done by an “expert,” as content validity is more appropriate.

Having face validity does not mean a test really measures what the researcher intends. It only means that raters judge it appears to do so. This makes face validity a crude, basic measure of validity.

A test item such as “I have recently thought of killing myself ” has obvious face validity as an item measuring suicidal cognitions and may be useful when measuring symptoms of depression.

However, the implication of items on tests with clear face validity is that they are more vulnerable to social desirability bias. Individuals may manipulate their responses to deny or hide problems or exaggerate behaviors to present a positive image of themselves.

It is possible for a test item to lack face validity but still have general validity and measure what it claims to measure. This is good because it reduces demand characteristics and makes it harder for respondents to manipulate their answers.

For example, the test item “ I believe in the second coming of Christ ” would lack face validity as a measure of depression (as the purpose of the item is unclear).

This item appeared on the first version of The Minnesota Multiphasic Personality Inventory (MMPI) and loaded on the depression scale.

Because most of the original normative sample of the MMPI were good Christians, only a depressed Christian would think Christ is not coming back. Thus, for this particular religious sample, the item does have general validity but not face validity.

Construct Validity

Construct validity assesses how well a test or measure represents and captures an abstract theoretical concept, known as a construct. It indicates the degree to which the test accurately reflects the construct it intends to measure. Researchers usually evaluate it through relationships with other variables and measures theoretically connected to the construct.

Construct validity was invented by Cronbach and Meehl (1955). This type of content-related validity refers to the extent to which a test captures a specific theoretical construct or trait, and it overlaps with some of the other aspects of validity

Aim: Cronbach and Meehl (1955) asked what makes a test a valid measure of an unobservable trait, like intelligence or anxiety. No independent gold standard exists to check such traits against.

Method: They proposed the nomological network: the web of theoretical links between a construct, other constructs, and observable measures. Researchers test each predicted link empirically.

Findings: A failed prediction is ambiguous. It could mean the test is flawed, the theory is flawed, or the comparison measure is flawed. Construct validation is therefore an ongoing process, not a single check.

Conclusion: Cronbach and Meehl’s 1955 paper is the founding paper of modern psychometrics. Later validation frameworks, including the multitrait-multimethod matrix, build on their nomological-network idea.

Construct validity does not concern the simple, factual question of whether a test measures an attribute.

Instead, it is about the complex question of whether test score interpretations are consistent with a nomological network involving theoretical and observational terms (Cronbach & Meehl, 1955).

To test for construct validity, it must be demonstrated that the phenomenon being measured actually exists. So, the construct validity of a test for intelligence, for example, depends on a model or theory of intelligence.

Construct validity entails demonstrating the power of such a construct to explain a network of research findings and to predict further relationships.

The more evidence a researcher can demonstrate for a test’s construct validity, the better. However, there is no single method of determining the construct validity of a test.

Instead, different methods and approaches are combined to present the overall construct validity of a test. For example, factor analysis and correlational methods can be used.

Convergent validity

Convergent validity is a subtype of construct validity. It assesses the degree to which two measures that theoretically should be related are related.

It demonstrates that measures of similar constructs are highly correlated. It helps confirm that a test accurately measures the intended construct by showing its alignment with other tests designed to measure the same or similar constructs.

For example, suppose there are two different scales used to measure self-esteem:

Take Scale A and Scale B. If both effectively measure self-esteem, high scorers on Scale A should also score high on Scale B. Low scorers on Scale A should likewise score low on Scale B.

A strong positive correlation between the two scores would support convergent validity. It would suggest both scales measure the same underlying construct of self-esteem.

Discriminant Validity

Discriminant validity is the extent to which measures of different constructs do not correlate with each other. It is the mirror image of convergent validity.

A self-esteem scale, for example, should not correlate strongly with an unrelated measure like reading speed. A strong correlation there would suggest the scale is picking up something other than self-esteem.

Campbell and Fiske (1959) formalized this two-sided convergent-discriminant test in their multitrait-multimethod matrix.

Aim: Campbell and Fiske (1959) argued that a valid test must both converge with related measures and diverge from unrelated ones. They warned that using the same method, such as self-report, for different traits can make unrelated traits look connected.

Method: They measured at least two different traits, each using at least two different methods, then compared every resulting correlation against the others.

Findings: A trait showed real convergent validity only when same-trait correlations beat both same-method and different-trait correlations. This showed the trait mattered more than the method used to measure it.

Conclusion: Campbell and Fiske’s two-sided test still shapes how psychologists design construct-validity studies in personality, clinical, and social psychology.

Concurrent Validity (i.e., occurring at the same time)

Concurrent validity evaluates how well a test’s results correlate with an established, accepted measure administered at the same time.

It helps in determining whether a new measure is a good reflection of an established one without waiting to observe outcomes in the future.

Very often, a new IQ or personality test might be compared with an older but similar test known to have good validity already.

Predictive Validity

Predictive validity assesses how well a test predicts a future outcome, such as later job performance or academic success.

For example, a new intelligence test might predict that high scorers at age 12 will more often earn a university degree years later. If that prediction holds true, the test has predictive validity.

Critical Evaluation of Validity Research

Modern methodology research has sharpened the classic validity framework above rather than replacing it.

Contemporary Research

Aim: Loken and Gelman (2017) asked why published psychology effect sizes are so often larger than the true effect.

Method: They combined classical measurement theory with the common practice of selecting only statistically significant results across many published studies.

Findings: Noisy measurement does not just shrink true effects, as older theory assumed. Combined with selecting for significance, it inflates published effect sizes upward.

Conclusion: A test’s low reliability threatens the validity of any published finding built on it, because the resulting bias is systematic rather than random.

References

Campbell, D. T., and Fiske, D. W. (1959) Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56, 81-105.

Cronbach, L. J., and Meehl, P. E. (1955) Construct validity in psychological tests. Psychological Bulletin, 52, 281-302.

Hathaway, S. R., & McKinley, J. C. (1943). Manual for the Minnesota Multiphasic Personality Inventory. New York: Psychological Corporation.

Kelley, T. L. (1927). Interpretation of educational measurements. New York: Macmillan.

Loken, E., and Gelman, A. (2017) Measurement error and the replication crisis. Science, 355(6325), 584-585.

Nevo, B. (1985). Face validity revisited. Journal of Educational Measurement, 22(4), 287-293.

Olivia Guy-Evans, MSc

BSc (Hons) Psychology, MSc Psychology of Education

Associate Editor for Simply Psychology

Olivia Guy-Evans is a writer and associate editor for Simply Psychology, where she contributes accessible content on psychological topics. She is also an autistic PhD student at the University of Birmingham, researching autistic camouflaging in higher education.


Saul McLeod, PhD

Chartered Psychologist (CPsychol)

BSc (Hons) Psychology, MRes, PhD, University of Manchester

Saul McLeod, PhD, is a qualified psychology teacher with over 18 years of experience in further and higher education. He has been published in peer-reviewed journals, including the Journal of Clinical Psychology.