Predictive validity is a subtype of criterion-related validity. It refers to how well scores from a psychological instrument predict a criterion measured in the future.
The test score is the predictor. The later outcome is the criterion. Researchers evaluate predictive validity by correlating the two, and higher correlations indicate stronger prediction. A coefficient of 0.60 shows stronger prediction than 0.30.
Test scores often drive decisions, such as predicting success on the job or in education. For example, college admissions tests predict academic success in college (the criterion).
Key Takeaways
- Definition: Predictive validity is how well test scores forecast a criterion (a later outcome), such as college grades or job performance.
- Correlation: It is measured by correlating test scores with the later criterion. A coefficient of 0.60 shows stronger prediction than 0.30.
- Timing: Predictive validity measures the criterion later. Concurrent validity measures it at the same time as the test.
- Five Steps: Choose a criterion, test a sample, collect criterion data, calculate the correlation, and interpret it in context.
- Challenges: Hard-to-measure criteria, range restriction, small samples, and the time and cost of following people up can weaken the evidence.
- Context: Results may not carry over to other samples or settings, and predictive validity should sit alongside other validity evidence.
Why is predictive validity important?
Predictive validity is important because it can provide evidence that a test is useful for its intended purpose.
For example, if a college admissions test has strong predictive validity, it can be used to help identify applicants who are most likely to succeed in college.
This helps colleges make more informed admissions decisions. It can also save students time and money by steering them away from programs that do not suit them.
Predictive validity is particularly important in areas such as employment selection, clinical diagnosis, and educational placement, where test scores are used to make decisions that have significant consequences for individuals.
Professional standards agree.
The Standards for Educational and Psychological Testing (American Educational Research Association et al., 2014) expect tests used for real decisions to carry documented validity evidence.
Examples include a clinical diagnosis, a special educational needs placement, or a university admission.
The Standards treat predictive evidence as one strand of a single argument that scores are valid for their intended use.
Predictive vs concurrent validity
Predictive validity and concurrent validity are both subtypes of criterion validity. Criterion validity is an instrument’s ability to predict an external variable, called the criterion, that it should in theory predict.
The key difference between the two subtypes lies in the temporal relationship between the test administration and the measurement of the criterion.
- Predictive validity: Examines the extent to which a test can predict a criterion that is measured in the future. In essence, it’s about forecasting future outcomes. For instance, a college admissions test (the predictor) is administered to predict how well a student will perform academically in their freshman year (the criterion).
- Concurrent validity: Examines the relationship between a test and a criterion measured at the same time. This type of validity helps understand how well a new test might correspond to an existing one or to a different type of assessment.
| Feature | Predictive validity | Concurrent validity |
|---|---|---|
| Criterion measured | In the future | At the same time |
| Question it answers | Do scores forecast a later outcome? | Do scores agree with an established measure now? |
| Typical use | Selection, diagnosis, prognosis and risk assessment | Validating a shorter or cheaper test against a trusted one |
| Example | An aptitude test at age eleven, checked against exam results years later | A five-minute screening questionnaire, checked against a full clinical interview the same day |
Prediction is what matters when a test informs a decision. A test that agrees only with a same-day criterion but forecasts nothing about tomorrow has limited practical value, however good its concurrent validity looks.
How to measure predictive validity
Measuring predictive validity takes five steps, set out below. Researchers select a relevant criterion, test a sample, collect criterion data later, calculate the correlation coefficient, and interpret it in context.
Predictive validity is just one aspect of test validity. Consider it alongside other types of validity evidence, especially construct validity, which asks whether a test captures the abstract psychological trait it is meant to represent.
1. Identify a Relevant Criterion
The first step is to identify a criterion that is meaningful and relevant to the purpose of the test.
For instance, if a test is designed to predict job success, then the criterion might be supervisor ratings of job performance or objective measures of productivity.
Choose a reliable criterion that can be measured accurately. Sometimes the ideal criterion lies too far in the future or is too hard to measure. In that case, researchers may use a proxy measure.
2. Administer the Predictor Test
Once a suitable criterion has been identified, the next step is to administer the predictor test to a sample of individuals.
Administer the test under standardized conditions to minimize the influence of extraneous variables.
3. Collect Criterion Data
After a suitable time interval, collect data on the chosen criterion for the same sample of individuals.
The length of the time interval will depend on the nature of the criterion and the purpose of the study.
For example, if the criterion is job performance, data might be collected after six months or a year on the job.
4. Calculate the Correlation Coefficient
The next step is to calculate the correlation coefficient between the predictor test scores and the criterion scores.
This statistic provides a quantitative measure of the strength and direction of the relationship between the two variables.
Higher correlations indicate stronger predictive validity.
Reliability, the consistency of a measure, sets a ceiling on predictive validity. In classical test theory, the correlation between two measures cannot exceed the square root of the product of their reliabilities (Spearman, 1904). In symbols: rxy ≤ √(rxx × ryy).
Poor reliability limits validity from the start. An unreliable test or criterion, such as inconsistent supervisor ratings, caps how well the test can appear to predict. Improving reliability raises the ceiling.
5. Interpret the Correlation in Context
It is crucial to interpret the correlation coefficient in the context of the specific study and its purpose.
Researchers should be cautious about assuming that a test that predicts a criterion in one context will necessarily do so in another context.
Validity is not a certificate awarded once. Messick (1995) argued that a test is valid only for a particular use, and that judgement stays open to revision as new evidence arrives.
A depression screening test validated for adults, for example, is not automatically valid for children.
Consider factors that may have influenced the correlation, such as sample characteristics, the reliability of the measures, and any restrictions of range in either the predictor or criterion variable.
Job requirements differ between employers and change over time. The applicant pool varies too. Any of these can alter a test’s predictive validity in a new situation.
Additionally, think about the practical implications of the findings.
For instance, a statistically significant correlation may still lack practical significance. The effect may be small, or the cost of using the test may outweigh the potential benefits.
Challenges in establishing predictive validity
Four practical problems limit what a predictive validity study can show:
- Selecting an appropriate criterion: Identifying and measuring the ideal criterion can be difficult. It may lie too far ahead or be too complex to measure, so proxy measures may capture the construct imperfectly.
- Range restriction: The sample may not cover the full range of predictor or criterion scores. This artificially reduces the correlation and underestimates the true predictive validity.
- Sample size: Reliable findings need an adequate sample. Small samples give unstable correlation coefficients and reduce statistical power.
- Time and cost: Longitudinal studies, where the criterion is measured later, are slow and expensive. Tracking participants and collecting complete criterion data is difficult.
A less obvious challenge is trusting the predictor itself. Flake et al. (2017) reviewed articles in the Journal of Personality and Social Psychology. They found that validity evidence for scales was often lacking. Coefficient alpha, an index of internal consistency, was frequently the only psychometric evidence reported.
Hussey and Hughes (2020) then assessed 15 widely used questionnaires (26 scales) using data from 144,496 experimental sessions.
Judged on internal consistency alone, 88% of the scales appeared to have good validity. Assessed comprehensively, only 4% did. The checks covered internal consistency, immediate and delayed test-retest reliability, factor structure, and measurement invariance, meaning whether a scale works equally across age and gender groups.
The lesson for predictive validity work is practical. Check the evidence for the predictor in your own sample, rather than relying on an earlier validation.
Ethical considerations
The use of tests for prediction, particularly in high-stakes decision-making contexts like employment or education, raises ethical considerations.
Fairness is crucial. Tests must not disadvantage particular groups.
Considerations include potential biases in test content or administration that might unfairly impact different groups.
Additionally, the consequences of testing, such as potential discrimination or labeling based on test scores, should be carefully evaluated.
Messick (1995) went further.
He argued that a test’s consequences are part of what validity means, not a separate ethical question raised afterwards.
Imagine a selection test that predicts job performance well but produces systematically unfair outcomes for a protected group. On Messick’s view, its validity argument is compromised. On the classical view, which keeps validity to its separate evidence types, fairness is a separate question.
Responsible test use involves minimizing negative consequences and ensuring that test scores are interpreted and used ethically and appropriately.
The law reinforces this. In the United States, the Uniform Guidelines on Employee Selection Procedures (Equal Employment Opportunity Commission et al., 1978) apply when a selection test produces unequal outcomes for protected groups.
The employer must then show the test is job-related. In practice, that means demonstrating its criterion or content validity before using it for hiring decisions.
Reading List
American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association.
Barrett, G. V., Phillips, J. S., & Alexander, R. A. (1981). Concurrent and predictive validity designs: A critical reanalysis. Journal of Applied Psychology, 66(1), 1–6. https://doi.org/10.1037/0021-9010.66.1.1
Eastwick, P. W., Eagly, A. H., Finkel, E. J., & Johnson, S. E. (2011). Implicit and explicit preferences for physical attractiveness in a romantic partner: A double dissociation in predictive validity. Journal of Personality and Social Psychology, 101(5), 993–1011. https://doi.org/10.1037/a0024061
Eastwick, P. W., Luchies, L. B., Finkel, E. J., & Hunt, L. L. (2014). The predictive validity of ideal partner preferences: A review and meta-analysis. Psychological Bulletin, 140(3), 623–665. https://doi.org/10.1037/a0032432
Equal Employment Opportunity Commission, Civil Service Commission, Department of Labor, & Department of Justice. (1978). Uniform guidelines on employee selection procedures. Federal Register, 43(166), 38290–38315.
Flake, J. K., Pek, J., & Hehman, E. (2017). Construct validation in social and personality research: Current practice and recommendations. Social Psychological and Personality Science, 8(4), 370–378. https://doi.org/10.1177/1948550617693063
Hussey, I., & Hughes, S. (2020). Hidden invalidity among 15 commonly used measures in social and personality psychology. Advances in Methods and Practices in Psychological Science, 3(2), 166–184. https://doi.org/10.1177/2515245919882903
Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741–749. https://doi.org/10.1037/0003-066X.50.9.741
Spearman, C. (1904). The proof and measurement of association between two things. American Journal of Psychology, 15(1), 72–101. https://doi.org/10.2307/1412159