Face validity is the degree to which a test appears to measure what it is intended to measure, judged by a test-taker or an expert. It is the weakest, most subjective form of validity evidence, but it still boosts test-taker cooperation, public acceptance, and clinical rapport.
Key Characteristics
- Subjective: It is based on a judgment or first impression, not statistical analysis.
- Surface-Level: It assesses the relevance and appropriateness of the test items at face value.
- Quick Assessment: It’s often the quickest and easiest way to initially check if a measure seems suitable.
- Aids Acceptance: Good face validity can increase the confidence and cooperation of the participants because the test seems relevant to them.
What Is Face Validity?
Face validity is the extent to which a test appears to measure what it is intended to measure.
In simpler terms, face validity is about whether a test looks like it is measuring what it claims to be measuring (Johnson, 2021).
- For example, asking someone if they feel sad or depressed would have face validity as a measure of depression.
- Asking someone to solve complex math problems would not have face validity as a measure of depression, even though there might be a correlation between mathematical ability and depression.
Face validity and genuine accuracy do not always move together. A driving exam’s hazard-perception test has excellent face validity: learners press a button the instant they spot a developing hazard in a video clip.
Some measures do the opposite on purpose.
The Implicit Association Test measures automatic racial or gender attitudes, and it has poor face validity by design. A respondent sorting words and faces into categories has no obvious sense that the test is revealing an attitude they might not consciously hold.
This opacity matters. A transparent test invites social desirability bias, the tendency to manage the impression it gives. That can make a measure less accurate even when it looks fine. The trade-off is a real one.
Face validity is not a technically rigorous form of validity, and does not rely on established theory for support (Fink, 2010).
It doesn’t guarantee that a test actually measures what it is supposed to measure.
However, it’s still an important consideration, particularly in applied settings, as it can impact test-taker cooperation, public acceptance of test results, and the establishment of rapport in clinical settings.
Importance of face validity
Face validity is not a technically rigorous form of validity, meaning it is not a guarantee that a test is actually measuring what it is supposed to. It is primarily a matter of public perception.
However, it is still an important consideration for several reasons:
- Test-taker cooperation: A test with good face validity is more likely to get honest, engaged answers, since test-takers who see the test as relevant take it more seriously.
- Public acceptance: When a test looks like it measures what it claims to, people trust its results more, which matters for decisions like educational placement or hiring.
- Clinical utility: In clinical settings, a test that looks relevant helps build rapport, so clients trust the clinician and engage more fully with treatment.
Tests that appear to be face valid can give participants and researchers alike confidence that the results of the assessment are fair and equitable (Johnson, 2021).
Face validity can be used to eliminate subpar research quickly.
For example, a researcher reviewing a paper on the link between vaccinations and autism in children might spot design flaws in the study. Those flaws could sink it. The paper might then be rejected for face validity alone.
Establishing face validity is also a useful first step toward other, more rigorous types of validity.
It feeds into content validity, defined as “the extent to which a test covers all important aspects of the domain being measured” (Siraj et al., 2021). It also feeds into construct validity: whether a test truly captures the psychological construct it claims to measure.
How can face validity be improved?
Here are ways to improve face validity:
- Clear items: Test-takers see a test as fairer and more valid when they understand exactly what each item is asking.
- Appropriate language: Items should use language the test-taker actually understands, simpler wording for children than for adults.
- Relevant items: Test-takers rate a test as more valid when its items clearly measure something that matters, such as the skills a job actually needs.
- Explained purpose: Test-takers cooperate more once they understand why they are being asked to take the test.
- Cultural fit: A test that looks valid in one culture, such as one asking about personal experience, may not in another.
Who should measure face validity?
Face validity is a subjective judgment about whether a test appears to measure what it is supposed to measure.
In general, however, it is best to have multiple people measure face validity, as different people may have different perspectives on what is important for measuring a construct.
Determining face validity often involves considering the opinions of both test-takers and experts:
- Test-takers: If a test doesn’t feel relevant, test-takers may lose motivation or even sabotage it. Older adults, for example, often disengage from memory tests they see as trivial.
- Experts: Content specialists judge whether a test’s items align with the construct being measured. They match each item to the domain it’s meant to cover.
There are two common approaches. Researchers might ask test-takers to rate the relevance and clarity of test items. Or they might convene a panel of experts to review the test content and give feedback on its face validity.
It is also important to note that face validity is not static; that is, what is considered face valid for measuring a construct can change over time.
Standards shift over time. For example, a personality test measuring “masculinity” and “femininity” was developed in the 1950s. It may no longer be considered valid today, since society’s understanding of gender has changed significantly since then.
As such, it is important to review and update measures of face validity regularly.
How to measure face validity
Face validity is a subjective judgment. It does not guarantee that a test actually measures what it is supposed to measure.
However, it can be a useful tool for improving the quality of a test by identifying items that are unclear or irrelevant to the test-takers.
- Gather judgments: Ask test-takers to rate how well each item seems to measure what it’s supposed to measure.
- Convene experts: A panel of experts reviews the test content and rates how well it aligns with the construct being measured.
- Systematic rating: Experts match each item to the domain facet it represents, then statistical methods like factor analysis document how much the judges agree.
When should you test face validity?
Face validity is often measured during the early stages of test development. It gives researchers an early signal of whether a test’s content and format suit the construct being measured.
One caveat applies. Face validity is only a preliminary step toward a test’s overall validity. Other validity types matter too, including content validity and predictive validity (whether a test’s scores forecast a later outcome, like exam performance). Only then can researchers determine whether a test actually works (Fink, 2010).
Face Validity vs Content Validity
Face validity and content validity are distinct but related concepts in the field of psychometrics. While both relate to the perceived appropriateness of a test, they differ in their scope and focus.
- Face validity is a superficial assessment of whether a test looks like it measures what it intends to measure. It is primarily based on the perceptions of test-takers and non-experts.
- Content validity, on the other hand, is a more rigorous evaluation that considers how well the test items represent the entire domain or universe of content that the test is designed to measure.
Content validity necessitates a careful examination of the test content by subject matter experts to determine its alignment with the construct being measured.
Face validity focuses more on appearances and perceptions than on a systematic evaluation of the test content.
For example, a depression questionnaire that asks about symptoms like sadness and loss of interest would have face validity because these symptoms are commonly associated with depression.
However, content validity would require ensuring that the questionnaire adequately covers all the key aspects of depression as defined by experts and diagnostic criteria.
Face validity is often considered a subtype of content validity, meaning that a test with good content validity will typically also have good face validity.
However, the reverse is not always true. A test can appear to measure what it is supposed to (face validity) but may not actually cover the full breadth of the construct (content validity).
Here’s a table summarizing the key differences between face validity and content validity:
| Feature | Face Validity | Content Validity |
|---|---|---|
| Definition | The degree to which a test appears to measure what it is supposed to measure. | The extent to which a test adequately samples the domain or universe of content that it is intended to measure. |
| Focus | Superficial appearance and perceptions of the test. | Systematic evaluation of test content. |
| Perspective | Test-takers, non-experts | Subject matter experts |
| Rigor | Subjective, less rigorous | Objective, more rigorous |
| Relationship | Often considered a subtype of content validity. | Can have good content validity without good face validity. |
| Examples | A depression questionnaire asking about sadness and loss of interest has face validity. | A math test that covers all the topics taught in a course has content validity. |
| Measurement | Subjective ratings from test-takers and experts. | Expert review of test items and their alignment with the construct; examination of item response consistencies and the empirical domain structure. |
| Importance | Important for test-taker cooperation, public acceptance, and clinical utility. | Crucial for ensuring that a test is a valid measure of the construct of interest. |
| Limitations | Does not guarantee actual validity; subjective judgments can vary. | Can be challenging to establish for complex constructs; requires careful consideration of the domain and its boundaries. |
| Other points | Widely regarded as the weakest form of validity evidence, since it rests on subjective impression rather than statistical testing. | Requires an explicit content specification, such as a syllabus or job-task analysis, against which the item pool can be checked for proportionate coverage. |
In essence, face validity is a useful initial indicator of a test’s appropriateness. Content validity provides a more robust and reliable assessment of whether a test truly measures what it is intended to measure.
Key Takeaways
- Definition: Face validity is whether a test appears, at a glance, to measure what it claims to measure.
- Subjective: It reflects a first impression from a test-taker or expert, not a statistical result, making it the weakest form of validity evidence.
- Can Mislead: A test can look right yet be flawed, or look wrong yet still work. The Implicit Association Test hides its purpose on purpose, limiting social desirability bias.
- Practical Value: Good face validity boosts test-taker cooperation and public trust in results, even though it proves nothing statistically.
- Not A Substitute: It never replaces content, criterion, or construct validity, the more rigorous, evidence-based checks.
- Best Checked Early: Researchers usually assess face validity during test development, then confirm it with harder validity evidence.
References
Fink, A. Peterson, P. L., Baker, E., & McGaw, B. (2010). International encyclopedia of education. Elsevier Ltd..
Johnson, E. (2021). Face validity. In Encyclopedia of autism spectrum disorders (pp. 1957-1957). Cham: Springer International Publishing.
McDermott, R. (2011). Internal and external validity. Cambridge handbook of experimental political science, 27-40.
Messick, S. (1995). Standards of validity and the validity of standards in performance assessment. Educational measurement: Issues and practice, 14(4), 5-8.
Rubio, D. M. (2005). Content validity.
Siraj, S., Stark, W., McKinley, S. D., Morrison, J. M., & Sochet, A. A. (2021). The bronchiolitis severity score: An assessment of face validity, construct validity, and interobserver reliability. Pediatric pulmonology, 56(6), 1739-1744.