Internal vs. External Validity In Psychology

Internal validity focuses on establishing cause and effect, while external validity addresses generalizability.

  • Internal validity refers to how well a study establishes a causal relationship between variables by minimizing confounding factors and bias.
  • External validity is the extent to which the study results can be generalized to other populations and settings beyond the specific research context.
Internal ValidityExternal Validity
DefinitionWhether conclusions about cause and effect relationships within a study are validThe extent study results apply to contexts beyond the original study
Main ConcernWere effects observed really caused by the independent variable or did flaws in the study design/conduct lead to that result?Can results be expected to apply to other settings, populations, times?
Key FactorsRandomization, control conditions, elimination of confounding variablesHaving a sample representative of the population of interest, testing variability in contexts
Examples ThreatsSelection bias, attrition, history effectsInteraction effects of setting and treatment, limited participant sample
How to ImproveUse control groups, randomization, blinding, account for confoundersDraw from heterogeneous, more representative samples, replicate across ranges of contexts
Balance ConsiderationControlling internal validity often means more artificial research contextBroader generalizability requires flexible, real-world applicable paradigms

Internal Validity 

Internal validity refers to the degree of confidence that the causal relationship being tested exists and is trustworthy.

Internal validity centers on the strength of the causal relationship between the independent and dependent variables.

It addresses whether the observed effects can be confidently attributed to the intervention or experimental manipulation rather than to extraneous factors or confounding variables.

Internal validity demands rigorous control. Researchers must isolate the effects of the independent variable from every extraneous or confounding variable in the design.

This often involves conducting research in laboratory settings with standardized procedures, random assignment, and control groups.

Such tight control helps ensure that the observed effects are truly due to the manipulation or intervention and not to confounding factors.

Threats can creep in at every stage. Campbell and Stanley (1963) catalogued eight recurring threats to internal validity: history, maturation, testing, instrumentation, statistical regression, selection, attrition, and selection-maturation interaction.

Example

Suppose you want to test whether a new weight-loss pill actually helps people lose weight. You randomly assign participants to one of two groups: one takes the pill, the other takes a placebo.

Random assignment is not enough on its own.

Blinding the research assistants keeps them unaware of which group each participant is in during the experiment. The participants themselves are blinded too. They do not know whether they are receiving the pill or the placebo.

If participants drop out, researchers check whether those who left differ systematically from those who stayed. Uneven dropout between groups can bias the results as much as poor random assignment.

This keeps threats to internal validity in check.

External Validity

External validity focuses on the extent to which the findings of a study can be generalized beyond the specific sample, setting, and time in which the research was conducted.

External validity matters for generalizability. When it is established, a study’s findings apply beyond the small group who took part, to a much larger population.

To generalize findings, researchers need to ensure that their samples, settings, and procedures reflect the real-world conditions to which they wish to apply their results.

However, the artificiality and constraints imposed by laboratory settings can limit the extent to which findings can be generalized to natural environments.

External validity is different from internal validity. It doesn’t establish causality or rule out confounders.

There are two types of external validity: ecological validity and population validity.

  • Ecological validity refers to whether a study’s findings can be generalized to other situations or settings. A high ecological validity means that there is a high degree of similarity between the experimental setting and another setting, and thus we can be confident that the results will generalize to that other setting.
  • Population validity refers to how well the experimental sample represents other populations or groups. Using random sampling techniques, such as stratified sampling or cluster sampling, significantly helps increase population validity. 

Example

Suppose you hypothesize that practicing mindfulness twice a week will improve the mental health of people diagnosed with depression.

You recruit people who have been diagnosed with depression for at least a year, aged 18–29. A clearly defined, representative sample like this helps ensure external validity.

You give participants a pre-test. A post-test then measures how often they experienced depression symptoms in the past week.

All participants receive individual mindfulness training. They practice for 15 minutes a day as part of their morning routine.

You can also replicate the study’s results using different methods of mindfulness or different samples of participants. 

Trade-off Between Internal and External Validity

This trade-off arises from research’s two fundamental goals. One is to establish causal relationships within a controlled setting (internal validity); the other is to generalize those findings to broader populations and contexts (external validity).

Internal validity is a prerequisite for external validity.

A weak causal claim doesn’t generalize. Without strong evidence that the relationship is real within the study itself, extending the findings to other contexts becomes meaningless.

Transition experiments help. They bridge the gap between fully controlled and real-world field settings.

They incorporate elements of both environments, gradually increasing ecological validity while maintaining some control.

This gradual shift can provide valuable insights into the robustness of the findings across different contexts.

  • Establishing causality in a controlled setting: This stage focuses on internal validity. It provides a solid foundation for understanding the basic mechanisms and generating hypotheses about how the phenomenon might work in different settings3.
  • Testing generalizability in a field experiment: This stage prioritizes external validity by assessing whether the causal relationship established in the controlled setting holds true in a more realistic environment.

Conducting two separate experiments requires more time, funding, and personnel. Researchers must weigh the potential benefits of the approach against these practical limitations.

Threats to Internal Validity

Attrition

Attrition refers to the loss of study participants over time. Participants might drop out or leave the study which means that the results are based solely on a biased sample of only the people who did not choose to leave.

Differential rates of attrition between treatment and control groups can skew results. Uneven drop-out affects the relationship between your independent and dependent variables and thus threatens a study’s internal validity.

Confounders

A confounding variable is an unmeasured third variable that influences, or “confounds,” the relationship between an independent and a dependent variable by suggesting the presence of a spurious correlation.

Confounders are threats to internal validity because you can’t tell whether the predicted independent variable causes the outcome or if the confounding variable causes it.

Participant Selection Bias

This is a bias that may result from the selection or assignment of study groups in such a way that proper randomization is not achieved.

If participants are not randomly assigned to groups, the sample obtained might not be representative of the population intended to be studied.

For example, some members of a population might be less likely to be included than others due to motivation, willingness to take part in the study, or demographics. 

Experimenter Bias

Experimenter bias occurs when an experimenter behaves in a different way with different groups in a study, impacting the results and threatening internal validity. This can be eliminated through blinding.

Aim: Rosenthal and Fode (1963) tested whether an experimenter’s expectations could bias results. They used rats, which cannot consciously guess or comply with a hypothesis.

Method: Five rats went to each psychology student. They ran a standard maze-learning task. Some students were told, falsely, that their rats were “maze-bright”; others were told theirs were “maze-dull.” In reality, all the rats came from the same stock and were assigned to condition at random.

Results: Rats labelled “maze-bright” were reported as learning faster and more accurately than rats labelled “maze-dull.” Yet the rats didn’t actually differ at all.

Conclusion: Rats cannot read an experimenter’s mind, so the result could not be explained by demand characteristics. Students who expected better performance seem to have unintentionally handled, timed, or scored their animals in ways that favored that outcome.

Social Interaction (Diffusion)

Diffusion refers to when the treatment in research spreads within or between treatment and control groups, usually through interaction or observation between the groups.

This is a real threat. Diffusion can lead to resentful demoralization, where the control group loses motivation because its members feel resentful about the group they are in.

Demand Characteristics

Demand characteristics are cues in the research situation that lead participants to guess a study’s purpose. They then adjust their behavior to fit it.

Aim: Orne (1962) argued a psychological experiment is a social situation, not a neutral stimulus. Participants actively guess what is being tested and shape their behavior around that guess.

Method: Participants performed tedious, meaningless tasks, such as adding long columns of numbers for hours, with no scientific justification given.

Results: Participants complied anyway, however absurd the task. Interviews afterward showed they had guessed the study’s real purpose and acted to confirm it.

Conclusion: Participants act as motivated helpers, not passive responders. Part of an apparent experimental effect can reflect this guessing, not the independent variable itself.

Historical Events

Historical events might influence the outcome of studies that occur over longer periods of time.

For example, changes in political leadership, natural disasters, or other unanticipated events might change the conditions of the study and influence the outcomes.

Instrumentation

Instrumentation refers to any change in the dependent variable in a study that arises from changes in the measuring instrument used. This happens when different measures are used in the pre-test and post-test phases. 

Maturation

Maturation refers to the impact of time on a study. Participants can change naturally over the course of a study, growing stronger, more coordinated, or simply more experienced with the testing procedure, independent of any treatment.

If study outcomes shift over time for these natural reasons, it becomes hard to tell whether the effect came from the treatment or simply from the passage of time.

Statistical Regression

Regression to the mean is a simple statistical fact. If one sample of a random variable is extreme, the next sampling of that variable will likely land closer to its average.

This threatens internal validity because participants at the extreme ends of a treatment can naturally drift back toward average scores over time. That drift can look like a treatment effect when it isn’t one.

Repeated Testing

Testing your research participants repeatedly with the same measures will influence your research findings because participants will become more accustomed to the testing.

Due to familiarity, or awareness of the study’s purpose, many participants might achieve better results over time.

Threats to External Validity 

Sample Features

If some feature(s) of the sample used were responsible for the effect, this could lead to limited generalizability of the findings.

Historical events matter for external validity too. The same events that threaten internal validity (see Historical Events above) can also limit how well a study’s findings generalize to other time periods.

The same selection bias threatens external validity too. It can produce a sample that fails to represent the wider population (see Participant Selection Bias above), limiting generalizability.

Factors such as the setting, time of day, location, researchers’ characteristics, noise, or the number of measures might affect the generalizability of the findings.

The same practice effects apply here too. As with internal validity (see Repeated Testing above), they can make results harder to generalize to people encountering the measure for the first time.

The Aptitude-Treatment Interaction refers to the concept that some treatments are more or less effective for particular individuals, depending on their specific abilities or characteristics.

Hawthorne Effect

The Hawthorne Effect refers to the tendency for participants to change their behaviors simply because they know they are being studied.

The same experimenter bias that threatens internal validity (see Experimenter Bias above) can also limit how well results generalize. That’s true whenever the experimenter’s own behavior was part of what produced the effect.

John Henry Effect

The John Henry Effect refers to the tendency for participants in a control group to work harder than usual. They know they are in an experiment and want to overcome the “disadvantage” of being in the control group.

Factors that Improve Internal Validity

Blinding

Blinding refers to a practice where the participants (and sometimes the researchers) are unaware of what intervention they are receiving.

This reduces the influence of extraneous factors and minimizes bias. Differences in outcome can then be linked to the intervention itself, not to whether participants knew they were receiving a new treatment.

Random Sampling

Using random sampling to obtain a sample that represents the population that you wish to study will improve internal validity. 

Random Assignment

Using random assignment to assign participants to control and treatment groups ensures that there is no systematic bias among the research groups. 

Strict Study Protocol

Highly controlled experiments tend to improve internal validity.

Experiments that occur in lab settings tend to have higher validity as this reduces variability from sources other than the treatment. 

Experimental Manipulation

Manipulating an independent variable in a study as opposed to just observing an association without conducting an intervention improves internal validity. 

Factors that Improve External Validity

Replication

Conducting a study more than once with a different sample or in a different setting to see if the results will replicate can help improve external validity.

If multiple studies test the same topic, a meta-analysis can pool their results. Replication makes a finding more trustworthy. It shows whether an independent variable’s effect holds up across studies.

Replication is the strongest method to counter threats to external validity by enhancing generalizability to other settings, populations, and conditions.

Field Experiments

Conducting a study outside the laboratory, in a natural, real-world setting will improve external validity (however, this will threaten the internal validity) 

Probability Sampling

Using probability sampling will counter selection bias by making sure everyone in a population has an equal chance of being selected for a study sample.

Recalibration

Recalibration is the use of statistical methods to maintain accuracy, standardization, and repeatability in measurements to assure reliable results.

Reweighting groups, if a study had uneven groups for a particular characteristic (such as age), is an example of calibration. 

Inclusion and Exclusion Criteria

Setting clear criteria for who can and cannot take part in the research ensures the population being studied is well defined. This, in turn, helps the sample stay representative of that population.

Psychological Realism

Psychological realism refers to making sure participants perceive the experimental manipulations as real events, without giving away the purpose of the study. This way, participants don’t behave differently than they would in real life just because they know the study’s goal.

Critical Evaluation

The internal/external validity framework offers a systematic way to critique almost any study. Check the causal claim first, then ask how far it generalizes.

The framework has real limits, though. One common criticism deserves scrutiny rather than automatic acceptance.

Simply noting that a study happened in a lab, and therefore “lacks ecological validity,” is a well-recognized but shallow objection. The criticism only carries weight when it names a specific feature of the setting that plausibly differs from real life.

A word-list memory experiment isn’t automatically invalid as a study of memory just because a supermarket feels more natural than a lab. The criticism needs more than that. It needs a concrete reason to think the two tasks work differently.

Validity is also not a single pass-or-fail property. Messick (1995) rejected treating face, content, criterion, and construct validity as separate boxes to tick.

He proposed a unified view instead. Every kind of evidence, including a test’s practical and social consequences, feeds into one overall judgment about whether a specific score-based inference is justified.

A test is never simply “valid” in the abstract. It is valid, or not, for a particular purpose and population, and that judgment can change as new evidence arrives.

Qualitative research applies the framework differently again. It isn’t usually meant to generalize beyond the group studied, the way population validity implies for quantitative work. Instead, it’s judged on credibility, coherence, and triangulation across sources of evidence.

Contemporary Research

A consistent theme in the past decade of methodological work is that psychology often assumed a validity case for its measures and samples rather than demonstrating one. Recent audits have found that assumption frequently unjustified.

Aim: Flake, Pek, and Hehman (2017) studied how often published psychology studies report real validity evidence. Many just cite an earlier paper’s validation instead.

Method: They coded a large sample of published articles for whether, and how, each measure’s validity was actually reported.

Results: Many measures carried no validity evidence for the current sample at all. Where evidence appeared, it was often just internal-consistency reliability, which measures consistency, not accuracy.

Conclusion: Construct validation here had become a citation ritual.

Hussey and Hughes (2020) went further. They directly tested fifteen widely used measures against a fuller battery of modern validity criteria. Most failed at least one basic check that had never actually been run before.

A parallel strand extends this to sampling. Rad, Martingano, and Ginges (2018) found that published samples had barely diversified since the WEIRD critique first appeared.

WEIRD stands for Western, Educated, Industrialized, Rich, and Democratic, the narrow slice of humanity psychology has long over-relied on. They proposed broader international sampling and clearer caveats about a claim’s intended population as practical fixes.

Key Takeaways

  • Internal Validity: Whether a study’s design supports the causal claim being made, by ruling out confounds and other alternative explanations.
  • External Validity: Whether findings generalize beyond the specific sample, setting, and time of the original study.
  • The Trade-off: Tighter control over confounds usually means a more artificial setting, so no single design maximizes both at once.
  • Threats Taxonomy: Campbell and Stanley (1963) catalogued eight recurring threats to internal validity, from history and maturation to attrition and selection-maturation interaction.
  • Demand Characteristics: Orne (1962) showed participants often guess a study’s purpose and shape their behavior to “help” confirm it, which can look like a genuine effect.
  • Modern Evidence: Recent audits (Flake et al., 2017; Hussey & Hughes, 2020) found many widely used psychology measures had never actually been checked against modern validity criteria.

References

Campbell, D. T., & Stanley, J. C. (1963). Experimental and quasi-experimental designs for research. Rand McNally.

Flake, J. K., Pek, J., & Hehman, E. (2017). Construct validation in social and personality research: Current practice and recommendations. Social Psychological and Personality Science, 8(4), 370-378. https://doi.org/10.1177/1948550617693063

Hussey, I., & Hughes, S. (2020). Hidden invalidity among 15 commonly used measures in social and personality psychology. Advances in Methods and Practices in Psychological Science, 3(2), 166-184. https://doi.org/10.1177/2515245919882903

Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741-749. https://doi.org/10.1037/0003-066X.50.9.741

Orne, M. T. (1962). On the social psychology of the psychological experiment: With particular reference to demand characteristics and their implications. American Psychologist, 17(11), 776-783. https://doi.org/10.1037/h0043424

Rad, M. S., Martingano, A. J., & Ginges, J. (2018). Toward a psychology of Homo sapiens: Making psychological science more representative of the human population. Proceedings of the National Academy of Sciences, 115(45), 11401-11405. https://doi.org/10.1073/pnas.1721165115

Rosenthal, R., & Fode, K. L. (1963). The effect of experimenter bias on the performance of the albino rat. Behavioral Science, 8(3), 183-189. https://doi.org/10.1002/bs.3830080302

Saul McLeod, PhD

BSc (Hons) Psychology, MRes, PhD, University of Manchester

Chartered Psychologist (CPsychol)

Saul McLeod, PhD, is a qualified psychology teacher with over 18 years of experience in further and higher education. He has been published in peer-reviewed journals, including the Journal of Clinical Psychology.


Olivia Guy-Evans, MSc

BSc (Hons) Psychology, MSc Psychology of Education

Associate Editor for Simply Psychology

Olivia Guy-Evans is a writer and associate editor for Simply Psychology, where she contributes accessible content on psychological topics. She is also an autistic PhD student at the University of Birmingham, researching autistic camouflaging in higher education.

Julia Simkus

Psychology Researcher and Writer

BA (Hons) Psychology, Princeton University

Julia Simkus is a Princeton University graduate in Clinical Psychology (Magna Cum Laude) and holds a Master of Arts in Applied Psychology from New York University. During her studies she worked as a research assistant to Professor Nicole Avena at Princeton, co-authoring three published works on food addiction and substance use disorders in peer-reviewed journals and Oxford University Press. She wrote and edited over 70 articles for Simply Psychology between 2021 and 2024.