Internal Validity In Psychology

Internal validity is how far a study’s results can be attributed to the independent variable, once rival explanations have been ruled out. The more confounding variables a design controls, the higher its internal validity.

Key Takeaways

  • Definition: Internal validity is how far a study shows that the independent variable, not a confounding factor, caused the change in the dependent variable.
  • Degree: It is not black-and-white. Internal validity reflects how much confidence a study’s controls justify in its cause-and-effect conclusion.
  • Confounds: The more confounding variables a study avoids, the higher its internal validity and the more we can trust its cause-effect finding.
  • Threats: Campbell and Stanley (1963) catalogued eight design-level threats, including history, maturation, testing, instrumentation, attrition and selection.
  • Safeguards: Random assignment, control groups, blinding and standardised protocols protect against these threats.
  • Trade-Off: Tight control raises internal validity but can lower ecological validity, so no single design maximises both.
Close-up view of university students discussing their group project while using tablet
Internal validity is crucial for being able to draw credible conclusions from research. It allows researchers to rule out alternative explanations for study findings besides the factor being tested. 

High internal validity is the gold standard for showing that a true cause-and-effect relationship exists between the independent variable (treatment) and the dependent variable (outcome).

What Is Internal Validity? Definition and Example

Internal validity is the quality of the research itself, especially in experiments that make a cause-and-effect claim. To earn it, the independent and dependent variables must be clearly, operationally defined.

Confounding variables, meaning any factors besides the independent variable that could explain a change in the dependent variable, must be eliminated or minimised. Only then can researchers attribute a change to the manipulation.

Three questions capture most of what a critical reader should ask about a study’s internal validity.

  • Measurement: Is the study really measuring what it claims to measure? A test can be reliable and well run yet miss the intended construct.
  • Participant behaviour: Did the location or nature of the research make participants act differently? This is the territory of demand characteristics and experimenter bias.
  • Researcher objectivity: Has the researcher stayed objective? Hopes about the “right” outcome can colour how ambiguous results are read.

A worked example shows these ideas in action.

A researcher hypothesizes that a new cognitive training program will improve memory and attention in elderly adults.

To test this, the researcher builds in six controls.

  • Random assignment: Older adults from one retirement community are randomly assigned to the training program (experimental group) or a health education program (control group). Both last 8 weeks.
  • Double-blinding: The researcher double-blinds the experiment, so neither participants nor test administrators know who is in each group.
  • Pre- and post-tests: Participants complete memory and attention tests before and after the 8 weeks.
  • Confound tracking: The researcher records baseline cognitive functioning, health status, age, education level and gender.
  • Standard delivery: Strict protocols deliver both programs identically, which minimises instructor bias.
  • Drop-out checks: Statistical analyses confirm that drop-outs do not affect the two groups differently.

With these controls, the researcher can be confident that any improvement reflects the new training rather than other variables. The design rules out rival explanations.

Controlling the Situation (Latané and Darley Seizure Study)

Key Method: Rigorous experimental control over all factors except the one being tested.

  • Goal: To prove that the number of bystanders (cause) directly affects helping behavior (effect).
  • High Internal Validity Achieved By: Using prerecorded voices for the victim and other participants. This meant every participant heard the exact same emergency presentation.
  • Why it Matters: This tight control eliminated extraneous variables (like the victim’s tone or visual cues), ensuring any difference in helping was due only to the perceived size of the group.

Threats to Internal Validity

Campbell and Stanley (1963) wrote the most influential account of what can go wrong. They did not run a single study.

Instead, they worked through common designs, such as pretest-posttest and non-equivalent control-group designs. For each, they asked which extraneous processes could change the outcome without any help from the independent variable.

The result was a diagnostic vocabulary of eight threats: history, maturation, testing, instrumentation, regression to the mean, selection, attrition and the selection-maturation interaction. That precision matters. A critic can now name the exact mechanism behind a spurious result instead of objecting that “there might have been confounds.”

The taxonomy is most useful in longitudinal, field and quasi-experimental research, where control is hardest. A randomised, single-session laboratory experiment controls most of these threats automatically.

ThreatImpact on ValiditySolution
HistoryAn external event, unrelated to the study, occurs during the study period and affects participants’ responses. This event provides an alternative explanation for the outcome.Random assignment helps distribute the external event’s effect across all groups.
MaturationChanges in the DV occur because of natural biological or psychological changes in participants over time (e.g., aging, fatigue, hunger, natural skill development), not due to the IV.Random assignment helps distribute these natural tendencies evenly across groups.
Testing (Repeated Testing Effects)Scores change simply because participants have taken the pre-test before, leading to familiarity or practice effects that bias post-test results.Use a control group or implement designs that counterbalance the order of conditions/measures.
Attrition (Mortality)Differential loss of participants from study groups. If dropout rate or dropout reasons differ between groups, the remaining groups may no longer be equivalent.Only poses a threat if it occurs after random assignment.
Regression to the MeanOccurs when participants are selected based on extreme scores (very high or low). Their scores naturally move closer to the mean on later tests, regardless of the intervention.Random assignment helps prevent this by distributing participants with extreme scores evenly.
InstrumentationThe measurement instrument or procedure changes during the study (e.g., observer drift, equipment degradation, scoring inconsistencies). This can confound changes in the DV with measurement error.Ensure consistent measurement procedures; calibrate equipment; train observers.
SelectionThe comparison groups differ systematically from the start (e.g., a self-selected treatment group is more motivated), so later differences may reflect these pre-existing differences rather than the IV.Random assignment makes the groups equivalent at the start.
Selection-Maturation InteractionNon-equivalent groups also change at different natural rates, so an apparent treatment effect is really two groups on different underlying trajectories.Random assignment where possible; quasi-experimental designs cannot rule this out automatically.

Confounding variables

Confounding variables are extraneous factors that influence the dependent variable in an experiment. They create a misleading association. That makes it hard to isolate the true effect of the independent variable.

They threaten internal validity because they offer alternative explanations for the results. Changes in the dependent variable may come from the independent variable or from the confounding variable. It is unclear which.

A failure to control extraneous variables undermines the ability of researchers to create causal inferences logically. Unfortunately, however, confounding variables are difficult to control outside of laboratory settings.

Nonetheless, Campbell (1957) identified several confounding variables that can threaten internal validity. 

Participant Factors

Participant reaction biases threaten internal validity because people may act differently when they know they are being observed. They take three main forms.

  • Participant expectancies: Participants try, consciously or not, to behave as the experimenter expects. A volunteer hoping to join a depression study may exaggerate symptoms on the screening questionnaire.
  • Participant reactance: Participants deliberately act against the hypothesis, often to protect their autonomy (Brehm, 1966). In a daylight and sleep study, someone might keep one bedtime regardless of exposure.
  • Evaluation apprehension: Participants worry about being judged, so answers drift toward social or group beliefs. Asked about a political issue in a group, they may conform to others.

The first two have well-known names. Participant expectancies are often called the Hawthorne effect: people perform as they think the researcher wants, usually better, because they know they are being watched.

Reactance is sometimes called the “screw-you effect”: people deliberately act against what they believe the researcher wants, out of resentment at being studied. Both fade in natural settings. Keeping participants naive to the hypothesis also helps.

Broadly, researchers can reduce these biases by guaranteeing participant anonymity, using cover stories, unobtrusive observations, and indirect measures.

Demand Characteristics

Demand characteristics are cues in the research situation that let participants guess a study’s purpose and then adjust their behaviour.

Some try to help confirm the hypothesis. A few try to sabotage it.

Orne (1962) argued that a psychological experiment is a social situation with its own unwritten contract. Participants know they are being observed, so they work out what is really being tested. The experimenter, who assumes people are simply responding to the independent variable, never sees this.

  • Aim: To show that participants actively guess an experiment’s purpose and shape their behaviour around that guess.
  • Method: Orne used demonstrations such as asking participants to add long columns of random numbers, then tear up the sheets, for hours. No scientific reason was given.
  • Results: Participants complied with absurd, effortful tasks for long periods. Interviews showed they had built their own theories of what the experimenter was “really” testing.
  • Conclusion: Part of an apparent experimental effect can be participants acting out their own theory of the study, not reacting to the independent variable.

Orne proposed quasi-control techniques to test this directly.

A separate group guesses what the researcher expects, without ever receiving the manipulation. If their guesses alone reproduce the results, demand characteristics are a likely explanation.

The idea reshaped methodology. Single- and double-blind procedures, cover stories and post-experiment suspicion checks became standard tools. The effect is easier to demonstrate in artificial, low-stakes laboratory tasks than to quantify in field research.

How Large Are Demand Effects?

Modern work asks how much demand really distorts findings. De Quidt, Haushofer and Roth (2018) built a way to bound the problem rather than assume it away.

  • Aim: To measure how far experimenter demand could distort findings from standard experiments and surveys.
  • Method: The researchers deliberately induced demand with “demand treatments” that changed what participants believed the researcher wanted. They applied the method to 11 classic tasks.
  • Results: Estimated bounds on demand effects averaged 0.13 standard deviations.
  • Conclusion: Typical demand effects are probably modest, though they are worth checking rather than assuming away.

Orne showed why demand matters. The newer evidence suggests that, across these tasks, demand effects are typically modest.

Sampling bias

Sampling bias occurs when the way participants are selected creates key differences between groups that could skew the results. In Campbell and Stanley’s (1963) taxonomy, this is the selection threat. It introduces systematic error into the comparison of an experimental and a control group.

Suppose a study tests a new math tutoring program. The researcher unknowingly draws the experimental group from advanced math classes and the control group from regular classes.

The groups now differ before the program begins. Students in the experimental group may already have higher math ability or motivation. Any gain could reflect those pre-existing differences rather than the program itself.

The same problem arises when participants choose their own group, because volunteers for a treatment tend to be more motivated than those who decline.

Attrition

According to Campbell (1957), attrition, also known as experimental mortality, is the differential loss of participants from experimental and control groups. It threatens internal validity when drop-out rates differ markedly between the groups.

Imagine a clinical trial of a new therapy for depression. Participants are randomly assigned to therapy (experimental group) or no therapy (control group) for 8 weeks.

Some participants drop out of both groups. However, twice as many leave the control group as the experimental group.

This differential attrition introduces bias. The participants remaining in each condition are no longer equivalent. The experimental group keeps more of its original members than the smaller control group does.

Any difference in depression levels at the end could reflect this imbalance. It need not reflect a real effect of the therapy.

Attrition does most damage when the reasons for dropping out relate to the intervention. A demanding treatment, for example, may retain only the most motivated participants. That biases the final comparison. The people who remain are not like the people who started.

Experimenter bias

Experimenter bias refers to when a researcher’s expectations, perceptions, or motivations influence the outcome of an experiment in unconscious ways. This threatens internal validity because it provides an alternative explanation for results besides the independent variable being tested.

For example, a psychologist is conducting an experiment on the effects of praise on child task performance. She expects praise to improve children’s performance.

During the experiment, she unconsciously provided more encouragement and positive body language when interacting with the praise group versus the neutral group.

Consequently, the praise group performs better. Yet was it the praise, or inadvertent experimenter bias? The children may simply have picked up on the researcher’s subtle supportive cues.

This demonstrates how a researcher’s cognitive bias can unknowingly impact participant responses and behavior in a way that distorts the causal relationship between variables.

Rosenthal and Fode (1963) tested this directly, in a setting where the participants could not guess the hypothesis.

  • Aim: To test whether an experimenter’s expectation can bias results when the participants, laboratory rats, cannot infer the hypothesis.
  • Method: Student experimenters each ran five rats through a maze. Some were told their rats were “maze-bright” and others “maze-dull”, though all came from the same stock.
  • Results: Students reported that “maze-bright” rats learned the maze faster and more accurately, even though the two groups of animals did not differ.
  • Conclusion: Rats cannot read minds, so students who expected better performance must have handled, timed or scored the animals in ways that favoured it.

The design rules out demand characteristics, which isolates the experimenter’s own behaviour. The same expectation effect extends to human participants and to teachers’ expectations of pupils.

Classic effect sizes are sometimes hard to reproduce in modern, more standardised laboratories. The lesson endures. Keep experimenters blind to condition, or automate data collection.

History

Historical events unrelated to the study can influence participants’ responses during the study period. That blurs the effect of the intervention.

A major news event, for instance, might raise anxiety levels across the general population and confound an anxiety study. History is a particular threat to experiments that run over longer periods.

Imagine a 12-month trial of a new anxiety therapy. Participants are randomly assigned to receive either the new therapy or an existing therapy.

Eight months in, the COVID-19 pandemic begins. This external event raises anxiety levels for people everywhere.

At the end of the trial, anxiety is reassessed. The new therapy group shows greater reductions than the existing therapy group.

Is the difference due to the new therapy’s effectiveness, or to the pandemic? Perhaps anxiety would have fallen similarly in both groups without it. History introduces confounds and alternative explanations that undermine internal validity.

Random assignment limits the damage, because the pandemic would affect both groups. History is most dangerous when there is no comparison group, or when only some participants are exposed.

Instrumentation

Instrumentation refers to the ability of experimental instruments to give consistent results throughout a study. Changes in the instruments or procedures used to collect data can create apparent changes in the dependent variable.

This introduces systematic measurement error. It offers an alternative explanation for any observed differences besides the independent variable.

Imagine a researcher measuring blood pressure with a battery-powered device in a trial of a drug for hypertension. As the battery decays, readings may fall, so post-test pressure looks lower than pre-test pressure.

Instrumentation is not limited to electronic or mechanical instruments.

A newly hired researcher rating participants’ mental health over a month may, with experience, rate more accurately at post-test than at pre-test (Flannelly et al., 2018).

Diffusion of information between participants

The diffusion of information and treatments between patients can call internal validity into question.

Information can leak in two main ways. In one form, participants adopt a different intervention from the one they were assigned because they believe it is more effective. A control participant in a weight-loss study may copy the treatment group’s intervention after learning that its members are losing more weight.

In another form, instructions go astray. Participants may be given different instructions, or instructions that those running the study misinterpret. Participants asked to take a medication biweekly may take it twice a week or once every two weeks (Flannelly et al., 2018; Campbell, 1957).

Maturation

Maturation covers any biological change linked to age or to the simple passage of time. Examples include becoming hungry, tired or fatigued, wound healing, recovering from surgery and disease progression.

This threatens internal validity. Natural change over time can explain study results instead of the independent variable.

In a year-long study of a new reading program, children may show reading gains. Some of that improvement could simply reflect neural development and the reading skills expected with age.

Maturation matters in short studies too. Children given a repetitive computer task may lose focus within an hour, so their performance worsens (Flannelly et al., 2018).

Testing

Repeatedly exposing participants to the same or similar measures can cause familiarity and practice effects. These influence performance independently of the intervention. The threat is greatest in studies that use cognitive tests or skills assessments.

A memory study illustrates this. A researcher tests a new method for improving memory in older adults. Participants take a memory assessment before and after the training program. Their post-test scores may rise partly because it was their second attempt at the exact same test.

Practice, not training, may explain the gain. Repeated testing on the same measures offers an alternative explanation: practice effects rather than a real result of the intervention.

How can we prevent threats to internal validity?

Some methods for increasing the internal validity of an experiment include:

1. Random Assignment (Random Allocation)

  • Action: Systematically choosing individuals for treatment groups so that every participant has an equal chance of being placed in any group (treatment or control).
  • Impact: This is the most crucial technique for internal validity. It helps ensure that pre-existing differences between participants (selection bias) are distributed evenly across groups by chance, allowing you to confidently attribute differences in outcomes to the independent variable.
  • Refinement: Allocation concealment is the process of hiding the assignment sequence from researchers and participants until the moment of intervention, which protects the integrity of the initial random assignment.
Random allocation

2. Control Groups

  • Action: Including a group of participants who receive no treatment or a standard/placebo treatment but are otherwise treated identically to the experimental group.
  • Impact: This design element is essential for ruling out threats like history (external events), maturation (natural changes over time), and the placebo effect. By comparing the experimental group’s outcome to the control group’s, you isolate the effect of the specific intervention.
  • Note: The combination of random assignment and a control group defines the Randomized Controlled Trial (RCT), the gold standard for establishing high internal validity.

3. Blinding (Masking)

  • Action: Keeping study participants, researchers, and/or data collectors unaware of which treatment assignment a participant received.
  • Impact: Blinding minimizes researcher bias and participant expectation (e.g., the Hawthorne or placebo effect), which can influence behavior or outcome reporting.
  • Strongest form: Keeping both participants and researchers unaware (double-blinding) is the most protective design. It addresses instrumentation drift and demand characteristics, cues that let participants guess a study’s purpose and adjust their behaviour.
  • Related tools: Cover stories and post-experiment suspicion checks work alongside blinding. Experimenters are kept blind to condition because their expectations alone can bias results (Rosenthal & Fode, 1963).

4. Standardized Protocols & Instrumentation

  • Action: Creating a detailed study protocol that specifies the exact procedures, materials, and measurement tools to be used across all groups and all phases of the study.
  • Impact: This standardization minimizes two major threats:
    • Instrumentation Threat: Ensures that any changes in results are not due to changes in the measurement tools or observers over time.
    • Differential Attrition: Consistent treatment administration and communication help keep participants motivated and engaged, reducing systematic dropout rates.

5. Manipulation Checks and Statistical Control

  • Action: Incorporating procedures to verify that the independent variable was delivered and perceived as intended (manipulation check) and using statistical techniques to adjust for known, measured pre-existing differences.
  • Impact: Manipulation checks confirm that the treatment was delivered as intended. They guard against diffusion of treatment and weak interventions.
  • Statistical control: Randomization is the best defense. Statistical control (e.g., ANCOVA or regression) can adjust for imbalances, a common issue in quasi-experimental designs.

Internal vs External Validity

Validity refers to how accurately a test measures what it claims to. Internal validity is a statement of causality and non-interference by extraneous factors. External validity is a statement of an experiment’s generalizability to different situations or groups.

Why Internal Validity Comes First

Internal validity concerns the robustness of an experiment in itself. An experiment with external but not internal validity cannot support a causal conclusion, so it is generally unreliable for scientific inference. An experiment with only internal validity can at least support causal claims in a narrow context.

Researchers therefore check the quality of the causal claim first, then ask how far it generalises.

The two properties also pull against each other. Tight control over extraneous variables is easiest in simplified, standardised settings. That same simplicity strips away the realism that supports generalisation to everyday life, so field studies gain realism at the cost of control.

No single design maximises both, and the right balance depends on the research question. For a fuller comparison, see internal vs. external validity.

Can Internal Validity Be Measured?

No single statistic measures internal validity. It is judged by how well a design rules out the threats above, and it is a matter of degree, not a pass-or-fail property.

Cronbach’s alpha (α) is sometimes mistaken for such a measure.

In fact, it assesses internal consistency, meaning how well the items within a test are intercorrelated. That is a form of reliability, not internal validity.

A high alpha does not guarantee that a test is unidimensional or that it measures the intended construct. Alpha also rises with the number of items, so a long scale can show a high alpha even when the average inter-item correlation is low.

Reliability is necessary but not sufficient for validity.

An unreliable measure gives a different number every time, so its accuracy cannot be assessed. A reliable measure can still measure the wrong thing, consistently.

Flake, Pek and Hehman (2017) found that coefficient alpha was often the only psychometric evidence reported for scales in social and personality research. Hussey and Hughes (2020) then tested how much that matters.

  • Aim: To assess the structural validity of 15 widely used self-report questionnaires (26 scales) in social and personality psychology.
  • Method: The researchers analysed 144,496 experimental sessions. They judged each scale on internal consistency, immediate and delayed test-retest reliability, factor structure, and measurement invariance for age and gender.
  • Results: Judged on internal consistency alone, 88% of scales appeared valid. Judged comprehensively, only 4% demonstrated good validity.
  • Conclusion: Underreporting may hide widespread invalidity in the measures researchers rely on, which could threaten many findings.

The lesson carries over to experiments: one reassuring number never settles whether a study’s conclusions are justified.

References

American Psychological Association. Internal Validity. American Psychological Association Dictionary.

Brehm, J. W. (1966). A theory of psychological reactance.

Campbell, D. T. (1957). Factors relevant to the validity of experiments in social settings. Psychological Bulletin, 54(4), 297–312. https://doi.org/10.1037/h0040950

Campbell, D. T., & Stanley, J. C. (1963). Experimental and quasi-experimental designs for research. Rand McNally.

de Quidt, J., Haushofer, J., & Roth, C. (2018). Measuring and bounding experimenter demand. American Economic Review, 108(11), 3266–3302. https://doi.org/10.1257/aer.20171330

Flake, J. K., Pek, J., & Hehman, E. (2017). Construct validation in social and personality research: Current practice and recommendations. Social Psychological and Personality Science, 8(4), 370–378. https://doi.org/10.1177/1948550617693063

Flannelly, K. J., Flannelly, L. T., & Jankowski, K. R. B. (2018). Threats to the internal validity of experimental and quasi-experimental research in healthcare. Journal of Health Care Chaplaincy, 24(3), 107–130. https://doi.org/10.1080/08854726.2017.1421019

Hussey, I., & Hughes, S. (2020). Hidden invalidity among 15 commonly used measures in social and personality psychology. Advances in Methods and Practices in Psychological Science, 3(2), 166–184. https://doi.org/10.1177/2515245919882903

Kim, J., & Shin, W. (2014). How to do random allocation (randomization). Clinics in orthopedic surgery, 6(1), 103-109.

Morse, G., & Graves, D. F. (2009). Internal Validity. The American Counseling Association Encyclopedia, 292-294.

Orne, M. T. (1962). On the social psychology of the psychological experiment: With particular reference to demand characteristics and their implications. American Psychologist, 17(11), 776–783. https://doi.org/10.1037/h0043424

Rosenthal, R., & Fode, K. L. (1963). The effect of experimenter bias on the performance of the albino rat. Behavioral Science, 8(3), 183–189. https://doi.org/10.1002/bs.3830080302

Saul McLeod, PhD

BSc (Hons) Psychology, MRes, PhD, University of Manchester

Chartered Psychologist (CPsychol)

Saul McLeod, PhD, is a qualified psychology teacher with over 18 years of experience in further and higher education. He has been published in peer-reviewed journals, including the Journal of Clinical Psychology.


Olivia Guy-Evans, MSc

BSc (Hons) Psychology, MSc Psychology of Education

Associate Editor for Simply Psychology

Olivia Guy-Evans is a writer and associate editor for Simply Psychology, where she contributes accessible content on psychological topics. She is also an autistic PhD student at the University of Birmingham, researching autistic camouflaging in higher education.

Charlotte Nickerson

Writer and Cognitive Engineer

AB History, Harvard University

Charlotte Nickerson is a Harvard graduate and cognitive engineer whose work sits at the intersection of social psychology, human behaviour, and technology design. She contributed over 100 articles to Simply Psychology and holds a Master's in Cognitive Engineering from ENSC.