Experimental design refers to how participants are allocated to different groups in an experiment. Types of design include repeated measures, independent groups, and matched pairs designs.
Researchers must decide how to allocate participants to the different conditions of the independent variable (IV) before running a study. This choice affects how many participants are needed and which statistical test analyzes the results.
For example, a study with 10 participants could test everyone in both conditions. This is a repeated measures design. Alternatively, the researcher could split the group in half, testing each half in only one condition.
Types
Three types of experimental designs are commonly used:
1. Independent Measures
Independent measures design, also known as between-groups, is an experimental design where different participants are used in each condition of the independent variable. Participants should be randomly allocated to conditions, giving each person an equal chance of being assigned to either group.
For example, in a study with two conditions, one group of participants completes each condition:
- Con: More people are needed than with the repeated measures design (i.e., more time-consuming).
- Pro: Avoids order effects, such as practice or fatigue, because each participant takes part in one condition only.
- Pro: Reduces demand characteristics, since participants who see only one condition find it harder to guess what the study is testing.
- Con: Differences between participants — called participant variables, a type of extraneous variable — such as age, gender, or social background, may affect results.
- Control: After the participants have been recruited, they should be randomly assigned to their groups. This should ensure the groups are similar, on average (reducing participant variables).
Fisher (1935): The Origins of Independent-Groups Design
This design descends from Ronald Fisher’s agricultural work at Rothamsted.
Aim: To formalize a general theory of designed experiments that would separate a manipulation’s effect from the many nuisance influences on the outcome (Fisher, 1935).
Method: Working from agricultural field trials, Fisher codified three principles: randomization, replication, and local control (blocking). He introduced completely randomized designs, randomized block designs, and Latin squares, linking each layout to the statistical test that analyzes it.
Results: Randomization licensed the significance test. Because allocation was made by chance, the null-hypothesis distribution could be justified by randomization itself.
Conclusion: Fisher’s framework moved from agriculture into psychology largely unchanged. An independent-groups experiment is, in essence, a completely randomized design; a matched-pairs experiment is a randomized block design with blocks of two. Fisher’s earlier analysis of variance (ANOVA) supplied the statistical engine that analyzes these designs.
2. Repeated Measures Design
Repeated measures design is an experimental design where the same participants take part in every condition of the independent variable. It is also known as within-groups or within-subjects design.
- Pro: As the same participants are used in each condition, participant variables (i.e., individual differences) are reduced.
- Con: Order effects can occur, where the order of the conditions affects participants’ behavior rather than the IV itself.
- Con: A practice effect can improve performance in a later condition as participants grow familiar with the task, while a fatigue effect can worsen performance as participants tire.
- Pro: Fewer people are needed as they participate in all conditions (i.e., saves time).
- Control: Researchers counterbalance the order of conditions across participants to spread any order effect evenly, as explained below.
Counterbalancing
Suppose we used a repeated measures design in which all of the participants first learned words in “loud noise” and then learned them in “no noise.”
We expect the participants to learn better in “no noise” because of order effects, such as practice. However, a researcher can control for order effects using counterbalancing.
The sample splits into two groups: experimental (A) and control (B). Group 1 tries A first and B second. Group 2 tries B first and A second. This balances out order effects.
Although order effects occur for each participant, they balance each other out in the results because they occur equally in both groups.
This example splits participants into two groups, but counterbalancing can also work within a single participant. In ABBA counterbalancing, each person completes the sequence A, B, B, A, so any steady practice effect is balanced across both conditions.
The logic works either way. Counterbalancing controls practice and fatigue effects well because they build up gradually. It is less effective against a carryover effect, where one treatment’s impact lingers into the next condition.
A lingering drug is one example. When carryover is likely, researchers space out the conditions or switch to an independent measures design.
3. Matched Pairs Design
A matched pairs design is an experimental design where pairs of participants are matched in terms of key variables, such as age or socioeconomic status.
One member of each matched pair is randomly assigned to the experimental group, and the other to the control group.
- Con: If one participant drops out, the researcher loses both members’ data.
- Pro: Reduces participant variables, since each condition contains people with similar abilities and characteristics.
- Con: Finding closely matched pairs is very time-consuming.
- Pro: Avoids order effects, so counterbalancing is not necessary.
- Con: Matching people exactly is impossible unless they are identical twins.
- Control: Members of each pair are randomly assigned to conditions, though this does not solve every problem.
Summary
Experimental design refers to how participants are allocated to an experiment’s different conditions, or IV levels. There are three types:
- Independent measures / between-groups: Different participants are used in each condition of the independent variable.
- Repeated measures / within groups: The same participants take part in each condition of the independent variable.
- Matched pairs: Each condition uses different participants, but they are matched in terms of important characteristics, e.g., gender, age, intelligence, etc.
Learning Check
Read about each of the experiments below. For each experiment, identify (1) which experimental design was used; and (2) why the researcher might have used that design.
Scenario 1: Comparing Two Therapies
To compare the effectiveness of two different types of therapy for depression, depressed patients were assigned to receive either cognitive therapy or behavior therapy for a 12-week period.
The researchers gave each participant a standardized depression test to assess symptom severity. They then paired participants with similar scores across the two therapy groups.
Scenario 2: Reading Comprehension by Age
To assess the difference in reading comprehension between 7 and 9-year-olds, a researcher recruited each group from a local primary school.
They were given the same passage of text to read and then asked a series of questions to assess their understanding.
Scenario 3: A Reading Intervention
Researchers wanted to compare two reading-teaching methods. A group of 5-year-olds was recruited from a primary school for the study, and their level of reading ability was assessed. They were then taught using scheme one for 20 weeks.
At the end of this period, their reading was reassessed, and a reading improvement score was calculated. They were then taught using scheme two for a further 20 weeks, and another reading improvement score for this period was calculated. The reading improvement scores for each child were then compared.
Scenario 4: Organization and Recall
To assess the effect of the organization on recall, a researcher randomly assigned student volunteers to two conditions.
Condition one attempted to recall a list of words organized into meaningful categories. Condition two attempted to recall the same words, randomly grouped on the page.
Experiment Terminology
Key terms used throughout this article are defined below.
- Ecological validity: the degree to which an investigation represents real-life experiences.
- Experimenter effects: the ways an experimenter can accidentally influence a participant through their appearance or behavior.
- Demand characteristics: clues in an experiment that lead participants to guess what the researcher is looking for, such as the experimenter’s body language.
- Independent variable (IV): the variable the experimenter manipulates, assumed to have a direct effect on the dependent variable.
- Dependent variable (DV): the variable the experimenter measures — the outcome, or result, of a study.
- Extraneous variables (EV): variables other than the IV that could affect the DV. Extraneous variables should be controlled where possible.
- Confounding variables: variable(s) that have affected the DV apart from the IV. A confounding variable is an extraneous variable that has not been controlled.
- Random allocation: assigning participants to conditions by chance, so everyone has an equal chance of any condition; this random allocation avoids bias and limits the effect of participant variables.
- Order effects: changes in performance caused by repeating a similar test, not by the IV itself.
- Practice effect: improved performance from growing familiar with the task.
- Fatigue effect: worse performance from boredom or tiredness.
- Carryover effect: a lasting change, such as a drug not fully cleared, that spills into a later condition.
Critical Evaluation
Each experimental design controls one threat to a study’s validity while leaving another active. Recognizing these trade-offs helps researchers choose the right design and judge how much to trust a study’s results.
Internal and External Validity
Independent groups designs control order effects but leave participant variables uncontrolled, since different people sit in each condition. Repeated measures designs control participant variables but introduce order effects, because the same people complete every condition.
Matched pairs designs sit in between. They control participant variables on the matched characteristics only, and, like independent groups, they avoid order effects entirely.
Choosing a design is a bet about which uncontrolled source of variance is most likely in a specific study. There is no design that is best for every study.
Within-subjects designs also create an unusual testing situation, since few real-world settings expose someone to every condition in sequence.
This trade-off between internal validity and external validity is a design-level property, not a flaw in any single study.
Randomised controlled trials illustrate the internal-validity end of this trade-off. A drug’s effect can persist in the body. Giving the same patient every condition would let one treatment carry over into the next, so independent groups is the default design in clinical trials.
Reaction-time and psychophysics research sits at the opposite end. Individual differences in speed are large and stable, and most laboratory manipulations are brief and reversible. So repeated measures with counterbalancing is standard there, trading a less natural testing sequence for statistical power (the probability of detecting a real effect).
Contemporary Research
The most-cited test of how well psychology findings hold up is the Reproducibility Project. It was a large-scale replication effort published in 2015.
Aim: To estimate what proportion of published psychology findings would replicate when the original methods were repeated with adequate statistical power (Open Science Collaboration, 2015).
Method: A team of 270 researchers repeated 100 studies from three major psychology journals, using a sample size large enough to detect the original effect with high power.
Results: Only 35 of the 97 original significant findings replicated with a significant result in the same direction. The average replication effect was about half the original size.
Conclusion: Many published findings do not survive a well-powered repeat attempt. Small samples and flexible analysis choices are more likely causes than fraud. The result reshaped how psychologists judge evidence, driving wider use of preregistration, larger samples, and multi-site replication.
Large multi-site projects have since tested how widely results generalize. Many Labs 2 (Klein et al., 2018) reran 28 classic effects across 125 samples in 36 countries. Most effects that replicated at all did so in most samples.
Many journals now expect researchers to preregister their design and analysis plan before data collection (Munafò et al., 2017). This practice limits undisclosed changes that can inflate false-positive results.
Choosing the Right Statistical Test
Experimental design is one of two factors that decide which inferential test fits the data. Level of measurement is the other. Independent-groups designs need unrelated-samples tests. Repeated-measures and matched-pairs designs need related-samples tests instead, because pairing treats the same person’s scores as linked, not independent.
| Level of measurement | Independent measures | Repeated measures / matched pairs |
|---|---|---|
| Nominal | Chi-squared (χ²) test | Sign test |
| Ordinal | Mann-Whitney U test | Wilcoxon signed-ranks test |
| Interval / ratio | Unrelated (independent) t-test | Related t-test |
The related and unrelated t-tests are parametric tests, since they use the mean and standard deviation and assume the sampling distribution is roughly normal. For more than two conditions, the parametric extension is one-way ANOVA, run in either an independent-groups or a repeated-measures form.
Key Takeaways
- Three Allocation Methods: Independent groups, repeated measures, and matched pairs are the three ways researchers assign participants to the levels of the independent variable.
- Independent Groups Trade-off: Using different people in each condition avoids order effects, but random allocation is needed to spread out participant variables like age or ability.
- Repeated Measures Trade-off: Testing the same people in every condition controls participant variables and needs fewer participants, but it risks order effects such as practice, fatigue, and carryover.
- Counterbalancing Has Limits: Reversing the order of conditions balances out practice and fatigue effects well, but it does not fully remove an asymmetric carryover effect.
- Matched Pairs Compromise: Pairing different participants on relevant characteristics avoids order effects while recovering some of the sensitivity lost by not testing the same person twice.
- Test Choice: Independent-groups data use unrelated-samples tests, while repeated-measures and matched-pairs data use related-samples tests.
- Replication Matters: A well-powered 2015 replication project found that only about a third of significant psychology findings replicated, pushing the field toward preregistration and larger samples.
References
Fisher, R. A. (1935). The design of experiments. Oliver and Boyd.
Klein, R. A., Vianello, M., Hasselman, F., Adams, B. G., Adams, R. B., Alper, S., Aveyard, M., Axt, J. R., Babalola, M. T., Bahník, Š., Batra, R., Berkics, M., Bernstein, M. J., Berry, D. R., Bialobrzeska, O., Binan, E. D., Bocian, K., Brandt, M. J., Busching, R., … Nosek, B. A. (2018). Many Labs 2: Investigating variation in replicability across samples and settings. Advances in Methods and Practices in Psychological Science, 1(4), 443–490. https://doi.org/10.1177/2515245918810225
Munafò, M. R., Nosek, B. A., Bishop, D. V. M., Button, K. S., Chambers, C. D., Percie du Sert, N., Simonsohn, U., Wagenmakers, E.-J., Ware, J. J., & Ioannidis, J. P. A. (2017). A manifesto for reproducible science. Nature Human Behaviour, 1(1), Article 0021. https://doi.org/10.1038/s41562-016-0021
Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), Article aac4716. https://doi.org/10.1126/science.aac4716


