In a within-subject design, each participant experiences all experimental conditions, whereas, in a between-subject design, different participants are assigned to each condition, with each experiencing only one condition.

Within-subjects (or repeated-measures) is an experimental design in which all study participants are exposed to the same treatments or independent variable conditions.
In within-subjects studies, each participant is compared to themselves, not to other participants. There is no separate control group. Each participant’s score in one condition becomes the baseline for comparing their score in the others.
In a between-subjects design (or between-groups, independent measures), the study participants are divided into groups, and each group is exposed to one treatment or condition.
Each participant is only assigned to a single treatment. This should be done by random allocation, ensuring that each participant has an equal chance of being assigned to one group.
The differences between the two groups are then compared to a control group that does not receive any treatment. The groups that undergo a treatment or condition are typically called the experimental groups.
To recap:
- In a within-subjects design, all participants receive every treatment.
- In a between-subjects design, participants only receive one treatment.
Design Similarities
- Both designs assess how a treatment or condition affects a study population, comparing several conditions within a single study.
- Both use a group of study participants who are exposed to the treatment or condition being tested.
- Both experimental designs are utilized in quantitative studies and aim to result in findings that are statistically likely to generalize to a whole population.
- Between-subjects and within-subjects design both have an independent variable that is manipulated or controlled by the study’s investigators and a dependent variable that is measured.
- Random allocation matters in both designs, but what gets randomized differs: between-subjects designs randomly assign participants to groups, while within-subjects designs randomize the order in which each participant completes the conditions.
Design Differences
- In a within-subjects design, all participants receive all treatments. In a between-subjects design, participants receive only one treatment.
- In a within-subjects design, each participant is compared to themselves across conditions, so there is no control group. In a between-subjects design, there is a control group that doesn’t receive any treatment and serves as a source of comparison for the treatment groups.
- Between-subjects designs need significantly more participants than within-subjects designs to detect a statistically significant difference between conditions.
Within-subjects designs need fewer. Each person provides a data point for every level of the independent variable.
A between-subjects experiment typically needs twice as many participants as an equivalent within-subjects study. That means more resources and funding to recruit a larger sample, run sessions, and cover costs. - Between-subjects designs tend to be easier and quicker to administer as each participant is only given one treatment. In contrast, within-subjects designs take longer to implement because every participant is given multiple treatments.
- Within-subjects designs are vulnerable to three order effects: practice, fatigue, and carryover.
Practice effects occur when participants improve simply from repeating the task. Fatigue effects occur when participants tire or lose motivation after multiple treatments.
Carryover is the most serious. It happens when one condition’s effects persist into the next, changing performance regardless of order. Counterbalancing controls all three by varying the order of conditions across participants.
Between-subjects designs avoid this problem entirely, since each participant experiences only one condition. - Because different participants provide data for each condition in between-subjects designs, individual differences among participants may threaten internal validity.
Within-subjects designs avoid this problem. Each participant is compared only to themselves, so individual differences cannot confound the result. This typically gives within-subjects designs higher statistical power: the probability that a study detects a real effect as statistically significant. Comparing each participant to their own scores removes individual differences from the error the treatment effect is judged against.
Order Effects and Counterbalancing
Giving the same participant every condition creates order effects: changes in a score caused by a condition’s position in the sequence rather than by the treatment itself.
Researchers control order effects with counterbalancing, a systematic variation of condition order that spreads any effect evenly across conditions instead of letting it favor one treatment.
Practice, Fatigue, and Carryover Effects
- Practice effects: Participants improve simply from repeating the task, growing familiar with the apparatus or timing, so scores in later conditions can rise for reasons unrelated to the treatment.
- Fatigue and boredom effects: Participants tire or lose motivation across a long session, so scores in later conditions can fall regardless of the treatment.
- Carryover effects: A treatment received in one condition changes the participant in a way that persists into the next, such as a drug not yet cleared from the body. Carryover is the most serious of the three because it can be stronger in one order than the other, and does not average out.
For example, in a memory study with two word lists, scores on the second list may rise simply because participants have practiced the task. This has nothing to do with the condition itself.
Repeated exposure to every condition also raises the risk of demand characteristics. A participant who sees more than one condition can more easily guess the study’s hypothesis and adjust their behavior to fit it.
Counterbalancing Methods
Three counterbalancing schemes are common in within-subjects research.
- ABBA counterbalancing: Each participant receives conditions in the order A, B, B, A (or B, A, A, B), so a steady practice effect cancels out because A and B occupy the same average position in the sequence.
- Between-subject counterbalancing: Half the sample receives A then B, and the other half receives B then A, so the order effect becomes ordinary variance within each condition rather than a bias favoring one of them.
- Latin square designs: With more than two conditions, a Latin square arranges the orders so each condition appears once in every position across the full set of participants, balancing every condition across every position.
Counterbalancing controls practice and fatigue effects well because they build up gradually across a session. It cannot fully correct a carryover effect that is asymmetric or builds quickly, such as a mood induction that has not faded.
When this kind of carryover is likely, researchers space conditions apart with a washout period or switch to a between-subjects design instead.
Critical Evaluation
Both within-subjects and between-subjects designs trade one weakness for another. Neither design is right for every study, and the correct choice depends on which threat to validity looms largest in a particular experiment.
Internal and External Validity
Each design controls one threat to internal validity while leaving another active. Between-subjects designs remove order effects but leave participant variables uncontrolled. Within-subjects designs control participant variables but introduce order effects instead.
Neither weakness disappears. Each design simply trades one for the other.
External validity runs the other way. A within-subjects study exposes participants to every condition, a sequence most real-world situations don’t resemble. Between-subjects studies preserve a more natural, one-off exposure, at the cost of needing more participants.
Ethics matters too. Some treatments cannot ethically be reversed or repeated in the same person, such as surgery or a one-off deception. These constraints rule out within-subjects testing regardless of its statistical advantages.
Independent-groups designs are also the only option for naturally distinct populations, such as clinical versus non-clinical groups. No participant can belong to both groups at once.
Modern reporting standards such as CONSORT and JARS require researchers to name the design explicitly and describe how participants were allocated. A study whose design is under-specified cannot be properly evaluated or replicated.
Contemporary Research
The most consequential recent shift in experimental design has been a new focus on replicability: whether a finding holds up when the study is repeated.
- Aim: To estimate what proportion of published psychology findings would replicate when the original study was repeated with adequate statistical power.
- Method: A team of 270 researchers repeated 100 studies from three major psychology journals, using each study’s original design and a sample size powered to detect the original effect.
- Results: Only 35 of 97 originally significant findings replicated significantly in the same direction, and the average replication effect was about half the original size.
- Conclusion: The project reframed design choices around replicability rather than statistical significance alone, and helped drive reforms such as preregistration and adequately powered sample sizes (Open Science Collaboration, 2015).
Choosing a well-powered design, and reporting it transparently, is now treated as part of getting the result right, not just detecting one. This applies equally to independent-groups and repeated-measures designs.
References
- Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), Article aac4716. https://doi.org/10.1126/science.aac4716
FAQs
What is a 2×2 within subject design?
A 2×2 within-subjects design is one in which there are two independent variables each having two different levels. This design allows researchers to understand the effects of two independent variables (each with two levels) on a single dependent variable.
When would you use a within-subjects design?
You typically use a within-subjects design when the manipulation is brief and reversible, such as a short cognitive or perceptual task. It also suits studies with large individual differences, like memory span or reaction time.
Fewer participants are needed. Because each person provides data in every condition, a within-subjects design needs fewer participants than a between-subjects design to detect the same effect.
When should a within-subjects design not be used?
A within-subjects design should not be used when the treatment could have a lasting or irreversible effect on participants, such as surgery or a one-shot deception.
Practice effects are also a concern. If earlier conditions teach participants the task, later scores rise for reasons unrelated to the treatment. Counterbalancing cannot fully separate a real treatment effect from this kind of improvement.
When should you use a between-subjects design?
Between-subjects designs are used when a treatment could permanently change participants, such as surgery or a one-shot deception, which rules out testing the same person twice.
They also suit studies with a large participant pool, or ones comparing naturally distinct groups, such as a clinical and a non-clinical sample.
When can a between-subjects design not be used?
Between-subjects cannot be used with small sample sizes because they will not be statistically powerful enough.
Between-subjects studies require at least twice as many participants as a within-subject design, which also means twice the cost and resources. When funding is limited, between-subjects design can likely not be used.
Can I use a within- and between-subjects design in the same study?
Yes. Between-subjects and within-subjects designs can be combined in a single study when you have two or more independent variables (a factorial design).
Factorial designs are a type of experiment where multiple independent variables are tested.
Each level of one independent variable (a factor) is combined with each level of every other independent variable to produce different conditions.
Is between-subjects or within-subjects design more powerful?
Within-subjects designs have more statistical power due to the lack of variation between the individuals in the study because participants are compared to themselves.
A between-subjects design would require a large participant pool in order to reach a similar level of statistical significance as a within-subjects design.
Key Takeaways
- Group Assignment: In within-subjects designs the same participants complete every condition; in between-subjects designs different participants are randomly assigned to just one.
- Order Effects: Within-subjects designs risk practice, fatigue, and carryover effects, controlled with counterbalancing.
- Participant Variables: Between-subjects designs risk uneven individual differences across groups; random allocation is the main safeguard.
- Statistical Power: Within-subjects designs generally need fewer participants because each person acts as their own control, reducing error variance.
- Sample Size: Between-subjects designs typically need roughly twice as many participants to detect the same effect.
- Mixed Designs: The two approaches can be combined in a single factorial study with two or more independent variables.
- Replicability: Large-scale replication projects now push researchers to plan design and power so a finding holds up when repeated, not just when first found.