A between-subjects study design, also called independent groups or between-participant design, allows researchers to assign test participants to different treatment groups.
In a between-subjects design, each participant is assigned to only one level of the independent variable (treatment condition), and researchers will compare group differences between participants in these various conditions.
Key Takeaways
- Definition: In a between-subjects design, each participant experiences only one level of the independent variable, and the different groups are compared against each other.
- Random Allocation: Participants are randomly assigned to conditions so that, on average, the groups start out comparable before the manipulation.
- Main Strength: Because no one repeats a condition, there are no order effects and less risk that participants guess the study’s aim.
- Main Limitation: Participant variables, pre-existing differences like age or mood, can differ between groups by chance and need a larger sample to average out.
- Vs Within-Subjects: A within-subjects design instead puts every participant through every condition, trading a bigger sample for the risk of order effects.
- Modern Standard: Since the 2015 Reproducibility Project, well-designed between-subjects studies are also expected to be adequately powered and preregistered before data collection.
How to Use
In a between-subjects design, participants are divided into separate groups. Each group experiences a different condition or level of the independent variable.
One of these groups is often a control group, receiving no treatment or a placebo. The other groups are experimental groups, each receiving a different level of the treatment or intervention.
Multi-arm studies are common too. Studies can include multiple experimental groups, each receiving a different level or type of treatment. For example, a drug-dose study might have three groups: a placebo control, a low-dose group, and a high-dose group.
The goal is simple. A between-subjects design compares the outcome (the dependent variable) across these groups, to see whether different levels of the independent variable produce significantly different results.
By manipulating the independent variable between groups, researchers can assess how effective different treatments or interventions are.
Minimizing Bias
To minimize bias, the participants should be randomly assigned to either the control group or one of the experimental conditions. They should not know which group they are assigned to.
Blinding (masking) in between-group designs prevents participants from knowing their group assignment. This reduces expectancy effects, demand characteristics, and placebo effects that could bias results.
By keeping subjects unaware of their condition, researchers can more confidently attribute any observed differences to the actual treatment effect rather than participant expectations or behavior changes.
Origins: Fisher and the Design of Experiments
The between-subjects design used in psychology today traces back to agricultural field trials at Rothamsted Experimental Station, run by the statistician Ronald A. Fisher.
In The Design of Experiments (Fisher, 1935), Fisher set out three principles that still govern between-subjects experiments today:
- Randomisation: allocating units to conditions by chance, so nuisance factors become part of the error term rather than a systematic bias.
- Replication: assigning several units to each condition so within-condition variance can be estimated.
- Local control (blocking): grouping similar units together and randomising within each group, removing a known nuisance variable from the error term.
Fisher had introduced this tool a decade earlier. His Statistical Methods for Research Workers (Fisher, 1925) presented the analysis of variance (ANOVA). ANOVA separates a study’s total variance into the part caused by the independent variable and the part caused by ordinary differences between people.
Independent-groups experiments in psychology are, structurally, Fisher’s completely randomised designs. The vocabulary survives too. Terms like “condition,” “treatment,” and “error term” all come from this same tradition.
Example
To test whether a new meditation app (your independent variable) can reduce anxiety levels (your dependent variable), you gather a sample of 100 participants who report high levels of anxiety.
You use a between-subjects design to divide the sample into two groups:
- A control group where the participants are instructed to continue their daily routines without using the meditation app,
- An experimental group where the participants are instructed to use the meditation app for 20 minutes daily for 4 weeks.
Before and after the 4-week period, you administer an anxiety assessment to all participants.
Then, you compare the change in anxiety levels between the two groups using statistical analysis to determine if the meditation app had a significant effect on reducing anxiety.
Research Studies
- Baeyens, Diaz, & Ruiz (2005) investigated the resistance to extinction of evaluative conditioning using separate groups of participants, which is a between-subjects approach.
- Ehrlichman et al. (1997) studied the modulation of startle reflex by pleasant and unpleasant odors, comparing the effects between different groups of subjects.
- Carey, Lester, & Valencia (2016) examined the effects of a fatal vision goggles intervention on attitudes toward drinking and driving and texting and driving among middle school children, using a between-subjects design to compare the intervention group to a control group.
- Chang and Kang (2018) investigated the impact of the 2018 North Korea-United States Summit on South Koreans’ altruism toward and trust in North Korean refugees, using a between-subjects design to compare responses before and after the summit.
- Egele, Kiefer, & Stark (2021) compared the faking of self-reported health behavior between a within-subjects and a between-subjects design, highlighting the use of both approaches in their study.
Between-subjects vs within-subjects design
In a between-subjects design, different groups of participants are exposed to different conditions, and the results are compared between these groups.
In contrast, a within-subjects design exposes each participant to all conditions, and the results are compared within the same group of participants.
The pretest in a within-subjects design serves as a baseline, similar to a control condition, while the posttest assesses the effects of the independent variable treatments.
Example: Between-subjects vs within-subjects design
You’re planning to study whether listening to classical music (your independent variable) while studying can improve memory retention (your dependent variable).
You can use either a between-subjects or a within-subjects design.
If you use a between-subjects design, you would split your sample into two groups of participants:
- A control group that studies in silence for 30 minutes
- An experimental group that studies while listening to classical music for 30 minutes
Then, you would administer the same memory test to all participants and compare the scores between the groups.
If you use a within-subjects design, everyone in your sample would undergo the same procedures:
- First, they would all study a list of words in silence for 30 minutes and take a memory test.
- After a break, they would study a new list of words while listening to classical music for 30 minutes.
- Finally, they would take another memory test on the second list of words.
You would compare the memory test scores from the silent condition and the classical music condition statistically.
These two types of designs can also be combined in a single study when you have two or more independent variables.
In factorial designs, multiple independent variables are tested simultaneously. Each level of one independent variable is combined with each level of every other independent variable to create different conditions.
For example, you could study the effects of both music (classical vs. no music) and study environment (library vs. café) on memory retention. This would create four conditions:
- Classical music in library
- Classical music in café
- No music in library
- No music in café
In a mixed factorial design, one variable is altered between subjects and another is altered within subjects.
For instance, you could have two groups of participants (between-subjects: classical music vs. no music) who each study in both the library and the café (within-subjects: study environment).
Advantages
Eliminates order effects
Order effects refer to the influence of the sequence or order in which conditions are presented on the results.
Here’s why order matters. In within-subjects designs, researchers ask the same participants to complete every condition, so order can distort results in three ways.
Practice effects make later conditions look better simply because participants have had more practice. Fatigue and boredom effects make later conditions look worse as participants tire or lose motivation.
Carryover is the trickiest. It happens when the experience of one condition lingers into the next, for example when a strategy learned in one task gets reused in another.
Within-subjects researchers control this with counterbalancing: varying the order conditions appear in so any order effect spreads evenly rather than piling up in one condition. Between-subjects designs sidestep the problem entirely.
Because each participant only ever experiences one condition, there is no order to control for in the first place.
Avoids carryover effect
Carryover effects refer to the influence of one experimental condition on a participant’s behavior or responses in a subsequent condition.
For example, learning a new skill in one condition might “carry over” and boost performance in a later condition. This happens even if that later condition was never designed to teach the skill.
Between-subjects designs avoid this. Because participants are split into separate treatment groups, one participant’s exposure will not affect the outcome of someone else’s condition.
That said, between-subjects designs don’t eliminate every type of carryover. Spillover between groups is still possible if participants in different conditions interact and share information.
Reduced Demand Characteristics
A participant who experiences only one condition has less to compare it against. This makes it harder for them to guess what the researcher is testing, or to change their behavior to please or thwart the study.
This benefit extends to how the design itself is built. Between-subjects layouts also extend naturally to many-condition factorial designs and drug-versus-placebo trials. They fit single- or double-blind studies too, since one participant logically cannot sit in more than one arm anyway.
Short and straightforward
Each participant is only assigned to one treatment group, so the experiments tend to be uncomplicated. Scheduling the testing groups is simple, and researchers tend to be able to receive and analyze the data quickly.
Reduced testing fatigue
Each participant is only tested in one condition. This helps between-subjects designs avoid the testing fatigue that can build up in within-subjects designs, where participants work through multiple conditions in a single session.
Limitations
A large participant pool is necessary
Because each subject is assigned to only one condition, this type of design requires a large sample. Thus, these studies also require more resources and budgeting to recruit participants and administer the experiments.
Individual differences
Differences between subjects within a given condition may be an explanation for results, introducing error and making the effects of an experimental condition less accurate.
Requires careful matching or random assignment
To help control for individual differences between groups, researchers must carefully match participants on key characteristics or use random assignment to conditions.
If groups differ from the outset, it can confound the results.
There’s a subtlety here. Random assignment does not guarantee the groups are matched on any single variable. What it guarantees is that any pre-existing differences are, on average, unrelated to condition assignment.
It matters. Skip random allocation and the study becomes a quasi-experiment. Any difference in the results could then be explained by the pre-existing grouping rather than the treatment.
Less statistical power
For the same sample size, between-subjects designs have less statistical power than within-subjects designs.
This means larger effect sizes are needed to detect significant differences between conditions, or larger sample sizes are required.
The reason is individual differences. In a within-subjects design, each participant is compared only to themselves, so stable personal traits like IQ or baseline mood cancel out of the comparison. In a between-subjects design, those same differences stay mixed into the error term, making a real effect harder to detect.
Critical Evaluation
Between-subjects design is not automatically the right choice for every study. It trades one set of strengths for another, and each trade-off is worth examining before you commit to it.
Internal and External Validity
A between-subjects design controls order effects well. It does not control participant variables as tightly as a repeated-measures design does, because different people sit in each condition.
This is the classic internal validity trade-off. No single design controls every threat at once, so choosing one is a bet about which source of error matters most in this specific study (Campbell & Stanley, 1963).
The trade-off runs both ways. Between-subjects studies also tend to preserve external validity better than repeated-measures studies.
Because each participant only ever encounters one condition, the testing situation looks more like a single real-world exposure. A repeated-measures participant, by contrast, works through an artificial sequence of multiple treatments (Campbell & Stanley, 1963).
Matched pairs sit in between the two extremes. They still use different people, but each is matched to their partner on traits chosen to resemble a repeated-measures comparison.
Contemporary Research
The most consequential recent development in experimental design has been a shift toward asking whether findings replicate, not just whether they were originally significant.
- Aim: To estimate what proportion of published psychology findings would replicate when the original methods were repeated with adequate statistical power.
- Method: The Open Science Collaboration (Nosek et al., 2015) recruited 270 contributing authors to replicate 100 studies from three top psychology journals. Each replication reused the original design, powered to detect the original effect size.
- Results: Only 35 of 97 original significant findings replicated in the same direction. The average replicated effect size was about half the size first reported.
- Conclusion: A well-designed independent-groups, repeated-measures, or matched-pairs study is not, on its own, enough evidence that an effect is real. Design must be paired with adequate power, a pre-specified analysis plan, and independent replication.
Choosing a between-subjects design does not remove this problem. For a fixed number of participants, between-subjects studies generally have less statistical power than repeated-measures studies. An under-powered study is especially likely to produce a significant result that overstates the true effect.
Researchers increasingly preregister their design and analysis plan before collecting data. This is now standard practice. It fixes the conditions, sample size, and analysis in advance, so they cannot be adjusted once the results are seen (Munafò et al., 2017).
Frequently Asked Questions
What’s the difference between a within-subjects versus a between-subjects design?
Between-subjects and within-subjects designs are two different methods for researchers to assign test participants to different treatments.
Researchers will assign each subject to only one treatment condition in a between-subjects design. In contrast, in a within-subjects design, researchers will test the same participants repeatedly across all conditions.
Between-subjects and within-subjects designs can be used in place of each other or in conjunction with each other.
Each type of experimental design has its own advantages and disadvantages, and it is usually up to the researchers to determine which method will be more beneficial for their study.
Can you use a between-subjects and within-subjects design in the same study?
Yes. Between-subject and within-subject designs can be combined in a single study when you have two or more independent variables (a factorial design).
Factorial designs are a type of experiment where multiple independent variables are tested. Each level of one independent variable (a factor) is combined with each level of every other independent variable to produce different conditions.
Each combination becomes a condition in the experiment. In a factorial experiment, the researcher has to decide for each independent variable whether to use a between-subjects design or a within-subjects design.
In a mixed factorial design, researchers will manipulate one independent variable between subjects and another within subjects.
What is between subject factorial design?
A between-subject factorial design is an experimental setup where participants are randomly assigned to different levels of two or more independent variables.
This design allows researchers to examine the individual effects of each independent variable and their interaction effect on the dependent variable, while each participant is exposed to only one combination of conditions.
References
Allen, M. (2017). The sage encyclopedia of communication research methods (Vols. 1-4). Thousand Oaks, CA: SAGE Publications, Inc doi: 10.4135/9781483381411
Baeyens, F., Díaz, E., & Ruiz, G. (2005). Resistance to extinction of human evaluative conditioning using a between-subjects design. Cognition & Emotion, 19(2), 245–268. https://doi.org/10.1080/02699930441000300
Birnbaum, M. H. (1999). How to show that 9> 221: Collect judgments in a between-subjects design. Psychological Methods, 4(3), 243.
Campbell, D. T., & Stanley, J. C. (1963). Experimental and quasi-experimental designs for research. Rand McNally.
Carey, A. A., Lester, T. G., & Valencia, R. M. (2016). The Effects of a Fatal Vision Goggles Intervention on Middle School Aged Children’s Attitudes toward Drinking and Driving and Texting and Driving as Related to Impulsivity: A Between Subjects Design (Doctoral dissertation, Brenau University).
Chang, H. I., & Kang, W. C. (2018). The Impact of the 2018 North Korea-United States Summit on South Koreans’ Altruism Toward and Trust in North Korean Refugees: Between-Subjects Design Around the Summit. Available at SSRN 3270334.
Egele, V. S., Kiefer, L. H., & Stark, R. (2021). Faking self-reports of health behavior: a comparison between a within-and a between-subjects design. Health psychology and behavioral medicine, 9(1), 895-916.
Ehrlichman, H., Brown Kuhl, S., Zhu, J., & Wrrenburg, S. (1997). Startle reflex modulation by pleasant and unpleasant odors in a between-subjects design. Psychophysiology, 34(6), 726–729. https://doi.org/10.1111/j.1469-8986.1997.tb02149.x
Fisher, R. A. (1925). Statistical methods for research workers. Oliver and Boyd.
Fisher, R. A. (1935). The design of experiments. Oliver and Boyd.
Jhangiani, R. S., Chiang, I.-C. A., Cuttler, C., & Leighton, D. C. (2019, August 1). Experimental Design. Research Methods in Psychology. Retrieved from https://kpu.pressbooks.pub/psychmethods4e/
Munafò, M. R., Nosek, B. A., Bishop, D. V. M., Button, K. S., Chambers, C. D., Percie du Sert, N., Simonsohn, U., Wagenmakers, E.-J., Ware, J. J., & Ioannidis, J. P. A. (2017). A manifesto for reproducible science. Nature Human Behaviour, 1(1), Article 0021. https://doi.org/10.1038/s41562-016-0021
Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), Article aac4716. https://doi.org/10.1126/science.aac4716