A controlled experiment is a research method in which a researcher manipulates one variable, the independent variable, while holding all other conditions constant. This tests a hypothesis about its effect.
In a controlled experiment, an independent variable (the cause) is systematically manipulated, and the dependent variable (the effect) is measured; any extraneous variables are controlled.
The researcher can operationalize (i.e., define) the studied variables so they can be objectively measured. The quantitative data can be analyzed to see if there is a difference between the experimental and control groups.

Key Takeaways
- Control: Comparing an experimental group against a control group is what licenses a cause-and-effect claim, ruling out rival explanations for a change in the dependent variable.
- Random Allocation: Assigning participants to conditions purely by chance spreads individual differences evenly, a principle formalised by the statistician Ronald Fisher in 1935.
- Order Effects: Repeated-measures designs risk practice and fatigue effects; counterbalancing (e.g., an ABBA design) spreads these evenly across conditions rather than eliminating them.
- Standardisation: Keeping instructions, materials and the environment constant controls extraneous variables and makes a study replicable.
- Modern Evidence: Trials that skip randomisation or blinding report measurably inflated effects, especially for subjectively judged outcomes (Page et al., 2016).
- Trade-Off: Tight control buys internal validity but can cost ecological validity, which is one reason quasi-experiments exist alongside true experiments.
What is the control group?
In experiments scientists compare a control group and an experimental group that are identical in all respects, except for one difference – experimental manipulation.
Unlike the experimental group, the control group is not exposed to the independent variable under investigation. It provides a baseline: the comparison that shows what would have happened without the manipulation.
Since experimental manipulation is the only planned difference between the experimental and control groups, any gap between them is attributed to that manipulation rather than to chance.
This is what makes the comparison meaningful. It licenses a cause-and-effect claim, not just a correlation.
Randomly allocating participants to independent variable groups means that all participants have an equal chance of ending up in either condition. This differs from random sampling. Random sampling concerns how participants are recruited from the wider population; random allocation concerns how they are then sorted into conditions.
The principle of random allocation is to avoid bias in how the experiment is carried out and limit the effects of participant variables. These are pre-existing differences between individuals, such as personality, mood or experience, that could otherwise confound the comparison between conditions.
Ronald Fisher formalised the principle in 1935. His book, The Design of Experiments, proposed randomisation as a defence against confounds researchers had not thought to measure (Fisher, 1935).
Placebo Controls
Not every control condition works the same way. A negative control is expected to produce no effect. It confirms that any change seen elsewhere is not just an artefact of the procedure.
Most simple psychology experiments use this negative-control logic: a no-treatment or “as normal” condition.
A placebo control tackles a subtler problem. Participants often show real change simply because they believe they have received a treatment, even when they have not.
An inert, inactive procedure that looks identical to the real one, such as a sugar pill, isolates this expectation effect from the treatment’s actual active ingredient.
Aim: Henry Beecher, an anaesthesiologist, aimed to establish how large and reliable placebo responses actually were. He saw this pattern often in clinical practice.
Method: Beecher (1955) reviewed and pooled 15 published clinical studies. They covered conditions including post-operative pain, seasickness, headache, cough and anxiety, in which patients had received a placebo alongside an active treatment.
Results: Placebos produced satisfactory relief in roughly 35% of patients on average across the pooled studies. Some studies found a comparable effect size.
Conclusion: Beecher concluded the placebo response was real and measurable, not simply “no effect.” He used the finding to argue for including placebo control groups in clinical trials.
Evaluation: The precise 35% figure has since been challenged. Beecher’s pooled studies mixed genuine placebo response with natural recovery and reporting bias. Even so, his core argument still stands: an apparent treatment effect cannot be trusted without a placebo comparison.
What are extraneous variables?
The researcher wants to be sure that any change in the dependent variable was caused by the manipulation of the independent variable, and not by something else.
Hence, all the other variables that could affect the dependent variable to change must be controlled. These other variables are called extraneous or confounding variables.
Extraneous variables should be controlled were possible, as they might be important enough to provide alternative explanations for the effects.
In practice, it would be difficult to control all the variables in a child’s educational achievement. For example, it would be difficult to control variables that have happened in the past.
A researcher can only control the current environment of participants, such as time of day and noise levels.
Why conduct controlled experiments?
Scientists use controlled experiments because they allow for precise control of extraneous and independent variables. This allows a cause-and-effect relationship to be established.
Controlled experiments also follow a standardized step-by-step procedure. This makes it easy for another researcher to replicate the study.
Applications of Controlled Experiments
The logic covered above is not confined to the psychology laboratory. The same core techniques, random allocation and a genuine control baseline, are used to test claims across medicine, education, technology and sport.
Medicine and Clinical Trials
The randomised, placebo-controlled, double-blind trial is the regulatory standard for testing whether a new drug or psychological therapy genuinely works. It combines every technique covered above in one design.
Random allocation equalises participant characteristics between groups. A placebo arm isolates the drug’s real effect from patients’ expectations, rather than from the belief that something has been given. Double-blinding then stops both patient hope and clinician enthusiasm from inflating the result.
Without these controls, an ineffective treatment could be licensed on the strength of expectation alone. Beecher (1955) demonstrated this at scale.
Kobak et al.’s (2005) double-blind trial of an alternative remedy for obsessive-compulsive disorder applied the same logic. Comparing the remedy against a placebo revealed no real difference between them.
Education
Randomised controlled trials have been argued for as a way of testing new teaching methods and curricula. The goal is the same causal rigour as a drug trial, rather than an uncontrolled classroom comparison (Torgerson & Torgerson, 2001).
A single school’s before-and-after comparison cannot rule out maturation, or a particularly motivated teacher, as the real explanation for any improvement. Randomly allocating classes or pupils removes that ambiguity.
This is now standard practice in large-scale education evaluation. It brings real practical difficulties, though. Schools, not individual pupils, are usually the unit that can realistically be randomised.
A “control” class can also be contaminated. Its teacher may have already picked up ideas from the new intervention being trialled elsewhere in the same school. Whole-school effects like this are a distinctive challenge for educational, rather than laboratory, control.
Workplace and Technology (A/B Testing)
The same logic now runs at enormous scale online. Large technology companies routinely allocate visitors at random to see one version of a webpage, feature or price against another, an “A/B test.”
Outcomes such as click-through or purchase rates are then compared between the two groups. Kohavi et al. (2013) describe organisations running thousands of such randomised online experiments every year.
Random allocation removes selection bias between visitors. A genuine “as before” control arm supplies the baseline for comparison.
Together, these let a company credit any difference in outcome to the change itself, rather than to which kind of customer happened to visit at which time. It is the same control logic covered above, now running at web scale across millions of visitors rather than dozens of lab participants.
Sport and Applied Performance
Control logic is harder to apply outside the laboratory or clinic. An elite athlete cannot ethically or practically be “randomly withheld” from a training method a coach believes works, the way a patient can be randomised to a placebo.
This is why sport science leans on within-athlete designs. The same athlete performs under multiple training conditions, rather than being compared against a separate control group.
Counterbalancing still matters here, exactly as it does in a repeated-measures laboratory study. It controls for order effects across training conditions.
Large randomised between-groups trials, of the kind common in clinical medicine, are used far less often in this field. Repeated-measures designs with counterbalancing are the more practical, and more common, alternative, letting each athlete act as their own baseline for comparison.
Key Terminology
The terms below define the core building blocks of a controlled experiment.
- Experimental Group: The group exposed to the manipulation under investigation, the independent variable. Their scores on the dependent variable are compared against the control group’s.
- Control Group: The baseline group not exposed to the independent variable, used for comparison against the experimental group.
- Ecological Validity: The extent to which a study’s findings generalise beyond the artificial conditions in which they were obtained, to real-life settings and behaviour. Tight experimental control protects internal validity, sometimes at the cost of ecological validity.
- Independent Variable (IV): The variable the experimenter manipulates (i.e., changes), assumed to have a direct effect on the dependent variable.
- Dependent Variable (DV): The variable the experimenter measures. This is the outcome, or result, of a study.
- Extraneous Variables (EV): All variables that are not the independent variable but could affect the results (DV) of the experiment. Extraneous variables should be controlled where possible.
- Confounding Variables: Variable(s) that have affected the results (DV), apart from the IV. A confounding variable is usually an extraneous variable that was not successfully controlled.
- Random Allocation: Assigning participants to conditions purely by chance, so individual differences are spread evenly rather than concentrated in one group.
Experimenter Effects
Ways an experimenter can unintentionally influence a participant, through appearance, behaviour or expectations. Also known as the experimenter (investigator) effect.
Aim: Rosenthal and Fode aimed to test whether an experimenter’s own expectations could influence a study’s results. No deliberate intent to bias the data was required.
Method: Student experimenters ran genetically identical rats through a maze-learning task. Half were falsely told their rats were bred to be fast learners, the other half that theirs were slow learners (Rosenthal & Fode, 1963).
Results: Rats whose handlers believed them fast learners learned the maze significantly faster. The animals did not differ in ability at all.
Conclusion: An experimenter’s expectations can unintentionally shape a study’s outcome, likely through subtle differences in handling, timing or recording. This gave direct evidence for the experimenter expectancy effect.
Demand Characteristics
Cues within an experimental situation, such as the setting, the instructions or the researcher’s manner. They try to guess the researcher’s hypothesis. They adjust their behaviour to fit it, a pattern Martin Orne named demand characteristics.
Aim: Orne aimed to show that participants are not passive responders to the IV, but active interpreters of the whole experimental situation.
Method: Orne (1962) drew on his own laboratory experience, including studies where participants performed pointlessly tedious tasks. He also reviewed the era’s social psychology research.
Results: Participants routinely tried to work out the experimenter’s hypothesis. They then adjusted their behaviour to confirm it. Using quasi-control groups, who imagined how they would respond without actually taking part, Orne showed many “results” from genuine experiments could be reproduced by demand characteristics alone.
Conclusion: An experiment that lets participants infer its hypothesis risks measuring compliance with that guess, not the real effect of the IV. This is why researchers use deliberate procedures, including blinding, to strip these cues away.
Order Effects
Order effects are changes in a participant’s performance caused by the order they complete conditions in, not by the independent variable itself. Two effects commonly appear:
- Practice effect: improved performance on a task due to repetition, for example because of growing familiarity with it.
- Fatigue effect: declining performance due to repetition, for example because of boredom or tiredness.
Counterbalancing, such as an ABBA design, spreads these effects evenly across conditions. It stops them from favouring whichever condition participants complete first.
Critical Evaluation of Controlled Experiments
For decades, the case for randomisation, blinding and placebo control rested mainly on logic: without them, a rival explanation cannot be ruled out. Newer research asks a more direct question. It also exposes real limits to what control can deliver.
Contemporary Research
The strongest evidence comes from Page et al. (2016), a systematic review pooling 24 meta-epidemiological studies. Each had compared trial results with and without a given design feature.
Trials with inadequate randomisation exaggerated effects by roughly 7% on average. Inadequate allocation concealment exaggerated effects by roughly 10%.
The distortion was far larger for blinding. Unblinded trials exaggerated subjectively judged outcomes, such as pain or symptom ratings, by around 23% on average. Objective outcomes like mortality showed little to no exaggeration.
The pattern is not new. It converges with older, smaller demonstrations of demand characteristics (Orne, 1962) and experimenter expectancy (Rosenthal & Fode, 1963), described above.
The finding also sits inside a wider reckoning. A large project directly replicated 100 published psychology studies. Only 36% reproduced a significant effect in the original direction, with the average effect roughly half its original size (Open Science Collaboration, 2015).
A subsequent manifesto responded directly to this. It argued for pre-registration, larger samples, blinded analysis and routine replication (Munafò et al., 2017).
Limits of Control
Randomisation is a necessary but not a sufficient safeguard. Krause and Howard (2003) showed it guarantees group equivalence only on average, across many hypothetical repeats of a study.
It does not guarantee that any one particular trial’s groups are actually equivalent. A randomised result licenses a claim about the average participant. It is not a guarantee about any specific one.
Tight control also buys internal validity at the potential cost of ecological validity. Blinding raises real ethical constraints too, not just practical ones.
A double-blind trial of an alternative remedy for obsessive-compulsive disorder illustrates the tension directly. Patients consented to a design in which, unknown to them, roughly half received no active treatment for its duration (Kobak et al., 2005).
Together, these strands move the case for control forward. It is no longer just “good practice”; it measurably changes the answer a study gets, while also showing why perfect control is neither fully achievable nor free of trade-offs.
References
Beecher, H. K. (1955). The powerful placebo. Journal of the American Medical Association, 159(17), 1602–1606. https://doi.org/10.1001/jama.1955.02960340022006
Fisher, R. A. (1935). The design of experiments. Oliver & Boyd.
Kobak, K. A., Taylor, L. V., Bystritsky, A., Kohlenberg, C. J., Greist, J. H., Tucker, P., Warnock, J., & Vapnik, T. (2005). St John’s wort versus placebo in obsessive-compulsive disorder: Results from a double-blind study. International Clinical Psychopharmacology, 20(6), 299–304. https://doi.org/10.1097/00004850-200511000-00003
Kohavi, R., Deng, A., Frasca, B., Walker, T., Xu, Y., & Pohlmann, N. (2013). Online controlled experiments at large scale. Proceedings of the 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1168–1176. https://doi.org/10.1145/2487575.2488217
Krause, M. S., & Howard, K. I. (2003). What random assignment does and does not do. Journal of Clinical Psychology, 59(7), 751–766. https://doi.org/10.1002/jclp.10170
Munafò, M. R., Nosek, B. A., Bishop, D. V. M., Button, K. S., Chambers, C. D., Percie du Sert, N., Simonsohn, U., Wagenmakers, E.-J., Ware, J. J., & Ioannidis, J. P. A. (2017). A manifesto for reproducible science. Nature Human Behaviour, 1, Article 0021. https://doi.org/10.1038/s41562-016-0021
Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. https://doi.org/10.1126/science.aac4716
Orne, M. T. (1962). On the social psychology of the psychological experiment: With particular reference to demand characteristics and their implications. American Psychologist, 17(11), 776–783. https://doi.org/10.1037/h0043424
Page, M. J., Higgins, J. P. T., Clayton, G., Sterne, J. A. C., Hróbjartsson, A., & Savović, J. (2016). Empirical evidence of study design biases in randomized trials: Systematic review of meta-epidemiological studies. PLOS ONE, 11(7), e0159267. https://doi.org/10.1371/journal.pone.0159267
Rosenthal, R., & Fode, K. L. (1963). The effect of experimenter bias on the performance of the albino rat. Behavioral Science, 8(3), 183–189. https://doi.org/10.1002/bs.3830080302
Torgerson, C. J., & Torgerson, D. J. (2001). The need for randomised controlled trials in educational research. British Journal of Educational Studies, 49(3), 316–328. https://doi.org/10.1111/1467-8527.t01-1-00178
FAQs
What is the control in an experiment?
In an experiment, the control is a standard or baseline group not exposed to the experimental treatment or manipulation. It serves as a comparison group to the experimental group, which does receive the treatment or manipulation.
The control group helps to account for other variables that might influence the outcome, allowing researchers to attribute differences in results more confidently to the experimental treatment.
Establishing a cause-and-effect relationship between the manipulated variable (independent variable) and the outcome (dependent variable) is critical in establishing a cause-and-effect relationship between the manipulated variable.
What is the purpose of controlling the environment when testing a hypothesis?
Controlling the environment when testing a hypothesis aims to eliminate or minimize the influence of extraneous variables. These variables other than the independent variable might affect the dependent variable, potentially confounding the results.
By controlling the environment, researchers can ensure that any observed changes in the dependent variable are likely due to the manipulation of the independent variable, not other factors.
This enhances the experiment’s validity, allowing for more accurate conclusions about cause-and-effect relationships.
It also improves the experiment’s replicability, meaning other researchers can repeat the experiment under the same conditions to verify the results.
Why are hypotheses important to controlled experiments?
Hypotheses are crucial to controlled experiments because they provide a clear focus and direction for the research. A hypothesis is a testable prediction about the relationship between variables.
It guides the design of the experiment, including what variables to manipulate (independent variables) and what outcomes to measure (dependent variables).
The experiment is then conducted to test the validity of the hypothesis. If the results align with the hypothesis, they provide evidence supporting it.
The hypothesis may be revised or rejected if the results do not align. Thus, hypotheses are central to the scientific method, driving the iterative inquiry, experimentation, and knowledge advancement process.
What is the experimental method?
The experimental method is a systematic approach in scientific research where an independent variable is manipulated to observe its effect on a dependent variable, under controlled conditions.


