A randomized control trial (RCT) is a type of study design that involves randomly assigning participants to either an experimental group or a control group to measure the effectiveness of an intervention or treatment.
Randomized Controlled Trials (RCTs) are considered the “gold standard” in medical and health research due to their rigorous design.

Key Takeaways
- Definition: An RCT randomly assigns participants to an intervention group or a control group to test whether a treatment actually causes an effect.
- Randomization: Random allocation balances both known and unknown differences between groups, which is why the design supports strong causal claims.
- Blinding: Hiding group assignments from participants, researchers, or both reduces bias from expectations and observer judgement.
- Landmark Trial: The first modern RCT, a 1948 test of streptomycin for tuberculosis, became the template for the design.
- Limitations: RCTs are expensive, can be impractical or unethical for some questions, and strict samples can limit how well results generalize.
- Modern Evidence: Since prospective trial registration became standard, far fewer trials report positive results, showing how much selective reporting had inflated earlier findings.
Control Group
A control group consists of participants who do not receive the intervention being tested. They serve as the comparison group, the yardstick against which the intervention’s effect is judged.
The choice of comparator shapes what question the trial answers:
- No-treatment control: Shows whether the intervention does anything beyond the natural course of the condition.
- Placebo control: An inert pill, sham procedure, or credible but inactive task. It holds constant the experience of being treated.
- Treatment as usual (or waiting list): Compares the new intervention with the standard care patients would ordinarily get.
- Active comparator: An established treatment. It asks the sharpest question: is the new intervention better than, or no worse than, the best current option?
Random allocation, not deliberate matching, keeps the control group comparable to the experimental group. That is the key mechanism. Assignment is left to chance.
As a result, age, gender, social class, ethnicity, and countless other characteristics end up balanced across both groups on average, including traits no researcher thought to measure. That is randomization’s real power.
This is what lets researchers attribute any later difference in outcome to the treatment itself. The two groups start out alike in every respect but one: the intervention. That is why scientists treat the RCT as the gold standard for clinical trials.
Random Allocation
Random allocation and random assignment are terms used interchangeably in the context of a randomized controlled trial (RCT).
Both refer to assigning participants to different groups in a study (such as a treatment group or a control group) in a way that is completely determined by chance.
The process of random assignment controls for confounding variables, ensuring differences between groups are due to chance alone.
Without randomization, researchers might consciously or subconsciously assign patients to a particular group for various reasons.
Random allocation is not the same as random sampling. Random sampling draws participants from the wider population by chance, which supports external validity (whether findings generalize). Random allocation divides an already-recruited sample into groups by chance, which supports internal validity (confidence that the intervention caused the outcome).
Chance balancing is probabilistic, not guaranteed. In a small trial, luck alone can leave the groups unbalanced, which is why larger samples and block or stratified randomization are used.
The allocation must also be genuinely random. Alternating patients, or assigning by date of birth or hospital number, is not random. Those methods are predictable, which lets bias creep back in.
Several methods can be used for randomization in a Randomized Control Trial (RCT). Here are a few examples:
- Simple Randomization: This is the simplest method, like flipping a coin. Each participant has an equal chance of being assigned to any group. This can be achieved using random number tables, computerized random number generators, or drawing lots or envelopes.
- Block Randomization: In this method, participants are randomized within blocks, ensuring that each block has an equal number of participants in each group. This helps to balance the number of participants in each group at any given time during the study.
- Stratified Randomization: This method is used when researchers want to ensure that certain subgroups of participants are equally represented in each group. Participants are divided into strata, or subgroups, based on characteristics like age or disease severity, and then randomized within these strata.
- Cluster Randomization: In this method, groups of participants (like families or entire communities), rather than individuals, are randomized.
- Adaptive Randomization: In this method, the probability of being assigned to each group changes based on the participants already assigned to each group. For example, if more participants have been assigned to the control group, new participants will have a higher probability of being assigned to the experimental group.
Computer software can generate random numbers or sequences that can be used to assign participants to groups in a simple randomization process.
For more complex methods like block, stratified, or adaptive randomization, computer algorithms can be used to consider the additional parameters and ensure that participants are assigned to groups appropriately.
Allocation concealment is a related safeguard. It hides the upcoming assignments from whoever recruits participants (see Allocation Concealment below).
Preventing that kind of foreknowledge helps avoid selection bias and protects the validity of the study results.
Allocation Concealment
Allocation concealment keeps the randomization process honest at the point of enrollment. It hides the upcoming sequence of group assignments from whoever is recruiting participants, so no one can steer a particular person into a particular group.
Concealment is not the same as randomization. Randomization already decides who gets which treatment. Concealment simply stops anyone from finding out the answer before a participant has enrolled.
In practice, the upcoming allocations are held in a central computer or in sequentially numbered, opaque, sealed envelopes until the participant is irrevocably entered. No one can peek ahead.
Without concealment, a researcher who knew the next slot was “treatment” might, consciously or not, enroll a more promising patient into it, or delay a sicker one. This would reintroduce exactly the selection bias randomization is meant to remove.
Concealment is also distinct from blinding:
- Concealment: Protects the assignment process before and at randomization, and is always possible.
- Blinding: Protects knowledge of the assignment afterwards, and is not always achievable.
Blinding (Masking)
Blinding, or masking, means withholding the group assignments during the study. Depending on the design, participants, researchers, or both are kept unaware of who is in which group. The reason is simple: bias.
A blinded study keeps participants from knowing about their treatment, which avoids bias in the research. Any information that could influence them is withheld until the study is complete.
Blinding can extend beyond participants to anyone involved in the research, including data collectors, evaluators, technicians, and data analysts.
Good blinding matters. It can reduce several experimental biases: expectation effects, observer bias, confirmation bias, researcher bias, and other distortions that creep in once people know the group assignments.
Trials differ in who is kept unaware:
- Single-blind: Participants do not know whether they receive the active intervention or the control. This neutralises the placebo effect and demand characteristics (cues that reveal a study’s aim).
- Double-blind: Neither participants nor researchers know who receives the drug or the placebo. This also guards against researcher expectancy, an investigator’s unconscious tendency to favour a treatment they believe in.
- Triple-blind: Masking extends to the statisticians or the outcome-adjudication committee, so even the analysis cannot be nudged toward a hoped-for result.
Each participant is randomly assigned to one of the two groups when they enroll, and the medication looks identical either way.
Blinding is easy in drug trials, where an identical-looking placebo can be made. It is much harder in psychological and behavioural research. A participant knows whether they are attending twelve weeks of psychotherapy or sitting on a waiting list. The therapist knows which treatment they deliver.
The usual remedy is partial: blind the outcome assessor. An independent clinician rates symptoms without knowing which group each participant was in, even when participants and therapists cannot be blinded.
This matters most for subjective outcomes. Hróbjartsson and Gøtzsche (2001) ran a large systematic review of trials that randomly assigned patients to a placebo or to no treatment.
Placebos had little effect on objective or binary outcomes. They did produce a modest apparent benefit on subjective, self-reported outcomes such as pain. Expectation-driven effects therefore concentrate in the soft, subjective measures that most psychological trials rely on.
Blinding cannot fix every investigator bias. Luborsky and colleagues (1999) pooled eight earlier reviews with a fresh analysis of 29 head-to-head psychotherapy comparisons.
Which treatment came out ahead tracked which treatment the study’s own research team was known to favour. Their paper called this allegiance a “wild card” in comparisons of treatment efficacy.
The effect need not be conscious. Allegiance can act through how skilfully a favoured therapy is delivered, which outcome measures are chosen, or how the comparison condition is built. Blinding an outcome assessor cannot correct any of that after the fact.
Figure 1. Evidence-based medicine pyramid. Each level of the pyramid reflects the quality of the research designs at that level.
Higher levels mean higher-quality evidence, but fewer studies of that kind exist in the published literature. Randomized controlled trials sit near the top, so fewer of them get published than weaker designs.
Research Designs
The choice of design should be guided by the research question and the nature of the treatments being studied. Practical considerations, such as sample size and resources, and ethical considerations, such as ensuring participants have access to potentially beneficial treatments, matter too.
The goal is to select a design that answers the research question validly and reliably, while minimizing bias and confounds.
1. Between-participants randomized designs
Between-participant design involves randomly assigning participants to different treatment conditions. In its simplest form, it has two groups: an experimental group receiving the treatment and a control group.
With more than two levels, multiple treatment conditions are compared. The key feature is that each participant experiences only one condition.
This design allows for clear comparison between groups without worrying about order effects or carryover effects.
It’s particularly useful for treatments that have lasting impacts or when experiencing one condition might influence how participants respond to subsequent conditions.
Example
A study testing a new antidepressant medication might randomly assign 100 participants to either receive the new drug or a placebo.
The researchers would then compare depression scores between the two groups after a specified treatment period to determine if the new medication is more effective than the placebo.
Use this design when:
- You want to compare the effects of different treatments or interventions
- Carryover effects are likely (e.g., learning effects or lasting physiological changes)
- The treatment effect is expected to be permanent
- You have a large enough sample size to ensure groups are equivalent through randomization
2. Factorial designs
Factorial designs investigate the effects of two or more independent variables simultaneously. They allow researchers to study both main effects of each variable and interaction effects between variables. Interactions matter.
These can be between-participants (different groups for each combination of conditions), within-participants (all participants experience all conditions), or mixed (combining both approaches).
Factorial designs let researchers see how different factors combine to influence outcomes, a fuller picture than testing one variable at a time. That efficiency matters.
They’re also more efficient than running separate studies for each variable, since a single design can reveal interactions a simpler one would miss.
Example
A study examining exercise intensity (high vs. low) and diet type (high-protein vs. high-carb) on weight loss might use a 2×2 factorial design. This crosses the two variables to create four groups.
Participants would be randomly assigned to one of four groups: high-intensity exercise with high-protein diet, or high-intensity exercise with high-carb diet. The other two groups paired low-intensity exercise with either a high-protein or a high-carb diet.
Use this design when:
- You want to study the effects of multiple independent variables simultaneously
- You’re interested in potential interactions between variables
- You want to increase the efficiency of your study by testing multiple hypotheses at once
3. Cluster randomized designs
In cluster randomized trials, groups or “clusters” of participants are randomized to treatment conditions, rather than individuals.
This is often used when individual randomization is impractical or when the intervention is naturally applied at a group level.
It’s particularly useful in educational or community-based research where individual randomization might be disruptive or lead to treatment diffusion.
Cluster designs carry costs. People within a cluster tend to resemble one another, which reduces the effective sample size and calls for special analysis.
Many clusters are also needed for randomization to balance the groups.
Example:
A study testing a new teaching method might randomize entire classrooms to either use the new method or continue with the standard curriculum.
The researchers would then compare student outcomes between the classrooms using the different methods, rather than randomizing individual students.
Use this design when:
- The intervention is naturally delivered to a whole group, such as a class, clinic, or community
- Randomizing individuals would be impractical or disruptive
- Treating some people in a group could contaminate others in the same group
- You can recruit enough clusters for randomization to balance the groups
4. Within-participants (repeated measures) designs
In these designs, each participant experiences all treatment conditions, serving as their own control.
Within-participants designs are more statistically powerful as they control for individual differences. They require fewer participants, making them more efficient.
However, they’re only appropriate when the treatment effects are temporary and when you can effectively counterbalance to control for order effects.
Example
A study on the effects of caffeine on cognitive performance might have participants complete cognitive tests on three separate occasions. Each time, they would have consumed a different amount of caffeine: none, a low dose, or a high dose.
The order of these conditions would be counterbalanced across participants to control for order effects.
Use this design when:
- You have a smaller sample size available
- Individual differences are likely to be large
- The effects of the treatment are temporary
- You can effectively control for order and carryover effects
5. Crossover designs
Crossover designs are a specific type of within-participants design where participants receive different treatments in different time periods.
This allows each participant to serve as their own control and can be more efficient than between-participants designs.
Crossover designs combine the benefits of within-participants designs (increased power, control for individual differences) with the ability to compare different treatments.
They’re particularly useful in clinical trials. Each participant experiences every treatment, but the effects of one must never bleed into the next.
Example:
A study comparing two different pain medications might have participants use one medication for a month, then switch to the other medication for another month after a washout period.
Pain levels would be measured during both treatment periods, allowing for within-participant comparisons of the two medications’ effectiveness.
Use this design when:
- You want to compare the effects of different treatments within the same individuals
- The treatments have temporary effects with a known washout period
- You want to increase statistical power while using a smaller sample size
- You want to control for individual differences in response to treatment
6. Superiority, Non-Inferiority, and Equivalence Trials
Trials also differ in the claim they are built to test:
- Superiority: Asks whether the intervention is better than the comparator.
- Non-inferiority: Asks whether a new treatment is not meaningfully worse than an established one. This suits treatments that are cheaper, safer, or easier to deliver.
- Equivalence: Asks whether two treatments are, within a specified margin, much the same.
These claims need different statistical frameworks, and confusing them is a common error. Failing to prove superiority is not the same as proving equivalence.
Advantages of RCTs
Prevents bias
In randomized control trials, participants must be randomly assigned to either the intervention group or the control group. Each individual has an equal chance of being placed in either group.
This is meant to prevent selection bias and allocation bias and achieve control over any confounding variables to provide an accurate comparison of the treatment being studied.
Chance does the balancing. Random assignment spreads patient characteristics that could influence the outcome evenly across groups, so any difference in outcome can be attributed to the treatment.
Strong internal validity
Because the participants are randomized, the characteristics between the two groups are balanced. If a significant difference in the primary outcome then appears, researchers can reasonably assume it reflects the intervention, not some other difference between the groups.
This is strong internal validity. It means confidence that the intervention itself produced the change. Random allocation balances measured and unmeasured confounders alike, which statistical adjustment in observational studies cannot match.
Power still depends on sample size.
Blinding
Blinding also reduces bias by hiding group assignments from participants, researchers, or both (see Blinding above for single-blind, double-blind, and triple-blind designs). Even partial blinding, such as an outcome assessor who does not know the allocation, helps keep results honest.
Limitations of RCTs
Costly and Timely
Some interventions require years or even decades to evaluate, rendering them expensive and time-consuming.
It might take an extended period of time before researchers can identify a drug’s effects or discover significant results.
Trials also need years of recruitment and follow-up. They also require substantial infrastructure for randomization, blinding, monitoring, and data management, and the population being studied can change while the trial is still running.
Requires large sample size
There must be enough participants in each group of a randomized control trial so researchers can detect any true differences or effects in outcomes between the groups.
Researchers cannot detect clinically important results if the sample size is too small.
That is one reason many important questions are simply never tested with an RCT: the scale required rules them out before a trial is ever designed.
Change in population over time
Because randomized control trials are longitudinal, it is almost inevitable that some participants will not complete the study. People drop out due to death, migration, non-compliance, or simply losing interest. That loss adds up.
Dropping out in this way is known as attrition, and it can threaten a trial’s statistical power. Selective attrition, where drop-out differs between the groups, raises a subtler danger: people who drop out of a demanding treatment are often those doing worst. Simply excluding them would flatter the intervention.
There is a safeguard for this.
Researchers guard against it by choosing how to analyse the data:
- Intention-to-treat: Everyone stays in the group they were randomized to, even if they never finished the treatment. Counting dropouts sounds paradoxical, but excluding them would bias the result.
- Per-protocol: Only participants who followed the protocol fully are included. This estimates the effect of the treatment as ideally delivered, but it is vulnerable to the same attrition bias.
Intention-to-treat is therefore generally the more conservative and trustworthy primary analysis.
Ethics
Randomized control trials are not always practical or ethical, and such limitations can prevent researchers from conducting their studies.
For example, a treatment could be too invasive to justify testing on healthy volunteers. Giving some participants a placebo instead of a real drug could also deny them their normal course of treatment for a serious illness. Without ethical approval, a randomized control trial cannot proceed.
Randomizing people is only ethical under clinical equipoise: genuine, honest uncertainty among experts about whether the intervention is better than the comparator.
Where that uncertainty is absent, the trial is unethical. Withholding a treatment believed to work cannot be justified. This rules out trials that would randomize adults to smoke or deny patients a treatment already known to save lives.
A placebo arm raises its own dilemma when an effective treatment already exists.
External Validity and Generalisability
Randomization gives an RCT strong internal validity: confidence that the intervention, not something else, produced the result. That strength can come at a cost. External validity, whether the finding generalizes beyond the trial, can suffer.
To keep a sample clean, an explanatory trial often excludes patients with other conditions, the very old, the very ill, or those taking other medications. This produces a tidy but unrepresentative sample.
A therapy that works in a pristine trial can still falter in messier, everyday practice. Researchers call this the efficacy-effectiveness gap.
Pragmatic trials restore some of that realism. They test the intervention under ordinary, real-world conditions. The trade-off is real: some of the tight control gets sacrificed along the way, and that is the price of realism.
The two ends of the spectrum differ in what they ask:
| Explanatory (efficacy) trial | Pragmatic (effectiveness) trial | |
|---|---|---|
| Question | Can this work? | Does this work in practice? |
| Conditions | Ideal and tightly controlled | Ordinary and real-world |
| Participants | Carefully selected | Typical patients, including those with other conditions |
| Clinicians | Expert therapists | Routine clinicians |
| Adherence | High | Imperfect |
Publication Bias and the Limits of RCT Evidence
An RCT is only as trustworthy as what gets published. Trials with a positive, significant result are more likely to see print than null ones. This is publication bias.
Within a published trial, the pattern can repeat. Outcomes that “worked” get emphasized. Pre-planned outcomes that did not are quietly dropped.
Some questions cannot ethically or practically go through an RCT at all: rare conditions, long-term population outcomes, or traits like age and trauma history that cannot be randomly assigned. Observational research remains essential here, not a lesser substitute.
The reverse failure is a standing joke among methodologists. A systematic review found that no randomized trial had ever tested whether parachutes prevent death and injury when jumping from an aircraft (Smith & Pell, 2003).
Some effects need no trial. Demanding one turns rigour into caricature.
Even where an RCT is possible, it is not the only tool worth trusting. It answers one question well: the average effect of a defined intervention in a defined population.
It says less about why an intervention works, or for whom, the kind of understanding mechanistic and qualitative research can supply.
Contemporary Research
Recent work on the RCT has shifted from defending the design to auditing it. Researchers now ask empirically how often trials are biased, by how much, and what actually fixes it.
Quantifying How Design Flaws Bias Results
Page and colleagues (2016) reviewed the accumulated “meta-epidemiological” literature: studies that compare effect sizes across many trials by methodological feature. They found one clear pattern. Trials with weak or unclear allocation concealment systematically exaggerate the treatment effect.
The bias is largest for subjective, self-reported outcomes. That is precisely the kind of measure psychology relies on most.
A similar pattern shows up within clinical psychology itself, where psychotherapy trials for adult depression vary widely in quality.
- Aim: To test whether the quality of psychotherapy trials for adult depression is associated with the effect sizes they report.
- Method: Cuijpers and colleagues (2010) coded 115 randomized trials against eight quality criteria, including intention-to-treat analysis, independent randomization, and blinded outcome assessors.
- Results: Only 11 studies met all eight criteria. Their average effect (d = 0.22) was less than a third of that in the other studies (d = 0.74).
- Conclusion: The effects of psychotherapy were still real, but methodological shortcuts had substantially overestimated them.
Trial quality and reported benefit are entangled, even inside clinical psychology.
Does Preregistration Change What Trials Find?
The clearest evidence that reform works comes from cardiology. It is a natural experiment in what happens when trials must declare their outcome in advance.
- Aim: To test whether the proportion of null results among large clinical trials increased once prospective registration became mandatory in 2000.
- Method: Kaplan and Irvin (2015) identified 55 large US cardiovascular trials from 1970 to 2012. They coded each as published before or after 2000, when ClinicalTrials.gov registration became mandatory.
- Results: Before 2000, 17 of 30 trials (57%) reported a significant benefit; after 2000, only 2 of 25 (8%) did. The drop tracked registration, not comparator choice or industry funding.
- Conclusion: When researchers had to declare their outcome in advance, positive findings largely dried up. Many earlier “positive” trials likely reflected flexible, after-the-fact outcome selection, not genuine effects.
This mirrors psychology’s own reckoning. When the Open Science Collaboration (2015) tried to replicate 100 published psychology studies, only around a third to a half produced a significant result the second time.
Alternatives to RCTs
The RCT is the strongest design for isolating an intervention’s causal effect, but it is not always the right tool. When randomization is impossible, unethical, or unnecessary, researchers turn to designs that trade some causal certainty for feasibility, long-term follow-up, or realism.
| Design | How groups are formed | Best suited to | Main limitation |
|---|---|---|---|
| Quasi-experimental | Existing membership, self-selection, or an administrative rule | Settings where randomizing is impractical or unethical | Groups can differ at the outset |
| Cohort | Naturally occurring exposure | Long-term or rare outcomes; exposures that cannot be assigned | Confounding by indication; reverse causation |
| Single-case | The person is their own control across phases | One person’s response; rare conditions | Limited generalizability |
| Naturalistic | No manipulation; everyday settings | Moment-to-moment, context-dependent dynamics | Cannot isolate a causal effect |
Quasi-Experimental Designs
Quasi-experimental designs keep the comparative logic of an experimental design but drop random allocation. An intervention group is still set against a comparison group.
Participants are assigned by existing group membership, self-selection, or an administrative rule. Chance plays no part. Examples include non-equivalent control-group designs, interrupted time-series, and regression-discontinuity designs around an eligibility cut-off.
Handley and colleagues (2011) note that these are often the only feasible option when randomizing patients or clinics would be impractical or unethical. With a pre-intervention baseline and a carefully chosen comparison group, they can still support reasonably strong causal inference.
The cost is exactly what they give up. Because chance does not decide group membership, the groups can differ at the outset.
Statistical adjustment can reduce, but never fully remove, confounding by variables the researcher failed to measure.
Cohort and Observational Studies
Cohort studies follow groups defined by a naturally occurring exposure, such as a lifestyle factor, an environmental risk, or a treatment chosen in ordinary practice. They compare outcomes by exposure, not by assignment.
Grimes and Schulz (2002) describe the design as suited to questions an RCT cannot answer ethically or practically. These include long-term risks, rare or delayed outcomes, and exposures such as smoking or early-life adversity that could never be assigned.
The trade-off is the one randomization removes. Because exposure was not allocated by chance, an association may reflect confounding by indication or reverse causation.
Confounding by indication means people who choose a treatment differ from those who do not. Reverse causation means the outcome drives the exposure. Cohort evidence therefore carries less causal certainty than a well-run RCT, even when its sample is far larger.
Single-Case Experimental Designs
Single-case (n-of-1) designs dispense with a separate comparison group. One participant is measured repeatedly across alternating phases, so the person serves as their own control over time.
Examples are the ABAB withdrawal (baseline, treatment, baseline, treatment) and the multiple-baseline design.
Kazdin (2019) sets out where this approach fits. It suits clinical settings where one person’s response to a specific treatment is the question. It also suits conditions too rare to recruit a group trial, and fields like applied behaviour analysis that work with small numbers.
The cost is generalizability. A change that coincides with the intervention phase could reflect natural fluctuation, history, or maturation. Single-case evidence therefore needs replication across many cases before it carries the weight of a well-powered RCT.
Naturalistic and Ecological Studies
Naturalistic and ecological studies record behaviour, mood, or physiology as it unfolds in everyday environments, often through repeated real-time sampling. They do not manipulate conditions inside a trial.
Shiffman, Stone, and Hufford (2008) reviewed ecological momentary assessment. Their review shows how it captures moment-to-moment, context-dependent dynamics. A between-groups trial cannot see these, because it compares averages at a few fixed time points.
It also avoids the artificiality a laboratory or clinic setting can introduce.
It cannot isolate a causal effect. Without a manipulated condition and a comparison group, a naturalistic study reveals patterns and correlations but cannot show that one factor caused a change in another.
None of these designs is a lesser substitute. Each is the better tool for a different kind of question, or the only tool where randomization is not possible.
The right choice matches the design to the question, the feasibility of randomization, and the ethics of the case.
Fictitious Example
An example of an RCT would be a clinical trial comparing a drug’s effect or a new treatment on a select population.
The researchers would randomly assign participants to either the experimental group or the control group. They would then compare outcomes between those who received the drug or treatment and those who did not.
Real-life Examples
RCTs in Psychology
Although the RCT was forged in medicine, it is now the design against which claims about psychological interventions are judged. Formal hierarchies of empirically supported treatments reserve their top tiers for interventions shown to beat a no-treatment, placebo, or alternative-treatment control.
Those trials must use randomized assignment, be adequately powered, and be replicated across independent settings. The design appears in several forms:
- Psychotherapy trials: Clients with a defined disorder, such as depression, anxiety, or PTSD, are randomized to a therapy like cognitive behavioural therapy or to a waiting-list, usual-care, or active-comparator condition.
- Drug trials: Psychiatry uses the classic double-blind, placebo-controlled parallel design to test psychoactive medication.
- Educational and behavioural interventions: School-based prevention and parenting programmes are often evaluated with cluster RCTs, such as Botvin and colleagues’ (2000) long-term follow-up of a school drug-prevention programme.
- Neuropsychological rehabilitation: Wilson and colleagues (2005) used a randomized design to test a paging system as a memory aid for people with traumatic brain injury.
The trials below show the range of questions the design can answer.
- Fabiano, G. A., Schatz, N. K., Merrill, B. M., Piscitello, J., Hayes, T. B., Jusko, M., Gnagy, E. M., Greiner, A. R., Tower, D., Boeckel, A., Gallo, R., Lupas, K., Gordon, C., Ramos, M., Sikov, J., Caron, S., & Pelham, W. E., Jr. (2025). A randomized, controlled trial to evaluate the efficacy of a daily report card intervention to enhance the efficacy of individualized education programs for children with attention-deficit/hyperactivity disorder. Journal of Consulting and Clinical Psychology, 93(7), 484–499.
- Preventing illicit drug use in adolescents: Long-term follow-up data from a randomized control trial of a school population (Botvin et al., 2000).
- A prospective randomized control trial comparing medical and surgical treatment for early pregnancy failure (Demetroulis et al., 2001).
- A randomized control trial to evaluate a paging system for people with traumatic brain injury (Wilson et al., 2005).
- Prehabilitation versus Rehabilitation: A Randomized Control Trial in Patients Undergoing Colorectal Resection for Cancer (Gillis et al., 2014).
- A Randomized Control Trial of Right-Heart Catheterization in Critically Ill Patients (Guyatt, 1991).
- Berry, R. B., Kryger, M. H., & Massie, C. A. (2011). A novel nasal excitatory positive airway pressure (EPAP) device for the treatment of obstructive sleep apnea: A randomized controlled trial. Sleep, 34, 479–485.
- Gloy, V. L., Briel, M., Bhatt, D. L., Kashyap, S. R., Schauer, P. R., Mingrone, G., . . . Nordmann, A. J. (2013, October 22). Bariatric surgery versus non-surgical treatment for obesity: A systematic review and meta-analysis of randomized controlled trials. BMJ, 347.
- Streeton, C., & Whelan, G. (2001). Naltrexone, a relapse prevention maintenance treatment of alcohol dependence: A meta-analysis of randomized controlled trials. Alcohol and Alcoholism, 36 (6), 544–552.
How Should an RCT be Reported?
Reporting of an RCT should be clear, transparent, and comprehensive. Readers need to understand the design, conduct, analysis, and interpretation of the trial, not just its headline result.
The Consolidated Standards of Reporting Trials (CONSORT) statement, updated by Schulz, Altman, and Moher (2010), is the international standard for reporting parallel-group RCTs.
Its flow diagram tracks participants through enrolment, allocation, follow-up, and analysis, letting a reader check whether randomization and blinding were real rather than taking the authors’ summary on trust.
Further Information
- Cocks, K., & Torgerson, D. J. (2013). Sample size calculations for pilot randomized trials: a confidence interval approach. Journal of clinical epidemiology, 66(2), 197-201.
- Kendall, J. (2003). Designing a research project: randomised controlled trials and their principles. Emergency medicine journal: EMJ, 20(2), 164.
History of the RCT: The Streptomycin Trial
Before 1948, claims that a treatment worked usually came from doctors comparing patients they had chosen to treat against those they had not. That method is wide open to exactly the confounding bias randomization was later built to solve.
The Medical Research Council’s trial changed that. Austin Bradford Hill, the statistician who designed it, insisted on a formal random allocation procedure specifically to remove that bias.
The trial almost universally credited as the first properly randomized controlled trial is the Medical Research Council’s 1948 test of streptomycin for tuberculosis, designed with Hill. It set the template.
- Aim: To test whether streptomycin added to bed rest improved recovery in acute pulmonary tuberculosis, using a design that removed allocation bias.
- Method: Just over a hundred tuberculosis patients were randomly allocated to streptomycin plus bed rest, or bed rest alone, with the schedule concealed from admitting clinicians. Outcomes were assessed over six months by assessors blind to each patient’s treatment.
- Results: The streptomycin group did substantially better: radiological improvement was far more common, and six-month mortality was markedly lower, roughly 7% versus 27%. The trial also reported a downside honestly, as streptomycin-resistant bacteria emerged rapidly.
- Conclusion: Streptomycin plus bed rest beat bed rest alone for tuberculosis. Combining random allocation, concealment, and blinded assessment gave a defensible way to compare treatments without bias. The trial became the template for the modern RCT.
Randomization was accepted as ethical here for a specific reason. Streptomycin was in genuinely scarce supply, so allocating it by lottery was arguably fairer than leaving the choice to physician preference.
That took real ethical thought.
That is an early, concrete example of clinical equipoise: genuine uncertainty about which treatment works better. It is the ethical condition every RCT still has to satisfy. That legacy still shapes practice.
The streptomycin trial’s design, random allocation combined with concealment and blinding, is still the template every RCT is judged against today.
It is also why the RCT sits near the top of the hierarchy of evidence for questions about whether a treatment actually works. That hierarchy places it above cohort studies, case-control studies, and expert opinion alone.
References
Akobeng, A.K., Understanding randomized controlled trials. Archives of Disease in Childhood, 2005; 90: 840-844.
Bell, C. C., Gibbons, R., & McKay, M. M. (2008). Building protective factors to offset sexually risky behaviors among black youths: a randomized control trial. Journal of the National Medical Association, 100 (8), 936-944.
Bhide, A., Shah, P. S., & Acharya, G. (2018). A simplified guide to randomized controlled trials. Acta obstetricia et gynecologica Scandinavica, 97 (4), 380-387.
Botvin, G. J., Griffin, K. W., Diaz, T., Scheier, L. M., Williams, C., & Epstein, J. A. (2000). Preventing illicit drug use in adolescents: Long-term follow-up data from a randomized control trial of a school population. Addictive Behaviors, 25 (5), 769-774.
Cuijpers, P., van Straten, A., Bohlmeijer, E., Hollon, S. D., & Andersson, G. (2010). The effects of psychotherapy for adult depression are overestimated: A meta-analysis of study quality and effect size. Psychological Medicine, 40 (2), 211-223. https://doi.org/10.1017/S0033291709006114
Demetroulis, C., Saridogan, E., Kunde, D., & Naftalin, A. A. (2001). A prospective randomized control trial comparing medical and surgical treatment for early pregnancy failure. Human Reproduction, 16 (2), 365-369.
Gillis, C., Li, C., Lee, L., Awasthi, R., Augustin, B., Gamsa, A., … & Carli, F. (2014). Prehabilitation versus rehabilitation: a randomized control trial in patients undergoing colorectal resection for cancer. Anesthesiology, 121 (5), 937-947.
Globas, C., Becker, C., Cerny, J., Lam, J. M., Lindemann, U., Forrester, L. W., … & Luft, A. R. (2012). Chronic stroke survivors benefit from high-intensity aerobic treadmill exercise: a randomized control trial.
Neurorehabilitation and Neural Repair, 26
(1), 85-95.
Grimes, D. A., & Schulz, K. F. (2002). Cohort studies: Marching towards outcomes. The Lancet, 359(9303), 341-345. https://doi.org/10.1016/S0140-6736(02)07500-1
Guyatt, G. (1991). A randomized control trial of right-heart catheterization in critically ill patients. Journal of Intensive Care Medicine, 6 (2), 91-95.
Handley, M. A., Schillinger, D., & Shiboski, S. (2011). Quasi-experimental designs in practice-based research settings: Design and implementation considerations. Journal of the American Board of Family Medicine, 24(5), 589-596. https://doi.org/10.3122/jabfm.2011.05.110067
Hróbjartsson, A., & Gøtzsche, P. C. (2001). Is the placebo powerless? An analysis of clinical trials comparing placebo with no treatment. New England Journal of Medicine, 344 (21), 1594-1602. https://doi.org/10.1056/NEJM200105243442106
Kaplan, R. M., & Irvin, V. L. (2015). Likelihood of null effects of large NHLBI clinical trials has increased over time. PLoS ONE, 10 (8), e0132382. https://doi.org/10.1371/journal.pone.0132382
Kazdin, A. E. (2019). Single-case experimental designs. Evaluating interventions in research and clinical practice. Behaviour Research and Therapy, 117, 3-17. https://doi.org/10.1016/j.brat.2018.11.015
Luborsky, L., Diguer, L., Seligman, D. A., Rosenthal, R., Krause, E. D., Johnson, S., Halperin, G., Bishop, M., Berman, J. S., & Schweizer, E. (1999). The researcher’s own therapy allegiances: A “wild card” in comparisons of treatment efficacy. Clinical Psychology: Science and Practice, 6(1), 95-106. https://doi.org/10.1093/clipsy.6.1.95
Medical Research Council. (1948). Streptomycin treatment of pulmonary tuberculosis: A Medical Research Council investigation. British Medical Journal, 2 (4582), 769-782.
MediLexicon International. (n.d.). Randomized controlled trials: Overview, benefits, and limitations. Medical News Today. Retrieved from https://www.medicalnewstoday.com/articles/280574#what-is-a-randomized-controlled-trial
Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349 (6251), aac4716. https://doi.org/10.1126/science.aac4716
Page, M. J., Higgins, J. P. T., Clayton, G., Sterne, J. A. C., Hróbjartsson, A., & Savović, J. (2016). Empirical evidence of study design biases in randomized trials: Systematic review of meta-epidemiological studies. PLoS ONE, 11 (7), e0159267. https://doi.org/10.1371/journal.pone.0159267
Schulz, K. F., Altman, D. G., & Moher, D. (2010). CONSORT 2010 statement: Updated guidelines for reporting parallel group randomised trials. BMJ, 340, c332. https://doi.org/10.1136/bmj.c332
Shiffman, S., Stone, A. A., & Hufford, M. R. (2008). Ecological momentary assessment. Annual Review of Clinical Psychology, 4, 1-32. https://doi.org/10.1146/annurev.clinpsy.3.022806.091415
Smith, G. C. S., & Pell, J. P. (2003). Parachute use to prevent death and major trauma related to gravitational challenge: Systematic review of randomised controlled trials. BMJ, 327(7429), 1459-1461. https://doi.org/10.1136/bmj.327.7429.1459
Wilson, B. A., Emslie, H., Quirk, K., Evans, J., & Watson, P. (2005). A randomized control trial to evaluate a paging system for people with traumatic brain injury. Brain Injury, 19 (11), 891-894.
