A case-control study is a research method where two groups of people are compared: those with the condition (cases) and those without (controls). By looking at their past, researchers try to identify what factors might have contributed to the condition in the ‘case’ group.
Key Takeaways
- Samples on Outcome: A case-control study starts with people who already have (cases) and do not have (controls) a condition, then looks back for differences in exposure.
- Speed and Cost: Because no one is followed forward in time, it is fast and cheap, and especially useful for rare or slow-developing conditions.
- Main Weaknesses: Recall bias, the difficulty of finding a truly comparable control group, and its inability to establish causation on its own.
- Hypothesis-Generating: Case-control evidence is usually a first step, later tested with a cohort study or a randomized controlled trial. Doll and Hill’s smoking research is a classic example.
- Wide Reach: The same logic scales from small clinical studies to genome-wide association studies with hundreds of thousands of participants, underpinning major findings in psychiatric genetics.
What is a Case Control Study?
A case-control study looks at people who already have a certain condition (the cases) and people who don’t (the controls). By comparing these two groups, researchers try to figure out what might have caused the condition. The comparison looks backward in time.
Definitions matter here.
The “cases” are the individuals with the disease or condition under study; the “controls” are similar individuals without it. Controls should share the cases’ characteristics, such as age, sex, demographic background, and health status.
This keeps the two groups similar in everything except the outcome being studied. Confounding variables are the real danger here. They are factors, other than the exposure under study, that could otherwise explain a difference between the groups.
Researchers follow the same basic sequence in every case-control study:
- Identify the cases – a group of people known to have the condition, typically drawn from a hospital register or clinical service.
- Identify the controls – a comparable group without the condition, drawn from the general population or the same clinical setting.
- Look back at both groups’ histories, through interviews, medical records, or existing datasets, to see who was exposed to the factor of interest.
- Compare how often each group was exposed. If the exposure is more common among the cases, the researcher can hypothesize that it may be linked to the outcome.
Case-control studies identify associations between an exposure and an outcome. They do not, on their own, prove that the exposure caused the outcome.
Figure: Schematic diagram of case-control study design. Kenneth F. Schulz and David A. Grimes (2002) Case-control studies: research in reverse. The Lancet Volume 359, Issue 9304, 431 – 434
Types of Case-Control Study
Not every case-control study is built the same way, and the design chosen affects how much confidence a reader should place in its findings.
Unmatched Studies
In an unmatched design, researchers sample cases and controls independently from their source populations.
Any differences in age, sex, or other characteristics are handled statistically, through regression, rather than through the sampling itself.
This approach is administratively simpler. It also allows a larger, more representative control pool.
But it depends on the researcher having measured and correctly modelled every relevant confounding variable, which is easier to state than to guarantee. An unmeasured confounder can quietly bias the whole result.
Large registries often use this design for exactly that reason.
Because nothing is matched at the sampling stage, an unmatched study can still test a candidate matching variable as a risk factor in its own right. A matched design rules that out by design.
Matched Studies
In a matched design, each case is paired with one or more controls who share specific characteristics. These are commonly age, sex, and sometimes location, chosen in advance as known or suspected confounders.
Two matching strategies exist.
Individual matching pairs one control to each case. Frequency matching instead gives the whole control group the same profile as the case group, without pairing individuals directly.
Matching improves statistical efficiency.
It also removes the matched variables as sources of confounding. But a matched variable can no longer be studied as a risk factor in its own right.
Matching on a variable that lies on the causal pathway can even bias the odds ratio toward the null (Grimes & Schulz, 2002). Over-matching is the risk to watch for.
Nested and Case-Cohort Studies
A nested case-control study is conducted within an existing cohort.
Researchers identify everyone in the cohort who develops the outcome (the cases) and sample members who did not (the controls), rather than assembling a fresh control group from scratch.
Because the cohort’s exposures were measured before anyone developed the outcome, a nested design keeps the case-control study’s speed while largely avoiding recall bias. Exposure was recorded prospectively, not reconstructed after the fact (Ernster, 1994).
A case-cohort study is a close relative.
It compares all cases in a cohort against a single random subsample of the whole cohort at baseline, rather than controls matched individually to each case. This lets the same control subsample be reused to study several outcomes from one cohort.
The logic scales either way.
Choosing among these designs is a trade-off between speed, cost, and how thoroughly confounding is controlled. The right choice depends on the research question and what data already exist.
Key Studies
Two studies show the case-control design at very different scales, from a handful of London hospitals to a multi-country genetic collaboration.
Doll and Hill (1950): Smoking and Lung Cancer
- Aim: Richard Doll and Austin Bradford Hill tested the then-controversial idea that smoking caused Britain’s sharply rising lung cancer rate. Lung cancer was still rare enough in 1950 to make a prospective cohort study impractical.
- Method: Doll and Hill recruited newly diagnosed lung cancer patients from 20 London hospitals as cases. Controls were patients admitted to the same hospitals with other conditions, matched on age, sex, and hospital.
- Results: Almost all male lung cancer patients in the study were smokers, and heavier smoking meant substantially higher risk in a clear dose-response pattern. Male patients with lung cancer who had never smoked were strikingly rare.
- Conclusion: Doll and Hill concluded that smoking was very likely a major cause of the rise in lung cancer. That conclusion ran directly against prevailing medical opinion and tobacco-industry-funded scepticism at the time.
Critics raised recall-bias and interviewer-bias objections, since newly diagnosed patients may search their history more thoroughly than a healthy control would (Schulz & Grimes, 2002).
The debate did not end there.
It was settled only when Doll and Hill’s own prospective cohort of 40,000 British doctors confirmed the finding without relying on retrospective recall (Doll & Hill, 1956).
Di Forti et al. (2019): Cannabis Potency and Psychosis (EU-GEI)
- Aim: The multi-site EU-GEI collaboration tested whether patterns of cannabis use could explain variation in new psychotic-disorder cases across different European cities. It compared daily use and high- versus low-potency cannabis specifically.
- Method: Across 11 European sites and one Brazilian site, researchers recruited 901 first-episode psychosis patients as cases and 1,237 population controls from the same areas. A structured interview established each participant’s pattern and potency of cannabis use before psychosis onset.
- Results: Daily cannabis use was linked to more than three times the odds of psychotic disorder compared with never using cannabis. The odds rose to nearly five times among daily users of high-potency cannabis.
- Conclusion: The authors concluded that differences in cannabis frequency and potency plausibly contributed to real geographic variation in psychotic-disorder rates, with clear implications for public health policy on cannabis potency.
The authors were explicit that their estimates depended on assuming cannabis use is causal, a limitation the design cannot resolve alone. Effect sizes also varied sharply across sites, so a single pooled odds ratio understates how much the finding depends on local context (Di Forti et al., 2019).
Advantages
Quick, inexpensive, and simple
Because these studies use already existing data and do not require any follow-up with subjects, they tend to be quicker and cheaper than other types of research. Case-control studies also do not require large sample sizes.
Beneficial for studying rare diseases
Researchers in case-control studies start with a population of people known to have the target disease instead of following a population and waiting to see who develops it.
This enables researchers to identify current cases and enroll a sufficient number of patients with a particular rare disease. That solves the sample-size problem.
Herbst, Ulfelder, and Poskanzer (1971) used this approach to link maternal use of a synthetic hormone, diethylstilbestrol, to an extremely rare vaginal cancer in young women.
A prospective study large enough to wait for that many rare cases to occur naturally would have been impossible to run.
Useful for preliminary research
Case-control studies are useful for an initial investigation of a suspected risk factor for a condition. The design is fast and inexpensive, so it lets researchers test whether an association is worth pursuing before committing to a larger follow-up study.
A case-control finding is a hypothesis, not a confirmed cause. Doll and Hill’s 1950 case-control study of smoking and lung cancer shows the pattern. The result was influential but stayed contested until their own prospective cohort study of over 40,000 British doctors confirmed it six years later.
Case-control evidence typically needs a cohort study or a randomized controlled trial to test it further.
Limitations
Subject to recall bias
Participants might be unable to remember when they were exposed or omit other details that are important for the study. In addition, those with the outcome are more likely to recall and report exposures more clearly than those without the outcome.
Memory is not neutral.
People who have just been diagnosed with a condition often search their own history more thoroughly for an explanation than a healthy control does (Coughlin, 1990).
Doll and Hill’s 1950 smoking-cancer interviews, described above, are a classic example of the concern.
Difficulty finding a suitable control group
It is important that the case group and the control group have almost the same characteristics, such as age, gender, demographics, and health status.
Forming an accurate control group can be challenging, so sometimes researchers enroll multiple control groups to bolster the strength of the case-control study.
Do not demonstrate causation
Case-control studies may prove an association between exposures and outcomes, but they can not demonstrate causation.
A case-control study cannot rule out reverse causation. It also cannot rule out a third factor producing both the exposure and the outcome.
Doll and Hill’s smoking-cancer finding remained contested until their own prospective cohort confirmed it. Di Forti et al. (2019) were explicit that their cannabis-psychosis estimates depended on an assumption of causality the case-control design cannot itself confirm.
Contemporary Research
Case-control methodology remains active today. Systematic pooling of case-control evidence has begun answering questions no single study could settle.
Favril, Yu, Uyar, and Sharpe (2022) pooled 37 psychological autopsy case-control studies from 23 countries, 5,633 cases and 7,101 controls in total. Their meta-analysis used random-effects models.
It ranked the strength of suicide risk factors across a pooled sample far larger than any single study. The presence of any mental disorder (pooled odds ratio around 13) and a history of self-harm (pooled odds ratio around 10) emerged as the strongest predictors.
Two factors led the field.
Both outweighed sociodemographic or life-event factors by a wide margin. This gives a quality-weighted answer that earlier single-site studies, such as the psychological autopsy research described above, could only address individually.
The genomic strand of this research works at a different scale entirely. Trubetskoy et al. (2022) compared the genomes of up to 76,755 people with schizophrenia against 243,649 controls, the largest case-control sample in this field.
That is genomic scale. The study identified genetic variants at 287 locations that occurred more often among cases, clustered in genes active in brain neurons.
No single variant causes schizophrenia on its own. Read together with the EU-GEI cannabis study above and this meta-analysis, these studies show the case-control design working at every level, from single-site research to genome-wide biology.
Examples
A case-control study is an observational study where researchers analyzed two groups of people (cases and controls) to look at factors associated with particular diseases or outcomes.
Below are some examples of case-control studies:
- Health Psychology: INTERHEART compared 11,119 heart-attack patients with 13,648 matched controls across 52 countries and found psychosocial stress linked to heart-attack risk (Rosengren et al., 2004).
- Suicide Research: The psychological autopsy method compares informant reports about people who died by suicide with reports about matched living controls (Cavanagh, Owens, & Johnstone, 1999).
- Environmental Health: Investigating the impact of exposure to daylight on the health of office workers (Boubekri et al., 2014).
- Neurology: Comparing serum vitamin D levels in individuals who experience migraine headaches with their matched controls (Togha et al., 2018).
- Respiratory Health: Analyzing correlations between parental smoking and childhood asthma (Strachan and Cook, 1998).
- Cardiovascular Health: Studying the relationship between elevated concentrations of homocysteine and an increased risk of vascular diseases (Ford et al., 2002).
- Gastroenterology: Assessing the magnitude of the association between Helicobacter pylori and the incidence of gastric cancer (Helicobacter and Cancer Collaborative Group, 2001).
- Oncology: Evaluating the association between breast cancer risk and saturated fat intake in postmenopausal women (Howe et al., 1990).
Frequently asked questions
1. What’s the difference between a case-control study and a cross-sectional study?
Case-control studies are different from cross-sectional studies in that case-control studies compare groups retrospectively while cross-sectional studies analyze information about a population at a specific point in time.
In cross-sectional studies, researchers are simply examining a group of participants and depicting what already exists in the population.
2. What’s the difference between a case-control study and a longitudinal study?
Case-control studies compare groups retrospectively, while longitudinal studies can compare groups either retrospectively or prospectively.
In a longitudinal study, researchers monitor a population over an extended period of time, and they can be used to study developmental shifts and understand how certain things change as we age.
In addition, case-control studies look at a single subject or a single case, whereas longitudinal studies can be conducted on a large group of subjects.
3. What’s the difference between a case-control study and a retrospective cohort study?
Case-control studies are retrospective as researchers begin with an outcome and trace backward to investigate exposure; however, they differ from retrospective cohort studies.
In a retrospective cohort study, researchers examine a group before any of the subjects have developed the disease, then examine any factors that differed between the individuals who developed the condition and those who did not.
Thus, the outcome is measured after exposure in retrospective cohort studies, whereas the outcome is measured before the exposure in case-control studies.
References
Boubekri, M., Cheung, I., Reid, K., Wang, C., & Zee, P. (2014). Impact of windows and daylight exposure on overall health and sleep quality of office workers: a case-control pilot study. Journal of Clinical Sleep Medicine: JCSM: Official Publication of the American Academy of Sleep Medicine, 10 (6), 603-611.
Cavanagh, J. T. O., Owens, D. G. C., & Johnstone, E. C. (1999). Suicide and undetermined death in south east Scotland: A case-control study using the psychological autopsy method. Psychological Medicine, 29(5), 1141–1149. https://doi.org/10.1017/S0033291799001038
Coughlin, S. S. (1990). Recall bias in epidemiologic studies. Journal of Clinical Epidemiology, 43(1), 87–91. https://doi.org/10.1016/0895-4356(90)90060-3
Di Forti, M., Quattrone, D., Freeman, T. P., Tripoli, G., Gayer-Anderson, C., Quigley, H., Rodriguez, V., Jongsma, H. E., Ferraro, L., La Cascia, C., La Barbera, D., Tarricone, I., Berardi, D., Szöke, A., Arango, C., Tortelli, A., Velthorst, E., Bernardo, M., Del-Ben, C. M., … Murray, R. M. (2019). The contribution of cannabis use to variation in the incidence of psychotic disorder across Europe (EU-GEI): A multicentre case-control study. The Lancet Psychiatry, 6(5), 427–436. https://doi.org/10.1016/S2215-0366(19)30048-3
Doll, R., & Hill, A. B. (1950). Smoking and carcinoma of the lung: Preliminary report. British Medical Journal, 2(4682), 739–748. https://doi.org/10.1136/bmj.2.4682.739
Doll, R., & Hill, A. B. (1956). Lung cancer and other causes of death in relation to smoking: A second report on the mortality of British doctors. British Medical Journal, 2(5001), 1071–1081. https://doi.org/10.1136/bmj.2.5001.1071
Favril, L., Yu, R., Uyar, A., & Sharpe, M. (2022). Risk factors for suicide in adults: Systematic review and meta-analysis of psychological autopsy studies. Evidence Based Mental Health, 25(4), 148–155. https://doi.org/10.1136/ebmental-2022-300549
Ford, E. S., Smith, S. J., Stroup, D. F., Steinberg, K. K., Mueller, P. W., & Thacker, S. B. (2002). Homocyst (e) ine and cardiovascular disease: a systematic review of the evidence with special emphasis on case-control studies and nested case-control studies. International journal of epidemiology, 31 (1), 59-70.
Helicobacter and Cancer Collaborative Group. (2001). Gastric cancer and Helicobacter pylori: a combined analysis of 12 case control studies nested within prospective cohorts. Gut, 49 (3), 347-353.
Herbst, A. L., Ulfelder, H., & Poskanzer, D. C. (1971). Adenocarcinoma of the vagina: Association of maternal stilbestrol therapy with tumor appearance in young women. New England Journal of Medicine, 284(16), 878–881. https://doi.org/10.1056/NEJM197104222841604
Howe, G. R., Hirohata, T., Hislop, T. G., Iscovich, J. M., Yuan, J. M., Katsouyanni, K., … & Shunzhang, Y. (1990). Dietary factors and risk of breast cancer: combined analysis of 12 case—control studies. JNCI: Journal of the National Cancer Institute, 82 (7), 561-569.
Lewallen, S., & Courtright, P. (1998). Epidemiology in practice: case-control studies. Community eye health, 11 (28), 57–58.
Rosengren, A., Hawken, S., Ôunpuu, S., Sliwa, K., Zubaid, M., Almahmeed, W. A., Blackett, K. N., Sitthi-amorn, C., Sato, H., & Yusuf, S. (2004). Association of psychosocial risk factors with risk of acute myocardial infarction in 11,119 cases and 13,648 controls from 52 countries (the INTERHEART study): Case-control study. The Lancet, 364(9438), 953–962. https://doi.org/10.1016/S0140-6736(04)17019-0
Schulz, K. F., & Grimes, D. A. (2002). Case-control studies: Research in reverse. The Lancet, 359(9304), 431–434. https://doi.org/10.1016/S0140-6736(02)07605-5
Strachan, D. P., & Cook, D. G. (1998). Parental smoking and childhood asthma: longitudinal and case-control studies. Thorax, 53 (3), 204-212.
Tenny, S., Kerndt, C. C., & Hoffman, M. R. (2021). Case Control Studies. In StatPearls. StatPearls Publishing.
Togha, M., Razeghi Jahromi, S., Ghorbani, Z., Martami, F., & Seifishahpar, M. (2018). Serum Vitamin D Status in a Group of Migraine Patients Compared With Healthy Controls: A Case-Control Study. Headache, 58 (10), 1530-1540.
Trubetskoy, V., Pardiñas, A. F., Qi, T., Panagiotaropoulou, G., Awasthi, S., Bigdeli, T. B., Bryois, J., Chen, C.-Y., Dennison, C. A., Hall, L. S., Lam, M., Watanabe, K., Frei, O., Ge, T., Harwood, J. C., Koopmans, F., Magnusson, S., Richards, A. L., Sidorenko, J., … Ripke, S. (2022). Mapping genomic loci implicates genes and synaptic biology in schizophrenia. Nature, 604, 502–508. https://doi.org/10.1038/s41586-022-04434-5
