Case-Control Study Design
Case-Control Study (Retrospective Case-Control Design) · Also known as: case-control study, retrospective study, matched case-control, nested case-control
A case-control study identifies individuals with a disease or outcome (cases) and a comparison group without the outcome (controls), then measures prior exposure retrospectively. Developed in the 1950s–1970s by epidemiologists like Schlesselman and MacMahon, case-control studies are especially efficient for rare diseases, as they sample cases enriched for the outcome, avoiding the need for enormous cohorts. They are a mainstay of clinical epidemiology, observational research, and outbreak investigations.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use case-control studies when: (1) the disease or outcome is rare (e.g., rare cancers, birth defects), making cohort studies prohibitively expensive, (2) a long latency period exists between exposure and outcome (years to decades), (3) studying multiple exposures for a single outcome, (4) conducting outbreak or epidemic investigations (comparing exposed vs unexposed), (5) generating hypotheses from existing data (medical records, disease registries), (6) resources are limited (case-control is cheaper and faster than cohort), (7) studying past exposures where follow-up is impossible.
Strengths & limitations
- Efficiency for rare outcomes: outcome-enriched sampling makes case-control far more efficient than cohort studies for rare diseases. A study of rare cancer might enroll 300 cases and 300 controls instead of 100,000 people followed for years.
- Speed and cost: retrospective data collection is faster and cheaper than prospective follow-up.
- Multiple exposures: a single case-control study can examine many exposures simultaneously, identifying new risk factors.
- Established disease registries: many diseases have registries (cancer, birth defects), providing ready access to well-characterized cases.
- Latency: case-control naturally handles long induction-latency periods between exposure and outcome diagnosis.
- Recall bias: cases, having experienced the outcome, may recall past exposures differently than controls. If cases exaggerate or minimize exposure reporting, OR is biased.
- Selection bias: how controls are chosen critically affects OR. Controls must come from the same source population that generated cases; using hospitalized controls when cases are community-derived introduces bias.
- Temporal ambiguity: retrospective exposure measurement can obscure causality. Did exposure precede outcome, or did outcome influence exposure recall?
- Cannot calculate incidence: case-control samples are outcome-enriched, so you cannot directly estimate disease incidence or absolute risk. OR is the only estimable association measure.
- Exposure misclassification: retrospective exposure assessment (especially for distant exposures) is error-prone, biasing OR toward the null.
Frequently asked
When is Odds Ratio a good approximation for Relative Risk?
The Odds Ratio (OR) approximates the Relative Risk (RR) when the disease is rare in the population (typically <10% prevalence). At low prevalence, the odds of disease ≈ risk of disease, so OR ≈ RR. When prevalence is >10%, OR overstates the effect relative to RR. For example, if RR = 2, the OR might be 2.5 or higher at higher prevalence. Always clarify that case-control studies yield OR, not RR. If clinical interpretation requires RR, you must convert OR using the baseline risk (prevalence of disease in the source population).
What is a nested case-control study, and why use it?
A nested case-control study selects cases and controls from within a prospective cohort. All case members are identified from the cohort as outcomes occur; controls are randomly sampled from non-cases in the cohort, often matched to cases on age and sex. Nested case-control combines cohort and case-control strengths: temporal sequence and incidence data are available (as in cohorts), yet cost is lower because only case and control subsets are measured for expensive exposures (e.g., biomarkers, genetic testing). Odds Ratio from nested case-control approximates cohort Relative Risk, even if disease is not rare.
How do I choose between matched and unmatched controls?
Matched controls (1:1 matched to cases on age, sex, etc.) increase precision and reduce confounding by matched variables. However, matching is statistically 'expensive': you must use conditional logistic regression, and you cannot estimate matched variables' independent effects. Unmatched controls allow adjustment for many confounders in logistic regression and may be more cost-effective. Matching is most useful for strong confounders (age, sex, calendar year) and when controls are hard to find. If potential confounders are numerous, consider unmatched design with regression adjustment.
How do I distinguish between selection bias and information bias in case-control studies?
Selection bias arises from how you select cases and controls. If cases and controls come from different source populations (e.g., hospitalized cases vs. community controls), the OR is biased. Information bias (misclassification) arises from errors in measuring exposure or outcome. Recall bias is a type of information bias: cases with a serious illness may recall past exposures differently than controls. Both bias OR, but in different ways. To minimize selection bias, ensure cases and controls are from the same source population. To minimize information bias, use objective records (medical charts, biomarkers) instead of recall.
Sources
- Schlesselman, J. J. (1982). Case-Control Studies: Design, Conduct, Analysis. Oxford University Press. ISBN: 978-0195027815
- Rothman, K. J., Lash, T. L., & Greenland, S. (2008). Modern Epidemiology (3rd ed.). Lippincott Williams & Wilkins. ISBN: 978-0781755657
- Greenland, S., & Thomas, D. C. (1990). On the need for the rare disease assumption in case-control studies. American Journal of Epidemiology, 132(2), 374–375. link ↗
How to cite this page
ScholarGate. (2026, June 4). Case-Control Study (Retrospective Case-Control Design). ScholarGate. https://scholargate.app/en/clinical-research/case-control-study-design
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Cohort Study DesignClinical Research↔ compare
- Cross-Sectional Study DesignClinical Research↔ compare