Process / pipelineEpidemiologyClinical / epidemiologyPipeline

Retrospective Cohort Study

Also known as: historical cohort study, non-concurrent cohort study, retrospective follow-up study, historical prospective study

OriginatorSystematic use attributed to early 20th-century occupational epidemiology; formalized in modern epidemiological theory by Brian MacMahon and othersYearMid-20th century (widely formalized 1950s–1970s)Sources2Related methods22

A retrospective cohort study assembles a group of individuals who share a common starting point and reconstructs their exposure history and subsequent outcomes entirely from pre-existing records. Because the data have already been collected before the study begins, the design is far faster and cheaper than a prospective cohort; however, the researcher must work with whatever information was recorded at the time rather than collecting purpose-built measurements.

Key highlights

  • Dramatically faster and less costly than a prospective cohort study because follow-up has already occurred.
  • Ideal for outcomes with long latency periods — decades of follow-up can be studied without waiting decades.
  • Allows estimation of incidence rates and relative risks — measures unavailable in case-control designs.
  • Large sample sizes are often feasible by linking routinely collected administrative or registry databases.
  • Can evaluate multiple outcomes from a single exposed cohort simultaneously.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use a retrospective cohort study when you need to estimate the relationship between an exposure and an outcome but cannot wait years for follow-up, and when high-quality historical records on exposure and outcome already exist. It is particularly well suited to occupational epidemiology, pharmacoepidemiology, and healthcare database research where administrative or registry data are routinely collected. Avoid it when the exposure of interest was not systematically recorded in historical records, when record quality is poor or non-standardized across sites, when individual-level confounders are largely unmeasured in the available data, or when the outcome is rare enough that a case-control design would be more efficient.

Strengths & limitations

Strengths
  • Dramatically faster and less costly than a prospective cohort study because follow-up has already occurred.
  • Ideal for outcomes with long latency periods — decades of follow-up can be studied without waiting decades.
  • Allows estimation of incidence rates and relative risks — measures unavailable in case-control designs.
  • Large sample sizes are often feasible by linking routinely collected administrative or registry databases.
  • Can evaluate multiple outcomes from a single exposed cohort simultaneously.
Limitations
  • Dependent entirely on the quality, completeness, and detail of pre-existing records — variables not recorded historically cannot be used.
  • Exposure assessment is typically cruder than in prospective studies, increasing the risk of misclassification and bias toward the null.
  • Unmeasured confounding is harder to address when the historical data were not collected for research purposes.
  • Selection bias can arise if membership in the cohort at baseline was influenced by early symptoms of the outcome (healthy worker effect in occupational settings).
  • Loss to follow-up is difficult to assess and may be differential if sicker individuals left the record system earlier.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

How is a retrospective cohort study different from a case-control study?

Both designs use existing data, but they differ in sampling logic. A retrospective cohort study starts with an exposed group and an unexposed group and compares outcome incidence between them; it yields relative risks and incidence rates. A case-control study starts with outcome cases and disease-free controls and compares past exposure prevalence; it yields odds ratios. When the outcome is rare, the case-control design is more efficient. When multiple outcomes are of interest or when the outcome is not rare, the retrospective cohort is preferred.

When is a retrospective cohort better than a prospective cohort?

A retrospective cohort is preferable when the outcome has a long latency (decades), when waiting for events to accumulate is impractical, or when the research question must be answered quickly and suitable historical records already exist. A prospective cohort is preferable when exposure measurement must be precise and standardized, when relevant confounders are not in existing records, or when the research question requires biological samples collected prospectively.

What is immortal time bias and how do I avoid it?

Immortal time bias occurs when person-time during which subjects could not have experienced the outcome — for example, the period between cohort entry and receiving an exposure that defines group membership — is incorrectly assigned to the exposed group. It makes the exposed group look artificially healthier. Avoid it by aligning cohort entry, exposure start, and follow-up start carefully, and by using time-varying exposure coding in survival models rather than fixed baseline exposure classification.

How should I handle missing exposure data in historical records?

First, characterise the extent and likely mechanism of missingness (missing completely at random, at random, or not at random). If missingness is limited and plausibly random, complete-case analysis with sensitivity analysis is acceptable. For larger proportions of missing data, multiple imputation is preferred, provided the variables needed to predict missingness are available. Worst-case scenario analyses help bound uncertainty. Avoid imputing the exposure variable itself using the outcome, as this can introduce bias.

Can a retrospective cohort study establish causality?

Like all observational designs, a retrospective cohort study can demonstrate association and temporal precedence of exposure before outcome, but it cannot rule out unmeasured confounding without additional evidence. Causal claims are strengthened by applying the Bradford Hill criteria (strength, consistency, specificity, dose-response, biological plausibility), by conducting negative control analyses, and by triangulating findings with other study designs including randomized trials where available.

Sources

  1. 1.
    Rothman, K. J., Greenland, S., & Lash, T. L. (2008). Modern Epidemiology (3rd ed.). Lippincott Williams & Wilkins.
    ISBN 978-0781755641
  2. 2.
    Retrospective cohort study. Wikipedia.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 3). Retrospective Cohort Study. ScholarGate. https://scholargate.app/epidemiology/retrospective-cohort-study