Process / pipelineResearch DesignSurvey / observational designPipeline

Longitudinal Ex Post Facto Design — Tracking Pre-Existing Groups Over Time

Also known as: longitudinal causal-comparative design, longitudinal after-the-fact design, longitudinal retrospective design, LEPF design

OriginatorFred N. Kerlinger (systematized); Donald T. Campbell & Julian C. Stanley (quasi-experimental framework)Year1964–1986 (Kerlinger 1964 first edition; Campbell & Stanley 1966)Sources2Related methods6

A longitudinal ex post facto design combines the time-depth of longitudinal research with the retrospective logic of ex post facto inquiry. Participants are grouped by a naturally occurring characteristic or past event — not randomly assigned — and then observed or measured at multiple points over time. The goal is to trace how pre-existing differences between groups unfold or predict outcomes across an extended period, without the researcher ever manipulating the independent variable.

Key highlights

  • Enables causal-direction inference by establishing temporal precedence — outcomes are measured after the pre-existing condition, not simultaneously.
  • Captures developmental and cumulative processes that cross-sectional designs cannot detect.
  • Practical and ethical for studying conditions that cannot be experimentally induced (adverse experiences, genetic traits, past policy exposure).
  • Can exploit rich administrative and registry datasets to study large representative samples over long periods.
  • Applicable across disciplines — from education and developmental psychology to public health and economics.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use a longitudinal ex post facto design when (1) the independent variable is a past event, naturally occurring condition, or ethically or practically unmanipulable characteristic; (2) the research question involves change, development, or cumulative effects over time; and (3) archival data, linked administrative records, or a committed cohort are available for repeated measurement. It is well suited to developmental psychology, educational research, epidemiology, sociology, and policy evaluation. Do NOT use it when random assignment is feasible — a true or quasi-experiment provides stronger causal inference. Do not use it when only a single time point is available (use cross-sectional ex post facto instead), or when attrition is expected to be severe and differential across groups, which would introduce selection bias that is difficult to correct.

Strengths & limitations

Strengths
  • Enables causal-direction inference by establishing temporal precedence — outcomes are measured after the pre-existing condition, not simultaneously.
  • Captures developmental and cumulative processes that cross-sectional designs cannot detect.
  • Practical and ethical for studying conditions that cannot be experimentally induced (adverse experiences, genetic traits, past policy exposure).
  • Can exploit rich administrative and registry datasets to study large representative samples over long periods.
  • Applicable across disciplines — from education and developmental psychology to public health and economics.
Limitations
  • Random assignment is absent, so unmeasured confounders may explain group differences — causal inference is always qualified.
  • Participant attrition across waves can introduce selection bias, particularly if dropout is related to the outcome.
  • Retrospective identification of the independent variable may rely on incomplete or inconsistently recorded archival data.
  • Long follow-up periods are costly and operationally demanding, requiring sustained funding and participant engagement.
  • Historical changes occurring between waves (history effects) may confound observed group differences.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

What is the core difference between a longitudinal ex post facto design and a prospective cohort study?

The terms overlap substantially in practice. Both follow participants over time and observe outcomes without experimental manipulation. The ex post facto label emphasizes that the defining group characteristic (the 'independent variable') is a past event or condition that pre-dates the study's measurement start, whereas prospective cohort studies measure baseline exposure status at enrollment and follow participants forward. In a longitudinal ex post facto design the exposure may be identified retrospectively from records, while a prospective cohort assembles participants before outcomes occur and measures exposure prospectively. Both face the same core limitation: no random assignment.

How many measurement waves do I need?

A minimum of two waves is required to call a design longitudinal. However, two waves only allow you to measure net change; they cannot distinguish different patterns of change (e.g., rapid early gain versus steady linear growth). Three or more waves are needed for growth-curve modeling. The required number of waves depends on the shape of the expected change trajectory — consult substantive theory and prior literature to determine an appropriate measurement schedule.

Can I use matching or propensity scores to make causal claims?

Propensity score matching and related methods reduce observed confounding and make the study more rigorous, but they only balance on measured covariates. If important confounders are unmeasured, matched estimates remain biased. Causal language ('X caused Y') is therefore not warranted; you may say the findings are consistent with a causal effect, or that a causal interpretation is plausible given the controls applied.

How should I handle missing data from participant dropout?

First, analyze and report attrition patterns — compare baseline characteristics of completers versus dropouts overall and by group. If dropout is not random (missing not at random), listwise deletion produces biased estimates. Use full information maximum likelihood (FIML) estimation or multiple imputation to handle missing data under the missing-at-random assumption, and consider sensitivity analyses for non-random dropout.

Is this design appropriate for dissertation research?

Yes, if existing archival or administrative data are available or if a short longitudinal window (6–18 months) is feasible within the dissertation timeline. Collecting primary longitudinal data over multiple years is rarely practical for a dissertation; exploiting existing panel datasets (e.g., PISA, NLSY, BHPS, national educational databases) is the standard strategy. Ensure your institution's IRB/ethics committee approves secondary data use and that data access agreements are in place before committing to the design.

Sources

  1. 1.
    Kerlinger, F. N. (1986). Foundations of Behavioral Research (3rd ed.). Holt, Rinehart and Winston.
    ISBN 978-0030417498
  2. 2.
    Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and Quasi-Experimental Designs for Generalized Causal Inference. Houghton Mifflin.
    ISBN 978-0395615560

You have read it. What now?

Cite this page

ScholarGate. (2026, June 3). Longitudinal Ex Post Facto Design. ScholarGate. https://scholargate.app/research-design/longitudinal-ex-post-facto-design

Longitudinal Ex Post Facto Design | ScholarGate