Counterfactual Impact Evaluation in Education Research
Also known as: CIE in education, counterfactual program evaluation, causal impact evaluation, education policy impact evaluation
Counterfactual impact evaluation (CIE) is the systematic application of causal inference designs — such as difference-in-differences, regression discontinuity, matching, and instrumental variables — to measure the genuine effect of education programs, policies, or interventions by constructing a credible counterfactual: what would have happened to participants had they not been treated.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use counterfactual impact evaluation when you need to move beyond descriptive comparisons to answer whether an education program caused improved outcomes — e.g., attendance, test scores, graduation rates, earnings. It is appropriate when administrative or survey data on participants and non-participants exist, and when at least one credible design (random assignment, cutoff, staggered rollout, or natural experiment) can be exploited. Do not use it when no plausible comparison group can be constructed, when data are too small to support the chosen design, or when the research question is descriptive rather than causal.
Strengths & limitations
- Provides credible causal evidence rather than mere correlations, which is essential for policy decisions in education.
- A family of designs rather than a single method — can be matched to the institutional context and available data.
- Effect sizes from well-designed CIE studies can be directly compared across programs and settings, supporting evidence synthesis.
- Can be applied to both experimental data (RCT) and non-experimental administrative records, making it broadly feasible.
- Transparent about identifying assumptions, enabling peer scrutiny and replication.
- Every design rests on a critical identifying assumption that is not directly testable from data alone; if violated, causal claims are invalid.
- Requires sufficient sample sizes — small programs or single-school studies often lack the power needed for reliable estimates.
- Typically estimates local effects (e.g., ATT or LATE) that may not generalize to other populations or contexts.
- Longitudinal administrative data, which CIE often requires, may not be accessible due to privacy restrictions.
- The method produces impact estimates for a specific program version; changes in implementation can alter the effect.
Frequently asked
Is counterfactual impact evaluation a single method or a family of methods?
It is a framework — the goal of constructing a credible counterfactual — that is implemented through different designs depending on context: RCT, DiD, RDD, matching, or IV. The choice of design is driven by the available data and the institutional setting of the education program.
How is this different from a simple pre-post comparison?
A pre-post comparison attributes all observed change to the program, ignoring secular trends, maturation, or concurrent events that would have caused change anyway. CIE uses a comparison group to estimate what would have happened without the program, subtracting that counterfactual change to isolate the program's causal contribution.
What sample size do I need?
It depends on the design and the expected effect size. RCTs and DiD designs for small education programs typically require hundreds of students per arm to detect moderate effects (d ≈ 0.2–0.3). RDD requires a sufficient density of observations around the eligibility cutoff. A power analysis before data collection is strongly advised.
Can I use CIE with administrative school data?
Yes — administrative records on enrollment, attendance, grades, and graduation are well-suited, especially for DiD and RDD designs. The key requirements are a consistent longitudinal identifier, a comparison cohort or eligibility cutoff, and outcome data measured both before and after the program.
What if my program was universally rolled out with no comparison group?
Universal rollouts with no untreated comparison are the hardest case for CIE. Possible workarounds include using a synthetic control method, exploiting staggered rollout timing across districts, or using a regression-discontinuity in an eligibility score. If none applies, causal inference is not credible and the study should be framed as descriptive.
Sources
- Blundell, R., & Costa Dias, M. (2002). Alternative approaches to evaluation in empirical microeconomics. Portuguese Economic Journal, 1(2), 91-115. DOI: 10.1007/s10258-002-0010-3 ↗
- Cerulli, G. (2015). Econometric Evaluation of Socio-Economic Programs: Theory and Applications. Springer. ISBN: 978-3-662-46400-2
How to cite this page
ScholarGate. (2026, June 3). Counterfactual Impact Evaluation in Education Research. ScholarGate. https://scholargate.app/en/causal-inference/counterfactual-impact-evaluation-in-education-research
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Difference-in-DifferencesEconometrics↔ compare
- Instrumental Variables in Health ResearchHealth Economics↔ compare
- Interrupted Time SeriesCausal inference↔ compare
- Propensity Score MatchingResearch Statistics↔ compare
- Synthetic Control MethodCausal inference↔ compare