Impact Evaluation Design
Also known as: Impact Evaluation, Causal Impact Evaluation Design, Counterfactual Evaluation Design
Impact evaluation design is the upstream task of structuring an evaluation so that it can credibly attribute changes in outcomes to a policy or program rather than to other factors. Its defining concern is the counterfactual: what would have happened to participants in the absence of the intervention. Codified in resources such as the World Bank's Impact Evaluation in Practice, the design process selects an identification strategy — randomised assignment, or a quasi-experimental method such as difference-in-differences, regression discontinuity, instrumental variables or matching — that constructs a valid comparison and yields an unbiased estimate of the intervention's effect.
Key highlights
- Forces explicit attention to the counterfactual, the crux of credible causal inference.
- Provides a menu of identification strategies matched to how programs are actually assigned.
- Front-loads power and measurement decisions so the study can answer its question.
- Yields quantified, defensible effect estimates suitable for high-stakes scaling and accountability decisions.
Intuition
This section is available to Pro members. Upgrade to Pro
How it works
This section is available to Pro members. Upgrade to Pro
When to use it
Use a formal impact evaluation design when the decision at stake genuinely requires a credible causal estimate of a program's effect — for scaling, continuation, or rigorous accountability — and when a valid counterfactual can be constructed, ideally by designing the evaluation before or alongside program rollout. It assumes outcomes are measurable, a comparison group is obtainable, and identifying assumptions are plausible. It is less appropriate when the program is universal with no possible comparison, when effects are too diffuse to attribute, or when the priority is understanding mechanisms and context rather than net effect — where theory-based approaches such as contribution analysis or realist evaluation are better. It is the design-stage scaffolding beneath specific estimators like difference-in-differences and regression discontinuity.
Strengths & limitations
- Forces explicit attention to the counterfactual, the crux of credible causal inference.
- Provides a menu of identification strategies matched to how programs are actually assigned.
- Front-loads power and measurement decisions so the study can answer its question.
- Yields quantified, defensible effect estimates suitable for high-stakes scaling and accountability decisions.
- A valid counterfactual is often hard or impossible to construct, especially for universal or system-wide policies.
- Strong designs can be costly, slow and demanding of data and analytic expertise.
- Estimates internal effects well but may have limited external validity beyond the study setting.
- Focuses on the net effect, telling little about why, for whom, or under what conditions a program works.
Common pitfalls
This section is available to Pro members. Upgrade to Pro
Applications
This section is available to Pro members. Upgrade to Pro
Frequently asked
Why is the counterfactual so central to impact evaluation?
Because impact is defined as the difference between what happened with the program and what would have happened without it — and the latter is never directly observable for those who received the program. Every impact evaluation design is, at heart, a strategy for credibly estimating that missing counterfactual using a comparison group or a source of as-good-as-random variation. If the counterfactual is poorly approximated, the estimate confuses the program's effect with pre-existing differences or concurrent trends, however sophisticated the statistics.
When is randomisation not the right choice?
Randomisation gives the cleanest counterfactual, but it is not always feasible or ethical: a program may already be running, be legally entitled to all eligible people, or be impossible to withhold. In such cases a quasi-experimental design is appropriate — exploiting an eligibility cutoff (regression discontinuity), a staggered or differential rollout (difference-in-differences), an instrument, or matching on observed characteristics. The design should match the way the program was actually assigned, and each quasi-experimental method carries its own identifying assumptions that must be defended.
How does impact evaluation design relate to specific methods like difference-in-differences?
Impact evaluation design is the overarching framework for choosing how to estimate a causal effect; methods such as difference-in-differences, regression discontinuity, propensity-score matching and synthetic control are the specific identification strategies it selects among. The design stage clarifies the causal question, the counterfactual and the available variation, then picks the estimator best suited to them. Choosing the method is therefore one decision within the broader design process, and the same program might be evaluable by several methods of differing credibility.
Sources
- 1.Gertler, P. J., Martinez, S., Premand, P., Rawlings, L. B., & Vermeersch, C. M. J. (2016). Impact Evaluation in Practice (2nd ed.). Washington, DC: World Bank.ISBN 9781464807794
You have read it. What now?
Cite this page
ScholarGate. (2026, June 22). Impact Evaluation Design. ScholarGate. https://scholargate.app/public-policy/impact-evaluation-design