Coarsened Exact Matching in Education Research
Coarsened Exact Matching for Causal Inference in Education Research · Also known as: CEM in education, CEM for educational studies, exact matching education, coarsened matching educational data
Coarsened Exact Matching (CEM) is a pre-processing matching strategy that reduces imbalance between treated and comparison groups before outcome analysis. In education research it is used to create balanced comparison groups from administrative records, survey data, or quasi-experimental study designs — for example comparing students who received an intervention against comparable students who did not, without relying on randomisation.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use CEM in education research when you have observational data — administrative records, national surveys, or programme participation files — and want to estimate the effect of an educational treatment (a tutoring programme, school choice policy, curriculum reform, or teacher professional development) on outcomes such as test scores, graduation, or attendance. It works well when covariates include a mixture of continuous measures (prior achievement, income) and categorical variables (race, school type) and when the analytic sample is large enough to sustain pruning. Do not use CEM when: (1) the dataset is very small and pruning would leave too few matched units for reliable inference; (2) you have many continuous covariates with no natural coarsening points, making exact matching across all combinations impossible; or (3) a key confounder is unobserved, since matching only balances measured characteristics.
Strengths & limitations
- Transparent and intuitive: researchers explicitly choose coarsening cutpoints and can inspect the matched strata, making the comparison logic easy to communicate to education policy audiences.
- Does not require estimating a propensity score model, so balance does not depend on correct specification of a logistic or probit model.
- Simultaneously balances all matched covariates rather than optimising a scalar distance, which can yield better joint covariate balance than nearest-neighbour propensity matching.
- Matched data can be analysed with any standard regression tool (OLS, logistic, multilevel models) to adjust for remaining within-stratum imbalance.
- Reduces researcher degrees of freedom: coarsening rules are set before seeing outcomes, limiting ad hoc data fishing.
- Pruning unmatched units can substantially reduce sample size, cutting statistical power — a particular concern in small education studies or when the treatment group is a niche subpopulation.
- The choice of coarsening cutpoints is subjective; different reasonable choices can lead to different matched samples and point estimates, requiring sensitivity analysis.
- Like all matching methods, CEM only controls for observed confounders; unobserved factors such as student motivation or family support may still bias the estimate.
- With many covariates, matching on all jointly can produce very sparse or empty strata, forcing researchers to either coarsen aggressively or drop covariates.
- Estimates apply only to the matched sample (the region of common support), which may not represent all students enrolled in the programme.
Frequently asked
How do I choose the coarsening cutpoints for prior test scores?
Common choices include quantile-based bins (quartiles or quintiles of the score distribution), substantively meaningful cut-scores (proficiency thresholds), or equal-width intervals. You should choose cutpoints before examining outcomes, report your choices transparently, and run a sensitivity analysis under at least one alternative coarsening to show results are not driven by the particular choice.
How much sample loss is acceptable after pruning?
There is no universal rule. A loss of 20-30% of comparison units is generally tolerable if balance improves substantially. If pruning eliminates more than half the treatment group, the matched sample may be too small for reliable inference or may represent only a narrow, unrepresentative subgroup of treated students.
Should I run a regression on the matched sample or just compare means?
Running a regression (OLS or multilevel, depending on data structure) on the matched sample is strongly recommended. It adjusts for any residual within-stratum covariate imbalance and accounts for clustering of students within schools, improving precision and reducing remaining bias.
How does CEM compare to propensity score matching in education data?
CEM does not require a correctly specified propensity model and tends to achieve better multivariate balance directly. However, propensity score matching retains more of the sample when covariates are continuous and high-dimensional. In large administrative datasets where sample loss from pruning is acceptable, CEM often outperforms propensity matching on balance.
Can I use CEM with multilevel data, such as students nested within schools?
Yes. After matching, you should use a multilevel or cluster-robust regression model to account for the school-level clustering. Some researchers also include school fixed effects in the post-matching regression to absorb school-level unobservables, provided enough within-school variation remains after matching.
Sources
- Iacus, S. M., King, G., & Porro, G. (2012). Causal inference without balance checking: Coarsened exact matching. Political Analysis, 20(1), 1-24. DOI: 10.1093/pan/mpr013 ↗
- Morgan, S. L., & Winship, C. (2015). Counterfactuals and Causal Inference: Methods and Principles for Social Research (2nd ed.). Cambridge University Press. ISBN: 978-1107065079
How to cite this page
ScholarGate. (2026, June 3). Coarsened Exact Matching for Causal Inference in Education Research. ScholarGate. https://scholargate.app/en/causal-inference/coarsened-exact-matching-in-education-research
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Coarsened Exact MatchingCausal inference↔ compare
- Difference-in-Differences in Education ResearchCausal inference↔ compare
- Propensity Score Matching in Education ResearchCausal inference↔ compare
- Regression discontinuity design in education researchCausal inference↔ compare