Regression modelCausal inferenceQuasi-experimental / causal inferenceModel

Coarsened Exact Matching in Education Research

Also known as: CEM in education, CEM for educational studies, exact matching education, coarsened matching educational data

OriginatorIacus, King, & PorroYear2012Sources2Related methods4

Coarsened Exact Matching (CEM) is a pre-processing matching strategy that reduces imbalance between treated and comparison groups before outcome analysis. In education research it is used to create balanced comparison groups from administrative records, survey data, or quasi-experimental study designs — for example comparing students who received an intervention against comparable students who did not, without relying on randomisation.

Key highlights

  • Transparent and intuitive: researchers explicitly choose coarsening cutpoints and can inspect the matched strata, making the comparison logic easy to communicate to education policy audiences.
  • Does not require estimating a propensity score model, so balance does not depend on correct specification of a logistic or probit model.
  • Simultaneously balances all matched covariates rather than optimising a scalar distance, which can yield better joint covariate balance than nearest-neighbour propensity matching.
  • Matched data can be analysed with any standard regression tool (OLS, logistic, multilevel models) to adjust for remaining within-stratum imbalance.
  • Reduces researcher degrees of freedom: coarsening rules are set before seeing outcomes, limiting ad hoc data fishing.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use CEM in education research when you have observational data — administrative records, national surveys, or programme participation files — and want to estimate the effect of an educational treatment (a tutoring programme, school choice policy, curriculum reform, or teacher professional development) on outcomes such as test scores, graduation, or attendance. It works well when covariates include a mixture of continuous measures (prior achievement, income) and categorical variables (race, school type) and when the analytic sample is large enough to sustain pruning. Do not use CEM when: (1) the dataset is very small and pruning would leave too few matched units for reliable inference; (2) you have many continuous covariates with no natural coarsening points, making exact matching across all combinations impossible; or (3) a key confounder is unobserved, since matching only balances measured characteristics.

Strengths & limitations

Strengths
  • Transparent and intuitive: researchers explicitly choose coarsening cutpoints and can inspect the matched strata, making the comparison logic easy to communicate to education policy audiences.
  • Does not require estimating a propensity score model, so balance does not depend on correct specification of a logistic or probit model.
  • Simultaneously balances all matched covariates rather than optimising a scalar distance, which can yield better joint covariate balance than nearest-neighbour propensity matching.
  • Matched data can be analysed with any standard regression tool (OLS, logistic, multilevel models) to adjust for remaining within-stratum imbalance.
  • Reduces researcher degrees of freedom: coarsening rules are set before seeing outcomes, limiting ad hoc data fishing.
Limitations
  • Pruning unmatched units can substantially reduce sample size, cutting statistical power — a particular concern in small education studies or when the treatment group is a niche subpopulation.
  • The choice of coarsening cutpoints is subjective; different reasonable choices can lead to different matched samples and point estimates, requiring sensitivity analysis.
  • Like all matching methods, CEM only controls for observed confounders; unobserved factors such as student motivation or family support may still bias the estimate.
  • With many covariates, matching on all jointly can produce very sparse or empty strata, forcing researchers to either coarsen aggressively or drop covariates.
  • Estimates apply only to the matched sample (the region of common support), which may not represent all students enrolled in the programme.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

How do I choose the coarsening cutpoints for prior test scores?

Common choices include quantile-based bins (quartiles or quintiles of the score distribution), substantively meaningful cut-scores (proficiency thresholds), or equal-width intervals. You should choose cutpoints before examining outcomes, report your choices transparently, and run a sensitivity analysis under at least one alternative coarsening to show results are not driven by the particular choice.

How much sample loss is acceptable after pruning?

There is no universal rule. A loss of 20-30% of comparison units is generally tolerable if balance improves substantially. If pruning eliminates more than half the treatment group, the matched sample may be too small for reliable inference or may represent only a narrow, unrepresentative subgroup of treated students.

Should I run a regression on the matched sample or just compare means?

Running a regression (OLS or multilevel, depending on data structure) on the matched sample is strongly recommended. It adjusts for any residual within-stratum covariate imbalance and accounts for clustering of students within schools, improving precision and reducing remaining bias.

How does CEM compare to propensity score matching in education data?

CEM does not require a correctly specified propensity model and tends to achieve better multivariate balance directly. However, propensity score matching retains more of the sample when covariates are continuous and high-dimensional. In large administrative datasets where sample loss from pruning is acceptable, CEM often outperforms propensity matching on balance.

Can I use CEM with multilevel data, such as students nested within schools?

Yes. After matching, you should use a multilevel or cluster-robust regression model to account for the school-level clustering. Some researchers also include school fixed effects in the post-matching regression to absorb school-level unobservables, provided enough within-school variation remains after matching.

Sources

  1. 1.
    Iacus, S. M., King, G., & Porro, G. (2012). Causal inference without balance checking: Coarsened exact matching. Political Analysis, 20(1), 1-24.
  2. 2.
    Morgan, S. L., & Winship, C. (2015). Counterfactuals and Causal Inference: Methods and Principles for Social Research (2nd ed.). Cambridge University Press.
    ISBN 978-1107065079

You have read it. What now?

Cite this page

ScholarGate. (2026, June 3). Coarsened Exact Matching in Education Research. ScholarGate. https://scholargate.app/causal-inference/coarsened-exact-matching-in-education-research

Coarsened Exact Matching in Education Research | ScholarGate