Propensity Weighting in Criminology
Also known as: IPTW for Justice Exposures, Inverse-Probability Weighting in Criminology, Propensity-Weighted Crime Effects, Observational Treatment-Effect Weighting
Propensity weighting estimates the causal effect of a justice exposure — incarceration, gang membership, a program, or a sanction — from observational data when randomization was impossible. It models each unit's probability of receiving the exposure given measured confounders (the propensity score) and then weights units by the inverse of that probability, creating a pseudo-population in which the exposure is unrelated to those confounders. Rosenbaum and Rubin introduced the propensity score in 1983, and Apel and Sweeten adapted it for criminology, where ethical and practical barriers make experiments rare.
Key highlights
- Uses all units rather than discarding unmatched cases, often giving more efficient estimates than one-to-one matching.
- Separates the design stage (achieving covariate balance) from outcome analysis, guarding against outcome-driven model tweaking.
- Can target the ATE or the ATT simply by changing the weighting scheme, fitting different policy questions.
- Extends naturally to doubly robust estimators and to time-varying exposures via marginal structural models.
- Provides transparent, checkable balance diagnostics that make the credibility of the comparison auditable.
Intuition
This section is available to Pro members. Upgrade to Pro
How it works
This section is available to Pro members. Upgrade to Pro
When to use it
Use propensity weighting when you want the causal effect of a binary justice exposure but cannot randomize, and you have measured the key variables that drive selection into that exposure. It is well suited to questions about the effect of incarceration versus community sanctions, gang membership, program participation, or arrest on later offending, employment, or recidivism. It is inappropriate when important confounders are unmeasured (weighting cannot fix unobserved selection), when there is little overlap between exposed and unexposed units (extreme weights destabilize estimates), or when the exposure is continuous or time-varying without the corresponding marginal structural model machinery. In those cases an experiment, a regression-discontinuity design, or sensitivity analysis is preferable.
Strengths & limitations
- Uses all units rather than discarding unmatched cases, often giving more efficient estimates than one-to-one matching.
- Separates the design stage (achieving covariate balance) from outcome analysis, guarding against outcome-driven model tweaking.
- Can target the ATE or the ATT simply by changing the weighting scheme, fitting different policy questions.
- Extends naturally to doubly robust estimators and to time-varying exposures via marginal structural models.
- Provides transparent, checkable balance diagnostics that make the credibility of the comparison auditable.
- Relies on the strong, untestable assumption of no unmeasured confounding; omitted selection variables bias the estimate.
- Extreme propensity scores near 0 or 1 produce very large weights that inflate variance and let a few units dominate.
- Requires good overlap (common support) between exposed and unexposed groups; without it, effects are extrapolated, not estimated.
- Results depend on correct specification of the propensity model, and balance can look good on included covariates while hiding others.
- Standard errors must account for the estimated weights, so naive variance estimates understate uncertainty.
Common pitfalls
This section is available to Pro members. Upgrade to Pro
Applications
This section is available to Pro members. Upgrade to Pro
Frequently asked
How is weighting different from propensity score matching?
Both use the propensity score to remove confounding, but matching pairs each treated unit with one or more similar controls and discards the rest, whereas weighting keeps all units and reweights them by the inverse probability of their observed exposure. Weighting typically uses the data more efficiently and targets the ATE or ATT cleanly, but it is more sensitive to extreme propensity scores, which matching sidesteps by trimming to the region of overlap.
What does inverse-probability weighting actually balance?
It balances only the covariates that enter the propensity model. By up-weighting units that were unlikely to receive the exposure they actually got, weighting constructs a pseudo-population in which the measured confounders are distributed identically across exposed and unexposed groups. It cannot balance variables you did not measure, which is why the no-unmeasured-confounding assumption is the design's Achilles heel and motivates sensitivity analysis.
Why do extreme weights cause problems?
When a treated unit has a propensity score near zero (it looked very unlikely to be treated) its inverse-probability weight becomes enormous, so that single observation can dominate the weighted average, inflate variance, and destabilize the estimate. Analysts address this by stabilizing weights, trimming or truncating extreme values, and restricting analysis to the region of common support where both exposed and unexposed units exist.
Sources
- 1.Apel, R. J., & Sweeten, G. (2010). Propensity score matching in criminology and criminal justice. In A. R. Piquero & D. Weisburd (Eds.), Handbook of Quantitative Criminology (pp. 543–562). Springer.
- 2.Rosenbaum, P. R., & Rubin, D. B. (1983). The central role of the propensity score in observational studies for causal effects. Biometrika, 70(1), 41–55.
You have read it. What now?
Cite this page
ScholarGate. (2026, June 22). Propensity Weighting in Criminology. ScholarGate. https://scholargate.app/criminology/propensity-weighting-crime