Capture-Recapture for Hidden Crime Populations
Also known as: Multiple Systems Estimation, Mark-Recapture for Hidden Populations, Dark-Figure Population Estimation, Lincoln-Petersen Crime Estimation
Capture-recapture, known in criminology and public health as multiple systems estimation, infers the size of a hidden population — undocumented homicide victims, trafficking victims, problem drug users, undetected offenders — that no single source counts completely. By examining how much two or more incomplete lists overlap, it estimates how many cases were missed by all of them: the 'dark figure' of crime. Borrowed from wildlife ecology, the method was synthesized for human populations by the International Working Group in 1995 and brought to criminal-justice policy by Bird and King.
Key highlights
- Estimates the true size of populations that are systematically under-counted by any single official source.
- Uses data that agencies already collect, requiring linkage rather than expensive new surveys.
- Log-linear and Bayesian formulations can model dependence between sources rather than assuming independence.
- Provides defensible national estimates (with uncertainty) for trafficking, drug use, and other hidden harms to guide policy.
- Grounded in a mature statistical theory shared with ecology and epidemiology, with well-developed software.
Intuition
This section is available to Pro members. Upgrade to Pro
How it works
This section is available to Pro members. Upgrade to Pro
When to use it
Use capture-recapture when the population of interest is hidden or under-recorded, no single source counts it fully, and you have two or more overlapping lists whose records can be reliably linked. It is well suited to estimating modern-slavery and trafficking victims, problem drug users, undocumented or missing homicide victims, and other dark-figure crime counts for policy and resource planning. It is inappropriate when lists cannot be linked accurately, when the population is not closed over the observation window (substantial entry, exit, or death), when list dependencies are strong and unidentifiable, or when there is only one source. In such cases victimization surveys or other estimation strategies may be more defensible.
Strengths & limitations
- Estimates the true size of populations that are systematically under-counted by any single official source.
- Uses data that agencies already collect, requiring linkage rather than expensive new surveys.
- Log-linear and Bayesian formulations can model dependence between sources rather than assuming independence.
- Provides defensible national estimates (with uncertainty) for trafficking, drug use, and other hidden harms to guide policy.
- Grounded in a mature statistical theory shared with ecology and epidemiology, with well-developed software.
- Estimates rest on assumptions — population closure, accurate linkage, and modeled list dependence — that are hard to verify.
- The key quantity is extrapolated to an unobserved category, so results carry wide and sometimes fragile uncertainty.
- Different list-dependence models can fit the observed overlaps equally well yet yield very different totals.
- Two-list estimators assume independence between sources, which is often violated for human populations.
- Heterogeneous capture probabilities (some individuals far more likely to appear on lists) bias estimates if unmodeled.
Common pitfalls
This section is available to Pro members. Upgrade to Pro
Applications
This section is available to Pro members. Upgrade to Pro
Frequently asked
What does capture-recapture assume about the population?
Classic estimators assume the population is closed over the study window (no entries, exits, births, or deaths), that individuals can be accurately matched across lists, that every individual has a non-zero chance of appearing on each list, and — in the simplest two-list case — that the lists are independent. Multiple-list log-linear and Bayesian models relax the independence assumption by estimating dependence between sources, but closure and accurate linkage remain essential.
Why are at least two lists required, and why are three or more better?
A single list cannot reveal what it missed, so at least two overlapping lists are needed to estimate coverage from their overlap. With only two lists you must assume independence, which is usually unrealistic for people. Three or more lists let you estimate the dependence between sources directly through interaction terms in a log-linear model, producing more credible totals and allowing model comparison and sensitivity analysis.
How reliable are multiple-systems-estimation numbers for policy?
They are often the best available estimates of otherwise uncountable hidden populations, but they are estimates with genuine uncertainty, not exact counts. Their credibility depends on linkage quality, the plausibility of the closure assumption, and how robust the total is across competing dependence models. Good practice reports a confidence interval, shows sensitivity across models, and frames the figure as an informed range rather than a precise headline number.
Sources
- 1.Bird, S. M., & King, R. (2018). Multiple systems estimation (or capture-recapture estimation) to inform public policy. Annual Review of Statistics and Its Application, 5, 95–118.
- 2.International Working Group for Disease Monitoring and Forecasting (1995). Capture-recapture and multiple-record systems estimation I: History and theoretical development. American Journal of Epidemiology, 142(10), 1047–1058.
You have read it. What now?
Cite this page
ScholarGate. (2026, June 22). Capture-Recapture for Hidden Crime Populations. ScholarGate. https://scholargate.app/criminology/capture-recapture-crime