Ecological Inference
Also known as: EI, Ecological regression, King's ecological inference, Aggregate-to-individual inference
Ecological inference is the problem of learning about individual behavior — such as how Black and white voters cast their ballots — when only aggregate data are available, like precinct-level turnout and racial composition. Because individual-level data are missing, the within-group rates are not directly observed; ecological inference recovers them by combining the deterministic accounting constraints that each precinct must satisfy with a statistical model of how the unobserved rates vary across precincts. Gary King's 1997 solution unified the deterministic method of bounds with Leo Goodman's classic ecological regression, sharply reducing the long-standing risk of the ecological fallacy.
Key highlights
- Recovers individual-level rates from aggregate data alone, the only option when no individual records exist.
- The method of bounds provides assumption-free constraints that every estimate must respect, guarding against impossible values.
- King's model lets within-unit rates vary across units and uses the bounds to discipline the statistical inference.
- Has a rigorous, well-developed inferential framework with established software and a clear application in voting-rights litigation.
Intuition
This section is available to Pro members. Upgrade to Pro
How it works
This section is available to Pro members. Upgrade to Pro
When to use it
Use ecological inference when individual-level data are unavailable and you must infer within-group behavior from aggregate units — classically estimating racially polarized voting from precinct returns in Voting Rights Act litigation, or reconstructing turnout and choice by demographic group. It is appropriate when you have many aggregate units with varying composition and credible assumptions about how within-unit rates relate to covariates. It is risky when aggregation bias is strong (within-unit rates correlate with composition in ways the model ignores), when units are few or homogeneous so the bounds are wide, or when the modeling assumptions cannot be checked against any individual-level benchmark.
Strengths & limitations
- Recovers individual-level rates from aggregate data alone, the only option when no individual records exist.
- The method of bounds provides assumption-free constraints that every estimate must respect, guarding against impossible values.
- King's model lets within-unit rates vary across units and uses the bounds to discipline the statistical inference.
- Has a rigorous, well-developed inferential framework with established software and a clear application in voting-rights litigation.
- Fundamentally a problem of partial identification: aggregate data cannot fully determine individual behavior without assumptions.
- Vulnerable to aggregation bias when within-unit rates are correlated with group composition in unmodeled ways.
- Estimates can be sensitive to the assumed distribution of within-unit rates, which is rarely verifiable from aggregate data alone.
- Performs poorly when units are few, homogeneous, or have non-overlapping bounds, leaving the quantities of interest weakly identified.
Common pitfalls
This section is available to Pro members. Upgrade to Pro
Applications
This section is available to Pro members. Upgrade to Pro
Frequently asked
What is the ecological fallacy and how does ecological inference avoid it?
The ecological fallacy, named after Robinson's 1950 finding, is the error of assuming that a correlation observed across aggregate units holds for individuals — for instance, inferring that immigrants are more literate because districts with more immigrants had higher literacy. Ecological inference avoids the fallacy by not equating aggregate associations with individual behavior; instead it uses the deterministic accounting identity and bounds in each unit, plus a model of how within-unit rates vary, to estimate the individual-level rates directly. It treats the within-group rates as the explicit unknowns rather than reading them off an aggregate slope.
How does King's method improve on Goodman's ecological regression?
Goodman's regression assumes the within-group rates are constant across all units, which is efficient when true but yields nonsensical estimates outside the zero-to-one interval when rates actually vary with composition, and it ignores the deterministic bounds. King's model lets each unit have its own pair of rates drawn from a truncated bivariate normal distribution confined to the unit square, estimates that distribution by maximum likelihood, and combines it with the unit-specific bounds. This respects the logical constraints, allows heterogeneity, and produces unit-level estimates and honest uncertainty, though it adds distributional assumptions that must be scrutinized.
When is ecological inference unreliable?
It is unreliable when aggregation bias is strong — that is, when within-unit rates are correlated with group composition in ways the model does not capture — because then no distributional assumption can be trusted. It also struggles when there are few units, when units are demographically homogeneous so the bounds are wide and uninformative, and when no individual-level data exist to validate the assumptions. In such cases the quantities of interest are only weakly identified, and analysts should report the bounds, conduct sensitivity analysis, and treat point estimates cautiously rather than as definitive.
Sources
- 1.King, G. (1997). A Solution to the Ecological Inference Problem: Reconstructing Individual Behavior from Aggregate Data. Princeton: Princeton University Press.ISBN 9780691012414
- 2.Goodman, L. A. (1953). Ecological Regressions and Behavior of Individuals. American Sociological Review, 18(6), 663–664.
You have read it. What now?
Cite this page
ScholarGate. (2026, June 22). Ecological Inference. ScholarGate. https://scholargate.app/political-science/ecological-inference