Regression modelSocial EpidemiologyCounterfactual mean decomposition / explained-unexplained partitionModel

Oaxaca-Blinder Health Decomposition

Also known as: Blinder-Oaxaca Decomposition for Health Inequalities, Threefold Decomposition of Health Disparities, Detailed Decomposition of Health Gaps, Nonlinear Oaxaca-Blinder for Binary Health Outcomes

OriginatorRonald Oaxaca; Alan Blinder (health extension popularized by Fairlie and others)Year1973Sources4Related methods6

The Oaxaca-Blinder decomposition partitions the mean difference in a health outcome between two groups into a portion explained by differences in their measured characteristics and a residual, unexplained portion attributed to differences in how those characteristics translate into health. Developed independently by Ronald Oaxaca (1973) and Alan Blinder (1973) to study labor-market wage gaps, the method was imported into social epidemiology to quantify, for example, how much of a Black-White, urban-rural, or rich-poor gap in self-rated health, BMI, hypertension, or mortality is accounted for by differences in socioeconomic exposures versus differences in returns to those exposures. Group-specific regressions are estimated, the gap in fitted means is written as a function of mean covariates and coefficients, and that gap is algebraically split into an explained (composition) component and an unexplained (coefficient) component, each of which can be further decomposed variable by variable.

Key highlights

  • Cleanly separates a health disparity into a composition (explained) part and a returns (unexplained) part, sharpening the interpretation of group gaps.
  • Supports a detailed, variable-by-variable accounting that identifies which specific exposures drive the explained portion of a health gap.
  • Familiar, widely implemented, and comparable across studies, enabling cumulative evidence on the sources of health inequalities.
  • Extends via Fairlie and related work to binary and count health outcomes, covering the many epidemiological outcomes that are not continuous.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use Oaxaca-Blinder health decomposition when you observe a mean difference in a health outcome between two well-defined groups and want to know how much of that disparity is attributable to differences in measured risk factors versus differences in the effects of those risk factors. It is appropriate when you have individual-level data with a rich, comparable covariate set in both groups, when the comparison is genuinely two-group (or can be reduced to one), and when the goal is descriptive accounting of a gap rather than estimation of a single causal effect. Reach for the nonlinear (Fairlie) variant when the outcome is binary or a count. The method is less suitable when the covariates are sparse or non-comparable across groups, when the outcome distributions barely overlap, or when stakeholders will misread the unexplained component as a clean estimate of discrimination; it should be paired with sensitivity analysis over the reference-coefficient choice and explicit caveats about omitted variables.

Strengths & limitations

Strengths
  • Cleanly separates a health disparity into a composition (explained) part and a returns (unexplained) part, sharpening the interpretation of group gaps.
  • Supports a detailed, variable-by-variable accounting that identifies which specific exposures drive the explained portion of a health gap.
  • Familiar, widely implemented, and comparable across studies, enabling cumulative evidence on the sources of health inequalities.
  • Extends via Fairlie and related work to binary and count health outcomes, covering the many epidemiological outcomes that are not continuous.
Limitations
  • The split depends on the choice of reference coefficients (the index-number problem), so the explained-unexplained partition is not unique.
  • The unexplained component absorbs all omitted-variable bias and is routinely over-interpreted as a pure measure of discrimination.
  • It is a descriptive decomposition of means, not a causal estimator, and rests on a (often unstated) ceteris-paribus counterfactual.
  • Detailed unexplained contributions for categorical variables are sensitive to the omitted reference category, complicating per-variable interpretation.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

What exactly do the explained and unexplained components mean for a health gap?

The explained (composition) component is the part of the mean health gap that would remain if the two groups had identical returns to their characteristics but kept their actual differences in those characteristics, for example differing average income, age, smoking, or insurance. The unexplained (coefficient) component is the part that would remain if the groups had identical covariate distributions but kept their different coefficients, meaning the same exposure maps to different health. In health research the unexplained part is often interpreted as reflecting discrimination, chronic stress, or unmeasured exposures, but because it also absorbs omitted-variable bias it should not be read as a clean causal quantity.

Why does the choice of reference coefficients matter, and which should I use?

The decomposition requires a reference coefficient vector against which to measure both composition and returns, and there is no uniquely correct choice, this is the classic index-number problem. Using group A's coefficients, group B's coefficients, or a pooled estimate shifts how much of the gap is labeled explained versus unexplained. A common modern practice is to use coefficients from a pooled model that includes a group indicator, but the right answer is to report results under more than one reference and treat the partition as a range rather than a single number. Transparency about the weighting is essential for honest interpretation.

Can I use Oaxaca-Blinder when my health outcome is binary, like having diabetes or not?

Not with the plain linear identity, because for nonlinear models the mean outcome is not simply mean covariates times coefficients and the exact algebraic split no longer holds. Fairlie's 2005 extension handles this by working with average predicted probabilities from a logit or probit model: the explained part is the change in mean predicted probability when one group's covariate distribution is passed through the other group's coefficients, and the unexplained part is the change from swapping the coefficients. This preserves the explained-unexplained logic for binary and count health outcomes, though results can depend on the order of substitution and how observations are paired across groups.

Sources

  1. 1.
    Oaxaca, R. (1973). Male-Female Wage Differentials in Urban Labor Markets. International Economic Review, 14(3), 693-709.
  2. 2.
    Blinder, A. S. (1973). Wage Discrimination: Reduced Form and Structural Estimates. The Journal of Human Resources, 8(4), 436-455.
  3. 3.
    Fairlie, R. W. (2005). An Extension of the Blinder-Oaxaca Decomposition Technique to Logit and Probit Models. Journal of Economic and Social Measurement, 30(4), 305-316.
  4. 4.
    Rahimi, E., & Hashemi Nazari, S. S. (2021). A detailed explanation and graphical representation of the Blinder-Oaxaca decomposition method with its application in health inequalities. Emerging Themes in Epidemiology, 18, 12.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 23). Oaxaca-Blinder Health Decomposition. ScholarGate. https://scholargate.app/social-epidemiology/oaxaca-blinder-health-decomposition