Process / pipelineEconomicsDistributional decompositionPipeline

Shapley Decomposition of Inequality

Also known as: Shapley Decomposition, Shorrocks Shapley Decomposition, Factor Decomposition of Inequality, Shapley Value Distributional Decomposition

OriginatorAnthony Shorrocks (working paper 1999; published 2013)Year2013Sources1Related methods7

The Shapley decomposition, formalized for distributional analysis by Anthony Shorrocks (in a widely circulated 1999 working paper, published in 2013), is a general procedure for attributing an inequality or poverty statistic to its contributing factors — income sources, population subgroups, or determinants. It borrows the Shapley value from cooperative game theory: each factor's contribution is its average marginal effect on the indicator across all possible orders in which factors could be eliminated. The result is an exact, symmetric, residual-free decomposition that applies to any indicator, even those (like the Gini) that have no natural analytic decomposition of their own.

Key highlights

  • Applies to any distributional indicator, including ones like the Gini that have no natural analytic decomposition.
  • Exact and residual-free by the efficiency property, so factor contributions always sum to the total with no leftover interaction term.
  • Symmetric and order-independent: it resolves the path-dependence that plagues naive sequential decompositions.
  • Unifies many previously ad hoc decompositions (income-source, subgroup, growth-redistribution) within one axiomatic framework.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use the Shapley decomposition when you want to attribute a distributional statistic to several factors that interact, and especially when the indicator has no natural built-in decomposition. Typical uses are decomposing inequality by income source, attributing a poverty or inequality change to determinants (growth, redistribution, prices, demographics), and apportioning inequality to covariates via a regression-based first stage. It is the method of choice for the Gini coefficient, which is not additively decomposable, and for unifying many ad hoc decompositions under one principled framework. The cost is computational: an exact Shapley decomposition requires evaluating the indicator over all 2^K factor subsets, so with many factors one uses sampling approximations. Define the reference state and elimination convention carefully, since they determine what the contributions mean.

Strengths & limitations

Strengths
  • Applies to any distributional indicator, including ones like the Gini that have no natural analytic decomposition.
  • Exact and residual-free by the efficiency property, so factor contributions always sum to the total with no leftover interaction term.
  • Symmetric and order-independent: it resolves the path-dependence that plagues naive sequential decompositions.
  • Unifies many previously ad hoc decompositions (income-source, subgroup, growth-redistribution) within one axiomatic framework.
Limitations
  • Computationally expensive: an exact decomposition evaluates the indicator over all 2^K factor subsets, so large K requires Monte Carlo approximation.
  • Results depend on the chosen reference state and on how a factor is 'eliminated' (zeroed, set to its mean, etc.), which is a modeling decision.
  • It is an accounting attribution, not a causal decomposition; contributions reflect correlations and the chosen functional form, not structural effects.
  • When factors are defined via a first-stage regression, the decomposition inherits any misspecification or omitted-variable bias of that model.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

How does the Shapley decomposition differ from the Theil within/between decomposition?

The Theil decomposition is an exact analytic split that exists only because the generalized-entropy class is additively decomposable; it works for Theil and the mean log deviation but not for the Gini. The Shapley decomposition is a general procedure that works for any indicator by averaging marginal contributions over all elimination orders. For indicators that already decompose analytically the Shapley result usually coincides with the natural decomposition, but Shapley extends to indices with no natural split, which is its main advantage.

Why is the order of eliminating factors a problem, and how does Shapley solve it?

Because factors interact, the marginal effect of removing one factor depends on which others have already been removed, so different elimination orders give different contributions and a leftover interaction term. The Shapley value removes this arbitrariness by averaging each factor's marginal contribution over all possible orders, weighting each order equally. The averaging makes the result symmetric across factors and, by the efficiency axiom, guarantees the contributions sum exactly to the total with no residual.

Is the Shapley decomposition computationally feasible for many factors?

An exact decomposition requires evaluating the indicator over all 2^K subsets of factors, which is fine for a handful of factors but explodes combinatorially. With many factors, practitioners use Monte Carlo sampling of random permutations to approximate the Shapley contributions, trading exactness for tractability. The approximation error shrinks with the number of sampled permutations, and software implementations typically expose the number of replications as a tuning parameter.

Sources

  1. 1.
    Shorrocks, A. F. (2013). Decomposition procedures for distributional analysis: a unified framework based on the Shapley value. Journal of Economic Inequality, 11(1), 99–126.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 22). Shapley Decomposition of Inequality. ScholarGate. https://scholargate.app/economics/shapley-decomposition-inequality