Regression modelDecision-makingRanking modelsModel

Bradley-Terry Model

Also known as: BT Model, Bradley-Terry-Luce Model, Paired Comparison Model, İkili Karşılaştırma Modeli

OriginatorRalph Bradley & Milton TerryYear1952Sources1Related methods9

The Bradley-Terry model is a probabilistic model for paired comparisons that assigns a latent strength parameter to each item and predicts the probability that one item beats another in a head-to-head contest. Introduced by Ralph A. Bradley and Milton E. Terry in 1952, it provides a principled statistical framework for ranking items from pairwise preference data, including incomplete comparison designs where not every pair is directly observed.

Key highlights

  • Provides a principled probabilistic foundation for ranking from pairwise data, yielding interpretable strength estimates with standard errors.
  • Handles incomplete comparison designs gracefully, requiring only that the comparison graph be connected rather than fully balanced.
  • Scales to large item sets through efficient maximum likelihood algorithms and has well-developed extensions for ties, home advantage, and covariates.
  • Produces a unique global ranking consistent with the law of comparative judgment, avoiding the cycles and contradictions of naive win-rate rankings.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use the Bradley-Terry model when your data consist of pairwise comparisons or head-to-head contest outcomes and your goal is to rank items or estimate relative preference strengths. It is appropriate for tournaments, consumer preference studies, sports analytics, and crowdsourced ranking tasks. Key assumptions include transitivity of preferences, independence of comparisons, and no ties (or tie-handling extensions). Avoid it when comparisons are highly intransitive, when context effects invalidate independence, or when ordinal rankings without a probabilistic model suffice.

Strengths & limitations

Strengths
  • Provides a principled probabilistic foundation for ranking from pairwise data, yielding interpretable strength estimates with standard errors.
  • Handles incomplete comparison designs gracefully, requiring only that the comparison graph be connected rather than fully balanced.
  • Scales to large item sets through efficient maximum likelihood algorithms and has well-developed extensions for ties, home advantage, and covariates.
  • Produces a unique global ranking consistent with the law of comparative judgment, avoiding the cycles and contradictions of naive win-rate rankings.
Limitations
  • Assumes strict transitivity of preferences, which may not hold in real-world scenarios involving intransitive cycles or context-dependent choices.
  • Does not accommodate ties natively in its basic formulation; extensions such as the Davidson model are required for tied outcomes.
  • Requires a connected comparison graph for identifiability; isolated items or disconnected groups cannot be jointly ranked.
  • Maximum likelihood estimates are undefined when some items win or lose all comparisons, necessitating regularization or Bayesian priors.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

How does the Bradley-Terry model differ from the Elo rating system?

Both models are grounded in the same paired comparison probability formula, but they differ in estimation approach. Elo updates ratings incrementally after each match using a fixed K-factor, making it well suited for online or sequential settings. Bradley-Terry uses batch maximum likelihood estimation over all observed outcomes simultaneously, yielding statistically efficient estimates with uncertainty quantification but requiring all data to be available at once.

Can the Bradley-Terry model handle more than two items in a single comparison?

The basic Bradley-Terry model is defined for pairwise comparisons only. Its generalization to rankings of three or more items simultaneously is the Plackett-Luce model, which decomposes a full ranking into a sequence of paired choices. If your data consist of complete or partial orderings rather than binary win-loss outcomes, the Plackett-Luce model is the appropriate extension.

What sample size is needed for reliable Bradley-Terry estimates?

Reliability depends on the number of comparisons per pair rather than on total item count alone. As a rule of thumb, at least five to ten comparisons per pair produce reasonably stable estimates. With fewer comparisons the likelihood surface becomes flat, yielding wide confidence intervals. Bayesian regularization with weakly informative priors on the strength parameters can improve stability when comparison data are sparse.

Sources

  1. 1.
    Bradley, R. A., & Terry, M. E. (1952). Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika, 39(3/4), 324–345.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 2). Bradley-Terry Model. ScholarGate. https://scholargate.app/decision-making/bradley-terry-model