Latent structureStatisticsPsychological scalingModel

Thurstone Scaling

Also known as: Law of Comparative Judgment, Thurstone's Method of Equal-Appearing Intervals, Case V Scaling, Thurstone Ölçekleme

OriginatorLouis Leon ThurstoneYear1927Sources1Related methods2

Thurstone Scaling, formally the Law of Comparative Judgment, is a psychometric model introduced by Louis Leon Thurstone in 1927 for deriving interval-level scale values from pairwise comparison data. By assuming that each stimulus evokes a normally distributed discriminal process on a psychological continuum, the method converts proportions of preference judgments into z-scores and recovers the latent positions of stimuli, enabling rigorous attitude and preference measurement.

Key highlights

  • Produces interval-level scale values from purely ordinal pairwise judgments, enabling parametric analysis
  • Grounded in a probabilistic latent-variable model with an explicit and testable fit criterion
  • Case V simplification is computationally straightforward and requires no iterative estimation
  • Historically foundational and widely cited, lending theoretical credibility to derived scales

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Thurstone Scaling is appropriate when stimuli need to be placed on an interval scale based on human judgments, the number of stimuli is small enough for exhaustive pairwise comparison (typically fewer than 30), and the researcher can assume that discriminal processes are approximately normally distributed. It suits attitude measurement, sensory evaluation, and preference research. Alternatives include the Bradley-Terry model for count-based paired comparisons, Likert scaling for larger item sets, and IRT models when individual-level measurement is required.

Strengths & limitations

Strengths
  • Produces interval-level scale values from purely ordinal pairwise judgments, enabling parametric analysis
  • Grounded in a probabilistic latent-variable model with an explicit and testable fit criterion
  • Case V simplification is computationally straightforward and requires no iterative estimation
  • Historically foundational and widely cited, lending theoretical credibility to derived scales
Limitations
  • Number of required judgments grows as O(n²) in stimuli, making large sets impractical without incomplete designs
  • Assumes normally distributed discriminal processes; violations of this assumption are difficult to detect with limited data
  • Case V assumption of equal dispersions and zero correlations is often untestable and may be unrealistic
  • Produces relative (not absolute) scale values; the zero point and unit are arbitrary and must be fixed externally

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

How many judges are needed for reliable Thurstone scale values?

Thurstone's original studies used panels of around 100 to several hundred judges. Smaller panels increase sampling error in the proportion estimates. As a practical rule, at least 50 to 100 judges per stimulus pair is recommended to produce stable z-score estimates, though bootstrap procedures can help quantify uncertainty when fewer judges are available.

What is the difference between Case I and Case V?

Case I is the fully general model that estimates separate discriminal dispersions and inter-stimulus correlations for each pair — requiring many parameters and specialized fitting routines. Case V imposes equality of dispersions and zero correlations, reducing estimation to a simple column-mean calculation. Case V is used in the vast majority of applications because it is tractable and usually provides an adequate fit.

Can Thurstone Scaling handle missing pairs in an incomplete design?

Yes. When not all pairs are judged, scale values can still be estimated from the available comparisons using least-squares or maximum-likelihood procedures, provided the comparison graph is connected (every stimulus can be linked to every other through a chain of observed pairs). Software implementations typically handle incomplete designs through matrix methods or iterative fitting.

Sources

  1. 1.
    Thurstone, L. L. (1927). A law of comparative judgment. Psychological Review, 34(4), 273–286.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 2). Thurstone Scaling. ScholarGate. https://scholargate.app/statistics/thurstone-scaling