Latent structurePsychometricsScale / measurementModel

Polytomous Rasch Model

Also known as: PRM, Rating Scale Model, Partial Credit Model, Polytomous IRT Rasch

OriginatorGerhard N. Masters (Partial Credit Model); David Andrich (Rating Scale Model)Year1978–1982Sources2Related methods7

The Polytomous Rasch Model extends the dichotomous Rasch framework to ordered response scales with three or more categories, such as Likert items or partial-credit tasks. It estimates person ability and item difficulty on the same interval-level logit scale, and it tests whether the response categories function as intended — prerequisites for rigorous ordinal measurement.

Key highlights

  • Places persons and items on a single interval-level logit scale, enabling meaningful comparisons of location regardless of which specific items were administered.
  • Provides explicit diagnostics for category functioning — disordered thresholds are detectable and actionable before finalising the scale.
  • Specific objectivity: if the model fits, person ability estimates are independent of the particular items used and item difficulty estimates are independent of the particular persons tested.
  • Supports differential item functioning (DIF) analysis, measurement invariance testing, and linking of multiple test forms on a common metric.
  • Both Rating Scale and Partial Credit variants are available, offering a choice between parsimony and flexibility.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use the Polytomous Rasch Model when your items have ordered response categories and your measurement goal is to place persons and items on a common interval scale with empirical tests of category functioning. It is especially appropriate for Likert-type attitude scales, partial-credit academic tests, and clinical rating instruments where verifying that response categories work as intended is critical. It is also the natural choice when you need to evaluate and justify the number of response categories, detect disordered thresholds, or establish measurement invariance across groups. Do NOT use it when sample size is very small (fewer than roughly 150–200 cases for stable threshold estimates), when response categories are truly nominal rather than ordered, or when item content varies so widely that a single latent dimension is implausible — in those cases, a multidimensional or mixture IRT model is more appropriate.

Strengths & limitations

Strengths
  • Places persons and items on a single interval-level logit scale, enabling meaningful comparisons of location regardless of which specific items were administered.
  • Provides explicit diagnostics for category functioning — disordered thresholds are detectable and actionable before finalising the scale.
  • Specific objectivity: if the model fits, person ability estimates are independent of the particular items used and item difficulty estimates are independent of the particular persons tested.
  • Supports differential item functioning (DIF) analysis, measurement invariance testing, and linking of multiple test forms on a common metric.
  • Both Rating Scale and Partial Credit variants are available, offering a choice between parsimony and flexibility.
Limitations
  • Assumes strict unidimensionality — all items must reflect a single latent trait; violations inflate misfit statistics and distort person estimates.
  • Requires larger samples than dichotomous Rasch models because each additional response category adds threshold parameters to estimate.
  • Strict Rasch fit criteria may be difficult to meet with real-world attitude scales, potentially requiring item revision or category collapsing before achieving acceptable fit.
  • Does not model item discrimination (unlike 2PL or graded response models), which may be a poor assumption when items vary substantially in how sharply they differentiate among persons.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

What is the difference between the Rating Scale Model and the Partial Credit Model?

Both are polytomous Rasch models. The Rating Scale Model (Andrich, 1978) assumes all items share the same threshold spacing — only the overall item difficulty varies. The Partial Credit Model (Masters, 1982) allows each item to have its own set of thresholds, which is more flexible but requires more data. If all your items use identical response options and you can justify uniform category spacing, RSM is more parsimonious; otherwise PCM is safer.

What are disordered thresholds and why do they matter?

A threshold is disordered when the estimated difficulty of moving to category k is lower than the difficulty of moving to category k-1, meaning a higher category is actually easier to reach — contradicting the intended ordinal structure. This usually signals that respondents cannot reliably distinguish between adjacent categories. The fix is to collapse adjacent categories or revise the response format before the instrument is used.

How large a sample do I need?

Rules of thumb vary. For the Rating Scale Model with a small number of thresholds, around 150–250 cases is often cited for stable item parameter estimates. The Partial Credit Model, having more parameters, typically needs 250 or more. Simulation studies suggest that the precision of threshold estimates is sensitive to sample size, especially for extreme categories that are rarely chosen.

Can I use the Polytomous Rasch Model for DIF analysis?

Yes. DIF in polytomous Rasch models can be investigated by comparing item parameters estimated separately in focal and reference groups, or by testing group-by-item interactions in a combined model. Some thresholds may show DIF even when the overall item location does not — this is called non-uniform or threshold-level DIF and has practical implications for scale fairness.

How does the Polytomous Rasch Model compare to the Graded Response Model?

The Graded Response Model (Samejima, 1969) also handles ordered categories but includes an item discrimination parameter, allowing items to differ in how sharply they distinguish among persons. The Rasch model fixes discrimination to 1 for all items, which is a testable constraint. If discrimination varies substantially across your items, the Graded Response Model will fit better; if you need the measurement properties (specific objectivity, person-free item calibration) that come with Rasch fit, invest effort in revising items until the equal-discrimination constraint holds.

Sources

  1. 1.
    Masters, G. N. (1982). A Rasch model for partial credit scoring. Psychometrika, 47(2), 149–174.
  2. 2.
    Andrich, D. (1978). A rating formulation for ordered response categories. Psychometrika, 43(4), 561–573.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 3). Polytomous Rasch Model. ScholarGate. https://scholargate.app/psychometrics/polytomous-rasch-model

Polytomous Rasch Model | ScholarGate