Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Psychometrics›Polytomous Rasch Model
Latent structureScale / measurement

Polytomous Rasch Model

Also known as: PRM, Rating Scale Model, Partial Credit Model, Polytomous IRT Rasch

The Polytomous Rasch Model extends the dichotomous Rasch framework to ordered response scales with three or more categories, such as Likert items or partial-credit tasks. It estimates person ability and item difficulty on the same interval-level logit scale, and it tests whether the response categories function as intended — prerequisites for rigorous ordinal measurement.

ScholarGate
  1. Latent structure
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Polytomous Rasch Model
Confirmatory factor anal…Differential Item Functi…Item Response TheoryMeasurement InvarianceOrdinal IRTRasch ModelPolytomous Construct Val…

When to use it

Use the Polytomous Rasch Model when your items have ordered response categories and your measurement goal is to place persons and items on a common interval scale with empirical tests of category functioning. It is especially appropriate for Likert-type attitude scales, partial-credit academic tests, and clinical rating instruments where verifying that response categories work as intended is critical. It is also the natural choice when you need to evaluate and justify the number of response categories, detect disordered thresholds, or establish measurement invariance across groups. Do NOT use it when sample size is very small (fewer than roughly 150–200 cases for stable threshold estimates), when response categories are truly nominal rather than ordered, or when item content varies so widely that a single latent dimension is implausible — in those cases, a multidimensional or mixture IRT model is more appropriate.

Strengths & limitations

Strengths
  • Places persons and items on a single interval-level logit scale, enabling meaningful comparisons of location regardless of which specific items were administered.
  • Provides explicit diagnostics for category functioning — disordered thresholds are detectable and actionable before finalising the scale.
  • Specific objectivity: if the model fits, person ability estimates are independent of the particular items used and item difficulty estimates are independent of the particular persons tested.
  • Supports differential item functioning (DIF) analysis, measurement invariance testing, and linking of multiple test forms on a common metric.
  • Both Rating Scale and Partial Credit variants are available, offering a choice between parsimony and flexibility.
Limitations
  • Assumes strict unidimensionality — all items must reflect a single latent trait; violations inflate misfit statistics and distort person estimates.
  • Requires larger samples than dichotomous Rasch models because each additional response category adds threshold parameters to estimate.
  • Strict Rasch fit criteria may be difficult to meet with real-world attitude scales, potentially requiring item revision or category collapsing before achieving acceptable fit.
  • Does not model item discrimination (unlike 2PL or graded response models), which may be a poor assumption when items vary substantially in how sharply they differentiate among persons.

Frequently asked

What is the difference between the Rating Scale Model and the Partial Credit Model?

Both are polytomous Rasch models. The Rating Scale Model (Andrich, 1978) assumes all items share the same threshold spacing — only the overall item difficulty varies. The Partial Credit Model (Masters, 1982) allows each item to have its own set of thresholds, which is more flexible but requires more data. If all your items use identical response options and you can justify uniform category spacing, RSM is more parsimonious; otherwise PCM is safer.

What are disordered thresholds and why do they matter?

A threshold is disordered when the estimated difficulty of moving to category k is lower than the difficulty of moving to category k-1, meaning a higher category is actually easier to reach — contradicting the intended ordinal structure. This usually signals that respondents cannot reliably distinguish between adjacent categories. The fix is to collapse adjacent categories or revise the response format before the instrument is used.

How large a sample do I need?

Rules of thumb vary. For the Rating Scale Model with a small number of thresholds, around 150–250 cases is often cited for stable item parameter estimates. The Partial Credit Model, having more parameters, typically needs 250 or more. Simulation studies suggest that the precision of threshold estimates is sensitive to sample size, especially for extreme categories that are rarely chosen.

Can I use the Polytomous Rasch Model for DIF analysis?

Yes. DIF in polytomous Rasch models can be investigated by comparing item parameters estimated separately in focal and reference groups, or by testing group-by-item interactions in a combined model. Some thresholds may show DIF even when the overall item location does not — this is called non-uniform or threshold-level DIF and has practical implications for scale fairness.

How does the Polytomous Rasch Model compare to the Graded Response Model?

The Graded Response Model (Samejima, 1969) also handles ordered categories but includes an item discrimination parameter, allowing items to differ in how sharply they distinguish among persons. The Rasch model fixes discrimination to 1 for all items, which is a testable constraint. If discrimination varies substantially across your items, the Graded Response Model will fit better; if you need the measurement properties (specific objectivity, person-free item calibration) that come with Rasch fit, invest effort in revising items until the equal-discrimination constraint holds.

Sources

  1. Masters, G. N. (1982). A Rasch model for partial credit scoring. Psychometrika, 47(2), 149–174. DOI: 10.1007/BF02296272 ↗
  2. Andrich, D. (1978). A rating formulation for ordered response categories. Psychometrika, 43(4), 561–573. DOI: 10.1007/BF02293814 ↗

How to cite this page

ScholarGate. (2026, June 3). Polytomous Rasch Model. ScholarGate. https://scholargate.app/en/psychometrics/polytomous-rasch-model

Related methods

Confirmatory factor analysisDifferential Item FunctioningItem Response TheoryMeasurement InvarianceOrdinal IRTRasch Model

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Confirmatory factor analysisPsychometrics↔ compare
  • Differential Item FunctioningPsychometrics↔ compare
  • Item Response TheoryPsychometrics↔ compare
  • Measurement InvariancePsychometrics↔ compare
  • Ordinal IRTPsychometrics↔ compare
  • Rasch ModelPsychometrics↔ compare
Compare side by side →

Referenced by

Polytomous Construct Validity

Similar methods

Ordinal Rasch ModelPolytomous item analysisPolytomous scale developmentPCM / GPCMOrdinal IRTRasch ModelPolytomous Construct ValidityGRM

Related reference concepts

Item Response TheoryStructural and Latent Variable ModelsLatent Class AnalysisRating ScalesMeasurementPatient-Reported Outcome Measures

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Polytomous Rasch Model (Polytomous Rasch Model). Retrieved 2026-07-21 from https://scholargate.app/en/psychometrics/polytomous-rasch-model · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Gerhard N. Masters (Partial Credit Model); David Andrich (Rating Scale Model)
Year
1978–1982
Type
Item response model
DataType
Ordered polytomous response categories (Likert, rating scales, partial-credit scored items)
Subfamily
Scale / measurement
Related methods
Confirmatory factor analysisDifferential Item FunctioningItem Response TheoryMeasurement InvarianceOrdinal IRTRasch Model
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account