Latent structurePsychometricsModel

Graded Response Model (GRM)

Also known as: Samejima's GRM, Derecelendirilmiş Tepki Modeli (GRM), graded IRT model

OriginatorFumiko SamejimaYear1969Sources2Related methods15

The Graded Response Model is an item response theory model developed by Fumiko Samejima in 1969 for ordered polytomous items such as Likert-type scales. It estimates both the discriminating power of each item and a set of threshold parameters marking the boundaries between adjacent response categories, while simultaneously placing persons on a continuous latent trait scale.

Key highlights

  • Directly models the ordered structure of polytomous items rather than collapsing categories or treating ordinal data as interval.
  • Estimates item discrimination alongside thresholds, allowing items that differ in how precisely they distinguish persons at each trait level.
  • Item parameters are sample-invariant and person parameters are item-invariant (under the IRT assumptions), enabling score comparisons across groups or occasions.
  • Provides category response function plots that make the psychometric behavior of each item visually transparent.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

The GRM is appropriate when items carry ordered response categories — most commonly Likert-type scales used in attitude measurement, personality inventories, or clinical rating forms — and when the goal is to obtain item parameters that are invariant across samples (given the same trait) and person scores that account for differential item discriminability. Several conditions must hold: items must form a unidimensional scale (checked with EFA or a fit-based unidimensionality test before GRM calibration), local independence must be satisfied, the ordered threshold assumption must hold (monotonically increasing boundaries), and there must be observations in every response category. The minimum recommended sample size is around 200, though items with many categories and low discrimination may require more.

Strengths & limitations

Strengths
  • Directly models the ordered structure of polytomous items rather than collapsing categories or treating ordinal data as interval.
  • Estimates item discrimination alongside thresholds, allowing items that differ in how precisely they distinguish persons at each trait level.
  • Item parameters are sample-invariant and person parameters are item-invariant (under the IRT assumptions), enabling score comparisons across groups or occasions.
  • Provides category response function plots that make the psychometric behavior of each item visually transparent.
Limitations
  • Requires substantially larger samples than classical test theory — at least 200 respondents, and often more when items have many categories.
  • Assumes strict unidimensionality and local independence; multidimensional constructs require more complex IRT models.
  • Threshold ordering is an assumption, not a guarantee; out-of-order thresholds signal categories that respondents do not discriminate, requiring category collapse or model revision.
  • Interpretation of item and person parameters demands familiarity with IRT logic, making results less immediately accessible to applied audiences than classical reliability statistics.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

How does the GRM differ from the 2PL model?

The 2PL model is designed for binary (right/wrong or yes/no) items and has one threshold per item. The GRM extends this to items with m ordered categories by assigning m−1 thresholds — one for each boundary between adjacent categories — while retaining a single discrimination parameter per item. Each boundary has its own logistic curve; the 2PL is the special case where m = 2.

What is the difference between the GRM and the Partial Credit Model?

Both handle ordered polytomous items, but they parameterize the response process differently. The GRM models cumulative probabilities — the chance of responding at or above each category — using a two-parameter logistic form with a common discrimination per item. The PCM (and its generalization, the GPCM) models step probabilities — the chance of choosing category k over k−1 — and is more natural for sequentially scored tasks such as essay rubrics or performance assessments.

What happens if thresholds are not in order?

Out-of-order thresholds, sometimes called threshold reversals or disordering, mean that two adjacent categories are not being reliably distinguished by respondents. The standard remedy is to collapse the disordered categories into a single category and re-run the calibration. Persistent disordering may indicate that the response format has too many categories for the construct being measured.

How large a sample do I need for GRM calibration?

A minimum of around 200 respondents is commonly cited as a practical floor, but the required sample grows with the number of items, the number of response categories, and low item discriminations. Items with five or more categories and moderate discrimination may need 300 to 500 respondents for stable parameter recovery. Simulation studies or pilot calibrations can clarify requirements for a specific instrument.

Sources

  1. 1.
    Samejima, F. (1969). Estimation of Latent Ability Using a Response Pattern of Graded Scores. Psychometrika Monograph Supplement, No. 17.
  2. 2.
    Embretson, S. E. & Reise, S. P. (2000). Item Response Theory for Psychologists. Lawrence Erlbaum Associates.
    ISBN 978-0805828191

You have read it. What now?

Cite this page

ScholarGate. (2026, June 1). GRM. ScholarGate. https://scholargate.app/psychometrics/graded-response-model

Graded Response Model (GRM) — Graded Response Model