Graded Response Model (GRM)
Graded Response Model · Also known as: Samejima's GRM, Derecelendirilmiş Tepki Modeli (GRM), graded IRT model
The Graded Response Model is an item response theory model developed by Fumiko Samejima in 1969 for ordered polytomous items such as Likert-type scales. It estimates both the discriminating power of each item and a set of threshold parameters marking the boundaries between adjacent response categories, while simultaneously placing persons on a continuous latent trait scale.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
+2 more
When to use it
The GRM is appropriate when items carry ordered response categories — most commonly Likert-type scales used in attitude measurement, personality inventories, or clinical rating forms — and when the goal is to obtain item parameters that are invariant across samples (given the same trait) and person scores that account for differential item discriminability. Several conditions must hold: items must form a unidimensional scale (checked with EFA or a fit-based unidimensionality test before GRM calibration), local independence must be satisfied, the ordered threshold assumption must hold (monotonically increasing boundaries), and there must be observations in every response category. The minimum recommended sample size is around 200, though items with many categories and low discrimination may require more.
Strengths & limitations
- Directly models the ordered structure of polytomous items rather than collapsing categories or treating ordinal data as interval.
- Estimates item discrimination alongside thresholds, allowing items that differ in how precisely they distinguish persons at each trait level.
- Item parameters are sample-invariant and person parameters are item-invariant (under the IRT assumptions), enabling score comparisons across groups or occasions.
- Provides category response function plots that make the psychometric behavior of each item visually transparent.
- Requires substantially larger samples than classical test theory — at least 200 respondents, and often more when items have many categories.
- Assumes strict unidimensionality and local independence; multidimensional constructs require more complex IRT models.
- Threshold ordering is an assumption, not a guarantee; out-of-order thresholds signal categories that respondents do not discriminate, requiring category collapse or model revision.
- Interpretation of item and person parameters demands familiarity with IRT logic, making results less immediately accessible to applied audiences than classical reliability statistics.
Frequently asked
How does the GRM differ from the 2PL model?
The 2PL model is designed for binary (right/wrong or yes/no) items and has one threshold per item. The GRM extends this to items with m ordered categories by assigning m−1 thresholds — one for each boundary between adjacent categories — while retaining a single discrimination parameter per item. Each boundary has its own logistic curve; the 2PL is the special case where m = 2.
What is the difference between the GRM and the Partial Credit Model?
Both handle ordered polytomous items, but they parameterize the response process differently. The GRM models cumulative probabilities — the chance of responding at or above each category — using a two-parameter logistic form with a common discrimination per item. The PCM (and its generalization, the GPCM) models step probabilities — the chance of choosing category k over k−1 — and is more natural for sequentially scored tasks such as essay rubrics or performance assessments.
What happens if thresholds are not in order?
Out-of-order thresholds, sometimes called threshold reversals or disordering, mean that two adjacent categories are not being reliably distinguished by respondents. The standard remedy is to collapse the disordered categories into a single category and re-run the calibration. Persistent disordering may indicate that the response format has too many categories for the construct being measured.
How large a sample do I need for GRM calibration?
A minimum of around 200 respondents is commonly cited as a practical floor, but the required sample grows with the number of items, the number of response categories, and low item discriminations. Items with five or more categories and moderate discrimination may need 300 to 500 respondents for stable parameter recovery. Simulation studies or pilot calibrations can clarify requirements for a specific instrument.
Sources
- Samejima, F. (1969). Estimation of Latent Ability Using a Response Pattern of Graded Scores. Psychometrika Monograph Supplement, No. 17. link ↗
- Embretson, S. E. & Reise, S. P. (2000). Item Response Theory for Psychologists. Lawrence Erlbaum Associates. ISBN: 978-0805828191
How to cite this page
ScholarGate. (2026, June 1). Graded Response Model. ScholarGate. https://scholargate.app/en/psychometrics/graded-response-model
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- 2PL IRTPsychometrics↔ compare
- 3PL IRTPsychometrics↔ compare
- CFAStatistics↔ compare
- EFAStatistics↔ compare
- LCAStatistics↔ compare
- PCM / GPCMPsychometrics↔ compare
- Rasch ModelPsychometrics↔ compare