Ordinal Item Response Theory
Also known as: polytomous IRT, ordinal IRT models, graded response models, ordinal latent trait models
Ordinal item response theory (ordinal IRT) comprises a family of probabilistic models — most notably the Graded Response Model and the Partial Credit Model — that relate a respondent's standing on a latent trait to the probability of choosing each ordered response category on a polytomous item. It extends classical IRT beyond dichotomous items to the Likert-type and rating-scale items that dominate psychometric measurement.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use ordinal IRT when your items have three or more ordered response categories — Likert scales, rating scales, partial-credit tasks — and you need precise, sample-independent estimates of item difficulty thresholds and item discrimination, or respondent ability scores with calibrated standard errors. It is especially valuable for scale development and refinement, computerised adaptive testing, and equating studies where understanding the measurement precision at each trait level matters. Do not use ordinal IRT when the sample is small (fewer than about 200–250 per group for the GRM); when unidimensionality is not plausible and you have not confirmed it with factor analysis; or when you simply need a quick scale score and classical reliability indices suffice.
Strengths & limitations
- Provides item-level diagnostic information — discrimination, multiple threshold locations — that classical item analysis cannot yield.
- Person ability estimates come with individualised standard errors, enabling precision-weighted scoring.
- Item and person parameters are theoretically sample-independent (invariant) once the model is correctly specified.
- Directly supports computerised adaptive testing, equating, and differential item functioning analyses.
- Category probability curves reveal whether adjacent response options are empirically distinguishable.
- More appropriate for Likert and partial-credit data than dichotomous IRT models, which collapse ordinal information.
- Requires substantially larger samples than classical methods; the GRM typically needs 250 or more respondents for stable parameter estimation.
- The unidimensionality assumption must hold reasonably well; violations bias parameters and ability estimates.
- Model selection among competing ordinal IRT models (GRM, PCM, RSM) requires fit comparisons and theoretical justification.
- Software output and interpretation of threshold parameters are less intuitive than classical item statistics for applied researchers.
- Threshold non-ordering (a higher threshold estimated below a lower one) indicates item problems that require recoding or item revision.
Frequently asked
What is the difference between the Graded Response Model and the Partial Credit Model?
Both handle ordered polytomous items but differ in parameterisation. The GRM uses cumulative logistic functions, estimating the probability of responding at or above each category; discrimination can vary across items. The PCM models adjacent-category transitions and constrains discrimination to be equal across categories within an item but allows it to vary across items. The Rating Scale Model further constrains thresholds to be identical across all items. Choice depends on whether your items share a common response format and your theoretical stance.
How large a sample do I need for ordinal IRT?
A commonly cited minimum for the GRM is about 200–250 respondents for stable threshold estimation when items have 4–5 categories and moderate discrimination. More complex models, more categories, or lower discrimination require larger samples. Simulation studies by Reise and Yu (1990) and others provide specific guidance for particular designs.
How do I choose between ordinal IRT and ordinal CFA?
If your goal is to obtain person ability estimates, item diagnostic parameters, or to support adaptive testing and equating, ordinal IRT is more appropriate. If you want to test a theorised factor structure, compare nested models, or compute factor scores within a structural equation context, ordinal CFA is better suited. Many scale development studies use both: EFA and CFA to establish structure, then IRT for item-level diagnostics.
What does a disordered threshold mean?
A disordered threshold (b_ik > b_{i,k+1} empirically) means that respondents skip an intermediate category, so the categories do not function as ordered steps on the trait continuum. This often indicates that adjacent categories are redundant or poorly labelled. Solutions include collapsing categories, revising labels, or removing the item.
Can ordinal IRT handle multidimensional constructs?
Standard ordinal IRT models assume unidimensionality. Multidimensional extensions (MIRT) exist and accommodate correlated or bifactor structures, but they require larger samples and more complex software. If your scale has a clear multidimensional structure, confirm it with CFA before applying MIRT.
Sources
- Samejima, F. (1969). Estimation of latent ability using a response pattern of graded scores. Psychometrika Monograph Supplement, 34(4, Pt. 2), 1–97. link ↗
- Embretson, S. E. & Reise, S. P. (2000). Item Response Theory for Psychologists. Lawrence Erlbaum Associates. ISBN: 978-0805828191
How to cite this page
ScholarGate. (2026, June 3). Ordinal Item Response Theory. ScholarGate. https://scholargate.app/en/psychometrics/ordinal-item-response-theory
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Confirmatory factor analysisPsychometrics↔ compare
- Differential Item FunctioningPsychometrics↔ compare
- GRMPsychometrics↔ compare
- Item Response TheoryPsychometrics↔ compare
- Ordinal CFAPsychometrics↔ compare
- PCM / GPCMPsychometrics↔ compare