Multilevel Item Response Theory
Also known as: Multilevel IRT, MLIRT, Hierarchical IRT, Explanatory Item Response Models
Multilevel item response theory (MLIRT) joins two powerful frameworks: an IRT measurement model that turns item responses into a latent ability, and a multilevel structural model that explains how that ability varies across nested groups such as classrooms, schools, or countries. Instead of first scoring a test and then running a multilevel regression on the scores, MLIRT does both at once, so that measurement error in ability is properly carried into the group-level analysis. It is the rigorous way to study how student and school characteristics relate to a latent trait measured by a test.
Key highlights
- Properly propagates measurement error from the item level into the multilevel analysis, avoiding two-step bias.
- Simultaneously calibrates items and models how the latent trait varies within and between groups.
- Yields more accurate between-group variance and standard errors than scoring then regressing.
- Unifies measurement and explanation in one framework, supporting rich covariate effects at each level.
Intuition
This section is available to Pro members. Upgrade to Pro
How it works
This section is available to Pro members. Upgrade to Pro
When to use it
Use multilevel IRT when you measure a latent trait with item responses and want to model how it varies across clustered groups — analyzing international assessments (PISA, TIMSS) with students nested in schools and countries, studying school or classroom effects on a tested construct, or building explanatory item response models with covariates at multiple levels. It is preferred over the two-step approach whenever measurement error would otherwise distort group-level estimates, particularly between-group variance. It is computationally intensive, requires item-level (not just scored) data, and like other multilevel models needs enough higher-level units to estimate variance components.
Strengths & limitations
- Properly propagates measurement error from the item level into the multilevel analysis, avoiding two-step bias.
- Simultaneously calibrates items and models how the latent trait varies within and between groups.
- Yields more accurate between-group variance and standard errors than scoring then regressing.
- Unifies measurement and explanation in one framework, supporting rich covariate effects at each level.
- Computationally demanding, especially with many items, levels, or a fully Bayesian treatment.
- Requires item-response data and correct specification of both the IRT and the multilevel parts.
- Estimation of between-group variance still needs adequate numbers of higher-level units.
- Greater complexity raises the risk of misspecification and convergence problems than simpler approaches.
Common pitfalls
This section is available to Pro members. Upgrade to Pro
Applications
This section is available to Pro members. Upgrade to Pro
Frequently asked
Why not just score the test and then run a multilevel model on the scores?
Because test scores are estimates of a latent ability and carry measurement error. Treating them as exact in a subsequent multilevel regression attenuates relationships and, importantly, distorts the estimated between-group variance and its standard error. Multilevel IRT keeps ability latent and estimates the measurement and multilevel models jointly, so the uncertainty in ability is propagated correctly. The two-step approach is convenient but biased, especially for the between-school variance that is often the quantity of interest.
How does multilevel IRT relate to educational hierarchical linear modeling?
Educational HLM models an observed outcome (often a test score) nested in schools. Multilevel IRT replaces that observed outcome with a latent trait measured by item responses, embedding the IRT measurement model inside the same multilevel structure. In effect, MLIRT is HLM where the dependent variable is measured by an IRT model rather than taken as a known score, which is why it gives unbiased group-level estimates. See the related Educational Hierarchical Linear Modeling entry.
What are explanatory item response models?
Explanatory item response models, framed by De Boeck and Wilson, recast IRT as generalized linear mixed models in which item and person effects are random or fixed effects that can be explained by covariates. This perspective makes multilevel IRT a natural special case: person ability becomes a multilevel random effect, and item difficulties can depend on item properties. It unifies measurement and explanation and lets analysts fit many IRT and multilevel IRT models with standard mixed-model machinery.
Sources
- 1.Fox, J.-P. (2010). Bayesian Item Response Modeling: Theory and Applications. Springer.DOI 10.1007/978-1-4419-0742-4ISBN 9781441907417
- 2.De Boeck, P., & Wilson, M. (Eds.). (2004). Explanatory Item Response Models: A Generalized Linear and Nonlinear Approach. Springer.ISBN 9780387402758
You have read it. What now?
Cite this page
ScholarGate. (2026, June 22). Multilevel Item Response Theory. ScholarGate. https://scholargate.app/education/multilevel-irt