Item Response Theory (IRT)
Item Response Theory · Also known as: IRT, latent trait theory, item characteristic curve theory, modern test theory
Item response theory models the probability that a respondent answers an item correctly (or endorses it) as a function of the respondent's latent trait level and the item's own statistical properties — difficulty, discrimination, and guessing. Unlike classical test theory, IRT places persons and items on the same scale, yielding measurement that is sample-independent for items and test-independent for persons.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
+74 more
When to use it
Use IRT when you need measurement that is robust to the particular sample or item set used — for example when developing licensure exams, clinical assessments, or large-scale survey instruments that will be administered across different populations or in multiple forms. IRT is especially valuable when detecting differential item functioning across groups, when building adaptive tests, or when linking scores across test forms. Minimum sample size depends on the model: the Rasch model can work with 150–200 respondents for stable estimates; the 2PL typically requires 300–500; the 3PL needs 500 or more. Do not use IRT when your sample is very small (n < 100), when items are poorly written and unlikely to fit any model, or when classical test theory statistics (Cronbach's alpha, corrected item-total correlations) are sufficient for the purpose at hand.
Strengths & limitations
- Provides sample-independent item parameter estimates and test-independent person estimates, enabling equating across forms.
- Yields item-level diagnostic information — fit statistics and ICCs — that classical test theory cannot provide.
- Supports adaptive testing by identifying items that maximise information at each trait level.
- Enables formal detection of differential item functioning (DIF) across demographic or cultural groups.
- Flexibly handles binary, polytomous, and mixed item formats through model extensions.
- Requires substantially larger samples than classical test theory for stable parameter estimation, especially for the 2PL and 3PL.
- Model fit must be checked — items and persons that misfit should be investigated, which adds analytic complexity.
- Unidimensionality is a core assumption; multidimensional data violate it and require more complex multidimensional IRT models.
- Software (e.g., R packages mirt, ltm; Mplus; IRTPRO) is more technical than classical test theory tools and results are harder to communicate to non-specialist audiences.
Frequently asked
What is the difference between IRT and classical test theory?
Classical test theory (CTT) focuses on total scores and characterises items by their average difficulty (p-value) and item-total correlation, both of which depend on the particular sample tested. IRT models the probability of each response as a function of a latent trait, estimating item parameters that are theoretically sample-independent and person parameters that are test-independent. IRT therefore enables equating across tests and more detailed item diagnostics, but requires larger samples.
What is the difference between the Rasch model, 2PL, and 3PL?
These are nested special cases of the general logistic IRT model. The Rasch (1PL) model uses only an item difficulty parameter (b), constraining all items to have equal discrimination. The 2PL adds a discrimination parameter (a) per item, allowing items to differ in how sharply they differentiate trait levels. The 3PL further adds a guessing parameter (c), modelling the probability of a correct answer by chance — appropriate for multiple-choice ability tests but not for personality scales.
How large a sample do I need for IRT?
It depends on the model. The Rasch model can produce stable item parameter estimates with roughly 150–200 respondents when there are at least 15–20 items. The 2PL generally needs 300–500 cases; the 3PL needs 500 or more. Sparse data with many items and few respondents (or vice versa) produce unstable estimates. Simulation studies and pilot testing help determine adequacy for specific designs.
Does IRT require unidimensionality?
Standard IRT models (1PL, 2PL, 3PL, GRM) assume that a single dominant latent trait drives the responses. This assumption should be checked before fitting the model, for example by examining eigenvalue ratios in an EFA, testing for essential unidimensionality, or using MIRT fit indices. When data are clearly multidimensional, a multidimensional IRT (MIRT) model is more appropriate.
What software is commonly used for IRT?
In R, the packages mirt, ltm, and TAM are widely used and freely available. Commercial options include IRTPRO, flexMIRT, and WINSTEPS (for Rasch models). Mplus supports IRT within a broader SEM framework. StatWise offers guided IRT analysis with automatic model selection and fit reporting.
Sources
- Lord, F. M. & Novick, M. R. (1968). Statistical Theories of Mental Test Scores. Addison-Wesley. link ↗
- Embretson, S. E. & Reise, S. P. (2000). Item Response Theory for Psychologists. Lawrence Erlbaum Associates. ISBN: 978-0805828191
How to cite this page
ScholarGate. (2026, June 3). Item Response Theory. ScholarGate. https://scholargate.app/en/psychometrics/item-response-theory
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Confirmatory factor analysisPsychometrics↔ compare
- Differential Item FunctioningPsychometrics↔ compare
- EFAStatistics↔ compare
- Rasch ModelPsychometrics↔ compare
- Scale developmentPsychometrics↔ compare