Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Psychometrics›Item Response Theory (IRT)
Latent structureScale / measurement

Item Response Theory (IRT)

Item Response Theory · Also known as: IRT, latent trait theory, item characteristic curve theory, modern test theory

Item response theory models the probability that a respondent answers an item correctly (or endorses it) as a function of the respondent's latent trait level and the item's own statistical properties — difficulty, discrimination, and guessing. Unlike classical test theory, IRT places persons and items on the same scale, yielding measurement that is sample-independent for items and test-independent for persons.

ScholarGate
  1. Latent structure
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Item Response Theory
Confirmatory factor anal…Differential Item Functi…EFARasch ModelScale developmentAdaptive screening test…Bayesian Differential It…Bayesian EFABayesian Item AnalysisBayesian Scale Developme…

+74 more

When to use it

Use IRT when you need measurement that is robust to the particular sample or item set used — for example when developing licensure exams, clinical assessments, or large-scale survey instruments that will be administered across different populations or in multiple forms. IRT is especially valuable when detecting differential item functioning across groups, when building adaptive tests, or when linking scores across test forms. Minimum sample size depends on the model: the Rasch model can work with 150–200 respondents for stable estimates; the 2PL typically requires 300–500; the 3PL needs 500 or more. Do not use IRT when your sample is very small (n < 100), when items are poorly written and unlikely to fit any model, or when classical test theory statistics (Cronbach's alpha, corrected item-total correlations) are sufficient for the purpose at hand.

Strengths & limitations

Strengths
  • Provides sample-independent item parameter estimates and test-independent person estimates, enabling equating across forms.
  • Yields item-level diagnostic information — fit statistics and ICCs — that classical test theory cannot provide.
  • Supports adaptive testing by identifying items that maximise information at each trait level.
  • Enables formal detection of differential item functioning (DIF) across demographic or cultural groups.
  • Flexibly handles binary, polytomous, and mixed item formats through model extensions.
Limitations
  • Requires substantially larger samples than classical test theory for stable parameter estimation, especially for the 2PL and 3PL.
  • Model fit must be checked — items and persons that misfit should be investigated, which adds analytic complexity.
  • Unidimensionality is a core assumption; multidimensional data violate it and require more complex multidimensional IRT models.
  • Software (e.g., R packages mirt, ltm; Mplus; IRTPRO) is more technical than classical test theory tools and results are harder to communicate to non-specialist audiences.

Frequently asked

What is the difference between IRT and classical test theory?

Classical test theory (CTT) focuses on total scores and characterises items by their average difficulty (p-value) and item-total correlation, both of which depend on the particular sample tested. IRT models the probability of each response as a function of a latent trait, estimating item parameters that are theoretically sample-independent and person parameters that are test-independent. IRT therefore enables equating across tests and more detailed item diagnostics, but requires larger samples.

What is the difference between the Rasch model, 2PL, and 3PL?

These are nested special cases of the general logistic IRT model. The Rasch (1PL) model uses only an item difficulty parameter (b), constraining all items to have equal discrimination. The 2PL adds a discrimination parameter (a) per item, allowing items to differ in how sharply they differentiate trait levels. The 3PL further adds a guessing parameter (c), modelling the probability of a correct answer by chance — appropriate for multiple-choice ability tests but not for personality scales.

How large a sample do I need for IRT?

It depends on the model. The Rasch model can produce stable item parameter estimates with roughly 150–200 respondents when there are at least 15–20 items. The 2PL generally needs 300–500 cases; the 3PL needs 500 or more. Sparse data with many items and few respondents (or vice versa) produce unstable estimates. Simulation studies and pilot testing help determine adequacy for specific designs.

Does IRT require unidimensionality?

Standard IRT models (1PL, 2PL, 3PL, GRM) assume that a single dominant latent trait drives the responses. This assumption should be checked before fitting the model, for example by examining eigenvalue ratios in an EFA, testing for essential unidimensionality, or using MIRT fit indices. When data are clearly multidimensional, a multidimensional IRT (MIRT) model is more appropriate.

What software is commonly used for IRT?

In R, the packages mirt, ltm, and TAM are widely used and freely available. Commercial options include IRTPRO, flexMIRT, and WINSTEPS (for Rasch models). Mplus supports IRT within a broader SEM framework. StatWise offers guided IRT analysis with automatic model selection and fit reporting.

Sources

  1. Lord, F. M. & Novick, M. R. (1968). Statistical Theories of Mental Test Scores. Addison-Wesley. link ↗
  2. Embretson, S. E. & Reise, S. P. (2000). Item Response Theory for Psychologists. Lawrence Erlbaum Associates. ISBN: 978-0805828191

How to cite this page

ScholarGate. (2026, June 3). Item Response Theory. ScholarGate. https://scholargate.app/en/psychometrics/item-response-theory

Related methods

Confirmatory factor analysisDifferential Item FunctioningEFARasch ModelScale development

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Confirmatory factor analysisPsychometrics↔ compare
  • Differential Item FunctioningPsychometrics↔ compare
  • EFAStatistics↔ compare
  • Rasch ModelPsychometrics↔ compare
  • Scale developmentPsychometrics↔ compare
Compare side by side →

Referenced by

Adaptive screening test evaluationBayesian Differential Item FunctioningBayesian EFABayesian Item AnalysisBayesian Scale DevelopmentBifactor ModelBookmark Standard SettingCAT Cronbach's AlphaCAT Generalizability TheoryCAT McDonald's OmegaCAT Scale DevelopmentCAT Test-Retest ReliabilityCAT-DIFCognitive Diagnostic ModelingComputerized Adaptive Test Content ValidityComputerized Adaptive Test Convergent ValidityComputerized adaptive test discriminant validityComputerized adaptive test item analysisComputerized adaptive test item response theoryComputerized adaptive test Rasch modelComputerized adaptive test reliability analysisConditional Standard Error of MeasurementDifferential Distractor FunctioningDifferential Item FunctioningDifferential Item Functioning in Educational TestingGeneralizability TheoryItem AnalysisLongitudinal DIFLongitudinal IRTLongitudinal Item AnalysisMulti-group Differential Item FunctioningMulti-group EFAMulti-group item analysisMulti-group item response theoryMulti-group Rasch modelMultidimensional Item Response TheoryMultilevel Differential Item FunctioningMultilevel Generalizability TheoryMultilevel Item Response TheoryMultilevel Rasch ModelOrdinal CFAOrdinal Differential Item FunctioningOrdinal EFAOrdinal IRTOrdinal Item AnalysisOrdinal McDonald's omegaOrdinal Measurement InvarianceOrdinal Rasch ModelOrdinal Reliability AnalysisOrdinal Scale DevelopmentPCM / GPCMPoint-Biserial CorrelationPolytomous Confirmatory Factor AnalysisPolytomous DIFPolytomous EFAPolytomous Rasch ModelPolytomous Reliability AnalysisPolytomous scale developmentRobust Cronbach's AlphaRobust Differential Item FunctioningRobust Exploratory Factor AnalysisRobust Item AnalysisRobust McDonald's OmegaRobust Rasch ModelScale developmentShort form differential item functioningShort Form Measurement InvarianceShort form Rasch modelShort-Form CFAShort-form Cronbach's alphaShort-Form IRTShort-form item analysisShort-form McDonald's omegaShort-form reliability analysisShort-Form Scale DevelopmentShort-form test-retest reliabilityStandardized Test AnalysisTest EquatingTestlet Response TheoryVertical ScalingWright Map Analysis

Similar methods

2PL IRTRasch Model3PL IRTOrdinal IRTMulti-group item response theoryLongitudinal IRTMulti-group Rasch modelMultidimensional Item Response Theory

Related reference concepts

Item Response TheoryPsychological Testing and PsychometricsPsychometrics & Statistics & MethodologyStructural and Latent Variable ModelsMeasurementPatient-Reported Outcome Measures

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Item Response Theory (Item Response Theory). Retrieved 2026-07-20 from https://scholargate.app/en/psychometrics/item-response-theory · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Frederic M. Lord (and Allan Birnbaum for the 2PL/3PL models)
Year
1952–1968
Type
Probabilistic measurement model
DataType
Binary or polytomous item responses
Subfamily
Scale / measurement
Related methods
Confirmatory factor analysisDifferential Item FunctioningEFARasch ModelScale development
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account