Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Psychometrics›Longitudinal Item Response Theory (LIRT)
Latent structureScale / measurement

Longitudinal Item Response Theory (LIRT)

Longitudinal Item Response Theory · Also known as: LIRT, longitudinal IRT, repeated-measures IRT, dynamic item response modeling

Longitudinal IRT extends classical item response theory to data collected at multiple time points, allowing researchers to model both the initial latent trait level and its change over time. It is used in educational assessment, clinical trials, and panel studies where the same items or item banks are administered repeatedly to the same individuals.

ScholarGate
  1. Latent structure
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Longitudinal IRT
Differential Item Functi…Item Response TheoryLongitudinal CFALongitudinal Measurement…Longitudinal DIF

When to use it

Use longitudinal IRT when you have panel data with repeated administration of the same items or item banks and your primary interest is in true latent-trait change — not merely observed score change. It is especially appropriate in educational growth studies, clinical symptom tracking over treatment phases, and longitudinal attitude research. Do not use standard IRT applied separately to each wave and then compare theta estimates: without constraining item parameters across waves the scale metrics are arbitrary and incomparable. Do not use longitudinal IRT when samples are small (fewer than roughly 200–300 per wave), when the number of items per wave is very low, or when the items clearly measure different constructs at different ages (as in developmental research where content validity shifts with age — here domain-referenced approaches or vertical scaling may be more appropriate).

Strengths & limitations

Strengths
  • Provides properly scaled change estimates by holding item parameters constant across waves, unlike naive observed-score comparisons.
  • Separates measurement error from true intra-individual change, yielding more reliable growth trajectories.
  • Can accommodate missing data across waves within MML or Bayesian frameworks, provided data are missing at random.
  • Allows formal testing of longitudinal measurement invariance, making the assumptions behind change scores transparent rather than implicit.
  • Compatible with complex item formats (polytomous, mixed response types) via appropriate IRT models (GRM, PCM, GPCM).
Limitations
  • Requires substantially larger samples than single-wave IRT to estimate the joint trait distribution across occasions reliably.
  • Model complexity increases rapidly with the number of waves, item types, and growth-curve specifications.
  • Software and analyst expertise requirements are higher than for standard CFA-based longitudinal models; accessible implementations remain limited.
  • Assumes that the items form a unidimensional scale at every wave; multidimensionality across time complicates both estimation and interpretation.

Frequently asked

How is longitudinal IRT different from running IRT separately at each wave?

Running IRT separately at each wave gives theta estimates on potentially different scales at each occasion — without constraining item parameters across waves there is no guarantee that a theta of 0.5 means the same thing at wave 1 and wave 2. Longitudinal IRT constrains item parameters to be equal (or partially equal) across waves, anchoring the scale so that change in theta reflects genuine trait change rather than arbitrary metric differences.

How is longitudinal IRT different from a longitudinal CFA with measurement invariance testing?

Both test whether items measure the same construct across time and model latent change, but they differ in how they handle item difficulty and score distributions. IRT explicitly models item difficulty and discrimination (and for polytomous items, threshold parameters), so it is particularly suited when items vary widely in difficulty or when the distribution of the latent trait is skewed. CFA is more common in practice and easier to implement; IRT is preferred when item-level precision and item banking are central, or when the data are sparse (few respondents per item-score combination).

What software can I use for longitudinal IRT?

The TAM and mirt packages in R offer longitudinal and multilevel IRT extensions. More flexible growth-curve IRT models can be programmed in Stan (Bayesian) or Mplus (using its mixture and growth-curve IRT modules). No single off-the-shelf software covers all longitudinal IRT specifications, so analysts often need to combine packages or write custom code.

What sample size do I need?

Longitudinal IRT is more demanding than single-wave IRT. Simulation studies suggest roughly 200–500 respondents per wave for stable item-parameter estimates under a 2PL model with 15–20 items, depending on the growth-curve complexity and amount of missing data. Models with more parameters (polytomous items, mixture components, many waves) require larger samples.

Can I handle missing responses between waves?

Yes, MML and Bayesian estimation handle missing responses between waves under a missing-at-random (MAR) assumption, which is more defensible than listwise deletion. However, if missingness is related to the latent trait itself (missing-not-at-random), additional modelling of the missing-data mechanism is needed to avoid biased change estimates.

Sources

  1. Embretson, S. E. (1991). A multidimensional latent trait model for measuring learning and change. Psychometrika, 56(3), 495–515. DOI: 10.1007/BF02294487 ↗
  2. von Davier, M. & Carstensen, C. H. (Eds.) (2007). Multivariate and Mixture Distribution Rasch Models: Extensions and Applications. Springer. ISBN: 978-0387329161

How to cite this page

ScholarGate. (2026, June 3). Longitudinal Item Response Theory. ScholarGate. https://scholargate.app/en/psychometrics/longitudinal-item-response-theory

Related methods

Differential Item FunctioningItem Response TheoryLongitudinal CFALongitudinal Measurement Invariance

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Differential Item FunctioningPsychometrics↔ compare
  • Item Response TheoryPsychometrics↔ compare
  • Longitudinal CFAPsychometrics↔ compare
  • Longitudinal Measurement InvariancePsychometrics↔ compare
Compare side by side →

Referenced by

Longitudinal DIF

Similar methods

Longitudinal Item AnalysisLongitudinal DIFLongitudinal scale developmentLongitudinal Measurement InvarianceLongitudinal CFAItem Response TheoryLongitudinal Construct ValidityLongitudinal Reliability Analysis

Related reference concepts

Item Response TheoryStructural and Latent Variable ModelsPsychological Testing and PsychometricsLatent Class AnalysisPsychometrics & Statistics & MethodologyMeasurement

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Longitudinal IRT (Longitudinal Item Response Theory). Retrieved 2026-07-21 from https://scholargate.app/en/psychometrics/longitudinal-item-response-theory · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Susan E. Embretson
Year
1991
Type
Latent trait / longitudinal psychometric model
DataType
Repeated dichotomous or polytomous item responses across time points
Subfamily
Scale / measurement
Related methods
Differential Item FunctioningItem Response TheoryLongitudinal CFALongitudinal Measurement Invariance
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account