Longitudinal Item Analysis
Also known as: LIA, repeated-measures item analysis, longitudinal item calibration, item parameter stability analysis
Longitudinal item analysis examines how the statistical properties of individual scale items — difficulty, discrimination, factor loadings, and fit — remain stable or change systematically across repeated measurement occasions. It is the item-level foundation of longitudinal measurement validity.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use longitudinal item analysis whenever a scale is administered at two or more time points and you wish to interpret change in latent means or latent relationships over time. It is mandatory before drawing conclusions about development, intervention effects, or trajectories, because non-invariant items cause biased comparisons. It is also appropriate for scale maintenance: checking that items function consistently across annual cohorts of a tracking study. Do not apply it when you have only one wave of data (cross-sectional designs call for DIF analysis by group, not time); when the sample at each wave is too small to support stable parameter estimation (fewer than about 200 per wave for ordinal data); or when the construct is theoretically expected to change in definition over time — in that case invariance failure may be substantively meaningful rather than a measurement flaw.
Strengths & limitations
- Provides direct evidence that score changes across occasions reflect true latent change rather than item drift.
- Works within both CTT (item-total correlations, alpha) and IRT frameworks, making it applicable across measurement traditions.
- Partial invariance solutions allow salvaging longitudinal comparisons even when some items drift, rather than abandoning the analysis.
- Identifies specific problematic items, enabling targeted scale revision rather than wholesale instrument replacement.
- Supports credible causal and growth-curve modelling by establishing the measurement preconditions.
- Requires the same instrument and, ideally, the same sample at each wave; panel attrition can introduce sample selection bias into item parameter estimates.
- In IRT frameworks, parameter linking across waves adds complexity and requires specialised software and expertise.
- Distinguishing true item drift from construct evolution is a conceptual challenge: non-invariance may signal a measurement problem or a substantively important change in the construct itself.
Frequently asked
Is longitudinal item analysis the same as measurement invariance testing?
They are closely related but not identical. Measurement invariance testing is the confirmatory analytic strategy — the sequence of constrained models. Longitudinal item analysis is the broader framework that includes descriptive item statistics, IRT parameter calibration, and linking, as well as invariance testing. Think of invariance testing as one key tool within the longitudinal item analysis process.
What sample size is needed?
For CTT-based analysis a common minimum is about 100–150 per wave. For IRT-based calibration with ordinal polytomous items, 200 or more per wave is advisable; for the 3-parameter logistic IRT model the recommendation rises to 500 or more. Smaller samples increase the risk of unstable parameter estimates and false non-invariance findings.
What if I find partial rather than full scalar invariance?
Partial scalar invariance means that at least some item intercepts are freely estimated across waves. You can still compare latent means meaningfully if at least two items per factor are fully invariant to anchor the scale. Acknowledge the limitation, identify which items drift, and interpret latent mean comparisons with appropriate caution.
Should I use CTT or IRT for longitudinal item analysis?
CTT is simpler and adequate for screening and scale maintenance in modest samples. IRT is preferred when item-level precision matters, when the test design allows adaptive administration, or when linking across alternate forms is required. Both frameworks can be applied longitudinally; the choice depends on sample size, software access, and the depth of item-level information needed.
How do I handle missing data across waves?
Full-information maximum likelihood (FIML) or multiple imputation are the recommended approaches. Both handle missing-at-random data without listwise deletion and are supported by major SEM software. Listwise deletion reduces sample size and can bias parameter estimates if dropout is related to the construct being measured.
Sources
- Meade, A. W., Johnson, E. C. & Braddy, P. W. (2008). Power and sensitivity of alternative fit indices in tests of measurement invariance. Journal of Applied Psychology, 93(3), 568–592. DOI: 10.1037/0021-9010.93.3.568 ↗
- Vandenberg, R. J. & Lance, C. E. (2000). A review and synthesis of the measurement invariance literature: Suggestions, practices, and recommendations for organizational research. Organizational Research Methods, 3(1), 4–70. DOI: 10.1177/109442810031002 ↗
How to cite this page
ScholarGate. (2026, June 3). Longitudinal Item Analysis. ScholarGate. https://scholargate.app/en/psychometrics/longitudinal-item-analysis
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Confirmatory factor analysisPsychometrics↔ compare
- Differential Item FunctioningPsychometrics↔ compare
- Item Response TheoryPsychometrics↔ compare
- Longitudinal CFAPsychometrics↔ compare
- Longitudinal Measurement InvariancePsychometrics↔ compare
- Test-Retest ReliabilityPsychometrics↔ compare