Longitudinal Measurement Invariance Testing
Also known as: LMI, longitudinal invariance, measurement equivalence across time, temporal measurement invariance
Longitudinal measurement invariance testing determines whether a psychological scale measures the same construct in the same way across two or more time points. It is a prerequisite for interpreting mean-level change scores in panel and repeated-measures studies, ensuring that observed change reflects true change in the construct rather than drift in the measurement instrument.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
+5 more
When to use it
Use longitudinal measurement invariance testing whenever you plan to compare latent means across time points in a panel or repeated-measures design, interpret the magnitude of change in a construct, or include a latent variable in a latent growth curve model. It should be reported as a formal analytic step before any longitudinal latent variable analysis. Do not use it when you have only a single time point, when you are working with a single observed variable rather than a latent construct, or when the scale has been substantially modified between waves — in that case the scale itself cannot be treated as equivalent and a different validation strategy is needed.
Strengths & limitations
- Provides formal statistical evidence that observed longitudinal change reflects genuine construct change rather than measurement drift.
- The hierarchical sequence of models (configural → metric → scalar) pinpoints exactly which parameters are non-invariant, enabling targeted remediation.
- Partial invariance frameworks allow mean comparisons to proceed even when full scalar invariance fails, as long as at least two items per factor are invariant and the exception is reported.
- Well-integrated into standard SEM software (R lavaan, Mplus, AMOS), making the tests readily accessible.
- Applicable to any multi-indicator latent variable measured at two or more occasions, regardless of sample size tier (with appropriate fit criteria).
- Requires a multi-indicator latent variable with the same items administered at every wave; single-item measures cannot be tested.
- The chi-square difference test is sensitive to sample size and can flag trivially small parameter differences as significant in large samples, making fit-change indices (ΔCFI, ΔRMSEA) preferable but not universally agreed upon.
- Partial invariance findings complicate interpretation and require researchers to make judgment calls about whether remaining comparisons are defensible.
- Testing assumes the same factor structure holds at every wave; if the construct itself is theoretically expected to reorganize over time (developmental restructuring), strict invariance testing may be inappropriate.
Frequently asked
Is metric invariance sufficient for comparing latent means across time?
No. Metric invariance (equal loadings) is required for comparing latent variances and covariances, but scalar invariance (equal loadings and intercepts) is required for comparing latent means. Without equal intercepts, individuals with the same latent trait level will appear to score differently on the observed items at different waves, making mean comparisons uninterpretable.
What should I do if full scalar invariance fails?
Examine modification indices to identify which item intercepts are non-invariant. If at least two intercepts per factor are invariant, partial scalar invariance can be accepted, the non-invariant intercepts can be freed, and latent mean comparisons can proceed with appropriate caveats. Report the specific non-invariant parameters and discuss what they imply substantively.
How does longitudinal invariance differ from multigroup invariance?
They use the identical analytic framework and model sequence, but they address different research questions. Longitudinal invariance asks whether the same scale measures the same construct equivalently across time within the same cohort. Multigroup invariance asks whether the scale measures the same construct equivalently across different groups (e.g., men vs. women) at the same time point.
How many time points are needed?
A minimum of two time points is required. With only two waves the scalar model is just-identified for the intercept constraints, so it is advisable to have at least three waves or a well-powered sample to obtain stable fit estimates. More waves increase the precision of the invariance tests and allow detection of gradual measurement drift.
Which software can run these tests?
The most common choices are R (lavaan package, with the measEq.syntax helper or semTools::compareFit), Mplus (using multiple-group syntax with time points as groups), and AMOS. All three produce the nested model sequence, chi-square difference tests, and modification indices needed for a complete longitudinal invariance analysis.
Sources
- Meredith, W. (1993). Measurement invariance, factor analysis and factorial invariance. Psychometrika, 58(4), 525–543. DOI: 10.1007/BF02294825 ↗
- Vandenberg, R. J. & Lance, C. E. (2000). A review and synthesis of the measurement invariance literature: Suggestions, practices, and recommendations for organizational research. Organizational Research Methods, 3(1), 4–70. DOI: 10.1177/109442810031002 ↗
How to cite this page
ScholarGate. (2026, June 3). Longitudinal Measurement Invariance Testing. ScholarGate. https://scholargate.app/en/psychometrics/longitudinal-measurement-invariance
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Confirmatory factor analysisPsychometrics↔ compare
- Measurement InvariancePsychometrics↔ compare
- Structural Equation ModelingResearch Statistics↔ compare