Longitudinal Test-Retest Reliability
Also known as: longitudinal stability reliability, repeated-measurement reliability, temporal stability across waves, longitudinal retest coefficient
Longitudinal test-retest reliability quantifies how consistently a scale or measure performs across two or more time points in a longitudinal study. It extends the classic test-retest paradigm by accounting for planned, often substantive, time lags between waves — making it essential for validating instruments used in panel, cohort, or growth-curve research.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use longitudinal test-retest reliability when you are validating an instrument intended for use across multiple waves in a panel, cohort, or intervention study, and you need to demonstrate that observed score changes reflect real change rather than measurement instability. It is particularly important for instruments tracking stable traits over long intervals and for clinical measures where absolute score consistency matters. Do not use a simple bivariate r_tt when the interval is long and genuine change is expected — instead, complement it with ICC or latent-variable auto-regressive models. Avoid this approach when only a single time point is available, when the study design does not permit the same participants to be re-measured, or when the construct is inherently labile and temporal instability is substantively meaningful rather than artifactual.
Strengths & limitations
- Directly evaluates an instrument's performance under the conditions of actual longitudinal use, rather than relying solely on internal consistency from a single occasion.
- The ICC provides an estimate of absolute score agreement across occasions, which is critical for clinical and applied settings where raw scores (not just rank orders) must be interpreted.
- Separates stable between-person variance from within-person occasion-specific fluctuation, enabling more precise reliability estimates in repeated-measures designs.
- Can be extended to multiple waves and modelled within SEM frameworks, yielding richer reliability decomposition than simple Pearson correlation.
- Results support or challenge the interpretation of longitudinal change scores: if test-retest reliability is low, apparent change may be artefact.
- Distinguishing true construct change from measurement unreliability requires additional modelling assumptions or design features (e.g., control groups, latent state-trait separation) that are not always feasible.
- Carry-over effects — where memory of previous responses influences later responses — can artificially inflate reliability estimates, particularly with short intervals.
- Sample attrition over waves can introduce bias: participants who drop out may differ systematically from those who remain, distorting the reliability estimate for the full population.
- A single r_tt or ICC value collapses complex temporal reliability dynamics into one number, masking whether reliability degrades monotonically, plateaus, or varies by subgroup.
- Requires the same participants to be measured at every wave, which is logistically demanding and not always possible.
Frequently asked
How is longitudinal test-retest reliability different from ordinary test-retest reliability?
Ordinary test-retest reliability typically uses a short interval (days to weeks) chosen specifically to minimise genuine construct change, so the coefficient reflects mainly measurement error. Longitudinal test-retest reliability is conducted over the planned waves of a longitudinal study — months or years — where genuine change is expected. This requires disentangling instrument error from real score change, often using ICC or latent-variable models rather than a plain Pearson r.
What ICC value indicates acceptable longitudinal reliability?
Conventional benchmarks place ICC values of 0.75–0.90 in the 'good' range and values above 0.90 in the 'excellent' range for clinical instruments. For research scales measuring stable traits, values above 0.70 are often acceptable over longer intervals, but the appropriate threshold depends on the construct's expected temporal stability and the study's decision stakes.
Should I test measurement invariance before computing longitudinal test-retest reliability?
Yes. If the factor structure, loadings, or intercepts change across waves (non-invariance), the correlation between wave scores conflates reliability with structural change. At minimum, metric invariance (equal loadings) should be established before interpreting r_tt or ICC as a reliability estimate.
Can I estimate longitudinal test-retest reliability with only two waves?
Yes — a two-wave design is the minimum. With two waves you can compute r_tt and a two-way mixed ICC. Adding a third or later wave allows modelling of reliability decay over time and better separation of trait, state, and error components using auto-regressive SEM frameworks.
Does attrition in a longitudinal study bias the reliability estimate?
It can. If participants who drop out differ systematically on the construct (e.g., those with more extreme or unstable scores are more likely to leave), the retained sample may appear more reliable than the population. Sensitivity analyses comparing completers and non-completers, or using full-information maximum likelihood estimation, help assess and mitigate this bias.
Sources
- Nunnally, J. C. & Bernstein, I. H. (1994). Psychometric Theory (3rd ed.). McGraw-Hill. ISBN: 978-0070478497
- MacKenzie, S. B., Podsakoff, P. M. & Podsakoff, N. P. (2011). Construct measurement and validation procedures in MIS and behavioral research: Integrating new and existing techniques. MIS Quarterly, 35(2), 293–334. DOI: 10.2307/23044045 ↗
How to cite this page
ScholarGate. (2026, June 3). Longitudinal Test-Retest Reliability. ScholarGate. https://scholargate.app/en/psychometrics/longitudinal-test-retest-reliability
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Longitudinal CFAPsychometrics↔ compare
- Longitudinal Measurement InvariancePsychometrics↔ compare
- Test-Retest ReliabilityPsychometrics↔ compare