Latent structurePsychometricsScale / measurementModel

Multilevel Test-Retest Reliability

Also known as: hierarchical test-retest reliability, multilevel ICC reliability, nested test-retest reliability, ML-TRT reliability

OriginatorShrout & Fleiss (ICC foundation); multilevel extension by Goldstein, Snijders, and othersYear1979 (ICC foundation); multilevel extension: 1990s–2000sSources2Related methods5

Multilevel test-retest reliability estimates how consistently a measurement instrument produces the same scores across repeated administrations when observations are nested within higher-level units — such as patients within clinics or students within classrooms. It partitions total score variance across levels using intraclass correlation coefficients derived from multilevel models.

Key highlights

  • Correctly partitions measurement error across levels, preventing inflation or deflation of reliability estimates caused by ignored clustering.
  • Yields level-specific reliability coefficients, revealing whether an instrument is more stable across persons or across occasions.
  • Handles unbalanced designs (unequal cluster sizes, missing occasions) through REML estimation without list-wise deletion.
  • Can incorporate covariates at each level to examine conditional reliability (e.g., reliability within demographic subgroups).
  • Directly extends to three or more levels, accommodating complex nested designs common in large-scale educational or clinical research.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use multilevel test-retest reliability when repeated measures are nested within identifiable higher-level clusters and you need a reliability estimate that honestly accounts for that structure. Typical applications include clinical instrument validation across sites, educational assessments within schools, and ecological momentary assessment data nested within persons. Do not use it as a simple substitute for single-level test-retest ICC when there is no meaningful clustering — the added complexity is unwarranted. Also avoid it when cluster sizes are very small (fewer than five units per cluster), as variance components become poorly estimated, or when the number of occasions is fewer than two per person.

Strengths & limitations

Strengths
  • Correctly partitions measurement error across levels, preventing inflation or deflation of reliability estimates caused by ignored clustering.
  • Yields level-specific reliability coefficients, revealing whether an instrument is more stable across persons or across occasions.
  • Handles unbalanced designs (unequal cluster sizes, missing occasions) through REML estimation without list-wise deletion.
  • Can incorporate covariates at each level to examine conditional reliability (e.g., reliability within demographic subgroups).
  • Directly extends to three or more levels, accommodating complex nested designs common in large-scale educational or clinical research.
Limitations
  • Requires sufficient cluster number (typically at least 20–30 clusters) and cluster size for stable variance component estimates; small designs produce wide confidence intervals around ICC.
  • REML estimation is iterative and can fail to converge in poorly specified or sparse models.
  • Interpreting multiple ICCs (between-person, within-person, cluster-level) simultaneously is more complex than reporting a single reliability coefficient.
  • Assumes that random effects are normally distributed; violations can bias variance components, especially with few clusters.
  • Software implementation varies across packages, and different default parameterisations can produce non-identical ICC values for the same data.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

How is multilevel test-retest reliability different from ordinary test-retest reliability?

Ordinary test-retest reliability (single-level ICC) treats all observations as independent. Multilevel test-retest reliability explicitly models the nesting of persons within clusters (e.g., clinics, schools), partitioning variance across levels and yielding separate reliability estimates at each level. Using a single-level ICC on clustered data inflates or deflates the estimate depending on the clustering structure.

Which ICC form should I report — consistency or agreement?

Agreement ICC is appropriate when the absolute value of scores matters, for example when a clinical decision depends on the score crossing a threshold. Consistency ICC is appropriate when only rank ordering is important. Agreement is the more conservative and usually more defensible choice in measurement validation contexts.

How many occasions and clusters do I need?

There is no universal rule, but simulation studies suggest at least two occasions per person, at least 20–30 clusters, and at least five to ten persons per cluster for stable variance component estimates. Fewer clusters lead to wide confidence intervals and potentially biased random-effect estimates.

Can I use multilevel test-retest reliability with ordinal data?

Standard REML-based multilevel reliability assumes continuous responses. For ordinal items, a multilevel polychoric or ordinal probit model is more appropriate. Alternatively, item-level reliability can be estimated within a multilevel confirmatory factor analysis framework using categorical indicators.

What software can estimate multilevel test-retest reliability?

Common options include the lme4 and psych packages in R (which allows manual ICC computation from variance components), Mplus for multilevel CFA-based reliability, and SPSS mixed models. Confidence intervals should always be obtained via bootstrap or likelihood-ratio profiling rather than asymptotic approximations in small samples.

Sources

  1. 1.
    Shrout, P. E. & Fleiss, J. L. (1979). Intraclass correlations: Uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428.
  2. 2.
    Liljequist, D., Elfving, B. & Skavberg Roaldsen, K. (2019). Intraclass correlation: A discussion and demonstration of basic features. PLOS ONE, 14(7), e0219854.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 3). Multilevel Test-Retest Reliability. ScholarGate. https://scholargate.app/psychometrics/multilevel-test-retest-reliability

Multilevel Test-Retest Reliability | ScholarGate