Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Psychometrics›Multilevel Test-Retest Reliability
Latent structureScale / measurement

Multilevel Test-Retest Reliability

Also known as: hierarchical test-retest reliability, multilevel ICC reliability, nested test-retest reliability, ML-TRT reliability

Multilevel test-retest reliability estimates how consistently a measurement instrument produces the same scores across repeated administrations when observations are nested within higher-level units — such as patients within clinics or students within classrooms. It partitions total score variance across levels using intraclass correlation coefficients derived from multilevel models.

ScholarGate
  1. Latent structure
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Multilevel Test-Retest Reliability
Confirmatory factor anal…Cronbach's AlphaGeneralizability TheoryMultilevel ModelingTest-Retest Reliability

When to use it

Use multilevel test-retest reliability when repeated measures are nested within identifiable higher-level clusters and you need a reliability estimate that honestly accounts for that structure. Typical applications include clinical instrument validation across sites, educational assessments within schools, and ecological momentary assessment data nested within persons. Do not use it as a simple substitute for single-level test-retest ICC when there is no meaningful clustering — the added complexity is unwarranted. Also avoid it when cluster sizes are very small (fewer than five units per cluster), as variance components become poorly estimated, or when the number of occasions is fewer than two per person.

Strengths & limitations

Strengths
  • Correctly partitions measurement error across levels, preventing inflation or deflation of reliability estimates caused by ignored clustering.
  • Yields level-specific reliability coefficients, revealing whether an instrument is more stable across persons or across occasions.
  • Handles unbalanced designs (unequal cluster sizes, missing occasions) through REML estimation without list-wise deletion.
  • Can incorporate covariates at each level to examine conditional reliability (e.g., reliability within demographic subgroups).
  • Directly extends to three or more levels, accommodating complex nested designs common in large-scale educational or clinical research.
Limitations
  • Requires sufficient cluster number (typically at least 20–30 clusters) and cluster size for stable variance component estimates; small designs produce wide confidence intervals around ICC.
  • REML estimation is iterative and can fail to converge in poorly specified or sparse models.
  • Interpreting multiple ICCs (between-person, within-person, cluster-level) simultaneously is more complex than reporting a single reliability coefficient.
  • Assumes that random effects are normally distributed; violations can bias variance components, especially with few clusters.
  • Software implementation varies across packages, and different default parameterisations can produce non-identical ICC values for the same data.

Frequently asked

How is multilevel test-retest reliability different from ordinary test-retest reliability?

Ordinary test-retest reliability (single-level ICC) treats all observations as independent. Multilevel test-retest reliability explicitly models the nesting of persons within clusters (e.g., clinics, schools), partitioning variance across levels and yielding separate reliability estimates at each level. Using a single-level ICC on clustered data inflates or deflates the estimate depending on the clustering structure.

Which ICC form should I report — consistency or agreement?

Agreement ICC is appropriate when the absolute value of scores matters, for example when a clinical decision depends on the score crossing a threshold. Consistency ICC is appropriate when only rank ordering is important. Agreement is the more conservative and usually more defensible choice in measurement validation contexts.

How many occasions and clusters do I need?

There is no universal rule, but simulation studies suggest at least two occasions per person, at least 20–30 clusters, and at least five to ten persons per cluster for stable variance component estimates. Fewer clusters lead to wide confidence intervals and potentially biased random-effect estimates.

Can I use multilevel test-retest reliability with ordinal data?

Standard REML-based multilevel reliability assumes continuous responses. For ordinal items, a multilevel polychoric or ordinal probit model is more appropriate. Alternatively, item-level reliability can be estimated within a multilevel confirmatory factor analysis framework using categorical indicators.

What software can estimate multilevel test-retest reliability?

Common options include the lme4 and psych packages in R (which allows manual ICC computation from variance components), Mplus for multilevel CFA-based reliability, and SPSS mixed models. Confidence intervals should always be obtained via bootstrap or likelihood-ratio profiling rather than asymptotic approximations in small samples.

Sources

  1. Shrout, P. E. & Fleiss, J. L. (1979). Intraclass correlations: Uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428. DOI: 10.1037/0033-2909.86.2.420 ↗
  2. Liljequist, D., Elfving, B. & Skavberg Roaldsen, K. (2019). Intraclass correlation: A discussion and demonstration of basic features. PLOS ONE, 14(7), e0219854. DOI: 10.1371/journal.pone.0219854 ↗

How to cite this page

ScholarGate. (2026, June 3). Multilevel Test-Retest Reliability. ScholarGate. https://scholargate.app/en/psychometrics/multilevel-test-retest-reliability

Related methods

Confirmatory factor analysisCronbach's AlphaGeneralizability TheoryMultilevel ModelingTest-Retest Reliability

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Confirmatory factor analysisPsychometrics↔ compare
  • Cronbach's AlphaStatistics↔ compare
  • Generalizability TheoryPsychometrics↔ compare
  • Multilevel ModelingResearch Statistics↔ compare
  • Test-Retest ReliabilityPsychometrics↔ compare
Compare side by side →

Similar methods

Multi-group test-retest reliabilityLongitudinal Test-Retest ReliabilityMultilevel Reliability AnalysisMultilevel Scale DevelopmentRobust Test-Retest ReliabilityTest-Retest ReliabilityShort-form test-retest reliabilityLongitudinal Reliability Analysis

Related reference concepts

Measurement Validity and ReliabilityInterrater ReliabilityPsychological Testing and PsychometricsPsychometrics & Statistics & MethodologyMultilevel and Partial Pooling ModelsItem Response Theory

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Multilevel Test-Retest Reliability (Multilevel Test-Retest Reliability). Retrieved 2026-07-21 from https://scholargate.app/en/psychometrics/multilevel-test-retest-reliability · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Shrout & Fleiss (ICC foundation); multilevel extension by Goldstein, Snijders, and others
Year
1979 (ICC foundation); multilevel extension: 1990s–2000s
Type
Reliability estimation under hierarchical data
DataType
Continuous or ordinal repeated-measures scores nested within higher-level units (e.g., individuals within clinics, items within raters)
Subfamily
Scale / measurement
Related methods
Confirmatory factor analysisCronbach's AlphaGeneralizability TheoryMultilevel ModelingTest-Retest Reliability
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account