Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Psychometrics›Multi-group Test-Retest Reliability
Latent structureScale / measurement

Multi-group Test-Retest Reliability

Multi-group Test-Retest Reliability Analysis · Also known as: multi-group temporal stability, cross-group test-retest reliability, group-comparative retest reliability, multi-sample temporal consistency

Multi-group test-retest reliability evaluates whether a measure produces stable scores across time separately for two or more defined groups — such as different genders, age cohorts, or clinical populations — and determines whether the degree of that temporal stability is equivalent across those groups.

ScholarGate
  1. Latent structure
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Multi-group test-retest reliability
Confirmatory factor anal…Multi-group confirmatory…Multi-group Cronbach's a…Multi-group measurement…Test-Retest Reliability

When to use it

Use multi-group test-retest reliability when validating a scale for use in multiple populations and equity of temporal consistency across those populations is a scientific or practical requirement. Typical scenarios include cross-cultural adaptation of instruments, clinical assessment tools used with demographically diverse samples, and educational tests compared across school cohorts. Do not use this procedure when only a single group is relevant, when the construct is expected to change substantially within the retest window (in which case test-retest reliability is inappropriate for any group), or when sample sizes within subgroups are too small (n < 30) to yield interpretable confidence intervals.

Strengths & limitations

Strengths
  • Reveals group-specific reliability that a single pooled ICC would conceal, enabling fairer scale evaluation.
  • ICC is grounded in a well-established variance-components framework (Shrout & Fleiss, 1979) that handles repeated measures appropriately.
  • The Fisher z-test for comparing ICCs across groups provides a formal statistical test rather than informal comparison of point estimates.
  • Supports equity auditing of measurement instruments by exposing differential temporal stability across demographic or clinical subgroups.
  • Applicable to any scored scale — total scores, subscale scores, or individual items — without requiring a factor-analytic model.
Limitations
  • Requires two data-collection waves, making it logistically more demanding and costly than single-administration reliability indices.
  • The optimal retest interval is construct-dependent and often unclear; too short invites memory contamination, too long risks real change masquerading as measurement error.
  • Group-specific ICC estimates can be imprecise when subgroup sample sizes are small, yielding wide confidence intervals that are uninformative.
  • Does not address whether items function the same way across groups — measurement invariance testing with CFA is needed for that.
  • Attrition between waves (participants lost to follow-up) can introduce bias if dropout is non-random or group-differential.

Frequently asked

How is multi-group test-retest reliability different from measurement invariance testing?

Test-retest reliability focuses on temporal stability: do scores from the same respondents correlate across two occasions? Measurement invariance testing uses confirmatory factor analysis to examine whether the factor loadings, intercepts, and error variances are equal across groups at a single point in time. Both are important for cross-group comparisons, but they address different validity questions and cannot substitute for each other.

Which ICC model should I use?

Use a two-way mixed-effects ICC (Model 3 in Shrout & Fleiss notation) when the two time points are fixed and constitute the full set of occasions you care about — the most common scenario. Use a two-way random-effects ICC (Model 2) when the occasions are a sample from a broader universe of possible measurement moments. One-way ICCs are generally inappropriate for test-retest because they cannot separate rater (occasion) variance from residual variance.

How large does each group need to be?

There is no universal minimum, but a practical guideline is at least 30 to 50 participants per group to obtain confidence intervals narrow enough to be informative. Smaller groups yield ICCs that may appear acceptable as point estimates while the confidence interval spans the full range from poor to excellent, making the estimate nearly worthless for decision-making.

What retest interval should I use?

The interval should be long enough that participants cannot recall their previous answers (typically at least two weeks for most psychological scales) but short enough that the construct being measured has not meaningfully changed. For stable traits such as personality, intervals of two to four weeks are common. For states or clinical symptoms that may fluctuate, the appropriate interval may be shorter, but the analyst must then argue that observed instability reflects measurement error rather than true change.

Can I compare more than two groups?

Yes. Group-specific ICCs and confidence intervals are computed independently for each group. Pairwise Fisher z-tests compare all pairs, or an omnibus test can evaluate whether the ICCs are homogeneous across all groups simultaneously. With many groups, adjust the significance threshold for multiple comparisons.

Sources

  1. Shrout, P. E. & Fleiss, J. L. (1979). Intraclass correlations: Uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428. DOI: 10.1037/0033-2909.86.2.420 ↗
  2. Vandenberg, R. J. & Lance, C. E. (2000). A review and synthesis of the measurement invariance literature: Suggestions, practices, and recommendations for organizational research. Organizational Research Methods, 3(1), 4–70. DOI: 10.1177/109442810031002 ↗

How to cite this page

ScholarGate. (2026, June 3). Multi-group Test-Retest Reliability Analysis. ScholarGate. https://scholargate.app/en/psychometrics/multi-group-test-retest-reliability

Related methods

Confirmatory factor analysisMulti-group confirmatory factor analysisMulti-group Cronbach's alphaMulti-group measurement invarianceTest-Retest Reliability

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Confirmatory factor analysisPsychometrics↔ compare
  • Multi-group confirmatory factor analysisPsychometrics↔ compare
  • Multi-group Cronbach's alphaPsychometrics↔ compare
  • Multi-group measurement invariancePsychometrics↔ compare
  • Test-Retest ReliabilityPsychometrics↔ compare
Compare side by side →

Similar methods

Multilevel Test-Retest ReliabilityLongitudinal Test-Retest ReliabilityMulti-group Reliability AnalysisTest-Retest ReliabilityShort-form test-retest reliabilityRobust Test-Retest ReliabilityLongitudinal Reliability AnalysisMulti-group measurement invariance

Related reference concepts

Psychological Testing and PsychometricsInterrater ReliabilityPsychometrics & Statistics & MethodologyMeasurement Validity and ReliabilityTests & TestingTest Reliability

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Multi-group test-retest reliability (Multi-group Test-Retest Reliability Analysis). Retrieved 2026-07-21 from https://scholargate.app/en/psychometrics/multi-group-test-retest-reliability · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Systematic multi-group extensions developed alongside measurement invariance frameworks (Vandenberg & Lance, 2000); intraclass correlation foundation in Shrout & Fleiss (1979)
Year
1979–2000
Type
Reliability estimation across groups
DataType
Repeated-measures scores from two or more defined groups
Subfamily
Scale / measurement
Related methods
Confirmatory factor analysisMulti-group confirmatory factor analysisMulti-group Cronbach's alphaMulti-group measurement invarianceTest-Retest Reliability
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account