Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Psychometrics›Robust Test-Retest Reliability
Latent structureScale / measurement

Robust Test-Retest Reliability

Also known as: robust temporal stability, outlier-resistant retest reliability, robust repeatability coefficient, robust intraclass correlation

Robust test-retest reliability quantifies how consistently a measure ranks or scores the same individuals across two occasions while protecting the estimate from distortion by outliers and non-normal score distributions. It replaces or supplements classical Pearson-based correlation and standard ICC formulas with robust estimators of location, scale, and association.

ScholarGate
  1. Latent structure
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Robust Test-Retest Reliability
Confirmatory factor anal…Cronbach's AlphaInterrater Reliability

When to use it

Use robust test-retest reliability when you administer the same measure twice to the same group and suspect — or wish to guard against — outliers, heavy tails, or skewed change-score distributions. It is particularly appropriate for clinical or community samples where extreme scorers are substantively common rather than data errors, for short scales with ceiling or floor effects, or for physiological measurements that are inherently variable. It is less necessary when the sample is large (n > 200), the scores are demonstrably normal, and outliers have been screened out by prior data-cleaning rules. Do not use it as a substitute for classical reliability when the research convention in your field requires a standard Pearson or ICC coefficient for comparability; in that case, report both.

Strengths & limitations

Strengths
  • Protects the reliability estimate from inflation or deflation caused by a small number of extreme or erroneous scores.
  • Appropriate for non-normal, skewed, or heavy-tailed score distributions common in clinical and applied psychometric samples.
  • Bootstrap confidence intervals avoid parametric assumptions about the sampling distribution of the statistic.
  • Yields an interpretable SEM on the original scale even when the error distribution is non-normal.
  • Can be combined with sensitivity analysis: comparing robust and classical estimates reveals how much outliers matter for a given dataset.
Limitations
  • Slightly less statistically efficient than the Pearson-based ICC when scores are genuinely normally distributed, because trimming or downweighting discards information.
  • Multiple robust estimators exist (percentage-bend, Winsorized, M-estimator-based) with no single universally agreed standard, which can make cross-study comparisons difficult.
  • Requires larger samples to achieve the same precision as classical methods because bootstrap intervals are wider when n is small.
  • Less familiar to applied reviewers than the standard ICC, so reporting must include additional justification and a comparison with the conventional estimate.

Frequently asked

Is robust test-retest reliability a completely different procedure from standard ICC?

No. The design and the conceptual goal — quantifying how consistently the same people score across two occasions — are identical. The difference is in the estimator: robust methods replace means and standard deviations with outlier-resistant counterparts before computing the reliability coefficient. The interpretation of the resulting index is the same.

Which robust estimator should I choose?

The percentage-bend correlation and the Winsorized correlation are the most widely cited. For ICC-type estimates, trimmed-mean-based methods described by Wilcox are a reasonable default. Report your choice and the trimming fraction or bending constant alongside the estimate so readers can evaluate sensitivity.

What sample size is needed?

Robust procedures generally require somewhat larger samples than classical ones to achieve equivalent precision because trimming or downweighting reduces the effective sample size. A minimum of 50 participants is a common rough floor; bootstrap confidence intervals become more stable around n = 100 or more.

Should I always prefer the robust estimate over the classical ICC?

Not necessarily. If the score distribution is close to normal and no substantive outliers are expected, the classical ICC is efficient and interpretable. Robust estimates are most valuable when outliers are plausible on substantive grounds. Reporting both allows readers to judge the influence of extreme cases.

Does a high robust test-retest coefficient prove that the measure is stable over longer periods?

No. The coefficient is bound to the specific interval used in the study. Temporal stability over a longer horizon must be evaluated in a separate study with an appropriate interval. A high coefficient over one week does not imply stability over six months.

Sources

  1. Wilcox, R. R. (2012). Introduction to Robust Estimation and Hypothesis Testing (3rd ed.). Academic Press. ISBN: 978-0123869838
  2. Test-retest reliability. Wikipedia. link ↗

How to cite this page

ScholarGate. (2026, June 3). Robust Test-Retest Reliability. ScholarGate. https://scholargate.app/en/psychometrics/robust-test-retest-reliability

Related methods

Confirmatory factor analysisCronbach's AlphaInterrater Reliability

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Confirmatory factor analysisPsychometrics↔ compare
  • Cronbach's AlphaStatistics↔ compare
  • Interrater ReliabilityPsychometrics↔ compare
Compare side by side →

Similar methods

Test-Retest ReliabilityLongitudinal Test-Retest ReliabilityShort-form test-retest reliabilityMultilevel Test-Retest ReliabilityMulti-group test-retest reliabilityOrdinal Test-Retest ReliabilityRobust Item AnalysisRobust Cronbach's Alpha

Related reference concepts

Psychological Testing and PsychometricsPsychometrics & Statistics & MethodologyMeasurement Validity and ReliabilityTests & TestingInterrater ReliabilityTest Reliability

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Robust Test-Retest Reliability (Robust Test-Retest Reliability). Retrieved 2026-07-21 from https://scholargate.app/en/psychometrics/robust-test-retest-reliability · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Built on classical test-retest reliability (Pearson, early 1900s); robust extensions formalized by Wilcox and colleagues from the 1990s onward
Year
1990s–2000s
Type
Reliability / measurement stability
DataType
Continuous or ordinal repeated measurements from two occasions
Subfamily
Scale / measurement
Related methods
Confirmatory factor analysisCronbach's AlphaInterrater Reliability
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account