Test-Retest Reliability
Also known as: stability reliability, temporal stability, repeatability coefficient, TRT reliability
Test-retest reliability quantifies the temporal consistency of a measure by correlating scores obtained from the same participants on two separate occasions. It is a cornerstone of psychometric validation, directly indicating whether a scale or instrument yields stable scores when the underlying construct has not changed.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
+9 more
When to use it
Use test-retest reliability when you need to demonstrate the temporal stability of a scale or measure as part of a psychometric validation study, particularly for constructs that are theoretically stable across the selected interval (e.g., personality traits, chronic symptoms, cognitive ability). It is the primary reliability evidence for performance-based or observer-rated measures where internal consistency is not estimable. Do not use test-retest as the sole reliability evidence for instruments measuring rapidly changing states (e.g., mood, pain intensity), and do not use it when a suitable retest interval cannot be justified theoretically or practically. It is also unsuitable when participant attrition between administrations is substantial, as selective dropout biases the correlation estimate.
Strengths & limitations
- Directly captures temporal stability, which is the most practically relevant form of reliability for stable constructs.
- Applicable to any format of measurement — questionnaires, performance tests, observer ratings, physiological indices — without requiring multiple parallel items.
- The standard error of measurement derived from r_tt enables clinically interpretable confidence intervals around individual scores.
- Straightforward to compute and to report, and universally understood by reviewers and practitioners.
- Intraclass correlation variants allow assessment of both rank-order consistency and absolute score agreement simultaneously.
- Choosing an appropriate inter-test interval is difficult: no single interval is universally correct, and the wrong choice can inflate or deflate the estimate.
- Reactive effects — memory, practice, fatigue, sensitisation — can bias the retest correlation up or down depending on the instrument and the interval.
- For constructs that genuinely change rapidly (states, acute symptoms), a low r_tt does not distinguish poor reliability from true score change.
- Participant attrition between time points introduces selection bias if dropout is related to the construct being measured.
- A single r_tt coefficient captures only one source of measurement error (time); it does not reflect error from item sampling or rater inconsistency.
Frequently asked
What is the difference between test-retest reliability and internal consistency?
Internal consistency (e.g., Cronbach's alpha) reflects how uniformly items within a single administration measure the same construct — it is estimated from one testing occasion. Test-retest reliability reflects temporal stability — it is estimated from two administrations and captures how much scores fluctuate over time due to measurement error. They are distinct sources of reliability evidence and should both be reported in a full psychometric validation.
Should I use Pearson r or the intraclass correlation coefficient (ICC)?
Pearson r measures only rank-order agreement and is insensitive to systematic shifts in the mean between occasions. The ICC measures both rank-order agreement and absolute agreement, making it more appropriate when the level of scores — not just the ordering — matters. For most clinical and health applications, the ICC with a two-way mixed model and absolute agreement is the preferred index.
How long should the interval between test and retest be?
There is no universally correct interval. It should be long enough to minimise memory and practice effects (typically at least two weeks for self-report scales) but short enough that the target construct is not expected to change meaningfully. The chosen interval must be reported and justified on theoretical grounds in any publication.
What is a minimum acceptable test-retest coefficient?
Conventional thresholds are r >= 0.70 for research instruments and r >= 0.80 for instruments used in applied or clinical decisions. These are guidelines, not hard rules; a lower value may be acceptable for highly dynamic constructs, and a higher value should be required for high-stakes individual assessment.
Can test-retest reliability be computed for a short scale with only a few items?
Yes. Unlike internal consistency, test-retest reliability does not require multiple items — it can be computed for a single-item measure or an index score. This makes it particularly valuable for demonstrating the stability of brief or single-item instruments.
Sources
- Nunnally, J. C. & Bernstein, I. H. (1994). Psychometric Theory (3rd ed.). McGraw-Hill. ISBN: 978-0070478497
- Anastasi, A. & Urbina, S. (1997). Psychological Testing (7th ed.). Prentice Hall. ISBN: 978-0023030857
How to cite this page
ScholarGate. (2026, June 3). Test-Retest Reliability. ScholarGate. https://scholargate.app/en/psychometrics/test-retest-reliability
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Confirmatory factor analysisPsychometrics↔ compare
- Generalizability TheoryPsychometrics↔ compare
- Interrater ReliabilityPsychometrics↔ compare
- Measurement InvariancePsychometrics↔ compare