Latent structurePsychometricsScale / measurementModel

Test-Retest Reliability

Also known as: stability reliability, temporal stability, repeatability coefficient, TRT reliability

OriginatorKarl PearsonYear1904Sources2Related methods20

Test-retest reliability quantifies the temporal consistency of a measure by correlating scores obtained from the same participants on two separate occasions. It is a cornerstone of psychometric validation, directly indicating whether a scale or instrument yields stable scores when the underlying construct has not changed.

Key highlights

  • Directly captures temporal stability, which is the most practically relevant form of reliability for stable constructs.
  • Applicable to any format of measurement — questionnaires, performance tests, observer ratings, physiological indices — without requiring multiple parallel items.
  • The standard error of measurement derived from r_tt enables clinically interpretable confidence intervals around individual scores.
  • Straightforward to compute and to report, and universally understood by reviewers and practitioners.
  • Intraclass correlation variants allow assessment of both rank-order consistency and absolute score agreement simultaneously.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use test-retest reliability when you need to demonstrate the temporal stability of a scale or measure as part of a psychometric validation study, particularly for constructs that are theoretically stable across the selected interval (e.g., personality traits, chronic symptoms, cognitive ability). It is the primary reliability evidence for performance-based or observer-rated measures where internal consistency is not estimable. Do not use test-retest as the sole reliability evidence for instruments measuring rapidly changing states (e.g., mood, pain intensity), and do not use it when a suitable retest interval cannot be justified theoretically or practically. It is also unsuitable when participant attrition between administrations is substantial, as selective dropout biases the correlation estimate.

Strengths & limitations

Strengths
  • Directly captures temporal stability, which is the most practically relevant form of reliability for stable constructs.
  • Applicable to any format of measurement — questionnaires, performance tests, observer ratings, physiological indices — without requiring multiple parallel items.
  • The standard error of measurement derived from r_tt enables clinically interpretable confidence intervals around individual scores.
  • Straightforward to compute and to report, and universally understood by reviewers and practitioners.
  • Intraclass correlation variants allow assessment of both rank-order consistency and absolute score agreement simultaneously.
Limitations
  • Choosing an appropriate inter-test interval is difficult: no single interval is universally correct, and the wrong choice can inflate or deflate the estimate.
  • Reactive effects — memory, practice, fatigue, sensitisation — can bias the retest correlation up or down depending on the instrument and the interval.
  • For constructs that genuinely change rapidly (states, acute symptoms), a low r_tt does not distinguish poor reliability from true score change.
  • Participant attrition between time points introduces selection bias if dropout is related to the construct being measured.
  • A single r_tt coefficient captures only one source of measurement error (time); it does not reflect error from item sampling or rater inconsistency.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

What is the difference between test-retest reliability and internal consistency?

Internal consistency (e.g., Cronbach's alpha) reflects how uniformly items within a single administration measure the same construct — it is estimated from one testing occasion. Test-retest reliability reflects temporal stability — it is estimated from two administrations and captures how much scores fluctuate over time due to measurement error. They are distinct sources of reliability evidence and should both be reported in a full psychometric validation.

Should I use Pearson r or the intraclass correlation coefficient (ICC)?

Pearson r measures only rank-order agreement and is insensitive to systematic shifts in the mean between occasions. The ICC measures both rank-order agreement and absolute agreement, making it more appropriate when the level of scores — not just the ordering — matters. For most clinical and health applications, the ICC with a two-way mixed model and absolute agreement is the preferred index.

How long should the interval between test and retest be?

There is no universally correct interval. It should be long enough to minimise memory and practice effects (typically at least two weeks for self-report scales) but short enough that the target construct is not expected to change meaningfully. The chosen interval must be reported and justified on theoretical grounds in any publication.

What is a minimum acceptable test-retest coefficient?

Conventional thresholds are r >= 0.70 for research instruments and r >= 0.80 for instruments used in applied or clinical decisions. These are guidelines, not hard rules; a lower value may be acceptable for highly dynamic constructs, and a higher value should be required for high-stakes individual assessment.

Can test-retest reliability be computed for a short scale with only a few items?

Yes. Unlike internal consistency, test-retest reliability does not require multiple items — it can be computed for a single-item measure or an index score. This makes it particularly valuable for demonstrating the stability of brief or single-item instruments.

Sources

  1. 1.
    Nunnally, J. C. & Bernstein, I. H. (1994). Psychometric Theory (3rd ed.). McGraw-Hill.
    ISBN 978-0070478497
  2. 2.
    Anastasi, A. & Urbina, S. (1997). Psychological Testing (7th ed.). Prentice Hall.
    ISBN 978-0023030857

You have read it. What now?

Cite this page

ScholarGate. (2026, June 3). Test-Retest Reliability. ScholarGate. https://scholargate.app/psychometrics/test-retest-reliability