Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Psychometrics›Test-Retest Reliability
Latent structureScale / measurement

Test-Retest Reliability

Also known as: stability reliability, temporal stability, repeatability coefficient, TRT reliability

Test-retest reliability quantifies the temporal consistency of a measure by correlating scores obtained from the same participants on two separate occasions. It is a cornerstone of psychometric validation, directly indicating whether a scale or instrument yields stable scores when the underlying construct has not changed.

ScholarGate
  1. Latent structure
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Test-Retest Reliability
Confirmatory factor anal…Generalizability TheoryInterrater ReliabilityMeasurement InvarianceCAT Generalizability The…CAT Test-Retest Reliabil…Computerized adaptive te…Longitudinal Construct V…Longitudinal Cronbach's…Longitudinal Item Analys…

+9 more

When to use it

Use test-retest reliability when you need to demonstrate the temporal stability of a scale or measure as part of a psychometric validation study, particularly for constructs that are theoretically stable across the selected interval (e.g., personality traits, chronic symptoms, cognitive ability). It is the primary reliability evidence for performance-based or observer-rated measures where internal consistency is not estimable. Do not use test-retest as the sole reliability evidence for instruments measuring rapidly changing states (e.g., mood, pain intensity), and do not use it when a suitable retest interval cannot be justified theoretically or practically. It is also unsuitable when participant attrition between administrations is substantial, as selective dropout biases the correlation estimate.

Strengths & limitations

Strengths
  • Directly captures temporal stability, which is the most practically relevant form of reliability for stable constructs.
  • Applicable to any format of measurement — questionnaires, performance tests, observer ratings, physiological indices — without requiring multiple parallel items.
  • The standard error of measurement derived from r_tt enables clinically interpretable confidence intervals around individual scores.
  • Straightforward to compute and to report, and universally understood by reviewers and practitioners.
  • Intraclass correlation variants allow assessment of both rank-order consistency and absolute score agreement simultaneously.
Limitations
  • Choosing an appropriate inter-test interval is difficult: no single interval is universally correct, and the wrong choice can inflate or deflate the estimate.
  • Reactive effects — memory, practice, fatigue, sensitisation — can bias the retest correlation up or down depending on the instrument and the interval.
  • For constructs that genuinely change rapidly (states, acute symptoms), a low r_tt does not distinguish poor reliability from true score change.
  • Participant attrition between time points introduces selection bias if dropout is related to the construct being measured.
  • A single r_tt coefficient captures only one source of measurement error (time); it does not reflect error from item sampling or rater inconsistency.

Frequently asked

What is the difference between test-retest reliability and internal consistency?

Internal consistency (e.g., Cronbach's alpha) reflects how uniformly items within a single administration measure the same construct — it is estimated from one testing occasion. Test-retest reliability reflects temporal stability — it is estimated from two administrations and captures how much scores fluctuate over time due to measurement error. They are distinct sources of reliability evidence and should both be reported in a full psychometric validation.

Should I use Pearson r or the intraclass correlation coefficient (ICC)?

Pearson r measures only rank-order agreement and is insensitive to systematic shifts in the mean between occasions. The ICC measures both rank-order agreement and absolute agreement, making it more appropriate when the level of scores — not just the ordering — matters. For most clinical and health applications, the ICC with a two-way mixed model and absolute agreement is the preferred index.

How long should the interval between test and retest be?

There is no universally correct interval. It should be long enough to minimise memory and practice effects (typically at least two weeks for self-report scales) but short enough that the target construct is not expected to change meaningfully. The chosen interval must be reported and justified on theoretical grounds in any publication.

What is a minimum acceptable test-retest coefficient?

Conventional thresholds are r >= 0.70 for research instruments and r >= 0.80 for instruments used in applied or clinical decisions. These are guidelines, not hard rules; a lower value may be acceptable for highly dynamic constructs, and a higher value should be required for high-stakes individual assessment.

Can test-retest reliability be computed for a short scale with only a few items?

Yes. Unlike internal consistency, test-retest reliability does not require multiple items — it can be computed for a single-item measure or an index score. This makes it particularly valuable for demonstrating the stability of brief or single-item instruments.

Sources

  1. Nunnally, J. C. & Bernstein, I. H. (1994). Psychometric Theory (3rd ed.). McGraw-Hill. ISBN: 978-0070478497
  2. Anastasi, A. & Urbina, S. (1997). Psychological Testing (7th ed.). Prentice Hall. ISBN: 978-0023030857

How to cite this page

ScholarGate. (2026, June 3). Test-Retest Reliability. ScholarGate. https://scholargate.app/en/psychometrics/test-retest-reliability

Related methods

Confirmatory factor analysisGeneralizability TheoryInterrater ReliabilityMeasurement Invariance

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Confirmatory factor analysisPsychometrics↔ compare
  • Generalizability TheoryPsychometrics↔ compare
  • Interrater ReliabilityPsychometrics↔ compare
  • Measurement InvariancePsychometrics↔ compare
Compare side by side →

Referenced by

CAT Generalizability TheoryCAT Test-Retest ReliabilityComputerized adaptive test reliability analysisGeneralizability TheoryLongitudinal Construct ValidityLongitudinal Cronbach's AlphaLongitudinal Item AnalysisLongitudinal McDonald's omegaLongitudinal Reliability AnalysisLongitudinal scale developmentLongitudinal Test-Retest ReliabilityMulti-group test-retest reliabilityMultilevel Test-Retest ReliabilityOrdinal Test-Retest ReliabilityShort-form reliability analysisShort-form test-retest reliability

Similar methods

Longitudinal Test-Retest ReliabilityShort-form test-retest reliabilityRobust Test-Retest ReliabilityOrdinal Test-Retest ReliabilityMulti-group test-retest reliabilityMultilevel Test-Retest ReliabilityCAT Test-Retest ReliabilityLongitudinal Reliability Analysis

Related reference concepts

Measurement Validity and ReliabilityPsychological Testing and PsychometricsPsychometrics & Statistics & MethodologyTest ReliabilityTests & TestingInterrater Reliability

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Test-Retest Reliability (Test-Retest Reliability). Retrieved 2026-07-21 from https://scholargate.app/en/psychometrics/test-retest-reliability · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Karl Pearson
Year
1904
Type
Reliability estimate
DataType
Continuous or ordinal scores from two administrations
Subfamily
Scale / measurement
Related methods
Confirmatory factor analysisGeneralizability TheoryInterrater ReliabilityMeasurement Invariance
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account