Short-form Test-Retest Reliability
Also known as: abbreviated scale temporal stability, short-form temporal consistency, retest reliability of short forms, SF test-retest
Short-form test-retest reliability quantifies how consistently an abbreviated version of a measurement instrument produces the same scores across two administrations separated by a defined time interval. It is a critical validation step whenever a full-length scale is shortened for practical use, confirming that item reduction has not degraded temporal stability.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Apply short-form test-retest reliability when you have developed or adopted an abbreviated version of an established scale and need to demonstrate that the reduction in item count has not compromised temporal consistency. It is especially important when the shortened tool will be used for repeated-measures designs, longitudinal monitoring, or clinical screening. Do not use this approach as a substitute for establishing the internal consistency of the short form; both forms of evidence are required. Avoid administering the two occasions so close together (under a week for stable traits) that memory rather than trait stability drives the correlation.
Strengths & limitations
- Directly assesses the most practically relevant form of reliability for repeated-measurement contexts.
- Reveals whether item reduction has introduced excess measurement error not visible from internal consistency indices alone.
- ICC provides an estimate of absolute agreement that detects score drift across occasions, not just rank-order consistency.
- Findings are interpretable to applied audiences and directly inform whether a short form is ready for field deployment.
- Confidence intervals for the ICC quantify estimation uncertainty, supporting transparent reporting.
- Requires two data collection waves, increasing burden on participants and researchers relative to single-session reliability methods.
- Choosing an appropriate retest interval involves a trade-off between memory contamination and genuine trait change that cannot be fully resolved by design alone.
- Temporal stability estimates depend on sample characteristics; a stable, homogeneous sample may overestimate reliability relative to a heterogeneous applied population.
- Reliability evidence from test-retest alone does not address construct validity; a well-replicated short form may still measure the wrong thing.
Frequently asked
What is an acceptable test-retest reliability coefficient for a short form?
Commonly cited benchmarks are .70 or above for research purposes and .80 or above for applied or clinical decisions. However, thresholds should be interpreted in context: a short form replacing a long instrument should ideally approach the parent scale's own test-retest reliability, and lower coefficients require explicit justification.
Should I report Pearson r or ICC?
Report the ICC when you want to assess absolute agreement, including any systematic mean difference between occasions. Report the two-way mixed-effects ICC model and specify whether you assessed consistency or absolute agreement. Pearson r captures rank-order stability only and will overestimate reliability if scores shift systematically between administrations.
How long should the retest interval be?
A two-to-four-week interval is conventional for stable trait constructs. The interval should reflect the expected stability of the underlying construct: shorter for state-like variables, longer only if very long-term stability is the claim. Always report the exact interval used so readers can judge its appropriateness.
Does high internal consistency guarantee adequate test-retest reliability for a short form?
No. Internal consistency (Cronbach's alpha, McDonald's omega) measures how well items correlate with each other within a single administration. Test-retest reliability measures consistency across time. A short form may have high alpha but poor temporal stability if its remaining items are sensitive to momentary fluctuations in state.
How large a sample do I need?
A minimum of 50 participants is often cited, but 100 or more is preferable to obtain stable ICC estimates with narrow confidence intervals. Sample size also affects the width of the 95% confidence interval around the ICC; reporting the interval helps readers assess the precision of your estimate.
Sources
- Smith, G. T., McCarthy, D. M., & Anderson, K. G. (2000). On the sins of short-form development. Psychological Assessment, 12(1), 102–111. DOI: 10.1037/1040-3590.12.1.102 ↗
- Nunnally, J. C., & Bernstein, I. H. (1994). Psychometric Theory (3rd ed.). McGraw-Hill. ISBN: 978-0070474659
How to cite this page
ScholarGate. (2026, June 3). Short-form Test-Retest Reliability. ScholarGate. https://scholargate.app/en/psychometrics/short-form-test-retest-reliability
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Confirmatory factor analysisPsychometrics↔ compare
- Cronbach's AlphaStatistics↔ compare
- Item Response TheoryPsychometrics↔ compare
- Test-Retest ReliabilityPsychometrics↔ compare