Ordinal Test-Retest Reliability
Also known as: rank-based test-retest reliability, ordinal temporal consistency, Spearman test-retest reliability, weighted kappa test-retest
Ordinal test-retest reliability quantifies how consistently an ordinal measurement instrument — such as a Likert-scale questionnaire or a rating tool — ranks or scores the same participants across two separate administrations separated by a stable interval, using correlation and agreement statistics suited to ordered categorical data.
Key highlights
- Directly assesses temporal stability, a distinct reliability dimension not captured by internal consistency alone.
- Rank-based statistics respect the ordinal level of measurement without imposing false interval assumptions.
- Weighted kappa provides item-level agreement evidence useful during early scale development.
- ICC allows decomposition into systematic and random error components, yielding richer reliability information.
- Results are interpretable on a standardised 0–1 scale with widely accepted benchmarks.
Intuition
This section is available to Pro members. Upgrade to Pro
How it works
This section is available to Pro members. Upgrade to Pro
When to use it
Use ordinal test-retest reliability when (1) the scale uses Likert-type or other ordered categorical response formats, (2) you need evidence that scores are stable over time under unchanged conditions, and (3) the underlying trait is theoretically stable across the chosen retest interval. Do NOT use standard Pearson-based test-retest when items are ordinal and distributions are skewed — the equal-interval assumption will inflate or distort the coefficient. Do not use test-retest as the sole reliability evidence when the construct is state-like and expected to change (e.g., mood, pain intensity today); internal consistency evidence is more appropriate in such cases.
Strengths & limitations
- Directly assesses temporal stability, a distinct reliability dimension not captured by internal consistency alone.
- Rank-based statistics respect the ordinal level of measurement without imposing false interval assumptions.
- Weighted kappa provides item-level agreement evidence useful during early scale development.
- ICC allows decomposition into systematic and random error components, yielding richer reliability information.
- Results are interpretable on a standardised 0–1 scale with widely accepted benchmarks.
- Requires two data-collection waves, increasing respondent burden, attrition, and study cost.
- Retest interval choice is inherently arbitrary; too short risks memory effects, too long risks true change confounding the estimate.
- Does not assess internal consistency or inter-rater agreement — other reliability facets must be evaluated separately.
- Stable traits in heterogeneous samples will yield inflated reliability estimates compared with homogeneous samples, limiting generalisability.
- For very short scales (fewer than four items) total-score rank correlations are highly sensitive to individual item outliers.
Common pitfalls
This section is available to Pro members. Upgrade to Pro
Applications
This section is available to Pro members. Upgrade to Pro
Frequently asked
Can I use Pearson correlation for test-retest if I have Likert data?
It is often done in practice, but it is not ideal. Pearson correlation assumes interval-level data and normally distributed differences, assumptions that Likert items frequently violate. Spearman rho or weighted kappa is preferable because they treat the data as ordinal ranks and are robust to skewed or bounded distributions.
What retest interval should I use?
The interval should be long enough that participants cannot recall their specific answers (typically at least one week) but short enough that the construct being measured is unlikely to have genuinely changed. For stable personality or attitude traits, two to four weeks is common. For fluctuating states, test-retest evidence is less meaningful and internal consistency should be emphasised instead.
What value of Spearman rho indicates acceptable reliability?
Conventional benchmarks vary by field and application. Values above 0.80 are generally considered good temporal stability for clinical and psychometric instruments. For high-stakes decisions, researchers often require 0.90 or above. Always report confidence intervals alongside the point estimate.
When should I use weighted kappa instead of Spearman rho?
Weighted kappa is most appropriate when analysing agreement at the individual item level (each item's category rating at time 1 vs. time 2) rather than total scores, or when the scale has very few ordered categories (e.g., 3 or 4 response options). For summed ordinal total scores with many possible values, Spearman rho or ICC is more natural.
Does a high test-retest coefficient prove that my scale is reliable?
Test-retest reliability is one facet of reliability; it shows the scale produces stable scores over time. It does not address whether items in the scale are internally consistent, whether different raters would assign the same scores, or whether the scale measures what it claims to measure. A complete reliability evaluation combines test-retest evidence with internal consistency (omega or alpha) and, where applicable, inter-rater agreement.
Sources
- 1.Shrout, P. E., & Fleiss, J. L. (1979). Intraclass correlations: Uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428.
- 2.Cohen, J. (1968). Weighted kappa: Nominal scale agreement with provision for scaled disagreement or partial credit. Psychological Bulletin, 70(4), 213–220.
You have read it. What now?
Cite this page
ScholarGate. (2026, June 3). Ordinal Test-Retest Reliability. ScholarGate. https://scholargate.app/psychometrics/ordinal-test-retest-reliability