Ordinal Test-Retest Reliability
Ordinal Test-Retest Reliability Analysis · Also known as: rank-based test-retest reliability, ordinal temporal consistency, Spearman test-retest reliability, weighted kappa test-retest
Ordinal test-retest reliability quantifies how consistently an ordinal measurement instrument — such as a Likert-scale questionnaire or a rating tool — ranks or scores the same participants across two separate administrations separated by a stable interval, using correlation and agreement statistics suited to ordered categorical data.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use ordinal test-retest reliability when (1) the scale uses Likert-type or other ordered categorical response formats, (2) you need evidence that scores are stable over time under unchanged conditions, and (3) the underlying trait is theoretically stable across the chosen retest interval. Do NOT use standard Pearson-based test-retest when items are ordinal and distributions are skewed — the equal-interval assumption will inflate or distort the coefficient. Do not use test-retest as the sole reliability evidence when the construct is state-like and expected to change (e.g., mood, pain intensity today); internal consistency evidence is more appropriate in such cases.
Strengths & limitations
- Directly assesses temporal stability, a distinct reliability dimension not captured by internal consistency alone.
- Rank-based statistics respect the ordinal level of measurement without imposing false interval assumptions.
- Weighted kappa provides item-level agreement evidence useful during early scale development.
- ICC allows decomposition into systematic and random error components, yielding richer reliability information.
- Results are interpretable on a standardised 0–1 scale with widely accepted benchmarks.
- Requires two data-collection waves, increasing respondent burden, attrition, and study cost.
- Retest interval choice is inherently arbitrary; too short risks memory effects, too long risks true change confounding the estimate.
- Does not assess internal consistency or inter-rater agreement — other reliability facets must be evaluated separately.
- Stable traits in heterogeneous samples will yield inflated reliability estimates compared with homogeneous samples, limiting generalisability.
- For very short scales (fewer than four items) total-score rank correlations are highly sensitive to individual item outliers.
Frequently asked
Can I use Pearson correlation for test-retest if I have Likert data?
It is often done in practice, but it is not ideal. Pearson correlation assumes interval-level data and normally distributed differences, assumptions that Likert items frequently violate. Spearman rho or weighted kappa is preferable because they treat the data as ordinal ranks and are robust to skewed or bounded distributions.
What retest interval should I use?
The interval should be long enough that participants cannot recall their specific answers (typically at least one week) but short enough that the construct being measured is unlikely to have genuinely changed. For stable personality or attitude traits, two to four weeks is common. For fluctuating states, test-retest evidence is less meaningful and internal consistency should be emphasised instead.
What value of Spearman rho indicates acceptable reliability?
Conventional benchmarks vary by field and application. Values above 0.80 are generally considered good temporal stability for clinical and psychometric instruments. For high-stakes decisions, researchers often require 0.90 or above. Always report confidence intervals alongside the point estimate.
When should I use weighted kappa instead of Spearman rho?
Weighted kappa is most appropriate when analysing agreement at the individual item level (each item's category rating at time 1 vs. time 2) rather than total scores, or when the scale has very few ordered categories (e.g., 3 or 4 response options). For summed ordinal total scores with many possible values, Spearman rho or ICC is more natural.
Does a high test-retest coefficient prove that my scale is reliable?
Test-retest reliability is one facet of reliability; it shows the scale produces stable scores over time. It does not address whether items in the scale are internally consistent, whether different raters would assign the same scores, or whether the scale measures what it claims to measure. A complete reliability evaluation combines test-retest evidence with internal consistency (omega or alpha) and, where applicable, inter-rater agreement.
Sources
- Shrout, P. E., & Fleiss, J. L. (1979). Intraclass correlations: Uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428. DOI: 10.1037/0033-2909.86.2.420 ↗
- Cohen, J. (1968). Weighted kappa: Nominal scale agreement with provision for scaled disagreement or partial credit. Psychological Bulletin, 70(4), 213–220. DOI: 10.1037/h0026256 ↗
How to cite this page
ScholarGate. (2026, June 3). Ordinal Test-Retest Reliability Analysis. ScholarGate. https://scholargate.app/en/psychometrics/ordinal-test-retest-reliability
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Ordinal Reliability AnalysisPsychometrics↔ compare
- Test-Retest ReliabilityPsychometrics↔ compare