Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Psychometrics›Ordinal Test-Retest Reliability
Latent structureScale / measurement

Ordinal Test-Retest Reliability

Ordinal Test-Retest Reliability Analysis · Also known as: rank-based test-retest reliability, ordinal temporal consistency, Spearman test-retest reliability, weighted kappa test-retest

Ordinal test-retest reliability quantifies how consistently an ordinal measurement instrument — such as a Likert-scale questionnaire or a rating tool — ranks or scores the same participants across two separate administrations separated by a stable interval, using correlation and agreement statistics suited to ordered categorical data.

ScholarGate
  1. Latent structure
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Ordinal Test-Retest Reliability
Ordinal Reliability Anal…Test-Retest Reliability

When to use it

Use ordinal test-retest reliability when (1) the scale uses Likert-type or other ordered categorical response formats, (2) you need evidence that scores are stable over time under unchanged conditions, and (3) the underlying trait is theoretically stable across the chosen retest interval. Do NOT use standard Pearson-based test-retest when items are ordinal and distributions are skewed — the equal-interval assumption will inflate or distort the coefficient. Do not use test-retest as the sole reliability evidence when the construct is state-like and expected to change (e.g., mood, pain intensity today); internal consistency evidence is more appropriate in such cases.

Strengths & limitations

Strengths
  • Directly assesses temporal stability, a distinct reliability dimension not captured by internal consistency alone.
  • Rank-based statistics respect the ordinal level of measurement without imposing false interval assumptions.
  • Weighted kappa provides item-level agreement evidence useful during early scale development.
  • ICC allows decomposition into systematic and random error components, yielding richer reliability information.
  • Results are interpretable on a standardised 0–1 scale with widely accepted benchmarks.
Limitations
  • Requires two data-collection waves, increasing respondent burden, attrition, and study cost.
  • Retest interval choice is inherently arbitrary; too short risks memory effects, too long risks true change confounding the estimate.
  • Does not assess internal consistency or inter-rater agreement — other reliability facets must be evaluated separately.
  • Stable traits in heterogeneous samples will yield inflated reliability estimates compared with homogeneous samples, limiting generalisability.
  • For very short scales (fewer than four items) total-score rank correlations are highly sensitive to individual item outliers.

Frequently asked

Can I use Pearson correlation for test-retest if I have Likert data?

It is often done in practice, but it is not ideal. Pearson correlation assumes interval-level data and normally distributed differences, assumptions that Likert items frequently violate. Spearman rho or weighted kappa is preferable because they treat the data as ordinal ranks and are robust to skewed or bounded distributions.

What retest interval should I use?

The interval should be long enough that participants cannot recall their specific answers (typically at least one week) but short enough that the construct being measured is unlikely to have genuinely changed. For stable personality or attitude traits, two to four weeks is common. For fluctuating states, test-retest evidence is less meaningful and internal consistency should be emphasised instead.

What value of Spearman rho indicates acceptable reliability?

Conventional benchmarks vary by field and application. Values above 0.80 are generally considered good temporal stability for clinical and psychometric instruments. For high-stakes decisions, researchers often require 0.90 or above. Always report confidence intervals alongside the point estimate.

When should I use weighted kappa instead of Spearman rho?

Weighted kappa is most appropriate when analysing agreement at the individual item level (each item's category rating at time 1 vs. time 2) rather than total scores, or when the scale has very few ordered categories (e.g., 3 or 4 response options). For summed ordinal total scores with many possible values, Spearman rho or ICC is more natural.

Does a high test-retest coefficient prove that my scale is reliable?

Test-retest reliability is one facet of reliability; it shows the scale produces stable scores over time. It does not address whether items in the scale are internally consistent, whether different raters would assign the same scores, or whether the scale measures what it claims to measure. A complete reliability evaluation combines test-retest evidence with internal consistency (omega or alpha) and, where applicable, inter-rater agreement.

Sources

  1. Shrout, P. E., & Fleiss, J. L. (1979). Intraclass correlations: Uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428. DOI: 10.1037/0033-2909.86.2.420 ↗
  2. Cohen, J. (1968). Weighted kappa: Nominal scale agreement with provision for scaled disagreement or partial credit. Psychological Bulletin, 70(4), 213–220. DOI: 10.1037/h0026256 ↗

How to cite this page

ScholarGate. (2026, June 3). Ordinal Test-Retest Reliability Analysis. ScholarGate. https://scholargate.app/en/psychometrics/ordinal-test-retest-reliability

Related methods

Ordinal Reliability AnalysisTest-Retest Reliability

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Ordinal Reliability AnalysisPsychometrics↔ compare
  • Test-Retest ReliabilityPsychometrics↔ compare
Compare side by side →

Similar methods

Test-Retest ReliabilityLongitudinal Test-Retest ReliabilityShort-form test-retest reliabilityRobust Test-Retest ReliabilityOrdinal Reliability AnalysisMulti-group test-retest reliabilityMultilevel Test-Retest ReliabilityPolytomous Reliability Analysis

Related reference concepts

Measurement Validity and ReliabilityPsychometrics & Statistics & MethodologyCorrelation and CovarianceInterrater ReliabilityPsychological Testing and PsychometricsRank-Based Methods

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Ordinal Test-Retest Reliability (Ordinal Test-Retest Reliability Analysis). Retrieved 2026-07-21 from https://scholargate.app/en/psychometrics/ordinal-test-retest-reliability · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Multiple contributors (Spearman, Cohen, Shrout & Fleiss)
Year
1904–1979
Type
Reliability / temporal stability
DataType
Ordinal (Likert, rating scales, ranked responses)
Subfamily
Scale / measurement
Related methods
Ordinal Reliability AnalysisTest-Retest Reliability
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account