Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Psychometrics›Computerized Adaptive Test Test-Retest Reliability
Latent structureScale / measurement

Computerized Adaptive Test Test-Retest Reliability

Also known as: CAT temporal stability, adaptive test retest reliability, CAT score consistency, computerized adaptive testing reliability

Computerized adaptive test (CAT) test-retest reliability quantifies the consistency of ability estimates obtained when the same examinees complete a CAT on two separate occasions. Because adaptive algorithms tailor each examinee's item set individually, traditional reliability frameworks must be adapted to account for non-overlapping item exposures across administrations.

ScholarGate
  1. Latent structure
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

CAT Test-Retest Reliability
Computerized Adaptive Te…Cronbach's AlphaItem Response TheoryTest-Retest Reliability

When to use it

Use CAT test-retest reliability when you need to demonstrate that an adaptive instrument yields stable scores across time, particularly before deploying a CAT in clinical screening, licensure, or educational placement contexts. It is especially appropriate when the item bank is large enough to limit item overlap between administrations and when a plausible stability interval can be justified for the target construct. Do not use test-retest as the sole reliability evidence when the measured trait is known to fluctuate rapidly over short periods — mood states, acute symptoms — as genuine construct change will suppress the coefficient and lead to misleading conclusions about instrument quality.

Strengths & limitations

Strengths
  • Provides direct evidence of score stability across time, which is the form of reliability most relevant to decisions made from a single CAT administration.
  • Accounts for the unique psychometric challenge of adaptive testing by working with theta estimates rather than raw scores, preserving the metric established by item response theory calibration.
  • When combined with the conditional standard error of measurement, disattenuation correction allows separation of true score variance from measurement error variance.
  • Widely understood by practitioners and acceptable to credentialing bodies and clinical regulators who require evidence of temporal consistency.
Limitations
  • Requires a second full administration, making data collection resource-intensive and dependent on participant retention.
  • Cannot distinguish true score change from measurement unreliability without an additional control or a longitudinal design with anchor points.
  • The choice of time interval between occasions is arbitrary and the optimal interval varies by construct, making comparisons across studies difficult.
  • Correlating theta estimates ignores the heterogeneity in precision across examinees that the adaptive design deliberately creates.

Frequently asked

Why can't I just use Cronbach's alpha or KR-20 to estimate CAT reliability?

Cronbach's alpha and KR-20 measure internal consistency within a single administration. In a CAT different examinees receive different items, so there is no single fixed item set whose internal consistency can be calculated across the full sample in the conventional way. You can compute internal consistency within each adaptive administration using the item information contributed by each selected item, but this does not capture temporal stability. Test-retest correlation is necessary to assess score consistency across occasions.

How large a sample do I need for a stable test-retest reliability estimate?

For a Pearson correlation, a sample of at least 100 examinees is commonly recommended to achieve acceptable precision. With fewer than 50 participants the 95% confidence interval around a correlation of 0.85 can span more than 0.20 points, making the estimate difficult to interpret. Power analysis for correlation testing can inform the required N for a given desired precision.

What is an appropriate retest interval for CAT reliability studies?

The interval should be long enough for memory effects on specific items to dissipate — typically at least two weeks — but short enough that true change in the measured construct is unlikely. For stable traits such as general cognitive ability or personality dimensions, intervals of two to four weeks are standard. For constructs sensitive to treatment or environmental change, a shorter interval or additional design controls are needed.

Does disattenuation correction always improve the estimate?

Disattenuation removes the attenuation due to measurement error in both theta estimates and therefore raises the reliability coefficient toward the true-score correlation. However, the correction amplifies any sampling error in the conditional standard errors of measurement, and in small samples the corrected coefficient can exceed 1.0 or become unstable. Report both the observed and corrected correlations with their confidence intervals.

Is test-retest reliability the same as agreement?

No. Pearson or Spearman correlation captures rank-order consistency but is insensitive to systematic mean shifts between occasions. If theta scores drift upward for all examinees between administrations — perhaps due to a practice effect or seasonal variation — the correlation can remain high while scores are systematically biased. Supplement the correlation with a paired-samples test for mean differences and, for clinical contexts, an intraclass correlation coefficient that is sensitive to both rank-order and absolute agreement.

Sources

  1. Weiss, D. J. (2004). Computerized adaptive testing for effective and efficient measurement in counseling and education. Measurement and Evaluation in Counseling and Development, 37(2), 70–84. DOI: 10.1080/07481756.2004.11909751 ↗
  2. Computerized adaptive testing. Wikipedia. link ↗

How to cite this page

ScholarGate. (2026, June 3). Computerized Adaptive Test Test-Retest Reliability. ScholarGate. https://scholargate.app/en/psychometrics/computerized-adaptive-test-test-retest-reliability

Related methods

Computerized Adaptive TestingCronbach's AlphaItem Response TheoryTest-Retest Reliability

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Computerized Adaptive TestingPsychometrics↔ compare
  • Cronbach's AlphaStatistics↔ compare
  • Item Response TheoryPsychometrics↔ compare
  • Test-Retest ReliabilityPsychometrics↔ compare
Compare side by side →

Similar methods

Computerized adaptive test reliability analysisCAT Cronbach's AlphaTest-Retest ReliabilityCAT Generalizability TheoryComputerized adaptive test construct validityComputerized Adaptive Test Convergent ValidityShort-form test-retest reliabilityComputerized adaptive test item analysis

Related reference concepts

Item Response TheoryPsychological Testing and PsychometricsPsychometrics & Statistics & MethodologyAdaptive TestingMeasurement Validity and ReliabilityTest Reliability

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — CAT Test-Retest Reliability (Computerized Adaptive Test Test-Retest Reliability). Retrieved 2026-07-21 from https://scholargate.app/en/psychometrics/computerized-adaptive-test-test-retest-reliability · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
David J. Weiss and colleagues (adaptive testing reliability literature)
Year
1970s–1980s
Type
Reliability estimation
DataType
Latent trait scores from adaptive item administrations
Subfamily
Scale / measurement
Related methods
Computerized Adaptive TestingCronbach's AlphaItem Response TheoryTest-Retest Reliability
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account