Latent structurePsychometricsScale / measurementModel

Computerized Adaptive Testing based on Item Response Theory (CAT-IRT)

Also known as: CAT-IRT, adaptive testing, IRT-based CAT, computerized adaptive testing

OriginatorLord, F. M.; further developed by Wainer, van der Linden, and othersYear1970s–1980sSources2Related methods14

Computerized adaptive testing based on item response theory is a sequential measurement procedure in which a computer algorithm selects successive test items tailored to each examinee's estimated ability level. Drawing on IRT to model item characteristics and ability estimation, CAT delivers precise scores with far fewer items than fixed-length tests, making it efficient for high-stakes assessments, clinical screening, and large-scale surveys.

Key highlights

  • Substantially reduces test length (often 40–60%) without sacrificing measurement precision, lowering examinee fatigue and administration time.
  • Provides item-level standard errors conditional on ability, enabling adaptive stopping rules calibrated to a precision target rather than a fixed item count.
  • Enables real-time score reporting immediately after test completion because scoring does not require group-level norming.
  • The difficulty of presented items is automatically tailored to the examinee, reducing floor and ceiling effects common in fixed tests.
  • Item exposure can be controlled algorithmically, protecting item-bank security across administrations.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

CAT-IRT is most valuable when test efficiency matters — when administering full fixed-form tests to every examinee is impractical, costly, or burdensome. It is appropriate when a well-calibrated item bank is available (typically requiring IRT calibration on several hundred examinees per item), when items can be presented and scored one-at-a-time on a computer, and when the construct is unidimensional or can be decomposed into unidimensional subscales. Avoid CAT when the item bank is small (fewer than 50–100 well-calibrated items), when the construct is strongly multidimensional and a single adaptive pass is insufficient, or when the testing context requires all examinees to see the same items for transparency or legal defensibility reasons.

Strengths & limitations

Strengths
  • Substantially reduces test length (often 40–60%) without sacrificing measurement precision, lowering examinee fatigue and administration time.
  • Provides item-level standard errors conditional on ability, enabling adaptive stopping rules calibrated to a precision target rather than a fixed item count.
  • Enables real-time score reporting immediately after test completion because scoring does not require group-level norming.
  • The difficulty of presented items is automatically tailored to the examinee, reducing floor and ceiling effects common in fixed tests.
  • Item exposure can be controlled algorithmically, protecting item-bank security across administrations.
Limitations
  • Requires a large, carefully calibrated item bank as a prerequisite; item calibration is a substantial methodological and logistical investment.
  • IRT calibration assumes local independence and a specified dimensionality; violations can bias ability estimates and information calculations.
  • Fixed-form score comparisons are complicated: different examinees see different items, which can raise fairness questions and complicate equating.
  • Multidimensional constructs require either unidimensional approximations or more complex multidimensional CAT algorithms.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

How large does the item bank need to be for a CAT?

A common guideline is at least five to ten items per desired test length, but operational programs typically maintain banks of 200–1000 items to support content balancing and exposure control. Smaller banks increase item overexposure and reduce the gains from adaptivity.

Can CAT work with polytomous items (e.g., Likert scales)?

Yes. Polytomous IRT models such as the Graded Response Model (Samejima, 1969) or the Partial Credit Model (Masters, 1982) extend CAT to ordered-category items, which are common in personality and clinical scales.

Is a CAT score comparable to a fixed-form score on the same test?

Because both are expressed on the same IRT theta scale and conditioned on the same item parameters, CAT and fixed-form scores are theoretically on the same scale. In practice, equating studies are recommended before mixing scores from different administration modes.

What is the difference between CAT and tailored testing?

The terms are often used interchangeably, but 'tailored testing' is an older, broader term covering any strategy that adjusts item selection to the examinee, including two-stage and branched formats. Modern CAT specifically refers to item-by-item computerized selection using IRT-based item information.

Does CAT require unidimensionality?

Standard CAT algorithms assume a single latent dimension. Multidimensional CAT (MCAT) algorithms exist and can adaptively estimate a vector of abilities, but they are considerably more complex and require a multidimensional item bank and calibration.

Sources

  1. 1.
    Wainer, H. (Ed.). (2000). Computerized Adaptive Testing: A Primer (2nd ed.). Lawrence Erlbaum Associates.
    ISBN 978-0805835113
  2. 2.
    van der Linden, W. J., & Glas, C. A. W. (Eds.). (2010). Elements of Adaptive Testing. Springer.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 3). Computerized adaptive test item response theory. ScholarGate. https://scholargate.app/psychometrics/computerized-adaptive-test-item-response-theory

Computerized Adaptive Testing based on Item Response Theory (CAT-IRT) | ScholarGate