Convergent Validity for Computerized Adaptive Tests
Convergent Validity Assessment for Computerized Adaptive Tests · Also known as: CAT convergent validity, adaptive test construct validation, CAT validity evidence, convergent validity in CAT
Convergent validity assessment for computerized adaptive tests (CATs) examines whether the ability or trait estimates produced by an adaptive algorithm correlate substantially with scores from other measures of the same construct. Because each examinee receives a different subset of items in a CAT, demonstrating that the resulting scores still converge with theoretically related external measures is a critical step in establishing construct validity evidence.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use this approach when a newly developed or adapted CAT instrument needs formal validity evidence before operational deployment, or when an existing CAT is being compared against a conventional fixed-form counterpart. It is especially important when the CAT covers a domain differently from legacy instruments — for example, using fewer items or a different response format — and stakeholders need assurance that scores are interchangeable or interpretable in the same construct framework. Do not rely solely on convergent correlations to establish validity; they are one source of evidence and must be supplemented with content validity evaluation, fit of the IRT model, and practical utility analyses. Also avoid this approach as the only check when the comparison criterion is itself of uncertain validity.
Strengths & limitations
- Directly addresses the construct interpretation of CAT scores in a way that IRT model-fit statistics alone cannot.
- The MTMM design simultaneously provides discriminant evidence, strengthening the overall validity argument.
- Can be conducted with the same validation sample used for IRT calibration if a parallel fixed-form is administered concurrently.
- Correlations are interpretable and communicable to practitioners and policy audiences unfamiliar with IRT.
- Attenuation corrections allow estimation of the true relationship between constructs, separating reliability shortfalls from true divergence.
- Requires a criterion measure of the same construct that is itself sufficiently reliable and valid, which is not always available.
- Convergent correlations are attenuated by measurement error in both the CAT and the criterion; uncorrected values can understate the true relationship.
- High convergent correlations with a flawed criterion instrument can provide false assurance about the CAT's validity.
- The method does not evaluate whether the CAT measures the construct with equal fidelity across different subgroups (that requires measurement invariance or DIF analysis).
Frequently asked
Why is convergent validity particularly important for CATs compared with fixed-form tests?
Because each examinee answers a different set of items, the content coverage of the CAT varies across people. Critics may question whether scores from variable-length, variable-content batteries reflect the same construct as scores from a standard instrument. Demonstrating convergence with an external criterion directly addresses this concern.
What correlation magnitude is sufficient for convergent validity?
There is no universal threshold, but correlations above approximately .50 (corrected for attenuation if possible) are frequently cited as supportive of convergent validity in psychological and educational measurement. The adequacy of any coefficient depends on the method similarity of the two instruments, the reliability of the criterion, and the theoretical closeness of the constructs.
Should I correct the convergent correlation for attenuation?
Reporting both the raw and attenuation-corrected coefficient is recommended. The corrected value estimates the true construct-level relationship; the uncorrected value reflects operational comparability. Relying only on the corrected value can obscure practical score-level discrepancies.
Can I use IRT model-fit statistics instead of convergent validity correlations?
IRT model fit evaluates internal consistency of item responses with the model, not whether the trait being measured matches an external construct of interest. Convergent validity and IRT fit address different questions and both are necessary components of a complete validity argument.
Do I need a separate validation sample, or can I use the calibration sample?
Using the same sample for IRT calibration and validity estimation is acceptable if the convergent measure was collected independently and the sample is sufficiently large, but cross-validation on a new sample is preferable. Mixing calibration and validation data can produce optimistically biased validity estimates.
Sources
- Wainer, H. (Ed.). (2000). Computerized Adaptive Testing: A Primer (2nd ed.). Lawrence Erlbaum Associates. ISBN: 978-0805835113
- Messick, S. (1989). Validity. In R. L. Linn (Ed.), Educational Measurement (3rd ed., pp. 13–103). American Council on Education / Macmillan. link ↗
How to cite this page
ScholarGate. (2026, June 3). Convergent Validity Assessment for Computerized Adaptive Tests. ScholarGate. https://scholargate.app/en/psychometrics/computerized-adaptive-test-convergent-validity
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Computerized adaptive test construct validityPsychometrics↔ compare
- Computerized adaptive test item response theoryPsychometrics↔ compare
- Convergent ValidityPsychometrics↔ compare
- Discriminant ValidityPsychometrics↔ compare
- Item Response TheoryPsychometrics↔ compare