Discriminant Validity in Computerized Adaptive Testing (CAT)
Discriminant Validity in Computerized Adaptive Testing · Also known as: CAT discriminant validity, adaptive test divergent validity, CAT scale differentiation, CAT construct separation
Discriminant validity in computerized adaptive testing (CAT) is the evaluation process confirming that a CAT-administered scale measures its intended construct distinctly from related but conceptually different constructs. Despite the adaptive item-selection mechanism varying each respondent's item set, evidence must be provided that CAT-derived scores do not overlap excessively with scores from theoretically distinct scales.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use discriminant validity evaluation whenever a CAT instrument is being validated and the target construct is conceptually adjacent to other measured traits (e.g., anxiety near depression, self-efficacy near motivation, work engagement near job satisfaction). It is essential in scale development studies where the CAT is proposed as a distinct measure alongside existing instruments. Do not apply this procedure when the CAT bank is brand-new and no suitable criterion scales exist yet — in that case, content validity and convergent validity should be established first. Also avoid treating a modest correlation between distinct constructs as discriminant validity failure; theoretical overlap between adjacent constructs is expected and should be acknowledged.
Strengths & limitations
- Directly addresses the differentiation of CAT-derived construct scores from related measures, supporting the interpretive value of adaptive scores.
- IRT-based theta estimates are on a common scale regardless of which items were administered, making cross-respondent comparisons meaningful.
- Can be embedded in established frameworks (MTMM, HTMT, AVE) that are familiar to psychometric reviewers.
- Adaptive administration reduces respondent burden while still yielding scores precise enough for validity analysis.
- Applicable across diverse CAT contexts: educational achievement, clinical symptom assessment, organizational behavior scales.
- Requires an external criterion sample measured on both the CAT and comparison scales simultaneously, which may be logistically demanding.
- Theta estimation precision depends on item bank quality; a sparse or poorly calibrated bank inflates measurement error and attenuates validity correlations artificially.
- HTMT and AVE criteria were developed for reflective CFA models; their application to IRT-based CAT scores requires careful justification.
- Variable item exposure across respondents can complicate standard error estimation in correlation analyses.
- When constructs are genuinely correlated in the population, achieving discriminant validity benchmarks (HTMT < 0.85) may be unrealistically strict.
Frequently asked
Why can I not just correlate raw CAT scores with criterion scores to assess discriminant validity?
Because different respondents answer different items in a CAT, raw scores are not directly comparable across individuals. IRT-based theta estimates place all respondents on the same latent trait scale regardless of which specific items they answered, making them the appropriate scores for validity correlation analyses.
What HTMT threshold should I use for CAT-based discriminant validity?
The commonly cited thresholds are HTMT < 0.85 (conservative) or HTMT < 0.90 (liberal). However, for constructs that are theoretically expected to correlate — such as anxiety and general distress — a threshold closer to 0.90 is more defensible. The choice should be theoretically motivated and reported transparently.
Is discriminant validity different from differential item functioning (DIF)?
Yes, they address entirely different questions. Discriminant validity asks whether the total CAT score (theta) distinguishes one construct from another at the scale level. DIF asks whether individual items within the CAT bank perform differently for subgroups (e.g., gender, ethnicity) after controlling for the latent trait, which is a fairness and item-level concern.
Do I need a fixed-form comparison scale to evaluate discriminant validity of a CAT?
Not necessarily. Discriminant validity can be assessed against another CAT scale, a fixed-form questionnaire, or observer-rated criterion measures. What matters is that the comparison instrument provides a valid score for a theoretically distinct construct, regardless of its format.
Can I assess discriminant validity during CAT item bank development, or only after the bank is finalized?
Formally, discriminant validity is assessed after the bank is calibrated and the CAT algorithm is operational, because you need actual theta estimates from adaptive administration. However, preliminary item-level analyses (e.g., factor structure of the full bank) during development can flag potential construct overlap early.
Sources
- Weiss, D. J. (2004). Computerized adaptive testing for effective and efficient measurement in counseling and education. Measurement and Evaluation in Counseling and Development, 37(2), 70–84. DOI: 10.1080/07481756.2004.11909751 ↗
- Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81–105. DOI: 10.1037/h0046016 ↗
How to cite this page
ScholarGate. (2026, June 3). Discriminant Validity in Computerized Adaptive Testing. ScholarGate. https://scholargate.app/en/psychometrics/computerized-adaptive-test-discriminant-validity
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Computerized adaptive test construct validityPsychometrics↔ compare
- Computerized Adaptive Test Convergent ValidityPsychometrics↔ compare
- Convergent ValidityPsychometrics↔ compare
- Discriminant ValidityPsychometrics↔ compare
- Item Response TheoryPsychometrics↔ compare