Computerized Adaptive Test Differential Item Functioning (CAT-DIF)
Computerized Adaptive Test Differential Item Functioning · Also known as: CAT DIF analysis, adaptive test DIF, DIF in computerized adaptive testing, CAT item bias detection
CAT-DIF identifies items in a computerized adaptive test that behave differently across demographic or group subpopulations after controlling for overall ability. Because adaptive algorithms select items non-randomly based on each examinee's estimated proficiency, standard DIF detection methods require adjustment before they can be validly applied in this context.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use CAT-DIF when you administer a large-scale computerized adaptive test to groups differing by gender, ethnicity, language background, or other demographic characteristics, and you need to ensure that item performance differences reflect true ability differences rather than item bias. It is especially important when high-stakes decisions — such as licensing, certification, college admission, or clinical assessment — are based on CAT scores. Do not apply standard fixed-form DIF procedures (e.g., Mantel–Haenszel based on number-correct scores) directly to CAT data without exposure correction, as the non-random item selection violates the assumptions of those methods. CAT-DIF is not needed when all examinees receive identical item sets, or when sample sizes within groups are too small to power DIF detection reliably (typically fewer than 200 per group).
Strengths & limitations
- Maintains measurement fairness in high-stakes adaptive assessments by detecting construct-irrelevant group differences at the item level.
- IRT-based ability conditioning is more precise than observed-score matching, reducing ability-confounding in DIF estimates.
- Applicable to large operational CAT banks with thousands of items, supporting systematic fairness review.
- Both statistical significance and effect-size criteria provide a two-stage filter that reduces false-positive DIF flags.
- Simulation-based exposure-correction methods can replicate realistic adaptive administration conditions without requiring new test administrations.
- Differential item exposure across groups complicates DIF estimation; ignoring exposure rates inflates Type I error.
- Large group sample sizes (typically 200 or more per group) are required for stable DIF statistics, which may be unavailable in small-scale CAT deployments.
- IRT model misfit at the bank level can propagate into misleading DIF estimates if item parameters are poorly calibrated.
- Content review to distinguish construct-irrelevant bias from legitimate group differences requires human expertise and is time-consuming.
- Removing DIF items from the bank may reduce item pool depth, limiting the CAT's ability to measure accurately at certain ability levels.
Frequently asked
Why can't I just use Mantel–Haenszel DIF on my CAT data?
Mantel–Haenszel matches examinees on their observed number-correct score and then compares group performance on each item. In a CAT, different examinees see different items, so the number-correct score is not comparable across examinees or groups. You must use IRT-based DIF statistics with ability estimates as the conditioning variable, and you must correct for differential item-exposure rates before comparing groups.
How large a sample do I need for CAT-DIF analysis?
A common minimum is around 200 examinees per group (focal and reference) for each item being evaluated. Because individual CAT items are not seen by all examinees, you may need a substantially larger total sample to accumulate sufficient responses per item per group, particularly for less frequently exposed items in the bank.
What do I do with items flagged for DIF?
Statistical flagging is the first step, not the final decision. Flagged items should be reviewed by content experts who assess whether the differential performance reflects a construct-irrelevant feature — such as culturally specific content or ambiguous wording — or a legitimate group difference on a related construct. Only items confirmed as biased should be removed or revised.
Can DIF analysis be conducted during test development versus operationally?
Both are important. Pre-operational DIF studies during item tryout help screen out biased items before they enter the live bank. Ongoing operational DIF monitoring is also necessary because item exposure patterns and examinee populations change over time, and items that appeared unbiased initially may show DIF as the test program matures.
What is the difference between DIF and differential test functioning in a CAT?
DIF is an item-level concept: it asks whether a single item performs differently across groups after ability matching. Differential test functioning (DTF) aggregates DIF across all items an examinee encounters, capturing whether the overall adaptive test systematically over- or underestimates ability for a group. In a CAT, DTF can emerge even when no single item shows large DIF, if small DIF effects accumulate across the adaptive path.
Sources
- Zwick, R., Thayer, D. T., & Mazzeo, J. (1997). Describing and categorizing DIF in polytomous items. Journal of Educational Measurement, 34(4), 261–285. DOI: 10.1002/j.2333-8504.1997.tb01726.x ↗
- Wainer, H. (Ed.). (2000). Computerized Adaptive Testing: A Primer (2nd ed.). Lawrence Erlbaum Associates. ISBN: 978-0805835113
How to cite this page
ScholarGate. (2026, June 3). Computerized Adaptive Test Differential Item Functioning. ScholarGate. https://scholargate.app/en/psychometrics/computerized-adaptive-test-differential-item-functioning
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Computerized adaptive test item analysisPsychometrics↔ compare
- Computerized adaptive test item response theoryPsychometrics↔ compare
- Computerized adaptive test measurement invariancePsychometrics↔ compare
- Differential Item FunctioningPsychometrics↔ compare
- Item Response TheoryPsychometrics↔ compare
- Multi-group Differential Item FunctioningPsychometrics↔ compare