Differential Item Functioning (DIF)
Differential Item Functioning · Also known as: DIF, item bias analysis, measurement non-equivalence, item-level measurement bias
Differential item functioning identifies test or survey items that behave differently for examinees from different groups — such as gender, ethnicity, or language background — after controlling for the underlying ability or trait being measured. DIF analysis is essential for fairness evaluation in educational testing and psychological scale development.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
+37 more
When to use it
Use DIF analysis whenever a test or scale will be administered to groups that may differ in cultural background, language, gender, age, or disability status, and whenever fairness or cross-group comparability is a concern. It is standard practice in high-stakes educational testing (e.g., college admissions, licensure exams) and is increasingly required for psychological scales claiming cross-cultural validity. Do not use DIF analysis as a substitute for measurement invariance testing at the factor level; the two approaches are complementary. DIF should also not be run when sample sizes in either group are very small (typically fewer than 100–200 per group for MH), as statistical power will be insufficient to detect meaningful bias.
Strengths & limitations
- Operates at the item level, identifying specific problematic items rather than evaluating overall scale equivalence only.
- Multiple established methods (MH, logistic regression, IRT-based) with different assumptions, allowing cross-validation.
- Distinguishes uniform DIF (one group consistently outperforms) from non-uniform DIF (performance gap depends on ability level), which has different fairness implications.
- Directly supports test fairness evaluation and can guide targeted item revision rather than wholesale scale revision.
- Integrates naturally into item response theory frameworks and large-scale test development pipelines.
- Requires adequate sample sizes in both focal and reference groups; small focal-group samples lead to underpowered tests.
- Matching on total score can be contaminated if the test itself contains many biased items, leading to purification iterations.
- Statistical DIF detection must be paired with expert content review; DIF flags are hypotheses, not verdicts about bias.
- Non-uniform DIF is harder to detect and may require larger samples or IRT-based approaches.
- Results are method-dependent; different procedures can disagree on which items are flagged.
Frequently asked
What is the difference between DIF and measurement invariance?
DIF operates at the individual item level and typically uses observed-score matching or IRT; measurement invariance testing (e.g., multi-group CFA) operates at the factor level and evaluates whether loadings, intercepts, and residuals are equal across groups. Both address the same underlying concern — that a scale measures the same construct equivalently across groups — but at different levels of analysis. Running both provides a more complete picture.
Which DIF method should I use?
The Mantel-Haenszel procedure is simple, well-validated, and appropriate for dichotomous items when groups are large. Logistic regression is more flexible (handles polytomous items and tests for non-uniform DIF). IRT-based methods are preferred when an IRT model is already being fitted to the data and when precise ability estimation matters. Using at least two methods and requiring convergent flagging reduces false positives.
How large do samples need to be for DIF analysis?
The reference group typically needs at least 200–500 examinees; the focal group at least 100–200, though larger is better. Very small focal groups (fewer than 50) have severely limited power to detect meaningful DIF, and results will be unstable. Sample-size requirements also increase for detecting non-uniform DIF.
What do I do after an item is flagged for DIF?
A flagged item should be reviewed by content experts who evaluate whether the item contains construct-irrelevant content that disadvantages one group. If bias is confirmed, revise or remove the item. If the DIF reflects real group differences on a secondary dimension the item legitimately taps (impact rather than bias), the item may be retained but should be treated with care in score interpretation.
Can DIF analysis be applied to Likert-scale items?
Yes. Logistic regression and polytomous IRT models (such as the graded response model or partial credit model) extend DIF analysis to ordinal polytomous items. The MH procedure in its basic form applies to dichotomous items only, but extensions for polytomous items exist.
Sources
- Holland, P. W. & Wainer, H. (Eds.) (1993). Differential Item Functioning. Lawrence Erlbaum Associates. ISBN: 978-0805809589
- Dorans, N. J. & Kulick, E. (1986). Demonstrating the utility of the standardization approach to assessing unexpected differential item performance on the Scholastic Aptitude Test. Journal of Educational Measurement, 23(4), 355–368. DOI: 10.1111/j.1745-3984.1986.tb00255.x ↗
How to cite this page
ScholarGate. (2026, June 3). Differential Item Functioning. ScholarGate. https://scholargate.app/en/psychometrics/differential-item-functioning
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Confirmatory factor analysisPsychometrics↔ compare
- Item Response TheoryPsychometrics↔ compare
- Measurement InvariancePsychometrics↔ compare
- Multi-group confirmatory factor analysisPsychometrics↔ compare
- Rasch ModelPsychometrics↔ compare