Differential Item Functioning (DIF) Analysis
Differential Item Functioning Analysis · Also known as: Madde Yanlılık Analizi (DIF — Differential Item Functioning), item bias analysis, Mantel-Haenszel DIF, Lord chi-square DIF, logistic regression DIF
Differential Item Functioning analysis examines whether examinees from different groups — such as gender, ethnicity, or language background — who have the same underlying ability respond differently to a test item. First formalised by Holland and Thayer in 1988 via the Mantel-Haenszel procedure, it is the principal tool in modern test development for detecting and removing item bias.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
DIF analysis is appropriate whenever a test or scale is intended to measure the same construct fairly across two or more distinct groups. The reference group (typically the majority or standardisation group) and the focal group (the group being evaluated for disadvantage) must be defined in advance. Ability matching relies on the total score or an IRT-estimated theta, so the test as a whole must provide a meaningful ordering of examinees. A minimum sample of roughly 200 examinees in each group is needed for stable Mantel-Haenszel estimates; smaller samples make DIF detection unreliable. Items must be scored as binary or ordinal responses, and at least five items should be present for the analysis to have adequate power. Multiple-comparison correction (Bonferroni) must be applied when testing many items simultaneously.
Strengths & limitations
- Provides a principled, statistically grounded method for identifying items that function inequitably across groups, supporting fairness in high-stakes testing.
- The Mantel-Haenszel procedure is non-parametric with respect to the ability distribution, making it robust when IRT model fit is uncertain.
- The ETS A/B/C classification translates statistical results into actionable decisions for test developers without requiring deep psychometric expertise.
- Logistic regression DIF detects both uniform and non-uniform DIF, providing a more complete picture of item behaviour than the Mantel-Haenszel odds ratio alone.
- DIF detection is purely statistical: a flagged item may show differential functioning for defensible content reasons rather than bias, so every flagged item requires substantive expert review before it is removed.
- Small samples in either group lead to low power and unstable estimates; a minimum of about 200 per group is widely recommended.
- The analysis requires a purified matching criterion — if many items have DIF, the total score used for matching is itself biased, distorting the DIF estimates for the remaining items.
- DIF analysis identifies group-level patterns; it says nothing about fairness for individual examinees.
Frequently asked
What is the difference between DIF and item bias?
DIF is a statistical finding: an item shows differential functioning when examinees of equal ability from different groups respond to it at different rates. Item bias is a substantive judgment: an item is biased when that differential functioning cannot be justified by the construct being measured. Every biased item has DIF, but not every DIF item is biased — differential functioning may reflect a legitimate multidimensional feature of the construct. Human expert review of flagged items is therefore essential.
Which DIF method should I use — Mantel-Haenszel, logistic regression, or Lord's chi-square?
Mantel-Haenszel is the most established and computationally simple method; it detects uniform DIF reliably and its delta-scale output maps directly to the ETS A/B/C classification used operationally. Logistic regression is preferred when you want to detect non-uniform DIF as well, or when you wish to include covariates such as language status alongside group membership. Lord's chi-square requires fitting an IRT model and is appropriate when IRT parameter estimates are already available and model fit is adequate. In practice, running Mantel-Haenszel and logistic regression together and comparing their flags is a robust strategy.
How large a sample do I need for DIF analysis?
The Mantel-Haenszel procedure and logistic regression DIF generally require at least 200 examinees in each group to achieve adequate statistical power for detecting moderate DIF (ETS class B). Smaller focal groups, such as 100 or fewer, result in unstable estimates and high rates of missed DIF. If your focal group is too small, consider descriptive item statistics as a preliminary diagnostic rather than formal DIF testing.
What does ETS A/B/C DIF classification mean?
The Educational Testing Service classification translates the Mantel-Haenszel log-odds ratio into an absolute delta-scale difference. Class A (negligible DIF) means the delta difference is less than 1.0, or the result is not statistically significant; items are generally retained without review. Class B (moderate DIF) covers delta differences between 1.0 and 1.5 with statistical significance; items are flagged for substantive expert review. Class C (large DIF) means a delta difference exceeding 1.5 with significance; items are typically eliminated or substantially revised before use in high-stakes contexts.
Sources
- Holland, P. W. & Thayer, D. T. (1988). Differential Item Performance and the Mantel-Haenszel Procedure. ETS Research Report Series. link ↗
- Magis, D., Beland, S., Tuerlinckx, F. & De Boeck, P. (2010). A General Framework and an R Package for the Detection of Dichotomous Differential Item Functioning. Behavior Research Methods, 42(3), 847–862. DOI: 10.3758/BRM.42.3.847 ↗
How to cite this page
ScholarGate. (2026, June 1). Differential Item Functioning Analysis. ScholarGate. https://scholargate.app/en/psychometrics/dif-analysis
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Confirmatory factor analysisPsychometrics↔ compare
- EFAStatistics↔ compare
- Item AnalysisPsychometrics↔ compare
- Measurement InvariancePsychometrics↔ compare