Latent structurePsychometricsScale / measurementModel

Differential Item Functioning (DIF)

Also known as: DIF, item bias analysis, measurement non-equivalence, item-level measurement bias

OriginatorWilliam H. Angoff and colleagues (ETS); systematized by Holland & WainerYear1970s–1993Sources2Related methods49

Differential item functioning identifies test or survey items that behave differently for examinees from different groups — such as gender, ethnicity, or language background — after controlling for the underlying ability or trait being measured. DIF analysis is essential for fairness evaluation in educational testing and psychological scale development.

Key highlights

  • Operates at the item level, identifying specific problematic items rather than evaluating overall scale equivalence only.
  • Multiple established methods (MH, logistic regression, IRT-based) with different assumptions, allowing cross-validation.
  • Distinguishes uniform DIF (one group consistently outperforms) from non-uniform DIF (performance gap depends on ability level), which has different fairness implications.
  • Directly supports test fairness evaluation and can guide targeted item revision rather than wholesale scale revision.
  • Integrates naturally into item response theory frameworks and large-scale test development pipelines.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use DIF analysis whenever a test or scale will be administered to groups that may differ in cultural background, language, gender, age, or disability status, and whenever fairness or cross-group comparability is a concern. It is standard practice in high-stakes educational testing (e.g., college admissions, licensure exams) and is increasingly required for psychological scales claiming cross-cultural validity. Do not use DIF analysis as a substitute for measurement invariance testing at the factor level; the two approaches are complementary. DIF should also not be run when sample sizes in either group are very small (typically fewer than 100–200 per group for MH), as statistical power will be insufficient to detect meaningful bias.

Strengths & limitations

Strengths
  • Operates at the item level, identifying specific problematic items rather than evaluating overall scale equivalence only.
  • Multiple established methods (MH, logistic regression, IRT-based) with different assumptions, allowing cross-validation.
  • Distinguishes uniform DIF (one group consistently outperforms) from non-uniform DIF (performance gap depends on ability level), which has different fairness implications.
  • Directly supports test fairness evaluation and can guide targeted item revision rather than wholesale scale revision.
  • Integrates naturally into item response theory frameworks and large-scale test development pipelines.
Limitations
  • Requires adequate sample sizes in both focal and reference groups; small focal-group samples lead to underpowered tests.
  • Matching on total score can be contaminated if the test itself contains many biased items, leading to purification iterations.
  • Statistical DIF detection must be paired with expert content review; DIF flags are hypotheses, not verdicts about bias.
  • Non-uniform DIF is harder to detect and may require larger samples or IRT-based approaches.
  • Results are method-dependent; different procedures can disagree on which items are flagged.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

What is the difference between DIF and measurement invariance?

DIF operates at the individual item level and typically uses observed-score matching or IRT; measurement invariance testing (e.g., multi-group CFA) operates at the factor level and evaluates whether loadings, intercepts, and residuals are equal across groups. Both address the same underlying concern — that a scale measures the same construct equivalently across groups — but at different levels of analysis. Running both provides a more complete picture.

Which DIF method should I use?

The Mantel-Haenszel procedure is simple, well-validated, and appropriate for dichotomous items when groups are large. Logistic regression is more flexible (handles polytomous items and tests for non-uniform DIF). IRT-based methods are preferred when an IRT model is already being fitted to the data and when precise ability estimation matters. Using at least two methods and requiring convergent flagging reduces false positives.

How large do samples need to be for DIF analysis?

The reference group typically needs at least 200–500 examinees; the focal group at least 100–200, though larger is better. Very small focal groups (fewer than 50) have severely limited power to detect meaningful DIF, and results will be unstable. Sample-size requirements also increase for detecting non-uniform DIF.

What do I do after an item is flagged for DIF?

A flagged item should be reviewed by content experts who evaluate whether the item contains construct-irrelevant content that disadvantages one group. If bias is confirmed, revise or remove the item. If the DIF reflects real group differences on a secondary dimension the item legitimately taps (impact rather than bias), the item may be retained but should be treated with care in score interpretation.

Can DIF analysis be applied to Likert-scale items?

Yes. Logistic regression and polytomous IRT models (such as the graded response model or partial credit model) extend DIF analysis to ordinal polytomous items. The MH procedure in its basic form applies to dichotomous items only, but extensions for polytomous items exist.

Sources

  1. 1.
    Holland, P. W. & Wainer, H. (Eds.) (1993). Differential Item Functioning. Lawrence Erlbaum Associates.
    ISBN 978-0805809589
  2. 2.
    Dorans, N. J. & Kulick, E. (1986). Demonstrating the utility of the standardization approach to assessing unexpected differential item performance on the Scholastic Aptitude Test. Journal of Educational Measurement, 23(4), 355–368.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 3). Differential Item Functioning. ScholarGate. https://scholargate.app/psychometrics/differential-item-functioning