Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Psychometrics›Differential Item Functioning (DIF)
Latent structureScale / measurement

Differential Item Functioning (DIF)

Differential Item Functioning · Also known as: DIF, item bias analysis, measurement non-equivalence, item-level measurement bias

Differential item functioning identifies test or survey items that behave differently for examinees from different groups — such as gender, ethnicity, or language background — after controlling for the underlying ability or trait being measured. DIF analysis is essential for fairness evaluation in educational testing and psychological scale development.

ScholarGate
  1. Latent structure
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Differential Item Functioning
Confirmatory factor anal…Item Response TheoryMeasurement InvarianceMulti-group confirmatory…Rasch ModelBayesian Differential It…Bayesian Item AnalysisBayesian Measurement Inv…CAT Scale DevelopmentCAT-DIF

+37 more

When to use it

Use DIF analysis whenever a test or scale will be administered to groups that may differ in cultural background, language, gender, age, or disability status, and whenever fairness or cross-group comparability is a concern. It is standard practice in high-stakes educational testing (e.g., college admissions, licensure exams) and is increasingly required for psychological scales claiming cross-cultural validity. Do not use DIF analysis as a substitute for measurement invariance testing at the factor level; the two approaches are complementary. DIF should also not be run when sample sizes in either group are very small (typically fewer than 100–200 per group for MH), as statistical power will be insufficient to detect meaningful bias.

Strengths & limitations

Strengths
  • Operates at the item level, identifying specific problematic items rather than evaluating overall scale equivalence only.
  • Multiple established methods (MH, logistic regression, IRT-based) with different assumptions, allowing cross-validation.
  • Distinguishes uniform DIF (one group consistently outperforms) from non-uniform DIF (performance gap depends on ability level), which has different fairness implications.
  • Directly supports test fairness evaluation and can guide targeted item revision rather than wholesale scale revision.
  • Integrates naturally into item response theory frameworks and large-scale test development pipelines.
Limitations
  • Requires adequate sample sizes in both focal and reference groups; small focal-group samples lead to underpowered tests.
  • Matching on total score can be contaminated if the test itself contains many biased items, leading to purification iterations.
  • Statistical DIF detection must be paired with expert content review; DIF flags are hypotheses, not verdicts about bias.
  • Non-uniform DIF is harder to detect and may require larger samples or IRT-based approaches.
  • Results are method-dependent; different procedures can disagree on which items are flagged.

Frequently asked

What is the difference between DIF and measurement invariance?

DIF operates at the individual item level and typically uses observed-score matching or IRT; measurement invariance testing (e.g., multi-group CFA) operates at the factor level and evaluates whether loadings, intercepts, and residuals are equal across groups. Both address the same underlying concern — that a scale measures the same construct equivalently across groups — but at different levels of analysis. Running both provides a more complete picture.

Which DIF method should I use?

The Mantel-Haenszel procedure is simple, well-validated, and appropriate for dichotomous items when groups are large. Logistic regression is more flexible (handles polytomous items and tests for non-uniform DIF). IRT-based methods are preferred when an IRT model is already being fitted to the data and when precise ability estimation matters. Using at least two methods and requiring convergent flagging reduces false positives.

How large do samples need to be for DIF analysis?

The reference group typically needs at least 200–500 examinees; the focal group at least 100–200, though larger is better. Very small focal groups (fewer than 50) have severely limited power to detect meaningful DIF, and results will be unstable. Sample-size requirements also increase for detecting non-uniform DIF.

What do I do after an item is flagged for DIF?

A flagged item should be reviewed by content experts who evaluate whether the item contains construct-irrelevant content that disadvantages one group. If bias is confirmed, revise or remove the item. If the DIF reflects real group differences on a secondary dimension the item legitimately taps (impact rather than bias), the item may be retained but should be treated with care in score interpretation.

Can DIF analysis be applied to Likert-scale items?

Yes. Logistic regression and polytomous IRT models (such as the graded response model or partial credit model) extend DIF analysis to ordinal polytomous items. The MH procedure in its basic form applies to dichotomous items only, but extensions for polytomous items exist.

Sources

  1. Holland, P. W. & Wainer, H. (Eds.) (1993). Differential Item Functioning. Lawrence Erlbaum Associates. ISBN: 978-0805809589
  2. Dorans, N. J. & Kulick, E. (1986). Demonstrating the utility of the standardization approach to assessing unexpected differential item performance on the Scholastic Aptitude Test. Journal of Educational Measurement, 23(4), 355–368. DOI: 10.1111/j.1745-3984.1986.tb00255.x ↗

How to cite this page

ScholarGate. (2026, June 3). Differential Item Functioning. ScholarGate. https://scholargate.app/en/psychometrics/differential-item-functioning

Related methods

Confirmatory factor analysisItem Response TheoryMeasurement InvarianceMulti-group confirmatory factor analysisRasch Model

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Confirmatory factor analysisPsychometrics↔ compare
  • Item Response TheoryPsychometrics↔ compare
  • Measurement InvariancePsychometrics↔ compare
  • Multi-group confirmatory factor analysisPsychometrics↔ compare
  • Rasch ModelPsychometrics↔ compare
Compare side by side →

Referenced by

Bayesian Differential Item FunctioningBayesian Item AnalysisBayesian Measurement InvarianceCAT Scale DevelopmentCAT-DIFComputerized adaptive test construct validityComputerized Adaptive Test Content ValidityComputerized adaptive test item analysisComputerized adaptive test measurement invarianceComputerized adaptive test Rasch modelComputerized adaptive test reliability analysisDifferential Distractor FunctioningDifferential Item Functioning in Educational TestingItem Response TheoryLongitudinal DIFLongitudinal IRTLongitudinal Item AnalysisMulti-group confirmatory factor analysisMulti-group Cronbach's alphaMulti-group Differential Item FunctioningMulti-group EFAMulti-group item analysisMulti-group item response theoryMulti-group measurement invarianceMulti-group Rasch modelMulti-group scale developmentMultilevel Differential Item FunctioningMultilevel Rasch ModelOrdinal Differential Item FunctioningOrdinal IRTOrdinal Item AnalysisOrdinal Measurement InvarianceOrdinal Rasch ModelPolytomous Construct ValidityPolytomous Measurement InvariancePolytomous Rasch ModelPolytomous scale developmentRobust Differential Item FunctioningRobust Item AnalysisRobust Rasch ModelShort form differential item functioningShort Form Measurement InvarianceShort form Rasch modelShort-Form IRT

Similar methods

Differential Item Functioning in Educational TestingRobust Differential Item FunctioningDIF AnalysisShort form differential item functioningMulti-group Differential Item FunctioningOrdinal Differential Item FunctioningBayesian Differential Item FunctioningMulti-group item analysis

Related reference concepts

Item Response TheoryPsychological Testing and PsychometricsEducational MeasurementPsychometrics & Statistics & MethodologyMeasurementCulture Fair Tests

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Differential Item Functioning (Differential Item Functioning). Retrieved 2026-07-20 from https://scholargate.app/en/psychometrics/differential-item-functioning · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
William H. Angoff and colleagues (ETS); systematized by Holland & Wainer
Year
1970s–1993
Type
Item-level bias detection
DataType
Ordinal / dichotomous item responses
Subfamily
Scale / measurement
Related methods
Confirmatory factor analysisItem Response TheoryMeasurement InvarianceMulti-group confirmatory factor analysisRasch Model
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account