Multi-Group Item Response Theory (MG-IRT)
Multi-Group Item Response Theory · Also known as: MG-IRT, multiple-group IRT, multi-group latent trait model, IRT across groups
Multi-group item response theory fits IRT models simultaneously across two or more defined groups — such as males and females, or different cultural samples — to determine whether item parameters are invariant across those groups. It is the primary IRT-based framework for testing measurement equivalence and detecting differential item functioning (DIF) at the model level.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use multi-group IRT when you have a test or scale already calibrated under IRT assumptions and you need to assess whether it measures the same construct with the same item characteristics across distinct groups — for instance, in cross-cultural validation, test fairness auditing, or clinical-vs-normative comparisons. At least two clearly defined groups with adequate sample sizes per group (commonly n >= 200 per group for stable 2PL estimation) are required, along with sufficient items to identify anchor sets. Do not use MG-IRT when groups have very small samples (< 100), when the overall unidimensionality assumption is violated across groups, or when only a handful of items are available and purification is impractical.
Strengths & limitations
- Provides item-level diagnostic information about DIF, identifying precisely which items and which parameters (difficulty vs. discrimination) are non-invariant.
- Handles both dichotomous and polytomous item formats through appropriate model selection (2PL, GRM, PCM).
- Accounts for differences in group trait distributions separately from item properties, allowing fair latent mean comparisons after anchor-based linking.
- Likelihood-ratio tests supply formal statistical inference with known distributional properties, complemented by effect-size measures such as signed/unsigned area between ICCs.
- Widely implemented in software (Mplus, R mirt package, flexMIRT, IRTPRO), making it practically accessible.
- Requires large samples per group (often >= 200) for stable parameter estimation, especially for the 2PL or 3PL; smaller samples produce imprecise item parameter estimates.
- Relies on the correct IRT model being selected; misfit of the underlying model (e.g., multidimensionality) can spuriously inflate DIF detection rates.
- The purification process for selecting anchor items is iterative and sensitive to starting assumptions; a contaminated anchor set can lead to incorrect linking and biased DIF conclusions.
- Interpreting non-uniform DIF (non-invariant discrimination) is more complex and less common in applied practice than uniform DIF.
- Computationally demanding compared to classical DIF methods such as Mantel-Haenszel, particularly for polytomous models with many items.
Frequently asked
How does MG-IRT differ from the Mantel-Haenszel method for detecting DIF?
Mantel-Haenszel is a classical, non-parametric contingency-table approach that conditions on total score as a proxy for the latent trait and detects uniform DIF only. MG-IRT directly models the latent trait with item parameters, detects both uniform and non-uniform DIF, and provides a fuller characterization of item functioning across the trait range — but requires larger samples and correct model specification.
How many anchor items are needed for linking in MG-IRT?
A common guideline is to use at least 20% of the items — or a minimum of four to six items — as anchors, chosen because they are believed to be DIF-free. Purification then iteratively refines this anchor set by removing items that show DIF in the initial calibration.
Can MG-IRT be applied to polytomous scales such as Likert items?
Yes. For ordered polytomous responses, the graded response model (GRM) or the partial credit model (PCM) replaces the dichotomous 2PL. The likelihood-ratio DIF test then compares category threshold parameters and, for the GRM, discrimination parameters across groups.
What sample size is required per group?
For the 2PL model a common recommendation is n >= 200 per group for stable discrimination and difficulty estimates. The 1PL (Rasch) requires smaller samples, sometimes n >= 100, while the 3PL typically needs n >= 500 per group. Polytomous models generally need larger samples because more parameters must be estimated per item.
What is the difference between MG-IRT and multi-group CFA for testing measurement invariance?
Both test whether item parameters are equivalent across groups, but they differ in framework. Multi-group CFA works in the covariance structure tradition, assuming (approximately) continuous indicators, and tests configural, metric, and scalar invariance in terms of loadings and intercepts. MG-IRT works directly with the item response probability as a function of the latent trait, accommodates binary and polytomous responses naturally, and provides item characteristic curves as diagnostic tools. For binary items, the two frameworks are closely related but not identical in their parameterization.
Sources
- Embretson, S. E. & Reise, S. P. (2000). Item Response Theory for Psychologists. Lawrence Erlbaum Associates. ISBN: 978-0805828191
- Kim, S.-H. & Cohen, A. S. (1998). Detection of differential item functioning under the graded response model with the likelihood ratio test. Applied Psychological Measurement, 22(4), 345–355. DOI: 10.1177/014662169802200403 ↗
How to cite this page
ScholarGate. (2026, June 3). Multi-Group Item Response Theory. ScholarGate. https://scholargate.app/en/psychometrics/multi-group-item-response-theory
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Confirmatory factor analysisPsychometrics↔ compare
- Differential Item FunctioningPsychometrics↔ compare
- Item Response TheoryPsychometrics↔ compare
- Multi-group confirmatory factor analysisPsychometrics↔ compare
- Multi-group measurement invariancePsychometrics↔ compare
- Multi-group Rasch modelPsychometrics↔ compare