Multi-Group Confirmatory Factor Analysis (MG-CFA)
Multi-Group Confirmatory Factor Analysis · Also known as: MG-CFA, multi-group CFA, measurement invariance testing, multi-sample CFA
Multi-group confirmatory factor analysis tests whether a measurement model holds equivalently across two or more groups — such as cultures, genders, or time points. By imposing increasingly stringent equality constraints and comparing model fit, it determines whether comparisons of latent mean scores are justified.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
+12 more
When to use it
Use MG-CFA whenever you plan to compare latent factor means or structural paths across groups and want to verify that such comparisons are valid. It is the standard method for establishing measurement equivalence before cross-cultural, multi-gender, or longitudinal comparisons. Minimum requirements are a reasonably large sample in each group (at least 100–200 is common guidance), a pre-specified factor structure supported by theory or prior EFA, and items measured at an interval or ordinal level appropriate for the estimator chosen. Do not use MG-CFA when you have no prior factor structure to test — use multi-group EFA instead — or when groups are too small to support stable SEM estimation.
Strengths & limitations
- Provides a rigorous, hierarchical framework for testing measurement equivalence at the level of loadings, intercepts, and residuals separately.
- Produces quantitative fit-difference statistics that allow partial invariance solutions, preserving some comparative information even when full invariance does not hold.
- Integrates naturally with structural equation modeling, so once measurement equivalence is established, latent mean comparisons and cross-group path analyses follow in the same framework.
- Can accommodate ordered-categorical (Likert) items through the WLSMV estimator with polychoric correlations, extending invariance testing beyond normally distributed continuous data.
- Results are directly interpretable in terms of measurement theory: each constraint level maps onto a theoretically meaningful claim about how the construct is measured.
- Requires a reasonably large sample in every group; sparse cells yield unstable parameter estimates and inflated chi-square differences.
- The chi-square difference test is sensitive to sample size: in large samples trivially small constraint violations become statistically significant, calling for CFI-based supplementary criteria.
- Achieving full scalar invariance is the exception rather than the rule in applied research; partial invariance requires additional analytical steps and complicates the interpretation of latent mean differences.
- The method presupposes that a sound factor structure has already been identified; a poorly specified baseline model propagates error through all subsequent invariance tests.
Frequently asked
What is the difference between metric and scalar invariance?
Metric invariance requires only that factor loadings are equal across groups, meaning the items relate to the latent factor with the same strength. Scalar invariance additionally requires equal item intercepts, meaning respondents with the same latent score tend to give the same raw item response regardless of group. Latent mean comparisons require scalar invariance; only correlational or path comparisons require metric invariance.
What should I do if only partial invariance holds?
Identify which loadings or intercepts are non-invariant using modification indices, free those specific parameters while keeping the others constrained, and re-run the model. Partial metric or scalar invariance still allows limited comparisons if at least two intercepts per factor are invariant, but conclusions must be qualified and the freed parameters reported transparently.
How large should each group sample be?
There is no single universal cutoff, but most recommendations suggest at least 100–200 participants per group for stable CFA estimation. With very complex models or many groups the requirements are higher. Simulation studies suggest that samples below 100 per group produce unreliable chi-square difference tests and unstable parameter estimates.
Can MG-CFA be used with ordinal Likert items?
Yes. The recommended approach is to treat the items as categorical, use polychoric correlations, and apply the WLSMV (or DWLS) estimator available in software such as Mplus, lavaan, or LISREL. Standard maximum likelihood on the raw ordinal scores tends to underestimate loadings and distort fit indices.
How does MG-CFA relate to differential item functioning?
Both assess whether items behave equivalently across groups, but they come from different traditions. DIF originated in item response theory and examines each item individually given the total score or latent trait. MG-CFA tests the entire factor model simultaneously and provides tests at both the loading and intercept level. Non-invariant intercepts in MG-CFA correspond broadly to uniform DIF in IRT.
Sources
- Vandenberg, R. J. & Lance, C. E. (2000). A review and synthesis of the measurement invariance literature: Suggestions, practices, and recommendations for organizational research. Organizational Research Methods, 3(1), 4–70. DOI: 10.1177/109442810031002 ↗
- Millsap, R. E. (2011). Statistical Approaches to Measurement Equivalence. Routledge. ISBN: 978-0805859447
How to cite this page
ScholarGate. (2026, June 3). Multi-Group Confirmatory Factor Analysis. ScholarGate. https://scholargate.app/en/psychometrics/multi-group-confirmatory-factor-analysis
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Confirmatory factor analysisPsychometrics↔ compare
- Differential Item FunctioningPsychometrics↔ compare
- EFAStatistics↔ compare
- Measurement InvariancePsychometrics↔ compare
- Multi-group EFAPsychometrics↔ compare
- Structural Equation ModelingResearch Statistics↔ compare