Multilevel Measurement Invariance
Multilevel Measurement Invariance Testing · Also known as: MLMI, multilevel factorial invariance, cross-level measurement invariance, multilevel CFA invariance
Multilevel measurement invariance testing evaluates whether a latent construct is measured equivalently both within clusters (e.g., individuals within teams) and between clusters (e.g., team-level aggregates). It extends standard measurement invariance procedures to nested data structures commonly encountered in organisational, educational, and cross-cultural research.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use multilevel measurement invariance testing whenever you have nested data and plan to compare latent means or relationships across clusters, aggregate individual scores to the group level, or examine cross-level interactions in a structural model. It is essential prior to multilevel SEM, and for any study making between-group inferences from clustered survey data. Do not use it when data are not hierarchically nested, when cluster sizes are too small to estimate a between-level model (at least 30 clusters with adequate within-cluster n is a common guideline), or when only a single level of analysis is theoretically relevant.
Strengths & limitations
- Provides a principled, model-based test of whether a construct retains its measurement properties across hierarchical levels.
- Separates within-level and between-level factor structures, revealing whether a construct operates differently at individual and group levels.
- Supports rigorous cross-level comparisons of latent means and relationships in multilevel structural models.
- Produces partial invariance solutions when only some items are non-invariant, allowing informed decisions about which items to retain or respecify.
- Widely implementable in major SEM software (Mplus, lavaan, OpenMx) with established fit criteria.
- Requires large and well-balanced cluster structures; sparse data at the between level (few clusters, small clusters) lead to unstable parameter estimates.
- Model complexity increases rapidly with the number of levels and constructs, raising convergence challenges.
- Standard chi-square difference tests are sensitive to sample size; fit-index-based criteria (ΔCFI, ΔRMSEA) are preferred but their thresholds are heuristic.
- Partial invariance at one level complicates substantive interpretation and requires careful reporting of which constraints were relaxed.
Frequently asked
How is multilevel measurement invariance different from multigroup CFA invariance?
Standard multigroup CFA tests whether a factor model is equivalent across observed, predefined groups (e.g., male vs. female). Multilevel measurement invariance tests whether the model is equivalent across the levels of a hierarchy (within-individual vs. between-cluster), without requiring predefined groups. The hierarchical nesting itself defines the two levels of analysis.
Do I need full scalar invariance to aggregate individual scores to the group level?
Yes. Scalar invariance (equal loadings and equal intercepts across levels) is the minimum requirement for meaningful comparison of latent means at the between level. Metric invariance alone permits comparison of factor variances and covariances but not latent means. Partial scalar invariance, where at least two items per factor have invariant intercepts, can support limited mean comparisons with acknowledged constraints.
How many clusters and how large must each cluster be?
Common recommendations suggest at least 30 clusters with at least 5–10 observations per cluster for stable between-level parameter estimates. Very small cluster numbers (fewer than 20) make the between-level model unreliable even when within-level estimates are accurate. Simulation studies suggest more clusters are generally more important than larger cluster size.
What software can estimate multilevel measurement invariance models?
Mplus is the most widely used and offers the most complete set of estimators (MLR, WLSMV) for this class of models. The R packages lavaan (with the multilevel extension) and OpenMx also support two-level CFA, though Mplus syntax is better documented in the methodological literature for this specific procedure.
What should I do if scalar invariance fails at the between level?
Inspect modification indices to identify which item intercepts are non-invariant across levels. If two or more items per factor retain invariant intercepts, partial scalar invariance can be claimed, and latent mean comparisons can proceed with the caveat that the non-invariant items are freed. Report the non-invariant items, re-examine their wording for level-specific interpretation differences, and consider revising or replacing them in future scale development.
Sources
- Muthén, B. O., & Asparouhov, T. (2009). Multilevel factor analysis of class and student achievement components. Journal of Educational and Behavioral Statistics, 34(2), 250–270. link ↗
- Ryu, E. (2014). Factorial invariance in multilevel confirmatory factor analysis. British Journal of Mathematical and Statistical Psychology, 67(1), 172–194. DOI: 10.1111/bmsp.12014 ↗
How to cite this page
ScholarGate. (2026, June 3). Multilevel Measurement Invariance Testing. ScholarGate. https://scholargate.app/en/psychometrics/multilevel-measurement-invariance
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Confirmatory factor analysisPsychometrics↔ compare
- Measurement InvariancePsychometrics↔ compare
- Structural Equation ModelingResearch Statistics↔ compare