Ordinal Measurement Invariance Testing
Also known as: ordinal MI, measurement invariance for ordinal data, ordinal CFA invariance, categorical measurement invariance
Ordinal measurement invariance testing evaluates whether a multi-group confirmatory factor model holds equivalent measurement properties across groups when scale items are ordinal — such as Likert-type response scales. It uses polychoric correlations and categorical estimators (WLSMV/DWLS) rather than Pearson-based methods, correcting the systematic bias that arises when ordinal data are treated as continuous.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use ordinal measurement invariance testing whenever you plan to compare latent means, factor correlations, or structural paths across groups using a scale composed of Likert-type or other ordered-categorical items. It is mandatory before any group comparison that claims to be more than descriptive. Do not use it when items are genuinely continuous (interval or ratio level), in which case standard metric/scalar CFA with ML is appropriate. Also avoid it if group sample sizes are small — polychoric correlations and WLSMV estimators require at least 200 cases per group, and fewer cases make threshold estimation unstable. It is not a substitute for differential item functioning analysis when item-level bias for individual response categories is the primary concern.
Strengths & limitations
- Respects the ordinal nature of Likert-type items by using polychoric correlations and categorical estimators, avoiding the bias of treating ordinal data as continuous.
- Provides a hierarchical framework — configural, metric, scalar — that precisely locates where group differences in measurement arise.
- Partial invariance solutions allow useful group comparisons even when full scalar invariance does not hold, provided the non-invariant items are clearly reported.
- WLSMV estimation is robust to non-normality and performs well with moderate sample sizes compared to full-information ML.
- Directly supports the validity of cross-group mean comparisons that are ubiquitous in survey and psychological research.
- Requires relatively large samples per group (at least 200 recommended) for stable polychoric correlation and threshold estimation.
- WLSMV chi-square difference tests require special procedures (DIFFTEST) and are not available in all software, adding implementation complexity.
- When many thresholds are non-invariant, interpreting partial invariance is complex and substantive conclusions about group differences become tentative.
- The ordinal CFA framework assumes the latent response continuum underlying each item is normally distributed, which may not hold for strongly skewed items.
Frequently asked
Why can't I just treat Likert items as continuous and run standard measurement invariance testing?
Treating ordinal responses as continuous data underestimates correlations — especially for items with few categories or skewed distributions — and produces biased factor loadings, incorrect standard errors, and inflated chi-square statistics. Polychoric correlations with a categorical estimator such as WLSMV correct these biases and provide more accurate model fit and parameter estimates.
What is the difference between ordinal measurement invariance and differential item functioning (DIF)?
Both assess whether items perform equivalently across groups, but from different frameworks. DIF (from item response theory) focuses on individual items and individual response categories, testing whether the probability of a given response differs between groups after conditioning on the latent trait. Ordinal measurement invariance testing works within a CFA framework, constraining entire loading and threshold parameters across groups and evaluating overall model fit. DIF is more item-centric and granular; ordinal MI testing evaluates the whole scale structure simultaneously.
What does partial invariance mean, and can I still compare groups?
Partial invariance means that some — but not all — loadings or thresholds are equal across groups. Latent mean comparisons remain possible if at least two items per factor have invariant loadings and thresholds, provided the non-invariant items are clearly identified and their parameters freed. The interpretation of group differences is then qualified: only the construct as defined by the invariant items is being compared.
Which software can run ordinal measurement invariance testing?
Mplus is the most widely used and supports WLSMV with the DIFFTEST option for model comparison. R packages lavaan (with estimator WLSMV and the lavTestScore/lavTestLRT functions) and semTools provide similar functionality. LISREL and OpenMx also support weighted least squares estimation for ordinal data.
Is full scalar invariance always required before comparing group means?
Full scalar invariance is the ideal standard, but partial scalar invariance — with at least two items per factor having invariant thresholds — is generally accepted as a workable basis for latent mean comparisons, provided non-invariant items are reported. Some researchers also recommend checking that non-invariant thresholds are randomly rather than systematically distributed across groups to avoid biased mean estimates.
Sources
- Millsap, R. E. (2011). Statistical Approaches to Measurement Invariance. Routledge. ISBN: 978-1848728936
- Muthén, B. O. (1984). A general structural equation model with dichotomous, ordered categorical, and continuous latent variable indicators. Psychometrika, 49(1), 115–132. DOI: 10.1007/BF02294210 ↗
How to cite this page
ScholarGate. (2026, June 3). Ordinal Measurement Invariance Testing. ScholarGate. https://scholargate.app/en/psychometrics/ordinal-measurement-invariance
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Confirmatory factor analysisPsychometrics↔ compare
- Differential Item FunctioningPsychometrics↔ compare
- Item Response TheoryPsychometrics↔ compare
- Measurement InvariancePsychometrics↔ compare
- Multi-group confirmatory factor analysisPsychometrics↔ compare
- Ordinal CFAPsychometrics↔ compare