Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Psychometrics›Ordinal Measurement Invariance Testing
Latent structureScale / measurement

Ordinal Measurement Invariance Testing

Also known as: ordinal MI, measurement invariance for ordinal data, ordinal CFA invariance, categorical measurement invariance

Ordinal measurement invariance testing evaluates whether a multi-group confirmatory factor model holds equivalent measurement properties across groups when scale items are ordinal — such as Likert-type response scales. It uses polychoric correlations and categorical estimators (WLSMV/DWLS) rather than Pearson-based methods, correcting the systematic bias that arises when ordinal data are treated as continuous.

ScholarGate
  1. Latent structure
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Ordinal Measurement Invariance
Confirmatory factor anal…Differential Item Functi…Item Response TheoryMeasurement InvarianceMulti-group confirmatory…Ordinal CFAOrdinal Nomological Vali…

When to use it

Use ordinal measurement invariance testing whenever you plan to compare latent means, factor correlations, or structural paths across groups using a scale composed of Likert-type or other ordered-categorical items. It is mandatory before any group comparison that claims to be more than descriptive. Do not use it when items are genuinely continuous (interval or ratio level), in which case standard metric/scalar CFA with ML is appropriate. Also avoid it if group sample sizes are small — polychoric correlations and WLSMV estimators require at least 200 cases per group, and fewer cases make threshold estimation unstable. It is not a substitute for differential item functioning analysis when item-level bias for individual response categories is the primary concern.

Strengths & limitations

Strengths
  • Respects the ordinal nature of Likert-type items by using polychoric correlations and categorical estimators, avoiding the bias of treating ordinal data as continuous.
  • Provides a hierarchical framework — configural, metric, scalar — that precisely locates where group differences in measurement arise.
  • Partial invariance solutions allow useful group comparisons even when full scalar invariance does not hold, provided the non-invariant items are clearly reported.
  • WLSMV estimation is robust to non-normality and performs well with moderate sample sizes compared to full-information ML.
  • Directly supports the validity of cross-group mean comparisons that are ubiquitous in survey and psychological research.
Limitations
  • Requires relatively large samples per group (at least 200 recommended) for stable polychoric correlation and threshold estimation.
  • WLSMV chi-square difference tests require special procedures (DIFFTEST) and are not available in all software, adding implementation complexity.
  • When many thresholds are non-invariant, interpreting partial invariance is complex and substantive conclusions about group differences become tentative.
  • The ordinal CFA framework assumes the latent response continuum underlying each item is normally distributed, which may not hold for strongly skewed items.

Frequently asked

Why can't I just treat Likert items as continuous and run standard measurement invariance testing?

Treating ordinal responses as continuous data underestimates correlations — especially for items with few categories or skewed distributions — and produces biased factor loadings, incorrect standard errors, and inflated chi-square statistics. Polychoric correlations with a categorical estimator such as WLSMV correct these biases and provide more accurate model fit and parameter estimates.

What is the difference between ordinal measurement invariance and differential item functioning (DIF)?

Both assess whether items perform equivalently across groups, but from different frameworks. DIF (from item response theory) focuses on individual items and individual response categories, testing whether the probability of a given response differs between groups after conditioning on the latent trait. Ordinal measurement invariance testing works within a CFA framework, constraining entire loading and threshold parameters across groups and evaluating overall model fit. DIF is more item-centric and granular; ordinal MI testing evaluates the whole scale structure simultaneously.

What does partial invariance mean, and can I still compare groups?

Partial invariance means that some — but not all — loadings or thresholds are equal across groups. Latent mean comparisons remain possible if at least two items per factor have invariant loadings and thresholds, provided the non-invariant items are clearly identified and their parameters freed. The interpretation of group differences is then qualified: only the construct as defined by the invariant items is being compared.

Which software can run ordinal measurement invariance testing?

Mplus is the most widely used and supports WLSMV with the DIFFTEST option for model comparison. R packages lavaan (with estimator WLSMV and the lavTestScore/lavTestLRT functions) and semTools provide similar functionality. LISREL and OpenMx also support weighted least squares estimation for ordinal data.

Is full scalar invariance always required before comparing group means?

Full scalar invariance is the ideal standard, but partial scalar invariance — with at least two items per factor having invariant thresholds — is generally accepted as a workable basis for latent mean comparisons, provided non-invariant items are reported. Some researchers also recommend checking that non-invariant thresholds are randomly rather than systematically distributed across groups to avoid biased mean estimates.

Sources

  1. Millsap, R. E. (2011). Statistical Approaches to Measurement Invariance. Routledge. ISBN: 978-1848728936
  2. Muthén, B. O. (1984). A general structural equation model with dichotomous, ordered categorical, and continuous latent variable indicators. Psychometrika, 49(1), 115–132. DOI: 10.1007/BF02294210 ↗

How to cite this page

ScholarGate. (2026, June 3). Ordinal Measurement Invariance Testing. ScholarGate. https://scholargate.app/en/psychometrics/ordinal-measurement-invariance

Related methods

Confirmatory factor analysisDifferential Item FunctioningItem Response TheoryMeasurement InvarianceMulti-group confirmatory factor analysisOrdinal CFA

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Confirmatory factor analysisPsychometrics↔ compare
  • Differential Item FunctioningPsychometrics↔ compare
  • Item Response TheoryPsychometrics↔ compare
  • Measurement InvariancePsychometrics↔ compare
  • Multi-group confirmatory factor analysisPsychometrics↔ compare
  • Ordinal CFAPsychometrics↔ compare
Compare side by side →

Referenced by

Ordinal Nomological Validity

Similar methods

Polytomous Measurement InvarianceOrdinal CFAMulti-group measurement invarianceRobust Measurement InvariancePolytomous Confirmatory Factor AnalysisMulti-group confirmatory factor analysisOrdinal EFAShort Form Measurement Invariance

Related reference concepts

Item Response TheoryStructural Equation ModelingStructural and Latent Variable ModelsFactor AnalysisPsychometrics & Statistics & MethodologyLatent Class Analysis

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Ordinal Measurement Invariance (Ordinal Measurement Invariance Testing). Retrieved 2026-07-20 from https://scholargate.app/en/psychometrics/ordinal-measurement-invariance · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Roger Millsap; Bengt Muthén
Year
1984–2011
Type
Multi-group model comparison
DataType
Ordinal / Likert-type item responses
Subfamily
Scale / measurement
Related methods
Confirmatory factor analysisDifferential Item FunctioningItem Response TheoryMeasurement InvarianceMulti-group confirmatory factor analysisOrdinal CFA
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account