Ordinal Item Analysis
Also known as: item analysis for ordinal data, polytomous item analysis, Likert item analysis, OIA
Ordinal item analysis evaluates each individual item in a rating-scale or Likert-type instrument using descriptive and correlational statistics suited to ordered categorical response formats. It guides item selection and refinement by flagging items with problematic difficulty, poor discrimination, or low corrected item-total correlations before reliability and validity studies proceed.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use ordinal item analysis in the item-refinement stage of scale development, after an initial item pool has been administered to a pilot sample of at least 100–200 respondents per subscale. It is appropriate for Likert-type items with three to seven response categories. Do not apply standard Pearson-based item statistics uncritically to items with only two or three categories — polychoric correlations are more appropriate for severely ordinal responses. Do not use ordinal item analysis as a substitute for factor analysis: it examines items one at a time within a total score and cannot reveal multidimensional structure. Avoid using it when items in the pool intentionally measure different constructs; running item analysis across a mixed-content item set will penalise items that are valid but measure a different dimension.
Strengths & limitations
- Simple, transparent, and widely understood by scale developers and reviewers across social and health sciences.
- Provides actionable item-level statistics — mean, SD, CITC, alpha-if-deleted — that directly guide item retention or revision decisions.
- Computationally inexpensive and feasible even in small pilot samples relative to factor analysis.
- Detects ceiling and floor effects that factor analysis may not highlight explicitly.
- Can be combined with factor analysis and IRT for a comprehensive item evaluation workflow.
- Classical item statistics such as mean difficulty and CITC are sample-dependent and may not generalise to populations with very different trait distributions.
- Standard Pearson-based CITC underestimates true item discrimination when item responses are coarsely ordinal; polychoric-based alternatives are more accurate but less commonly implemented.
- Item analysis evaluates items within a total-score framework and therefore assumes unidimensionality; it cannot correctly evaluate items in a multidimensional scale without analysing each subscale separately.
- Relies on the total score as a proxy for the latent trait, which is itself imperfect, especially in short scales.
- The conventional CITC threshold of 0.30 is a rule of thumb with no formal statistical basis.
Frequently asked
What is an acceptable corrected item-total correlation for an ordinal item?
A common guideline is CITC ≥ 0.30. Items between 0.20 and 0.30 are borderline and warrant review for content importance. Items below 0.20 are generally poor candidates for retention unless they carry essential content that cannot be replaced. These thresholds assume a reasonably large item pool and pilot sample.
Should I use Pearson or polychoric correlations for the CITC?
Pearson correlations between item scores and the total score are used in most software by default and are adequate when items have five or more response categories and are approximately symmetric. For items with three or four categories, or when distributions are markedly skewed, polychoric correlations between each item and a factor score or latent total provide a more appropriate ordinal estimate of the underlying linear relationship.
Can ordinal item analysis replace confirmatory factor analysis for scale validation?
No. Item analysis evaluates items within a single total-score framework. It cannot assess whether a hypothesised factor structure fits the data, test measurement invariance across groups, or distinguish between competing dimensional models. CFA is a necessary follow-up step for scale validation.
How large a pilot sample do I need for stable item statistics?
As a rough minimum, 100–200 respondents per subscale are recommended for initial item analysis. With fewer than 50 respondents the CITC estimates are highly unstable. Larger samples are needed for subsequent factor analysis and IRT calibration.
What if deleting an item raises alpha but that item is content-critical?
Retain the item and document the decision. Reliability is one criterion among several in scale development. If an item represents a facet of the construct that cannot be captured by other items, it should be kept even if it slightly lowers alpha. Consider revising the item wording rather than deleting it.
Sources
- Nunnally, J. C. & Bernstein, I. H. (1994). Psychometric Theory (3rd ed.). McGraw-Hill. ISBN: 978-0070474659
- Drasgow, F., Levine, M. V., Tsien, S., Williams, B. & Mead, A. D. (1995). Fitting polytomous item response theory models to multiple-choice tests. Applied Psychological Measurement, 19(2), 143–165. DOI: 10.1177/014662169501900203 ↗
How to cite this page
ScholarGate. (2026, June 3). Ordinal Item Analysis. ScholarGate. https://scholargate.app/en/psychometrics/ordinal-item-analysis
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Differential Item FunctioningPsychometrics↔ compare
- Item Response TheoryPsychometrics↔ compare
- Ordinal Cronbach's AlphaPsychometrics↔ compare
- Ordinal EFAPsychometrics↔ compare
- Ordinal Reliability AnalysisPsychometrics↔ compare
- Scale developmentPsychometrics↔ compare