Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Psychometrics›Polytomous Item Analysis
Latent structureScale / measurement

Polytomous Item Analysis

Also known as: ordered-category item analysis, graded response analysis, polytomous IRT, rated-scale item analysis

Polytomous item analysis examines the psychometric behavior of items that have more than two ordered response categories — such as Likert-type scales or partial-credit tasks. It evaluates each item's difficulty thresholds, discriminating power, and category functioning to determine whether the full response scale is being used as intended and whether each item contributes reliably to measuring the underlying construct.

ScholarGate
  1. Latent structure
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Polytomous item analysis
Confirmatory factor anal…EFAGRMPCM / GPCM

When to use it

Use polytomous item analysis whenever you are developing or evaluating a scale whose items carry three or more ordered response options — Likert attitude scales, partial-credit achievement items, or observer rating scales. It is the appropriate successor to simple dichotomous item analysis once the response format is ordinal rather than binary. Do NOT apply standard polytomous IRT models when responses are purely nominal (unordered categories) — use nominal response models instead. Avoid when sample size is small (fewer than 200–300 respondents for stable IRT parameter estimation); CTT-based indices are more robust under small samples but yield less detailed diagnostics.

Strengths & limitations

Strengths
  • Models the full response scale, revealing whether each ordered category functions distinctly and contributes to measurement.
  • Provides item-level discrimination, difficulty thresholds, and information functions that are invariant across (representative) subpopulations under correct model fit.
  • Identifies disordered thresholds and underused categories early in scale development, preventing wasted items in final instruments.
  • Enables test information analysis — showing precisely where on the trait continuum an item or scale is most precise.
  • Compatible with both IRT frameworks (GRM, PCM, RSM) and CTT diagnostics (corrected item-total correlations, alpha-if-deleted), offering flexibility.
Limitations
  • Stable IRT parameter estimation typically requires samples of 200 or more; small samples produce imprecise discrimination and threshold estimates.
  • Model selection — GRM versus PCM versus RSM — requires subject-matter judgment and comparative fit testing; the wrong model can obscure true item properties.
  • Interpretation of multiple threshold parameters per item is more complex than simple p-values and point-biserials from dichotomous analysis.
  • Assumes local independence and unidimensionality; violations inflate reliability estimates and distort item parameter recovery.

Frequently asked

What is the difference between the Graded Response Model and the Partial Credit Model?

Both handle ordered polytomous items but model different transition probabilities. The GRM (Samejima, 1969) models cumulative boundary response functions — the probability of scoring at or above each category — and allows item-specific discrimination. The PCM (Masters, 1982) models adjacent-category transitions — the probability of choosing category k over k-1 — and constrains discrimination to be equal across items (Rasch-family). Choose GRM when items vary in discriminating power and when a cumulative model fits the construct; choose PCM when equal discrimination is defensible and Rasch-family measurement properties are desired.

How do I detect disordered thresholds and what should I do?

Disordered thresholds occur when a higher category boundary has a lower difficulty estimate than a lower boundary — meaning respondents are, in effect, more likely to skip a middle category. Examine the threshold estimates and the Category Response Functions: if CRF peaks are absent or very low for a middle category, it is rarely used. The remedy is to collapse the underused category with an adjacent one and re-run the analysis, checking that the revised scale has ordered thresholds.

Can polytomous item analysis be conducted without IRT software?

Yes. Classical test theory approaches — corrected item-total (polyserial) correlations, item means, item standard deviations, and coefficient alpha-if-deleted — are simpler diagnostics that do not require IRT software and are robust with smaller samples. They do not, however, provide threshold estimates, information functions, or formal model-fit statistics. A combined CTT-and-IRT approach is common in applied scale development.

How large a sample do I need for polytomous IRT analysis?

A minimum of about 200 respondents is generally recommended for the simpler Rasch-family polytomous models (PCM, RSM), and 300–500 for the GRM with multiple response categories and many items. Smaller samples produce unstable discrimination and threshold estimates. When sample size is limited, use CTT-based diagnostics and treat IRT results as exploratory.

Should I always use IRT for polytomous items?

Not necessarily. IRT provides richer diagnostics — information functions, invariant parameters, model-data fit — but requires larger samples and more technical expertise. For routine scale development with 200+ respondents and standard Likert formats, an IRT analysis adds substantial value. For small studies or preliminary item screening, CTT indices (item-total correlations, alpha-if-deleted) are efficient and sufficient as a first pass.

Sources

  1. Samejima, F. (1969). Estimation of latent ability using a response pattern of graded scores. Psychometrika Monograph Supplement, 34(4, Pt. 2), 1–97. DOI: 10.1007/BF03372160 ↗
  2. Embretson, S. E. & Reise, S. P. (2000). Item Response Theory for Psychologists. Lawrence Erlbaum Associates. ISBN: 978-0805828191

How to cite this page

ScholarGate. (2026, June 3). Polytomous Item Analysis. ScholarGate. https://scholargate.app/en/psychometrics/polytomous-item-analysis

Related methods

Confirmatory factor analysisEFAGRMPCM / GPCM

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Confirmatory factor analysisPsychometrics↔ compare
  • EFAStatistics↔ compare
  • GRMPsychometrics↔ compare
  • PCM / GPCMPsychometrics↔ compare
Compare side by side →

Similar methods

Polytomous scale developmentPolytomous Rasch ModelOrdinal IRTPolytomous Construct ValidityOrdinal Rasch ModelGRMPCM / GPCMPolytomous DIF

Related reference concepts

Item Response TheoryItem AnalysisPsychometrics & Statistics & MethodologyStructural and Latent Variable ModelsPsychological Testing and PsychometricsLatent Class Analysis

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Polytomous item analysis (Polytomous Item Analysis). Retrieved 2026-07-21 from https://scholargate.app/en/psychometrics/polytomous-item-analysis · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Fumiko Samejima (graded response model, 1969); David Andrich (rating scale model, 1978); Geoffrey Masters (partial credit model, 1982)
Year
1969–1982
Type
Item-level psychometric analysis
DataType
Ordinal polytomous responses (e.g., Likert-scale items)
Subfamily
Scale / measurement
Related methods
Confirmatory factor analysisEFAGRMPCM / GPCM
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account