Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Psychometrics›Item Analysis (Classical Test Theory)
Latent structure

Item Analysis (Classical Test Theory)

Also known as: Madde Analizi (Klasik Test Kuramı), CTT item analysis, classical item analysis

Item analysis is the foundational psychometric procedure for evaluating the quality of individual test or scale items within the Classical Test Theory (CTT) framework, as systematised by Allen and Yen (1979) and Crocker and Algina (1986). It produces an item difficulty index, an item discrimination index, and a distractor analysis for each item, enabling test developers to identify items that are too easy, too hard, or failing to separate high- and low-ability respondents.

ScholarGate
  1. Latent structure
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Item Analysis
Confirmatory factor anal…Cronbach's AlphaEFAItem Response TheoryTest EquatingDIF AnalysisOrdinal Content Validity

When to use it

Item analysis is appropriate at the pilot or pre-operational stage of any test or scale development project where items are binary (correct/incorrect) or polytomously scored. Three conditions should be met. Each item must have a recorded correct or scored response code. The sample, while a minimum of 30 is feasible, should be as large as is practical to yield stable estimates — 100 or more respondents is a common practical target. The analysis should use correlations with the total score that already includes the item (corrected item-total correlations are preferred to avoid spurious inflation). Item analysis is an appropriate first step before reliability estimation (Cronbach's alpha) and before committing to a factor-analytic or item-response-theory examination of the item pool.

Strengths & limitations

Strengths
  • Simple, transparent, and universally understood by test developers, reviewers, and accreditation bodies.
  • Requires no distributional assumptions and is applicable to very small pilot samples.
  • Provides actionable item-level diagnostics — difficulty, discrimination, and distractor functioning — in a single pass.
Limitations
  • Difficulty and discrimination indices are sample-dependent; they may shift substantially across groups that differ in ability level.
  • The framework does not model the probability of a correct response as a function of latent ability, so item parameters are not invariant across samples of different ability distributions.
  • Distractor analysis is meaningful only for multiple-choice items; it offers no guidance for open-ended or rating-scale formats beyond the basic difficulty and discrimination statistics.

Frequently asked

What is the difference between item analysis and item response theory?

Item analysis within Classical Test Theory computes sample-dependent statistics — difficulty p and discrimination r_pb — that describe how each item performed in a particular group. Item response theory (IRT) models the probability of a correct response as a mathematical function of a person's latent ability, yielding item parameters that are theoretically invariant across groups of different ability distributions. CTT item analysis is simpler and requires smaller samples; IRT is more powerful but demands larger samples and stronger assumptions. Item analysis is commonly the first step, with IRT applied to well-screened item pools.

What discrimination index should I use?

The point-biserial correlation r_pb between the binary item score and the continuous total score is the most widely recommended discrimination index for items scored dichotomously, with a minimum acceptable value of about 0.30. An alternative is the D index (upper 27% minus lower 27% correct), which is simpler to compute by hand and easier to explain, but r_pb is preferred in modern psychometric practice because it uses all the data.

What should I do with an item whose p value is outside the 0.20–0.80 range?

Very easy items (p > 0.80) or very hard items (p < 0.20) have limited variance and therefore limited capacity to discriminate. They should generally be revised or dropped, unless they serve a specific diagnostic purpose — for example, a very easy item at the start of a test to settle respondents, or a very hard item included deliberately to challenge the highest performers.

How large a sample do I need for stable item statistics?

A minimum of 30 examinees is workable for a rough screening, but estimates of p and r_pb are considerably more stable with 100 or more respondents. For high-stakes tests or large item banks, pilot samples of 200–500 or more are recommended to ensure that item statistics will generalise to the operational examinee population.

Sources

  1. Allen, M. J. & Yen, W. M. (1979). Introduction to Measurement Theory. Brooks/Cole. ISBN: 978-0818501333
  2. Crocker, L. & Algina, J. (1986). Introduction to Classical and Modern Test Theory. Holt, Rinehart & Winston. ISBN: 978-0030616341

How to cite this page

ScholarGate. (2026, June 1). Item Analysis (Classical Test Theory). ScholarGate. https://scholargate.app/en/psychometrics/item-analysis

Related methods

Confirmatory factor analysisCronbach's AlphaEFAItem Response TheoryTest Equating

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Confirmatory factor analysisPsychometrics↔ compare
  • Cronbach's AlphaStatistics↔ compare
  • EFAStatistics↔ compare
  • Item Response TheoryPsychometrics↔ compare
  • Test EquatingPsychometrics↔ compare
Compare side by side →

Referenced by

DIF AnalysisOrdinal Content Validity

Similar methods

Ordinal Item AnalysisMulti-group item analysisRobust Item AnalysisStandardized Test AnalysisItem Response TheoryBayesian Item AnalysisComputerized adaptive test item analysisShort-form item analysis

Related reference concepts

Item AnalysisItem Response TheoryPsychometrics & Statistics & MethodologyPsychological Testing and PsychometricsMeasurementTest Construction

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Item Analysis (Item Analysis (Classical Test Theory)). Retrieved 2026-07-21 from https://scholargate.app/en/psychometrics/item-analysis · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Classical Test Theory tradition; foundational texts by Allen & Yen (1979) and Crocker & Algina (1986)
Year
1979
Type
Descriptive / psychometric screening
Outcome
Item difficulty index (p), item discrimination index (r pb or D), distractor analysis
Data
Binary or ordinal item scores
Min Sample
30
Difficulty
1
Related methods
Confirmatory factor analysisCronbach's AlphaEFAItem Response TheoryTest Equating
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account