Item Analysis (Classical Test Theory)
Also known as: Madde Analizi (Klasik Test Kuramı), CTT item analysis, classical item analysis
Item analysis is the foundational psychometric procedure for evaluating the quality of individual test or scale items within the Classical Test Theory (CTT) framework, as systematised by Allen and Yen (1979) and Crocker and Algina (1986). It produces an item difficulty index, an item discrimination index, and a distractor analysis for each item, enabling test developers to identify items that are too easy, too hard, or failing to separate high- and low-ability respondents.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Item analysis is appropriate at the pilot or pre-operational stage of any test or scale development project where items are binary (correct/incorrect) or polytomously scored. Three conditions should be met. Each item must have a recorded correct or scored response code. The sample, while a minimum of 30 is feasible, should be as large as is practical to yield stable estimates — 100 or more respondents is a common practical target. The analysis should use correlations with the total score that already includes the item (corrected item-total correlations are preferred to avoid spurious inflation). Item analysis is an appropriate first step before reliability estimation (Cronbach's alpha) and before committing to a factor-analytic or item-response-theory examination of the item pool.
Strengths & limitations
- Simple, transparent, and universally understood by test developers, reviewers, and accreditation bodies.
- Requires no distributional assumptions and is applicable to very small pilot samples.
- Provides actionable item-level diagnostics — difficulty, discrimination, and distractor functioning — in a single pass.
- Difficulty and discrimination indices are sample-dependent; they may shift substantially across groups that differ in ability level.
- The framework does not model the probability of a correct response as a function of latent ability, so item parameters are not invariant across samples of different ability distributions.
- Distractor analysis is meaningful only for multiple-choice items; it offers no guidance for open-ended or rating-scale formats beyond the basic difficulty and discrimination statistics.
Frequently asked
What is the difference between item analysis and item response theory?
Item analysis within Classical Test Theory computes sample-dependent statistics — difficulty p and discrimination r_pb — that describe how each item performed in a particular group. Item response theory (IRT) models the probability of a correct response as a mathematical function of a person's latent ability, yielding item parameters that are theoretically invariant across groups of different ability distributions. CTT item analysis is simpler and requires smaller samples; IRT is more powerful but demands larger samples and stronger assumptions. Item analysis is commonly the first step, with IRT applied to well-screened item pools.
What discrimination index should I use?
The point-biserial correlation r_pb between the binary item score and the continuous total score is the most widely recommended discrimination index for items scored dichotomously, with a minimum acceptable value of about 0.30. An alternative is the D index (upper 27% minus lower 27% correct), which is simpler to compute by hand and easier to explain, but r_pb is preferred in modern psychometric practice because it uses all the data.
What should I do with an item whose p value is outside the 0.20–0.80 range?
Very easy items (p > 0.80) or very hard items (p < 0.20) have limited variance and therefore limited capacity to discriminate. They should generally be revised or dropped, unless they serve a specific diagnostic purpose — for example, a very easy item at the start of a test to settle respondents, or a very hard item included deliberately to challenge the highest performers.
How large a sample do I need for stable item statistics?
A minimum of 30 examinees is workable for a rough screening, but estimates of p and r_pb are considerably more stable with 100 or more respondents. For high-stakes tests or large item banks, pilot samples of 200–500 or more are recommended to ensure that item statistics will generalise to the operational examinee population.
Sources
- Allen, M. J. & Yen, W. M. (1979). Introduction to Measurement Theory. Brooks/Cole. ISBN: 978-0818501333
- Crocker, L. & Algina, J. (1986). Introduction to Classical and Modern Test Theory. Holt, Rinehart & Winston. ISBN: 978-0030616341
How to cite this page
ScholarGate. (2026, June 1). Item Analysis (Classical Test Theory). ScholarGate. https://scholargate.app/en/psychometrics/item-analysis
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Confirmatory factor analysisPsychometrics↔ compare
- Cronbach's AlphaStatistics↔ compare
- EFAStatistics↔ compare
- Item Response TheoryPsychometrics↔ compare
- Test EquatingPsychometrics↔ compare