Latent structurePsychometricsModel

Item Analysis (Classical Test Theory)

Also known as: Madde Analizi (Klasik Test Kuramı), CTT item analysis, classical item analysis

OriginatorClassical Test Theory tradition; foundational texts by Allen & Yen (1979) and Crocker & Algina (1986)Year1979Sources2Related methods7

Item analysis is the foundational psychometric procedure for evaluating the quality of individual test or scale items within the Classical Test Theory (CTT) framework, as systematised by Allen and Yen (1979) and Crocker and Algina (1986). It produces an item difficulty index, an item discrimination index, and a distractor analysis for each item, enabling test developers to identify items that are too easy, too hard, or failing to separate high- and low-ability respondents.

Key highlights

  • Simple, transparent, and universally understood by test developers, reviewers, and accreditation bodies.
  • Requires no distributional assumptions and is applicable to very small pilot samples.
  • Provides actionable item-level diagnostics — difficulty, discrimination, and distractor functioning — in a single pass.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Item analysis is appropriate at the pilot or pre-operational stage of any test or scale development project where items are binary (correct/incorrect) or polytomously scored. Three conditions should be met. Each item must have a recorded correct or scored response code. The sample, while a minimum of 30 is feasible, should be as large as is practical to yield stable estimates — 100 or more respondents is a common practical target. The analysis should use correlations with the total score that already includes the item (corrected item-total correlations are preferred to avoid spurious inflation). Item analysis is an appropriate first step before reliability estimation (Cronbach's alpha) and before committing to a factor-analytic or item-response-theory examination of the item pool.

Strengths & limitations

Strengths
  • Simple, transparent, and universally understood by test developers, reviewers, and accreditation bodies.
  • Requires no distributional assumptions and is applicable to very small pilot samples.
  • Provides actionable item-level diagnostics — difficulty, discrimination, and distractor functioning — in a single pass.
Limitations
  • Difficulty and discrimination indices are sample-dependent; they may shift substantially across groups that differ in ability level.
  • The framework does not model the probability of a correct response as a function of latent ability, so item parameters are not invariant across samples of different ability distributions.
  • Distractor analysis is meaningful only for multiple-choice items; it offers no guidance for open-ended or rating-scale formats beyond the basic difficulty and discrimination statistics.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

What is the difference between item analysis and item response theory?

Item analysis within Classical Test Theory computes sample-dependent statistics — difficulty p and discrimination r_pb — that describe how each item performed in a particular group. Item response theory (IRT) models the probability of a correct response as a mathematical function of a person's latent ability, yielding item parameters that are theoretically invariant across groups of different ability distributions. CTT item analysis is simpler and requires smaller samples; IRT is more powerful but demands larger samples and stronger assumptions. Item analysis is commonly the first step, with IRT applied to well-screened item pools.

What discrimination index should I use?

The point-biserial correlation r_pb between the binary item score and the continuous total score is the most widely recommended discrimination index for items scored dichotomously, with a minimum acceptable value of about 0.30. An alternative is the D index (upper 27% minus lower 27% correct), which is simpler to compute by hand and easier to explain, but r_pb is preferred in modern psychometric practice because it uses all the data.

What should I do with an item whose p value is outside the 0.20–0.80 range?

Very easy items (p > 0.80) or very hard items (p < 0.20) have limited variance and therefore limited capacity to discriminate. They should generally be revised or dropped, unless they serve a specific diagnostic purpose — for example, a very easy item at the start of a test to settle respondents, or a very hard item included deliberately to challenge the highest performers.

How large a sample do I need for stable item statistics?

A minimum of 30 examinees is workable for a rough screening, but estimates of p and r_pb are considerably more stable with 100 or more respondents. For high-stakes tests or large item banks, pilot samples of 200–500 or more are recommended to ensure that item statistics will generalise to the operational examinee population.

Sources

  1. 1.
    Allen, M. J. & Yen, W. M. (1979). Introduction to Measurement Theory. Brooks/Cole.
    ISBN 978-0818501333
  2. 2.
    Crocker, L. & Algina, J. (1986). Introduction to Classical and Modern Test Theory. Holt, Rinehart & Winston.
    ISBN 978-0030616341

You have read it. What now?

Cite this page

ScholarGate. (2026, June 1). Item Analysis. ScholarGate. https://scholargate.app/psychometrics/item-analysis

Item Analysis (Classical Test Theory) | ScholarGate