Three-Parameter Logistic IRT Model (3PL)
Three-Parameter Logistic Item Response Theory Model · Also known as: 3PL IRT — Üç Parametreli Madde Tepki Modeli, three-parameter logistic model, 3PLM, Birnbaum model
The three-parameter logistic (3PL) model, introduced by Allan Birnbaum in 1968, is an item response theory model that describes the probability of a correct response to a binary test item as a function of three item-level parameters — difficulty, discrimination, and a lower asymptote representing guessing — and one person-level parameter representing latent ability.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
The 3PL model is appropriate when you have a set of binary (right/wrong) items — most commonly a multiple-choice achievement or ability test — and you have theoretical or empirical reason to believe that examinees with very low ability can still guess correctly at a non-trivial rate. Several conditions must hold: unidimensionality (all items reflect a single latent trait) and local independence (item responses are independent given ability) are foundational assumptions that should be evaluated before fitting the model. A large sample of at least 500 respondents is required because the guessing parameter c is estimated with considerable uncertainty in small samples; with fewer than 100 respondents or fewer than five items the model cannot be reliably calibrated and a simpler 2PL or Rasch model should be preferred instead. The c parameter should theoretically hover near 1/K, where K is the number of response options; if it drifts far from that value for many items the model may be misspecified.
Strengths & limitations
- Explicitly accounts for the guessing floor on multiple-choice items, preventing the ability estimates from being artificially pulled toward the centre by lucky guesses.
- Places all items and all persons on a single common logit scale, enabling item banking, test equating, and adaptive testing.
- Richer than the 2PL in situations where a meaningful guessing component exists, and richer than the Rasch model in that it also allows items to vary in discriminating power.
- The guessing parameter c is the hardest of the three to estimate reliably and requires large samples (≥ 500); with smaller samples estimates are noisy and may not be meaningful.
- The model is overparameterised relative to the Rasch model and 2PL; model comparison with AIC and BIC is essential to justify the added complexity.
- Assumes strictly unidimensional structure: if the item pool is multidimensional, parameter estimates are distorted. Unidimensionality must be checked, for instance with EFA or confirmatory IRT, before calibration.
Frequently asked
What is the difference between the Rasch model, 2PL, and 3PL?
All three are logistic IRT models for binary items, but they differ in how many item parameters they estimate. The Rasch (1PL) model assumes all items are equally discriminating and ignores guessing, estimating only difficulty. The 2PL adds a discrimination parameter, allowing items to vary in how sharply they differentiate ability levels. The 3PL further adds the guessing parameter, modelling the non-zero floor of the item characteristic curve. Each additional parameter makes the model more flexible but also harder to estimate and more demanding of sample size.
Why is a large sample necessary for the 3PL?
The guessing parameter c is weakly identified: it is estimated from behaviour at the low tail of the ability distribution, which is sparsely populated in most samples. Even with 200 or 300 respondents the c estimates are unstable and their standard errors are large. A conventional minimum of around 500 respondents is needed to obtain reasonably precise estimates; larger samples of 1,000 or more are common in operational testing programmes.
How do I check whether the 3PL is better than the 2PL for my data?
Compare the two models using AIC and BIC computed from their marginal maximum likelihoods. If the 3PL achieves a meaningfully lower AIC or BIC, the guessing parameter is earning its place. Also inspect the estimated c values: if most items show c near zero, the guessing parameter is not contributing and the 2PL is the more parsimonious choice. A likelihood ratio test can be used if both models are nested and the sample is large.
Can the 3PL be used for Likert-type or rating-scale items?
No. The 3PL is defined for dichotomous (right/wrong, yes/no) responses only. For ordered polytomous items, such as Likert scales or partial-credit scoring, the appropriate IRT models are the Graded Response Model (for items with ordered categories of varying step widths) or the Partial Credit Model (for step-based scoring).
Sources
- Birnbaum, A. (1968). Some latent trait models and their use in inferring an examinee's ability. In F. M. Lord & M. R. Novick (Eds.), Statistical theories of mental test scores (pp. 397–479). Addison-Wesley. link ↗
- Baker, F. B. & Kim, S. H. (2004). Item response theory: Parameter estimation techniques (2nd ed.). Marcel Dekker. ISBN: 978-0824758172
How to cite this page
ScholarGate. (2026, June 1). Three-Parameter Logistic Item Response Theory Model. ScholarGate. https://scholargate.app/en/psychometrics/three-pl-irt
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- 2PL IRTPsychometrics↔ compare
- CFAStatistics↔ compare
- Cronbach's AlphaStatistics↔ compare
- EFAStatistics↔ compare
- Rasch ModelPsychometrics↔ compare