Angoff Standard Setting
Also known as: Angoff Method, Modified Angoff Method, Yes/No Angoff, Angoff Cut-Score Procedure
The Angoff method is a test-centered procedure for establishing a passing score (cut score) on an examination. A panel of content experts conceptualizes a 'borderline' or minimally competent examinee and, for each item, estimates the probability that such an examinee would answer it correctly. Summing those probabilities yields a recommended cut score for each panelist, and averaging across panelists and discussion rounds produces the performance standard. It is among the most widely used standard-setting methods in licensure, certification, and K-12 testing.
Key highlights
- Conceptually transparent and legally defensible, with a long track record in high-stakes licensure and certification.
- Decomposes a single high-stakes judgment into many tractable item-level judgments grounded in test content.
- Flexible: the original probability and the simpler Yes/No variants accommodate different panel capacities.
- Naturally incorporates normative and impact feedback across rounds to improve panelist calibration and consensus.
Intuition
This section is available to Pro members. Upgrade to Pro
How it works
This section is available to Pro members. Upgrade to Pro
When to use it
Use the Angoff method when a defensible cut score is needed for a dichotomously scored test — licensure and certification exams, end-of-course tests, and proficiency classifications — and a panel of qualified content experts is available. It is most appropriate for multiple-choice and other right/wrong items, where item-level probability judgments are meaningful. It is less suited to performance assessments, essays, or polytomous items, where examinee-centered or holistic methods (e.g., Body of Work) or the Bookmark method may fit better. Strong facilitation, training, and feedback rounds are essential to defensibility.
Strengths & limitations
- Conceptually transparent and legally defensible, with a long track record in high-stakes licensure and certification.
- Decomposes a single high-stakes judgment into many tractable item-level judgments grounded in test content.
- Flexible: the original probability and the simpler Yes/No variants accommodate different panel capacities.
- Naturally incorporates normative and impact feedback across rounds to improve panelist calibration and consensus.
- Panelists are notoriously poor at estimating item difficulty for the borderline examinee, biasing probability judgments.
- Cut scores can be sensitive to panel composition, training quality, and the borderline conceptualization.
- Designed for dichotomously scored items; extensions to polytomous or performance items are awkward.
- Treats items as independent, ignoring local dependence and the actual conditional probabilities implied by examinee data.
Common pitfalls
This section is available to Pro members. Upgrade to Pro
Applications
This section is available to Pro members. Upgrade to Pro
Frequently asked
What is the difference between the original Angoff and the modified (Yes/No) Angoff?
In the original method, panelists estimate a continuous probability (0 to 1) that the borderline examinee answers each item correctly. In the Yes/No modification by Impara and Plake, panelists make a simpler binary judgment — would the borderline examinee get this item right or not — and the cut score is the count of 'yes' items. The binary version reduces cognitive load and is easier for panelists, though it discards the finer gradation of probability judgments.
How does the Angoff method compare to the Bookmark method?
Both set cut scores via expert panels, but they differ in materials and judgment. Angoff asks for item-by-item probability judgments and works directly with raw test content. The Bookmark method presents items ordered by IRT difficulty in an ordered item booklet and asks panelists to place a 'bookmark' where the borderline examinee's mastery ends. Bookmark scales better to large item pools and polytomous items; Angoff is simpler conceptually but more demanding per item. See the related Bookmark Standard Setting entry.
How many panelists and rounds are needed?
There is no universal rule, but operational panels commonly use 10 to 20 content experts and two or three rating rounds. Multiple rounds with normative and impact feedback let panelists recalibrate and converge. More important than exact counts are panelist qualifications, representativeness, thorough training on the borderline examinee, and documented evaluation of the procedure's consistency and validity.
Sources
- 1.Cizek, G. J., & Bunch, M. B. (2007). Standard Setting: A Guide to Establishing and Evaluating Performance Standards on Tests. Sage.ISBN 9781412916820
- 2.Angoff, W. H. (1971). Scales, norms, and equivalent scores. In R. L. Thorndike (Ed.), Educational Measurement (2nd ed., pp. 508–600). American Council on Education.ISBN 9780827230309
You have read it. What now?
Cite this page
ScholarGate. (2026, June 22). Angoff Standard Setting. ScholarGate. https://scholargate.app/education/angoff-standard-setting