Process / pipelineEducationStandard settingPipeline

Bookmark Standard Setting

Also known as: Bookmark Method, Bookmark Procedure, Item Mapping Standard Setting, Ordered Item Booklet Method

OriginatorHoward Mitzel, Daniel Lewis, Richard Patz & Donald Ross Green (CTB/McGraw-Hill)Year2001Sources2Related methods4

The Bookmark method is an item-response-theory-based standard-setting procedure in which test items are arranged in a booklet ordered from easiest to hardest. Panelists page through this ordered item booklet and place a 'bookmark' at the point separating items a borderline examinee would likely master from those they would not, judged against a fixed response probability (commonly two-thirds). The latent ability at the bookmark defines the cut score. Developed at CTB/McGraw-Hill, it became one of the dominant methods for large-scale K-12 assessments.

Key highlights

  • Scales efficiently to large item pools and naturally sets multiple performance-level cuts in one ordered booklet.
  • Handles mixed-format tests, ordering multiple-choice and polytomous score points on a common difficulty scale.
  • Reduces panelist burden: one bookmark judgment per cut replaces hundreds of item-by-item probability estimates.
  • Grounded in the operational IRT scale, so cut scores map directly back to the reporting metric.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use the Bookmark method when items are already calibrated on an IRT scale, the item pool is large, multiple performance levels (e.g., Basic, Proficient, Advanced) must be set at once, or the test mixes multiple-choice and constructed-response items. It is the workhorse of many statewide K-12 standard-setting efforts because it handles big pools and polytomous items gracefully. It requires a defensible IRT calibration and a justified response-probability criterion; when items are not IRT-calibrated, or for small credentialing exams, the Angoff method may be simpler and equally defensible.

Strengths & limitations

Strengths
  • Scales efficiently to large item pools and naturally sets multiple performance-level cuts in one ordered booklet.
  • Handles mixed-format tests, ordering multiple-choice and polytomous score points on a common difficulty scale.
  • Reduces panelist burden: one bookmark judgment per cut replaces hundreds of item-by-item probability estimates.
  • Grounded in the operational IRT scale, so cut scores map directly back to the reporting metric.
Limitations
  • Depends entirely on the quality and fit of the IRT calibration; poor calibration propagates into the standards.
  • Cut scores shift with the chosen response-probability criterion, a consequential but somewhat arbitrary decision.
  • Items of similar difficulty cluster, so a single bookmark move can jump the cut score substantially.
  • Panelists may misunderstand the RP-based 'mastery' notion, conflating it with simply getting an item right.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

What is the response probability (RP) criterion and why is it usually 0.67?

The RP criterion is the probability of a correct response that defines 'mastery' for placing items in the ordered booklet. A value of two-thirds (0.67) is conventional because it corresponds to a reasonable level of expected success and aligns with item-mapping practices, but values such as 0.5 or 0.80 are also used. The choice matters: a higher RP places items at higher abilities, shifting all cut scores upward, so it should be set deliberately and documented.

How does the Bookmark method handle constructed-response items?

Each score point of a polytomous item is treated as a separate 'item step' and assigned its own location on the latent scale using the RP criterion. These step locations are interleaved with multiple-choice item locations in the same ordered booklet, so a single bookmark can span both formats. This is a key reason Bookmark is favored for mixed-format statewide tests.

Why can the cut score jump when a panelist moves the bookmark only a few pages?

Items tend to cluster at similar difficulties, so several booklet pages can share nearly the same location while a few pages span a wide difficulty gap. Moving the bookmark across a sparse region of the scale changes the cut θ a lot, while moving it within a dense cluster barely changes it. This nonuniform spacing is why item-location maps and impact data are reviewed during feedback rounds.

Sources

  1. 1.
    Cizek, G. J., & Bunch, M. B. (2007). Standard Setting: A Guide to Establishing and Evaluating Performance Standards on Tests. Sage.
    ISBN 9781412916820
  2. 2.
    Mitzel, H. C., Lewis, D. M., Patz, R. J., & Green, D. R. (2001). The bookmark procedure: Psychological perspectives. In G. J. Cizek (Ed.), Setting Performance Standards: Concepts, Methods, and Perspectives (pp. 249–281). Lawrence Erlbaum.
    ISBN 9780805835586

You have read it. What now?

Cite this page

ScholarGate. (2026, June 22). Bookmark Standard Setting. ScholarGate. https://scholargate.app/education/bookmark-standard-setting