Many-Facet Rasch Measurement
Also known as: MFRM, Many-Faceted Rasch Model, Facets Model, Linacre Facets Model
Many-facet Rasch measurement (MFRM) extends the basic Rasch model to assessments mediated by raters. Beyond examinee ability and item difficulty, it adds explicit parameters for rater severity and for any other facet of the rating situation — task, occasion, rating criterion — placing them all on one common logit scale. Developed by John Michael Linacre, MFRM lets analysts estimate and adjust for the fact that some raters are systematically harsh and others lenient, producing 'fair' ability estimates that do not penalize an examinee for happening to draw a severe judge.
Key highlights
- Estimates and adjusts for individual rater severity, yielding fair examinee measures despite uneven rater assignment.
- Places persons, items, raters, and other facets on one interpretable logit scale for direct comparison.
- Provides rater-level fit statistics for ongoing monitoring of rating quality and detection of erratic judges.
- Handles incomplete, connected rating designs efficiently, requiring far fewer ratings than fully crossed designs.
Intuition
This section is available to Pro members. Upgrade to Pro
How it works
This section is available to Pro members. Upgrade to Pro
When to use it
Use many-facet Rasch measurement whenever scores depend on human raters and you need to separate examinee ability from rater severity and other facets — writing assessments, oral and clinical examinations, performance tasks, sports judging, and rated portfolios. It requires a connected (linked) rating design so that the facets are jointly estimable; fully nested designs where every examinee is rated by a unique rater cannot identify severity separately from ability. It complements generalizability theory: G-theory estimates variance components attributable to facets, while MFRM models and adjusts for individual elements within each facet.
Strengths & limitations
- Estimates and adjusts for individual rater severity, yielding fair examinee measures despite uneven rater assignment.
- Places persons, items, raters, and other facets on one interpretable logit scale for direct comparison.
- Provides rater-level fit statistics for ongoing monitoring of rating quality and detection of erratic judges.
- Handles incomplete, connected rating designs efficiently, requiring far fewer ratings than fully crossed designs.
- Requires a connected rating design; disconnected subsets cannot be placed on the same scale without anchors.
- Assumes the modeled facets act additively and that severity is stable, missing rater-by-examinee interactions and drift unless explicitly modeled.
- As a Rasch model, it imposes equal item discrimination, which can misfit data where discrimination genuinely varies.
- Adjustment for severity assumes severity is correctly estimated; sparse or biased designs undermine the fair scores.
Common pitfalls
This section is available to Pro members. Upgrade to Pro
Applications
This section is available to Pro members. Upgrade to Pro
Frequently asked
How does many-facet Rasch measurement differ from generalizability theory?
Both address rater and task effects, but at different levels. Generalizability theory partitions score variance into components attributable to facets (raters, tasks, occasions) and estimates how reliable scores are under various designs. MFRM goes further by estimating a parameter for each individual element — this specific rater's severity, this specific task's difficulty — and adjusting examinee measures accordingly. G-theory characterizes the facet as a whole; MFRM models the elements within it. They are complementary, not competing. See the related Generalizability Theory entry.
What is a 'fair score' in MFRM?
A fair score (or fair average) is the examinee ability estimate adjusted to what it would have been under raters and tasks of average severity and difficulty. Because examinees in realistic designs are not all judged by the same raters, raw averages are not comparable; the fair score corrects for the particular, possibly harsher or easier, raters each examinee happened to encounter, making cross-examinee comparison defensible.
Why does the rating design need to be 'connected'?
To place all raters on a common severity scale, the design must link them — for example, by having overlapping raters score common performances or anchor tasks. If the data split into disconnected subsets with no shared elements, the model cannot tell whether one subset scored low because its examinees were weaker or its raters harsher. A connected design ensures every facet parameter is estimable on the same scale.
Sources
- 1.Linacre, J. M. (1989). Many-Facet Rasch Measurement. MESA Press.ISBN 9780941938020
- 2.Engelhard, G., & Wind, S. A. (2018). Invariant Measurement with Raters and Rating Scales: Rasch Models for Rater-Mediated Assessments. Routledge.ISBN 9781848725812
You have read it. What now?
Cite this page
ScholarGate. (2026, June 22). Many-Facet Rasch Measurement. ScholarGate. https://scholargate.app/education/many-facet-rasch-measurement