Rasch Analysis of Disability Measures
Also known as: Rasch Measurement Model for Disability, Rasch Analysis of Outcome Measures, Rasch Modeling in Rehabilitation, Rasch Calibration of Disability Scales
Rasch analysis is a psychometric method, based on Georg Rasch's probabilistic measurement model, used to test and refine the disability, function, and participation scales that pervade disability and rehabilitation research. As set out for clinicians by Alan Tennant and Philip Conaghan in 2007, fitting the Rasch model checks whether a scale's items genuinely measure a single underlying trait at interval level, so that summing item scores into a total is justified. Because so many disability outcome measures simply add ordinal item ratings — assuming items are equally difficult and that ordinal categories behave like interval data — Rasch analysis provides the rigorous test of whether that common practice is actually valid.
Key highlights
- Tests whether summing ordinal items into a total score is actually justified, rather than assuming it.
- Converts ordinal scores to interval-level logit measures, legitimizing change scores and parametric analysis.
- Detects misfitting items, disordered response categories, multidimensionality, and local dependence.
- Identifies differential item functioning, supporting fair comparison across groups, cultures, and conditions.
Intuition
This section is available to Pro members. Upgrade to Pro
How it works
This section is available to Pro members. Upgrade to Pro
When to use it
Use Rasch analysis when developing, validating, refining, or scoring a disability, function, or participation scale and you need to know whether its items form a unidimensional, interval-level measure that can be fairly summed and compared across groups. It is central to outcome-measure development in rehabilitation, to checking instruments for differential item functioning across cultures or conditions, and to converting ordinal scores to interval estimates for analysis. It is less appropriate when the construct is deliberately multidimensional and not meant to yield one score, when the sample is too small for stable estimation, or when only a quick reliability check rather than fundamental measurement evaluation is required.
Strengths & limitations
- Tests whether summing ordinal items into a total score is actually justified, rather than assuming it.
- Converts ordinal scores to interval-level logit measures, legitimizing change scores and parametric analysis.
- Detects misfitting items, disordered response categories, multidimensionality, and local dependence.
- Identifies differential item functioning, supporting fair comparison across groups, cultures, and conditions.
- Requires adequate sample sizes for stable item and person estimates.
- Demands psychometric expertise and specialized software to run and interpret correctly.
- Strict model can reject useful items, and decisions about removing or splitting items involve judgment.
- Focuses on a single latent trait, so genuinely multidimensional constructs need separate analyses.
Common pitfalls
This section is available to Pro members. Upgrade to Pro
Applications
This section is available to Pro members. Upgrade to Pro
Frequently asked
Why not just add up the items on a disability scale?
Because adding ordinal items assumes that all items are equally difficult and that response categories like none-mild-moderate-severe are evenly spaced, assumptions that are rarely true and seldom tested. A summed total can then distort measurement, making a given change mean different things in different parts of the scale. Rasch analysis tests whether the items actually conform to a single interval-level trait and, when they do, provides a conversion from the raw total to an interval estimate, so that summing remains convenient but the resulting score is measurement-justified rather than merely an index.
What is differential item functioning and why does it matter for disability measures?
Differential item functioning (DIF) occurs when an item has a different difficulty for two groups who are otherwise at the same level of the underlying trait — for example, a self-care item that is systematically harder to endorse for one gender, diagnosis, or culture. DIF biases comparisons: differences in total scores may reflect item bias rather than true differences in disability. Detecting DIF is essential when a measure is used across groups or translated for cross-cultural research, and Rasch analysis can flag biased items so they can be split, removed, or interpreted with caution.
What does it mean to convert ordinal scores to interval level?
Ordinal scores tell you only the order of categories, not the distances between them, so arithmetic on them — averaging, computing change scores, running parametric tests — is technically invalid. When data fit the Rasch model, the model places persons and items on a common interval scale measured in logits, where equal differences mean equal amounts of the trait. This transformation, often delivered as a raw-score-to-logit conversion table, lets clinicians keep summing items while analysts treat the converted measure as interval-level, which is the practical reason Rasch analysis is valued in outcome measurement.
Sources
- 1.Tennant, A., & Conaghan, P. G. (2007). The Rasch measurement model in rheumatology: What is it and why use it? When should it be applied, and what should one look for in a Rasch paper? Arthritis Care & Research, 57(8), 1358-1362.
You have read it. What now?
Cite this page
ScholarGate. (2026, June 23). Rasch Analysis of Disability Measures. ScholarGate. https://scholargate.app/disability-studies/rasch-analysis-disability-measures