Vertical Scaling
Also known as: Developmental Scaling, Vertical Linking, Cross-Grade Scaling, Growth Scale Construction
Vertical scaling places tests written for different grade levels onto a single continuous score scale so that growth from one grade to the next can be measured in common units. Unlike horizontal equating, which links alternate forms intended to be interchangeable, vertical scaling deliberately links tests of differing difficulty and content to build a developmental continuum spanning, for example, grades 3 through 8. It is the measurement foundation that lets a fourth-grade and a fifth-grade score be subtracted to express how much a student grew.
Key highlights
- Creates a single developmental scale on which year-to-year academic growth can be quantified directly.
- Provides the measurement substrate for growth percentiles, value-added models, and longitudinal reporting.
- Leverages IRT to handle tests of differing difficulty and length while preserving a common metric.
- Supports flexible linking designs (common item or common person) that fit operational testing constraints.
Intuition
This section is available to Pro members. Upgrade to Pro
How it works
This section is available to Pro members. Upgrade to Pro
When to use it
Use vertical scaling whenever scores from tests built for different grades must be compared on a common metric — most importantly to measure academic growth across grades, to support growth and value-added models, and to report progress on a single developmental scale. It is not appropriate for linking tests of fundamentally different constructs, and it is distinct from horizontal equating, which links interchangeable forms of the same-grade test. Because grade-level content changes, vertical scales rest on the assumption that a single underlying construct runs through the grades, an assumption that weakens when curricula diverge sharply.
Strengths & limitations
- Creates a single developmental scale on which year-to-year academic growth can be quantified directly.
- Provides the measurement substrate for growth percentiles, value-added models, and longitudinal reporting.
- Leverages IRT to handle tests of differing difficulty and length while preserving a common metric.
- Supports flexible linking designs (common item or common person) that fit operational testing constraints.
- Rests on a strong unidimensionality assumption that one construct spans all grades, which is questionable when grade content shifts substantially.
- Resulting growth patterns are sensitive to scaling method (concurrent vs. separate calibration, choice of linking), so the metric is partly a modeling artifact.
- Anchor-item performance can drift across grades, and weak or unrepresentative anchors bias the linkage.
- Interval-scale interpretation of growth (e.g., 'equal growth means equal learning') is difficult to justify substantively.
Common pitfalls
This section is available to Pro members. Upgrade to Pro
Applications
This section is available to Pro members. Upgrade to Pro
Frequently asked
How is vertical scaling different from test equating?
Equating links alternate forms of a test built to the same specifications and difficulty so the forms are interchangeable; the linked scores are meant to be used as if from one form. Vertical scaling links tests deliberately built to differ in difficulty and content across grades, producing a developmental scale on which growth can be measured. Equated forms are exchangeable; vertically scaled grade tests are not. See the related Test Equating entry.
Why do growth curves from vertical scales often flatten in higher grades?
Across many vertical scales, average grade-to-grade gains shrink in higher grades. This deceleration is partly substantive but is also strongly influenced by scaling choices — the linking method, the calibration approach, and the spread of item difficulties. Because the pattern can change with the method, decelerating growth should be interpreted cautiously rather than as a hard fact about learning.
What linking design should I use?
Common-item (anchor) designs embed shared items across adjacent grades and are operationally convenient. Common-person (scaling test) designs have some students take items spanning grades, which can give cleaner links but adds testing burden. The choice depends on testing logistics, the strength and representativeness of available anchors, and whether concurrent or separate calibration is planned.
Sources
- 1.Kolen, M. J., & Brennan, R. L. (2014). Test Equating, Scaling, and Linking: Methods and Practices (3rd ed.). Springer.ISBN 9781493903160
- 2.Yen, W. M., & Fitzpatrick, A. R. (2006). Item Response Theory. In R. L. Brennan (Ed.), Educational Measurement (4th ed., pp. 111–153). American Council on Education / Praeger.ISBN 9780275981259
You have read it. What now?
Cite this page
ScholarGate. (2026, June 22). Vertical Scaling. ScholarGate. https://scholargate.app/education/vertical-scaling