Latent structureEducationScaling and linkingModel

Vertical Scaling

Also known as: Developmental Scaling, Vertical Linking, Cross-Grade Scaling, Growth Scale Construction

OriginatorEducational measurement tradition (Thurstone; Kolen & Brennan synthesis)Year2014Sources2Related methods5

Vertical scaling places tests written for different grade levels onto a single continuous score scale so that growth from one grade to the next can be measured in common units. Unlike horizontal equating, which links alternate forms intended to be interchangeable, vertical scaling deliberately links tests of differing difficulty and content to build a developmental continuum spanning, for example, grades 3 through 8. It is the measurement foundation that lets a fourth-grade and a fifth-grade score be subtracted to express how much a student grew.

Key highlights

  • Creates a single developmental scale on which year-to-year academic growth can be quantified directly.
  • Provides the measurement substrate for growth percentiles, value-added models, and longitudinal reporting.
  • Leverages IRT to handle tests of differing difficulty and length while preserving a common metric.
  • Supports flexible linking designs (common item or common person) that fit operational testing constraints.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use vertical scaling whenever scores from tests built for different grades must be compared on a common metric — most importantly to measure academic growth across grades, to support growth and value-added models, and to report progress on a single developmental scale. It is not appropriate for linking tests of fundamentally different constructs, and it is distinct from horizontal equating, which links interchangeable forms of the same-grade test. Because grade-level content changes, vertical scales rest on the assumption that a single underlying construct runs through the grades, an assumption that weakens when curricula diverge sharply.

Strengths & limitations

Strengths
  • Creates a single developmental scale on which year-to-year academic growth can be quantified directly.
  • Provides the measurement substrate for growth percentiles, value-added models, and longitudinal reporting.
  • Leverages IRT to handle tests of differing difficulty and length while preserving a common metric.
  • Supports flexible linking designs (common item or common person) that fit operational testing constraints.
Limitations
  • Rests on a strong unidimensionality assumption that one construct spans all grades, which is questionable when grade content shifts substantially.
  • Resulting growth patterns are sensitive to scaling method (concurrent vs. separate calibration, choice of linking), so the metric is partly a modeling artifact.
  • Anchor-item performance can drift across grades, and weak or unrepresentative anchors bias the linkage.
  • Interval-scale interpretation of growth (e.g., 'equal growth means equal learning') is difficult to justify substantively.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

How is vertical scaling different from test equating?

Equating links alternate forms of a test built to the same specifications and difficulty so the forms are interchangeable; the linked scores are meant to be used as if from one form. Vertical scaling links tests deliberately built to differ in difficulty and content across grades, producing a developmental scale on which growth can be measured. Equated forms are exchangeable; vertically scaled grade tests are not. See the related Test Equating entry.

Why do growth curves from vertical scales often flatten in higher grades?

Across many vertical scales, average grade-to-grade gains shrink in higher grades. This deceleration is partly substantive but is also strongly influenced by scaling choices — the linking method, the calibration approach, and the spread of item difficulties. Because the pattern can change with the method, decelerating growth should be interpreted cautiously rather than as a hard fact about learning.

What linking design should I use?

Common-item (anchor) designs embed shared items across adjacent grades and are operationally convenient. Common-person (scaling test) designs have some students take items spanning grades, which can give cleaner links but adds testing burden. The choice depends on testing logistics, the strength and representativeness of available anchors, and whether concurrent or separate calibration is planned.

Sources

  1. 1.
    Kolen, M. J., & Brennan, R. L. (2014). Test Equating, Scaling, and Linking: Methods and Practices (3rd ed.). Springer.
    ISBN 9781493903160
  2. 2.
    Yen, W. M., & Fitzpatrick, A. R. (2006). Item Response Theory. In R. L. Brennan (Ed.), Educational Measurement (4th ed., pp. 111–153). American Council on Education / Praeger.
    ISBN 9780275981259

You have read it. What now?

Cite this page

ScholarGate. (2026, June 22). Vertical Scaling. ScholarGate. https://scholargate.app/education/vertical-scaling