Test Equating
Test Equating, Scaling, and Linking · Also known as: Test Eşitleme (Test Equating), score equating, equipercentile equating, IRT true-score equating, linear equating
Test equating is a family of statistical methods that converts scores earned on one test form onto the score scale of another form, so that scores from different administrations or versions can be compared and reported on a common metric. The foundational modern treatment is Kolen and Brennan (2004/2014); Holland and Dorans (2006) provide the authoritative chapter-length overview within the field of educational measurement.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Test equating is appropriate when scores from two or more test forms must be compared or reported on a single scale — typically in large-scale standardised testing programmes (national exams, licensure tests, admissions tests) that produce new forms each year. Four conditions are required. First, both forms must measure the same construct; equating scores from conceptually different tests is called linking or concordance, not equating, and yields weaker comparability guarantees. Second, the sample must be large enough for stable distribution estimation — at minimum around 300 examinees per form, and substantially more for IRT-based equating. Third, the data-collection design must support equating: single-group and random-groups designs allow direct distribution comparison; the NEAT (Non-Equivalent groups with Anchor Test) design, the most common operational design, requires an anchor set of common items present in both forms to control for group ability differences. Fourth, for IRT equating, each form must first be calibrated with a fitting IRT model (Rasch, 2PL, or 3PL) and the model fit must be adequate.
Strengths & limitations
- Provides a rigorous, defensible basis for comparing scores across test forms administered at different times, enabling fair score reporting and longitudinal tracking of performance.
- Equipercentile equating makes no parametric assumptions about the score distributions, making it robust to non-normality.
- IRT true-score equating is the most theoretically sound method when an adequate IRT model fits both forms, because it operates on the latent ability scale rather than on observed raw scores.
- Requires large samples — at minimum 300 per form for equipercentile equating, and considerably more for stable IRT parameter estimation — limiting applicability to small-scale assessments.
- The NEAT design assumes that the anchor items function identically in both groups (anchor item invariance); violations from differential item functioning or context effects can bias the equating.
- Equipercentile equating in the score-distribution tails is unstable without smoothing, and the choice of smoothing method introduces subjective decisions.
Frequently asked
What is the difference between equating and linking?
Equating is the technically strongest form of score comparability: it applies only when both forms measure the same construct to the same degree, are built to the same content and statistical specifications, and are administered under equivalent conditions. Linking (or concordance) is a broader term covering any score transformation between tests that do not fully meet those conditions — for example, placing scores from two different admissions tests on a common scale. Equating supports stronger comparability claims; linking supports weaker, context-dependent ones.
What is the NEAT design and why is it widely used?
NEAT stands for Non-Equivalent groups with Anchor Test. In operational testing it is rarely possible to give the same examinees two complete forms or to assign examinees randomly to forms. Instead, a set of anchor (common) items is embedded in both forms and administered to the different groups. The anchor items provide the statistical bridge needed to separate true form-difficulty differences from group-ability differences. Most large-scale testing programmes use the NEAT design because it is practical and accommodates non-equivalent examinee groups across administrations.
When should IRT true-score equating be preferred over equipercentile equating?
IRT true-score equating is preferred when an IRT model fits both forms adequately and the sample is large enough for stable item parameter estimation (often 500 or more per form). It is theoretically superior because it works at the level of the latent ability scale, handles missing data naturally in a computer-adaptive context, and produces smoother concordance tables. Equipercentile equating is preferred when IRT model fit is questionable, when the sample is moderate in size, or when a simpler and more transparent method is needed for stakeholder communication.
How large a sample does test equating require?
A common operational guideline is at least 300 examinees per form for equipercentile equating with smoothing. IRT-based equating typically requires at least 500 examinees per form for the 3PL model and somewhat fewer for the Rasch or 2PL models. Below these thresholds, sampling error in the score distributions or item parameter estimates makes the concordance table unreliable for high-stakes decisions.
Sources
- Kolen, M.J. & Brennan, R.L. (2014). Test Equating, Scaling, and Linking: Methods and Practices (3rd ed.). Springer. ISBN: 978-1-4939-0316-6
- Holland, P.W. & Dorans, N.J. (2006). Linking and Equating. In R.L. Brennan (Ed.), Educational Measurement (4th ed., pp. 187–220). American Council on Education / Praeger. link ↗
How to cite this page
ScholarGate. (2026, June 1). Test Equating, Scaling, and Linking. ScholarGate. https://scholargate.app/en/psychometrics/test-equating
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Confirmatory factor analysisPsychometrics↔ compare
- Generalizability TheoryPsychometrics↔ compare
- Item Response TheoryPsychometrics↔ compare
- Rasch ModelPsychometrics↔ compare