Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Psychometrics›Test Equating
Latent structure

Test Equating

Test Equating, Scaling, and Linking · Also known as: Test Eşitleme (Test Equating), score equating, equipercentile equating, IRT true-score equating, linear equating

Test equating is a family of statistical methods that converts scores earned on one test form onto the score scale of another form, so that scores from different administrations or versions can be compared and reported on a common metric. The foundational modern treatment is Kolen and Brennan (2004/2014); Holland and Dorans (2006) provide the authoritative chapter-length overview within the field of educational measurement.

ScholarGate
  1. Latent structure
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Test Equating
Confirmatory factor anal…Generalizability TheoryItem Response TheoryRasch ModelAngoff Standard SettingConditional Standard Err…Item AnalysisStandardized Test Analys…Vertical Scaling

When to use it

Test equating is appropriate when scores from two or more test forms must be compared or reported on a single scale — typically in large-scale standardised testing programmes (national exams, licensure tests, admissions tests) that produce new forms each year. Four conditions are required. First, both forms must measure the same construct; equating scores from conceptually different tests is called linking or concordance, not equating, and yields weaker comparability guarantees. Second, the sample must be large enough for stable distribution estimation — at minimum around 300 examinees per form, and substantially more for IRT-based equating. Third, the data-collection design must support equating: single-group and random-groups designs allow direct distribution comparison; the NEAT (Non-Equivalent groups with Anchor Test) design, the most common operational design, requires an anchor set of common items present in both forms to control for group ability differences. Fourth, for IRT equating, each form must first be calibrated with a fitting IRT model (Rasch, 2PL, or 3PL) and the model fit must be adequate.

Strengths & limitations

Strengths
  • Provides a rigorous, defensible basis for comparing scores across test forms administered at different times, enabling fair score reporting and longitudinal tracking of performance.
  • Equipercentile equating makes no parametric assumptions about the score distributions, making it robust to non-normality.
  • IRT true-score equating is the most theoretically sound method when an adequate IRT model fits both forms, because it operates on the latent ability scale rather than on observed raw scores.
Limitations
  • Requires large samples — at minimum 300 per form for equipercentile equating, and considerably more for stable IRT parameter estimation — limiting applicability to small-scale assessments.
  • The NEAT design assumes that the anchor items function identically in both groups (anchor item invariance); violations from differential item functioning or context effects can bias the equating.
  • Equipercentile equating in the score-distribution tails is unstable without smoothing, and the choice of smoothing method introduces subjective decisions.

Frequently asked

What is the difference between equating and linking?

Equating is the technically strongest form of score comparability: it applies only when both forms measure the same construct to the same degree, are built to the same content and statistical specifications, and are administered under equivalent conditions. Linking (or concordance) is a broader term covering any score transformation between tests that do not fully meet those conditions — for example, placing scores from two different admissions tests on a common scale. Equating supports stronger comparability claims; linking supports weaker, context-dependent ones.

What is the NEAT design and why is it widely used?

NEAT stands for Non-Equivalent groups with Anchor Test. In operational testing it is rarely possible to give the same examinees two complete forms or to assign examinees randomly to forms. Instead, a set of anchor (common) items is embedded in both forms and administered to the different groups. The anchor items provide the statistical bridge needed to separate true form-difficulty differences from group-ability differences. Most large-scale testing programmes use the NEAT design because it is practical and accommodates non-equivalent examinee groups across administrations.

When should IRT true-score equating be preferred over equipercentile equating?

IRT true-score equating is preferred when an IRT model fits both forms adequately and the sample is large enough for stable item parameter estimation (often 500 or more per form). It is theoretically superior because it works at the level of the latent ability scale, handles missing data naturally in a computer-adaptive context, and produces smoother concordance tables. Equipercentile equating is preferred when IRT model fit is questionable, when the sample is moderate in size, or when a simpler and more transparent method is needed for stakeholder communication.

How large a sample does test equating require?

A common operational guideline is at least 300 examinees per form for equipercentile equating with smoothing. IRT-based equating typically requires at least 500 examinees per form for the 3PL model and somewhat fewer for the Rasch or 2PL models. Below these thresholds, sampling error in the score distributions or item parameter estimates makes the concordance table unreliable for high-stakes decisions.

Sources

  1. Kolen, M.J. & Brennan, R.L. (2014). Test Equating, Scaling, and Linking: Methods and Practices (3rd ed.). Springer. ISBN: 978-1-4939-0316-6
  2. Holland, P.W. & Dorans, N.J. (2006). Linking and Equating. In R.L. Brennan (Ed.), Educational Measurement (4th ed., pp. 187–220). American Council on Education / Praeger. link ↗

How to cite this page

ScholarGate. (2026, June 1). Test Equating, Scaling, and Linking. ScholarGate. https://scholargate.app/en/psychometrics/test-equating

Related methods

Confirmatory factor analysisGeneralizability TheoryItem Response TheoryRasch Model

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Confirmatory factor analysisPsychometrics↔ compare
  • Generalizability TheoryPsychometrics↔ compare
  • Item Response TheoryPsychometrics↔ compare
  • Rasch ModelPsychometrics↔ compare
Compare side by side →

Referenced by

Angoff Standard SettingConditional Standard Error of MeasurementItem AnalysisStandardized Test AnalysisVertical Scaling

Similar methods

Vertical ScalingStandardized Test AnalysisMulti-group item response theoryComputerized adaptive test measurement invarianceMulti-group Rasch modelItem Response TheoryComputerized adaptive test item response theoryMulti-group measurement invariance

Related reference concepts

Equated ScoresMeasurementEducational MeasurementItem Response TheoryStandard Setting (Scoring)Educational Testing

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Test Equating (Test Equating, Scaling, and Linking). Retrieved 2026-07-21 from https://scholargate.app/en/psychometrics/test-equating · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Kolen & Brennan (foundational treatise, 2004/2014); Holland & Dorans (2006)
Year
1984 (modern statistical treatment)
Type
Score transformation / latent-scale calibration
Outcome
Concordance table mapping raw scores on one test form to the score scale of another
Data
Raw or IRT-scaled scores from two or more test forms
Min Sample
300
Design
NEAT (Non-Equivalent groups with Anchor Test) or single-group / random-groups
Difficulty
3
Related methods
Confirmatory factor analysisGeneralizability TheoryItem Response TheoryRasch Model
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account