Process / pipelineHistorical DemographyRecord-linkagePipeline

Historical Nominal Record Linkage

Also known as: Record linkage, Census linking, Fellegi-Sunter matching, Historical individual linkage

OriginatorIvan Fellegi and Alan Sunter (probabilistic theory); James Feigenbaum, Ran Abramitzky, Leah Boustan (historical ML methods)Year2016Sources2Related methods9

Historical nominal record linkage is the task of recognising when records in different sources, two censuses, a census and a draft register, a baptism and a marriage, refer to the same person, even though no shared identifier exists and names are misspelled, ages misreported, and places renamed. Linkage is the engine behind longitudinal historical micro-data: it builds the life-course panels that underpin studies of migration, mobility, mortality, and the long-run effects of early-life conditions. Three families of methods dominate. Deterministic linkage applies hand-crafted rules; the probabilistic Fellegi-Sunter framework weights field agreements and disagreements by their discriminating power; and supervised machine learning, trained on hand-linked examples, learns to classify candidate pairs. Modern historical practice, led by Abramitzky, Boustan, Feigenbaum, and collaborators, emphasises transparent, replicable algorithms and, crucially, explicit measurement of linkage error, since false matches and missed links can bias every downstream estimate.

Key highlights

  • Builds longitudinal panels from otherwise cross-sectional historical sources
  • Fellegi-Sunter weighting principled in rewarding rare, discriminating agreements
  • Machine-learning classifiers exploit complex patterns and scale to millions of records
  • Modern practice quantifies linkage error explicitly, enabling bias correction

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use nominal record linkage whenever a research question requires following the same individuals across separate historical sources, to build life-course or intergenerational panels for studies of migration, social mobility, mortality, fertility, or the long-run effects of early-life exposures. Probabilistic and machine-learning methods are preferred at scale and when transparent error measurement is needed; simple deterministic rules suffice only for small, clean, high-quality sources. Always pair linkage with evaluation against a labelled sample and with analysis of selection into the linked set. Avoid relying on linked data without quantifying linkage error, since false matches and selective non-linkage can severely bias every downstream estimate.

Strengths & limitations

Strengths
  • Builds longitudinal panels from otherwise cross-sectional historical sources
  • Fellegi-Sunter weighting principled in rewarding rare, discriminating agreements
  • Machine-learning classifiers exploit complex patterns and scale to millions of records
  • Modern practice quantifies linkage error explicitly, enabling bias correction
Limitations
  • Common names and sparse identifiers make many individuals unlinkable
  • Blocking can discard true matches split across blocks
  • Linked samples are selective, over-representing stable, distinctively named people
  • Requires labelled training data for probabilistic or ML calibration

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

Deterministic, probabilistic, or machine learning, which is best?

It depends on data quality and scale. Deterministic rules are transparent and adequate for small, clean sources but brittle to error. The probabilistic Fellegi-Sunter model handles noise principledly and is well understood. Supervised machine learning typically achieves the best precision-recall trade-off at scale but needs hand-linked training data. Modern historical practice often combines them and, regardless of choice, insists on measuring linkage error explicitly.

Why does linkage error matter for results?

False matches link different people, injecting noise and bias, while selective non-linkage means the linked sample is unrepresentative, typically over-representing people with stable, distinctive names who did not migrate. Both distort downstream estimates of mobility or migration. Reporting precision and recall on a labelled sample, and analysing who gets linked, lets researchers bound or correct these biases instead of presenting linked-data results as if exact.

What is blocking and why is it necessary?

Blocking partitions records into groups, by name sound, birth decade, sex, within which candidate pairs are compared, avoiding the impossible task of comparing every record to every other across millions of entries. It makes linkage computationally feasible but risks discarding true matches that fall into different blocks. Designers therefore choose blocking keys that are both stable across sources and discriminating, sometimes running several blocking passes to recover otherwise lost links.

Sources

  1. 1.
    Abramitzky, R., Boustan, L., Eriksson, K., Feigenbaum, J., & Perez, S. (2021). Automated Linking of Historical Data. Journal of Economic Literature, 59(3), 865-918.
  2. 2.
    Feigenbaum, J. J. (2016). Automated Census Record Linking: A Machine Learning Approach. Working paper, Boston University.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 23). Historical Nominal Record Linkage. ScholarGate. https://scholargate.app/historical-demography/nominal-record-linkage-historical