Process / pipelineLibrary Information ScienceLibrary collection assessmentPipeline

Collection Overlap Analysis

Also known as: Collection Overlap Study, Holdings Overlap Analysis, Title Duplication Analysis, Collection Comparison Analysis

OriginatorLibrary collection-management literature; Thomas E. Nisonger (synthesis)Year1998Sources2Related methods4

Collection overlap analysis measures the degree to which two or more library collections hold the same titles, quantifying how much of each collection is shared, how much is unique, and how much in total the collections cover together. By treating holdings as sets and computing intersection, union, and overlap coefficients on matched identifiers such as ISBN, ISSN, or OCLC number, the method turns a vague sense of duplication into reproducible figures. These figures drive concrete decisions: where consortial partners can rely on one another, which titles are uniquely held and so must be preserved, and where duplicate purchasing or storage can be reduced. The technique is a workhorse of cooperative collection development and shared-print retention, summarized across the serials and collection-management literature including Nisonger's syntheses.

Key highlights

  • Turns a vague sense of duplication into reproducible quantities, supporting concrete cooperative-collection decisions.
  • Identifies uniquely held titles that require preservation, directly informing shared-print and weeding choices.
  • Reports combined reach across collections, clarifying how much partners cover together versus individually.
  • Rests on simple, transparent set mathematics that is easy to audit and to scale to large holdings with modern matching tools.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use collection overlap analysis when you need to quantify duplication and uniqueness across two or more collections to support cooperative collection development, shared-print retention, consortial purchasing, or weeding decisions. It is appropriate whenever you can obtain comparable, deduplicated holdings lists with reliable shared identifiers, and when the practical question is about coverage relationships between collections rather than the absolute quality of any single collection. It is less suitable when holdings records lack stable identifiers (so matching is unreliable), when the relevant question is about use or demand rather than holdings (use studies serve better), or when comparing collections at very different scales without choosing an overlap measure suited to that asymmetry. The method also assumes that a title held is a title available, so it should be paired with retention and access data before irreversible weeding.

Strengths & limitations

Strengths
  • Turns a vague sense of duplication into reproducible quantities, supporting concrete cooperative-collection decisions.
  • Identifies uniquely held titles that require preservation, directly informing shared-print and weeding choices.
  • Reports combined reach across collections, clarifying how much partners cover together versus individually.
  • Rests on simple, transparent set mathematics that is easy to audit and to scale to large holdings with modern matching tools.
Limitations
  • Results are only as good as identifier matching; missing or inconsistent identifiers bias both overlap and uniqueness.
  • Measures holdings, not use, so high overlap does not by itself justify deduplication without demand and retention data.
  • Edition, format, and bound-with complications make title matching genuinely ambiguous in many cases.
  • Choice of coefficient (overlap vs. Jaccard) can substantially change the apparent result when collections differ greatly in size.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

What is the difference between the overlap coefficient and the Jaccard coefficient?

Both summarize shared holdings, but they normalize differently. The overlap (Szymkiewicz-Simpson) coefficient divides the intersection by the size of the smaller collection, so it measures how completely the smaller collection is contained in the larger; it can reach 1 even when the larger collection is far bigger. The Jaccard coefficient divides the intersection by the union, measuring shared titles relative to combined titles, so it stays low when collections differ greatly in size. For asymmetric comparisons, report whichever matches your question and state which you used.

Why is identifier matching so important?

Every overlap and uniqueness figure depends on correctly deciding whether two records describe the same title. Matching on stable identifiers such as ISBN, ISSN, or OCLC control number is far more reliable than matching on title and author strings, which break on punctuation, transliteration, edition statements, and format differences. Poor matching either misses true matches (inflating uniqueness) or merges distinct editions (inflating overlap). Because retention and weeding decisions hinge on these counts, investing in clean, identifier-based deduplication is the single most important quality step.

Does high overlap mean a library can safely weed duplicated titles?

Not on its own. High overlap shows that other collections hold the same titles, but a responsible decision also needs use data and a group-wide count of retained copies. Shared-print programs typically commit to keeping a minimum number of copies of each title so that overlap does not silently erode the last accessible copies. Overlap analysis should therefore feed into a retention framework that weighs duplication against demand, access guarantees, and the risk of losing the only remaining copy, rather than driving weeding directly.

Sources

  1. 1.
    Nisonger, T. E. (1998). Management of Serials in Libraries. Englewood, CO: Libraries Unlimited.
    ISBN 9781563084782
  2. 2.
    IFLA Section on Acquisition and Collection Development (2001). Guidelines for a Collection Development Policy Using the Conspectus Model. The Hague: IFLA.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 23). Collection Overlap Analysis. ScholarGate. https://scholargate.app/library-information-science/collection-overlap-analysis