Co-word Analysis — Keyword Co-occurrence Network Analysis
Co-word Analysis · Also known as: keyword co-occurrence analysis, co-word mapping, keyword co-word network, CWA
Co-word analysis is a scientometric technique that quantifies how often pairs of keywords, subject terms, or title words appear together across a corpus of publications. By treating simultaneous occurrence as a proxy for conceptual relatedness, it constructs networks and clusters that reveal the intellectual structure, dominant themes, and emerging sub-fields of a research domain.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
+14 more
When to use it
Co-word analysis is appropriate when the research goal is to map the thematic structure of a research domain, identify intellectual clusters, or trace how topics have evolved over time. It suits bibliometric or scientometric reviews, systematic literature reviews with a mapping component, and journal editorial studies. The technique requires a reasonably large corpus — typically at least 100 documents, with several hundred or more yielding more stable cluster solutions. It is not appropriate as a substitute for content reading: it identifies what terms are associated, not what researchers actually claim or argue. Do not apply co-word analysis when the corpus is small (fewer than 50 records), when the keyword metadata is sparse or unstandardised, or when the research question requires understanding the quality or findings of individual studies rather than the overall shape of a field.
Strengths & limitations
- Scales to corpora of thousands of documents that cannot be read manually, enabling comprehensive field-level analysis.
- Produces transparent, replicable outputs that can be exactly reproduced given the same dataset and parameters.
- Simultaneously reveals both the thematic content and the relational structure of a research field.
- Temporal co-word analysis allows identification of emerging, motor, niche, and declining research themes.
- Software tools (VOSviewer, bibliometrix) make implementation accessible without custom programming.
- Complements citation-based methods (co-citation, bibliographic coupling) by capturing conceptual rather than intellectual linkage.
- Results depend heavily on keyword quality: inconsistent indexing, missing keywords, or non-standardised author keywords introduce noise and can produce misleading clusters.
- High-frequency generic terms dominate co-occurrence counts if not filtered carefully, obscuring substantive thematic signal.
- Tells you which topics co-occur but not what researchers say about them — thematic interpretation still requires expert reading of representative papers.
- Cluster labels and interpretations are subjective; different analysts may label the same cluster differently.
- Limited to terms present in metadata fields; important concepts discussed only in full text are invisible to the analysis.
Frequently asked
What is the difference between co-word analysis and co-citation analysis?
Co-citation analysis links documents (or authors) based on being cited together in subsequent papers, capturing intellectual influence and the structure of knowledge use. Co-word analysis links keywords based on appearing together in the same documents, capturing conceptual content. Co-citation reveals the shoulders on which a field stands; co-word reveals what the field is currently talking about. The two methods are complementary and are often reported together in science-mapping studies.
How many documents do I need for reliable co-word analysis?
There is no universal minimum, but below about 100 documents the co-occurrence matrix tends to be too sparse to form meaningful clusters, and results are highly sensitive to individual records. Most published co-word studies work with several hundred to several thousand documents. If your corpus is smaller, consider a qualitative content analysis or a thematic analysis of individual papers instead.
Should I use author keywords or index keywords?
Both have trade-offs. Author keywords reflect the authors' own framing and are closer to current terminology, but they are inconsistent and may lack standardisation across papers. Index keywords (e.g., Web of Science KeyWords Plus, MeSH in PubMed) are controlled vocabulary applied by professional indexers, offering greater consistency but sometimes lagging behind emerging terminology. Many studies use both, or combine them after deduplication and harmonisation.
Which software is best for co-word analysis?
VOSviewer is the most widely used tool for generating keyword co-occurrence maps and is freely available; it handles data import, threshold setting, clustering, and visualisation in a single interface. The R package bibliometrix provides scripted, reproducible co-word analysis with additional flexibility for time-sliced and comparative studies. CiteSpace and SciMAT are alternatives with stronger temporal evolution features. The choice depends on your need for reproducibility, customisation, and familiarity with R.
How do I decide on the minimum co-occurrence threshold?
The threshold controls network density. A threshold of 2 typically retains most keywords and produces an unwieldy, over-connected network. Thresholds of 3–10 are common in practice, calibrated so that the resulting network has a manageable number of nodes (roughly 50–300) while retaining the most informative keywords. VOSviewer and bibliometrix both report the number of keywords that meet each threshold, allowing iterative adjustment before committing to a final value.
Sources
- Callon, M., Courtial, J. P., Turner, W. A., & Bauin, S. (1983). From translations to problematic networks: An introduction to co-word analysis. Social Science Information, 22(2), 191–235. DOI: 10.1177/053901883022002003 ↗
- He, Q. (1999). Knowledge discovery through co-word analysis. Library Trends, 48(1), 133–159. link ↗
How to cite this page
ScholarGate. (2026, June 3). Co-word Analysis. ScholarGate. https://scholargate.app/en/scientometrics/co-word-analysis
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Bibliographic CouplingBibliometrics↔ compare
- Bibliometric AnalysisScientometrics↔ compare
- Co-Citation AnalysisBibliometrics↔ compare
- Science MappingBibliometrics↔ compare
- Scientometric AnalysisScientometrics↔ compare
- Thematic Evolution AnalysisScientometrics↔ compare