Keyword Co-Occurrence Analysis
Also known as: term co-occurrence, keyword network analysis, thematic analysis, term clustering
Keyword co-occurrence analysis is a text mining and bibliometric method that identifies research themes and their relationships by analyzing how frequently terms or keywords appear together in abstracts, titles, or indexed keywords of scientific publications. When two keywords appear together frequently, they are considered co-occurring, indicating a shared thematic or conceptual relationship. This method rapidly reveals the topical structure of a research field without relying on formal classifications, making it particularly useful for detecting emerging research areas and understanding disciplinary boundaries.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use keyword co-occurrence analysis when you need rapid mapping of a research field without waiting for citation accumulation, when detecting emerging research trends or interdisciplinary intersections, when tracking terminology evolution, or when identifying research gaps (combinations of keywords not yet co-occurring). It is ideal for exploring recent literature (last 2–5 years) where citation counts are still low. Keyword analysis is particularly valuable in fast-moving fields (computer science, medicine, environmental science) where terminology and priorities shift rapidly. Combine it with citation-based methods for complete picture: keywords show current trends, citations show validated influence.
Strengths & limitations
- Real-time detection: keywords appear immediately upon publication; no citation lag required.
- Terminology sensitivity: detects semantic shifts, buzz words, and emerging conceptual frameworks.
- Scalable: keyword networks scale to tens of thousands of papers without computational burden.
- Intuitive interpretation: keyword clusters and relationships are more directly interpretable than citation patterns.
- No database selection bias: citation databases (WoS, Scopus) may miss certain publication types; keyword analysis can include preprints, conference papers, and institutional repositories.
- Terminology inconsistency: the same concept has many names ('deep learning'='neural networks'='artificial intelligence'—not identical but overlapping); manual standardization is needed.
- Author keyword bias: authors select keywords to maximize visibility; fashionable keywords are over-used regardless of relevance.
- Noise in automated extraction: keywords extracted from text (not author-assigned) include non-substantive terms and noise.
- High co-occurrence of stop keywords: universal keywords ('methods', 'study', 'analysis') appear in all papers, generating spurious co-occurrences; filtering is essential.
- Temporal trends may reflect indexing changes: sudden appearance of a new keyword may reflect a database change rather than research emergence.
Frequently asked
Should I use author-assigned keywords or extract keywords from text?
Author keywords are higher quality but sparse (not all papers have them). Text extraction is comprehensive but noisier. Best practice: use both. Author keywords as primary (quality), then supplement with extracted keywords for papers without assigned keywords. Always standardize and filter extracted keywords to remove stop words and noise. Report which approach was used.
How do I standardize keywords without over-merging distinct concepts?
Create a synonym dictionary manually for major keywords. For 'machine learning', group: ['machine learning', 'ML', 'statistical learning', 'supervised learning' (if context is ML not formal statistics)]. For 'neural networks', group: ['neural networks', 'deep learning', 'artificial neural networks']. Be conservative: 'neural networks' and 'artificial intelligence' overlap but are distinct—keep them separate unless your analysis specifically requires merging. Use domain expertise; do not automate standardization without expert review.
What co-occurrence frequency threshold should I use?
It depends on corpus size. For a corpus of 1000 papers, a threshold of 3–5 co-occurrences is typical. For 10,000 papers, 10–20 is reasonable. For 100,000+ papers, 50+ may be needed. The goal is to reduce noise while retaining signal. A rule of thumb: use a threshold that yields 100–500 keyword pairs; fewer pairs suggests your threshold is too high, more suggests too low. Start conservative (threshold=5) and relax if needed based on network density.
Can I use keyword co-occurrence to forecast which research areas will become influential?
Not directly. Keyword co-occurrence shows what researchers are currently studying, not what will become important. However, temporal analysis (keyword growth trends) can suggest momentum: rapidly rising co-occurrence suggests growing community interest. Combine with citation analysis: rising keywords + rising citations = strong indicator of emerging prominence. Keywords alone cannot predict influence—many trending keywords prove to be temporary fads.
Sources
- Cobo, M. J., López-Herrera, A. G., Herrera-Viedma, E., & Herrera, F. (2011). An approach for detecting, quantifying, and visualizing the evolution of a research field: A practical application to the fuzzy sets theory field. Journal of Informetrics, 5(1), 146–166. DOI: 10.1016/j.joi.2010.10.002 ↗
- Van Eck, N. J., & Waltman, L. (2010). Software survey: VOSviewer, a computer program for bibliometric mapping. Scientometrics, 84(2), 523–538. DOI: 10.1007/s11192-009-0146-3 ↗
How to cite this page
ScholarGate. (2026, June 4). Keyword Co-Occurrence Analysis. ScholarGate. https://scholargate.app/en/bibliometrics/keyword-co-occurrence
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Bibliographic CouplingBibliometrics↔ compare
- Co-Citation AnalysisBibliometrics↔ compare
- Science MappingBibliometrics↔ compare
- VOSviewer and CiteSpace ToolsBibliometrics↔ compare