Process / pipelineScience Technology StudiesBibliometrics / technology analyticsPipeline

Tech Mining

Also known as: Technology mining, S&T text mining, Technical intelligence mining

OriginatorAlan L. Porter & Scott W. CunninghamYear2005Sources2Related methods4

Tech mining is the text mining of science and technology information—the publication, patent, and proposal databases that record the world's research and invention—to extract competitive technical intelligence. Coined by Alan Porter and Scott Cunningham, it turns large, fielded bibliographic corpora into actionable answers about who is doing what, where, with whom, and along which trajectories. By extracting entities such as authors, institutions, countries, keywords, and assignees and analysing their co-occurrence over time, tech mining profiles emerging technologies, maps research landscapes, and supports R&D management and innovation policy decisions.

Key highlights

  • Scales to tens or hundreds of thousands of records, surfacing patterns no manual literature review could detect.
  • Exploits the rich fielded structure of S&T databases (authors, affiliations, keywords, classifications, assignees, dates) for multi-angle analysis.
  • Directly serves decision-making—R&D management, forecasting, competitive intelligence—rather than producing analysis for its own sake.
  • Combines quantitative bibliometric indicators with visual maps (co-word, co-author, science overlay) that communicate findings to non-specialists.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use tech mining when you face a large body of science and technology records and need to extract intelligence about a technology's actors, topics, dynamics, and trajectory faster and more systematically than manual review allows. It suits competitive technical intelligence, technology forecasting and roadmapping, R&D portfolio and landscape analysis, and emerging-technology scoping. It assumes access to well-structured bibliographic or patent databases and the analytical effort to clean and interpret them. It is less appropriate when the relevant knowledge is tacit or undocumented, when the corpus is too small for statistical patterning, when fields are too noisy to consolidate reliably, or when deep conceptual rather than bibliometric understanding is the goal.

Strengths & limitations

Strengths
  • Scales to tens or hundreds of thousands of records, surfacing patterns no manual literature review could detect.
  • Exploits the rich fielded structure of S&T databases (authors, affiliations, keywords, classifications, assignees, dates) for multi-angle analysis.
  • Directly serves decision-making—R&D management, forecasting, competitive intelligence—rather than producing analysis for its own sake.
  • Combines quantitative bibliometric indicators with visual maps (co-word, co-author, science overlay) that communicate findings to non-specialists.
Limitations
  • Field cleanup and term consolidation are laborious and error-prone, and poor normalisation silently biases every downstream result.
  • Coverage and indexing limits of the source databases (missing venues, language bias, lag, inconsistent affiliations) constrain completeness.
  • Co-occurrence and frequency patterns are correlational and require expert interpretation; they do not by themselves explain causes or quality.
  • Bibliometric counts can be gamed and may reward volume over impact, so naive reading of the indicators can mislead strategic decisions.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

How does tech mining differ from ordinary bibliometrics or a systematic literature review?

Tech mining is bibliometrics oriented explicitly toward decision-relevant technical intelligence and built on software-supported text mining of fielded S&T records. Compared with a systematic literature review, which reads and synthesises content to answer a research question, tech mining analyses the metadata and term structure of very large corpora to reveal actors, topics, trends, and relationships at scale. It complements rather than replaces close reading, which is still needed to interpret and validate the patterns.

Why is field cleanup so important?

Bibliographic and patent records are entered inconsistently—one organisation may appear under dozens of name variants, authors are abbreviated differently, and keywords overlap. Without thesaurus-based consolidation and fuzzy matching ('list cleanup'), counts fragment and co-occurrence networks distort, so leading players and emerging topics are misidentified. Practitioners typically spend the largest share of a tech-mining project on this normalisation step because the validity of every downstream analytic depends on it.

What is a science overlay map?

A science overlay map projects a corpus onto a fixed global base map of science—built from journal-to-journal citation relations and partitioned into disciplines—so that the analyst can see, at a glance, which fields a body of research or a patent portfolio draws on and how interdisciplinary it is. It is a powerful tech-mining visualisation for positioning an organisation's or technology's footprint within the wider landscape of science.

Sources

  1. 1.
    Porter, A. L., & Cunningham, S. W. (2005). Tech Mining: Exploiting New Technologies for Competitive Advantage. Wiley.
    ISBN 9780471475675
  2. 2.
    Porter, A. L. (2007). How tech mining can enhance R&D management. Research-Technology Management, 50(2), 15-20.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 22). Tech Mining. ScholarGate. https://scholargate.app/science-technology-studies/tech-mining-analysis