Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Survey Methodology›Online Document Collection — Digital Records as Research Data
Process / pipelineData collection

Online Document Collection — Digital Records as Research Data

Online Document Collection Method · Also known as: digital document collection, web document gathering, online archival data collection, digital records collection

Online document collection is the systematic process of identifying, retrieving, and compiling digital documents — including web pages, institutional publications, social media posts, policy documents, and digital archives — as primary or supplementary research data. It extends classical document analysis into internet-mediated environments, enabling researchers to access large, geographically dispersed corpora without fieldwork travel or physical archive access.

ScholarGate
  1. Process / pipeline
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Online Document Collection
API-based Data CollectionContent AnalysisDocument CollectionOnline Non-participant O…Web ScrapingRemote Document Collecti…

When to use it

Use online document collection when the research question can be addressed through publicly available or accessible digital texts and when primary data collection (interviews, surveys) is impractical, expensive, or ethically sensitive. It is well suited to historical or longitudinal studies, policy analysis, organisational research, and studies of public discourse. It is a strong supplementary method for triangulation in mixed-methods designs. Do not use it as the sole method when the research question concerns personal experience, attitudes, or behaviour that is not captured in public documents; and avoid it when document authenticity or completeness cannot be verified, or when the phenomenon of interest leaves little or no online documentary trace.

Strengths & limitations

Strengths
  • Provides access to large, geographically dispersed corpora without travel or physical access constraints.
  • Documents are naturally occurring data — not produced under research conditions — reducing reactivity bias.
  • Cost-efficient: many sources are freely accessible through open-access portals, databases, or public websites.
  • Well suited to historical and longitudinal research because digital archives preserve dated versions of documents.
  • Easily combined with other methods (interviews, surveys) for triangulation in mixed-methods designs.
  • Scalable — computational text analysis tools can process thousands of documents that manual reading could not.
Limitations
  • Online documents may be edited, removed, or moved after collection, making replication or follow-up difficult.
  • Access restrictions, paywalls, or authentication barriers can exclude important sources and introduce selection bias.
  • Document quality and authenticity can be difficult to verify, particularly for user-generated or anonymous content.
  • Represents only what has been digitised and published online — marginalised voices and informal knowledge are often underrepresented.
  • Large corpora require data management infrastructure and may require computational skills beyond many researchers' training.

Frequently asked

How is online document collection different from web scraping?

Web scraping is an automated technical process for extracting text or data from web pages using scripts or crawlers. Online document collection is a broader methodological category that may include scraping but also encompasses manual retrieval, database downloads, and API calls. The defining feature of online document collection is its purposive, criteria-driven approach to building a research corpus; scraping is one tool within that process.

Are social media posts legitimate research documents?

Yes, in many research contexts they are. Social media posts are naturally occurring texts that express opinions, experiences, and interactions. However, their use raises ethical questions about informed consent and individual privacy even when content is publicly accessible. Researchers should follow their institution's ethics guidelines and disciplinary norms, anonymise individuals where appropriate, and be transparent about collection and analysis procedures.

How do I handle documents that disappear after collection?

Record the exact URL and date of access for every document at the time of collection. Use archiving tools such as the Internet Archive Wayback Machine or save local copies in PDF format. Zotero and similar reference managers can capture page snapshots. This provenance record is essential for citation accuracy and for responding to peer reviewers who cannot locate a source.

Do I need ethical approval for online document collection?

It depends on the source and content. Institutional reports, government documents, and published organisational texts are generally low-risk. Collecting identifiable personal content from social media or health forums typically requires ethical review, particularly in medical or psychological research. Always check your institutional research ethics policy before beginning collection.

How many documents constitute an adequate corpus?

There is no universal minimum. Adequacy is judged by saturation — the point at which additional documents no longer introduce new themes or evidence relevant to the research question — and by representativeness of the source landscape. A focused policy study might be well served by 20–50 key documents; a large discourse study might require hundreds. Define your criteria explicitly and report the final corpus size transparently.

Sources

  1. Bowen, G. A. (2009). Document analysis as a qualitative research method. Qualitative Research Journal, 9(2), 27–40. DOI: 10.3316/QRJ0902027 ↗
  2. Prior, L. (2003). Using Documents in Social Research. Sage Publications. ISBN: 978-0761965497

How to cite this page

ScholarGate. (2026, June 3). Online Document Collection Method. ScholarGate. https://scholargate.app/en/survey-methodology/online-document-collection

Related methods

API-based Data CollectionContent AnalysisDocument CollectionOnline Non-participant ObservationWeb Scraping

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • API-based Data CollectionSurvey Methodology↔ compare
  • Content AnalysisQualitative↔ compare
  • Document CollectionSurvey Methodology↔ compare
  • Online Non-participant ObservationSurvey Methodology↔ compare
  • Web ScrapingSurvey Methodology↔ compare
Compare side by side →

Referenced by

Online Non-participant ObservationRemote Document Collection

Similar methods

Remote Document CollectionDigital Document AnalysisDocument CollectionMulti-source Document CollectionWeb ScrapingDocument AnalysisRemote Web ScrapingDigital Content analysis

Related reference concepts

Corpus Building and CurationQualitative Research MethodsCorpus Linguistics and Web CorporaText ClusteringDigital Archives and Cultural HeritageSystematic Review

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Online Document Collection (Online Document Collection Method). Retrieved 2026-07-21 from https://scholargate.app/en/survey-methodology/online-document-collection · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Adapted from traditional document analysis; digital form emerged with widespread internet adoption
Year
1990s–2000s (digital / web era)
Type
Qualitative / mixed-methods data collection technique
DataType
Digital text documents, web pages, PDFs, social media posts, institutional records, digital archives
Subfamily
Data collection
Related methods
API-based Data CollectionContent AnalysisDocument CollectionOnline Non-participant ObservationWeb Scraping
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account