Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Research Ethics›Data Fabrication and Falsification
Process / pipelineethical-violations

Data Fabrication and Falsification

Definition, Detection, and Prevention of Research Data Fabrication and Falsification · Also known as: FFP Data Violations, Data Integrity Violations

Data fabrication and falsification are serious forms of research misconduct involving intentional misrepresentation of research data. Fabrication means inventing data that were never actually collected; falsification means altering authentic data to change the meaning. Both undermine scientific integrity, waste research resources, and can harm research subjects and the public. Federal policy (42 CFR Part 93) formally defines these violations; detection is improving through statistical analysis tools and data transparency practices; prevention requires robust data governance and culture of accountability.

ScholarGate
  1. Process / pipeline
  2. v1
  3. 3 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Data Fabrication and Falsification
Conflict of Interest in…Research Integrity Princ…Research Misconduct

When to use it

Data integrity verification is essential for: (1) clinical trials (high stakes—safety and efficacy claims affect patient health); (2) large observational studies (large N makes patterns of fabrication easier to detect); (3) research with significant positive/surprising findings (heightened scrutiny); (4) research with rapid publication trajectory (researcher publishing far more than peers, raising suspicion); (5) research funded by industry (incentive for positive results); (6) research with ethical concerns (historically, fabrication sometimes emerged in ethically questionable studies—lack of oversight enabled fraud). Data verification should be routine in high-risk areas; good practice extends verification to all research.

Strengths & limitations

Strengths
  • Advances in detection: statistical software can now identify suspicious patterns suggesting fabrication or falsification (Benford's Law analysis, variance testing, p-value distribution analysis); makes fraud harder to hide.
  • Data transparency enabling verification: open science practices (data archiving, sharing) allow independent researchers to access raw data and re-analyze; catches fabrication/falsification post-publication.
  • Audit trails: modern electronic data systems automatically record who made changes and when; prevents silent modification of data without trace.
  • Institutional oversight: data audits, monitoring procedures, and statistical checks during research conduct (not after) catch problems earlier.
  • Career consequences: proven fabrication/falsification results in funding bans, loss of authorship, retractions, sometimes termination—strong deterrent.
  • Literature correction: when fabrication detected, journals issue retractions; prevents continued reliance on false findings.
  • Prevention training: explicit instruction on data integrity helps early-career researchers understand expectations and avoid misconduct.
Limitations
  • Sophisticated fraud hard to catch: researcher with statistical expertise can fabricate data that pass statistical anomaly checks (create realistic distributions, appropriate variance, avoid p-value clustering). Requires multiple detection methods.
  • Falsification harder to detect than fabrication: if researcher has legitimate raw data but selectively reports some values and omits others, statistical checks may not catch this (distributions still realistic, just selected subset reported).
  • Large datasets complicate audit: as datasets grow (big data, electronic health records with millions of records), auditing all values becomes computationally expensive; typically only random samples are checked.
  • Detection after publication: many fraud cases discovered years after publication (e.g., Diederik Stapel's psychology fraud emerged 10+ years later). By then, other researchers have built on fraudulent findings.
  • Burden on researchers: requiring detailed documentation, audit trails, data sharing increases researcher workload and time required for data management.
  • Conflicts with privacy: sharing raw data to enable verification can compromise subject privacy despite de-identification; some sensitive datasets (genetic, mental health) difficult to share while protecting privacy.
  • Institutional variability: small institutions, low-resource settings may lack capacity for robust data audits; creates inconsistent oversight.

Frequently asked

If I accidentally report an incorrect statistic in a paper (arithmetic error), is this considered falsification?

PROBABLY NOT—if it is honest error. Falsification requires intent or recklessness to misrepresent data. Arithmetic error (1+1=3 mistake) is not falsification; it is negligence/error. However, REPEATED uncorrected errors that systematically bias results in desired direction suggest recklessness (deliberate carelessness to achieve desired outcome), which may constitute falsification. Prevent by: (1) Double-check all calculations; use software to verify formulas. (2) If error discovered, promptly publish correction/erratum. (3) Review published data carefully for accuracy before submission. If significant error discovered post-publication, contact journal and request erratum. Transparency about error demonstrates integrity.

I want to exclude some data points from analysis (e.g., from subject who didn't follow protocol). Is this data falsification?

Excluding data is NOT falsification if done transparently with pre-specified criteria. Falsification is HIDING exclusion or MISREPRESENTING which data analyzed. APPROPRIATE DATA EXCLUSION: (1) Establish in advance (protocol, pre-registration) what data will be excluded and why. Examples: subject withdrew consent, failed eligibility check, protocol violation. (2) Document reason for exclusion in data file (do not silently delete; add note 'excluded: consent withdrawn'). (3) Report in paper how many excluded and why. (4) Conduct sensitivity analysis (report both with and without excluded data to show results are not driven by exclusion). INAPPROPRIATE DATA EXCLUSION (FALSIFICATION): (1) Silently delete inconvenient data points without documentation or reporting. (2) Exclude data post-hoc (after seeing results) because they don't support desired conclusion. (3) Report different numbers of subjects in different analyses without explaining why (suggests data were selectively included/excluded). Transparent, pre-specified exclusion is legitimate; hidden, post-hoc exclusion is falsification.

Can I analyze my data multiple ways and report only the result I believe is most valid?

Depends on whether you're being transparent. APPROPRIATE ANALYSIS CHOICES: (1) Pre-specify primary analysis method in protocol/pre-registration. Analyze data using primary method, report primary result prominently. (2) If you use secondary analyses (exploratory, sensitivity checks), label them clearly as secondary/exploratory. (3) Justify why you changed analysis method if you did (discovered data violation, learned better method, etc.). Report both original and revised method. INAPPROPRIATE (FALSIFICATION): (1) Conduct many analyses, report only 1 that gives desired result, hide other analyses. This is 'p-hacking' or 'HARKing' (Hypothesizing After Results Known). (2) Change analysis method post-hoc because first method gave null result, report revised method result without mentioning you changed methods. Best practice: pre-register primary analysis; report ALL primary and major secondary outcomes regardless of significance.

What statistical patterns suggest data fabrication that I should investigate?

RED FLAGS for possible fabrication: (1) TOO-PERFECT DISTRIBUTIONS: Data show no outliers, extremely low variance, or perfectly symmetric distributions (real data are messier). (2) SUSPICIOUS PRECISION: Reported measurements have impossible precision (blood pressure to 0.0001 mmHg; temperature of 98.6000°F). (3) BENFORD'S LAW VIOLATION: First digits not distributed per Benford's law (real data show characteristic distribution; fabricated data often uniform). (4) P-VALUE CLUSTERING: Disproportionate number of p-values just below .05 (suggests p-hacking or fabrication). (5) IDENTICAL VALUES: Many subjects with identical measurements in supposedly variable trait (e.g., exactly 10 subjects with identical weight). (6) INCONSISTENT SUMMARIES: Reported statistics don't match raw data (N, mean, SD inconsistent; totals don't add up). (7) IMPOSSIBLE VALUES: Negative counts, proportions >100%, measurements outside plausible range. (8) DUPLICATE RECORDS: Multiple subjects with identical demographics and identical measurements across multiple variables. If you observe these patterns, investigate—request raw data verification, check calculations, document audit procedures.

How can I tell if a published paper I'm citing had fabricated data?

SIGNS SUGGESTING POSSIBLE FABRICATION: (1) Paper has been retracted or expressions of concern issued—journal editor flagged concerns. (2) Paper's findings are inconsistent with other research on same topic and were never replicated successfully (if fraudulent, replication attempts fail). (3) Retraction Watch or PubPeer comments flagging statistical impossibilities (e.g., Carlisle identified anesthesia papers with data inconsistent with claimed methods). (4) Author of paper subsequently investigated for misconduct or had previous retractions. (5) Statistical anomalies noted by independent reanalyzes (e.g., Carlisle's analysis of anesthesia literature identified papers with mathematically impossible distributions). WHAT TO DO: (1) Check Retraction Watch database for retraction status. (2) Search PubPeer (post-publication peer review platform) for comments about paper. (3) If paper is foundational to your own work, consider replicating key findings or seeking confirmation from other sources. (4) If you suspect fabrication, report to journal editor and/or Office of Research Integrity. (5) When citing papers, note their quality and limitations; avoid over-relying on single suspect study.

Sources

  1. U.S. Office of Research Integrity. (2005). Public Health Service Policy on Research Misconduct. 42 CFR Part 93. Definitions of fabrication and falsification. link ↗
  2. Carlisle, J.B. (2017). Data Fabrication and Deviation in Statistics in Anesthesia Articles. Anesthesia, 72(2), 221–237. link ↗
  3. Nuijten, M.B., Hartgerink, C.H., van Assen, M.A., et al. (2015). The Prevalence of Statistical Reporting Errors in Psychology (1985-2013). Behavior Research Methods, 48(4), 1205–1226. DOI: 10.3758/s13428-015-0664-2 ↗

How to cite this page

ScholarGate. (2026, June 3). Definition, Detection, and Prevention of Research Data Fabrication and Falsification. ScholarGate. https://scholargate.app/en/research-ethics/data-fabrication-falsification

Related methods

Conflict of Interest in ResearchResearch Integrity PrinciplesResearch Misconduct

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Conflict of Interest in ResearchResearch Ethics↔ compare
  • Research Integrity PrinciplesResearch Ethics↔ compare
  • Research MisconductResearch Ethics↔ compare
Compare side by side →

Referenced by

Research Integrity PrinciplesResearch Misconduct

Similar methods

Research MisconductResearch Integrity PrinciplesArticle Retraction ProcessData Sharing and Open SciencePlagiarism in Academic ResearchDuplicate Publication and Salami SlicingDeception and Debriefing in ResearchConflict of Interest in Research

Related reference concepts

Reproducible ResearchMissing Data and AttritionType I and Type II ErrorsStatistical Power and Sample SizeStatistical Hypothesis TestingInternal Validity

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Data Fabrication and Falsification (Definition, Detection, and Prevention of Research Data Fabrication and Falsification). Retrieved 2026-07-21 from https://scholargate.app/en/research-ethics/data-fabrication-falsification · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
U.S. Office of Research Integrity; definitions in federal policy 42 CFR 93
Subfamily
ethical-violations
Year
2005
Type
Standard
Related methods
Conflict of Interest in ResearchResearch Integrity PrinciplesResearch Misconduct
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account