Data Fabrication and Falsification
Definition, Detection, and Prevention of Research Data Fabrication and Falsification · Also known as: FFP Data Violations, Data Integrity Violations
Data fabrication and falsification are serious forms of research misconduct involving intentional misrepresentation of research data. Fabrication means inventing data that were never actually collected; falsification means altering authentic data to change the meaning. Both undermine scientific integrity, waste research resources, and can harm research subjects and the public. Federal policy (42 CFR Part 93) formally defines these violations; detection is improving through statistical analysis tools and data transparency practices; prevention requires robust data governance and culture of accountability.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Data integrity verification is essential for: (1) clinical trials (high stakes—safety and efficacy claims affect patient health); (2) large observational studies (large N makes patterns of fabrication easier to detect); (3) research with significant positive/surprising findings (heightened scrutiny); (4) research with rapid publication trajectory (researcher publishing far more than peers, raising suspicion); (5) research funded by industry (incentive for positive results); (6) research with ethical concerns (historically, fabrication sometimes emerged in ethically questionable studies—lack of oversight enabled fraud). Data verification should be routine in high-risk areas; good practice extends verification to all research.
Strengths & limitations
- Advances in detection: statistical software can now identify suspicious patterns suggesting fabrication or falsification (Benford's Law analysis, variance testing, p-value distribution analysis); makes fraud harder to hide.
- Data transparency enabling verification: open science practices (data archiving, sharing) allow independent researchers to access raw data and re-analyze; catches fabrication/falsification post-publication.
- Audit trails: modern electronic data systems automatically record who made changes and when; prevents silent modification of data without trace.
- Institutional oversight: data audits, monitoring procedures, and statistical checks during research conduct (not after) catch problems earlier.
- Career consequences: proven fabrication/falsification results in funding bans, loss of authorship, retractions, sometimes termination—strong deterrent.
- Literature correction: when fabrication detected, journals issue retractions; prevents continued reliance on false findings.
- Prevention training: explicit instruction on data integrity helps early-career researchers understand expectations and avoid misconduct.
- Sophisticated fraud hard to catch: researcher with statistical expertise can fabricate data that pass statistical anomaly checks (create realistic distributions, appropriate variance, avoid p-value clustering). Requires multiple detection methods.
- Falsification harder to detect than fabrication: if researcher has legitimate raw data but selectively reports some values and omits others, statistical checks may not catch this (distributions still realistic, just selected subset reported).
- Large datasets complicate audit: as datasets grow (big data, electronic health records with millions of records), auditing all values becomes computationally expensive; typically only random samples are checked.
- Detection after publication: many fraud cases discovered years after publication (e.g., Diederik Stapel's psychology fraud emerged 10+ years later). By then, other researchers have built on fraudulent findings.
- Burden on researchers: requiring detailed documentation, audit trails, data sharing increases researcher workload and time required for data management.
- Conflicts with privacy: sharing raw data to enable verification can compromise subject privacy despite de-identification; some sensitive datasets (genetic, mental health) difficult to share while protecting privacy.
- Institutional variability: small institutions, low-resource settings may lack capacity for robust data audits; creates inconsistent oversight.
Frequently asked
If I accidentally report an incorrect statistic in a paper (arithmetic error), is this considered falsification?
PROBABLY NOT—if it is honest error. Falsification requires intent or recklessness to misrepresent data. Arithmetic error (1+1=3 mistake) is not falsification; it is negligence/error. However, REPEATED uncorrected errors that systematically bias results in desired direction suggest recklessness (deliberate carelessness to achieve desired outcome), which may constitute falsification. Prevent by: (1) Double-check all calculations; use software to verify formulas. (2) If error discovered, promptly publish correction/erratum. (3) Review published data carefully for accuracy before submission. If significant error discovered post-publication, contact journal and request erratum. Transparency about error demonstrates integrity.
I want to exclude some data points from analysis (e.g., from subject who didn't follow protocol). Is this data falsification?
Excluding data is NOT falsification if done transparently with pre-specified criteria. Falsification is HIDING exclusion or MISREPRESENTING which data analyzed. APPROPRIATE DATA EXCLUSION: (1) Establish in advance (protocol, pre-registration) what data will be excluded and why. Examples: subject withdrew consent, failed eligibility check, protocol violation. (2) Document reason for exclusion in data file (do not silently delete; add note 'excluded: consent withdrawn'). (3) Report in paper how many excluded and why. (4) Conduct sensitivity analysis (report both with and without excluded data to show results are not driven by exclusion). INAPPROPRIATE DATA EXCLUSION (FALSIFICATION): (1) Silently delete inconvenient data points without documentation or reporting. (2) Exclude data post-hoc (after seeing results) because they don't support desired conclusion. (3) Report different numbers of subjects in different analyses without explaining why (suggests data were selectively included/excluded). Transparent, pre-specified exclusion is legitimate; hidden, post-hoc exclusion is falsification.
Can I analyze my data multiple ways and report only the result I believe is most valid?
Depends on whether you're being transparent. APPROPRIATE ANALYSIS CHOICES: (1) Pre-specify primary analysis method in protocol/pre-registration. Analyze data using primary method, report primary result prominently. (2) If you use secondary analyses (exploratory, sensitivity checks), label them clearly as secondary/exploratory. (3) Justify why you changed analysis method if you did (discovered data violation, learned better method, etc.). Report both original and revised method. INAPPROPRIATE (FALSIFICATION): (1) Conduct many analyses, report only 1 that gives desired result, hide other analyses. This is 'p-hacking' or 'HARKing' (Hypothesizing After Results Known). (2) Change analysis method post-hoc because first method gave null result, report revised method result without mentioning you changed methods. Best practice: pre-register primary analysis; report ALL primary and major secondary outcomes regardless of significance.
What statistical patterns suggest data fabrication that I should investigate?
RED FLAGS for possible fabrication: (1) TOO-PERFECT DISTRIBUTIONS: Data show no outliers, extremely low variance, or perfectly symmetric distributions (real data are messier). (2) SUSPICIOUS PRECISION: Reported measurements have impossible precision (blood pressure to 0.0001 mmHg; temperature of 98.6000°F). (3) BENFORD'S LAW VIOLATION: First digits not distributed per Benford's law (real data show characteristic distribution; fabricated data often uniform). (4) P-VALUE CLUSTERING: Disproportionate number of p-values just below .05 (suggests p-hacking or fabrication). (5) IDENTICAL VALUES: Many subjects with identical measurements in supposedly variable trait (e.g., exactly 10 subjects with identical weight). (6) INCONSISTENT SUMMARIES: Reported statistics don't match raw data (N, mean, SD inconsistent; totals don't add up). (7) IMPOSSIBLE VALUES: Negative counts, proportions >100%, measurements outside plausible range. (8) DUPLICATE RECORDS: Multiple subjects with identical demographics and identical measurements across multiple variables. If you observe these patterns, investigate—request raw data verification, check calculations, document audit procedures.
How can I tell if a published paper I'm citing had fabricated data?
SIGNS SUGGESTING POSSIBLE FABRICATION: (1) Paper has been retracted or expressions of concern issued—journal editor flagged concerns. (2) Paper's findings are inconsistent with other research on same topic and were never replicated successfully (if fraudulent, replication attempts fail). (3) Retraction Watch or PubPeer comments flagging statistical impossibilities (e.g., Carlisle identified anesthesia papers with data inconsistent with claimed methods). (4) Author of paper subsequently investigated for misconduct or had previous retractions. (5) Statistical anomalies noted by independent reanalyzes (e.g., Carlisle's analysis of anesthesia literature identified papers with mathematically impossible distributions). WHAT TO DO: (1) Check Retraction Watch database for retraction status. (2) Search PubPeer (post-publication peer review platform) for comments about paper. (3) If paper is foundational to your own work, consider replicating key findings or seeking confirmation from other sources. (4) If you suspect fabrication, report to journal editor and/or Office of Research Integrity. (5) When citing papers, note their quality and limitations; avoid over-relying on single suspect study.
Sources
- U.S. Office of Research Integrity. (2005). Public Health Service Policy on Research Misconduct. 42 CFR Part 93. Definitions of fabrication and falsification. link ↗
- Carlisle, J.B. (2017). Data Fabrication and Deviation in Statistics in Anesthesia Articles. Anesthesia, 72(2), 221–237. link ↗
- Nuijten, M.B., Hartgerink, C.H., van Assen, M.A., et al. (2015). The Prevalence of Statistical Reporting Errors in Psychology (1985-2013). Behavior Research Methods, 48(4), 1205–1226. DOI: 10.3758/s13428-015-0664-2 ↗
How to cite this page
ScholarGate. (2026, June 3). Definition, Detection, and Prevention of Research Data Fabrication and Falsification. ScholarGate. https://scholargate.app/en/research-ethics/data-fabrication-falsification
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Conflict of Interest in ResearchResearch Ethics↔ compare
- Research Integrity PrinciplesResearch Ethics↔ compare
- Research MisconductResearch Ethics↔ compare