Data Sharing and Open Science
Data Sharing, Reproducibility, and Open Science Practices · Also known as: Open Data, Research Data Sharing, Research Reproducibility
Data sharing and open science are practices that maximize research transparency and reproducibility by making raw data, analysis code, and methods publicly available alongside publications. The replication crisis (widespread failure to reproduce published findings in psychology, medicine, and other fields) revealed that traditional publication—focusing on novel results—incentivizes selective reporting and p-hacking. Open science practices (preregistration, data sharing, code sharing, open materials) aim to reduce bias and enable independent verification. Major funders (NIH, NSF, EU) now mandate open science practices, and many journals require data availability statements or code repositories.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Implement open science practices from the start of any research project, especially empirical studies with human subjects, clinical trials, or studies testing multiple hypotheses. Preregister studies in Clinicaltrials.gov (required for clinical trials) or OSF (encouraged for all research). Share data and code with every publication, at minimum in a public repository. Use open science to increase credibility, enable replication, and comply with funder mandates.
Strengths & limitations
- Enables independent verification and replication, strengthening confidence in findings.
- Reduces researcher degrees of freedom (p-hacking, HARKing, selective reporting) by making analysis plans public.
- Increases impact: papers with open data are cited more and have higher reproducibility.
- Supports meta-analyses and new analyses; data become a research resource for future work.
- Aligns with public funding principles; publicly-funded research should have publicly-accessible data.
- Catches errors and fraud early; suspicious patterns are visible to community.
- Privacy concerns: de-identifying sensitive data is complex; some data cannot be shared (e.g., video, genetic data with identifiable traits).
- Intellectual property: researchers may delay sharing proprietary data pending patent or commercial use.
- Data documentation: preparing codebooks and metadata requires effort and expertise.
- Technical barriers: not all researchers are comfortable with GitHub, data repositories, or DOIs.
- Incomplete implementation: preregistration does not prevent all problematic practices; enforcement is weak.
Frequently asked
Does sharing data mean I lose credit for my work?
No. You are cited as the data provider. Sharing data increases impact and citations, not decreases them. Papers with open data receive ~25% more citations on average. You can also impose a brief embargo (e.g., 1 year before others access data) to allow you time to complete analyses before others use the data.
How do I de-identify human subjects data for sharing?
Remove direct identifiers (names, IDs, dates of birth). Use aggregate or binned data where possible (age ranges instead of exact dates). Check for quasi-identifiers that could re-identify subjects (e.g., rare disease + hospital = identifiable). Consult your IRB for guidance; some data cannot be safely de-identified and should not be shared publicly. Alternative: share through a data enclave where researchers access data on-site under restrictions.
Is preregistration the same as clinical trial registration?
Similar but not identical. Clinical trial registration (Clinicaltrials.gov) is mandatory for clinical trials and includes basic info (hypothesis, primary outcome). Preregistration (OSF, AsPredicted) is more detailed and includes analysis plan, hypotheses, and sample size. Both are prospective registration; both prevent HARKing. Clinical trials require registration as a condition of publication (ICMJE requirement). Other research can voluntarily preregister on OSF.
What if I discover unexpected findings after preregistration?
Exploratory findings are allowed. In your paper, clearly label findings as 'preregistered' (planned) or 'exploratory' (unexpected). Report exploratory findings but note they are hypothesis-generating, not confirmatory. This transparency prevents HARKing; readers can distinguish planned findings from post-hoc discoveries. Exploratory findings should be replicated in new data before being treated as confirmed.
Can I charge for access to my shared data?
Generally no. Open science principles require data to be freely accessible. However, if data are proprietary or part of a commercial product, you can share a subset, use a limited-use license, or place data in a controlled-access repository (requiring approval to access). Always disclose restrictions clearly. Fully open data maximize impact and compliance with funders.
Sources
- Open Science Framework (2023). OSF. Center for Open Science. link ↗
- Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J., et al. (2016). The FAIR Guiding Principles for Scientific Data Management and Stewardship. Scientific Data, 3, 160018. DOI: 10.1038/sdata.2016.18 ↗
- Cohen, S. A., Cox, R. P., Favor, T. K., & Glover, S. C. (2016). The Role of Preregistration in Psychological Research. Psychological Science Agenda (American Psychological Association). link ↗
How to cite this page
ScholarGate. (2026, June 3). Data Sharing, Reproducibility, and Open Science Practices. ScholarGate. https://scholargate.app/en/publication-ethics/data-sharing-open-science
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Open Access Publishing ModelsPublication Ethics↔ compare
- Peer Review ProcessPublication Ethics↔ compare
- Plagiarism in Academic ResearchPublication Ethics↔ compare
- Preprint Servers in SciencePublication Ethics↔ compare