← Back to Blog

Data sharing · ~8 min read

Data-Sharing Requirements in Clinical Research

Mandatory data-sharing policies in clinical research: what ICMJE and NIH require, and how often the data is actually shared.

In short

Since 1 July 2018, the ICMJE has required a data-sharing statement in every clinical trial manuscript, and since 25 January 2023 most NIH-funded research has needed an approved data management and sharing plan.12 The declared intent to share has gone up sharply as a result. What has not moved is actual reuse: a cross-sectional audit of ten leading surgical journals found data were obtainable for reanalysis in exactly 3.1% of trials both before and after the ICMJE policy took effect, and a separate audit of JAMA, Lancet, and NEJM found 68.6% of trials declared they would share data but only 0.6% had individual-participant data actually deposited and publicly available two years later.34 The burden is real too, investigators who fulfilled at least one data request spent a median of 18 person-hours preparing it, uncompensated,5 and so is the privacy risk in small surgical cohorts, where fewer than a dozen patients per group is enough to make re-identification a realistic concern once indirect identifiers are combined.6 Mandates have changed what gets written in a manuscript; they have not yet changed what gets reused.

What the ICMJE and NIH policies actually require

Individual participant data (IPD) sharing means making the de-identified, patient-level dataset behind a published analysis available to other researchers, as distinct from sharing only the aggregate results reported in a paper. The International Committee of Medical Journal Editors (ICMJE), the body whose "Recommendations" govern manuscript requirements at most major medical and surgical journals, began requiring a data-sharing statement in any manuscript reporting clinical trial results submitted on or after 1 July 2018.1 The statement has to say whether individual de-identified participant data will be shared, exactly what will be shared (the dataset alone, or also the protocol and statistical analysis plan), when it becomes available and for how long, and by what mechanism and to whom. The ICMJE is explicit that "undecided is not an acceptable answer." For trials that began enrolling participants on or after 1 January 2019, the same data-sharing plan has to be registered up front, at the time of trial registration, not written retrospectively at submission.

The NIH's Data Management and Sharing (DMS) Policy took effect on 25 January 2023 and applies to essentially all NIH-funded research that generates scientific data, extramural grants, contracts, and intramural projects submitted, proposed, or executed on or after that date, regardless of award size.2 Investigators must submit a DMS plan at the application stage describing how data will be managed and shared, and then comply with the plan NIH approves as part of the award; training grants, fellowships, and awards that do not generate scientific data are excluded. Neither policy asks a researcher to hand data to a funder or journal outright, both put the obligation on the investigator to make data available to other researchers on the stated terms, after the fact.

The case for the mandate

The argument for compulsory sharing is straightforward and not seriously disputed in principle: a result that can be independently reanalysed from the original dataset is a stronger form of verification than a result that can only be checked against the numbers printed in the paper, and pooled IPD is what feeds individual-patient-data meta-analyses, which are widely regarded as more reliable than meta-analyses built from published summary statistics alone. Where reanalysis of shared data has actually happened, it has tended to confirm rather than overturn the original published findings, which is itself the point of the exercise: a mandate that produces even occasional independent confirmation of a trial's primary result is doing something a results section alone cannot do. That is a genuine point in the mandate's favour, and it is the reason funders and editors have kept pushing the requirement rather than abandoning it in the face of the compliance problems below.

What data availability looks like once the mandate is in force

The gap between the policy and its effect is the part of this debate that gets under-reported. A 2022 cross-sectional study examined 130 randomised controlled trials published in ten leading surgical journals, 65 before and 65 after the ICMJE policy came into force in July 2018, and requested the underlying data for each.3 Data sufficient to reanalyse the primary outcome were obtained for 2 of 65 trials (3.1%) before the policy and 2 of 65 (3.1%) after it, no measurable change (odds ratio 1.00; 95% CI, 0.07-14.19). A data-sharing statement appeared far more often after the policy (16.9% of post-policy trials versus none before it), but none of the four trials whose data were actually obtainable were among the ones that had carried a data-sharing statement in the first place. In other words, the paperwork increased and the underlying availability did not.

The same pattern shows up outside surgery. An audit of clinical trials published in JAMA, Lancet, and NEJM between July 2018 and April 2020 found that 68.6% declared an intention to share data in their data-sharing statement, but only two IPD datasets (0.6%) were actually de-identified and publicly deposited by the time of the audit; most of the rest were nominally "available on request" through a repository or the authors, and among trials that specifically promised repository deposit, fewer than one in five had actually deposited anything, largely because of unresolved embargoes.4 A broader scoping review of the IPD-sharing literature reached the same conclusion from a different angle: willingness to share, as expressed in a statement, is consistently high, but actual data-sharing rates are "suboptimal," enforcement by publishers is inconsistent, and most datasets that are shared are never requested at all.

The burden the mandate creates

None of this happens for free on the investigator's side. A survey of clinical trial investigators who had published under a journal data-sharing policy and had granted at least one data request found a median of 18 person-hours (range 3-125) spent preparing a single dataset for release, de-identifying it, documenting variables, and packaging it in a usable form, almost always without dedicated funding or staff time allocated to the task.5 For a surgical trial team without a dedicated data manager, that person-hour cost lands on whoever ran the study, typically the same resident or junior faculty member already stretched across the clinical and academic sides of the job.

The privacy problem in small cohorts

Surgical datasets are frequently small relative to the pharmaceutical trials the sharing mandates were originally designed around, and small sample size is itself a driver of re-identification risk. A pharmaceutical-industry privacy review of 530 scientific publications used a k-anonymity threshold and found that treatment groups or cohorts with fewer than 12 participants carried materially elevated re-identification risk once indirect identifiers, age, sex, geographic origin, procedure type, were combined; 13% of the publications reviewed required changes before release specifically to reduce that risk.6 A single-surgeon case series or a pilot RCT with 10 patients per arm, common in surgical literature, sits squarely inside that risk band. De-identification techniques exist, but they take expertise and time that most surgical research teams do not have in-house, which folds the privacy problem back into the burden problem: doing this properly, rather than just stripping obvious identifiers, is itself a further cost the mandate does not fund.

What this means for the manuscript you're writing

None of the evidence above argues that IPD sharing is worthless, the reanalyses that do happen mostly confirm the original results, and the case for pooling data into future meta-analyses is sound. What it argues is that a data-sharing statement, on its own, is not doing the verification work its presence implies. A "data available on reasonable request" line satisfies the ICMJE checkbox without making the dataset any more likely to be reused, and the actual reproducibility of a result should not rest on whether a future reader happens to submit a request that gets fulfilled. Writing an honest, specific data-sharing statement, naming exactly what will be shared, for how long, and under what access terms, rather than a vague promissory line, is worth doing regardless of enforcement, both because ICMJE explicitly asks for it and because it is the more defensible position if a request does come in. But it should not stand in as the only claim in a manuscript that its results can be checked.

Frequently asked questions

Does declaring a data-sharing plan mean the data will actually be shared?

Not reliably. A cross-sectional audit of ten leading surgical journals found data obtainable for reanalysis in only 3.1 percent of trials both before and after the ICMJE policy took effect, and a separate audit of JAMA, Lancet, and NEJM found 68.6 percent of trials declared they would share data but only 0.6 percent had actually deposited it two years later.

How much work does fulfilling a data-sharing request take?

A survey of investigators who had granted at least one data request found a median of 18 person-hours, with a range of 3 to 125 hours, spent de-identifying, documenting, and packaging a single dataset, almost always without dedicated funding or staff time.

Is patient re-identification a real risk in small surgical datasets?

Yes. A pharmaceutical-industry privacy review found that cohorts with fewer than 12 participants carried materially elevated re-identification risk once indirect identifiers such as age, sex, and procedure type were combined, and a single-surgeon case series or a small pilot RCT commonly falls inside that risk band.

References

  1. International Committee of Medical Journal Editors. Clinical Trials, Data Sharing. In: Recommendations for the Conduct, Reporting, Editing, and Publication of Scholarly Work in Medical Journals. Updated 2023. https://www.icmje.org/recommendations/browse/publishing-and-editorial-issues/clinical-trial-registration.html
  2. National Institutes of Health, Office of Science Policy. NIH Data Management and Sharing Policy Overview. Effective January 25, 2023. https://grants.nih.gov/policy-and-compliance/policy-topics/sharing-policies/dms/policy-overview
  3. Bergeat D, Lombard N, Gasmi A, Le Floch B, Naudet F. Data Sharing and Reanalyses Among Randomized Clinical Trials Published in Surgical Journals Before and After Adoption of a Data Availability and Reproducibility Policy. JAMA Netw Open. 2022;5(6):e2215209. https://doi.org/10.1001/jamanetworkopen.2022.15209
  4. Danchev V, Min Y, Borghi J, Baiocchi M, Ioannidis JPA. Evaluation of Data Sharing After Implementation of the International Committee of Medical Journal Editors Data Sharing Statement Requirement. JAMA Netw Open. 2021;4(1):e2033972. https://doi.org/10.1001/jamanetworkopen.2020.33972
  5. Tannenbaum S, Ross JS, Krumholz HM, et al. Early Experiences With Journal Data Sharing Policies: A Survey of Published Clinical Trial Investigators. Ann Intern Med. 2018;169(8):586-588. https://doi.org/10.7326/M18-0723
  6. Maritsch F, Cil I, McKinnon C, Potash J, Baumgartner N, Philippon V, Pavlova BG. Data privacy protection in scientific publications: process implementation at a pharmaceutical company. BMC Med Ethics. 2022;23:65. https://doi.org/10.1186/s12910-022-00804-w

Sharing your raw dataset isn't the only route to a verifiable result, and the evidence above suggests it's currently an unreliable one, most shared datasets are never requested, and preparing one costs a median 18 uncompensated hours when they are. StatsPlease's deterministic engine gives you a second, lower-burden layer of verification that doesn't depend on anyone's data request being fulfilled: the exact test selected, the assumptions checked, and the resulting statistic are stated precisely enough in your Methods and Results that any reviewer can rerun the identical calculation in SPSS and get the identical number, before a data-sharing request is ever made.

Try StatsPlease free