Preprints · ~7 min read
Preprints vs Peer Review in Clinical Research
A fraudulent ivermectin preprint reached meta-analyses before anyone checked it. Most preprints are not that story, but knowing the difference is the point.
Published
In short
Preprint servers such as medRxiv let clinical and surgical researchers share findings months before formal peer review, and the pandemic made that mainstream: medRxiv, launched in June 2019, had received roughly 41,000 submissions by 2022, with about 54% of them COVID-19-related.1 Most of the time this works: one large comparison of over 1,000 epidemiological estimates found preprint point estimates changed by an average of only 6% after peer review,2 and conclusions changed materially in only a minority of cases.3 But the exceptions matter. A fraudulent ivermectin preprint was folded into meta-analyses that shaped clinical advocacy for the drug before its data problems were identified and it was retracted.4 Preprints also get retracted faster than journal articles, a median of 29 days versus 139,5 which is a real advantage for catching errors quickly, but only if a reader knows to check.
A server that barely existed before 2020
medRxiv (pronounced "med-archive") launched in June 2019 as a joint effort of Cold Spring Harbor Laboratory, Yale University, and BMJ, built on the model that bioRxiv had already established for biology. It posts manuscripts that have not yet been through formal peer review, with a screening step for obvious harms but no assessment of methodological quality. By the time a group of researchers involved in running the server, including its co-founders, presented a full accounting of its first roughly three years of operation, medRxiv had received about 41,000 submissions and posted around 34,000 of them, drawing in some 185,000 authors from about 155 countries.1 Roughly 54% of everything posted was related to COVID-19.1 A server that had existed for barely seven months before the pandemic began became, almost overnight, one of the main channels through which new clinical findings on a fast-moving disease reached other researchers, journalists, and the public, well before any of it had been vetted by peer reviewers.
The case for posting before review
The argument for preprints in clinical and surgical research is straightforward: peer review is slow, and slow has a cost when the question is time-sensitive. A surgical trial finishing enrolment during a device shortage, or a treatment comparison relevant to an active outbreak, can sit in a journal's review queue for months while surgeons and physicians make decisions without it. Posting a preprint establishes priority, makes the full methodology available for scrutiny by anyone, not just two or three assigned reviewers, and lets other groups begin replicating or extending the work immediately. Of the roughly 34,000 preprints medRxiv had posted through its first three years, about 39% went on to be published in a peer-reviewed journal, with a median gap of around five months between posting and that eventual publication.1 For a clinician trying to make a decision in month one, that five-month gap is the entire argument for reading the preprint rather than waiting.
What actually changes between preprint and published version
The strongest evidence on whether this trade-off is reasonable comes from studies that tracked the same paper through both stages. One analysis matched more than 1,000 epidemiological estimates across 100 preprint-publication pairs and found that point estimates shifted by an average of just 6% during peer review, with a correlation of 0.99 between the preprint and published values; confidence intervals narrowed by about 7% on average as review tightened the analysis.2 A separate analysis of bioRxiv and medRxiv preprints from the earliest phase of the pandemic looked specifically at whether a paper's stated conclusions changed between versions, and found a discrete change in about 7.2% of non-COVID-19 abstracts and 17.2% of COVID-19 abstracts, though the authors noted that most of even these changes did not qualitatively alter the paper's conclusions.3 Read together, these findings are reassuring for the median preprint: the numbers mostly hold up. They are not reassuring about the tail, because a 17% rate of some change to a COVID-19 paper's conclusions, across thousands of pandemic-era preprints, is not a small number of papers in absolute terms, and neither study can tell a reader in advance which preprint they are looking at.
When an unreviewed claim reaches practice
The clearest documented case of that tail risk in the pandemic literature is the Elgazzar ivermectin preprint. Posted to the preprint server Research Square, the study reported a substantial survival benefit for ivermectin in COVID-19 and was, before any of its problems were identified, incorporated into multiple meta-analyses used to argue for the drug's clinical benefit; in the most prominent of these, the Elgazzar data alone accounted for an estimated 12.6% of the pooled effect on survival.4 It was retracted from Research Square in July 2021 after reviewers found that data for around 79 participants appeared duplicated, that some recorded participant deaths predated the trial's own start date, and that sections of the text had been plagiarised.4 By the time the retraction happened, the underlying preprint had already been cited in pooled analyses that fed public and clinical advocacy for ivermectin use well beyond what any single, unreviewed, and ultimately fraudulent dataset should have supported. The episode is not a story about preprints being uniquely dangerous; the two retracted hydroxychloroquine papers that shaped early pandemic guidance, in the Lancet and New England Journal of Medicine, had both gone through formal peer review.6 It is a story about what happens when a single unreviewed dataset enters a synthesis before anyone has checked it, in a research environment moving too fast for the normal safeguards to catch up.
Retraction happens faster on preprint servers, if anyone is watching
One genuine advantage preprints have shown during the pandemic is speed of correction, not just speed of dissemination. A comparison of COVID-19 papers retracted between January 2020 and March 2022 found that preprints were retracted after a median of 29 days, compared with a median of 139 days for peer-reviewed journal articles,5 and preprints in that sample also drew substantially more social media attention on average, which may partly explain why problems surfaced faster.5 That speed depends entirely on scrutiny actually happening: open commentary, replication attempts, and journalists or clinicians checking the underlying data. A preprint nobody reads critically corrects just as slowly as a journal article nobody reads critically, and unlike a retracted journal article, a withdrawn preprint has historically lacked consistent labelling standards across servers, so a stale, unmarked copy can keep circulating on secondary sites after the version of record has been pulled.
What this means for the manuscript you are about to write
None of this argues against posting a preprint of a surgical or clinical study. It argues for being precise about what a preprint is and is not. It is not peer-reviewed evidence, and citing one in your own manuscript's introduction or discussion should say so explicitly rather than treating it as an equivalent source to a published article. If you are citing preprint data in a systematic review or meta-analysis, that choice needs its own justification and, ideally, a note about what would change if the preprint is later revised, exactly the kind of change the tracking studies above show does happen in a meaningful minority of cases.23 And if you are posting your own findings early, precision cuts the other way too: the Methods and Results a reviewer or reader encounters in your preprint should be reproducible exactly as written, because for however many months it sits in front of readers before formal review catches up, it is the only version of your findings anyone has.
Frequently asked questions
Should I cite a preprint the same way I cite a peer-reviewed article?
No. A preprint is not peer-reviewed evidence, and citing one in your manuscript's introduction or discussion should say so explicitly rather than treating it as an equivalent source to a published article. If you are citing preprint data in a systematic review or meta-analysis, that choice needs its own justification.
How much do preprint results typically change after peer review?
Not much, on average. One analysis of over 1,000 epidemiological estimates found point estimates shifted by an average of just 6% during peer review, with a correlation of 0.99 between the preprint and published values. Stated conclusions changed in a minority of cases, about 7% of non-COVID-19 abstracts and 17% of COVID-19 abstracts in one pandemic-era analysis.
Are preprints retracted faster than peer-reviewed journal articles?
Yes, in the COVID-19 literature: a median of 29 days for preprints versus 139 days for peer-reviewed articles. That speed depends entirely on scrutiny actually happening, open commentary, replication attempts, and readers checking the underlying data, so a preprint nobody reads critically corrects just as slowly as a journal article nobody reads critically.
You might also read
References
- Ross JS, Sever R, Bloom T, Hindle S, Yunusov D, Roeder T, Inglis JR, Krumholz HM. medRxiv preprint submissions, posts, and key metrics, 2019-2021. Presented at: 9th International Congress on Peer Review and Scientific Publication; September 2022.
- Nelson LJ, Cortés-Corrales S, Piper JT, et al. Robustness of evidence reported in preprints during peer review. Lancet Glob Health. 2022;10(12):e1684-e1687. https://doi.org/10.1016/S2214-109X(22)00368-0
- Brierley L, Nanni F, Polka JK, et al. Tracking changes between preprint posting and journal publication during a pandemic. PLoS Biol. 2022;20(2):e3001285. https://doi.org/10.1371/journal.pbio.3001285
- Hill A, Mirchandani M, Pilkington V. Ivermectin for COVID-19: addressing potential bias and medical fraud. Open Forum Infect Dis. 2022;9(2):ofab645. https://doi.org/10.1093/ofid/ofab645
- Sra M, Baraskar B, Vishwakarma H, et al. Comparative analysis of retracted pre-print and peer-reviewed articles on COVID-19. medRxiv. Preprint posted online July 15, 2022. https://doi.org/10.1101/2022.07.12.22277529
- Mehra MR, Desai SS, Ruschitzka F, Patel AN. RETRACTED: hydroxychloroquine or chloroquine with or without a macrolide for treatment of COVID-19: a multinational registry analysis. Lancet. 2020;395(10240):1820. https://doi.org/10.1016/S0140-6736(20)31180-6
Whether your next manuscript is heading to a preprint server first or straight into peer review, the numbers in it should not depend on which stage a reader catches it at. StatsPlease's deterministic engine computes every statistic from the dataset you upload and states exactly which test, model, and assumptions it used, so a reviewer, a co-author, or a reader who finds your preprint months before formal review can rerun the same analysis in SPSS and get the identical result. That reproducibility matters most in exactly the window this post is about: after you post, before anyone has formally checked your work.
Try StatsPlease free