Statistical philosophy · ~9 min read
ASA Statement on Statistical Significance (2019)
In 2019 the ASA's own journal called for retiring the phrase “statistically significant.” Other statisticians, and the ASA itself, pushed back within two years.
Published
In short
In 2019, the American Statistical Association published a 43-paper special issue of The American Statistician whose lead editorial, "Moving to a World Beyond 'p<0.05,'" went further than the ASA's widely cited 2016 statement and called for retiring the phrase "statistically significant" altogether.2 Statisticians did not close ranks behind it: John Ioannidis argued in JAMA that some form of significance filter remains necessary to separate signal from noise,3 and Deborah Mayo and David Hand argued in Synthese that an account with no threshold for what counts as evidence is not a test at all.4 By 2021, the ASA's own president's task force stepped in to say the 2019 editorial did not represent official ASA policy and that p-values and significance tests, properly applied, "should not be abandoned."5 Meanwhile the New England Journal of Medicine's 2019 statistical guidelines moved in the opposite direction from the "abandon significance" call, keeping significance thresholds and tightening requirements around them.6 This is a live, unresolved methodological disagreement among statisticians, not a settled convention your manuscript is failing to follow.
Two ASA documents, not one
It is easy to conflate the ASA's 2016 statement with its 2019 follow-up, because both concern the same six-decade-old habit of treating a P value of .05 as a bright line. They are not the same document, and they do not make the same claim. The 2016 statement, led by Ronald Wasserstein and Nicole Lazar, was a consensus document built by more than two dozen statisticians over several months of deliberation, and its core message was narrower than it is often remembered as: a P value measures how incompatible the data are with a specified statistical model, it is not a measure of the probability that a hypothesis is true, and a P value near .05 by itself offers only weak evidence against the null.1 It criticised the misuse of the threshold. It did not call for abandoning the word "significant."
The 2019 document went further. Wasserstein, Allen Schirm, and Lazar wrote the lead editorial for a special supplement to The American Statistician titled "Moving to a World Beyond 'p<0.05,'" introducing 43 papers from statisticians debating what should replace dichotomous testing.2 The editorial's own recommendation was explicit and considerably more sweeping than 2016: stop using the phrase "statistically significant" entirely, whether attached to a P value, a confidence interval, or any other statistical measure, because the label has become disconnected from the underlying evidence it is supposed to summarise. That distinction matters for a manuscript, because reviewers and editors who cite "the ASA position" are frequently thinking of one of these two documents while meaning the other.
The case for retiring the word
The argument behind the 2019 editorial is not that P values are useless. It is that the word "significant" invites a false binary: a result either clears an arbitrary threshold or it does not, and everything about the magnitude, precision, and context of the finding gets compressed into that single pass-or-fail label. A P of .049 and a P of .051 usually sit on nearly identical underlying effect estimates, yet the label sorts them into opposite categories. The editorial's proposed alternative was to describe what the data are compatible with directly, reporting the point estimate and its interval rather than a verdict, and to stop treating .05 as a discovery threshold at all, in any direction.
Where other statisticians drew the line
The pushback did not come from defenders of P-hacking or careless threshold use. It came from statisticians who accepted the diagnosis in the 2016 statement but thought the 2019 prescription overcorrected. Writing in JAMA, John Ioannidis argued that "a significance filter in some form is essential" for separating signal from noise, particularly in fields such as genome-wide association studies where the overwhelming majority of tested associations are null by design; removing the filter without improving statistical training, he wrote, leaves researchers to describe their data according to unstated biases instead.3 He also pointed out that a 1990 attempt at exactly this policy, when the journal Epidemiology banned significance testing outright, was not widely followed by other journals over the following three decades, which he read as evidence the reform does not travel well in practice.
Deborah Mayo and David Hand made a more structural version of the same objection in Synthese: if an account of evidence cannot specify outcomes that would count against a claim, because every threshold has been abandoned, then it is not functioning as a test at all.4 Their concern was less about the word "significant" itself and more about what tends to accompany calls to drop it, namely the loosening of any pre-specified boundary between what a study's data support and what they do not. Both critiques share a common thread: the 2016 statement's complaint was about how P values get misused, while the 2019 editorial's remedy was aimed at the vocabulary, and the two are not obviously the same fix for the same problem.
The ASA's own board stepped back in
The disagreement was not confined to outside commentators. Concerned that the 2019 editorial was being read as official ASA policy, then-ASA president Karen Kafadar convened a president's task force, and in 2021 the group, including statisticians such as Yoav Benjamini, Bradley Efron, Xiao-Li Meng, and Kafadar herself, published its own statement in the Annals of Applied Statistics.5 It stated plainly that p-values and significance tests, "when properly applied and interpreted," are important tools that should not be abandoned, and that thresholds should be set contextually rather than eliminated wholesale. That a sitting ASA president felt the need to convene a task force specifically to clarify that the 2019 editorial did not speak for the association is itself a fairly direct signal of how contested the "retire the word" position was within the field, not just outside it.
What actually changed in the journals
Whatever the outcome inside the statistics profession, it is worth checking what happened where researchers actually submit manuscripts. The New England Journal of Medicine issued new statistical reporting guidelines in 2019, and they did not follow the "abandon significance" recommendation.6 The journal kept P-value thresholds and continued to permit "statistically significant" language for pre-specified primary analyses, while tightening rather than loosening its handling of multiplicity: secondary and exploratory analyses that were not adjusted for multiple comparisons were restricted to confidence intervals with an explicit caveat that they should not be used to infer significance. That is closer to the opposite of what the 2019 editorial proposed than a confirmation of it. This is one journal's policy, not evidence of a field-wide shift, but it is a concrete, citable data point against assuming the "retire significance" position simply became the new normal in clinical publishing.
What to do with this in your own manuscript
None of this settles which side is right, and it is not this post's job to settle it either. What it should settle is a narrower question: whether your Discussion section needs to defend a specific stance on the word "significant," or whether it needs to follow whatever your target journal's own instructions for authors specify. Check that first, since journals differ and at least one major one has explicitly kept the conventional language. Independent of that choice, the underlying reporting habits both sides of this debate agree on are worth following regardless: report the exact P value rather than only whether it cleared .05, report the effect size and its confidence interval alongside it, and avoid hedge phrases like "trending toward significance" that neither camp defends. Those practices satisfy the spirit of the 2016 statement, the 2019 editorial, and its critics all at once, because the disagreement between them is about vocabulary and thresholds, not about whether readers deserve the full estimate.
Frequently asked questions
Did the ASA officially ban the phrase "statistically significant"?
No. The call to retire the phrase came from the lead editorial of a 2019 special issue of The American Statistician, written by three of its editors. In 2021 the ASA's own president's task force stated that the 2019 editorial did not represent official ASA policy, and that p-values and significance tests, properly applied, should not be abandoned.
What is the difference between the ASA's 2016 and 2019 statements?
The 2016 statement was a consensus document that criticized the misuse of P values and the .05 threshold, without calling for the word "significant" to be dropped. The 2019 editorial went further and explicitly recommended retiring the phrase "statistically significant" entirely. They are two different documents making two different claims, and are often conflated.
Should I still use the word "significant" in my manuscript?
This is unresolved among statisticians, so follow your target journal's own instructions for authors rather than either side of the debate. Independent of that choice, report the exact P value, the effect size, and its confidence interval, and avoid hedge phrases like "trending toward significance," since both sides of the debate agree on those practices.
You might also read
References
- Wasserstein RL, Lazar NA. The ASA Statement on p-Values: Context, Process, and Purpose. Am Stat. 2016;70(2):129-133. https://doi.org/10.1080/00031305.2016.1154108
- Wasserstein RL, Schirm AL, Lazar NA. Moving to a World Beyond "p<0.05". Am Stat. 2019;73(sup1):1-19. https://doi.org/10.1080/00031305.2019.1583913
- Ioannidis JPA. The Importance of Predefined Rules and Prespecified Statistical Analyses: Do Not Abandon Significance. JAMA. 2019;321(21):2067-2068. https://doi.org/10.1001/jama.2019.4582
- Mayo DG, Hand D. Statistical Significance and Its Critics: Practicing Damaging Science, or Damaging Scientific Practice? Synthese. 2022;200:220. https://doi.org/10.1007/s11229-022-03692-0
- Benjamini Y, De Veaux RD, Efron B, et al. The ASA President's Task Force Statement on Statistical Significance and Replicability. Ann Appl Stat. 2021;15(3):1084-1085. https://doi.org/10.1214/21-AOAS1501
- Harrington D, D'Agostino RB Sr, Gatsonis C, et al. New Guidelines for Statistical Reporting in the Journal. N Engl J Med. 2019;381(3):285-286. https://doi.org/10.1056/NEJMe1906559
Whichever convention your target journal follows on the word "significant," StatsPlease's deterministic engine reports the exact P value your data produced, alongside the effect size and confidence interval, every time, so your Results section states what the test actually found rather than a borrowed side of an unsettled argument.
Try StatsPlease free