Why is the p-value cutoff 0.05? Ronald Fisher, a British statistician, called it convenient in a 1925 textbook, and nearly everyone followed him. It is a habit with some arithmetic behind it. It is not a constant that anyone derived from anything deeper.
The habit matters because the line makes decisions. It decides whether an A/B test gets called a winner, whether a study gets written up as a finding, and whether a change on a dashboard gets treated as real. A p-value is the probability of seeing data at least as extreme as yours if there were truly no effect. The figure 0.05 is simply the point where people agreed to start calling that probability small.
Where 1 in 20 came from
Fisher set out the cutoff in Statistical Methods for Research Workers in 1925. His reasoning, as quoted in a University of Tennessee explainer, was this: "p = 0.05, or 1 in 20, is 1.96 or nearly 2... deviations exceeding twice the standard deviation are thus formally regarded as significant."
A standard deviation measures how far results typically scatter from the average. On a bell curve, about 95% of results land within roughly two standard deviations of the center. Fisher wanted a rule of thumb for when a result was worth a second look, and "more than about two standard deviations out" was a round, easy one. The 0.05 is that rule stated as a probability.
Nothing in the mathematics picks 2 over 2.5 or 3. Fisher picked a number that was easy to remember. He also didn't treat it as fixed. He used different thresholds depending on the research, and in 1956 he wrote that "no scientific worker has a fixed level of significance at which from year to year, and in all circumstances, he rejects hypotheses."
How a rule of thumb became a rule
The rigidity came from somewhere else. In 1933 Jerzy Neyman and Egon Pearson built a competing system for hypothesis testing. In their version you set an error rate, called alpha, before collecting any data, and then you accept or reject. Alpha is the share of false alarms you are willing to tolerate when there is really nothing there.
Fisher saw the p-value as a sliding measure of evidence. Neyman and Pearson saw a pass or fail decision. A 2025 paper in PLOS Mental Health traces how the two approaches were merged in the mid-20th century, even though they conflict. The result was a pass or fail decision with Fisher's number as the bar. Researchers kept his 0.05 and treated it as a fixed alpha, and the habit of sorting results into "significant" and "not significant" followed.
The 0.05 line, from rule of thumb to argument
-
1925Fisher calls 1 in 20 convenient
Proposed in Statistical Methods for Research Workers as a guide to which results deserve attention.
-
1933Neyman and Pearson formalize alpha
A preset error rate and an accept or reject decision, set before the data come in.
-
2015A journal bans p-values
Basic and Applied Social Psychology removes them from its pages entirely.
-
2016The ASA issues a statement
Its first formal position on a specific statistical method warns against bright-line thresholds.
-
201772 researchers propose 0.005
The stricter bar would apply to claims of new discoveries, with results between 0.05 and 0.005 labeled "suggestive."
-
201943 papers and a sharper editorial
A special issue, "Moving to a World Beyond p<0.05," is published alongside an editorial saying "statistically significant: don't say it and don't use it."
-
2021The ASA walks it back
A Presidential Task Force clarifies that the 2019 editorial was not official ASA policy.
Is anyone trying to move it?
Yes, from two directions. One camp wants a stricter line. The 2017 proposal by 72 researchers would move claims of new discoveries to 0.005, which is ten times harder to reach.
The other camp objects to having any line at all. The American Statistical Association's 2016 statement, by Ronald Wasserstein and Nicole Lazar, said that "scientific conclusions and business or policy decisions should not be based only on whether a p-value passes a specific threshold," and that widespread use of p < 0.05 "has resulted, among other things, in a large number of false discoveries." It did not call for banning p-values.
A working paper by Patrick Vu and Stefan Faridani puts a number on that worry. They analyzed three large replication projects and estimated that a finding with a p-value of exactly 0.05 has a 0.10 to 0.25 chance of replicating, depending on the field. A same-sized repeat of a just-significant result was about three times as likely to come back insignificant as significant.
The argument is not settled. Philosopher Deborah Mayo argued in March 2026 that error-control thresholds still help science catch failed claims. Ten years after the ASA statement, no agreed replacement for 0.05 is in common use. Lazar's own position is that "there should be no one-size-fits-all solution."
What 0.05 does not mean
It does not mean there is a 5% chance the result is a fluke. As Kyle Hewitt put it in The Conversation, the p-value "tells you how likely a difference this big is if nothing real is going on, not how likely it is that nothing real is going on." The distinction is spelled out in what a p-value tells you.
It also does not mark the point where evidence switches on. Results at p = 0.049 and p = 0.051 are nearly identical, yet the first gets reported as a discovery and the second as nothing. Rosnow and Rosenthal joked about this in 1989: "surely, God loves the .06 nearly as much as the .05."
Reading a result that sits near the line
Say a checkout test lands at p = 0.051 and someone wants to call it a loss, or at 0.049 and someone wants to ship it. Treat both as the same weak result. Then ask the questions the line skips. How big is the effect, and is the confidence interval around it narrow enough to act on? How many variants or metrics were tested before this one came in under 0.05? Was the sample size fixed in advance, or did the test run until it crossed?
The last question matters most. Results that clear the bar get reported and the rest get dropped, which is how clean, dramatic results crowd out the messy ones. If your team uses 0.05, use it knowingly. It is a convention you are choosing because a statistician found it convenient a century ago. It is not a property of your data.
Comments
No comments yet. Be the first to comment!
Leave a Comment