There is no single good p-value. What counts as "good enough" depends on your field and on what a false alarm would cost you. The conventions run from 0.10 at the lenient end to 5×10⁻⁸ at the strict end.

A p-value is the probability of seeing data at least as extreme as yours if there were truly no effect. A smaller number means the data fit less comfortably with "nothing is going on." Below are the seven thresholds you will run into, ordered from most lenient to most demanding, with who uses each one and where the number comes from.

Threshold Where you see it Source
0.10 Business tests where missing a real effect is the bigger risk Atticus Li, Better Decisions
0.05 Default in social science, business analytics and A/B testing platforms Fisher, 1925; Statsig
0.01 Decisions where a false positive is costly; some clinical research Simply Psychology; Atticus Li
0.005 Proposed bar for new discovery claims Benjamin et al., Nature Human Behaviour
0.001 Labeled "highly significant"; high-stakes research Simply Psychology
5 sigma (about 3×10⁻⁷) High-energy physics Benjamin et al.
5×10⁻⁸ Genomics Benjamin et al.

0.10, when missing a real effect is the bigger risk

The loosest threshold on the list has no academic standing. It shows up in business practice as a deliberate trade-off. In an analyst-oriented explainer, Atticus Li argues that the cutoff should move with the stakes. He suggests raising it to 0.10 when missing a true effect is riskier than acting on a false one. His example is a 15% lift at p = 0.08, which is practically important but not significant at 0.05.

Know what you are giving up. On the evidence scale used by Statistics Fundamentals, values from 0.05 to 0.10 count as weak evidence against "no effect." If you use 0.10, set it before the test starts. Mida.so's A/B testing glossary recommends choosing the threshold in advance so the result can't nudge the rule.

0.05, the default almost everywhere

This is the line most analysts mean when they say "significant," and it is the standard cutoff across A/B testing platforms. Choosing it means accepting a 5% false-positive rate. Out of every 20 tests where nothing real is happening, about one will look like a winner by chance.

The number is a habit, not a law. Ronald Fisher proposed it in 1925 as a flexible rule of thumb.

How strong is the evidence at exactly p = 0.05? The authors of the 0.005 proposal (item 4) put it at a Bayes factor of roughly 2.5 to 1 up to 3.4 to 1 against "no effect." A Bayes factor is how many times better the data fit one explanation than the other, and they call that weak evidence. The share of real winners matters too. Li estimates that if only 10 of 100 changes you test truly work, about 35 to 40% of your "significant" results could be false positives.

One more trap: peeking. If you stop a test the moment p drops below 0.05, your false-positive rate rises above the 5% you thought you were accepting.

0.01 for expensive mistakes

Tightening to 0.01 means accepting a one-in-100 false-positive rate instead of one in 20. Simply Psychology notes that stricter thresholds like 0.01 are used in high-stakes research such as clinical trials. Li's business rule mirrors this: drop to 0.01 when acting on a false positive is costly. On the Statistics Fundamentals scale, values from 0.001 to 0.01 count as strong evidence.

The price is sample size. A stricter bar needs more data to detect the same effect, so decide this before you size the test, not after.

0.005: the proposed bar for new discoveries

In a paper titled "Redefine Statistical Significance" in Nature Human Behaviour, 72 researchers led by Daniel J. Benjamin proposed a change. For fields that use 0.05, they wanted the threshold for claiming a new discovery to drop to 0.005. Results between 0.005 and 0.05 would be called "suggestive" instead of "significant."

Their case rests on the evidence each line represents. At p = 0.005 the Bayes factor is roughly 14 to 1 up to 26 to 1 against "no effect," compared with about 3 to 1 at 0.05.

The proposal did not win the field. Harry Crane argued that its benefits disappear once you model p-hacking, the practice of trying analyses until one crosses the line. A separate camp, including Amrhein and McShane, argued for dropping thresholds altogether. A February 2026 review in the Journal of Visceral Surgery lists 0.005 or 0.001 as one option among several, not the answer.

0.001, labeled "highly significant"

Below 0.001 is where Simply Psychology applies the label "highly statistically significant." Statistics Fundamentals calls it very strong evidence against "no effect." Like 0.01, it appears in clinical trials and other high-stakes work.

There is also a reporting rule here. Under APA style, values this small are written as "p < .001" and never "p = .000." A p-value can never equal zero. If a dashboard or report shows 0.000, it has rounded the number away.

Five sigma in particle physics

High-energy physics states its standard in sigma, meaning standard deviations away from what "no effect" would predict. The discovery standard is 5 sigma, which works out to a p-value of about 3×10⁻⁷, or roughly three in ten million.

Benjamin and colleagues exempted physics from their proposal because it already uses a far stricter bar. They treated it as a field that had solved its own version of the problem.

If you read a physics result quoted in sigma, translate it this way. Five is the discovery standard, and anything less is not being claimed as a discovery.

5×10⁻⁸ in genomics

The strictest common threshold is 5×10⁻⁸, the convention in genomics. Benjamin et al. also named it as a field already beyond their proposal. It is a million times stricter than 0.05.

The 0.05 arithmetic shows why a field might go this far. If one in 20 tests with no real effect comes up significant, a study that runs a vast number of tests at 0.05 would produce false alarms by the thousand. Each extra test is another chance at a fluke, and a very strict line is one way to keep the total down.

Which line should you use?

For most business work, start with 0.05. Move to 0.01 when a false winner is expensive, or to 0.10 when missing a real one is. Make the choice before the data arrive.

No official body tells you any of these numbers is correct. The American Statistical Association's 2016 statement on p-values declined to endorse any cutoff, 0.05 and 0.005 included. A February 2026 review in the Journal of Evaluation in Clinical Practice found that 82.6% of the studies it covered criticized pass/fail use of p-values, and none defended the p-value as a standalone basis for decisions.

Whatever threshold you pick, report the size of the effect and its confidence interval alongside it. Even a tiny p-value can't tell you whether the effect is big enough to matter, which is why a small p-value still isn't proof.

The p-value was never intended to be a substitute for scientific reasoning.

Ron Wasserstein
Executive Director, American Statistical Association