If you have both a p-value and a confidence interval for the same result, read the interval first. The two come from the same arithmetic, so a 95% confidence interval that leaves out zero is the same finding as p < .05. The interval also keeps two things the p-value throws away: how big the effect is and how precisely it was measured.
That is the whole difference. A p-value answers one question: would data like this be surprising if nothing were going on? An interval answers that question too. It also shows the range of effect sizes the data are consistent with, in the units you care about, such as conversion points, dollars or days.
The two side by side
| P-value | 95% confidence interval | |
|---|---|---|
| What it reports | One probability between 0 and 1 | A range of values on the original measurement scale |
| Formal meaning | Probability of data at least this extreme if the null hypothesis were true (MetricGate) | A range from a method that captures the true value in about 95% of repeated experiments (Munro) |
| Shows effect size | No | Yes, through the estimate at its center |
| Shows precision | No | Yes, through its width |
| Shows direction of effect | No, for a two-sided test | Yes, through the sign of its bounds |
| Signals significance at .05 (two-sided) | Yes, when p < .05 | Yes, when the interval excludes zero (StatsTest) |
| Most common misreading | "The probability the null hypothesis is true" | "A 95% chance the true value is in this range" |
| Hypothetical tests A and B below | p ≈ 0.046 for both | 0.01 to 0.99 points vs 0.1 to 9.9 points |
The p-value
A p-value is the probability of getting data at least as extreme as yours if the null hypothesis, the assumption that there is no effect, were true. It is a single number, and it is easy to misread; the pillar page covers each way it gets misread.
Its strength is that it is compact and everywhere. A meta-research preprint co-authored by John Ioannidis, covering more than 22 million PubMed abstracts from 1990 to 2025, found that the share of PubMed Central full-text articles reporting p-values rose from 5.2% to 53.3%. The median number per article rose from 2 to 7. That growth happened despite decades of calls to report intervals instead.
Its weakness is that it blends two things into one figure: the size of the effect and the size of the sample. A huge sample can make a trivial effect look impressive, and a small sample can make a real effect look like nothing. An article in PLOS Mental Health gives one example of each. A meditation study with 500 participants produced p < 0.001 for an effect of d = 0.10, which is negligible. (Cohen's d measures an effect in standard deviations.) A depression treatment study produced p = 0.08 for an effect of d = 0.45, which is clinically meaningful, and it was dismissed as ineffective because it was too small to detect that effect reliably.
The p-value is the right tool when all you need is a quick screen for whether a result clears a pre-agreed bar. It is the wrong tool for deciding whether an effect is worth acting on.
The 95% confidence interval
A confidence interval is a range around your estimate. It is built by a method that, over many repeated experiments, would capture the true value about 95% of the time. Because it is stated in the original units, it answers the question a manager actually asks: how much?
Professor Nadeem Shafique Butt, a biostatistician, puts its advantage plainly: "Intervals reveal effect magnitude and uncertainty, helping distinguish statistically significant but clinically trivial findings."
It has limits too. The 95% level means about one interval in 20 will miss the true value purely by chance. An interval also reflects only the uncertainty its model accounts for. In polling, for example, Nate Silver notes that real-world error has historically been wider than the theoretical margin of error. And even careful readers underuse intervals. A review of 64 papers from three leading epidemiology journals in 2022, published in Global Epidemiology, found no outright misinterpretations. It did find that discussion of intervals was often incomplete and rarely considered the full range of plausible effects.
The interval is the right tool for anyone deciding whether to ship, spend or change something.
Why a 95% interval that skips zero means p below .05
For most everyday tests, a 95% interval is the estimate plus or minus about 1.96 standard errors. The standard error is how much the estimate would bounce around if you reran the experiment on a fresh sample.
The interval excludes zero exactly when the estimate sits more than 1.96 standard errors away from zero. That is the same condition that gives a two-sided p-value below .05. The two are not separate pieces of evidence that happen to agree. They are one calculation reported two ways. The same logic links a 99% interval to p < .01. If you want the mechanics, see the steps for turning a test statistic into a p-value.
This equivalence holds when both numbers come from the same estimate and the same standard error. It does not extend to comparing two separate intervals. The StatsTest blog warns that two groups' intervals can overlap even when the difference between the groups is statistically significant.
Same p-value, very different news
Two A/B tests on one dashboard
Say your dashboard shows two finished tests, each measuring the lift in conversion rate in percentage points. Test A ran on a high-traffic homepage banner. Test B ran on a low-traffic checkout page. Both report the same p-value.
| Test A estimated lift | 0.5 points |
|---|---|
| Test A standard error | 0.25 points |
| Test A: estimate ÷ standard error | 2.0 |
| Test A 95% interval (0.5 ± 1.96 × 0.25) | 0.01 to 0.99 points |
| Test B estimated lift | 5 points |
| Test B standard error | 2.5 points |
| Test B: estimate ÷ standard error | 2.0 |
| Test B 95% interval (5 ± 1.96 × 2.5) | 0.1 to 9.9 points |
| Two-sided p-value, both tests | about 0.046 |
The p-values are identical because both estimates sit 2.0 standard errors from zero. The intervals tell two different stories. Test A is measured tightly: the lift is real but almost certainly under 1 point, so the decision is whether a small, reliable gain is worth the work. Test B could be nearly nothing or nearly 10 points. That is a promising result that needs more traffic before anyone builds a forecast on the 5.
Width matters even when the center stays the same. Silver's election model offers a clean illustration from outside A/B testing. A candidate with a 3-point lead and a 5-point margin of error wins about 88% of the time. With the same lead and a 10-point margin, she wins only 72% of the time. The point estimate did not move; the decision-relevant probability did.
The interval has a misreading of its own
The p-value's famous error is treating it as the probability the null hypothesis is true. The interval's parallel error is reading "95% confidence" as "a 95% chance the true value is inside this range." Pediatrician and researcher Alasdair Munro's answer to that belief is "This is false."
The 95% describes the method, not any single interval. Run the experiment many times, and about 95% of the intervals you build will contain the true value. Your particular interval either does or does not.
Munro's example is a trial of vitamin C for sepsis that reported a relative risk of 0.8, with a 95% interval of 0.7 to 0.9. Strong prior evidence says the therapy does not work, so he argues that the realistic chance the true effect lies in that range is "closer to 0% than 95%." A safer reading of any interval is the range of effects most compatible with the data you collected.
Which number to trust when they are both on the screen
Read the confidence interval, and treat the p-value as a summary of one fact the interval already contains.
If the whole interval sits in a range you would act on, the result is strong enough to act on. If the whole interval sits in a range too small to matter, a p-value below .05 does not change that. If the interval spans both trivial and important values, as Test B does, you do not have an answer yet. The fix is more data, not a second look at the p-value.
When you report a result to someone else, give the estimate and the interval in the original units. Add the p-value if your audience expects it, but never give it alone.
Comments
No comments yet. Be the first to comment!
Leave a Comment