Why does the same small lift look like noise in a small test and a sure thing in a big one? Because a p-value depends on two things, not one. It reflects how big the difference is and how much data you used to measure it. Add enough data and almost any real difference, however tiny, will eventually produce a small p-value.

A p-value is the probability of seeing data at least this extreme if there were truly no difference between the versions. The American Statistical Association said what follows from that in its 2016 statement on p-values: "Any effect, no matter how tiny, can produce a small p-value if the sample size or measurement precision is high enough, and large effects may produce unimpressive p-values if the sample size is small or measurements are imprecise." This is why an A/B test can move from "nothing here" to "significant" while the product stays exactly the same.

How sample size gets into the p-value

Most tests a dashboard runs come down to one ratio. The top of the ratio is the difference you observed. The bottom is the standard error, which is how much that difference would bounce around if you reran the test on fresh visitors. The ratio is the test statistic (a z or t score), and the p-value is read from it. A bigger ratio gives a smaller p.

Sample size affects only the bottom of the ratio. The standard error shrinks with the square root of the number of observations. Collect 100 times the data and the standard error falls to a tenth of its size, so the same difference produces a test statistic ten times larger. Researchers writing in Scientific Reports stated it bluntly: "The larger the sample size, the smaller the p-value."

One caveat for real tests. In practice the observed lift wobbles from one sample to the next. The example below holds it fixed so you can see the sample size working on its own.

One lift, three sample sizes

Say a checkout test runs at a 5.0% baseline conversion rate, and the variant converts at 5.4%. That is a lift of 0.4 percentage points. Pooled across both arms, the conversion rate is 5.2%. For a two-proportion test, the standard error of the difference is √(2 × 0.052 × 0.948 ÷ n), where n is the number of visitors per arm.

500 per arm: 25 vs 27 conversionsSE = 1.40 pts, z = 0.29, p ≈ 0.78
5,000 per arm: 250 vs 270 conversionsSE = 0.44 pts, z = 0.90, p ≈ 0.37
50,000 per arm: 2,500 vs 2,700 conversionsSE = 0.14 pts, z = 2.85, p ≈ 0.004

The lift is 0.4 points in every row. The jump from 500 to 50,000 visitors per arm multiplied the data by 100, cut the standard error by a factor of 10, and raised z by a factor of 10. That alone moved the result from "nothing to see" to well under 0.05.

If you want the step from a z or t score to the p-value itself, it is covered in calculating a p-value by hand.

Where it shows up

Two studies with the same p-value and very different effects. MetricGate simulated two studies that both landed at p ≈ 0.04. Study A had 12 people per group and a large effect: Cohen's d of 0.85, where d is the difference in means divided by the standard deviation. Study B had 400 people per group and a tiny effect, d = 0.13. The effects differ by a factor of 6.5, yet the p-values match. Statistician Richard Morey made the general point in 2016: "A moderate discrepancy with a moderate sample size can yield the same p-value as a very tiny discrepancy with a large sample size."

Resources moved on the strength of sample size. A paper in PLOS Mental Health describes a meditation study with 500 participants that reported p < 0.001. The effect was d = 0.10, a 1.5-point drop on a 100-point anxiety scale. Institutions still shifted resources toward meditation programs. The same paper describes the reverse case. A depression intervention had p = 0.08 but a meaningful d = 0.45, it was written off as ineffective, and its funding was eliminated.

The coin that is technically unfair. MetricGate also notes that with 10 million flips you can detect a coin landing heads 50.01% of the time. "The p-value will be tiny, but no one would call the coin unfair."

Two things this gets confused with

"Big samples make everything significant." They do not. Big samples make real effects detectable. They cannot create an effect that is not there. A Monte Carlo study in the Journal of Evaluation in Clinical Practice ran 1,000 simulations at each of 11 sample sizes, from 5 to 800 per group. When the two groups truly had no difference, the share of results with p < 0.05 stayed near 5% at every sample size. When a small real difference existed (means of 4.5 and 4.4, standard deviation 0.5), power, the chance of detecting the difference, rose from 5.4% to 98.1%.

"A smaller p-value means a bigger effect." It does not. The ASA lists this among its six principles: a p-value does not measure the size of an effect or the importance of a result. In the worked example above, the lift was the same in every row. Only the certainty about it changed.

What to check when a test turns significant

Before the test starts, decide the smallest lift that would be worth shipping. When a result comes in, read the lift itself first and the p-value second. A p of 0.004 on a lift too small to matter tells you that you measured something small very precisely. That is all it tells you.

Ask for three numbers together, as MetricGate recommends: the estimated lift, its confidence interval, and the p-value. In the AJPM Focus simulations, effect estimates stayed fairly stable as samples grew while the intervals narrowed. The interval shows you both how big the effect is and how sure you can be about it.

Read a small test with care in the other direction too. A large p-value from a few hundred visitors per arm does not show that there is no effect. It may only mean the test was too small to see one.

If huge samples make tiny effects significant, what does significance tell you?

The pillar page explains what a p-value measures, the common ways it gets misread, and why statistical significance is not the same as an effect worth acting on.

Read what a p-value tells you