← All posts

Real or luck? P-values, significance and statistical tests

How to tell a real difference from random noise. We start with ten coin flips you can count by hand, then use exactly the same idea on an A/B test.

You change the colour of a Buy button. With the old button, 1,000 out of 10,000 people buy. With the new one, 1,080 out of 10,000 do. The new button looks better.

But if you’d shown the old button to two different groups of 10,000 people, you wouldn’t have got the same number twice either. People are random. One week 1,000 buy, another week 1,040. So is the extra 80 the button, or just the usual noise?

That’s the question statistical tests answer, and p-values are how they report the answer. I’ve read the textbook definition of a p-value many times and it still slipped out of my head every time. So in this post I build it up from something small enough to count by hand, ten coin flips, and then take exactly the same idea back to the button.

If words like distribution, standard deviation or normal curve are new to you, my probability post covers them from scratch. You can also just read on. I explain each one briefly when it first comes up.

The idea in one paragraph

Before any maths, here’s the whole trick. Assume nothing interesting is going on. Work out how often pure chance would give you a result at least as extreme as yours. If that’s rare, start doubting the “nothing is going on” story.

It’s the same logic as a courtroom. The defendant is presumed innocent. The question the jury asks is: if they really were innocent, how surprising would this evidence be? Their fingerprints on the window: a bit surprising, but there could be an innocent reason. Fingerprints, CCTV footage and the stolen laptop in their car: so surprising that “innocent” stops being believable. A p-value is a number for exactly that: how surprising your data would be if nothing were going on.

Everything else in this post is detail around that one move.

Ten coin flips

A friend hands you a coin and says it’s fair. You flip it 10 times and get 8 heads. Do you believe them?

Step 1: assume the boring explanation

The boring explanation is “the coin is fair”: a 50% chance of heads on every flip, and each flip ignores the ones before it. In statistics the boring explanation is called the null hypothesis, written \(H_0\). “Null” as in no effect, nothing going on.

We’re not saying we believe it. We’re adopting it for a moment so we can ask a question we can actually answer: what does a fair coin do?

Step 2: watch what a fair coin does

The quickest way to get a feel for it is to watch. Below, one “trial” is 10 flips of a fair coin. Each finished trial adds one to the bar for its number of heads. Run a few hundred.

A few things jump out. Five heads is the most common result, but it only happens about a quarter of the time. Four and six are close behind. Eight or more is rare, but not unheard of: a perfectly fair coin produces it every now and then.

We can also get the exact numbers instead of simulating. Each flip has 2 outcomes, so ten flips can come out in 2 × 2 × … × 2 = 2¹⁰ = 1,024 different orders. With a fair coin, every one of those orders is equally likely.

Ten flips create 1024 sequences One flip gives two sequences, two gives four, three gives eight and ten gives 1024. The displayed third-flip sequences are only the four beginning with H. Each extra flip doubles the sequence list1 flip: 2 sequencesH · T2 flips: 4 sequencesHH · HT · TH · TT3 flips: 8 sequencesHHH · HHT · HTH · HTT… keep doubling …10 flips: 1,024 sequences2¹⁰ = multiply ten copies of 2
Each extra flip doubles the number of possible orders. With ten flips there are 1,024, all equally likely if the coin is fair.

Now sort those 1,024 orders by how many heads they contain. 252 of them have exactly 5 heads. 45 have exactly 8. 10 have 9, and only 1 has 10 (HHHHHHHHHH). Divide each count by 1,024 and you get the chart below, which is what your simulation was heading towards.

How often a fair coin gives each number of heads in 10 flips Bar chart of the binomial distribution for 10 fair flips. 5 heads is most likely at 24.6%. The bars for 8, 9 and 10 heads are highlighted; together they have probability 5.5%. The mirror bars for 0, 1 and 2 heads are also marked, bringing the two-sided total to 10.9%. 0 heads: 0.1% 0 1 heads: 1.0% 1 2 heads: 4.4% 2 4.4% 3 heads: 11.7% 3 11.7% 4 heads: 20.5% 4 20.5% 5 heads: 24.6% 5 24.6% 6 heads: 20.5% 6 20.5% 7 heads: 11.7% 7 11.7% 8 heads: 4.4% 8 4.4% 9 heads: 1.0% 9 10 heads: 0.1% 10 number of heads in 10 flips 8, 9 or 10 heads: 5.5% 0, 1 or 2 heads: 5.5% if the coin is fair · 1,024 equally likely sequences
Everything a fair coin can do in 10 flips. Five heads is the most likely result, at 24.6%. The blue bars are the results at least as far from 5 as our 8 heads.

Step 3: decide what “at least this extreme” means

We got 8 heads, which is 3 away from the 5 a fair coin would give on average. So the question becomes: how often does a fair coin land 3 or more away from 5? That includes 8, 9 and 10 heads, and also 2, 1 and 0. Two heads would be just as suspicious as eight. It’s the same distance from fair, just in the other direction.

Extreme means at least the observed distance Number line of heads counts from 0 through 10. Counts at least 3 away from 5 are highlighted: 0, 1, 2, 8, 9 and 10. Observed: 8 heads, distance 3 from 50123456789103 awayCount: 0, 1, 2 and 8, 9, 10Ignore: 3, 4, 5, 6, 7
Distance from 5 is what counts, not direction. Two heads is exactly as lopsided as eight.

Why not just ask how likely exactly 8 heads is? Because any single exact result is unlikely. Even the most common result, 5 heads, only happens a quarter of the time, and with a thousand flips every exact count would have a tiny chance. What we want to know is how far out towards the edges our result sits, so we count it together with everything even further out.

Step 4: add them up

8, 9 or 10 heads: 45 + 10 + 1 = 56 orders. 0, 1 or 2 heads: another 56. So:

\[p = \frac{56 + 56}{1024} = \frac{112}{1024} \approx 0.109\]

That’s the p-value, about 11%. In plain words: if the coin is fair, about 11 out of every 100 sets of ten flips would look at least this lopsided.

Try other results below. Click a bar to pretend that’s what you got.

Step 5: decide

Is 11% surprising enough to say the coin isn’t fair? Something that happens about one time in nine isn’t very rare.

To avoid arguing about it after the fact, you pick a cut-off before looking at the data. The usual cut-off is 5%. It’s called the significance level and written \(\alpha\) (alpha). If p comes out below \(\alpha\), the result is called statistically significant and you reject the null hypothesis.

Our 10.9% is above 5%, so we don’t reject “fair”. Be careful with what that means, though. It does not prove the coin is fair. Ten flips is very little evidence, and a coin that lands heads 70% of the time would often give results like ours too. Not significant means “not enough evidence”, which is a very different statement from “no effect”.

The 5% isn’t magic either. Ronald Fisher suggested it in the 1920s as a convenient line, and it stuck. A p-value of 4.9% and one of 5.1% are practically the same evidence, even though one gets called significant and the other doesn’t.

One-sided or two-sided

We counted both edges because a coin biased either way would be a problem. That’s a two-sided test. If, before flipping, you only cared whether the coin favours heads, you’d count only 8, 9 and 10 and get half the p-value, 5.5%. That’s a one-sided test.

The rule is to choose before you see the data. Choosing the side afterwards because it gives the smaller number is cheating, even if it doesn’t feel like it.

The formula, now that you’ve done it

The chance of exactly \(k\) heads in 10 fair flips is:

\[P(X = k) = \frac{\binom{10}{k}}{2^{10}}\]

\(X\) is the number of heads. \(\binom{10}{k}\), read “10 choose k”, counts the orders with \(k\) heads: 45 for \(k = 8\). The bottom counts all the orders. And the p-value is the sum of the bars we picked:

\[p = P(X \ge 8) + P(X \le 2)\]

“8 or more, plus 2 or fewer”. That’s the 112/1024 you already worked out, written shorter.

What a p-value is, and what it isn’t

Here’s the general version of what we just did:

\[p = P(\text{a result at least this extreme} \mid H_0)\]

Read the bar as “assuming”: the chance of a result at least this extreme, assuming the null hypothesis is true.

That “assuming” is where almost everyone goes wrong, me included for a long time. The p-value starts by assuming \(H_0\) is true, so it can’t then tell you the chance that \(H_0\) is true. A p-value of 3% does not mean:

  • a 3% chance there’s no real effect,
  • a 3% chance the result is a fluke,
  • a 97% chance your change works.

It means: if there were no real effect, results this extreme would turn up 3% of the time.

That can sound like hair-splitting, so here’s a case where the difference is obvious. A smoke alarm rarely goes off when there’s no fire: say 1% of the time. That’s like a p-value of 1%. Does it mean that when the alarm goes off, there’s a 99% chance of fire? Not in my kitchen. If you burn toast most mornings, most alarms are toast. “How often does it ring when there’s no fire?” and “now that it’s ringing, how likely is a fire?” are different questions. The p-value only answers the first one. (If you read the probability post, this is the “order matters when you say given” trap.)

Two ways to be wrong

Any rule that makes decisions from noisy data will sometimes get it wrong. There are exactly two ways it can happen:

Two ways a test can be wrong A two by two grid. Columns: the truth is no effect, or a real effect. Rows: you say it works, or you say nothing happened. No effect but you say it works is a false positive, with chance alpha. Real effect and you say it works is correct, with chance equal to power. No effect and you say nothing is correct. Real effect but you say nothing is a false negative, with chance beta. What is true vs. what you decideIn reality:No effectReal effectYou say"it works"You say"nothing"False positiveType I, chance αCorrectchance = powerCorrectchance 1 − αFalse negativeType II, chance β
The test only sees noisy data, never the truth. Orange cells are the two mistakes. You choose α, the false-positive rate, directly. The false-negative rate β depends on how much data you have and how big the real effect is.
  • A false positive (also called a Type I error) is declaring an effect that isn’t there. If there’s really nothing going on, this happens with chance \(\alpha\). At 5%, about 1 in 20 tests of useless changes will still come out “significant”.
  • A false negative (a Type II error) is missing an effect that is there. Its chance is written \(\beta\) (beta).

Power is the other side of a false negative: the chance your test catches a real effect of a given size.

\[\text{power} = 1 - \beta\]

Power depends on how big the real effect is and how much data you have. Big effects are easy to spot. Small ones need a lot of data. We’ll work it out for the button shortly.

Back to the button

Here’s the button experiment. Visitors were randomly split into two groups, often called arms: group A saw the old button, group B the new one.

Group Visitors Bought Conversion rate
A: old button 10,000 1,000 10.0%
B: new button 10,000 1,080 10.8%

B converts 0.8 percentage points better: 10.8% against 10.0%. As a relative change, that’s 8% better, because 0.8 is 8% of 10. Both are correct. Just be clear which one you mean, because people mix them up all the time.

The null hypothesis is that the button makes no difference: both groups have the same true conversion rate, and the 0.8-point gap is noise. We’ll use a two-sided test with \(\alpha = 5\%\), decided before the experiment started.

With the coin we could list every possible outcome. With 20,000 people that’s hopeless. But there’s a neat way to build the “nothing going on” world anyway.

Shuffle the labels

If the button truly makes no difference, then the labels A and B are meaningless. Someone who bought would have bought with either button, and someone who didn’t, wouldn’t have. So:

  1. Take all 20,000 people with their outcomes: 2,080 bought, 17,920 didn’t.
  2. Throw the A and B labels away, shuffle, and deal everyone into two new groups of 10,000.
  3. Work out the difference in conversion rate between the two new groups.
  4. Repeat thousands of times.
Permutation keeps the data and changes group labels The pooled 20,000 conversion outcomes stay fixed. Each shuffle assigns 10,000 outcomes to each group, calculates a new difference and counts absolute differences at least 0.008. Shuffle labels, not the outcomesKeep all 20,000 outcomes2,080 ones + 17,920 zerosRandomly split into two groups10,000 in A; 10,000 in BRecord shuffled B rate − A rateRepeat to build the null distributionCount |difference| ≥ 0.008Both negative and positive tails
The outcomes never change, only which group each person is dealt into. If the button doesn't matter, every deal is as plausible as the real one.

Each shuffle gives you a difference that came purely from how people happened to be split into groups. The pile of shuffled differences is the “nothing going on” world, just like the coin chart was. Then you count: how often is a shuffled difference at least as far from zero as our real +0.8?

After a few thousand shuffles you’ll settle at around 6.7%. That’s the p-value. This method is called a permutation test, and I like it a lot, because there’s no formula hiding anything. It’s the coin’s count-and-divide, done by a computer. Here it is in Python:

import numpy as np

rng = np.random.default_rng(0)
a = np.r_[np.ones(1000), np.zeros(9000)]   # old button: 1 = bought
b = np.r_[np.ones(1080), np.zeros(8920)]   # new button
observed = b.mean() - a.mean()             # 0.008
everyone = np.concatenate([a, b])

shuffles, extreme = 10_000, 0
for _ in range(shuffles):
    rng.shuffle(everyone)
    diff = everyone[10_000:].mean() - everyone[:10_000].mean()
    extreme += abs(diff) >= abs(observed) - 1e-12   # tiny tolerance for rounding

p = (extreme + 1) / (shuffles + 1)   # count the real split as one of the shuffles
print(p)  # about 0.067; it moves a little from run to run

One caution: shuffling only works when people were randomly assigned and are independent of each other. If the same person shows up several times, or whole households were assigned together, you have to shuffle those units together, or the “nothing going on” world you build is wrong.

The shortcut: the z-test

Shuffling works, but before computers it was impossible, and on huge datasets it’s still slow. Look at the shape the shuffled differences made, though: a bell. That’s the central limit theorem at work. A conversion rate is an average of lots of 0s and 1s, and averages come out bell-shaped.

If we know the null world is a bell centred on zero, we only need one more number to draw it: how wide it is. That width is the standard error (SE), the amount the difference between two groups typically wobbles from random sampling alone. Then we measure our result in “standard errors away from zero” and read the answer off the bell. That’s the two-proportion z-test. One step at a time:

1. Find the shared conversion rate. Under the null, both groups have the same rate, so pool them: 2,080 buyers out of 20,000.

\[\hat p = \frac{1000 + 1080}{10000 + 10000} = 0.104\]

The hat means “estimated from data”. Annoyingly, this \(\hat p\) is a conversion rate, not a p-value. Same letter, different job.

2. Work out the standard error of the difference.

\[SE = \sqrt{\hat p\,(1-\hat p)\left(\frac{1}{n_A} + \frac{1}{n_B}\right)}\]

You don’t need to memorise it, but you can read it. \(\hat p(1-\hat p)\) is how unpredictable a single buy-or-not outcome is. The \(1/n\) parts shrink it as the groups get bigger, where \(n_A\) and \(n_B\) are the group sizes. The square root turns it back into ordinary units. With our numbers:

\[SE = \sqrt{0.104 \times 0.896 \times \frac{2}{10000}} \approx 0.00432\]

So two identical buttons would typically differ by about 0.43 percentage points just by chance. Our gap is 0.8.

3. Count how many standard errors our gap is.

\[z = \frac{\hat p_B - \hat p_A}{SE} = \frac{0.108 - 0.100}{0.00432} \approx 1.85\]

This is the test statistic, one number for “how far out is my result”. It’s the same shape as a z-score: (what I saw − what the null expects) / (typical wobble). Here the null expects 0.

4. Read the p-value off the bell. If the null is true, \(z\) follows a standard normal curve: a bell centred on 0 with a standard deviation of 1. We want the area beyond ±1.85:

Where the observed z-score falls on the null distribution A standard normal curve: the distribution of the z-score if the button change did nothing. The observed z of 1.85 is marked. The area beyond plus and minus 1.85 is shaded and totals 6.4%, the p-value. The cut-offs at plus and minus 1.96, which leave 5% in the tails, are drawn as dashed lines just outside it. Right tail beyond z = 1.85: 3.2% Left tail beyond z = −1.85: 3.2% observed z = 1.85 cut-off 1.96 −1.96 3.2% 3.2% -4 -3 -2 -1 0 1 2 3 4 z-score z-scores you would see if the change did nothing · shaded tails = p-value = 6.4%
The bell is where z would land if the button did nothing. The shaded tails beyond ±1.85 add up to 6.4%: that's the p-value. The dashed lines at ±1.96 are the 5% cut-off, and our z stops just short of them.
\[p = 2 \times \big(1 - \Phi(\lvert z\rvert)\big) \approx 0.064\]

In pieces:

  • \(\lvert z\rvert\) is z without its sign: 1.85.
  • \(\Phi\) (“phi”) is the area of the standard normal curve to the left of a point. \(\Phi(1.85) \approx 0.968\).
  • \(1 - 0.968 = 0.032\) is the area in the right tail.
  • Double it for the left tail: 0.064.

So p ≈ 6.4%, very close to the 6.7% from shuffling. Two completely different methods, nearly the same answer. (The small gap is because the bell is a smooth approximation of what are really lumpy, whole-number counts.) Tick the overlay box in the shuffle widget and you’ll see the bell sitting right on top of the shuffled bars.

5. Decide. 6.4% is above our 5% line, so the result is not significant. We haven’t shown the new button is better. We also haven’t shown it isn’t. The dashed lines at ±1.96 in the chart are where the 5% cut-off sits: \(z\) has to get past 1.96 to count as significant, and ours stopped at 1.85.

You won’t do this by hand in practice. In Python it’s one line:

from statsmodels.stats.proportion import proportions_ztest

z, p = proportions_ztest(count=[1080, 1000], nobs=[10000, 10000])
print(f"z = {z:.2f}, p = {p:.3f}")  # z = 1.85, p = 0.064

Same effect, more data

Now imagine we’d run the test with 20,000 people per group instead, and got exactly the same rates, 10.0% and 10.8%.

The effect is the same. But the standard error shrinks with more data, by a square root: double the people and the wobble gets divided by √2 ≈ 1.41.

\[SE \approx 0.00305, \qquad z \approx 2.62, \qquad p \approx 0.009\]

Now it’s significant. Nothing about the button changed. We just measured it more precisely. This is worth sitting with for a second: a p-value mixes up how big an effect is with how much data you collected. With enough users, a tiny, useless effect becomes “significant”. With too few, a big, valuable one doesn’t.

That’s a comparison between two separate experiments, by the way. It is not permission to keep adding users to a finished test until it crosses the line. More on that in the traps section.

Power: how many users do you need?

Before you run the experiment, ask: if the new button really does add 0.8 points, how likely is this test to notice?

With 10,000 people per group, the answer is only about 46%. That’s worse than a coin flip. You would miss a real, worthwhile improvement more often than you’d find it. With 20,000 per group it’s about 75%.

The widget shows why. Imagine running the same experiment many times. The blue curve is where your measured lift would land if the button did nothing. The purple curve is where it would land if the true lift really is +0.8 points. Anything outside the dashed lines counts as significant. Power is the share of the purple curve that ends up outside them.

The sample size formula

You can also run that backwards: pick the power you want, usually 80%, and solve for the number of people. For two equal-sized groups, a standard approximation is:

\[n \approx \frac{\left(z_{1-\alpha/2} + z_{1-\beta}\right)^2 \left[p_A(1-p_A) + p_B(1-p_B)\right]}{(p_B - p_A)^2}\]

It looks scary, so piece by piece:

  • \(n\) is people per group, not in total.
  • \(p_A\) and \(p_B\) are the rates you’re planning around: 10% and 10.8%. You choose them before the experiment, based on the smallest improvement you’d care about.
  • \(z_{1-\alpha/2} \approx 1.96\) is the cut-off for a two-sided test at 5%. \(z_{1-\beta} \approx 0.84\) is the extra distance needed for 80% power.
  • The bracket is how noisy each group is.
  • The bottom is the effect, squared. So halving the effect you want to detect means four times as many people.

With our numbers:

\[n \approx \frac{(1.96 + 0.84)^2 \times (0.090 + 0.096)}{0.008^2} \approx 22{,}800\]

So roughly 23,000 people per group. Put 23,000 into the widget and you’ll see power land at about 80%.

Power is a planning tool. After the experiment, working out “observed power” from your own result doesn’t tell you anything the p-value hasn’t already. Look at the confidence interval instead.

Confidence intervals: how big is the effect?

A p-value tells you whether a result is surprising. It doesn’t tell you how big the effect is, and that’s usually what you actually want to know. For that, use a confidence interval: your estimate, plus or minus a margin for the wobble.

\[\text{interval} = (\hat p_B - \hat p_A) \pm 1.96 \times SE\]

For our button that’s +0.80 ± 0.85 points, so the 95% confidence interval runs from −0.05 to +1.65 percentage points.

Read it as: the data are consistent with anything from “very slightly worse” to “1.65 points better”. That’s far more useful than “not significant”. It tells you zero is still possible, but so is a big win, and that the honest next step is to collect more data.

(Small detail: for the interval, the SE is worked out from each group’s own rate instead of the pooled one, because we’re no longer assuming the two rates are equal. With these numbers it comes out at 0.00432 either way.)

What does the “95%” mean? It’s about the method, not this one interval. If you repeated the experiment many times and built an interval each time, about 95% of those intervals would contain the true difference.

Three results with their 95% confidence intervals Estimated lift in conversion rate, in percentage points, with 95% confidence intervals. With 10,000 users per arm: +0.80, interval −0.05 to +1.65, which crosses zero. With 20,000 per arm: +0.80, interval +0.20 to +1.40. With 2 million per arm: +0.10, interval +0.04 to +0.16, far from zero but also far below the +0.5 point lift that would matter to the business. no effect worth shipping: +0.5 10,000 per arm 10,000 per arm: +0.80 points, 95% CI -0.05 to +1.65, p = 0.064 p = 0.064 20,000 per arm 20,000 per arm: +0.80 points, 95% CI +0.20 to +1.40, p = 0.009 p = 0.009 2,000,000 per arm 2,000,000 per arm: +0.10 points, 95% CI +0.04 to +0.16, p = 0.0009 p = 0.0009 -0.5 +0.0 +0.5 +1.0 +1.5 +2.0 lift in conversion rate, percentage points
Three experiments. The first can't rule out zero. The second can. The third is hugely "significant", yet its whole interval sits below the +0.5-point lift that would be worth shipping.

Before you run a test, decide the smallest lift that would be worth the work, then compare the interval against that line, not just against zero. Statistically significant and worth doing are different questions.

Three ways to fool yourself

1. Testing lots of things at once

Say you check 20 metrics on a button that changes nothing. Each test has a 5% chance of a false alarm. The chance that at least one of them comes out significant, if the tests are independent, is:

\[1 - (1 - 0.05)^{20} \approx 64\%\]

Each test has a 95% chance of staying quiet. All 20 staying quiet is 0.95 multiplied by itself 20 times, about 36%. Everything else, 64%, is at least one false alarm. Look at enough metrics and something will always light up.

Here’s what p-values look like across many experiments. When nothing is going on, they’re spread evenly between 0 and 1, so 5% land below 0.05 by pure chance. When there’s a real effect, they pile up near zero.

Distribution of p-values with and without a real effect Two histograms of p-values from 2,000 simulated experiments with 10,000 users per arm. Left, an A/A test where nothing changed: the p-values are spread evenly from 0 to 1, and 5.0% fall below 0.05. Right, a real lift from 10% to 10.8%: p-values pile up near zero, and 48% fall below 0.05. no real effect (A/A) p between 0.00 and 0.05: 99 of 2,000 p between 0.05 and 0.10: 111 of 2,000 p between 0.10 and 0.15: 113 of 2,000 p between 0.15 and 0.20: 105 of 2,000 p between 0.20 and 0.25: 99 of 2,000 p between 0.25 and 0.30: 95 of 2,000 p between 0.30 and 0.35: 113 of 2,000 p between 0.35 and 0.40: 109 of 2,000 p between 0.40 and 0.45: 92 of 2,000 p between 0.45 and 0.50: 98 of 2,000 p between 0.50 and 0.55: 83 of 2,000 p between 0.55 and 0.60: 103 of 2,000 p between 0.60 and 0.65: 97 of 2,000 p between 0.65 and 0.70: 91 of 2,000 p between 0.70 and 0.75: 102 of 2,000 p between 0.75 and 0.80: 107 of 2,000 p between 0.80 and 0.85: 84 of 2,000 p between 0.85 and 0.90: 105 of 2,000 p between 0.90 and 0.95: 113 of 2,000 p between 0.95 and 1.00: 81 of 2,000 p < 0.05: 5.0% 0 0.5 1 p-value real lift 10% → 10.8% p between 0.00 and 0.05: 962 of 2,000 p between 0.05 and 0.10: 247 of 2,000 p between 0.10 and 0.15: 149 of 2,000 p between 0.15 and 0.20: 102 of 2,000 p between 0.20 and 0.25: 87 of 2,000 p between 0.25 and 0.30: 68 of 2,000 p between 0.30 and 0.35: 51 of 2,000 p between 0.35 and 0.40: 47 of 2,000 p between 0.40 and 0.45: 39 of 2,000 p between 0.45 and 0.50: 43 of 2,000 p between 0.50 and 0.55: 31 of 2,000 p between 0.55 and 0.60: 30 of 2,000 p between 0.60 and 0.65: 18 of 2,000 p between 0.65 and 0.70: 19 of 2,000 p between 0.70 and 0.75: 13 of 2,000 p between 0.75 and 0.80: 24 of 2,000 p between 0.80 and 0.85: 19 of 2,000 p between 0.85 and 0.90: 17 of 2,000 p between 0.90 and 0.95: 15 of 2,000 p between 0.95 and 1.00: 19 of 2,000 p < 0.05: 48.1% 0 0.5 1 p-value 2,000 simulated experiments per panel · 10,000 users per arm
P-values from 2,000 simulated experiments each. Left: no real effect, and the p-values are flat, with 5% below 0.05 by chance. Right: a real lift, and the p-values pile up near zero, but at this sample size only about half get below 0.05. That's the low power from earlier.

The fix: pick one primary metric before the experiment. If you must make claims about many tests, tighten the cut-off. The simplest way is the Bonferroni correction: divide \(\alpha\) by the number of tests, so 0.05 / 20 = 0.0025 each. The Benjamini-Hochberg procedure is less strict and controls the share of your “discoveries” that are false, rather than the chance of any false alarm at all.

2. Peeking

You plan a two-week test. On day 3 you check: p = 0.04! You stop and ship.

The trouble is that the 5% false alarm rate only holds if you look once, at the end. Every peek is another chance for noise to wander across the line, and noise wanders a lot early on, when there’s little data. In a quick simulation of a button that does nothing, checked 20 times along the way and stopped at the first p below 0.05, about a quarter of the experiments ended in a false “win”.

Either decide the sample size in advance and look once, or use a method built for repeated looks (sequential testing). An ordinary p-value doesn’t protect you from peeking.

3. Reading “not significant” as “no effect”

Back to the button: estimated +0.80 points, interval −0.05 to +1.65, p ≈ 0.064. The data leave room for no improvement and for a useful one. “Not significant” just means this experiment couldn’t tell. If you need to show that two things are effectively the same, that’s a different test (an equivalence test), with its own planning.

Choosing a test

The coin and the button used tests for counts and rates. Other data needs other tests, but they mostly share the same shape: (the difference you saw) / (how much it would wobble by chance), compared against what that ratio looks like when nothing is going on.

Two questions get you most of the way to the right one:

  1. What is one observation? One independent user? The same person measured twice? A whole household? Getting this wrong breaks most tests.
  2. What are you measuring? A yes/no outcome gives you a rate. Revenue or time on page gives you numbers with an average.
Situation Common test Watch out for
Two independent conversion rates Two-proportion z-test Very small counts need an exact test
Two independent averages Welch’s t-test Big outliers, such as a few huge orders
The same people measured twice Paired t-test Test the per-person differences
Two groups, skewed numbers Mann-Whitney U test It compares rankings, not averages
Three or more groups ANOVA, or similar Then adjust for multiple comparisons
Randomised groups you can reshuffle Permutation test Shuffle the unit that was randomised

For averages, Welch’s t-test has the familiar shape:

\[t = \frac{\bar x_B - \bar x_A}{\sqrt{s_A^2/n_A + s_B^2/n_B}}\]

\(\bar x_A\) and \(\bar x_B\) are the two group averages. \(s_A^2\) and \(s_B^2\) are their variances, how spread out each group is. Difference on top, wobble underneath, same as before. The result is compared against a t-distribution, which is a bell with slightly fatter tails to account for the extra uncertainty of small samples.

from scipy import stats

# one revenue value per independent user in each array
t, p = stats.ttest_ind(revenue_b, revenue_a, equal_var=False)  # Welch's t-test

The checklist

When I read or run a test now, I go through it in this order:

  1. What’s the boring explanation? Write down the null hypothesis, and decide one- or two-sided, before looking at the data.
  2. How big is the difference? Keep the units: +0.8 percentage points, not just “significant”.
  3. How much could it wobble by chance? Check that observations are independent and the standard error makes sense.
  4. How surprising is it if nothing is going on? That’s the p-value, and it is not the chance the null is true.
  5. What rule did we agree on? Compare against \(\alpha\) without moving the goalposts, switching metrics or stopping early.
  6. What would we actually do? Compare the confidence interval with the smallest lift worth shipping, and plan the next test’s sample size for an effect that matters.

The coin is the whole idea in miniature: assume fair, count how often fair looks this lopsided, add up the blue bars. Every test after that is the same question with a different way of counting.

References

  1. Wasserstein, R. L. and Lazar, N. A. The ASA Statement on p-Values: Context, Process, and Purpose. The American Statistician, 2016. https://doi.org/10.1080/00031305.2016.1154108. What p-values do and don’t mean.
  2. Greenland, S. et al. Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations. European Journal of Epidemiology, 2016. https://doi.org/10.1007/s10654-016-0149-3. A long list of common misreadings.
  3. Neyman, J. and Pearson, E. S. On the problem of the most efficient tests of statistical hypotheses. Philosophical Transactions of the Royal Society A, 1933. https://doi.org/10.1098/rsta.1933.0009. Where Type I and Type II errors come from.
  4. Fisher, R. A. Statistical Methods for Research Workers. Oliver and Boyd, 1925. https://psychclassics.yorku.ca/Fisher/Methods/. The origin of the 0.05 convention.
  5. Welch, B. L. The generalization of “Student’s” problem when several different population variances are involved. Biometrika, 1947. https://doi.org/10.1093/biomet/34.1-2.28. The unequal-variance t-test.
  6. statsmodels developers. statsmodels.stats.proportion.proportions_ztest. https://www.statsmodels.org/stable/generated/statsmodels.stats.proportion.proportions_ztest.html.
  7. SciPy developers. scipy.stats.ttest_ind. https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.ttest_ind.html. Welch’s t-test via equal_var=False.
  8. Benjamini, Y. and Hochberg, Y. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society: Series B, 1995. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x.
  9. Johari, R. et al. Peeking at A/B Tests. Proceedings of KDD, 2017. https://doi.org/10.1145/3097983.3097992. Why repeated looks inflate false positives, and what to do instead.
  10. Kohavi, R., Tang, D. and Xu, Y. Trustworthy Online Controlled Experiments. Cambridge University Press, 2020. https://doi.org/10.1017/9781108653985. The practical guide to A/B testing.