Statistical power is the probability that an experiment will detect an effect of a stated size, given that the effect is really there, and sample size is the quantity a team buys that probability with.
An agreed criterion decides how a result will be read. Whether the test can produce a readable result at all is settled earlier, by arithmetic that takes half an hour and saves a fortnight of measuring nothing.
The framework came from Jerzy Neyman and Egon Pearson in the early 1930s, who separated the two ways a test can be wrong. A false positive reports a difference that is not there, and a false negative misses a difference that is. Conventional practice fixes the first at 5 per cent and the second at 20 per cent, which is where the familiar pair of 95 per cent confidence and 80 per cent power comes from. Neither figure is a law of nature, and both are choices a team could make differently.
This lands on a product manager as a scheduling problem before it lands as a statistical one. The question is always how long the test has to run, and the answer depends on how small an effect the team wants to be able to see. Asking for a smaller effect is the most expensive request in experimentation, and almost nobody making it realises the price.
The sections below define the three quantities a calculation needs. The relationship between them follows, with the rule of thumb that makes it arithmetic anybody can do. A table then gives sizes for a worked baseline, and the last sections cover what all of this means for a small product and why an underpowered win overstates itself.
The three quantities a calculation needs
The baseline rate. What the measure does today in the control group. A conversion of 4 per cent and a conversion of 40 per cent need very different sample sizes for the same relative improvement, because the variance of a rate depends on the rate itself.
The minimum detectable effect. The smallest difference the team wants the test to be able to see. Stated as a relative figure it is usually a percentage improvement on the baseline, so a 10 per cent minimum detectable effect on a 4 per cent baseline means a treatment rate of 4.4 per cent.
Power. The probability that a test detects an effect of exactly that size if it is really present. At 80 per cent power, a change delivering exactly the minimum detectable effect will be found four times in five and missed once in five.
Two further choices sit underneath and are usually left at their defaults. The significance level fixes the false positive rate at 5 per cent, and a two sided test spends that 5 per cent on both directions, since a change can make things worse as easily as better.
The relationship between effect size and sample size
The number of people a test needs rises with the variance of the measure and falls with the square of the effect the team wants to detect. That square is where the cost lives. Halving the effect a test can see multiplies the people required by four, and dividing it by ten multiplies them by a hundred.
Kohavi and colleagues published a workable rule of thumb at a conference in 2014 and repeated it in the 2020 book. For 80 per cent power and a two sided test at 5 per cent significance, each group needs roughly sixteen times the variance divided by the square of the difference the team wants to detect.
The sixteen comes from the two conventions. The critical value for 5 per cent two sided significance is 1.96, the value for 80 per cent power is 0.84, and their sum squared is 7.84. Doubling that for two groups gives 15.68, which rounds to 16 and is close enough for planning.
For a rate, the variance is the baseline rate multiplied by one minus the baseline rate. At a 4 per cent baseline that is 0.04 times 0.96, or 0.0384. Sixteen times 0.0384 is 0.6144, and every figure in the table below is that number divided by the square of the absolute difference.
A table of effect against users required
The table takes a 4 per cent baseline conversion, 80 per cent power and 5 per cent two sided significance. The treatment rate column shows what the improvement means in absolute terms, and the last two columns are per group and in total across a fifty fifty split.
| Improvement sought | Treatment rate | Absolute difference | Users per group | Users in total |
|---|---|---|---|---|
| 1 per cent relative | 4.04 per cent | 0.0004 | 3,840,000 | 7,680,000 |
| 2 per cent relative | 4.08 per cent | 0.0008 | 960,000 | 1,920,000 |
| 5 per cent relative | 4.20 per cent | 0.0020 | 153,600 | 307,200 |
| 10 per cent relative | 4.40 per cent | 0.0040 | 38,400 | 76,800 |
| 20 per cent relative | 4.80 per cent | 0.0080 | 9,600 | 19,200 |
| 50 per cent relative | 6.00 per cent | 0.0200 | 1,536 | 3,072 |
Any row can be checked in one division. For the 10 per cent row the absolute difference is 0.004, its square is 0.000016, and 0.6144 divided by 0.000016 is 38,400. For the 5 per cent row the difference is 0.002, its square is 0.000004, and the same division gives 153,600.
Reading down the table shows the square at work. Going from a 20 per cent improvement to a 10 per cent one multiplies the requirement from 9,600 to 38,400, which is four times. Going from 10 per cent to 5 per cent multiplies it again, and going from 2 per cent to 1 per cent takes it from 960,000 to 3,840,000 per group.
The last row is the one worth dwelling on. A change that lifts conversion by half needs 1,536 people per group, which almost any product can supply in a fortnight. Large effects are cheap to measure and small effects are ruinous, which inverts the usual assumption that careful measurement is for small refinements.
What a product with modest traffic can test
Those requirements turn into a calendar the moment a team divides by its own weekly traffic. Take a product receiving twelve thousand new signups a week and running tests on all of them at a fifty fifty split, which gives six thousand people per group each week. The table converts directly into a calendar.
A 20 per cent relative improvement needs 9,600 per group, which is 1.6 weeks and rounds up to two. A 10 per cent improvement needs 38,400 per group, which is 6.4 weeks and rounds up to seven. A 5 per cent improvement needs 153,600 per group, which is 25.6 weeks, or roughly six months of the entire product's traffic spent on one question.
Four responses follow honestly from a calendar like that, and the fourth is the one teams skip.
- Test larger changes. A product with limited traffic can afford to ask whether a rebuilt flow beats the old one and cannot afford to ask about button copy. That constraint is a useful filter on the backlog.
- Use a more sensitive measure. A criterion closer to the change, such as completion of the step being altered, has a higher baseline rate and therefore a smaller requirement than a downstream conversion.
- Narrow the population. Running the test only on the people who reach the changed screen removes everybody the change could never affect, which raises the observable effect without changing the product at all.
- Decide without an experiment. Some questions cannot be settled by a test at this traffic, and the professional answer is to say so and use judgement, research or a reversible release.
Deciding without an experiment needs stating plainly, because the alternative teams reach for is running the underpowered test anyway and treating whatever comes back as evidence.
The exaggeration in an underpowered result
An underpowered test does not simply fail more often. When it does produce a significant result, that result is systematically too large, and the reason is mechanical. A small sample only clears the significance threshold when the observed difference is big, so the significant results are drawn from the extreme tail of what the experiment could have measured.
Andrew Gelman and John Carlin described this in 2014 as a magnitude error, and they proposed reporting the exaggeration ratio alongside power. Their second finding is worse. In a very noisy study, a significant result carries a real chance of pointing in the wrong direction entirely.
For a product team the consequence is concrete. An underpowered test that reports a 12 per cent lift is most likely measuring a true effect far smaller than 12 per cent, so the forecast built on that number will not arrive and nobody will know why. Running the calculation first is what prevents the whole sequence.
A correctly sized test still watches one number. Nothing in this arithmetic notices that the winning change doubled page load time or trebled support contacts, and catching that damage takes a second set of measures agreed in advance.