The five ways a correctly run experiment still misleads the team that ran it. What significance does and does not claim, why checking early inflates the false positive rate, the split that proves the plumbing is broken, effects that fade once novelty wears off, and what happens when enough metrics are tested at once.
A supermarket loyalty app tested a shorter signup form. Each group received 120,000 visitors, the original produced 5,400 signups at 4.50 per cent and the shorter form produced 5,640 at 4.70 per cent, and the p value for the 0.20 point gap is 0.019. A colleague reports a 98 per cent chance that the shorter form is better. What does the p value actually say?
AThat the result would come back the same way in about 98 tests out of every hundred if the experiment were repeated.
BThat if the shorter form changed nothing at all, a gap this wide or wider would turn up in about two tests in every hundred.
CThat there is a 98 per cent chance the shorter form is better, which is what a p value below 0.05 licenses a team to claim.
DThat the shorter form is 98 per cent likely to deliver a gain of at least 0.20 percentage points.