Domain 5 of 6

Reading a Result Honestly

The five ways a correctly run experiment still misleads the team that ran it. What significance does and does not claim, why checking early inflates the false positive rate, the split that proves the plumbing is broken, effects that fade once novelty wears off, and what happens when enough metrics are tested at once.

5
Concepts
~16%
Of the exam
14
Practice questions
Concepts in this domain
01Statistical significanceWhat a p value claims and what it cannot claim, what a confidence interval adds that a verdict of significant throws away, and why a difference can be statistically significant and still too small to fund.02The peeking problemWhat repeated checking does to the false positive rate of a running test, the published figures for one check against ten, why the habit is invisible from inside a team and the two legitimate ways to stop a test early.03Sample ratio mismatchThe validity check that runs before any result is interpreted, the chi squared arithmetic behind it, the five stages of an experiment where a mismatch is created and why a mismatched test is discarded.04Novelty and primacy effectsThe effect that fades once a change stops being new, the effect that grows once people have learned their way around it, the two readings that separate them from noise and the published evidence on how often either one decides anything.05Multiple comparisonsWhy examining many measures at the five per cent threshold produces winners out of nothing, the arithmetic for twenty measures and for ninety six segment cuts, the two corrections available and the practice that avoids needing either.
Try a question from this domain

A supermarket loyalty app tested a shorter signup form. Each group received 120,000 visitors, the original produced 5,400 signups at 4.50 per cent and the shorter form produced 5,640 at 4.70 per cent, and the p value for the 0.20 point gap is 0.019. A colleague reports a 98 per cent chance that the shorter form is better. What does the p value actually say?

  • AThat the result would come back the same way in about 98 tests out of every hundred if the experiment were repeated.
  • BThat if the shorter form changed nothing at all, a gap this wide or wider would turn up in about two tests in every hundred.
  • CThat there is a 98 per cent chance the shorter form is better, which is what a p value below 0.05 licenses a team to claim.
  • DThat the shorter form is 98 per cent likely to deliver a gain of at least 0.20 percentage points.
14 questions on this domain.

One per page, with a worked explanation.

Start the set