Domain 4 of 6

Designing an Experiment

Turning a question into a test that can answer it. What randomisation buys, the hypothesis and the criterion a result is judged against, how many users are needed and for how long, the guardrails that catch damage a winning metric hides, and what to do when an experiment is impossible.

6
Concepts
~18%
Of the exam
16
Practice questions
Concepts in this domain
01Controlled experimentsRandom assignment as the only practical route from correlation to cause, what the control and treatment groups are for, how the assignment unit is chosen and where it goes wrong, and the question an experiment answers that no amount of observation can.02Hypotheses and the overall evaluation criterionThe hypothesis as a prediction specific enough to fail, the four parts it needs, the overall evaluation criterion set out by Kohavi, Tang and Xu in 2020, and a worked composite where the weights agreed in advance decide whether the same result reads as a win or a loss.03Sample size and statistical powerThe minimum detectable effect, statistical power and the relationship that ties them to the number of people a test needs, a table running from a one per cent improvement to a fifty per cent one with the arithmetic shown, and what follows honestly for a product with modest traffic.04Guardrail metricsMeasures watched for damage rather than for improvement, the two kinds Kohavi, Tang and Xu separate, a table of common guardrails with what each one protects, and the arithmetic showing how a two per cent win can be a loss once latency is priced.05Experiment duration and exposureThe three things that set how long a test runs, why the answer comes in whole weeks, the exposure point that decides who is counted, and the arithmetic showing that a one per cent ramp needs about twenty five times the duration of an even split.06Alternatives to a controlled experimentThe changes that cannot be randomised into two groups, the four designs available when a split is impossible, what each one assumes and how much smaller the claim becomes once the randomisation has gone.
Try a question from this domain

A car parking app wants to know whether a rebuilt pricing screen raises the share of sessions ending in a paid booking. An analyst proposes building two comparison groups matched on city, device and length of time as a customer, rather than assigning drivers by a coin toss. What does the matched design give up?

  • ABalance on every attribute nobody recorded, since matching controls only the attributes somebody thought to collect and leaves the rest free to differ.
  • BNothing, as long as the two matched groups are the same size and cover the same weeks.
  • CThe ability to report a confidence interval, since an interval can only be computed for a randomly assigned comparison.
  • DSome sensitivity, since a matched design needs a larger sample than a randomised one to detect the same effect.
16 questions on this domain.

One per page, with a worked explanation.

Start the set