Turning a question into a test that can answer it. What randomisation buys, the hypothesis and the criterion a result is judged against, how many users are needed and for how long, the guardrails that catch damage a winning metric hides, and what to do when an experiment is impossible.
A car parking app wants to know whether a rebuilt pricing screen raises the share of sessions ending in a paid booking. An analyst proposes building two comparison groups matched on city, device and length of time as a customer, rather than assigning drivers by a coin toss. What does the matched design give up?
ABalance on every attribute nobody recorded, since matching controls only the attributes somebody thought to collect and leaves the rest free to differ.
BNothing, as long as the two matched groups are the same size and cover the same weeks.
CThe ability to report a confidence interval, since an interval can only be computed for a randomly assigned comparison.
DSome sensitivity, since a matched design needs a larger sample than a randomised one to detect the same effect.