The part that changes the product. Deciding to ship, iterate or abandon on the evidence available, what happens to a measure once a team is judged against it, the questions behavioural data cannot answer at all, writing a finding somebody who was absent can act on, and building a practice that accumulates knowledge rather than dashboards.
A crane hire booking product ran two tests. The first reports no significant difference with an interval running from a 0.4 per cent loss to a 0.5 per cent gain. The second reports no significant difference with an interval running from a 9 per cent loss to a 10 per cent gain. A summary slide lists both as flat. What does that slide hide?
ANothing that matters, since neither test cleared its threshold and both changes should therefore be removed.
BThat the second test is the stronger evidence of the two, since a wider interval has covered more of the possible outcomes.
CThat the two support opposite actions. The first has settled the question and the code can be deleted, and the second failed to answer it, so the honest report says the test could not detect an effect rather than saying the change did nothing.
DThat neither interval means anything without its p value, since an interval containing zero cannot be read on its own.