Choosing a product metric is the work of deriving one measure from a claim the product strategy makes, then testing the candidate before a team commits a quarter of its time to it.
The alternative arrives by default and costs nothing to adopt, which is why so many teams end up with it. Every analytics tool reports a standard set on the day somebody installs it. Sessions, page views, a bounce rate and a count of daily active users all appear without anybody asking for them. A team that starts from that list is steering by whatever a vendor found easiest to collect across every product it sells to, and none of those numbers knows what this particular product promised anybody.
The cost of that shortcut shows up two quarters later. A team reports that daily active users rose six per cent, a director asks whether the product is better at the thing it was funded to do, and nobody in the room can answer from the number on the screen. The measure was never connected to the claim the strategy makes, so no movement in it can confirm or deny the claim.
The sections below derive a measure from a strategy with a worked case. They then separate the three roles a measure can play, which are goal, driver and guardrail. Finally they give the five tests a candidate has to pass and the two ways a chosen measure turns out to be useless.
Deriving a measure from the claim a strategy makes
A strategy makes a claim about why customers will choose this product, and a measure is derived by asking what would be observably true if the claim held.
The worked case is a product sold to housing associations for managing repairs to their properties. Its strategy says the product wins mid sized associations by cutting the number of repair visits that fail because the engineer turned up without the right part. That sentence is the claim. Everything below comes out of it.
The derivation follows the three steps Rodden, Hutchinson and Fu set out in 2010 as Goals Signals Metrics, applied to the claim rather than to a feature.
- The goal. An engineer arrives at a property able to finish the repair on that visit.
- The signal. A repair job is closed on the first visit and no second visit is booked for the same fault inside the next thirty days.
- The metric. The share of repair jobs closed on the first visit, counted monthly, over all jobs opened in that month and resolved within thirty days of opening.
Naming the metric is where most of the argument happens, because a measure is only as good as its denominator. Counting the share over all jobs opened in a month would include jobs still open on the last day, which would drag the figure down for reasons that have nothing to do with parts. Resolving the denominator to jobs opened and closed within thirty days makes the figure stable and states the limit of what it covers. A measure with no denominator written down is an argument waiting to happen, because the first person to rebuild the query will choose a different one and report a different number.
Goal measures, driver measures and guardrail measures
A single measure rarely travels alone, and the set around it has a structure worth naming.
Ron Kohavi, Diane Tang and Ya Xu set out that structure in Trustworthy Online Controlled Experiments in 2020, separating three roles. A goal measure states what the organisation is ultimately after, which is usually slow moving and hard for any one team to shift. A driver measure sits closer to the work and moves sooner, which is what makes it usable inside a quarter. A guardrail measure is watched for damage, so it never improves as a result of good work and only ever warns that something got worse.
The repairs product fills all three. Its goal measure is the share of associations renewing their contract. Its driver measure is the share of repairs closed on the first visit, which the team believes moves renewals. Its guardrails are the time an engineer spends inside the app per job, the crash rate on the older handsets many engineers carry and the number of support calls from engineers in the field.
Holding the three roles apart prevents the commonest argument in a metrics review, which is a team defending a measure it cannot move against a director asking about a measure that pays the bills. Both are right about their own measure. Naming one a driver and the other a goal, and stating the belief that connects them, turns the argument into a question anybody can test later.
Five tests a candidate measure has to pass
A candidate measure survives five tests before it goes on a wall, and the worked case above runs through them against three alternatives.
| Candidate | Follows from the strategy | The product can emit it | The team can move it in a quarter | Hard to move without delivering value |
|---|---|---|---|---|
| Share of repairs closed on the first visit | Yes | Yes, the engineer app records an outcome per job | Yes | Yes, a fault reopened inside thirty days takes the job out of the numerator |
| Repair visits scheduled | No, it counts work done | Yes | Yes | No, booking more visits raises it |
| Parts lookups performed in the app | No | Yes | Yes | No, one extra prompt raises it |
| Contract value renewed | Yes, eventually | Yes, from the billing records | No, contracts renew once a year | Yes |
The five tests themselves are worth stating in full, because a candidate that fails one of them usually fails it quietly.
- It follows from a claim in the strategy. Somebody can trace a straight line from the measure back to a sentence the organisation has committed to.
- The product can emit it. The events the arithmetic needs already exist, or a team can add them in a few weeks and knows what that costs.
- The team can move it. A release the team ships changes the figure inside a period the team is willing to wait for.
- It resists gaming. No cheap change raises the number without a customer being better off, which is the test the two middle rows of the table fail.
- Somebody owns it. One named person reports on it, defends the definition and says when it has changed enough to act on.
The fifth test looks like administration and prevents more waste than the other four together. A measure owned by everybody is checked by nobody, and its definition drifts as each analyst rebuilds the query from memory.
A measure nobody on the team can move
A measure a team cannot move gives it no information during the period the team is working.
Contract value renewed is the clearest case in the repairs product. Contracts run for a year and almost all of them renew in April. A change shipped in October reaches the renewal figure six months later, by which time the team has shipped forty other changes and a competitor has cut its price. The feedback loop is longer than the planning cycle, so the number arrives after every decision it could have informed.
The same failure appears in a second form, which is a measure moved mostly by somebody else. A team of six working on onboarding cannot move total revenue, because sales coverage, pricing and enterprise renewals move it far harder than onboarding does. Handing that team revenue as its measure guarantees two outcomes. Good work will be invisible inside the noise, and a quarter when sales closes three large accounts will read as an onboarding success.
A measure that moves on its own
At the other end sits a measure that moves without anybody touching the product, and it is the harder of the two to spot because it feels responsive.
Repairs completed each month is the example in this product. The figure rose 22 per cent in January, and the reason was a cold snap that burst pipes across the north of England. The team had shipped an onboarding change that month and reported the rise beside it in the monthly review. The weather did all the work. Nothing in the count could separate one cause from the other.
Turning the count into a ratio removes the problem. The share of repairs closed on the first visit is unaffected by how many repairs there were, so a cold January raises the denominator and the numerator together and leaves the ratio where it was. That is the general repair for a measure that drifts on its own, and it is why the useful product measures are mostly ratios and rates.
A chosen measure ends up as a short written definition, which holds the name, the exact arithmetic, the denominator, the owner, the value it stands at today and the movement that would count as worth acting on. That page is a decision and nothing more. The share of repairs closed on the first visit does not exist as a number until the engineer app records an outcome for every job, with a property naming the fault and another saying whether a second visit followed. Making a product emit what a team decided to measure is where the next module starts.