Concept 2 of 5

Goodhart's law and metric gaming

3 questions test this

Goodhart's law states that a measure used to control behaviour stops being a good measure of the thing it was originally chosen to represent.

A decision rule that ships a change only when a named measure clears a threshold gives that measure real authority inside an organisation. Authority is the condition under which a measure starts behaving differently, because people now have a reason to move it by whichever route is cheapest, and the cheapest route is rarely the one the measure was meant to detect.

The observation came from monetary policy. Charles Goodhart, then an economist at the Bank of England, made it in a 1975 paper on the problems of monetary management in the United Kingdom, written for a Reserve Bank of Australia conference. His wording was about statistics and control.

Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.

Goodhart was describing what happened when the authorities began targeting particular measures of the money supply. The relationships those measures had held with inflation for years came apart as soon as the measures became instruments of policy. Marilyn Strathern restated the idea for institutions in a 1997 paper on university audit, in the form most people now quote, which is that a measure ceases to be a good measure once it becomes a target.

The sections below show the three shapes gaming takes in product work, explain why most of it happens without intent, and then set out the defences. The page closes on the part of the problem that measurement cannot reach.

The three shapes metric gaming takes in product work

Moving the number without moving the behaviour. A support organisation judged on first response time introduces an automated acknowledgement that fires within two minutes of a ticket arriving. The median first response falls from 47 minutes to 2. Time to resolution rises from 9 hours to 14, because the acknowledgement satisfied the clock and nothing else changed. Both numbers are correct and one of them now measures a robot.

Changing what counts as the event. A team judged on activation defines activation as connecting a data source. Somebody proposes connecting a sample data set automatically on first login, which is genuinely useful for demonstrating the product. Activation moves from 34 per cent to 91 per cent in one release. Week four retention does not move at all, because the definition changed and the customers did not.

Taking the gain from somewhere the measure cannot see. A growth team judged on signups adds a prompt to a flow people were already in, and signups rise by 22 per cent. Paid conversion falls, because the additional signups came from people who were not ready and now arrive at a paywall they resent. The measure went up and the business went down, which is the version of the problem that costs the most and gets noticed the latest.

Why most gaming is done in good faith

Donald Campbell described the same effect in social policy in a 1976 paper, arguing that the more a quantitative indicator is used to make decisions, the more it comes under pressure to distort and the more it corrupts the process it was meant to monitor. His framing is useful because it locates the fault in the system of measurement and never in the people inside it.

Each of the three cases above was built by somebody solving the problem in front of them. An engineer asked to reduce first response time built the acknowledgement and reduced it. A designer adding a sample data set made an empty product easier to understand. Nobody set out to deceive, and a review that treats gaming as a conduct question will find no misconduct and change nothing.

What all three share is a measure sitting some distance from the outcome it stands for. First response time stands for a customer being helped, activation stands for a customer reaching value, and signups stand for revenue. Each of these moves lives in the gap between the proxy and the outcome, and the width of that gap decides how gameable a measure is.

The defences that hold a measure honest

Four practices reduce the gap or make it visible, and a programme usually needs all four.

  1. A measure close to the customer outcome. A measure of what a customer achieved is harder to move by a shortcut than a measure of what the product did. Time to resolution sits closer than first response time, and second week usage sits closer than a connected data source.
  2. A quality measure paired with every target. Response time pairs with resolution rate. Signups pair with paid conversion. Activation pairs with retention at four weeks. The pair makes the cheap route visible in the same report, so the shortcut and its cost arrive together.
  3. Guardrails owned outside the team being judged. A guardrail the team controls is a guardrail that gets redefined. Latency, error rate, unsubscribes and support contacts belong to somebody with no stake in the result.
  4. A versioned definition with a review date. A measure that jumps by 57 points in one release has almost always had its definition changed. A dated record of how each measure is calculated turns that jump from a mystery into a one line answer.

Pairing a target with a quality measure is the defence product teams adopt most easily, because it needs no new tooling. John Doerr describes paired key results in Measure What Matters, published in 2018, on exactly this reasoning, which is that a quantity target with no quality target beside it will be met in whichever way is fastest.

What measurement cannot fix

The defences above make gaming visible and they do not remove the reason for it. A measure that decides promotions, bonuses and headcount will be moved by the cheapest available means, and adding a second measure only means two numbers now need moving.

The part that has to change is what the organisation rewards. A team rewarded for shipping features ships features. A team rewarded for moving a number moves the number. A team rewarded for a customer outcome, and asked to show evidence that its work caused the outcome, is being asked for something much harder to produce by a shortcut, since that evidence has to survive the same reading every experiment result gets.

Whether that is reasonable to ask of a team is a judgement about the organisation, and the honest answer is that most organisations reward what they can see. That brings the problem back to the analyst, because a finding nobody outside the team understands cannot change what anybody is rewarded for.

Common misconceptions

Goodhart's law is an argument against setting numeric targets at all.

It is an argument for choosing which measure carries a target and for watching what moves alongside it. A measure sitting close to the customer outcome, paired with a quality measure and surrounded by guardrails, survives being targeted. A proxy chosen because it was easy to collect does not.

Metric gaming means somebody is behaving dishonestly.

Most gaming is a reasonable response to what an organisation rewards, carried out by people who believe they are doing their job. The automated acknowledgement that halves a response time was built by somebody solving the problem they were given. Treating the behaviour as a character fault hides the design fault that produced it.

3 questions test this concept

A support organisation at a pension provider is judged on first response time. An engineer builds an automated acknowledgement that fires within two minutes of a ticket arriving. The median first response falls from 47 minutes to 2 and time to resolution rises from 9 hours to 14. A director opens a conduct review. What has the review got wrong?

  • AThe figures, since a median first response of 2 minutes is not achievable by any support team and the measurement must itself be at fault.
  • BWhere the fault sits. The engineer solved the problem they were given, so the repair is a measure closer to the customer outcome with a quality measure paired to it, since time to resolution sits closer than first response time does.
  • CThe premise, since Goodhart's law is an argument against setting numeric targets at all and the target is what should be withdrawn.
  • DThe ownership, since first response time is held by the team being judged on it and a guardrail in that position always gets redefined.
Check whether it stuck.

One per page, with a worked explanation.

Start the set
Related material
Book
Measure What Matters, On pairing a target with a second measure that shows what it costs.
Book
Outcomes Over Output, On measuring a change in customer behaviour, which is harder to move by any route other than the intended one.