What analytics cannot answer

Behavioural data records what somebody did inside a product and holds no record of why they did it.

Goodhart's law described what happens to a measure once an organisation rewards it, and the four defences against gaming all work on that problem. A measure can pass every one of those four defences and still answer only the question behavioural data is able to answer. Sitting close to the customer outcome, carrying a quality measure, belonging to somebody outside the team being judged and holding a dated definition all make a measure harder to move by a shortcut. None of the four makes it able to report a reason.

The limit has nothing to do with pressure, gaming or bad measurement. It is a property of the evidence. An event is an actor, an action, a time and some properties, and none of those four fields is a reason. A product observes fingers on glass and records the consequences.

A product manager meets the limit at the moment a chart produces a meeting instead of a decision. Everybody agrees the number fell. Nobody can say why, five explanations are offered in twenty minutes, and each of the five implies a different quarter of work. The chart settled where and it settled nothing else.

The sections below set out what an event holds and what it leaves out. Session replay, surveys and interviews then take a section each, covering what every one of them recovers and how each misleads, followed by a table putting the four methods beside each other. The page closes on the population no product measure contains and on the decisions no number settles.

What behavioural data records and what it leaves out

An event carries four things. Somebody did something, at a moment, with some properties attached. Every measure in this course is built from those four fields, which is why the measures are reliable about counts and silent about motives.

Three questions arrive constantly and none of them is answerable from events. Why somebody left. What they expected before they arrived. What they would have done if the product had behaved differently. The first asks for a reason, the second asks about a state of mind that preceded any event, and the third asks about a world that did not happen.

One checkout shows how little a precise number settles. A retailer finds that 38 per cent of the people reaching the address step never reach payment, and the figure is stable across four weeks and three devices. Five explanations fit it exactly.

  1. The form asks for a company name, so private customers assume the product is sold to businesses.
  2. The delivery estimate appears for the first time at that step and says eleven days.
  3. The address lookup rejects valid postcodes in one country, so people from there cannot proceed at all.
  4. The step takes nine seconds to load on a mid range phone.
  5. A third of the people who reach it are comparing prices and never intended to buy anything today.

Each of those five produces a drop of the same size at the same step, and each one implies different work. Copy, logistics, a validation bug, performance and acquisition mix are five different teams. No amount of extra behavioural data separates them, because the chart already contains everything the events can say, and running it for another year returns a more precise version of the same silence.

What does separate them is another kind of evidence, and three kinds are available. Each answers part of the question and each carries a failure mode that puts a team somewhere worse than where it started.

Session replay and the visit it reconstructs

Session replay records the events a browser emits and plays them back as a reconstruction of what the screen showed, so an analyst can watch one person's visit unfold in order.

Four things live inside a single visit that no summary event captures. Repeated clicking on something that is not a control. A field somebody completed, cleared and completed again three times. An error message that rendered below the visible part of the page. The order in which somebody actually read a page, which is rarely the order it was designed in. All four are visible in a recording and absent from every chart.

Session replay misleads in four ways and the first is the one that costs most.

  1. The analyst chooses the recordings. A filter built from the hypothesis somebody already holds returns sessions that support it, so the sample is selected by the conclusion.
  2. A recording carries no rate. Ten sessions feel like proof and are ten people out of forty thousand. Nothing in a replay says how common anything in it is.
  3. It records the screen and never the person. A three minute pause is a phone call, a colleague at the desk, a kettle or genuine confusion, and the recording gives no way to tell those apart.
  4. It captures whatever was on the page. Form contents, customer records and message text all land in the recording, which makes replay the heaviest personal data exposure of any tool in this course and puts it squarely inside the consent rules from module two.

The useful position is to treat replay as a generator of candidates. Watching twenty sessions around a drop produces two or three plausible explanations in an afternoon, faster than any other method. Establishing that one of them happens often enough to matter means going back to the events with a new question and a new measure.

Surveys and what people report about their own reasons

A survey reaches people a recording cannot, since it can be sent to somebody who cancelled last month, who abandoned a trial in March or who never registered at all. It also returns a number from many people, comparable against the same question asked a quarter later, which is the property no qualitative method has.

Three failures come with that reach, and the first two are about who answers.

Non response. The people who complete a survey differ from the people who ignore it, and the difference runs in a predictable direction. Enthusiasts answer, the recently annoyed answer, and the large middle group who felt nothing in particular never replies. A satisfaction figure computed over respondents describes respondents, and the response rate belongs beside the figure every time it is published.

Wording. The question decides the answer more than most teams expect. A cancellation question offering six reasons to choose from returns one of those six, whatever the actual reason was, because somebody wanting to finish the form picks the closest option available.

The why question in particular. Richard Nisbett and Timothy Wilson published Telling More Than We Can Know in Psychological Review in May 1977, reviewing evidence that people have little direct access to the processes behind their own judgements. Their argument was that a person asked to explain a choice produces an account built from a plausible theory of what would usually cause such a choice, and that the account can be confidently wrong about what actually happened in their head.

That finding changes which question is worth asking. Asking somebody what they did, when and with what, asks for a memory of behaviour, which people report reasonably well. Asking why asks for something they cannot observe in themselves. An exit survey is therefore at its most useful when it asks what the person was trying to achieve and what they used instead, with a free text box underneath, since the box is where a reason nobody listed can appear.

Interviews and the question nobody thought to ask

An interview is the only method here that can turn up a question nobody had. Events answer a question somebody defined in advance and instrumented. Survey options were written in advance. A conversation can go somewhere nobody planned, which is the entire reason to spend an hour on it.

Interviews mislead in three ways. Numbers are small, so nothing in them supports a rate. The interviewer's framing shapes the answer, and the more invested the interviewer is in a direction the more the answers follow it. And people report intentions they will not act on, which is why a question about whether somebody would pay for a capability returns a far more encouraging answer than a price ever does.

The question of how many conversations are enough has a published answer that is quoted more often than it is read. Jakob Nielsen and Thomas Landauer modelled the discovery of usability problems as a Poisson process at the INTERACT and CHI conference in 1993. That model is the source of the familiar claim that five participants find about 85 per cent of the problems in an interface. The arithmetic behind the claim assumes each problem has roughly a 31 per cent chance of appearing with any one participant. Five sessions therefore find a fault that a third of people hit. A fault catching one person in twenty needs far more sessions, and quoting the figure for rare problems misreads the model it came from.

What each method is good for

Four methods now sit on the table and the choice between them is a choice about which question is open.

MethodThe question it answersHow it misleadsWhat it costs
Behavioural eventsWhat happened, to how many people, in what orderSilent about motive, and blind to anybody it never recordedInstrumentation, a tracking plan and the discipline to keep both
Session replayWhat one visit looked like from the screenThe analyst picks which sessions to watch, and ten of them carry no rateA licence, and the heaviest privacy exposure of the four
SurveyHow many people report something, comparably over timeOnly respondents answer, and the wording supplies the replyA response rate, and the goodwill spent asking
InterviewWhat somebody was trying to do, including things nobody asked aboutSmall numbers, framing by the interviewer, stated intentionsTwo people's time for an hour per conversation

The four methods pair up in one direction and the order matters. Behavioural data locates where something happens and how often it happens. Replay and interviews then propose candidate reasons for it. A survey estimates how widespread a proposed reason is across the whole population. A controlled experiment settles whether acting on that reason changes anything, which is the only step of the four that establishes cause.

A team running those in the other order pays for a month of research into a problem a chart would have located in an hour, or ships a change based on three conversations and finds out a quarter later that the three were unusual. The order is the whole of the method.

The people who never arrived

Every method above shares one blind spot and it is the largest thing missing from a product's evidence. Behavioural data covers people who reached the product, got far enough to be instrumented and were recorded. Replay and surveys inside the product reach the same population. Interviews reach whoever agreed to talk.

The reasoning that names this comes from aircraft. In 1943 Abraham Wald worked at the Statistical Research Group at Columbia University on where to add armour to bombers, and the damage charts in front of him were drawn from aircraft that came back. Wald pointed out that the areas with the fewest recorded hits were the areas where a hit stopped an aircraft returning, so the armour belonged exactly where the data looked clean. His memoranda on estimating vulnerability from the damage of survivors reversed the recommendation the charts appeared to support.

The product version of that missing population is large and entirely invisible.

  1. People who never heard of the product.
  2. People who read a pricing page and left before any event fired.
  3. People whose browser blocked the analytics script.
  4. People who declined the consent banner.
  5. Buyers a sales team never reached, in segments nobody targeted.
  6. Companies that evaluated the product carefully and bought a competitor.

None of those six is a data quality fault. Nothing broke, no event failed and no monitoring would catch any of it, because the shape of the sample is the shape of the business. A product that decides what to build from what its current customers do will improve steadily for the people it already has and will keep ruling out everybody else, one reasonable decision at a time.

Reaching that population means going outside the product entirely. Lost deal reasons recorded in the sales system name the competitors and the missing capabilities. A single question at cancellation catches people on their way out. Panel research and market surveys reach people who never came. Talking to companies that chose somebody else is the most uncomfortable of the four and the most informative, because they evaluated the product and can describe what was missing. Each route is weak on its own, and each reaches somebody the events will never contain.

The decisions that no number settles

A population nobody measured is one reason a decision rests on judgement, and several other decisions have the same property whatever the data covers.

  1. Whether to serve a segment at all. Nothing in the behaviour of current customers says anything about a market the product has never entered.
  2. What the product will refuse to do. A limit accepted on safety, legal or ethical grounds costs measurable revenue and returns nothing a chart can show.
  3. The price of something nobody has bought. Willingness to pay for a capability that does not exist has no behavioural evidence anywhere.
  4. The first version of anything. A new product has no users, so every method here has an empty population until after somebody has decided.
  5. A consequence arriving after the measurement window. A change that lifts a quarterly number and costs trust over three years is measured correctly and read wrongly.

Each of those five is a decision somebody has to make with an argument rather than a measurement, and the useful contribution from an analyst is to say so plainly. An answer of the form "the evidence available cannot settle this, and here is what would" is a finding, and it is worth more than a number computed over the wrong population and presented as though it answered the question.

That kind of finding is the hardest one to communicate. A number looks like an answer and needs no defending, while a stated limit needs an argument made in front of people who wanted the number. Writing either of them so that somebody who was absent for the whole analysis can act correctly is the subject of the next page.

Common misconceptions

Enough behavioural data will eventually explain why people behave as they do.

An event records an actor, an action, a time and some properties. A motivation emits nothing, so adding a further year of events produces a more precise description of the same silence. Several different reasons produce identical charts, and no volume of the chart separates them.

An exit survey asking why somebody cancelled explains the churn.

Richard Nisbett and Timothy Wilson argued in Psychological Review in 1977 that people have little direct access to the processes behind their own judgements, so a reason offered afterwards is a plausible account built after the event. Exit answers are a list of candidate explanations, useful for deciding what to investigate and never a measurement of cause.

Where this is examined
Product Metrics and Analytics
Turning Analysis into Decisions, 17 per cent of the exam.
Related material
Book
Continuous Discovery Habits, On asking somebody for the story of what they did instead of asking them to explain themselves.
Book
Lean Customer Development, On reaching the people who considered the product and bought something else.
Concepts