Product metric frameworks

A product metric framework is a published set of measurement categories a team fills in, offered so that nobody has to derive a complete set of measures from first principles.

Two of them turn up in almost every product conversation. AARRR came from the investor Dave McClure, who presented it in 2007 under the title Startup Metrics for Pirates, and it describes the journey a customer makes from arriving to paying. HEART came from Google's research team. Kerry Rodden, Hilary Hutchinson and Xin Fu published it at the CHI conference in 2010, in a paper called Measuring the User Experience on a Large Scale. It describes the quality of the experience somebody has once they are inside the product.

Both are useful and both are misread the same way. A framework is a checklist of categories, so it tells a team which questions to ask and never which number answers one. The work of turning acquisition into a defined event, or task success into a completion rate with a stated denominator, belongs to the team and takes longer than choosing the framework did.

The sections below take AARRR and then HEART, with what each one was built to solve. They compare the two against the north star in a table. Finally they set out the criticism each deserves and the way most teams end up using more than one.

AARRR and the five stages it names

AARRR splits the life of a customer into five stages and asks a team to measure each one separately.

  1. Acquisition. Somebody arrives at the product from an advertisement, a search result, a link from a colleague or a sales conversation.
  2. Activation. Somebody has a first experience good enough to come back from, which a team has to define as a specific action.
  3. Retention. Somebody returns and keeps returning.
  4. Referral. Somebody brings another person to the product.
  5. Revenue. Somebody pays, or somebody in their organisation does.

The initials spell the noise a pirate makes, which is where the nickname pirate metrics comes from. McClure built it for an early stage startup where the whole company is one funnel and the only question worth asking is which stage leaks hardest. A founder with five numbers on a page can see that 40,000 people arrive, 3,200 finish setup, 900 return in week two, 40 invite anybody and 110 pay, and can work out which of those five figures deserves the next month.

That example also shows the framework doing its best work, which is making a comparison possible between stages that were previously discussed separately. Marketing owned the 40,000. Product owned the 3,200. Setting them in one column is what turns two departments into one funnel.

HEART and the Goals Signals Metrics process

HEART answers a different question, which is whether the experience of using a feature is any good.

Rodden, Hutchinson and Fu wrote the paper because the two things Google's research teams could do at the time did not meet in the middle. Small scale methods such as interviews and observed sessions produced rich evidence about the experience of a handful of people. Large scale behavioural data covered millions of people and counted usage without saying whether anybody had a good time. HEART was the attempt to build large scale measures of the experience itself.

The five categories are happiness, engagement, adoption, retention and task success. Happiness covers how people feel about the product and usually comes from a survey. Engagement covers how much somebody uses it inside a period. Adoption covers people starting to use a feature for the first time. Retention covers the existing population still there later. Task success covers whether somebody completed what they came to do, which is measured through completion rates, error rates and the time a task took.

The part of the paper that transfers best is not the five categories at all. Rodden, Hutchinson and Fu paired HEART with a three step process called Goals Signals Metrics, which runs in that order. State the goal for the feature, which is what would count as success for the person using it. Name the signal, which is the behaviour that would appear if the goal were being met or missed. Then choose the metric, which is the arithmetic that counts the signal. A team that runs those three steps derives its own measures and needs no framework at all.

What each framework was built for

FrameworkWhat it optimises forThe stage it suitsWhat it leaves out
AARRRGrowth through a funnel from first arrival to payment and referralAn early product still proving that people arrive, return and payThe quality of the experience, and most of what happens after the second visit
HEARTThe experience of somebody already using a featureAn established product with research capacityAcquisition, revenue and the cost of serving anybody
North starOne measure of delivered value with the inputs that move itA team that knows who it serves and needs one directionCost, risk and everything the single measure excludes

The row worth reading twice is the last column, because a framework causes damage through what it leaves out. A team measuring only AARRR ships a product that acquires and converts well and is unpleasant to use. A team measuring only HEART polishes an experience nobody is arriving at.

The honest criticism of each framework

AARRR carries four weaknesses, and three of them follow from the funnel shape it assumes.

The first weakness is that AARRR names categories and never a measure. Activation means nothing until somebody decides which action counts as activated, and that decision is the whole of the work. Two teams using AARRR on the same product can report activation figures that differ by a factor of four.

The second weakness is that a one way sequence describes a product passed through once. A product people use for four years spends nearly all of its life inside retention, and a five box model gives that four year period one box. Retention then gets a third of the attention it deserves because it has a fifth of the diagram.

The third weakness is that the order is wrong for a business selling to companies. Revenue sits at the end of the sequence, and in enterprise software a contract is signed before anybody logs in for the first time. The stages then run backwards, with revenue first and activation somewhere in month three, which makes the diagram actively misleading for the teams reading it.

The fourth weakness is that nothing in AARRR asks whether the product is any good. All five stages can improve while the experience gets worse, because the model counts arrivals, returns and payments and never asks whether anybody finished what they came for.

HEART carries three weaknesses of its own, and they come from the setting it was built in.

The first weakness is that happiness needs a survey, which needs a sample, a running instrument and somebody to read the free text. A team of six shipping every week usually has none of that, so happiness goes uncounted and the first letter is decoration. A team in that position should say so out loud and leave the letter off its own template.

The second weakness is that engagement and retention overlap for most products. A weekly active count and a week four retention figure rise and fall together, so a team reports two numbers that carry one piece of information and feels better covered than it is.

The third weakness is that adoption and retention both need a defined population to count against, and defining that population is a genuinely hard problem for a product where most visitors are anonymous. HEART assumes that question is settled, and the answer arrives later in this course.

Using more than one framework at a time

Most teams end up holding two of these at once, and the combination is more useful than any single one.

A north star sits at the top and settles the direction. AARRR supplies the vocabulary for the path a new customer takes, which is what a growth team needs. HEART supplies the questions for a feature already in use, which is what a team improving search needs. And Goals Signals Metrics is the derivation route underneath all three, because it works whether or not a team has chosen a framework at all.

Naming a category is the cheap half. Choosing the measure that goes inside it decides whether a quarter of work pointed at anything, and that choice starts from the strategy rather than from the framework. The next page works through how a team makes it.

Common misconceptions

AARRR and HEART are competing frameworks, so a team picks one of them.

The two answer different questions and most teams end up holding both. AARRR describes the journey a new customer makes from arriving to paying. HEART describes the experience of somebody already using a feature. A team running a signup redesign and a search improvement in the same quarter needs the first for one and the second for the other.

Where this is examined
Product Metrics and Analytics
Choosing What to Measure, 15 per cent of the exam.
Related material
Book
Lean Analytics, On matching the measure a team watches to the stage it has reached.
Book
Outcomes Over Output, On why a category of measure means nothing until somebody names the behaviour underneath it.
Concepts