Cohort analysis

Cohort analysis is the practice of grouping people by something they share, usually the period in which they arrived, and then following each group forward through the same number of days or weeks of its own life.

A funnel measured once gives a single number over a single window, so it cannot say whether a drop is new. Cohort analysis is the repair, because it replaces one number with a series of groups that are each measured from their own starting line.

The method came from demography long before it reached software. Norman Ryder set out the cohort as a tool for studying social change in the American Sociological Review in 1965, arguing that a population is best understood as a succession of groups born into different circumstances and carrying those circumstances forward. Product teams inherited the idea with the units swapped. A birth cohort becomes a signup week, and a lifetime becomes the ninety days after a person first used the product.

The reason a product manager needs this is that a blended number moves for reasons that have nothing to do with the product. A marketing campaign, a seasonal peak, a price change or a single large customer rolling out to eight hundred staff will all move a weekly retention figure while the software sits unchanged. A team without cohorts spends the following Monday explaining a number that was never about their work.

The sections below separate the two kinds of cohort a product team builds. One cohort table is then read in the three directions that carry meaning, and a worked case shows a blended figure falling while every group inside it holds steady. Finally they set out what the method leaves unanswered.

Acquisition cohorts and behavioural cohorts

An acquisition cohort groups people by when they arrived. If somebody signed up in the week beginning 19 January, they belong to that group for as long as the product measures them. Membership is settled once, at the start, and nothing a person does afterwards moves them into a different group. This is what makes the comparison across weeks fair, since every group is measured from day zero of its own life.

A behavioural cohort groups people by something they did. One group connected a data source in their first week and another group did not. Membership here depends on behaviour, so the two groups differ in more than the one action used to define them. When a person connects a data source early, they usually had data worth connecting in the first place. That is a fact about the person before it is a fact about the product.

Both are useful and they answer different questions. An acquisition cohort answers whether the product is treating each new intake better than the last. A behavioural cohort answers what separates the people who stayed from the people who left, which is the question the next stage of this course opens on.

One cohort table read in three directions

The table below follows four weekly signup cohorts at a product for small finance teams, read on 2 February. Each percentage is the share of that cohort still active in the given week of its own life, so week one means the seven days after the signup week and never a calendar week.

Week joinedPeopleWeek 1Week 2Week 3Week 4
5 January1,20038 per cent27 per cent23 per cent21 per cent
12 January1,40040 per cent29 per cent25 per cent
19 January2,60026 per cent17 per cent
26 January1,50041 per cent

Down a column the table compares groups at the same age. Week one retention runs 38, 40, 26 and 41 per cent, and the 19 January figure is the one that stands out. Something about that intake differed from the three around it.

Across a row the table follows one group through its own life. The 5 January cohort fell from 38 to 27 per cent in a week, then to 23 and then to 21, so most of the leaving happened immediately and the rate of loss slowed after that. Reading the same row a month later is how a team learns whether the fall ever stops.

Along a diagonal the table picks out one calendar moment. Three cells fall inside the same seven days, which are the 5 January cohort's week three, the 12 January cohort's week two and the 19 January cohort's week one. An outage or a broken release in that week therefore appears as a dip running diagonally across otherwise healthy rows, and a blended chart hides the pattern completely.

A blended figure falling while every cohort holds

The 19 January cohort came from a paid campaign on a deals site. That is why it is nearly twice the size of the weeks around it, and why its week one figure sits at 26 per cent while its neighbours sit in the low forties. Nothing in the product changed that week, and the arithmetic below shows what one intake of that size does to the headline figure.

Across all four cohorts the number of people still active after one week is 456 from the 5 January group, 560 from 12 January, 676 from 19 January and 615 from 26 January. Those add to 2,307 people out of the 6,700 who signed up, which is a blended week one retention of 34.4 per cent. Take the campaign cohort out and the same arithmetic gives 1,631 of 4,100, which is 39.8 per cent.

A team reporting only the blended figure has a five point fall to explain and no way to explain it. A team reporting the cohort table has a campaign that bought unqualified signups, a decision to make about whether to run it again and three untouched cohorts confirming the product is working as it did in December. The underlying numbers are identical. Only the grouping differs.

What cohort analysis leaves unanswered

Cohorts separate a change in the product from a change in who arrived, and they say nothing about why either happened. The 19 January group retained badly, and whether those people wanted something different, landed on a mismatched page or simply came for a discount is outside the table.

The method also needs patience that quarterly reporting rarely allows. A cohort defined this week has one usable column, and its second column arrives next week whatever anybody decides in the meantime. A team that wants a thirteen week answer waits thirteen weeks, and the strongest pressure on any analytics practice is the pressure to answer sooner than the data can.

There is a further limit worth stating plainly. A small cohort produces a percentage that swings on a handful of people, so a group of eighty signups moving from 25 to 30 per cent is four people changing their minds. Weekly cohorts suit a product with thousands of signups a week, and monthly cohorts suit everything smaller.

None of those limits removes what the table shows most clearly, which is the shape of a row. Each cohort loses a large share immediately and then loses less each period, and the interesting question is whether that slowing ever reaches a floor. Plotted as a line, that shape is the single most diagnostic picture in product analytics.

Common misconceptions

A cohort table shows whether the product is getting better over time.

A cohort table shows whether the experience each arriving group meets is getting better, which is a narrower claim. A column read downwards compares groups acquired at different moments through different channels, so a change in that column can come from the product or from the mix of people arriving. Separating the two needs the acquisition source on the table as well.

Where this is examined
Product Metrics and Analytics
Reading Product Behaviour, 22 per cent of the exam.
Related material
Book
Lean Analytics, On cohorts as the way to tell a product change from a market change.
Book
Hacking Growth, On following each intake of new users forward as its own group.
Concepts