Adaptive estimating sizes work by comparison rather than by duration. An item is judged against other items already sized, the team assigns it a number on a relative scale, and how long that number takes is discovered by observation rather than declared in advance.
The reason for the indirection is that people are unreliable at absolute duration and quite good at comparison. Asked how long a task takes, an estimator produces a number shaped by optimism and by what they think the answer should be. Asked whether this task is bigger than that one, the same person is consistent. Relative sizing uses the judgement that works and measures the part that does not.
What a size is made of
A relative size folds together volume of work, complexity and uncertainty. Two items can share a size for different reasons, one being a large amount of straightforward work and the other a small amount of work nobody has done before.
Scales are usually nonlinear, so the gaps widen as the numbers grow. That is deliberate. The difference between an item sized two and one sized three is meaningful, and the difference between forty and forty five is invented precision on work nobody understands yet. A widening scale forces the estimator to say roughly how big rather than exactly how big, which is all the information that exists at that point.
Sizing belongs to the people who will do the work. An estimate produced by somebody else is a target wearing the vocabulary of an estimate, and it carries none of the knowledge that made the technique worth using.
Techniques worth recognising
| Technique | How it works | What it is good for |
|---|---|---|
| Planning poker | Everyone estimates privately, all reveal at once, differences are discussed before re estimating | Surfacing hidden disagreement, since two people with the same number for different reasons is the risk it catches |
| Affinity grouping | Items are sorted into piles of similar size, quickly, then the piles are given values | A large backlog being sized for the first time, where relative order matters more than precision |
| Shirt sizes | Small, medium, large, extra large, converted to numbers later or not at all | Early portfolio level work, and audiences who read numbers as commitments |
| Comparison to a reference | Each item is compared to one agreed item of known size | Keeping a scale stable over months, since drift is the main way a scale stops meaning anything |
| Counting items | No sizing at all, forecast from the number of items finished per cycle | Teams whose items are broken down to a similar size anyway, where sizing adds ceremony without adding accuracy |
Counting items deserves more attention than it usually gets. Where a team already splits work until each piece is small, the number of items finished per cycle forecasts about as well as the sum of their sizes, and the sizing session was overhead. The exam does not require this position, but it does test whether you understand that the framework asks for a size and not for a particular way of arriving at one.
Velocity, capacity and the difference
Velocity is the amount of work a team actually completed in a cycle, measured in that team's own units. It is observed after the fact.
Capacity is how much the team expects to be able to take on in the next cycle, adjusted for what is known about it. A cycle containing two public holidays and a training day has less capacity than the last one, and the adjustment is arithmetic rather than judgement.
| Term | Definition | What it answers |
|---|---|---|
| Velocity | Sizes of items completed in a cycle | How fast has this team actually been |
| Average velocity | Mean over several recent cycles | What is a reasonable central expectation |
| Velocity range | Lowest and highest of recent cycles | How much does this team vary, which is what a forecast needs |
| Capacity | Available effort for the coming cycle | How much should be planned this time |
A single cycle's velocity is nearly useless on its own. Three or more gives a range, and the range is what makes a forecast honest, because the spread is the part that determines whether a date is safe.
Forecasting a date
Take the backlog remaining, in sizes. Divide by velocity. That gives cycles remaining, and multiplying by cycle length gives a date.
Doing it once with the average produces a single date and a false impression. Doing it three times, with the lowest recent velocity, the average and the highest, produces a range that can be reported with the confidence it deserves.
| Input | Calculation | What the answer says |
|---|---|---|
| Remaining size divided by pessimistic velocity | 240 divided by 20, giving 12 cycles | The date this is unlikely to slip past |
| Remaining size divided by average velocity | 240 divided by 30, giving 8 cycles | The date to plan around |
| Remaining size divided by optimistic velocity | 240 divided by 40, giving 6 cycles | The date this could reach if nothing goes wrong, which is not a date to promise |
Two assumptions sit under all three and both need stating when the forecast is reported. The backlog is assumed not to grow, which it will, and the team is assumed to stay the same, which it may not. A forecast that names its assumptions can be corrected when they break. One that arrives as a date gets treated as a commitment and then defended.
Why the number must not leave the team
Velocity is denominated in units the team invented. It has no meaning outside that team, and it responds instantly to being used as a target.
A team measured on velocity will produce more velocity, by sizing more generously, by splitting items to inflate the count, or by finishing work to a lower standard. None of those deliver anything more, and all of them destroy the forecasting value of the number, which was the only reason to collect it. This is the most reliably examined point on the topic, because it is the one practitioners get wrong in real organisations.