The Product Owner is accountable for maximising the value of the product resulting from the work of the Scrum Team. The Guide states the accountability and then stops, saying nothing about what value is or how anyone should measure it.
That silence is deliberate rather than an oversight. Value depends on the product, the market and who the product is for, so a framework that named a measure would be wrong for most of the products using it. What Scrum does supply is where the evidence comes from, which is observation of a working Increment by the people it is meant to serve.
Output and outcome are different questions
A Sprint that completes every selected item has produced output. Output is countable and it is inside the team's control, which is exactly why it is what tends to get reported. Whether any of it produced value is a separate question, answered by what changed for the people using the product and for the organisation building it.
An outcome is that change in behaviour or result. Ten features shipped is output, and support calls falling by a third because one of them worked is an outcome. The two can move in opposite directions, and a team measured on output has every reason not to notice when they do.
An Increment nobody can react to yields no evidence
The Increment matters here because it has to be usable. A step towards the Product Goal that nobody can try tells you nothing about whether it was worth building, and value remains an assertion until somebody reacts to something real.
This is why the Definition of Done concerns a Product Owner and not only the Developers. Work that does not meet it cannot be released, so it cannot be inspected by the people whose reaction would settle the question, and the Sprint has produced no evidence at all. Value held back for a distant release is also value unmeasured for that whole period, while the plan carries on resting on assumptions nobody has tested.
Where the evidence is gathered
The Sprint Review is the scheduled point at which the Scrum Team and stakeholders inspect the Increment and discuss what has changed in the environment. The Product Backlog is adapted from what that discussion reveals, which is what makes the event part of measurement rather than a report on the Sprint.
Scheduled does not mean exclusive. Usage data, support volume, sales conversations and anything else the product generates all count as observation, and a Product Owner who waits for the Sprint Review to learn how the last release landed has given up most of the feedback available. What the event adds is that the evidence gets acted on in front of the people affected by it.
A measure adopted as a target stops measuring
Any measure that becomes a target gets optimised directly, and the shortest route to moving a number is rarely the thing the number was standing in for. Time on page rises when the interface becomes confusing. Support tickets fall when the support form gets harder to find, and both look like progress on a dashboard.
The defence is to keep asking what the measure stood for and to treat a sudden improvement as a claim to be checked. Several measures read together are harder to game than one, and a measure attached to a stated expectation is harder still, because it arrives with a prediction that can turn out to be wrong.
The two things that actually help
Faced with an accountability for value and no method, a Product Owner tends to reach for tooling. Two moves genuinely help, and neither of them is a tool.
The first is validating assumptions of value through frequent releases. Value is a claim until somebody uses the thing, and the only way to convert the claim into evidence is to put a usable Increment in front of the people it was built for and watch what they do with it. Releasing more often does more than deliver value sooner. It shortens the distance between a decision and the evidence about that decision, which is what makes the next decision better than a guess.
The second is the order of the Product Backlog. Ordering is where a Product Owner's judgement about value is actually expressed, because it is the only place that judgement has consequences. An item everybody agreed was valuable and left in fortieth position has been valued, whatever was said about it in the room. Putting the judgement in the order also makes it arguable, since a visible sequence can be challenged by a stakeholder or a Developer in a way that a private conviction cannot.
The popular alternatives disappoint for related reasons.
A scoring formula presented as objective is nothing of the kind. Every weighting in it is a judgement somebody made, and running the inputs through arithmetic moves that judgement out of the conversation and into the coefficients, where nobody argues with it because it now looks like a result. A scoring model used to structure a discussion earns its place. The same model used to settle one has hidden the decision rather than made it.
Setting a value figure on individual items through a group estimation game produces precision without evidence. A room of people who have not spoken to a customer this month will still converge on a number, because converging is what the game is designed to do, and the number is then carried around as fact for the rest of its life on the strength of having a decimal point in it.
Neither practice is forbidden, and neither is a substitute for shipping something and looking at what happened.
Scrum.org's Evidence-Based Management
Beyond the Guide, Scrum.org publishes Evidence-Based Management, which groups measures into four Key Value Areas. These are Current Value, Unrealised Value, Ability to Innovate and Time-to-Market. It is not part of the Scrum Guide, but it is part of Scrum.org's published material on measuring value and is worth recognising by name, along with the fact that two of its four areas describe the organisation building the product rather than the product itself. The detail belongs to its own page.