Concept 2 of 4

Human oversight

2 questions test this

Human oversight is the control everybody claims and few implement. It is also the one the EU AI Act names directly for high risk systems, so it is moving from good practice to obligation.

The three arrangements

A human decides, the system advises. The model produces a suggestion and a person makes the call. Safest, most expensive, and the right arrangement where a wrong decision is costly and irreversible.

A human reviews, the system decides. The system acts and a person checks, either every case or a sample. Cheaper, and it only works while the reviewer retains the ability and the time to disagree.

A human can stop it. The system runs unsupervised and somebody is watching the aggregate, able to intervene or switch it off. Appropriate at high volume and low stakes, and it depends on having something worth watching, which is a monitoring problem rather than an oversight one.

The mistake is claiming the first, staffing the second and operating the third.

What makes review real

The reviewer can judge the output. Somebody checking a legal summary needs to be able to read the contract. A reviewer without the expertise is a delay rather than a control.

They have the time. Review capacity is the binding constraint, and it is what quietly turns oversight into rubber stamping. If the volume is a hundred an hour, the review is a glance.

Disagreeing is expected and cheap. If overriding the system is slower, requires justification, or is treated as an error, people stop. What is measured decides this more than what is written in the policy.

They see what the system is unsure about. Confidence, the sources used, or the cases flagged as unusual. A bare answer gives a reviewer nothing to work with.

Designing for it

Route by risk rather than reviewing everything. Send the low confidence, the unusual and the high consequence to a person, and let the rest through. That concentrates a fixed amount of attention where it changes outcomes.

Make the override the easy path in the interface. Show the reasoning, or at least the sources, so review is possible. And measure the override rate, because a rate near zero usually means the reviewer has stopped reading rather than that the system is perfect.

The obligation

The EU AI Act requires that high risk systems be designed so a person can oversee them effectively, which explicitly includes understanding the system's limits, staying alert to automation bias, being able to disregard the output, and being able to stop it. That is a design requirement rather than a staffing one, and it lands on the product rather than on the operations team.

Common misconceptions

A person reviewing the output means there is human oversight.

Only where that person can judge it, has time to, and is expected to disagree. Somebody approving four hundred suggestions a shift is supplying a signature, and the system has oversight on paper and none in practice.

Better model accuracy makes oversight less necessary.

It makes it harder. The more often a system is right, the less carefully people check it, so the rare wrong answer is the one most likely to pass through. Accuracy and complacency rise together.

Human oversight means a human makes every decision.

It means a human can intervene meaningfully. That may be reviewing each case, sampling, or holding the ability to stop the system, and which one is appropriate depends on what a wrong decision does.

2 questions test this concept

A claims system produces a recommendation and an assessor approves or rejects it. Each assessor handles around four hundred cases a shift and the override rate is under one per cent. How should the oversight arrangement be described?

  • AEffective oversight, since a qualified person reviews every case.
  • BEffective oversight, because the low override rate shows the system is accurate.
  • COversight in name only, because the volume leaves no time to judge and the override rate suggests review has stopped.
  • DNot oversight at all, since the assessor cannot see the model's reasoning.
Check whether it stuck.

One per page, with a worked explanation.

Start the set
Related material
Book
Weapons of Math Destruction, On automated decisions with no route of appeal.
Template
Launch checklist, Everything that has to happen from two weeks out to one week after release, grouped by when it falls due, each line with a named owner and a go or no go decision on the day.