A person checks the output is the mitigation everybody writes down. It goes in the design doc, the risk register and the compliance answer, and it is treated as a fixed quantity — set up review, get a catch rate.
Over one whole stretch of the improvement path, a model with half the error rate ships MORE wrong answers to the customer. Same reviewer, same skill, same policy, same coverage, same everything in the runbook.
Move the slider: https://dev48.infy.uk/ai/days/day76-human-oversight.html
Nothing here simulates language
The engine simulates the decision. A reviewer is a detector with a sensitivity and a threshold, and the threshold that minimises their expected cost is a function of how often they actually meet an error — which is a function of the model.
model gets better
-> errors get rarer
-> the reviewer's rational threshold rises
-> the reviewer catches a smaller fraction
Two curves move in opposite directions and the product is not monotone. Over part of the range, the second effect wins.
Why this is not a story about lazy reviewers
The threshold shift is rational. A reviewer who flags at the same rate when errors are ten times rarer is spending almost all of their flags on correct output, and someone will eventually tell them to stop. The behaviour that makes oversight degrade is exactly the behaviour you would ask for.
The line worth taking away
The catch rate is not a property of the reviewer. It is a property of the pair.
So "a human reviews it" is not a mitigation you can state once and carry forward. Re-measure it every time the model changes — the number you measured against the old model does not describe the new system, and it can be worse in the direction nobody checks.
Verifier 2,236 assertions, 0 failures.
Top comments (0)