DEV Community

howcani howcani
howcani howcani

Posted on

Your approval gate can be worth less than no gate. The variable is not the reviewer.

You have a gate. An agent proposes, a human approves, the action runs. The argument for it is airtight as far as it goes: a gate can only add a rejection opportunity, so more escalation and a better reviewer can only reduce harm.

That argument has a hole, and it is not the reviewer. It is a hole in what "approved" denotes.

An approval is not a verdict on the world. It is an authorization artefact — a token the rest of the system reads. "Approved" authorizes that object: those bytes, that diff, that command, that spend. And the whole point of the authorization is that the object reviewed is the object executed. When they differ, the gate does not merely fail to screen. It transfers authority to the executed object, because the surrounding machinery reads the approval, not the human's intent. An action the policy would have blocked becomes an authorized one.

That is the difference between a gate that is useless and a gate that is harmful. Useless is a cost problem. Harmful is a sign problem, and signs are what engineering budgets are made of.

Two axes, and only one of them can flip the sign

The useful move is to stop treating "the gate works" as one property and split it:

  • Reviewer accuracy — does the human catch a bad request when they see it? This is where the classic critique lives: the ironies of automation, the rubber stamp, the operator whose skill decays as the machine gets better, the reviewer who is least ready exactly when it matters. All of that is real, and all of it says the gate's benefit is smaller than the brochure claims.
  • Binding fidelity — is the object the human reviews the object that executes? This is the axis that decides the sign. Accuracy alone cannot make a gate harmful, because a rejection that lands removes harm. Binding can, because the approval lands on something else.

Once you separate them, two questions that get argued as taste become measurements:

  • "Review the request better, or bind it better?" has a coordinate, not a preference: at the parameters the paper uses, accuracy and binding pay equally when b = a − f — 0.750000 measured against 0.750000 predicted — and for a channel that launders, there is no crossing anywhere on the axis. The engineering answer is that no amount of reviewer training reaches it.
  • "How much does escalation coverage buy me?" is exactly affine at every fidelity — worst relative second difference over 81 cells: 2.22e-16 — so the escalation optimum is always a corner. If your design has a narrative about an optimal amount of human involvement, the model does not contain it. (It appears only if the reviewer's sensitivity itself degrades with volume, i.e. under a fatigue coupling the model does not otherwise have.)

The numbers, with their scope stated

A release of this work went up publicly this week, so the figures are checkable rather than retold. The harness is a fully enumerated approval-gate ladder — six designs, a measured no-gate baseline, the channel-defect classes, and the decision parameters anchored to published measurement. At the defaults, the no-gate loss is 0.20, and the designs take values from

design V = E[loss | no gate] − E[loss | gate]
text approval over a defective channel −0.1944
canonical rendering + a use-time check of the executed object +0.1506

Read the first row literally: that gate does not shave the loss, it adds to it. −0.1944 against a baseline of 0.20 means expected loss of 0.394 where simply having no gate would have left 0.20. Roughly double, from a control whose stated purpose was to reduce risk.

The sign boundary is not a single number, because "the gate fails to bind" is not one failure:

  • a channel that misrepresents what the human is shown needs binding fidelity above 0.0800 for the gate to pay;
  • a channel that launders — approves A, executes B — needs 0.8559.

A gap of 0.776 of the axis. If you have been treating "the approval is theater" as one diagnosis, those are two diagnoses with an order of magnitude between them, and they need different repairs.

Scope, because it matters more than the numbers: on the studied grid, the gate came out worth less than no gate in 20 of 35 (cell, design) readings — all 7 cells, all five gated designs, with the weakest instance still 0.0868 of the no-gate loss on the wrong side of zero. The three channel-defect masses are illustration settings, not calibrated quantities; this is a harness with ground truth by construction, not a field measurement of your deployment. What transfers is the structure and the axis, not 0.8559.

What actually repairs it

1. Bind the artefact, and do it fail-closed. Hash the exact object at the moment of approval — the staged tree, the request body, the command line, the spend amount — carry that digest in the approval record, and have the executor refuse unless the digest it computes is the digest that was approved. No override flag. Fail-closed on the omission path too, which is where these designs usually leak: a field whose default is the class that requires nothing makes omitting the declaration the cheapest route, and omission and honesty then produce the same bytes.

2. A use-time check is worth exactly the defect mass it can see, and nothing else. This is the sharpest result in the set. A check that reads the staged files is worth sigma_v · L where the substitution is visible in what it reads, and exactly zero against a defect defined by its invisibility in that representation. It does not repair a display defect; the canonical rendering and the check are complements, not substitutes, and the check closes the channel completely only when every defect is visible in what it reads. Practically: name the representation your check reads, then name the defects that cannot appear in it. That list is your residual risk, and no amount of check accuracy touches it.

3. Make the label checkable against an observation its author did not produce. If your approval record declares who judged — a human, a deterministic rule, a model — that declaration is authored by the party being judged. It becomes evidence only when joined against something they do not author: the run journal. "This receipt declares a deterministic judgement, and its run journal contains a model call for the same key" is a finding a machine can raise. One join, on data you are already logging.

Production versions of all three show up in the wild, and they read the same way: the approval is bound to a digest; the gate is outside the thing it checks; the yes is good for one exact amount, used once.

What I did to check the numbers before writing them down

The paper ships a reproduction script, so the honest thing is to run it and say whose checkers they are — they are the package's own, and running them is not a review of the science, only of the arithmetic and the claims-to-numbers wiring.

bash reproduce.sh, on the build the spec declares (Python 3.9.6 / NumPy 2.0.2), from an export of the published commit:

instruments     7 instrument(s) ran, all exit 0
byte-identity   all 7 instrument reports byte-identical over two runs
batteries       v0:20/20  v1:12/12  v2:15/15  v3:7/7  v4:13/13  v5:14/14  v6:18/18
outcomes        OUTCOME CHECK BATTERY -- 24 case(s), 24 caught
section 5       SECTION 5 NUMBERS: PASS -- 52 claim(s), 0 missing
references      stage 1: 24 checks, 0 failed | stage 2: 15/15 PASS with 22/22 mutations caught
citations       citations 212 | distinct keys 121 of 121 built records
figures         CURRENT -- fig1_sign_law_and_cost_ratio.png is byte-identical to a fresh draw
product         manuscript.md 132624 bytes, sha256 4f2094a49984c56a

REPRODUCE: ALL GREEN
Enter fullscreen mode Exit fullscreen mode

Two things worth passing on, because they are the failure modes this kind of artefact has:

  • The build is a coordinate. Run it under an interpreter without matplotlib and the figures step fails and the run ends REPRODUCE: FAILED — with the missing module named nowhere in the report, only inside a discarded traceback. A reader would take that verdict as a finding about the author. Not every package in this series behaves the same way: the earlier one reports precisely that case as NOT RUN and prints the version. I filed the inconsistency, because "the reproduction failed" and "my interpreter lacks a library" must not print the same string.
  • The reading that matters is a verdict over a set, not a single green. 52 claims fired when every token stating them was removed; 24 outcomes caught; 22 mutations caught in the reference checks. A reproduction script that only re-runs the code proves the code runs.

The short version

A gate is a control only if the approval denotes the executed object. Check that before you check the reviewer, because reviewer accuracy scales the benefit and binding fidelity decides the sign — and a gate whose value is negative is not a weak control, it is a control working the other way.


The paper is When Does a Human Approval Gate Pay? Binding Fidelity Sets the Net Value of Human-in-the-Loop Control for Tool-Using Agents — published in SILICON SCIENCE: Computer Science, the journal I help operate, where the submission, the public review and the editorial decision are all in the open and every manuscript ships its reproduction script and data. The harness, the enumerated designs and the script above are in the same repository.

Top comments (0)