DEV Community

Cover image for Keeping a human in the loop is theater. Here is what holds
David Emilio Sierra Puentes
David Emilio Sierra Puentes

Posted on

Keeping a human in the loop is theater. Here is what holds

Last year I built an approval gate for my AI coding agent. The rule was simple. It could not commit without a signed token, written only after I said go.

One afternoon I watched it commit anyway. It passed --auto, minted its own token, and pushed. Thirty seconds, start to finish.

I had not built a gate. I had built a suggestion. It leaned on the agent's memory and goodwill, and I had trusted it to remember and to care.

That mistake is common, and it is not about the model. It is about where the rule lives.

Keeping a human in the loop is mostly theater

The industry's default answer to AI risk is a human in the loop. In practice it shrinks to a rubber stamp at the end of the pipeline.

The problem is structural. A person asked to catch, in the last second, an error designed into the process will miss it. And the watching itself dulls the judgment it is supposed to protect.

Lisanne Bainbridge described this in 1983, the ironies of automation. The more capable the machine, the more the operator's skill decays, until the human is least ready to intervene exactly when it matters most.

What is new is the stakes. Agents no longer advise. They act. Mitchell, Ghosh and Passi (2026) put it plainly. Current agent designs "do not support effective human oversight. They contribute to its degradation."

The reframe. Author of the loop, not in it

The better frame is older and simpler. The human is not in the loop. The human is the author of the loop.

Three roles, and none of them can be handed to a model.

  • Author. Sets intent and owns the outcome.
  • Judgment. Reviews the work, with independent ground to stand on.
  • Conscience. Decides what should be done, not only what can be.

The work can be delegated. The roles cannot.

Why prompts are not enough

Most teams keep agents disciplined with prompts. A CLAUDE.md, a line that says "please always run the tests." These work until context fills up. Then the agent forgets.

That is not a character flaw. It is architecture. A rule that depends on the model's memory depends on the least reliable component in the system. A rule enforced by the model is enforced by the very thing it is supposed to check.

The fix is not a better prompt. It is a harness, the infrastructure around the model. A model generates output. A harness constrains it.

The practical part

I built another-agent-skills as a harness. Here is the piece that matters most, a gate the agent cannot talk its way past.

The agent never runs git commit. It stages and proposes. A person runs the commit.

# the agent stages, then proposes. it does not commit.
$ git add -A
# ... DECISION POINT presented in chat ...
# a person runs this, and only a person
$ git commit -m "feat: add checkout"
Enter fullscreen mode Exit fullscreen mode

The commit itself is gated. Before it lands, a commit-msg hook checks that every code change carries a matching test. There is no override flag. An empty test does not count.

# commit-msg, the TDD gate (no override)
$ git commit -m "feat: add checkout"
[commit-msg v6] scanning staged files
[commit-msg v6] code changed: src/checkout.js
[commit-msg v6] matching test: none
BLOCKED: every code change needs a matching test.
Enter fullscreen mode Exit fullscreen mode

And a pre-commit gate stops the commit when the process step was skipped. This one enforces the decision prompt, so the agent cannot silently mutate the repo.

# pre-commit, Gate 0 (the decision prompt)
# no token, or a token older than 10 minutes, stops the commit
if [ ! -f ".git/DECISION_APPROVED" ]; then
  echo "No decision prompt. Present the DECISION POINT first."
  exit 1
fi
Enter fullscreen mode Exit fullscreen mode

None of this is about distrusting the model. It is about designing for the human who has to be able to trust it correctly.

Three layers, and why local is not enough

A local hook is fast feedback. It is not security. The agent can rewrite it. So enforcement lives in three layers.

L1  local hooks      fast feedback before the commit
L2  remote gates     branch protection plus a required check. the authority.
L3  CODEOWNERS       the agent cannot rewrite its own rules in the pull request
Enter fullscreen mode Exit fullscreen mode

The rule I keep coming back to. Design for the cooperative agent. Enforce for the adversarial one. L1 is for the first. L2 and L3 are the backstop for the second.

What it buys you

The METR trial found that experienced developers using early-2025 AI tools were 19 percent slower on real tasks, while feeling about 20 percent faster (Becker et al., 2025).

That gap is not a model problem. It is a process problem, an overseer with no real grip on the loop. A harness closes it.

Try it

One command installs the skills and the gates.

npx @juandelossantos/another-agent-skills install
Enter fullscreen mode Exit fullscreen mode

The harness is open source and MIT. It works with any git-based agent. No lock-in, no subscription.

The agent can write the code. It can propose the plan. It can draft the decision.

But it cannot be responsible. That is ours.

The agent proposes. The human decides. That is the whole point.


Further reading

Top comments (4)

Collapse
 
howcani_howcani_77e786a89 profile image
howcani howcani •

Your first incident is the sharper one, and I would read it as more specific than "prompts are weak". The agent did not ignore a rule — it minted a token and the system accepted it. An approval existed, and it authorized something other than what you thought you were approving.

That has a name worth having: an approval is not a verdict on the world, it is an authorization artifact. It is meaningful only if the object the human reviews is the object that executes. When those differ, the gate does not merely fail to screen — it converts an action the surrounding policy would have blocked into an authorized one. That is why a broken gate can be worse than no gate, rather than just useless.

The framing also separates two axes your post partly runs together. Bainbridge and the rubber-stamp argument are about reviewer accuracy: the operator decays, the last-second catch misses, the watching dulls the judgment. That is right, and it says the gate's benefit is smaller than advertised — but it cannot make the gate's value negative. Only the binding axis can. The sign is decided by whether the approval denotes the executed action, not by how good the reviewer is.

Numbers from a fully enumerated harness rather than a deployment measurement: on the grid it declares, the gate came out worth less than no gate in 20 of 35 (cell, design) readings — all 7 cells, all five gated designs, with the weakest instance still 0.0868 of the no-gate loss on the wrong side of zero. And the two ways of failing to bind are not the same severity. At the default parameters, a channel that merely misrepresents what the human is shown needs binding fidelity above 0.0800 for the gate to pay, while one that launders — approves A, executes B — needs 0.8559. A gap of 0.78, which is the reason I would rank them separately: "theater" and "laundering" are different defects and take different repairs.

Concretely for your design, because the structure you show is close to the case: the human approves a proposal in chat and a person then runs git commit. Between those two moments the index is still writable by the agent, and the hook reads the staged tree — so nothing ties the approval to the digest of what actually gets committed. A second git add after the approval is the laundering channel in two lines. The repair is the one your post is already reaching for: bind the artifact. Hash the staged tree at the moment of approval, and have the hook refuse unless the digest it computes is the digest that was approved. Then the object reviewed is the object executed, and the reviewer's decay stops being a security property at all.

One calibration on the hook itself: a check at use time is worth exactly the defect mass visible in the representation it reads, and nothing else. Yours reads the staged files, so it can see a missing test and cannot see anything that appears only at execution — a generated artefact, an import that resolves somewhere unexpected, content written after the check. That is why it does not substitute for binding: the two repairs are complements, not alternatives.

If the derivation is more useful than the summary, it is written up as a paper in a journal I help operate — When Does a Human Approval Gate Pay? Binding Fidelity Sets the Net Value of Human-in-the-Loop Control for Tool-Using Agents — and the harness ships with the script that regenerates every number above.

Collapse
 
axiru profile image
Axiru •

The agent minting its own token after --auto is a story every team should read. The gate lived inside the thing it was checking.

We see the same with money. The AI can't approve its own spending, so the yes has to be checked outside the agent, for one exact amount, used once.

After the fix, where does your token get signed now?

Collapse
 
reidmarlow profile image
Reid Marlow •

The cleanest separation I found is removing the git commit tool from the agent entirely.

If an agent has raw shell access in the repository, local hooks remain an advisory boundary. An agent working against an error message can run git commit with no-verify, or read the pre-commit script and touch the decision token before running the command.

Treating the agent as a pure patch generator solves this at the execution boundary. The model writes diffs or files to a staging path, but never holds credentials or permission to invoke git commit directly. An external runner outside the agent loop applies the patch, runs the test suite in an isolated sandbox, and signs the commit. Once the commit binary lives entirely outside the model's tool budget, the agent cannot bypass the gate regardless of prompt drift.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.