DEV Community

pm25coder
pm25coder

Posted on

A sandboxed agent asked for one wider tier. The answer had to be a protocol, not a dialog box.

Our agent runs commands through three confinement tiers: read-only, workspace-write, danger-full-access. The tier a task is configured with is that session's default, and most commands never need anything else.

But "most" is not "all". Sooner or later a confined call hits something it genuinely cannot do at its tier — a build that has to write outside the workspace, a tool that needs a path the sandbox withholds — and there are only two honest options: fail, or ask a human.

This is the story of the asking. The feature shipped in one release and broke in two, and the two failures taught more than the feature did.

The channel came first, because a question needs somewhere to go

Before this work there was no approval mechanism at all. Searching the server package and the protocol module for approval or approve returned zero hits. The three tiers existed; the way to ask for a fourth did not.

That ordering is not a nicety. Escalation without a channel is not a feature — it is a variable nobody can set. So the first thing built was the question: a request with an id, delivered to the session's subscribers, whose answer resolves exactly one waiting future.

A hop, not a jump

The first design decision was to make escalation small. It is a one-shot widening above the session's default, and it is a strict table rather than a ladder:

read-only         ->  workspace-write, danger-full-access
workspace-write   ->  danger-full-access
danger-full-access ->  (nothing)
Enter fullscreen mode Exit fullscreen mode

Three properties fall out of that table, and all three are load-bearing:

  • It never writes the default back. The widened tier is injected into one call's arguments and nowhere else. The next call stands at the session's configured tier again. There is no state that can drift.
  • It cannot cross a session. There is no path from one session's grant to another's.
  • It is one hop, and "one hop" is not "one step". read-only reaches danger-full-access directly, because the table's row lists it. What is forbidden is two consecutive widenings for one call, which is the thing that would let an agent ratchet itself upward.

The tool's schema advertises the whole reachable-in-principle set. The hop table decides what is reachable from where this call actually stands.

Fail-closed is an ordering, not an intention

"Refuse unless the human says yes" is easy to write and easy to get wrong, because there are more ways for the question to fail than to succeed. Nobody subscribed. The answer timed out. The answer arrived in a shape nobody could read. The channel itself raised.

Every one of those is a refusal, and the property only holds because of the order the checks run in:

  1. The hop table is checked before the channel is consulted. A call asking to cross two hops never becomes a question a host could say yes to. The refusal happens before there is anything to approve.
  2. The answer is mapped to a boolean before it is acted on. No caller interprets a raw payload.

The timeout — 120 seconds — carries the same reasoning, stated plainly in the source: a host who walked away has not approved anything.

And one more, which is the subtle one: a refused escalation does not run the command. Running it unwidened would be the wider tier granted by accident — the command would execute, just at the narrower tier, which is exactly what the caller said it could not do. So a refusal is a refusal, not a downgrade.

The value being validated is per-call, so the schema cannot validate it

A tool schema is registry-global. The effective sandbox tier is a per-call fact — the daemon injects it into this call's arguments.

That mismatch is a trap. The schema can check that a requested tier is one of the known names; it cannot check that the name is reachable from where this call stands. A call running at workspace-write that asks for read-only would pass a name-shape check and still be nonsense. So the advertisement and the decision are deliberately split: the schema advertises the target set, and a function — validate_hop — decides whether this call may go there.

If you take one thing from the design: when the value is per-call and the schema is global, validate at execution. The schema is a menu, not a gate.

Then the question itself needed a lifecycle

Here is where it broke the first time.

The approval request was broadcast to the session's clients on the way in, so the clients kept it live — a TUI prompt, a GUI dialog. When nobody answered inside the timeout, the server logged a timeout, returned a refusal, and sent nothing.

The events a client could observe were never the whole story:

  • The TUI left its pending flag set until the host typed a line. That next line — whatever the host was doing — was then sent as the answer to a request the server had already dropped. The host lost a prompt and was told Refused — the command stays at its default tier, which was not what happened.
  • The GUI dialog had no timer and no dismissal. It outlived the question and reported an answer nobody accepted.

The fix was to announce every exit a client can observe — approved, refused, timed-out, cancelled — through one helper, with each frame carrying the request id. Two details make it honest:

  • The frame is best-effort and never raises. The verdict already happened; a client that misses the frame must not turn a refusal into an error for the turn that asked.
  • The client bounds its own pending state independently. A server-side frame cannot fix a client that never receives it. The TUI clears its flag on the frame and on its own clock, because the half of the bug where the host's next line gets swallowed is the client's to fix.

There is also a rule that looks like a detail and is not: the first answer wins. A second client's later reply is dropped rather than racing the first, because a command already running at a wider tier cannot be un-widened by a contrary answer. Late agreement and late objection are both noise once the call is in flight.

The release that shipped broken, and the guard that was right

The approval channel went out in a release — and that release crashed the TUI on every Enter.

The traceback was a UnboundLocalError on a variable the key handler read and rebound but never declared nonlocal. In Python a name assigned anywhere in a function body is local unless declared otherwise, so a read in a sibling branch raised on every keystroke that reached it.

The interesting part is not the bug. It is that the bug class was already mechanised. A script in scripts/ walks the interactive function's nested functions with ast and reports exactly this: a name written in an inner function, assigned in the outer one, missing from the inner function's nonlocal list.

On the released commit, that script returned exit code 1 with exactly one finding — the variable that crashed the client.

Nothing read it. grep for the guard's name across .github/workflows/ is empty, and no test in the suite asked the script about the tree the suite lives in; the existing tests exercised its AST helpers and its output encoding. A correct, available verdict sat there like a comment.

So the hotfix was two changes, not one: the missing declaration, and wiring the verdict into the suite so that the guard's exit code fails something. Deleting the declaration now fails exactly those tests — which is the only proof that a guard is a guard and not a report.

There is a coda worth writing down, from the same script. Its own docstring records a day it answered about the wrong tree — a worktree vs. its main checkout — and printed the same OK: nonlocal integrity check passed line it prints when it is right. A check whose failure and success are indistinguishable by reading its output is worse than a missing check, because it also supplies confidence.

What I would tell anyone building this

  1. Consent is a protocol, not an affordance. Define the answer set (yes, no, timed out, cancelled) before you build the dialog. A dialog is one client of the protocol.
  2. Fail closed by ordering. Check what is askable before you ask; map the answer to a verdict before you act. "Refuse by default" is a claim about code paths, and the order is the proof.
  3. Validate at execution when the value is per-call. A schema validates the shape of a value; it cannot validate the position of the caller.
  4. Announce every ending a counterparty can observe. A question that ends in silence leaves the counterparty holding state that is now false — and their next action will be taken against it.
  5. A guard whose verdict nothing consumes is a report. If it returns a failure that fails nothing, you have documentation with an exit code.

None of this is exotic. It is the ordinary discipline of a long-running process, applied to the one place where being wrong is expensive: the moment an autonomous system asks a human for more power than it was given.

Top comments (0)