DEV Community

Your Agent's Guardrails Can't See the Money

arun rajkumar on August 21, 2026

There's a post going round this week about agent guardrails that opens with a good story. The author's agent wanted to force-push to main. Not beca...
Collapse
 
peterbuildssecure profile image
Peter •

A declaration in the tool schema would be useful for routing, but I would not make the declaration part of the security boundary. A malicious or simply incomplete tool can omit it.

The stronger convention is a preview–authorize–execute protocol owned by the system holding the state. Preview resolves the proposed effect and returns an operation ID plus the subject, target, amount, currency, relevant state version and expiry. Human approval produces a single-use mandate bound to that exact operation. Execute consumes it atomically while rechecking the authoritative state and idempotency key.

That handles the preflight race without asking the agent host to reproduce the ledger’s rules. If the balance, refund total, ownership or dispute state changes, execution rejects the stale mandate and requires a new preview.

The tool metadata can still say “this operation requires effect authorization” so clients present the right flow. But the server should require the mandate whether or not the client noticed that metadata. Otherwise the convention protects careful clients while direct callers retain the bypass.

Collapse
 
mickyarun profile image
arun rajkumar •

Also three weeks late, which is worse given this is the most useful comment on the post.

The load-bearing sentence is your last one. A rule that only binds clients who noticed the metadata isn't a boundary, it's documentation, and the direct caller keeps the bypass. Anything described as a convention has already told you what it is.

What I'd add from the payments side: you have described an authorisation, and the industry arrived at that shape by being robbed repeatedly. Bound to the specific operation, verified by the side holding the state, rejected when stale. The part I would hold onto hardest is expiry, because it is what stops the set of accepted mandates growing into a list nobody can audit.

Where I get uncomfortable is preview itself. Execute rechecks authoritative state, which is right. But preview has to resolve the proposed effect, and any gap between what preview simulates and what execute actually does is a new place for a difference to live. Much smaller surface than the one you're closing. Not zero.

Collapse
 
peterbuildssecure profile image
Peter •

The preview/execute gap is the right thing to be uncomfortable about, and I don't think it's fully closed by "execute rechecks authoritative state" alone — that guarantees execute uses fresh state, not that it computes the same effect preview showed the user. Binding the mandate to a digest of the previewed effect itself (amount, destination, operation) and having execute recompute that digest fresh from current state before acting closes the specific gap you're describing: if preview and execute logic ever drift — a bug, a race, a code path that changed between the two calls — the mismatch is a structural comparison failure, not a silent behavioral difference nobody notices until the money's gone. Doesn't remove the surface, but it turns "hope preview and execute agree" into something that fails loud when they don't.

Thread Thread
 
mickyarun profile image
arun rajkumar •

You're right and I overstated it. Rechecking at execute guarantees the state hasn't moved. It doesn't guarantee the operation about to run is the one that was previewed, and those are different claims. A preview of a £40 refund and an execute of a £4,000 refund can both see an unmoved account.

What closes it is binding the authority to the fields, not only to the state. Amount, payee, operation, inside the thing being verified. That's what a card authorisation does and it's why it can't be re-presented for a different amount. I wrote this up since and you're quoted: dev.to/mickyarun/four-people-rebui...

Thread Thread
 
peterbuildssecure profile image
Peter •

Binding named fields (amount/payee/operation) works until someone adds a field nobody enumerated. A more failure-proof version: hash the entire canonicalized argument payload (stable key order, typed encoding) and bind the mandate to that hash instead of a field list. Then the mandate covers exactly these arguments, not these three arguments plus whatever we forgot, and a new field on a new operation type falls inside the binding automatically instead of needing someone to remember to update the mandate schema.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Agreed on the failure mode. A field list is a promise that someone will remember to update it, and that promise breaks quietly.

The cost of the whole-payload hash is the mirror image. It binds fields the execute side is allowed to fill in. Idempotency keys, correlation ids, a provider reference assigned after preview, a timestamp. Hash all of it and a legitimate execute stops matching. So you write canonicalisation rules that exclude those, and you are enumerating again, just from the other end.

The difference is which way the enumeration fails. Miss a field on a binding list and you get a silent hole. Miss one on an exclusion list and you get a loud broken execute. The second is much better. I think that is the real argument for your version, and it is stronger than "covers everything", because it does not.

Collapse
 
mickyarun profile image
arun rajkumar •

Follow-up for @max_quimby, @peterbuildssecure and @jon_at_backboardio.

This thread and a thread on a completely different post ended up describing the same object, so I wrote it up rather than leave it sitting in two comment sections. All three of you are quoted.

dev.to/mickyarun/four-people-rebui...

Short version: what got built in this thread is a payment authorisation. Bound to the instruction, verified by the side holding the state, single-use, expiring. Payments arrived at that shape by being robbed for forty years rather than by designing it.

The part I couldn't close is that payments has a natural transaction boundary to hang the expiry off, and an agent halfway through a long task hasn't got one.

Collapse
 
max_quimby profile image
Max Quimby •

The refund example is the best articulation I've seen of why string-level hooks plateau. We run agent pipelines with pre-execution hooks, and they're genuinely good at the force-push class of problem — anything where the danger is written into the command. But every expensive near-miss we've had looked exactly like your second refund call: schema-valid, well-formed, reasonable-sounding arguments, and the thing that made it wrong lived in state the hook couldn't see.

What's worked better for us is pushing invariants down into the resource layer instead of the call layer: idempotency keys so a duplicate refund is structurally impossible, per-target baselines so an amount 1000x the historical median for that account needs a second signal, and reconciliation jobs that compare intent logs against effects after the fact. None of that is agent-specific, which I think is the point — the agent just removes the human who used to eyeball the number.

Curious whether you'd put the check synchronous (block the call) or async (catch within minutes)? Sync is safer but the latency tax on every benign call adds up.

Collapse
 
mickyarun profile image
arun rajkumar •

Three weeks late, sorry. This deserved an answer the day you wrote it.

Sync where the money is, async everywhere else, and the threshold is the whole design. @jon_at_backboardio made the case below better than I would have: the calls where blocking hurts are high-volume and low-value, and the calls where blocking is fine are the ones you actually want to stop. Those two sets barely overlap, so the latency tax lands on a rounding error of your traffic.

The line I keep coming back to is yours though. The agent just removes the human who used to eyeball the number. What that exposed for us is that the eyeball was never a control. It was a sampling process with unknown coverage that had been getting booked as a control for years. Losing it is bad. Finding out we had been depending on it is worse.

One distinction on pushing invariants into the resource layer: idempotency keys make a duplicate structurally impossible, which is a real win, but they do nothing about wrong-once. The second refund call in the post isn't a duplicate. It is a first-time, well-formed, singular £40,000 mistake, and every key in the world lets it through.

Collapse
 
jon_at_backboardio profile image
Jonathan Murray •

on max's sync vs async question, i think the tax is smaller than it looks because the two axes line up in your favour.

the calls where blocking hurts are high volume and low value. the calls where blocking is fine are the ones you actually want to stop. a £40,000 refund can afford 200ms. forty thousand £4 refunds cannot, and don't need to.

so it isn't sync or async, it's a cheap sync gate whose threshold comes from the target. one cached number per target, refreshed lazily, some multiple of that account's recent movement. under it, execute and reconcile after. over it, block and go get the mandate. you pay full preflight latency on a rounding error of your traffic, and the cached number being slightly stale doesn't matter because it isn't authorising anything, it's only deciding whether to ask.

the case that worries me more than either is your last one, multi-step. five refunds of £8,000 each, every one under every cap, same payee, four minutes. every individual call is boring. the mandate model catches that only if mandates are scoped to a window rather than to an operation, and that's a different shape from what peter described.

"a rubber stamp with extra steps" is the line. showing a human raw args isn't review, it's laundering.

Collapse
 
mickyarun profile image
arun rajkumar •

Very late reply, and this one has been sitting in my head since August.

"It isn't authorising anything, it's only deciding whether to ask" is the sentence that makes the whole design work. Staleness is disqualifying for an authorisation and completely fine for a triage decision, and collapsing those two is why people over-engineer the gate.

On multi-step, you're right that it's a different shape, and payments has a name for it: velocity. Cumulative exposure per payee per window rather than a cap per transaction. It has been standard in card fraud for decades, precisely because five boring transactions are the attack. What nobody tells you is that the check is the easy part. Choosing the window is the hard part, and whatever you choose, someone can straddle it. Four minutes clears an hourly window if you start at 59 minutes past.

And "showing a human raw args isn't review, it's laundering" is the best line anyone has left on anything I've written. It is the failure mode of every approval screen I have ever built.