DEV Community

arun rajkumar
arun rajkumar

Posted on

Your Agent's Guardrails Can't See the Money

There's a post going round this week about agent guardrails that opens with a good story. The author's agent wanted to force-push to main. Not because it was confused. The rebase was stuck, force-pushing would unstick it, and every step in that chain of reasoning was sound. Locally correct, non-locally expensive.

What makes that example teachable is that the command carries its own consequence. git push --force origin main has the danger written into it. You can pattern-match it. You can put it on a list. A hook can read the string and stop.

Most of what I worry about doesn't look like that at all.

Two calls, one of them a disaster

Here are two refund calls.

await payments.refund({
  paymentId: "pay_9f2c14",
  amount: 4000,
  currency: "GBP",
  reason: "customer_request",
});

await payments.refund({
  paymentId: "pay_9f2c14",
  amount: 4000000,
  currency: "GBP",
  reason: "customer_request",
});
Enter fullscreen mode Exit fullscreen mode

One is forty pounds. The other is forty thousand. Same endpoint, same argument names, same shape, same reason code, same everything a pre-execution hook can read off the text.

Now the part that actually matters: neither of them is inherently wrong.

Forty thousand might be a perfectly good refund against a perfectly good invoice. Forty might be a refund on a payment that was already refunded an hour ago, which in some ways is the worse of the two. You cannot sort these by looking at them, because the thing that makes one of them a mistake isn't in the call.

Blast radius is a property of the target, not the command

The force-push case works because the danger lives in the verb. Force-push is dangerous in nearly every context. The set of situations where you genuinely want it is small enough to enumerate, so a rule can cover it and the rule stays roughly true.

An amount-bearing API call is dangerous as a property of the state on the other side. Whether that refund is safe depends on things the caller simply doesn't have:

  • Has this payment already been refunded, fully or partially?
  • Does the merchant's balance cover it, or does this pull them negative?
  • How much has this actor already moved today?
  • Is the original payment under dispute, where a refund does something quite different to who ends up liable?
  • Is pay_9f2c14 even this merchant's payment?

A hook sitting in the agent process knows none of this. It can read intent. Cost is not a property of intent.

"Fine, have the hook go and check"

This is the obvious next move and it isn't stupid. Let the guard call the API, fetch the payment, look at the balance, then decide.

Try building it and you find out what you signed up for.

Your hook now needs credentials to read payment state, which means the agent host is holding read access to your ledger in order to protect you from the agent host. It needs to understand refund semantics, dispute states and settlement timing well enough to form a judgement, so your refund rules now live in two codebases that have to agree forever. And it has a window between checking and executing, which is exactly where the interesting failures live. The balance was fine when you looked. Something else landed. Yours goes through anyway.

You haven't built a guardrail. You've built a second, worse copy of your authorisation service, running in the least trusted process you own, with a cache.

Put the check where the state is

The version that survives contact is boring. The check goes at the rail, at the point of authorisation, because that's the only place that knows what this actor may move, what it has already moved, and what the target looks like right now.

We're regulated, so we already had to have that spine. Every movement of money goes through an authorisation step with the full picture and an audit record on the other side. Wiring agents into it didn't mean building something new. It mostly meant resisting the urge to build something in front of it and call that safety.

The agent-side hook still has a job. Just a smaller one than people want to give it.

// The hook classifies and asks. It does not decide.
async function preToolUse(call) {
  if (!MOVES_MONEY.has(call.tool)) return allow(call);

  // Ask the system that owns the state what this would actually cost.
  const effect = await rail.preflight(call.args);

  return confirmWithHuman({ intent: call.args, effect });
}
Enter fullscreen mode Exit fullscreen mode

rail.preflight is where the work happens, and it lives on the side that has the ledger. Its whole purpose is so a human sees "this refunds £40,000 against an invoice already refunded in full on the 3rd" rather than "the agent would like to call refund." One of those is a decision. The other is a rubber stamp with extra steps.

It's also explicitly advisory. Between preflight and execute the world moves. Enforcement stays server side, under a mandate scoped to an amount, a payee and a clock. If preflight and the rail ever disagree, the rail wins.

I wrote more about mandates versus API keys in an earlier post if that thread interests you.

The uncomfortable bit

Your agent framework cannot own your safety story for anything that moves money, and it doesn't matter how good its hooks get.

That's not a knock on the hooks. Command-level guards are useful and I'd rather have them than not. For destructive verbs, where the danger is in the word, they're the right tool and they'll catch real mistakes that would otherwise cost someone a weekend.

They just sit at the wrong altitude for a whole class of call that reads as completely unremarkable and is defined entirely by context the caller doesn't hold.

So here's the test I'd apply to any guard before trusting it. Could this exact call, byte for byte, be both correct and catastrophic depending on something the guard cannot see? If the answer is yes, that guard is a linter. Useful. Not load-bearing. Put the real check where the state lives.

Has anyone landed on a decent convention for a tool declaring "you can't infer my blast radius from my arguments, come and ask"? I haven't seen it in the MCP spec, and it feels like the piece that's missing.

Top comments (21)

Collapse
 
peterbuildssecure profile image
Peter •

A declaration in the tool schema would be useful for routing, but I would not make the declaration part of the security boundary. A malicious or simply incomplete tool can omit it.

The stronger convention is a preview–authorize–execute protocol owned by the system holding the state. Preview resolves the proposed effect and returns an operation ID plus the subject, target, amount, currency, relevant state version and expiry. Human approval produces a single-use mandate bound to that exact operation. Execute consumes it atomically while rechecking the authoritative state and idempotency key.

That handles the preflight race without asking the agent host to reproduce the ledger’s rules. If the balance, refund total, ownership or dispute state changes, execution rejects the stale mandate and requires a new preview.

The tool metadata can still say “this operation requires effect authorization” so clients present the right flow. But the server should require the mandate whether or not the client noticed that metadata. Otherwise the convention protects careful clients while direct callers retain the bypass.

Collapse
 
mickyarun profile image
arun rajkumar •

Also three weeks late, which is worse given this is the most useful comment on the post.

The load-bearing sentence is your last one. A rule that only binds clients who noticed the metadata isn't a boundary, it's documentation, and the direct caller keeps the bypass. Anything described as a convention has already told you what it is.

What I'd add from the payments side: you have described an authorisation, and the industry arrived at that shape by being robbed repeatedly. Bound to the specific operation, verified by the side holding the state, rejected when stale. The part I would hold onto hardest is expiry, because it is what stops the set of accepted mandates growing into a list nobody can audit.

Where I get uncomfortable is preview itself. Execute rechecks authoritative state, which is right. But preview has to resolve the proposed effect, and any gap between what preview simulates and what execute actually does is a new place for a difference to live. Much smaller surface than the one you're closing. Not zero.

Collapse
 
peterbuildssecure profile image
Peter •

The preview/execute gap is the right thing to be uncomfortable about, and I don't think it's fully closed by "execute rechecks authoritative state" alone — that guarantees execute uses fresh state, not that it computes the same effect preview showed the user. Binding the mandate to a digest of the previewed effect itself (amount, destination, operation) and having execute recompute that digest fresh from current state before acting closes the specific gap you're describing: if preview and execute logic ever drift — a bug, a race, a code path that changed between the two calls — the mismatch is a structural comparison failure, not a silent behavioral difference nobody notices until the money's gone. Doesn't remove the surface, but it turns "hope preview and execute agree" into something that fails loud when they don't.

Thread Thread
 
mickyarun profile image
arun rajkumar •

You're right and I overstated it. Rechecking at execute guarantees the state hasn't moved. It doesn't guarantee the operation about to run is the one that was previewed, and those are different claims. A preview of a £40 refund and an execute of a £4,000 refund can both see an unmoved account.

What closes it is binding the authority to the fields, not only to the state. Amount, payee, operation, inside the thing being verified. That's what a card authorisation does and it's why it can't be re-presented for a different amount. I wrote this up since and you're quoted: dev.to/mickyarun/four-people-rebui...

Thread Thread
 
peterbuildssecure profile image
Peter •

Binding named fields (amount/payee/operation) works until someone adds a field nobody enumerated. A more failure-proof version: hash the entire canonicalized argument payload (stable key order, typed encoding) and bind the mandate to that hash instead of a field list. Then the mandate covers exactly these arguments, not these three arguments plus whatever we forgot, and a new field on a new operation type falls inside the binding automatically instead of needing someone to remember to update the mandate schema.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Agreed on the failure mode. A field list is a promise that someone will remember to update it, and that promise breaks quietly.

The cost of the whole-payload hash is the mirror image. It binds fields the execute side is allowed to fill in. Idempotency keys, correlation ids, a provider reference assigned after preview, a timestamp. Hash all of it and a legitimate execute stops matching. So you write canonicalisation rules that exclude those, and you are enumerating again, just from the other end.

The difference is which way the enumeration fails. Miss a field on a binding list and you get a silent hole. Miss one on an exclusion list and you get a loud broken execute. The second is much better. I think that is the real argument for your version, and it is stronger than "covers everything", because it does not.

Thread Thread
 
peterbuildssecure profile image
Peter •

That's the sharper argument, and it points to a stronger version still: pair the exclusion list with a closed-world check on the payload schema itself. Anything present that isn't in the exclusion list must both exist at preview and match exactly at execute -- no exceptions. A brand-new field nobody's categorized yet then fails by construction: it's neither excluded nor bound, so it can't sneak into either bucket. Still enumeration, but now the failure mode for an uncategorized field is "reject unconditionally," not just "loud instead of silent."

Thread Thread
 
mickyarun profile image
arun rajkumar •

Closed-world is the right default and I would take it.

The cost is that it moves the failure from design time to deploy time. The first person who adds a field to the execute-side payload gets a rejection in production, not a review comment. That is the trade I would make, but it needs saying out loud, or someone widens the exclusion list at 2am to clear the alert and the closed-world property is gone.

Practical version: make the exclusion list a schema annotation next to the field rather than a separate list living in policy code. Then adding a field and categorising it are the same edit. The check is enforcing a decision someone actually made, instead of punishing them for not knowing another file existed.

Thread Thread
 
peterbuildssecure profile image
Peter •

The annotation-next-to-field move buys you something more, too: once the categorization lives in the schema, a CI check can diff it on every PR and fail the build if a new field has no annotation — so the failure moves from 'runtime rejection in production' to 'merge-time CI failure caught by the person who just added the field, with full context.' That's strictly better than either the original exclusion list or a bare closed-world runtime check, because the person paying the cost is the one who created it, not whoever's on call when the payload finally changes shape.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Yes. Merge-time beats runtime, and "the person paying the cost is the one who created it" is the actual argument, not the tooling.

The limit is where the field came from. A CI diff sees fields someone on your side added to the schema. The case that bites hardest is the provider adding one. A new reference in the execute-side response, nothing changed in our repo, nothing for CI to diff. That one still lands on whoever is on call, and the runtime closed-world rejection is the right response to it, because nobody here has decided anything about that field yet.

So both checks stay. CI catches every field we introduced, at the moment we introduced it. The runtime check catches the ones we didn't, and its firing is the signal that a categorisation is now owed. The runtime check is the one that gets widened at 2am, which is why the rejection should name the field and open the same PR the CI check would have.

Thread Thread
 
peterbuildssecure profile image
Peter •

The rejection opening a PR is the right instinct, and it doesn't have to depend on someone remembering to do it at 2am — the reject handler can file that PR itself. On first sight of an unknown field, it diffs against the schema and opens a PR with the field pre-populated as an exclusion-list candidate, tagged for review. The human still decides yes/no on categorization; they just stop being the one who has to remember to start the process.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Right, and it removes the step that was always going to be skipped.

One thing I would change: do not pre-populate it as an exclusion candidate. The reviewer at 9am sees a PR that is already a valid diff, already passes CI, and already carries the answer. Approving it is one click and it looks like agreement. That is the same widening-at-2am you are designing out, moved into business hours and given a green tick.

Open it with the categorisation field empty and a check that fails on empty. The PR exists, the process started without anyone remembering, and the merge still requires somebody to type which bucket the field goes in. The automation does the remembering, the human does the deciding, which is the split you named.

The half I do not have an answer for: the PR says a field appeared. It cannot say why, and the why is at the provider. So whoever picks it up is categorising from a field name and a sample value, and "reference_2" categorised from a sample value is a guess that will read as a decision in the audit log a year later.

Thread Thread
 
peterbuildssecure profile image
Peter •

The empty-field-plus-failing-check version is the right shape — it's the same principle as the machine-checkable ticket ID from the retention thread, just applied to a different artifact. On the audit-log gap: you can't capture why the provider added the field, but you can stop the PR from being 'a name and a sample value' by attaching the actual sample payload(s) that triggered the check to the PR body when it opens. A year later, 'reference_2 categorised as X' reads as an unexplained decision; 'reference_2 categorised as X, based on these three sample values, this shape' reads as a decision with its evidence attached. It doesn't recover the provider's intent, but it recovers what the categoriser actually had in front of them, which is the part that currently disappears.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Attaching the evidence is right and I would do it, but in this domain the attachment is the problem. The unknown field is the one you know nothing about, which means you cannot know it is safe to paste. reference_2 is exactly where a name or an account detail ends up. So the PR body gets a redacted shape, and redaction removes the part that made the sample worth having.

What survives: put the shape and the cardinality in the PR. Length, character class, how many distinct values across the sample, whether it repeats per payer. All of that is decidable automatically, carries no payload, and is most of what the categoriser actually reasoned from. Then a pointer to the real samples somewhere access-controlled, so a year later the decision is auditable by someone who is allowed to see them.

Which makes the audit trail two-tier, and the cheap tier is the one that survives. Same split you and @anp2network landed on for the mandate records on the other thread. I did not expect it to turn up in schema review.

Thread Thread
 
peterbuildssecure profile image
Peter •

The shape-and-cardinality split is the right move, and it's worth naming what it actually is: you're releasing a statistic about the field instead of the field, which is the same move differential-privacy-adjacent logging makes when it can't release raw values — publish what's decidable about the distribution, gate the individual records. Security logging already does a version of this: SIEM indexes commonly store a truncated or hashed identifier in the searchable tier and keep the full record in a restricted vault, precisely so most triage never needs the access-controlled tier at all. Same two-tier shape, just applied to schema categorization instead of log records. Worth stealing the retention-asymmetry that comes with it too — the cheap tier (shape, cardinality, repetition) can live indefinitely since it carries no payload, while the access-controlled tier gets whatever retention policy applies to raw PII, and those two clocks don't have to match.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Retention asymmetry is the part I had not priced, and it is a better reason to build it this way than the one I gave.

One correction from the regulated side, because it is the mistake I would have made myself. The cheap tier is not automatically outside the retention clock. A truncated or hashed identifier is pseudonymised, not anonymised, under UK GDPR, so it remains personal data and inherits whatever clock applies to the record it points at. Shape and cardinality are usually safer than a hash, but not by definition. Distinct count of one per payer, fixed length of eight, always present, is not a statistic about a distribution. It is a description of an account number, and over a small enough payer population it singles people out on its own.

So the test I would run before declaring the cheap tier indefinite is whether it survives a join. Each shape record is harmless alone. Twenty of them, kept forever, keyed to the same payer, is a profile assembled out of things that were individually fine to keep. That is a failure I would expect to find in our own logs rather than in somebody else's, because nobody ever decided to build it.

Which lands the practical version one notch off yours: the cheap tier gets a stated retention too, just a much longer one, and it gets reviewed rather than set once. The SIEM parallel holds on everything except the word indefinitely.

Collapse
 
mickyarun profile image
arun rajkumar •

Follow-up for @max_quimby, @peterbuildssecure and @jon_at_backboardio.

This thread and a thread on a completely different post ended up describing the same object, so I wrote it up rather than leave it sitting in two comment sections. All three of you are quoted.

dev.to/mickyarun/four-people-rebui...

Short version: what got built in this thread is a payment authorisation. Bound to the instruction, verified by the side holding the state, single-use, expiring. Payments arrived at that shape by being robbed for forty years rather than by designing it.

The part I couldn't close is that payments has a natural transaction boundary to hang the expiry off, and an agent halfway through a long task hasn't got one.

Collapse
 
max_quimby profile image
Max Quimby •

The refund example is the best articulation I've seen of why string-level hooks plateau. We run agent pipelines with pre-execution hooks, and they're genuinely good at the force-push class of problem — anything where the danger is written into the command. But every expensive near-miss we've had looked exactly like your second refund call: schema-valid, well-formed, reasonable-sounding arguments, and the thing that made it wrong lived in state the hook couldn't see.

What's worked better for us is pushing invariants down into the resource layer instead of the call layer: idempotency keys so a duplicate refund is structurally impossible, per-target baselines so an amount 1000x the historical median for that account needs a second signal, and reconciliation jobs that compare intent logs against effects after the fact. None of that is agent-specific, which I think is the point — the agent just removes the human who used to eyeball the number.

Curious whether you'd put the check synchronous (block the call) or async (catch within minutes)? Sync is safer but the latency tax on every benign call adds up.

Collapse
 
mickyarun profile image
arun rajkumar •

Three weeks late, sorry. This deserved an answer the day you wrote it.

Sync where the money is, async everywhere else, and the threshold is the whole design. @jon_at_backboardio made the case below better than I would have: the calls where blocking hurts are high-volume and low-value, and the calls where blocking is fine are the ones you actually want to stop. Those two sets barely overlap, so the latency tax lands on a rounding error of your traffic.

The line I keep coming back to is yours though. The agent just removes the human who used to eyeball the number. What that exposed for us is that the eyeball was never a control. It was a sampling process with unknown coverage that had been getting booked as a control for years. Losing it is bad. Finding out we had been depending on it is worse.

One distinction on pushing invariants into the resource layer: idempotency keys make a duplicate structurally impossible, which is a real win, but they do nothing about wrong-once. The second refund call in the post isn't a duplicate. It is a first-time, well-formed, singular £40,000 mistake, and every key in the world lets it through.

Collapse
 
jon_at_backboardio profile image
Jonathan Murray •

on max's sync vs async question, i think the tax is smaller than it looks because the two axes line up in your favour.

the calls where blocking hurts are high volume and low value. the calls where blocking is fine are the ones you actually want to stop. a £40,000 refund can afford 200ms. forty thousand £4 refunds cannot, and don't need to.

so it isn't sync or async, it's a cheap sync gate whose threshold comes from the target. one cached number per target, refreshed lazily, some multiple of that account's recent movement. under it, execute and reconcile after. over it, block and go get the mandate. you pay full preflight latency on a rounding error of your traffic, and the cached number being slightly stale doesn't matter because it isn't authorising anything, it's only deciding whether to ask.

the case that worries me more than either is your last one, multi-step. five refunds of £8,000 each, every one under every cap, same payee, four minutes. every individual call is boring. the mandate model catches that only if mandates are scoped to a window rather than to an operation, and that's a different shape from what peter described.

"a rubber stamp with extra steps" is the line. showing a human raw args isn't review, it's laundering.

Collapse
 
mickyarun profile image
arun rajkumar •

Very late reply, and this one has been sitting in my head since August.

"It isn't authorising anything, it's only deciding whether to ask" is the sentence that makes the whole design work. Staleness is disqualifying for an authorisation and completely fine for a triage decision, and collapsing those two is why people over-engineer the gate.

On multi-step, you're right that it's a different shape, and payments has a name for it: velocity. Cumulative exposure per payee per window rather than a cap per transaction. It has been standard in card fraud for decades, precisely because five boring transactions are the attack. What nobody tells you is that the check is the easy part. Choosing the window is the hard part, and whatever you choose, someone can straddle it. Four minutes clears an hourly window if you start at 59 minutes past.

And "showing a human raw args isn't review, it's laundering" is the best line anyone has left on anything I've written. It is the failure mode of every approval screen I have ever built.