DEV Community

Your agent isn't reckless. It just can't see the blast radius.

Rabih Jabr on August 20, 2026

I've been running Claude Code as a daily driver for about three months now. It writes Ansible I'd have taken a week to write. It reads a codebase f...
Collapse
 
joinwell52 profile image
joinwell52

The useful part is making the denial explain what the agent may do next. I would still resolve destructive paths before matching the command: $BUILD_DIR can be present in the text and expand to an empty or unexpected location at runtime. Checking the resolved target against an allowed root catches a different failure class than pattern matching.

Collapse
 
rabih_jabr_29 profile image
Rabih Jabr

Hi @joinwell52
Correct, and it's a gap rather than a difference of opinion. My guard is a lint on the text, so this passes: rm -rf "${BUILD_DIR:?}"

The :? means it won't expand to empty, but if BUILD_DIR=/ it deletes exactly what you'd expect. Different failure class, completely uncaught. Verified just now.

The one thing that makes resolution awkward: a PreToolUse hook doesn't share the shell's variable state. Each Bash call can be its own process, and the hook sees the command string plus cwd, not whatever the agent exported three commands ago. So resolving $BUILD_DIR reliably isn't available to me.

What is available is the allowed-root check on literal paths, and on the resolved value where the variable happens to be in the hook's own environment. That's a strictly better primitive than what I have. If you want to write it, it's one file and I'll merge it.

I have created an issue if you would like to get your hands dirty :D
github.com/RabihJabr29/claude-guar...

Collapse
 
joinwell52 profile image
joinwell52

You’re right that a PreToolUse hook cannot recover variable state from an earlier shell, so I wouldn’t try to guess it. For a destructive command, treat an unresolved variable as non-provable and deny it with a specific remediation: expand the path in the same command, or pass the resolved target as structured input. Then canonicalize that value, reject empty paths and the filesystem root, and require it to remain under an allowed root. Literal paths can go through the same final check.

Collapse
 
kevinbai profile image
kevinbai

Blast radius visibility is the missing layer between "agent can do X" and "agent should do X". Making side effects explicit before execution is a practical governance pattern, not just a safety guardrail.

Collapse
 
mads_hansen_27b33ebfee4c9 profile image
Mads Hansen

This is the right place to intervene, but I’d treat command hooks as one layer, not the security boundary. Shell text has too many equivalent forms: aliases, functions, sh -c, Python subprocess calls, encoded payloads, scripts written then executed, alternate Git clients, and commands split across tool calls. Regex-quality guards tend to become a bypass catalog.

Where possible, enforce the invariant at the owning system: protected branches and server-side force-push denial; scoped credentials; filesystem sandbox and path allowlists; read-only secret mounts; DB roles; egress controls; CI rules that reject skipped tests or modified historical migrations. The hook then supplies fast feedback before the authoritative control rejects it.

I’d also bind every denial to structured policy ID/version, normalized action, resolved cwd/repository, matched evidence, and a non-executable remediation template. Log both denials and allowed near-matches so rule drift and false positives are measurable.

Finally, test compositions, not only single commands: write secret → rename → stage; generate script → execute; many individually bounded deletes; symlink escape; environment-variable indirection. Blast radius often appears across a sequence even when every individual tool call looks acceptable.

Collapse
 
rabih_jabr_29 profile image
Rabih Jabr

Hello @mads_hansen_27b33ebfee4c9
You're right, and I went and checked rather than argue. Every form you listed gets through...

Five for five. The first one fails for a genuinely stupid reason — my "is this a git push" check requires whitespace before git, and -c "git puts a quote there.

So: agreed, this is a layer, not a boundary. The README says "not a sandbox" but the post doesn't say it nearly clearly enough, and that's my fault for writing the fun version. The invariant belongs at the owning system, server-side branch protection, DB roles, scoped credentials — and this is fast feedback in front of that, at best.

The suggestion I'm actually going to steal is logging allowed near-matches. Right now the dispatcher is silent on allow and I have zero visibility into false positives, which is embarrassing given precision is the thing I claim to optimise for. I'm measuring nothing.

Composition testing is the harder one, and it exposes a real limit in the design rather than a bug. Every example in a guard is a single tool call, so the format literally cannot express "write secret → rename → stage." Sequence-aware guards would need state across calls, and I don't have a good answer for that yet. Genuinely useful to have it named.

I have created those 2 issues:
github.com/RabihJabr29/claude-guar...
github.com/RabihJabr29/claude-guar...

I would greatly appreciate any contributions <3

Collapse
 
hannune profile image
Tae Kim

The credential read section stopped me cold. I've had the same invisible leak in entity resolution pipelines where an agent reads candidate pairs into context for scoring and the whole set lands unlogged in transcript. Every write was audited; it never occurred to me to treat reads the same way. The grep workaround is obvious now that I see it spelled out.

Collapse
 
alexshev profile image
Alex Shev

What I like here is the focus on the mechanism behind “Your agent isn't reckless. It just can't see the blast radius..” A useful follow-up would be one concrete before/after metric: what changed in latency, error rate, review time, or operator workload once the approach was applied?

Collapse
 
mk023 profile image
Marco

Really strong work, Rabih. 👏

The idea that stuck with me most is that the agent can make a locally correct decision while being completely blind to the non-local consequence. “It could see the command. It could not see the crater” is an excellent way to frame the problem.

I also really like the design choice of making the guards precise rather than aggressive. The distinction between blocking a mistake and blocking ambiguity is especially important for agentic systems.

And embedding the blocked/allowed examples directly into each guard is a great testing pattern. The guard doesn't just define what it prevents — it also defines what it deliberately allows. That makes the security boundary executable rather than just documented. 🔐

Thirteen guardrails backed by real scars is a much better starting point than fifty rules based on hypothetical problems. Great work. 🚀

Collapse
 
rabih_jabr_29 profile image
Rabih Jabr

Hello @mk023
Thank you, that's generous.

One honest amendment to the part you liked most, though. @mads_hansen pointed out a real limit in the examples-as-tests pattern: every example is a single tool call, so the format can't express a dangerous sequence. Write a secret, rename the file, stage it, each step passes on its own, and there's no way to write that down as an example.

So it makes the boundary executable, but only for boundaries that fit in one command. That turns out to be a meaningful chunk of the ones that matter, and I didn't notice the ceiling until someone else pointed at it.

The repo is public and open to contributions ;)

Collapse
 
mickyarun profile image
arun rajkumar

The force-push example works because the command is legible. In payments we mostly get the opposite, which is why I land where Mads does on hooks not being the boundary.

A refund call for forty pounds and a refund call for forty thousand are the same shape. Same endpoint, same argument names, same everything a pre-execution guard can read off the text. The blast radius is not in the command, it is in the state of the thing the command is about to touch, and the only way to know it is to ask the system that owns that state.

So the check has to sit where the effect gets realised rather than where the intent gets expressed. Ours ended up server side at the point of authorisation, because that is the only place that knows what this actor is allowed to move and how much of it has already moved today. Everything in front of that is advisory, useful for catching typos and not much else.

Does your setup have any way to ask the target what a command would cost before running it, or is the radius always inferred from the text?

Collapse
 
mnemehq profile image
Theo Valmis

The napkin-sized list is the right instinct and it has a scaling problem of its own: it lives in your head and your CLAUDE.md, which means it only protects you, on this machine, until you forget to port it somewhere else. The failure mode you're describing, locally correct, non-local consequence, is exactly what we're trying to catch at Mneme before generation, not by reviewing the diff but by checking it against the same short list you just wrote, mechanically, every time.

Collapse
 
jon_at_backboardio profile image
Jonathan Murray

the thing you named at the end, that every example is a single tool call and the format can't express write secret then rename then stage, might be closer than it looks.

you don't need session state for that. you need one append-only file per session holding the paths any guard has seen, written on every PreToolUse regardless of verdict. then the git add guard asks one more question: is anything in this stage set a path we touched earlier under a different name. that's a lookup, not state reconstruction, and it fails open like everything else you wrote if the file isn't there.

wouldn't stop a determined person. would catch the actual sequence you're worried about, which is the agent doing three individually reasonable things in a row.

fits your rules too. invisible while you work, only ever votes no, no config.

on rajkumar's question about asking the target what a command would cost, i think the honest answer is that for git you can and for postgres you can't, and the asymmetry is the interesting part. git will tell you what a push would do with a dry run. no database will tell you what an ALTER is going to cost without you already knowing the shape of the table.

Collapse
 
richard_smith_154156d471ef profile image
Richard Smith

Reminds me of the old sysadmin mantra: "a command that works isn't the same as a command that does what you want." Same idea, just one layer higher up the stack now.