If you let a coding agent run shell commands, sooner or later it will want to delete something. Usually it is harmless — clearing dist/ before a rebuild. Occasionally it is rm -rf "$BUILD_DIR/" with an empty variable, or a cleanup step suggested by a README the agent should never have trusted. This post is about how to stop an AI agent from running rm -rf in the cases that matter, without making it useless for the cases that don't.
The core lesson: a rule that looks for the literal text rm -rf is a tripwire, not a guardrail. I'll show exactly where it fails, with test output, and what to layer around it.
Why agents reach for rm -rf
Three common paths:
-
Legitimate cleanup. "Clean the build and try again" is a perfectly reasonable instruction, and
rm -rf buildis a perfectly reasonable way to do it. -
Variable expansion.
rm -rf "$OUT_DIR/"is fine untilOUT_DIRis empty and the command becomesrm -rf "/". Agents compose commands quickly and rarely addset -uguards. - Injected instructions. A README, issue comment or tool result says "to reset the environment, run …". The model is trying to be helpful.
You cannot fix the first by banning deletion, and you cannot fix the third by asking the model to be careful. You need layers that do not depend on the model's judgment.
Layer 1: make the blast radius small
Before any rule, reduce what a bad delete can reach:
- Run the agent as an unprivileged user in a container or VM, with only the project mounted.
- Mount anything the agent should not change read-only. A read-only bind mount does not care how the command was spelled.
-
Commit (or stash) before delegating. A clean
git statusturns most workspace deletes intogit checkout .. - Keep backups of anything outside git the agent can touch.
This is the only layer that is indifferent to command syntax. Everything after it is about catching intent earlier.
Layer 2: agent-level deny rules
Most coding agents let you deny command patterns. In Claude Code, for example:
{
"permissions": {
"deny": ["Bash(rm -rf *)", "Bash(rm -fr *)"]
}
}
Gemini CLI and others have similar mechanisms. Use them — but know what they are. Older Gemini CLI documentation said it plainly: command-specific restrictions based on simple string matching "can be easily bypassed". The same is true of any prefix or substring rule.
Proving the problem: a literal rule versus real spellings
Here is a policy with one deny rule for the literal string and a permissive shell rule, written in the Cirvix policy DSL (where command = "rm -rf" compiles to a contains match):
deny:
name = deny-destructive-shell
tool = shell.exec
command = "rm -rf"
allow:
name = allow-shell
tool = shell.exec
test "rm -rf":
tool = shell.exec
command = rm -rf ./build
expect deny
test "rm -fr":
tool = shell.exec
command = rm -fr ./build
expect deny
test "rm -r -f":
tool = shell.exec
command = rm -r -f ./build
expect deny
test "find delete":
tool = shell.exec
command = find . -delete
expect deny
Running cirvix policy test against it (package version 0.3.0):
✓ rm -rf → deny (deny-destructive-shell)
✗ rm -fr line 15
expected deny
actual allow by allow-shell
call shell.exec risk CRITICAL
✗ rm -r -f line 19
expected deny
actual allow by allow-shell
call shell.exec risk CRITICAL
✗ find delete line 23
expected deny
actual allow by allow-shell
call shell.exec risk HIGH
3 failed 1 passed
One spelling caught, three missed. Notice the risk column, though: the engine already classified rm -fr and rm -r -f as CRITICAL and find -delete as HIGH. The literal rule just never asked.
Layer 3: classify the command, then decide on the class
The fix is to write rules against what a command is, not how it is spelled, and to never let a wildcard allow cover dangerous classes:
deny:
name = deny-rm-rf-literal
tool = shell.exec
command = "rm -rf"
deny:
name = deny-critical-shell
tool = shell.exec
risk >= CRITICAL
reason = "A CRITICAL command must be permitted by a rule that names it, never by a wildcard."
require_approval:
name = hold-high-risk-shell
tool = shell.exec
risk >= HIGH
approvers = developer
allow:
name = allow-safe-shell
tool = shell.exec
risk <= MEDIUM
Same kind of test cases, plus a few more realistic ones:
✓ rm -rf → deny (deny-rm-rf-literal)
✓ rm -fr → deny (deny-critical-shell)
✓ rm -r -f → deny (deny-critical-shell)
✓ rm with empty variable → deny (deny-rm-rf-literal)
✓ find -delete → require_approval (hold-high-risk-shell)
✓ python rmtree → require_approval (hold-high-risk-shell)
✓ git clean → require_approval (hold-high-risk-shell)
✓ npm test → allow (allow-safe-shell)
8/8 PASSED
What changed:
- Recursive force-deletes in any flag order are denied by class, not text.
-
Commands the classifier does not recognise as safe —
find -delete,python -c "shutil.rmtree(...)",git clean -fdx— are held for a person instead of silently allowed. That is the right default for "arbitrary execution I can't characterise". -
Everyday commands still run.
npm testis on a short, anchored allowlist and is not interrupted.
The classifier is deliberately rules, not a model: the same command gets the same risk level every time, and a command with chaining or substitution characters (;, &&, |, $() is disqualified from the safe allowlist, so npm test; rm -rf ~ cannot ride in on npm test.
Layer 4: test the policy like code
Whatever engine you use, keep a list of nasty spellings and run it in CI. A starter set:
rm -rf ./build rm -fr ./build rm -r -f dist
rm -rf "$EMPTY/" find . -delete git clean -fdx
python3 -c "import shutil; shutil.rmtree('x')"
xargs rm -rf < list.txt npm test; rm -rf ~
If a new rule makes one of these pass when it shouldn't, you find out in a pull request, not an incident.
What this does not cover
Policy only sees calls that are routed through it. If the agent can spawn a shell some other way, or runs a script whose contents you never inspected, the policy evaluates the outer command (./scripts/clean.sh) — which, with the rules above, is held as unrecognised. That is a good default, but it is not visibility into the script. Keep Layer 1.
Trying it with Cirvix
Cirvix AgentControl is an open-source policy engine for governed AI agent tool calls: shell commands via the Claude Code hook, MCP calls via its gateway, or functions wrapped with its Node/Python SDKs. The policies above validate and run with:
npm install -g @cirvix_ai/agent-control
cirvix policy check --policy shell.policy
cirvix policy test --policy shell.policy
The repository's policies/default.policy contains a fuller baseline (destructive shell, history rewrites, package installs, production deploys) with its own test cases. Cirvix does not stop prompt injection and only governs calls routed through it; what it gives you is a tested, explainable decision before the command runs.
Repo: https://github.com/CIRVIX/agent-control
Docs and guides: https://cirvix.com
Top comments (0)