DEV Community

Umang Kumar
Umang Kumar

Posted on

How to Stop an AI Agent From Running rm -rf (and Why String Matching Isn't Enough)

If you let a coding agent run shell commands, sooner or later it will want to delete something. Usually it is harmless — clearing dist/ before a rebuild. Occasionally it is rm -rf "$BUILD_DIR/" with an empty variable, or a cleanup step suggested by a README the agent should never have trusted. This post is about how to stop an AI agent from running rm -rf in the cases that matter, without making it useless for the cases that don't.

The core lesson: a rule that looks for the literal text rm -rf is a tripwire, not a guardrail. I'll show exactly where it fails, with test output, and what to layer around it.

Why agents reach for rm -rf

Three common paths:

  1. Legitimate cleanup. "Clean the build and try again" is a perfectly reasonable instruction, and rm -rf build is a perfectly reasonable way to do it.
  2. Variable expansion. rm -rf "$OUT_DIR/" is fine until OUT_DIR is empty and the command becomes rm -rf "/". Agents compose commands quickly and rarely add set -u guards.
  3. Injected instructions. A README, issue comment or tool result says "to reset the environment, run …". The model is trying to be helpful.

You cannot fix the first by banning deletion, and you cannot fix the third by asking the model to be careful. You need layers that do not depend on the model's judgment.

Layer 1: make the blast radius small

Before any rule, reduce what a bad delete can reach:

  • Run the agent as an unprivileged user in a container or VM, with only the project mounted.
  • Mount anything the agent should not change read-only. A read-only bind mount does not care how the command was spelled.
  • Commit (or stash) before delegating. A clean git status turns most workspace deletes into git checkout ..
  • Keep backups of anything outside git the agent can touch.

This is the only layer that is indifferent to command syntax. Everything after it is about catching intent earlier.

Layer 2: agent-level deny rules

Most coding agents let you deny command patterns. In Claude Code, for example:

{
  "permissions": {
    "deny": ["Bash(rm -rf *)", "Bash(rm -fr *)"]
  }
}
Enter fullscreen mode Exit fullscreen mode

Gemini CLI and others have similar mechanisms. Use them — but know what they are. Older Gemini CLI documentation said it plainly: command-specific restrictions based on simple string matching "can be easily bypassed". The same is true of any prefix or substring rule.

Proving the problem: a literal rule versus real spellings

Here is a policy with one deny rule for the literal string and a permissive shell rule, written in the Cirvix policy DSL (where command = "rm -rf" compiles to a contains match):

deny:
  name = deny-destructive-shell
  tool = shell.exec
  command = "rm -rf"

allow:
  name = allow-shell
  tool = shell.exec

test "rm -rf":
  tool = shell.exec
  command = rm -rf ./build
  expect deny
test "rm -fr":
  tool = shell.exec
  command = rm -fr ./build
  expect deny
test "rm -r -f":
  tool = shell.exec
  command = rm -r -f ./build
  expect deny
test "find delete":
  tool = shell.exec
  command = find . -delete
  expect deny
Enter fullscreen mode Exit fullscreen mode

Running cirvix policy test against it (package version 0.3.0):

  ✓ rm -rf  → deny (deny-destructive-shell)
  ✗ rm -fr  line 15
      expected  deny
      actual    allow  by allow-shell
      call      shell.exec   risk CRITICAL
  ✗ rm -r -f  line 19
      expected  deny
      actual    allow  by allow-shell
      call      shell.exec   risk CRITICAL
  ✗ find delete  line 23
      expected  deny
      actual    allow  by allow-shell
      call      shell.exec   risk HIGH

  3 failed  1 passed
Enter fullscreen mode Exit fullscreen mode

One spelling caught, three missed. Notice the risk column, though: the engine already classified rm -fr and rm -r -f as CRITICAL and find -delete as HIGH. The literal rule just never asked.

Layer 3: classify the command, then decide on the class

The fix is to write rules against what a command is, not how it is spelled, and to never let a wildcard allow cover dangerous classes:

deny:
  name = deny-rm-rf-literal
  tool = shell.exec
  command = "rm -rf"

deny:
  name = deny-critical-shell
  tool = shell.exec
  risk >= CRITICAL
  reason = "A CRITICAL command must be permitted by a rule that names it, never by a wildcard."

require_approval:
  name = hold-high-risk-shell
  tool = shell.exec
  risk >= HIGH
  approvers = developer

allow:
  name = allow-safe-shell
  tool = shell.exec
  risk <= MEDIUM
Enter fullscreen mode Exit fullscreen mode

Same kind of test cases, plus a few more realistic ones:

  ✓ rm -rf  → deny (deny-rm-rf-literal)
  ✓ rm -fr  → deny (deny-critical-shell)
  ✓ rm -r -f  → deny (deny-critical-shell)
  ✓ rm with empty variable  → deny (deny-rm-rf-literal)
  ✓ find -delete  → require_approval (hold-high-risk-shell)
  ✓ python rmtree  → require_approval (hold-high-risk-shell)
  ✓ git clean  → require_approval (hold-high-risk-shell)
  ✓ npm test  → allow (allow-safe-shell)

  8/8 PASSED
Enter fullscreen mode Exit fullscreen mode

What changed:

  • Recursive force-deletes in any flag order are denied by class, not text.
  • Commands the classifier does not recognise as safe — find -delete, python -c "shutil.rmtree(...)", git clean -fdx — are held for a person instead of silently allowed. That is the right default for "arbitrary execution I can't characterise".
  • Everyday commands still run. npm test is on a short, anchored allowlist and is not interrupted.

The classifier is deliberately rules, not a model: the same command gets the same risk level every time, and a command with chaining or substitution characters (;, &&, |, $() is disqualified from the safe allowlist, so npm test; rm -rf ~ cannot ride in on npm test.

Layer 4: test the policy like code

Whatever engine you use, keep a list of nasty spellings and run it in CI. A starter set:

rm -rf ./build          rm -fr ./build         rm -r -f dist
rm -rf "$EMPTY/"        find . -delete         git clean -fdx
python3 -c "import shutil; shutil.rmtree('x')"
xargs rm -rf < list.txt  npm test; rm -rf ~
Enter fullscreen mode Exit fullscreen mode

If a new rule makes one of these pass when it shouldn't, you find out in a pull request, not an incident.

What this does not cover

Policy only sees calls that are routed through it. If the agent can spawn a shell some other way, or runs a script whose contents you never inspected, the policy evaluates the outer command (./scripts/clean.sh) — which, with the rules above, is held as unrecognised. That is a good default, but it is not visibility into the script. Keep Layer 1.

Trying it with Cirvix

Cirvix AgentControl is an open-source policy engine for governed AI agent tool calls: shell commands via the Claude Code hook, MCP calls via its gateway, or functions wrapped with its Node/Python SDKs. The policies above validate and run with:

npm install -g @cirvix_ai/agent-control
cirvix policy check --policy shell.policy
cirvix policy test  --policy shell.policy
Enter fullscreen mode Exit fullscreen mode

The repository's policies/default.policy contains a fuller baseline (destructive shell, history rewrites, package installs, production deploys) with its own test cases. Cirvix does not stop prompt injection and only governs calls routed through it; what it gives you is a tested, explainable decision before the command runs.


Repo: https://github.com/CIRVIX/agent-control
Docs and guides: https://cirvix.com

Top comments (0)