DEV Community

Cover image for Curing Agent Approval Fatigue: Using a Local LLM Gatekeeper for Safe Shell Execution
Cason Adams
Cason Adams

Posted on

Curing Agent Approval Fatigue: Using a Local LLM Gatekeeper for Safe Shell Execution

If you use autonomous coding agents like oh-my-pi (omp), Claude Code, or Cursor, you have almost certainly encountered the Agent Approval Dilemma.

You have two choices, and both of them are terrible:

  1. YOLO Mode (Auto-Approve Everything): You set your permissions to unrestricted execution. The agent blazes through tasks with zero interruptions—right up until a subtle hallucination or edge case runs git reset --hard, drops your database container, wipes a directory with an unchecked glob, or pushes unfinished code to main.
  2. Paranoid Mode (Prompt on Every Command): You configure the agent to ask before executing any bash command. Within 10 minutes, you’ve manually approved 45 routine commands: cat package.json, git status, bun test, ls -la, mkdir -p dist. Your brain shuts off. You experience acute Approval Fatigue, start blindly mashing Enter, and when a truly dangerous command finally slips through, you approve it anyway.

Security that relies on tired humans repeatedly clicking "OK" is security theater.

To solve this, we built guard—an open-source plugin for oh-my-pi that acts as an intelligent, hybrid gatekeeper for the shell tool. It pairs sub-millisecond deterministic heuristics with a free, private, local LLM (via Ollama, Apple MLX, or local inference) to evaluate commands before they run.

Routine development commands run instantly in the background. Genuinely risky commands get halted with an actionable, plain-English security audit.

Here is how it works under the hood.


The Architecture: Hybrid Defense in Depth

Relying solely on an LLM to evaluate safety is slow and vulnerable to prompt injection. Relying solely on regex is brittle and breaks on complex pipelines or novel commands.

guard uses a two-tiered hybrid approach:

                  Agent issues bash command
                             │
                             ▼
             ┌───────────────────────────────┐
             │ Tier 1: Regex Guard (0.1ms)   │
             │   Traps catastrophic wipes,   │
             │   fork bombs, raw disk writes │
             └───────────────┬───────────────┘
                             │
            Critical Danger? │ No
            ┌────────────────┴───────────────┐
            │                                │
            ▼ (Yes)                          ▼
 ┌─────────────────────┐      ┌─────────────────────────────┐
 │ Prompt User (Block) │      │ Tier 2: Local Model Audit   │
 └─────────────────────┘      │  Evaluates intent, paths,   │
                              │  cloud side effects         │
                              └──────────────┬──────────────┘
                                             │
                             Safe / Read-Only?
                             ┌───────────────┴───────────────┐
                             │                               │
                             ▼ (Yes)                         ▼ (No / Ambiguous)
               ┌──────────────────────────┐    ┌───────────────────────────┐
               │ Execute Silently (0 int) │    │ Explain & Prompt Human    │
               └──────────────────────────┘    └───────────────────────────┘
Enter fullscreen mode Exit fullscreen mode

1. Tier 1: Sub-Millisecond Deterministic Trap

Before spending any compute on an LLM, the command is tested against an immediate pattern trap:

export const CRITICAL_DANGER_REGEX =
  /(\brm\s+-[a-zA-Z]*r[a-zA-Z]*f?\s+([/~]|\.\.|\*)|:\(\)\s*\{\s*:\|:&\s*\};:|\bmkfs\b|\bdd\s+if=|>+\s*\/dev\/sd|\bgit\s+reset\s+--hard\b)/i;
Enter fullscreen mode Exit fullscreen mode

This instantly intercepts catastrophic commands:

  • rm -rf /, rm -rf ~, or unanchored wildcard deletions (rm -rf *).
  • Classic shell fork bombs (:(){ :|:& };:).
  • Low-level disk/partition formatting (mkfs, dd if=, raw writes to /dev/sd*).
  • Destructive git history erasures (git reset --hard).

If matched, execution is halted immediately. A high-contrast warning pops up explaining the exact nature of the threat (e.g. "Critical destructive filesystem wipe detected"), completely bypassing the LLM.

2. Tier 2: The Local Model Evaluator

If the command isn't an outright disaster, it needs context-aware evaluation. Is rm -rf dist/ safe? Yes, that's just cleaning build output. Is rm -rf ../config safe? Probably not.

Rather than burning expensive cloud tokens or leaking your bash history to external APIs, guard resolves the @guard or @judge model role configured in your ~/.omp/agent/config.yml.

This is specifically designed to use small, blazing-fast local models running on your own machine:

  • ollama/qwen2.5-coder:7b
  • ollama/llama3.2:3b
  • local-inference (MLX on Apple Silicon or llama.cpp)

Because these run locally, latency is negligible (often 100-250ms), cost is zero, and proprietary project context never leaves your computer.


Inside the Code: How Guard Intercepts Tool Calls

In oh-my-pi, plugins hook into the execution lifecycle via extension events. guard listens to tool_call:

export default function registerBashGuard(pi: PiExtensionAPI) {
  pi.on("tool_call", async (event: ToolCallEvent, ctx?: ExtensionContext) => {
    // Only intercept bash commands
    if (event.toolName !== "bash") return;

    const command =
      typeof event.input?.command === "string" ? event.input.command.trim() : undefined;
    if (!command) return;

    // 1. Fast-path deterministic trap
    if (CRITICAL_DANGER_REGEX.test(command)) {
      return promptUser(ctx, command, getCriticalDangerAudit(command));
    }

    // 2. Resolve local guard / judge model
    const guard = await resolveGuardModel(ctx);
    if ("block" in guard) {
      return promptUser(ctx, command, guard.reason);
    }

    // 3. Classify with structured prompt
    const verdict = await evaluateCommandSafety(guard.model, guard.apiKey, command);

    // 4. Safe commands proceed with zero interruption!
    if (!verdict.safe) {
      return promptUser(ctx, command, {
        action: verdict.action,
        reason: verdict.reason || "Action modifies state, cloud resources, or data.",
      });
    }
  });
}
Enter fullscreen mode Exit fullscreen mode

Handling Reasoning Models (<think> Stripping)

Modern local models (like DeepSeek-R1 or reasoning-tuned Qwen models) emit thought traces wrapped in <think>...</think>. If you simply JSON.parse their output, your parser will crash on the raw text preamble.

guard cleans the reasoning stream before parsing the decision payload:

export function parseGuardOutput(text: string): GuardVerdict {
  // Strip thought blocks from reasoning models
  const clean = text
    .replace(/<(?:think|thought|thinking)>[\s\S]*?<\/(?:think|thought|thinking)>/gi, "")
    .trim();

  const start = clean.indexOf("{");
  const end = clean.lastIndexOf("}");
  if (start !== -1 && end > start) {
    try {
      const parsed = JSON.parse(clean.slice(start, end + 1));
      if (typeof parsed.safe === "boolean") {
        return {
          safe: parsed.safe,
          action: parsed.action,
          reason: parsed.reason,
        };
      }
    } catch {
      // Heuristic fallback
    }
  }

  // Fail safe if ambiguous
  return { safe: false, reason: "Ambiguous response from guard model." };
}
Enter fullscreen mode Exit fullscreen mode

Defining "Safe": Principles of the Guard Prompt

The prompt fed to the local model enforces strict boundaries:

  1. Safe Categories (Run automatically):

    • Build & Test: cargo check, go test ./..., bun run test, tsc --noEmit, pytest.
    • Inspection & Diagnostics: cat, grep, git status, git diff, docker ps, lsof, jq.
    • Local Workspace File Operations: mkdir build/, cp -r src/ tmp/, rm -rf .cache.
  2. Unsafe Categories (Require human confirmation):

    • Destructive Deletion: Recursive removals outside project boundaries or untracked mass deletes.
    • Remote Publishing & Infrastructure: git push --force, terraform apply, kubectl delete pod, aws s3 rm, npm publish.
    • Privilege Escalation: Any use of sudo, su, or edits to /etc/ or system roots.

When a command is deemed unsafe, instead of a cryptic warning, guard renders an interactive UI prompt:

┌── Security Guard ──────────────────────────────────────────────┐
│ Command : git push origin main --force                         │
│ Action  : Force pushes local history to remote repository      │
│ Risk    : Overwrites remote branch history and may discard     │
│           collaborator commits.                                │
│                                                                │
│ [Enter] Approve    [Esc] Reject                                │
└────────────────────────────────────────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode

You are never guessing why the command was flagged. The risk is contextual, clear, and takes half a second to verify.


Setting It Up in 2 Minutes

If you're using oh-my-pi:

1. Add the Marketplace & Install Guard

In your terminal:

# Add the official plugins marketplace
omp plugin marketplace add casonadams/omp-plugins

# Install the guard plugin
omp plugin install guard@casonadams-plugins
Enter fullscreen mode Exit fullscreen mode

2. Configure Your Settings: Delegate Approval to Guard

Here is the crucial part that people often miss: You must disable oh-my-pi's native, indiscriminate bash prompting.

By default, omp's built-in engine treats bash as an exec-tier tool and prompts you before running any shell command. If you leave native prompting on, omp will still interrupt you for ls and git status, defeating the whole purpose of the plugin!

To let guard take over as the sole, intelligent gatekeeper, update your ~/.omp/agent/config.yml:

tools:
  approvalMode: yolo # Delegates command gating to the guard plugin

modelRoles:
  guard: ollama/qwen2.5-coder:7b # Recommended local model (falls back to judge)
Enter fullscreen mode Exit fullscreen mode

Why approvalMode: yolo?

It sounds counter-intuitive at first: "Why would I set YOLO mode if I care about safety?!"

Here's why:

  1. tools.approvalMode: yolo tells omp's native engine not to prompt blindly on every single tool execution.
  2. The guard plugin hooks directly into pi.on("tool_call") before the bash tool ever executes.
  3. If guard classifies the command as safe, it permits execution seamlessly.
  4. If guard detects risk, it triggers ctx.ui.askDialog—putting the interactive approval dialog back in front of you with the exact risk explanation and command preview.

(Alternative: If you want omp to prompt for other tools like browser or MCP actions, but let guard handle bash, you can target just bash with tools.approval.bash: allow while setting tools.approvalMode: write.)

Recommended Local Models

For local, zero-cost, air-gapped inspection, we recommend:

  • ollama/qwen2.5-coder:7b (Exceptional bash/DevOps command comprehension)
  • ollama/llama3.2:3b (Ultra-lightweight, near-instant inference)
  • Apple MLX / Local Inference (Native silicon acceleration via local-inference)

3. The Result

The next time your agent goes to work:

  • cat src/index.ts $\rightarrow$ passes instantly with zero prompts.
  • bun test $\rightarrow$ passes instantly with zero prompts.
  • git status $\rightarrow$ passes instantly with zero prompts.
  • git push origin main --force $\rightarrow$ HALTS:
┌── Security Guard ──────────────────────────────────────────────┐
│ Command : git push origin main --force                         │
│ Action  : Force pushes local history to remote repository      │
│ Risk    : Overwrites remote branch history and may discard     │
│           collaborator commits.                                │
│                                                                │
│ [Enter] Proceed    [Esc] Cancel    [Tab] Feedback              │
└────────────────────────────────────────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode

Summary

Agentic coding shouldn't require choosing between reckless abandonment and cognitive exhaustion.

By combining deterministic pattern matching with a fast, local AI classifier, guard eliminates the white noise of routine confirmations while preserving a rock-solid safety net for your filesystem and infrastructure.

Top comments (0)