TL;DR
I run a fully autonomous implementation system built on Claude Code that ships code while I sleep. The thing that made it usable wasn't a smarter model, it was a three-tier permission model plus an approval queue so the agent only pings me for actions that are destructive or outward-facing. Interruptions went from ~40 a day to ~3, and I stopped rubber-stamping things I didn't read. Here's how I built it, with the code that matters. π
The Problem
When I first let my agent run unattended, I had two settings: ask for everything or ask for nothing.
"Ask for everything" was a disaster. Every npm test, every file edit, every git commit popped a confirmation. After a day I was approving 40+ prompts and reading none of them. That's worse than no gate at all, because it feels safe.
"Ask for nothing" was the opposite disaster. It was fast right up until the agent force-pushed over a branch a teammate was working on, and another time it replied to a GitHub issue with a half-baked answer under my name. Nobody died, but I lost a weekend and some trust.
The insight that took me embarrassingly long to reach: not all actions carry the same risk, and the risk is mostly knowable before you run the action. Deleting a temp file and deleting a database are both rm-shaped, but only one of them needs a human.
So the constraint became: the agent must be able to run for hours without me, but must never do something irreversible or public without a human saying yes. And the "yes" had to be async, because I'm not sitting there watching.
How I Solved It
Step 1: Classify every action into three tiers
I settled on three tiers. Two felt too coarse, four felt like bureaucracy.
| Tier | Meaning | Examples | Behavior |
|---|---|---|---|
| Green | Reversible, local | edit files, run tests, commit to a feature branch, read logs | Auto-approve, log it |
| Yellow | Reversible but expensive or noisy | install a new dependency, run a migration on a dev DB, open a PR | Auto-approve, but batch into a daily digest I review |
| Red | Irreversible or outward-facing | force push, delete branch, post a comment on someone else's issue, send a message, touch prod config, spend money | Stop and enqueue. Wait for a human. |
The rule I use to place something in Red is dead simple: "If this goes wrong, can I undo it alone in under five minutes?" If no, it's Red. If it involves another human seeing it (a comment, an email, a Slack message, a public post), it's Red regardless.
Step 2: The risk classifier is boring on purpose
I initially wanted the LLM to classify its own actions. Bad idea. Asking the agent "is this dangerous?" is asking the fox to rate the henhouse. It's not malicious, it's just optimistic when it's mid-task and wants to finish.
So the classifier is a deterministic function that runs outside the model, on the raw tool call. Rules first, model second, and the model only gets to escalate risk, never lower it.
# risk.py β deterministic tier assignment for tool calls
import re
from enum import IntEnum
class Tier(IntEnum):
GREEN = 0
YELLOW = 1
RED = 2
RED_PATTERNS = [
r"\bgit\s+push\b.*(--force|-f\b)",
r"\bgit\s+(branch|push)\b.*(-D|--delete|:refs/)",
r"\brm\s+-rf?\s+(/|~|\$HOME)",
r"\b(DROP|TRUNCATE)\s+(TABLE|DATABASE)\b",
r"\bkubectl\s+(delete|apply).*prod",
r"\bterraform\s+(apply|destroy)\b",
]
YELLOW_PATTERNS = [
r"\b(npm|pnpm|yarn)\s+(add|install)\s+\S",
r"\bpip\s+install\b",
r"\bgh\s+pr\s+create\b",
r"\balembic\s+upgrade\b",
]
# Any tool that talks to a human or the public is Red, full stop.
OUTWARD_TOOLS = {"slack.send", "github.comment", "email.send", "publish"}
def classify(tool_name: str, args: dict) -> Tier:
if tool_name in OUTWARD_TOOLS:
return Tier.RED
cmd = args.get("command", "") if tool_name == "bash" else ""
for pat in RED_PATTERNS:
if re.search(pat, cmd, re.IGNORECASE):
return Tier.RED
for pat in YELLOW_PATTERNS:
if re.search(pat, cmd, re.IGNORECASE):
return Tier.YELLOW
return Tier.GREEN
Is this list complete? No. It's a floor. Anything the regexes miss falls through to Green, which sounds scary, but remember Green actions are by definition reversible in my environment: the agent works in a git worktree on a throwaway branch, on a dev database I can rebuild in a minute. The blast radius of a missed classification is bounded by the sandbox, not by the classifier.
The agent can also self-escalate. I gave it a request_approval(reason) tool and told it in the system prompt: "If you're unsure whether an action is reversible, call this. Being asked to wait is never a failure." In practice it uses this maybe once a day, and it's been right every time.
Step 3: The approval queue is just a directory of JSON files
I resisted the urge to build anything clever here. The queue is a folder. Each pending request is one JSON file. That's it.
{
"id": "20260929-2214-a3f9",
"created_at": "2026-09-29T22:14:03+09:00",
"task": "Close stale issues older than 180 days",
"tool": "github.comment",
"args": {"issue": 412, "body": "Closing as stale. Reopen if still relevant."},
"tier": "RED",
"reason": "Outward-facing: comment visible to issue author",
"expires_at": "2026-09-30T22:14:03+09:00",
"status": "pending"
}
Why files instead of a database or a message broker?
- I can
lsit. When something looks off at 2 a.m., I want tocatthe request, not open a dashboard. - It survives the agent crashing. The agent process can die and restart, and the queue is still there.
- It's trivially syncable. I mirror the folder to my phone with a file sync app, so approving from the couch is editing one field.
The approval action is literally changing "status": "pending" to "approved" or "rejected". A tiny watcher loop polls the folder and resumes the agent's task with the verdict.
sequenceDiagram
participant A as Agent
participant C as Classifier
participant Q as Queue (folder)
participant H as Human
A->>C: tool call (github.comment)
C-->>A: RED
A->>Q: write request.json (pending)
A->>A: park task, pick next Green task
H->>Q: read, set status=approved
Q-->>A: watcher sees change
A->>A: resume parked task, execute
Step 4: Don't block, park
This was the second-biggest win after tiering. Early on, a Red request would block the whole agent. It sat idle for six hours because I was asleep, over one GitHub comment.
Now the agent parks the blocked task, saves its state to disk (which branch it's on, what it's done, what it was about to do), and moves on to the next task in its list. When the approval lands, it reloads that state and continues. Most nights it finishes three or four other things while one request waits for me.
The state file is intentionally tiny: task ID, branch name, a one-paragraph summary of progress, and the pending tool call. Enough to resume, not so much that it goes stale. I learned from a previous war story that giant state dumps rot fast.
Step 5: Timeouts and defaults
Every queued request has an expires_at. If I don't answer within 24 hours, it's auto-rejected, not auto-approved. The task gets marked "needs human" and the agent writes a note explaining what it wanted to do and why.
Fail-closed by default felt obvious in hindsight, but my first version auto-approved after a timeout "so the pipeline wouldn't stall." One week later I found the agent had closed 30 issues while I was on a trip, and nobody had said yes. Fail closed. Always.
Step 6: The daily digest for Yellow
Yellow actions run immediately but get appended to a Markdown digest I read with my morning coffee. Each line is one action: what, why, and a link to the diff or PR.
## Yellow actions β 2026-09-29
- 09:14 `pnpm add zod@3.23.8` β needed for input validation in the webhook handler (PR #88)
- 11:40 `gh pr create` β "Fix race in retry loop" (PR #89)
- 14:02 `alembic upgrade head` on dev DB β added index on events.created_at
Reading this takes two minutes and catches things like "why did you add a 40 MB dependency for a date parser." I've reverted maybe one in twenty of these. That ratio tells me the tier is set right: high enough to not interrupt, low enough that I still see it.
Lessons Learned
- Rubber-stamping is worse than no gate. If you're approving 40 things a day, you're approving zero things. Reduce prompts until every one is worth reading.
- Classify actions, not intentions. The classifier runs on the raw tool call with regexes and allow-lists. The model can push risk up but never down. The fox does not rate the henhouse.
- "Outward-facing" is its own axis. Reversibility isn't enough. A comment can be deleted, but the human who got the notification already read it. Anything another person sees is Red.
- Park, don't block. An agent waiting on a human should do other work. Save minimal resume state and move on.
- Fail closed on timeout. A rejected-by-default request costs you one task. An approved-by-default request costs you trust. These are not symmetric.
What's Next
Two things I'm working on:
- Tier learning. I log every Yellow action I revert and every Red action I approve. If a Red pattern is approved 20 times in a row with no incident, I want a suggestion to demote it to Yellow, with me still holding the pen. The goal is for the queue to get quieter over time without getting dumber.
- Explain-your-request. Right now a Red request includes a one-line reason. I want the agent to attach the exact diff or message it's about to send, rendered, so I can approve from my phone without opening a laptop. Most rejections happen because I couldn't see the payload, not because the action was wrong.
I'm on Claude Code as of late September 2026, with the system running in Python 3.13 on macOS. The tiering code above is deliberately framework-free so you can drop it in front of any agent loop.
Wrap-up
If you're running an AI coding agent unattended and it's either annoying you or scaring you, the fix is probably not a better model. It's a better contract about when it's allowed to act alone. Three tiers, a folder of JSON files, and fail-closed timeouts got me from "babysitting" to "checking a digest with coffee."
If this was useful, follow me here on Dev.to. I write build logs like this from running an autonomous coding system full-time, and the next one is about how the tier-learning experiment goes. And if you've solved the "when should the agent ask" problem differently, tell me in the comments. I'd genuinely like to steal your approach. π‘
Top comments (0)