Your agent's safety check now runs a model call per command
(Correcting the plan I started this post with: I originally titled it "2 of the 4 options on your approval prompt permanently widen permissions." I couldn't source "permanently" for the second option, so I dropped it. What I found while checking was better anyway.)
An approval prompt is a thing you click. You read the command, you decide, you move on, and the whole gate costs you one second of attention and nothing at all on your invoice.
That stopped being true. As of Claude Code v2.1.283, auto mode is the built-in starting permission mode on every plan and provider. There is no prompt in the default path any more. In its place is a background classifier running on Claude Sonnet 5, and for some plans those calls count toward your token usage.
What it costs, in the vendor's own words
From Anthropic's permission-modes documentation:
The classifier runs on Claude Sonnet 5 by default rather than on your
/modelselection.Each check sends a portion of the transcript plus the pending action, adding a round-trip before execution.
On Enterprise plans and on accounts that use the Claude API, Claude Platform on AWS, Amazon Bedrock, Google Cloud's Agent Platform, or Microsoft Foundry, classifier calls count toward your token usage.
Read that third line again, because it is the one that matters for a bill. On a consumer subscription, this is not charging you per keystroke. On an API key or an enterprise agreement, every shell command and every network request now drags a slice of your transcript into a second model before the first one is allowed to finish.
And the overhead is not spread evenly:
Reads and working-directory edits outside protected paths skip the classifier, so the overhead comes mainly from shell commands and network operations.
So the cheap-looking activity, a loop of shell commands, is exactly the expensive one. Our cost work has said for a while that the loop is the bill, not the model. The safety check turns out to be inside that loop now, and it was the one part of the loop nobody budgeted for.
I am not going to print a dollar figure, and you should be suspicious of anyone who does. The docs never say how much of the transcript goes into a check, so any total would be arithmetic I performed on a number I made up. What is documented is the mechanism: a second model, per action, on your meter.
Three things it changes without telling you
This is the part I would want someone to have told me.
One: entering auto mode quietly throws away your broad allow rules. You spent an afternoon building Bash(npm *) so you would stop being interrupted. Then you accept one prompt and switch to auto:
On entering auto mode, broad allow rules that grant arbitrary code execution are dropped: Blanket
Bash(*)orPowerShell(*), Wildcarded interpreters likeBash(python*), Package-manager run commands,Agentallow rules,Monitorallow rules, because Claude Code runs Monitor commands through the shell.Claude Code restores the dropped rules when you leave auto mode.
So your ruleset depends on which mode you are in, it changes on a single click, and nothing in the interface narrates it. Leave auto and they come back. That is a much better design than dropping them, and it is still a ruleset that varies underneath you.
Two: a missing verdict is a silent refusal. The failure modes are not symmetrical:
A blocked action: Claude Code shows a notification and lists the action in
/permissionsunder the Recently denied tab, where you can pressrto retry it with a manual approval.No verdict from the classifier: Claude Code denies the action without the notification or the Recently denied entry.
Blocked because it thought you were in danger, you see it. Denied because the classifier had an off day, you do not. Same outcome for your run, one is recoverable in a keystroke and the other is a mystery you will spend an hour reconstructing. There is a related ceiling: the turn stops after ten responses in a row with no verdict.
Three: boundaries you type in chat can be evicted. This is the one that should worry me most, and it connects straight to why agents forget by tomorrow.
You can tell the agent in conversation that this repo must never deploy to production. The docs are explicit that the classifier honours that:
The classifier treats boundaries you state in the conversation as a block signal. … A boundary stays in force until you lift it in a later message. Claude's own judgment that a condition was met does not lift it.
And then:
Boundaries are not stored as rules. The classifier re-reads them from the transcript on each check, so a boundary can be lost if context compaction removes the message that stated it. For a hard guarantee, add a deny rule instead.
You stated it in English, it is enforced by re-reading English, and a routine housekeeping operation can delete the sentence that made it true. Anthropic's own advice is to write a deny rule instead, and they are right. But note what happened: the soft way to express a constraint is the way that quietly fails, and it fails in the direction of removing a protection you believed you had.
The part that is better than what I warned about
I wrote prompt injection is your agent's tools being the real attack surface five weeks ago, and in a comment last week I told a reader the fix is that the approval screen should carry the provenance of the instruction. A page the agent just fetched should not be able to speak with your authority.
Auto mode attacks that from the other direction, and it works:
Tool results are stripped from those requests, so hostile content in a file or web page can't manipulate the classifier directly.
The gate never sees the attacker's text, so the attacker's text cannot vote. That is a better mechanism than anything I proposed. I was arguing for showing provenance to the human; Anthropic chose to withhold it from the judge instead. Neither is complete, since the classifier is blind to command output and cannot check one deletion against another, but the second one is strictly harder to attack.
So my earlier advice was aimed at the wrong layer. Provenance should be enforced where the decision is made, not merely displayed where the human is asked to agree.
The approval screen is still there. It is just the fallback.
Worth saying plainly, because the docs are unusually honest about this: "Auto mode reduces permission prompts but does not guarantee safety. Use it for tasks where you trust the general direction, not as a replacement for review on sensitive operations."
And the human prompt it replaced is still reachable, and it is still the sharpest tool you have. Which is where the part I got wrong first time comes back in.
When Claude Code shows you a Bash prompt, it lists four options. Here is the second one:
Yes, and don't ask again for: npm test *
You read npm test. You are offered npm test *. The wildcard is the artifact that gets written, and "a * in a Bash rule matches any text, including spaces, so one rule covers a family of commands." One keystroke turns a decision about today's command into a standing permission over the family. If you take that option often enough, the wildcard becomes your real policy and you never composed it.
The saving is also lopsided in a way that is easy to miss. A Bash approval is "Permanently per repository and command." A file modification approval is "Until session end." Every permission you accumulate across sessions is therefore a shell permission, while edit permissions evaporate overnight. Anyone who approved both reasonably concluded that edits are now allowed repo-wide, and the prompt did not say otherwise because the prompt is not where that asymmetry is recorded.
The rules are narrower than they read, too. A rule that refuses git push does not catch git -C . push, which appears in the docs as an illustration of the gap rather than as a bug report.
What I would do on Monday
If you are on an API key and your agent runs a lot of shell commands, this is the first line item to look at, because it is invisible in the token dashboard unless you go looking for it.
Three things, none clever:
-
Treat a rule you saved as a policy statement, and audit
settings.local.jsonwhen they change. If you cannot say out loud whatnpm test *covers, it covers more than you think. - Put anything you must never have broken into a deny rule. Not in conversation. In the file.
- Run in auto mode on tasks where you would be bored by the prompts, and leave it off everywhere the prompts are the point. The cost overhead lands almost entirely on shell and network calls, which is roughly the same work either way, so you get the speed and the ledger entry and not much else.
Bottom line: the safety check did not get weaker so much as it got quieter and metered. That is a good trade for most work. It is a bad trade for the constraint you meant in English and never wrote down.
If you want the rest of this series when it drops: subscribe via Buttondown and follow so the next one lands in your feed. I am curious what is the widest rule in your settings file right now, because the honest answer is usually a wildcard nobody typed on purpose.
Related on The Agent Loop
- Your agent's cost problem isn't the model. It's the loop.
- Why your MCP approval gate never fires (and what to do instead)
- Your agent's approval needs an expiry date
- Agent memory in 60 seconds: why your agent forgets by tomorrow
Sources
- Claude Code: Configure permissions: permission modes, rule precedence, what a prompt shows, what persists where.
- Claude Code: Permission modes: auto mode, the classifier, cost and latency, fallbacks, boundaries stated in conversation.
- Claude Code: Edit auto-mode rules: block and allow rule overrides.
-
Forem
skip_indexing?: why some posts carrynoindex.
FAQ
Does this mean Claude Code ignores my permissions now?
No. The docs are explicit that "Permission rules are enforced by Claude Code, not by the model," and rules are still evaluated deny, then ask, then allow. What changed is that with auto mode on by default, fewer calls reach a human, and a new classifier sits in front of the ones that would have.
Am I billed per command?
Only in the plans named above: Enterprise, and the Claude API plus Claude Platform on AWS, Bedrock, Google Cloud's Agent Platform and Microsoft Foundry. Consumer subscriptions are not itemized per classifier request. I am not publishing a dollar figure because the transcript portion size is not documented, and I would rather leave a gap than fill it with a guess.
Does auto mode stop prompt injection better than my manual gate?
For the specific vector where fetched page content tries to talk its way into an action, yes, and clearly. Tool results are stripped from classifier requests, so hostile content cannot influence it. Manual review was never protected from that either, because a human skimming a screen has the same problem.
Why not just turn it off?
You can. Manual is the default in older builds, and permissions.defaultMode in ~/.claude/settings.json will hold a mode across sessions. You trade the per-action classifier call for prompt fatigue, which is its own failure mode.
Where do I see what is blocked?
/permissions, on the Recently denied tab. Blocks show up there and r retries with a manual approval. Missing verdicts do not appear at all, which is the gap worth knowing about.

Top comments (0)