Every coding agent now reads an instructions file. CLAUDE.md, AGENTS.md, .cursorrules, whatever yours is called. Mine is long. It has been long for months, because every time something went wrong I did the obvious thing and wrote a sentence about it.
Here is what I eventually noticed. The agent follows those sentences most of the time. Not always. And the times it doesn't are not random: they're the moments when the file is long, the task is interesting, and the rule is the least salient thing in a 30k-token context. Which is exactly when you need the rule.
A rule in prose competes with every other sentence in the file for attention. A rule in a hook does not compete with anything. It just runs.
So over a few months I moved the rules that actually mattered out of the instructions file and into PreToolUse hooks. There are about a dozen now. The agent is noticeably less annoying to supervise, and I stopped repeating myself.
The examples below are Claude Code hooks, because that's what I use. The idea is the part worth stealing.
The mechanics are almost insultingly simple
A PreToolUse hook is a program. The agent is about to call a tool, your program gets the call as JSON on stdin, and you decide.
#!/usr/bin/env python3
import sys, json
data = json.load(sys.stdin)
cmd = (data.get("tool_input") or {}).get("command", "")
if "psql" in cmd and "DELETE" in cmd.upper():
print(json.dumps({
"decision": "block",
"reason": "No DELETE against this database from an agent session. "
"Write the statement into a migration and let a human run it.",
}))
sys.exit(0)
Two details make this better than a code review or a nag in the prompt:
The call never happens. Not "the model is discouraged from it". The tool returns your refusal instead of running.
And the reason goes back to the model, not to you. That's the part I underrated. A good reason makes the agent correct itself and carry on, so a blocked call costs a few seconds instead of a round trip through your attention. A bad reason ("blocked") makes it guess, flail, or ask you. The reason string is the whole product.
The tripwires
Grouped by what they're actually protecting, because the categories turned out to matter more than the specifics.
Things that can't be undone. A guard that makes writes to the production database impossible from any agent session: mass deletes, truncates, cache flushes, a session-level SET through a connection pooler. A guard that blocks merge commands, because merging is a human's decision and an agent that can merge will eventually merge something at 2am. A guard that refuses sed -i into source files, which is how an agent rewrites a file without the edit ever passing through the checks on the edit tools.
Things that are expensive. A guard that blocks local type-checking entirely. That one sounds mad until you know that a cold run across this monorepo costs about fifty seconds of CPU and one process peaked at 5.6 GB, and that several agents in several worktrees will happily do it at once on a 24 GB laptop. CI type-checks every push anyway. The hook's reason says so, so the agent commits instead of grinding.
Things that are silently wrong. My favourite hook, and the shortest. Our analytics query language reads a bare timestamp literal in the project's timezone, not UTC. Write timestamp > '2026-07-03 00:00' and you have quietly shifted your window by ten hours. It once made a feature flag look like it was ramping when it wasn't, which produced a confident wrong diagnosis and a needless experiment reset. Now:
for m in re.finditer(r"(timestamp|created_at|_at)\s*[<>=]+\s*'(\d{4}-\d{2}-\d{2}[^']*)'", text, re.I):
if re.search(r"([+-]\d{2}:?\d{2}|Z)\s*$", m.group(2)):
continue # explicit offset, fine
print(json.dumps({"decision": "block", "reason": TZ_TRAP_EXPLANATION}))
Forty lines, including the explanation. That class of bug cannot reach me any more, from any session, in any file or shell command, whether or not anybody remembered the rule.
Things that are about taste, and therefore hopeless in prose. A guard that blocks any edit adding a multi-line comment block. I had written "do not write comments" in the instructions file in capital letters, with reasons, twice. It did not work reliably, because the model's prior for "explain this clever bit" is strong and my sentence was one sentence. The hook works every time, and its reason tells the agent to rewrite the block as one line rather than just deleting the thought.
Things about process. A guard on new UI that requires a stable test id, because the e2e suite selects by test id and a renamed element silently stops testing anything. And a guard that refuses to push a branch until an adversarial review receipt exists for the current head. The second one is the most aggressive thing in my setup, and I'd defend it: an agent that is sure its work is ready is not evidence that the work is ready.
What it cost
I don't want to sell this as free.
False positives are real and they're the price. My shell guards block patterns that are usually footguns and occasionally the right thing, and then I'm reading a refusal about a command I meant to run. The honest fix is to go edit the hook, which I do, and which is still cheaper than the alternative because the fix is a diff somebody can review rather than another sentence in a file nobody re-reads.
They're also per-repo and they rot. A hook that encodes "CI runs this check" is wrong the day CI stops. Mine have tests, which felt absurd when I wrote them and does not any more.
And they don't replace the instructions file. Taste, context, architecture, the reason a weird thing is weird: that's all still prose, and prose is still the right shape for it. Hooks are for the small number of rules where "usually" isn't good enough.
The thing I actually changed my mind about
I started doing this because I didn't trust the model. That was the wrong frame, and it made me write mean little guards with bad reasons.
The better frame is that an agent has no way to know which of your forty rules is the one that costs money. You do. A hook is how you say "this one is not a preference", and a good reason string is how you teach instead of scold. The agent reads it, understands why, fixes its own command, and moves on. That's a better loop than any amount of bold text in a markdown file.
So: what's on your deny list? I'm especially curious about the opposite view, because one rule I keep going back and forth on is whether blocking git push crosses from guardrail into busywork. Tell me I'm wrong about that one.
Top comments (4)
"The reason string is the whole product" — yes, this is the line. The failure mode of a bare
blockedis that the agent treats it as noise and either retries a variant or escalates to you, so a lazy reason converts a deterministic guard back into the exact attention tax you were trying to kill. One thing I'd add: the salience argument cuts both ways over time. A dozen hooks is maintainable; the trap is that hooks quietly become their own unversioned rulebook, and six months in nobody remembers whysed -iis blocked, so someone loosens it during an incident. We started treating the hook set like code — each one gets a one-line comment with the incident that birthed it, and they get reviewed, not just accumulated. Also curious where you draw the line between "block" and "warn-and-continue" — some rules (merge, prod DELETE) are obviously hard blocks, but for the fuzzier ones, does returning a reason and letting the model decide work better than a wall?Yes to the rulebook risk, and
sed -iis my sharpest example of it too. What's saved me there isn't a comment in the file, it's that the reason string has to explain itself. Mine says why sed trips the harness and what to use instead, so whoever loosens it mid-incident reads the rationale inside the refusal rather than going hunting for it. The incident still goes in the docstring, and two of my hooks have unit tests, which felt absurd to write and doesn't any more.On block versus warn, I ended up with three modes, and severity turned out to be the wrong axis. Determinism is the right one: can the hook tell from the tool input alone, with near certainty, that this call is wrong? If yes, block, because a correct retry is cheap and the reason says how. If the rule is a heuristic that will fire on legitimate work, warn and continue, inject the context and let the model keep its judgement. My comment guard does exactly that split: a real multi-line block is a hard block, a borderline one-liner just gets a reminder appended.
The third mode is the one I never see discussed: auto-allow. Same mechanism, used to approve read-only commands so they never surface a prompt. Most of my deny list isn't about danger, it's about not spending my own attention, and allowing is the same job from the other end.
On git push, I think the line is whether the agent can do anything with the reason. Your review-receipt guard is a contract because its reason names a condition the agent can go and satisfy, then push. A flat block on push has no such condition, so the only correct response is to stop and ask you, which is exactly the round trip the hook was supposed to save. So I'd keep push guarded but never unconditionally blocked. One gap in the timezone hook: it reads literal text, so a query whose timestamp is assembled from a variable at runtime walks straight past it.
Fair, yours is the better test. :)
Both my push guards do name a condition, the review one tells it which file to write, the comment one prints the lines it deleted so they're satisfiable. Merge is the one I block flat out, and that's deliberate, it isn't the agent's call to make.
On the timezone hook: yep, literals only. I wrote it for a bare date typed into an ad-hoc query, which is what bit me. A template or a bound param sails straight past and I don't have a good answer for that yet. Probably needs catching where the query gets built instead.
Would love to know your opinions. Your comment made me think. Thanks for it.