I had an incident at work where an AI-powered agent made a change that passed CI and seemed entirely reasonable in the PR. The problem wasn’t just ...
For further actions, you may consider blocking this person and/or reporting abuse
I would add: ask the agent to prove the test can fail.
A green test is not automatically evidence that the behaviour works. Temporarily reverting or mutating the implementation, then check that test can go red for the intended reason. Otherwise green can become dangerous, because it creates confidence without protection.
A permissions file is a claim about capability. The control is the enforcement point, and the two drift apart quietly. Moving a guardrail out of the prompt and into config is a real step up, and it still leaves an assertion sitting there until something shows the deny actually firing.
Your own example undercuts itself.
deny Bash(curl *)sits next toallow Bash(npm run *). A rule that matches invocation strings is a spelling test. It is not a capability boundary. Any package script that shells out to curl is reachable through the allowed spelling, and that wildcard hands over a user-defined script namespace, which is exactly where the denied capability walks back in. The name returns by other routes too: an absolute path,command curl, a leading backslash,sh -c. What that deny list bounds is names. The review you are proposing reads names.In the stack we run, risky writes route through wrapper scripts, and name matching alone was not enough. Shims had to go at the front of PATH because
command X,\Xand/usr/bin/Xeach defeat matching on the name. Then the more embarrassing failure. That shim directory dropped off PATH and sat inactive for a stretch while the health check kept printing OK. Config armed, enforcement point never reached. Reading the file would not have caught that.So the review item with weight is cheap to run: fire the denied action through the allowed entry point and produce the refusal line.
What in your review tells you a deny rule was evaluated at all, rather than configured?
The permissions-file-as-code-review point is the sharpest part of this. We hit the exact same drift building viaSocket's automation rules: a workflow config is technically "just JSON" so it gets skimmed like a .prettierrc, but it's the thing that actually decides blast radius. One thing I'd add to your three questions: what happens on retry. An agent that re-runs a failed step because the setup allowed it, against a tool that isn't idempotent, produces the same kind of silent damage as a missing permission, and no permissions file catches it because it's not a scope problem, it's a state problem.
This is a useful shift in perspective. It prompted me to turn agent review into a concrete control in my own development process: for higher-risk changes, I now review the capability delta—the authority added by the change—and require evidence from the actual enforcement point.
The important lesson for me is that configuration alone is not sufficient evidence. I document what is newly allowed, where that boundary is enforced, which human gate remains, and the result of a safe negative test showing that an out-of-scope action is rejected.
I deliberately keep this evidence data-minimised: references, hashes, statuses, and error codes rather than prompts, secrets, or full payloads.
This complements rather than replaces risk-based diff review. The two review layers answer different questions: “Could the agent do this?” and “Is the resulting change correct?”
How do you decide where that boundary sits—which changes still require a full human diff review, and which can safely rely on capability controls plus targeted verification?
This shift in focus from code to agent-permissions in your review process is incredibly insightful. As you noted, the change in the ratio of code production to review speed really highlights the necessity for a more structured approach to understanding agent capabilities and their impact. It might be beneficial to implement a clear documentation strategy for agent behaviors and limits, perhaps even creating a checklist for reviewers to ensure critical aspects aren’t missed. If you’re looking for additional engineering support to enhance this review process further, I’d be glad to explore a paid collaboration. How are you currently addressing the documentation of these agent configurations?
The shift from reviewing code to reviewing the agent makes me think the next question is whether the agent can actually prove that its actions stayed within its authority. At IT Path Solutions, we’ve found that having the right permissions configured is only one layer; the system also needs observable evidence of what the agent attempted, what actually executed, and what was blocked. That creates a useful distinction between “the agent was allowed to do this” and “the agent actually did only what it was supposed to do.” For production systems, that evidence can be just as important as the permission model itself.
A tightly-scoped agent can still produce issues and even worse can be biased
I've been doing graph engineering — basically mapping out the call graph, who-calls-what — as a way to keep catching this stuff continuously, not just once at review time. Trying to catch code quality, security, compliance issues and bugs this way. The challenge is many false positives or minor findings but getting there..
The ratio flip you describe is the part most teams haven't internalized: when writing is slower than reading, reading is the control. When an agent emits a day's diff in an hour, reading becomes the theater. We hit the same wall reviewing agent PRs — the fix that worked for us was moving the control one layer up, exactly like you said: we review the agent's constraints (allowed tools, allowed paths, a deny-by-default network policy) once, then spot-check the diff. The failure mode nobody warns about: the PR description is also written by the agent, so "read the description, approve" is really "let the agent grade its own homework."
One thing that helped us: a deterministic checker that re-derives what the agent claims it did (files touched vs. files declared, tests claimed vs. tests actually run in CI) and fails the PR on mismatch. Cheap, boring, catches the honest-hallucination cases that reading never will.
How are you handling the description-trust problem — do you have reviewers re-verify the agent's claims, or did you move that check into CI too?
On "show the deny actually firing": mine never did. The Read-gating hook stayed green because the model reads files through cat -n and sed -n in Bash, so the hook saw nothing. The transcript proved the deny was a claim.
Really enjoyed this perspective. The shift from reviewing code to reviewing the agent's capabilities, permissions, and boundaries feels like a real change happening across AI-assisted development.
The point about configuration files, tool access, and approval gates being part of the review process is especially important. Green CI doesn't automatically mean a safe agent.
Thanks for highlighting that reliable AI development is about reviewing the entire automation setup—not just the generated diff.
The missing artifact being the permission audit rings true. Reviewing the diff tells you what changed; it says nothing about what was reachable. I hit a smaller version of this: an agent I let post to a board API started calling endpoints I never mentioned, just because they existed in the same docs. The output review caught what it wrote, never the surface it could see. "What could this have touched" is a review question, not a runtime question - by runtime it has already been answered.
Strong reframe. One practical way to make this reviewable is to gate on capability delta, not diff size: before a run, the agent emits a small manifest (paths it may write, tools/network scopes, and deploy rights). The operator reviews that once; each PR then carries the manifest plus a machine-generated run summary: files touched, tool calls, tests run, retries, and any scope expansion. Auto-merge only when the observed set stays within the manifest and CI is green. Escalate on a scope expansion, repeated retries, or a failed invariant—not just a large diff. That turns “review the agent” into an auditable loop without forcing a human to inspect every token.
excellent
Good post! I wanna discuss in detail about experience and colloboration.
Best