Agents are no longer only answering questions. They open PRs, run queries and send emails. Developer forums and regulators are now openly asking who is responsible when an agent does something nobody approved.
Whatever the answer turns out to be, the engineering answer is the same: give real capability, but bound the blast radius. Three patterns do most of that work.
1. Approval gates: the agent proposes, a human decides
The agent still reasons and plans. For consequential tools, the final "actually do it" step waits for a person.
async function runAgentLoopWithApproval(goal, tools, consequentialTools,
callModel, requestHumanApproval) {
const observations = [];
for (let step = 1; step <= 8; step++) {
const decision = await callModel({ goal, observations });
if (decision.type === "final_answer") return decision.content;
if (consequentialTools.includes(decision.tool)) {
const approved = await requestHumanApproval(decision); // blocks
if (!approved) {
observations.push({ step, tool: decision.tool,
result: { rejected: true, reason: "not approved" } });
continue; // the agent sees the rejection and re-plans
}
}
const result = await tools[decision.tool](decision.arguments);
observations.push({ step, tool: decision.tool, result });
}
}
Don't gate everything. An agent that asks permission for every search is slower than doing the task yourself. Gate the tools that send, write, delete or spend.
2. Scoped permissions: make the bad action impossible
A gate depends on a human noticing a bad proposal. A scoped permission means the bad action was never possible in the first place.
- A support agent's DB tool can only
SELECTthis customer's rows. NoDELETE, no other customers. - A deploy agent's token can deploy to staging, not production.
This is plain least privilege, and it's also the strongest defense against prompt injection, because an injected instruction can't call a capability the agent doesn't have.
3. Dry runs: show the effect before committing
A dry run executes the planning logic without the side effect:
"This would delete 3 rows:
id IN (812, 813, 977)."
A gate asks should this happen? A dry run answers what exactly would happen? They work best together. Approving a concrete dry-run result is far more meaningful than approving "delete some records."
Choosing the combination
| Action | Example | Guardrail |
|---|---|---|
| Low stakes, reversible | search, read-only lookup | none needed |
| Moderate, harder to reverse | email a customer, update a record | approval gate + dry run |
| High stakes, irreversible | payments, prod data deletion | scoped permissions and gate. Never the gate alone |
These are the direct mitigations for "Excessive Agency" in the OWASP Top 10 for LLM applications. It's the risk that grows fastest as agents get more capable.
This is from the AI/LLM Engineering pillar on discoveringCode, a free, ad-free notebook that goes from beginner to expert on AI/LLM engineering and four other pillars. Related: why prompt injection can't simply be patched.
Top comments (0)