DEV Community

Gaurang
Gaurang

Posted on Originally published at discoveringcode.in

Your AI Agent Needs a Brake Pedal: Approval Gates, Scoped Permissions and Dry Runs

Agents are no longer only answering questions. They open PRs, run queries and send emails. Developer forums and regulators are now openly asking who is responsible when an agent does something nobody approved.

Whatever the answer turns out to be, the engineering answer is the same: give real capability, but bound the blast radius. Three patterns do most of that work.

1. Approval gates: the agent proposes, a human decides

The agent still reasons and plans. For consequential tools, the final "actually do it" step waits for a person.

async function runAgentLoopWithApproval(goal, tools, consequentialTools,
                                        callModel, requestHumanApproval) {
  const observations = [];
  for (let step = 1; step <= 8; step++) {
    const decision = await callModel({ goal, observations });
    if (decision.type === "final_answer") return decision.content;

    if (consequentialTools.includes(decision.tool)) {
      const approved = await requestHumanApproval(decision); // blocks
      if (!approved) {
        observations.push({ step, tool: decision.tool,
                            result: { rejected: true, reason: "not approved" } });
        continue; // the agent sees the rejection and re-plans
      }
    }

    const result = await tools[decision.tool](decision.arguments);
    observations.push({ step, tool: decision.tool, result });
  }
}
Enter fullscreen mode Exit fullscreen mode

Don't gate everything. An agent that asks permission for every search is slower than doing the task yourself. Gate the tools that send, write, delete or spend.

2. Scoped permissions: make the bad action impossible

A gate depends on a human noticing a bad proposal. A scoped permission means the bad action was never possible in the first place.

  • A support agent's DB tool can only SELECT this customer's rows. No DELETE, no other customers.
  • A deploy agent's token can deploy to staging, not production.

This is plain least privilege, and it's also the strongest defense against prompt injection, because an injected instruction can't call a capability the agent doesn't have.

3. Dry runs: show the effect before committing

A dry run executes the planning logic without the side effect:

"This would delete 3 rows: id IN (812, 813, 977)."

A gate asks should this happen? A dry run answers what exactly would happen? They work best together. Approving a concrete dry-run result is far more meaningful than approving "delete some records."

Choosing the combination

Action Example Guardrail
Low stakes, reversible search, read-only lookup none needed
Moderate, harder to reverse email a customer, update a record approval gate + dry run
High stakes, irreversible payments, prod data deletion scoped permissions and gate. Never the gate alone

These are the direct mitigations for "Excessive Agency" in the OWASP Top 10 for LLM applications. It's the risk that grows fastest as agents get more capable.


This is from the AI/LLM Engineering pillar on discoveringCode, a free, ad-free notebook that goes from beginner to expert on AI/LLM engineering and four other pillars. Related: why prompt injection can't simply be patched.

Top comments (0)