DEV Community

Umang Kumar
Umang Kumar

Posted on

Why approval prompts don't work as a security boundary for coding agents

When a coding agent asks a human to approve a file change, a database call, or a deploy, the approval prompt feels like a security boundary. It's not.

In practice, approval prompts leak authority for three reasons: approval fatigue, approvals that never expire, and stale pending approvals that outlive the decision they capture.

Approval fatigue

An agent that runs an entire PR lifecycle can surface dozens of prompts in a single session. Most of them are low-risk: formatting a config, reading a log, re-running tests. A few are high-risk: writing to a deployment manifest, touching secrets, changing network rules.

When everything looks the same in the prompt, humans stop reading. They learn to click "Approve" to keep the agent moving and catch up later. That later is when the deploy went to the wrong environment and the rollback took thirty minutes.

The fix isn't more prompts. It's making every prompt distinguishable.

Approvals that never expire

Most agent toolkits record an approval as a boolean decision on a category of action: "approve kubectl apply" or "approve writes to /etc". There is no time bound attached.

An approval granted at 9:00 AM should not authorize the same action at 4:00 PM, after the incident context has changed, the on-call has handed off, and the agent's mission scope has drifted.

Without an expiry, the approval becomes a standing credential. The human reviewed a point-in-time description and unknowingly minted a persistent grant.

Stale pending approvals

The flip side: a prompt sits in a queue while the human is on a call or asleep. Three hours later, the agent replays the pending request against a newer snapshot of the environment.

The description the human approved no longer matches what would execute. The approval was sound when issued; it is unsound when fulfilled.

This is why "pending approvals" and "authorized actions" are not the same thing. A pending approval captures a decision about a specific action at a specific time. Once that time window closes, the decision must be withdrawn.

A better shape: expiring HOLD approvals bound to the exact action hash

A practical fix that keeps humans in the loop without minting standing credentials:

  1. Halt the action, don't queue it. The agent emits a HOLD with a deterministic hash of exactly what would execute.
  2. Bind the approval to the hash, not the category. The human's "yes" signs off on that hash. If anything about the inputs, target, or arguments changes, the hash changes and the approval doesn't apply.
  3. Attach a short, non-extendable expiry. Five to fifteen minutes is enough for a focused review; long enough that the human isn't forced to answer while a deployment countdown hits zero.
  4. Fail closed on expiry. When the timer expires, the HOLD transitions to DENY. The agent re-emits a fresh request with a fresh hash if it still wants the action.

This is how Cirvix AgentControl structures an approval:

import { createHash } from "crypto";

function actionHash(action: Record<string, unknown>): string {
  return createHash("sha256")
    .update(JSON.stringify(action, null, 0))
    .digest("hex");
}

interface HoldApproval {
  id: string;
  actionHash: string;
  expiresAt: number; // epoch ms
  approvedBy: string;
  approvedAt: number;
  status: "pending" | "approved" | "denied" | "expired";
}

function verifyApproval(
  hold: HoldApproval,
  requestedAction: Record<string, unknown>
): boolean {
  const freshHash = actionHash(requestedAction);

  return (
    hold.status === "approved" &&
    hold.actionHash === freshHash &&
    Date.now() < hold.expiresAt
  );
}

const call = { tool: "kubectl_apply", manifest: "deploy.yaml", namespace: "prod" };
const hold: HoldApproval = {
  id: "hold-001",
  actionHash: actionHash(call),
  expiresAt: Date.now() + 10 * 60 * 1000,
  approvedBy: "human@example.com",
  approvedAt: Date.now(),
  status: "approved",
};

console.log(verifyApproval(hold, call)); // true

const drifted = { ...call, namespace: "prod-backup" };
console.log(verifyApproval(hold, drifted)); // false
Enter fullscreen mode Exit fullscreen mode

The three checks — status approved, hash matches the exact action, and current time is still inside the expiry — make the approval a verifiable ticket for one execution, not a standing credential.

What changes for the agent

  • The agent's loop gains an explicit HOLD/DENY transition instead of silently retrying a pending queue.
  • Tool wrappers compute the action hash before execution and refuse to run if the hash doesn't match the approval that cleared them.
  • Observability becomes simpler: every approval event carries the hash and the expiry, so you can diff a denied execution against what was approved and see exactly which field drifted.

What doesn't change

Humans still review. The difference is that their review is scoped: a short window, a specific action, and a clear failure mode when the scope shifts.

Prompt-based approvals are still useful as a UX. They just shouldn't be the boundary. The boundary is the hash-bound, time-bound decision that the enforcement layer verifies at execution.

I'm building Cirvix AgentControl, an open-source default-deny policy layer for agent tool calls: https://github.com/CIRVIX/agent-control (try npx @cirvix_ai/agent-control scan).

Top comments (0)