DEV Community

Umang Kumar
Umang Kumar

Posted on

Human Approval for AI Agent Tool Calls: What to Hold, Who Approves, and How the Agent Waits

"Just add a human in the loop" is the most common answer to agent risk, and the most commonly botched one. Approve everything and people click through; approve nothing and the first bad command runs unattended. Human approval for AI agent tool calls works when it is narrow, specific and boring: a small set of actions, a named person, an approval that means yes to this exact call, and an agent that knows how to wait.

I've written before about why generic approval prompts fail as a security boundary. This post is the constructive half: how to design the hold itself.

Step 1: decide what deserves a human

A good test for each action class: if this ran wrongly, could we undo it in five minutes without telling anyone? If yes, don't hold it. If no, consider it.

Usually hold:

  • Production deploys and infrastructure changes
  • Database migrations and writes to shared data
  • Publishing: git push to shared branches, opening PRs, posting comments, sending messages
  • Installing third-party packages (install scripts run with your permissions)
  • Shell commands the policy can't classify as safe

Usually don't hold:

  • Reads inside the workspace
  • Edits to workspace files under version control
  • Running the test suite, linters, git status / git diff

Never hold — deny instead:

  • Reading credential files
  • Recursive force-deletes, history rewrites on shared branches
  • Anything you would refuse even if someone asked nicely

That last category matters. A hold is for actions that are sometimes right. If the answer is always no, a hold just trains people to approve the request that should have been refused.

Step 2: name the approver

"Someone should approve this" turns into "whoever is around clicks yes". Put the approver in the rule: platform-oncall for migrations and deploys, developer (the person running the agent) for unrecognised shell commands, release-manager for production. The approver list is part of the policy, reviewed like any other code.

Step 3: write it as policy and test it

Here are holds written in the Cirvix policy DSL, alongside the denies and allows they sit between:

deny:
  name = deny-destructive
  tool = shell.exec
  risk >= CRITICAL
  reason = "Destructive or remote-exec commands are never run by an agent."

require_approval:
  name = hold-db-migrate
  tool = database.migrate
  approvers = platform-oncall
  reason = "A migration changes the shape of data every other system reads."

require_approval:
  name = hold-prod-deploy
  tool = deploy.apply
  env = production
  approvers = platform-oncall, release-manager
  reason = "Changes what is serving live traffic."

require_approval:
  name = hold-unknown-shell
  tool = shell.exec
  risk >= HIGH
  approvers = developer
  reason = "Unrecognised command; a person should see it first."

allow:
  name = allow-safe-shell
  tool = shell.exec
  risk <= MEDIUM

allow:
  name = allow-workspace-write
  tool = filesystem.write
  workspace = true
Enter fullscreen mode Exit fullscreen mode

With test cases in the same file, cirvix policy test prints:

  ✓ migration waits  → require_approval (hold-db-migrate)
  ✓ staging deploy is not held by the prod rule  → deny
  ✓ prod deploy waits  → require_approval (hold-prod-deploy)
  ✓ unknown script waits  → require_approval (hold-unknown-shell)
  ✓ tests just run  → allow (allow-safe-shell)
  ✓ workspace edit just runs  → allow (allow-workspace-write)

  6/6 PASSED
Enter fullscreen mode Exit fullscreen mode

Note the second test: a staging deploy is denied, not allowed, because no rule permits it and the set is default-deny. Tests like this catch the moment someone assumes a hold rule also grants things it doesn't.

Two ordering rules make holds safe to compose: a matching deny always beats a hold, and a hold beats a permit. Adding a broad allow later can't silently skip the human.

Step 4: bind the approval to the exact call

An approval that means "the agent may run database.write" is a blank cheque: the agent asks to update one row, gets a yes, and spends it on something else. An approval should be bound to the precise call: agent, action, canonical resource, command, and the arguments.

Cirvix's local approval store does this with a fingerprint over those fields, where the arguments are hashed rather than stored (they often contain credentials). A grant is matched by fingerprint, not by tool name, so a yes to one call cannot be reused for a different one.

Step 5: make approvals expire and single-use

"Yes, do that" means yes to the situation in front of the approver now. In Cirvix's local store, an unanswered request expires after 15 minutes by default, and a granted approval stays spendable for 10 minutes and is consumed when used. Shorter grant lifetimes are deliberate: the state the approver reasoned about changes quickly.

The state machine is small on purpose — pending, then approved, denied or expired — and terminal states are terminal. An approved request cannot later be denied, so "who approved this?" has one answer.

Step 6: teach the agent the difference between hold and deny

If a hold looks like an error, agents learn to give up on work a person was about to approve. Make it a distinct signal. In the Cirvix SDKs, a hold raises CirvixHeld (a subclass of CirvixDenied) carrying the approvers and approval id:

import { guard, CirvixDenied, CirvixHeld } from "@cirvix_ai/agent-control";

try {
  await tools.run_migration({ name: "2026_10_add_index" });
} catch (err) {
  if (err instanceof CirvixHeld) {
    // Not a failure: report who needs to approve, then continue other work.
    console.log(`Waiting on ${err.approvers.join(", ")} (approval ${err.approvalId})`);
  } else if (err instanceof CirvixDenied) {
    // Re-plan: the remediation often names the legitimate path.
    console.log(err.policy, err.remediation);
  } else {
    throw err;
  }
}
Enter fullscreen mode Exit fullscreen mode

Holds are non-blocking by default: the call returns immediately with the approval id instead of hanging. An agent running inside an editor with nobody watching a second terminal would otherwise look exactly like a frozen tool call.

Step 7: give approvers a fast queue

Approvers work from the CLI against the same state directory the agent's process uses:

cirvix approvals                      # list held calls
cirvix approve <id> --by alice        # grant (records the reviewer name)
cirvix deny <id> --by alice
Enter fullscreen mode Exit fullscreen mode

After approval, the agent retries the call and the grant is consumed.

Honest limits

  • Local approval records name a reviewer; they are not authenticated signatures. --by alice is a claim, not proof of identity. Team workflows that need authenticated approvers need an identity layer on top.
  • An approval authorises a retry. Nothing automatically resumes the original call, and there is no transaction tying approval, execution and outcome together.
  • Holds only apply to calls routed through the policy layer (MCP gateway, Claude Code hook, or SDK-wrapped tools).

Where Cirvix fits

Cirvix AgentControl is an open-source (Apache-2.0) authorization layer for AI agent tool calls: permit, hold for named approvers, or deny, decided before a governed call executes. Install and test the policy above:

npm install -g @cirvix_ai/agent-control
cirvix policy check --policy approvals.policy
cirvix policy test  --policy approvals.policy
Enter fullscreen mode Exit fullscreen mode

GitHub: https://github.com/CIRVIX/agent-control
Website: https://cirvix.com

Top comments (0)