"Just add a human in the loop" is the most common answer to agent risk, and the most commonly botched one. Approve everything and people click through; approve nothing and the first bad command runs unattended. Human approval for AI agent tool calls works when it is narrow, specific and boring: a small set of actions, a named person, an approval that means yes to this exact call, and an agent that knows how to wait.
I've written before about why generic approval prompts fail as a security boundary. This post is the constructive half: how to design the hold itself.
Step 1: decide what deserves a human
A good test for each action class: if this ran wrongly, could we undo it in five minutes without telling anyone? If yes, don't hold it. If no, consider it.
Usually hold:
- Production deploys and infrastructure changes
- Database migrations and writes to shared data
- Publishing:
git pushto shared branches, opening PRs, posting comments, sending messages - Installing third-party packages (install scripts run with your permissions)
- Shell commands the policy can't classify as safe
Usually don't hold:
- Reads inside the workspace
- Edits to workspace files under version control
- Running the test suite, linters,
git status/git diff
Never hold — deny instead:
- Reading credential files
- Recursive force-deletes, history rewrites on shared branches
- Anything you would refuse even if someone asked nicely
That last category matters. A hold is for actions that are sometimes right. If the answer is always no, a hold just trains people to approve the request that should have been refused.
Step 2: name the approver
"Someone should approve this" turns into "whoever is around clicks yes". Put the approver in the rule: platform-oncall for migrations and deploys, developer (the person running the agent) for unrecognised shell commands, release-manager for production. The approver list is part of the policy, reviewed like any other code.
Step 3: write it as policy and test it
Here are holds written in the Cirvix policy DSL, alongside the denies and allows they sit between:
deny:
name = deny-destructive
tool = shell.exec
risk >= CRITICAL
reason = "Destructive or remote-exec commands are never run by an agent."
require_approval:
name = hold-db-migrate
tool = database.migrate
approvers = platform-oncall
reason = "A migration changes the shape of data every other system reads."
require_approval:
name = hold-prod-deploy
tool = deploy.apply
env = production
approvers = platform-oncall, release-manager
reason = "Changes what is serving live traffic."
require_approval:
name = hold-unknown-shell
tool = shell.exec
risk >= HIGH
approvers = developer
reason = "Unrecognised command; a person should see it first."
allow:
name = allow-safe-shell
tool = shell.exec
risk <= MEDIUM
allow:
name = allow-workspace-write
tool = filesystem.write
workspace = true
With test cases in the same file, cirvix policy test prints:
✓ migration waits → require_approval (hold-db-migrate)
✓ staging deploy is not held by the prod rule → deny
✓ prod deploy waits → require_approval (hold-prod-deploy)
✓ unknown script waits → require_approval (hold-unknown-shell)
✓ tests just run → allow (allow-safe-shell)
✓ workspace edit just runs → allow (allow-workspace-write)
6/6 PASSED
Note the second test: a staging deploy is denied, not allowed, because no rule permits it and the set is default-deny. Tests like this catch the moment someone assumes a hold rule also grants things it doesn't.
Two ordering rules make holds safe to compose: a matching deny always beats a hold, and a hold beats a permit. Adding a broad allow later can't silently skip the human.
Step 4: bind the approval to the exact call
An approval that means "the agent may run database.write" is a blank cheque: the agent asks to update one row, gets a yes, and spends it on something else. An approval should be bound to the precise call: agent, action, canonical resource, command, and the arguments.
Cirvix's local approval store does this with a fingerprint over those fields, where the arguments are hashed rather than stored (they often contain credentials). A grant is matched by fingerprint, not by tool name, so a yes to one call cannot be reused for a different one.
Step 5: make approvals expire and single-use
"Yes, do that" means yes to the situation in front of the approver now. In Cirvix's local store, an unanswered request expires after 15 minutes by default, and a granted approval stays spendable for 10 minutes and is consumed when used. Shorter grant lifetimes are deliberate: the state the approver reasoned about changes quickly.
The state machine is small on purpose — pending, then approved, denied or expired — and terminal states are terminal. An approved request cannot later be denied, so "who approved this?" has one answer.
Step 6: teach the agent the difference between hold and deny
If a hold looks like an error, agents learn to give up on work a person was about to approve. Make it a distinct signal. In the Cirvix SDKs, a hold raises CirvixHeld (a subclass of CirvixDenied) carrying the approvers and approval id:
import { guard, CirvixDenied, CirvixHeld } from "@cirvix_ai/agent-control";
try {
await tools.run_migration({ name: "2026_10_add_index" });
} catch (err) {
if (err instanceof CirvixHeld) {
// Not a failure: report who needs to approve, then continue other work.
console.log(`Waiting on ${err.approvers.join(", ")} (approval ${err.approvalId})`);
} else if (err instanceof CirvixDenied) {
// Re-plan: the remediation often names the legitimate path.
console.log(err.policy, err.remediation);
} else {
throw err;
}
}
Holds are non-blocking by default: the call returns immediately with the approval id instead of hanging. An agent running inside an editor with nobody watching a second terminal would otherwise look exactly like a frozen tool call.
Step 7: give approvers a fast queue
Approvers work from the CLI against the same state directory the agent's process uses:
cirvix approvals # list held calls
cirvix approve <id> --by alice # grant (records the reviewer name)
cirvix deny <id> --by alice
After approval, the agent retries the call and the grant is consumed.
Honest limits
- Local approval records name a reviewer; they are not authenticated signatures.
--by aliceis a claim, not proof of identity. Team workflows that need authenticated approvers need an identity layer on top. - An approval authorises a retry. Nothing automatically resumes the original call, and there is no transaction tying approval, execution and outcome together.
- Holds only apply to calls routed through the policy layer (MCP gateway, Claude Code hook, or SDK-wrapped tools).
Where Cirvix fits
Cirvix AgentControl is an open-source (Apache-2.0) authorization layer for AI agent tool calls: permit, hold for named approvers, or deny, decided before a governed call executes. Install and test the policy above:
npm install -g @cirvix_ai/agent-control
cirvix policy check --policy approvals.policy
cirvix policy test --policy approvals.policy
GitHub: https://github.com/CIRVIX/agent-control
Website: https://cirvix.com
Top comments (0)