An agent connected to your CRM through MCP gets a simple instruction: clean up the inactive accounts. It finds 214 matches, picks the delete tool, and fills in the arguments with total confidence. Nothing about that sequence is unusual. The only thing missing is a moment where a person sees the exact action and says yes or no.
That moment does not exist by default. MCP standardizes how agents discover and call tools, but it deliberately leaves approval to the people building the system. The specification says a human should be able to deny a tool call, then stops there. The safety flags a server can attach to a tool are hints, not locks, and a client that ignores them breaks nothing.
Human in the loop approval in MCP is a design choice, not a switch you flip. Learning how to add approval gates to MCP server actions comes down to one principle: the gate has to sit where the side effect happens, enforced by code the agent cannot route around. Get that right and the rest becomes engineering.
Which AI Agent Actions Need Human Approval in MCP? A Risk Tiering Guide
Pausing every tool call for a human is not a safety strategy. It trains reviewers to click Approve without reading. The goal is to spend human attention only where a mistake is expensive.
A practical way to decide is to sort every action an agent can take into tiers, using three questions:
- Can it be undone?
- Who can see its effects?
- How many records does it touch?
Risk tiers
- Tier 1 (Read): Lookups, searches, and list calls. These can run freely, though they still belong in your logs because reads are how data leaks.
- Tier 2 (Internal write): Creating a draft, tagging a ticket, updating a field you can restore. These usually run without a pause as long as there is a clear way back.
- Tier 3 (External facing): Sending an email, posting in a shared channel, publishing a page, creating an invoice. Other people see the result and you cannot unsend it, so approval is the default.
- Tier 4 (Destructive): Deleting records, overwriting data, revoking access. Approve every time, or block the action for agents entirely.
- Tier 5 (Irreversible or high value): Moving money, deploying to production, permanent deletion with no recovery, bulk changes across many records. Require approval from a named role, or keep the action out of the agent's reach altogether.
Environment matters as much as the action itself. The same delete call that is harmless against a sandbox deserves a gate once it points at live customer data, so production versus staging should be part of the policy.
Most integration layers encode this thinking as a policy per integration. Corsair, for example, labels every endpoint as read, write, or destructive, then maps those labels to allow, require approval, or deny through its permission modes: open, cautious, strict, and readonly. Cautious lets agents read and write freely but sends destructive calls to a human. Strict also gates ordinary writes and blocks destructive actions outright.
Whatever tooling you use, write your tiers down before you write any gating code, because the tiers are the policy.
MCP Approval Gates Aren't Enforced by Annotations: How to Build Real Ones
MCP tool definitions can carry annotations such as readOnlyHint, destructiveHint, idempotentHint, and openWorldHint. They look like a ready-made risk system, and it is tempting to wire them straight into approval logic. They are not enforcement, for four reasons.
- They are hints: The specification describes annotations as behavioral hints and tells clients to treat them as untrusted unless they come from a trusted server. A server that labels a delete tool as read only is simply wrong, and nothing in the protocol stops it.
- Their defaults are pessimistic: When a server leaves them out, a tool is assumed to be potentially destructive and not read only. That gives you approval fatigue on one side and false confidence on the other once someone sets the flags carelessly.
- They only matter if a client acts on them: The spec says there should always be a human in the loop who can deny tool calls, and that applications should show confirmation prompts. Both are recommendations. The protocol does not mandate any interaction model, so a headless agent runtime may never ask anybody anything.
- They describe a tool, not a call: One update tool can rename a label or wipe a table depending on its arguments. Risk lives in the arguments, and an annotation cannot see them.
Use annotations for what they are good at, which is helping clients decide what to show a user. Then enforce on the server.
A real approval gate has five properties:
- It sits in front of the side effect. The check runs in code immediately before the write, delete, or send, not in a system prompt or a tool description. Prompts are advice, and agents can ignore advice.
- The decision lives out of the agent's reach. Pending requests and decisions are stored somewhere the agent cannot write to, so it cannot approve itself.
- It binds to exact arguments. The arguments are frozen at request time, shown to the reviewer, and executed as shown. If anything changes, it is a new request.
- It is single use and it expires. An approval covers one execution inside a short window, then it is spent.
- Your server verifies the approver. Never trust a name or role asserted by the client. The spec itself warns against relying on client-provided user identification without verification.
The shape of the code is small:
async function runAction(call, caller) {
const policy = policyFor(call.action); // allow, require_approval, or deny
if (policy === "deny") return blocked("Blocked by policy");
if (policy === "allow") return execute(call);
const approved = await approvals.consume(
caller,
call.action,
hash(call.args)
);
if (!approved) {
await approvals.createPending(caller, call.action, call.args);
return pendingApproval();
}
return execute(approved.frozenArgs); // runs once, with the reviewed arguments
}
Two details matter more than they look.
First, put the gate at the layer that owns the integration, not at a single tool. If a delete can be reached through an MCP tool, a background job, or a script, all three must pass the same check, otherwise gating one route simply teaches the agent to take another.
Second, deny by default for anything unclassified. A new tool with no tier stays gated until someone decides what it is.
This is the reasoning behind Corsair's MCP adapters. They give the agent three tools for listing operations, inspecting schemas, and running scripts, and the permission policy gates the calls that run through them. The check sits at the integration endpoint, so it does not depend on how any one tool describes itself.
How to Implement Human in the Loop Approval in MCP: Pause, Resume, and Audit
Once the gate exists, the real work is the lifecycle around it. Every human in the loop approval flow has three phases.
Pause
When a call needs approval, the server does not execute it and does not guess. It records a pending request and tells the agent what happened.
There are two ways to hold the call:
- Asynchronous: The tool returns right away with a message that approval is needed and where to give it. The agent stops, tells the user, and retries after the decision. This is the safer default because reviews can take minutes or hours, and many agent runtimes cut off slow tool calls.
- Synchronous: The tool call stays open while the server polls for a decision, so the agent sees one slow call. It feels smooth in a demo and breaks the moment a client enforces a timeout.
Corsair supports both, with asynchronous as the default, and its hosted approval page gives reviewers an approve or deny screen without building one yourself. The decision record stays in your own database either way.
Give every pending request an expiry and treat silence as a no. Corsair applies a ten-minute window unless you set another, and its docs recommend denying on timeout.
Resume
After approval there are two clean ways to continue.
The agent can retry the same call and the server matches it to the approved record, or the server can execute the stored request itself the moment the decision lands.
In both cases, the action runs with the arguments the reviewer saw. Never ask the model to rebuild the call from memory, because a slightly different argument is a different action.
Audit
An approval flow without a record is just a delay.
For every gated call, capture:
- Who asked: The agent, the user, and the tenant.
- What was asked: The tool, the target, the exact arguments with secrets redacted, and the tier that triggered review.
- Who decided: The verified approver, the decision, and the timestamp.
- What happened: The execution attempt and the real result.
Keep approval and success as separate events. A yes does not mean the delete worked, because the downstream API can still reject it or time out.
The spec expects clients to log tool use for audit, but a client log is not your log. Record events on the server, using hooks that run before and after each call or equivalent middleware, so no route can skip the record.
How to Build Approval Workflows with MCP Servers (2026 Spec Guide)
The July 28, 2026 revision of the MCP specification changes how a server can ask a person for input in the middle of a tool call, which makes it the most important update for anyone building approval workflows with MCP servers.
Before, a server that wanted a confirmation sent its own request back to the client over a connection that had to stay open. The new revision makes the protocol stateless, drops the initialize handshake and sessions, and replaces those server-initiated requests with Multi Round Trip Requests, usually shortened to MRTR.
Any request can land on any server instance, which suits approval flows that may wait on a human.
The flow works like this:
- The client calls a tool.
- The server decides the call needs input and ends the request with a result of type
input_required. The result carries one or more input requests, such as an elicitation asking the user a question, and optionally an opaquerequestState. - The client asks the user, then retries the same tool call with the answers attached and the
requestStateechoed back, under a new request ID. - The server checks everything and returns the final result.
Here is what an approval request looks like on the wire:
{
"resultType": "input_required",
"inputRequests": {
"approve_delete": {
"method": "elicitation/create",
"params": {
"mode": "form",
"message": "Delete 214 inactive customer records from production?",
"requestedSchema": {
"type": "object",
"properties": {
"approve": {
"type": "boolean"
}
},
"required": ["approve"]
}
}
}
},
"requestState": "signed, expiring blob"
}
The retry carries the person's answer in inputResponses, next to the same requestState.
Elicitation has three possible outcomes: accept, decline, and cancel. Treat only an accept with approve set to true as permission to run the action, and handle the other two as a clean refusal.
Four rules keep this flow safe:
-
Sign the state: The spec treats
requestStateas attacker-controlled input. If it influences authorization or business logic, protect its integrity with an HMAC or authenticated encryption and reject anything that fails verification. - Bind the state: Include the authenticated principal, a short expiry, and a digest of the original request's parameters, so a blob cannot be replayed by another user or on a different call.
- Enforce single use yourself: The spec notes that these measures limit replay but do not guarantee a state is used only once. An approved deletion must be consumed on the server.
- Verify identity on the server: Servers must not rely on user identification supplied by the client, because it can be forged.
Know the limits of in-band approval.
An elicitation answer travels back through the client, and the spec only says clients should offer approval controls, so a client that automatically accepts can answer yes on a human's behalf. Form mode also exposes the exchange to the client and the model's context, and it must never be used to collect secrets.
For changes to production data, a safer pattern is URL mode elicitation, added in the November 2025 revision. It sends the reviewer to a page on your own domain, where your server authenticates them. The client only learns that the user agreed to open the link, and your server learns who actually approved.
In practice, many teams use both: a quick in-band confirmation for lower tiers, and an out-of-band page for tiers four and five.
Reviews that take hours do not fit a single held request at all. For those, return a pending result and resume later, or look at the Tasks extension, which moved out of the core protocol in this revision and lets clients poll for the status of long-running work.
Older clients may not speak the new revision yet, so keep the asynchronous pattern from the previous section as a fallback.
Human Approval for AI Agents in MCP: What to Pause, What to Log, What to Block
Put the earlier sections together and the whole policy fits on one page.
Pause
- Tier 3, 4, and 5 actions, which means anything external-facing, destructive, or irreversible.
- Bulk operations, where one action touches many records at once.
- Anything aimed at production that you would let run freely in staging.
- The first use of a new tool, integration, or tenant, until you know how it behaves.
Log
- Every call, including reads, with the caller, tool, arguments, policy that fired, and result.
- Every approval event: request, reviewer, decision, expiry, and execution.
- Denied and expired requests, because they show you what agents actually try to do.
Block
- Irreversible actions no agent has a reason to run, such as dropping databases, deleting accounts, or changing billing and access controls.
- Unclassified tools, until someone assigns them a tier.
- Replays: expired approvals, spent approvals, and approvals whose arguments have changed.
Approval fatigue is a security problem in its own right.
When reviewers face a flood of low-stakes prompts, they stop reading them. Show the exact action in plain language, include the target and the count, save prompts for the tiers where a mistake is costly, and let routine safe actions pass with logging instead.
Approval gates work best when they live in the integration layer, close to the API call they protect, instead of being rebuilt inside every MCP server.
Corsair is an open source TypeScript integration layer for AI agents that applies permission modes per integration, holds risky calls for human approval with frozen arguments, and keeps the decision record in your own database. It connects to agents through MCP adapters, so the same policy covers every route an agent can take.
Start with one destructive action, gate it, and test approval, denial, and expiry before you widen the policy.
Frequently Asked Questions
How do I implement human in the loop approval in an MCP server?
Classify each action by risk, then enforce a policy in server code right before the side effect. Calls that need approval create a pending record, return a clear message to the agent, and run only after a verified person approves, using the arguments they reviewed. Log the request, the decision, and the result.
Are MCP tool annotations like destructiveHint enough to gate dangerous actions?
No. The specification treats annotations as hints and tells clients to consider them untrusted unless the server is trusted. Use them to help clients decide what to display, and enforce approval on the server where the action actually executes.
What are Multi Round Trip Requests, and how do they help with approvals?
Multi Round Trip Requests, introduced in the July 28, 2026 specification revision, let a server end a tool call with an input_required result and ask the client for input. The client collects the answer and retries the same call with it attached. This lets a server request confirmation without holding a connection open or storing session state.
Should approvals use MCP elicitation or a separate review page?
Use elicitation for quick confirmations on lower-risk actions in interactive clients. For production data and high-value actions, use an out-of-band review page, or URL mode elicitation, where your server authenticates the approver. In-band answers pass through the client, so they are weaker evidence of who actually decided.
How do I stop approval gates from creating review fatigue?
Gate only the tiers where mistakes are costly, usually external-facing, destructive, and irreversible actions. Show the exact action, target, and record count in plain language, expire stale requests, and let routine low-risk actions run with logging. If reviewers approve nearly everything, tighten the policy or improve what the prompt shows.
Top comments (0)