Human in the loop for AI agents means a person approves, reviews or is notified about specific actions the agent takes, instead of trusting it end to end. It works best as a decision made per action rather than per agent: which actions the agent runs on its own, which it reports, which wait for approval and which it never takes.
Most vendor conversations treat it as a checkbox: yes, there is a human in the loop. That answer hides the two ways oversight fails. If a person reviews everything, the agent becomes a suggestion engine feeding a queue nobody clears. If nobody reviews anything, you learn what the agent misunderstood only after it has caused a problem. This article gives you a method for placing each action, and then covers the part that usually gets skipped, which is designing the review so the person stays a real check instead of approving by reflex.
Human in the loop is a decision you make per action
An AI agent in a business workflow takes a series of distinct actions. It reads a message, classifies it, updates a record, drafts a reply, sends it, closes a ticket or moves money. Those actions carry different levels of risk, so they need different levels of oversight. Each action can sit in one of four modes.
| Mode | What the agent does | What the person does | Fits when |
|---|---|---|---|
| Act and log | Completes the action and records what it did and why | Reviews a sample of the log on a schedule | The action is easy to reverse, stays internal, and happens often |
| Act and notify | Completes the action and tells a named person | Reads the notice and can reverse it inside a short window | The action is reversible but someone should know it happened |
| Approve before | Prepares the action with its evidence and stops | Approves, edits, or rejects, usually with one click | The action is hard to undo or reaches customers, money, access, or people |
| Never | Cannot perform the action, because the permission was not granted | Does the action themselves, with the agent's preparation | The action is irreversible and consequential, or outside the agent's job |
Most of the agent's value comes from the first two modes. Approve before is where the agent builds trust with the people who depend on it, and the never list is what keeps the worst week survivable.
OpenAI's practical guide to building agents names two triggers that warrant human intervention: exceeding a failure threshold, such as too many retries, and high-risk actions that are "sensitive, irreversible, or have high stakes," which it says should trigger human oversight "until confidence in the agent's reliability grows." That last clause matters, because it makes approve before a starting position for an action rather than a permanent one.
You can see the same structure in the tools engineers use to let agents work on code. Anthropic's Claude Code permission system is tiered. Reads inside the working directory need no approval, while shell commands, file edits and most web fetches ask first, and an approval can be saved as a standing rule. Its documentation reserves the mode that skips prompts for "isolated environments" where the agent cannot cause damage. Per action oversight already ships there as a product default.
How to decide which mode an action gets
Placing an action does not require understanding the model. It takes two answers that a business can give in a meeting.
The first is reversibility. Can the action be undone, by whom, at what cost and inside what window? Relabeling an email can be undone in a second, and a sent email cannot be undone at all.
The second is blast radius. Who and what does the action reach? A note on an internal record touches one row, and a message to a customer touches your reputation. A changed payment detail touches money, a change to someone's access touches security, and a decision about a person's employment touches their life.
OpenAI's guide suggests rating each tool an agent can use as low, medium or high risk on exactly these kinds of factors, including "read-only vs. write access, reversibility, required account permissions, and financial impact," and using the rating to pause or escalate. The placement matrix below applies the same idea for the person who owns the workflow.
| Reversibility | Small radius: internal, one record | Medium radius: a team's work, a vendor, an internal message | Large radius: customers, money, access, employment, legal |
|---|---|---|---|
| Undone in seconds by anyone | Act and log | Act and notify | Approve before |
| Undone with effort, or only inside a window | Act and notify | Approve before | Approve before |
| Cannot be undone | Approve before | Approve before, or never | Never at launch |
Here is where some common actions land. Classifying an inbound message and assigning an owner in the CRM are act and log. Posting a daily status summary to an internal channel is act and notify. Drafting a reply to a customer needs no gate, because a draft changes nothing, but sending that reply is approve before, since the send is the action. Issuing a refund or changing a vendor's bank details is never at launch: the agent prepares the case and a person does it.
Which systems the agent can reach at all is a separate decision, covered in what access an AI agent should have to your business systems. This article assumes access is settled and looks at what happens inside it.
The actions that stay behind a person, whatever the matrix says
Anything that commits money, changes access or permissions, deletes records, sends something a customer or vendor will treat as a promise, or affects someone's employment or legal position belongs on the human side of the line at launch. That is the default in our discovery sessions, and it matches the approval line in the rollout guide.
The agent can still help with these actions. It gathers the evidence and drafts the action, and a person commits it. In most workflows, the preparation is where the hours were going in the first place.
Design the review so it stays real
Vendor pages usually skip this half of human in the loop, and it decides whether the design keeps working or slowly stops meaning anything.
Name a person as the approver
An approver has to be a person, so "operations approves" does not count. Name the approver and a backup, and record both. If nobody can be named, the action is not ready for approve before and belongs in never.
Use the channel the approver already works in
Approval requests go to the tool the approver already uses, whether that is a chat channel or an inbox. They arrive complete, with the proposed action, the evidence the agent used, its reasoning in one or two lines and the options to approve, edit or reject. If checking an approval means opening four systems, it will be approved without checking.
A deadline and a default
Every approve before action needs a response window and a stated outcome if the window closes with no response. The safe default is almost always "do nothing and escalate to the backup." A default of "proceed if nobody objects" turns the approval gate into a timer, which provides no oversight at all. Who approves, in what tool and by when is one of the seven inputs the implementation guide asks you to bring to a first working session.
A review load a person can actually carry
Most designs ignore this constraint, and human factors research on people supervising automation is direct about it. In their review of the empirical studies, Parasuraman and Manzey found that automation complacency occurs when the supervised task competes with other work for attention, that it shows up in novices and experts alike, and that it "cannot be overcome with simple practice." Automation bias, acting on the machine's suggestion when it is wrong, "cannot be prevented by training or instructions." (Human Factors, 2010)
In practice, if you route hundreds of low-stakes approvals a day to one busy person, you get hundreds of clicks and very few real reviews. A more diligent person will not fix that, but a shorter approve before list will. Put the reversible, internal, high-volume actions in act and log, keep approve before for the actions that need it, and treat the daily approval count as a design number you check every week.
Sample the autonomous tier
Act and log still means somebody looks. A person checks a sample on a schedule instead of every instance in real time. Pull a fixed number of logged actions each week, check them against what a person would have done and record the miss rate. That miss rate tells you whether the action belongs in its current mode.
Enforce approval in the system instead of the prompt
An instruction in the agent's prompt such as "never send an email without approval" is only a request. OWASP's guidance on excessive agency names excessive autonomy as one of three root causes of damaging agent actions, and its mitigation is explicit. It calls for human approval of high-impact actions, implemented either in the downstream system or in the tool that performs the action, and for complete mediation, so that the systems the agent calls enforce authorization instead of the model deciding whether it is allowed. In plain terms, the send button should wait for a person because the integration is built that way, whatever the agent has been told.
Log everything and cap the rate
OWASP lists two measures that limit the damage of a bad action without preventing it: log and monitor what the agent's tools do, and rate limit them so monitoring catches an undesirable pattern before significant damage occurs. Anthropic's guidance on building effective agents makes the same point from the other side, noting that autonomy brings "the potential for compounding errors." Every action in every mode should get a record of the input, the decision, the evidence and the approver if there was one, and every action type should have a ceiling per hour. If an agent starts doing the wrong thing two hundred times, the limit should stop it at ten.
Move actions between modes on evidence, in both directions
For most actions, approve before is meant to be temporary, so you need a rule for what earns a move.
Use the evidence from the review design above. For an action in approve before, track how often the approver changed or rejected what the agent proposed. For an action in act and log, track the sampled miss rate. When both the edit rate and the miss rate have stayed low across enough consecutive cycles to mean something for that workflow, move the action one step: approve before to act and notify, or act and notify to act and log. For illustration only, and not as a benchmark, here is the kind of rule to write down: an internal action graduates after a full month of cycles with no approver edits and no sampled misses. Set your own numbers and put them in writing before launch.
Moves go the other way too. If the workflow's rules change, an integration changes, the underlying model changes or the edit rate climbs, the action goes back a step until the evidence recovers. Customer-facing sends and anything touching money do not graduate on schedule alone. They graduate when a named person signs off on the evidence, and some of them never do.
Anthropic's guidance describes agents that "pause for human feedback at checkpoints or when encountering blockers." Graduation is how you decide, deliberately and with a record, where those checkpoints sit this quarter.
A worked example: an inbox triage AI employee
Take a common role family: an AI employee that monitors a shared inbound inbox, keeps it organized and prepares replies. Here is one reasonable launch configuration. Your numbers and owners will differ.
| Action | Mode at launch | Why |
|---|---|---|
| Read new messages and classify them by type | Act and log | Internal, reversible in seconds, high volume |
| Assign an owner in the helpdesk or CRM | Act and log | Internal, reversible, the owner sees it immediately |
| Post a morning summary of open items to the team channel | Act and notify | Internal message, useful for everyone to see, easy to correct |
| Draft a reply to a customer | No gate needed | Drafting changes nothing until someone sends it |
| Send a reply to a customer | Approve before | Reaches a customer, cannot be unsent |
| Close a ticket as resolved | Act and notify | Reversible, and the owner should know |
| Issue a refund or credit | Never at launch | Money, and the case is prepared for a person |
The review design that goes with it has the support lead approving customer sends, with the operations manager as backup. Requests arrive in the team's chat tool with the customer's message, the draft and the reason for the proposed reply. The window is the same business day, and if it closes, the item escalates to the backup and nothing is sent. Each week the lead reviews a fixed sample of classifications and assignments and records the miss rate. After a month of cycles with no approver edits on a narrow class of replies, such as order status confirmations that quote the record exactly, that class becomes a candidate to move to act and notify, and the lead decides.
The review design worksheet
Fill in one row per action before launch. If a row cannot be completed, the action is not ready for any mode except never.
| Action | Mode at launch | Approver and backup | Channel | Deadline and default | What is logged | Evidence to move it | Next review date |
|---|---|---|---|---|---|---|---|
| Send customer reply | Approve before | Support lead, ops manager | Team chat | Same business day, then escalate and hold | Message, draft, evidence, decision, approver | No edits across a full month of cycles, lead signs off | First of next month |
This worksheet is the human side of the scorecard in the AI agent use cases guide. A provider should ask you for it rather than write it for you.
Frequently asked questions
Does human in the loop make an AI agent slower than doing the work ourselves?
Only if you gate the wrong actions. Approve before on a customer send costs one click on a draft that is already written and already checked against the record. Approve before on every classification costs a person their afternoon and produces reflex approvals. If you gate by reversibility and blast radius, most of the workflow runs without waiting.
Can the agent decide for itself when to ask a person?
An agent can be instructed to escalate when it is unsure, and it should be, but that instruction is a request inside a prompt rather than a control. The actions that must wait for a person should wait because the system enforces it, following the complete mediation principle OWASP describes. Use the agent's own judgment as a second safety net and never as the only one.
What is the difference between human in the loop and human on the loop?
People use the terms loosely. Broadly, in the loop means a person approves an action before it happens, and on the loop means the action happens while a person supervises the results and can intervene. In the four mode model, that is approve before versus act and notify, or act and log with sampling. Most real workflows use both, for different actions.
When can we remove an approval step?
When the evidence supports it: a low edit rate on the approvals and a low miss rate on the samples, sustained across enough cycles for that workflow, with a named person signing off. Write the rule down before launch, move one step at a time and keep the ability to move back.
Is approve before enough on its own?
No. It is one of three layers. The agent's access has to be scoped so it cannot reach what it should not, every action has to be logged and rate limited, and the approval gates sit on top of both. An approval gate does little if the agent behind it has unrestricted access.
Decide the gates before the tools
Whether to have a human in the loop is rarely the open question. What you need to settle is which actions get which kind of oversight, who checks them and what evidence would change that. Answer it per action, write down the review design and revisit it monthly. Your agent will do more on its own over time, and every expansion will have a record behind it.
If you buy a managed service or build your own, the approval rules are still yours to own. Nobody outside your business can decide what your customers should hear from a machine, or who signs off on a refund. The BankrBot incident shows what happens when nobody makes that decision.
Originally published at vantasoft.com.
Top comments (1)
Treating oversight as a decision per action instead of per agent is the right framing. The four modes map neatly onto reversibility and blast radius.
The part I'd stress is the one you flag at the end: approval by reflex. If 98% of approval requests are fine, people stop reading by the second week. A few things that help: show the diff of what will change rather than the agent's explanation of it; rotate in occasional known-bad test actions and measure whether reviewers catch them; and track approval time. When the median drops to two seconds, the step has become a rubber stamp.
How do you decide when an action can move down a level, from "approve" to "act and notify"? A track record of N clean approvals, or something more formal?