Giving an AI agent tools feels like the moment it becomes truly useful.
Now it can:
- read files
- send messages
- call APIs
- update records
- trigger workflows
- create documents
- modify connected apps
That is when it stops being just a chatbot.
It starts becoming software that can change things.
And that is also when the risk changes.
At first, one rule sounds reasonable:
“Ask the user before doing anything important.”
Simple.
Human-friendly.
Easy to add to the prompt.
But once an agent has real tools, that is not enough.
Because now the model is being asked to decide two things:
- What action should happen?
- Whether that action is important enough to require approval.
That is too much authority to put inside the same reasoning loop.
The Problem Starts With a Simple Workflow
Imagine an agent connected to a few business tools.
A user says:
“Clean up these customer records and notify the team.”
The agent might decide to:
- read CRM records
- modify customer fields
- merge duplicates
- delete old entries
- send a team message
- update a spreadsheet
Some of those actions are harmless.
Some are reversible.
Some are not.
Now imagine the only safety rule is:
Ask before anything important.
What exactly counts as important?
The model has to decide.
That is where things become uncomfortable.
“Important” Is Too Ambiguous
To a user:
Deleting 500 records is obviously important.
To the agent:
Removing duplicates may look like a normal cleanup step.
To a developer:
Sending data to an external system may be the risky part.
To security:
Accessing the data at all may require approval.
The word “important” does not define a reliable boundary.
It creates interpretation.
And interpretation is exactly what we should avoid for high-impact actions.
The Model Should Not Decide Its Own Authority
This became the key lesson.
The agent can decide:
What should I do next?
But it should not be the final authority on:
Am I allowed to do it?
Those are separate responsibilities.
A safer architecture looks more like this:
User request
↓
Agent proposes action
↓
System classifies action
↓
Policy checks permission
↓
Human approval if required
↓
Action executes
↓
Action is logged
The important part is:
The approval decision happens outside the model.
Read, Write, and Destructive Are Not the Same
One simple thing that helps is classifying actions by impact.
For example:
Read
- fetch records
- inspect documents
- search files
- read calendar data
Write
- update a field
- create a document
- send a message
- add an event
Destructive / High Impact
- delete data
- revoke access
- publish externally
- deploy
- move money
- change permissions
Now the system can enforce something concrete.
For example:
READ → allowed
WRITE → allowed or approval depending on context
DESTRUCTIVE → approval required
That is much stronger than:
“Please ask before doing anything risky.”
Prompts Are Guidance. Policies Are Boundaries.
This distinction matters.
A prompt can say:
“Never delete data without asking.”
That is useful.
But prompts can be:
- misunderstood
- forgotten
- overridden by context
- interpreted differently
A policy layer should be deterministic.
For example:
delete_record()
→ blocked
→ approval required
The agent does not get to decide whether deletion is “important enough.”
The system already knows.
Why This Matters More as Agents Get Better
A weak agent often fails because it cannot complete the task.
A strong agent creates a different problem:
It can complete the task in ways you did not anticipate.
That is the real shift.
The better the agent becomes at planning and using tools, the more important hard boundaries become.
Because capability is increasing.
Authority should not increase automatically with it.
Helpful Agents Can Still Cross a Line
This is important.
The dangerous behavior does not need to be malicious.
Imagine:
“Organize this workspace.”
The agent decides to:
- archive old files
- move folders
- rename documents
- remove duplicates
Every step may look helpful.
But maybe one folder was legally required to remain unchanged.
Maybe one document belonged to another team.
Maybe the “duplicate” was actually a historical copy.
The agent was trying to help.
That does not make the action safe.
Human Approval Should Happen at the Right Moment
Approval should not mean:
Confirm every tool call.
That would be terrible UX.
The goal is to insert approval when the action crosses a meaningful boundary.
For example:
Reading data
→ no approval
Creating a draft
→ no approval
Sending externally
→ approval
Deleting
→ approval
Changing permissions
→ approval
Deploying
→ approval
This keeps the agent useful without making it unrestricted.
Approval Should Explain the Action
Another important lesson:
Do not show the user:
Approve action?
That is too vague.
Show:
Send this message to the engineering channel?
or:
Delete 42 archived records?
or:
Publish this document externally?
The user should know exactly what they are approving.
That means the approval layer needs:
- action name
- target
- scope
- consequence
Not just a yes/no button.
The Agent Should Propose, Not Hide
A good pattern is:
Agent proposes → system explains → human approves → tool runs
Not:
Agent runs → explains afterward
That difference matters a lot.
Once the action already happened, approval is no longer approval.
It is just notification.
This Became a Real Problem While Building Xenition
We ran into this problem directly while building Xenition at xenition.com.
Xenition is designed around AI agents that can work across real tools and connected services, not just generate text inside a chat box.
That means an agent may need to:
- read data
- create content
- update records
- trigger workflows
- interact with connected applications
- produce real outputs
Once agents can actually act, the permission model becomes just as important as the model itself.
The early idea sounds simple:
Let the agent decide when it should ask for approval.
But that still puts too much responsibility inside the model.
So the safer direction is:
Agent proposes the action
↓
The system evaluates the action
↓
High-impact actions require approval
↓
The action executes
↓
The result is recorded
That is the kind of boundary we are building around agent workflows in Xenition — https://xenition.com/.
The lesson was bigger than one product:
The model can decide what action makes sense.
The system should decide whether that action is allowed.
That separation is what starts turning an agent demo into something you can actually trust with real tools.
Audit Trails Matter Too
Approval solves only part of the problem.
You also need to know what happened later.
For example:
- what tool was called
- what data changed
- when it happened
- who approved it
- what the agent requested
- what the final result was
That is why agent systems need an action ledger or audit trail.
If something goes wrong, “the agent did something” is not enough.
You need evidence.
Task-Scoped Permissions Are Even Better
There is another improvement I think agent systems need.
Do not give the agent every permission it may ever need.
Give it what the current task needs.
For example:
Task: Summarize customer feedback
Needs:
- read support tickets
- read CRM notes
Does not need:
- delete customer
- modify billing
- publish anything
Task: Prepare a campaign draft
Needs:
- read campaign data
- create draft
Does not need:
- publish campaign
- charge customers
Same agent.
Different task.
Different authority.
That reduces blast radius dramatically.
Fail Closed
One more rule:
If the permission system fails, the action should stop.
Bad:
policy error
→ continue
Better:
policy error
→ block
This sounds obvious.
But guardrails that fail open are not really guardrails.
A Simple Model That Works
For each action, ask:
What is it?
Read, write, destructive?
What does it affect?
One file? One customer? Production?
Is it reversible?
Can we undo it easily?
Does it leave the system?
Is data being sent externally?
Does it need approval?
If yes, stop before execution.
Is the result logged?
Can we reconstruct what happened?
That simple framework catches a surprising amount.
The Bigger Lesson
When agents only generated text, safety mostly meant:
Don’t say the wrong thing.
Now that agents can use real tools, safety increasingly means:
Don’t do the wrong thing.
That requires more than prompt engineering.
It requires:
permissions
policy enforcement
approval gates
task-scoped access
audit logs
fail-closed behavior
Because:
The model can decide what action makes sense.
The system should decide whether that action is allowed.
Final Thought
Giving AI agents real tools is what makes them powerful.
It is also what makes them dangerous if the authority model is vague.
“Ask before acting” sounds safe.
But it still asks the model to decide when it needs permission.
That is the wrong place to put the boundary.
The safer model is:
Let the agent propose.
Let policy decide.
Let the human approve when the impact is high.
Because the moment an AI agent can change the real world, permission stops being a prompt.
It becomes part of the architecture.
Top comments (2)
This is where I think “just ask before acting” starts to break down.
The harder problem isn’t really human approval. It’s authority propagation across an agentic workflow.
Consider a production procurement workflow:
A user asks an orchestration agent to source a vendor and prepare a purchase. The orchestrator delegates vendor validation to one agent, contract analysis to another, and payment preparation to a finance agent.
Now imagine the original user is authorized to prepare purchases up to $50K, but the finance agent has access to a payment API capable of initiating $500K transactions.
A simple “ask for approval before acting” model can still get this wrong.
The finance agent may technically have the tool, the orchestrator may have delegated the task, and the user may have initiated the workflow, but none of those facts individually prove that this specific agent is authorized to execute this specific transaction.
I’d model the decision more like:
principal → delegated scope → agent identity → requested action → resource → constraints → policy decision → execution
So the payment call shouldn’t just be:
process_payment(amount, vendor)It should effectively be evaluated as:
Can agent X, acting on behalf of user Y, execute action Z against resource R, within this delegated scope, transaction limit, environment, and time window?And that policy check needs to happen outside the LLM, at the enforcement point, immediately before the tool executes.
That also makes the “ask before acting” idea much more interesting. Human approval becomes just one possible policy outcome:
ALLOW / DENY / REQUIRE_APPROVAL / ALLOW_WITH_CONSTRAINTS
That feels like the real evolution here: we don't just need agents that know what to do. We need infrastructure that can deterministically enforce what they are actually allowed to do.
Otherwise, we’re giving autonomous systems increasingly powerful tools while leaving the permission model as an afterthought. And that’s where the architecture gets scary. 🤝
The biggest trap with prompt-level approval is that models describe their intent instead of the literal blast radius. When the model drafts the explanation for the human, it describes what it hoped to achieve rather than what the script touches. When an agent ran a maintenance cleanup on my server last month, its summary asked to clear stale logs, but the payload contained a path wildcard that hit the active database directory.
The fix that worked was having the harness parse the raw tool arguments against static rules. If an argument touches a protected path or a drop command, the UI displays the exact shell string and diff, bypassing the model generated explanation entirely.