AI agents are moving from request-response tools to systems that can keep working after the user leaves.
That shift changes the security model.
A chatbot produces text for a person to review. An agent may read email, edit files, call an API, deploy code or initiate a transaction. Once a model can take actions, prompt quality is no longer enough.
The system needs a permission architecture.
Prompts are not access control
Instructions such as “never send an email without permission” are useful, but they are not a reliable security boundary.
The model is still interpreting language. It may encounter ambiguous instructions, untrusted content or conflicting context. A privileged action should therefore be enforced outside the model by a deterministic execution layer.
The model may propose an action. The surrounding system should decide whether that action is allowed.
1. Give each agent its own identity
Avoid letting multiple agents share one broad service account.
A separate identity makes it possible to answer basic operational questions:
- Which agent requested the action?
- Which user or workflow initiated it?
- What resources was it allowed to access?
- Can its access be revoked without affecting other agents?
Agent identity should be treated like workload identity, not like a convenient shared API key.
2. Separate reading from writing
Reading a calendar and deleting an event are not equivalent permissions.
Tool definitions should separate capabilities according to their consequences. For example:
email.reademail.draftemail.sendemail.delete
An agent may receive read access by default while send or delete operations remain behind an approval step.
This also makes logs and security reviews much easier to understand.
3. Gate actions by consequence
Not every tool call needs human approval. Requiring confirmation for everything removes much of the value of automation.
Approval should be based on impact.
A useful starting point is:
- Reversible, internal action: allow automatically
- External but recoverable action: allow with notification
- Irreversible or high-impact action: require approval
- Action outside the agent’s defined scope: deny
The policy should live in code or configuration, not only in the system prompt.
An illustrative policy could look like this:
action: send_external_email
default: deny
requires:
- scoped_agent_identity
- human_approval
log:
- initiating_user
- destination
- proposed_content
- policy_decision
4. Keep an audit trail that explains actions
Logging only the final API request is not enough.
A useful audit record should connect:
- The initiating user or scheduled workflow
- The agent identity
- The requested objective
- The tool and resource accessed
- The policy decision
- The resulting state change
The goal is not to store every hidden reasoning token. It is to preserve enough evidence to reconstruct what the system attempted, what was authorized and what changed.
5. Design revocation and rollback before launch
Every long-running agent needs a clear stop mechanism.
Teams should be able to:
- Revoke the agent’s credentials
- Disable an individual tool
- Cancel active work
- Restore affected data where possible
- Prevent retries after a failed or denied action
A kill switch without tested recovery procedures only solves half of the problem.
6. Avoid coupling permission policy to one model
Model providers and model versions change. Permissions should not have to be redesigned every time the underlying model changes.
Keep these layers separate:
- Model selection
- Business workflow
- Tool execution
- Permission policy
- Audit and recovery
This separation also makes it easier to evaluate another model without accidentally changing the agent’s authority.
A pre-launch checklist
Before enabling an always-on agent, ask:
- Does it have a unique identity?
- Are read and write permissions separated?
- Are irreversible actions approval-gated?
- Can every external action be traced to a user and policy decision?
- Can access be revoked immediately?
- Has rollback actually been tested?
- Can the model be replaced without rewriting the permission layer?
The most capable agent is not automatically the most useful one.
For production systems, trust comes from knowing what the agent can do, what it cannot do and exactly where it must stop and ask.
Which action in your current workflow would require a mandatory human approval gate?
AI disclosure: This article was developed with AI assistance and reviewed by the author before publication.
Top comments (0)