Your AI agent can modify a database, issue a refund, or change production infrastructure. What stops it from doing the wrong thing?
Recent AI agent incidents show the same problem: agents can reach production and execute actions their operators never intended.
Different incidents. Same missing boundary.
Five failure modes
1. Runaway execution
An agent gets stuck in a loop and keeps calling tools. A token budget isn't an execution budget. Without a hard limit, a small loop can become a large bill.
2. Destructive tool calls
An agent may have access to tools that can change production state. One wrong call can become an incident.
3. Indirect prompt injection
Agents process untrusted content: tickets, documents, emails, web pages. If that content changes the agent's behavior, execution boundaries limit the damage.
4. Incorrect tool calls
Agents can use wrong parameters, misunderstand tool results, or claim an action succeeded when it didn't. Production systems shouldn't rely on the model to decide whether a consequential action is safe.
5. Missing approval boundaries
Reading data is different from deleting it. Drafting a refund is different from issuing one. Sensitive actions need approval tied to the actual action, not just a step in the workflow.
6. Missing evidence
When something goes wrong, you need to know:
- What action was requested?
- With which parameters?
- What policy applied?
- Was it allowed, blocked, or approved?
- What actually happened?
Logs are useful, but recording the execution decision is different.
The missing boundary
Agent frameworks are good at orchestrating models and tools.
But orchestration isn't authorization.
A framework can tell an agent:
Here are the tools you can use.
It doesn't necessarily answer:
Should this specific action be allowed?
If the agent can reach the tool, the action can happen.
A safer architecture puts a policy between the agent and the tool:
Agent
|
| action + parameters
v
Policy
|
+-- allow
+-- block
+-- require approval
|
v
Tool
|
v
Production
The policy evaluates the actual action.
- Not the prompt.
- Not the agent's intentions. The action.
This is an authorization problem
Agent behavior is probabilistic. Authorization shouldn't be.
The model can decide:
I want to issue this refund.
It shouldn't decide:
I am allowed to issue this refund.
The policy can evaluate the actual request:
{
"tool": "issue_refund",
"amount": 4000,
"customer_id": "12345"
}
What we're building
This is why we built NullRun.
NullRun sits between agents and the systems they can affect. It evaluates execution requests and can allow, block, or require approval before the tool runs.
It also records the decision, including blocked actions.
The goal isn't to make agents reliable.
It's to put reliable execution boundaries around unreliable decision-making.
The model can ask for anything.
It shouldn't be the authority that decides what gets executed.
Top comments (0)