Here is a rule that a container cannot enforce, and that almost every team running an agent eventually wants:
Let the agent post progress updates to the incident channel, but never more than three times in ten minutes.
Nothing about that rule is about reach. The agent already has the Slack tool. The container has already allowed it. The rule is about behavior over time, and a sandbox has no opinion about that. So does this next one:
Let the agent push, but only if the tests are passing.
That is not an isolation rule either. It is a rule about the state of the world at the moment of the action. A fence cannot express it.
This post is about the gap between those two things, and what a policy layer that actually closes it looks like. The concrete example throughout is AWS's Strands Box, released in developer preview on October 7–8, 2026, plus the policy language it embeds — but the gap it addresses is not vendor-specific. If you only remember one sentence, make it this: isolation answers "what can it reach"; policy answers "what may it do". It is easy to have the first and believe you have the second.
The gap, stated plainly
The usual answer to "my agent did something I did not want" is a sandbox. Sandboxes are good, and you should have one. But look at what a sandbox actually decides: which paths are mounted, which hosts resolve, which syscalls are permitted. All of that is decided before the agent runs, and none of it depends on what the agent is doing.
Now look at the failures that actually hurt. An agent investigating a production incident reads the right logs and posts updates — fine — then posts forty more and buries the human responder. An agent with a payments tool refactors something and calls the API two thousand times. An agent with a git push tool pushes before the tests finish. In each case the sandbox worked exactly as designed: the agent reached only what it was allowed to reach. The problem was never whether it could reach the tool. It was what it did with it.
That is the class of bug isolation is structurally unable to fix, because isolation is a statement about the environment, not about the trajectory.
Two classes of rule a sandbox cannot hold
Strip it down and there are two, and they are worth naming separately because they fail for different reasons.
1. Cumulative and rate-limited constraints. "At most three per ten minutes." "At most one hundred dollars per day." These require memory of previous actions. A sandbox is stateless by construction; it evaluates the current call against static rules and has no idea what the agent did an hour ago. You can sometimes approximate a rate limit with a proxy or a quota on the tool itself, but then the constraint lives in the tool, scattered across as many places as you have tools.
2. Conditional, cross-tool constraints. "Block outbound HTTP after the agent has read from the customer-data directory." This one is worse, because the two halves go through different tools. The read happens through the filesystem; the request goes through the network. A sandbox sees two unrelated operations. Expressing the rule requires a layer that watches both and can connect them.
If your policy language can only talk about a single call in isolation, it cannot express either class. Which means your "policy" is really just more isolation with better branding.
What replaces it: policy evaluated at the point of action
The design that closes the gap puts a policy engine in the path of every action, and — this is the important part — gives it a shared history.
Strands Box does this by using the OS sandbox not as the policy layer but as a funnel. The kernel-level isolation makes sure the policy layer is the only way out, and then the policy itself is evaluated in user space, where it can be aware of the system, the protocol, and the language. Concretely, it intercepts at four points: network egress, the shell interpreter, the Python interpreter, and a broker for Model Context Protocol (MCP) tools. Those four cover "what a coding agent actually does all day": run a command, run some generated code, call an API, call a tool.
The piece that makes the stateful rules possible is a unified event vocabulary across the enforcement points. A file read through a shell command and a file read from a Python script both report as the same kind of event. An HTTP request from curl and one from generated Python are both the same kind of event. Once every interpreter reports actions in the same shape, you can write a rule like "after the agent reads a file from the customer-data directory, block further outbound HTTP requests" without naming which tool did the read. That single property is what turns a pile of per-tool checks into one coherent policy.
And because the history is shared across enforcement points, a rule can reach from an earlier action through one tool to a later action through another. "Three Slack posts per ten minutes" is enforceable not by trusting the agent to count, but because the layer itself counts.
Note the shape of the enforcement, too. The decision is made outside the agent's own reasoning, which is the entire point: the agent cannot be talked out of it. A model under prompt injection can be persuaded that a rule does not apply. A policy engine that sees only "the agent is asking to make an outbound HTTP request, and it already read the customer directory" cannot be persuaded of anything. This is why moving the check from the prompt to the operating system is a category change, not a tuning change.
Where the credentials go
One more design point worth stealing even if you never run this specific tool: keep secrets out of the agent's environment.
If a tool needs an API token, the naive design puts that token in the agent's context so the tool call can carry it. That is a token the model can leak — through a log, a summary, an outbound request it was not supposed to make. The better pattern is substitution at the gateway: the agent's context holds a placeholder, and the enforcement layer swaps in the real secret at the moment the request passes through it. The agent never holds the credential, so the credential cannot be exfiltrated by the agent.
That is the same principle as the stateful rules, applied to data instead of actions: the thing that must be trusted should be small, and the agent should never be the thing holding it.
Three things that will bite you
1. Not everything the agent touches goes through the policy layer. If you grant a path directly in the configuration, it is bounded by containment and does not appear in the policy history. So a rule of the form "after the agent reads X, deny Y" will silently not fire if X was reached through a directly granted path rather than through an intercepted tool. Silent non-application is the worst failure mode for a policy, because it looks like it is working. Audit which of your rules depend on events that might be bypassing the engine.
2. Platform coverage is the constraint, not the idea. At launch this runs on macOS only, with Linux and Windows framed as roadmap. If your agent runs in Linux containers in production — and most do — you are looking at a development-machine control today, not a production control. Do not let a developer-preview demo make you believe your production agents are governed.
3. In developer preview, you own the maintenance. Open source here means you take on configuration, update, and patching as your own responsibility. That is not a theoretical worry: a related AWS agent tooling component had a documented bypass of its human-consent gate in the past. Treat agent infrastructure as a dependency with a patch cadence and an owner, not as a set-and-forget library.
What to do this week
You can act on all of this without adopting any specific product.
- Inventory your "approve everything" agents. Which of your agents currently run in a mode where every action is auto-approved? That list is your exposure surface, and it is usually longer than people expect.
- Write down the one stateful rule you actually need. Not a policy framework — one sentence. "No more than N per window." "Deny outbound after reading directory X." If you cannot name it, you do not have a policy problem yet; you have an observability problem.
- Check whether your logs distinguish attempted from permitted. A log that records what the agent did is half a log. You want to see what it tried and whether a policy allowed or blocked it. If a block leaves no line, you have no evidence that your control exists.
- Rehearse the deny path. Pick a rule, trigger the denial on purpose, and confirm it actually stops the run rather than merely writing a line. A deny that logs and proceeds is not a control; it is a note to self.
The one-sentence version
A sandbox tells you what your agent could touch. A policy layer tells you what it was allowed to do, and refuses the rest. The failures that hurt are the second kind, and they were never expressible in a fence.
This post was written with AI assistance. The author is responsible for its content.
Top comments (0)