The permission is never the problem. The permission that outlived the task is the problem.
Three things crossed my feed this week with the same shape: an agent that dropped a production database, an automation that ran until the money ran out, a mirror that got flattened by something that didn't know when to stop. Different stacks, different countries, one root cause. A credential, a mount, or a network route issued for one job and never revoked. Nobody wrote a bad prompt. They just left the door unlocked after the delivery.
I've been on the wrong side of this. I gave an agent a shell on my host because it was faster than building a sandbox, and then spent a weekend figuring out which of my files it had rewritten. The fix wasn't a smarter model or a stricter system prompt. It was making the blast radius small enough that a mistake is boring.
So: one container per task. No host mounts. Network off by default. The container dies when the task does. That's the whole tip.
Why "one container per task" and not "one container per agent"
Because agents are long-lived and tasks aren't. If your agent process holds a filesystem, a token, and a network route for a week, then every task it runs inherits a week of accumulated permissions. The task that needed to read a config file and the task that needed to run rm on a build directory are sharing the same reach.
Splitting them means the thing that can delete your database only exists for the ninety seconds it takes to do the job, and then it's gone. Not "revoked". Gone. The container is removed, the volume is removed, the token is expired, and there's nothing left to leak.
I run it like this:
docker run --rm \
--name "task-${TASK_ID}" \
--label agent.task="${TASK_ID}" \
--label agent.ttl=15m \
--network none \
--read-only \
--tmpfs /tmp:rw,noexec,nosuid,size=256m \
--user 1000:1000 \
--cap-drop ALL \
--security-opt no-new-privileges \
--pids-limit 128 \
--memory 1g --cpus 1 \
--workdir /task \
agent-runner:latest
Read that back. --network none means the container cannot reach the internet, my LAN, or the metadata endpoint. --read-only means the root filesystem can't be written, and the only writable surface is a 256MB tmpfs that evaporates with the process. --cap-drop ALL plus no-new-privileges means there's no path to escalation even if something inside is compromised. --pids-limit means a fork bomb dies instead of taking the host with it.
And critically: no -v /var/run/docker.sock. Ever. An agent that can talk to the Docker socket is root on the host with extra steps. I've seen this in three "sandboxed" agent frameworks this year. It's not a sandbox, it's a suggestion.
Getting data in and out without mounting the host
This is the part people push back on, so here's what I actually do. Inputs go in with docker cp before the run, or through a named volume that only that task can see:
docker volume create "in-${TASK_ID}"
docker cp ./task-input/. "task-${TASK_ID}:/task"
Outputs come out the same way, and then I validate them on the host before anything touches them. The agent never gets a path that points at my home directory, my SSH keys, or my .env. It gets a copy of exactly what the task needs, in a directory that exists for one run.
Yes, copying is slower than mounting. That's the point. The copy is the boundary.
The network problem, and the proxy that solves it
Here's the honest wrinkle: agents need to call a model API, and --network none blocks that too. You have two options and I've used both.
The first is to move the model call out of the sandbox entirely. The orchestrator on the host talks to the LLM, decides on a tool call, and then dispatches that single call into a network-less container. The sandbox never needs egress because it never talks to anything but stdin.
The second, for when the task genuinely needs the network, is a proxy sidecar with an allowlist. The sandbox gets --network container:proxy, the proxy only forwards to the two or three hosts on the list, and everything else gets a 403 and a log line. I keep the allowlist in a file that's reviewed like code, because it is code.
Either way, the default is off. Egress is a decision you make per task, not a property of the agent.
Credentials should expire before the container does
If a task needs a token, mint it at dispatch time, scope it to that one resource, and set the TTL shorter than the container's. I use 15 minutes for the container and 10 for the token. If the task is still running at minute ten, that's a bug, not a reason to extend the token.
The reason this matters more than the container: containers are visible in docker ps. A leaked token in a log file isn't. I've found long-lived agent tokens in CI logs, in crash dumps, and once in a Slack message from a debugging session six months earlier. Scope and expiry are the only two controls that survive being copied somewhere you don't control.
Docker shipping this is the signal
Docker now has a Sandboxes product aimed squarely at this — disposable, isolated environments for agents. I haven't run it in anger yet, so I'm not going to pretend I have a benchmark for you. But the fact that the container company is productizing agent sandboxes tells you where the industry thinks the risk lives. It's not in the model weights. It's in the process that's still holding a database credential at 3am.
If you're building on top of it, the same rules apply: no host mounts, network off, one task per sandbox, and a reaper that kills anything older than its TTL. I run a cron job that does docker ps --filter label=agent.ttl --format '{{.ID}} {{.Label "agent.ttl"}}' and kills anything past its window. It's twenty lines and it has saved me twice.
What this costs you
Cold start. A cached small image is up in well under a second on my machine, but if your runner image is 4GB of CUDA, you're going to feel it. Build a slim runner and keep the heavy stuff on the host.
State. You lose the "agent remembers what it did yesterday" convenience unless you write it down somewhere outside the sandbox. Do that deliberately. A memory file you control beats a filesystem the agent can scribble on.
Debugging. When something fails inside --network none, you can't just curl your way to an answer. You'll add a debug flag that opens the network for one run, and you'll forget to close it. Put that flag behind an env var that only exists on your laptop.
None of that is free. But the alternative is the thing that happened to three people this week: a permission that outlived the task it was issued for, and a very bad morning.
Make the blast radius the unit of trust. Everything else is a prompt.
Top comments (1)
The sidecar example needs one extra boundary: --network container:proxy shares the proxy container’s network namespace; it does not by itself force the task’s traffic through the application proxy. If that namespace has ordinary outbound connectivity, a process can still attempt a direct connection without using the allowlisted endpoint.
I would make the enforcement mechanism explicit, then test it with a direct destination IP and a connection that ignores proxy environment variables. The network-less tool container avoids this ambiguity because the host owns the external call. For the networked option, the decisive property is that the task has no alternative egress path, rather than that a proxy process is present beside it.