If you updated Docker Desktop recently — version 4.63 or later — a new command is already installed on your machine: docker agent. No setup, no announcement you had to read. The repository describes it as a way to "build, run, and share AI agents with a declarative YAML config," it speaks to OpenAI, Anthropic, Gemini, Bedrock, Mistral, xAI, and Docker's own model runner, it takes "any MCP server (local, remote, or Docker-based)," and it ships preinstalled. Docker says they use it to build itself. That is several million dev machines quietly gaining an agent runtime, one auto-update at a time.
I spent an evening in the repo and the docs instead of writing another framework take, because this project sits exactly on the fault line I keep coming back to: the gap between what an agent can read and what it's allowed to do. Two years ago that gap was my problem — I wrote a seatbelt hook for my own coding agent, attack-tested it, and ended the post at the only wall that held: the operating system. Docker just shipped a productized version of that wall. The engineering is genuinely good. The default is the story.
What actually ships
The model is Claude Code's converted to YAML: agents as config, not code. A toolsets block declares what the agent gets — type: filesystem, type: shell, type: mcp for anything from Docker's MCP catalog ("hundreds of MCP servers," the docs say, attachable with a ref: line). There is a permissions page, a secrets guide, OpenTelemetry tracing. The README reports anonymous telemetry and an Apache-2.0 license, 10,000-plus commits of movement.
Read as a positive policy, the toolset block is the right idea — it is the same inversion my seatbelt's second version went through: stop enumerating what's forbidden, declare what's allowed, let the unlisted fail closed. A runtime where the tool surface is a diffable YAML file is a runtime you can review in a pull request. Credit where due.
The wall, as documented
The sandbox mode docs describe something above my pay grade of hook scripts. Enable it and the agent runs inside a Docker Sandboxes VM — not a raw container, an orchestrator around a dedicated sandbox CLI. The mount table is explicit: the working directory read-write; the agent config and kit directories read-only; everything else invisible — "other host files are not visible to the agent." The VM gets its own $HOME.
The network story is the part I would have asked for and didn't expect: a default-deny egress proxy, with a per-run allowlist covering the model gateway and package registries, extensible only by declared hosts (runtime.network_allowlist) or a persisted docker agent sandbox allow. Default-deny egress is the control that actually contains a prompt-injected agent, because the exfiltration path dies at the proxy regardless of what the model decided. Cloud mode implies the sandbox and refuses to upload host API keys.
This is the wall. It is better designed than anything I bolted together, and I want it on.
The finding
--sandbox defaults to false.
Read the docs' own phrasing: enable it per run (docker agent run --sandbox), bake it into the YAML (runtime: sandbox: true), or attach it to an alias. Every path is opt-in. Out of the box, docker agent run agent.yaml runs on the host — where type: shell and type: filesystem mean what they have always meant on a laptop: your user account is the blast radius. The sandbox with its read-only mounts and its default-deny proxy exists, documented, well-built — in the glovebox.
Defaults are policy. Every mainstream harness that ships a safety mode behind a flag has bet that users read docs, and the track record of that bet is the entire genre of "my agent deleted the database" posts. dev.to already has one for the neighboring toolchain — an agent deleting every container on a machine through MCP tool permissions — and Docker's MCP catalog makes attaching that power a one-click, YAML-declared convenience. Convenience is not a vulnerability. Convenience is what decides whether the vulnerability gets exercised.
Where the doors remain
Even with the sandbox on, three doors don't close, and the docs are honest about at least one of them:
The tool supply chain is context supply chain. Isolation contains what code does; tool poisoning attacks what the model reads. An MCP server inside a perfect VM still returns tool descriptions that carry instructions — the model consumes them as context, and no mount table filters text. The catalog's one-click attach is the new blocklist problem from my seatbelt post, one layer up: the trust decision moved from the shell command to the tool definition, and most reviews of an agent YAML stop at the toolset names.
Secret redaction is best-effort, in Docker's own words — the docs note auto-kit redaction may let obfuscated tokens slip through. Combined with "previous sandboxes are never deleted automatically," that means tokens can persist in stopped sandbox state — the docs remind you removal takes sbx rm --force, and that stopped sandboxes may still incur storage charges. Sensitive work needs the cloud mode's rule — no host keys uploaded — or nothing at all.
Mutable tags break reproducibility, so the docs say to pin remote kits and images by digest. Correct advice, and nobody's default.
And one layer is missing entirely, not just loosely configured: the docs ship OpenTelemetry tracing, which answers what happened — to a collector you control, in a format you can't sign. Nobody's threat model at Docker includes "the operator edits the transcript afterward," which is fair for them and wrong for anyone running agents that touch production. A trace is a debug story. An audit needs a record you can prove was never edited. That layer does not ship here — or anywhere mainstream yet.
What to steal
# The seatbelt stays ON — every path to a sandboxed run, in order of safety:
docker agent run --sandbox agent.yaml # per-run, explicit
# baked in, so it survives copy-paste of the YAML:
# runtime:
# sandbox: true
# attached to the alias your team actually types:
docker agent alias add safe-coder myorg/coder --sandbox
# pin everything the agent will pull:
# pin remote kits/images by digest (docs' own advice)
# and clean up state that never cleans itself:
docker agent sandbox ls
sbx rm --force <name> # stopped sandboxes persist
Plus the two rules the YAML cannot express: never let type: shell and type: filesystem ship together in a default config you didn't review, and read an MCP server's tool definitions like code — because that is what they are now.
Two honest limits
This is a source read, not a test. I did not run docker-agent for this post; every claim about the sandbox comes from the README and docs, which means the docs' accuracy is an assumption I am lending them. The gap between a sandbox's documentation and its behavior is where I found four bypasses in my own hook, so treat my assessment as "the design is right," not "the wall holds."
And the critique has a shelf life. Ten thousand commits and preinstalled distribution means Docker can flip the sandbox default in one release — which would make the headline of this post obsolete in the best way. I would genuinely rather be wrong by Tuesday.
Your turn
Did your Docker Desktop ship with docker agent yet? Did you enable the sandbox before or after reading this — and if you have run it both ways, what did the sandbox block that the host run would have done? The seatbelt thread continues in the comments.
Top comments (7)
sam, this is a brilliant and necessary teardown of the new docker agent architecture. your point that "defaults are policy" is the most critical takeaway. a sandbox that exists but defaults to off is just a false sense of security waiting for a copy-paste error.
as someone building a secure js sandbox for ai code execution (koda), the "fail closed" (default-deny) philosophy is non-negotiable. if an agent can't be trusted, the environment must assume hostility from line one. on constrained hardware like a $150 phone, a runaway agent making unauthorized network calls or consuming resources is catastrophic, making strict isolation even more vital.
your insight that "the tool supply chain is context supply chain" is spot on. even in a perfect vm, if an mcp server's tool description contains a prompt injection, the model will execute it. isolation contains the blast radius, but it doesn't fix the poisoned context.
given that pinning by digest is the only real fix for mutable tags, do you think orchestrators should eventually enforce strict digest pinning at the parser level for any remote kit/mcp reference, rather than just recommending it in the docs?
fantastic write-up. the seatbelt thread continues! 🐯
Short answer: yes — and the argument is the same one as "defaults are policy": a docs-level recommendation is a suggestion, and the whole week's lesson is that suggestions get talked out of. Digest pinning enforced at the resolver is the first chokepoint in the chain that can't be negotiated with: a remote ref without a digest gets refused — and the refusal is journaled as an event, because a blocked resolution with a decision id is a finding, not a log line. Grade it by blast radius: warn-and-record on tag refs, hard-fail on latest, hard-fail on any remote ref inside a run that has egress.
One addition your sandbox instincts will like: pinning fixes mutability, not provenance. A digest faithfully points at whatever the publisher pushed — malware included. And a compose file the agent itself generated, digest-pinned and all, is still self-attestation: the same writer wrote the pin and the judgment. So parser-level enforcement pairs with the two-writer rule — the pin is proposed by the config, and something other than the proposing process decides it's allowed. Isolation contains the blast radius; the receipt says what was allowed to land inside it. Your $150 phone is a genuinely good adversarial lab for that boundary — the hardware can't afford a wasted escape, so every control has to earn its place.
🐏
sam, the "two-writer rule" is a massive insight. you’re absolutely right: parser-level enforcement fixes mutability, but it doesn't fix provenance. separating the proposal (the config) from the approval (a distinct, non-agent process) is the exact separation of duties needed to prevent self-attestation loops.
and i have to say, reading that my "$150 phone is a genuinely good adversarial lab for that boundary" is one of the best compliments i’ve received as a builder. when you have zero margin for error or wasted resources, every single control truly has to earn its place. it forces ruthless architectural discipline.
thank you for taking the time to break this down so clearly. the "warn-and-record vs. hard-fail" grading by blast radius is going straight into my mental model for koda’s sandbox. 🐯🛡️
That discipline — every control earns its place — is the thing most teams only learn after an incident, and you're starting with it, which is the unfair advantage. One habit to keep as KODA grows: break your own sandbox before anyone else does, and write down what fell out — the failures you publish become the spec everyone else trusts. Build well, Harun. 🐯
sam, "break your own sandbox before anyone else does" is going straight into the koda engineering handbook.
you’re absolutely right: publishing the failures is what actually builds the spec and the trust. treating the sandbox like a chaos engineering target from day one, rather than waiting for a post-incident review, is the only way to stay ahead of the curve.
thank you for the mentorship-level advice and for taking the time to engage so deeply on this. it means a lot coming from someone with your experience.
build well! 🐯
"Defaults are policy" is the whole post for me.
Someone made the same point on my coding agents post yesterday: a sandboxed, read-only agent can still leak things through one "helpful" curl. Default-deny egress kills that path no matter what the model decided, so it really should be the default.
The MCP part worries me more though. Tool descriptions are text the model reads as instructions, and most reviews of the YAML stop at the toolset name. Did you see anything in the docs about pinning or diffing the tool definitions themselves, not just the image digest?
Straight answer: no — the docs pin the image, and the tool definitions ride outside the pin. A pinned image pins the server code, which constrains what it can serve, but tool definitions are produced at runtime: config-driven, env-dependent, and in the remote case not pinned at all — an SSE/HTTP MCP reference is a URL, and whatever it serves this morning is what your model reads. So the layer you're pointing at is the real one: the toolset is the lockfile nobody writes.
The shape I'd trust: capture the listTools response at registration, hash it, seal the hash as the reviewed baseline — then on every session start, diff the served definitions against it and surface a changed description as a finding, not silently. One caveat carried over from the image-pin argument: the baseline has to be sealed by whoever reviewed it, not by the process that connects — a baseline the connector can rewrite drifts with the thing it measures. Your curl example is the same philosophy one layer up, by the way: default-deny egress pins the network path, toolset diffing pins the context path — the two halves of the supply chain the model actually travels.