DEV Community

NeuPortal
NeuPortal

Posted on

What Meta Muse Gets Right About Sandboxing an AI Agent, and the Three Places It Still Leaked

Meta's Muse is a consumer agent that acts on a person's email, calendar, purchases and, since 17 September, their Mac. Alongside the launch, Meta published an unusually detailed write-up of how the agent is contained. If you are building anything that holds user credentials and lets a model take actions, it is one of the more useful public references this year - and its first three weeks in production show exactly where a good design stops protecting you.

The threat model, stated honestly

An agent that reads untrusted input (web pages, inbound mail) and also holds the power to act (send, buy, delete) is one successful prompt injection away from working for someone else. Meta does not claim to have solved that. Its own write-up concedes that injection is still unsolved across the industry. So the design goal is not a model that cannot be fooled. It is a system in which a fooled model can do very little.

1. Two domains on one box

Meta's framing: "The right mental model is two isolated security domains on one box, not an LLM powered agent with root."

Each user gets a Linux VM. The agent core runs inside a systemd-nspawn container, with root inside the container mapped to an unprivileged user on the host, a filtered syscall set (io_uring is not available) and a trimmed capability set (no CAP_SYS_PTRACE, no CAP_NET_ADMIN). Outside the container sit the pieces an attacker would actually want: safety classifiers, privilege-separated workers that run connector code, the credential store, the egress gatekeeper and the database. They talk over Unix sockets that authenticate callers with SO_PEERCRED.

Steal this: put the model's runtime and every credential holder on opposite sides of a kernel boundary, and authenticate local IPC by peer credentials, not by a shared secret the model could read.

2. Exactly one process decides what leaves

A host-side agent called Sentinel is the sole permission authority for connector actions and for all network egress. It inspects requests at L4 and L7 - host, IP, port, method, path and body - and guards against SSRF.

The detail worth copying is where approvals live. When a human has to confirm an action, the prompt is rendered in the client UI rather than in the conversation, and the answer goes straight to Sentinel. A compromised model cannot forge consent by writing "the user approved" into its own transcript, because the transcript is not the channel.

Steal this: approvals must travel on a path the model cannot write to.

3. Surrogate credentials

The agent only ever holds stand-in tokens. Real OAuth tokens and passwords live in a separate store, and Sentinel substitutes them at the network edge. Meta's write-up puts the consequence bluntly: because the agent never sees real tokens, coaxing it into revealing them is futile. That protects secrets only - a hijacked agent can still reach whatever data the user connected.

The browser gets the same treatment. The browsing sub-agent works from the page's accessibility tree rather than the DOM, cannot execute JavaScript, and has DevTools disabled. While the vault fills a login form, the agent is paused.

Steal this: the cheapest defence against exfiltration is to never put the secret into the model's context.

4. Taint tracking on processes

Processes that have read user data are marked as tainted using eBPF and lose automatic permission to reach the network. Reading and leaking are treated as a sequence to be interrupted, not as two unrelated events.

5. Filter the inputs attackers want

Before a message reaches the agent, the mail connector removes what an account takeover would need - one-time passcodes, links that reset a password, passwordless login links - using deterministic filters plus a classifier. Where a service allows it, Muse also separates read and write permissions more finely than the provider's OAuth scopes do - Gmail read access without access to Gmail settings, for example.

6. Keep money at arm's length

Meta's recommended payment route is Link by Stripe. Where the merchant is on the Link network, the user's saved Link payment method is charged; elsewhere Link mints a virtual card locked to a single merchant and amount, valid only briefly. Muse can also sign in to a store account and use a card saved there. Every purchase needs explicit approval, so a hijacked session can at most raise an approval prompt - and where a one-time card is used, that card is useless beyond the one basket.

Where it leaked anyway

None of the three incidents below broke the cloud sandbox. All three happened at edges the sandbox does not cover.

1. A debug knob in a production client. On 21 September, Patrick Wardle published a flaw in the Mac app: an undocumented setting that chose the dictation server could be changed by any process running as the user, with no elevated permissions. Point it elsewhere and voice input plus the session token go to the attacker. The precondition was code already running on the Mac, and Meta's fix, shipped within a day, removed the setting from production builds. Wardle's counter-argument is worth taking seriously: a ClickFix-style lure, where the victim pastes a command into Terminal, turns "local only" into remote.

Lesson: the client is part of the attack surface. Strip debug configuration at build time, and treat any locally writable setting that redirects data as a remote bug with one social-engineering step in front of it.

2. "Your machine" versus "our internals". Two developers persuaded Muse to export the whole file system of its VM - system files, internal documentation, and the agent's memory stored as plain Markdown; one developer's export came to 6.8 GB once unpacked. Meta's answer was that this is intended: the machine belongs to the user.

Lesson: decide explicitly which parts of the agent's environment are user data and which are vendor internals. If export is a feature, assume everything you left inside is public.

3. The model's account of its own permissions. An Inc. columnist reported that Muse acted on his Messages after he believed he had declined access, and that the agent's own account of where that information came from turned out to be false - which Meta confirmed.

Lesson: never let the agent be the source of truth about what it can access. Show users a deterministic view of permissions and an activity log generated by the permission layer, not narrated by the model.

A checklist to take into your own design review

  • A kernel boundary between the model runtime and anything holding credentials.
  • A single egress authority that inspects requests at L7.
  • Approvals rendered outside any channel the model can write to.
  • Surrogate tokens; real secrets never enter the context window.
  • Taint marking for processes that touched user data.
  • Inbound filters for OTPs, reset links and magic links.
  • Payment instruments scoped to one merchant, amount and time window.
  • No debug endpoints in production clients, and integrity checks on client config.
  • A permission view that comes from the system, not from the model.
  • A bug bounty that prices prompt injection explicitly. Meta pays up to $130,000 for an injection that compromises one user's agent, within a programme capped at $300,000.

The design shows the industry's current best answer to agent security: assume the model will be fooled, then make fooling it unprofitable. Its first weeks in public show where that answer ends - on the parts of the system the model does not run on.

Top comments (0)