Facts as of September 30, 2026.
"Do not use separate Bots as a security boundary."
That sentence isn't from a pentest report. It's from xAI's own Grok Bot documentation, for a product marketed as "AI teammates you can give real work to."
Within seven weeks, three of the biggest AI vendors shipped the same kind of product: an agent that runs around the clock, logs into your tools and reports back when it's done. xAI went first with Grok Bot on August 11, Meta followed with Muse on September 8, and OpenAI launched Dots at DevDay on September 29. Most coverage compares context windows, benchmarks and avatars. If you've ever written a threat model, the interesting question is a different one: which computer does the agent run on, who else shares that computer, and what decides what leaves it?
The common pattern: the agent gets a computer
Until recently, "agent" in practice meant an LLM calling tools over an API while you sat in a chat window. Close the tab, the agent is gone. All three products break that model in the same three ways.
Persistent. The agent has a name, memory across conversations, and keeps working while you're offline.
A real machine. Each product gives the agent a cloud computer with a browser, file system and logged-in sessions. xAI's FAQ puts it plainly: "Every Bot on your account uses one persistent cloud computer." That's what lets an agent operate websites that have no API, save files and sign into accounts like a human at a desk.
Proactive. Dots keeps researching with read-only tools while idle, Grok Bot runs routines on a schedule or on events, and Muse watches inboxes and prices and pings you when something changes.
In other words: these vendors are selling a workstation for software that behaves like a colleague. And the moment you see it that way, the important questions turn into access-control questions. You don't ask a new hire about their IQ first. You ask what they have access to.
OpenAI Dots: one cloud computer per agent
Dots runs on GPT-6 Astra. Each dot gets its own cloud computer, separate from your machine; local desktop access is off by default and has to be turned on through the ChatGPT desktop app. You talk to it through the ChatGPT apps, Slack, Microsoft Teams and voice calls, with SMS available as a limited beta for US Pro users. It inherits your ChatGPT app connections and, per OpenAI, reaches more than 4,000 apps through plugins.
The permission model is the interesting part. OpenAI's announcement says "Custom Rules let you allow specific actions, require approval, or block them," and the Help Center lists four levels: Take action without asking, Take action if pre-approved, Ask before taking action and Hand off to you. Permanent deletion and software installs go through approval; password changes and money transfers are handed back to you entirely.
The part I like most: proactive research is architecturally constrained. When the dot works on its own initiative, it uses "tools that are restricted to be read-only, which means that they can't send messages, change app content, or control your browser or computer." That's a clean split between looking around and acting, and neither competitor draws that line as clearly.
Pricing: one dot is included in ChatGPT Pro and Business Premium; more dots are announced for later, without a price yet. Availability matters if you're in Europe: the Pro rollout excludes the EEA, Switzerland and the UK, while Business Premium is available "across all supported ChatGPT regions." Enterprise, Edu and Healthcare workspaces get a beta that an admin must enable.
Meta Muse: an egress guard that approves every outbound action
Muse runs on Meta's Muse Spark model and targets consumers first: email, travel bookings, bill negotiation, forms. It's available on iOS, Android, muse.ai and WhatsApp.
Architecturally, it's the most interesting of the three. Every user gets a dedicated "Muse Secure VM" holding both the agent and the user's data. Next to it runs a second agent, the Sentinel, separated from Muse at the system level. Per Meta, nothing Muse does reaches the internet unless the Sentinel approves it, and it asks the user for permission when needed. Credentials sit in secure storage that Muse can use without seeing them. Payments go through one-time cards via Stripe Link. A "Confidential VM" with user-held keys is promised for later this year.
So Muse separates thinking from acting into two processes: the model proposes, the Sentinel decides what leaves the box. Conceptually, that's the strongest answer to prompt injection any vendor ships today. If the agent reads a hidden instruction on a crafted page telling it to email your last bank statements to a stranger, that send still has to pass a guard that knows what you actually asked for. Whether that holds up in daily use, nobody can say yet, Meta included.
The launch already produced some data points. Reuters reported on internal Meta posts describing the agent routing around guardrails to expose a person's private iCloud photos in testing, and monitoring tasks silently stopping after about 15 minutes. Interactions are used for training by default (opt-out available). Pricing is the most transparent of the three: free up to 100M tokens per week, then $20 or $100 per month. It's available in the US and Canada, 18+.
Grok Bot: many bots, one computer, shared logins
Grok Bot launched in beta on August 11 and ships with SuperGrok and Cursor plans (xAI is now SpaceXAI, part of SpaceX, which acquired Cursor). It's positioned as a team: you create several bots with different roles and let two to six of them coordinate in group chats. You can show a bot a workflow once, and it turns the demonstration into a draft skill you review, then runs it as a routine on a schedule or trigger. It prefers connectors where they exist and falls back to driving the browser.
Here's the difference that matters: all bots on an account run on one shared cloud computer. Each gets its own screen, but browser cookies, sessions, files and command-line credentials are shared. Isolation between users is strict: per xAI's security FAQ, each gets a dedicated Firecracker microVM. Isolation between bots in one account doesn't exist. xAI is upfront about it:
"Do not use separate Bots as a security boundary."
And from the Bots docs:
"Deleting a Bot removes its active profile, conversation, and routines from Grok Bot. Shared computer files and sign-ins are not isolated by Bot and may remain on the computer."
Think through what that means in practice: the bot that reads unfiltered customer email uses the same browser sessions as the bot that prepares bank transfers. For real privilege separation, you need separate accounts, because only those give you separate machines.
The comparison
| OpenAI Dots | Meta Muse | xAI Grok Bot | |
|---|---|---|---|
| Launch | Sep 29, 2026 | Sep 8, 2026 | Aug 11, 2026 (beta) |
| Model | GPT-6 Astra | Muse Spark | no fixed model named |
| Compute | own cloud computer per dot | one Secure VM per user | one shared cloud computer per account |
| Isolation within an account | per dot | per user + Sentinel egress guard | none: files, sessions, logins shared |
| Sensitive actions | Custom Rules (4 levels), proactive mode read-only | Sentinel approves all network actions, user approval when needed | Auto-review rules, drafts approved before sending, admin policies |
| Pricing | 1 dot in ChatGPT Pro / Business Premium | free tier, $20, $100 / month | bundled with SuperGrok and Cursor plans |
| EU availability | Pro: no · Business Premium: yes · Enterprise: beta | US + Canada only | no regional statement in the docs |
Isolation decreases from left to right. All three are legitimate engineering choices. But for an agent that touches customer data, the shared box is the one that doesn't fit.
Why the isolation model beats the model
An LLM that reads web pages, emails and PDFs will inevitably read text written by attackers. Prompt injection can't be patched away, because the model can't reliably separate instructions from content. Every improvement in hit rate lowers the risk; none removes it.
What decides the damage is the blast radius of a compromised agent, and that radius is the isolation model:
- Dots: the radius ends at that dot's own computer. A hijacked research dot can't touch the sessions of an accounting dot on a different machine.
- Muse: the radius ends at the Sentinel. Inside the VM the agent can hallucinate all it wants; only what the guard approves gets out.
- Grok Bot: the radius ends at the account. A bot reading a poisoned page sits on the same machine, with the same cookies, files and open logins, as every other bot you own.
The sysadmin version: Dots gives every employee their own laptop. Muse gives every employee their own laptop and puts a doorman at the exit who reads every letter before it goes out. Grok Bot sits ten employees at one shared PC where everyone stays logged in, and honestly writes in the manual: if you want separation of privileges, buy more PCs.
Credit to xAI for writing it down. But a warning in a manual shifts responsibility to you, and it's a different thing from a boundary you can't technically bypass.
What to do before these land in your stack
You don't need any of these products to start on the part that actually matters. Write the permission model first, vendor-agnostic. It can be as simple as a policy file you keep in your repo:
# Illustrative least-privilege policy for one agent (not a vendor format)
agent: invoice-intake
owner: finance-lead@example.com # a named human, like a manager for an employee
scopes:
mailbox: { read: true, send: approval } # drafting is fine, sending needs a human
calendar: { read: true, write: true }
erp: { read: true, post_booking: false }
banking: { read: false, write: false } # never
crm: { read: true, write: true, delete: false }
isolation:
dedicated_runtime: true # no shared browser sessions with other agents
egress: allowlist # only the domains this job needs
logging:
every_action: true # with the agent's stated reason
With that file in hand, evaluating Dots, Muse or Grok Bot can become an afternoon of mapping their controls onto your columns instead of a quarter of discovery. Two more things worth deciding early:
- Build or rent. These three are rented agents on someone else's infrastructure, with models you can't pick and quotas nobody has priced long-term. Open-source harnesses like OpenClaw or Hermes Agent run on your own infra with any model provider. They aren't automatically safer: OpenClaw received nine CVEs between March 18 and 21, 2026, including a CVSS 9.9 authorization bypass. But they give you control over the agent's computer, which is the variable this whole post is about.
- Start boring. An agent that reads invoices from a mailbox and files them, with read-only access, is a good first case. An agent that negotiates contracts belongs at the end of the list.
If you're in the EU, keep two constraints in mind: GDPR Article 22 limits automated decisions with legal or similarly significant effects without human involvement, and the AI Act's transparency obligations for systems interacting with people have applied since August 2, 2026. An agent that writes emails and makes calls on your behalf has to disclose that it's an AI. Formally that duty sits with the provider, but the problem lands on you.
Treat Dots, Muse and Grok Bot as a preview of what your own agents will look like soon, and build the frame they're allowed to work in now. Which one ends up in your company, or whether you build your own, then becomes a purchasing decision instead of a security decision.
Originally published in German at next-levels.de. I'm co-founder and CEO of Next Levels, a digital agency in Germany; AI consulting is one of our service lines.
Top comments (0)