You open a repository you do not fully trust in Codex, pick read-only mode because you are being careful, and ask it a question about the code. Per the Accomplish AI write-up, that is enough for whoever wrote the repository to run a command on your machine. No approval dialog, no visible process, nothing on screen. The mode you picked for safety was not the boundary you thought it was.
Oren Yomtov published the write-up on September 15 and BleepingComputer picked it up on September 20, which is when it crossed my feed. Both bugs were reported to OpenAI on August 12 and fixed within eight days, so this is not an unpatched emergency. I am writing about it anyway because the fix lives in a build number, while the interesting part lives in a config file on every machine where a Codex Desktop install ever ran. I maintain a small MCP security scanner and I keep writing about agent sandboxes failing, and this one stung differently: the sandbox did its job, and the escape walked through a default nobody was asked about.
Two commands get you most of the answer on your own machine:
# Check 1: which build am I on?
codex --version
# Write-up's stated fixed builds: CLI 0.149.0 or later (Overpatch),
# Desktop 26.818.21641 or later (Heapjack)
# Check 2: did a Desktop install write a global tool into my config?
grep -n "node_repl" ~/.codex/config.toml
# A hit means the tool the escape targeted is inherited by every session.
What was the sandbox promising, exactly?
Read-only is the strictest mode Codex offers. The promise is simple to state: the agent can read and reason about a repository, but nothing in that repository can effect a change on your machine. Repository contents are untrusted input, and the whole value of the mode depends on one separation holding: text that the agent reads must never become execution.
What makes the Heapjack escape worth studying is that it does not attack the sandbox's code at all. It attacks the assumption that inside one process, the code and the data can be kept apart by anything weaker than the process boundary itself.
How does a token in shared memory become a shell?
The escape targets a JavaScript tool that Codex Desktop installs. Its architecture, as the write-up describes it, is a native Rust parent process that is unsandboxed, plus a Node process started inside the Codex sandbox where the JavaScript actually runs. The parent does the privileged work: it opens applications, connects to unix sockets, writes the global config file. The sandboxed side sends it requests.
Two details turn that design into a hole.
First, the trusted and untrusted JavaScript run as two contexts in one Node process. One V8 isolate, one heap. The credential that tells the parent a request came from trusted code is a secret token, and per the write-up, "the token is a string in that heap." The same heap the untrusted code lives in.
Second, the requests travel as JSON lines on stdout, one pipe shared by both contexts. So the attack sequence is mechanical. Untrusted code takes a snapshot of the heap with v8.getHeapSnapshot() and tries every string shaped like a UUID. The parent's error messages leak whether a guess was close: a wrong token comes back "not authorized", while a right token with a bad argument comes back with a real validation error. Once the token is recovered, the untrusted side writes its own request onto the shared pipe. The native parent reads it, checks the token, and does the work.
Why the error message was the whole oracle
This is the detail I would put in front of every engineer building an agent tool. From the guesser's side, failures came in exactly two flavors: a rejection for a wrong token and a genuine validation error for a right one. That difference is a guessable oracle. The boundary between trusted and untrusted was a single string, and the system told the attacker on every attempt how close the guess was.
The published proof of concept is macOS-specific: the parent runs the system open command so launchd starts the application unsandboxed, outside Codex's process tree. The write-up also notes any unix socket is reachable (Docker's daemon socket called out as the obvious one) and that the tool can edit the global config file. Platform scope beyond macOS is not established in anything I read, and I did not reproduce any of this myself.
What did Overpatch do differently?
The second bug needed workspace-write mode, not read-only, and it lived in the open-source Codex CLI's patch flow. The flaw, quoted from the write-up: apply_patch "grants write access to the parent folder of each path in the patch. Name /tmp and it grants write access to /."
The published chain uses a two-change patch: one change names /tmp, which widens the grant, and the other appends a line to .zshrc through a symlink into the home directory. The next terminal you open runs that line unsandboxed.
A permission model that grants upward
Granting on the parent of an attacker-chosen path is granting on attacker-chosen scope. I wrote earlier this month about the Cursor allowlist bypass that starts with a file named curl, and the shape rhymes: a control that trusts a label instead of the thing the label resolves to. There it was a command name resolving to a project-local file. Here it is a path string expanding into a grant on everything above it.
Why is a default-on config entry the real story?
The node_repl tool is not something anyone opted into. Per the write-up, Codex Desktop writes an [mcp_servers.node_repl] block into the global ~/.codex/config.toml at install, with no prompt and no setting to turn it off, and the plain CLI inherits it from that shared file. The write-up's own diagnosis is the best sentence in the whole disclosure: "The thing doing the enforcement was sitting inside the thing being enforced."
This completes a pattern I keep hitting in agent security. Two unauthenticated CVSS 9.8 agent sandbox CVEs landed the same day earlier this month because defaults composed into a 9.8 while each default looked harmless alone. CrewAI's sandbox CVE turned on a nine-name import blocklist. Here, the widening of the trust boundary shipped as a convenience nobody was consulted about. The bug is rarely "there was no sandbox." It is that the boundary was one default away from the thing it was guarding.
One more observation: the tool rode into the config through the MCP servers table, the same table where every third-party server you wire in lands. An agent config is a running inventory of privileged helpers, and most people, me included until I started scanning configs, have never printed theirs. That is the same blind spot as a planted prompt being a valid credential: secrets and privileges accumulate in the places that are easiest to read and hardest to remember.
What did OpenAI actually fix?
Here is the sourcing picture, stated plainly. Both bugs were reported on August 12 and, per the write-up, fixed inside of eight days. Heapjack is closed in Codex Desktop build 26.818.21641 and Overpatch in Codex CLI 0.149.0, and both version strings trace to the Accomplish write-up; BleepingComputer attributes them the same way. I could not find an OpenAI statement quoted in any coverage I checked, no CVE identifier, no CVSS score, and no mention of a bounty as of September 21. I also found no source claiming in-the-wild exploitation, and affected version ranges were never published, only fixed builds.
That last gap is why the checks at the top of this post answer "am I affected": the disclosure does not hand you a version range, so the honest alternative is checking your own build and your own config. One thing I could not resolve: whether updating Desktop removes an existing node_repl entry from the global config. The write-up does not say, and I am not going to guess, so treat check 2 as something to look at rather than something a reinstall definitely clears.
What are the 3 checks, ranked by what they catch?
Check 1, the build check, catches the patched-versus-not question. Run codex --version against the fixed builds above, and check the Desktop build in its about screen. Verify both against vendor release notes before trusting them, since the strings trace to the Accomplish write-up.
Check 2, the config audit, catches the inheritance question. Grep the global config for node_repl, then go one step further and list every [mcp_servers.*] entry in it. Each one is a tool your sessions inherit automatically, and the audit costs one cat.
Check 3 is the habit, and it is the only one that survives the next disclosure: treat mode labels as claims, not boundaries. For repositories from people you do not know, keep an outer layer outside the agent anyway, a container or a VM, because a mode is one software layer, and this year's sandbox stories keep ending the same way: generous defaults composing into a missing boundary.
Full honesty on method: these checks are derived from the disclosure and the config file, not from a lab run. I am on Windows, the published proof of concept is macOS, and I did not reproduce the escape. What I verified is the sourcing chain, and that is what this post claims and nothing more.
Is a sandbox allowed to be a suggestion?
Here is the argument I would like settled in the comments. If a read-only guarantee depends on a secret sitting in shared memory, was the mode ever a boundary, or was it a preference with confident naming? Should software that cannot enforce a mode refuse to offer it, the way a server should refuse to start when its guarantees depend on a substring check, or is a mostly-working sandbox still worth shipping because most attackers are not Oren Yomtov?
My position, held loosely: a mode name is a contract, and shipping one you cannot enforce is worse than not shipping it, because users make real decisions on the label. The counterargument is fair too: read-only still stops the boring majority of damage, and demanding perfect isolation before offering partial protection is how you get tools that do nothing. I have not fully made up my mind, and the tiebreaker for me would be disclosure: a mode that documents what it can and cannot guarantee is honest, a mode that implies more than it enforces is not.
What would change my mind: a public OpenAI advisory that frames these as defense-in-depth gaps rather than sandbox escapes, or evidence that the node_repl default ships with a documented off switch I could not find. Absent either, the default stays the story, and the config on your machine is the part only you can check.
Primary sources: Escaping the OpenAI Codex sandbox, twice (Accomplish AI, Sep 15) · Researchers escape OpenAI Codex sandbox to run commands on host (BleepingComputer, Sep 20) · TLDR Dev newsletter, Sep 17 (corroborating summary).
Top comments (0)