In January 2026, Andrej Karpathy posted a thread about going from 80 percent manual coding to 80 percent agent coding in a single month. It pulled 40,000 likes in a week, and the replies split exactly down the middle: half pointing at Claude Code, half pointing at Codex. Six months later, that split is still the live debate, and most of the content about it is trash. It is either written by someone affiliate-linked to one answer, or it repeats folklore numbers that controlled tests have already debunked.
This article is the opposite of that. I have not run a paid 100-hour head-to-head myself, and I will not pretend I did. Every number below comes from public benchmarks, vendor announcements, and long-form comparisons by developers who did run the tests, all linked inline. The short version: the model layer is a tie, the pricing is not, and the right answer depends on the kind of work you are doing, not on which company you like.
The pricing asymmetry nobody prices honestly
Start with the number that decides most purchases before any benchmark gets read.
Both tools sit on top of a chat subscription. Codex comes with ChatGPT plans, Claude Code comes with Claude plans, and the tiers look similar on the surface: roughly $20, then $100, then $200. Under the surface they are not similar at all.
- Codex at $20 (ChatGPT Plus): meaningful daily coding-agent runtime. That is 15 to 80 GPT-5.5 messages per 5-hour window, 30 to 150 GPT-5.3-Codex messages, and 10 to 60 cloud tasks, per OpenAI's Codex pricing page. Long-form comparisons, including a 100-hour head-to-head on r/ClaudeCode, consistently report that a $20 Plus plan absorbs daily agent work that on Anthropic's side would require the $100 Max tier.
- Claude Code at $20 (Claude Pro): Anthropic's own support docs describe it as suited for light usage. The credible-volume tier is Max 5x at $100, and that quota is shared between claude.ai chat and Claude Code, so your agent and your chatbot compete for the same budget.
- The counter-case is real: Codex's $20 tier ships tight 5-hour rolling windows, and the limits have been revised mid-stream more than once. The most-discussed r/codex thread this spring was about users reporting four-times tighter weekly quotas without notice.
So the honest statement is: Codex is cheaper to try, and both get expensive at scale. If you are on a $20 budget, this section is probably the whole decision for you. If money is not the constraint, read on, because the free-tier math is not where the interesting differences live.
Tokens per task: the real cost lives here
Plan pricing is the visible number. The number that decides whether you stay inside your limits is tokens per task, and Claude Code consistently spends more of them. It reads more files, plans before writing, and verifies tools before calling them. That behavior has a price.
The cleanest controlled test I found is a head-to-head from Composio that ran the same two prompts (a PR-triage system and a real-time code review UI) against Claude Code on Opus 4.7 and Codex on GPT-5.5, same machine, same MCP setup. The results kill two myths at once:
- The gap is not 5 to 10 times. Claude Code used roughly 192,000 tokens at about $2.50. Codex used 136,000 tokens at about $2.04. That is a 1.4x token gap and a 23 percent cost gap. The folklore number is wrong.
- The premium bought something. In that controlled run, the extra tokens got a more decomposed architecture (twelve components versus seven) and an unprompted smoke test. Codex left one task empty because its MCP path was misconfigured and it shipped anyway.
There is a pattern underneath: the gap widens when the agent is doing tool-heavy work. If your session talks to Linear, GitHub, and a database through MCP, Claude Code's "check tools first, then plan" loop runs the bill up faster. For a self-contained refactor with no tool calls, the gap nearly closes. If you use Claude Code and want to cut consumption, Firecrawl's token efficiency guide documents 12 techniques that benchmarks show cutting costs by 77 to 91 percent.
The benchmarks: a tie at the frontier
On public benchmarks, the two model families are within noise of each other. From the vendors' own published numbers:
- SWE-bench Verified: Anthropic's Opus 4.6 scored 78.3 percent. OpenAI's GPT-5.1-Codex-Max scored 77.9 percent at its highest reasoning setting.
- Community results tilt by task type: multiple developer comparisons report Claude Code winning consistently on frontend and UI work, which people attribute to its richer skill and MCP ecosystem around frontend tooling. Codex wins consistently on long-running backend tasks, because its cloud agent runs unattended for hours and it is the first model family trained for compaction, which holds long-context coherence over 24-plus hour runs.
- Adoption is real on both sides: OpenAI says 95 percent of its engineers use Codex weekly, shipping roughly 70 percent more PRs since adopting it. Anthropic does not publish an equivalent number, but Claude Code is the tool most agent-harness features in the industry have been copied from.
Arguing about which model is 0.4 benchmark points ahead is a waste of your attention. All these numbers are vendor-reported, they move every quarter, and the interesting variable is not the model. It is the harness around it.
The harness: depth versus surface
This is where the two products genuinely diverge, and it is the part most comparisons gloss over.
Claude Code is the deepest programmable harness on the market. The stack includes Subagents (specialized agents in .claude/agents/ with their own context window and tool allowlist), Skills (the SKILL.md format Anthropic open-sourced in December 2025), Hooks (26 lifecycle events you can intercept with shell scripts), a plugin marketplace, and Dynamic Workflows, where Claude orchestrates tens of subagents in one session. If you want a custom policy like "run the linter before every commit and reject commits that fail tests without human approval," hooks give you exactly that knob.
Codex takes the opposite bet: one product across six surfaces. CLI, IDE extension, Codex Cloud, the ChatGPT app sidebar, mobile (GA May 2026), and a Chrome extension. All of them share your ChatGPT account, your session history, and your AGENTS.md config. You can start a refactor on your phone during a commute, pick it up in VS Code at your desk, and review the PR from the Chrome extension without losing state. Codex is now used by "more than 5 million people every week" per OpenAI's June 2026 announcement. It is also open source under Apache-2.0, so you can read and modify the harness itself.
The sandboxing philosophies are the deepest split. Codex enforces isolation at the kernel layer: Seatbelt on macOS, bubblewrap with Landlock on Linux, the Windows sandbox in PowerShell, with network off by default. The OS enforces the boundary before the model ever reaches it, which is deterministic. Claude Code enforces policy at the application layer through those hooks plus an Auto mode classifier shipped in March 2026, which reviews tool calls. Less deterministic, far more expressive. If you are running an agent on a sensitive codebase and want a hard guarantee it cannot touch the network, Codex's model is what you want. If you want to encode your team's policy as code, Claude Code's is.
AGENTS.md versus CLAUDE.md: a small thing that matters
Codex uses AGENTS.md, the community-defined repo instruction file. Cursor, Windsurf, OpenCode, and most other agents respect it. One file, many agents. Claude Code uses CLAUDE.md, which is proprietary to Anthropic's ecosystem but more powerful inside it: hierarchical resolution where the most specific file wins, @path imports so you can compose instruction files, and auto-memory that writes back when you tell Claude to remember something. The practical catch, noted in Builder.io's comparison: Claude Code still does not read AGENTS.md, so multi-agent repos end up maintaining two files.
The decision matrix
Here is the scannable version, so you can save this and skip the rest of the internet's take:
- Budget is $20 a month → Codex. The Plus tier absorbs daily agent work; Claude Pro will not.
- You live in the terminal and want maximum control → Claude Code. Skills, hooks, subagents, plan mode.
- Your work is long-running and autonomous → Codex. Cloud runs unattended for hours; compaction holds coherence past 24 hours.
- You do frontend and UI work → Claude Code. The frontend skill and MCP ecosystem consistently produces better results.
- You work across phone, desktop, and web in one day → Codex. The six-surface continuity is real and nothing on Anthropic's side matches it.
- You are on a sensitive codebase and want kernel-level sandbox guarantees → Codex.
- You want to encode custom policy, linting, and approval gates → Claude Code hooks.
- You run multiple different agents on one repo → Codex, for AGENTS.md. Keep a synced CLAUDE.md alongside.
- You want to read or fork the harness itself → Codex. It is Apache-2.0 open source.
What I would actually do
If I were setting up from scratch today, I would not pick one. The people running both consistently report the same hybrid pattern: Claude Code for interactive work in the terminal, where the fast feedback loop and the hook system earn their keep, and Codex Cloud for background tasks, long autonomous runs, and PR review through @codex mentions. Both speak MCP, so one MCP server setup works in either. The cost of running both at entry level is one $20 subscription, which is less than the hourly rate of the debugging session you will need after picking a tool by brand loyalty instead of by task.
The debate is loud because the models are tied and the differences are architectural. Architectural differences are good news. It means the answer is not "which one is better," it is "which one matches the work in front of you this week."
I write about AI tooling, backend engineering, and the developer workflow every week. Subscribe, it is free.
Which one is in your terminal right now, and what made you pick it? Did the $20 pricing asymmetry factor in, or did you choose on harness features? Tell me in the comments.
Top comments (0)