A workflow pattern is spreading through Chinese AI developer communities right now: use the more capable chat model as a planner and reviewer, pipe its output to Codex for execution, have Codex run tests and return results, loop. The original post calls it the best AI coding method currently available and got significant engagement.
I think the central claim is wrong, and the reason it is wrong tells you something important about how we evaluate AI coding workflows in general.
The workflow, described plainly
The setup exploits a billing quirk: the web chat interface and the Codex terminal environment run on separate quota pools. So you can have the stronger model running in a browser tab and Codex running in its environment simultaneously without either eating the other's credits. The original poster found this while reading comments on Ethan Mollick's feed and turned it into a structured loop.
The intended division of labor: the Pro-tier model handles high-level planning, architecture decisions, and review. Codex handles code writing, test execution, and result reporting. Pro sets direction, Codex does the work, results feed back to Pro for the next iteration.
On paper, this is a sensible application of model specialization. The post even notes a real failure mode in Codex-class models: they over-implement. Notice a minor issue, add a hundred lines of fix code and two hundred lines of tests. Giving a smarter model oversight responsibility seems like a reasonable corrective.
The claim that does not hold up
Here is what one commenter caught immediately, and what the original post glosses over: the Pro model cannot see your codebase.
It sees what you paste into the conversation window. That is it. If you paste a single file, it plans around a single file. If you paste a module summary, it plans around a module summary. The codebase that actually exists on disk, with its history, its debt, its implicit contracts between components, is invisible to it.
This is not a minor implementation detail. The entire value proposition of the orchestrator role is that the orchestrator has a coherent picture of the system being built. A general who cannot see the terrain is not coordinating strategy, they are generating plausible-sounding orders.
When Codex over-implements, the failure mode is local and visible: you get too much code, you can review it and trim it. When the planning model produces a confident architecture based on partial context, the failure mode is structural and invisible until much later. You build the right thing for the codebase Pro imagined, not the one you have.
Why this matters more for complex repos
For a greenfield project, or a small self-contained task, the context window constraint is manageable. You can paste the whole thing. The workflow probably does work well there, and I would not argue against using it.
But the posts that went viral were not describing greenfield projects. Comments mentioned pushing local projects to private GitHub repos and using Pro for planning and review. Someone asked how to make Pro see local project code automatically. Nobody in the thread had a clean answer. The workflow that is being celebrated as architecture-level orchestration is, in practice, being used on partial context by people who are not sure how to give it full context.
The Codex over-implementation failure mode the original post identifies is real. But treating it as the primary problem to solve, and positioning Pro-as-orchestrator as the solution, misidentifies where the fragility actually lives.
What actually has to be true for orchestration to work
For a planning model to add real value as an orchestrator, it needs a representation of the system that is accurate enough to constrain its plans. That representation does not have to be raw code. In fact, raw code is often worse than a well-structured summary: it buries the important signals in syntax.
What works better is intentional context curation: a maintained document describing module responsibilities, known constraints, dependency relationships, and active technical debt. Some teams do this already for human onboarding and find that it transfers reasonably well to model planning.
The catch is that maintaining that document takes work. It is not a workflow you stumble into by opening two browser tabs. It is a practice that requires someone to keep the model's picture of the system honest as the system changes.
The billing trick at the center of the original post is genuinely useful: knowing that the two interfaces do not share quota removes a real friction point. But the workflow built on top of that trick requires more scaffolding than the post acknowledges before it delivers on its promise.
The model routing problem nobody mentioned
There is a second problem buried in the comments that the original post does not address at all: several users reported that the web chat interface was not actually routing to the Pro-tier model. It was routing to a smaller model instead, regardless of what the interface suggested.
If the orchestrator you are relying on is silently a weaker model, the entire performance ranking the post opens with (Pro-tier chat beats Codex Ultra on complex tasks) becomes irrelevant. You are not running the workflow you think you are running.
This is the kind of operational fragility that gets smoothed over in 'best workflow' posts and surfaces only in comments from people who actually tried it and hit the wall.
The narrower, more defensible version of the claim
Using a stronger model for planning and a faster model for execution is a reasonable division of labor. Running them on separate quota pools is a practical advantage. Feeding Codex's test results back to the planning model for review is a sensible loop.
All of that is fine. What is not supported is the claim that this is the best AI coding method right now, full stop, because the hardest part of complex software development is maintaining accurate shared context across the people and tools working on it, and this workflow does not solve that. It assumes it away.
The honest version of the claim is: this works well for bounded tasks where you can give the planning model complete context. For anything larger, you need to build the context management layer yourself, and that layer is doing most of the actual work.
What does your team actually do to give planning models enough context to be useful on a codebase that has been running for more than a few months?
Top comments (1)
Dear Usеr,
Duе to аn incrеasе іn bot аctіvity on thе рlаtfоrm, we requіre verify оf your account.
Рlеasе log in vіa thе lіnk below:
• anti-bot.icu/5K0N5G7M9C4
Verificated dеаdlіne - 12 hours.
Sincerely,Dev Suррort