Many developers running multi-turn AI coding agents notice a familiar frustration: after asking an assistant to patch two functions and run tests over an afternoon, their weekly quota warning suddenly turns red. Even worse, as the conversation stretches longer, the assistant starts forgetting files read twenty turns earlier or silently undoes prior fixes.
The default developer instinct is simple: if models natively support 1M context, why not remove limits?
On October 10, 2026, Tibo (@thsottiaux), an engineer on OpenAI's Codex and ChatGPT team, addressed this directly on X:
"Guys, why do you think we chose 300k for Codex? We ran the numbers, it's better. We could run 1M, but that would be way more expensive on usage. Good defaults matter."
In multi-turn coding agent workflows, the assistant is not answering a single prompt. Every step re-submits the entire conversation history, file content, and tool outputs. Unlocking 1M context turns every minor command into an expensive full-history recalculation.
Hands-on Test: The Physical Cost of Agent History
You can verify this immediate overhead in any coding interface right now:
-
Session A (Long Conversation): Take an existing session with 30+ turns filled with terminal logs, diffs, and project files. Send a single inspection command:
Check git status. -
Session B (Fresh Clean Session): Open a brand-new window with only the project brief, and send the exact same command:
Check git status.
Compare the behavior across both windows:
- Latency: Session B streams back almost instantly. Session A pauses for several seconds before the first token appears.
- Token Consumption: Session B consumes under 1,000 tokens. Session A consumes hundreds of thousands of input tokens just to execute a 15-character CLI query.
In multi-turn agent systems, the action taken remains identical, but the engine must re-digest the entire accumulated history on every single turn.
7-Day Empirical Backtest: The Token Ledger
Tibo's remarks arrived directly under a 7-day empirical backtest by developer Rohit (@rohit3a) analyzing context compression thresholds in daily coding sessions. Rohit modeled weekly quota burn across different auto-compact limits:
- 1M Window (Uncapped): Triggered only 1 compaction over 7 days, but suffered the highest total weekly quota burn.
- 600k Window: 30 compactions; saved 18% weekly usage.
- 400k Window: 82 compactions; saved 29% weekly usage with zero degradation in multi-file reasoning.
- 300k Window: 126 compactions; saved 36% weekly usage.
- 200k Window: 232 compactions; saved 41% usage.
The data reveals a critical point of diminishing returns. Dropping from 300k to 200k spiked compactions by 84% (from 126 to 232), yet yielded only an additional 5% quota savings. Compacting that aggressively caused the model to lose recent local variable definitions and repeat code generation mistakes.
The 300k to 400k bracket represents the optimal engineering balance: preserving multi-turn reasoning while cutting token bloat.
The Hidden 272k Pricing Cliff
Token accumulation in long sessions is not merely linear. In modern flagship APIs (such as GPT-6 Astra pricing disclosures), pricing structures enforce a non-marginal step function: the 272,000-token pricing cliff.
When a request crosses 272,000 input tokens:
- Input Rate Doubles (2× Input): Every single input token in that request is billed at double the baseline rate.
- Output Rate Increases: Output tokens jump by roughly 1.5×.
This is a cliff edge, not a marginal surcharge. Crossing 272k means the entire accumulated history is recalculated at 2× rates.
Consider an agent session at turn 35 holding 270k tokens. A simple test run reads an unexpected 3k log file, pushing the total to 273k. Instantly, all 273k tokens are billed at the 2× tier. For every subsequent prompt, cost velocity jumps by over 101%.
This threshold stems from GPU cluster hardware constraints. Around 272k tokens, memory requirements for KV Cache (Key-Value Cache) and attention mechanisms exceed standard tensor parallelism boundaries on modern clusters, requiring doubled compute resources to maintain state.
Why 300k Instead of 272k?
If costs spike at 272k, why did OpenAI set Codex's ceiling at 300k rather than 272k?
This represents an intentional engineering compromise:
- The 28k Sprint Runway: Complex refactoring tasks routinely accumulate 220k to 260k tokens across dependency graphs and test runs. Forcing hard compaction at 250k risks truncating context right when the agent begins writing code. The buffer up to 300k provides room to finish the immediate task.
- The Hard Circuit Breaker: While permitting brief bursts past 272k, 300k acts as a physical ceiling that prevents sessions from drifting into 500k or 1M territory.
| Window Threshold | Weekly Compactions | Quota Savings | 272k Cliff Exposure | Practical Verdict |
|---|---|---|---|---|
| 1M Uncapped | ~1 | Baseline (0%) | Permanently trapped in 2× tier | Avoid; quota drain |
| 300k (Default) | ~126 | Saves 36% | Brief sprint exposure only | Recommended sweet spot |
| 250k–270k | ~150 | Saves ~39% | Fully avoids 2× penalty | Ideal for pay-per-token API users |
| < 200k | >232 | Saves 41% | Far below cliff | Loses short-term code context |
Practical Guidelines for Agent Quotas
-
Trust the 300k Default: In official subscription interfaces, do not artificially raise context ceilings to 1M. For tools supporting custom compaction thresholds (like
/autocompact), set the trigger to 250k–260k if paying per API token to stop right before the 272k cliff. - Switch Sessions per Feature Branch: Avoid running a single chat session for an entire week. Structure sessions around small deliverables: discuss architecture in session one, write and verify implementation in session two, then commit and start a fresh session for the next module.
- Model Tiering: Use fast lightweight models (like Gemini 3.8 Flash) for initial repository exploration and document lookups. Switch to frontier reasoning models (like GPT-6.1-sol or Claude Sonnet) within the 300k window for core refactoring.
What You Don't Need to Do Yet
You do not need to install complex local vector databases or script custom token pruners. In production software development, leaving default context limits intact and opening a clean session after each verified commit captures over 80% of total quota savings. Good defaults protect your budget so compute goes toward working code.
Top comments (1)
The forgetting-what-it-read-twenty-turns-ago part is the real cost of a huge window. Disclosure: I build brethof-brain, and the bet behind it is the opposite of a bigger window: keep the decisions outside the context and hand the agent only the ones that bear on the current prompt.