GPT-6 Astra landed on September 3 with a headline spec everyone repeated: 1.1 million tokens of context. Dump your whole repo in, your whole knowledge base, the entire PDF — no more chunking.
Here's the line item nobody put on the slide: input is $10 per million tokens. So filling that window costs $10 per call, before Astra generates a single token of answer. Add output at $50 per million, and a chunky response pushes one interaction past $12–15.
Do that in a loop — an agent that re-reads context every step — and you've built a very smart, very expensive way to burn a cloud budget. Let's do the actual math so you know when it's worth it.
The numbers, on one napkin
Astra's public pricing:
| Price / 1M tokens | |
|---|---|
| Input | $10.00 |
| Output | $50.00 |
| Cached input | $1.00 |
| Cache write | $12.50 |
So a few reference points:
- Fill 1M tokens of context, once: ~$10.
- An agent that re-sends ~1M context across 20 steps: ~$200, just for reading. Output is extra.
- The same context, cached after the first call: ~$12.50 to write it, then $1 per subsequent read instead of $10. That 10× drop is the whole game.
That last row is the one people miss.
When the big window is genuinely worth it
Paying to fill 1M tokens earns its keep when the task truly needs global reasoning over the whole thing at once:
- Cross-file reasoning where the answer depends on how twelve files interact — refactors, "why is this slow," "trace this bug across the stack." Retrieval that fetches 5 chunks will miss the connection that lives in the 6th.
- One-shot synthesis over a big document where you can't know in advance which parts matter — a contract review, a long incident timeline, a whole spec.
- When engineering time costs more than tokens. If a good RAG pipeline takes you two weeks to tune and the giant context "just works" for a low-volume internal tool, $10 a call can be cheaper than your salary.
When it's a trap
The 1M window seduces you into paying to re-read things that didn't change:
- You knew which 5% mattered. If retrieval could have found the relevant section, you just paid 20× to make the model find it for you.
- Agent loops without caching. Re-sending the full context every step is the single fastest way to a surprise bill. Cache it, or don't loop over it.
- "Context rot." Beyond a few hundred K tokens, more context often lowers answer quality — the model gets distracted by the irrelevant 95%. You can pay more and get worse.
- High volume. At scale, $10 a call is a business model, not a line item. Do the multiplication before you ship it.
The rule I'd actually follow
Reach for the giant context when the task needs the whole thing and you can't predict which part. Reach for retrieval when you can. And the moment you call the same big context twice, cache it — that's a 10× cut hiding in plain sight.
The 1M window isn't a scam. It's a genuinely useful tool with a price tag that punishes lazy usage. Treat "just put it all in context" as a decision with a dollar figure attached, not a free default, and it's great. Treat it as free, and Astra will happily teach you otherwise on the invoice.
What's the most surprising AI bill you've gotten — and was it context, output, or an agent loop that did it? Name and shame your own architecture below. 👇
I write about building with AI and the honest economics of it. Follow me here if that's your lane. 👋
Top comments (0)