Three companies shipped top-tier models within a month of each other, and the headline price gap is large: Google's Gemini 4 Argon launched at $2 per million input tokens, while OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 cost $10.
If you're picking a model for something you're building, price is only half the question. The other half is whether you can actually call it today. Here's both, from each company's own announcement.
This is not a hands-on test. Every price and benchmark below comes from the vendor's own page, linked at the end.
TL;DR
| Model | Company | Input / output per 1M tokens | Can you use it today? |
|---|---|---|---|
| Gemini 4 Argon | Google (Sep 30 2026) | $2 / $10 introductory, then $4 / $20 | Only trusted cyber defenders (Fairwind). Paid API and Google AI Ultra next, no date |
| GPT-6 Astra | OpenAI (Sep 3 2026) | $10 / $50 | Yes: ChatGPT paid plans + API |
| GPT-6.1 Sol | OpenAI | $2 / $10 | Yes: API |
| Claude Fable 5.1 | Anthropic | $10 / $50 | Yes |
| Claude Opus 5.5 | Anthropic | $4 / $20 | Yes |
| Claude Sonnet 5.5 | Anthropic | $2 / $10 | Yes |
What it costs per month
Take a workload of 10M input tokens and 2M output tokens a month (a busy internal tool or a small production feature). Using the list prices above:
| Model | Input | Output | Monthly |
|---|---|---|---|
| Gemini 4 Argon (introductory) | 10 × $2 = $20 | 2 × $10 = $20 | $40 |
| Gemini 4 Argon (later price) | 10 × $4 = $40 | 2 × $20 = $40 | $80 |
| GPT-6 Astra | 10 × $10 = $100 | 2 × $50 = $100 | $200 |
| GPT-6.1 Sol | $20 | $20 | $40 |
| Claude Fable 5.1 | $100 | $100 | $200 |
| Claude Opus 5.5 | $40 | $40 | $80 |
| Claude Sonnet 5.5 | $20 | $20 | $40 |
Two things stand out:
- The $2 tier isn't unique to Google. GPT-6.1 Sol and Claude Sonnet 5.5 cost the same $2 / $10, and you can use both today.
- Argon's price is introductory. At the announced later price of $4 / $20, it costs the same as Claude Opus 5.5.
What each company says its model is for
Gemini 4 Argon. Google positions it for long, complex work: coding, legal and finance, and cybersecurity defense. Output goes up to 1M tokens (from 64K). Google reports DeepSWE v1.1 77.9% (a record), No. 1 on AutomationBench (51.3%), LVBench 91.7% and CWE-bench 68% (tied first).
GPT-6 Astra. OpenAI's pitch: "Anything you can do on a computer, Astra can do for you." OpenAI says it's state of the art on FrontierMath Tier 4, ARC-AGI 3 and TerminalBench 4.0. Its cheaper sibling GPT-6.1 Sol matches Astra on DeepSWE, according to OpenAI.
Claude. Fable 5.1 is Anthropic's most capable generally available model, built for long-running work. Opus 5.5 "performs at the level of Fable 5.1 on most work", according to Anthropic, and is 40% cheaper to run than Opus 5 and over 30% faster.
Why the benchmarks don't settle it
Every number above is the company's own result on its own setup. Different harnesses, prompts and tool access can move a score a lot, so a Google number and an OpenAI number for different benchmarks can't be ranked against each other. Treat them as each vendor's claim about where its model is strong, not as a leaderboard.
If you need a real answer for your use case, run your own eval: 20 to 50 tasks from your actual product, scored the same way on each model.
Which one should you build on today?
- You need it in production this week: Argon isn't an option yet. Compare GPT-6.1 Sol and Claude Sonnet 5.5 at $2 / $10, or Claude Opus 5.5 at $4 / $20 if you need more capability.
- You need the top tier now: GPT-6 Astra or Claude Fable 5.1, both $10 / $50.
- You're planning for later and cost matters most: watch for Argon's paid API. Budget at the $4 / $20 price, not the introductory one.
Sources
- Google: Gemini 4 Argon
- OpenAI: GPT-6 Astra · GPT-6.1 Sol
- Anthropic: Claude Fable · Claude Opus 5.5 · Pricing
I make clear tech and AI explainers on RainyTech. Which model are you building on right now? Tell me in the comments.
Top comments (5)
rainytech, this is a brilliantly clear breakdown of the llm pricing and availability landscape! 🐯
your point about "run your own eval: 20 to 50 tasks from your actual product" is the absolute golden rule of llm engineering. vendor benchmarks are marketing; real-world performance on your specific edge cases is what actually matters.
this perfectly mirrors the model routing logic we built into koda. we don't fire up a heavy, expensive model for a simple syntax check. we route to a fast, cheap model (like a 20b) for quick tasks, and reserve the heavy compute (120b) only for complex architectural debugging.
since you've been digging into these apis, i'm curious: have you noticed any significant differences in rate-limiting behavior or context-window handling between gpt-6.1 sol and claude sonnet 5.5 at that $2/$10 tier? for indie devs, those operational limits often matter more than the raw token price.
fantastic, highly practical analysis. thanks for keeping the community grounded in reality! 🛡️📊
Thanks Harun, and thanks for flagging that phishing comment 🙏 I've hidden it.
Good question. I haven't tested rate limits or long-context behavior for GPT-6.1 Sol vs Claude Sonnet 5.5 myself. This piece only compares list prices and what each company publishes. Both set rate limits by account tier, so check each provider's rate-limit docs for your plan. A hands-on test of exactly that would make a good follow-up.
rainytech, you're very welcome! glad i could help keep the feed clean. 🛡️
that makes total sense regarding the rate limits. i'll definitely keep an eye on the provider docs and maybe run a small load test on my end to see how gpt-6.1 sol handles sustained, 3g-friendly payloads compared to sonnet.
thanks again for the top-tier, no-fluff breakdown!
Official Platform Update
Security protocols have been updated for all developer accounts.
THIS IS A PHISHING SCAM 🚨 Do not click this link. Dev.to will never ask you to verify your account via a third-party link in the comments.
Some comments have been hidden by the post's author - find out more