Gemini 4 Argon leads Google’s benchmark table on 13 of 19 rows against GPT-6 Astra, Claude Opus 5.5, and Claude Fable 5.1. But availability matters more than any benchmark: you can buy Astra and Opus 5.5 today, while Argon is available only to Fairwind Program defenders. Its introductory price is $2/$10 per million input/output tokens, increasing to $4/$20 afterward—the same list price as Opus 5.5.
This comparison breaks down pricing and limits, groups Google’s benchmark results by winner, highlights evaluation caveats, and provides a practical routing guide. It also shows how to compare the available models with your own prompts in Apidog. If you are new to Argon, start with what is Gemini 4 Argon. For availability details, see the Argon release date guide.
Price and limits side by side
Google’s comparison includes Claude Fable 5.1, so it belongs in the table. Prices below are per 1M tokens and come from Google’s launch post, OpenAI’s GPT-6 Astra model page, and Anthropic’s pricing page.
| Gemini 4 Argon | GPT-6 Astra | Claude Opus 5.5 | Claude Fable 5.1 | |
|---|---|---|---|---|
| Can you call it today? | No, Fairwind only | Yes | Yes | Yes |
| API model ID | Not published | gpt-6-astra |
claude-opus-5-5 |
claude-fable-5-1 |
| Input | $2 intro, then $4 | $10 | $4 | $10 |
| Cached input | $0.10 intro, then $0.20 | $1 | $0.20 | $0.25 |
| Output | $10 intro, then $20 | $50 | $20 | $50 |
| Max output | 1M (Google’s stated limit) | 128K | 128K (300K on Batch, beta) | 128K |
| Context window | Not published | 1,050,000 (922K max input) | 1M | 1M |
| Long-prompt surcharge | Not stated | Over 272K input: 2x input and cache, 1.5x output, on the full request | None | None |
Three implementation implications stand out:
- Argon and Opus 5.5 have the same standard rates. Argon’s cost advantage over Opus lasts only during Google’s introductory period, which has no published end date.
- Astra and Fable 5.1 cost more per token. Their listed input and output rates are 2.5x Argon’s standard rates and 5x Argon’s introductory rates.
- Treat Argon’s 1M output limit carefully. Google states a 1M-token output limit, but Vals AI lists a 262K maximum output for the Argon configuration it tested.
For per-request calculations, see Gemini 4 Argon pricing.
Where each model leads in Google’s table
Google published this table, so interpret it accordingly. Argon’s scores are Google-reported, using a mix of self-computed results and leaderboard results. Competitor scores are mostly vendor-reported figures or public leaderboard results, often at maximum reasoning settings. Different harnesses mean small score gaps may not be meaningful.
| Who leads | Benchmark | Argon | Astra | Fable 5.1 | Opus 5.5 |
|---|---|---|---|---|---|
| Argon: knowledge work | Vals Index | 68.9% | 63.1% | 65.8% | 67.0% |
| Argon: knowledge work | AutomationBench | 51.3% | 41.4% | 31.4% | 42.5% |
| Argon: knowledge work | Harvey’s Legal Agent Benchmark | 19.6% | 5.4% | 6.7% | 3.8% |
| Argon: coding | DeepSWE v1.1 | 77.9% | 74.1% | 67.4% | 74.2% |
| Argon: long context | GraphWalks 256K to 1M (F1) | 84.2% | 71.8% | 65.0% | 66.8% |
| Argon: multimodal | LVBench | 91.7% | 87.5% | 79.7% | 83.7% |
| Argon: multimodal | Chartography | 71.6% | 71.0% | 46.2% | 66.3% |
| Astra | FrontierSWE v2 | 55.0% | 65.5% | 56.3% | 62.3% |
| Astra | Terminal-Bench Science 0.1 | 57.6% | 68.1% | 52.6% | 63.3% |
| Astra | OSWorld-2.0 (offline, partial) | 69.2% | 72.6% | not reported | not reported |
| Opus 5.5 | Terminal-bench 4.0 | 57.4% | 58.2% | 57.9% | 66.4% |
| Opus 5.5 | PostTrainBench | 45.3% | 44.3% | 40.2% | 49.3% |
| Tie | CWE-bench v1 | 68.0% | 68.0% | 58.0% | 67.0% |
Argon’s other outright wins are Vals Finance Agent v2, Vibe Code Bench, LABBench 2, RiemannBench, GraphWalks up to 128K, and Agent’s Last Exam.
Fable 5.1 does not lead a row in Google’s table. On CWE-bench v1, DeepMind’s cyber leaderboard shows a three-way 68% tie between Argon, Astra, and Grok 4.7, with Opus 5.5 at 67%.
For Argon versus Fable 5.1, Argon scores higher on 15 of the 17 rows where both report a result. Fable exceeds Argon only on:
- FrontierSWE v2: 56.3% vs. 55.0%
- Terminal-bench 4.0: 57.9% vs. 57.4%
See Claude Fable 5.1 benchmarks for Anthropic’s reported numbers and Gemini 4 Argon benchmarks for the full 19-row table.
Three caveats before you trust the gaps
1. Competitor scores are not Google reruns
Google’s methodology says non-Gemini results are “sourced from providers’ self reported numbers unless otherwise mentioned.” Several rows also come from public leaderboards operated by Vals AI, Proximal, and Surge.
Use these numbers for model selection hypotheses, not as final production evidence.
2. Benchmark harnesses differ
On DeepSWE v1.1, Google calculated Argon’s 77.9% using a mini-swe-agent harness. Astra’s number comes from a public leaderboard, while Anthropic’s scores come from system cards.
On LVBench, Gemini sampled video at one frame per second. Astra received 800 frames, Opus 5.5 received 600, and Fable 5.1 received 300, reportedly because of API limits.
Do not treat these as strictly equivalent runs.
3. A model can score differently across charts
Opus 5.5’s Terminal-bench 4.0 score changes between charts based on the harness and effort setting. Anthropic reports its 66.4% result at xhigh effort.
When creating an internal evaluation spreadsheet, keep the following fields alongside each score:
model
model_version
reasoning_or_effort_setting
benchmark_version
agent_harness
tool_configuration
date_tested
source
Never combine benchmark results from different harnesses into a single ranking without labeling the differences.
What third-party evaluators say
Independent leaderboards narrow the gap.
Artificial Analysis lists Argon at #8 out of 223 entries, although that ranking counts each reasoning setting separately. Grouped by distinct model, Argon scores 53, tying GPT-6 Astra and Claude Fable 5.1. It trails Claude Opus 5.5 at 58 at maximum settings and Claude Sonnet 5.5 at 56.
Artificial Analysis also lists Argon’s hallucination rate on AA-Omniscience at 15%, compared with 51% for Astra at its maximum setting.
On Arena, Argon ranks first in Text at 1525, marked Preliminary with 4,942 votes. Opus 5.5 ranks fourth.
Vals AI ranks Argon first among 41 models on the Vals Index, making it the first Gemini model to top that ranking.
There are also reports of internal skepticism. Bloomberg reported that some Google employees with access found Argon less impressive on certain coding and front-end design tasks than its benchmark results suggest. Google called that characterization inaccurate.
Which model should you route to?
There is no universal best frontier model. Route requests by workload, quality requirements, latency, output limits, and cost.
| Task | Route today | When Argon ships |
|---|---|---|
| Legal, finance, and office automation agents | Opus 5.5 (67.0% Vals Index) | Test Argon: it leads all four knowledge-work rows |
| Terminal-heavy coding agents | Opus 5.5 (66.4% Terminal-bench 4.0) | Opus 5.5 still leads |
| Agentic software engineering | Astra (65.5% FrontierSWE v2) or Opus 5.5 | Argon leads DeepSWE but trails FrontierSWE; test both |
| Computer use and GUI agents | Astra (72.6% OSWorld-2.0) | Astra leads OSWorld; Argon leads Agent’s Last Exam |
| Reasoning over 256K-token prompts | Opus 5.5 (no surcharge) or Astra (surcharge over 272K) | Argon (84.2% GraphWalks 256K to 1M) |
| Video and chart understanding | Astra (87.5% LVBench) | Argon (91.7%), with the frame-count caveat |
| Science and ML engineering | Astra (Terminal-Bench Science), Opus 5.5 (PostTrainBench) | Argon leads LABBench 2 and RiemannBench; test it |
| Single responses over 128K tokens | Opus 5.5 on Batch (300K, beta) | Argon, up to Google’s stated 1M |
| High-volume, cost-sensitive work | Opus 5.5 ($4/$20) | Argon at $2/$10 while the intro lasts |
If price matters more than peak benchmark results, compare OpenAI’s lower-cost line with Opus using GPT-6 Sol vs Claude Opus 5.5.
Compare the models on your own prompts
Vendor benchmarks cannot tell you how models handle your repository, documents, tools, response formats, or failure modes. Build a small repeatable evaluation instead.
You can set this up today in Apidog.
Step 1: Create three requests
Create one project with three requests:
- GPT-6 Astra
- Claude Opus 5.5
- Gemini with a configurable model ID
For the Gemini request, use an environment variable:
{
"model": "{{GEMINI_MODEL}}",
"contents": [
{
"role": "user",
"parts": [
{
"text": "{{PROMPT}}"
}
]
}
]
}
Set GEMINI_MODEL to gemini-3.8-flash until Google publishes Argon’s API model ID.
Use the GPT-6 Astra API guide and what is Claude Opus 5.5 for the vendor-specific request formats.
Step 2: Store credentials and shared inputs as variables
Create environment variables for each vendor key:
OPENAI_API_KEY
ANTHROPIC_API_KEY
GOOGLE_API_KEY
GEMINI_MODEL
PROMPT
Put the test prompt in PROMPT so every request receives exactly the same input.
For multi-case testing, add a dataset with fields such as:
{
"case_id": "bugfix-001",
"prompt": "Fix the failing test without changing the public API.",
"max_output_tokens": 4000,
"max_cost_usd": 0.25
}
Step 3: Add assertions
At minimum, validate:
- HTTP status is
200 - The response contains non-empty output
- Output tokens remain below your configured ceiling
- Estimated cost stays within budget
- Required structured fields are present, if using JSON output
For example, your assertions should enforce the equivalent of:
assert(response.status === 200);
assert(answer.length > 0);
assert(outputTokens <= maxOutputTokens);
assert(estimatedCostUsd <= maxCostUsd);
Step 4: Track cost per request
The basic cost formula is:
cost =
(input_tokens / 1_000_000 × input_rate) +
(output_tokens / 1_000_000 × output_rate)
For a request with 20,000 input tokens and 4,000 output tokens:
Astra:
20,000 × $10 / 1M + 4,000 × $50 / 1M
= $0.20 + $0.20
= $0.40
Opus 5.5:
20,000 × $4 / 1M + 4,000 × $20 / 1M
= $0.08 + $0.08
= $0.16
Argon at intro rates:
20,000 × $2 / 1M + 4,000 × $10 / 1M
= $0.04 + $0.04
= $0.08
Argon at standard rates:
20,000 × $4 / 1M + 4,000 × $20 / 1M
= $0.08 + $0.08
= $0.16
Per-token pricing is not the same as per-task cost. Artificial Analysis measured Argon at roughly 62K output tokens per task versus about 27K for Astra at maximum effort. Measure actual token use for your own tasks.
Step 5: Run one scenario and compare results
Run all three requests as one test scenario. Compare:
- Task completion rate
- Output validity
- Tool-call correctness
- Latency
- Input and output token usage
- Estimated request cost
- Human review score for a sample of outputs
When Argon becomes generally available, update only this variable:
GEMINI_MODEL=<published-argon-model-id>
Then rerun the same scenario. You will have a same-prompt comparison against Astra and Opus 5.5 without rebuilding your test setup.
FAQ
Is Gemini 4 Argon better than GPT-6 Astra?
They tie at 53 on Artificial Analysis. In Google’s table, Argon scores higher on most rows, but Astra wins FrontierSWE v2, Terminal-Bench Science 0.1, and OSWorld-2.0. Astra is also available today.
Gemini 4 Argon vs Claude Opus 5.5: which is better?
Google’s table favors Argon on most rows. Opus 5.5 wins Terminal-bench 4.0 and PostTrainBench, and it leads Artificial Analysis 58 to 53. At Argon’s standard rates, the two models cost the same.
Is Gemini 4 Argon cheaper than Claude Opus 5.5?
During the introductory period, Argon is half the price: $2/$10 versus $4/$20 per million input/output tokens. After the introductory period, their list prices are identical. Google has not announced when the introductory pricing ends.
Which model has the largest output limit?
Argon has Google’s stated 1M-token output limit, although Vals AI lists 262K for the configuration it tested. The other models cap synchronous output at 128K. See the 1M output tokens guide for client-side implications.
Can I use Gemini 4 Argon today?
Only through Google’s Fairwind Program, which is available to a set of vetted cyber defense partners. Paid API customers and Google AI Ultra subscribers are expected later, but no date has been published.
Pick by workload, then test
Argon looks strongest in Google’s results for knowledge work, long context, and multimodal tasks. Astra leads on FrontierSWE, Terminal-Bench Science, and OSWorld. Opus 5.5 leads on terminal work and ML engineering while matching Argon’s eventual standard price.
Build production workflows on the models you can call now. Keep the Gemini model ID in an environment variable, record task-level quality and cost, and rerun the same scenario when Argon opens access.
To build the three-request comparison in one project, download Apidog.
Top comments (0)