DEV Community

Cover image for Gemini 4 Argon Beats GPT-6 Astra and Claude on Most Benchmarks
Khasky
Khasky

Posted on

Gemini 4 Argon Beats GPT-6 Astra and Claude on Most Benchmarks

Google DeepMind has announced Gemini 4 Argon, a frontier model aimed at long software engineering jobs, legal and finance work, and cyber defense. In Google's own evaluation it beats GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 on most of the benchmarks it reports. Almost nobody can use it yet.

Where Argon comes first

The comparison table covers 19 rows across knowledge work, agentic coding, ML engineering, science and math, long context, computer use, multimodal understanding and cybersecurity. Argon finishes first or tied for first on 14 of them.

The number Google leads with is DeepSWE v1.1, a set of real-world software engineering tasks. Argon scores 77.9% there, against 74.2% for Opus 5.5, 74.1% for Astra and 67.4% for Fable 5.1. The widest gap sits on Harvey's Legal Agent Benchmark, where Argon reaches 19.6% and none of the other three gets past 6.7%. 🎯

Benchmark                          Argon   Astra   Fable 5.1  Opus 5.5
Vals Index                         68.9    63.1    65.8       67.0
AutomationBench                    51.3    41.4    31.4       42.5
Vals Finance Agent v2              65.4    53.5    58.9       58.6
Harvey's Legal Agent Benchmark     19.6     5.4     6.7        3.8
DeepSWE v1.1                       77.9    74.1    67.4       74.2
FrontierSWE v2                     55.0    65.5    56.3       62.3
Vibe Code Bench                    91.9    89.6    90.3       90.3
Terminal-bench 4.0                 57.4    58.2    57.9       66.4
PostTrainBench                     45.3    44.3    40.2       49.3
Terminal-Bench Science 0.1         57.6    68.1    52.6       63.3
LABBench 2                         88.8    85.4    68.6       73.1
RiemannBench                       76.0    72.0    65.6       69.6
GraphWalks BFS, up to 128k         99.7    98.7    91.4       90.6
GraphWalks BFS, 256k to 1M         84.2    71.8    65.0       66.8
Agent's Last Exam (pass rate)      39.5    34.2    -          38.2
OSWorld-2.0 (offline, partial)     69.2    72.6    -          -
Chartography                       71.6    71.0    46.2       66.3
LVBench                            91.7    87.5    79.7       83.7
CWE-bench v1                       68.0    68.0    58.0       67.0
Enter fullscreen mode Exit fullscreen mode

All values are percentages from Google's table. GraphWalks reports F1, and OSWorld-2.0 uses an offline subset with partial scores.

Where it does not

Five rows go to someone else, and one is a tie. Astra keeps FrontierSWE v2 with 65.5% against Argon's 55.0%, Terminal-Bench Science 0.1 with 68.1% against 57.6%, and the OSWorld-2.0 offline subset with 72.6% against 69.2%. Opus 5.5 keeps Terminal-bench 4.0 with 66.4% against 57.4%, and PostTrainBench with 49.3% against 45.3%. On CWE-bench v1 Argon and Astra share first place at 68.0%.

Every figure here comes from Google's own runs, and the methodology page is the place to check how the rival models were set up.

One million tokens in one answer

Argon can write up to 1M output tokens in a single response. The previous limit was 64K, so one answer can now run about 15 times longer. Google ties this to the work the model targets: long, multi-step jobs where the reasoning has to hold across the whole run.

The price now and later

The introductory API price is $2 per 1M input tokens and $10 per 1M output tokens, and cached input costs 95% less than regular input. When the introductory period ends, the rate becomes $4 and $20.

GPT-6 Astra lists $10 per 1M input tokens and $50 per 1M output tokens at its standard short-context rate. That puts Argon at 5x cheaper than Astra on both sides today and 2.5x cheaper after the increase. The comparison is token price against token price, and a model that writes longer answers can still cost more per finished task. 💸

What Google already used it for

Inside Google, Argon has been working on real projects before release. Quantum computing researchers used it to optimize spacetime resources, counted as qubits times gates, and in one case it beat the published baseline by 40% within minutes.

A team of Argon agents read fleet-wide profiling telemetry across Google's data centers and applied memory optimizations on its own, freeing more than 300 TiB of memory.

Argon agents are also moving C and C++ codebases to Rust. Early reports put that at tens of thousands of lines, which holds for core libraries such as re2 and libgav1, while the largest job, the Fuchsia Zircon kernel, runs past 800K lines. The libgav1 result is a memory-safe video decoder with identical output that runs 2.7x faster than the earlier Rust port.

Who can use it

Right now Argon is going to a set of trusted cyber defenders through Google's Fairwind Program. Paid API customers and Google AI Ultra subscribers come next, followed by developers, enterprises and consumers more broadly. Google has given no date.

Until Argon reaches the paid API, the table and the $2 price are something to plan a test around, and nobody outside Fairwind can yet check whether those 14 first places hold up.

References


Follow me for more on AI and Software Development:

khasky — LinkedIn / GitHub / Patreon / Bluesky / Mastodon / Medium / Devto

khaskydev — X / Threads / Instagram / Pinterest / Tumblr / Facebook / VK

Top comments (0)