Google DeepMind has announced Gemini 4 Argon, a frontier model aimed at long software engineering jobs, legal and finance work, and cyber defense. In Google's own evaluation it beats GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 on most of the benchmarks it reports. Almost nobody can use it yet.
Where Argon comes first
The comparison table covers 19 rows across knowledge work, agentic coding, ML engineering, science and math, long context, computer use, multimodal understanding and cybersecurity. Argon finishes first or tied for first on 14 of them.
The number Google leads with is DeepSWE v1.1, a set of real-world software engineering tasks. Argon scores 77.9% there, against 74.2% for Opus 5.5, 74.1% for Astra and 67.4% for Fable 5.1. The widest gap sits on Harvey's Legal Agent Benchmark, where Argon reaches 19.6% and none of the other three gets past 6.7%. 🎯
Benchmark Argon Astra Fable 5.1 Opus 5.5
Vals Index 68.9 63.1 65.8 67.0
AutomationBench 51.3 41.4 31.4 42.5
Vals Finance Agent v2 65.4 53.5 58.9 58.6
Harvey's Legal Agent Benchmark 19.6 5.4 6.7 3.8
DeepSWE v1.1 77.9 74.1 67.4 74.2
FrontierSWE v2 55.0 65.5 56.3 62.3
Vibe Code Bench 91.9 89.6 90.3 90.3
Terminal-bench 4.0 57.4 58.2 57.9 66.4
PostTrainBench 45.3 44.3 40.2 49.3
Terminal-Bench Science 0.1 57.6 68.1 52.6 63.3
LABBench 2 88.8 85.4 68.6 73.1
RiemannBench 76.0 72.0 65.6 69.6
GraphWalks BFS, up to 128k 99.7 98.7 91.4 90.6
GraphWalks BFS, 256k to 1M 84.2 71.8 65.0 66.8
Agent's Last Exam (pass rate) 39.5 34.2 - 38.2
OSWorld-2.0 (offline, partial) 69.2 72.6 - -
Chartography 71.6 71.0 46.2 66.3
LVBench 91.7 87.5 79.7 83.7
CWE-bench v1 68.0 68.0 58.0 67.0
All values are percentages from Google's table. GraphWalks reports F1, and OSWorld-2.0 uses an offline subset with partial scores.
Where it does not
Five rows go to someone else, and one is a tie. Astra keeps FrontierSWE v2 with 65.5% against Argon's 55.0%, Terminal-Bench Science 0.1 with 68.1% against 57.6%, and the OSWorld-2.0 offline subset with 72.6% against 69.2%. Opus 5.5 keeps Terminal-bench 4.0 with 66.4% against 57.4%, and PostTrainBench with 49.3% against 45.3%. On CWE-bench v1 Argon and Astra share first place at 68.0%.
Every figure here comes from Google's own runs, and the methodology page is the place to check how the rival models were set up.
One million tokens in one answer
Argon can write up to 1M output tokens in a single response. The previous limit was 64K, so one answer can now run about 15 times longer. Google ties this to the work the model targets: long, multi-step jobs where the reasoning has to hold across the whole run.
The price now and later
The introductory API price is $2 per 1M input tokens and $10 per 1M output tokens, and cached input costs 95% less than regular input. When the introductory period ends, the rate becomes $4 and $20.
GPT-6 Astra lists $10 per 1M input tokens and $50 per 1M output tokens at its standard short-context rate. That puts Argon at 5x cheaper than Astra on both sides today and 2.5x cheaper after the increase. The comparison is token price against token price, and a model that writes longer answers can still cost more per finished task. 💸
What Google already used it for
Inside Google, Argon has been working on real projects before release. Quantum computing researchers used it to optimize spacetime resources, counted as qubits times gates, and in one case it beat the published baseline by 40% within minutes.
A team of Argon agents read fleet-wide profiling telemetry across Google's data centers and applied memory optimizations on its own, freeing more than 300 TiB of memory.
Argon agents are also moving C and C++ codebases to Rust. Early reports put that at tens of thousands of lines, which holds for core libraries such as re2 and libgav1, while the largest job, the Fuchsia Zircon kernel, runs past 800K lines. The libgav1 result is a memory-safe video decoder with identical output that runs 2.7x faster than the earlier Rust port.
Who can use it
Right now Argon is going to a set of trusted cyber defenders through Google's Fairwind Program. Paid API customers and Google AI Ultra subscribers come next, followed by developers, enterprises and consumers more broadly. Google has given no date.
Until Argon reaches the paid API, the table and the $2 price are something to plan a test around, and nobody outside Fairwind can yet check whether those 14 first places hold up.
References
- Google: Gemini 4 Argon, our next era of frontier intelligence
- Google DeepMind: Gemini 4 Argon evaluation methodology
- OpenAI API pricing
Follow me for more on AI and Software Development:
khasky — LinkedIn / GitHub / Patreon / Bluesky / Mastodon / Medium / Devto
khaskydev — X / Threads / Instagram / Pinterest / Tumblr / Facebook / VK
Top comments (0)