Google announced Gemini 4 Argon on September 30, 2026, with a benchmark table that puts it highest in 13 of 19 rows. If you choose models at a developer-tool or API company, you now have to decide whether that changes this quarter's evaluation plan, budget, or roadmap, working from a launch post, because almost nobody can call the model yet.
Access is what limits that decision. Google's announcement says Argon is "rolling out to a set of trusted cyber defenders through our Fairwind Program." The Gemini API models page, last updated October 1, 2026, lists no Gemini 4 model. There is no API model ID, no date for wider access, and no stated end to the introductory price of $2 per million input tokens and $10 per million output tokens.
This week, decide which workloads to queue for evaluation once a model ID appears. Going by where Argon leads and trails in the published results, queue long-output generation and repo-scale code migration first, and keep terminal-driven agent loops on the model you use now. We have not run Argon. Every performance figure below is Google's or comes from a named third party, as of October 1, 2026.
TL;DR
- Gemini 4 Argon is announced but not callable: rollout is limited to Google's Fairwind Program, and no Gemini 4 model appears on the Gemini API models or pricing pages as of October 1, 2026.
- Google's own table has Argon highest in 13 of 19 rows, with clear leads on DeepSWE, long context from 256k to 1M tokens, and multimodal (LVBench). It trails GPT-6 Astra and Claude Opus 5.5 on the terminal-driven agent benchmarks.
- Independent scoring is less flattering. On Artificial Analysis's Intelligence Index, Argon ties GPT-6 Astra at 53 and trails Claude Opus 5.5 and Sonnet 5.5. It has the lowest hallucination rate of any model scoring 45+ there, but lower accuracy than GPT-6 Astra.
- The $2/$10 price is introductory and undated. The regular $4/$20 price matches Claude Opus 5.5, and Argon used about 2.3 times Astra's output tokens per task in Artificial Analysis's runs.
What Google announced
The announcement, "Gemini 4 Argon: our next era of frontier intelligence", signed by "Koray Kavukcuoglu, SVP, Google DeepMind and Chief AI Architect, Google," makes these claims:
- An output token limit of "an industry-leading 1M tokens, up from the previous 64K tokens," which Google ties to "longer, more complex use cases."
- Introductory pricing of "$2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price," followed by "$4 per 1M input tokens and $20 per 1M output tokens" once the introductory period expires. Google does not say when that is.
- Internal coding use: "Google engineers have been using Argon for their daily tasks, from everyday debugging to large-scale codebase migrations and algorithm designs."
- Operations work: Argon agents "autonomously identify and apply memory optimizations across Google's data centers, freeing up over 300 TiB of memory once rolled out."
- Security: Argon can "autonomously find, validate, and patch critical software vulnerabilities," with safety resting on "misalignment mitigations that monitor Argon's chain-of-thought and actions and stop execution when necessary."
Google's own take on timing is cautious: "We'll continue to gather feedback from early testers as we iterate on guardrails before making Argon available to developers, enterprises, and consumers as soon as possible." For the wider release it names an order, "starting with paid API customers and Google AI Ultra subscribers," but no timeframe for the API, Vertex AI, AI Studio, or the Gemini app.
The post does not state the context window or the input modalities. Artificial Analysis lists a "Context Window: 1M tokens" in its launch article. On modalities, that article and its model page disagree, so treat both facts as unconfirmed until Google publishes a model card.
What the 1M output limit means, and how it differs from the context window
The context window caps what goes in; the output limit caps what one response can return.
The output limit is the less familiar of the two. One Hacker News reader wrote on launch day: "I don't understand the point or meaning of an output token limit." The context window is what the model reads in; the output limit caps what it writes back in one response.
A 64K output cap means a long artifact, such as a multi-file migration diff, a generated SDK, or a long report, has to be split across several calls and stitched back together. A 1M cap should remove much of that splitting for the workloads Google names: "longer, more complex use cases." DataCamp notes there are "no published rate limits in the announcement," so how often you could request a response that long is unknown.
Long outputs cost money. Calculated from Google's list prices, a response using the full 1M output tokens would cost $10 in output at the introductory price and $20 at the regular price. Argon also writes a lot per task: in Artificial Analysis's runs it averaged "62k output tokens per task, compared with 27k for GPT-6 Astra (max)."
Where Argon leads and trails in Google's table
The Google DeepMind model page compares Argon with GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5. Argon is highest outright in 13 of 19 rows, tied in one, and behind in five. The rows that matter most for a developer-tool team:
| Benchmark (Google's table) | Argon's position | Gemini 4 Argon | Compared with |
|---|---|---|---|
| DeepSWE v1.1 | Leads | 77.9% | Claude Opus 5.5 74.2%, GPT-6 Astra 74.1% |
| GraphWalks 256k to 1M | Leads | 84.2% | GPT-6 Astra 71.8% |
| Harvey's Legal Agent Benchmark | Leads | 19.6% | Claude Fable 5.1 6.7% |
| AutomationBench | Leads | 51.3% | Claude Opus 5.5 42.5% |
| LVBench | Leads | 91.7% | GPT-6 Astra 87.5% |
| CWE-bench v1 | Tied | 68.0% | GPT-6 Astra 68.0% (tie) |
| FrontierSWE v2 | Trails | 55.0% | GPT-6 Astra 65.5% |
| Terminal-bench 4.0 | Trails | 57.4% | Claude Opus 5.5 66.4% |
| PostTrainBench | Trails | 45.3% | Claude Opus 5.5 49.3% |
| Terminal-Bench Science 0.1 | Trails | 57.6% | GPT-6 Astra 68.1% |
| OSWorld-2.0 (Offline subset) | Trails | 69.2% | GPT-6 Astra 72.6% |
"Agentic coding" splits in two. Argon leads DeepSWE by about four points and trails FrontierSWE v2 by more than ten. If you run coding agents, look at the benchmark that resembles your loop.
Two third-party results favor Argon. The Vals Index leaderboard, updated 9/30/2026, ranks Argon first at "68.90%," ahead of Claude Sonnet 5.5 at "67.04%."
On AA-Omniscience, Artificial Analysis found that "Gemini 4 Argon has a 15% hallucination rate, the lowest of any model scoring 45+ on the Intelligence Index."
Where Gemini Argon falls short
Terminal-driven agents are the clearest weak spot. The two sources measuring it disagree on details, so keep their figures apart. In Google's table, Argon scores 57.4% on Terminal-bench 4.0 against 66.4% for Claude Opus 5.5 and 58.2% for GPT-6 Astra.
In its own runs, Artificial Analysis reports that "Gemini 4 Argon achieves 57%, ... only behind Claude Sonnet 5.5 (max, 64%), Claude Opus 5.5 (max, 60%) and GPT-6 Astra (59%)." The harnesses and sources differ, so the competitor numbers do not line up, but both sources place Argon behind Opus 5.5.
On the independent index, Argon is not in front. Artificial Analysis says "Gemini 4 Argon (high) scores 53 on the Artificial Analysis Intelligence Index, matching GPT-6 Astra (max, 53)." The Decoder adds: "Anthropic's models still lead. Claude Opus 5.5 sits at 58 points and Claude Sonnet 5.5 at 56." Its headline sums it up: Argon "closes the gap with OpenAI and Anthropic but doesn't take a clear lead."
A low hallucination rate comes with lower accuracy. On the same AA-Omniscience evaluation, Argon scores 50%, which Artificial Analysis calls "a 5 point decrease from Gemini 3.1 Pro Preview."
GPT-6 Astra scores 63% on that test. For a product that answers factual questions, Argon's two numbers point in different directions.
Argon uses more tokens per task. The 62k versus 27k output-token gap means a per-token price comparison overstates the saving. The Decoder concludes: "The price advantage comes from lower token rates, not from efficiency."
Whose numbers these are
Read the methodology note in Google's model evaluation PDF. It says: "All the results for non-Gemini models are sourced from providers' self reported numbers unless otherwise mentioned below." Two of Argon's own coding results are Google's runs: "DeepSWE v1.1 results for Gemini 4 Argon are self computed, using a mini-swe agent harness" and "Terminal-Bench 4.0 results for Gemini 4 Argon are self computed." Argon's scores are "pass @1" at "the highest thinking settings" unless the PDF notes otherwise.
So the table mixes Google's own runs (of Argon on most rows, and of every model on GraphWalks, PostTrainBench, and LVBench) with providers' self-reported numbers and public leaderboards such as Vals AI, and the real-world vulnerability benchmark uses "an internal dataset." One commenter on the Hacker News launch thread, nonethewiser, asked: "How is it even possible for every model to release benchmark results where they are #1 in 75% of categories?" We found no independent hands-on coding reports as of October 1, 2026. With no public access, DataCamp notes, "Almost nobody outside the Fairwind cohort has used Argon yet."
Gemini 4 Argon vs GPT-6 Astra, Claude Fable, Opus, and the model you use now
The table compares every option on the same criteria, read on October 1, 2026. Prices are list prices per 1M tokens from OpenAI API pricing (standard tier, short context is "≤272K input tokens") and Claude pricing.
| Option | Public API status, Oct 1, 2026 | List price per 1M input / output tokens | Google's benchmark table | Artificial Analysis Intelligence Index | Vals Index |
|---|---|---|---|---|---|
| Gemini 4 Argon | No API model ID; Fairwind Program only | $2 / $10 introductory, then $4 / $20 | Highest in 13 of 19 rows, tied in 1 | 53 (high) | 68.90% |
| GPT-6 Astra | Listed on OpenAI API pricing | $10 / $50 (short context) | Leads FrontierSWE v2, Terminal-Bench Science 0.1, OSWorld-2.0; ties CWE-bench v1 | 53 (max) | 63.13% |
| GPT-6.1 Sol | Listed on OpenAI API pricing | $2 / $10 (short context) | Not in Google's table | 52 (max) | 61.15% |
| Claude Fable 5.1 | Listed on Claude pricing | $10 / $50 | In the table; leads no row | Ties Argon and Astra, per The Decoder | 65.83% |
| Claude Opus 5.5 | Listed on Claude pricing | $4 / $20 | Leads Terminal-bench 4.0 and PostTrainBench | 58, per The Decoder | 66.97% |
| Claude Sonnet 5.5 | Listed on Claude pricing | $2 / $10 | Not in Google's table | 56, per The Decoder | 67.04% |
| Status quo: the model you run today | In production now | What you pay today | Your own results on your own tasks | Not applicable | Not applicable |
Argon leads on Google's scoreboard and tops the Vals Index. On Artificial Analysis's index it ties GPT-6 Astra and Claude Fable 5.1, behind two Anthropic models. At its regular price it costs the same as Claude Opus 5.5, and it is the only option in the table you cannot run today.
The status quo is an option too. Hacker News commenter deanc: "I can open codex or claude code apps or CLI and get real work done today with the latest models (even on the cheapest plans)." Argon has no public API, so anything you ship this quarter runs on a model you can call today.
The case for planning a migration now
The strongest opposing view: Google's table has Argon ahead, the Vals Index puts it ahead of Claude Sonnet 5.5 and Opus 5.5, and at $2/$10, Argon's introductory price is a fifth of GPT-6 Astra's $10/$50 short-context price and Claude Fable 5.1's $10/$50. If that holds the day the API opens, Argon becomes the obvious default, and a team with a migration plan ready saves first.
Four facts limit that argument:
- The lead depends on the scoreboard. It holds in Google's table, which mixes Google's own runs with providers' self-reported numbers and public leaderboards. On Artificial Analysis's index, Argon ties GPT-6 Astra and trails Claude Opus 5.5 and Sonnet 5.5.
- The price is temporary. Google gives no end date for the introductory period, and Vals already lists Argon at "$4 / $20" with a cost of "$15.68" per test. At $4/$20, Argon costs exactly what Claude Opus 5.5 costs.
- The low price is not unique. The introductory $2/$10 matches GPT-6.1 Sol's short-context price and Claude Sonnet 5.5, which you can call today.
- Cost per task depends on token use. Artificial Analysis puts Argon at "$1.99 per Intelligence Index task" while the price is "currently discounted 50%." Argon also writes about 2.3 times as many output tokens per task as Astra.
What survives is a narrower claim: Argon deserves a place in your evaluation queue, budgeted at the regular price. Prepare the tests now; a roadmap rewrite can wait for access and your own results. VentureBeat's Carl Franzen wrote that "model choice is still workload-dependent, even if Argon now gives Google its strongest claim yet to overall frontier leadership by benchmark count."
Which workloads to queue for Argon
Use one rule for every workload: queue it when published results show Argon leading on that kind of task and the result would change what you ship; leave it where it is when they show Argon behind. Budget every test at $4/$20 and compare cost per task.
| Workload | Published evidence | What to do |
|---|---|---|
| Long-output generation: large generated files, long reports, multi-file changes in one response | 1M output token limit, up from 64K (Google) | Queue first |
| Repo-scale code migration and long-context code work | DeepSWE v1.1 lead (Google's own run); GraphWalks 256k to 1M lead; Google reports internal use for "large-scale codebase migrations" | Queue first |
| Factual answers where a wrong answer is costly | Lowest hallucination rate of any model scoring 45+ on its index, but 50% accuracy against Astra's 63% (Artificial Analysis) | Test against your own question set |
| Terminal-driven agent loops | Trails on Terminal-bench 4.0, FrontierSWE v2, and Terminal-Bench Science 0.1 (Google); behind three models on Terminal Bench 4 (Artificial Analysis) | Stay on your current model |
| Vulnerability finding and patching | Google's headline claim, from an internal dataset with no counts or false-positive rates; Fairwind access only | Watch; you cannot test it outside the program |
This matches DataCamp's Argon guide: "When it opens up to regular users, I would reach for Argon on long-context document and repository work and stay with Claude Opus 5.5 or GPT-6 Astra for terminal-driven agent loops." The guide describes reactions so far as "first impressions of the announcement, not hands-on tests."
By situation:
- If your product generates long artifacts, such as SDKs or migration tooling, build your Argon evaluation set now.
- If your core loop is a coding agent driving a terminal, keep it on whatever wins your own tests today, and recheck when independent hands-on reports appear.
- If cost is your reason to look, compare Argon with Claude Opus 5.5 at the same $4/$20, counting output tokens.
- If you defend software and are in the Fairwind Program, you are among the few who can test Google's security claims now.
How to evaluate Argon once access opens
- Freeze a task set from your own backlog for the queued workloads, with the expected result for each task.
- Run Argon through the harness you use for your current model, and record the thinking level. Google's DeepSWE result came from "a mini-swe agent harness," and Argon's scores use "the highest thinking settings" unless the PDF notes otherwise, so your numbers may differ.
- Log output tokens and cost per task at the regular $4/$20 price, next to the same figures for your current model.
- Count wrong answers separately from refusals or abstentions, because Artificial Analysis's results separate hallucination rate from accuracy.
What Argon means for developer-tool companies
For a developer-tool company, Argon is a model you might build with, and also one more agent that reads your API and developer docs. When Argon reaches developers, coding agents built on it can read the same pages your users read and repeat whatever is out of date in them.
EkLine works on that side of the problem. It "keeps product, API, and developer docs accurate enough for AI engines, search systems, and agents to find, trust, and cite," as EkLine's post on AI visibility puts it. EkLine Journeys (in beta), described on EkLine's AI visibility page, run the questions people ask through Perplexity, ChatGPT, Gemini, and Claude, and let Claude Code or Codex try your product in a sandbox. You see what each assistant or coding agent says about your product and which pages it cited, so you know which pages to fix.
EkLine is not built to choose or benchmark models, and we have not tested Argon. The page names Gemini, not Argon specifically. EkLine cannot make an assistant cite your page; it shows you what the assistants say and keeps the pages they read accurate.
FAQ
Why is Gemini 4 Argon being released to cybersecurity teams first?
Google says Argon is "rolling out to a set of trusted cyber defenders through our Fairwind Program" while it iterates on guardrails. Security is also Google's headline claim: it says Argon can "autonomously find, validate, and patch critical software vulnerabilities." Google adds that it is "actively engaged in the U.S. government's voluntary process for pre-release model access while we gradually expand access."
How do I get access to Gemini 4 Argon?
As of October 1, 2026, the announcement gives no sign-up or waitlist link outside the Fairwind Program. DataCamp reports that "There is no free tier, no waitlist link, and no published rate limits in the announcement." Google says the wider rollout will start with paid API customers and Google AI Ultra subscribers, without a date. Watch the Gemini API models page for a Gemini 4 model ID.
Did Google catch up after being behind for months?
Partly. Artificial Analysis titled its review "Google is back as one of the top three labs in intelligence achieved" and scores Argon 53, matching GPT-6 Astra, while The Decoder notes that Claude Opus 5.5 and Sonnet 5.5 still score higher. Google's own table, which combines Google's runs with providers' and leaderboards' numbers, shows a lead in 13 of 19 rows.
To check what AI assistants and coding agents read and repeat from your developer docs, see how EkLine shows what AI assistants say about your product.

Top comments (0)