Gemini 4 Argon is Google’s new frontier model, announced on September 30, 2026. It is currently available only to Fairwind Program defenders: a vetted group of cyber defense partners. Introductory pricing is $2 per million input tokens and $10 per million output tokens, increasing to $4 and $20 respectively after the introductory period. Argon raises the output limit to 1M tokens per response. There is no public API, published model ID, or release date yet.
This guide covers what Google has and has not confirmed, where Argon fits in the Gemini lineup, expected access, pricing, benchmarks, and what you can build now. Choosing a model today? Read our Gemini 4 Argon vs GPT-6 Astra vs Claude Opus 5.5 comparison. If you test model APIs in Apidog, the implementation section shows how to make Argon a one-variable model switch.
What Google has confirmed vs. what it has not
Early coverage often mixes Google statements with third-party estimates. The table below separates confirmed information from unknowns using Google’s launch post and evals methodology.
| Item | Status | What we know |
|---|---|---|
| Announcement | Confirmed | September 30, 2026, by Koray Kavukcuoglu of Google DeepMind |
| Access today | Confirmed | A set of Fairwind Program partners |
| Next in line | Confirmed, undated | Paid API customers and Google AI Ultra subscribers |
| Price | Confirmed | $2/$10 intro, then $4/$20 per 1M tokens; cached input is 95% off |
| Output limit | Confirmed | 1M tokens, up from 64K |
| Benchmarks | Confirmed, Google-reported | A 19-row table plus a methodology PDF |
| Model ID | Not published | Strings online are third-party placeholders |
| Input context window | Not published | Google's long-context eval used prompts up to 1M tokens |
| Knowledge cutoff | Not published | Nothing from Google |
| Intro period length | Not published | No date for the switch to $4/$20 |
| Rate limits, free tier | Not published | Nothing from Google |
| Release date | Not published | “As soon as possible” |
If a site lists an Argon context window, model string, or speed figure, treat it as unconfirmed until Google publishes it in official documentation.
What Gemini 4 Argon is and where it sits
Argon is the only Gemini 4 model announced so far and Google’s new top-tier Gemini model. Google calls it “our new frontier model,” built “to sustain deep reasoning across complex, long-horizon workflows.” It claims frontier performance in software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense.
A Google spokesperson told Reuters that Argon is larger than the company’s previous “Pro” models.
Until now, the top Gemini API model was gemini-3.1-pro-preview, which remains in Preview:
- Input: 1,048,576 tokens
- Output: 65,536 tokens
- Pricing: $2 input / $12 output per million tokens for prompts up to 200K
Argon’s introductory output price of $10 per million tokens undercuts that output rate. Google has not said whether Argon will have a separate long-prompt tier like 3.1 Pro or whether more Gemini 4 models will follow.
Google says it already uses Argon internally for coding, research, and writing. It reports that Argon agents freed more than 300 TiB of fleet memory and are migrating C and C++ code to Rust, including up to 800K+ lines for the Fuchsia Zircon kernel. Those rewrites remain under audit before production use.
Who can use Gemini 4 Argon today, and who is next
Today: a subset of Fairwind partners.
Fairwind is Google’s limited-access cyber defense program with more than 650 partners. However, only “a set of Fairwind Program partners” can currently access Argon, either directly or through CodeMender, Google’s code security agent.
Direct access is provided as a managed model on Gemini Enterprise with zero data retention. Partners must commit to:
- Phishing-resistant MFA
- Access restricted to internal security, incident-response, or penetration-testing teams
- Per-employee usage tracking
Governments, critical infrastructure operators, and core technology platforms can apply through the Fairwind Program page.
Next: paid API customers and Google AI Ultra subscribers.
Google says it will gather feedback from early testers while iterating on guardrails. It is also participating in the U.S. government’s voluntary pre-release access process. The rollout section is titled “Rolling out soon,” but Google provides no date.
Track updates in our Gemini 4 Argon release date and access guide.
Not yet: everyone else.
As of October 1, Argon does not appear on the Gemini API models page, pricing page, Gemini app, Antigravity, or OpenRouter.
Google AI Ultra, starting at $99.99 per month in the US, is a consumer subscription rather than an API key. Google has not specified which Ultra tier will receive Argon first. There is no free access path today; see Is Gemini 4 Argon free? for available alternatives.
Gemini 4 Argon pricing
| Token type | Intro price, per 1M | Price after intro, per 1M |
|---|---|---|
| Input | $2.00 | $4.00 |
| Cached input, 95% off | $0.10 | $0.20 |
| Output | $10.00 | $20.00 |
| One full 1M-token response, output only | $10.00 | $20.00 |
Google has not announced how long the introductory period will last.
At standard rates, Argon costs the same as Claude Opus 5.5. At introductory rates, it matches GPT-6.1 Sol and Claude Sonnet 5.5. Cached-input prices are calculated from Google’s stated 95% discount.
For scenario-based estimates, see Gemini 4 Argon pricing.
The 1M output limit
Google says it is “significantly expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens.”
For developers, this means a single response can theoretically contain hundreds of thousands of generated tokens. Google’s argument is that this lets the model work through complex problems in one trajectory.
For comparison:
| Model | Maximum synchronous output |
|---|---|
| Gemini 4 Argon | 1M tokens |
| GPT-6 Astra | 128K tokens |
| Claude Opus 5.5 | 128K tokens |
| Claude Fable 5.1 | 128K tokens |
Two implementation caveats matter:
- This is an output limit, not a published context-window specification. Google has not published Argon’s input window.
- Use the limit reported by your endpoint. Vals AI, for example, lists a 262K maximum output for the configuration it tested.
When Argon becomes available, set explicit client-side timeout and output-cost limits rather than assuming every request can safely generate 1M tokens. See the guide to Argon’s 1M output tokens for timeout, streaming, and cost-cap guidance.
Gemini 4 Argon benchmarks: the headline
All benchmark figures below come from Google’s table. Argon’s scores are Google-reported, with some self-computed and others sourced from leaderboards such as Vals AI. Competitor scores are primarily vendor-reported or public-leaderboard results, and benchmark harnesses differ.
By our recount, Argon leads 13 of 19 rows outright, ties one, and trails on five.
| Benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Fable 5.1 | Claude Opus 5.5 |
|---|---|---|---|---|
| Vals Index | 68.9% | 63.1% | 65.8% | 67.0% |
| DeepSWE v1.1 | 77.9% | 74.1% | 67.4% | 74.2% |
| GraphWalks 256K to 1M, F1 | 84.2% | 71.8% | 65.0% | 66.8% |
| FrontierSWE v2 | 55.0% | 65.5% | 56.3% | 62.3% |
| Terminal-bench 4.0 | 57.4% | 58.2% | 57.9% | 66.4% |
| PostTrainBench | 45.3% | 44.3% | 40.2% | 49.3% |
The rows Argon loses include:
- FrontierSWE v2
- Terminal-Bench Science 0.1
- OSWorld-2.0
- Terminal-bench 4.0
- PostTrainBench
Third-party evaluations place Argon at the frontier but do not show a clear overall lead. Artificial Analysis scores Argon at 53 on its Intelligence Index, tied with GPT-6 Astra and Claude Fable 5.1, behind Claude Opus 5.5 at 58 and Claude Sonnet 5.5 at 56 using maximum settings. Vals ranks Argon first among 41 models on the Vals Index.
For the full benchmark table and methodology notes, see Gemini 4 Argon benchmarks.
Cyber defense in one paragraph
Google says Argon can “autonomously find, validate, and patch critical software vulnerabilities.” For trusted defenders and internal teams, Google says it will release Argon without cyber guardrails so they can use its cybersecurity defense capabilities.
On DeepMind’s cyber page, Argon scores:
- 85.8% on real-world vulnerability discovery, compared with 71.0% for Gemini 3.8 Flash Cyber
- 68% on CWE-bench v1, tied with Grok 4.7 and GPT-6 Astra
- 0.7% Gray Swan indirect prompt-injection attack success rate at k=15, the lowest score on that chart
For API implications, read what Argon’s cyber release means for the APIs you run and the overview of Gemini 3.8 Flash Cyber.
What to do now as a developer
Do not block your implementation on Argon. Build with a model you can call today, and isolate the model name in one environment variable.
Use gemini-3.8-flash as the current stand-in:
- Stable model
- 1,048,576 input tokens
- 65,536 output tokens
- Free tier
- $0.75 input / $3.75 output per million tokens through December 31, 2026
Google’s Interactions API documentation says that all new models, multimodal capabilities, tools, and agentic features will launch on the Interactions API. Build against that API now, then replace the model variable when Google publishes Argon’s model ID.
MODEL="${GEMINI_MODEL:-gemini-3.8-flash}"
curl -s "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "'"$MODEL"'",
"input": "List the breaking changes in this OpenAPI diff.",
"generation_config": {"thinking_level": "medium"}
}'
Make the model switch testable
In Apidog, save this call as a reusable request and define these environment variables:
GEMINI_API_KEY=your_api_key
GEMINI_MODEL=gemini-3.8-flash
Then add response assertions for usage fields:
usage.total_output_tokensusage.total_thought_tokens
Thinking tokens bill as output tokens, so use both fields when enforcing a token budget. Keep the Gemini 3.8 Flash response as a baseline. Once Argon appears in the official models list:
- Change
GEMINI_MODEL. - Re-run the same request.
- Compare latency, output quality, and token usage.
- Keep the token assertion enabled to catch unexpectedly large responses.
For request shapes, see the Gemini 4 Argon API guide. To decide whether waiting is worthwhile, read Argon vs. Gemini 3.8 Flash.
FAQ
Can I use Gemini 4 Argon in the Gemini API?
Not yet. Only a set of Fairwind partners have access, and the public models list does not include it.
What is the Gemini 4 Argon model ID?
Google has not published one. Strings currently in circulation, including evaluator slugs, are third-party placeholders. Do not hard-code them.
When is the Gemini 4 Argon release for developers?
Google has not provided a date. Paid API customers and Google AI Ultra subscribers are expected to come first, “as soon as possible.”
Is Gemini 4 Argon better than GPT-6 Astra and Claude Opus 5.5?
It leads most rows in Google’s table, but it ties GPT-6 Astra on Artificial Analysis and trails Claude Opus 5.5 there. For more detail, see what is Claude Opus 5.5 and the GPT-6 Astra API guide.
Does Gemini 4 Argon have a 1M context window?
Google has not said. The published 1M figure is an output limit. Google’s long-context evaluation used prompts up to 1M tokens, but that is not a published context-window specification.
Get ready before Argon opens
Argon is announced, priced, and benchmarked, but you cannot call it through the public API yet.
Build with Gemini 3.8 Flash using the Gemini 3.8 Flash API guide, keep the model ID in one environment variable, and add a token-usage assertion to every request.
To prepare your request collection, environment variables, and baseline responses for the future model swap, download Apidog and configure it now.


Top comments (0)