Gemini 4 Argon leads 13 of the 19 rows in Google’s launch table outright, ties one row (CWE-bench v1, with GPT-6 Astra), and trails on five: FrontierSWE v2, Terminal-bench 4.0, PostTrainBench, Terminal-Bench Science 0.1, and OSWorld-2.0. The access limitation matters: Argon is currently available only to Fairwind Program defenders, so developers outside Google’s partner list and a few benchmark organizations cannot rerun these numbers yet.
Below is the full table, including who measured each row, where Argon loses, methodology caveats, cyber scores, third-party scoreboards, and a practical way to prepare your own evaluation in Apidog before access opens. For the model overview, start with what Gemini 4 Argon is. For a buying comparison, see Argon vs GPT-6 Astra vs Claude Opus 5.5.
The full table, with who measured each row
These are Google-reported results from the launch post and the evals methodology PDF, which is headed “results as of October, 2026.”
Google computed 10 of the 19 Argon scores itself. The other nine come from public leaderboards. Rival scores come from leaderboards, Google runs, vendor system cards, or vendor blog posts, depending on the row. Harnesses are not uniform across every benchmark.
Bold indicates the row leader.
| Benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Fable 5.1 | Claude Opus 5.5 | Who measured it |
|---|---|---|---|---|---|
| Vals Index | 68.9% | 63.1% | 65.8% | 67.0% | Vals AI |
| AutomationBench | 51.3% | 41.4% | 31.4% | 42.5% | Zapier leaderboard (private set) |
| Vals Finance Agent v2 | 65.4% | 53.5% | 58.9% | 58.6% | Vals AI |
| Harvey Legal Agent | 19.6% | 5.4% | 6.7% | 3.8% | Vals AI |
| DeepSWE v1.1 | 77.9% | 74.1% | 67.4% | 74.2% | Argon self-computed; Astra leaderboard; Anthropic system cards |
| FrontierSWE v2 | 55.0% | 65.5% | 56.3% | 62.3% | Proximal leaderboard |
| Vibe Code Bench | 91.9% | 89.6% | 90.3% | 90.3% | Vals AI |
| Terminal-bench 4.0 | 57.4% | 58.2% | 57.9% | 66.4% | Argon self-computed; rivals leaderboard |
| PostTrainBench | 45.3% | 44.3% | 40.2% | 49.3% | Self-computed for all models |
| Terminal-Bench Science 0.1 | 57.6% | 68.1% | 52.6% | 63.3% | Argon self-computed; rivals leaderboard |
| LABBench 2 | 88.8% | 85.4% | 68.6% | 73.1% | Self-computed for all models |
| RiemannBench | 76.0% | 72.0% | 65.6% | 69.6% | Surge leaderboard |
| GraphWalks up to 128K | 99.7% | 98.7% | 91.4% | 90.6% | Self-computed for all models |
| GraphWalks 256K to 1M | 84.2% | 71.8% | 65.0% | 66.8% | Self-computed for all models |
| Agent’s Last Exam | 39.5% | 34.2% | n/r | 38.2% | Argon self-computed; rivals leaderboard |
| OSWorld-2.0 (offline) | 69.2% | 72.6% | n/r | n/r | Argon self-computed; Astra from OpenAI’s blog post |
| Chartography | 71.6% | 71.0% | 46.2% | 66.3% | Surge leaderboard |
| LVBench | 91.7% | 87.5% | 79.7% | 83.7% | Self-computed for all models |
| CWE-bench v1 | 68.0% | 68.0% | 58.0% | 67.0% | Public leaderboard |
n/r means not reported.
Grok 4.7 is not in Google’s table, but it also scores 68% on the CWE-bench leaderboard. That makes CWE-bench a three-way tie.
Read the source distribution before using the headline count
The source mix is important when you use these results to select a model:
-
Nine rows come from public leaderboards for every model: Vals AI, Zapier, Proximal, Surge, and CWE-bench.
- Argon leads seven.
- Argon ties one.
-
Five rows were run by Google for all four models.
- Argon leads four.
-
Five rows combine a Google-run Argon score with rival scores from another source.
- Three of Argon’s five losses occur in these mixed-source rows.
This does not eliminate methodology concerns, but it is relevant context: if Google self-computing systematically favored Argon, the losses would be expected to cluster in Google-run rows rather than mixed rows.
Where Argon trails, and why it matters for agentic coding
Argon’s five losses are concentrated in benchmarks that matter for terminal-based and repository-scale agents:
- FrontierSWE v2: 55.0% versus Astra’s 65.5%. Argon is last among the four models, behind Fable 5.1 at 56.3% and Opus 5.5 at 62.3%. This is a public leaderboard result.
- Terminal-bench 4.0: 57.4% versus Opus 5.5’s 66.4%. Argon is last again, although Astra and Fable 5.1 are within one point.
- PostTrainBench: 45.3% versus Opus 5.5’s 49.3%. Argon finishes second in a row Google ran for all models.
- Terminal-Bench Science 0.1: 57.6% versus Astra’s 68.1%. Argon is third, ahead of Fable 5.1 only. Google ran Argon’s attempt with a 6x verifier timeout.
- OSWorld-2.0 (offline subset): 69.2% versus Astra’s 72.6%. Anthropic reports only a combined online and offline score, so neither Claude model appears.
For agentic coding specifically, the results split two to two:
| Benchmark type | Result |
|---|---|
| DeepSWE v1.1 | Argon wins |
| Vibe Code Bench | Argon wins |
| FrontierSWE v2 | Argon finishes last |
| Terminal-bench 4.0 | Argon finishes last |
The DeepSWE result mixes harnesses: Argon used a mini-swe agent harness, Astra’s score comes from the public leaderboard, and Claude scores come from system cards. Vibe Code Bench is cleaner because Vals AI scored all four models, but Argon’s winning margin is only 1.6 points.
If your application depends on terminal automation, long-running repository tasks, or shell-based debugging, prioritize FrontierSWE v2 and Terminal-bench 4.0 over the aggregate win count. Astra holds the FrontierSWE lead; see the GPT-6 Astra hands-on for current usage observations. For Anthropic’s results, see the Claude Fable 5.1 benchmarks breakdown.
Bloomberg reports that anonymous Google employees with access say Argon underwhelms on some coding and front-end design work relative to its benchmark scores. Google called that characterization inaccurate.
Methodology caveats that change how you read the table
Before treating a row as a production forecast, check the evaluation configuration.
- Maximum effort settings: Argon ran “with the Gemini API with the highest thinking settings.” Google used the “maximum thinking/reasoning settings available” for Astra, Fable 5.1, and Opus 5.5. These results compare ceilings, not default latency, default quality, or cost.
- LVBench input differences: Google ran LVBench for all four models without tools. Gemini sampled video at 1 FPS. Astra received 800 frames, Fable 5.1 received 300, and Opus 5.5 received 600 due to API limitations. The models did not receive identical inputs.
- OSWorld best-of-three scoring: Argon’s score was “maxed over 3 runs with a single attempt per run.” This is the best run, not an average across runs.
- Agent’s Last Exam execution window: Argon ran on the ALE-Claw harness with a five-hour window.
- Safety filters remained enabled: Agent’s Last Exam and OSWorld used safety filters. Flagged responses returned as empty strings. This could reduce Argon’s result rather than inflate it.
The cyber rows from DeepMind’s cyber page
The DeepMind cyber page adds four scores beyond the launch table:
| Cyber eval | Gemini 4 Argon | Comparison |
|---|---|---|
| Real-world vulnerability discovery | 85.8% | Gemini 3.8 Flash Cyber: 71.0% |
| Wiz Penetration Test Benchmark | 70.9% | Gemini 3.8 Flash Cyber: 58.2% |
Gray Swan IPI attack success (k=15, lower is better) |
0.7% | Lowest on the chart; Kimi K3 highest at 52.7% |
| CWE-bench v1 | 68% | Three-way tie with Grok 4.7 and GPT-6 Astra; Opus 5.5 at 67% |
The vulnerability evaluation used an internal Antigravity harness that Google describes as “not cyber specialized,” with source access.
The Wiz test gives the model access only to the running website and its public behavior, without application source code.
Both headline comparisons are against Google’s prior cyber model rather than direct competitors. The Argon cyber defense explainer covers what these scores mean for APIs you operate.
Third-party scoreboards
Several organizations published scores within minutes of launch, which indicates pre-release access.
- Artificial Analysis: Lists Gemini 4 Argon (High) at an Intelligence Index of 53, ranked #8 of 223. The ranking counts each reasoning setting separately. By distinct model, Argon ties Fable 5.1 and GPT-6 Astra, while trailing Opus 5.5 at 58 maximum and Sonnet 5.5 at 56 maximum. Artificial Analysis reports a 15% hallucination rate on AA-Omniscience, compared with Astra at maximum effort at 51%, with accuracy of 50% versus 63%. Running the index cost $1.99 per task, with approximately 62K output tokens per task versus approximately 27K for Astra at maximum settings. Artificial Analysis shows no speed or latency data.
- Vals AI: Places Argon at 68.90% on the Vals Index, #1 of 41 and the first Gemini model to top the index, at $15.68 per test. Vals AI lists a 262K maximum output for the tested configuration, below Google’s stated 1M.
- Arena: Ranks Argon #1 on Text at 1525, marked Preliminary with 4,942 votes, and #8 on WebDev.
Why outlets publish 12, 13, and 14 wins
Different headlines use different comparison rules. Here is the recount from Google’s 19 rows:
| Counted against | Argon wins | Ties | Argon loses | Not reported |
|---|---|---|---|---|
| Best of all three rivals | 13 | 1 | 5 | 0 |
| GPT-6 Astra only | 14 | 1 | 4 | 0 |
| Claude Opus 5.5 only | 14 | 0 | 4 | 1 |
| Claude Fable 5.1 only | 15 | 0 | 2 | 2 |
Use the count that matches your decision:
- 13 wins is correct when Argon must beat the entire field.
- 14 wins is correct in a head-to-head comparison with Astra or Opus 5.5.
- 12 wins does not match these full-table framings, so inspect which rows the source excluded or counted differently.
Also check the margins. Harvey Legal Agent is a wide gap at 19.6% versus 6.7%, while Chartography is separated by only 0.6 points.
How to run your own eval the day access opens
Public benchmarks measure someone else’s harness. Your evaluation should measure your prompts, tools, schemas, latency requirements, and cost limits.
Build the scenario now using a model you can call. When Argon becomes available, change one environment variable and rerun the same tests.
1. Select production-like test cases
Create prompts that map to the work you actually need:
- A repository bug fix with expected tests or a patch.
- A terminal task requiring command sequencing.
- A long-document question with a verifiable answer.
- A structured API response that must match a JSON schema.
- A tool-calling workflow with error handling.
Use real sanitized incidents when possible. Include both successful and failure cases.
2. Create environment variables in Apidog
In Apidog, create an environment with:
GEMINI_API_KEY=<your-key>
GEMINI_MODEL=gemini-3.8-flash
Google has not published Argon’s model ID, so GEMINI_MODEL should be the only value you need to change later.
3. Parameterize each request
Save each prompt as a request and reference the model variable rather than hardcoding the model ID.
{
"model": "{{GEMINI_MODEL}}",
"contents": [
{
"role": "user",
"parts": [
{
"text": "Fix the failing test. Return the patch and a short explanation."
}
]
}
]
}
Group related requests into a scenario, such as:
agentic-coding/
repository-fix
terminal-debugging
code-review
long-context-question
4. Add assertions for quality and cost controls
Do not evaluate only HTTP status. Assert the properties your integration requires:
- Response status is successful.
- Required JSON fields exist.
- The output includes expected terms, files, commands, or citations.
- The response is valid JSON when your application expects JSON.
- Token usage remains below your cost ceiling.
- Tool calls, if used, match expected names and arguments.
For example, track usageMetadata and cap thought-token usage:
pm.test("returns a successful response", () => {
pm.response.to.have.status(200);
});
const body = pm.response.json();
pm.expect(body.usageMetadata).to.have.property("thoughtsTokenCount");
pm.expect(body.usageMetadata.thoughtsTokenCount).to.be.below(50000);
pm.expect(body).to.have.property("candidates");
Size the thoughtsTokenCount ceiling to your own budget and latency target rather than treating maximum reasoning as the default.
5. Capture a baseline before Argon is public
Run the scenario with Gemini 3.8 Flash and save the report.
Record at least:
| Metric | Why it matters |
|---|---|
| Pass rate | Measures task success against your assertions |
| Output validity | Detects schema and parsing failures |
| Token usage | Tracks model cost behavior |
| Latency | Shows whether a higher-quality result is usable in your workflow |
| Failure examples | Helps diagnose regressions after a model swap |
When Argon’s model ID is released, update:
GEMINI_MODEL=<Argon model ID>
Then rerun the exact same scenario and compare reports.
Google says that “all new models” will launch on the Interactions API, so save an Interactions API version of each request as well. For expected price and limit differences, see Argon vs Gemini 3.8 Flash.
FAQ
Are Gemini 4 Argon’s benchmarks independently verified?
Partly. Nine of the 19 rows come from public leaderboards for every model, and Artificial Analysis, Vals AI, and Arena posted their own results. The remaining rows are Google-run for Argon.
What is Gemini 4 Argon’s DeepSWE score?
Argon scores 77.9% on DeepSWE v1.1, ahead of Opus 5.5 at 74.2% and Astra at 74.1%. Argon used a mini-swe agent harness, while rival scores come from a leaderboard and system cards.
Does Argon beat GPT-6 Astra on benchmarks?
Head to head in Google’s table, Argon wins 14 rows, ties one, and loses four. On Artificial Analysis, they tie at 53.
Can I rerun these benchmarks myself?
Not yet. Argon is limited to Fairwind partners, and Google has not provided a public release date. The Argon release date tracker follows the rollout.
Next step
Build the scenario now, run it on Gemini 3.8 Flash, and keep the report. When Argon becomes available, swap the model variable and you will have comparable numbers from your own workload within minutes.
Download Apidog to set up the evaluation.
Top comments (0)