DEV Community

Cover image for Gemini 4 Argon Benchmarks: All 19 Rows, How Google Ran Them, and the 5 It Loses
Hassann
Hassann

Posted on Originally published at apidog.com

Gemini 4 Argon Benchmarks: All 19 Rows, How Google Ran Them, and the 5 It Loses

Gemini 4 Argon leads 13 of the 19 rows in Google’s launch table outright, ties one row (CWE-bench v1, with GPT-6 Astra), and trails on five: FrontierSWE v2, Terminal-bench 4.0, PostTrainBench, Terminal-Bench Science 0.1, and OSWorld-2.0. The access limitation matters: Argon is currently available only to Fairwind Program defenders, so developers outside Google’s partner list and a few benchmark organizations cannot rerun these numbers yet.

Try Apidog today

Below is the full table, including who measured each row, where Argon loses, methodology caveats, cyber scores, third-party scoreboards, and a practical way to prepare your own evaluation in Apidog before access opens. For the model overview, start with what Gemini 4 Argon is. For a buying comparison, see Argon vs GPT-6 Astra vs Claude Opus 5.5.

The full table, with who measured each row

These are Google-reported results from the launch post and the evals methodology PDF, which is headed “results as of October, 2026.”

Google computed 10 of the 19 Argon scores itself. The other nine come from public leaderboards. Rival scores come from leaderboards, Google runs, vendor system cards, or vendor blog posts, depending on the row. Harnesses are not uniform across every benchmark.

Bold indicates the row leader.

Benchmark Gemini 4 Argon GPT-6 Astra Claude Fable 5.1 Claude Opus 5.5 Who measured it
Vals Index 68.9% 63.1% 65.8% 67.0% Vals AI
AutomationBench 51.3% 41.4% 31.4% 42.5% Zapier leaderboard (private set)
Vals Finance Agent v2 65.4% 53.5% 58.9% 58.6% Vals AI
Harvey Legal Agent 19.6% 5.4% 6.7% 3.8% Vals AI
DeepSWE v1.1 77.9% 74.1% 67.4% 74.2% Argon self-computed; Astra leaderboard; Anthropic system cards
FrontierSWE v2 55.0% 65.5% 56.3% 62.3% Proximal leaderboard
Vibe Code Bench 91.9% 89.6% 90.3% 90.3% Vals AI
Terminal-bench 4.0 57.4% 58.2% 57.9% 66.4% Argon self-computed; rivals leaderboard
PostTrainBench 45.3% 44.3% 40.2% 49.3% Self-computed for all models
Terminal-Bench Science 0.1 57.6% 68.1% 52.6% 63.3% Argon self-computed; rivals leaderboard
LABBench 2 88.8% 85.4% 68.6% 73.1% Self-computed for all models
RiemannBench 76.0% 72.0% 65.6% 69.6% Surge leaderboard
GraphWalks up to 128K 99.7% 98.7% 91.4% 90.6% Self-computed for all models
GraphWalks 256K to 1M 84.2% 71.8% 65.0% 66.8% Self-computed for all models
Agent’s Last Exam 39.5% 34.2% n/r 38.2% Argon self-computed; rivals leaderboard
OSWorld-2.0 (offline) 69.2% 72.6% n/r n/r Argon self-computed; Astra from OpenAI’s blog post
Chartography 71.6% 71.0% 46.2% 66.3% Surge leaderboard
LVBench 91.7% 87.5% 79.7% 83.7% Self-computed for all models
CWE-bench v1 68.0% 68.0% 58.0% 67.0% Public leaderboard

n/r means not reported.

Grok 4.7 is not in Google’s table, but it also scores 68% on the CWE-bench leaderboard. That makes CWE-bench a three-way tie.

Read the source distribution before using the headline count

The source mix is important when you use these results to select a model:

  • Nine rows come from public leaderboards for every model: Vals AI, Zapier, Proximal, Surge, and CWE-bench.
    • Argon leads seven.
    • Argon ties one.
  • Five rows were run by Google for all four models.
    • Argon leads four.
  • Five rows combine a Google-run Argon score with rival scores from another source.
    • Three of Argon’s five losses occur in these mixed-source rows.

This does not eliminate methodology concerns, but it is relevant context: if Google self-computing systematically favored Argon, the losses would be expected to cluster in Google-run rows rather than mixed rows.

Where Argon trails, and why it matters for agentic coding

Argon’s five losses are concentrated in benchmarks that matter for terminal-based and repository-scale agents:

  • FrontierSWE v2: 55.0% versus Astra’s 65.5%. Argon is last among the four models, behind Fable 5.1 at 56.3% and Opus 5.5 at 62.3%. This is a public leaderboard result.
  • Terminal-bench 4.0: 57.4% versus Opus 5.5’s 66.4%. Argon is last again, although Astra and Fable 5.1 are within one point.
  • PostTrainBench: 45.3% versus Opus 5.5’s 49.3%. Argon finishes second in a row Google ran for all models.
  • Terminal-Bench Science 0.1: 57.6% versus Astra’s 68.1%. Argon is third, ahead of Fable 5.1 only. Google ran Argon’s attempt with a 6x verifier timeout.
  • OSWorld-2.0 (offline subset): 69.2% versus Astra’s 72.6%. Anthropic reports only a combined online and offline score, so neither Claude model appears.

For agentic coding specifically, the results split two to two:

Benchmark type Result
DeepSWE v1.1 Argon wins
Vibe Code Bench Argon wins
FrontierSWE v2 Argon finishes last
Terminal-bench 4.0 Argon finishes last

The DeepSWE result mixes harnesses: Argon used a mini-swe agent harness, Astra’s score comes from the public leaderboard, and Claude scores come from system cards. Vibe Code Bench is cleaner because Vals AI scored all four models, but Argon’s winning margin is only 1.6 points.

If your application depends on terminal automation, long-running repository tasks, or shell-based debugging, prioritize FrontierSWE v2 and Terminal-bench 4.0 over the aggregate win count. Astra holds the FrontierSWE lead; see the GPT-6 Astra hands-on for current usage observations. For Anthropic’s results, see the Claude Fable 5.1 benchmarks breakdown.

Bloomberg reports that anonymous Google employees with access say Argon underwhelms on some coding and front-end design work relative to its benchmark scores. Google called that characterization inaccurate.

Methodology caveats that change how you read the table

Before treating a row as a production forecast, check the evaluation configuration.

  • Maximum effort settings: Argon ran “with the Gemini API with the highest thinking settings.” Google used the “maximum thinking/reasoning settings available” for Astra, Fable 5.1, and Opus 5.5. These results compare ceilings, not default latency, default quality, or cost.
  • LVBench input differences: Google ran LVBench for all four models without tools. Gemini sampled video at 1 FPS. Astra received 800 frames, Fable 5.1 received 300, and Opus 5.5 received 600 due to API limitations. The models did not receive identical inputs.
  • OSWorld best-of-three scoring: Argon’s score was “maxed over 3 runs with a single attempt per run.” This is the best run, not an average across runs.
  • Agent’s Last Exam execution window: Argon ran on the ALE-Claw harness with a five-hour window.
  • Safety filters remained enabled: Agent’s Last Exam and OSWorld used safety filters. Flagged responses returned as empty strings. This could reduce Argon’s result rather than inflate it.

The cyber rows from DeepMind’s cyber page

The DeepMind cyber page adds four scores beyond the launch table:

Cyber eval Gemini 4 Argon Comparison
Real-world vulnerability discovery 85.8% Gemini 3.8 Flash Cyber: 71.0%
Wiz Penetration Test Benchmark 70.9% Gemini 3.8 Flash Cyber: 58.2%
Gray Swan IPI attack success (k=15, lower is better) 0.7% Lowest on the chart; Kimi K3 highest at 52.7%
CWE-bench v1 68% Three-way tie with Grok 4.7 and GPT-6 Astra; Opus 5.5 at 67%

The vulnerability evaluation used an internal Antigravity harness that Google describes as “not cyber specialized,” with source access.

The Wiz test gives the model access only to the running website and its public behavior, without application source code.

Both headline comparisons are against Google’s prior cyber model rather than direct competitors. The Argon cyber defense explainer covers what these scores mean for APIs you operate.

Third-party scoreboards

Several organizations published scores within minutes of launch, which indicates pre-release access.

  • Artificial Analysis: Lists Gemini 4 Argon (High) at an Intelligence Index of 53, ranked #8 of 223. The ranking counts each reasoning setting separately. By distinct model, Argon ties Fable 5.1 and GPT-6 Astra, while trailing Opus 5.5 at 58 maximum and Sonnet 5.5 at 56 maximum. Artificial Analysis reports a 15% hallucination rate on AA-Omniscience, compared with Astra at maximum effort at 51%, with accuracy of 50% versus 63%. Running the index cost $1.99 per task, with approximately 62K output tokens per task versus approximately 27K for Astra at maximum settings. Artificial Analysis shows no speed or latency data.
  • Vals AI: Places Argon at 68.90% on the Vals Index, #1 of 41 and the first Gemini model to top the index, at $15.68 per test. Vals AI lists a 262K maximum output for the tested configuration, below Google’s stated 1M.
  • Arena: Ranks Argon #1 on Text at 1525, marked Preliminary with 4,942 votes, and #8 on WebDev.

Why outlets publish 12, 13, and 14 wins

Different headlines use different comparison rules. Here is the recount from Google’s 19 rows:

Counted against Argon wins Ties Argon loses Not reported
Best of all three rivals 13 1 5 0
GPT-6 Astra only 14 1 4 0
Claude Opus 5.5 only 14 0 4 1
Claude Fable 5.1 only 15 0 2 2

Use the count that matches your decision:

  • 13 wins is correct when Argon must beat the entire field.
  • 14 wins is correct in a head-to-head comparison with Astra or Opus 5.5.
  • 12 wins does not match these full-table framings, so inspect which rows the source excluded or counted differently.

Also check the margins. Harvey Legal Agent is a wide gap at 19.6% versus 6.7%, while Chartography is separated by only 0.6 points.

How to run your own eval the day access opens

Public benchmarks measure someone else’s harness. Your evaluation should measure your prompts, tools, schemas, latency requirements, and cost limits.

Build the scenario now using a model you can call. When Argon becomes available, change one environment variable and rerun the same tests.

1. Select production-like test cases

Create prompts that map to the work you actually need:

  • A repository bug fix with expected tests or a patch.
  • A terminal task requiring command sequencing.
  • A long-document question with a verifiable answer.
  • A structured API response that must match a JSON schema.
  • A tool-calling workflow with error handling.

Use real sanitized incidents when possible. Include both successful and failure cases.

2. Create environment variables in Apidog

In Apidog, create an environment with:

GEMINI_API_KEY=<your-key>
GEMINI_MODEL=gemini-3.8-flash
Enter fullscreen mode Exit fullscreen mode

Google has not published Argon’s model ID, so GEMINI_MODEL should be the only value you need to change later.

3. Parameterize each request

Save each prompt as a request and reference the model variable rather than hardcoding the model ID.

{
  "model": "{{GEMINI_MODEL}}",
  "contents": [
    {
      "role": "user",
      "parts": [
        {
          "text": "Fix the failing test. Return the patch and a short explanation."
        }
      ]
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

Group related requests into a scenario, such as:

agentic-coding/
  repository-fix
  terminal-debugging
  code-review
  long-context-question
Enter fullscreen mode Exit fullscreen mode

4. Add assertions for quality and cost controls

Do not evaluate only HTTP status. Assert the properties your integration requires:

  • Response status is successful.
  • Required JSON fields exist.
  • The output includes expected terms, files, commands, or citations.
  • The response is valid JSON when your application expects JSON.
  • Token usage remains below your cost ceiling.
  • Tool calls, if used, match expected names and arguments.

For example, track usageMetadata and cap thought-token usage:

pm.test("returns a successful response", () => {
  pm.response.to.have.status(200);
});

const body = pm.response.json();

pm.expect(body.usageMetadata).to.have.property("thoughtsTokenCount");
pm.expect(body.usageMetadata.thoughtsTokenCount).to.be.below(50000);

pm.expect(body).to.have.property("candidates");
Enter fullscreen mode Exit fullscreen mode

Size the thoughtsTokenCount ceiling to your own budget and latency target rather than treating maximum reasoning as the default.

5. Capture a baseline before Argon is public

Run the scenario with Gemini 3.8 Flash and save the report.

Record at least:

Metric Why it matters
Pass rate Measures task success against your assertions
Output validity Detects schema and parsing failures
Token usage Tracks model cost behavior
Latency Shows whether a higher-quality result is usable in your workflow
Failure examples Helps diagnose regressions after a model swap

When Argon’s model ID is released, update:

GEMINI_MODEL=<Argon model ID>
Enter fullscreen mode Exit fullscreen mode

Then rerun the exact same scenario and compare reports.

Google says that “all new models” will launch on the Interactions API, so save an Interactions API version of each request as well. For expected price and limit differences, see Argon vs Gemini 3.8 Flash.

FAQ

Are Gemini 4 Argon’s benchmarks independently verified?

Partly. Nine of the 19 rows come from public leaderboards for every model, and Artificial Analysis, Vals AI, and Arena posted their own results. The remaining rows are Google-run for Argon.

What is Gemini 4 Argon’s DeepSWE score?

Argon scores 77.9% on DeepSWE v1.1, ahead of Opus 5.5 at 74.2% and Astra at 74.1%. Argon used a mini-swe agent harness, while rival scores come from a leaderboard and system cards.

Does Argon beat GPT-6 Astra on benchmarks?

Head to head in Google’s table, Argon wins 14 rows, ties one, and loses four. On Artificial Analysis, they tie at 53.

Can I rerun these benchmarks myself?

Not yet. Argon is limited to Fairwind partners, and Google has not provided a public release date. The Argon release date tracker follows the rollout.

Next step

Build the scenario now, run it on Gemini 3.8 Flash, and keep the report. When Argon becomes available, swap the model variable and you will have comparable numbers from your own workload within minutes.

Download Apidog to set up the evaluation.

Top comments (0)