Google dropped Gemini 4 Argon on September 30 and Hacker News lost its mind: 1,300+ points, 870+ comments, the top of the front page.
Then everyone tried to use it and hit a wall. Argon is not generally available. Let's separate the signal from the launch-day noise.
What shipped
Argon is Google's new frontier model, pitched at long-running workflows: software engineering, enterprise knowledge work (law, finance), and cyber defense.
Access today:
- Now: vetted cyber defenders via Google's Fairwind Program, plus Googlers (thousands use it internally in Antigravity).
- Next: paid API customers and Google AI Ultra subscribers.
- Date for "next": none published.
Google says it is iterating on guardrails first and coordinating with the US government's voluntary pre-release access process. Translation: the cyber capability is the reason it's gated. Wiz reportedly used it to find a critical healthcare-software vulnerability that earlier frontier models missed.
The numbers (Google's own evals, so salt accordingly)
| Benchmark | Argon | GPT-6 Astra | Claude Opus 5.5 |
|---|---|---|---|
| DeepSWE v1.1 | 77.9% | 74.1% | 74.2% |
| Terminal-Bench 4.0 | 57.4% | n/a | 66.4% |
| FrontierSWE v2 | 55.0% | 65.5% | n/a |
| Vals Index | 68.9% | n/a | 67.0% |
| GraphWalks (256K-1M) | 84.2% | 71.8% | 66.8% |
| OSWorld-2.0 | 69.2% | 72.6% | n/a |
Independent Vals puts Argon #1 on its index (68.9%) at about $15.68 per test, versus $32.14 for Opus 5.5.
Read that table carefully. Argon wins on repo-level SWE tasks and long-context retrieval. It loses on terminal-driven agent work (Terminal-Bench 4.0, by 9 points to Opus) and on FrontierSWE v2 (10 points to Astra). "Beats everyone on most benchmarks" is true. "Best coding model" is not.
The price is the real story
Introductory pricing is $2 / $10 per million tokens (input/output), with cached input 95% off. Reported comparisons:
- GPT-6 Astra: $10 / $50
- Claude Opus 5.5: $4 / $20
- Argon intro: $2 / $10, reportedly doubling to $4 / $20 later
At intro pricing, that's 5x cheaper than Astra with a higher score on DeepSWE. Even at the doubled price it matches Opus. If it holds up outside Google's evals, this is a margin story for anyone running agents at volume: the cost per resolved task is what your CFO sees, not the leaderboard.
What HN developers actually said
The thread is less about Argon and more about the harness:
- Antigravity gets called "a terrible harness": forced auto-compaction around 250k tokens even though the models support far larger windows, no persistent memory, no stdin/stdout piping, weak Neovim/Emacs story.
- Google restricts the model to its own harness, while competitors let you bring your own. Several people said this prices them out of API-based workflows.
- Account-ban fear is real. Commenters cited bans for accessing models outside official channels.
- Hallucination complaints on research tasks, even from fans of the speed.
Peter Yang's take is the cleanest summary: Google cooked on the model, now it needs to compete on the coding harness and the personal-agent product.
My take
Model quality is converging. What differentiates now is three boring things:
- Cost per solved task, not cost per token.
- Harness quality: compaction, memory, permissions, scripting. Claude Code and Codex win developers here today.
- Access policy: if I can't swap my own harness or API client, I'm renting, not building.
Argon scores well on 1, is unproven on 2, and currently fails 3 for almost everyone.
What to do this week
Don't rewrite anything. Do this instead:
# Keep your model behind one seam so swapping is a config change
MODELS = {
"default": "claude-opus-5-5",
"cheap_bulk": "gemini-4-argon", # flip on when GA
}
def pick(task):
return MODELS["cheap_bulk"] if task.is_bulk_swe else MODELS["default"]
- Build a private eval from 20 of your own real tickets. Public benchmarks disagree with each other by 10 points; yours won't.
- Measure cost per merged PR, not tokens.
- Wait for independent numbers. Right now every Argon score is first-party or from a gated cohort.
- Watch for the GA date and whether third-party harnesses are allowed.
Bottom line
Argon might be the best price-to-performance frontier model on the board. It's also a model you can't call yet, in a harness developers don't like, with a pricing promise that expires. Respect the benchmark, ignore the hype, and keep your abstraction layer clean.
Sources: Hacker News discussion, Vals.ai, OfficeChai, Droid Life, tbreak. Figures reported September 30, 2026; Google's benchmarks are unverified by independent testers.
Top comments (0)