DEV Community

Yong Yu
Yong Yu

Posted on Originally published at yongboyu.hashnode.dev

Gemini 4 Argon Just Dropped — But You Can't Call It Yet: Pricing, 1M Output Tokens, and What to Do Until the API Lands

Attributed compile (not original research)

Primary source: Gemini 4 Argon: our next era of frontier intelligence by Koray Kavukcuoglu (Google DeepMind / Google blog), published 2026-09-30

Secondary roundup: Google announces Gemini 4 Argon as its new frontier model (9to5Google), published 2026-09-30

Optional context: Google's first Gemini 4 model is 'Argon' (Engadget), published 2026-10-01

Contrast only (not the headline): Introducing GPT-6.1 Sol (OpenAI) and GPT-6.1 Sol in GitHub Copilot (GitHub Changelog), both 2026-09-29

This post restates Googles launch packaging in my own words, with a builder checklist for what to do before a public model ID exists. I did not run Gemini 4 Argon it is not publicly available yet. Every price, benchmark, and access claim below is vendor-reported (or clearly attributed to a named secondary source).


Google announced Gemini 4 Argon on September 30, 2026, as its next frontier model after the Gemini 3.x line. The capability story is long-horizon software engineering, enterprise knowledge work, and defensive cybersecurity. The access story matters more for builders this week: Argon is live only inside Google’s Fairwind Program for trusted cyber defenders. Broader rollout is “soon,” starting with paid API customers and Google AI Ultra subscribers. There is no public model ID and no first-call snippet for general developers yet.

That phased release is the difference between bookmark the blog” and “ship a production default tomorrow.”

What Google announced (and why Fairwind-first matters)

Per Kavukcuoglu’s post, Argon is built for deep reasoning across complex, multi-step workflows. Google says it is already using Argon internally (thousands of Googlers, vendor claim) for specialized coding, research, and writing. External access starts with Fairwind — and for trusted defenders plus Google’s own teams, Google will ship a version without cyber guardrails so they can use full defensive capability.

For everyone else, Google is still iterating on safeguards (misuse including CBRN/cyber, prompt injection, misalignment monitoring of chain-of-thought and actions, and hardened sandboxes) while engaging the U.S. government’s voluntary pre-release process. Treat Fairwind-only as a hard product constraint: do not plan a customer-facing pipeline on Argon until you have a documented API model ID and a paid access path.

Pricing trap: intro $2/$10 vs post-intro $4/$20

Googles stated launch pricing (vendor):

Per 1M tokens (Google) Introductory After intro period
Input $2 $4
Output $10 $20
Cached input 95% off input price same discount structure stated at launch

Intro sticker matches the familiar $2 / $10 band many teams already budget for mid-frontier coding models. After the introductory period, Google says rates move to $4 / $20 — Opus-class list territory. Cached input at 95% off input price is the quiet lever: long agent loops that re-send the same system prompt, repo summary, or tool schema will live or die on cache hit rate once Argon is in the API.

Build cost models for both columns now. An eval that looks cheap at intro rates can double on input/output when the promo ends. Do not bake “Argon is $2/$10 forever” into a board deck.

Why 1M output tokens changes agent design

Google expands Argon’s output token limit to 1M, up from 64K — framed as industry-leading generation headroom for long trajectories. Context-window talk is saturated. Output headroom is different: one trajectory can think, tool-call, revise, and emit a large artifact (migration plan, multi-file patch set, research memo) without mid-flight truncation.

Loops designed for 64K128K completions assume early summarization or multi-turn “continue. A 1M ceiling (once callable) means fewer artificial breakpoints — and more runaway-spend risk. Add budgets, stop conditions, and per-step logs before you get an API key.

Vendor benchmarks (labeled as such)

These figures come from Google’s announcement and secondary reporting. They are not my harness results.

Benchmark (vendor-reported) Gemini 4 Argon Notes
DeepSWE v1.1 77.9% Long-horizon SWE; 9to5Google cites Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1%
AutomationBench (Zapier) 51.3% Google claims #1
LVBench (long video) 91.7% Google claims SOTA
CWE-bench v1 68% Google claims tied for first (vuln remediation)

Google also points to leading placement on the Vals Index and strong domain results on Vals Finance Agent v2 and Harvey’s Legal Agent Benchmark. Orientation only: different benches measure different jobs. Your golden set still wins.

Internal Google anecdotes (vendor examples, not your SLA)

Capability hints, not SLAs: quantum subroutine optimization (40% better than a published baseline in one example); fleet memory work freeing 300+ TiB; CRust migrations from re2 / libgav1 up to 800K+ lines for Fuchsia Zircon (under heavy audit); and a libgav1 Rust SIMD rewrite claimed 2.7× faster than the prior Rust port with identical video output.

Cyber angle: Argon is trained for defense; Wiz’s Scan for Good is cited for a critical healthcare exposure previous frontier models missed. Vendor narrative — not an independent pen-test for compliance.

What builders should do this week

OpenAI shipped GPT-6.1 Sol on September 29 at the same sticker band ($2 / $10) and it is already in the API and GitHub Copilot per OpenAI and GitHub’s changelog. That contrast is actionable: do not stall shipping while Fairwind widens.

  1. Keep shipping on models you can call today: GPT-6.1 Sol, Claude Opus / Sonnet 5.5, Grok where it fits.
  2. Prepare long-horizon SWE evals now — multi-file edits, flaky tests, migration-style tasks. When Argon’s model ID lands, you want a 2050 trajectory harness ready.
  3. Watch for the public model ID and paid API / Ultra notes. No ID → no production default.
  4. Do not block a release on Fairwind-only access unless you are in that program.
  5. Price for the post-intro column ($4/$20) in any Argon ROI slide; assume cache misses until you measure them.
  6. Security posture: if you ever get a no-guardrail cyber build, isolate it. Fairwind is defense capability plus staged safeguards not “same sandbox as the public chatbot.”

When Argon might become your default (once the API lands)

Only after a callable model ID and billing access:

  1. Workloads are long-horizon and you are hitting today’s completion caps.
  2. A golden eval shows Argon beating your current default on your failures — not DeepSWE screenshots alone.
  3. Cost model survives post-intro $4/$20 and realistic cache hit rates.
  4. Prompt-injection / tool-exfiltration tests pass under your threat model (Google claims Gray Swan IPI leadership — verify on your agents).
  5. You can rollback to Sol / Opus / Sonnet without rewriting orchestration.

Until then, Argon is a roadmap signal. Treat the blog as planning input; treat GPT-6.1 Sol and the Claude 5.5 pair as this week’s shipping surface.

Why this post exists

Frontier launches now split into two products: the capability blog and the availability calendar. Argon is a serious capability claim with a Fairwind-first calendar. The durable skill is refusing to confuse them — and having harnesses, budgets, and rollbacks ready for the day a public model ID appears.


Byline: YongBo Yu — Toronto AI engineer (agents, LLM workflows). GitHub: YongBoYu1.

Top comments (0)