Gemini 4 Pro leads GPT-6 Astra and Claude Fable 5.1 in a circulating benchmark table. That makes it interesting to evaluate. It does not establish that Google has shipped a better model.
The reporting dated September 20, 2026 describes an unreleased model without an official model card, public API endpoint, or confirmed pricing. The scores come from community posts, alleged internal checkpoints, and Arena.ai sightings. Some earlier sheets were reportedly flagged as predictions or partially copied data.
My read: the claimed improvements in coding, computer use, and multimodal reasoning deserve attention. I would still keep production decisions tied to reproducible evaluations and actual access.
Start with the provenance
Several different kinds of evidence have become bundled into the Gemini 4 story.
The strongest context cited in the reporting is Google’s July 2026 earnings discussion of investment in a “larger Gemini 4 base model.” That supports an account of increased training investment. It does not validate a particular checkpoint, benchmark result, context limit, or token price.
The remaining claims fall into three groups:
| Source | What it claims | What remains unresolved |
|---|---|---|
| Posts attributed to leaker Pankaj Kumar and subsequent community reporting | An internal Pro checkpoint, an October release, possibly an earlier Flash-Lite variant | Checkpoint identity and launch timing |
Arena.ai encounters under names such as gemini-3.8-flash
|
Strong SVG, Three.js, structured output, and agent behavior | Whether the model was Gemini 4 |
| A circulating comparison sheet | Benchmark scores, prices, and context limits across several model families | Authenticity, evaluation settings, and reproducibility |
The comparison sheet includes Gemini 4 Pro, Flash, and Flash-Lite alongside GPT-6 Astra, Claude Fable 5.1, Claude Opus 5, and GPT-5.6 Sol.
Google has not confirmed the table or the reported Arena identities. The reporting points to a history of stealth evaluations before launches, but that pattern cannot identify a particular anonymous model.
What the benchmark sheet actually says
I find the numbers easier to assess when the comparison and its limitations sit together. Every result below is an unverified claim from the circulating sheet.
| Benchmark | Gemini 4 Pro | GPT-6 Astra | Claude Fable 5.1 |
|---|---|---|---|
| DeepSWE v1.1 | 88.7% | 86.9% | 69.1% |
| Terminal-bench 2.1 | 95.3% | 94.1% | 92.8% |
| Terminal-bench 4.0 | 69.7% | 66.4% | 57.9% |
| GDPval-AA v2 | 2064 Elo | 1994 Elo | 1853 Elo |
| OSWorld-2.0 | 86.8% | 84.5% | 77.9% |
| HLE-Verified | 72.1% | 67.3% | 62.1% |
| LVBench | 95.8% agentic / 95.0% static | 93.8% | 90.4% |
The OSWorld-2.0 number is described as a partial score with batch tools enabled. LVBench gives separate agentic and static results for Gemini, while the comparison supplies only one figure for each competitor. Those details matter when deciding whether the rows represent equivalent evaluation conditions.
Coding and sustained agent work
DeepSWE v1.1 is presented as a long-horizon software engineering evaluation. The sheet puts Gemini 4 Pro at 88.7%, ahead of Astra’s 86.9%, Fable’s 69.1%, and Claude Opus 5’s 75.0%.
The two Terminal-bench versions measure different workloads in the reporting: version 2.1 covers agentic terminal coding, while version 4.0 covers general agent capabilities. I would keep their results separate rather than treat them as interchangeable measures of coding quality.
If reproduced, these results would make Gemini 4 Pro a serious candidate for repository work: refactoring across files, debugging, running tools, and recovering from failed steps. They would still leave the practical question of how reliably it finishes my tasks.
Knowledge work and specialist evaluations
The sheet’s 2064 GDPval-AA v2 rating would make Gemini 4 Pro the only listed model above 2000 Elo. It also claims:
- Vals Finance Agent v2: 74.2%.
- Harvey’s Legal Agent Benchmark: 18.7% on complex legal workflows, measured by all-pass rate.
- CharXiv Reasoning: 94.7% for reasoning over complex charts without tools.
- BioMysteryBench and LABBench2: leading results on human-solvable and difficult subsets, without exact scores supplied in the article.
I would resist flattening these into a single “professional work” score. An all-pass legal workflow result, a chart reasoning accuracy, and an Elo rating describe different properties.
The same applies to HLE-Verified’s claimed 72.1%: it suggests strong multidisciplinary expert reasoning under that evaluation, subject to verification. It does not establish reliability for every specialist application.
Price and context contain their own uncertainty
The leaked pricing is attractive enough to influence deployment plans—if it survives launch.
| Model | Input per 1M tokens | Output per 1M tokens | Claimed maximum input |
|---|---|---|---|
| Gemini 4 Pro | $2.25 | $11.25 | 2M tokens |
| Gemini 4 Flash | $0.75 | $3.75 | 2M tokens |
| Gemini 4 Flash-Lite | $0.35 | $1.75 | 2M tokens |
| GPT-6 Astra | $12.00 | $60.00 | 1M tokens |
| Claude Fable 5.1 | $10.00 | $50.00 | 200k tokens in the sheet |
The Gemini Pro figures are quoted without caching. They amount to roughly one-fifth of Astra’s listed token prices. Fable 5.1’s reported cache-read discounts complicate any comparison based solely on uncached rates.
The reporting also gives Gemini 3.x Pro pricing as $2–$4 per million input tokens and $12–$18 per million output tokens, depending on context length. That makes the leaked prices look plausible, but plausibility is a weak substitute for a published rate card.
The context numbers disagree
The sheet lists 2M input tokens for the Gemini 4 family, 1M for Astra, and 200k for the Claude Fable/Opus 5 line.
However, the same reporting says official Anthropic documentation commonly gives Fable 5.1 a 1M context window. The 200k entry may describe a particular evaluation configuration or simply be wrong. I would not use it to claim a confirmed tenfold context advantage.
Other community accounts mention targets around 1.5M tokens or higher, and separate “ghost-routing” claims suggest 10M ceilings. These are distinct claims, not a coherent specification.
For repository agents, I care about retrieval fidelity and task completion as the context fills. A maximum token count alone cannot answer either question.
What the Arena demos can tell us
Reported encounters with the anonymous model include SVG scenes such as a pelican riding a bicycle, mechanical butterflies built in Three.js, long JSON outputs without obvious repetition, and capable agent behavior.
Community comparisons reportedly judged some of these outputs superior or competitive with Astra. I understand why those examples spread: complex visual output makes differences immediately visible.
Their evidentiary limit is substantial, though. A compelling result from a placeholder model does not establish its identity, typical performance, or production configuration.
Claims about multi-million-token context, 256k output limits, cross-session memory, and native internet or tool access also need separate confirmation. A demo may combine model behavior with capabilities supplied by the surrounding application.
I would save the prompts as future evaluation cases. I would avoid using the screenshots as a deployment ranking.
Architecture and reasoning: separate reports from inference
The reporting attributes “most ambitious pre-training run yet” language to Google and describes a significantly larger base model. It then infers possible increases in parameter count, effective capacity, training compute, and multimodal coverage.
Those are reasonable possibilities. They do not reveal whether the architecture uses denser scaling, mixture-of-experts, different sparsity, or another approach.
The alleged argon checkpoint
Descriptions of an early checkpoint called argon mention a High thinking-effort mode, roughly 2.4 minutes for one generation, and an output ceiling of 256k tokens, compared with a reported previous limit of approximately 64k.
If accurate, those details suggest room for more computation and longer generations. They do not establish typical latency, effective output quality, or production limits.
The reporting expects reasoning controls comparable to current Gemini Flash low, medium, and high thinking levels. Until documented, I would treat that as an interface expectation.
The expected capability direction
The proposed focus is consistent across the leaks: sustained software engineering, tool use, context reliability, and multimodal understanding.
The article connects those expectations to agentic features attributed to Gemini 3.5/3.8 Flash and the September 2026 Gemini 3.8 Live models. Expected improvements include function calling, code execution, search grounding, structured JSON/XML outputs, video comprehension, spatial reasoning, and voice interactions.
For application development, I would turn those expectations into test cases. “Better tool use” becomes successful function selection and argument construction. “Long-context reliability” becomes retrieving the right evidence from a large input. “Stronger agents” becomes completing a task despite intermediate failures.
The RSI claim needs different evidence
Recursive self-improvement has become attached to the launch speculation, but the reporting provides no independently verified evidence that Gemini 4 autonomously improved its own model weights through a closed loop.
It references executive comments connecting AI investment to longer-term self-improvement goals and research such as Dream-RSI, where agents improve exploration strategies without changing model weights.
That distinction is central. Improving an agent’s strategy, using AI in model research, and autonomously producing better model generations are different claims. The leaked benchmark table cannot establish the last one.
How I’d prepare an evaluation
The September 20 account describes Astra and Fable 5.1 as publicly available and independently evaluated, with different strengths and closely matched overall results on third-party trackers. Gemini 4 Pro remains an alleged upcoming option in that snapshot.
I would keep building against available models and prepare a repeatable comparison:
- Freeze representative tasks. Include repository changes, terminal work, structured outputs, and the multimodal inputs the application actually receives.
- Record execution conditions. Capture tool access, thinking settings, context size, caching, retry policy, and agent configuration.
- Measure completed work. Track correctness, latency, failures, and total spend across retries and tool loops.
- Re-run when public access exists. Verify the documented model ID, limits, and pricing before interpreting results.
- Route by observed performance. A cheaper Flash variant may suit one workload while a Pro model earns its cost on another.
A unified multi-model API such as CometAPI can simplify that comparison through an OpenAI-compatible interface, provided the required models and features are available. I would still check feature support rather than assume that changing a model parameter preserves every provider-specific behavior.
October 2026 is the reported release expectation, with some late-September speculation and a possible earlier Flash-Lite release. None is an official commitment in the supplied reporting.
My threshold for switching is straightforward: public access, documented operating limits, and better results on the workloads I need to ship. The leaked numbers give me reasons to run those tests. They do not supply the results.
Originally published at cometapi.com
Top comments (0)