Yesterday SpaceXAI shipped Grok 4.7 with a benchmark table in which Grok 4.7 wins. Three weeks ago OpenAI shipped GPT-6 Astra with a table in which Astra wins. Anthropic did the same with Claude Fable 5.1. None of these tables is fake — but every one of them was put together by the team that ships the model.
So before renewing my own subscriptions I did the boring thing: pulled numbers from eight sources into one matrix. Three independent leaderboards — Artificial Analysis, Vals AI and Arena (blind human voting) — plus the launch posts from OpenAI, Anthropic, Google and SpaceXAI, which, put side by side, check each other surprisingly well.
Snapshot date: September 22, 2026. Every number below is labeled as independent or vendor-reported.
The lineup: six models, four subscriptions
The only rule: the model must be available in a regular paid plan an individual can buy.
| Model | Released | Subscription | Price / month |
|---|---|---|---|
| Claude Fable 5.1 | Sep 1 | Claude Pro / Max | $20 / $100–200 |
| Claude Opus 5 | before Sep | Claude Pro / Max | $20 / $100–200 |
| GPT-6 Astra | Sep 3 | ChatGPT Plus / Pro | $20 / $200 |
| GPT-5.6 Sol | before Sep | ChatGPT Plus / Pro | $20 / $200 |
| Gemini 3.8 Flash | Sep 2 | Google AI Pro / Ultra | $19.99 / $250 |
| Grok 4.7* | Sep 21 | SuperGrok | $30 |
* At launch Grok 4.7 is available in Cursor, Grok Build and the API; SpaceXAI hasn't yet confirmed it in the SuperGrok consumer docs. Check before paying.
Left out on purpose: Meta Muse Spark 1.3 (free in the Meta AI app — not a subscription), Claude Mythos 5.1 (gated to verified security and life-science programs), and open-weight models like DeepSeek, Qwen, Kimi and GLM — they get their own article.
Why you shouldn't take a launch table at face value
Not because vendors lie. Because the same benchmark can be run several ways, and every vendor picks the way that flatters them. Four examples from September alone:
- ARC-AGI-3. OpenAI reports 99.9% for GPT-6 Astra — in a harness that preserves reasoning state between steps. In the neutral harness everyone uses, Astra scores 62.7%.
- HealthBench Professional. SpaceXAI's table gives Claude Fable 5.1 62.1%. The length-adjusted independent leaderboard gives the same model 56.6% — a gap larger than first-to-third place.
- OSWorld 2.0. OpenAI lists Claude Opus 5 at 70.2%. Anthropic lists the same model at 39.6% — in "strict" mode. Both are correct; they're just not comparable.
- CursorBench. SpaceXAI uses v4.0, Anthropic uses v3.2. Fable 5.1 scores 51.8% and 73.4% respectively.
Three rules I now apply: independent rankings beat vendor tables even when they have fewer rows; always check who is missing from a table (Grok 4.7's chart has no GPT-6 Astra); and a number without its mode (max, xhigh, with tools, strict) is not a number.
Independent rankings: a tie at the top
- Artificial Analysis Intelligence Index: Fable 5.1 and GPT-6 Astra — 53 each, Opus 5 — 51, GPT-5.6 Sol — 47, Grok 4.7 — 46.
- Vals Index (finance, legal research, code migration, medical documentation): Fable 5.1 68.8%, Opus 5 67.2%, Astra 66.6%, Gemini 3.8 Flash 62.3%, Grok 4.7 60.2%.
- Arena Text (blind voting, Sep 13): Fable 5.1 1498, Opus 5 and Gemini 3.8 Flash 1493 each — within the confidence interval. Arena Code is a different story: Astra 1800, Fable 1758, Opus 1687.
If you're picking one subscription "for everything", Claude and ChatGPT are effectively equal right now. The decision comes down to the tasks where they diverge — and there are plenty.
The full matrix: 6 models × 19 tests
Leader of each row in bold. "—" means the model wasn't tested or the result isn't published. I = independent, V = vendor-reported.
| Task · test | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol | Gemini 3.8 Flash | Grok 4.7 | Source |
|---|---|---|---|---|---|---|---|
| Overall · AA Index | 53 | 51 | 53 | 47 | <46 | 46 | I · Artificial Analysis |
| Professional tasks · Vals Index | 68.8 | 67.2 | 66.6 | — | 62.3 | 60.2 | I · Vals AI |
| Blind chat · Arena Text | 1498 | 1493 | — | — | 1493 | — | I · Arena |
| Code, blind · Arena Code | 1758 | 1687 | 1800 | — | — | — | I · Arena |
| Terminal agent · Terminal-Bench 4.0 | 55.8 | 52.3 | 57.7 | 37.3 | 19.1 | 38.0 | V · OpenAI, Anthropic |
| Long-horizon SWE · DeepSWE v1.1 | 70.0 | 74.0 | — | 72.7 | 73.7 | 71.0 | V · Google, SpaceXAI |
| In-editor coding · CursorBench 4.0 | 51.8 | — | — | 41.7 | — | 46.3 | V · SpaceXAI |
| Clinical reasoning · HealthBench Pro | 56.6 | — | 63.4 | 60.5 | — | 56.7 | I · aggregated leaderboard |
| Medical coding · Vals MedCode | #7 | 63.6 | #23 | — | — | — | I · Vals AI |
| Legal agent · Harvey LAB | 6.7 | — | — | 2.5 | — | 19.6 | V · SpaceXAI |
| Legal research · Vals Legal Research | — | 55.3 | #21 | — | — | — | I · Vals AI |
| Electrical engineering · EEBench | 56.4 | — | — | 39.4 | — | 64.0 | V · SpaceXAI |
| Documents · GDPval-AA (Elo) | 1853 | 1824 | — | 1711 | — | — | V · Anthropic |
| Multi-hour office work · AA Briefcase | 1678 | — | — | 1487 | — | 1657 | V · SpaceXAI |
| Computer use · OSWorld 2.0 | — | 70.2 | 72.6 | 65.7 | — | — | V · OpenAI |
| Research math · FrontierMath T4 | 87.8 | 73.2 | 97.6 | 83.0 | — | — | V · OpenAI |
| Expert knowledge · HLE with tools | 65.0 | 63.6 | 57.2 | — | — | — | V · Anthropic, OpenAI |
| Grad-level Q&A · GPQA Diamond | 93.7 | 93.7 | 96.0 | 94.6 | 95.3 | — | V · OpenAI |
| Images & documents · MMMU Pro | 90.6 | 89.9 | — | — | — | — | I · Vals AI |
| API price, $ / 1M tokens (in / out) | 10 / 50 | — | 10 / 50 | 4 / 20 | 0.75 / 3.75 | 2 / 6 | V · price lists |
Percentages unless noted; Elo and indices in points. For Terminal-Bench, Fable 5.1 uses Anthropic's figure (55.8); SpaceXAI's max-effort run shows 57.9. For HealthBench, the length-adjusted version is used; SpaceXAI's table shows 62.1 for Fable 5.1.
What it means, task by task
Coding depends on what you call coding. If the model works alone in a terminal, it's a tie between Astra and Fable. If you sit next to it in the editor, Fable leads on real Cursor sessions. If you need a lot of code cheaply, Gemini 3.8 Flash is within four points of the leaders on DeepSWE at roughly 1/13 of the token price.
Medicine: ChatGPT for clinical reasoning, Claude for documentation. Astra tops HealthBench Professional (an OpenAI-built benchmark — worth keeping in mind). On medical coding and scribing, independent Vals AI puts Opus 5 and Fable 5.1 first. None of this replaces a doctor: even the leader meets physician criteria in fewer than two thirds of hard cases.
Law and electrical engineering: the only rows Grok wins. 19.6% on Harvey's Legal Agent Benchmark vs 6.7% for Fable; 64.0% on EEBench vs 56.4%. Both numbers are SpaceXAI's own. On independent legal research, Opus 5 leads.
Office work: Claude for documents, ChatGPT for clicking around. Fable leads GDPval-AA and AA Briefcase; Astra leads OSWorld 2.0 and Vals CUA-bench.
Math: Astra by a mile — 97.6% on FrontierMath Tier 4, ten points ahead. But on Humanity's Last Exam with tools, Fable (65.0%) beats Astra (57.2%).
Task → subscription cheat sheet
| Task | Model | Plan |
|---|---|---|
| Pair-coding in the editor | Fable 5.1 | Claude Pro / Max |
| Autonomous coding agent | Astra or Fable 5.1 | ChatGPT or Claude |
| Lots of code, low budget | Gemini 3.8 Flash | Google AI Pro |
| Clinical questions | GPT-6 Astra | ChatGPT Plus |
| Medical documentation | Opus 5 / Fable 5.1 | Claude Pro |
| Legal research | Opus 5 | Claude Pro |
| Legal agent work, circuits | Grok 4.7 | SuperGrok (once confirmed) |
| Reports, spreadsheets, decks | Fable 5.1 | Claude Pro / Max |
| UI automation | GPT-6 Astra | ChatGPT Plus |
| Math | GPT-6 Astra | ChatGPT Plus / Pro |
What I actually use
Two subscriptions, and this table didn't change them. Claude for code — I build iOS and macOS apps and spend most of the day in the editor with the model, which is exactly where Fable 5.1 leads. ChatGPT Pro for everything else — images for my site, calculations, tasks where the model has to finish the job on its own.
The main takeaway: in 2026, "which AI is the best" is the wrong question. The top three differ less than the same model does across modes. You don't pick the best model — you pick the right one for the task.
Next up: the same task-by-task table for open-weight models you can run locally — where they've already caught up with the $20 subscriptions, and where they're still a year behind.
Originally published at klukyanov.ru.
Shorter weekly write-ups (in Russian) — on Telegram.
Top comments (0)