DEV Community

Cover image for Which AI Subscription Should You Pick for a Task: a Benchmark Table (September 2026)
Kirill Lukyanov
Kirill Lukyanov

Posted on Originally published at klukyanov.ru

Which AI Subscription Should You Pick for a Task: a Benchmark Table (September 2026)

Yesterday SpaceXAI shipped Grok 4.7 with a benchmark table in which Grok 4.7 wins. Three weeks ago OpenAI shipped GPT-6 Astra with a table in which Astra wins. Anthropic did the same with Claude Fable 5.1. None of these tables is fake — but every one of them was put together by the team that ships the model.

So before renewing my own subscriptions I did the boring thing: pulled numbers from eight sources into one matrix. Three independent leaderboards — Artificial Analysis, Vals AI and Arena (blind human voting) — plus the launch posts from OpenAI, Anthropic, Google and SpaceXAI, which, put side by side, check each other surprisingly well.

Snapshot date: September 22, 2026. Every number below is labeled as independent or vendor-reported.

The lineup: six models, four subscriptions

The only rule: the model must be available in a regular paid plan an individual can buy.

Model Released Subscription Price / month
Claude Fable 5.1 Sep 1 Claude Pro / Max $20 / $100–200
Claude Opus 5 before Sep Claude Pro / Max $20 / $100–200
GPT-6 Astra Sep 3 ChatGPT Plus / Pro $20 / $200
GPT-5.6 Sol before Sep ChatGPT Plus / Pro $20 / $200
Gemini 3.8 Flash Sep 2 Google AI Pro / Ultra $19.99 / $250
Grok 4.7* Sep 21 SuperGrok $30

* At launch Grok 4.7 is available in Cursor, Grok Build and the API; SpaceXAI hasn't yet confirmed it in the SuperGrok consumer docs. Check before paying.

Left out on purpose: Meta Muse Spark 1.3 (free in the Meta AI app — not a subscription), Claude Mythos 5.1 (gated to verified security and life-science programs), and open-weight models like DeepSeek, Qwen, Kimi and GLM — they get their own article.

Why you shouldn't take a launch table at face value

Not because vendors lie. Because the same benchmark can be run several ways, and every vendor picks the way that flatters them. Four examples from September alone:

  • ARC-AGI-3. OpenAI reports 99.9% for GPT-6 Astra — in a harness that preserves reasoning state between steps. In the neutral harness everyone uses, Astra scores 62.7%.
  • HealthBench Professional. SpaceXAI's table gives Claude Fable 5.1 62.1%. The length-adjusted independent leaderboard gives the same model 56.6% — a gap larger than first-to-third place.
  • OSWorld 2.0. OpenAI lists Claude Opus 5 at 70.2%. Anthropic lists the same model at 39.6% — in "strict" mode. Both are correct; they're just not comparable.
  • CursorBench. SpaceXAI uses v4.0, Anthropic uses v3.2. Fable 5.1 scores 51.8% and 73.4% respectively.

Three rules I now apply: independent rankings beat vendor tables even when they have fewer rows; always check who is missing from a table (Grok 4.7's chart has no GPT-6 Astra); and a number without its mode (max, xhigh, with tools, strict) is not a number.

Independent rankings: a tie at the top

  • Artificial Analysis Intelligence Index: Fable 5.1 and GPT-6 Astra — 53 each, Opus 5 — 51, GPT-5.6 Sol — 47, Grok 4.7 — 46.
  • Vals Index (finance, legal research, code migration, medical documentation): Fable 5.1 68.8%, Opus 5 67.2%, Astra 66.6%, Gemini 3.8 Flash 62.3%, Grok 4.7 60.2%.
  • Arena Text (blind voting, Sep 13): Fable 5.1 1498, Opus 5 and Gemini 3.8 Flash 1493 each — within the confidence interval. Arena Code is a different story: Astra 1800, Fable 1758, Opus 1687.

If you're picking one subscription "for everything", Claude and ChatGPT are effectively equal right now. The decision comes down to the tasks where they diverge — and there are plenty.

The full matrix: 6 models × 19 tests

Leader of each row in bold. "—" means the model wasn't tested or the result isn't published. I = independent, V = vendor-reported.

Task · test Fable 5.1 Opus 5 GPT-6 Astra GPT-5.6 Sol Gemini 3.8 Flash Grok 4.7 Source
Overall · AA Index 53 51 53 47 <46 46 I · Artificial Analysis
Professional tasks · Vals Index 68.8 67.2 66.6 62.3 60.2 I · Vals AI
Blind chat · Arena Text 1498 1493 1493 I · Arena
Code, blind · Arena Code 1758 1687 1800 I · Arena
Terminal agent · Terminal-Bench 4.0 55.8 52.3 57.7 37.3 19.1 38.0 V · OpenAI, Anthropic
Long-horizon SWE · DeepSWE v1.1 70.0 74.0 72.7 73.7 71.0 V · Google, SpaceXAI
In-editor coding · CursorBench 4.0 51.8 41.7 46.3 V · SpaceXAI
Clinical reasoning · HealthBench Pro 56.6 63.4 60.5 56.7 I · aggregated leaderboard
Medical coding · Vals MedCode #7 63.6 #23 I · Vals AI
Legal agent · Harvey LAB 6.7 2.5 19.6 V · SpaceXAI
Legal research · Vals Legal Research 55.3 #21 I · Vals AI
Electrical engineering · EEBench 56.4 39.4 64.0 V · SpaceXAI
Documents · GDPval-AA (Elo) 1853 1824 1711 V · Anthropic
Multi-hour office work · AA Briefcase 1678 1487 1657 V · SpaceXAI
Computer use · OSWorld 2.0 70.2 72.6 65.7 V · OpenAI
Research math · FrontierMath T4 87.8 73.2 97.6 83.0 V · OpenAI
Expert knowledge · HLE with tools 65.0 63.6 57.2 V · Anthropic, OpenAI
Grad-level Q&A · GPQA Diamond 93.7 93.7 96.0 94.6 95.3 V · OpenAI
Images & documents · MMMU Pro 90.6 89.9 I · Vals AI
API price, $ / 1M tokens (in / out) 10 / 50 10 / 50 4 / 20 0.75 / 3.75 2 / 6 V · price lists

Percentages unless noted; Elo and indices in points. For Terminal-Bench, Fable 5.1 uses Anthropic's figure (55.8); SpaceXAI's max-effort run shows 57.9. For HealthBench, the length-adjusted version is used; SpaceXAI's table shows 62.1 for Fable 5.1.

What it means, task by task

Coding depends on what you call coding. If the model works alone in a terminal, it's a tie between Astra and Fable. If you sit next to it in the editor, Fable leads on real Cursor sessions. If you need a lot of code cheaply, Gemini 3.8 Flash is within four points of the leaders on DeepSWE at roughly 1/13 of the token price.

Medicine: ChatGPT for clinical reasoning, Claude for documentation. Astra tops HealthBench Professional (an OpenAI-built benchmark — worth keeping in mind). On medical coding and scribing, independent Vals AI puts Opus 5 and Fable 5.1 first. None of this replaces a doctor: even the leader meets physician criteria in fewer than two thirds of hard cases.

Law and electrical engineering: the only rows Grok wins. 19.6% on Harvey's Legal Agent Benchmark vs 6.7% for Fable; 64.0% on EEBench vs 56.4%. Both numbers are SpaceXAI's own. On independent legal research, Opus 5 leads.

Office work: Claude for documents, ChatGPT for clicking around. Fable leads GDPval-AA and AA Briefcase; Astra leads OSWorld 2.0 and Vals CUA-bench.

Math: Astra by a mile — 97.6% on FrontierMath Tier 4, ten points ahead. But on Humanity's Last Exam with tools, Fable (65.0%) beats Astra (57.2%).

Task → subscription cheat sheet

Task Model Plan
Pair-coding in the editor Fable 5.1 Claude Pro / Max
Autonomous coding agent Astra or Fable 5.1 ChatGPT or Claude
Lots of code, low budget Gemini 3.8 Flash Google AI Pro
Clinical questions GPT-6 Astra ChatGPT Plus
Medical documentation Opus 5 / Fable 5.1 Claude Pro
Legal research Opus 5 Claude Pro
Legal agent work, circuits Grok 4.7 SuperGrok (once confirmed)
Reports, spreadsheets, decks Fable 5.1 Claude Pro / Max
UI automation GPT-6 Astra ChatGPT Plus
Math GPT-6 Astra ChatGPT Plus / Pro

What I actually use

Two subscriptions, and this table didn't change them. Claude for code — I build iOS and macOS apps and spend most of the day in the editor with the model, which is exactly where Fable 5.1 leads. ChatGPT Pro for everything else — images for my site, calculations, tasks where the model has to finish the job on its own.

The main takeaway: in 2026, "which AI is the best" is the wrong question. The top three differ less than the same model does across modes. You don't pick the best model — you pick the right one for the task.

Next up: the same task-by-task table for open-weight models you can run locally — where they've already caught up with the $20 subscriptions, and where they're still a year behind.

Originally published at klukyanov.ru.

Shorter weekly write-ups (in Russian) — on Telegram.

Top comments (0)