State of ai devto body · TXT
TL;DR: Anthropic (Claude Opus 5.5, Sonnet 5.5, Fable 5.1) and OpenAI (GPT-6 Astra, GPT-6.1 Sol) currently lead; open-weights models such as Kimi K3 and DeepSeek V4 trail by three to eight months. What decides your productivity is no longer the score, but the price per task and the real usage limits of your subscription.
Snapshot as of 30 September 2026. The original of this article is a living post: models, benchmarks, prices and limits are re-verified every week. For the latest numbers and the dated changelog, check the up-to-date version on next-levels.de.
The benchmark table you just read in some newsletter is already out of date. Over the last four weeks, the top spots have been reshuffled five times. And yet the question that’s really burning in the developer forums is a different one. A ChatGPT Pro user writes: “Weekly limit reset yesterday, only 20% left today.” The model question has become a side issue. Your productivity is determined by the cost per task and your weekly limit; that last percentage point on the leaderboard is just noise.
So this post answers four questions: which models are leading the pack, which benchmarks can still measure anything, what the $20, $100 and $200 subscriptions really include, and what developers experience with Claude Max, Codex and Kimi in their day-to-day work.
At a glance
- Top performers: Claude Opus 5.5 and Sonnet 5.5 lead the independent coding and agent indices, while GPT-6 Astra leads in knowledge and interactive reasoning.
- SWE-bench Verified, GPQA and AIME are saturated. Terminal-Bench 4.0, ARC-AGI-3, GDPval-AA and the Artificial Analysis Index with private test sets are meaningful.
- The mid-range (US$2 to US$4 per million input tokens) has seen a collapse in price: GPT-6.1 Sol and Sonnet 5.5 cost 2/10, Opus 5.5 4/20. The top tier remains at 10/50.
- The subscription tiers differ in volume, not in model access; even entry tiers get the top models: Claude effectively reduced its weekly limits by 17% in September, while OpenAI paused Pro 20× and reopened it with half the quota.
- Recommendation: Opus 5.5 at medium reasoning effort as the default, top-tier models only for review and security, Kimi K3 as a cheap second track for code without personal data.
AI models compared: who’s currently in the lead
The real news this summer comes from the open-weights camp. Since July, four Chinese labs have released models with between one and three trillion parameters and one million tokens of context under open licences: Kimi K3 from Moonshot (2.8 trillion parameters, 104 billion active, weights available since 27 July), Alibaba’s Qwen3.8-Max, DeepSeek V4-Pro under the MIT licence and, most recently, Xiaomi’s MiMo-V2.6-Pro, which, with an index score of 46, is the strongest open model in the Artificial Analysis ranking.
How big is the gap? The US institute CAISI estimates it for DeepSeek V4 Pro at around eight months behind the US frontrunners; for GLM-5.3, it measures on cyber benchmarks around four months. Nathan Lambert from Interconnects sees the gap narrowed by K3 from six to nine months to three to five months. Three to eight months: That’s the head start you are buying today when you pay for a closed frontier model.
And the leaders? A duopoly. In June, Anthropic moved the “Fable/Mythos 5” class above Opus; Fable 5.1 and Mythos 5.1 were launched on 1 September; Mythos remains reserved for security-cleared programmes. They were followed by Claude Opus 5.5 at 4/20 US dollars per million tokens, which, according to Anthropic, is “on a par with Fable 5.1 for most tasks”, and Sonnet 5.5 at the unchanged price of 2/10.
OpenAI has simultaneously replaced its entire range. GPT-6 Astra is the company’s largest training run (over 100,000 GPUs), with 1.05 million tokens of context and a price of 10/50 US dollars per million. Below it sit GPT-6 Sol and Luna, at half the price of the 5.6 generation, and on 29 September, GPT-6.1 Sol, which, according to OpenAI, delivers Astra-level performance on DeepSWE at a fifth of the cost.
Google has spent the year without a new Pro model. Gemini 3.5 Pro was announced at I/O and never launched; instead, three Flash models were released in six weeks, most recently Gemini 3.8 Flash on 2 September. According to Google, Gemini 4 is set to be released well before the end of the year. xAI releases new models monthly (Grok 4.7 on 21 September, US$2/6). Meta has effectively shelved the Llama line: since April, its only model has been the proprietary Muse Spark, currently at version 1.3.
| Model | Provider | Release | Context | API price (1M tokens in/out) | Open weights |
|---|---|---|---|---|---|
| Claude Fable 5.1 | Anthropic | 1 September 2026 | 1M | 10 $ / 50 $ | no |
| Claude Opus 5.5 | Anthropic | 22 September 2026 | 1M | 4 $ / 20 $ | no |
| Claude Sonnet 5.5 | Anthropic | 28 September 2026 | 1M | 2 $ / 10 $ | no |
| GPT-6 Astra | OpenAI | 4 September 2026 | 1.05M | 10 $ / 50 $ | no |
| GPT-6.1 Sol | OpenAI | 29 September 2026 | 1.05M | 2 $ / 10 $ | no |
| GPT-6 Luna | OpenAI | 22 September 2026 | 1.05M | $0.10 / $0.50 | no |
| Gemini 3.8 Flash | 2 September 2026 | N/A | $0.75 / $3.75 (introductory price until 31 December) | no | |
| Grok 4.7 | xAI | 21 September 2026 | 500K | $2 / $6 | no |
| Kimi K3 | Moonshot | 16 July 2026 (weights 27 July) | 1M | 3 $ / 15 $ | yes (revenue-based licence) |
| DeepSeek V4-Pro (0813) | DeepSeek | 13 August 2026 | 1M | N/A | yes (MIT) |
| Qwen3.8-Max | Alibaba | 3 August 2026 | 1M | N/A | yes (own licence; Apache 2.0 only for Qwen3.8-27B) |
| MiMo-V2.6-Pro | Xiaomi | 21 September 2026 | 1M | $0.43 / $0.87 | yes (MIT) |
Vendor list prices as of 30 September 2026; batch and cache discounts come on top. And that is precisely where the most important figure of the month lies: Anthropic has reduced the cache read prices by 75% for Fable 5.1 (to $0.25 per million) and by 60% for Opus 5.5 (to $0.20). In an agent run that re-reads the same 200,000-token context fifty times, that amounts to ten million cache tokens: with Opus 5.5, the cost is two US dollars instead of five. The list price for fresh tokens is almost irrelevant for such workloads.
Benchmarks: what still has room for growth and what is saturated
If you pick models from tables, you have a problem: the benchmarks you know are obsolete. OpenAI removed SWE-bench Verified from its own reporting in February because 59% of the tested tasks were flawed and top models were able to reproduce the human reference solution verbatim. GPQA Diamond was removed from the index by Artificial Analysis in September because all top models “simply solve” it. On AIME 2026, the top 3 sit between 97% and 99%. On FrontierMath, Epoch AI had to re-release the benchmark in June after 42% of the tasks contained errors.
What counts instead are interactive, economically grounded tests that are hard to contaminate. Terminal-Bench 4.0 comprises 66 curated tasks; eight were removed (saturated, publicly solved or with quality issues), 19 were revised. ARC-AGI-3 is interactive; the model must learn the rules of an environment through trial and error, so there is no solution that could be found in the training dataset. GDPval-AA measures real-world professional tasks from 44 occupations in blind pair comparisons. And Artificial Analysis weights private, never-before-published test sets at 45% in the Index v4.3, compared with 40% in v4.2 and 20% previously.
| Benchmark | What it measures | Top 3 (as at 30 September) | Status |
|---|---|---|---|
| AA Intelligence Index v4.3 | 10 evaluations, 45% private sets, independent | Opus 5.5 (58), Opus 5.5 xhigh / Sonnet 5.5 (56), Fable 5.1 / GPT-6 Astra (53) | current |
| Terminal Bench 4.0 | agent-based terminal tasks, independent | Sonnet 5.5 (63.6%), Opus 5.5 max (59.6%), Opus 5.5 xhigh (59.6%) | current |
| Arena WebDev | Web apps in a blind comparison, 795,000 votes | Opus 5.5 (1,827), GPT-6 Astra (1,792), Fable 5.1 (1,751) | current |
| GDPval-AA v2.1 | 220 professional tasks, 44 professions, Elo, independent | Opus 5.5 (1846), Sonnet 5.5 (1844), Opus 5.5 xhigh (1820) | current |
| Humanity’s Last Exam | Expert knowledge, no tools, independent (Scale) | GPT-6 Astra (54.2%), Gemini 3.1 Pro (47.3%), Fable 5.1 (46.8%) | current |
| ARC-AGI-3 | Interactive reasoning, independent (ARC Prize) | GPT-6 Astra 62.7% on the standard harness (99.9% on the provider’s harness), Opus 5 30.2%; Opus 5.5 not yet measured | current, harness dispute |
| SWE-bench Verified | Coding tickets | Top score 96%; ranks 11 to 15 within half a point at 80% | Saturated; discontinued by OpenAI |
| GPQA Diamond | Natural sciences | Provider figures: Astra 96.0%, GPT-5.6 Sol 94.6% | saturated, removed from the AA Index |
| AIME 2026 | Math olympiad | Top 3 within 2.1 points at 97–99% | saturated |
“xhigh”/“max” denote the reasoning effort used in the measurement. “Harness” refers to the test environment: which tools the model is given, how many attempts, and what compute budget.
The figures from providers’ blogs are almost universally higher than the independent measurements. Anthropic reports 70.6% for Sonnet 5.5 on Terminal Bench 4.0, while Artificial Analysis measures 63.6%. OpenAI reports 99.9% for Astra on ARC-AGI-3, while the standard ARC Prize harness yields 62.7%, with computing costs of around 26,000 US dollars for the test run. Neither figure is a lie; they are different harnesses and reasoning levels. They are just not comparable. In short: trust the figure measured by someone who isn’t selling the model.
What has shifted in recent weeks: Since the release of Opus and Sonnet 5.5, Anthropic has been leading the way in agentic coding tests and professional task indices, while OpenAI leads in knowledge, maths and ARC. Both providers are now promoting the same metric. Anthropic calculates that Opus 5.5 at medium reasoning effort beats Astra at maximum effort, at roughly a fifth of the cost per task. OpenAI countered on 29 September with GPT-6.1 Sol and exactly the same formula.
AI subscriptions compared: what you get for $20, $100 and $200
Comparing AI subscriptions is simpler than it looks. All providers have settled on three tiers (around 20, 100 and 200 US dollars), and all control the price via limits. Euro prices are for Germany and include VAT where the provider states them:
| Provider | Entry level | Mid-range | Top tier | What distinguishes the tiers |
|---|---|---|---|---|
| Claude | Pro €21.42 (Opus, Sonnet, Claude Code) | Max 5× €107.10 | Max 20× €214.20 | 5-hour window plus weekly limit; Fable 5.1 available only on Max and Team Premium, up to 50% of the weekly limit |
| ChatGPT | Plus €23 (15 to 150 Sol messages per 5 hours) | Pro 5× €103 | Pro 20× €229 (new sign-ups since 29 September with halved quota) | Weekly quotas at all tiers; Astra Ultrafast only on Pro for $500 |
| Gemini | AI Pro $19.99 (3.1 Pro, Antigravity entry-level) | AI Ultra $99.99 | AI Ultra $199.99 | Gemini CLI disabled for subscriptions since 18 June; replaced by the closed Antigravity CLI |
| Kimi Code | Andante ¥49 (K2.7 Code) | Moderato ¥99 / Allegretto ¥199 (K3, 1M) | Allegro ¥699 | Monthly pool plus weekly allowance plus 5-hour rate; processed by Moonshot AI Pte. Ltd. in Singapore |
| Cursor | Pro $20 ($20 credit) | Pro+ $60 ($70 credit) | Ultra $200 ($400 credit) | Credits instead of requests since 2025; Auto mode draws on credits too |
| GitHub Copilot | Pro $10 ($10 credit) | Pro+ $39 | Business $19 / Enterprise $39 per user | From 1 June 2026: usage-based pricing according to API rates |
Euro prices for Claude according to ssdnodes, for ChatGPT according to heise; Google and Kimi do not quote euro prices for Germany. For teams, the number of user seats is relevant: Claude Team Premium with Claude Code costs 100 US dollars per seat on an annual subscription, ChatGPT Business Premium also 100; with a VAT number, Claude Max 20× comes to 180 euros net.
The difference between the tiers is not a model upgrade. With Claude, even the Pro tier at 21 euros gets Opus 5.5 and Claude Code. With OpenAI, the Plus tier gets GPT-6 Astra, but only between five and 45 Astra messages every five hours. You pay by volume. So the more interesting question is: how fast will I hit the limit?
In practice: what developers really experience with Claude Max, Codex and Kimi
The pricing page says “5× Pro”. In recent weeks, both big providers have revised downward what that actually means.
On 14 September, Anthropic replaced the “+50% weekly limit” promotion, launched in May, with a permanent “+25% over the old baseline”. That sounds like a bonus, but compared with the last four months it is a 17% cut, which Anthropic, after criticism on Hacker News, confirmed. The 5-hour windows, which were doubled in May, remain in place.
At the same time, a class action lawsuit has been running in the US since June, alleging that Max 20× actually delivers six to eight times the usage of Pro, not twenty times. A Max 20× user who has been logging his usage for months reports 20% of his weekly quota used on the first day after the reset. His verdict: worth it for $200, “but I wouldn’t bet on it a year from now”.
An effect that hardly anyone takes into account: Fable 5.1 is capped at 50% of the weekly limit on Max and costs 2.5 times as much per token as Opus 5.5. According to a widely shared estimate by Theo Browne, making Opus 5.5 your default gives you roughly four times the effective quota.
OpenAI tackled the problem from the other side. On 10 September, new registrations for Pro 20× were paused because, according to Codex head Tibo Sottiaux, this segment places the greatest load on the systems; on 29 September, sign-ups reopened with halved quota; existing customers are spared until 29 October. The quote from the intro is from this phase: “Weekly limit reset yesterday, only 20% left today”, writes a Pro-5× user on the OpenAI forum.
Common practice is to split the work: Codex for the bulk of small changes and cloud parallelisation, Claude Code for long, continuous sessions. In an analysis of 500 Reddit comments, 65% preferred Codex, mainly because of the limits, while Claude Code won eight out of twelve blind tests, which is not a contradiction.
Kimi is the third track, and it deserves to be taken seriously. Moonshot’s Kimi K2.7 Code runs as a drop-in inside Claude Code (point ANTHROPIC_BASE_URL at https://api.moonshot.ai/anthropic, done) and, at 0.95/4 US dollars per million tokens, costs around a quarter to a fifth of Opus 5.5.
A sample calculation against the then-current Opus 4.8 came to 0.69 US dollars instead of 4.00 US dollars for a 50-turn agent run. A practitioner who ran K2.7 in Claude Code for about two weeks describes it as disciplined in following instructions, with no scope creep, but weaker at extended thinking and with less community know-how about the best prompts.
The catch lies elsewhere. The international Kimi platform is operated by Moonshot AI Pte. Ltd. in Singapore, where the data is processed; a GDPR-compliant data processing agreement (DPA) is not publicly available. Singapore is a third country without an adequacy decision from the EU, and the parent company is based in Beijing. For a European company with customer data in the repo, that is a deal-breaker as long as there is no EU entity offering a DPA.
For code without personal data, it is a cheap second track. Ramp reports that 6.1% of US companies paying for AI are now using platforms that provide open-source or Chinese models.
Two numbers show how far enterprise reality is from the forums. According to Menlo Ventures, around 54% of enterprise spending on AI coding in 2025 went to Anthropic and 21% to OpenAI; in the Ramp AI Index from August, 43.5% of US companies pay Anthropic and around 40% pay OpenAI.
Usage numbers tell a different story than budgets. The Stack Overflow survey 2026, with over 49,000 responses, shows that 84% use AI tools, 29% trust their accuracy (3% “strongly”), and 66% say the code is “almost right, but not quite”. Among the tools, Copilot leads with 68% of AI users, Cursor is at 18%, and Claude Code at just under 10%.
What this means for your setup
I think it’s a mistake to rely on a single model for a team. The reason is economic: the cost per task depends more on the reasoning level than on the model. Opus 5.5 scores an identical 59.6% on Terminal Bench 4.0 at “max” and “xhigh”; the higher level just burns more tokens. My recommendation as of today:
Default: a mid-range model at medium reasoning effort. Today, that means Opus 5.5 medium or GPT-6.1 Sol; both deliver top results at a fifth of the cost. Review: frontier models, i.e. each provider’s most expensive tier (Fable 5.1, Astra), are reserved for cases where a mistake gets expensive: security reviews, architecture decisions, the last pair of eyes before the merge.
Bulk: for high-volume work with clear instructions (generating tests, migration scripts, documentation), a third, low-cost track pays off, provided data protection and DPAs are sorted. That could be Luna, Gemini Flash or an open-weights model on European infrastructure.
As I described in the post “Agent = Harness + Model”, the harness (context, memory, tools and evals) determines whether an agent is productive. In a properly built setup, swapping the model is a single config line; ANTHROPIC_BASE_URL above is an example.
If switching from Opus 5 to 5.5 takes days, the dependency is in the wrong place. And anyone sending customer data through a model outside the EU because it’s five times cheaper should read my post on the European AI stack first.
When deciding on a subscription: a team of five developers is better off with Claude Team Premium (US$100 per seat, including Claude Code) or ChatGPT Business Premium than with five individual Max subscriptions. SSO and audit logs are one reason.
The other: on the Team and Business plans, chats are not used for training by default and belong to the company account; with private subscriptions, they are linked to the employee and move with them. If you stick with Cursor or Copilot: both now bill at API rates, and the $20 plan is really a credit balance.
What I’m watching next
Four open threads that will shape the coming weeks.
Anthropic announced Haiku 5.5 on 22 September, for release “in the coming weeks”, but it isn’t here yet. At OpenAI, the grace period for existing Pro-20× customers ends on 29 October; after that, the halved quota will apply to everyone. Gemini 4 is expected to launch well before the end of the year and would be Google’s first new Pro model since February. And the class action lawsuit against Anthropic’s Max limits is pending; a judgement or settlement would directly affect the subscription table above.
Frequently Asked Questions
Which AI model is currently the best? For agentic coding and professional tasks, Claude Opus 5.5 leads the independent indices (Artificial Analysis, Terminal-Bench 4.0, GDPval-AA); for knowledge and interactive reasoning, GPT-6 Astra (Humanity’s Last Exam, ARC-AGI-3). The performance gap is smaller than the price gap.
Is Claude Max 20× still worth it? If you work in Claude Code for several hours a day and set Opus 5.5 as your default instead of Fable, then yes. Expect to pay €214 incl. VAT in Germany and bear in mind that a single large agent run will make a noticeable dent in your weekly limit. For teams, Team Premium at 100 US dollars per seat is usually the better option.
Can I use Kimi K3 for business in the EU? Technically yes, via API or as a drop-in in Claude Code. Legally, it depends on what data you send: processing is carried out by Moonshot AI Pte. Ltd. in Singapore, a third country without an EU adequacy decision, and a data processing agreement is not publicly available. For code without personal data, it’s a cost-effective alternative; for customer data, it isn’t.
Why do vendor blogs show different benchmark numbers than this post? Because vendors measure with their own harness, at the highest reasoning level and sometimes with multiple attempts. Independent evaluators such as Artificial Analysis, ARC Prize or Scale use a fixed harness and budget. This post prefers the independent numbers and labels vendor numbers as such.
Changelog
The original post logs every change to models, prices, limits and top-3 rankings with a date. This dev.to copy is a snapshot; the live changelog has everything since.
- 30 September 2026 — First published. Status: Opus 5.5, Sonnet 5.5, GPT-6 Astra, GPT-6.1 Sol, Gemini 3.8 Flash, Grok 4.7, Kimi K3, MiMo-V2.6-Pro. Claude weekly limit −17% since 14 September; ChatGPT Pro 20× has been open again (halved) since 29 September.
I’m Slawa, CEO of Next Levels, a German agency building Shopware commerce, apps and AI agents for mid-sized companies. This article was originally published on next-levels.de, where it is re-verified every week. If a number here looks stale, the original has the current one.
Which split are you running in your team: one model for everything, or separate tracks for default, review and bulk work? Let me know in the comments.
Top comments (0)