Two data points landed within days of each other this week, and together they gut an assumption the entire AI industry has been quietly built on: that customers will always pay up for the smartest available model.
They won't. And the numbers are now public enough that you don't have to take anyone's word for it.
The Fable 5 Problem
The Financial Times got hold of spending data from 70,000 companies via Ramp, the corporate card and expense platform, and it's ugly reading for Anthropic. Fable 5 — Anthropic's largest, most expensive model, launched in early June — has plateaued at roughly 11% of the company's total tool spend, more than two months after release. That breaks a pattern that's held since the ChatGPT era started: businesses defaulting to whatever model sits at the top of the leaderboard, price be damned.
It gets more specific. Anthropic's own Opus 5 — smaller, cheaper, launched in late July — has already overtaken Fable 5 in business spending. A company's newer, non-flagship model is cannibalizing its own flagship, less than two months after the flagship shipped. Miles Clements, a partner at Accel (which has ~$1bn invested in Anthropic), put it bluntly to the FT: "Most people don't need to operate at the frontier... [that] was not a durable era."
Anthropic isn't cratering — revenue is still up almost sevenfold since January, hit $65bn annualized in July (up from $47bn in May), and the company posted its first adjusted operating profit in Q2. But growth undershot the most bullish investor projections ($80bn annualized), right as Anthropic heads into what could be the largest IPO in history, expected to value it at $2 trillion or more. Meanwhile OpenAI's annualized revenue jumped 35% this quarter to over $40bn on the back of the cheaper GPT-5.6, after a sluggish start to the year. Ramp's chief economist Ara Kharazian summed up why nobody should trust their own trendlines right now: "If you impute previous trends you expect Anthropic to own the market. But because [OpenAI's newest model] was so good and Fable underperformed, it's been the reverse."
The Benchmark That Should Worry Every Frontier Lab
The FT story is about enterprise spending behavior. The second data point is about raw capability, and it's arguably scarier for the labs charging premium prices.
An independent benchmark called the Ed-o-meter — 28 real-world tasks across coding, data work, tool use, security, and general reasoning, run identically across 17 models via OpenRouter — added four new models this week. The headline: GLM-5.3, an open-weight Chinese model, is the first model on the board to clear all five task categories at 100%. It posted a 9.3/10 rubric score (third-highest overall) for $0.0101 per task.
Fable 5? 79% pass rate — joint-bottom on the board — at $0.0748 per task, more than seven times the cost for worse results. It also refused five of the 28 tasks outright, tripped up by an overly aggressive safety classifier that also nailed Opus 5 on identical grounds (Opus scored only 43% on coding for the same reason — benign debugging tasks blocked before a single token generated). Anthropic's newest models are, per this benchmark, getting outperformed by a model that costs a fraction as much and being penalized by their own guardrails on top of it.
A related post making this exact case — "GLM-5.3 beat Anthropic/OpenAI models for 1/5 the cost" — hit Hacker News this week and pulled 234 points and 100+ comments before getting flagged (HN's mods regularly flag anything smelling of an ad, deserved or not — worth noting, not dismissing). Whatever you think of the framing, the underlying benchmark numbers are independently reproducible and hold up.
Why This Is Happening Now
Three things converged:
- Open-weight models caught up faster than labs priced for. GLM-5.3, DeepSeek-v4-pro, and others are now benchmarking competitively with frontier proprietary models on real tasks, not just synthetic leaderboards, at a fraction of the inference cost.
- Most business workloads don't need frontier intelligence. Summarization, code review, CRM data mapping, ticket triage — the bulk of enterprise AI spend — was never a task that required the most expensive model available. It just used to be the default because nobody had reason to check.
- Regulatory friction made the expensive option worse, not better. Fable 5's launch was disrupted by a Trump administration national-security intervention that forced a temporary withdrawal, and lingering data-retention rules imposed as a condition of its relaunch have kept dampening adoption even after the political heat receded.
What This Means If You're Buying AI Tools
If you or your team have been defaulting to the "best" model in every workflow — Fable 5, Opus, GPT-5.6-sol, whatever your top-tier default is — this week's data is your prompt to actually measure instead of assume.
Audit by task, not by vendor. The Ed-o-meter's category breakdown is the useful part: a model can be excellent at reasoning-heavy realworld tasks and mediocre at coding, or vice versa. Route work to the model that's actually good at that category, not the one with the highest sticker price.
Budget for a multi-model stack, not a single vendor. Anthropic's own customers are already doing this — dropping down from Fable to Opus mid-contract because the cheaper option handles their actual workload. If your AI spend is concentrated on one frontier model "because it's the best," you're very likely overpaying for capability you don't use on the majority of your requests.
Watch the safety-classifier tax. Both Fable 5 and Opus 5 got dinged on this benchmark not for being incapable, but for provider-side refusal filters blocking benign requests before generating output. If your workflow touches anything that could plausibly look like a "coding-debug" edge case or contains injected text (emails, scraped docs, PDFs), test for false-positive refusals specifically — it's now a documented, cross-model pattern at Anthropic, not a one-off.
The Bigger Signal
None of this means frontier models stop mattering — someone still needs to push the ceiling up, and Anthropic, OpenAI, and Google are the ones doing it. But the assumption that pushing the ceiling automatically translates into revenue is now visibly cracking, right as Anthropic walks into a $2 trillion IPO expecting investors to buy that story. The "biggest model wins" era priced a lot of tools higher than the market was actually willing to pay for most of its daily workload. This week, the spending data and the benchmark data both said so, independently, on the same days.
Top comments (0)