For about a year now, every model launch has followed the same script. A lab drops a new flagship, it’s a little smarter than the last one, and it costs more. You brace for it. New tier, new price, same shrug. So when Anthropic put out Claude Opus 5 on July 24, I opened the pricing page already wincing.
Then I read it twice, because it was half the price.
Not half the price of some bloated competitor. Half the price of Anthropic’s own current king, Fable 5. And on the Artificial Analysis Intelligence Index, Opus 5 edged out Fable 5 for the number one spot. Smarter and cheaper, in the same launch. That’s not how any of this has gone for a year.
And then, three days later, the biggest open-weights model in human history dropped on Hugging Face, and I sat there doing the math on whether I could run it. Spoiler: I cannot. Nobody with fewer than eight datacenter GPUs can. Welcome to the week.
- The Flagship Got Cheaper (Wait, What?)
- The 2.8 Trillion Parameter Paperweight
- The Actually Useful Release Nobody Tweeted About
- Cheapskate Picks: What I’d Actually Pay For
- The Horror Show
- The Still-Waiting Room
- The Takeaway
The Flagship Got Cheaper (Wait, What?)
Here’s the part I keep re-reading. Claude Opus 5 costs $5 per million input tokens and $25 per million output. That’s the exact same price as Opus 4.8, the model it replaces. Meanwhile Fable 5, Anthropic’s big expensive brain, sits at $10 in and $50 out. So Opus 5 is literally half the sticker of the model it just beat on the intelligence leaderboard.
On the Artificial Analysis Intelligence Index, Opus 5 (max) scores 61. Fable 5 lands at 60. GPT-5.6 Sol comes in at 59, and Moonshot’s Kimi K3 at 57. It’s a photo finish at the top, but the point is that the cheaper Anthropic model is the one holding the trophy this week.
The number that actually matters, though, isn’t the Index rank. It’s cost per task. Sticker price per token lies to you, because different models burn different amounts of tokens to finish the same job. Artificial Analysis measured it: running their full Intelligence Index costs about $2.03 per task on Opus 5 versus $2.75 on Fable 5. That’s roughly 26 percent cheaper to do the same work, on top of the lower per-token rate. For anyone running agent loops where the bill compounds, that’s the real headline.
The benchmarks back up the “it’s genuinely good” claim, not just the “it’s cheap” one. On Frontier-Bench v0.1 it hit 43.3 percent, against Fable 5’s 33.7 and GPT-5.6 Sol’s 34.4. On ARC-AGI-3 it scored 30.2 percent while Opus 4.8 managed a sad 1.5. It’s the new default on Claude Max and the strongest model available on Claude Pro.
So what’s the catch, because there’s always a catch. The catch is buried in the launch chart. On the Frontier-Bench numbers, Anthropic notes that Opus 4.8 “stood in as a fallback” whenever a safety classifier refused an Opus 5 request. Fine. Except they never said how often that happened. So the flagship benchmark quietly folds in a weaker model’s answers by an amount nobody will tell you. If you remember the refuse-and-reroute mess that followed Fable 5 around, this is the same species of problem wearing a nicer suit.
The 2.8 Trillion Parameter Paperweight
While Anthropic was cutting prices, Moonshot was flexing. On July 27 they released the open weights for Kimi K3, and this thing is a monster: 2.8 trillion parameters, the largest open-weight model ever shipped. On the Artificial Analysis Index it scores 57, which makes it the highest-scoring open model on the board, ahead of everything else you can actually download.
Here’s where the dream meets the parking lot. The download is about 1.4 terabytes of weights even at MXFP4 quantization. The architecture is a mixture of experts, 896 experts total with 16 active per token, so roughly 50 billion parameters are actually doing work on any given pass. To load it you need something like four to eight H100 80GB GPUs. Your 4090 can’t touch it. Your maxed-out Mac Studio can’t touch it. “Own your weights” is a beautiful slogan right up until you price the hardware to hold them.
So in practice, for almost everyone, Kimi K3 is still just an API you rent at $3 per million in and $15 per million out. The weights being open is great for labs, cloud providers, and the three guys on Reddit with a GPU rack in the garage. For the rest of us it’s a philosophical victory, not a practical one.
And it’s not a clean win even on quality. Accuracy went up about 13 points over the K2.6 generation, which sounds great, but the hallucination rate also climbed about 12 points. Moonshot frames that as the model being “more willing to answer.” Cute. For regulated, legal, medical, or financial work, “more willing to answer” is a polite way of saying “more confidently wrong more often,” and there’s no dial to trade it back.
The Actually Useful Release Nobody Tweeted About
Buried under the Opus 5 launch and the Kimi K3 spectacle, Google shipped the model I’d actually reach for on a Tuesday. Gemini 3.6 Flash landed July 21 at $1.50 in and $7.50 out. That output price is down from the $9 that Gemini 3.5 Flash charged, and Google says it uses about 17 percent fewer output tokens on top of that. Cheaper rate, fewer tokens, same 1 million token context, knowledge cutoff pushed to March 2026.
The funny part: Artificial Analysis gives 3.6 Flash the same Intelligence Index score as 3.5 Flash, a 50. So the tech press mostly shrugged. No leap, no headline. But it gained on the benchmarks that matter for real work, SWE-Bench Pro up to 58.7 percent from 55.1, OSWorld computer use up to 83 from 78.4. And it sits inside the competitive band of five of the six Arena categories I track. Cheaper than the thing it replaces, and it beats the mid-tier of the pricier labs. That’s the whole pitch, and it’s a good one.
Cheapskate Picks: What I’d Actually Pay For
Here’s the trick I run every week, because I’m cheap and I have to be. The Arena leaderboards cluster tight at the top. The entire Overall top 20 this week fits inside 32 rating points, from Fable 5 at 1508 down to a pack at 1476. When the whole visible field is that compressed, paying the leader’s price buys you almost nothing over something a fraction of the cost. So the game is: find the cheapest model still inside spitting distance of the category leader.
Here’s where that landed this week:
| Category | Leader | $ out | Cheapskate pick | $ out | Δ rating | Cheaper by |
|---|---|---|---|---|---|---|
| Overall | Fable 5 (1508) | $50 | MiMo v2.5 Pro (1465, #37) | $0.87 | −43 | ~57x |
| Coding | Opus 4.7-thinking (1553) | ~$25 | MiMo v2.5 Pro (1519, #25) | $0.87 | −34 | ~29x |
| Creative Writing | Fable 5 (1507) | $50 | Gemini 3-Flash (1458, #21) | $3 | −49 | ~16.7x |
| Instruction Following | Fable 5 (1515) | $50 | MiMo v2.5 Pro (1469, #23) | $0.87 | −46 | ~57x |
| Hard Prompts | Fable 5 (1535) | $50 | MiMo v2.5 Pro (1494, #25) | $0.87 | −41 | ~57x |
| Math | Fable 5 (1539) | $50 | Qwen3.7 Max (1490, #14) | $4.42 | −49 | ~11.3x |
MiMo v2.5 Pro, Xiaomi’s open-weights model, is the actual story this week, not Gemini 3.6 Flash. It sweeps four of six categories, Overall, Coding, Instruction Following, and Hard Prompts, all at the same $0.87 output price, and it does it by sitting at rank 23 to 37 in every one of those boards. A top-20 read misses it every single time, which is exactly the mistake I made drafting this section the first time around: I caught it on Coding, went back and checked the rest, and found the same miss in three more categories. Nobody’s talking about it. It’s just sitting there, MIT-licensed, being the actual answer.
Gemini 3.6 Flash, the model with a whole section above this one, doesn’t win a single category outright once you look past the top 20. It’s still inside the band in five of six categories, so it’s a perfectly good pick if you want something you don’t have to go hunting for or self-host, but the honest cheapest option in four of six categories is a phone company’s open-weights model most readers have never heard of.
Creative Writing breaks the pattern in both directions. MiMo doesn’t even show up in this category’s band, because Arena’s Creative Writing top end closes off faster (around rank 22 here) than the other categories do. And the actual cheapest thing inside that narrower band isn’t Gemini 3.6 Flash either, it’s the older, cheaper Gemini 3-Flash at $3, forty-nine points back and sitting right at the edge of the cutoff.
Math is the one category where nothing changed. Qwen3.7 Max is still the cheapest model in the band, and it’s still clinging to the very edge at 49 points back. Grok 4.5 and Gemini 3.6 Flash both sit closer to the leader for a bit more money if you’d rather not ride the edge.
One caveat on Overall and Hard Prompts: I pulled 40 rows deep for each category and the band technically hadn’t closed yet at row 40 in those two (the cutoff-adjacent rows were still inside the window). MiMo’s $0.87 is close to the price floor this cycle, so it’s very unlikely anything cheaper is sitting a few rows further down, but “very unlikely” isn’t the same as “confirmed,” so treat those two picks as high-confidence rather than fully closed.
If you’re keeping score: MiMo v2.5 Pro is the boring, correct answer this week, not Gemini 3.6 Flash. Four out of six categories, one price tag, and it took reading past rank 20 in every single one of them to find that out.
The Horror Show
Every roundup needs a section where I tell you what broke. This week was generous.
The big one was DeepSeek’s migration cliff. DeepSeek is the single most-used vendor on OpenRouter right now, about 17.6 percent of all routed tokens. On July 24 at 15:59 UTC they hard-retired the old model names deepseek-chat and deepseek-reasoner. Not deprecated with a grace period. Retired. Calls to those names now return errors with no fallback, so any service that didn’t repoint to deepseek-v4-flash or deepseek-v4-pro started throwing user-facing failures the moment the clock hit. Worse, reasoning moved from being a model name to being a request parameter, so lazy integrations silently lost their thinking mode before the hard cutoff even arrived. If your app went weird last Friday afternoon, there’s your answer.
Then there’s the Opus 5 fallback I already whined about. A benchmark that quietly swaps in a different model when the safety filter trips, by an amount nobody discloses, is exactly the kind of asterisk that gets left out of the headline.
And Kimi K3 has a confidence problem. The hallucination rate climbing 12 points while the marketing calls it “more willing to answer” is going to bite somebody who wired it into a pipeline that trusts its output. Retrieval checks and citation verification are not optional with this one.
The Still-Waiting Room
The upcoming section is short and it’s mostly one name. Gemini 3.5 Pro missed its target again, and I’ve lost count, but this is at least the fourth slip. It’s still stuck in Vertex AI enterprise preview. Instead of the Pro model everyone actually wanted, Google shipped a fistful of Flash variants on July 21, which is how we got 3.6 Flash. TechCrunch’s headline said it plainly: three new Gemini models, but no 3.5 Pro. Reporting says the model keeps failing to hit Google’s own internal performance bar. At some point “delayed” starts to read as “in trouble.”
On the open side, expect the Kimi K3 community quants and finetunes to start rolling now that the weights are public, assuming you have the hardware to do anything with them.
The Takeaway
For most of the past year, the story in these roundups was rent versus own. Big closed models you pay for by the token, versus open Chinese weights you could theoretically host yourself. This week broke that frame in both directions at once. Anthropic made the closed flagship cheaper than its predecessor, and Moonshot made the open flagship so enormous that owning it is meaningless unless you run a datacenter.
So the real question quietly changed. It’s not rent versus own anymore. It’s how much intelligence can you actually afford to run. Opus 5 answers it one way, by dropping the price of the top shelf. Gemini 3.6 Flash answers it another, by being 85 percent cheaper and good enough. Kimi K3 answers it by being technically free and practically out of reach.
For the record, the market keeps drifting east while all this happens. Chinese models hit a record 58 percent of tokens processed by US firms on OpenRouter this month, peaking around 63 percent earlier in July. DeepSeek alone is that 17.6 percent, Qwen another 13.9, and Anthropic is the last US lab standing in the top 10. Make of that what you will.
Me, I’m going to keep running Gemini 3.6 Flash for the boring stuff and paying up for Opus 5 when the task actually needs a brain. And I’m going to keep not running Kimi K3, because I do not, in fact, own a rack of H100s. Maybe next week.

Top comments (0)