DEV Community

Roman Shumyatsky
Roman Shumyatsky

Posted on Originally published at modelpricewatch.com

Frontier LLM prices didn't move for 5 months. In August, they moved three times, and one lab tripled its rate.

On August 1 I published a report whose headline finding was that frontier LLM API prices are structurally sticky. Across 40 daily readings of an equal-weight index of ten flagship models — one per lab — not one lab had ever changed the price of an existing model. Every move in the index had come from a new model replacing an old one.

August made that sentence false in three weeks.

Here's what moved, why the index still ended the month lower, and what happened 72 hours after the cutoff that dwarfs all of it.

The month in one table

The index is the equal-weight average of ten flagships' blended price per million tokens (3 parts input to 1 part output, list prices as printed on the vendor's own pricing page).

Date What happened Index ($/Mtok)
Aug 1 Opening level $4.39
Aug 4 Alibaba's slot: Qwen3.7-Max → Qwen3.8-Max ($3.75 → $3.00 blended) $4.32
Aug 16 DeepSeek V4 Pro repriced: flat $0.435/$0.87 → peak $1.32/$3.96 (+264% blended) $4.46
Aug 21 GPT-5.6 Sol repriced: $5/$30 → $4/$20, labelled promotional (−29%) $4.14
Sep 1 Closing level $4.14

Net for the month: −5.7%. Since the first reading on February 23: −9.4%.

Three other flagship handovers happened in August (Muse Spark 1.1 → 1.2, Grok 4.5 → 4.6, GLM-5.2 → 5.3) and moved nothing, because each successor kept its predecessor's list price. That's the pattern I described in August. The two bolded rows are the pattern breaking.

Move 1: DeepSeek turned "list price" into a schedule

Until 16:00 UTC on August 16, DeepSeek V4 Pro billed a single flat rate: $0.435 in / $0.87 out. Then the pricing page split it in two:

  • Peak (01:00–04:00 and 06:00–10:00 UTC): $1.32 / $3.96
  • Off-peak (every other hour): exactly half — $0.66 / $1.98

The index tracks the peak rate as the list price. Two reasons. DeepSeek defines off-peak as a discount from peak, not the other way round, so peak is the published number. And a caller who doesn't schedule around the clock needs a ceiling, not a floor. But note that even the off-peak rate ($0.99 blended) is 82% above the old flat price. This wasn't a discount scheme layered on the old price. It was a 3x increase with a discount window attached.

If you run DeepSeek workloads and can batch them: 10:00–01:00 UTC is a long off-peak window. If you can't, your V4 Pro bill roughly tripled in mid-August and the vendor did not send you an email about it.

Move 2: OpenAI's cut, with an asterisk

On August 21 GPT-5.6 Sol went from $5/$30 to $4/$20 — $11.25 → $8.00 blended, −29%. It's the first list-price cut by any constituent in the index's history.

The asterisk is a sentence that appeared on the pricing page the same day: "GPT-5.6 Sol's promotional pricing is available at least through November 21, 2026."

That's a floor, not a reversion date. OpenAI publishes no date on which the price goes back up. The index takes list prices as printed, so the cut is in — but if you've modelled Sol at $4/$20 past November 21, that input is contingent on a sentence, and the sentence can change.

Two of the ten flagships now publish a list price that's really the top of a range: one has a time-of-day schedule under it, the other has a promo clock over it.

Why the cheapest GPT-4-class model didn't get more expensive (yet)

The other line I track is the floor: the cheapest model that clears a fixed capability bar (GPQA Diamond ≥ 70, externally measured — roughly GPT-4-class reasoning). It's the deflation story; it fell 17x between March and July.

In August it did not move. It stayed at $0.113 per million tokens, set by DeepSeek V4 Flash since July 25.

Here's the thing. On August 16 DeepSeek also raised its own V4 Flash peak rate, from $0.14/$0.28 to $0.44/$1.32 — $0.175 → $0.66 blended, 3.8x. The model setting the floor got nearly four times more expensive from its maker, and the floor didn't budge.

That's because the floor reads the cheapest listed price for a base model across every host that serves it, and DeepInfra kept serving V4 Flash at $0.09/$0.18. So the cheapest GPT-4-class token on September 1 was priced not by the lab that trained the model but by a third party hosting it, 5.9x below the lab's own peak rate.

That's a new kind of floor, and a more fragile one. A lab's list price is a policy; a reseller's list price is a margin. On September 1 GPT-4-class capability was 37x cheaper than the flagship ceiling. How much of that gap is one host's pricing decision is a question I didn't have to ask in July.

The spread narrowed from both ends

"Frontier flagship" spanned a 21x price range on August 1 ($0.544 to $11.25 blended). On September 1 it spanned 13x ($0.75 to $10.00) — and it narrowed from both directions at once. Sol's cut brought the top down to Claude Opus 5's $10.00. DeepSeek's move to peak pricing lifted the bottom off V4 Pro and onto Mistral Large 3's $0.75.

One casualty of the month's launches: with GLM-5.3 arriving API-first (open weights announced, not yet released), Kimi K3 is the basket's only open-weights constituent. The "proprietary premium" comparison I ran in August is now nine numbers versus one, so I'm not running it.

Then September 3 happened

This report is dated September 1. I need to tell you about the following 72 hours anyway, because the live index no longer says $4.14.

On September 3 OpenAI released GPT-6 Astra at $10/$50 — $20.00 blended — and named it the default flagship. It took the OpenAI slot in the basket from GPT-5.6 Sol on September 4, and the index went from $4.14 to $5.34 in a day. That's +29.0%, the largest single move in the index's history, and the first time the index has been above its February 23 first reading (+16.8%).

The spread that had narrowed to 13x reopened to 27x. At a published benchmark score of 96, Astra costs $0.208 per point of measured intelligence, 3.5x the September 1 frontier average.

Two days earlier, on September 1, Anthropic had launched Claude Fable 5.1 at the same $10/$50, with cached input at $0.25 — 2.5% of the input rate, the deepest cache discount on any flagship card I track.

Two $10/$50 flagships in 72 hours, at the end of a month in which the ceiling fell 5.7%. I'm not going to call a turn on one reading. But "the frontier price never moves" was August's finding, and September's first four days have already had a reprice up, a reprice down, and the biggest step up on record.

What I'd do with this if I were paying the bills

  • Write the tier next to the price, not just the date. Peak or off-peak. Standard or promotional. Short-context or long-context. Two of ten flagships now have a list price that means "up to."
  • Check the UTC clock on DeepSeek workloads. Off-peak is 17 hours a day and exactly half price. If your pipeline runs at 08:00 UTC, you're paying double for no reason.
  • Treat Sol's $4/$20 as expiring November 21 until OpenAI says otherwise. Model the reversion; be pleasantly surprised.
  • Know who's actually pricing your cheapest model. If it's a host and not the lab, check that the host still lists it at that price before you renew the budget. The lab's own rate for the same model may be several times higher.

Every figure above is as of September 1, 2026, and is fixed in the dated report at modelpricewatch.com/reports/state-of-llm-pricing-2026-09. The live index is at modelpricewatch.com/price-index, with the methodology and the full price history dataset on Hugging Face (CC-BY-4.0). Every price links to the vendor page it came from, with a capture timestamp.

Disclosure: I maintain ModelPriceWatch.

If you run DeepSeek in production: did you notice the August 16 change, and when? I'm trying to work out how many teams found out from the invoice.

Top comments (0)