Originally published on techpill.de. Data: a16z “LLMflation”, Epoch AI, Stanford AI Index.
The phrase "AI keeps getting cheaper" is a bit of a double-edged sword. While it's true that the cost of achieving a fixed capability has plummeted by about 1,000× since 2021, the price tag for the top-tier model you can buy has only dropped around 25×. So yes, cheap AI is out there, but it’s usually lagging about two years behind the cutting edge. Grasping this distinction is key to understanding the economics driving the AI boom.
The key numbers
- 1,000× cheaper to reach GPT-3-level ability (2021 → 2024)
- ~50× median annual price drop across six benchmarks (ranging from 9× to 900×)
- ~25× cheaper for the leading model — over six years
Two prices you must not confuse
When you hear "is AI getting cheaper?", remember there are two distinct prices at play:
- The price of a capability: This is what it costs to operate a model that consistently meets a certain standard — like the knowledge test that GPT-3 just managed to pass. This price is dropping fast.
- The price of the frontier: This refers to the cost of the best model available right now. This price barely budges.
The reason for this discrepancy is straightforward: the frontier is always advancing. Each new flagship model sets a new standard and tends to cost about the same as its predecessor. Meanwhile, capabilities that were once pricey are now becoming accessible through cheaper models. So, both statements — "AI got 1,000× cheaper" and "the best model stays expensive" — hold true at once. They simply measure two different things.
1,000× cheaper: the GPT-3 case
The clearest example comes from a16z, which even coined a name for the trend: "LLMflation." Back in November 2021, GPT-3 went public as the only model scoring 42 on the MMLU benchmark, and it cost roughly $60 per million tokens. Fast-forward three years, and a tiny open model — Llama 3.2 3B — hit that same score for about $0.06 per million tokens. That’s a 1,000× gap, or put another way, roughly 10× cheaper every single year.
Here’s the part people miss: nobody discounted a single model by 1,000×. What fell was the cost of reaching a capability level at all, as smaller and more efficient models kept clearing the same bar.
What one dollar buys
Let’s make that tangible. In 2021, a dollar spent on GPT-3 got you around 12,500 words — about five pages. Today that same dollar buys roughly 12.5 million words at the same capability level: think a shelf holding some 125 novels. The graphic below draws the ratio at true scale — one dot against a thousand.
The decline is uneven
The drop is real, but it’s far from uniform. Epoch AI tracked pricing across six benchmarks and found a wild spread: anywhere from 9× to 900× per year, with the median landing near 50× (GPT-4-level PhD-science questions fell about 40× annually). The pattern makes sense — simple, well-defined tasks get cheap first, because small models pick them up quickly. The truly hard stuff stays expensive longer, since it still leans on big models.
Faster than Moore's Law
For context: Moore’s Law roughly doubles compute every two years. Fixed-capability AI pricing halves a cheap tier in about a year — and the fastest tiers in a matter of months. As a16z frames it, nothing in computing history has ever deflated this quickly — not compute during the microprocessor era, not bandwidth during the dotcom boom.
Why the frontier stays expensive
Now the flip side. While a fixed capability got 1,000× cheaper in three years, the best-available model has only come down about 25× in six. If you insist on the newest model, you’re chasing a target that keeps sprinting away — effectively paying 2020 prices. Willing to sit two years back? You pay next to nothing.
What it means for you
Whether you use AI directly or build products on top of it, one rule falls out of all this:
- Do you genuinely need the frontier? For the hardest problems, yes — that’s the cost of being out front.
- Or is "excellent, from two years ago" good enough? For most real work — summaries, translation, classification, chat, code suggestions — a model that led the pack a year or two back is more than enough, at a sliver of the price.
- So the question shifts from "can I afford AI?" to "how much lag can my product afford?" In practice: tier your models. Route the easy calls to something cheap, and hand only the genuinely hard cases to the frontier.
The whole picture in one graphic
Sources: a16z, Welcome to LLMflation (2024) · Epoch AI, LLM inference prices have fallen rapidly but unequally across tasks (2025) · Stanford HAI, AI Index 2025.
Full article with the data table and FAQ: techpill.de/how-ai-got-1000x-cheaper



Top comments (0)