DEV Community

Cover image for The 90% Price Cut That Changes How You Build AI Products
Yano.AI Technologies Inc.
Yano.AI Technologies Inc.

Posted on Originally published at yanoai.tech

The 90% Price Cut That Changes How You Build AI Products

In January 2025, a model scoring 75% on GPQA Diamond cost roughly 30 cents per question. Eighteen months later a model matched that score for four hundredths of a cent (Source: Epoch AI, 2026). That is a 725-fold drop in the price of machine reasoning, inside the window most product teams spent arguing about which model to pick.

Infographic

The cost curve is steeper than the capability curve

The economists at Epoch AI tracked five hard benchmarks across three years of model releases and found the cost of a fixed level of accuracy fell about 47% per quarter, or roughly 13 times per year (Source: Epoch AI, 2026). For comparison, that rate is about four times faster than DNA sequencing, six times faster than compute, and 18 times faster than lithium batteries.

The decline is not uniform across a model's life, and that has an architectural consequence. A performance level drops 66% per quarter right after it first debuts as state of the art, then slows to 32% per quarter two years later (Source: Epoch AI, 2026). Anything built on the newest capability is a depreciating asset. Anything built on a solved problem gets cheaper every quarter.

Two vendors cut small model prices in the same week

Anthropic released Claude Haiku 5.5 on October 7, 2026, priced 90% lower than Haiku 4.5 for prompts under 100,000 tokens (Source: Anthropic, 2026). Input dropped from $1.00 to $0.10 per million tokens and output from $5.00 to $0.50 (Source: Anthropic, 2026).

OpenAI moved in the same direction weeks earlier, halving API prices for GPT-6 Sol and Luna against their GPT-5.6 promotional rates. Luna landed at $0.10 input and $0.50 output per million tokens, matching Haiku 5.5 to the cent (Source: OpenAI, 2026).

Two labs arriving at the same price point on the same class of work is not coincidence, it is the market saying the small tier is now the default tier for anything repetitive.

Capability moved too, which is why routing matters

A cheap model that cannot do the job is not a cost saving, it is a rework loop. The benchmark gap that used to make small models unusable on agentic work has largely closed. On Terminal-Bench 4.0, Haiku 5.5 scores 39.2% where Haiku 4.5 scored 0.0% (Source: Anthropic, 2026).

Haiku 5.5 is also the first model in Anthropic's small class with an adjustable effort setting, letting callers trade intelligence for cost per call (Source: Anthropic, 2026). That setting matters more than the headline price. On AutomationBench, GPT-6 Sol at xhigh effort scored 33.2% at $0.27 per task, while Claude Opus 5 at max effort scored 26.9% at 11.1 times the cost (Source: OpenAI, 2026).

What production teams measured

Vendor benchmarks describe possibility. Customer deployments describe economics, and the gap is usually where your own architecture lives.

Asana reported over 30% lower latency on task completions after moving short triage work to Haiku 5.5 (Source: Anthropic, 2026). AlphaSense runs roughly 8 million such calls per week and measured an accuracy improvement from 0.76 to 0.84 on its document question workload (Source: Anthropic, 2026).

Box scored 11 points higher than Haiku 4.5 at about half the latency and moved the model into analytical work at scale (Source: Anthropic, 2026). Cognition paired Haiku 5.5 as a sidekick subagent under Opus 5.5 and reported a FrontierCode score of 66.2 while cutting cost and latency (Source: Anthropic, 2026).

The pattern across all four is identical. Nobody replaced the expensive model. They gave it less work.

Why this matters most outside Silicon Valley

The Philippines has a specific version of this problem. Only 14.9% of firms use AI technologies, and the country's AI market is projected to grow from about $772 million in 2024 to $3.49 billion by 2030 (Source: U.S. International Trade Administration, 2026). In the IT-BPM sector, which ended 2025 with roughly $40 billion in export revenues and 1.9 million workers, 67% of surveyed member firms have already adopted AI tools (Source: U.S. International Trade Administration, 2026).

For a firm at the 14.9% mark, the honest blocker is rarely model quality. It is that a frontier-tier stack priced per token makes a high-volume use case uneconomic before anyone measures the benefit.

A summary pass on every inbound ticket, a classifier on every transaction, a retrieval step on every customer question. These are the workloads that determine whether a deployment pays back, and they are exactly the workloads where per-unit cost has been falling fastest.

Three routing rules that hold regardless of vendor

First, split by task shape rather than by model tier. Classification, summarization, extraction, and compaction are high-volume and narrow. Plan and long-horizon reasoning are low-volume and wide. The first group belongs on the small tier, the second does not.

Second, treat the cache as infrastructure. Anthropic halved Sonnet 5.5 cache reads, cutting about 20% off the cost of most agentic work (Source: Anthropic, 2026). Cache reads make up a large share of agent token consumption, so a caching gap is a recurring invoice.

Third, re-run your routing rules quarterly. Given a 47% quarterly decline in the cost of fixed performance, a routing table tuned six months ago is mispriced now.

A caveat worth taking seriously

Epoch AI's authors note their estimates come from benchmark data, and that benchmark performance is an imperfect proxy for useful work (Source: Epoch AI, 2026). Price declines on your own workload will not match the headline rate. Measure cost per completed task on your own data before redesigning anything.

FAQ

Q: Is a small model accurate enough for production customer work?
A: On narrow, repetitive tasks, current customer deployments report accuracy gains over the previous small model rather than losses. On open-ended reasoning, the larger model is still the better choice (Source: Anthropic, 2026).

Q: How much can I actually save by re-routing?
A: Anthropic reports around 75% lower average cost to run Haiku 5.5 versus Haiku 4.5, with 90% of prior requests falling in the cheaper under-100k-token bracket (Source: Anthropic, 2026).

Q: Does a bigger model always cost more per task?
A: No. At matched effort settings, a mid-tier model can outperform a frontier model at a fraction of the per-task cost, which is what the AutomationBench results show (Source: OpenAI, 2026).

Q: Should I switch models every time a new one launches?
A: Price declines fastest right after a capability first debuts, then slow over roughly two years. Evaluate on your own workload, then stay put if the numbers hold.

Key Takeaway

The most expensive decision in an AI stack is no longer which model answers the hard question. It is which model is doing the hundred easy ones per hard question. For most organizations, that routing decision is where the next order of magnitude of cost reduction lives.

Take one workflow this month, measure its cost per completed task, and ask whether the model handling it needs frontier capability at all.

Sources

Top comments (0)