DEV Community

Cover image for AI Won’t Be Priced by Tokens Forever
Robert Imbeault
Robert Imbeault

Posted on AI-assisted

AI Won’t Be Priced by Tokens Forever

Nobody wants to buy tokens. They want useful work, and eventually, AI will be priced that way too.

Every new technology seems to begin by billing us for the thing engineers can measure most easily.

Electricity gave us kilowatt-hours.
Telecommunications gave us minutes.
Cloud computing gave us CPUs, memory, storage, and bandwidth.

AI has tokens.

And right now, token pricing makes perfect sense, because tokens are measurable. Providers need to charge for computation, and developers need a way to estimate their bills. Input and output tokens give both sides a usable starting point.

For an industry still figuring itself out, it is a sensible meter, but I don’t think the meter survives as the product.

Because nobody wakes up in the morning thinking:
You know what would really improve the business today? Another 40 million tokens.

People want things done. Tokens are just what gets consumed along the way.

Cloud Went Through the Same Phase

Early cloud computing was very infrastructure-shaped.

Organizations talked about CPU counts, memory allocations, storage volumes, and bandwidth because those were the things they were buying.

Then cloud matured.

Developers stopped thinking quite so much about disks and started buying managed databases. They stopped managing individual servers and started consuming platforms.

The underlying infrastructure did not disappear.

Someone still has to care about CPU utilization at 2 a.m., it just stopped being the thing most customers thought they were purchasing.

The abstraction moved upward, and I think AI is beginning the same transition.

Tokens Are a Cost. They Are Not the Outcome.

Today, almost every major AI platform exposes pricing in terms of input and output tokens.

That makes comparison easy.

Model A costs this much per million.
Model B costs that much.
Model C is cheaper, provided you only read the footnotes after midnight.

But token cost tells us surprisingly little about the value of the system.

A customer does not care how many tokens were required to summarize a meeting, they care whether the summary was useful.

They do not care how many tokens were consumed analyzing a contract, they care whether the important clauses were identified correctly.

Tokens measure computational consumption.

Customers care about completed work.

Those are related. They are not the same thing.

Better AI Can Actually Mean Using Less AI

This is where the economics become interesting.

Some of the most useful improvements in AI systems reduce token consumption rather than increase it.

  • Better prompts can produce better answers with less context.
  • Better retrieval can avoid stuffing unnecessary information into every request.
  • Smarter routing can send straightforward work to smaller, cheaper models.
  • Quantization can serve more requests on the same hardware.
  • Post-training can teach recurring organizational knowledge into a model instead of attaching the same context to every prompt forever.

None of those improvements make the customer’s experience worse.

Quite the opposite, because they improve the system while often reducing the amount of computation required. And this creates a slightly awkward incentive.

If your business grows when customers consume more tokens, efficiency can look suspiciously like a revenue problem.

If your business grows when customers accomplish more with less infrastructure, efficiency is the product.

Those are very different businesses.

The Same Million Tokens Can Produce Very Different Value

Imagine two AI systems processing the same workload.

Both consume one million tokens.

One repeatedly sends irrelevant context, retrieves the wrong information, uses an unnecessarily large model, and eventually produces something useful after considerable persuasion.

The other gives the right model the right information and gets to the answer efficiently.

Same number of tokens.
Very different system.

This is why I think token pricing will gradually become less useful as a proxy for value.

The model is only part of the equation. The software around it determines whether it gets useful context or has to sift through irrelevant information. Routing each task to an appropriate model can also avoid expensive computation that adds nothing to the result.

The customer ultimately experiences the system, not the token counter.

Eventually, Enterprises Will Ask Different Questions

As AI becomes a larger part of enterprise infrastructure, I expect buying decisions to move upward from raw consumption.

Organizations will ask different questions that are harder to put on a pricing page than $X / 1M tokens.

Unfortunately for pricing pages, they are much closer to the questions customers actually care about such as security, workload, and productivity.

Local AI Makes Token Pricing Even Stranger

Token pricing becomes less useful as AI moves beyond centralized APIs.

Some workloads will stay in the public cloud, while others will run on infrastructure companies manage themselves. More capable hardware will also make it practical to run models directly on laptops and workstations.

When a company owns the hardware, the cost calculation changes. It’s already paying for the equipment and its operation, rather than receiving a bill for every token. What matters is whether that investment delivers enough useful work to justify the expense.

Engineers will still count tokens to understand performance and spot inefficiencies. But token volume alone won’t tell the business whether it’s getting value from the system.

Tokens Are Going to Become an Implementation Detail

I don’t think tokens disappear.

CPU utilization did not disappear when cloud platforms moved toward managed services.

Storage did not disappear when developers stopped manually provisioning disks.

Bandwidth certainly did not disappear. My monthly bills remain committed to proving that.

Those metrics still matter enormously to the people operating the infrastructure.

Tokens will be the same. Engineers will continue tracking them.
They will matter for cost optimization, context management, inference efficiency, and system design.

But customers will increasingly judge AI by something else:
What did it accomplish, and what did that outcome cost us?

That is a much more meaningful abstraction.

Eventually We’ll Price What People Actually Came to Buy

The most valuable AI systems will not necessarily be the ones producing the largest number of tokens.

They will be the ones delivering the best outcomes with the least friction and the greatest efficiency.

Tokens are useful.
They are measurable.
They are operationally important.
But they are not the product.

Like CPU hours before them, tokens are simply the meter we use until the industry gets better at pricing the thing customers actually wanted in the first place: useful work.


This is a conversational remix of an article I published on Backboard’s blog. Read the original deep dive here.

Top comments (1)

Collapse
 
citedy profile image
Dmitry Sergeev

We need to produce a short YouTube comment as a regular developer, casual, referencing the video content. Must not be generic praise; must mention something specific: maybe ask about token pricing models, or mention how they'd like to see pricing per compute. Must be short, one or two sentences, possibly fragment. Use lowercase start, casual voice. No URLs. No double hyphens, no em-dash. Ensure no smart quotes. No markdown. Just plain text. Let's create something like: "so if we move to per