DEV Community

Cover image for Tokens Are Not the Metric
Andrés Clúa
Andrés Clúa

Posted on

Tokens Are Not the Metric

When developers start using AI seriously, one metric appears almost immediately: token usage.

  • How many tokens did this agent consume?
  • How much did this coding session cost?
  • Can we reduce the context window?
  • Can we use a cheaper model?

Those are reasonable questions. But they can also lead us toward the wrong optimization.

The goal should not be to use fewer tokens.
The goal should be to waste fewer tokens.

There is a big difference.

A million tokens can be cheap

Imagine two AI agents. These are hypothetical examples, not benchmark results.

Agent A Agent B
Tokens 50 million 5 million
What it does Implements a feature Repeatedly scans the same repository
Writes tests Retries failed tasks
Fixes regressions Summarizes information nobody reads
Opens a pull request Runs every 30 minutes
Generates documentation Produces no meaningful output
Result Saves an engineer two days of work Nothing

Which one is expensive?

Looking only at token consumption, Agent A looks ten times worse. Looking at actual value, Agent B may be the worse investment.

This is why I increasingly think about AI usage in three buckets: learning, production, and spin.

1. Learning tokens

These are the tokens we spend figuring things out:

  • Trying a new coding agent
  • Experimenting with MCP
  • Building a better system prompt
  • Testing whether an agent can translate a Figma design into production code
  • Letting a model explore an unfamiliar codebase

Some of this work will fail. That's fine.

In traditional software development, we don't expect every minute spent by an engineer to produce production code. We research. We prototype. We read documentation. We throw things away.

AI systems need the same space. Trying to aggressively eliminate this type of token usage can actually slow down adoption.

The important question isn't "How many tokens did we spend?"

It's "What did we learn from spending them?"

2. Production tokens

These are the easiest tokens to justify. They produce something useful:

  • A pull request
  • A migration
  • A test suite
  • A component
  • A technical document
  • A content update
  • A bug fix
  • An accessibility audit
  • A production deployment

Once AI moves beyond autocomplete and starts operating as an agent, token consumption becomes closer to compute.

You wouldn't measure a CI pipeline only by how many CPU seconds it consumed. You care about what the pipeline accomplished. AI should be evaluated the same way.

A useful metric might be:

AI cost / completed task
AI cost / pull request merged
AI cost / engineering hour saved
Enter fullscreen mode Exit fullscreen mode

Suddenly, token usage becomes much more meaningful.

These metrics still need context: task complexity, quality, review time, and downstream rework. A merged PR is useful evidence, but it isn't proof that the work saved time.

3. Spin

This is where the real problem lives.

Spin is AI activity that looks productive but produces little or no value.

Some common examples:

  • Agents repeatedly reading the same files without a reason
  • Enormous context windows being resent unnecessarily
  • Automated jobs running far more often than needed
  • Multiple agents duplicating the same research
  • Retry loops without proper termination conditions
  • Long reasoning chains for trivial operations
  • Generated reports nobody reads
  • Background agents running simply because they can

This is the AI equivalent of an infinite loop with a cloud bill attached.

And as AI systems become more autonomous, spin becomes much easier to create. A developer using an AI assistant manually will eventually notice that something isn't working. A scheduled agent can waste money quietly for weeks.

The problem with "make it cheaper"

When AI costs start increasing, the first reaction is usually: use a cheaper model.

Sometimes that's correct. But it's often treating the symptom instead of the cause.

Suppose an agent performs a task 48 times per day when it only needs to run twice. Switching to a model that costs half as much sounds like an optimization. But reducing the execution frequency would cut the execution cost by roughly 96%, assuming the cost per run stays the same.

const costPerRun = 1;

const cheaperModel = 48 * (costPerRun * 0.5); // 24
const fewerRuns = 2 * costPerRun;            // 2

const savingsCheaperModel = 1 - cheaperModel / 48; // 0.50
const savingsFewerRuns = 1 - fewerRuns / 48;       // ~0.96
Enter fullscreen mode Exit fullscreen mode

Architecture can matter more than model pricing.

Before changing models, I would ask:

  • Does this agent need to run this often?
  • Does it need the entire repository as context?
  • Can it retrieve only the relevant files?
  • Does every step need the most capable model?
  • Can deterministic code replace part of the workflow?
  • Does the output lead to an actual action?
  • Does the agent know when to stop?

These questions can produce much larger savings.

AI efficiency is becoming an engineering discipline

We already have mature ways of measuring traditional software systems: latency, CPU, memory, error rates, throughput, availability.

AI systems need their own operational metrics.

Cost per successful task

total AI cost / successfully completed tasks
Enter fullscreen mode Exit fullscreen mode

Include failed runs in the numerator, and define success before running the workflow.

Retry rate

executions requiring at least one retry / total executions
Enter fullscreen mode Exit fullscreen mode

A high retry rate can reveal brittle prompts, bad tools, or poor agent architecture.

Human intervention rate

tasks requiring human correction / total tasks
Enter fullscreen mode Exit fullscreen mode

Track correction time too. A quick approval and a two-hour repair have very different costs.

Useful output rate

How often does an agent produce something that actually gets used?

Context efficiency

How much of the provided context was actually relevant? This is harder to measure directly, but comparing focused retrieval against larger context inputs can help.

Agent idle activity

How often does the system execute when nothing meaningful changed?

These are more useful indicators of AI efficiency than raw token counts alone.

The cheapest token is not always the best token

Developers have spent decades learning that premature optimization can make systems worse. AI is no different.

If a stronger model uses 30% more tokens but solves the problem correctly on the first attempt, it may be cheaper overall than a smaller model requiring five retries. The result depends on pricing and the work each attempt performs.

Likewise, giving an agent more context can sometimes reduce total cost if it prevents it from exploring blindly.

Token efficiency is not the same thing as token minimization. The objective is productive computation.

Think in terms of return

A useful starting point is:

value-to-cost ratio = value produced / total workflow cost
Enter fullscreen mode Exit fullscreen mode

Total workflow cost should include AI execution, tools, human review, and rework. Estimating the value produced is harder, but ignoring it doesn't make the cost metric more useful.

  • A $20 agent run that saves four hours of engineering time is probably excellent.
  • A $0.20 automated task that runs thousands of times without producing anything useful is not.

As AI becomes embedded in engineering workflows, teams will need to become comfortable with this distinction. Otherwise, we risk optimizing systems for impressive-looking token dashboards rather than actual productivity.

The real optimization

So when your AI bill starts growing, don't immediately ask "How do we use fewer tokens?"

Ask:

Which spending produces useful work, which helps us learn, and which keeps the system busy without moving anything forward?

Here is how I would start:

  1. Pick one workflow.
  2. Define what success looks like.
  3. Record the full cost of getting there.
  4. Check whether anyone actually uses the output.
  5. Remove unnecessary runs, bound retries, and make the stopping condition explicit.

Protect experiments that teach you something. Improve workflows that deliver useful work. Shut down activity that does neither.

That's a much better optimization target than a smaller number on a token dashboard.


Inspired by The Token Gym and its distinction between learning, shipped work, and spin. The agent scenarios and workflow examples above are illustrative.

What do you measure in your AI workflows today: token consumption, cost per task, or whether the output actually gets used? Let me know in the comments.

Top comments (0)