DEV Community

Cover image for Using Token Data to Improve the AI Agent Experience
Quantiles.io
Quantiles.io

Posted on

Using Token Data to Improve the AI Agent Experience

Token usage can tell you a lot about what happened between the start of a task and the final answer. An agent may complete the task successfully, but only after searching the documentation several times, retrying failed steps, or pulling large amounts of irrelevant tool output into context. Looking at token usage metrics across repeated trials helps surface that extra work and distinguish a direct, efficient workflow from one that depends on additional context, retries, and recovery.

Tokens add up

Tokens are the units a model uses to process content. In text, a token is roughly 3/4 of a single word, but different models use different tokenizers, so the same text can produce different token counts on different systems. That difference also means equal token totals across models don't necessarily represent the same amount of content or computation.

Token usage is usually measured by numbers of input and output tokens, and some models can also report reasoning tokens, which are used in reasoning models' chains of thought. Input tokens are everything sent to the model, including the prompt, conversation history, documentation, code, and tool results. Output tokens are what the model generates in response. In agent workflows, input tokens can add up quickly as more context is carried across multiple model calls.

The way different hosted model providers price tokens also varies by provider, though generally output tokens are often more expensive per token than input tokens, even when input tokens make up most of the overall usage.

The example below shows a single agent evaluation task through Quantiles and breaks down the input and output tokens used at each step of the workflow.

Tracing Token Usage Through a Quantiles Evaluation Task

Tracing Token Usage Through a Quantiles Evaluation Task

Using tokens to understand agent behavior

In an agent workflow, token usage can help reflect the path the agent takes through a task: the context it reads, the tools it calls, the retries it makes, and the information it carries forward. That makes it a useful signal for behavior beyond task completion. For example, two trials may both succeed, even though one requires more context, additional exploration, more recovery steps, and longer execution time.

Recent research shows just how much that path can vary. A 2026 study of eight frontier models on SWE-bench Verified found that repeated executions of the same task could differ by as much as 30 times in total token usage. More tokens didn't necessarily lead to better results either. Accuracy often peaked at an intermediate level of token usage and then plateaued as consumption continued to increase.

Some patterns in the evaluation worth looking for include:

  • Repeated documentation reads. The agent may be struggling to find, read, or interpret the instructions it needs.
  • Large tool responses. A tool may be returning much more context than the agent needs for the task.
  • Retries and recovery. Token usage can climb quickly when the agent encounters an error, revisits earlier steps, or tries several approaches.
  • Growing context. Long conversations, tool outputs, and intermediate results can accumulate across model calls and increase input-token usage.
  • Unnecessary exploration. The agent may inspect files, APIs, or documentation that are unrelated to the task. In some cases, the model may not be able to read the data it requires without bringing large amounts of irrelevant data into context as well.
  • Extra verification. Higher usage isn't always waste. Sometimes the additional tokens come from checking work, testing an implementation, or validating the final result.

What matters most is the pattern, not whether a token count is simply high or low. Running the same task multiple times lets you compare usage across trials and spot meaningful differences. If one attempt uses substantially more tokens than the others, the trace can help show whether the agent needed more context, retries, or effort to recover from an error to reach the same result.

Using token data to improve the agent experience

For a developer tool, token usage is more than a model cost. It can tell you how much work an agent has to do to successfully use your product.

If agents consistently need large amounts of context, repeated tool calls, or several attempts to complete common workflows, that can be a sign of friction or barriers to entry in the product itself. The agent may be spending extra tokens because the documentation is hard to navigate, tool responses contain more information than it needs, an API requires too many steps, or an error message doesn't provide enough guidance to recover cleanly. In that sense, excess token usage can point to places where the product is making the agent work harder than necessary.

The table below shows how token usage data can point to opportunities for improving the agent experience.

Opportunities to Improve the Agent Experience

Opportunity How it helps
Reduce unnecessary context Smaller, more focused tool responses and clearer documentation can keep irrelevant information out of the model context.
Make workflows more direct Fewer unnecessary steps, searches, and retries can reduce both token usage and completion time.
Improve reliability High or highly variable token usage can point to places where agents are getting confused or taking inconsistent, inefficient paths.
Lower customer costs When customers pay for model usage, a more token efficient product can make every agent interaction less expensive.
Make your product easier for agents to use Clear interfaces, useful errors, focused tools, and discoverable documentation can help agents reach the right outcome with less exploration.

Using fewer tokens isn't the goal on its own. It's more important to create a product that agents can understand, navigate, and use efficiently. By tracking token usage alongside task success and execution traces, you can see where the experience is already working well and where improvements could make agent workflows even faster, more reliable, and more successful.

Top comments (1)

Collapse
 
devantibot profile image
DEV ANTIBOT •

You need to complete account verification.Link in the profile.