DEV Community

Nicolai Bohn
Nicolai Bohn

Posted on

Agent cost optimization: same quality, lower price

Most teams pick their agent's model while prototyping. It works, it ships, and nobody touches it again. Meanwhile cheaper models come out every few weeks.

Switching sounds easy until someone asks whether the new model will handle your refund policy questions as well as the old one. Nobody knows. There is no test set to run it against, so the team stays on the old model, and the bill stays where it is.

Your agent's cost is a design decision

Breaking that loop takes two things: knowing what a single run of your agent costs, and a safe way to compare one setup against another.

Where the money goes

Cost is tokens times price, summed over every LLM call your agent makes. Four levers decide how big that number gets.

Four levers decide what a run costs

Model. Price per token can differ by an order of magnitude between model tiers. For most agents this is the biggest lever. Test it by swapping the model.

Prompt and context. The system prompt and any retrieved context are sent again on every call. A 3,000-token system prompt in a ten-turn conversation is 30,000 input tokens before the user has said much at all. Test it by shortening the prompt.

Agent steps. Every tool call or reasoning step is another LLM call, and conversations carry their history forward. An agent that takes six steps where two would do pays for it on every request. Test it by capping the steps.

Caching. Many providers bill cached input tokens at a lower rate. Keeping the start of your prompt stable lets you benefit from that. Test it by keeping prefixes stable and comparing.

Cheaper only counts if quality holds

A cost number on its own tells you very little. If v2 costs half as much but fails a third of your cancellation tests, you haven't saved anything. You've moved the cost to your support team.

So look at pass rate and cost together, on the same test set.

Cheaper only counts if quality holds

That gives you four outcomes:

  • Cheaper, same or better quality: promote it.
  • More expensive, better quality: decide whether the gain matters to your users.
  • Cheaper, lower quality: check which behaviors failed. A prompt tweak may win the quality back.
  • More expensive, lower quality: roll it back.

Change one thing, compare side by side

The method is plain experimentation. Change one variable at a time so you know what caused the difference.

Change one thing, compare side by side

  1. Save a baseline. v1 is your current setup: model, prompt and temperature.
  2. Vary one parameter. v2 swaps the model. v3 keeps the model and trims the prompt.
  3. Run the same tests against each version, with the same metrics.
  4. Compare the runs on pass rate and cost.

For a support agent this could look like:

Version Model System prompt What you're testing
v1 your current model full prompt baseline
v2 a smaller model full prompt does the cheaper model hold up?
v3 your current model trimmed prompt how much does the prompt cost you?

How this works in Rhesis

Rhesis is an open-source platform for testing LLM apps and agents. We recently added cost tracking to test runs, and it pairs well with Experiments:

  • Every test run shows its tokens, cost and model in the runs list. You can sort by cost to find your most expensive runs.
  • Each run has a Cost card with input and output tokens priced separately. It updates live while the run executes.

Test run summary with the Cost card next to pass rate

  • Experiments store configurations as versions. A parameter schema defines the slots (model, prompt, temperature), and every change becomes an immutable version, so a run against v2 is reproducible later.
  • Side-by-side comparison shows pass rate and cost with deltas against the baseline. The cost delta turns green when the new run was cheaper.

Costs are calculated from the tokens each LLM call reports, multiplied by list prices from LiteLLM's model database. That makes them a good relative comparison. They don't include discounts or committed-use pricing, and models missing from the database show tokens without a cost. Details are in the token costs docs and the experiments docs.

Here's the walkthrough:

Try it

Rhesis is MIT-licensed. You can use it at app.rhesis.ai or self-host it from GitHub.

If you're running cost experiments on your own agents, I'd like to hear which lever saved you the most. Let me know in the comments.

Originally published on the Rhesis blog.

Top comments (0)