DEV Community

Cover image for What Can $10 Buy Across Different LLM APIs?
Adam Z
Adam Z

Posted on

What Can $10 Buy Across Different LLM APIs?

LLM API pricing pages usually quote prices per million tokens.

That is useful for billing, but it is not always the easiest way to think about cost when building an application.

A question I find more intuitive is:

If I have a $10 API budget, how many real requests can I make?

The answer depends heavily on the ratio between input and output tokens.

Let's use one simple workload and compare a few models.

The workload

Assume each API request contains:

4,000 input tokens
1,000 output tokens
Enter fullscreen mode Exit fullscreen mode

That gives us a 5,000-token request with an 80/20 input-output split.

For this comparison, I am using the pricing currently recorded in my dataset:

Model Input / 1M tokens Output / 1M tokens
GPT-6 Luna $0.10 $0.50
Gemini 3.8 Flash $0.75 $3.75
Claude Sonnet 5.5 $2.00 $10.00

These are standard token rates and do not include special long-context tiers or other processing modes.

Now let's turn those prices into something more concrete.

GPT-6 Luna

For one request:

Input:
4,000 / 1,000,000 × $0.10
= $0.0004

Output:
1,000 / 1,000,000 × $0.50
= $0.0005

Total:
$0.0009 per request
Enter fullscreen mode Exit fullscreen mode

With a $10 budget:

$10 / $0.0009
≈ 11,111 requests
Enter fullscreen mode Exit fullscreen mode

That is roughly:

44.4 million input tokens

and

11.1 million output tokens

for the $10 budget.

Gemini 3.8 Flash

Now run exactly the same workload through Gemini 3.8 Flash.

Input:
4,000 / 1,000,000 × $0.75
= $0.003

Output:
1,000 / 1,000,000 × $3.75
= $0.00375

Total:
$0.00675 per request
Enter fullscreen mode Exit fullscreen mode

With $10:

$10 / $0.00675
≈ 1,481 requests
Enter fullscreen mode Exit fullscreen mode

Same workload.

Same $10.

Very different request capacity.

Claude Sonnet 5.5

For Claude Sonnet 5.5:

Input:
4,000 / 1,000,000 × $2
= $0.008

Output:
1,000 / 1,000,000 × $10
= $0.01

Total:
$0.018 per request
Enter fullscreen mode Exit fullscreen mode

A $10 budget gives approximately:

$10 / $0.018
≈ 556 requests
Enter fullscreen mode Exit fullscreen mode

So for this particular workload:

Model Approx. requests for $10
GPT-6 Luna 11,111
Gemini 3.8 Flash 1,481
Claude Sonnet 5.5 556

That difference is large enough to matter when an application moves from experimentation to production traffic.

But this table does not mean GPT-6 Luna is automatically the best choice.

Cheap tokens do not necessarily mean cheap tasks

Imagine Model A needs one request to solve a coding task.

Model B is cheaper per token but needs:

  • three retries
  • a longer system prompt
  • more tool calls
  • more generated tokens
  • an additional verification pass

Model B may still produce a higher total cost per completed task.

This is why I think there are two separate questions developers should ask.

Question 1: What does one request cost?

That is mostly arithmetic.

Question 2: How many requests does the application need to complete the task?

That is an evaluation problem.

Token price answers the first question.

It does not answer the second.

Output tokens deserve more attention

There is another interesting detail in the example.

The request contains four times as many input tokens as output tokens:

4,000 input
1,000 output
Enter fullscreen mode Exit fullscreen mode

Yet for all three models in this example, the output portion costs more per token.

For Claude Sonnet 5.5, the single request costs:

Input:  $0.008
Output: $0.010
Enter fullscreen mode Exit fullscreen mode

So only 20% of the tokens account for more than half of the cost.

This makes output control surprisingly important.

If an agent produces verbose intermediate reasoning, oversized summaries, unnecessary code explanations, or repeated generated content, reducing output may save more money than aggressively shortening the prompt.

What happens at production scale?

A $10 experiment may not sound important.

Now imagine the application processes 100,000 of these requests.

Using the same 4,000-input / 1,000-output workload:

GPT-6 Luna:
100,000 × $0.0009
≈ $90

Gemini 3.8 Flash:
100,000 × $0.00675
≈ $675

Claude Sonnet 5.5:
100,000 × $0.018
≈ $1,800
Enter fullscreen mode Exit fullscreen mode

At that point model choice is no longer a minor implementation detail.

It becomes part of the product economics.

Caching can change the calculation

The examples above assume all input is charged at the standard input rate.

Real applications can be different.

Many agent workloads repeatedly send information such as:

  • system instructions
  • coding conventions
  • repository context
  • long reference documents
  • tool definitions
  • shared conversation context

When an API supports discounted cached input, part of that input may become significantly cheaper.

So a more complete calculation looks like:

Total cost =
  normal input tokens × input price
+ cached input tokens × cached price
+ output tokens × output price
Enter fullscreen mode Exit fullscreen mode

For applications with large reusable prompts, caching can materially change the economics.

Context-window size is not a price guarantee

Another thing I try not to assume is:

This model supports a one-million-token context window, therefore one million tokens always cost the normal input rate.

Providers can have special pricing rules for large contexts, different processing tiers, batch workloads, caching, or other modes.

A model's context window tells you what is technically possible.

It does not necessarily tell you what that workload will cost.

That is why pricing comparisons should preserve the provider source and the date when the pricing was verified.

Turning the calculation into a small tool

I was repeatedly doing calculations like these manually, so I built a simple LLM API Cost Calculator.

Instead of asking only for a model price, it lets you enter a workload:

input tokens / request
cached input tokens / request
output tokens / request
number of requests
Enter fullscreen mode Exit fullscreen mode

There is also a budget-oriented way to think about the problem: start with a fixed amount of money and estimate the workload capacity.

I deliberately keep the pricing source and verification date next to the results because a perfectly accurate calculation based on outdated pricing is still a wrong answer.

The underlying rates are also available separately in the LLM API Pricing reference.

The metric I actually care about

After looking at enough pricing tables, I have become less interested in:

Which model has the cheapest token?

The metric I would rather know is:

How much does it cost to successfully complete one useful task?

For some applications, a very inexpensive model will win.

For others, paying more for a model that completes the task reliably in fewer steps may be cheaper overall.

A useful evaluation therefore needs both:

cost per request
×
requests per successful task
Enter fullscreen mode Exit fullscreen mode

That gets much closer to the economics of a real AI application than a simple price-per-million-token leaderboard.


Pricing changes frequently. The figures above reflect the pricing data I was using in October 2026. Always verify current provider pricing and pricing conditions before making production cost decisions.

Top comments (0)