<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Adam Z</title>
    <description>The latest articles on DEV Community by Adam Z (@adam_z).</description>
    <link>https://dev.to/adam_z</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4167616%2Fb535db69-52a5-449a-9ba6-57ae80e77d03.jpg</url>
      <title>DEV Community: Adam Z</title>
      <link>https://dev.to/adam_z</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/adam_z"/>
    <language>en</language>
    <item>
      <title>What Can $10 Buy Across Different LLM APIs?</title>
      <dc:creator>Adam Z</dc:creator>
      <pubDate>Wed, 07 Oct 2026 02:37:32 +0000</pubDate>
      <link>https://dev.to/adam_z/what-can-10-buy-across-different-llm-apis-4i1n</link>
      <guid>https://dev.to/adam_z/what-can-10-buy-across-different-llm-apis-4i1n</guid>
      <description>&lt;p&gt;LLM API pricing pages usually quote prices per million tokens.&lt;/p&gt;

&lt;p&gt;That is useful for billing, but it is not always the easiest way to think about cost when building an application.&lt;/p&gt;

&lt;p&gt;A question I find more intuitive is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If I have a $10 API budget, how many real requests can I make?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The answer depends heavily on the ratio between input and output tokens.&lt;/p&gt;

&lt;p&gt;Let's use one simple workload and compare a few models.&lt;/p&gt;

&lt;h2&gt;
  
  
  The workload
&lt;/h2&gt;

&lt;p&gt;Assume each API request contains:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;4,000 input tokens
1,000 output tokens
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That gives us a 5,000-token request with an 80/20 input-output split.&lt;/p&gt;

&lt;p&gt;For this comparison, I am using the pricing currently recorded in my dataset:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input / 1M tokens&lt;/th&gt;
&lt;th&gt;Output / 1M tokens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Luna&lt;/td&gt;
&lt;td&gt;$0.10&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;
&lt;td&gt;$0.75&lt;/td&gt;
&lt;td&gt;$3.75&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 5.5&lt;/td&gt;
&lt;td&gt;$2.00&lt;/td&gt;
&lt;td&gt;$10.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These are standard token rates and do not include special long-context tiers or other processing modes.&lt;/p&gt;

&lt;p&gt;Now let's turn those prices into something more concrete.&lt;/p&gt;

&lt;h2&gt;
  
  
  GPT-6 Luna
&lt;/h2&gt;

&lt;p&gt;For one request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Input:
4,000 / 1,000,000 × $0.10
= $0.0004

Output:
1,000 / 1,000,000 × $0.50
= $0.0005

Total:
$0.0009 per request
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With a $10 budget:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$10 / $0.0009
≈ 11,111 requests
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is roughly:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;44.4 million input tokens&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;and&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;11.1 million output tokens&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;for the $10 budget.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gemini 3.8 Flash
&lt;/h2&gt;

&lt;p&gt;Now run exactly the same workload through Gemini 3.8 Flash.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Input:
4,000 / 1,000,000 × $0.75
= $0.003

Output:
1,000 / 1,000,000 × $3.75
= $0.00375

Total:
$0.00675 per request
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With $10:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$10 / $0.00675
≈ 1,481 requests
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same workload.&lt;/p&gt;

&lt;p&gt;Same $10.&lt;/p&gt;

&lt;p&gt;Very different request capacity.&lt;/p&gt;

&lt;h2&gt;
  
  
  Claude Sonnet 5.5
&lt;/h2&gt;

&lt;p&gt;For Claude Sonnet 5.5:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Input:
4,000 / 1,000,000 × $2
= $0.008

Output:
1,000 / 1,000,000 × $10
= $0.01

Total:
$0.018 per request
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A $10 budget gives approximately:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$10 / $0.018
≈ 556 requests
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So for this particular workload:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Approx. requests for $10&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Luna&lt;/td&gt;
&lt;td&gt;11,111&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;
&lt;td&gt;1,481&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 5.5&lt;/td&gt;
&lt;td&gt;556&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That difference is large enough to matter when an application moves from experimentation to production traffic.&lt;/p&gt;

&lt;p&gt;But this table does &lt;strong&gt;not&lt;/strong&gt; mean GPT-6 Luna is automatically the best choice.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cheap tokens do not necessarily mean cheap tasks
&lt;/h2&gt;

&lt;p&gt;Imagine Model A needs one request to solve a coding task.&lt;/p&gt;

&lt;p&gt;Model B is cheaper per token but needs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;three retries&lt;/li&gt;
&lt;li&gt;a longer system prompt&lt;/li&gt;
&lt;li&gt;more tool calls&lt;/li&gt;
&lt;li&gt;more generated tokens&lt;/li&gt;
&lt;li&gt;an additional verification pass&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Model B may still produce a higher total cost per completed task.&lt;/p&gt;

&lt;p&gt;This is why I think there are two separate questions developers should ask.&lt;/p&gt;

&lt;h3&gt;
  
  
  Question 1: What does one request cost?
&lt;/h3&gt;

&lt;p&gt;That is mostly arithmetic.&lt;/p&gt;

&lt;h3&gt;
  
  
  Question 2: How many requests does the application need to complete the task?
&lt;/h3&gt;

&lt;p&gt;That is an evaluation problem.&lt;/p&gt;

&lt;p&gt;Token price answers the first question.&lt;/p&gt;

&lt;p&gt;It does not answer the second.&lt;/p&gt;

&lt;h2&gt;
  
  
  Output tokens deserve more attention
&lt;/h2&gt;

&lt;p&gt;There is another interesting detail in the example.&lt;/p&gt;

&lt;p&gt;The request contains four times as many input tokens as output tokens:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;4,000 input
1,000 output
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Yet for all three models in this example, the output portion costs more per token.&lt;/p&gt;

&lt;p&gt;For Claude Sonnet 5.5, the single request costs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Input:  $0.008
Output: $0.010
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So only 20% of the tokens account for more than half of the cost.&lt;/p&gt;

&lt;p&gt;This makes output control surprisingly important.&lt;/p&gt;

&lt;p&gt;If an agent produces verbose intermediate reasoning, oversized summaries, unnecessary code explanations, or repeated generated content, reducing output may save more money than aggressively shortening the prompt.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happens at production scale?
&lt;/h2&gt;

&lt;p&gt;A $10 experiment may not sound important.&lt;/p&gt;

&lt;p&gt;Now imagine the application processes 100,000 of these requests.&lt;/p&gt;

&lt;p&gt;Using the same 4,000-input / 1,000-output workload:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GPT-6 Luna:
100,000 × $0.0009
≈ $90

Gemini 3.8 Flash:
100,000 × $0.00675
≈ $675

Claude Sonnet 5.5:
100,000 × $0.018
≈ $1,800
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At that point model choice is no longer a minor implementation detail.&lt;/p&gt;

&lt;p&gt;It becomes part of the product economics.&lt;/p&gt;

&lt;h2&gt;
  
  
  Caching can change the calculation
&lt;/h2&gt;

&lt;p&gt;The examples above assume all input is charged at the standard input rate.&lt;/p&gt;

&lt;p&gt;Real applications can be different.&lt;/p&gt;

&lt;p&gt;Many agent workloads repeatedly send information such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;system instructions&lt;/li&gt;
&lt;li&gt;coding conventions&lt;/li&gt;
&lt;li&gt;repository context&lt;/li&gt;
&lt;li&gt;long reference documents&lt;/li&gt;
&lt;li&gt;tool definitions&lt;/li&gt;
&lt;li&gt;shared conversation context&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When an API supports discounted cached input, part of that input may become significantly cheaper.&lt;/p&gt;

&lt;p&gt;So a more complete calculation looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Total cost =
  normal input tokens × input price
+ cached input tokens × cached price
+ output tokens × output price
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For applications with large reusable prompts, caching can materially change the economics.&lt;/p&gt;

&lt;h2&gt;
  
  
  Context-window size is not a price guarantee
&lt;/h2&gt;

&lt;p&gt;Another thing I try not to assume is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This model supports a one-million-token context window, therefore one million tokens always cost the normal input rate.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Providers can have special pricing rules for large contexts, different processing tiers, batch workloads, caching, or other modes.&lt;/p&gt;

&lt;p&gt;A model's context window tells you what is technically possible.&lt;/p&gt;

&lt;p&gt;It does not necessarily tell you what that workload will cost.&lt;/p&gt;

&lt;p&gt;That is why pricing comparisons should preserve the provider source and the date when the pricing was verified.&lt;/p&gt;

&lt;h2&gt;
  
  
  Turning the calculation into a small tool
&lt;/h2&gt;

&lt;p&gt;I was repeatedly doing calculations like these manually, so I built a simple &lt;a href="https://iilib.com/workbench/llm-api-cost" rel="noopener noreferrer"&gt;LLM API Cost Calculator&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Instead of asking only for a model price, it lets you enter a workload:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;input tokens / request
cached input tokens / request
output tokens / request
number of requests
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is also a budget-oriented way to think about the problem: start with a fixed amount of money and estimate the workload capacity.&lt;/p&gt;

&lt;p&gt;I deliberately keep the pricing source and verification date next to the results because a perfectly accurate calculation based on outdated pricing is still a wrong answer.&lt;/p&gt;

&lt;p&gt;The underlying rates are also available separately in the &lt;a href="https://iilib.com/research/llm-api-pricing" rel="noopener noreferrer"&gt;LLM API Pricing reference&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The metric I actually care about
&lt;/h2&gt;

&lt;p&gt;After looking at enough pricing tables, I have become less interested in:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Which model has the cheapest token?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The metric I would rather know is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How much does it cost to successfully complete one useful task?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For some applications, a very inexpensive model will win.&lt;/p&gt;

&lt;p&gt;For others, paying more for a model that completes the task reliably in fewer steps may be cheaper overall.&lt;/p&gt;

&lt;p&gt;A useful evaluation therefore needs both:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cost per request
×
requests per successful task
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That gets much closer to the economics of a real AI application than a simple price-per-million-token leaderboard.&lt;/p&gt;




&lt;p&gt;Pricing changes frequently. The figures above reflect the pricing data I was using in October 2026. Always verify current provider pricing and pricing conditions before making production cost decisions.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>webdev</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
