<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Andrea Susic</title>
    <description>The latest articles on DEV Community by Andrea Susic (@asymm).</description>
    <link>https://dev.to/asymm</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3956731%2F6b688374-37b8-4fb7-aa52-ceb57a5ae04d.JPG</url>
      <title>DEV Community: Andrea Susic</title>
      <link>https://dev.to/asymm</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/asymm"/>
    <language>en</language>
    <item>
      <title>How to estimate your LLM inference bill before you ship (with Python)</title>
      <dc:creator>Andrea Susic</dc:creator>
      <pubDate>Thu, 01 Oct 2026 08:26:29 +0000</pubDate>
      <link>https://dev.to/asymm/how-to-estimate-your-llm-inference-bill-before-you-ship-with-python-1j6m</link>
      <guid>https://dev.to/asymm/how-to-estimate-your-llm-inference-bill-before-you-ship-with-python-1j6m</guid>
      <description>&lt;p&gt;Most unexpected LLM bills start with a misleadingly cheap prototype.&lt;/p&gt;

&lt;p&gt;Someone tests a feature with a handful of short prompts, checks the usage dashboard and concludes that inference costs almost nothing. Then the application reaches production. Prompts get longer, retrieval adds context, conversations stretch across multiple turns and thousands of users start making requests.&lt;/p&gt;

&lt;p&gt;A month later, the invoice looks very different.&lt;/p&gt;

&lt;p&gt;Fortunately, inference cost is relatively straightforward to model. If your provider charges by token, you can estimate the cost of a realistic request before launch and then model what happens as traffic grows.&lt;br&gt;
The arithmetic is simple. Getting realistic inputs is the important part.&lt;/p&gt;
&lt;h2&gt;
  
  
  How LLM inference is usually billed
&lt;/h2&gt;

&lt;p&gt;Many hosted LLM APIs charge separately for input and output tokens.&lt;br&gt;
Input tokens include everything sent to the model: the system prompt, retrieved context, conversation history, tool information and the user's current message.&lt;br&gt;
Output tokens are the tokens generated by the model.&lt;br&gt;
Providers commonly quote both prices per million tokens, and output tokens are often more expensive than input tokens.&lt;br&gt;
At its simplest, the calculation is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cost = (input_tokens × input_price + output_tokens × output_price) / 1,000,000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The formula is rarely the source of forecasting errors.&lt;br&gt;
The problem is estimating &lt;code&gt;input_tokens&lt;/code&gt;.&lt;br&gt;
A production request is usually much larger than the text the user sees in the chat box.&lt;/p&gt;
&lt;h2&gt;
  
  
  Step 1: Count a realistic request
&lt;/h2&gt;

&lt;p&gt;Start with something that resembles the request your application will actually send.&lt;br&gt;
If you are building a retrieval-augmented support assistant, for example, include the system instructions, retrieved documents and user message rather than measuring the user message alone.&lt;/p&gt;

&lt;p&gt;Install &lt;code&gt;tiktoken&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;tiktoken
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then create a representative request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;tiktoken&lt;/span&gt;

&lt;span class="c1"&gt;# o200k_base is a modern BPE tokenizer.
# Different model families can use different tokenizers, so use the
# target model's tokenizer when you need model-specific counts.
&lt;/span&gt;&lt;span class="n"&gt;enc&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tiktoken&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_encoding&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;o200k_base&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;enc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="n"&gt;SYSTEM_PROMPT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;You are a support assistant for an online store.
Answer only from the provided context. If the answer is not in the
context, say you don&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;t know and offer to connect a human agent.
Keep answers under 120 words and use a friendly, plain tone.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

&lt;span class="n"&gt;RETRIEVED_CONTEXT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Returns: items can be returned within 30 days of delivery.
Refunds go back to the original payment method within 5 business days.
Shipping: standard delivery takes 2-4 business days. Express is next day if ordered before 14:00. Damaged items: send a photo to support and we replace the item at no cost.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

&lt;span class="n"&gt;USER_MESSAGE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Hi, my blender arrived with a cracked jug.
Can I get a new one, and how long will it take?&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

&lt;span class="n"&gt;SAMPLE_ANSWER&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Sorry to hear the jug arrived cracked. Please reply with a photo
of the damage and your order number, and we&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;ll send a replacement at no cost. Replacements ship with standard delivery, which takes 2-4 business days. If you&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;d prefer a refund instead, that&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s possible too, and it reaches your original payment method within 5 business days.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

&lt;span class="n"&gt;input_tokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="nf"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;SYSTEM_PROMPT&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nf"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;RETRIEVED_CONTEXT&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nf"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;USER_MESSAGE&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;output_tokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;SAMPLE_ANSWER&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbkhrql5gt8oz9m8ddvtk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbkhrql5gt8oz9m8ddvtk.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is still an estimate. Chat APIs can add formatting and message-structure overhead, and different model families use different tokenizers.&lt;br&gt;
For an open-weight model, use the tokenizer associated with the model when possible. With models distributed through Hugging Face, that might look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;transformers&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;AutoTokenizer&lt;/span&gt;

&lt;span class="n"&gt;tok&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AutoTokenizer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_pretrained&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;&amp;lt;model-repo-id&amp;gt;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;count_model_tokens&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tok&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The closer your test request is to production traffic, the more useful the estimate becomes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Turn tokens into money
&lt;/h2&gt;

&lt;p&gt;Once you have representative token counts, pricing different models is straightforward.&lt;/p&gt;

&lt;p&gt;The rates below are deliberately illustrative. Replace them with the current input and output prices from the providers you are evaluating.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Illustrative prices per 1M tokens.
# Replace these with your provider's current rates.
&lt;/span&gt;&lt;span class="n"&gt;PRICES&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;small model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;input&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;0.20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;output&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;0.80&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;mid model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;   &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;input&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;1.00&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;output&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;4.00&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;large model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;input&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;3.00&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;output&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;15.00&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="n"&gt;REQUESTS_PER_DAY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2_000&lt;/span&gt;
&lt;span class="n"&gt;DAYS_PER_MONTH&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;

&lt;span class="c1"&gt;# Allow for retries, longer-than-average responses,
# prompt changes and other production variance.
&lt;/span&gt;&lt;span class="n"&gt;SAFETY_MARGIN&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;1.3&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;cost_per_request&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;price&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="nf"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;input_tokens&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;price&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;input&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;output_tokens&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;price&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;output&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1_000_000&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Input tokens per request:  &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;input_tokens&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Output tokens per request: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;output_tokens&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;per request&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;per month&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;price&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;PRICES&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;per_req&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;cost_per_request&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;price&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;monthly&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;per_req&lt;/span&gt;
        &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;REQUESTS_PER_DAY&lt;/span&gt;
        &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;DAYS_PER_MONTH&lt;/span&gt;
        &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;SAFETY_MARGIN&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;per_req&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="mf"&gt;12.6&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;monthly&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="mf"&gt;12.2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Using the example token counts, you might get something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Input tokens per request:  149
Output tokens per request: 74

model         per request    per month
small model      0.000089         6.94
mid model        0.000445        34.71
large model      0.001557       121.45
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The exact figures are not important because both tokenisation and provider pricing will vary.&lt;br&gt;
The comparison is what matters.&lt;/p&gt;

&lt;p&gt;The same application and traffic assumptions can produce substantially different monthly costs depending on model choice. At the same time, the calculation may reveal that inference is a relatively small part of the application's total operating cost.&lt;br&gt;
That is useful information too. If the projected token bill is €30 per month, spending weeks engineering around it probably makes little sense.&lt;/p&gt;
&lt;h2&gt;
  
  
  Step 3: Model the multipliers your prototype missed
&lt;/h2&gt;

&lt;p&gt;The simple calculation assumes every request looks roughly the same.&lt;br&gt;
Production applications rarely behave that way.&lt;/p&gt;

&lt;p&gt;Three factors in particular can make real token consumption substantially higher than the first estimate.&lt;/p&gt;
&lt;h3&gt;
  
  
  1. Conversation history grows with every turn
&lt;/h3&gt;

&lt;p&gt;Many chat implementations send previous messages back to the model so that it has the context required to answer the next message.&lt;br&gt;
That means later requests can contain much more input than earlier ones.&lt;br&gt;
Consider this simplified example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;conversation_input_tokens&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;system_tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;user_tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;answer_tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;turns&lt;/span&gt;
&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    Estimate total input tokens billed across a conversation
    when previous messages are included on each new turn.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="n"&gt;history&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;system_tokens&lt;/span&gt;

    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;turns&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;history&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;user_tokens&lt;/span&gt;
        &lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;history&lt;/span&gt;
        &lt;span class="n"&gt;history&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;answer_tokens&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;total&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;turns&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;turns&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nf"&gt;conversation_input_tokens&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="mi"&gt;400&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="mi"&gt;40&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="mi"&gt;80&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;turns&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1 440
3 1680
5 3400
10 9800
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzlvw9hrfxilrvwjax51d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzlvw9hrfxilrvwjax51d.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Ten turns therefore do not necessarily cost ten times as much input as the first turn. In this simplified example, total input consumption across the conversation is more than 22 times the first request's input.&lt;br&gt;
Why?&lt;/p&gt;

&lt;p&gt;Because previous messages are repeatedly included as context.&lt;br&gt;
Real implementations vary. Some APIs and providers offer prompt caching, and applications can summarise, truncate or selectively retrieve previous conversation history rather than continually resending everything.&lt;/p&gt;

&lt;p&gt;The important lesson is to estimate chat costs per conversation, not simply by multiplying the first message by the expected number of turns.&lt;/p&gt;
&lt;h3&gt;
  
  
  2. Language can materially change token counts
&lt;/h3&gt;

&lt;p&gt;Do not assume that a 100-word prompt has the same token count in every language.&lt;br&gt;
Tokenisation efficiency varies by language, writing system, vocabulary and tokenizer. The same underlying meaning can therefore require noticeably different numbers of tokens depending on both the language being used and the model's tokenizer.&lt;br&gt;
That matters commercially.&lt;/p&gt;

&lt;p&gt;If you benchmark an application using English prompts but most production users communicate in Serbian, Polish, German or another language, the English benchmark may not accurately represent production token consumption.&lt;br&gt;
The safest approach is not to apply a generic "non-English multiplier."&lt;/p&gt;

&lt;p&gt;Build a representative sample of actual user prompts in the languages your application supports and run each sample through the tokenizer of the models you are considering.&lt;br&gt;
For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;samples&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;english&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;My order arrived damaged. How can I replace it?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;serbian&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Moja porudžbina je stigla oštećena. Kako mogu da je zamenim?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;language&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;samples&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;language&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a multilingual product, repeat this with dozens or hundreds of representative prompts rather than relying on a single sentence.&lt;br&gt;
The resulting numbers are much more useful than assuming one language will always cost a fixed percentage more than another.&lt;/p&gt;
&lt;h3&gt;
  
  
  3. Context tends to grow after launch
&lt;/h3&gt;

&lt;p&gt;Prompt growth is the easiest of the three to overlook.&lt;br&gt;
The first system prompt contains five rules. Then a production problem appears and someone adds three more.&lt;br&gt;
Retrieval initially returns three document chunks. Later it returns five because answer quality improves.&lt;/p&gt;

&lt;p&gt;A few examples are added to handle a difficult edge case. Tool descriptions become longer. More metadata is inserted into every request.&lt;/p&gt;

&lt;p&gt;Each change looks small in isolation.&lt;br&gt;
But repeated input is multiplied by every request the application makes.&lt;/p&gt;

&lt;p&gt;Suppose you add 500 tokens of permanent instructions to an application handling 100,000 requests per month. That change alone creates:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;500 × 100,000 = 50,000,000 additional input tokens per month
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At scale, prompt design is therefore also infrastructure design.&lt;br&gt;
This is one reason the example calculator includes a safety margin. More importantly, token measurement should become part of the development process rather than a calculation performed once before launch.&lt;/p&gt;
&lt;h2&gt;
  
  
  Step 4: Turn the estimate into an engineering decision
&lt;/h2&gt;

&lt;p&gt;Once you know the approximate cost per request and per month, the numbers can guide architecture choices.&lt;/p&gt;

&lt;p&gt;Start with the smallest model that meets your quality requirements. Run the same evaluation set against several model tiers. There is little reason to pay for a larger model on every request if a smaller one reliably handles the workload.&lt;/p&gt;

&lt;p&gt;Route requests by complexity. Straightforward classification, extraction or support questions may work well on a smaller model, while difficult requests can be escalated to a more capable one.&lt;br&gt;
Reduce unnecessary input. System prompts, retrieval results and conversation history are all recurring costs. Removing irrelevant context can improve both economics and, in some cases, model performance.&lt;/p&gt;

&lt;p&gt;Test caching where the workload allows it. Repeated prefixes, instructions and context may qualify for discounted or cached processing depending on the provider.&lt;/p&gt;

&lt;p&gt;Measure actual production distributions. An average request is useful for planning, but averages hide expensive outliers. Track median and high-percentile token consumption as well as the mean.&lt;/p&gt;

&lt;p&gt;Compare metered inference with dedicated infrastructure only when demand is sufficiently predictable. Usage-based APIs are attractive when traffic is uncertain or variable. Dedicated GPU capacity becomes easier to evaluate once you know how much compute the application consistently consumes.&lt;/p&gt;
&lt;h2&gt;
  
  
  Build a small pricing table
&lt;/h2&gt;

&lt;p&gt;You do not need a sophisticated cost platform to make the initial decision.&lt;br&gt;
A simple table with model name, input price, output price, average input tokens, average output tokens and expected request volume is enough to compare several scenarios.&lt;br&gt;
Populate the &lt;code&gt;PRICES&lt;/code&gt; dictionary with current provider rates rather than estimates. Providers serving open-weight models commonly publish separate rates for input and output tokens; for example, current &lt;a href="https://orionfactory.ai/pay-as-you-go-inference.php" rel="noopener noreferrer"&gt;per-token prices for open-weight models&lt;/a&gt; can be used to populate the same calculator for models such as DeepSeek, GLM and Kimi.&lt;/p&gt;

&lt;p&gt;Then run several scenarios rather than one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;TRAFFIC_SCENARIOS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pilot&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;expected&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;2_000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;high growth&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;10_000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;scenario&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;requests_per_day&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;TRAFFIC_SCENARIOS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;scenario&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;upper&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;price&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;PRICES&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="n"&gt;per_req&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;cost_per_request&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;price&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;monthly&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;per_req&lt;/span&gt;
            &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;requests_per_day&lt;/span&gt;
            &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;DAYS_PER_MONTH&lt;/span&gt;
            &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;SAFETY_MARGIN&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;monthly&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="mf"&gt;10.2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That gives you something more useful than a single forecast: a range showing what happens if adoption is much lower or much higher than expected.&lt;br&gt;
Wrapping up&lt;br&gt;
Estimating an LLM inference bill does not require sophisticated financial modelling.&lt;/p&gt;

&lt;p&gt;You need realistic traffic, realistic prompts and current token prices.&lt;br&gt;
The basic process is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Build representative production requests, including system instructions, retrieved context and conversation history.&lt;/li&gt;
&lt;li&gt;Count input and output tokens using the tokenizer appropriate for the model.&lt;/li&gt;
&lt;li&gt;Apply the current input and output prices for each model you are considering.&lt;/li&gt;
&lt;li&gt;Model conversations, supported languages, retries and likely prompt growth.&lt;/li&gt;
&lt;li&gt;Run multiple traffic scenarios instead of relying on a single forecast.&lt;/li&gt;
&lt;li&gt;Measure actual token consumption after launch and update the model with production data.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The biggest forecasting mistake is usually not getting the price of a token wrong.&lt;/p&gt;

&lt;p&gt;It is counting only the tokens that are obvious.&lt;/p&gt;

&lt;p&gt;A user's ten-word question may arrive at the model wrapped in thousands of tokens of instructions, retrieved documents, tool definitions and conversation history. Once you measure the complete request rather than the visible message, LLM inference costs become much easier to predict, and much harder to be surprised by.&lt;/p&gt;

&lt;p&gt;I work with AI infrastructure providers, including the one linked above.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>tutorial</category>
      <category>python</category>
      <category>llm</category>
    </item>
    <item>
      <title>The hidden cost of cloud GPU training: egress, idle time, and lock-in</title>
      <dc:creator>Andrea Susic</dc:creator>
      <pubDate>Thu, 28 May 2026 20:20:53 +0000</pubDate>
      <link>https://dev.to/asymm/the-hidden-cost-of-cloud-gpu-training-egress-idle-time-and-lock-in-15f5</link>
      <guid>https://dev.to/asymm/the-hidden-cost-of-cloud-gpu-training-egress-idle-time-and-lock-in-15f5</guid>
      <description>&lt;p&gt;The GPU hourly rate is the number everyone compares. It is also the number that tells you the least about what a training run actually costs.&lt;/p&gt;

&lt;p&gt;The sticker price, say $2 to $3.50 an hour for an H100 on a specialized cloud, is the visible tip. The real bill is built from three things almost nobody puts on the comparison spreadsheet: the GPU sitting idle, the data you have to move, and the cost of ever leaving. This post breaks down each one, with 2026 numbers, and what you can actually do about it.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Idle time: paying full price for nothing
&lt;/h2&gt;

&lt;p&gt;The most expensive line item in most setups is not compute. It is compute you pay for but never use.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The 5 percent problem&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A 2026 Cast AI report found average GPU utilization across Kubernetes clusters on major clouds sits around 5 percent. Other analyses are kinder, Anyscale puts sustained production utilization below 50 percent, FinOps studies land at 20 to 30 percent, but the conclusion holds: most of every GPU-hour you pay for produces no useful work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it happens&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It is structural, not laziness. Workloads bounce between CPU preprocessing, GPU training, and CPU postprocessing. Python dataloaders on the GPU node starve the accelerator. Teams overprovision to dodge out-of-memory errors and default to the biggest instance "just in case."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why it hurts more than CPU waste&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;An idle CPU costs cents per hour. An idle GPU costs dollars per hour. A single AWS p4d.24xlarge left idle over one weekend burns about $1,573 for nothing. A month of overnight and weekend idling typically wastes $3,000 to $8,000 per instance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What to do&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Add idle detection.&lt;/strong&gt; A script watching &lt;code&gt;nvidia-smi&lt;/code&gt; that scales down an instance after utilization stays below ~5 percent for 30 minutes is the highest-ROI thing most teams can ship. Commonly cuts 20 to 35 percent off GPU spend.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Right-size the hardware.&lt;/strong&gt; Not every job needs an H100. Running on the biggest card when a smaller one delivers the same result is pure burn.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fix the pipeline first.&lt;/strong&gt; If dataloaders are starving the GPU, a bigger GPU does not help. Profile the input pipeline before upgrading hardware.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  2. Egress: the cost of moving your own data
&lt;/h2&gt;

&lt;p&gt;Uploading data is free everywhere. Moving it out is not, and the asymmetry is deliberate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The 2026 rates&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Verified across multiple pricing surveys: AWS charges about $0.09 per GB outbound, roughly $90 per TB. Google Cloud is higher at $0.12 per GB. Azure sits in the same range. Hetzner includes large free allowances and charges on the order of $1 per TB beyond them. Some object-storage options are zero egress entirely.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why training amplifies it&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The volumes are large and recurring. Datasets, checkpoints, and exported weights all move. On a workload pulling 10 TB out per month, the gap between a hyperscaler and a zero-egress provider is the difference between a four-figure line item and almost nothing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What to do&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Co-locate compute and storage.&lt;/strong&gt; Keep training data in the same region and provider as the GPUs. Intra-zone transfer is usually free; internet egress is not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compress before transfer.&lt;/strong&gt; gzip or zstd cuts checkpoint and dataset volume 30 to 60 percent, and egress is billed per byte.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Price egress into provider comparisons.&lt;/strong&gt; A cheap GPU hour on a provider with expensive egress can lose to a pricier hour on one with none.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  3. Lock-in: the bill you pay to leave
&lt;/h2&gt;

&lt;p&gt;The third cost is the one you do not see until you try to escape it. It is the same mechanism as egress, viewed over a longer horizon.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Data gravity is the anchor&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Once terabytes of data and checkpoints accumulate in one region, moving them is slow and expensive. The egress fee is not just a per-transfer cost, it is an exit tax that grows with every gigabyte you store. By design, the more your data piles up, the less likely you are to leave.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is not only data&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Proprietary services, custom tooling, and provider-specific orchestration all raise the cost of moving. But for training workloads, raw data gravity is the heaviest anchor.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What to do&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Prefer portable formats and open tooling.&lt;/strong&gt; Standard container images, open checkpoint formats, and provider-agnostic orchestration keep your options open.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model the exit cost up front.&lt;/strong&gt; Calculate what it would cost to move everything out at a year's expected data volume. If the number is alarming, factor it in now, not later.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Favor zero or low egress for data-heavy work.&lt;/strong&gt; When leaving is cheap, lock-in mostly evaporates, and you keep leverage over your own infrastructure.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Putting it together
&lt;/h2&gt;

&lt;p&gt;The headline GPU rate is the smallest part of the story. A realistic cost model looks more like this:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;(GPU rate x hours x utilization gap) + storage + egress + the eventual cost of leaving&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Two things follow. First, fixing utilization is usually the fastest win, because you are already paying for that waste today. Second, egress and lock-in are decisions you make once, at the start, that compound for as long as the project runs.&lt;/p&gt;

&lt;p&gt;This is why the provider landscape is shifting. Specialized GPU clouds and regional providers increasingly compete on exactly these hidden costs: transparent hourly billing instead of a maze of ancillary fees, and zero egress so your data, and your freedom to move it, stays yours. &lt;a href="https://orionfactory.ai/" rel="noopener noreferrer"&gt;Orion AI Factory&lt;/a&gt; in Europe is one example of that model, and the same logic shows up across a growing set of regional and specialized providers. The common thread is pricing the things that used to hide in the footnotes.&lt;/p&gt;

&lt;p&gt;None of this needs exotic tooling. Watch your utilization, keep data close to compute, compress what you move, and know your exit cost before you are locked in. The teams that win on cost are not the ones with the biggest budgets. They are the ones who read past the hourly rate.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;References&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Cast AI, GPU utilization report, 2026&lt;/li&gt;
&lt;li&gt;Anyscale, production GPU utilization analysis, January 2026&lt;/li&gt;
&lt;li&gt;GPUPerHour, data egress pricing across 44+ providers (&lt;a href="https://gpuperhour.com/reference/data-egress" rel="noopener noreferrer"&gt;https://gpuperhour.com/reference/data-egress&lt;/a&gt;), April 2026&lt;/li&gt;
&lt;li&gt;LeanOps, AI cloud cost optimization guide, 2026&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>database</category>
    </item>
  </channel>
</rss>
