Most unexpected LLM bills start with a misleadingly cheap prototype.
Someone tests a feature with a handful of short prompts, checks the usage dashboard and concludes that inference costs almost nothing. Then the application reaches production. Prompts get longer, retrieval adds context, conversations stretch across multiple turns and thousands of users start making requests.
A month later, the invoice looks very different.
Fortunately, inference cost is relatively straightforward to model. If your provider charges by token, you can estimate the cost of a realistic request before launch and then model what happens as traffic grows.
The arithmetic is simple. Getting realistic inputs is the important part.
How LLM inference is usually billed
Many hosted LLM APIs charge separately for input and output tokens.
Input tokens include everything sent to the model: the system prompt, retrieved context, conversation history, tool information and the user's current message.
Output tokens are the tokens generated by the model.
Providers commonly quote both prices per million tokens, and output tokens are often more expensive than input tokens.
At its simplest, the calculation is:
cost = (input_tokens × input_price + output_tokens × output_price) / 1,000,000
The formula is rarely the source of forecasting errors.
The problem is estimating input_tokens.
A production request is usually much larger than the text the user sees in the chat box.
Step 1: Count a realistic request
Start with something that resembles the request your application will actually send.
If you are building a retrieval-augmented support assistant, for example, include the system instructions, retrieved documents and user message rather than measuring the user message alone.
Install tiktoken:
pip install tiktoken
Then create a representative request:
import tiktoken
# o200k_base is a modern BPE tokenizer.
# Different model families can use different tokenizers, so use the
# target model's tokenizer when you need model-specific counts.
enc = tiktoken.get_encoding("o200k_base")
def count(text: str) -> int:
return len(enc.encode(text))
SYSTEM_PROMPT = """You are a support assistant for an online store.
Answer only from the provided context. If the answer is not in the
context, say you don't know and offer to connect a human agent.
Keep answers under 120 words and use a friendly, plain tone."""
RETRIEVED_CONTEXT = """Returns: items can be returned within 30 days of delivery.
Refunds go back to the original payment method within 5 business days.
Shipping: standard delivery takes 2-4 business days. Express is next day if ordered before 14:00. Damaged items: send a photo to support and we replace the item at no cost."""
USER_MESSAGE = """Hi, my blender arrived with a cracked jug.
Can I get a new one, and how long will it take?"""
SAMPLE_ANSWER = """Sorry to hear the jug arrived cracked. Please reply with a photo
of the damage and your order number, and we'll send a replacement at no cost. Replacements ship with standard delivery, which takes 2-4 business days. If you'd prefer a refund instead, that's possible too, and it reaches your original payment method within 5 business days."""
input_tokens = (
count(SYSTEM_PROMPT)
+ count(RETRIEVED_CONTEXT)
+ count(USER_MESSAGE)
)
output_tokens = count(SAMPLE_ANSWER)
This is still an estimate. Chat APIs can add formatting and message-structure overhead, and different model families use different tokenizers.
For an open-weight model, use the tokenizer associated with the model when possible. With models distributed through Hugging Face, that might look like this:
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("<model-repo-id>")
def count_model_tokens(text: str) -> int:
return len(tok.encode(text))
The closer your test request is to production traffic, the more useful the estimate becomes.
Step 2: Turn tokens into money
Once you have representative token counts, pricing different models is straightforward.
The rates below are deliberately illustrative. Replace them with the current input and output prices from the providers you are evaluating.
# Illustrative prices per 1M tokens.
# Replace these with your provider's current rates.
PRICES = {
"small model": {"input": 0.20, "output": 0.80},
"mid model": {"input": 1.00, "output": 4.00},
"large model": {"input": 3.00, "output": 15.00},
}
REQUESTS_PER_DAY = 2_000
DAYS_PER_MONTH = 30
# Allow for retries, longer-than-average responses,
# prompt changes and other production variance.
SAFETY_MARGIN = 1.3
def cost_per_request(price):
return (
input_tokens * price["input"]
+ output_tokens * price["output"]
) / 1_000_000
print(f"Input tokens per request: {input_tokens}")
print(f"Output tokens per request: {output_tokens}\n")
print(f"{'model':<12} {'per request':>12} {'per month':>12}")
for name, price in PRICES.items():
per_req = cost_per_request(price)
monthly = (
per_req
* REQUESTS_PER_DAY
* DAYS_PER_MONTH
* SAFETY_MARGIN
)
print(f"{name:<12} {per_req:>12.6f} {monthly:>12.2f}")
Using the example token counts, you might get something like:
Input tokens per request: 149
Output tokens per request: 74
model per request per month
small model 0.000089 6.94
mid model 0.000445 34.71
large model 0.001557 121.45
The exact figures are not important because both tokenisation and provider pricing will vary.
The comparison is what matters.
The same application and traffic assumptions can produce substantially different monthly costs depending on model choice. At the same time, the calculation may reveal that inference is a relatively small part of the application's total operating cost.
That is useful information too. If the projected token bill is €30 per month, spending weeks engineering around it probably makes little sense.
Step 3: Model the multipliers your prototype missed
The simple calculation assumes every request looks roughly the same.
Production applications rarely behave that way.
Three factors in particular can make real token consumption substantially higher than the first estimate.
1. Conversation history grows with every turn
Many chat implementations send previous messages back to the model so that it has the context required to answer the next message.
That means later requests can contain much more input than earlier ones.
Consider this simplified example:
def conversation_input_tokens(
system_tokens,
user_tokens,
answer_tokens,
turns
):
"""
Estimate total input tokens billed across a conversation
when previous messages are included on each new turn.
"""
total = 0
history = system_tokens
for _ in range(turns):
history += user_tokens
total += history
history += answer_tokens
return total
for turns in (1, 3, 5, 10):
print(
turns,
conversation_input_tokens(
400,
40,
80,
turns
)
)
Output:
1 440
3 1680
5 3400
10 9800
Ten turns therefore do not necessarily cost ten times as much input as the first turn. In this simplified example, total input consumption across the conversation is more than 22 times the first request's input.
Why?
Because previous messages are repeatedly included as context.
Real implementations vary. Some APIs and providers offer prompt caching, and applications can summarise, truncate or selectively retrieve previous conversation history rather than continually resending everything.
The important lesson is to estimate chat costs per conversation, not simply by multiplying the first message by the expected number of turns.
2. Language can materially change token counts
Do not assume that a 100-word prompt has the same token count in every language.
Tokenisation efficiency varies by language, writing system, vocabulary and tokenizer. The same underlying meaning can therefore require noticeably different numbers of tokens depending on both the language being used and the model's tokenizer.
That matters commercially.
If you benchmark an application using English prompts but most production users communicate in Serbian, Polish, German or another language, the English benchmark may not accurately represent production token consumption.
The safest approach is not to apply a generic "non-English multiplier."
Build a representative sample of actual user prompts in the languages your application supports and run each sample through the tokenizer of the models you are considering.
For example:
samples = {
"english": "My order arrived damaged. How can I replace it?",
"serbian": "Moja porudžbina je stigla oštećena. Kako mogu da je zamenim?",
}
for language, text in samples.items():
print(language, count(text))
For a multilingual product, repeat this with dozens or hundreds of representative prompts rather than relying on a single sentence.
The resulting numbers are much more useful than assuming one language will always cost a fixed percentage more than another.
3. Context tends to grow after launch
Prompt growth is the easiest of the three to overlook.
The first system prompt contains five rules. Then a production problem appears and someone adds three more.
Retrieval initially returns three document chunks. Later it returns five because answer quality improves.
A few examples are added to handle a difficult edge case. Tool descriptions become longer. More metadata is inserted into every request.
Each change looks small in isolation.
But repeated input is multiplied by every request the application makes.
Suppose you add 500 tokens of permanent instructions to an application handling 100,000 requests per month. That change alone creates:
500 × 100,000 = 50,000,000 additional input tokens per month
At scale, prompt design is therefore also infrastructure design.
This is one reason the example calculator includes a safety margin. More importantly, token measurement should become part of the development process rather than a calculation performed once before launch.
Step 4: Turn the estimate into an engineering decision
Once you know the approximate cost per request and per month, the numbers can guide architecture choices.
Start with the smallest model that meets your quality requirements. Run the same evaluation set against several model tiers. There is little reason to pay for a larger model on every request if a smaller one reliably handles the workload.
Route requests by complexity. Straightforward classification, extraction or support questions may work well on a smaller model, while difficult requests can be escalated to a more capable one.
Reduce unnecessary input. System prompts, retrieval results and conversation history are all recurring costs. Removing irrelevant context can improve both economics and, in some cases, model performance.
Test caching where the workload allows it. Repeated prefixes, instructions and context may qualify for discounted or cached processing depending on the provider.
Measure actual production distributions. An average request is useful for planning, but averages hide expensive outliers. Track median and high-percentile token consumption as well as the mean.
Compare metered inference with dedicated infrastructure only when demand is sufficiently predictable. Usage-based APIs are attractive when traffic is uncertain or variable. Dedicated GPU capacity becomes easier to evaluate once you know how much compute the application consistently consumes.
Build a small pricing table
You do not need a sophisticated cost platform to make the initial decision.
A simple table with model name, input price, output price, average input tokens, average output tokens and expected request volume is enough to compare several scenarios.
Populate the PRICES dictionary with current provider rates rather than estimates. Providers serving open-weight models commonly publish separate rates for input and output tokens; for example, current per-token prices for open-weight models can be used to populate the same calculator for models such as DeepSeek, GLM and Kimi.
Then run several scenarios rather than one:
TRAFFIC_SCENARIOS = {
"pilot": 200,
"expected": 2_000,
"high growth": 10_000,
}
for scenario, requests_per_day in TRAFFIC_SCENARIOS.items():
print(f"\n{scenario.upper()}")
for name, price in PRICES.items():
per_req = cost_per_request(price)
monthly = (
per_req
* requests_per_day
* DAYS_PER_MONTH
* SAFETY_MARGIN
)
print(f"{name:<12} {monthly:>10.2f}")
That gives you something more useful than a single forecast: a range showing what happens if adoption is much lower or much higher than expected.
Wrapping up
Estimating an LLM inference bill does not require sophisticated financial modelling.
You need realistic traffic, realistic prompts and current token prices.
The basic process is:
- Build representative production requests, including system instructions, retrieved context and conversation history.
- Count input and output tokens using the tokenizer appropriate for the model.
- Apply the current input and output prices for each model you are considering.
- Model conversations, supported languages, retries and likely prompt growth.
- Run multiple traffic scenarios instead of relying on a single forecast.
- Measure actual token consumption after launch and update the model with production data.
The biggest forecasting mistake is usually not getting the price of a token wrong.
It is counting only the tokens that are obvious.
A user's ten-word question may arrive at the model wrapped in thousands of tokens of instructions, retrieved documents, tool definitions and conversation history. Once you measure the complete request rather than the visible message, LLM inference costs become much easier to predict, and much harder to be surprised by.
I work with AI infrastructure providers, including the one linked above.


Top comments (0)