Why token prices are only a small part of the cost of running AI in production
When an AI feature becomes expensive, the model is an obvious place to look.
But as software engineers, we should ask a more uncomfortable question first: how much of the work are we asking the model to do unnecessarily?
A request may trigger several model calls, carry thousands of tokens of irrelevant context, and repeat work after a failure. Replacing the model with a cheaper one might reduce the bill, but it leaves those architectural decisions untouched.
Before optimizing the price of a call, I would examine the workflow that created it.
1. The Model Price Is Only the Beginning
Model pricing matters, but it does not explain the entire cost of an AI feature.
A production workflow may involve:
Input and output tokens sent to a model.
Multiple model calls for a single user request.
Document retrieval, search, and reranking.
Tool execution and external API calls.
Retries, fallbacks, and additional validation.
Application infrastructure, storage, and monitoring.
Some costs appear directly on a model provider’s invoice. Others belong to the surrounding infrastructure. Either way, looking only at token prices can hide the architectural decisions driving the total.
Consider a customer-support assistant that answers questions about invoices.
A simple implementation retrieves relevant invoice information and makes one model call. A more elaborate implementation classifies the request, summarizes the conversation, searches multiple data sources, asks a model to plan an action, calls a tool, and invokes another model to produce the final answer.
Both implementations may solve the same problem. Their cost structures can be very different.
A useful way to think about the total is:
Total Workflow Cost
=
Model Calls
+
Retrieval and Search
+
Tool Execution
+
Retries and Recovery
+
Allocated Infrastructure
The first engineering task is to understand which of these components actually contributes to the bill.
2. Context Is Not Free
One of the easiest ways to increase AI costs is to keep adding information to the prompt.
A feature starts with a system instruction and the user’s question. Later, the team adds conversation history, customer details, internal documentation, previous tool results, policy documents, and examples.
Each addition seems reasonable in isolation. Together, they can create a large context that is repeatedly processed even when much of it is irrelevant.
Imagine an employee asking:
Why is the March invoice still open?
The application may need the invoice status, payment history, and the latest relevant support interaction.
It probably does not need every invoice belonging to the customer, the entire support history, and every internal policy document related to billing.
Yet a naive retrieval implementation might send all of them to the model.
The solution is to distinguish useful context from merely available context.
Retrieve only information relevant to the current question.
Limit conversation history to what the current task needs.
Summarize older history when a full transcript is unnecessary.
Avoid repeating stable instructions and reference material when caching can reuse them.
Set explicit limits on retrieved documents and output length.
Measure whether additional context improves the answer enough to justify its cost.
Removing a critical policy or a necessary piece of evidence can increase mistakes, retries, and human intervention.
Keep the context that improves the result, and remove the context that merely adds volume.
Caching can help, but it cannot fix poor context design
When a workflow repeatedly sends the same instructions or reference material, prompt caching can reduce the cost of processing repeated input.
Providers implement caching differently, and eligibility, cache lifetime, pricing, and supported models vary. The OpenAI API pricing documentation distinguishes standard input tokens from cached input tokens. Google also documents context caching in its Vertex AI documentation.
Caching is useful when repeated context is necessary. But if an application sends thousands of irrelevant tokens on every request, making those tokens cheaper does not make them useful.
First reduce unnecessary context. Then cache the information that genuinely needs to be reused.
3. Count the Calls, Not Just the Tokens
A single user request does not necessarily mean a single model call.
Consider this workflow:
User Request
↓
Classify Intent
↓
Retrieve Documents
↓
Plan Next Action
↓
Call Tool
↓
Interpret Tool Result
↓
Generate Final Answer
Depending on the implementation, several of these stages may invoke a model.
Some workflows genuinely need that separation. Complex tasks involving ambiguous instructions, multiple data sources, or decisions that depend on earlier results may benefit from additional reasoning steps.
Other workflows accumulate model calls because every new requirement becomes another AI step.
An application might ask a model to classify a request even though deterministic rules can handle the supported categories. It might ask another model to extract a customer identifier already available in the authenticated application context. It might summarize a short tool response before passing it to a final generation step.
Each call adds cost and latency. It also creates another opportunity for an incorrect interpretation.
For every model call, ask:
What specific responsibility does this call have?
What does it add that the previous step cannot provide?
What changes if this call is removed?
If removing a call does not meaningfully change the result, it may not belong in the workflow.
This does not mean putting every responsibility into one enormous prompt. Combining unrelated responsibilities can make a system harder to test, constrain, and debug.
The objective is to remove unnecessary calls while preserving the boundaries that provide real value.
Every model call should earn its place in the architecture.
4. Not Every Request Needs the Same Model
Extracting a date from a well-defined sentence is not the same problem as interpreting an ambiguous customer complaint. Classifying a request into a small set of known categories is not the same as analyzing conflicting evidence across several documents.
Using the same high-capability model for every task may simplify implementation, but it can also be unnecessarily expensive.
A more deliberate design routes work according to its requirements.
Incoming Request
|
Determine Task Type
|
+-----------+-----------+
| |
Well-defined Complex or
task ambiguous task
| |
Smaller model or More capable
deterministic logic model
| |
+-----------+-----------+
|
Validate Result
|
Continue Workflow
A smaller model is a useful replacement only if it meets the task’s quality requirements.
If it produces more invalid outputs, triggers more retries, or requires frequent escalation, its lower token price may not translate into lower operating costs.
The same applies to fallback strategies. Sending every failure to a larger model can hide a weak first-stage design. Sometimes the right response is to improve the input, correct a validation problem, or ask the user for clarification rather than invoke another model.
Define task categories, evaluate candidate models against representative examples, and route only the categories where the less expensive option meets the required quality and latency thresholds.
The decision should be based on observed task performance, not model size alone.
5. Calculate the Difference: A Worked Example
Let’s make the cost difference concrete.
Suppose an invoice assistant processes a request using six model calls and consumes 12,000 tokens in total. After reviewing the workflow, the team removes redundant classification and summarization steps, reduces irrelevant context, and retains the calls needed for interpretation and response generation.
The optimized version uses two model calls and 3,000 tokens.
These figures are illustrative assumptions, not measurements from a production system. For simplicity, assume both versions use the same model and that their combined input and output tokens have an average effective price of $2 per million tokens. Actual input and output prices usually differ, and cached tokens may have separate pricing.
The calculations are:
Initial:
12,000 / 1,000,000 × $2 = $0.024
Optimized:
3,000 / 1,000,000 × $2 = $0.006
Under these assumptions, the optimized design reduces model-call count by 66.7%, token consumption by 75%, and model-token cost by 75%.
But the model bill is only part of the calculation.
The Cost of Being Wrong
Now suppose the initial design answers correctly 98% of the time, while the optimized version answers correctly only 94% of the time. Assume each incorrect answer requires $2 worth of employee correction.
The expected cost per request becomes:
Initial:
$0.024 + (2% × $2.00) = $0.064
Optimized:
$0.006 + (6% × $2.00) = $0.126
The optimized version has reduced model-token costs by 75%, yet its expected total cost per request is nearly twice as high.
This is why token savings alone are a poor measure of AI efficiency. A cheaper workflow can create more expensive work elsewhere in the business.
The calculation assumes that each incorrect answer incurs an additional $2 correction cost and that the stated accuracy rates hold. Real systems would need measured error rates and correction costs to establish their actual economics.
6. Optimize Without Breaking the System
Once the workflow is understood, several optimizations become easier to evaluate.
Reduce unnecessary output
If an application needs a status and a short explanation, asking for a long narrative adds little value.
Use structured output where appropriate, define clear response limits, and avoid generating information that downstream components will discard.
However, do not truncate output blindly. An incomplete structured object or missing evidence can trigger failures and additional work.
Make retries deliberate
A timeout, a rate-limit response, malformed JSON, and a semantically incorrect answer are different failure modes. They require different handling.
Use bounded retries, appropriate backoff, and failure-specific recovery strategies. Do not repeatedly send the same request when the underlying problem is invalid input or a deterministic validation failure.
For operations with external side effects, retry behavior must also account for idempotency. Repeating an operation should not accidentally send the same invoice or create the same business transaction twice.
Cache only equivalent work
Caching can avoid repeated computation when inputs and the meaning of the result remain equivalent.
A cached answer about a public technical concept may be reusable for many users. A cached answer about an invoice, account balance, or customer permission may become incorrect or expose information if identity, authorization, or underlying data changes.
Cache keys, invalidation rules, and access-control boundaries are part of the design.
Use asynchronous processing when possible
Some tasks do not require an immediate response. Document classification, scheduled summarization, and bulk extraction may be suitable for batch processing.
Provider offerings differ, but the OpenAI API pricing documentation describes separate pricing options for supported batch workloads.
The broader engineering principle is to match execution mode to business requirements rather than pay for immediate responsiveness when it is not needed.
Measure every change
Before changing a production workflow, establish a baseline for cost, latency, and task quality. Change one major factor at a time where practical, then compare the new behavior against representative requests.
A cost reduction is not a success if it pushes quality below the required threshold or increases the risk of unauthorized actions.
For important workflows, define acceptance criteria in advance. A drafting assistant and a system that sends invoices or updates customer records should not have the same tolerance for failure.
7. Measure the Cost of a Successful Outcome
The invoice assistant example shows why the cost of a successful outcome matters more than the price of an individual request. The model bill is only one part of the calculation; the cost of correcting failures can outweigh the savings.
A dashboard showing monthly model spend can tell you that costs have increased. It cannot, by itself, tell you whether the system has become less efficient.
Perhaps usage doubled because twice as many customers are using the feature. Perhaps a recent change increased the number of calls per request. Perhaps a cheaper model reduced the price of each call but also reduced the percentage of tasks completed successfully.
These scenarios require different responses.
For that reason, AI efficiency should be evaluated using the cost of a successfully completed task.
Cost per Successful Task
=
Total Cost of All Attempts
/
Number of Successful Tasks
The numerator should include the relevant costs across the complete workflow: model calls, retries, retrieval, tool use, and allocated infrastructure where those costs can be measured reliably.
The denominator needs a meaningful definition of success.
For an invoice assistant, generating a plausible answer is not necessarily success. The answer may need to be correct, grounded in authorized data, compliant with business rules, and useful enough for the employee to complete the task.
A workflow that produces fluent responses but frequently requires correction may be less efficient than a simpler one. Likewise, a workflow that returns an answer quickly but violates a permission boundary is not a successful outcome.
Cost per successful task complements other operational metrics such as cost per request, latency, throughput, error rate, and quality. It connects them to the result the application is supposed to deliver.
Make the workflow measurable
For each execution, collect enough information to understand where the cost came from.
A useful trace may include:
Workflow and task category.
Model and provider used at each step.
Input, output, and cached token usage where available.
Number of model calls and retries.
Retrieval and tool execution costs where measurable.
Latency by stage.
Validation outcomes, escalations, and final task status.
Connect these records using a workflow or request identifier. Otherwise, the model provider’s usage report and the application’s quality metrics remain separate views of the system.
Be careful with sensitive data in traces. Measuring a workflow does not require indiscriminately storing complete prompts, private documents, or customer records. Capture what is needed for diagnosis while applying suitable access controls, retention policies, and redaction.
The purpose of observability is to make the cost and outcome of a workflow explainable.
Conclusion
An AI feature can be expensive because the model is expensive. But it can also be expensive because the application sends too much information, makes too many calls, repeats failed work, chooses the wrong execution strategy, or measures activity instead of outcomes.
Those are architecture decisions.
Choosing a model is part of building an AI system, but it is not the whole job. The surrounding software determines when the model is called, what it receives, what happens after it responds, and how much additional work is required before the task is complete.
The goal is not to buy the cheapest tokens. It is to build the simplest reliable workflow that delivers the required result at a sustainable cost.

Top comments (0)