The first month of a new AI project is cheap. By month three, when real traffic hits and your prompts have grown to handle edge cases, the API bills start to look like a SaaS subscription you didn't sign up for.
This is mostly avoidable. Most LLM cost problems come from a small set of patterns: oversized prompts, unnecessary calls, not using caching, and picking the wrong model for the job. Each of these is fixable without touching output quality.
Here's what I've found that actually works.
Profile before you optimize
The first step is always measurement. You need to know where the tokens are going before you start cutting.
import anthropic
from dataclasses import dataclass, field
from typing import Any
@dataclass
class TokenTracker:
"""Track token usage across API calls."""
calls: list[dict] = field(default_factory=list)
def record(self, call_name: str, response: Any) -> None:
usage = getattr(response, "usage", None)
if usage:
self.calls.append({
"name": call_name,
"input_tokens": usage.input_tokens,
"output_tokens": usage.output_tokens,
"total_tokens": usage.input_tokens + usage.output_tokens,
})
def summary(self) -> dict:
if not self.calls:
return {}
total_input = sum(c["input_tokens"] for c in self.calls)
total_output = sum(c["output_tokens"] for c in self.calls)
by_call = sorted(
self.calls,
key=lambda c: c["total_tokens"],
reverse=True,
)
return {
"total_input_tokens": total_input,
"total_output_tokens": total_output,
"top_consumers": by_call[:5],
}
tracker = TokenTracker()
client = anthropic.Anthropic()
def tracked_call(name: str, **kwargs):
response = client.messages.create(**kwargs)
tracker.record(name, response)
return response
Run this for a week. The summary usually reveals one or two calls that account for 60-70% of tokens. Optimize those first.
Prompt caching
Anthropic's prompt caching is the single highest-leverage optimization available. When your prompt contains a large static section — a system prompt, a long document, a schema — you can mark it as cacheable. Anthropic stores it for up to 5 minutes (or 1 hour with extended caching) and charges ~10% of normal input pricing for cache hits.
import anthropic
client = anthropic.Anthropic()
# A system prompt with a large static section
LARGE_SYSTEM_CONTEXT = """
You are a senior Python developer reviewing code for a financial services company.
Coding standards (applies to all reviews):
- All functions must have type annotations
- Error handling must be explicit — no bare except clauses
- All database queries must use parameterized inputs
- No logging of PII (customer IDs, account numbers, SSNs)
- All external API calls must have timeout configuration
- Test coverage must exceed 80% for new code
[... 2000 more tokens of standards ...]
"""
def review_code(code: str) -> str:
response = client.messages.create(
model="claude-sonnet-5",
max_tokens=1024,
system=[
{
"type": "text",
"text": LARGE_SYSTEM_CONTEXT,
"cache_control": {"type": "ephemeral"}, # Cache this block
}
],
messages=[
{"role": "user", "content": f"Review this code:\n\n```
{% endraw %}
python\n{code}\n
{% raw %}
```"}
],
)
return response.content[0].text
The cache_control: ephemeral flag tells Anthropic to cache everything before that point. On the first call, you pay normal pricing. On subsequent calls within the TTL, you pay ~10% for the cached tokens. For a 2,000-token system prompt with 100 calls per hour, this is roughly a 90% cost reduction on input tokens.
What to cache:
- Large system prompts with detailed instructions
- Reference documents (API schemas, codebases, knowledge bases)
- Few-shot examples that don't change per-request
What not to cache:
- User-specific context (personalization overrides the caching benefit)
- Prompts that change frequently
Model routing
Not every task needs your best model. Routing to the right model size is often worth 3-5x cost reduction with no quality loss on the routed tasks.
from enum import Enum
from anthropic import Anthropic
class TaskComplexity(Enum):
SIMPLE = "simple" # Classification, extraction, yes/no
MEDIUM = "medium" # Summarization, reformatting, generation with constraints
COMPLEX = "complex" # Reasoning, code review, multi-step analysis
MODEL_MAP = {
TaskComplexity.SIMPLE: "claude-haiku-4-5-20251001",
TaskComplexity.MEDIUM: "claude-sonnet-5",
TaskComplexity.COMPLEX: "claude-sonnet-5", # or claude-opus-5-5 when needed
}
client = Anthropic()
def classify_intent(user_message: str) -> str:
"""Simple classification — Haiku is more than enough."""
response = client.messages.create(
model=MODEL_MAP[TaskComplexity.SIMPLE],
max_tokens=10,
messages=[{
"role": "user",
"content": f"Classify this as 'question', 'request', or 'complaint': {user_message}"
}]
)
return response.content[0].text.strip().lower()
def summarize_document(document: str) -> str:
"""Summarization — Sonnet handles this well."""
response = client.messages.create(
model=MODEL_MAP[TaskComplexity.MEDIUM],
max_tokens=500,
messages=[{
"role": "user",
"content": f"Summarize in 3 bullet points:\n\n{document}"
}]
)
return response.content[0].text
def review_architecture(design_doc: str) -> str:
"""Architecture review — worth paying for the better model."""
response = client.messages.create(
model=MODEL_MAP[TaskComplexity.COMPLEX],
max_tokens=2000,
messages=[{
"role": "user",
"content": f"Review this architecture for scalability issues:\n\n{design_doc}"
}]
)
return response.content[0].text
The classification call in a customer support pipeline might run 10,000 times a day. Haiku vs. Sonnet is a 10x price difference. For a yes/no classification, the quality difference is immaterial.
Output compression
A common source of bloat: asking for more output than you need and then parsing it down. If you only need the extracted value, ask for only the value.
Instead of:
# Bad: asking for explanation when you only need the answer
prompt = "What is the sentiment of this review? Explain your reasoning."
# Gets 300 tokens of explanation + the answer
Do:
# Good: ask only for what you'll use
prompt = "What is the sentiment? Reply with exactly one word: positive, negative, or neutral."
# Gets 1 token
For structured extraction:
from pydantic import BaseModel
from anthropic import Anthropic
import json
class ExtractedFields(BaseModel):
company_name: str
contact_email: str
urgency: str # "low", "medium", "high"
client = Anthropic()
def extract_lead(email_text: str) -> ExtractedFields:
response = client.messages.create(
model="claude-haiku-4-5-20251001",
max_tokens=100, # Tight limit — structured output is short
messages=[{
"role": "user",
"content": f"""Extract from this email. Return ONLY valid JSON, no other text:
{{"company_name": "...", "contact_email": "...", "urgency": "low|medium|high"}}
Email:
{email_text}"""
}]
)
return ExtractedFields.model_validate_json(response.content[0].text)
Setting max_tokens tight for structured extraction forces short outputs and creates a natural guard against runaway generation.
Response caching (semantic deduplication)
Some queries are effectively the same question asked in slightly different ways. If you're serving many users, you can cache responses by semantic similarity rather than exact string match.
import hashlib
import json
import time
from functools import lru_cache
# Simple exact-match cache (works for API endpoints with predictable inputs)
_cache: dict[str, tuple[str, float]] = {}
CACHE_TTL_SECONDS = 300
def cached_completion(
prompt: str,
model: str = "claude-sonnet-5",
ttl: int = CACHE_TTL_SECONDS,
) -> str:
cache_key = hashlib.sha256(f"{model}:{prompt}".encode()).hexdigest()
now = time.time()
if cache_key in _cache:
cached_value, cached_at = _cache[cache_key]
if now - cached_at < ttl:
return cached_value # Cache hit — no API call
response = client.messages.create(
model=model,
max_tokens=1024,
messages=[{"role": "user", "content": prompt}]
)
result = response.content[0].text
_cache[cache_key] = (result, now)
return result
For production, use Redis instead of an in-process dict:
import redis
import hashlib
import json
r = redis.Redis(host="localhost", port=6379, decode_responses=True)
def cached_completion_redis(prompt: str, model: str, ttl: int = 300) -> str:
cache_key = f"llm:{hashlib.sha256(f'{model}:{prompt}'.encode()).hexdigest()}"
cached = r.get(cache_key)
if cached:
return cached
response = client.messages.create(
model=model,
max_tokens=1024,
messages=[{"role": "user", "content": prompt}]
)
result = response.content[0].text
r.setex(cache_key, ttl, result)
return result
The TTL to use depends on how stale a response can be. For FAQ responses: hours. For stock price summaries: seconds. For code review: invalidate on file change.
Batching and async calls
If you're making multiple independent API calls sequentially, you're paying in time. Run them concurrently:
import asyncio
import anthropic
async_client = anthropic.AsyncAnthropic()
async def analyze_many(texts: list[str]) -> list[str]:
"""Analyze all texts concurrently."""
tasks = [
async_client.messages.create(
model="claude-haiku-4-5-20251001",
max_tokens=200,
messages=[{"role": "user", "content": f"Summarize: {text}"}]
)
for text in texts
]
responses = await asyncio.gather(*tasks)
return [r.content[0].text for r in responses]
# Sequential (bad for cost/time at scale):
# results = [analyze(text) for text in texts]
# Concurrent (good):
results = asyncio.run(analyze_many(texts))
For OpenAI, the Batch API is even more cost-effective for non-time-sensitive work — 50% off on inputs and outputs, with 24-hour turnaround:
import openai
import json
client = openai.OpenAI()
def create_batch_analysis(texts: list[str]) -> str:
"""Submit a batch of analyses at 50% cost."""
requests = [
{
"custom_id": f"req-{i}",
"method": "POST",
"url": "/v1/chat/completions",
"body": {
"model": "gpt-4o-mini",
"messages": [{"role": "user", "content": f"Summarize: {text}"}],
"max_tokens": 200,
}
}
for i, text in enumerate(texts)
]
# Write requests to a JSONL file
with open("batch_input.jsonl", "w") as f:
for req in requests:
f.write(json.dumps(req) + "\n")
# Upload and submit
with open("batch_input.jsonl", "rb") as f:
batch_file = client.files.create(file=f, purpose="batch")
batch = client.batches.create(
input_file_id=batch_file.id,
endpoint="/v1/chat/completions",
completion_window="24h",
)
return batch.id # Poll this later
The Batch API is ideal for nightly report generation, data pipeline processing, or any workload where you're not waiting on a human.
Putting it together: a cost-aware client
Here's a wrapper that applies multiple optimizations transparently:
import anthropic
from dataclasses import dataclass, field
import hashlib
import time
@dataclass
class CostAwareClient:
model_simple: str = "claude-haiku-4-5-20251001"
model_standard: str = "claude-sonnet-5"
cache_ttl: int = 300
_cache: dict = field(default_factory=dict)
_token_usage: list = field(default_factory=list)
def __post_init__(self):
self._client = anthropic.Anthropic()
def complete(
self,
prompt: str,
max_tokens: int = 1024,
use_cache: bool = True,
simple_task: bool = False,
system: str | None = None,
) -> str:
model = self.model_simple if simple_task else self.model_standard
cache_key = hashlib.sha256(f"{model}:{system}:{prompt}".encode()).hexdigest()
if use_cache and cache_key in self._cache:
value, ts = self._cache[cache_key]
if time.time() - ts < self.cache_ttl:
return value
kwargs = {
"model": model,
"max_tokens": max_tokens,
"messages": [{"role": "user", "content": prompt}],
}
if system:
kwargs["system"] = system
response = self._client.messages.create(**kwargs)
result = response.content[0].text
self._token_usage.append({
"model": model,
"input": response.usage.input_tokens,
"output": response.usage.output_tokens,
})
if use_cache:
self._cache[cache_key] = (result, time.time())
return result
def usage_report(self) -> dict:
if not self._token_usage:
return {}
return {
"total_calls": len(self._token_usage),
"total_input_tokens": sum(u["input"] for u in self._token_usage),
"total_output_tokens": sum(u["output"] for u in self._token_usage),
"by_model": {
m: sum(u["input"] + u["output"] for u in self._token_usage if u["model"] == m)
for m in set(u["model"] for u in self._token_usage)
}
}
The numbers
On a typical project, applying these patterns gets you:
- Prompt caching: 60-90% reduction on cached input tokens
- Model routing: 5-10x reduction on routed tasks (simple → small model)
- Output compression: 30-70% reduction on output tokens for structured extraction
- Response caching: Highly workload-dependent — 40-80% for FAQ-style workloads
- Batching/async: Primarily a latency win, but avoids rate limit retries
None of these are free — they add code complexity. The ones worth doing are the ones that hit your top 2-3 cost consumers from the profiling step. Start with those and stop when the cost is acceptable.
The prompts that work best with AI tools tend to be the ones that are specific, structured, and constrained. If you're spending time writing good prompts, the AI Dev Toolkit has 272 organized examples across code review, debugging, testing, and architecture — one-time purchase.
Top comments (0)