DEV Community

Peyton Green
Peyton Green

Posted on

LLM Cost Optimization in Python: Cutting API Bills Without Cutting Quality

The first month of a new AI project is cheap. By month three, when real traffic hits and your prompts have grown to handle edge cases, the API bills start to look like a SaaS subscription you didn't sign up for.

This is mostly avoidable. Most LLM cost problems come from a small set of patterns: oversized prompts, unnecessary calls, not using caching, and picking the wrong model for the job. Each of these is fixable without touching output quality.

Here's what I've found that actually works.


Profile before you optimize

The first step is always measurement. You need to know where the tokens are going before you start cutting.

import anthropic
from dataclasses import dataclass, field
from typing import Any

@dataclass
class TokenTracker:
    """Track token usage across API calls."""
    calls: list[dict] = field(default_factory=list)

    def record(self, call_name: str, response: Any) -> None:
        usage = getattr(response, "usage", None)
        if usage:
            self.calls.append({
                "name": call_name,
                "input_tokens": usage.input_tokens,
                "output_tokens": usage.output_tokens,
                "total_tokens": usage.input_tokens + usage.output_tokens,
            })

    def summary(self) -> dict:
        if not self.calls:
            return {}
        total_input = sum(c["input_tokens"] for c in self.calls)
        total_output = sum(c["output_tokens"] for c in self.calls)
        by_call = sorted(
            self.calls,
            key=lambda c: c["total_tokens"],
            reverse=True,
        )
        return {
            "total_input_tokens": total_input,
            "total_output_tokens": total_output,
            "top_consumers": by_call[:5],
        }

tracker = TokenTracker()
client = anthropic.Anthropic()

def tracked_call(name: str, **kwargs):
    response = client.messages.create(**kwargs)
    tracker.record(name, response)
    return response
Enter fullscreen mode Exit fullscreen mode

Run this for a week. The summary usually reveals one or two calls that account for 60-70% of tokens. Optimize those first.


Prompt caching

Anthropic's prompt caching is the single highest-leverage optimization available. When your prompt contains a large static section — a system prompt, a long document, a schema — you can mark it as cacheable. Anthropic stores it for up to 5 minutes (or 1 hour with extended caching) and charges ~10% of normal input pricing for cache hits.

import anthropic

client = anthropic.Anthropic()

# A system prompt with a large static section
LARGE_SYSTEM_CONTEXT = """
You are a senior Python developer reviewing code for a financial services company.

Coding standards (applies to all reviews):
- All functions must have type annotations
- Error handling must be explicit — no bare except clauses
- All database queries must use parameterized inputs
- No logging of PII (customer IDs, account numbers, SSNs)
- All external API calls must have timeout configuration
- Test coverage must exceed 80% for new code
[... 2000 more tokens of standards ...]
"""

def review_code(code: str) -> str:
    response = client.messages.create(
        model="claude-sonnet-5",
        max_tokens=1024,
        system=[
            {
                "type": "text",
                "text": LARGE_SYSTEM_CONTEXT,
                "cache_control": {"type": "ephemeral"},  # Cache this block
            }
        ],
        messages=[
            {"role": "user", "content": f"Review this code:\n\n```
{% endraw %}
python\n{code}\n
{% raw %}
```"}
        ],
    )
    return response.content[0].text
Enter fullscreen mode Exit fullscreen mode

The cache_control: ephemeral flag tells Anthropic to cache everything before that point. On the first call, you pay normal pricing. On subsequent calls within the TTL, you pay ~10% for the cached tokens. For a 2,000-token system prompt with 100 calls per hour, this is roughly a 90% cost reduction on input tokens.

What to cache:

  • Large system prompts with detailed instructions
  • Reference documents (API schemas, codebases, knowledge bases)
  • Few-shot examples that don't change per-request

What not to cache:

  • User-specific context (personalization overrides the caching benefit)
  • Prompts that change frequently

Model routing

Not every task needs your best model. Routing to the right model size is often worth 3-5x cost reduction with no quality loss on the routed tasks.

from enum import Enum
from anthropic import Anthropic

class TaskComplexity(Enum):
    SIMPLE = "simple"   # Classification, extraction, yes/no
    MEDIUM = "medium"   # Summarization, reformatting, generation with constraints
    COMPLEX = "complex" # Reasoning, code review, multi-step analysis

MODEL_MAP = {
    TaskComplexity.SIMPLE: "claude-haiku-4-5-20251001",
    TaskComplexity.MEDIUM: "claude-sonnet-5",
    TaskComplexity.COMPLEX: "claude-sonnet-5",  # or claude-opus-5-5 when needed
}

client = Anthropic()

def classify_intent(user_message: str) -> str:
    """Simple classification — Haiku is more than enough."""
    response = client.messages.create(
        model=MODEL_MAP[TaskComplexity.SIMPLE],
        max_tokens=10,
        messages=[{
            "role": "user",
            "content": f"Classify this as 'question', 'request', or 'complaint': {user_message}"
        }]
    )
    return response.content[0].text.strip().lower()

def summarize_document(document: str) -> str:
    """Summarization — Sonnet handles this well."""
    response = client.messages.create(
        model=MODEL_MAP[TaskComplexity.MEDIUM],
        max_tokens=500,
        messages=[{
            "role": "user",
            "content": f"Summarize in 3 bullet points:\n\n{document}"
        }]
    )
    return response.content[0].text

def review_architecture(design_doc: str) -> str:
    """Architecture review — worth paying for the better model."""
    response = client.messages.create(
        model=MODEL_MAP[TaskComplexity.COMPLEX],
        max_tokens=2000,
        messages=[{
            "role": "user",
            "content": f"Review this architecture for scalability issues:\n\n{design_doc}"
        }]
    )
    return response.content[0].text
Enter fullscreen mode Exit fullscreen mode

The classification call in a customer support pipeline might run 10,000 times a day. Haiku vs. Sonnet is a 10x price difference. For a yes/no classification, the quality difference is immaterial.


Output compression

A common source of bloat: asking for more output than you need and then parsing it down. If you only need the extracted value, ask for only the value.

Instead of:

# Bad: asking for explanation when you only need the answer
prompt = "What is the sentiment of this review? Explain your reasoning."
# Gets 300 tokens of explanation + the answer
Enter fullscreen mode Exit fullscreen mode

Do:

# Good: ask only for what you'll use
prompt = "What is the sentiment? Reply with exactly one word: positive, negative, or neutral."
# Gets 1 token
Enter fullscreen mode Exit fullscreen mode

For structured extraction:

from pydantic import BaseModel
from anthropic import Anthropic
import json

class ExtractedFields(BaseModel):
    company_name: str
    contact_email: str
    urgency: str  # "low", "medium", "high"

client = Anthropic()

def extract_lead(email_text: str) -> ExtractedFields:
    response = client.messages.create(
        model="claude-haiku-4-5-20251001",
        max_tokens=100,  # Tight limit — structured output is short
        messages=[{
            "role": "user",
            "content": f"""Extract from this email. Return ONLY valid JSON, no other text:
{{"company_name": "...", "contact_email": "...", "urgency": "low|medium|high"}}

Email:
{email_text}"""
        }]
    )
    return ExtractedFields.model_validate_json(response.content[0].text)
Enter fullscreen mode Exit fullscreen mode

Setting max_tokens tight for structured extraction forces short outputs and creates a natural guard against runaway generation.


Response caching (semantic deduplication)

Some queries are effectively the same question asked in slightly different ways. If you're serving many users, you can cache responses by semantic similarity rather than exact string match.

import hashlib
import json
import time
from functools import lru_cache

# Simple exact-match cache (works for API endpoints with predictable inputs)
_cache: dict[str, tuple[str, float]] = {}
CACHE_TTL_SECONDS = 300

def cached_completion(
    prompt: str,
    model: str = "claude-sonnet-5",
    ttl: int = CACHE_TTL_SECONDS,
) -> str:
    cache_key = hashlib.sha256(f"{model}:{prompt}".encode()).hexdigest()
    now = time.time()

    if cache_key in _cache:
        cached_value, cached_at = _cache[cache_key]
        if now - cached_at < ttl:
            return cached_value  # Cache hit — no API call

    response = client.messages.create(
        model=model,
        max_tokens=1024,
        messages=[{"role": "user", "content": prompt}]
    )
    result = response.content[0].text
    _cache[cache_key] = (result, now)
    return result
Enter fullscreen mode Exit fullscreen mode

For production, use Redis instead of an in-process dict:

import redis
import hashlib
import json

r = redis.Redis(host="localhost", port=6379, decode_responses=True)

def cached_completion_redis(prompt: str, model: str, ttl: int = 300) -> str:
    cache_key = f"llm:{hashlib.sha256(f'{model}:{prompt}'.encode()).hexdigest()}"
    cached = r.get(cache_key)
    if cached:
        return cached

    response = client.messages.create(
        model=model,
        max_tokens=1024,
        messages=[{"role": "user", "content": prompt}]
    )
    result = response.content[0].text
    r.setex(cache_key, ttl, result)
    return result
Enter fullscreen mode Exit fullscreen mode

The TTL to use depends on how stale a response can be. For FAQ responses: hours. For stock price summaries: seconds. For code review: invalidate on file change.


Batching and async calls

If you're making multiple independent API calls sequentially, you're paying in time. Run them concurrently:

import asyncio
import anthropic

async_client = anthropic.AsyncAnthropic()

async def analyze_many(texts: list[str]) -> list[str]:
    """Analyze all texts concurrently."""
    tasks = [
        async_client.messages.create(
            model="claude-haiku-4-5-20251001",
            max_tokens=200,
            messages=[{"role": "user", "content": f"Summarize: {text}"}]
        )
        for text in texts
    ]
    responses = await asyncio.gather(*tasks)
    return [r.content[0].text for r in responses]

# Sequential (bad for cost/time at scale):
# results = [analyze(text) for text in texts]

# Concurrent (good):
results = asyncio.run(analyze_many(texts))
Enter fullscreen mode Exit fullscreen mode

For OpenAI, the Batch API is even more cost-effective for non-time-sensitive work — 50% off on inputs and outputs, with 24-hour turnaround:

import openai
import json

client = openai.OpenAI()

def create_batch_analysis(texts: list[str]) -> str:
    """Submit a batch of analyses at 50% cost."""
    requests = [
        {
            "custom_id": f"req-{i}",
            "method": "POST",
            "url": "/v1/chat/completions",
            "body": {
                "model": "gpt-4o-mini",
                "messages": [{"role": "user", "content": f"Summarize: {text}"}],
                "max_tokens": 200,
            }
        }
        for i, text in enumerate(texts)
    ]

    # Write requests to a JSONL file
    with open("batch_input.jsonl", "w") as f:
        for req in requests:
            f.write(json.dumps(req) + "\n")

    # Upload and submit
    with open("batch_input.jsonl", "rb") as f:
        batch_file = client.files.create(file=f, purpose="batch")

    batch = client.batches.create(
        input_file_id=batch_file.id,
        endpoint="/v1/chat/completions",
        completion_window="24h",
    )
    return batch.id  # Poll this later
Enter fullscreen mode Exit fullscreen mode

The Batch API is ideal for nightly report generation, data pipeline processing, or any workload where you're not waiting on a human.


Putting it together: a cost-aware client

Here's a wrapper that applies multiple optimizations transparently:

import anthropic
from dataclasses import dataclass, field
import hashlib
import time

@dataclass
class CostAwareClient:
    model_simple: str = "claude-haiku-4-5-20251001"
    model_standard: str = "claude-sonnet-5"
    cache_ttl: int = 300
    _cache: dict = field(default_factory=dict)
    _token_usage: list = field(default_factory=list)

    def __post_init__(self):
        self._client = anthropic.Anthropic()

    def complete(
        self,
        prompt: str,
        max_tokens: int = 1024,
        use_cache: bool = True,
        simple_task: bool = False,
        system: str | None = None,
    ) -> str:
        model = self.model_simple if simple_task else self.model_standard
        cache_key = hashlib.sha256(f"{model}:{system}:{prompt}".encode()).hexdigest()

        if use_cache and cache_key in self._cache:
            value, ts = self._cache[cache_key]
            if time.time() - ts < self.cache_ttl:
                return value

        kwargs = {
            "model": model,
            "max_tokens": max_tokens,
            "messages": [{"role": "user", "content": prompt}],
        }
        if system:
            kwargs["system"] = system

        response = self._client.messages.create(**kwargs)
        result = response.content[0].text
        self._token_usage.append({
            "model": model,
            "input": response.usage.input_tokens,
            "output": response.usage.output_tokens,
        })

        if use_cache:
            self._cache[cache_key] = (result, time.time())

        return result

    def usage_report(self) -> dict:
        if not self._token_usage:
            return {}
        return {
            "total_calls": len(self._token_usage),
            "total_input_tokens": sum(u["input"] for u in self._token_usage),
            "total_output_tokens": sum(u["output"] for u in self._token_usage),
            "by_model": {
                m: sum(u["input"] + u["output"] for u in self._token_usage if u["model"] == m)
                for m in set(u["model"] for u in self._token_usage)
            }
        }
Enter fullscreen mode Exit fullscreen mode

The numbers

On a typical project, applying these patterns gets you:

  • Prompt caching: 60-90% reduction on cached input tokens
  • Model routing: 5-10x reduction on routed tasks (simple → small model)
  • Output compression: 30-70% reduction on output tokens for structured extraction
  • Response caching: Highly workload-dependent — 40-80% for FAQ-style workloads
  • Batching/async: Primarily a latency win, but avoids rate limit retries

None of these are free — they add code complexity. The ones worth doing are the ones that hit your top 2-3 cost consumers from the profiling step. Start with those and stop when the cost is acceptable.


The prompts that work best with AI tools tend to be the ones that are specific, structured, and constrained. If you're spending time writing good prompts, the AI Dev Toolkit has 272 organized examples across code review, debugging, testing, and architecture — one-time purchase.

Top comments (0)