DEV Community

Cover image for How I Cut My LLM API Costs by 70% Without Touching My Code
Shaw Sha
Shaw Sha

Posted on

How I Cut My LLM API Costs by 70% Without Touching My Code

I was spending $200 a month on AI APIs. Now I'm down to $60, and my application works exactly the same. Same latency, same response quality, same user experience. The only thing that changed was how I route my requests.

Let me walk you through what I did, because it took me about three hours to set up, and it's been saving me money every single month since.

The Starting Point

My application uses GPT-4o for chat completions, Claude for long-form content generation, and a custom fine-tuned model for classification tasks. When I first built this, I picked the models based on quality and just wired them up directly. Every call went straight to whatever provider I chose at build time.

The problem is that direct integration locks you into one provider's pricing. And the pricing differences between providers for similar quality outputs are honestly wild if you actually sit down and compare them.

I was running about 500,000 API calls per month. Not huge, but not trivial either. At roughly $0.40 per 1M tokens average across my workloads, with most of that being output tokens which are always more expensive, I was looking at a $200 average monthly bill. When my usage spiked to 700,000 calls, that hit $280.

The pain point wasn't just the absolute number. It was watching that number grow while knowing I couldn't do much about it without rewriting large chunks of my integration code.

The Realization

Here's what I discovered: model prices change constantly, and the "best" model for a task shifts over time. But my code had hardcoded endpoints and model names everywhere. Switching from one provider to another meant touching dozens of files.

I remember sitting there at 11 PM, looking at a Git diff that was essentially a global find-and-replace of API endpoints, thinking there has to be a better way.

There is. It's called an AI gateway.

What Actually Worked

Instead of calling OpenAI, Anthropic, or any provider directly, my code now calls one endpoint. That endpoint handles the routing, load balancing, and fallbacks. My application code doesn't know or care which provider is actually serving the request.

The setup looks like this:

import requests

# Before: direct call to OpenAI
# response = openai.chat.completions.create(model="gpt-4o", messages=messages)

# After: gateway call, same interface
response = requests.post(
    "https://my-gateway.example.com/v1/chat/completions",
    headers={"Authorization": "Bearer my-gateway-key"},
    json={
        "model": "gpt-4o",  # still specify what you want
        "messages": messages,
        "max_tokens": 500
    }
)
Enter fullscreen mode Exit fullscreen mode

The gateway takes my request, checks which provider currently offers the best price for that model category, and routes accordingly. It also handles rate limits, retries, and load balancing.

The key insight is that the interface matches the provider format. I didn't have to change any of my existing logic. The gateway presents the exact same API convention that the big providers use, so my code, my prompts, my logging all stayed intact.

The Numbers That Matter

Here's the breakdown of where the savings actually came from:

1. Dynamic Provider Selection

The biggest win was letting the gateway pick which provider serves the request. For casual chat completion, the price difference between providers for comparable models was around 30-40%. By routing to the cheapest provider with adequate quality at any given moment, I cut that cost instantly.

2. Model Tiering

I had been using the same high-end model for everything. Turns out, a lot of my requests didn't need the smartest model available. Simple classification, basic extraction, or quick summarization tasks work perfectly fine on smaller, faster models that cost about 70% less per token.

The gateway let me set rules: if the request is classified as "simple," route it to a cheaper model. Only complex reasoning tasks go to the top-tier models. My quality metrics didn't budge, but the cost did.

3. Batch Processing and Caching

This is where I got a bit nerdy. The gateway caches identical requests, which I didn't even realize I was making so many of. When the same request comes in twice, it serves the cached response. My cache hit rate is around 18% now, which sounds low, but those are 18% of calls I'm paying exactly zero for.

4. Pay-As-You-Go Over Subscriptions

This was the biggest mindset shift. I was holding standing API credits with multiple providers, paying monthly minimums, and letting that money just sit there. When I moved to a pay-as-you-go routing model, I stopped paying for unused capacity.

The cost comparison after a full month:

Cost Category Before After
Provider fees $185 $52
Gateway usage — $8
Total $185 $60

That's a 68% reduction, and it's been consistent for three months now.

The Implementation Details

If you're thinking about doing this, here's the honest breakdown of what it takes:

What You Need

  • A gateway service (open-source options exist)
  • Read access to provider pricing (check their published rate cards)
  • Your existing API keys

The Core Config

Most gateways are configured with YAML or JSON rules. My config has two critical sections:

  1. Provider pool: which providers are available and their base URLs
  2. Routing rules: when to use which provider/model
providers:
  - name: openai
    base_url: https://api.openai.com/v1
    api_key: env.OPENAI_KEY

  - name: anthropic
    base_url: https://api.anthropic.com/v1
    api_key: env.ANTHROPIC_KEY

routes:
  - pattern: "classification"
    provider: openai
    model: gpt-3.5-turbo

  - pattern: "long-form"
    provider: anthropic
    model: claude-3-5-sonnet

  - pattern: "general"
    provider: cheapest_available
    model: auto
Enter fullscreen mode Exit fullscreen mode

The Migration

Here's the part that surprised me: migration took me about three hours for an application with 60+ integration points.

The steps were:

  1. Spin up the gateway locally
  2. Point a test environment at it
  3. Run my existing test suite (which mocked API calls) against the gateway
  4. Swap the base URL in production config
  5. Watch logs for a week

No code changes. No prompt changes. The switch was literally a configuration change.

Edge Cases I Hit

I want to be upfront about the complications, because they're real:

Token Counting Difference

Different providers count tokens slightly differently for billing. What's a "token" to Anthropic isn't always identical to OpenAI's definition. Your billing metering might show different usage numbers than what the gateway reports. I had to build a normalization layer for my internal cost tracking.

Rate Limit Headaches

Free or cheaper tiers often have rate limits that premium tiers don't. When I routed more traffic to budget providers, I hit rate limits on a few high-traffic days. The gateway's retry logic handled it, but I had to tune the retry backoff to avoid spamming.

Quality Consistency

No two models are identical, even if they score similarly on benchmarks. My classification task (which uses GPT-3.5-turbo on the budget side) gave slightly different confidence scores than the premium model. I had to adjust my confidence threshold by about 5% to maintain the same precision.

What I'm Doing Now

Since setting this up, I've built a habit of checking my gateway analytics weekly. Some providers have dropped prices by 20-30% in the months since. The gateway picks that up automatically.

I'm also experimenting with speculative routing, where the gateway sends the first few tokens to a cheap model and evaluates whether to upgrade mid-stream for complex responses. Early results suggest another 10-15% savings on long generation tasks, though the infrastructure is a bit more involved.

One more thing: I moved to a pay-as-you-go gateway service instead of running my own infrastructure for this. I don't want to maintain a load balancer, handle failover, and monitor uptime on top of my actual application. The per-request fee is small enough that it's worth it.

By the way, the gateway I moved to is tai.shadie-oneapi.com — it's a pay-as-you-go aggregator that doesn't lock you into a monthly subscription. You only pay for the tokens you actually use across whatever providers it routes to. I found it while comparing aggregate API pricing, and the no-vendor-lock-in model was what sold me.

The whole experience has shifted how I think about API costs. I used to treat them as a fixed overhead. Now they're an optimizable variable, and it's been a significant chunk of money back every month. If you're spending over $100/month on API calls, I'd genuinely recommend spending an afternoon setting up routing. The math works out.

Top comments (0)