DEV Community

rolltok
rolltok

Posted on

How I Switched Between 5 LLM Providers Without Changing My Code

Last Tuesday, I was debugging a production issue at 2pm when I got a 503 error from my LLM provider. Again.

This was the third time that week. My app was down for 15 minutes while I manually switched to a backup provider, updated environment variables, and redeployed. By the time I finished, I'd lost three customers.

I thought: there has to be a better way.

The Problem

Most LLM providers have their own API format. OpenAI uses one structure, Anthropic uses another, Google has yet another. If you want to switch providers (say, when one goes down), you need to rewrite your entire request logic.

Here's what I mean:

# OpenAI format
response = openai.ChatCompletion.create(
    model="gpt-4",
    messages=[{"role": "user", "content": "Hello"}]
)

# Anthropic format
response = anthropic.messages.create(
    model="claude-3-opus-20240229",
    messages=[{"role": "user", "content": "Hello"}]
)
Enter fullscreen mode Exit fullscreen mode

Different SDKs. Different request structures. Different response formats. If you want to support multiple providers, you're writing adapter layers for each one.

What I Tried First

I started by building my own abstraction layer:

class LLMProvider:
    def __init__(self, provider_name):
        self.provider = provider_name

    def chat(self, message):
        if self.provider == "openai":
            return self._call_openai(message)
        elif self.provider == "anthropic":
            return self._call_anthropic(message)
        # ... 500 lines of adapter code later
Enter fullscreen mode Exit fullscreen mode

This worked, but it was fragile. Every time a provider updated their API, I had to update my adapter. I was spending more time maintaining adapters than building features.

The Unified API Approach

Then I found a different approach: use a single API endpoint that's compatible with the OpenAI format, but routes to different providers behind the scenes.

Here's what the code looks like:

import os
import requests

api_key = os.getenv("LLM_API_KEY")

def chat(model, message):
    response = requests.post(
        "https://api.example.com/v1/chat/completions",
        headers={
            "Authorization": f"Bearer {api_key}",
            "Content-Type": "application/json"
        },
        json={
            "model": model,  # Just change this to switch providers
            "messages": [
                {"role": "user", "content": message}
            ]
        }
    )
    return response.json()["choices"][0]["message"]["content"]

# Use Qwen
print(chat("qwen3.7-plus", "Hello!"))

# Switch to DeepSeek - no code changes
print(chat("deepseek-chat", "Hello!"))
Enter fullscreen mode Exit fullscreen mode

That's it. No adapter layers. No SDK switching. Just change the model name.

Handling Errors (Because APIs Fail)

Here's where things get tricky. APIs fail. A lot.

Last Thursday, I was testing this approach and got a 429 error (rate limit exceeded). My first thought was to just retry immediately. Bad idea. I got another 429. Then another. I was stuck in a retry loop.

Here's what actually worked:

import time

def chat_with_retry(model, message, max_retries=3):
    for attempt in range(max_retries):
        try:
            response = requests.post(
                "https://api.example.com/v1/chat/completions",
                headers={
                    "Authorization": f"Bearer {api_key}",
                    "Content-Type": "application/json"
                },
                json={
                    "model": model,
                    "messages": [{"role": "user", "content": message}]
                }
            )
            response.raise_for_status()
            return response.json()["choices"][0]["message"]["content"]

        except requests.exceptions.HTTPError as e:
            if response.status_code == 429:
                wait_time = 2 ** attempt  # Exponential backoff: 1s, 2s, 4s
                print(f"Rate limit hit. Waiting {wait_time}s...")
                time.sleep(wait_time)
            elif response.status_code == 401:
                print("Invalid API key. Check your credentials.")
                return None
            else:
                print(f"HTTP error: {e}")
                return None

    print("Max retries exceeded")
    return None
Enter fullscreen mode Exit fullscreen mode

The key is exponential backoff. Don't retry immediately. Wait 1 second, then 2 seconds, then 4 seconds. This gives the API time to recover.

Streaming Responses

If you want real-time output (like ChatGPT), you need streaming. Here's how I did it:

def chat_stream(model, message):
    response = requests.post(
        "https://api.example.com/v1/chat/completions",
        headers={
            "Authorization": f"Bearer {api_key}",
            "Content-Type": "application/json"
        },
        json={
            "model": model,
            "messages": [{"role": "user", "content": message}],
            "stream": True
        },
        stream=True
    )

    for line in response.iter_lines():
        if line:
            # Parse Server-Sent Events format
            # Format: data: {"choices": [{"delta": {"content": "..."}}]}
            decoded = line.decode("utf-8")
            if decoded.startswith("data: "):
                data = decoded[6:]  # Remove "data: " prefix
                if data != "[DONE]":
                    import json
                    chunk = json.loads(data)
                    content = chunk["choices"][0]["delta"].get("content", "")
                    print(content, end="", flush=True)
Enter fullscreen mode Exit fullscreen mode

This prints each token as it arrives, giving you that typewriter effect.

What I Learned

After two weeks of using this approach in production, here's what I've learned:

  1. Unified APIs save time. I spent 3 days building my own adapter layer. I could have spent 3 hours integrating a unified API.

  2. Error handling is critical. Don't just retry immediately. Use exponential backoff. Log errors. Monitor rate limits.

  3. Streaming is worth it. Users prefer seeing tokens appear in real-time rather than waiting 10 seconds for a full response.

  4. Model switching should be trivial. If you need to change 50 lines of code to switch providers, your abstraction is too complex.

Full Disclosure

I'm now working on a project called RollTok, which is a unified API gateway for multiple LLM providers. The code examples above are based on what I learned while building it. If you want to try it out, you can get an API key at rolltok.com.

But honestly, the approach works with any unified API. The specific service doesn't matter as much as the pattern.

What's Next

I'm planning to write about:

  • How to track token usage across multiple providers
  • How to implement request/response logging
  • How to set up connection pooling for better performance

If you have questions or suggestions, drop them in the comments.

Top comments (4)

Collapse
 
rinrinrinrin profile image
Rin Katsuragi •

This is insightful and interesting. I had the same issue with my partner with custom function bundles for robotics, where we had to swap AI models, but without tinkering with the function or schemas or adapters. For us Skillware worked, but it required quite some adjustments from our side, mostly cause we wanted our own custom skills for specific arms. Now this is an interesting take that reminds me of Skillware for APIs. I will try it.

Collapse
 
rolltok profile image
rolltok •

Hey Rin, thanks for this! Yeah, that's exactly the frustration - you just want to swap models without touching anything else. I haven't actually heard of Skillware before (I might be wrong, but is it more for robotics-specific use cases?). Sounds like you went through a ton of customization work for different arms.

Curious - what made you and your partner decide to build your own setup instead of using something like this? Always trying to understand how others solve this problem.

Collapse
 
rinrinrinrin profile image
Rin Katsuragi •

Skillware is not pop, it's like a niche ai agent toolkit that goes beyond simple agents skills that are just markdown files. It packages scripts, prompts, schemas, and everything a function call needs to execute deterministic code, which is what we need for robotics exactly. My partner found this framework cause they plan to expand to robotic but haven't, so we took their framework and adjusted it for robotics to use for our project, eg. udnerstanding vision, weights of objects and more without relying on the llm, but on hardcoded skill bundles.

The API approach is close or similar to MCPs I guess, but I am making connections already how this can be used for specific use cases.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.