Last Tuesday, I was debugging a production issue at 2pm when I got a 503 error from my LLM provider. Again.
This was the third time that week. My app was down for 15 minutes while I manually switched to a backup provider, updated environment variables, and redeployed. By the time I finished, I'd lost three customers.
I thought: there has to be a better way.
The Problem
Most LLM providers have their own API format. OpenAI uses one structure, Anthropic uses another, Google has yet another. If you want to switch providers (say, when one goes down), you need to rewrite your entire request logic.
Here's what I mean:
# OpenAI format
response = openai.ChatCompletion.create(
model="gpt-4",
messages=[{"role": "user", "content": "Hello"}]
)
# Anthropic format
response = anthropic.messages.create(
model="claude-3-opus-20240229",
messages=[{"role": "user", "content": "Hello"}]
)
Different SDKs. Different request structures. Different response formats. If you want to support multiple providers, you're writing adapter layers for each one.
What I Tried First
I started by building my own abstraction layer:
class LLMProvider:
def __init__(self, provider_name):
self.provider = provider_name
def chat(self, message):
if self.provider == "openai":
return self._call_openai(message)
elif self.provider == "anthropic":
return self._call_anthropic(message)
# ... 500 lines of adapter code later
This worked, but it was fragile. Every time a provider updated their API, I had to update my adapter. I was spending more time maintaining adapters than building features.
The Unified API Approach
Then I found a different approach: use a single API endpoint that's compatible with the OpenAI format, but routes to different providers behind the scenes.
Here's what the code looks like:
import os
import requests
api_key = os.getenv("LLM_API_KEY")
def chat(model, message):
response = requests.post(
"https://api.example.com/v1/chat/completions",
headers={
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json"
},
json={
"model": model, # Just change this to switch providers
"messages": [
{"role": "user", "content": message}
]
}
)
return response.json()["choices"][0]["message"]["content"]
# Use Qwen
print(chat("qwen3.7-plus", "Hello!"))
# Switch to DeepSeek - no code changes
print(chat("deepseek-chat", "Hello!"))
That's it. No adapter layers. No SDK switching. Just change the model name.
Handling Errors (Because APIs Fail)
Here's where things get tricky. APIs fail. A lot.
Last Thursday, I was testing this approach and got a 429 error (rate limit exceeded). My first thought was to just retry immediately. Bad idea. I got another 429. Then another. I was stuck in a retry loop.
Here's what actually worked:
import time
def chat_with_retry(model, message, max_retries=3):
for attempt in range(max_retries):
try:
response = requests.post(
"https://api.example.com/v1/chat/completions",
headers={
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json"
},
json={
"model": model,
"messages": [{"role": "user", "content": message}]
}
)
response.raise_for_status()
return response.json()["choices"][0]["message"]["content"]
except requests.exceptions.HTTPError as e:
if response.status_code == 429:
wait_time = 2 ** attempt # Exponential backoff: 1s, 2s, 4s
print(f"Rate limit hit. Waiting {wait_time}s...")
time.sleep(wait_time)
elif response.status_code == 401:
print("Invalid API key. Check your credentials.")
return None
else:
print(f"HTTP error: {e}")
return None
print("Max retries exceeded")
return None
The key is exponential backoff. Don't retry immediately. Wait 1 second, then 2 seconds, then 4 seconds. This gives the API time to recover.
Streaming Responses
If you want real-time output (like ChatGPT), you need streaming. Here's how I did it:
def chat_stream(model, message):
response = requests.post(
"https://api.example.com/v1/chat/completions",
headers={
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json"
},
json={
"model": model,
"messages": [{"role": "user", "content": message}],
"stream": True
},
stream=True
)
for line in response.iter_lines():
if line:
# Parse Server-Sent Events format
# Format: data: {"choices": [{"delta": {"content": "..."}}]}
decoded = line.decode("utf-8")
if decoded.startswith("data: "):
data = decoded[6:] # Remove "data: " prefix
if data != "[DONE]":
import json
chunk = json.loads(data)
content = chunk["choices"][0]["delta"].get("content", "")
print(content, end="", flush=True)
This prints each token as it arrives, giving you that typewriter effect.
What I Learned
After two weeks of using this approach in production, here's what I've learned:
Unified APIs save time. I spent 3 days building my own adapter layer. I could have spent 3 hours integrating a unified API.
Error handling is critical. Don't just retry immediately. Use exponential backoff. Log errors. Monitor rate limits.
Streaming is worth it. Users prefer seeing tokens appear in real-time rather than waiting 10 seconds for a full response.
Model switching should be trivial. If you need to change 50 lines of code to switch providers, your abstraction is too complex.
Full Disclosure
I'm now working on a project called RollTok, which is a unified API gateway for multiple LLM providers. The code examples above are based on what I learned while building it. If you want to try it out, you can get an API key at rolltok.com.
But honestly, the approach works with any unified API. The specific service doesn't matter as much as the pattern.
What's Next
I'm planning to write about:
- How to track token usage across multiple providers
- How to implement request/response logging
- How to set up connection pooling for better performance
If you have questions or suggestions, drop them in the comments.




Top comments (4)
This is insightful and interesting. I had the same issue with my partner with custom function bundles for robotics, where we had to swap AI models, but without tinkering with the function or schemas or adapters. For us Skillware worked, but it required quite some adjustments from our side, mostly cause we wanted our own custom skills for specific arms. Now this is an interesting take that reminds me of Skillware for APIs. I will try it.
Hey Rin, thanks for this! Yeah, that's exactly the frustration - you just want to swap models without touching anything else. I haven't actually heard of Skillware before (I might be wrong, but is it more for robotics-specific use cases?). Sounds like you went through a ton of customization work for different arms.
Curious - what made you and your partner decide to build your own setup instead of using something like this? Always trying to understand how others solve this problem.
Skillware is not pop, it's like a niche ai agent toolkit that goes beyond simple agents skills that are just markdown files. It packages scripts, prompts, schemas, and everything a function call needs to execute deterministic code, which is what we need for robotics exactly. My partner found this framework cause they plan to expand to robotic but haven't, so we took their framework and adjusted it for robotics to use for our project, eg. udnerstanding vision, weights of objects and more without relying on the llm, but on hardcoded skill bundles.
The API approach is close or similar to MCPs I guess, but I am making connections already how this can be used for specific use cases.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.