The moment you add a second AI model to your app, you inherit a new class of production bug: the cascade. One provider has a bad day, your retry logic melts, and a simple 503 turns into a user-facing outage.
I learned this the hard way running agents that call Claude, GPT-4o, and DeepSeek. Here's the pattern that fixed it.
The problem with naive fallback
Most "fallback" implementations look like this:
try:
return call_claude(prompt)
except Exception:
return call_gpt(prompt)
This fails in three ways:
- It retries the same failing provider until the user gives up.
-
It ignores rate-limit headers (
retry-after), so it hammers a throttled endpoint. - It has no circuit breaker — one slow provider drags your whole p95 latency up.
The cascading router pattern
A production router needs four things: ordered providers, per-provider timeouts, a circuit breaker, and structured logging.
import time
from dataclasses import dataclass, field
@dataclass
class Provider:
name: str
fn: callable
failures: int = 0
open_until: float = 0.0
def available(self) -> bool:
return time.time() >= self.open_until
def trip(self, cooldown: float = 30.0):
self.failures += 1
if self.failures >= 3:
self.open_until = time.time() + cooldown
self.failures = 0
Then the router walks the chain, skipping tripped providers:
def route(prompt, providers, timeout=20):
errors = []
for p in providers:
if not p.available():
errors.append(f"{p.name}: circuit open")
continue
try:
return p.fn(prompt, timeout=timeout)
except Exception as e:
p.trip()
errors.append(f"{p.name}: {e}")
raise RuntimeError("all providers failed: " + "; ".join(errors))
The four rules that matter
- Skip, don't retry, a tripped provider. A circuit breaker is cheaper than a timeout.
- Fail fast, fall down. Each provider gets its own timeout (I use 20s) so a hung request can't block the cascade.
- Log the whole chain. When all providers fail, the error message should tell you why each one failed — not just the last.
- Make the order configurable. Provider health changes daily; hardcoding the order means redeploying to adapt.
What this buys you
Since I switched to this pattern, my multi-model agents went from "up when OpenAI is up" to effectively always-up. A provider outage now shows up as a slightly higher cost line, not a page at 3am.
The whole thing is ~60 lines of Python with zero dependencies. You don't need a framework — you need a circuit breaker and the discipline to skip instead of retry.
I packaged the full router + n8n workflow JSON (importable) if you'd rather not rebuild it from scratch — it's linked from my profile.
Top comments (0)