DEV Community

Aman Kumar
Aman Kumar

Posted on

Model Fallback Chains on One OpenAI-Compatible Base URL: When to Switch Models and How (Python)

Most apps I see call exactly one model. That works until the model is overloaded, deprecated, or just returns something your parser can't use, and then the whole feature is down. If your endpoint exposes several models behind one OpenAI-compatible base URL, a fallback chain is cheap insurance. The hard part isn't the loop. It's deciding which errors deserve a fallback and which ones should fail loudly.

My examples use APIClaw, the flat-rate OpenAI-compatible gateway I build, so weigh my choice of endpoint accordingly. Nothing below is specific to it; any provider that speaks the Chat Completions format works the same way.

Step 1: get real model IDs, don't guess them

The most common fallback bug I see is a chain whose second or third entry was never a valid model ID. It sits there for weeks and fails the one time you need it. Pull the IDs from the endpoint itself:

from openai import OpenAI

client = OpenAI(base_url="https://apiclaw.biz/v1", api_key="YOUR_KEY")
for m in client.models.list().data:
    print(m.id)
Enter fullscreen mode Exit fullscreen mode

Copy IDs exactly as printed. Gateways often name models differently from the original vendor, and a near-miss like a dot instead of a dash gives you a 404 "model not found", not a helpful suggestion.

I'd also run this check in CI or at startup and refuse to boot if a configured fallback model isn't in the list. It takes one request and catches the whole category of silent misconfiguration.

Step 2: classify errors before you write the loop

Here is how I split failures:

Fall back to the next model:

  • 404 or "model not found" for that specific model (it was removed or renamed)
  • 429 that persists after your normal retry with backoff
  • 5xx and 529-style "overloaded" errors after one retry
  • Timeouts on a model that is usually fast
  • A response that fails your own validation (empty content, invalid JSON when you asked for JSON)

Do not fall back, raise immediately:

  • 401 or 403. A bad key is bad for every model; trying three models just triples the noise in your logs.
  • 400 errors about your request shape, like an invalid parameter or a message that is too long. The next model will usually reject it too, and if it doesn't, you've hidden a bug.
  • Content you need a human to look at.

That second list is the part people skip. A fallback chain that swallows 400s turns a one-line fix into a mystery about why outputs got worse.

Step 3: the loop

import time
import openai
from openai import OpenAI

client = OpenAI(
    base_url="https://apiclaw.biz/v1",
    api_key="YOUR_KEY",
    timeout=60,
    max_retries=1,   # SDK retries once per model; the chain handles the rest
)

CHAIN = ["PRIMARY_MODEL_ID", "SECOND_MODEL_ID", "CHEAP_MODEL_ID"]

FALLBACK_ERRORS = (
    openai.NotFoundError,
    openai.RateLimitError,
    openai.InternalServerError,
    openai.APITimeoutError,
    openai.APIConnectionError,
)

def complete(messages, validate=None, **kwargs):
    last_err = None
    for model in CHAIN:
        start = time.monotonic()
        try:
            resp = client.chat.completions.create(
                model=model, messages=messages, **kwargs
            )
            text = resp.choices[0].message.content or ""
            if validate and not validate(text):
                raise ValueError(f"validation failed on {model}")
            log(model, "ok", time.monotonic() - start, resp.usage)
            return text, model
        except (openai.AuthenticationError,
                openai.PermissionDeniedError,
                openai.BadRequestError):
            raise
        except FALLBACK_ERRORS + (ValueError,) as e:
            log(model, type(e).__name__, time.monotonic() - start, None)
            last_err = e
            continue
    raise RuntimeError("all models in chain failed") from last_err

def log(model, status, seconds, usage):
    print(f"{model} {status} {seconds:.1f}s {usage}")
Enter fullscreen mode Exit fullscreen mode

A few deliberate choices in there:

  • max_retries=1 on the client. The SDK defaults to retrying twice with backoff. Multiply that by three models and a single bad request can turn into nine calls. One retry per model is enough for a transient blip.
  • The function returns which model answered. Store it with the output. When someone reports a weird answer, you want to know it came from the third model in the chain, not the first.
  • Validation failures count as fallback triggers, but only for checks you'd actually reject in production. "Is this valid JSON with the keys I need?" is a good check. "Is this answer good?" is not something to automate inside a retry loop.

Step 4: order the chain on purpose

There are two sensible orderings, and they solve different problems.

Quality first: strongest model, then a similar-tier model from a different family, then a cheaper one. Use this for user-facing answers where a degraded reply beats an error page. Putting a different model family second matters: if the first failure is a vendor-wide outage, a sibling model from the same vendor often fails too.

Cheap first: a fast, inexpensive model, then escalate to a stronger one only when validation fails. This works well for extraction, classification and formatting jobs, where the cheap model is right most of the time and your validator can tell when it isn't.

Whatever the order, keep chains short. Three models is usually the most I'd use. Past that, latency on the worst-case path gets bad: three 60-second timeouts is a three-minute request.

Step 5: watch the prompts, not just the errors

Different models follow the same prompt differently. Things that have bitten me:

  • System prompt strength. A terse instruction one model obeys, another treats as a suggestion. If the fallback model regularly fails validation, the fix is often a more explicit prompt, not a different model.
  • Tool calling. Tool schemas usually carry over on an OpenAI-compatible endpoint, but models differ in how eagerly they call tools and how they format arguments. Parse arguments defensively and validate them with the same schema you advertised.
  • Max tokens. If you set max_tokens tight for the primary model, a wordier fallback can get cut off mid-JSON. Check finish_reason; if it's length, treat that as a validation failure.
  • Streaming. Falling back mid-stream is messy, because you've already shown partial output. For streamed UIs I only fall back before the first token arrives, and I fail visibly after that.

Step 6: test the chain on purpose

A fallback path nobody exercises is a fallback path that doesn't work. Two cheap tests:

  1. Put a fake model ID first in the chain in a test and assert the call succeeds via the second model. This proves the 404 path works.
  2. Pass a validator that always returns False and assert you get RuntimeError after exactly len(CHAIN) attempts. This proves you aren't looping forever or swallowing the final error.

In production, count how often each model answers. If the second model is answering 20% of requests, that isn't resilience anymore; that's your primary model being the wrong choice.

A note on cost

On per-token pricing, every fallback attempt that produced output costs money, and validation-triggered fallbacks can quietly double the bill for a feature. On a flat-rate plan the cost question becomes a request-count question instead, so the thing to watch is how many attempts each user action triggers. Either way, the log line in the loop above is what lets you see it.

Summary

  • Pull model IDs from /v1/models and verify them at startup.
  • Fall back on 404, persistent 429, 5xx, timeouts and validation failures.
  • Raise immediately on 401, 403 and 400.
  • Keep the SDK's own retries low so attempts don't multiply.
  • Record which model answered, keep chains to about three models, and test the fallback path deliberately.

It's maybe forty lines of code, and it turns "the model is down" from an outage into a log entry.

Top comments (0)