DEV Community

Cover image for Building a Multi-Provider Fallback Router for Resilient AI Agents
Tanu Sri
Tanu Sri

Posted on

Building a Multi-Provider Fallback Router for Resilient AI Agents

If your autonomous agent relies on a single AI provider, it is an outage waiting to happen.

Rate limits, regional outages, and API failures can bring your system down at the worst possible moment.

We learned this during a live demo.

Then we made our agent provider-agnostic.

The problem

Our first implementation used a single provider for all natural-language explanations.

During a demo to a small audience, the API returned a 429 rate-limit error.

The entire AI Analysis page went blank.

We finished the demo with an empty dashboard.

That was the last time we trusted a single provider.

The fix — a fallback router in the AI Copilot

We rewrote platform/ai-copilot/main.py to support multiple providers through a fallback router.

The agent tries Groq first for low-latency inference, then falls back to Gemini or OpenAI if the primary provider is unavailable.

Here is the core logic:


# platform/ai-copilot/main.py

GROQ_API_KEY = os.getenv("GROQ_API_KEY", "")
GROQ_API_URL = "https://api.groq.com/openai/v1/chat/completions"

try:
    async with httpx.AsyncClient() as client:
        res = await client.post(
            GROQ_API_URL,
            json={
                "model": "openai/gpt-oss-20b",
                "messages": [
                    {
                        "role": "user",
                        "content": prompt
                    }
                ],
                "temperature": 0.2,
                "response_format": {
                    "type": "json_object"
                },
            },
            headers={
                "Authorization": f"Bearer {GROQ_API_KEY}",
                "Content-Type": "application/json",
            },
            timeout=15.0,
        )

        res.raise_for_status()

        content = res.json()["choices"][0]["message"]["content"]
        parsed = json.loads(content)

        return parsed

except json.JSONDecodeError:
    logger.error(
        f"Failed to parse Groq response as JSON: {content}"
    )

    return fallback_response()

except Exception as e:
    logger.error(
        f"groq_api_failed: {str(e)}"
    )

    return await fallback_to_gemini(incident_data)
Enter fullscreen mode Exit fullscreen mode

The important part is not the individual provider.

It is the fallback contract.

No matter which provider serves the request, the caller receives the same response structure:


{
    "executive_summary": "...",
    "technical_summary": "...",
    "postmortem": "..."
}
Enter fullscreen mode Exit fullscreen mode

This keeps the downstream application independent of the provider that generated the response.

Why the contract matters more than the provider

Every provider has different API behavior, limits, models, and response characteristics.

If provider-specific formatting leaks into your downstream code, changing providers becomes a refactoring exercise.

Instead, we enforce a single response contract.

Conceptually:

                 ┌─────────────┐
                 │  AI Copilot │
                 └──────┬──────┘
                        │
                Provider Router
                        │
          ┌─────────────┼─────────────┐
          ↓             ↓             ↓
        Groq         Gemini         OpenAI
          │             │             │
          └─────────────┼─────────────┘
                        ↓
              Common Response Contract
                        ↓
          ┌─────────────┼─────────────┐
          ↓             ↓             ↓
       Dashboard      AI Analysis    Audit
Enter fullscreen mode Exit fullscreen mode

Our downstream consumers — the dashboard BFF, the AI Analysis page, and the audit engine — do not need to know which provider served the request.

They simply receive a well-formed summary object.

That separation is what makes the providers interchangeable.

Before vs after

The following results are from our project testing:

Metric Before After
API timeout failures per week 12 0
Providers supported 1 3
Mean time to failover N/A 1.2s
Code changes to swap providers Full refactor Environment variable

The goal was not to eliminate provider failures.

The goal was to prevent a single provider failure from becoming an application failure.

The honest lesson

Always decouple the reasoning layer from the specific API implementation.

Your agent should not become unavailable simply because one provider becomes unavailable.

A provider abstraction allows the system to switch between providers without rewriting the downstream application.

The fallback mechanism is simple, but its impact is significant:

Primary Provider
      ↓
   Available?
    ↙     ↘
  Yes      No
   ↓        ↓
Response   Fallback
              ↓
       Secondary Provider
              ↓
           Response
Enter fullscreen mode Exit fullscreen mode

A small amount of resilience engineering can prevent a provider outage from becoming an application outage.

How this fits with agent memory

A multi-provider router handles availability.

Agent memory handles knowledge.

Combined with agent memory, the fallback router allows the reasoning layer to remain available when the primary provider fails, while the memory layer allows the agent to recall information from previous incidents.

The Hindsight documentation explains how persistent memory complements generative reasoning.

The Hindsight GitHub repository also provides concrete examples of the recall/retain pattern.

The overall architecture becomes:

Incident
   ↓
Agent
   ↓
Memory Recall
   ↓
Known Incident?
  ↙       ↘
Yes        No
 ↓          ↓
Recall    AI Reasoning
              ↓
       Provider Router
         ↙    ↓    ↘
      Groq Gemini OpenAI
         \     |     /
          \    |    /
           Response
              ↓
        Recovery Workflow
              ↓
          Retain Result
              ↓
          Agent Memory
Enter fullscreen mode Exit fullscreen mode

Memory reduces unnecessary reasoning for recurring incidents, while the provider-agnostic reasoning layer provides resilience when new reasoning is required.

Together, they give the agent both continuity and availability.

Top comments (0)