DEV Community

Srijan Verma
Srijan Verma

Posted on

Stop Letting Flaky APIs Crash Your AI Agents

How to combine exponential backoff, circuit breakers, and graceful fallbacks for production-grade agentic workflows.

The Bottleneck in Production

AI agents are only as reliable as the tools they invoke. When an LLM decides to search the web, scrape a URL, or fetch database records, it depends entirely on network stability.

In production, external APIs fail constantly. A sudden surge causes 429 rate limits, a third-party microservice throws a 504 timeout, or a target endpoint goes down entirely. The naive approach—executing raw tool calls directly inside the agent loop—is a ticking time bomb:

# The Naive Anti-Pattern: Fragile Tool Execution
def execute_agent_tool(tool_name: str, payload: dict):
    # One 500 error here kills the entire multi-step reasoning chain
    response = requests.post(f"https://api.service.internal/{tool_name}", json=payload)
    return response.json()
Enter fullscreen mode Exit fullscreen mode

When this call breaks, the unhandled exception crashes the runtime. You lose the entire reasoning graph, waste LLM tokens, and degrade the user experience.


The System Architecture: Layered Tool Defense

To keep multi-step agents alive, you need a defensive execution pipeline wrapped around every tool. Instead of allowing errors to bubble up and kill the agent, we handle failures across three distinct layers:

  1. Exponential Backoff: Mitigate transient network glitches and minor rate spikes by retrying with increasing delays.
  2. Circuit Breaker: Detect persistent downtime. If an API fails three times consecutively, trip the breaker to stop sending doomed requests.
  3. Graceful Fallbacks & Partial Degradation: When a primary service is down, route the query to a replica, cached store, or lightweight fallback (e.g., cached search index instead of a live browser scrape).
[ Agent Core ] 
      │
      ▼
┌───────────────────────────────┐
│     Circuit Breaker Check     │
│   (Is Primary Service Up?)    │
└──────────────┬────────────────┘
       OPEN    │   CLOSED (Healthy)
       ┌───────┴────────┐
       ▼                ▼
┌─────────────┐  ┌───────────────────────────┐
│  Fallback   │  │ Retry Engine (Backoff)    │
│  Provider   │  │ └──> Primary API Endpoint │
└──────┬──────┘  └──────────────┬────────────┘
       │                        │ SUCCESS
       │ (Degraded)             │
       ▼                        ▼
┌────────────────────────────────────────────┐
│      Structured Response + Metadata        │
│   (LLM receives context of degradation)    │
└────────────────────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode

By returning a degraded result accompanied by metadata (e.g., status: "degraded", source: "fallback_cache"), the LLM can adjust its downstream reasoning rather than hallucinating over missing data.


The Implementation

We combine tenacity for retry logic with circuitbreaker to isolate failing services. The following production-ready pattern ensures failures are caught and handled before reaching the LLM orchestrator.

from tenacity import retry, stop_after_attempt, wait_exponential
from circuitbreaker import circuit, CircuitBreakerError

# 1. Protect external call with circuit breaker and exponential backoff
@circuit(failure_threshold=3, recovery_timeout=60)
@retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=2, max=10))
def fetch_primary_data(query: str) -> dict:
    resp = requests.get("https://api.flaky-service.com/v1/search", params={"q": query}, timeout=3)
    resp.raise_for_status()
    return {"data": resp.json(), "source": "primary", "degraded": False}

# 2. Resilient fallback wrapper for the agent runtime
def execute_tool_safely(query: str) -> dict:
    try:
        return fetch_primary_data(query)
    except (CircuitBreakerError, Exception) as err:
        # Fallback to internal cache or secondary tool
        cached_result = local_cache.get(query) or "No fresh data available."
        return {
            "data": cached_result,
            "source": "cache_fallback",
            "degraded": True,
            "error_context": str(err)
        }
Enter fullscreen mode Exit fullscreen mode

This snippet ensures three critical guarantees:

  • Idempotent Retry Safety: The request backs off exponentially up to 10 seconds.
  • Fail-Fast Protection: Once the breaker opens, execution drops straight to the fallback without waiting for timeouts.
  • Agent Loop Continuity: The orchestrator receives a structured dictionary containing error metadata instead of an uncaught exception.

Production Lessons & Takeaways

  • Isolate Circuit Breakers Per Tool: Never use a single global circuit breaker. Wrap each external integration independently so a failing weather API doesn't disable your payment tool.
  • Pass Degradation Metadata to the Prompt: When serving fallback data, explicitly inform the model via context ("Note: Live search is offline. Using cached results from 2 hours ago."). This keeps the model's responses

Top comments (0)