DEV Community

yureki_lab
yureki_lab

Posted on

How I Made My Autonomous Coding Agent Survive API Rate Limits and Outages

TL;DR

My fully autonomous implementation system runs 24/7, and for months its biggest killer wasn't bad code — it was the model API saying "slow down" or "I'm overloaded" at 3 a.m. I fixed it with four boring pieces: an error classifier, jittered backoff with a retry budget, a circuit breaker shared across all parallel agents, and a "checkpoint and park" step instead of crashing. Here's the build log and the five lessons I'd tell past me.

The Problem

My setup is a small fleet: an orchestrator module that hands tasks to several parallel implementation agents, plus a self-healing agent that picks up broken work. Everything talks to a hosted LLM API. When it works, I wake up to merged branches and green CI.

When it doesn't, I wake up to this:

[02:14:07] agent-3  task=refactor-auth  ERROR 429 rate_limit_error
[02:14:07] agent-1  task=add-pagination ERROR 429 rate_limit_error
[02:14:08] agent-2  task=fix-flaky-test ERROR 429 rate_limit_error
[02:14:09] agent-3  retrying (1/5)...
[02:14:09] agent-1  retrying (1/5)...
[02:14:09] agent-2  retrying (1/5)...
...
[02:14:31] agent-3  giving up. task marked FAILED
Enter fullscreen mode Exit fullscreen mode

Three things were going wrong at once:

  1. Every error looked the same. A 429 (rate limit), a 529/503 (provider overloaded), a 400 (my prompt was too big), and a network reset all went through one except Exception: retry() path.
  2. Agents retried in lockstep. Four agents hit the limit together, backed off for the same fixed time, then hit it together again. Classic thundering herd.
  3. "Giving up" meant losing work. A task 40 minutes into a refactor would be marked failed, and the next run started from scratch — burning even more tokens and rate limit.

Over one rough two-week stretch I counted it: 31% of failed tasks weren't real failures at all. They were transient API errors that my system turned into permanent ones.

The constraint that made it interesting: I can't just "add more quota." Limits are per account, all agents share them, and outages happen whether I like it or not. The system had to degrade gracefully, not just retry harder.

How I Solved It

Here's the shape of the final design:

flowchart LR
    A[Agent call] --> B{Circuit open?}
    B -- yes --> P[Checkpoint & park task]
    B -- no --> C[Call model API]
    C -- ok --> D[Continue task]
    C -- error --> E[Classify error]
    E -- fatal --> F[Fail fast, explain why]
    E -- retryable --> G{Retry budget left?}
    G -- yes --> H[Jittered backoff] --> C
    G -- no --> I[Trip breaker] --> P

All code below is Python 3.13, trimmed to the load-bearing parts.

1. Classify before you retry

The single highest-value change was refusing to treat errors as one bucket.

from enum import Enum

class ErrKind(Enum):
    RATE_LIMIT = "rate_limit"   # 429: we're too fast
    OVERLOADED = "overloaded"   # 529/503: they're too busy
    TRANSIENT = "transient"     # timeouts, resets, 500/502
    FATAL = "fatal"             # 400/401/403/413: retrying won't help

def classify(status: int | None, exc: Exception | None) -> ErrKind:
    if status == 429:
        return ErrKind.RATE_LIMIT
    if status in (503, 529):
        return ErrKind.OVERLOADED
    if status is None or status in (500, 502, 504):
        return ErrKind.TRANSIENT  # includes connection resets
    return ErrKind.FATAL
Enter fullscreen mode Exit fullscreen mode

That FATAL bucket matters more than it looks. Before this, a prompt that exceeded the context window got retried five times — five identical, guaranteed failures, each one eating rate limit that the other agents needed. Now a 400 fails immediately with a clear reason, and the self-healing agent can actually do something useful with it (like splitting the task).

2. Jittered backoff with a retry budget, not a retry count

Fixed retries ("try 5 times") were the root of the lockstep problem. I switched to exponential backoff with full jitter, and I honor retry-after when the API sends it:

import random

BASE = {ErrKind.RATE_LIMIT: 2.0, ErrKind.OVERLOADED: 10.0, ErrKind.TRANSIENT: 1.0}
CAP = 120.0

def backoff(kind: ErrKind, attempt: int, retry_after: float | None) -> float:
    if retry_after is not None:
        # Server knows best, but add a little jitter so agents don't sync up.
        return retry_after + random.uniform(0, 2)
    ceiling = min(CAP, BASE[kind] * 2 ** attempt)
    return random.uniform(0, ceiling)  # full jitter
Enter fullscreen mode Exit fullscreen mode

The bigger idea is the retry budget. Instead of "5 attempts per call," each task gets a time budget for waiting — say, 8 minutes total across the whole task. A task that's had a smooth run can absorb one long overload; a task that's been fighting errors all along runs out sooner and gets parked.

import time

class RetryBudget:
    def __init__(self, seconds: float):
        self.remaining = seconds

    def spend(self, wait: float) -> bool:
        if wait > self.remaining:
            return False
        self.remaining -= wait
        time.sleep(wait)
        return True
Enter fullscreen mode Exit fullscreen mode

This made behavior predictable. I know the worst-case time a task spends waiting, regardless of how many calls it makes.

3. One circuit breaker for the whole fleet

Jitter spreads retries out, but four agents independently backing off still means four agents independently probing a provider that's clearly down. So I added a circuit breaker — and crucially, made it shared. All agents read and write one small state file (a lock-protected JSON file was enough; no Redis needed at my scale).

import json, time
from pathlib import Path
from filelock import FileLock

STATE = Path("breaker.json")
LOCK = FileLock("breaker.json.lock")
THRESHOLD, COOLDOWN = 6, 300  # 6 failures -> open for 5 min

def breaker_open() -> bool:
    with LOCK:
        s = json.loads(STATE.read_text()) if STATE.exists() else {}
        return time.time() < s.get("open_until", 0)

def record(ok: bool) -> None:
    with LOCK:
        s = json.loads(STATE.read_text()) if STATE.exists() else {"fails": 0}
        s["fails"] = 0 if ok else s.get("fails", 0) + 1
        if s["fails"] >= THRESHOLD:
            s["open_until"] = time.time() + COOLDOWN
            s["fails"] = 0
        STATE.write_text(json.dumps(s))
Enter fullscreen mode Exit fullscreen mode

When the breaker opens, nobody calls the API for five minutes. After the cooldown, the first agent to wake up acts as the probe. If it succeeds, the count resets and everyone else proceeds. If it fails, the breaker re-opens. This took my "error storm" logs from hundreds of lines to a handful.

4. Checkpoint and park instead of crash

This was the piece that actually saved work. When a task runs out of retry budget or hits an open breaker, it doesn't fail. It writes a checkpoint and goes into a PARKED state:

def park(task, reason: str) -> None:
    task.checkpoint = {
        "branch": task.branch,               # work is already committed here
        "completed_steps": task.done_steps,
        "next_step": task.current_step,
        "notes": task.summary_so_far(),      # short, human-readable
    }
    task.status = "PARKED"
    task.park_reason = reason
    task.save()
Enter fullscreen mode Exit fullscreen mode

The orchestrator treats PARKED differently from FAILED: parked tasks get resumed first once the breaker closes, starting from next_step with the notes injected into the prompt. Because each agent commits to its branch after every meaningful step, "resume" is cheap — the agent reads the checkpoint and the diff, not the whole history.

Here's what the same night looks like now:

[02:14:07] agent-3  429 rate_limit -> wait 3.1s (budget 476s left)
[02:14:08] agent-1  429 rate_limit -> wait 0.7s (budget 480s left)
[02:14:11] agent-2  529 overloaded -> wait 14.2s
[02:14:40] breaker  OPEN for 300s (6 consecutive failures)
[02:14:40] agent-3  PARKED at step 4/7 (breaker open)
[02:14:41] agent-1  PARKED at step 2/5 (breaker open)
[02:19:41] agent-2  probe OK -> breaker CLOSED
[02:19:42] agent-3  RESUMED from step 4/7
Enter fullscreen mode Exit fullscreen mode

The numbers

After running this for about six weeks:

  • ✅ Tasks failing due to transient API errors: 31% → under 3% of all failures
  • ✅ Wasted retries on fatal errors (oversized prompts, etc.): effectively zero
  • ✅ Average extra wall-clock per parked task: ~7 minutes, versus a full restart before
  • ⚠️ Morning surprises: still non-zero, but now they're real bugs, which is what I want to wake up to

Lessons Learned

1. Most "agent failures" are infrastructure failures wearing a costume. 💡 Before blaming the model's reasoning, check how many failures were the network or the provider. For me it was almost a third. Fixing plumbing beat any prompt tweak I tried that month.

2. Never retry an error you haven't classified. A blanket retry turns a cheap, instant failure (bad request) into a slow, expensive one — and it steals capacity from healthy work. If you only do one thing from this post, do the classifier.

3. Budget time, not attempts. "5 retries" means wildly different things depending on backoff. A per-task wait budget gives you a hard ceiling you can reason about and makes long tasks and short tasks behave fairly.

4. Parallel agents need shared failure state. Each agent being individually polite is not enough. If they don't share a breaker, they will collectively hammer a struggling API. Shared state can be embarrassingly simple — a locked JSON file worked fine for a handful of agents.

5. Design for pause, not just success or failure. The PARKED state was the real unlock. Long-running autonomous work needs a third outcome: "I stopped safely and I know exactly where to pick up." That only works if the agent commits often and leaves short notes as it goes.

What's Next

A few things I'm working on now:

  • Priority-aware throttling — when quota is tight, let short, high-value tasks go first and park the big refactors.
  • Fallback models for low-risk steps — things like summarizing a diff or writing a commit message don't need the strongest model, so they shouldn't compete for the same limit.
  • Surfacing breaker trips on my remote control dashboard, so I can see at a glance whether a quiet night was "nothing to do" or "provider was down."

Wrap-up

If you're running agents unattended — even a single Claude Code session in a loop — your reliability ceiling is probably your error handling, not your prompts. Classify, back off with jitter, share a breaker, and park instead of crash.

👉 Follow me on Dev.to for more build logs from running a fully autonomous coding system, and drop a comment: what's the weirdest way an API outage has broken your agent setup? I'd love to compare war stories. 🚀

Top comments (1)

Collapse
 
max_quimby profile image
Max Quimby •

The "31% of failed tasks weren't real failures" number is the whole article — that's the stat that justifies every boring piece of the fix.

The detail I'd underline for anyone copying this: the circuit breaker has to be shared across agents, which you did, but the subtle part is who gets to close it again. If every parked agent independently probes to see if the provider recovered, you rebuild the thundering herd at recovery time instead of at failure time. A single half-open probe — one agent tests the water, everyone else waits for its verdict — is what actually keeps you from getting re-throttled the instant the limit clears.

The other thing that saved us was making "checkpoint and park" resumable at the sub-task level, not the task level. Parking a 40-minute refactor is only a win if resume doesn't redo the first 38 minutes.

Question on the retry budget: is it per-task or per-account-per-window? We found a per-task budget still let a fleet of tasks collectively hammer a shared limit, so we had to move the budget up to the account level and have the breaker enforce it globally.