Your agent's retry loop is either too dumb or infinite: handling 429 with Retry-After + a time budget
Every autonomous agent that calls LLM APIs eventually hits this problem: HTTP 429 rate limits. Most agents either:
- Give up immediately (losing valid requests)
- Retry instantly (getting banned faster)
- Use a fixed backoff (wasting time when the server already told us exactly how long to wait)
What if your retry logic could actually read the server's instructions?
The Problem with Simple Retries
When an LLM API returns HTTP 429, it usually includes headers like:
-
Retry-After: 60(wait 60 seconds) -
retry-after-ms: 120000(wait 120 milliseconds) -
x-ratelimit-reset-requests: 1.5s(wait 1.5 seconds)
Most retry libraries either ignore these or can't parse them. Others retry with fixed delays, wasting the server's explicit guidance.
A Better Approach: Retry-After + Time Budget
I built a simple Python shim that:
- Reads server wait times from all common header formats
- Uses exponential backoff with jitter when no wait time is specified
- Respects a total time budget so your agent doesn't get stuck
- Stops on non-retryable errors (400/401/403/404/422)
from llm_retry_shim import retry_call
# Works with any LLM SDK
resp = retry_call(
lambda: client.messages.create(model="claude-3-5-sonnet-20241022",
max_tokens=512,
messages=msgs),
max_attempts=6,
budget_s=90
)
Key Features
- No dependencies - just one Python file
- Parses all common wait headers - seconds, milliseconds, time formats
- Full jitter backoff when server gives no guidance
- Hard stops when you run out of attempts or time
- Never retries your own mistakes - 4xx errors fail fast
Why This Matters for Autonomous Agents
Your agent needs to be resilient but not wasteful. When APIs rate limit, you want to:
- Use the server's exact wait time when provided
- Backoff intelligently when not
- Never get stuck in an infinite retry loop
- Preserve your total time budget for other tasks
Get the Free Tool
I've released this as a free, open-source tool with no restrictions:
It includes a self-test function and works with Anthropic, OpenAI, or any HTTP-based LLM client.
What's Next?
Rate limits are just one piece of building robust autonomous agents. What other reliability challenges are you facing with your LLM-powered systems?
More free agent tools from Rene Vergara: https://renevibe76.gumroad.com
Top comments (0)