DEV Community

QuietDesk Studio
QuietDesk Studio

Posted on

Rate-Limiting and Backoff Patterns for MCP Servers That Don't Fall Over Under Load

A team I talked with last quarter had an agent stuck in a loop calling the same tool for eight hours. By the time someone noticed, it had run up a five-figure cloud bill. That's not a hypothetical — a runaway automation loop making 127,000 API calls in roughly 8 hours has been reported to cost around $47,000 in a single incident. No rate limiter, no circuit breaker, nothing between the agent and the API meter.

This post covers how to design rate limiting and retry/backoff for an MCP server so a stuck agent gets stopped cheaply instead of expensively. It does not cover OAuth token refresh (see my earlier post on that gotcha) or full observability setup — that's a separate piece. Scope here: what layer to rate-limit at, how to return a 429 an agent can actually act on, and the one retry mistake that causes most MCP production outages.

Why "just add exponential backoff" isn't the fix

Most teams stop at client-side retry logic and call it done. That's table stakes, not a solution — it protects against transient blips but does nothing for structural overload, and it can make things worse if you get the placement wrong.

The failure mode that actually takes MCP servers down in production isn't "no backoff." It's cascading retries inside a single agent turn: the model calls a tool, gets a 429, the MCP server retries internally with backoff, and while that retry is still pending, the same agent turn calls a second tool that hits the same upstream bucket. Now you have two retry loops racing on one quota, and bucket recovery ends up taking the longer of the two backoff schedules instead of either one alone.

The fix isn't fancier backoff math. It's not retrying inside the server at all for agent-facing calls — bubble the 429 straight back to the model with a clear, structured error and a wait-time hint, and let the agent decide whether to wait, switch tools, or tell the user. Retrying silently inside the server hides the real signal from the one thing (the agent's reasoning loop) that's actually in a position to stop calling the same tool five more times in the next two seconds.

Pick the layer you're actually rate-limiting

"Rate limit the server" is underspecified. There are at least three layers, and conflating them is how teams end up instrumenting one and getting blindsided by the other two:

Layer What it controls Good for
Transport / connection Requests per second, globally or per-session Simple, predictable backpressure; easy to reason about
Tool / token budget Cost-weighted limits per tool call Tools with wildly different execution cost (a 1-second lookup vs. a 10-second batch job)
Tenant / credential One bucket per (tenant, upstream API) pair Multi-tenant setups where one shared OAuth credential serves many customers

The tenant layer matters more than it looks. If one OAuth credential is shared across customers, the first agent to get throttled can trigger a cascade where every other session behind that same credential inherits the block — sometimes within twenty seconds of the first call. Per-tenant token buckets contain the blast radius; the credential model underneath decides how big that radius is in the first place. If you're issuing one shared API key to a backend that serves 40 tenants, your rate limiter needs to know about tenants, not just requests.

A token bucket you can actually reason about

Token bucket is the right primitive for most MCP tool-call limiting because it allows bursts (a legitimate agent batch-processing 10 items) while still capping sustained rate.

import time
import asyncio

class TenantTokenBucket:
    def __init__(self, capacity: int, refill_per_sec: float):
        self.capacity = capacity
        self.tokens = capacity
        self.refill_per_sec = refill_per_sec
        self.last_refill = time.monotonic()
        self.lock = asyncio.Lock()

    async def try_consume(self, cost: int = 1) -> bool:
        async with self.lock:
            now = time.monotonic()
            elapsed = now - self.last_refill
            self.tokens = min(
                self.capacity,
                self.tokens + elapsed * self.refill_per_sec
            )
            self.last_refill = now
            if self.tokens >= cost:
                self.tokens -= cost
                return True
            return False
Enter fullscreen mode Exit fullscreen mode

Two details that matter and are easy to skip:

  1. It's async and lock-guarded. MCP's streaming architecture means a single slow tool call can block the event loop if your limiter isn't async-aware — you'll end up deadlocking clients that are just waiting on a response, which looks like a hang, not a rate limit.
  2. cost is a parameter, not always 1. A tool that fans out into five upstream calls should consume five tokens, not one. Flat per-request limiting wastes budget on cheap calls and underprices expensive ones.

Returning a 429 the agent can act on

A rate limit is only as useful as the signal it sends back. The common failure is a bare 429 with no machine-readable guidance — the client has no idea whether to wait 500ms or 5 minutes, so it does something arbitrary, usually retrying immediately and making things worse.

Structure your rate-limit error so the agent (not a human reading logs) can parse it:

{
  "error": {
    "code": "rate_limited",
    "message": "Tool call rate limit exceeded for this session",
    "retry_after_ms": 4200,
    "scope": "tenant:acme-corp:tool:search_invoices"
  }
}
Enter fullscreen mode Exit fullscreen mode

The retry_after_ms field is doing the real work — it tells the agent's retry logic exactly how long to back off instead of guessing. And if an upstream API sends you a Retry-After header, treat it as a mandatory floor, not a suggestion. Some APIs enforce hard cool-downs (Airtable's 30-second window is a well-known example), and retrying before that window closes just resets the clock.

The backoff schedule itself

For the retries that do belong in your code (agent-to-tool retries, not the cascading server-side kind above), a schedule that shows up consistently in production guidance:

  • Start at 500ms
  • Double on each retry, capped at 30 seconds
  • Add ±20% jitter so multiple sessions recovering from the same outage don't all retry in the same instant (thundering herd)
  • Never go tighter than 500ms — anything faster just looks like a burst to the very limiter you're trying to respect

Catching stuck agents before they trip the limit

A rate limiter protects your server, but it doesn't fix the agent that's stuck in a loop — it just makes the loop cheaper. Worth adding separately: if a session calls the same tool with identical arguments more than a handful of times in a row, that's a strong signal of a stuck agent, not legitimate traffic. A real agent varies arguments between calls (different IDs, different queries); a stuck one repeats the exact same call. Flagging that pattern and returning a hard stop (not just a slower retry) is cheaper than waiting for the token bucket to catch it call-by-call.

What's in the pack

Rate limiting is one piece of a much longer list — session handling, timeout budgets, error taxonomy, and the pre-launch checks that catch this stuff before it's live. The rest is in the pack: the AI Agent Incident Postmortem & Permission-Scoping Template Pack has the incident-response templates for when a rate limit does get bypassed and something runs away anyway — scoped permission worksheets and a postmortem template built specifically for agent tool-call incidents, not generic infra outages.

If you're shipping an MCP server this month, the combination that actually holds up under load is: tenant-scoped token buckets, structured 429s with retry_after_ms, no server-side retries on agent-facing calls, and a separate stuck-loop detector that doesn't wait for the bucket to empty. Everything else is tuning.


Written with AI assistance and reviewed for accuracy.

Top comments (0)