DEV Community

ushiro
ushiro

Posted on

8 of the 24 Codes Behind HTTP 429 Are Not Rate Limits. Backoff Will Never Clear Them.

The retry-on-429 loop is the most copied piece of LLM plumbing there is. Catch the status, sleep,
double the sleep, try again. Every SDK ships one and every tutorial recommends it.

I read the published error documentation of 9 model vendors and pulled out every code they put
behind HTTP 429. There are 24 of them. Here is what they actually mean — the right-hand column is
the vendor's own stated cause, shortened to fit; nothing below is my paraphrase of their intent.

anthropic  rate_limit_error                    A rate limit, OR the monthly spend cap. Same code.
aws        ThrottlingException                 Exceeding the account quotas for Amazon Bedrock.
aws        ModelNotReadyException              The model is not ready to serve inference requests.
cohere     Too Many Requests                   The rate limit has been exceeded.
deepseek   Rate Limit Reached                  You are sending requests too quickly.
google     rate_limit_exceeded                 Per-minute or per-second request or token limit.
google     too_many_requests                   Too many requests in a short period of time.
google     quota_exceeded                      You have exceeded your daily quota.
groq       Too Many Requests                   Too many requests in a given timeframe.
mistral    Too many requests                   Exceeded the rate limit for your subscription tier.
openai     Rate limit reached for requests     You are sending requests too quickly.
openai     Slow down                           Your request rate increased too quickly.
qwen       Throttling                          (no cause published)
qwen       Throttling.AllocationQuota          (no cause published)
qwen       Throttling.BurstRate                Call frequency spiked suddenly.
qwen       Throttling.RateQuota                Call frequency (RPS/RPM) triggered rate limiting.

openai     Credit balance exhausted            Your organization has no prepaid credits remaining.
openai     Organization spend limit reached    Your organization reached its enforced spend limit.
openai     Organization usage limit reached    Reached its OpenAI-assigned usage limit.
openai     Project spend limit reached         Your project reached its enforced spend limit.
qwen       CommodityNotPurchased               Workspace subscription not purchased.
qwen       PrepaidBillOverdue                  Workspace prepaid bill expired.
qwen       PostpaidBillOverdue                 Model inference service expired.
qwen       BudgetLimitExceeded                 The configured budget has been exhausted.
Enter fullscreen mode Exit fullscreen mode

The blank line is mine — the vendors publish one list. Everything under it is documented as a
billing or account state, not a rate.

The vendors say so themselves

I did not classify these. The vendors did, in the remedy they print next to the code:

openai / Credit balance exhausted — "Add credits to continue using the API."

openai / Project spend limit reached — "Increase or remove the spend limit in your project
settings."

qwen / CommodityNotPurchased — "Purchase workspace service first."

qwen / BudgetLimitExceeded — "You will be unable to make further API calls until the budget
limit is increased or reset."

All four of the OpenAI ones and all four Qwen ones need a human to spend money or raise a
limit. No amount of sleeping changes them. A loop that retries a 429 six times with exponential
backoff spends about a minute confirming that your credit card is still not on file, and then
raises the same error it would have raised immediately — except now it looks like a transient
capacity problem in your logs.

The same condition is a 402 at four other vendors

This is the part that makes the status code useless as a signal.

condition vendor status
out of money DeepSeek 402 Insufficient Balance
out of money Anthropic 402 billing_error
out of money Cohere 402 Payment Required
out of money Google 402 payment_required
out of money OpenAI 429 Credit balance exhausted
out of money Qwen 429 PrepaidBillOverdue

"You have run out of balance" is DeepSeek's 402. "Your organization has no prepaid credits
remaining" is OpenAI's 429. Same sentence, different half of the 4xx range.

If you run more than one vendor behind one client — which is the entire point of most gateway
code — a handler that branches on the status code will fail-fast on DeepSeek and retry-loop on
OpenAI, for the identical business condition.

Google, which serves both, spells out the consequence in the remedy on its own 402:

"Add credits to your billing account, or turn on auto-reload. Don't retry: the request won't
succeed until credits are added.
"

That instruction is correct, and it is printed on the status code that nobody retries anyway. The
codes that actually end up inside retry loops are the 429s, and that is where it is missing.

Qwen does it inside a single vendor. "Workspace subscription not purchased" is
CommodityNotPurchased, a 429. "Model Studio service is not activated" is
AccessDenied.Unpurchased, a 403.

Anthropic puts both meanings on one code

Branching on the vendor's code instead of the status is the obvious fix, and for one vendor it is
not enough. rate_limit_error is Anthropic's only 429, and its documented cause now reads:

"Your organization has hit a rate limit, reached its usage tier's monthly spend cap, or reached
a spend limit on the Claude Code workspace. A tier spend-cap 429 has no retry-after header and
keeps failing until access resumes.
"

One code, one status, and the two cases need opposite handling. Anthropic tells you how to tell
them apart — the absence of Retry-After — which means the discriminator is a missing header.

Two more that aren't what they look like

aws / ModelNotReadyException is a 429 and has nothing to do with your request rate. The
model is warming. AWS documents that its own SDK "will automatically retry the operation up to 5
times" — so if you wrapped the SDK in your own retry loop, you have 5×N attempts and a timeout
budget you did not plan.

google / quota_exceeded is a 429 that says "You have exceeded your daily quota." Google's
own remedy is "Wait until the quota resets or request a quota increase." An exponential backoff
capped at a few minutes will exhaust itself against a counter that resets tomorrow. Google
publishes a separate code — rate_limit_exceeded — for the per-minute case, and that one's remedy
does say "Wait and retry with exponential backoff." One vendor, one status, two opposite
instructions.

What the 24 actually split into

14   clear on their own if you back off
 1   clears at the daily reset, not on your timer      google / quota_exceeded
 1   clears when the model finishes warming            aws / ModelNotReadyException
 8   do not clear until a person spends money
Enter fullscreen mode Exit fullscreen mode

The first line includes Anthropic's, which is in it only half the time.

Only 3 of the 9 vendors mention the Retry-After header anywhere in their error documentation —
OpenAI, Mistral, and Anthropic, whose mention is that the spend-cap case does not send one. The
header may well be on the wire elsewhere; it is not in the docs you would read to write the
handler.

What I would actually do about it

Branch on the vendor's code, not the status. Every vendor here returns a machine-readable code
in the body. The status tells you which half of the 4xx range someone chose; the code usually
tells you what happened — with Anthropic's as the exception you have to special-case.

Give the money errors their own path. They are not retryable and they are not transient —
they are an outage with a billing cause, and they should page someone rather than dissolve into a
retry counter. Alerting on "429 rate" makes exactly this class invisible.

Cap unknown 429s, don't trust them. Treating an unrecognized 429 as retryable is fine.
Retrying it 6 times with a 60-second ceiling because "429 means slow down" is what turns a
five-second failure into a minute of held connections.

What this does and does not measure

This is the vendors' published documentation, read on 2026-09-28 — not traffic. If a vendor
returns a code it never documented, I do not have it, and several of these docs are visibly
incomplete: Qwen publishes no cause text at all for two of its four throttling codes.

It moves, too. Six days before publishing this I counted 21 codes, not 24: Google has since
documented too_many_requests and a 402, OpenAI a second rate code, Qwen a budget one, and
Anthropic rewrote its single cause to name the spend cap. A post about error semantics going
stale in under a week is not an aside — it is the same reason the handler you wrote last year
is describing a set of conditions that no longer matches.

That cuts one way only. Every code above is one the vendor chose to write down, and the eight that
need a credit card are documented as clearly as the ones that need patience. They are behind the
same three digits anyway.


I track these pages and keep what they overwrite. The 183 documented error codes across 9
vendors, one page each, with the vendor's own cause and remedy:
aichangewatch.com/errors

Top comments (1)

Collapse
 
makeev profile image
Mikhail Makeev •

The spend-cap 429 that "has no retry-after header" is the detail I'd build a handler around, which only works if the header is honest. I shipped a limiter where it wasn't. The minute window was sliding, and Retry-After was computed as if the whole overage decayed at once, but the requests that caused it sat in the current minute and didn't. So after a burst the header said 2 seconds, the SDK waited 2, got another 429, waited again and gave up, while the window actually reopened about 30 seconds later. Fixed that day, now it names the real wait. Did any vendor's 429s come with a Retry-After that you checked against how long the wait actually was?