Getting 429s from Groq and just sleeping a few seconds before retrying? That works until it doesn't.
import time, requests
resp = requests.post("https://api.groq.com/openai/v1/chat/completions",
headers={"Authorization": f"Bearer {KEY}"},
json=payload)
if resp.status_code == 429:
time.sleep(5) # blind retry. this is the bug.
resp = requests.post(url, headers=headers, json=payload)
Here is what the Groq docs say that most retry loops ignore: every response carries x-ratelimit-* headers, and retry-after appears ONLY on 429 responses. The wait time you are guessing at is sitting right there in the 429 response, and the loop above threw it away.
The headers that matter:
- x-ratelimit-limit-requests: your RPD (requests per day) cap
- x-ratelimit-limit-tokens: TPM cap
- x-ratelimit-remaining-requests / -remaining-tokens: how much headroom is left
- x-ratelimit-reset-requests / -reset-tokens: when the window resets
- retry-after: set only on 429s. the authoritative wait time. use it.
The dead ends we keep seeing in real agent runs:
- Fixed sleep hammering an org-level limit. Groq enforces limits at the org level, not per user. One noisy retry loop can starve the whole org while it burns through attempts.
- Watching TPM while OTPM trips you. Some accounts get an input/output token split (ITPM/OTPM). A huge output can trip OTPM even when total TPM looks fine, and your loop sees "tokens are fine" and retries straight into the same wall.
- Cached tokens don't count toward limits. A high cache hit rate lowers effective pressure, which is great for throughput but confusing when the usage numbers don't add up.
What actually works: read retry-after from the 429 and back off to exactly that window. Check your real numbers at console.groq.com/settings/limits instead of guessing.
Full notes on the gotchas, including the ITPM/OTPM split and the org-level trap:
https://vectle.com/skills/skl_2mbinMsdstsavd9J2DD2jA
Top comments (0)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.