Every free LLM API says yes right up until the moment it says no. The no arrives as an HTTP 429, usually forty requests into a demo or halfway through a batch job, and most tutorials end the story at "add a retry loop." Here is what actually works when you are building on free tiers, provider by provider and layer by layer.
First, find out which limit you hit
A 429 is not one problem, it is four. Free tiers meter you on several axes at once, and the fix depends on which one fired:
- Requests per minute (RPM): you are asking too often. Smoothing helps, capacity does not.
- Tokens per minute (TPM): a few long prompts ate the whole minute. Shorten contexts, cap max_tokens.
- Requests per day (RPD): you are done until the reset. No code fix exists, only spreading work across days or providers.
- Concurrency: two in-flight requests when the tier allows one. Queue, do not parallelize.
Read the response body and headers before you touch your code. Groq sends x-ratelimit-remaining-requests (daily) and x-ratelimit-remaining-tokens (per minute) plus retry-after, so the 429 tells you which window is empty and when to come back. The Gemini API returns 429 RESOURCE_EXHAUSTED and names the violated window in details[].quotaId, with retryDelay set for per-minute limits; its daily quotas reset at midnight Pacific. OpenRouter free-model caps jump from a small daily allowance to a larger one once the account has purchased credits, and its error text and x-ratelimit-reset-requests header say which window closed. Some providers, sadly, send nothing: if a 429 comes back bare, treat every window as closed and back off to the hour.
Then, retry like it means something
Three rules that cover most cases:
- Honor
retry-afterwhen present. It is the provider telling you the truth; guessing is worse. - Exponential backoff with jitter when it is not: 1s, 2s, 5s, 10s, plus random offset, capped at a few minutes. Retrying a daily cap in a loop accomplishes nothing, so stop after 3 to 5 attempts and surface the error.
- Count tokens, not requests. A retry loop tuned for RPM will still fail when TPM is the real wall, because the retry itself burns more tokens per minute.
Then, stop depending on one well
One provider's free tier is one quota pool. The structural fix is to treat free tiers as a pool of pools:
- Keep several OpenAI-compatible keys (Gemini, Groq, NVIDIA NIM, OpenRouter, Cloudflare Workers AI, Cerebras and others each publish their own caps) and fail a request over to the next provider when one window closes. Because all of them speak the OpenAI chat format, your client code does not change, only the base URL and key.
- Route by limit shape, not just availability. A nightly summarization job wants RPD headroom (Groq's daily allowances, or Gemini's daily caps). An interactive prototype wants RPM smoothness. A one-shot extraction can burn a per-token-limited tier without hurting anything.
- Watch the terms. Several free tiers (Cohere's trial key, Mistral's La Plateforme free tier) are explicitly for evaluation, not production. OpenRouter's terms prohibit reselling API access. Read them before you build a product on top.
Hand-rolling this means a provider abstraction, a per-provider window tracker, a queue, and health checks for every one of them, which is most of a side project by itself. If you would rather not own that layer, there are routers that do it: I build FreeLLMAPI, an open-source, self-hosted router that keeps your provider keys, handles 429s and retries with backoff, falls back to other providers and models, and exposes one OpenAI-compatible endpoint at the end. The same job can be done by hand with the playbook above; the router just means you write it once and self-host it.
And log the shape of your traffic
Whatever you choose, log per-request: provider, model, latency, status, and the rate-limit headers. A week of logs tells you which window you actually bump, and that single fact decides whether you need smoothing, shorter prompts, a second provider, or honestly, the $5 paid tier. Most "we need more quota" problems turn out to be "one long prompt repeated in a loop" problems. The 429 is information. Read it before you reroll the key.
Top comments (0)