DEV Community

unicodefawn
unicodefawn

Posted on

Free LLM API tiers in 2026: the actual rate limits, and why they keep changing

Disclosure: I'm a co-founder of nopaywall, the catalog this post draws from.

If you've ever built a side project on a "free" LLM API, you know the problem: the free tier is real, but the limits are buried in docs, they change without notice, and the blog post you found last year is wrong now. I'll go through what the free tiers look like at the moment and how we keep the data current.

The numbers (as listed on 2026-10-07, rounded)

Provider Free limit Notes
Groq ~30 RPM, ~1K requests/day, ~8K tokens/min GPT-OSS 120B/20B, Qwen, Whisper
Google AI Studio (Gemini API) ~15 RPM, ~500 requests/day, ~250K tokens/min
Mistral Codestral ~30 RPM, ~2K requests/day code-focused
Hugging Face Inference ~300 requests/hour
Puter.js ~30 requests per 10 s, 3 concurrent runs from the browser

These are caps, not promises. Eligibility can vary by country, so check each provider's terms before you ship anything that depends on them.

Practical takeaways

  • Prototype vs. production. Around 500 to 1K requests a day is enough for a demo, a personal tool or a hackathon. A public app with real users will hit it on day one.
  • Watch tokens per minute, not just RPM. An 8K TPM limit allows only a few long-context calls per minute, even if you're well under the request cap.
  • Build a fallback chain. Since several providers serve the same open model families (Llama, Qwen, DeepSeek, Mistral, Gemma, GPT-OSS…), you can put an OpenAI-compatible router in front and fail over when you hit a 429.
  • Pin the date. When you write down a limit in your README, write the date you checked it next to it.

How the catalog stays current

nopaywall lists 100+ free API providers, with a model finder so you can search by model family instead of by vendor. Collection and checking are automated with source evidence: sources (GitHub repos and topic lists) are collected, then the official pricing or docs page is inspected. A card goes public only when it has a reachable destination and enough evidence for the limit. API providers and limits are re-checked within 72 hours. Cards are removed after a confirmed expiry or a failed verification. Unclear records stay off public pages.

Free, freemium, open-source and paid access are labeled separately. That way "free trial, card required" doesn't get mixed up with "free tier".

Links

Top comments (0)