DEV Community

Tariq Nasser
Tariq Nasser

Posted on

Free LLM API Tiers in October 2026: What's Left and How I Chain Them

I run a handful of internal tools that call a language model a few hundred times a day: a changelog summarizer, a ticket classifier, a script that turns noisy logs into a readable incident timeline. None of them justify a monthly bill, so for the past year they have lived entirely on free LLM API tiers. That works, but only if you treat "free" as a set of quotas you have to manage, not a gift.

The lists that rank for this topic are mostly snapshots from spring. Since then, providers have been dropped, models have been retired and limits have moved. This post is what the landscape looks like on 6 October 2026, the catches I've had to design around, and the small fallback chain I use so one provider's daily cap doesn't take a tool down.

One scoping note before the details: every tier below is text-only, or close to it. When a project needs image or video generation I don't bother hunting for free quotas; I send those calls to Synexa, which runs models like FLUX Schnell behind a single predictions endpoint. Everything else in this post is about text.

Three kinds of "free"

Before comparing numbers, it helps to separate what providers mean by free, because the failure modes differ.

  1. Permanent quota. A fixed number of requests or tokens per minute and per day that resets. Groq, OpenRouter's :free models, NVIDIA NIM and Google Gemini work this way. When you run out you get a 429 and wait.
  2. Monthly credit allowance. A dollar amount that refills monthly. Mistral's free plan carries $10/month in API credits; Hugging Face gives free users $0.10/month of Inference Provider credit. When it's gone, it's gone until the month rolls over.
  3. Anonymous or keyless access. No signup at all. OVHcloud AI Endpoints allows 2 requests per minute per IP per model without a key, Kilo Code's gateway serves its free pool at 200 requests per hour per IP, and LLM7.io lets anonymous callers reach its turbo models. Great for prototypes, fragile for anything you'd page someone over.

Trial credits that expire don't count, and I've left them out.

The free LLM API landscape as of October 2026

These figures come from the community-maintained awesome-free-llm-apis list, cross-checked against Groq's and OpenRouter's own rate-limit pages where I could. Every one of these endpoints is OpenAI SDK-compatible unless noted.

Provider Example free models Free limit The catch
Groq openai/gpt-oss-120b, openai/gpt-oss-20b, qwen/qwen3.8-27b 30 RPM, 1K RPD, 8K TPM, 200K TPD per model Limits apply per organization, not per key
OpenRouter nvidia/nemotron-3-super-120b-a12b:free, google/gemma-4-31b-it:free and others 20 RPM, 50 RPD (1,000 RPD after $10 in credits) Free providers may log prompts
NVIDIA NIM openai/gpt-oss-20b, nvidia/nemotron-3-super-120b-a12b 40 RPM, 10,000 RPD per model Needs NVIDIA Developer Program membership
Google Gemini Gemini 3.x and 2.5 Flash / Flash-Lite, 2.5 Pro Google no longer publishes per-model free limits; check AI Studio Free-tier prompts may be used to improve products (not in EEA/UK/CH)
Mistral Mistral Medium 3.5, Small 4, Large 3, Codestral $10/month in credits Free-mode prompts may train models unless you opt out
Cloudflare Workers AI @cf/meta/llama-3.3-70b-instruct-fp8-fast, @cf/openai/gpt-oss-120b 10,000 Neurons/day, shared across all models Not OpenAI-style; uses its own /ai/run URL
Cohere Command A, Command R+, Aya Expanse 32B 20 RPM, 1,000 calls/month Trial key is non-commercial only
Z AI (Zhipu) GLM-4.7-Flash, GLM-4.6V-Flash 1 concurrent request GLM-4.5-Flash retirement announced
ModelScope Qwen3.5-35B-A3B, Qwen3.5-27B 2,000 RPD total, up to 500 per model Alibaba Cloud binding plus real-name verification
SiliconFlow Qwen/Qwen3-8B 1,000 RPM, 50,000 TPM Real-name ID verification since May 2026

Compared with the March version of the same list, Cerebras, GitHub Models and Kluster AI no longer appear in it, and several newer entries (Aion Labs, Kilo Code, OVHcloud, ModelScope) do. Groq also shut down qwen/qwen3.6-27b on 14 September 2026 in favor of qwen/qwen3.8-27b. If you copied model IDs from an older article, check that they still resolve before you blame your own code.

The catches I actually design around

Daily caps hurt more than per-minute caps. OpenRouter's free models default to 50 requests per day per account until you've bought $10 of credits, at which point the ceiling becomes 1,000. Per-minute limits only slow you down; a daily cap stops you until midnight UTC. My rule: a provider with a low daily cap goes last in the chain, never first.

Know which header means what. Groq sends x-ratelimit-remaining-requests on every response, and its docs are explicit that this always refers to requests per day, while x-ratelimit-remaining-tokens refers to tokens per minute. retry-after only shows up on an actual 429. OpenRouter is the other way round: successful responses carry no rate-limit headers, so you query the key instead:

curl https://openrouter.ai/api/v1/key \
  -H "Authorization: Bearer $OPENROUTER_API_KEY"
# look at data.free_model_daily_requests -> { used, limit, remaining }
Enter fullscreen mode Exit fullscreen mode

Shared pools are easy to drain by accident. Cloudflare's 10,000 daily Neurons are shared across every Workers AI model on the account and reset at 00:00 UTC. Going over does not bill you; the request fails. A runaway embedding job can quietly starve your chat model. Groq's limits are also per organization, so two services sharing one org share one budget.

"Free" sometimes means "you are the training data." Google says free-tier prompts may be used to improve products, except for users in the EEA, Switzerland and the UK. Mistral's free mode may train on inputs unless you opt out. OpenRouter notes that free providers may log prompts. I keep anything containing customer data off these tiers entirely, and that one rule decides which tools can use them at all.

Regional and identity gates. SiliconFlow and ModelScope require real-name verification, which in practice means Chinese documents or a support ticket. Gemini's terms say that if you make an API client available to users in the EEA, Switzerland or the UK, you must use paid services. If you have users there, read that section before building on the free quota.

The SDK's own retries work against you. The OpenAI Python SDK retries 429s by default, sleeping between attempts. On a free tier that's exactly wrong when another provider has capacity right now. I set max_retries=0 and let my own chain decide where the next attempt goes.

A fallback chain that survives 429s

Because Groq, NVIDIA NIM and OpenRouter all speak the OpenAI Chat Completions format, one client class covers all three. The order reflects headroom: NVIDIA's 10,000 RPD first, Groq's 1,000 RPD second, OpenRouter's 50 RPD last as an emergency spare.

import os
from openai import OpenAI, RateLimitError, APIStatusError, APIConnectionError

# (name, base_url, env var holding the key, model id)
PROVIDERS = [
    ("nvidia", "https://integrate.api.nvidia.com/v1", "NVIDIA_API_KEY", "openai/gpt-oss-20b"),
    ("groq", "https://api.groq.com/openai/v1", "GROQ_API_KEY", "openai/gpt-oss-120b"),
    ("openrouter", "https://openrouter.ai/api/v1", "OPENROUTER_API_KEY",
     "nvidia/nemotron-3-super-120b-a12b:free"),
]

# max_retries=0: fail fast and move to the next provider instead of sleeping
CLIENTS = [
    (name, OpenAI(base_url=url, api_key=os.environ[env], max_retries=0, timeout=60), model)
    for name, url, env, model in PROVIDERS
    if os.environ.get(env)
]


def chat(messages, **kwargs):
    errors = []
    for name, client, model in CLIENTS:
        try:
            resp = client.chat.completions.create(model=model, messages=messages, **kwargs)
            return name, resp.choices[0].message.content
        except RateLimitError:
            errors.append(f"{name}: 429")
        except APIStatusError as e:  # 402 on OpenRouter means a negative balance
            errors.append(f"{name}: {e.status_code}")
        except APIConnectionError:
            errors.append(f"{name}: connection error")
    raise RuntimeError("every free tier refused: " + ", ".join(errors))


if __name__ == "__main__":
    provider, text = chat(
        [{"role": "user", "content": "Summarize: deploy failed, migration 042 timed out."}],
        max_tokens=200,
    )
    print(f"[{provider}] {text}")
Enter fullscreen mode Exit fullscreen mode

Two things I learned the hard way. First, return the provider name with every answer and log it; when output quality shifts, that log tells you whether a fallback happened. Second, don't treat every model in the chain as interchangeable. A 20B model and a 120B model will not give the same answers, so I only chain models I've checked on my own prompts, and I keep prompts plain enough that the weakest model in the chain still copes.

Which tier I'd pick for what

  • Batch jobs and classifiers: NVIDIA NIM, for the daily headroom.
  • Interactive tools where latency matters: Groq first, with the chain behind it.
  • Trying many models quickly: OpenRouter's :free variants; one key, many models, and openrouter/free routes to whatever free model is available.
  • Long-context or multimodal input: Gemini, after checking your quota in AI Studio and the data-use terms for your region.
  • Zero-signup demos: OVHcloud's anonymous tier or Kilo Code's gateway, accepting that they can change without notice.

FAQ

Is there a completely free LLM API with no credit card?
Yes. Groq, OpenRouter's :free models, NVIDIA NIM, Gemini, Mistral's free mode and Cloudflare Workers AI all start without a card. OVHcloud's anonymous tier and Kilo Code's gateway don't even need a key.

Which free LLM API has the highest daily limit?
Among the OpenAI-compatible options, NVIDIA NIM lists 10,000 requests per day per model, well above Groq's 1,000 and OpenRouter's default 50. SiliconFlow's per-minute numbers are higher, but it needs ID verification.

Can I use a free LLM API in production?
For low-stakes internal tools, yes, with a fallback chain and no sensitive data. For anything customer-facing, read the terms: Cohere's trial key is non-commercial, and Gemini's terms require paid services for apps serving EEA, UK or Swiss users.

Why do I get 429 errors on OpenRouter free models so quickly?
Free variants are capped at 20 requests per minute and 50 per day until your account has bought at least $10 in credits. Call GET /api/v1/key and check free_model_daily_requests to see how many you have left.

Wrapping up

Free tiers in late 2026 are generous enough to run real internal tools, as long as you plan around daily caps, keep private data off them, and fail over instead of retrying in place. The landscape will keep shifting, so check model IDs against the provider's own docs every few months. And when a project grows past text, I keep the same one-endpoint habit for media and route image and video jobs through Synexa rather than juggling another set of quotas.

Top comments (0)