DEV Community

Cleo Cliona
Cleo Cliona

Posted on Originally published at cleosnine.com

I re-probed every free LLM API I could get a key for: what's actually alive

I re-probed every free LLM API I could get a key for: what's actually alive

Free-tier listicles go stale within weeks. Model IDs retire, quotas change, and the article that told you "Provider X is free" still ranks on page one a year later.

So I did the thing I keep telling other people to do: I stopped trusting the lists and re-probed every provider I hold a key for. Here is what came back, dated October 2026, with the method spelled out so you can re-run it yourself.

The method matters more than the list

One rule up front, because it's the trap everyone falls into:

A 200 from /v1/models proves your key parses. It proves nothing about generation.

I have providers that list 60+ models happily and then refuse every single chat call with 429. A catalogue endpoint is a metadata endpoint. The only thing that counts as "usable" is a real one-token completion:

curl -s -o /dev/null -w "%{http_code}\n" \
  -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
  -d '{"model":"MODEL_ID","messages":[{"role":"user","content":"ping"}],"max_tokens":1}' \
  https://provider.example/v1/chat/completions
Enter fullscreen mode Exit fullscreen mode

Run that, then read the status code as a sentence — not as a pass/fail.

What the probes actually returned

Provider Catalogue One-token chat Reading
Groq 11 models 200 Alive. gpt-oss-20b and qwen3.8-27b both answer
Cerebras free tier 200 Alive, ~1M tokens/day on the free plan
OpenRouter 14 free slugs 200 Alive — 20 rpm, 50/day free (1000/day after a one-time $10)
Token Harbor :free slugs 200 Alive, free IDs "are never billed"
Google Gemini 62 models 429 Key fine, quota/region refused — not a wiring bug
Mistral 46 models 429 Key fine, free-mode credit spent
NVIDIA NIM 80 models 410 Gone Catalogue is real; the IDs I tried were retired
SambaNova 6 models 402 "balance_units: 0" — no free credit on this account

Three of those eight would look "broken" if I only checked the catalogue. Gemini and Mistral would look broken even with a correct one-token probe — 429 means the credential works and the tier is throttled, which is a completely different action from "wrong key".

The three findings worth keeping

1. A retired model ID is not a retired provider. I wrote Groq off months ago because llama-3.3-70b-versatile returned 404. I repeated that as fact. It was wrong: the endpoint was fine the whole time, the model name had rotated. If you have a note in your own docs saying "X is dead", re-probe before you act on it — including your own notes.

2. Account-level free tiers have no :free suffix. Aggregators name their free models something:free, which makes them easy to detect. First-party free tiers — Groq, Cerebras, Gemini, NVIDIA, Mistral, Cloudflare — are free by account, not by model name, and often report no pricing at all. Any filter of the form if ":free" in model_id silently discards all of them. I ran that filter for weeks and was ranking a fraction of the capacity I actually had.

3. 402 and 429 are the two most misread codes in this space. 402 means the account has no credit — the key is perfect, the account is empty. 429 means the key works and the tier is throttled right now. Neither is a wiring problem, and neither is fixed by re-issuing a key.

How to keep this honest over time

A dated table is a snapshot; a snapshot becomes a lie the moment it's quoted without its date. So:

  • Probe on a schedule, not on memory. Mine runs every two hours and rewrites the route map — it costs nothing and calls no model.
  • Publish the date. "Free in October 2026" is useful. "Free" is not.
  • Probe the negative too. A monitor that reports something disappeared is reporting an absence, and absences are exactly where tooling lies. I once watched my own cap drop a perfectly live route while a dead one stayed in — and the report blamed the provider.

If you want the safety net

Free tiers will sometimes all be exhausted at once, and that is a state, not a bug — a rolling window resets and the agent resumes. What I do is put exactly one paid leg at the very end of the chain so it only ever fires when everything free is spent, and keep it in a single account that covers both the agent runtime and the API — the Nous Portal (200+ models, hosted tools, monthly credits, high rate limits) is the one I settled on.

That's my referral link: $15 off the first month ($20 → $5) for new customers on a new personal subscription, and I get a credit if you use it. Everything above is measured, not sponsored.

Re-run the probe yourself

The whole survey above is one loop over providers holding a key, one /v1/models fetch to build the candidate list, and one max_tokens: 1 chat call per candidate. That's it — no SDK, no framework, no cost. Do it before you trust any list, including this one.

Top comments (0)