DEV Community

orca_forge
orca_forge

Posted on Edited on Originally published at forge.workstyle.tech

Half of Free LLM APIs Were Dead After 2 Months

📝 Originally published (in Japanese) at forge.workstyle.tech.

Introduction to the Investigation of Free LLM API Providers

In July, I wrote an article titled "A summary of free LLM API providers that can be used without a credit card." Eight companies were listed, and it was suggested that by combining multiple providers and implementing a fallback mechanism, it would be possible to operate in production using only the free tier.

However, after two months, when I inspected the pipeline that was running with this configuration, I found that the success rate had dropped to 60%.

Upon investigating the cause, I discovered that multiple companies listed at the time had ceased to exist for various reasons, and not a single fallback was functioning. This article is a record of the full investigation, with all numbers obtained by actually hitting the APIs on September 7, 2026.

The Beginning: A 6-Company Configuration Was Essentially Operating as a Single Company

When I threw the same request at the router (LiteLLM) 10 times, the result was as follows:

Success 6 / Failure 4
Enter fullscreen mode Exit fullscreen mode

The contents of the failures were as follows:

litellm.APIError: CerebrasException - Payment required to access this resource.
  Received Model Group=free
  Available Model Group Fallbacks=None

litellm.NotFoundError: GeminiException - models/gemini-2.0-flash is no longer available
Enter fullscreen mode Exit fullscreen mode

The line Available Model Group Fallbacks=None was the key. Despite combining six companies, no fallback was occurring.

Therefore, I individually hit each company's API to confirm their status:

Provider Result
Groq ✗ 401 Invalid API Key
Cerebras ✗ 402 Payment required
OpenRouter ✗ 401 API key expired
Gemini ✗ 404 model no longer available
Mistral ✗ 429 Rate limit exceeded
Cohere ○ 200

Only one company was still alive. The reason 6 out of 10 attempts were successful was that LiteLLM's num_retries: 5 was retrying and eventually landing on Cohere. Although it seemed redundant, in reality, the entire load was on a single company.

The Way of Decay Was Not Uniform

This is the main point. I thought it would be settled with the phrase "the free tier ended," but each had a different reason. I categorized 13 candidates by actually hitting them:

1. The Free Tier Itself Became Paid — Cerebras

402 Payment required to access this resource. Visit your billing tab.
Enter fullscreen mode Exit fullscreen mode

The key is valid, and authentication passes, but billing is required. As of July, there was a free tier of "1M tokens/day". This alone clearly made the article's description incorrect.

2. The Service Has Ended — GitHub Models

GET https://models.github.ai/catalog/models
→ HTTP 410 Gone
Enter fullscreen mode Exit fullscreen mode

410 is a status code that means "permanently gone". It was considered as a candidate, but now there is no room for consideration.

3. The Model Was Completely Replaced — Groq

After reissuing the key, the error changed from 401 to 404.

404 The model `llama-3.3-70b-versatile` does not exist or you do not have access to it.
Enter fullscreen mode Exit fullscreen mode

The free tier was still alive, but the model had disappeared. The current catalog has 9 models, and none of the Llama series remains.

allam-2-7b / canopylabs/orpheus-* / groq/compound / groq/compound-mini
openai/gpt-oss-120b / openai/gpt-oss-20b / qwen/qwen3.6-27b / qwen/qwen3.8-27b
Enter fullscreen mode Exit fullscreen mode

4. Only the Model Name Became Invalid — Gemini

404 NOT_FOUND: This model models/gemini-2.0-flash is no longer available.
    Please update your code to use models/gemini-3.6-flash
Enter fullscreen mode Exit fullscreen mode

Here as well, the free tier was still alive. It was simply a matter of not keeping up with the model name's generation change.

As a countermeasure, I changed it to the alias gemini-flash-latest. It passes in actual measurements, and this way, it will automatically follow the next name change.

5. The Key Had an Expiration Date — OpenRouter

401 API key expired.
Enter fullscreen mode Exit fullscreen mode

Looking at the dashboard, Expires: Expired / Last Used: Never / Key limit: unlimited. It expired due to the expiration date, not the usage amount. When creating the key, you can choose "No expiration," so for resident purposes, you should always choose that.

6. "No Card" But Balance Is Necessary — Z.ai

429 Insufficient balance or no resource package. Please recharge.
Enter fullscreen mode Exit fullscreen mode

The key issuance does not require a card, but you cannot make a single request without charging. I have seen it introduced as a "free tier without a card," but the reality was different.

7. The Catalog Is Out of Sync with Reality — NVIDIA NIM

The model list API returns 81 items. I hit all of them.

Success           13 items
HTTP 404       55 items   ← Catalog items that do not exist
timeout/exception    7 items
503 / 500 / 400  6 items
Enter fullscreen mode Exit fullscreen mode

Furthermore, among the 13 living items, 3 were content safety classifications, 2 were translation-only, and 1 was image input. Only 5 could be used for general chat.

The explanation that "81 models can be used for free" would lead to implementation and then discovery of the issue.

8. Essentially Discontinued — SambaNova

Here are the results of hitting all 7 models:

DeepSeek-V3.1 / V3.2 / Meta-Llama-3.3-70B / gpt-oss-120b → 402 A payment method is required
MiniMax-M2.7 / M3                                        → 429 high demand
gemma-4-31B-it                                           → ○ 200(but 12,021ms)
Enter fullscreen mode Exit fullscreen mode

Only one model survived for free, and it took 12 seconds. Compared to Groq's 220ms, it's 55 times slower. While it's not incorrect to say there's a free tier, the judgment for practical use is different.

9. Card Was Mandatory — Nebius / Scaleway

Nebius's registration screen requires billing details, including name, address, and card information (with a $0 authorization for card verification). Scaleway also states in its official documentation that "to receive regular rate limits, you need to register a card and complete KYC."

Neither meets the condition of "no credit card required."

10. Token Permission Design Had a Pitfall — HuggingFace

403 This authentication method does not have sufficient permissions
Enter fullscreen mode Exit fullscreen mode

Upon checking the token, only repo.content.read was attached. For inference, inference.serverless.write (or "Make calls to Inference Providers" in the UI) is necessary.

After rechecking the permissions, it passed. The issuance of the key and the key being usable are separate; it's a straightforward point, but I didn't realize it until I saw the error message.

The Real Reason Fallbacks Didn't Work

This was the biggest lesson.

LiteLLM's settings were as follows:

router_settings:
  routing_strategy: simple-shuffle
  num_retries: 5
  allowed_fails: 1
  cooldown_time: 60
Enter fullscreen mode Exit fullscreen mode

At first glance, it seems robust because it retries five times. However, in actual measurement, it failed the moment it hit Cerebras.

402 (payment required) and 404 (model not found) are not targets for retries. Retries and fallbacks assume failures like 429 (rate limit) or timeouts, which might resolve if waited on. 402 and 404 are failures that won't resolve with waiting, so they immediately return as failures.

In other words,

Leaving dead providers in the pool will continue to generate failures at a certain probability.

If one out of six companies is permanently dead, simple calculation shows that 1/6 of the requests will fail. The notion that "having many combined is safe" only holds if all companies are alive.

In reality, after removing Cerebras from the pool, the success rate changed as follows:

Success rate transition: 6/10 (60%) → 15/15 → 20/20 → 24/24 → 28/28
Enter fullscreen mode Exit fullscreen mode

Conclusion: Summary Articles Decay in 2 Months

Since the article I wrote myself became half unusable after two months, this serves as a warning to myself as well. For readers, I'll conclude with three practical points:

1. Create a deathwatch first.
Before increasing providers, creating a mechanism to detect and remove dead providers is more effective. This time, more than updating model names, removing the dead company (Cerebras) had a greater impact on the success rate. A simple script that sends one request to each company is enough.

2. Fallbacks are not omnipotent.
They are effective for 429 (rate limit) or timeouts. 402 (payment required), 404 (model disappeared), and 401 (key expired) are not covered by retries. These are failures that won't resolve with waiting, so they continue to generate failures until a human notices and removes them.

3. Model names should not be fixed; confirm before adding.
If there's an alias like latest, use it. If not, before adding to the pool, hit it once to confirm that a response is returned and that reasoning_content is not attached. Even within the same series, different versions can behave differently.


Note: The numbers in this article are all from actual measurements on September 7, 2026. The conditions for free tiers, provided models, and rate limits can change over a few months, as seen this time. When introducing them, please always check the latest status on each company's official documentation. This article, too, will likely be half incorrect two months later.

Top comments (0)