DEV Community

Cover image for 9 free LLM APIs compared: rate limits, card rules and the catch (2026)
aifree.dev
aifree.dev

Posted on

9 free LLM APIs compared: rate limits, card rules and the catch (2026)

You can prototype on LLMs without paying anything - but "free tier" means very different things depending on the provider. Some give permanent rate-limited access, some give a monthly credit, and some only free a handful of models.

Here are nine LLM APIs that need no credit card at signup, with what you actually get and the catch for each. Terms were last checked on September 13, 2026; free lineups rotate, so check the provider page before you build on one.

The short version

API What is free Card The catch
Google Gemini Flash models + Gemini 2.5 Pro, rate-limited No No image (Nano Banana), Veo, Imagen or Pro previews
Groq Hosted open models, OpenAI-compatible No ~30 RPM / 1K requests per day on many models
OpenRouter A rotating set of free models No 50 requests/day until you buy 10 credits
NVIDIA NIM Four preview models No Needs a verified NVIDIA account; no published limits
Mistral $10/month in API credits No Limited messages and coding sessions on the Free plan
Cohere Rate-limited trial key No No production or commercial use
Z.AI GLM Flash models at $0 No Free pricing only on the Flash models
LLM7.io Up to 1M tokens/day No Input + output count together; needs a free token for the full 1M
Pollinations Text, image, audio, video over plain HTTP No 1 request per 15 s anonymous, watermarks possible

Google Gemini Developer API

Free input and output tokens within per-model rate limits across the Gemini Flash models, Gemini 2.5 Pro, Gemma 4, embeddings, and TTS/live previews, via the Gemini API and Google AI Studio. Google Search grounding is free up to 500 requests per day on Gemini 2.5 models.

The catch: image generation (Nano Banana), Veo, Imagen, Lyria and the Pro previews are not on the free tier.

Groq

Every account starts on the Free plan: rate-limited access to hosted open models through an OpenAI-compatible API. For gpt-oss-120b, gpt-oss-20b and the Qwen models that is 30 requests per minute, 1K requests per day, 8K tokens per minute and 200K tokens per day. Whisper is included too (20 RPM, 2K requests per day).

The catch: limits apply per organization, not per key, and Groq notes they may change.

OpenRouter

A curated set of free model variants at zero cost - 20 requests per minute on all of them.

The catch: only 50 requests per day until you have bought at least 10 credits (then 1,000 per day). Extra accounts or keys don't raise the limit, and the free lineup rotates - at last check it was three models.

NVIDIA NIM

Free inference endpoints for four preview models: kimi-k3, deepseek-v4-pro-0813, nemotron-3.5-lightning-30b-a3b and nemotron-3-ultra-550b-a55b.

The catch: you need to create and verify an NVIDIA account before you get a key, and no rate limits or SLA are published for the free tier.

Mistral

The Free plan includes $10 per month in API credits, plus Studio for testing models and the Vibe agent on web and mobile.

The catch: limited messages, web searches and coding sessions on the Free plan.

Cohere

A free, rate-limited Trial API key on signup, for development and non-commercial prototyping.

The catch: trial keys are explicitly not allowed for production or commercial use - going live requires paid billing.

Z.AI (GLM)

GLM-4.7-Flash, GLM-4.5-Flash and GLM-4.6V-Flash are priced at $0 for input, cached input and output. No usage caps are stated on the pricing page.

The catch: only the Flash models are free.

LLM7.io

Up to 1,000,000 tokens per day through an OpenAI-compatible API with a free token (2 requests/second, 40/minute, 100/hour). Without any key you still get 500,000 tokens per day at lower request rates.

The catch: input and output tokens count together against the daily cap.

Pollinations

Text, image, audio and video generation over plain HTTP, no signup required. Anonymous use gets basic models at 1 request per 15 seconds; free registration raises it to 1 request per 5 seconds and unlocks standard models.

The catch: the anonymous rate is slow, and free-tier images may carry watermarks (registering removes them).

Switching between them

Groq and OpenRouter both speak the OpenAI API, so most SDKs and coding agents work by swapping the base URL and key:

from openai import OpenAI

# Groq
client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key="GROQ_API_KEY")

# OpenRouter (pick a model with the :free suffix)
# client = OpenAI(base_url="https://openrouter.ai/api/v1", api_key="OPENROUTER_API_KEY")

reply = client.chat.completions.create(
    model="openai/gpt-oss-20b",
    messages=[{"role": "user", "content": "Hello"}],
)
print(reply.choices[0].message.content)
Enter fullscreen mode Exit fullscreen mode

Which one should you start with?

  • Most capable model for free: Gemini (2.5 Pro is on the free tier).
  • Fastest responses: Groq.
  • Most tokens per day: LLM7.io (1M/day).
  • Monthly credit you can spend on any model: Mistral ($10/month).
  • No signup at all: Pollinations, or LLM7.io without a key.

None of these are meant for production traffic - treat them as prototyping budgets.


The live, sortable version of this comparison - updated when providers change their terms - is at aifree.dev/compare/free-llm-apis. For one-off signup credits, see the free API credits list.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to