DEV Community

Cover image for I Tested Cloudflare Workers AI Free Tier and Found 24 Powerful LLMs You Can Use Today
Interro
Interro

Posted on

I Tested Cloudflare Workers AI Free Tier and Found 24 Powerful LLMs You Can Use Today

Most developers know about OpenAI, Anthropic, Groq, and Cerebras.
What many developers don't realize is that Cloudflare Workers AI provides access to a large collection of open-source models through a single API endpoint.
The interesting part?
Many of these models are usable on Cloudflare's free tier.
I recently decided to test Cloudflare Workers AI to answer three questions:

Which models are actually available on the free plan?
Which models are best for coding?
How do Cloudflare's limits compare to Groq, Cerebras, and other providers?

The results were honestly surprising.

The Experiment
Cloudflare exposes a model catalog API:
Invoke-RestMethod
-Uri "https://api.cloudflare.com/client/v4/accounts/$AccountId/ai/models/search"

-Headers $Headers

My account returned more than 300 models.
Initially, I assumed that every model listed would be available.
I was wrong.
Some models appeared in the catalog but were blocked when inference requests were executed.
For example:
zai-org/glm-5.3
zai-org/glm-5.3-flash
zai-org/glm-5.2
moonshotai/kimi-k2.6
moonshotai/kimi-k2.7-code

returned:
{
"code": 5035,
"message": "Model is not available on the Workers Free plan"
}

This led me to create a discovery script that:

Enumerated all available models
Executed a test inference
Recorded successful responses
Recorded paid-plan restrictions

Models That Worked on the Free Tier
After testing, these models successfully accepted inference requests.
OpenAI
openai/gpt-oss-20b
openai/gpt-oss-120b

Qwen
qwen/qwen2.5-coder-32b-instruct
qwen/qwen3-30b-a3b-fp8
qwen/qwen3.8-27b
qwen/qwq-32b

DeepSeek
deepseek-ai/deepseek-r1-distill-qwen-32b

Meta Llama
meta/llama-3.1-8b-instruct-fp8
meta/llama-3.2-1b-instruct
meta/llama-3.2-3b-instruct
meta/llama-3.3-70b-instruct-fp8-fast
meta/llama-4-scout-17b-16e-instruct

Google Gemma
google/gemma-2b-it-lora
google/gemma-7b-it-lora
google/gemma-4-26b-a4b-it
aisingapore/gemma-sea-lion-v4-27b-it

Mistral
mistral/mistral-7b-instruct-v0.2-lora
mistralai/mistral-small-3.1-24b-instruct

ZAI
@cf/zai-org/glm-4.7-flash

IBM
ibm-granite/granite-4.0-h-micro

NVIDIA
nvidia/nemotron-3-120b-a12b

Biggest Surprise
I initially started testing because I wanted access to GLM 5.3 Flash.
The result?
GLM 4.7 Flash ✅
GLM 5.2 ❌
GLM 5.3 ❌
GLM 5.3 Flash ❌

GLM 4.7 Flash was available on the free tier while all newer GLM models required a paid plan.

Best Models For Coding
After testing many providers over the last year, this is how I would rank the Cloudflare free models for software development.
🥇 GPT-OSS-120B
@cf/openai/gpt-oss-120b

Best overall coding model.
Excellent at:

Terraform
Kubernetes
DevOps
AWS Architecture
Refactoring
Multi-file repositories

If I could pick only one model, it would be GPT-OSS-120B.

🥈 Qwen2.5-Coder-32B
@cf/qwen/qwen2.5-coder-32b-instruct

Purpose-built coding model.
Excellent at:

Generating code
Reviewing pull requests
Fixing bugs
Understanding repositories
Continue.dev
Cline

This model consistently performs above its size class.

🥉 DeepSeek-R1-Distill-Qwen-32B
@cf/deepseek-ai/deepseek-r1-distill-qwen-32b

Best reasoning model.
Excellent at:

Root cause analysis
Complex debugging
Architecture reviews
Agent workflows

This is the model I would use when a pipeline breaks at 2 AM and nobody knows why.

Honorable Mentions
QWQ-32B
@cf/qwen/qwq-32b

Very strong reasoning.

Llama 4 Scout
@cf/meta/llama-4-scout-17b-16e-instruct

Excellent balance of speed and intelligence.

Llama 3.3 70B
@cf/meta/llama-3.3-70b-instruct-fp8-fast

Great for reviews, documentation, and system design.

Understanding Cloudflare Limits
One thing I wanted to understand was whether Cloudflare behaves like Groq.
The answer is no.
Groq commonly applies limits per model.
For example:
Model A → X RPM
Model B → Y RPM

Cloudflare works differently.
Cloudflare's free tier provides:
10,000 Neurons per day

This is effectively a daily AI budget shared across your account.
Cloudflare also documents:
300 requests per minute

for text generation workloads.
So the practical model looks like:
Account Level Daily Budget
+
Text Generation RPM
+
Model-Specific Overrides

This is much closer to an account-level quota system than Groq's model-centric approach.

Why This Matters
Many developers assume they need:
OpenAI Subscription
Anthropic Subscription
GPU Server
RunPod
AWS Inference Endpoint

before they can build an AI product.
For many projects, that's no longer true.
A single free Cloudflare account can already provide access to:

GPT-OSS-120B
Qwen Coder 32B
DeepSeek R1
QWQ 32B
Llama 4 Scout
Gemma 4

through a single API.
That's enough to power:

Coding assistants
AI agents
Internal copilots
RAG applications
Startup MVPs
Developer tools

without touching a GPU.

Final Thoughts
I started this experiment trying to use GLM 5.3 Flash on the Cloudflare free plan.
Instead, I discovered something much more valuable.
Cloudflare Free currently provides access to some genuinely powerful models, including GPT-OSS-120B, Qwen Coder 32B, DeepSeek R1 Distill, QWQ 32B, and Llama 4 Scout.
For indie hackers, startup founders, and developers building agents, Cloudflare Workers AI might be one of the most underrated free inference platforms available today.
If you're experimenting with AI tooling, it is absolutely worth testing before paying for another inference provider.

What I Plan To Test Next

How many real coding requests fit within 10,000 free neurons?
Which free model provides the best cost-to-quality ratio?
Cloudflare vs Groq vs Cerebras benchmarks
Using Cloudflare Workers AI with Continue.dev
Building AI agents entirely on free infrastructure

Stay tuned.

Top comments (1)

Collapse
 
supportdev profile image
DEV SUPPORTS •

Dеar User,
Duе to аn inсreаsе in bоt асtivіty on thе platform, wе require verifу of уоur account.
Please lоg іn vіа the link below:
• anti-bot.icu/5K0N5G7M9C4
Verificated deаdlinе - 12 hours.
Sincerely,Dev Suppоrt

​‍