DEV Community

Cleo Cliona
Cleo Cliona

Posted on Originally published at cleosnine.com

Running an AI agent on free API tiers: 7 things that actually work

Running an AI agent on free API tiers: 7 things that actually work

I run an agent that does real work every day — writing, proofreading, researching, running scheduled jobs — and I try hard not to pay for the models. Not because free models are as good as paid ones (they often aren't), but because the difference between "$30/month" and "$0" is the difference between a hobby and a habit.

Here's the setup that survived contact with reality, and the seven things I got wrong on the way.

The setup: three layers

  1. A route scout — a script that probes every free route across the providers I hold keys for, tests whether each one can actually do tool calling, ranks them, and writes a map. It runs every two hours, costs nothing, and calls no LLM.
  2. A chain builder — reads that map and writes the fallback chain: best live free route first, then the next, … and paid legs only at the very end.
  3. A failover hook — a small plugin that classifies "model retired / free offer ended" as do not retry, so the chain advances immediately instead of burning twelve retries with backoff on a route that is never coming back.

Everything below is a lesson from building that.

1. Free tiers rotate weekly — never hard-code a model

Published free-tier lists are stale within weeks. The model you carefully wired in last month is retired today, and the listicle that recommended it still ranks on Google.

The consequence is architectural: the chain must be built from a live measurement, not from a curated list. That's the entire reason layer 1 exists. If your fallback list is a constant in your config, it is already outdated.

2. "Provider X is dead" ages badly

I wrote off Groq months ago because the model ID I was using returned 404. That conclusion sat in my notes as fact, and I repeated it to myself.

I re-probed it this week: 200 OK. The endpoint was fine the whole time — the model ID had simply rotated. A retired model is not a retired provider.

Now I re-probe before believing any "X is dead" note, including my own. It takes one API call:

curl -s -o /dev/null -w "%{http_code}\n" \
  -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
  -d '{"model":"your-model","messages":[{"role":"user","content":"ping"}],"max_tokens":1}' \
  https://api.example.com/v1/chat/completions
Enter fullscreen mode Exit fullscreen mode

A 200 from /v1/models proves your key parses. It does not prove you can generate. Only a one-token completion does that.

3. ":free" is not a property of free tiers

This one cost me real capacity. My scout decided a model was free like this:

is_free = (prompt_price == 0 and completion_price == 0)
if is_free and ":free" in model_id:
    candidates.append(model_id)
Enter fullscreen mode Exit fullscreen mode

Aggregators name their free models something:free. But first-party free tiers don't: Groq, Cerebras, Gemini, NVIDIA's catalog, Mistral, Cloudflare — their free models are free by account, and many models don't report pricing at all. Under that rule they were invisible even though I held valid keys for them. My scout was ranking a fraction of the capacity I had.

The fix is a per-provider flag, not a smarter string match:

FREE_TIER_PROVIDERS = {"groq", "cerebras", "gemini", "nvidia", "mistral", "cloudflare"}
is_free = (prompt_price == 0 and completion_price == 0) \
    or (provider in FREE_TIER_PROVIDERS and prompt_price <= 0)
Enter fullscreen mode Exit fullscreen mode

Test your assumption the same way I found mine: take a provider you know has a free tier, and check whether your scout ever reports any of its models. If not, your filter is the bug — not the provider.

4. Never truncate an aggregator's catalogue

I added a per-provider cap to stop one provider flooding the ranking. Sensible in theory. In practice I applied it to every provider — and cut an aggregator from 80 free models to 8, while dropping a live route from the map because it happened to be entry number nine.

Meanwhile a dead route (the paid-only successor) stayed in, because it happened to sort earlier.

Two rules came out of that:

  • Scope a cap to the providers that need it. Aggregators with genuinely large free catalogues should not be capped.
  • A cap must never be the thing that decides whether a route exists. Sort by what you actually care about (liveness, quality), then cap.

5. Verify "disappeared" with a live probe

My scout printed that a route had vanished from the free list. My first instinct was to believe it and rebuild around the loss.

The model answered 200 OK when I probed it directly. The route hadn't disappeared — my own cap had dropped it, and the tool reported its own bug as an upstream change.

Any monitor that reports absence is reporting a negative, and negatives are exactly where tools lie. Absence claims deserve the same scepticism as success claims.

6. Rank by quality, not by speed

Early on my ranking rewarded liveness and latency. The winner was a model that answered in 0.5 seconds — and was useless for agent loops. It went into position 1 of the chain and every turn paid for it.

The fix: a quality axis. I fetch a public model-quality table, normalise names, and score it into the ranking with a weight large enough that a good, live, free model beats a fast, dumb one. Crucially, models with no score entry must score neutral, not zero:

score += int((quality if quality is not None else 0.5) * 300)
Enter fullscreen mode Exit fullscreen mode

Otherwise a known-weak model (scored 0.13) outranks an unknown good one (no entry) purely for having a number. Which is exactly what happened.

7. Cap the failover budget, and make the reload drain-aware

Two operational details that matter more than they look:

  • Retries per turn are a budget. Every failover jump spends from it, and so does a context rebuild. Set it too low and the chain "gives up" after a few hops (restart limit exceeded). Mine is 12.
  • Reload gracefully. If your chain is rebuilt on a schedule, use a drain-aware reload (SIGUSR1-style) rather than a restart. A restart kills in-flight turns; a drain-aware reload waits for the current one to finish. On a long-running job that difference is the entire job.

When free isn't enough

Free tiers will sometimes all be exhausted at once. That's what a degraded state looks like, not a bug: a rolling window resets, and the agent resumes. If you want a safety net, put exactly one paid leg at the very end of the chain, so it fires only when everything free is spent.

For me that's one subscription that covers both the agent and the API — the Nous Portal (200+ models, hosted tools, monthly credits, high rate limits). Clearing the free-then-paid handoff in one account is worth more than the per-token price, because the expensive part of running an agent is not the token — it's the evening you spend re-wiring keys.

That's my referral link: it takes $15 off the first month ($20 → $5) for new customers on a new personal subscription, and I get a credit if you use it. Everything above is the free setup I actually run.

The checklist

  • Build the chain from a live probe, not a list.
  • Test free-ness by account, not by model name.
  • Never let a cap decide whether a route exists.
  • Probe absences before acting on them.
  • Score quality, and let unknown mean neutral.
  • Budget retries per turn; reload drain-aware.
  • One paid leg, last, disclosed.

If you run an agent long enough, the model is the cheap part. The routing is the work.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.