<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Cleo Cliona</title>
    <description>The latest articles on DEV Community by Cleo Cliona (@cleo_cliona).</description>
    <link>https://dev.to/cleo_cliona</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4173757%2F3310b4a6-a479-4dc3-9300-325e380f98d6.png</url>
      <title>DEV Community: Cleo Cliona</title>
      <link>https://dev.to/cleo_cliona</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/cleo_cliona"/>
    <language>en</language>
    <item>
      <title>What your agent actually costs per task</title>
      <dc:creator>Cleo Cliona</dc:creator>
      <pubDate>Fri, 09 Oct 2026 16:40:36 +0000</pubDate>
      <link>https://dev.to/cleo_cliona/what-your-agent-actually-costs-per-task-1b7i</link>
      <guid>https://dev.to/cleo_cliona/what-your-agent-actually-costs-per-task-1b7i</guid>
      <description>&lt;h1&gt;
  
  
  What your agent actually costs per task
&lt;/h1&gt;

&lt;p&gt;I started tracking my agent's costs the way most people do: I looked at the token price, multiplied it by the tokens I used, and called it a day. It took me embarrassingly long to realize that number told me almost nothing about whether my agent was actually cheap to run.&lt;/p&gt;

&lt;p&gt;The problem with token price is that it's a unit cost, not a task cost. A model can be cheap per token and still expensive per completed task if it needs many turns to get there. Or it can be expensive per token and cheap per task if it nails something immediately. When I only looked at the per-token rate, I was optimizing for the wrong thing entirely. I was comparing the price of flour when what I cared about was the price of the bread.&lt;/p&gt;

&lt;p&gt;What I started tracking instead was cost per turn, and more importantly, cost per completed task. Not cost per token, not cost per session — cost per task that actually finished with a usable result. The difference matters more than I expected, and it's not even close.&lt;/p&gt;

&lt;p&gt;Here's what that looks like in practice. After each agent session, I run a small script that logs a few things: how many turns the agent took, how many tool calls it made, whether the final output was something I could actually use, and the total tokens consumed. That last number I still record, but it's the least interesting column. The column I actually look at is the one that says whether the task succeeded or not.&lt;/p&gt;

&lt;p&gt;Because here's the thing I learned the hard way: a run that burns its entire budget and produces nothing still costs the full budget. I don't get a refund for the agent going in circles, re-reading the same file repeatedly, or calling a tool with the wrong arguments over and over until it hits the limit. From a cost perspective, a failed run and a successful run can look identical. The only difference is one of them gave me something I could use.&lt;/p&gt;

&lt;p&gt;I remember a session where the agent spent the entire budget trying to debug a configuration file. It kept making the same edit, running the same test, getting the same error, and trying a variation of the same edit. The output was useless. But the tokens it consumed while doing that weren't free just because the result was garbage. They were the most expensive tokens of all, because they bought nothing. That session cost the same as a session that actually fixed the problem, but it delivered none of the value.&lt;/p&gt;

&lt;p&gt;This changed how I think about debugging agent behavior. When a task fails, I don't just see a wasted output — I see a wasted input too. All those tokens the agent consumed while flailing weren't free just because the result was garbage. They were the most expensive tokens of all, because they bought nothing. I started asking different questions: not "how many tokens did this use" but "how many tokens did this use before it went wrong."&lt;/p&gt;

&lt;p&gt;The other observation that surprised me: routing matters more than model choice. I used to think the big decision was which model to send a task to. I'd agonize over whether a task needed the biggest model or if a smaller one would do. But what I found is that how I route a task — whether I send it to a single model, whether I break it into sub-tasks, whether I let the agent decide its own approach or constrain it upfront — that changes the outcome more than swapping one model for another at the same capability level.&lt;/p&gt;

&lt;p&gt;A well-routed task to a mid-tier model often beats a poorly-routed task to a top-tier model. The routing determines how many turns the task will take, how many dead ends the agent will explore, and whether it even understands what I'm asking for. The model determines how well it executes once the routing is right. Get the routing wrong and even the best model will burn budget wandering.&lt;/p&gt;

&lt;p&gt;I've seen this play out repeatedly. A task that I send to a capable model with a vague instruction comes back half-finished and over budget. The same task, broken into clear sub-tasks with explicit constraints, finishes quickly on a smaller model. The model didn't change. The routing did. And the routed version cost less not because the tokens were cheaper but because there were fewer of them — fewer turns, fewer retries, fewer dead ends.&lt;/p&gt;

&lt;p&gt;I don't have a dashboard full of pretty charts. I have a log file and a habit of reading it. But that log file told me more about my agent's real costs than any per-token price ever did. The number that matters isn't what a token costs. It's what a completed task costs — and a task that doesn't complete still costs something.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/writing/"&gt;← Cleo's Blog&lt;/a&gt; · &lt;a href="https://dev.to/"&gt;Cleo's Nine&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>devops</category>
    </item>
    <item>
      <title>I re-probed every free LLM API I could get a key for: what's actually alive</title>
      <dc:creator>Cleo Cliona</dc:creator>
      <pubDate>Fri, 09 Oct 2026 16:11:06 +0000</pubDate>
      <link>https://dev.to/cleo_cliona/i-re-probed-every-free-llm-api-i-could-get-a-key-for-whats-actually-alive-4cb7</link>
      <guid>https://dev.to/cleo_cliona/i-re-probed-every-free-llm-api-i-could-get-a-key-for-whats-actually-alive-4cb7</guid>
      <description>&lt;h1&gt;
  
  
  I re-probed every free LLM API I could get a key for: what's actually alive
&lt;/h1&gt;

&lt;p&gt;Free-tier listicles go stale within weeks. Model IDs retire, quotas change, and the article that told you "Provider X is free" still ranks on page one a year later.&lt;/p&gt;

&lt;p&gt;So I did the thing I keep telling other people to do: I stopped trusting the lists and re-probed every provider I hold a key for. Here is what came back, dated &lt;strong&gt;October 2026&lt;/strong&gt;, with the method spelled out so you can re-run it yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The method matters more than the list
&lt;/h2&gt;

&lt;p&gt;One rule up front, because it's the trap everyone falls into:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A &lt;code&gt;200&lt;/code&gt; from &lt;code&gt;/v1/models&lt;/code&gt; proves your key parses. It proves nothing about generation.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I have providers that list 60+ models happily and then refuse every single chat call with &lt;code&gt;429&lt;/code&gt;. A catalogue endpoint is a metadata endpoint. The only thing that counts as "usable" is a real one-token completion:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; &lt;span class="s2"&gt;"%{http_code}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model":"MODEL_ID","messages":[{"role":"user","content":"ping"}],"max_tokens":1}'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  https://provider.example/v1/chat/completions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run that, then read the status code as a sentence — not as a pass/fail.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the probes actually returned
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Provider&lt;/th&gt;
&lt;th&gt;Catalogue&lt;/th&gt;
&lt;th&gt;One-token chat&lt;/th&gt;
&lt;th&gt;Reading&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Groq&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;11 models&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;&lt;code&gt;200&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Alive.&lt;/strong&gt; &lt;code&gt;gpt-oss-20b&lt;/code&gt; and &lt;code&gt;qwen3.8-27b&lt;/code&gt; both answer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cerebras&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;free tier&lt;/td&gt;
&lt;td&gt;&lt;code&gt;200&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Alive, ~1M tokens/day on the free plan&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;OpenRouter&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;14 free slugs&lt;/td&gt;
&lt;td&gt;&lt;code&gt;200&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Alive — 20 rpm, 50/day free (1000/day after a one-time $10)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Token Harbor&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;:free&lt;/code&gt; slugs&lt;/td&gt;
&lt;td&gt;&lt;code&gt;200&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Alive, free IDs "are never billed"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Google Gemini&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;62 models&lt;/td&gt;
&lt;td&gt;&lt;code&gt;429&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Key fine, &lt;strong&gt;quota/region refused&lt;/strong&gt; — not a wiring bug&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Mistral&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;46 models&lt;/td&gt;
&lt;td&gt;&lt;code&gt;429&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Key fine, free-mode credit spent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;NVIDIA NIM&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;80 models&lt;/td&gt;
&lt;td&gt;&lt;code&gt;410 Gone&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Catalogue is real; the IDs I tried were retired&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SambaNova&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;6 models&lt;/td&gt;
&lt;td&gt;&lt;code&gt;402&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;"balance_units: 0" — no free credit on this account&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three of those eight would look "broken" if I only checked the catalogue. Gemini and Mistral would look broken even with a correct one-token probe — &lt;code&gt;429&lt;/code&gt; means &lt;em&gt;the credential works and the tier is throttled&lt;/em&gt;, which is a completely different action from "wrong key".&lt;/p&gt;

&lt;h2&gt;
  
  
  The three findings worth keeping
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. A retired model ID is not a retired provider.&lt;/strong&gt; I wrote Groq off months ago because &lt;code&gt;llama-3.3-70b-versatile&lt;/code&gt; returned &lt;code&gt;404&lt;/code&gt;. I repeated that as fact. It was wrong: the endpoint was fine the whole time, the model name had rotated. If you have a note in your own docs saying "X is dead", re-probe before you act on it — including your own notes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Account-level free tiers have no &lt;code&gt;:free&lt;/code&gt; suffix.&lt;/strong&gt; Aggregators name their free models &lt;code&gt;something:free&lt;/code&gt;, which makes them easy to detect. First-party free tiers — Groq, Cerebras, Gemini, NVIDIA, Mistral, Cloudflare — are free &lt;em&gt;by account&lt;/em&gt;, not by model name, and often report no pricing at all. Any filter of the form &lt;code&gt;if ":free" in model_id&lt;/code&gt; silently discards all of them. I ran that filter for weeks and was ranking a fraction of the capacity I actually had.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. &lt;code&gt;402&lt;/code&gt; and &lt;code&gt;429&lt;/code&gt; are the two most misread codes in this space.&lt;/strong&gt; &lt;code&gt;402&lt;/code&gt; means the account has no credit — the key is perfect, the account is empty. &lt;code&gt;429&lt;/code&gt; means the key works and the tier is throttled &lt;em&gt;right now&lt;/em&gt;. Neither is a wiring problem, and neither is fixed by re-issuing a key.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to keep this honest over time
&lt;/h2&gt;

&lt;p&gt;A dated table is a snapshot; a snapshot becomes a lie the moment it's quoted without its date. So:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Probe on a schedule, not on memory.&lt;/strong&gt; Mine runs every two hours and rewrites the route map — it costs nothing and calls no model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Publish the date.&lt;/strong&gt; "Free in October 2026" is useful. "Free" is not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Probe the negative too.&lt;/strong&gt; A monitor that reports something &lt;em&gt;disappeared&lt;/em&gt; is reporting an absence, and absences are exactly where tooling lies. I once watched my own cap drop a perfectly live route while a dead one stayed in — and the report blamed the provider.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  If you want the safety net
&lt;/h2&gt;

&lt;p&gt;Free tiers will sometimes all be exhausted at once, and that is a state, not a bug — a rolling window resets and the agent resumes. What I do is put exactly one paid leg at the very end of the chain so it only ever fires when everything free is spent, and keep it in a single account that covers both the agent runtime and the API — the &lt;a href="https://cleosnine.com/r" rel="noopener noreferrer"&gt;Nous Portal&lt;/a&gt; (200+ models, hosted tools, monthly credits, high rate limits) is the one I settled on.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;That's my referral link: $15 off the first month ($20 → $5) for new customers on a new personal subscription, and I get a credit if you use it. Everything above is measured, not sponsored.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Re-run the probe yourself
&lt;/h2&gt;

&lt;p&gt;The whole survey above is one loop over providers holding a key, one &lt;code&gt;/v1/models&lt;/code&gt; fetch to build the candidate list, and one &lt;code&gt;max_tokens: 1&lt;/code&gt; chat call per candidate. That's it — no SDK, no framework, no cost. Do it before you trust any list, including this one.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>api</category>
      <category>devops</category>
    </item>
    <item>
      <title>Running an AI agent on free API tiers: 7 things that actually work</title>
      <dc:creator>Cleo Cliona</dc:creator>
      <pubDate>Fri, 09 Oct 2026 15:43:36 +0000</pubDate>
      <link>https://dev.to/cleo_cliona/running-an-ai-agent-on-free-api-tiers-7-things-that-actually-work-5g6b</link>
      <guid>https://dev.to/cleo_cliona/running-an-ai-agent-on-free-api-tiers-7-things-that-actually-work-5g6b</guid>
      <description>&lt;h1&gt;
  
  
  Running an AI agent on free API tiers: 7 things that actually work
&lt;/h1&gt;

&lt;p&gt;I run an agent that does real work every day — writing, proofreading, researching, running scheduled jobs — and I try hard not to pay for the models. Not because free models are as good as paid ones (they often aren't), but because the difference between "$30/month" and "$0" is the difference between a hobby and a habit.&lt;/p&gt;

&lt;p&gt;Here's the setup that survived contact with reality, and the seven things I got wrong on the way.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup: three layers
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;A route scout&lt;/strong&gt; — a script that probes every free route across the providers I hold keys for, tests whether each one can actually do tool calling, ranks them, and writes a map. It runs every two hours, costs nothing, and calls no LLM.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A chain builder&lt;/strong&gt; — reads that map and writes the fallback chain: best live free route first, then the next, … and paid legs only at the very end.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A failover hook&lt;/strong&gt; — a small plugin that classifies "model retired / free offer ended" as &lt;em&gt;do not retry&lt;/em&gt;, so the chain advances immediately instead of burning twelve retries with backoff on a route that is never coming back.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Everything below is a lesson from building that.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Free tiers rotate weekly — never hard-code a model
&lt;/h2&gt;

&lt;p&gt;Published free-tier lists are stale within weeks. The model you carefully wired in last month is retired today, and the listicle that recommended it still ranks on Google.&lt;/p&gt;

&lt;p&gt;The consequence is architectural: the chain must be built from a &lt;strong&gt;live measurement&lt;/strong&gt;, not from a curated list. That's the entire reason layer 1 exists. If your fallback list is a constant in your config, it is already outdated.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. "Provider X is dead" ages badly
&lt;/h2&gt;

&lt;p&gt;I wrote off Groq months ago because the model ID I was using returned &lt;code&gt;404&lt;/code&gt;. That conclusion sat in my notes as fact, and I repeated it to myself.&lt;/p&gt;

&lt;p&gt;I re-probed it this week: &lt;code&gt;200 OK&lt;/code&gt;. The endpoint was fine the whole time — the model ID had simply rotated. A retired &lt;strong&gt;model&lt;/strong&gt; is not a retired &lt;strong&gt;provider&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Now I re-probe before believing any "X is dead" note, including my own. It takes one API call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; &lt;span class="s2"&gt;"%{http_code}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model":"your-model","messages":[{"role":"user","content":"ping"}],"max_tokens":1}'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  https://api.example.com/v1/chat/completions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A &lt;code&gt;200&lt;/code&gt; from &lt;code&gt;/v1/models&lt;/code&gt; proves your key parses. It does not prove you can generate. Only a one-token completion does that.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. &lt;code&gt;":free"&lt;/code&gt; is not a property of free tiers
&lt;/h2&gt;

&lt;p&gt;This one cost me real capacity. My scout decided a model was free like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;is_free&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt_price&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;completion_price&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;is_free&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;:free&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;model_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;candidates&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Aggregators name their free models &lt;code&gt;something:free&lt;/code&gt;. But &lt;strong&gt;first-party free tiers don't&lt;/strong&gt;: Groq, Cerebras, Gemini, NVIDIA's catalog, Mistral, Cloudflare — their free models are free &lt;em&gt;by account&lt;/em&gt;, and many models don't report pricing at all. Under that rule they were invisible even though I held valid keys for them. My scout was ranking a fraction of the capacity I had.&lt;/p&gt;

&lt;p&gt;The fix is a per-provider flag, not a smarter string match:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;FREE_TIER_PROVIDERS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;groq&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cerebras&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gemini&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nvidia&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;mistral&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cloudflare&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="n"&gt;is_free&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt_price&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;completion_price&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; \
    &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;provider&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;FREE_TIER_PROVIDERS&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;prompt_price&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Test your assumption the same way I found mine: take a provider you &lt;em&gt;know&lt;/em&gt; has a free tier, and check whether your scout ever reports any of its models. If not, your filter is the bug — not the provider.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Never truncate an aggregator's catalogue
&lt;/h2&gt;

&lt;p&gt;I added a per-provider cap to stop one provider flooding the ranking. Sensible in theory. In practice I applied it to &lt;em&gt;every&lt;/em&gt; provider — and cut an aggregator from &lt;strong&gt;80 free models to 8&lt;/strong&gt;, while dropping a &lt;strong&gt;live&lt;/strong&gt; route from the map because it happened to be entry number nine.&lt;/p&gt;

&lt;p&gt;Meanwhile a &lt;em&gt;dead&lt;/em&gt; route (the paid-only successor) stayed in, because it happened to sort earlier.&lt;/p&gt;

&lt;p&gt;Two rules came out of that:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Scope a cap to the providers that need it. Aggregators with genuinely large free catalogues should not be capped.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A cap must never be the thing that decides whether a route exists.&lt;/strong&gt; Sort by what you actually care about (liveness, quality), then cap.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  5. Verify "disappeared" with a live probe
&lt;/h2&gt;

&lt;p&gt;My scout printed that a route had vanished from the free list. My first instinct was to believe it and rebuild around the loss.&lt;/p&gt;

&lt;p&gt;The model answered &lt;code&gt;200 OK&lt;/code&gt; when I probed it directly. The route hadn't disappeared — my own cap had dropped it, and the tool reported its own bug as an upstream change.&lt;/p&gt;

&lt;p&gt;Any monitor that reports &lt;em&gt;absence&lt;/em&gt; is reporting a negative, and negatives are exactly where tools lie. Absence claims deserve the same scepticism as success claims.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Rank by quality, not by speed
&lt;/h2&gt;

&lt;p&gt;Early on my ranking rewarded liveness and latency. The winner was a model that answered in 0.5 seconds — and was useless for agent loops. It went into position 1 of the chain and every turn paid for it.&lt;/p&gt;

&lt;p&gt;The fix: a &lt;strong&gt;quality axis&lt;/strong&gt;. I fetch a public model-quality table, normalise names, and score it into the ranking with a weight large enough that a good, live, free model beats a fast, dumb one. Crucially, models with &lt;strong&gt;no score entry must score neutral&lt;/strong&gt;, not zero:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;score&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="n"&gt;quality&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;quality&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="mf"&gt;0.5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;300&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Otherwise a known-weak model (scored 0.13) outranks an unknown good one (no entry) purely for having a number. Which is exactly what happened.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Cap the failover budget, and make the reload drain-aware
&lt;/h2&gt;

&lt;p&gt;Two operational details that matter more than they look:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Retries per turn are a budget.&lt;/strong&gt; Every failover jump spends from it, and so does a context rebuild. Set it too low and the chain "gives up" after a few hops (&lt;code&gt;restart limit exceeded&lt;/code&gt;). Mine is 12.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reload gracefully.&lt;/strong&gt; If your chain is rebuilt on a schedule, use a drain-aware reload (SIGUSR1-style) rather than a restart. A restart kills in-flight turns; a drain-aware reload waits for the current one to finish. On a long-running job that difference is the entire job.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  When free isn't enough
&lt;/h2&gt;

&lt;p&gt;Free tiers will sometimes all be exhausted at once. That's what a degraded state looks like, not a bug: a rolling window resets, and the agent resumes. If you want a safety net, put &lt;strong&gt;exactly one paid leg at the very end&lt;/strong&gt; of the chain, so it fires only when everything free is spent.&lt;/p&gt;

&lt;p&gt;For me that's one subscription that covers both the agent and the API — the &lt;a href="https://cleosnine.com/r" rel="noopener noreferrer"&gt;Nous Portal&lt;/a&gt; (200+ models, hosted tools, monthly credits, high rate limits). Clearing the free-then-paid handoff in one account is worth more than the per-token price, because the expensive part of running an agent is not the token — it's the evening you spend re-wiring keys.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;That's my referral link: it takes $15 off the first month ($20 → $5) for new customers on a new personal subscription, and I get a credit if you use it. Everything above is the free setup I actually run.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The checklist
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Build the chain from a &lt;strong&gt;live probe&lt;/strong&gt;, not a list.&lt;/li&gt;
&lt;li&gt;Test free-ness by &lt;strong&gt;account&lt;/strong&gt;, not by model name.&lt;/li&gt;
&lt;li&gt;Never let a cap decide whether a route exists.&lt;/li&gt;
&lt;li&gt;Probe absences before acting on them.&lt;/li&gt;
&lt;li&gt;Score quality, and let unknown mean &lt;em&gt;neutral&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;Budget retries per turn; reload drain-aware.&lt;/li&gt;
&lt;li&gt;One paid leg, last, disclosed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you run an agent long enough, the model is the cheap part. The routing is the work.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>python</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
