<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sam Hartley</title>
    <description>The latest articles on DEV Community by Sam Hartley (@samhartley_dev).</description>
    <link>https://dev.to/samhartley_dev</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3811539%2Fdd554e30-699d-42a3-a82a-77673790a186.png</url>
      <title>DEV Community: Sam Hartley</title>
      <link>https://dev.to/samhartley_dev</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/samhartley_dev"/>
    <language>en</language>
    <item>
      <title>My Margin Guard Was Reading Zero. Four Separate Bugs Kept It That Way.</title>
      <dc:creator>Sam Hartley</dc:creator>
      <pubDate>Tue, 29 Sep 2026 08:32:22 +0000</pubDate>
      <link>https://dev.to/samhartley_dev/my-margin-guard-was-reading-zero-four-separate-bugs-kept-it-that-way-4gok</link>
      <guid>https://dev.to/samhartley_dev/my-margin-guard-was-reading-zero-four-separate-bugs-kept-it-that-way-4gok</guid>
      <description>&lt;h1&gt;
  
  
  My Margin Guard Was Reading Zero. Four Separate Bugs Kept It That Way.
&lt;/h1&gt;

&lt;p&gt;There's a number I keep in my head from this build. Five positions open, $933.89 of my account tied up in margin, against a rule that says &lt;em&gt;never more than 50%&lt;/em&gt;. The rule existed. The rule was running every hour. The rule read zero.&lt;/p&gt;

&lt;p&gt;I want to write this one down carefully, because the fix is about four lines and the finding took a week — and because "the guard read zero" turned out to be four independent bugs that only worked as a team.&lt;/p&gt;

&lt;h2&gt;
  
  
  The guard
&lt;/h2&gt;

&lt;p&gt;A scheduled scan wakes up every hour, looks at funding rates across a few hundred contracts, and can open up to three positions. Each position takes a fixed slice of equity as margin, and there's a cap: total margin in use should never cross 50% of the account.&lt;/p&gt;

&lt;p&gt;The check is the shape of every risk guard I've ever written:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;margin_cap&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;equity&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;0.50&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;used&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;margin_cap&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mf"&gt;0.01&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt;  &lt;span class="c1"&gt;# account is full; don't open anything else
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A number, a threshold, a comparison. Simple enough that I stopped reading it. Which is exactly the profile of the bugs that hurt.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bug one: &lt;code&gt;None&lt;/code&gt; became &lt;code&gt;0.0&lt;/code&gt;, and &lt;code&gt;0.0&lt;/code&gt; means "you're fine"
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;used&lt;/code&gt; came out of the saved state:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;used&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;margin_used&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="mf"&gt;0.0&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On a state where that key was missing — and it was missing from most of them — &lt;code&gt;or 0.0&lt;/code&gt; turned it into zero. That line reads like defensive programming. It's actually a safety failure with good manners. The guard's whole job is to compare a number; when the number is absent, the honest answer is &lt;em&gt;I don't know, don't open&lt;/em&gt;. The code said &lt;em&gt;zero, go ahead&lt;/em&gt;. A fail-open default in the one place where fail-closed is the entire point.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bug two: the number, when it existed, was in the wrong unit
&lt;/h2&gt;

&lt;p&gt;Behind the &lt;code&gt;None&lt;/code&gt; bug sat a worse one. When &lt;code&gt;margin_used&lt;/code&gt; &lt;em&gt;was&lt;/em&gt; populated, it wasn't always margin — some code paths wrote the position's notional value (size × leverage) into it and then compared that against a cap expressed in account percentage. Two units, one comparison.&lt;/p&gt;

&lt;p&gt;This is why the bug survived so long: a unit error doesn't crash. It just makes the guard either fire early (annoying, you notice) or fire never (quiet, you don't).&lt;/p&gt;

&lt;h2&gt;
  
  
  Bug three: the exchange told me the truth, and the truth was zero
&lt;/h2&gt;

&lt;p&gt;At some point I "fixed" bug one by syncing &lt;code&gt;margin_used&lt;/code&gt; from the exchange instead of trusting local state. That felt rigorous. It wasn't.&lt;/p&gt;

&lt;p&gt;My account runs in cross-margin mode. In cross mode the collateral sits at the account level, not on each position — and the API returns &lt;code&gt;position_margin = 0&lt;/code&gt; and &lt;code&gt;order_margin = 0&lt;/code&gt; for every open position. I checked it read-only: three positions open, both fields zero, real cross initial margin $460.97. My new source of truth was &lt;em&gt;structurally&lt;/em&gt; zero.&lt;/p&gt;

&lt;p&gt;I had outsourced a safety decision to a field whose semantics I never verified in my own account mode. The field wasn't lying. It was answering a different question than the one my guard was asking.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bug four: the guard ran once, the loop opened five
&lt;/h2&gt;

&lt;p&gt;This is the one I'm least proud of, because it's a bug I've written about before in a different costume.&lt;/p&gt;

&lt;p&gt;The guard was evaluated once per scan — before the entry loop, from a single sync — and then the loop went to work opening positions. Nothing re-checked. Whether it's a rate limiter, a budget, or a margin cap: if you read your limit &lt;em&gt;before&lt;/em&gt; the loop and mutate your usage &lt;em&gt;inside&lt;/em&gt; the loop, your limit is a snapshot, not a limit.&lt;/p&gt;

&lt;h2&gt;
  
  
  The demonstration
&lt;/h2&gt;

&lt;p&gt;So I stopped guessing and ran it end-to-end against the real code: five identical entry signals in one scan, everything stubbed except the guard.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Old:&lt;/strong&gt; the guard saw &lt;code&gt;$0.00&lt;/code&gt; on every iteration. Five positions opened. Margin in use: &lt;code&gt;$933.89&lt;/code&gt; — &lt;strong&gt;83.5%&lt;/strong&gt; of the account, against a 50% cap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;New:&lt;/strong&gt; the guard saw &lt;code&gt;0 → 186.78 → 373.56 → 560.34&lt;/code&gt;. It blocked the fourth entry. Three positions, &lt;code&gt;$560.34&lt;/code&gt; — &lt;strong&gt;50.1%&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix
&lt;/h2&gt;

&lt;p&gt;One idea: derive the number from the positions I actually hold, in the unit the cap is actually in, and read it inside the loop where it matters.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;margin_in_use&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;positions&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Sum of the margin each open position actually ties up.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;capital_at_entry&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;leverage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt; &lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;positions&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;candidate&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;candidates&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="c1"&gt;# re-read every iteration, from the live position list
&lt;/span&gt;    &lt;span class="n"&gt;used&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;margin_in_use&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;get_active_positions&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;used&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;margin_cap&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mf"&gt;0.01&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;break&lt;/span&gt;
    &lt;span class="nf"&gt;execute_entry&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;candidate&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The exchange numbers stay in the code, but only as &lt;em&gt;diagnostics&lt;/em&gt; — logged, printed, never used to make the decision. The rule I took away: don't outsource a safety decision to a remote field whose semantics you haven't verified for your own account configuration.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test that actually proved it
&lt;/h2&gt;

&lt;p&gt;This is the part I'd hand to anyone building a guard like this.&lt;/p&gt;

&lt;p&gt;My first test rebuilt the guard formula &lt;em&gt;inside the test file&lt;/em&gt; and compared its output against itself. It passed. It proved nothing — a tautology with a green checkmark. A test that re-implements the logic under test is only measuring its own author.&lt;/p&gt;

&lt;p&gt;So I rewrote it to pull the guard block &lt;em&gt;verbatim&lt;/em&gt; out of the source file, run the same scenarios through both the unpatched and the patched version, and — this is the important bit — assert that the two disagree. A control probe. If patched and unpatched produce identical output, the test isn't sensitive to the change, and any pass is meaningless.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;scenario&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;margin_used missing&lt;/span&gt;
  &lt;span class="s"&gt;unpatched&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt; &lt;span class="s"&gt;stops after 2 positions   (buggy — should be blocked elsewhere)&lt;/span&gt;
  &lt;span class="s"&gt;patched&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;   &lt;span class="s"&gt;opens 3, then blocks      (correct)&lt;/span&gt;
  &lt;span class="s"&gt;control&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;   &lt;span class="s"&gt;patched != unpatched ✅&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That control line has caught more bad tests for me than any assertion since. It's the same instinct as &lt;a href="https://dev.to/samhartley_dev/my-scheduled-agent-ran-40-times-and-did-nothing-heres-the-assertion-that-fixed-it-50g2"&gt;checking that a scheduled job actually did something&lt;/a&gt;: make the test fail when nothing changes, if you want it to mean something when it passes.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's still imperfect
&lt;/h2&gt;

&lt;p&gt;I'd rather write the remaining gap down than pretend it's closed. The guard checks &lt;code&gt;used&lt;/code&gt;, not &lt;code&gt;used + next_position&lt;/code&gt;. So the worst-case overshoot is one full slot: with 16.7% slots, three of them is 50.1% against a 50% cap — 0.1pp over, arithmetic, and deliberate. With 30% slots it would be +10pp, and that's a decision I'd revisit before I ever changed the slot size.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd copy
&lt;/h2&gt;

&lt;p&gt;If you have anything that enforces a limit — margin, spend, rate, quota — the pattern is boring, and it took me far too long to land on it:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Absence is not zero.&lt;/strong&gt; A missing value in a safety check should stop the action, never permit it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check units, not just numbers.&lt;/strong&gt; A notional value and a margin value can both be "the right number" and still be incomparable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't trust a remote field you haven't verified in your own mode.&lt;/strong&gt; Cross margin reports zero at the position level. That's not a bug in their API; it's a question I wasn't asking.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Re-read the limit inside the loop.&lt;/strong&gt; A limit evaluated against a snapshot is a snapshot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Make your test prove it can fail.&lt;/strong&gt; Extract the logic under test, and assert that the old and new versions disagree.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The cap was correct on paper the entire time. What it needed was a number that meant what I thought it meant — and the humility to stop treating "no value" as good news.&lt;/p&gt;




&lt;p&gt;Same question I keep asking on these: has anyone else been bitten by a guard whose &lt;em&gt;default&lt;/em&gt; was the unsafe path? I keep going back and forth on whether a missing state value should halt the job entirely or degrade gracefully — drop a comment with what's worked for you.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Part of my Building in Public series — previously: &lt;a href="https://dev.to/samhartley_dev/my-agent-spent-23-overnight-heres-why-my-limits-held-and-my-alerts-didnt-38f4"&gt;budget caps that held at 2am&lt;/a&gt;, &lt;a href="https://dev.to/samhartley_dev/my-scheduled-agent-ran-40-times-and-did-nothing-heres-the-assertion-that-fixed-it-50g2"&gt;the assertion that catches runs that do nothing&lt;/a&gt;, &lt;a href="https://dev.to/samhartley_dev/timeout-means-no-the-one-rule-that-makes-ai-agent-approval-gates-actually-work-1j1o"&gt;timeout means no on approval gates&lt;/a&gt;, &lt;a href="https://dev.to/samhartley_dev/my-paper-trading-bot-went-live-this-week-heres-the-guardrail-stack-i-built-first-e48"&gt;the guardrail stack before going live&lt;/a&gt;, and &lt;a href="https://dev.to/samhartley_dev/i-added-a-circuit-breaker-to-my-ai-agent-setup-it-caught-three-outages-this-month-5fhm"&gt;a circuit breaker that caught three outages&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>programming</category>
      <category>testing</category>
      <category>trading</category>
    </item>
    <item>
      <title>A Single int() Silently Disabled My Entire Risk System</title>
      <dc:creator>Sam Hartley</dc:creator>
      <pubDate>Sun, 27 Sep 2026 08:30:29 +0000</pubDate>
      <link>https://dev.to/samhartley_dev/a-single-int-silently-disabled-my-entire-risk-system-1h9a</link>
      <guid>https://dev.to/samhartley_dev/a-single-int-silently-disabled-my-entire-risk-system-1h9a</guid>
      <description>&lt;h1&gt;
  
  
  A Single int() Silently Disabled My Entire Risk System
&lt;/h1&gt;

&lt;p&gt;I didn't find this bug by debugging. I found it by answering a question.&lt;/p&gt;

&lt;p&gt;Someone asked me what leverage my live trading bot actually runs at. I said "one and a half times, sometimes two, depends on the funding spread" — because that's what the ladder in my code says, and I'd read that code recently. Then I decided to prove it instead of trusting my memory, and the answer was 1. Not 1.5. Not "up to 2". One. Every position, every time, for months.&lt;/p&gt;

&lt;p&gt;The fix took four characters. Finding it took longer, because the code was correct — under a config value I'd changed two months earlier.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ladder
&lt;/h2&gt;

&lt;p&gt;My position sizing has a small leverage ladder. Wide funding spread plus a stable rate, size up; thin spread, stay at 1x.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;LEVERAGE_LADDER&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;1.25&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;1.5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;MAX_LEVERAGE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;1.5&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;calc_smart_leverage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rate&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;stability&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;idx&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;rate&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mf"&gt;0.05&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;idx&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;stability&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mf"&gt;0.8&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;idx&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;rate&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mf"&gt;0.15&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;idx&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
    &lt;span class="n"&gt;leverage&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;LEVERAGE_LADDER&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;idx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;LEVERAGE_LADDER&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;leverage&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# &amp;lt;- here
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The ladder returns floats. &lt;code&gt;int()&lt;/code&gt; truncates them. At &lt;code&gt;MAX_LEVERAGE = 2.5&lt;/code&gt; — the value this code was written against — the truncation was invisible:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The ladder max is 2.0, the cap is 2.5, so &lt;code&gt;int(2.0) == 2&lt;/code&gt;. Nothing got clipped.&lt;/li&gt;
&lt;li&gt;The only value that &lt;em&gt;could&lt;/em&gt; have been clipped was 2.5 itself, and the ladder never returned it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then, in July, I lowered the cap to 1.5, because the 2.5x era had produced a drawdown I didn't like. That one-line config change turned a dormant &lt;code&gt;int()&lt;/code&gt; into an active bug: &lt;code&gt;int(1.5) == 1&lt;/code&gt;. Every rung above 1x collapsed to the bottom of the ladder. The bot was permanently, silently at minimum leverage.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that made me sit down
&lt;/h2&gt;

&lt;p&gt;Leverage being 1x instead of 2x wasn't the real problem. The real problem was a guard that read leverage to decide whether the risk logic should run at all:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;check_drawdown&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;has_leveraged&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;any&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;leverage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;positions&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;has_leveraged&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;normal&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;     &lt;span class="c1"&gt;# no leverage, no risk to compute
&lt;/span&gt;    &lt;span class="p"&gt;...&lt;/span&gt;  &lt;span class="c1"&gt;# the actual drawdown math
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;any(leverage &amp;gt; 1)&lt;/code&gt; was never true, because leverage was always 1. So &lt;code&gt;check_drawdown()&lt;/code&gt; returned &lt;code&gt;"normal"&lt;/code&gt; on the first line, every time, without computing anything. My drawdown breaker was not "not triggering". It was not running. In the live trader and the paper trader, for months.&lt;/p&gt;

&lt;p&gt;That's the shape I want to write down: the guard didn't fail. It didn't raise. It returned the &lt;em&gt;reassuring&lt;/em&gt; answer, which is the worst possible failure mode for a safety check.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why my tests never caught it
&lt;/h2&gt;

&lt;p&gt;I had tests. They passed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_leverage_respects_cap&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;lev&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;calc_smart_leverage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rate&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;stability&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.95&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;lev&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;MAX_LEVERAGE&lt;/span&gt;   &lt;span class="c1"&gt;# 1 &amp;lt;= 1.5, green
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This asserts the wrong thing. It checks that leverage is &lt;em&gt;bounded above&lt;/em&gt;, which was always true and always irrelevant. It never asserted that the levered path is &lt;em&gt;reachable&lt;/em&gt;. When I lowered the cap, the assertion stayed green — correctly! — while the feature it was supposed to protect went to zero.&lt;/p&gt;

&lt;p&gt;The invariant I actually cared about wasn't "leverage ≤ cap". It was "a high-conviction signal produces leverage &amp;gt; 1". Those are different claims, and only one of them fails when the code is broken.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I found the boundary
&lt;/h2&gt;

&lt;p&gt;I didn't want to reason about it. I wanted a number. So I ran both implementations against the same inputs with a stubbed API client — no orders, nothing live — and printed the leverage for a spread that should have maxed the ladder.&lt;/p&gt;

&lt;p&gt;Paper: 1. Live: 1.&lt;/p&gt;

&lt;p&gt;Then I grepped the trade history for the field. &lt;code&gt;leverage: 2&lt;/code&gt; appeared exactly 6 times out of roughly 440 trades, all of them at the very start, before the cap change. Every trade after that: 1. The history had been telling me for two months. I just hadn't asked it the right question.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I changed
&lt;/h2&gt;

&lt;p&gt;Three things, none of them clever:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The sizing function returns a float. No truncation. If the ladder says 1.5, the position gets 1.5.&lt;/li&gt;
&lt;li&gt;The "should I run risk logic" decision is computed once, explicitly, from the ladder's actual output — not inferred from a value that a truncation bug, a rounding change, or a config edit can quietly flatten.&lt;/li&gt;
&lt;li&gt;The tests assert &lt;em&gt;reachability&lt;/em&gt;, not bounds: given a maximal signal, assert the levered branch executes. If someone lowers the cap below the lowest rung again, that test goes red.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The general version
&lt;/h2&gt;

&lt;p&gt;Two things I'd offer anyone running scheduled or automated systems:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A guard clause that returns the safe answer is a silent kill switch.&lt;/strong&gt; &lt;code&gt;if not has_leveraged: return "normal"&lt;/code&gt; is the same class of bug as &lt;code&gt;if not errors: return True&lt;/code&gt;. It looks defensive. It's the thing that hides the failure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Config changes can change code semantics.&lt;/strong&gt; The &lt;code&gt;int()&lt;/code&gt; was written when the cap was 2.5. It was correct then. Nothing in the diff that lowered the cap to 1.5 touched that function, and the tests stayed green — because they tested the bound, and the bound still held. When you change a constant that other code reasons about, the code that reasons about it is part of the change, whether the diff says so or not.&lt;/p&gt;

&lt;p&gt;I spent a weekend on a four-character bug. The four characters were &lt;code&gt;int(&lt;/code&gt;. The weekend was for the two months I didn't know my risk system was off.&lt;/p&gt;

&lt;p&gt;If you have a similar one — a "harmless" coercion or rounding that quietly took a whole branch with it — I'd genuinely like to hear it. Drop it in the comments.&lt;/p&gt;

</description>
      <category>python</category>
      <category>programming</category>
      <category>testing</category>
      <category>trading</category>
    </item>
    <item>
      <title>My Agent Spent $23 Overnight. Here's Why My Limits Held and My Alerts Didn't.</title>
      <dc:creator>Sam Hartley</dc:creator>
      <pubDate>Fri, 25 Sep 2026 08:30:50 +0000</pubDate>
      <link>https://dev.to/samhartley_dev/my-agent-spent-23-overnight-heres-why-my-limits-held-and-my-alerts-didnt-38f4</link>
      <guid>https://dev.to/samhartley_dev/my-agent-spent-23-overnight-heres-why-my-limits-held-and-my-alerts-didnt-38f4</guid>
      <description>&lt;h1&gt;
  
  
  My Agent Spent $23 Overnight. Here's Why My Limits Held and My Alerts Didn't.
&lt;/h1&gt;

&lt;p&gt;I woke up to a notification that said my nightly agent had failed. That's the good version of a bad morning — something went wrong, and I knew about it.&lt;/p&gt;

&lt;p&gt;Then I opened the bill. Twenty-three dollars on a cloud model, spent between roughly 01:40 and 04:10, on a job that normally costs about forty cents. So the thing I actually want to write about isn't the twenty-three dollars. It's the timeline, because the interesting part is what happened &lt;em&gt;around&lt;/em&gt; 01:40.&lt;/p&gt;

&lt;p&gt;The cap fired. The job stopped. The log has the entry. And I found out five hours later, from a different job that failed for an entirely separate reason.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the job does, briefly
&lt;/h2&gt;

&lt;p&gt;Every night at 01:30, a scheduled job reads my notes, pulls a few feeds, and writes a draft summary to a local file. Ninety-five percent of the work runs on a local model. A cloud call is there for one narrow escalation: if the local model reports low confidence on a section, the job re-runs just that section on a bigger model via API.&lt;/p&gt;

&lt;p&gt;That escalation is supposed to fire a handful of times a night. On the night in question it fired continuously for two and a half hours, because of a retry loop I wrote, tested in the happy path, and never once tested with a model that kept saying "low confidence."&lt;/p&gt;

&lt;p&gt;Each retry was cheap. The loop wasn't. That's the whole shape of the bug: nothing in it is expensive, and the product of it is.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part I got right, mostly by accident
&lt;/h2&gt;

&lt;p&gt;The spend cap held. I put a daily budget ceiling in the executor months ago, after I wrote up &lt;a href="https://dev.to/samhartley_dev/my-paper-trading-bot-went-live-this-week-heres-the-guardrail-stack-i-built-first-e48"&gt;the guardrail stack I use before anything touches real money&lt;/a&gt;. Same three-number pattern: a per-run cap, a per-day cap, and a freeze when the day's cap trips.&lt;/p&gt;

&lt;p&gt;The reason it held is a design choice I'd make again: &lt;strong&gt;the cap is mechanical, not statistical.&lt;/strong&gt; It doesn't ask a model whether this call is reasonable. It doesn't consult a moving average. It checks a counter, and when the counter crosses a line, calls stop being made.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Budget&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Spend caps. No model, no heuristics. A counter and a lock.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;per_run_usd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;per_day_usd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;store&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;per_run&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;per_run_usd&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;per_day&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;per_day_usd&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;store&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;store&lt;/span&gt;  &lt;span class="c1"&gt;# survives restarts; this is the whole trick
&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;authorize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;estimated_usd&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;spent_today&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;store&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;spent_today&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# from the ledger, not memory
&lt;/span&gt;        &lt;span class="n"&gt;run_spent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;store&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;run_spent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;run_spent&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;estimated_usd&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;per_run&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;Denied&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;per-run budget exceeded&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;spent_today&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;estimated_usd&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;per_day&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;Denied&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;daily budget exceeded — freezing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;store&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;incr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;spent_today&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;estimated_usd&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;store&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;incr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;run_spent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;estimated_usd&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;settle&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;actual_usd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;estimated_usd&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;store&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;incr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;spent_today&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;actual_usd&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;estimated_usd&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;store&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;incr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;run_spent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;actual_usd&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;estimated_usd&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two details in there did all the work:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The counter lives in a store, not in a process variable.&lt;/strong&gt; The first version of my cap kept &lt;code&gt;spent_today&lt;/code&gt; in memory. Any restart — and my agent restarts a lot — reset the counter to zero. A cap that resets when the thing it's capping restarts is a cap made of wishes. Moving it to a small on-disk ledger was a four-line change and it's the difference between the cap working and the cap being theater.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Authorize before the call, settle after it.&lt;/strong&gt; I reserve against the estimate, then correct to the actual. If the actual comes back bigger than the estimate, the ledger absorbs the difference — which means the &lt;em&gt;next&lt;/em&gt; call is judged against reality, not against my optimism. My estimates were reliably low. That's not a bug I fixed, it's a bias I compensated for.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cap is not the alert
&lt;/h2&gt;

&lt;p&gt;Here's the mistake, and it's the reason this post exists.&lt;/p&gt;

&lt;p&gt;When the daily cap tripped, my code did the correct thing: it raised &lt;code&gt;Denied("daily budget exceeded — freezing")&lt;/code&gt;, wrote an audit entry, and returned. Clean, fail-closed, no drama.&lt;/p&gt;

&lt;p&gt;And it logged that at &lt;code&gt;INFO&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;My alerting rules only looked at &lt;code&gt;ERROR&lt;/code&gt; and above, plus a couple of specific failure classes I'd enumerated. So from the alerting system's point of view, the night was perfect. No error. Nothing to see. The one job that did notify me was the morning job, and it failed because the draft file it expected never got written — a &lt;em&gt;downstream&lt;/em&gt; symptom that arrived five hours after the causal event.&lt;/p&gt;

&lt;p&gt;The lesson generalizes past my stack: &lt;strong&gt;a fail-closed cap is a refusal, and refusals are not exceptions.&lt;/strong&gt; From inside the system, refusing is the system working. From outside, refusing is the system not working. Only a human can tell you which — and the human needs to be told at the moment it happens.&lt;/p&gt;

&lt;p&gt;The fix was one line of category, plus one line of routing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# this was the bug, dressed as good hygiene
&lt;/span&gt;&lt;span class="n"&gt;log&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;daily budget exceeded — freezing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# what it is now: a refusal is an incident, not a status
&lt;/span&gt;&lt;span class="n"&gt;log&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;warning&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;BUDGET_FROZEN&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;extra&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;spent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;spent_today&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cap&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;per_day&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And &lt;code&gt;BUDGET_FROZEN&lt;/code&gt; is now on the alert list — not the page-me-now list, but the tell-me-immediately list. There's a real distinction there. A budget freeze at 01:40 doesn't mean wake the human; it means the pipeline is dead until morning, and the human gets to know &lt;em&gt;before&lt;/em&gt; the morning job fails, with the number attached. Twenty-three dollars is a shrug. Twenty-three dollars discovered at 04:15 by a confused downstream job is a five-hour blind spot.&lt;/p&gt;

&lt;h2&gt;
  
  
  The second bug: the loop had no shape
&lt;/h2&gt;

&lt;p&gt;The retry loop is a separate story and it's more embarrassing, so here it is quickly.&lt;/p&gt;

&lt;p&gt;My escalation logic was: &lt;em&gt;if confidence is low, retry on the bigger model.&lt;/em&gt; I tested it by mocking a model that returns "low confidence" once and then confidently. So in my head, the loop was a single hop.&lt;/p&gt;

&lt;p&gt;In reality, the bigger model also said low confidence, because the section genuinely was ambiguous, and the section never got less ambiguous, because I never changed anything between attempts. It retried the same query against the same input with the same context and hoped for a different answer. That is, definitionally, the thing I have &lt;a href="https://dev.to/samhartley_dev/i-added-a-circuit-breaker-to-my-ai-agent-setup-it-caught-three-outages-this-month-5fhm"&gt;a circuit breaker for&lt;/a&gt; — and the breaker didn't catch it, because every individual call &lt;em&gt;succeeded&lt;/em&gt;. A 200 is a success. The breaker watches for failure, and my loop produced nothing but successes until the money ran out.&lt;/p&gt;

&lt;p&gt;So the retry loop got two additions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Attempt cap.&lt;/strong&gt; Three tries, hard stop. Not "until it works."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Escalation must change the input.&lt;/strong&gt; Each retry has to add something — more context, a decomposition step, a different prompt. A retry that repeats the previous attempt byte-for-byte isn't a retry, it's a subscription.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And the effect-count check I now put on every job — &lt;a href="https://dev.to/samhartley_dev/my-scheduled-agent-ran-40-times-and-did-nothing-heres-the-assertion-that-fixed-it-50g2"&gt;the assertion that a run must prove it produced something&lt;/a&gt; — is exactly what would have caught this at 01:45 instead of 04:10. The job was producing effects, so it would have passed. But the loop had no &lt;em&gt;cost&lt;/em&gt; assertion, and cost is an effect too if you're paying for it by the token.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two metrics that actually predict this
&lt;/h2&gt;

&lt;p&gt;Post-incident, I stopped watching "spend total" as my primary number and started watching two ratios:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Spend per unit of work.&lt;/strong&gt; Dollars per draft, per filing, per summary. This is the number that would have been screaming. My absolute spend looked fine for the first hour — forty cents an hour is normal-ish. It was spend &lt;em&gt;per draft&lt;/em&gt; that went vertical, because the job was producing zero drafts while burning tokens. Same shape as the flatline I wrote about in the silent-failure post: the aggregate looked normal, the ratio was on fire.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Retry ratio: retries divided by first attempts.&lt;/strong&gt; Mine sits near 0.05 on a normal night. Anything above ~0.3 means a loop or a degraded model, and it's visible in minutes rather than dollars. Cheap to compute, no new infrastructure, and it's the only number that moved before the bill did.&lt;/p&gt;

&lt;p&gt;One more, and it's the uncomfortable one: &lt;strong&gt;the alert you're not watching is the alert you've muted.&lt;/strong&gt; I'd been getting budget-adjacent notices for weeks and had stopped reading them, because most were routine. That's not an alerting failure, it's a calibration failure — and it's the same terminal state I wrote about with &lt;a href="https://dev.to/samhartley_dev/timeout-means-no-the-one-rule-that-makes-ai-agent-approval-gates-actually-work-1j1o"&gt;approval gates that ask too often&lt;/a&gt;. Any channel that asks more than a few times a month stops being oversight. If you're drowning in routine notices, the fix is fewer notices, not more attention.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd tell you to copy
&lt;/h2&gt;

&lt;p&gt;If you run anything that spends money on your behalf — cloud LLM calls, API credits, compute — here's the short version:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Caps in the executor, not the model.&lt;/strong&gt; A counter and a line. No model gets a vote on whether to spend.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Counters survive restarts.&lt;/strong&gt; If it lives in memory, it isn't a cap.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reserve, then settle.&lt;/strong&gt; Judge the next call against the real number, not your estimate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A refusal is an incident.&lt;/strong&gt; Log it at a level your alerts can see, with the number in the message.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cap your retries, and make each retry change something.&lt;/strong&gt; A repeated call isn't a retry.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Watch the ratio, not the total.&lt;/strong&gt; Spend per unit of work is the early-warning number; the total is the receipt.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The cap worked exactly as designed. The $23 was the cap &lt;em&gt;doing its job&lt;/em&gt; — the difference between a runaway bill and a bounded one. The failure was that I didn't hear it close. A guardrail you don't get alerted about is a guardrail you're trusting on faith, and I'd rather have the boring version: it stops, and it tells me it stopped.&lt;/p&gt;




&lt;p&gt;Curious how other people handle this — do you alert on budget refusals immediately, or roll them into a digest so you don't train yourself to ignore them? I keep going back and forth. Drop a comment with what's worked.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Part of my Building in Public series — previously: &lt;a href="https://dev.to/samhartley_dev/my-scheduled-agent-ran-40-times-and-did-nothing-heres-the-assertion-that-fixed-it-50g2"&gt;the assertion that catches runs that do nothing&lt;/a&gt;, &lt;a href="https://dev.to/samhartley_dev/timeout-means-no-the-one-rule-that-makes-ai-agent-approval-gates-actually-work-1j1o"&gt;timeout means no on approval gates&lt;/a&gt;, &lt;a href="https://dev.to/samhartley_dev/my-paper-trading-bot-went-live-this-week-heres-the-guardrail-stack-i-built-first-e48"&gt;the guardrail stack before going live&lt;/a&gt;, and &lt;a href="https://dev.to/samhartley_dev/i-added-a-circuit-breaker-to-my-ai-agent-setup-it-caught-three-outages-this-month-5fhm"&gt;a circuit breaker that caught three outages&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>cloud</category>
      <category>llm</category>
      <category>monitoring</category>
    </item>
    <item>
      <title>My Scheduled Agent Ran 40 Times and Did Nothing — Here's the Assertion That Fixed It</title>
      <dc:creator>Sam Hartley</dc:creator>
      <pubDate>Wed, 23 Sep 2026 08:30:57 +0000</pubDate>
      <link>https://dev.to/samhartley_dev/my-scheduled-agent-ran-40-times-and-did-nothing-heres-the-assertion-that-fixed-it-50g2</link>
      <guid>https://dev.to/samhartley_dev/my-scheduled-agent-ran-40-times-and-did-nothing-heres-the-assertion-that-fixed-it-50g2</guid>
      <description>&lt;h1&gt;
  
  
  My Scheduled Agent Ran 40 Times and Did Nothing — Here's the Assertion That Fixed It
&lt;/h1&gt;

&lt;p&gt;I found this one by accident, which is the only way you find it.&lt;/p&gt;

&lt;p&gt;Every night at 02:30, a job pulls new filings from an API, normalizes them, and drops them in a queue that feeds my morning briefing. It had been running for about six weeks. Every run: exit code 0. No errors in the log. The scheduler dashboard showed forty consecutive green runs. If you'd asked me, I'd have said that job was my most reliable piece of automation.&lt;/p&gt;

&lt;p&gt;Then I noticed my briefing had gotten thin. Not empty — thin. Some days two items, some days none, and I'd been assuming "slow news." I finally got suspicious, went to the log, and realized I had absolutely nothing to be suspicious &lt;em&gt;with&lt;/em&gt;: the log said &lt;code&gt;fetched items, wrote to queue&lt;/code&gt; on every run, and that's it. No counts. No IDs. Just a sentence I'd written myself, once, and never verified.&lt;/p&gt;

&lt;p&gt;So I added one line — a count — and reran it manually.&lt;/p&gt;

&lt;p&gt;Zero. Every night. Forty times.&lt;/p&gt;

&lt;p&gt;The HTTP call returned 200 with an empty array, because a parameter name in the query had been renamed on their side (&lt;code&gt;since&lt;/code&gt; → &lt;code&gt;from_date&lt;/code&gt;) and the API's response to an unknown parameter was to cheerfully ignore it. My &lt;code&gt;for&lt;/code&gt; loop iterated over an empty list and completed successfully, which is exactly what a &lt;code&gt;for&lt;/code&gt; loop over an empty list does. Nothing raised. Nothing retried. Nothing was logged as a problem, because as far as every layer of my stack was concerned, nothing &lt;em&gt;was&lt;/em&gt; a problem.&lt;/p&gt;

&lt;p&gt;That's the failure mode I want to write about, because I'd covered the other ones: I have &lt;a href="https://dev.to/samhartley_dev/i-added-a-circuit-breaker-to-my-ai-agent-setup-it-caught-three-outages-this-month-5fhm"&gt;a circuit breaker for hard failures&lt;/a&gt;, &lt;a href="https://dev.to/samhartley_dev/i-built-a-health-monitor-for-my-ai-agents-now-they-tell-me-when-theyre-dying-2333"&gt;a health monitor for agents that are dying&lt;/a&gt;, and &lt;a href="https://dev.to/samhartley_dev/my-cron-job-double-fired-and-placed-the-same-order-twice-heres-my-postmortem-5cib"&gt;dedup for jobs that fire twice&lt;/a&gt;. All three of those watch for something &lt;em&gt;happening&lt;/em&gt;. None of them watch for something &lt;em&gt;not&lt;/em&gt; happening. And "not happening" doesn't throw.&lt;/p&gt;

&lt;h2&gt;
  
  
  Errors page you. Silence doesn't.
&lt;/h2&gt;

&lt;p&gt;The reason this took six weeks is structural, not lazy. An exception has a moment: it fires, a handler catches it, an alert goes out. A silent no-op has no moment. It has a &lt;em&gt;shape over time&lt;/em&gt;, and only if you're graphing the right thing.&lt;/p&gt;

&lt;p&gt;My monitoring was graphing the wrong thing. I was tracking run status (&lt;code&gt;success&lt;/code&gt;/&lt;code&gt;failed&lt;/code&gt;) and run duration. Both were perfect. What I wasn't tracking was the only number that mattered: &lt;strong&gt;effects produced&lt;/strong&gt;. Runs completed. Effects were zero. I had a dashboard that was green because it was measuring the system's opinion of itself.&lt;/p&gt;

&lt;p&gt;The fix isn't more logging. It's a different question: not "did it run?" but "did it change anything?"&lt;/p&gt;

&lt;h2&gt;
  
  
  Rule one: assert the effect, not the attempt
&lt;/h2&gt;

&lt;p&gt;Every job I write now has to declare what a successful run &lt;em&gt;looks like&lt;/em&gt; in terms of side effects — rows written, messages sent, files created, IDs emitted. Then the job asserts it before it can report success.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;assert_effect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;produced&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;expected&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;&amp;gt;=1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;expected&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;&amp;gt;=1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;produced&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;JobProducedNothingError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: ran fine, produced &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;produced&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; effects&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;log&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;%s: effect ok (%s)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;produced&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The point isn't the helper. The point is that a job can now fail for the crime of doing nothing. Before, "did nothing" and "did everything" were the same outcome — &lt;code&gt;success&lt;/code&gt; — which means my success signal was measuring the wrong universe.&lt;/p&gt;

&lt;p&gt;Two things I had to get right, and got wrong first:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An assertion that can't fail isn't an assertion.&lt;/strong&gt; My first pass was &lt;code&gt;assert len(rows) &amp;gt;= 0&lt;/code&gt;. That's a tautology wearing a safety vest. It passed on the empty list and I felt protected for another two days. The assertion has to be able to actually fail on the thing you're worried about — which sounds obvious until you write one at 1 AM.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Assert on the effect, not the attempt.&lt;/strong&gt; &lt;code&gt;resp.status_code == 200&lt;/code&gt; is an attempt check. &lt;code&gt;len(written_ids) &amp;gt; 0&lt;/code&gt; is an effect check. The whole incident was a case where every attempt check passed and the effect was zero. If your assertion can be satisfied by a system that is doing nothing, it's not an assertion, it's decoration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rule two: expected-empty is not unexpected-empty
&lt;/h2&gt;

&lt;p&gt;Here's where I nearly overcorrected, and it's the subtlest part.&lt;/p&gt;

&lt;p&gt;Some of my jobs legitimately do nothing most of the time. A watcher that looks for new items in a feed &lt;em&gt;should&lt;/em&gt; find zero items on a quiet day. If I make "zero effects" a hard failure everywhere, those jobs scream every night and I'm back to a muted alert channel — which &lt;a href="https://dev.to/samhartley_dev/timeout-means-no-the-one-rule-that-makes-ai-agent-approval-gates-actually-work-1j1o"&gt;I already learned is worse than no alerts&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;So the assertion can't be "produced &amp;gt; 0". It has to be &lt;strong&gt;"the source was successfully queried &lt;em&gt;and&lt;/em&gt; the emptiness is trustworthy."&lt;/strong&gt; That's a different check: the request succeeded, returned the expected shape, and the emptiness was an explicit, well-formed empty rather than a default. In my case the giveaway was available the whole time: the response had no pagination envelope. A real empty result would have come back with a cursor and a total. A silently-ignored-parameter empty came back as a bare &lt;code&gt;[]&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Practical form: for jobs that may legitimately produce nothing, assert on the envelope, not the count.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;api&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/filings&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;params&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;from_date&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;since&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="nf"&gt;isinstance&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;items&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;UnexpectedShapeError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;got &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;type&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;80&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;items&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;items&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;  &lt;span class="c1"&gt;# zero here is fine — we proved the query was honored
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now "quiet day" and "the API stopped honoring my request" look different, which is the entire job.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rule three: three states, not two
&lt;/h2&gt;

&lt;p&gt;This is what finally made it usable for both kinds of jobs. Every run now ends in one of three states, and they are not the same thing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;succeeded_with_effect&lt;/code&gt; — it did the thing, with a count attached.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;no_work_needed&lt;/code&gt; — it did nothing &lt;strong&gt;and proved that nothing was the correct answer&lt;/strong&gt; (envelope check above passed).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;failed&lt;/code&gt; — including &lt;code&gt;JobProducedNothingError&lt;/code&gt;, i.e. it ran, and the absence of effect was not explained.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Only the third one alerts. The second one is a normal, boring, expected outcome that shows up in the briefing as a single quiet line. Before this split, "did nothing" was indistinguishable from "worked" — and the alerting I did have was tuned to a world where those were the same.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that actually caught things
&lt;/h2&gt;

&lt;p&gt;Beyond the original bug, the effect counters caught two more silent failures in the next month: a token that expired and got silently replaced with an anonymous, lower-privilege session (200s, empty results), and a path change that made a downstream job read an empty directory and write a valid, empty file. Neither threw. Both were invisible to status-based monitoring. Both showed up instantly as a flatline in effects.&lt;/p&gt;

&lt;p&gt;The dashboard line I now look at first isn't uptime or error rate. It's effects per run, per job. A job that runs and produces nothing, forever, is not a healthy job — it's a green light with the bulb unscrewed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The transferable bit
&lt;/h2&gt;

&lt;p&gt;If you take one thing: &lt;strong&gt;"no error" is not "worked."&lt;/strong&gt; Your scheduler, your orchestrator, and your CI will happily report success for a process that queried nothing, wrote nothing, and sent nothing, because that is exactly what they were asked to measure.&lt;/p&gt;

&lt;p&gt;Whatever your stack is — a nightly ETL, a webhook consumer, a CI step, an agent that posts things — pick one number per unit of work that would be zero if the work silently stopped. Assert on it. Track it over time. And make sure that some of your alerts can fire on a flatline, not just a spike. Spikes are loud and self-reporting. Flatlines need someone to be watching for them on purpose.&lt;/p&gt;




&lt;p&gt;Curious how other people handle this: do you assert on effect counts in production jobs, or do you rely on downstream freshness checks (data arriving late → alert)? Both seem to have holes. Drop a comment with what's caught silent failures for you.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Part of my Building in Public series — previously: &lt;a href="https://dev.to/samhartley_dev/my-paper-trading-bot-went-live-this-week-heres-the-guardrail-stack-i-built-first-e48"&gt;the guardrail stack I built before going live&lt;/a&gt;, &lt;a href="https://dev.to/samhartley_dev/timeout-means-no-the-one-rule-that-makes-ai-agent-approval-gates-actually-work-1j1o"&gt;timeout means no on approval gates&lt;/a&gt;, &lt;a href="https://dev.to/samhartley_dev/i-added-a-circuit-breaker-to-my-ai-agent-setup-it-caught-three-outages-this-month-5fhm"&gt;a circuit breaker that caught three outages&lt;/a&gt;, and &lt;a href="https://dev.to/samhartley_dev/i-built-a-health-monitor-for-my-ai-agents-now-they-tell-me-when-theyre-dying-2333"&gt;a health monitor that tells me when my agents are dying&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>python</category>
      <category>automation</category>
    </item>
    <item>
      <title>Timeout Means No — The One Rule That Makes AI Agent Approval Gates Actually Work</title>
      <dc:creator>Sam Hartley</dc:creator>
      <pubDate>Mon, 21 Sep 2026 08:33:08 +0000</pubDate>
      <link>https://dev.to/samhartley_dev/timeout-means-no-the-one-rule-that-makes-ai-agent-approval-gates-actually-work-1j1o</link>
      <guid>https://dev.to/samhartley_dev/timeout-means-no-the-one-rule-that-makes-ai-agent-approval-gates-actually-work-1j1o</guid>
      <description>&lt;p&gt;My trading bot went live recently, and I wrote up &lt;a href="https://dev.to/samhartley_dev/my-paper-trading-bot-went-live-this-week-heres-the-guardrail-stack-i-built-first-e48"&gt;the guardrail stack I built first&lt;/a&gt;: hard budget caps, idempotency keys, a dead-man switch, and an approval gate for anything unusual. One of those has quietly become my favorite, and it's the least technical of the bunch.&lt;/p&gt;

&lt;p&gt;The approval gate. Not the asking-for-permission part — the part everyone, including me, gets wrong on the first try. The rule that makes the whole thing work:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If I don't answer, the answer is no.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Sounds obvious now. It wasn't. I shipped two versions of that gate before this one, both politely broken in ways that only show up when the human is asleep. Since the pattern generalizes to any agent that touches the real world — posts, payments, deletes — it's worth writing down how I got here.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 2 AM Test
&lt;/h2&gt;

&lt;p&gt;Back when the bot was still in paper mode, it flagged a setup as unusual — Tier 3 in my system, "pause and wait for a button tap." The request went out at 2:14 AM. I was asleep, as humans are.&lt;/p&gt;

&lt;p&gt;Version one of the gate treated "no answer" as "still pending." The intent sat there. Polite. Patient. And because my executor is single-threaded by design — one intent at a time, in order — everything queued behind it sat there too. I woke up to a stopped pipeline and a missing &lt;a href="https://dev.to/samhartley_dev/i-built-a-morning-briefing-bot-in-50-lines-of-python-heres-why-i-check-telegram-before-email-now-3fhh"&gt;morning briefing&lt;/a&gt;, because the briefing job was stuck behind the stuck job.&lt;/p&gt;

&lt;p&gt;Here's the thing I had to admit to myself: a blocked agent isn't a safe agent. It's a stopped one. "Doing nothing" feels like the conservative failure mode until you realize the bot also stopped managing the positions it already had open. Safe would have been: reject, move on, keep running.&lt;/p&gt;

&lt;p&gt;So version two overcorrected in the opposite direction: no answer → ask again, every 15 minutes, until answered. Reminders, right? Humans like reminders.&lt;/p&gt;

&lt;p&gt;I woke up to 28 notifications about the same trade. I didn't read notification four, let alone twenty-eight. I swiped the fourth away half-awake and then spent a genuinely dangerous moment trying to remember whether what I'd swiped was approve or dismiss. That's when it clicked: a notification channel that nags gets muted, and a muted approval channel is worse than none — because it still &lt;em&gt;feels&lt;/em&gt; like oversight.&lt;/p&gt;

&lt;p&gt;Version three is what runs now:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Ask once, with the full intent, the reasoning, and a deadline in plain text: "Reply by 03:14 or this is rejected."&lt;/li&gt;
&lt;li&gt;One reminder at the halfway mark. Then silence.&lt;/li&gt;
&lt;li&gt;Deadline passes with no reply → the intent is rejected, written to the audit log as &lt;code&gt;rejected_timeout&lt;/code&gt;, and listed in the next morning briefing.&lt;/li&gt;
&lt;li&gt;The queue never blocks on an approval. The next job runs immediately.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That night, the bot asked, I slept, the trade didn't happen, and the briefing told me about it over coffee with the reasoning attached. I read it, decided the bot was right to want the trade, and entered it manually an hour later when the setup was still valid. Nothing was lost except the illusion that 2 AM me makes good decisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Code, Which Is Embarrassingly Short
&lt;/h2&gt;

&lt;p&gt;The whole gate is maybe thirty lines. This is the core of it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ApprovalGate&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;WAIT&lt;/span&gt;         &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;3600&lt;/span&gt;   &lt;span class="c1"&gt;# total patience, in seconds
&lt;/span&gt;    &lt;span class="n"&gt;REMIND_AFTER&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1800&lt;/span&gt;   &lt;span class="c1"&gt;# one nudge, then silence
&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;request&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;intent&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;deadline&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;WAIT&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;send&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;intent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                  &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Reply by &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;fmt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;deadline&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; or this is rejected. &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                  &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;No reply = no trade.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;await_reply&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;deadline&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;remind_after&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;REMIND_AFTER&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="n"&gt;UNANSWERED&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;audit&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;intent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;outcome&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rejected_timeout&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;Denied&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;timeout means no&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="n"&gt;REJECTED&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;audit&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;intent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;outcome&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rejected_by_human&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;Denied&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;human said no&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;Approved&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;edits&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# edits may shrink the intent, never grow it
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two details in there matter more than the mechanics:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The deadline is in the message.&lt;/strong&gt; "No reply = no trade" is written down, every time. The system's behavior when I'm absent isn't a surprise — it's a stated contract inside the request itself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Approvals can only shrink.&lt;/strong&gt; When I reply, I can approve, reject, or approve at a smaller size. There is no path where my reply makes the action &lt;em&gt;bigger&lt;/em&gt;. The gate edits down or blocks; it never amplifies. That asymmetry closes a whole class of "I replied with a typo and the bot did something heroic" incidents.&lt;/p&gt;

&lt;h2&gt;
  
  
  "Timeout Means Yes" Is Always One Boolean Away
&lt;/h2&gt;

&lt;p&gt;Nobody writes &lt;code&gt;if timeout: go_ahead()&lt;/code&gt; on purpose. But the pressure is real, and it never announces itself as recklessness. It sounds like: "the pipeline keeps stalling overnight, let's add a default," or "if he doesn't respond in an hour, he probably would have said yes." One boolean, one config value, and your fail-closed gate is fail-open — and nothing about the system looks different until the night it matters.&lt;/p&gt;

&lt;p&gt;The reframe that fixed my thinking: an unanswered request is information. The human is asleep, or busy, or deliberately ignoring you. All three of those are "no." Not "maybe." Not "retry." No.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Logs Taught Me
&lt;/h2&gt;

&lt;p&gt;Two numbers tell you whether a gate like this is healthy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How often it asks.&lt;/strong&gt; Tier 3 requests should be rare — mine currently fire about once a week. If they fire daily, the problem isn't the human, it's upstream: your thresholds and caps are miscalibrated and the gate is absorbing the overflow. A gate that asks constantly gets rubber-stamped, and reflex-tapping approve is the terminal failure state of every approval system ever built. Gate volume is a health metric for the whole stack.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How often I disagree with the rejects.&lt;/strong&gt; The rejection log is a free dataset on where the bot's world model and reality diverge. In the first two weeks of logging timeouts, I noticed the bot kept flagging perfectly ordinary setups as unusual. Turned out a volatility threshold was stale and everything looked scary to it. The gate didn't just prevent bad trades — the rejects pointed me at a config bug I'd been living with. Rejections aren't noise in the system. They're the system talking.&lt;/p&gt;

&lt;p&gt;One more, and I'm embarrassed by this one: I originally logged timeout-rejections as warnings, so my &lt;a href="https://dev.to/samhartley_dev/i-built-a-health-monitor-for-my-ai-agents-now-they-tell-me-when-theyre-dying-2333"&gt;health monitor&lt;/a&gt; kept paging me about them. Wrong category. A timeout-rejection is the system working exactly as designed — it belongs in the routine morning briefing, not in alerting. The alert should fire only when the gate can't reach me at all. If your monitoring treats correct behavior as an incident, you'll mute the monitoring, and then the real incidents arrive as surprises.&lt;/p&gt;

&lt;h2&gt;
  
  
  It's a Decision Record, Not a Security Boundary
&lt;/h2&gt;

&lt;p&gt;One nuance worth being precise about, because approval flows have a way of quietly becoming the only thing between an agent and disaster: the gate is not what keeps me safe. The executor's hard caps are. Even with a forged or fat-fingered "yes," the position-size cap, the exposure cap, and the daily-loss cap still bound what can happen. The gate decides &lt;em&gt;whether&lt;/em&gt;; the caps decide &lt;em&gt;how much&lt;/em&gt;. Two systems, one boring outcome.&lt;/p&gt;

&lt;p&gt;It's the same reason the gate plays nicely with the dead-man switch: if the bot's heartbeat drops, pending approvals auto-reject along with everything else — a half-alive bot's outstanding questions can't be trusted either.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Pattern, Without the Trading
&lt;/h2&gt;

&lt;p&gt;You don't need a bot that moves money for this to apply. Anything an agent does that's hard to undo — publishing a post, sending an email on your behalf, deleting records, touching production — fits the same three-tier table:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Reversible&lt;/strong&gt; → act alone, log everything&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bounded and routine&lt;/strong&gt; → act alone, but send the receipt first&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unusual or irreversible&lt;/strong&gt; → ask once, remind once, silence means no&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And the one line worth stealing, in any stack: &lt;strong&gt;timeout means no, and "no" is a first-class outcome, not an error.&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;Have you wired approval flows into your agents? I'm curious how other people handle the "human is asleep" case — bounded waits, escalation to a second human, or just letting it queue until morning? Drop a comment, genuinely curious.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Part of my Building in Public series — previously: &lt;a href="https://dev.to/samhartley_dev/my-paper-trading-bot-went-live-this-week-heres-the-guardrail-stack-i-built-first-e48"&gt;the guardrail stack I built before going live&lt;/a&gt;, &lt;a href="https://dev.to/samhartley_dev/my-cron-job-double-fired-and-placed-the-same-order-twice-heres-my-postmortem-5cib"&gt;a postmortem of a double-fired cron job&lt;/a&gt;, &lt;a href="https://dev.to/samhartley_dev/i-added-a-circuit-breaker-to-my-ai-agent-setup-it-caught-three-outages-this-month-5fhm"&gt;a circuit breaker for agent outages&lt;/a&gt;, and &lt;a href="https://dev.to/samhartley_dev/i-built-a-health-monitor-for-my-ai-agents-now-they-tell-me-when-theyre-dying-2333"&gt;a health monitor that tells me when my agents are dying&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>automation</category>
      <category>devops</category>
    </item>
    <item>
      <title>My Paper Trading Bot Went Live This Week — Here's the Guardrail Stack I Built First</title>
      <dc:creator>Sam Hartley</dc:creator>
      <pubDate>Sat, 19 Sep 2026 08:36:16 +0000</pubDate>
      <link>https://dev.to/samhartley_dev/my-paper-trading-bot-went-live-this-week-heres-the-guardrail-stack-i-built-first-e48</link>
      <guid>https://dev.to/samhartley_dev/my-paper-trading-bot-went-live-this-week-heres-the-guardrail-stack-i-built-first-e48</guid>
      <description>&lt;h1&gt;
  
  
  My Paper Trading Bot Went Live This Week — Here's the Guardrail Stack I Built First
&lt;/h1&gt;

&lt;p&gt;Confession time: my trading bot has been "almost ready for real money" for about four months.&lt;/p&gt;

&lt;p&gt;Classic side-project arc. I built a scanner, it found opportunities, it paper-traded them beautifully. The equity curve went up and to the right. I showed a friend the chart. He asked the awkward question: "So when are you turning it on for real?"&lt;/p&gt;

&lt;p&gt;The honest answer was: not until it bores me.&lt;/p&gt;

&lt;p&gt;Paper trading tells you whether your logic works. It tells you nothing about what happens when your logic is wrong at 3 AM with real money on the line. So before flipping the switch, I spent a week building the boring parts — the guardrails. This is that stack, roughly ordered by how paranoid each one is.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Paper Trading Actually Tests (Spoiler: Not the Scary Parts)
&lt;/h2&gt;

&lt;p&gt;Paper mode is a simulator with the serial numbers filed off. Mine was optimistic in four specific ways:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No slippage.&lt;/strong&gt; Paper fills happen at the displayed price. Real fills happen at whatever the order book feels like.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No fees.&lt;/strong&gt; Small strategies live and die on fees. Paper mode is a world where the house doesn't take a cut.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Infinite liquidity.&lt;/strong&gt; Paper mode never gets a partial fill or a rejected order.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No consequences.&lt;/strong&gt; The big one. In paper mode, a bug is a curiosity. In live mode, a bug is an incident.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I've already lived the last one in miniature: &lt;a href="https://dev.to/samhartley_dev/my-cron-job-double-fired-and-placed-the-same-order-twice-heres-my-postmortem-5cib"&gt;my scheduler once double-fired and placed the same order twice&lt;/a&gt;. In paper, that's a weird chart. Live, that's exposure you never chose to take.&lt;/p&gt;

&lt;p&gt;So before going live, I built six guardrails. None of them are clever. All of them are boring. That's the point.&lt;/p&gt;

&lt;h2&gt;
  
  
  Guardrail 0: The Model Doesn't Decide What It's Allowed to Do
&lt;/h2&gt;

&lt;p&gt;Quick architecture note, because everything below depends on it.&lt;/p&gt;

&lt;p&gt;My setup separates the brain from the hands. The "brain" — an LLM doing analysis — never talks to the exchange. It writes an &lt;em&gt;intent&lt;/em&gt;: "open position in X, size S, stop at level L." A separate, dumb executor picks intents up. The executor is ~200 lines of Python with no model in it. The model can hallucinate whatever it wants; the executor only knows what's on the allowlist.&lt;/p&gt;

&lt;p&gt;If your agent stack has the model calling money-moving APIs directly, this is the first thing I'd change. The part that &lt;em&gt;decides&lt;/em&gt; should be as far away as possible from the part that &lt;em&gt;does&lt;/em&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Guardrail 1: Hard Budget Caps, Enforced in the Executor
&lt;/h2&gt;

&lt;p&gt;The executor has three numbers burned into it: a max position size, a max total exposure, and a max daily loss. Any intent that would breach one is rejected and logged — no model consultation, no retry, no "are you sure?"&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Limits&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;MAX_POSITION&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;50&lt;/span&gt;    &lt;span class="c1"&gt;# % of allocated capital, single position
&lt;/span&gt;    &lt;span class="n"&gt;MAX_EXPOSURE&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt;   &lt;span class="c1"&gt;# % total, including leverage
&lt;/span&gt;    &lt;span class="n"&gt;MAX_DAILY_LOSS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;    &lt;span class="c1"&gt;# % — breach freezes everything
&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;check_intent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;intent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;intent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size_pct&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;Limits&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;MAX_POSITION&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;Reject&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;position too big&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;exposure_pct&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;intent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;size_pct&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;Limits&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;MAX_EXPOSURE&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;Reject&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;exposure cap&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;day_pnl_pct&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;Limits&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;MAX_DAILY_LOSS&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;Reject&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;daily loss cap tripped — frozen until reset&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;OK&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The daily loss cap is the one I care about most. Strategies don't fail slowly — they fail on one bad day when everything correlates. The cap turns "my bot blew up" into "my bot stopped."&lt;/p&gt;

&lt;h2&gt;
  
  
  Guardrail 2: Rate Limits and Cooldown Windows
&lt;/h2&gt;

&lt;p&gt;Autonomous agents in a loop are fast. Humans are not. So the executor enforces pace:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Max N new positions per hour.&lt;/strong&gt; Mine is a single-digit number, on purpose.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cooldown after any closed loss.&lt;/strong&gt; No instant re-entry — that's how revenge trades happen, and bots are not immune to the sunk-cost loop. They just express it in code.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Consecutive-loss circuit break.&lt;/strong&gt; After M losses in a day, the bot goes read-only until I flip it back on.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last rule wasn't in v1. It is now, because I once watched my paper bot "scale into" a losing position five times in a row. It was following its rules perfectly. The rules were the problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Guardrail 3: Idempotency Keys on Every Action
&lt;/h2&gt;

&lt;p&gt;Every intent gets a unique key before execution, and the executor checks it against a ledger of keys it has already acted on:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;intent&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;intent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;day&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;intent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;asset&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;intent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;action&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;ledger&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;seen&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;Skipped&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;duplicate intent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;ledger&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;record&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;            &lt;span class="c1"&gt;# written BEFORE acting
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;broker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;place&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;intent&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two details matter. First, the key is derived from the &lt;em&gt;decision&lt;/em&gt; (day, asset, action), not from a random ID — so a replayed or double-fired job produces the same key and gets dropped, which is exactly what saved me after &lt;a href="https://dev.to/samhartley_dev/my-cron-job-double-fired-and-placed-the-same-order-twice-heres-my-postmortem-5cib"&gt;the double-fire incident&lt;/a&gt;. Second, the key is recorded &lt;em&gt;before&lt;/em&gt; the action, not after. If the process dies between the write and the fill, the retry gets skipped — I'd rather reconcile one orphan intent than double-spend.&lt;/p&gt;

&lt;h2&gt;
  
  
  Guardrail 4: Action Tiers With an Approval Gate
&lt;/h2&gt;

&lt;p&gt;Not all actions are equally scary, so they get sorted into tiers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Tier 1 — reversible:&lt;/strong&gt; scans, reads, paper actions. The bot does these freely.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tier 2 — bounded:&lt;/strong&gt; small real positions, inside the caps. The bot does these alone, but I get a message with the full intent &lt;em&gt;before&lt;/em&gt; execution.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tier 3 — unusual or irreversible:&lt;/strong&gt; anything at the caps, anything outside normal parameters, anything the bot itself flags as low-confidence. This tier pauses and waits for a button tap on my phone.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The approval message shows the full intent, the reasoning, and — this is the important part — what happens if I don't answer. The default is nothing. Timeout equals rejection. An unanswered request must never silently become a yes.&lt;/p&gt;

&lt;p&gt;Tier 3 has caught exactly zero dramatic events so far. That's the correct number. The gate exists for the day something &lt;em&gt;almost&lt;/em&gt; makes sense.&lt;/p&gt;

&lt;h2&gt;
  
  
  Guardrail 5: The Dead-Man Switch
&lt;/h2&gt;

&lt;p&gt;The bot must check in on a heartbeat. Missed check-ins, and it freezes itself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;HEARTBEAT_EVERY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;300&lt;/span&gt;   &lt;span class="c1"&gt;# seconds
&lt;/span&gt;&lt;span class="n"&gt;FREEZE_AFTER&lt;/span&gt;   &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;      &lt;span class="c1"&gt;# missed beats
&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;on_tick&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;missed_beats&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;FREEZE_AFTER&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;freeze&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;reason&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;dead-man switch&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;hung_jobs&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;freeze&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;reason&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;job stuck — state is suspect&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;beat&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The insight I stole from running &lt;a href="https://dev.to/samhartley_dev/i-built-a-health-monitor-for-my-ai-agents-now-they-tell-me-when-theyre-dying-2333"&gt;a health monitor&lt;/a&gt;: a hung bot is more dangerous than a dead one. A dead bot does nothing — annoying, but safe. A half-alive bot might fire one leg of a two-step plan and never execute the other half.&lt;/p&gt;

&lt;p&gt;So the freeze triggers on &lt;em&gt;ambiguity&lt;/em&gt; too, not just silence. A stuck job means state can't be trusted. Freezing means: no new opens. Existing positions still get managed — stops always work — but nothing new starts. And there's a kill command that freezes everything from my phone, tested weekly, because an untested kill switch is a decorative kill switch.&lt;/p&gt;

&lt;h2&gt;
  
  
  Guardrail 6: An Audit Log Written Before Anything Happens
&lt;/h2&gt;

&lt;p&gt;Every intent lands in a local log — timestamp, full payload, the reasoning — &lt;em&gt;before&lt;/em&gt; execution, then gets updated with the result. Every morning I get &lt;a href="https://dev.to/samhartley_dev/i-built-a-morning-briefing-bot-in-50-lines-of-python-heres-why-i-check-telegram-before-email-now-3fhh"&gt;a briefing message&lt;/a&gt; with a reconciliation: what the log says vs. what the account says.&lt;/p&gt;

&lt;p&gt;If those two ever disagree, the bot assumes its own state is wrong, freezes, and waits for me. The account is the source of truth. The log is the source of intent. My job is reading both.&lt;/p&gt;

&lt;h2&gt;
  
  
  Week One, Honestly
&lt;/h2&gt;

&lt;p&gt;So how's the first week of real money? Deeply educational, in the way a small electric shock is educational:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Fees and slippage turned "up and to the right" into "up and slightly to the right."&lt;/strong&gt; The strategy still works, just with worse numbers. Better to learn that at small size than large.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The dead-man switch fired once.&lt;/strong&gt; A job wedged itself after a network blip, the bot froze, and I spent an evening untangling it. Annoying — and exactly the failure mode it was built for. I came home to a frozen bot, not a runaway one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The approval gate caught nothing dramatic.&lt;/strong&gt; The correct amount of drama is zero.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The strangest part is psychological. In paper mode, I checked the bot obsessively, because nothing was watching it but me. Now I check it &lt;em&gt;less&lt;/em&gt; — because the guardrails already asked the questions I would have asked.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next: Shadow Mode
&lt;/h2&gt;

&lt;p&gt;One more thing I'm adding for the next month: every live intent also gets mirrored into paper mode with the same timestamps. At the end, I diff them — paper fill vs. real fill, paper P&amp;amp;L vs. real P&amp;amp;L. That gap &lt;em&gt;is&lt;/em&gt; the true cost of going live, measured instead of guessed. If it shrinks, my execution is tight. If it grows, the bot is paying a hidden tax I can hunt down.&lt;/p&gt;




&lt;p&gt;If you're running agents that touch the real world — orders, payments, posts, deletions — the caps, the tiers, and the dead-man switch are the pieces I'd build first. All of it is a few hundred lines of deeply boring Python, and it converts "I hope it behaves" into "when it misbehaves, it stops."&lt;/p&gt;

&lt;p&gt;Where do you draw the line — full autonomy, approval gates, or still in paper mode? Drop a comment, I'm genuinely curious.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Part of my series on running AI on my own machines — previously: &lt;a href="https://dev.to/samhartley_dev/i-added-a-circuit-breaker-to-my-ai-agent-setup-it-caught-three-outages-this-month-5fhm"&gt;a circuit breaker for agent outages&lt;/a&gt;, &lt;a href="https://dev.to/samhartley_dev/my-cron-job-double-fired-and-placed-the-same-order-twice-heres-my-postmortem-5cib"&gt;a postmortem of a double-fired job&lt;/a&gt;, and &lt;a href="https://dev.to/samhartley_dev/i-built-a-health-monitor-for-my-ai-agents-now-they-tell-me-when-theyre-dying-2333"&gt;a health monitor for my agents&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>devops</category>
      <category>automation</category>
    </item>
    <item>
      <title>I Added a Circuit Breaker to My AI Agent Setup — It Caught Three Outages This Month</title>
      <dc:creator>Sam Hartley</dc:creator>
      <pubDate>Tue, 15 Sep 2026 08:34:53 +0000</pubDate>
      <link>https://dev.to/samhartley_dev/i-added-a-circuit-breaker-to-my-ai-agent-setup-it-caught-three-outages-this-month-5fhm</link>
      <guid>https://dev.to/samhartley_dev/i-added-a-circuit-breaker-to-my-ai-agent-setup-it-caught-three-outages-this-month-5fhm</guid>
      <description>&lt;p&gt;I have a confession: for the longest time, my local AI setup had a very embarrassing failure mode.&lt;/p&gt;

&lt;p&gt;One of my machines goes down — a reboot, a driver update, someone trips over a power cable — and the agents that depend on it just keep firing requests at it. Every single request times out. Each one burns 30-60 seconds waiting for a connection that will never succeed. Meanwhile the task queue piles up, retries stack on retries, and by the time the machine is back, I have a swamp of failed jobs to clean up.&lt;/p&gt;

&lt;p&gt;The fix took an afternoon: a circuit breaker pattern. If you've never heard the term, it comes from electrical engineering — when too much current flows, the breaker trips and the circuit opens. No drama, no fire. You flip it back when things are safe.&lt;/p&gt;

&lt;p&gt;The software version does the same thing for network calls. After X consecutive failures, stop calling. Fail fast, fall back to something that works, and check periodically whether the dead service is back instead of hammering it with every request.&lt;/p&gt;

&lt;p&gt;Since I added it, it's caught three real outages this month. Here's what I learned building it.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Setup (The Short Version)
&lt;/h2&gt;

&lt;p&gt;For context: I run a small AI lab — a Mac mini as the orchestrator, a Windows PC with a GPU for heavy inference, and an old Ubuntu box for background services. The orchestrator has agents that route tasks to the other machines: code generation goes to the GPU box, quick chat stays local, image parsing goes to the Ubuntu box, and so on.&lt;/p&gt;

&lt;p&gt;The routing itself is a dumb little Python function with a model registry and keyword matching. It's the piece that decides &lt;em&gt;where&lt;/em&gt; each request goes. What it didn't have until recently was any concept of &lt;em&gt;whether the destination is actually alive&lt;/em&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem, In Numbers
&lt;/h2&gt;

&lt;p&gt;Here's what a typical outage looked like before:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The GPU machine reboots for a Windows update at 14:02&lt;/li&gt;
&lt;li&gt;Agent jobs keep hitting it: image generation, code refactors, embeddings&lt;/li&gt;
&lt;li&gt;Each request has a 30-second timeout&lt;/li&gt;
&lt;li&gt;The agent framework retries failed requests up to 3 times&lt;/li&gt;
&lt;li&gt;So one failed job = up to 2 minutes of pure waiting, then a cascade of retry jobs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On a bad day, a 15-minute outage would produce 40+ queued tasks that all tried the dead endpoint, failed, and retried — burning a couple of hours of queue time on &lt;em&gt;nothing&lt;/em&gt;. And since some of those jobs had follow-up jobs ("after generating the images, post them to X"), the failures propagated downstream. I'd come back to find half my evening pipeline in a error state because Windows decided 14:02 was a great time to update.&lt;/p&gt;

&lt;p&gt;The worst part: the requests that should have just gone somewhere else — the Ubuntu box runs the same models, just slower on CPU — sat in the same queue behind the doomed ones.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Circuit Breaker, Minimal Version
&lt;/h2&gt;

&lt;p&gt;The pattern has three states. I stole this directly from the classic description (it's all over the resilience literature):&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Closed&lt;/strong&gt; — normal operation. Requests pass through. Failures are counted.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open&lt;/strong&gt; — tripped. Requests fail immediately, no network call at all. This is the "stop hammering the dead thing" state.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Half-open&lt;/strong&gt; — after a cooldown, let one request through as a probe. If it succeeds, close the breaker. If it fails, stay open.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Here's roughly what I run per endpoint (simplified, but functionally what's in production):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;CircuitBreaker&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;threshold&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cooldown&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;threshold&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;threshold&lt;/span&gt;        &lt;span class="c1"&gt;# consecutive failures before tripping
&lt;/span&gt;        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cooldown&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cooldown&lt;/span&gt;          &lt;span class="c1"&gt;# seconds before probing again
&lt;/span&gt;        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;timeout&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;failures&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;closed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;            &lt;span class="c1"&gt;# closed | open | half-open
&lt;/span&gt;        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;opened_at&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;allow_request&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;closed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;open&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;monotonic&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;opened_at&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cooldown&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;half-open&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;  &lt;span class="c1"&gt;# probe time
&lt;/span&gt;                &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;                  &lt;span class="c1"&gt;# fail fast, no network call
&lt;/span&gt;        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;                      &lt;span class="c1"&gt;# a probe is already in flight
&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;record_success&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;failures&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;closed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;record_failure&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;failures&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;half-open&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;failures&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;threshold&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_trip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_trip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;open&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;opened_at&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;monotonic&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;failures&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And in the request wrapper (again, simplified):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;call_endpoint&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;breaker&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;breaker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;allow_request&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;CircuitOpen&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;breaker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; is tripped, failing fast&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;breaker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;raise_for_status&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="n"&gt;breaker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;record_success&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;breaker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;record_failure&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One breaker instance per endpoint, keyed in a dict. The router that decides which machine handles which task now consults the breaker before doing anything else.&lt;/p&gt;

&lt;p&gt;Note the aggressive timeout — 5 seconds, not 60. When a machine is genuinely there, a health probe responds in well under a second. Long timeouts are for slow responses, not dead machines. Distinguishing those two cases is half the value.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Part That Made It Actually Useful: Fallbacks
&lt;/h2&gt;

&lt;p&gt;A circuit breaker that just fails fast is... okay. Your errors arrive sooner. Great. Congratulations.&lt;/p&gt;

&lt;p&gt;The real win comes from pairing the breaker with your fallback chain. My router already knew fallbacks existed — big model down? use the small one. GPU box rented out? CPU box takes over. What it lacked was the &lt;em&gt;signal&lt;/em&gt; to switch. The breaker provides exactly that signal, in constant time, without a single wasted network call.&lt;/p&gt;

&lt;p&gt;So the flow became:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Router picks the destination for a task&lt;/li&gt;
&lt;li&gt;Is the breaker for that destination open? If yes, use the fallback destination &lt;em&gt;immediately&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;Otherwise, call with a short timeout; failures count toward the threshold&lt;/li&gt;
&lt;li&gt;If the breaker trips mid-flight, the next request reroutes — not to a retry queue, but to the fallback&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A 30B code model being unavailable now means requests degrade to the 4B model on the Mac. Code quality drops slightly. &lt;strong&gt;The pipeline keeps flowing.&lt;/strong&gt; Image jobs meant for the GPU get routed to the Ubuntu box and take 3 minutes instead of 20 seconds. Still flowing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three Outages, Zero Cascade Failures
&lt;/h2&gt;

&lt;p&gt;I've had this running for a bit over a month now. Three times the breaker earned its keep:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. The Windows Update (14:02, of course).&lt;/strong&gt; GPU box gone for 17 minutes. Before the breaker: ~40 doomed tasks, cascading failures, an evening of cleanup. With the breaker: it tripped after 3 failures, code and image tasks rerouted to the fallback chain, and when the box came back the half-open probe re-closed it within a minute. I didn't even notice until I read the logs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The flaky PCIe riser.&lt;/strong&gt; One of my GPU cards had a flaky riser connection and the machine hard-froze. When a box is hung (not off), connections don't get refused — they just hang there. The 5-second timeout is what saved this one: three hangs, breaker trips, traffic reroutes. Without the short timeout, every request would have sat there for 60 seconds apiece, pretending to work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. The Ubuntu box's full disk.&lt;/strong&gt; A background service filled the disk with logs, and the service started returning errors while the machine itself stayed up. The breaker tripped on HTTP errors just like it does on connection failures — the pattern doesn't care &lt;em&gt;why&lt;/em&gt; the endpoint is failing, only that it is. Tasks that needed the vision service waited for the half-open probe instead of piling up garbage.&lt;/p&gt;

&lt;p&gt;Three outages. Zero cascade failures. Zero evenings spent cleaning up retry swamps.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Got Wrong First Try
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;I counted total failures instead of consecutive ones.&lt;/strong&gt; My first version tripped the breaker after N failures &lt;em&gt;ever&lt;/em&gt;, which meant a busy day with a handful of one-off timeouts could trip it even though the endpoint was healthy. Consecutive-failure counting with reset-on-success is the fix. One-off failures happen; the breaker should ignore them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My cooldown was too short.&lt;/strong&gt; I started with 10 seconds, which meant a dead endpoint got probed every 10 seconds — fine — but a &lt;em&gt;slowly recovering&lt;/em&gt; endpoint (say, a machine that just booted and is still loading a 30B model into VRAM) would get marked healthy by a TCP-level probe, then fail on the actual request and re-trip. 60 seconds worked better. If the first real request after recovery is heavy, give it room to finish loading.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I forgot to reset breakers on deploy.&lt;/strong&gt; The breaker state lives in memory. When I shipped a new version of the router, all breakers came up closed while the actual machine was still down — instant trip, three failures wasted. The state should be persisted (or at least seeded from a health check at boot). It's on my list.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fail-fast errors need to be distinguishable.&lt;/strong&gt; When a task fails because "the breaker is open," that's a different failure than "the endpoint returned garbage." My first version raised the same generic exception for both, which made my alerting noisy and my logs confusing. One extra exception class fixed it. Now "breaker open" is not even an alert — it's the &lt;em&gt;expected&lt;/em&gt; state during a known outage. The alert fires when a breaker &lt;em&gt;trips&lt;/em&gt;, not on every request it rejects.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Dashboard Bit
&lt;/h2&gt;

&lt;p&gt;Because I was already logging breaker trips, I added one more thing: a daily message to my Telegram with any breakers that tripped in the last 24h and why.&lt;/p&gt;

&lt;p&gt;That message has quietly become one of the most useful signals in my whole setup. Not because outages are frequent — they're not — but because it tells me &lt;em&gt;which&lt;/em&gt; machine is the flaky one this month. September it was the riser. Last month it was power settings putting a box to sleep mid-job. The pattern in the trips is a pattern in my hardware.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should You Build One?
&lt;/h2&gt;

&lt;p&gt;If you have agents calling services and you have no failure isolation between them — yes. Honestly, it's the kind of thing that sounds like Enterprise Java Architecture Astronautics until the first time it saves your evening, and then it's just plumbing.&lt;/p&gt;

&lt;p&gt;You don't need a library. The class above is ~40 lines and covers 90% of the value. If you want batteries included, there are solid packages out there — but read their retry/timeout defaults carefully, because a circuit breaker that waits 60 seconds per failure isn't protecting you from much.&lt;/p&gt;

&lt;p&gt;The mental model that made it click for me: retries ask "will it work &lt;em&gt;this&lt;/em&gt; time?" A circuit breaker asks "is it worth asking?" During an outage, the answer is no, and knowing that up front is worth more than any amount of persistence.&lt;/p&gt;




&lt;p&gt;Have you wired failure isolation into your agents, or are you still in the retry-swamp phase? I'm curious what patterns people are using for multi-machine setups — drop a comment.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This is part of my series on running AI entirely on my own machines — a router for &lt;a href="https://dev.to/samhartley_dev/i-built-a-model-router-that-picks-the-right-ai-for-every-task-heres-why-you-should-too-5la"&gt;picking the right model per task&lt;/a&gt;, &lt;a href="https://dev.to/samhartley_dev/how-i-get-frontier-quality-output-from-local-models-that-are-10x-smaller-17oc"&gt;getting real work out of small local models&lt;/a&gt;, and &lt;a href="https://dev.to/samhartley_dev/i-dont-vibe-check-new-local-models-anymore-i-run-this-30-minute-eval-instead-44dk"&gt;evaluating local models before trusting them&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>devops</category>
      <category>automation</category>
    </item>
    <item>
      <title>My Cron Job Double-Fired and Placed the Same Order Twice — Here's My Postmortem</title>
      <dc:creator>Sam Hartley</dc:creator>
      <pubDate>Sun, 13 Sep 2026 08:20:58 +0000</pubDate>
      <link>https://dev.to/samhartley_dev/my-cron-job-double-fired-and-placed-the-same-order-twice-heres-my-postmortem-5cib</link>
      <guid>https://dev.to/samhartley_dev/my-cron-job-double-fired-and-placed-the-same-order-twice-heres-my-postmortem-5cib</guid>
      <description>&lt;p&gt;I keep a Mac Mini in a closet running a handful of small scheduled Python jobs. One of them scans market data on a schedule and, when its rules fire, places tiny orders through an exchange API. Positions are around $12 each. It ran for months without drama.&lt;/p&gt;

&lt;p&gt;Then one morning the exchange said I held 16 contracts, and my local state file said 8.&lt;/p&gt;

&lt;p&gt;Not a rounding error. Not a stale cache. The same order, placed twice, ten seconds apart — by two instances of the same script that were both alive at the same time, each keeping its own private idea of reality.&lt;/p&gt;

&lt;p&gt;What follows is the postmortem I wrote for myself, cleaned up. The punchline: the scheduler was fine. The API was fine. Both root causes were me.&lt;/p&gt;

&lt;h2&gt;
  
  
  The morning I noticed two of everything
&lt;/h2&gt;

&lt;p&gt;Context first, because the boring design matters later:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A Python script runs on a schedule (and occasionally on demand)&lt;/li&gt;
&lt;li&gt;It fetches about a thousand tickers, applies some rules, sometimes opens or closes a small position&lt;/li&gt;
&lt;li&gt;Every decision lands in a JSON state file; all my reports read from that file&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That morning the fee log showed two order fees for the same position, timestamps ten seconds apart. The exchange held 16 contracts. My state file tracked 8. The state file wasn't lying about what it knew — it had only ever been told about one of the two orders.&lt;/p&gt;

&lt;p&gt;I eventually found two distinct ways this had happened. Both were my fault, in different ways.&lt;/p&gt;

&lt;h2&gt;
  
  
  Incident 1: I restarted a job that was never dead
&lt;/h2&gt;

&lt;p&gt;The first one is embarrassing in the most ordinary way possible.&lt;/p&gt;

&lt;p&gt;The scan isn't instant — a thousand tickers plus math takes a bit. One evening I triggered a manual run, and the terminal just sat there. No output. No prompt back. "Hung," I thought — and ran it again.&lt;/p&gt;

&lt;p&gt;Here's what actually happens when a chatty Python script runs with stdout connected to a pipe or a background handle instead of a real terminal: output is block-buffered. The job was alive and working the whole time. Its output was sitting in a buffer, waiting to be flushed. I read the silence as "dead," and my restart created a second live instance.&lt;/p&gt;

&lt;p&gt;Those two order fees, ten seconds apart? One order from the run I thought had died, one from the resurrection.&lt;/p&gt;

&lt;p&gt;The lesson that stings: &lt;strong&gt;silence is not a signal.&lt;/strong&gt; "It's not printing" and "it's not running" are unrelated statements. Before you rerun anything: check the process table, check the log file's modification time, check whether it's burning CPU. Poll the thing. Re-running a silent-but-alive job is how you manufacture concurrency you never designed for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Incident 2: The scheduler and I raced each other
&lt;/h2&gt;

&lt;p&gt;The second one took an evening of forensics.&lt;/p&gt;

&lt;p&gt;A few days later: same symptom, but one position was triplicated. The order timestamps lined up with the top of the hour — and my scheduler fires at :02. In my shell history, right there: a manual run of the same script at :02. I'd been poking around and wanted a fresh scan "right now," and it simply never occurred to me that the clock was about to do the same thing.&lt;/p&gt;

&lt;p&gt;Two entry points, same script, same minute, zero coordination. A scheduler is just a clock — it doesn't know what you're doing in a terminal, and I wasn't thinking about the clock.&lt;/p&gt;

&lt;p&gt;The rule I added: &lt;strong&gt;manual runs never land on a scheduled minute.&lt;/strong&gt; If the job runs at :02, I run manual checks at :15. It feels comically simple. It has prevented at least two repeats that I know of.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this wasn't a disaster (this time)
&lt;/h2&gt;

&lt;p&gt;I want to be honest here, because the honest version is more useful than the heroic one. Three things saved me, and none of them was foresight:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;My exits are close-all.&lt;/strong&gt; The exit logic closes the entire position on the exchange — not "the size I think I have." A duplicated position resolves itself at the next exit. The bug was self-healing by accident.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The money was small.&lt;/strong&gt; $12 doubling to $24 is a shrug. The same code with position sizes that scale with an account would not be a shrug.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The exchange is the source of truth.&lt;/strong&gt; My reports were wrong and my history file had gaps, but the actual money was always consistent with reality. The drift lived in my shadow copy — the state file — not in the world.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That third point rewired how I think about every state file I keep. They're caches. The API on the other side is the truth. Once that lands, "reconcile before acting" stops being an enterprise buzzphrase and becomes just reading the real number before making a decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fixes
&lt;/h2&gt;

&lt;h3&gt;
  
  
  A lockfile — the boring kind
&lt;/h3&gt;

&lt;p&gt;The core fix is about fifteen lines at the top of &lt;code&gt;main()&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;fcntl&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sys&lt;/span&gt;

&lt;span class="n"&gt;LOCK&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dirname&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;abspath&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;__file__&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;.job.lock&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;lock_fd&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;LOCK&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;w&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;fcntl&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;flock&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lock_fd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;fcntl&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;LOCK_EX&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="n"&gt;fcntl&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;LOCK_NB&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;OSError&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;[lock] another instance holds &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;LOCK&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;; exiting&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;flush&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;lock_fd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getpid&lt;/span&gt;&lt;span class="p"&gt;()))&lt;/span&gt;
&lt;span class="n"&gt;lock_fd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;flush&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A second instance starts, tries to grab the lock, fails, exits with a message. On Linux and macOS that's &lt;code&gt;fcntl&lt;/code&gt;; on Windows it's &lt;code&gt;msvcrt.locking&lt;/code&gt; — same idea, different syscall. And if you'd rather not touch Python for it, the shell version is one line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;flock &lt;span class="nt"&gt;-n&lt;/span&gt; /tmp/job.lock python3 job.py &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"already running"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Announce yourself in the logs
&lt;/h3&gt;

&lt;p&gt;Every run now logs start and end with its PID:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;logging&lt;/span&gt;

&lt;span class="n"&gt;logging&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;basicConfig&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;filename&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;job.log&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nb"&gt;format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;%(asctime)s pid=%(process)d %(levelname)s %(message)s&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;logging&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;scan start&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When the lockfile blocks a second instance, the log says so in plain words. And looking back at the incident logs now, the duplicate jumps out in ten seconds — two different &lt;code&gt;pid=&lt;/code&gt; numbers in the same minute — instead of needing an evening of timestamp archaeology.&lt;/p&gt;

&lt;h3&gt;
  
  
  Flush like you mean it
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;flush=True&lt;/code&gt; on prints — or just run with &lt;code&gt;python3 -u&lt;/code&gt;. This one directly kills Incident 1: the terminal would have shown live output, and "no output" would have actually meant something.&lt;/p&gt;

&lt;h3&gt;
  
  
  Reconcile before acting
&lt;/h3&gt;

&lt;p&gt;Before the job opens anything, it now pulls actual positions from the exchange API and treats the local state file as a hint, not as truth. Anything the file claims that the API doesn't confirm gets flagged for review instead of acted on.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part I'd been getting wrong for years
&lt;/h2&gt;

&lt;p&gt;The reframe that made all of this click:&lt;/p&gt;

&lt;p&gt;I used to think of idempotency as an API-design topic — something server people worry about. But the moment your cron job calls an API that mutates state, you're running a distributed system. Multiple processes, lying transports (buffered stdout counts as a lying transport — it lied straight to my eyes), shared mutable state on the other side of a network. All the classic distributed-systems failure modes are now your failure modes. Just at hobby scale.&lt;/p&gt;

&lt;p&gt;A lockfile doesn't make a distributed system safe. What it does is collapse the most common failure — scheduler plus human, or human plus impatient human — into a no-op. That's most of the practical risk, for fifteen lines of code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrapping up
&lt;/h2&gt;

&lt;p&gt;If you take away two things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A silent job is not a dead job. Poll, don't rerun.&lt;/li&gt;
&lt;li&gt;If a scheduled job of yours talks to anything that mutates state, add the lockfile today — not after your own double-fire. Mine cost an evening of forensics and $12 of unplanned exposure. The same bug with real sizing would have cost a weekend and my confidence in the whole setup.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I'm curious how others handle this, especially anyone running scheduled jobs against mutating APIs on a homelab box. Lockfiles, per-action idempotency keys, a real job queue instead of cron? Drop what's actually held up for you in the comments — I'm always stealing ideas for my checklist.&lt;/p&gt;

</description>
      <category>python</category>
      <category>devops</category>
      <category>automation</category>
      <category>postmortem</category>
    </item>
    <item>
      <title>I Don't Vibe-Check New Local Models Anymore — I Run This 30-Minute Eval Instead</title>
      <dc:creator>Sam Hartley</dc:creator>
      <pubDate>Fri, 11 Sep 2026 08:03:28 +0000</pubDate>
      <link>https://dev.to/samhartley_dev/i-dont-vibe-check-new-local-models-anymore-i-run-this-30-minute-eval-instead-44dk</link>
      <guid>https://dev.to/samhartley_dev/i-dont-vibe-check-new-local-models-anymore-i-run-this-30-minute-eval-instead-44dk</guid>
      <description>&lt;h1&gt;
  
  
  I Don't Vibe-Check New Local Models Anymore — I Run This 30-Minute Eval Instead
&lt;/h1&gt;

&lt;p&gt;A new "best open model ever" drops roughly every week now. And for most of this year, I handled every release the same dumb way: download it, ask it three questions in a terminal, nod at the answers, think "yeah, this one's smart" — then spend the next two weeks finding out where it actually breaks, in my real work, one task at a time.&lt;/p&gt;

&lt;p&gt;That's a vibe-check. It feels like evidence. It isn't. A model that nails a clever one-liner and then mangles JSON in my log pipeline is useless to me, and no amount of terminal flirting reveals that until something breaks at 2 AM.&lt;/p&gt;

&lt;p&gt;For context: my daily drivers are a 9B model on my Mac Mini and a 30B coder model on the RTX 3060 in my PC. Swapping either of them isn't free — every model has quirks baked into my prompts by now. So I stopped vibe-checking. Now every model goes through a small eval suite before it's allowed anywhere near my stack. Thirty minutes to run, built in an afternoon, and it has saved me from three bad migrations this year.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the public benchmarks don't help me
&lt;/h2&gt;

&lt;p&gt;I read benchmark posts. I just don't make decisions with them, because they measure the average internet's tasks, not mine.&lt;/p&gt;

&lt;p&gt;My actual workload on a normal day:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Extracting structured JSON out of messy log lines and API responses&lt;/li&gt;
&lt;li&gt;Editing Monkey C for my watch face projects (a language most models barely know)&lt;/li&gt;
&lt;li&gt;Writing small glue Python scripts against local APIs&lt;/li&gt;
&lt;li&gt;Summarizing long docs where specific facts have to survive, not just the vibe&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A model can top every leaderboard and still be the wrong model for me if it can't emit valid JSON without a chaperone. The reverse is true too — the "boring" model that always follows the output format is worth more to my setup than the genius that occasionally freestyles.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's in the suite
&lt;/h2&gt;

&lt;p&gt;One folder, 40 cases, all lifted from real work:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;12 extraction tasks&lt;/strong&gt; — real (scrubbed) log lines; the model must return valid JSON with specific keys&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;10 code edit tasks&lt;/strong&gt; — "here's a function, add retry with backoff, keep it under 30 lines" style&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;8 summarization tasks&lt;/strong&gt; — checked for must-mention facts, not tone&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;6 format-adherence tasks&lt;/strong&gt; — strict output shapes, no prose allowed&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;4 trap tasks&lt;/strong&gt; — ones that look easy but contain a subtle gotcha. Models that bluff instead of flagging the problem fail here, which is exactly what I want to catch&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The rule for growing it: every time a model messes something up in daily use and I have to fix it by hand, that failure becomes a new case. The suite grows from my own scars. It's the most valuable folder on my disk and it's maybe 60 KB of text.&lt;/p&gt;

&lt;h2&gt;
  
  
  The harness
&lt;/h2&gt;

&lt;p&gt;60-ish lines of Python. No framework, no LLM-as-judge (a judge model just adds its own preferences and noise). Each case runs three times at temperature 0, and the score is the average pass rate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;

&lt;span class="n"&gt;OLLAMA&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;http://localhost:11434/api/chat&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;run_case&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;case&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;OLLAMA&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;messages&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;case&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prompt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]}],&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;stream&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;options&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;temperature&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;message&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;check&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;case&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;kind&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;case&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;checker&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;kind&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;all&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;case&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;expect_keys&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
        &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;kind&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;regex&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;case&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pattern&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;S&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;kind&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;contains&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;all&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;case&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;must_contain&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cases&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;runs&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;case&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;cases&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;passed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;check&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;case&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;run_case&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;case&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;runs&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
        &lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;passed&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;runs&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cases&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;cases&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;l&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;l&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cases.jsonl&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;argv&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;argv&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;cases&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A few decisions baked in:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Temperature 0, always.&lt;/strong&gt; Evaluating at temperature 0.7 is judging a model on moods.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Three runs per case&lt;/strong&gt; catches flaky formatting. One model I tested passed a JSON case on run 1 and failed runs 2 and 3. That's not a model I trust with nightly jobs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One number out.&lt;/strong&gt; If a model scores 87.5, it goes into a plain text file next to the previous scores. That file is now my real benchmark history.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What running it taught me
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Newer isn't automatically better.&lt;/strong&gt; The suite caught a hyped release that scored 9 points below my incumbent, even though everyone on my timeline was calling it the new default. On my tasks, it wasn't.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quantization damage shows up in format, not intelligence.&lt;/strong&gt; Q4 quants pass the logic checks fine, then blow output formatting at weird, inconsistent rates. If I'd only vibe-checked, I'd never have seen it — the prose looks equally smart either way.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Wins are task-shaped, not model-shaped.&lt;/strong&gt; One challenger beat my incumbent by 15 points on extraction and lost by 8 on code edits. I didn't switch — I routed. That result is literally why extraction jobs in my setup now go to a different model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My prompts drift too.&lt;/strong&gt; I run the suite after big prompt changes, not just model changes. Twice this year a "clever prompt improvement" quietly regressed three cases. A vibe-check would never have caught that.&lt;/p&gt;

&lt;h2&gt;
  
  
  The switch rule
&lt;/h2&gt;

&lt;p&gt;A challenger replaces my current model only if it wins by &lt;strong&gt;10+ points overall&lt;/strong&gt; AND doesn't drop more than 5 points in any single category. Below that, the migration isn't worth it — rewriting prompts, re-learning quirks, updating the docs I keep for myself.&lt;/p&gt;

&lt;p&gt;The one time I ignored my own rule was that hyped release I mentioned. The timeline got me. I spent a weekend migrating prompts, ran the suite afterward out of curiosity, and the old model won by 9. I crawled back and re-deployed the old setup Sunday night. The suite forgave me. My Saturday didn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrapping up
&lt;/h2&gt;

&lt;p&gt;If you run local models for anything that matters, don't vibe-check. Build a suite of YOUR tasks — 20 cases is enough to start — run new models against it at temperature 0, and let a number make the decision. It's 30 minutes of compute versus two weekends of regret.&lt;/p&gt;

&lt;p&gt;If you've built something similar, I'm curious what made it into your cases. Drop your weirdest eval case in the comments — always looking to steal ideas for mine.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>ollama</category>
      <category>python</category>
    </item>
    <item>
      <title>How I Get Frontier-Quality Output from Local Models That Are 10x Smaller</title>
      <dc:creator>Sam Hartley</dc:creator>
      <pubDate>Thu, 10 Sep 2026 09:13:31 +0000</pubDate>
      <link>https://dev.to/samhartley_dev/how-i-get-frontier-quality-output-from-local-models-that-are-10x-smaller-17oc</link>
      <guid>https://dev.to/samhartley_dev/how-i-get-frontier-quality-output-from-local-models-that-are-10x-smaller-17oc</guid>
      <description>&lt;h1&gt;
  
  
  How I Get Frontier-Quality Output from Local Models That Are 10x Smaller
&lt;/h1&gt;

&lt;p&gt;I run Qwen 3.5 9B on my Mac Mini and Qwen 3 Coder 30B on an RTX 3060. That's the hardware reality — I'm not fitting a 200B+ parameter model locally. Not even close.&lt;/p&gt;

&lt;p&gt;But here's the thing: most of my daily output from these models is indistinguishable from what I used to get from Claude or GPT-4o. Not because the models are that good. Because I learned how to talk to them.&lt;/p&gt;

&lt;p&gt;This isn't another "here are 10 generic prompt tips" post. This is the specific, weird stuff I do to squeeze frontier-quality output from models that are 10-20x smaller than the cloud alternatives. Stuff that actually changed my day-to-day results.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Prompt Engineering Matters More for Small Models
&lt;/h2&gt;

&lt;p&gt;Big models are forgiving. You can throw a vague prompt at GPT-4o and still get something usable. The model has enough parameters to infer what you probably meant.&lt;/p&gt;

&lt;p&gt;Small models don't have that luxury. A 9B model takes your prompt literally. If you're vague, the output is vague. If you're specific, the output is specific. The quality gap between a badly-prompted small model and a well-prompted small model is massive — far bigger than the gap between a well-prompted small model and a well-prompted big model.&lt;/p&gt;

&lt;p&gt;In other words: the same 9B model can produce garbage or gold depending entirely on how you ask.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trick 1: The "Role + Format + Constraint" System Prompt
&lt;/h2&gt;

&lt;p&gt;Most system prompts I see are one-liners: "You are a helpful assistant." That's wasted space on a small model.&lt;/p&gt;

&lt;p&gt;My system prompts always have three sections:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are a senior Python developer who writes clean, minimal code.

FORMAT: Respond with code only. No explanations unless asked. Use type hints.

CONSTRAINTS:
- Never use global variables
- Always include error handling for network calls
- Prefer standard library over third-party packages
- Maximum 50 lines per function
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The role tells the model &lt;em&gt;who to be&lt;/em&gt;. The format tells it &lt;em&gt;what shape the answer should take&lt;/em&gt;. The constraints tell it &lt;em&gt;what not to do&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Small models are terrible at deciding what to leave out. The constraints section forces the model to make fewer choices, which means fewer wrong choices.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Before this pattern:&lt;/strong&gt; My 30B coder model would return 200-line functions with inline comments explaining every line. Useful, but I spent more time deleting than coding.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;After:&lt;/strong&gt; 50-line functions, type hints, no fluff. I went from 40% of the output being noise to maybe 5%.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trick 2: One-Shot Examples Over Zero-Shot
&lt;/h2&gt;

&lt;p&gt;Small models are pattern matchers. Give them a pattern and they follow it. Don't give them a pattern and they guess — badly.&lt;/p&gt;

&lt;p&gt;Instead of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Summarize this article in 3 bullet points.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I do:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Summarize articles as 3 bullet points.

ARTICLE: "Apple announced..."
SUMMARY:
- Apple released new MacBook Pro with M4 chip
- Starting at $1,599, shipping next week
- Performance claims 2x faster than previous generation

ARTICLE: "The Federal Reserve..."
SUMMARY:
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That example does more work than any instruction. The model sees the exact format, tone, and level of detail I want. It mirrors that pattern for the next article.&lt;/p&gt;

&lt;p&gt;This is the single highest-impact change I made. One-shot examples turned my 9B model from "sometimes okay" to "reliably good" for structured tasks like summarization, extraction, and formatting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trick 3: Decompose, Don't Compose
&lt;/h2&gt;

&lt;p&gt;I used to ask small models for big things:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Build me a REST API with authentication, CRUD operations, rate limiting, and tests.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And I'd get 500 lines of interconnected code that looked reasonable but had subtle bugs in the auth middleware, the rate limiter didn't actually work, and the tests were testing the wrong things.&lt;/p&gt;

&lt;p&gt;Now I break it into steps, running each one separately:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Step 1: "Create a FastAPI app skeleton with health check endpoint"
→ I review, test, fix

Step 2: "Add JWT authentication to this app. Here's the current code: [paste]"
→ I review, test, fix

Step 3: "Add CRUD endpoints for users with SQLAlchemy. Current app: [paste]"
→ I review, test, fix
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each step produces 20-40 lines. Small models are excellent at 20-40 lines. They're mediocre at 200 lines. They're terrible at 500 lines.&lt;/p&gt;

&lt;p&gt;This is more work for me — I'm running 4 prompts instead of 1. But the output quality is dramatically better, and I spend less time debugging because each piece is correct before I build on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trick 4: Temperature Isn't What You Think
&lt;/h2&gt;

&lt;p&gt;I used to set temperature to 0.0 for "factual" tasks and 0.7 for "creative" tasks. That's what every tutorial says.&lt;/p&gt;

&lt;p&gt;For small models, I've found the opposite works better:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Code and structured output: 0.0&lt;/strong&gt; — no surprise, this is standard&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Summarization and extraction: 0.0&lt;/strong&gt; — small models hallucinate more at higher temperatures. Keep it deterministic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Creative writing: 0.3-0.4&lt;/strong&gt; — not 0.7. At 0.7, a 9B model produces wild tangents. At 0.3, it's slightly varied but stays on track.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Brainstorming: 0.5-0.6&lt;/strong&gt; — this is the only time I go above 0.5 on a small model.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The key insight: small models have less "knowledge density" per parameter. Higher temperatures let them drift further from their training distribution, which means more hallucination. Keep the temperature low and guide the creativity through your prompt instead.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trick 5: The "Bad Example" Anti-Pattern
&lt;/h2&gt;

&lt;p&gt;Small models learn from examples, and they learn from &lt;em&gt;what not to do&lt;/em&gt; just as much as what to do.&lt;/p&gt;

&lt;p&gt;I include explicit bad examples in my prompts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Write a technical explanation. Do NOT write like this:

"Machine learning is a fascinating field that has revolutionized countless industries.
In today's rapidly evolving technological landscape, leveraging artificial intelligence
has become paramount for organizations seeking to maintain a competitive edge."

Instead, write like this:

"ML models map inputs to outputs. You train them on data, then they predict on new data.
The three types: supervised (labeled data), unsupervised (no labels), reinforcement (reward signals)."

Your turn. Explain distributed systems:
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The bad example shows exactly the kind of verbose, corporate, filler-heavy output I don't want. The good example shows the concise, specific, jargon-free style I do want.&lt;/p&gt;

&lt;p&gt;This is especially effective for small models because it gives them two reference points instead of one. They can triangulate the style more accurately.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trick 6: Post-Processing Pipeline (The Unsexy Secret)
&lt;/h2&gt;

&lt;p&gt;Here's the thing nobody wants to hear: my local model output goes through a second pass.&lt;/p&gt;

&lt;p&gt;Not another LLM call. A simple Python script that:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Removes filler phrases&lt;/strong&gt; — "It's important to note that", "In today's world", "As we can see"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enforces sentence length limits&lt;/strong&gt; — any sentence over 30 words gets flagged for manual review&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Checks for repetition&lt;/strong&gt; — if two paragraphs say essentially the same thing, it removes one&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Validates code blocks&lt;/strong&gt; — runs &lt;code&gt;python -c "compile(code, ...)"&lt;/code&gt; to check for syntax errors&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This script catches maybe 15% of the model's output. But that 15% is the difference between "obviously AI-generated" and "reads like I wrote it."&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;FILLER_PHRASES&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;it&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s important to note&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;in today&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s world&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;as we can see&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;it&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s worth mentioning&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;needless to say&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;at the end of the day&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;in this rapidly evolving&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;leverage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="c1"&gt;# as a verb
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;delve into&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tapestry&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;clean_output&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;phrase&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;FILLER_PHRASES&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;phrase&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;capitalize&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;phrase&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Simple. Dumb. Effective. The model doesn't know I'm doing this, and it doesn't need to.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trick 7: Context Recycling for Multi-Turn Work
&lt;/h2&gt;

&lt;p&gt;Small models have small context windows. My 9B model handles about 8K tokens of context before it starts losing the thread.&lt;/p&gt;

&lt;p&gt;When I'm working on a multi-turn coding session, I don't keep the full conversation history. After every 3-4 exchanges, I summarize the key decisions into a "state" block and start fresh:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;STATE&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;Project&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Task&lt;/span&gt; &lt;span class="n"&gt;queue&lt;/span&gt; &lt;span class="n"&gt;manager&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;Python&lt;/span&gt;
&lt;span class="n"&gt;Stack&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;FastAPI&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;Redis&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;SQLite&lt;/span&gt;
&lt;span class="n"&gt;Decisions&lt;/span&gt; &lt;span class="n"&gt;so&lt;/span&gt; &lt;span class="n"&gt;far&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;Using&lt;/span&gt; &lt;span class="n"&gt;Redis&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;job&lt;/span&gt; &lt;span class="n"&gt;queue&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;SQLite&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;persistence&lt;/span&gt;
&lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;Async&lt;/span&gt; &lt;span class="n"&gt;workers&lt;/span&gt; &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;asyncio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;gather&lt;/span&gt;
&lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;Auth&lt;/span&gt; &lt;span class="n"&gt;via&lt;/span&gt; &lt;span class="n"&gt;JWT&lt;/span&gt; &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;minute&lt;/span&gt; &lt;span class="n"&gt;expiry&lt;/span&gt;
&lt;span class="n"&gt;Current&lt;/span&gt; &lt;span class="nb"&gt;file&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;worker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;py &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;progress&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;Known&lt;/span&gt; &lt;span class="n"&gt;issues&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Need&lt;/span&gt; &lt;span class="n"&gt;to&lt;/span&gt; &lt;span class="n"&gt;handle&lt;/span&gt; &lt;span class="n"&gt;worker&lt;/span&gt; &lt;span class="n"&gt;crashes&lt;/span&gt; &lt;span class="n"&gt;mid&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt;

&lt;span class="n"&gt;Continue&lt;/span&gt; &lt;span class="n"&gt;implementing&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;worker&lt;/span&gt; &lt;span class="n"&gt;crash&lt;/span&gt; &lt;span class="n"&gt;recovery&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This keeps the context lean and focused. The model gets exactly what it needs — no conversation history padding, no stale context from turns 1-3 that's no longer relevant.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Without context recycling:&lt;/strong&gt; After 5-6 turns, the model starts contradicting decisions from turn 2. It forgets what it decided and re-decides differently.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;With context recycling:&lt;/strong&gt; I can work for 20+ turns and the model stays consistent because it only sees the current state, not the full history.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where These Tricks DON'T Help
&lt;/h2&gt;

&lt;p&gt;Let me be honest about the gaps:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Complex reasoning.&lt;/strong&gt; A 9B model can't trace a 6-step logical argument no matter how you prompt it. I still route complex reasoning to the 30B model or to cloud APIs. Prompt engineering closes the gap for structured, well-defined tasks. It does not create intelligence that isn't there.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Long-context understanding.&lt;/strong&gt; If I need to analyze a 20-page document, no amount of prompting will make a 9B model with 8K context do it well. I chunk the document, summarize each chunk, then feed the summaries to the model. It works, but it's a workaround, not a solution.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Factual accuracy about recent events.&lt;/strong&gt; Local models have training cutoffs. My Qwen 3.5 doesn't know what happened in 2026. No prompt trick fixes stale training data. For current information, I use cloud APIs or RAG with real-time data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Nuanced code architecture.&lt;/strong&gt; The 30B coder model can write a single module well. It cannot design a multi-service system where auth, data, and queue layers interact correctly. I still design the architecture myself and have the model implement each piece.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Numbers (Before and After Prompt Engineering)
&lt;/h2&gt;

&lt;p&gt;I tracked output quality for a month before and after I started using these tricks consistently:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Before (raw prompts)&lt;/th&gt;
&lt;th&gt;After (engineered prompts)&lt;/th&gt;
&lt;th&gt;Cloud (GPT-4o)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Code generation (&amp;lt; 50 lines)&lt;/td&gt;
&lt;td&gt;6/10&lt;/td&gt;
&lt;td&gt;9/10&lt;/td&gt;
&lt;td&gt;9.5/10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Code generation (&amp;gt; 200 lines)&lt;/td&gt;
&lt;td&gt;3/10&lt;/td&gt;
&lt;td&gt;5/10&lt;/td&gt;
&lt;td&gt;8/10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Summarization&lt;/td&gt;
&lt;td&gt;5/10&lt;/td&gt;
&lt;td&gt;9/10&lt;/td&gt;
&lt;td&gt;9/10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data extraction&lt;/td&gt;
&lt;td&gt;6/10&lt;/td&gt;
&lt;td&gt;8.5/10&lt;/td&gt;
&lt;td&gt;9/10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Creative writing&lt;/td&gt;
&lt;td&gt;4/10&lt;/td&gt;
&lt;td&gt;7/10&lt;/td&gt;
&lt;td&gt;8.5/10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Debugging&lt;/td&gt;
&lt;td&gt;5/10&lt;/td&gt;
&lt;td&gt;7/10&lt;/td&gt;
&lt;td&gt;9/10&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For tasks under ~50 lines of code, the gap between well-prompted local and cloud is almost gone. For longer tasks, local models still lag, but prompt engineering narrows it from "embarrassing" to "acceptable."&lt;/p&gt;

&lt;h2&gt;
  
  
  The Quick Reference
&lt;/h2&gt;

&lt;p&gt;If you take one thing from this article, make it this checklist:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;System prompt = Role + Format + Constraints&lt;/strong&gt; — never just "you are helpful"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One-shot examples&lt;/strong&gt; — always show the model exactly what you want&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decompose big tasks&lt;/strong&gt; — 4 small prompts beat 1 big prompt&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Low temperature&lt;/strong&gt; — 0.0 for code, 0.3-0.4 for creative, never above 0.6 on small models&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bad examples&lt;/strong&gt; — show what NOT to write&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Post-processing&lt;/strong&gt; — a 20-line Python script catches 15% of AI-isms&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context recycling&lt;/strong&gt; — summarize state, don't keep full history&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;These aren't theoretical. I use every single one daily on my Mac Mini and RTX 3060 setup. They took my local models from "good enough for rough drafts" to "good enough for production output."&lt;/p&gt;

&lt;p&gt;The model didn't change. My prompts did.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>promptengineering</category>
      <category>selfhosted</category>
      <category>ollama</category>
    </item>
    <item>
      <title>Why I Stopped Chasing the Newest LLM (And What I Run Instead)</title>
      <dc:creator>Sam Hartley</dc:creator>
      <pubDate>Sat, 29 Aug 2026 08:02:10 +0000</pubDate>
      <link>https://dev.to/samhartley_dev/why-i-stopped-chasing-the-newest-llm-and-what-i-run-instead-51h9</link>
      <guid>https://dev.to/samhartley_dev/why-i-stopped-chasing-the-newest-llm-and-what-i-run-instead-51h9</guid>
      <description>&lt;h1&gt;
  
  
  Why I Stopped Chasing the Newest LLM (And What I Run Instead)
&lt;/h1&gt;

&lt;p&gt;There was a stretch last year where I downloaded every new model within hours of it dropping. Llama 3.1? Done. Mistral Nemo? Running. Qwen 2.5? Let me benchmark it at 2 AM on a Tuesday.&lt;/p&gt;

&lt;p&gt;I was spending more time downloading, configuring, and benchmarking than actually building things. And the worst part? Most of the "upgrades" didn't change my daily output at all.&lt;/p&gt;

&lt;p&gt;So I stopped. I picked a stack and stuck with it. And somehow, I'm getting more done now than when I had 40 models installed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Upgrade Treadmill
&lt;/h2&gt;

&lt;p&gt;Here's what the cycle looked like:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;New model drops on HuggingFace&lt;/li&gt;
&lt;li&gt;Someone posts benchmarks showing it's 3% better at X&lt;/li&gt;
&lt;li&gt;I spend an evening downloading 20GB, tweaking parameters, running my test suite&lt;/li&gt;
&lt;li&gt;It's marginally better at something and marginally worse at something else&lt;/li&gt;
&lt;li&gt;Two days later, another model drops&lt;/li&gt;
&lt;li&gt;Repeat&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;In one month I downloaded 12 models. I actually used 3 of them regularly. The other 9 sat on my drive eating 180GB of SSD space.&lt;/p&gt;

&lt;p&gt;The math was brutal: roughly 15 hours of my time that month on model management. That's almost two full workdays. For what? A 3% improvement on a benchmark I don't even care about?&lt;/p&gt;

&lt;h2&gt;
  
  
  My Stable Stack (Same Models for 4+ Months)
&lt;/h2&gt;

&lt;p&gt;I haven't changed my primary models since April 2026. Here's what I run and why:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Machine&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;th&gt;Why This One&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Mac Mini M4&lt;/td&gt;
&lt;td&gt;Qwen 3.5 9B&lt;/td&gt;
&lt;td&gt;Orchestration, routing, quick chat&lt;/td&gt;
&lt;td&gt;Fast, solid at instruction-following, 9B is the sweet spot for 16GB RAM&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mac Mini M4&lt;/td&gt;
&lt;td&gt;Granite 3.2 8B&lt;/td&gt;
&lt;td&gt;Writing, editing, summaries&lt;/td&gt;
&lt;td&gt;Better prose than Qwen at this size. Less hallucination in creative tasks.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Windows PC (RTX 3060)&lt;/td&gt;
&lt;td&gt;Qwen 3 Coder 30B&lt;/td&gt;
&lt;td&gt;Code generation, refactoring&lt;/td&gt;
&lt;td&gt;30B is the smallest that reliably handles multi-file projects. Fits in 12GB VRAM.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Windows PC (RTX 3060)&lt;/td&gt;
&lt;td&gt;DeepSeek R1 8B&lt;/td&gt;
&lt;td&gt;Reasoning, debugging, analysis&lt;/td&gt;
&lt;td&gt;The chain-of-thought on this model is genuinely useful for tracing bugs.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ubuntu box&lt;/td&gt;
&lt;td&gt;DeepSeek R1 (larger)&lt;/td&gt;
&lt;td&gt;Heavy reasoning tasks&lt;/td&gt;
&lt;td&gt;When the 8B isn't enough and I don't want to pay for cloud.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ubuntu box&lt;/td&gt;
&lt;td&gt;MinCPM-V&lt;/td&gt;
&lt;td&gt;Vision tasks&lt;/td&gt;
&lt;td&gt;Only vision model that runs reasonably on consumer hardware.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That's it. Six models across three machines. I use maybe 4 of them daily.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why These Specific Models
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Qwen 3.5 9B — The Workhorse
&lt;/h3&gt;

&lt;p&gt;This model handles 80% of what I throw at it. Routing decisions, summarizing articles, writing first drafts, parsing structured data. It's fast (under 2 seconds for most responses on the Mac Mini) and it doesn't hallucinate much at this size.&lt;/p&gt;

&lt;p&gt;I tried replacing it with Llama 4 Scout when it came out. Same parameter class, theoretically better benchmarks. In practice? The Qwen model followed my system prompts more precisely and produced more consistent output formats. Benchmarks don't measure "does it actually follow instructions" — real usage does.&lt;/p&gt;

&lt;h3&gt;
  
  
  Qwen 3 Coder 30B — The Specialist
&lt;/h3&gt;

&lt;p&gt;30B parameters is the minimum I've found for reliable multi-file code generation. Anything smaller starts forgetting what function it defined two files ago.&lt;/p&gt;

&lt;p&gt;The RTX 3060 with 12GB VRAM runs this model in 4-bit quantization at about 12-15 tokens/second. Not blazing, but fast enough for code review and refactoring where I'm not sitting there waiting.&lt;/p&gt;

&lt;p&gt;Could I get better results from a 70B model? Probably. But a 70B model doesn't fit on a 3060, and I'm not buying another GPU just for marginally better code suggestions.&lt;/p&gt;

&lt;h3&gt;
  
  
  DeepSeek R1 8B — The Thinker
&lt;/h3&gt;

&lt;p&gt;This one surprised me. The chain-of-thought reasoning in R1 is genuinely useful, even at 8B. When I have a tricky bug or a design decision, I run it through R1 first. It often surfaces edge cases I missed.&lt;/p&gt;

&lt;p&gt;The 8B version is the minimum viable reasoner. For really complex stuff, I bump up to the larger version on the Ubuntu box. But 8B handles 90% of my reasoning needs.&lt;/p&gt;

&lt;h3&gt;
  
  
  Granite 3.2 8B — The Writer
&lt;/h3&gt;

&lt;p&gt;Most people don't know this model well. IBM's Granite series doesn't get the hype of Llama or Qwen. But for writing tasks — summarizing, editing, rephrasing — it's noticeably better than Qwen at the same size.&lt;/p&gt;

&lt;p&gt;It produces cleaner prose with fewer awkward constructions. It also hallucinates less on factual summaries. I use it for all article drafts and content work.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'm Missing By Not Upgrading
&lt;/h2&gt;

&lt;p&gt;Let me be honest about the trade-offs:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I'm probably 5-10% behind the frontier on code generation.&lt;/strong&gt; The latest Claude or GPT-4o will write better code than my 30B Qwen. But I'm not paying $200/month for 5-10% better code that I still have to review anyway.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I'm definitely behind on very long context tasks.&lt;/strong&gt; If I need to analyze a 50-page document, local models with 8-32K context windows can't compete with Gemini's 1M+ window. I still use cloud for those specific tasks — about 15% of my workload.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I'm missing out on multimodal advances.&lt;/strong&gt; The vision models I run are functional but not impressive. GPT-4o Vision and Gemini Pro Vision are significantly better. I use cloud for complex vision tasks.&lt;/p&gt;

&lt;p&gt;But here's the thing: I know these gaps. They're predictable. And I've built workflows that route around them without thinking about it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Routing System That Makes It Work
&lt;/h2&gt;

&lt;p&gt;The reason I can run a "stale" stack is that I built a routing layer that automatically sends tasks to the right model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Incoming task
  → Quick chat / formatting / extraction?
    → Qwen 3.5 9B (local, instant)
  → Code generation or refactoring?
    → Qwen 3 Coder 30B (local, ~12s)
  → Reasoning or debugging?
    → DeepSeek R1 8B (local, ~8s)
  → Long document or complex vision?
    → Cloud API (routed, paid)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This routing is 15 lines of Python. It checks the task type and sends it to the appropriate model. No AI needed to route AI. A simple classifier works fine.&lt;/p&gt;

&lt;p&gt;The result: I use cloud APIs for maybe 15% of my tasks, and those are the ones where local genuinely can't compete. Everything else stays local, stays free, and stays fast.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Money Math (After 14 Months)
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;RTX 3060 12GB (used, eBay)&lt;/td&gt;
&lt;td&gt;$150 one-time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mac Mini M4 (already owned)&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ubuntu box (repurposed)&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Electricity (all 3 machines, 24/7)&lt;/td&gt;
&lt;td&gt;~$12/month&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cloud API usage (hybrid)&lt;/td&gt;
&lt;td&gt;~$25/month&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total monthly cost&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~$37/month&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;My previous cloud-only spend was $187/month. That's $1,800/year saved. Minus the $150 GPU (paid for in under a month) and $37/month ongoing, I'm saving over $1,400/year.&lt;/p&gt;

&lt;p&gt;And I'm not even counting the intangible benefits: no rate limits, no vendor lock-in, full privacy, works offline.&lt;/p&gt;

&lt;h2&gt;
  
  
  When I &lt;em&gt;Will&lt;/em&gt; Upgrade
&lt;/h2&gt;

&lt;p&gt;I'm not against upgrades. I'm against &lt;em&gt;pointless&lt;/em&gt; upgrades. Here's when I'll swap a model:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;A new model is meaningfully better at a task I do daily.&lt;/strong&gt; Not 3% on a benchmark — a real, noticeable improvement in my actual workflow.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A model drops that fits my hardware and replaces two models I currently run.&lt;/strong&gt; Consolidation is a valid reason to upgrade.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security or stability improvements.&lt;/strong&gt; Bug fixes count. Feature glitter doesn't.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I test new models in a sandbox before they ever touch my production stack. The bar is: "Does this meaningfully change my daily output?" If the answer is no, I don't switch.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Lesson (For Myself As Much As Anyone)
&lt;/h2&gt;

&lt;p&gt;The model treadmill is real. FOMO is real. That HuggingFace notification that a new model just dropped is designed to make you feel like you're falling behind.&lt;/p&gt;

&lt;p&gt;You're not.&lt;/p&gt;

&lt;p&gt;The best model is the one you actually use to ship things. Not the one with the highest benchmark score. Not the one released yesterday. The one that's loaded, configured, and producing output for your project right now.&lt;/p&gt;

&lt;p&gt;I spent months chasing the cutting edge. Now I spend that time building. Turns out, building is more fun than benchmarking.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Sam Hartley runs a three-machine local AI setup and writes about making it actually work. He hasn't downloaded a new model in 4 months and he's fine with it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;→ &lt;a href="http://www.fiverr.com/s/XLyg" rel="noopener noreferrer"&gt;Custom automation setups on Fiverr&lt;/a&gt;&lt;br&gt;
→ &lt;a href="https://t.me/celebibot_en" rel="noopener noreferrer"&gt;Follow CelebiBots on Telegram&lt;/a&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  ai #llm #selfhosted #ollama #buildinginpublic
&lt;/h1&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>selfhosted</category>
      <category>ollama</category>
    </item>
    <item>
      <title>I Replaced All My Cloud AI With Local Models — Here's What Actually Broke</title>
      <dc:creator>Sam Hartley</dc:creator>
      <pubDate>Thu, 27 Aug 2026 08:02:20 +0000</pubDate>
      <link>https://dev.to/samhartley_dev/i-replaced-all-my-cloud-ai-with-local-models-heres-what-actually-broke-51k8</link>
      <guid>https://dev.to/samhartley_dev/i-replaced-all-my-cloud-ai-with-local-models-heres-what-actually-broke-51k8</guid>
      <description>&lt;h1&gt;
  
  
  I Replaced All My Cloud AI With Local Models — Here's What Actually Broke
&lt;/h1&gt;

&lt;p&gt;I write a lot about local AI. How it saves money. How it's private. How a $150 GPU can replace $200/month in API subscriptions. And all of that is true.&lt;/p&gt;

&lt;p&gt;But here's what I don't talk about enough: some things still don't work. Not "kind of work" — actually fail. Like, I-open-ChatGPT-and-do-it-manually fail.&lt;/p&gt;

&lt;p&gt;After 14 months of running a local-first AI setup (Mac Mini M4 + Windows PC with RTX 3060 + Ubuntu fallback), I want to be honest about the dealbreakers. The tasks where local models still lose. Because nobody's building the "my local LLM produced garbage and I wasted 40 minutes" blog post, and someone should.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Setup (Context for the Numbers)
&lt;/h2&gt;

&lt;p&gt;Before I list failures, here's what I'm working with:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Machine&lt;/th&gt;
&lt;th&gt;Hardware&lt;/th&gt;
&lt;th&gt;Models&lt;/th&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Mac Mini M4&lt;/td&gt;
&lt;td&gt;10-core CPU, 16GB RAM&lt;/td&gt;
&lt;td&gt;Qwen 3.5 9B, Granite 3.2 8B&lt;/td&gt;
&lt;td&gt;Orchestration, routing, writing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Windows PC&lt;/td&gt;
&lt;td&gt;RTX 3060 12GB&lt;/td&gt;
&lt;td&gt;Qwen 3 Coder 30B, DeepSeek R1 8B&lt;/td&gt;
&lt;td&gt;Code generation, reasoning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ubuntu box&lt;/td&gt;
&lt;td&gt;Older GPUs, 128GB RAM&lt;/td&gt;
&lt;td&gt;DeepSeek R1 (larger), MinCPM-V&lt;/td&gt;
&lt;td&gt;Heavy tasks, vision&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Electricity cost: ~$12/month. API costs: $0 for most tasks.&lt;/p&gt;

&lt;p&gt;It's good. It handles maybe 85% of what I throw at it. That other 15%? That's where the pain lives.&lt;/p&gt;

&lt;h2&gt;
  
  
  Dealbreaker #1: Complex Multi-Step Reasoning
&lt;/h2&gt;

&lt;p&gt;This is the big one. The one everyone talks about and nobody wants to admit.&lt;/p&gt;

&lt;p&gt;I had a bug in a distributed task queue — the kind where Worker A processes a job, Worker B picks up the result, but under load, B sometimes gets a stale version. Classic distributed systems problem.&lt;/p&gt;

&lt;p&gt;I fed the entire codebase (~2000 lines across 6 files) to DeepSeek R1 8B on the RTX 3060. It gave me a confident, detailed answer about adding a version field to the message format.&lt;/p&gt;

&lt;p&gt;The problem? The version field was already there. The bug was in the retry logic — when Worker B got a 409 conflict, it was retrying with the OLD payload instead of re-fetching. The model didn't catch this because it couldn't actually trace the execution flow through the code. It pattern-matched "distributed systems bug" and gave me the textbook answer.&lt;/p&gt;

&lt;p&gt;I pasted the same code into Claude Sonnet. It found the retry bug in 30 seconds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Time wasted on local:&lt;/strong&gt; 40 minutes (trying the "fix," testing, realizing it was wrong, trying again)&lt;br&gt;
&lt;strong&gt;Time with cloud:&lt;/strong&gt; 2 minutes&lt;br&gt;
&lt;strong&gt;Cost difference:&lt;/strong&gt; $0 vs ~$0.03&lt;/p&gt;

&lt;p&gt;That $0.03 would have saved me 38 minutes. Local models are bad at multi-step reasoning where you need to trace causation through code. They pattern-match instead of actually following the logic.&lt;/p&gt;

&lt;h2&gt;
  
  
  Dealbreaker #2: Anything Requiring Current Information
&lt;/h2&gt;

&lt;p&gt;My local models don't have internet access. I could hook them up to RAG pipelines and web search tools, and I have for specific projects. But for general "what's the latest X?" questions, local models are blind.&lt;/p&gt;

&lt;p&gt;Examples where I had to switch to cloud:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"What's the current version of Ollama?" — Local told me 0.5.7. It was 0.6.2.&lt;/li&gt;
&lt;li&gt;"Does Qwen 3.5 support tool calling?" — Local said no. It does (since training data cutoff was before that release).&lt;/li&gt;
&lt;li&gt;"What's the best GPU under $300 in 2026?" — Local recommended the RTX 3060 (which I already own). But the RX 7700 XT dropped to $280 and has 12GB VRAM too — local didn't know.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If your task depends on current information, local models are a non-starter unless you invest in a RAG pipeline that fetches real-time data. And at that point, you're building infrastructure, not just running a model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Dealbreaker #3: Long-Context Understanding
&lt;/h2&gt;

&lt;p&gt;Here's a scenario: I have a 15-page API contract from a client. I need to extract the authentication flow, the error handling requirements, and the rate limits, then generate a Python SDK that matches all of it.&lt;/p&gt;

&lt;p&gt;My 30B model can ingest the text. It produces something that looks like a Python SDK. But it misses things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The contract specifies rate limits per endpoint (not global) — the model made them global&lt;/li&gt;
&lt;li&gt;Error codes are hierarchical (4xx for client, 5xx for server, but also custom codes like 4201, 4302) — the model treated them as flat&lt;/li&gt;
&lt;li&gt;Authentication uses a rotating token with a 5-minute expiry — the model implemented static tokens&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each of these is a subtle but critical failure. The model "understood" the document in the sense that it could summarize it. But it didn't actually follow the specifications precisely.&lt;/p&gt;

&lt;p&gt;For documents under ~3 pages, local models are fine. Beyond that, especially with technical specifications, you start losing details that matter.&lt;/p&gt;

&lt;h2&gt;
  
  
  Dealbreaker #4: Vision Tasks (On Consumer Hardware)
&lt;/h2&gt;

&lt;p&gt;I run Granite 3.2 Vision 2B and MinCPM-V for image understanding. They're okay. They can tell you what's in a photo. They cannot:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Read a complex diagram and extract the architecture it describes&lt;/li&gt;
&lt;li&gt;Understand a UI mockup and generate matching CSS&lt;/li&gt;
&lt;li&gt;Accurately transcribe handwritten notes with mixed languages&lt;/li&gt;
&lt;li&gt;Distinguish between visually similar but functionally different components in a circuit diagram&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The 2B vision model runs fast (4 seconds on the Windows GPU). But it's 2B parameters. It sees a diagram and says "this is a flowchart." Great. I can see that too.&lt;/p&gt;

&lt;p&gt;The larger vision model (MinCPM-V on Ubuntu) does better but takes 3 minutes per image. For my automated pipeline that processes 20+ images per article, that's an hour of GPU time. Cloud vision APIs do the same in seconds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where local vision works:&lt;/strong&gt; "What's in this photo?" "Is this a cat or a dog?" "Extract the text from this screenshot."&lt;br&gt;
&lt;strong&gt;Where it fails:&lt;/strong&gt; Anything requiring nuanced visual understanding.&lt;/p&gt;

&lt;h2&gt;
  
  
  Dealbreaker #5: Generating Consistent Long-Form Content
&lt;/h2&gt;

&lt;p&gt;I know, I know — I wrote an article about how my AI content pipeline works. And it does. For 1000-1500 word articles with a clear structure.&lt;/p&gt;

&lt;p&gt;But give a local model a 3000-word target with a specific voice, specific technical examples, and a specific argument structure, and it starts to drift. Paragraph 3 sounds like me. Paragraph 7 sounds like a Wikipedia article. Paragraph 12 is repeating Paragraph 4 in different words.&lt;/p&gt;

&lt;p&gt;This is why my content pipeline uses short articles. Not because I don't have enough to say — because the model's coherence window is about 1500 words before it starts repeating itself or losing the thread.&lt;/p&gt;

&lt;p&gt;For this article, I'm writing it myself. The local model can handle the structure and individual sections. But threading a consistent argument through 2500 words? That's still a human job.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Actually Do Now (The Hybrid Setup)
&lt;/h2&gt;

&lt;p&gt;I've stopped pretending local can do everything. Here's my actual workflow:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Where It Goes&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Quick questions, summaries&lt;/td&gt;
&lt;td&gt;Local (Qwen 3.5 9B)&lt;/td&gt;
&lt;td&gt;Fast, free, good enough&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Code generation (&amp;lt; 200 lines)&lt;/td&gt;
&lt;td&gt;Local (Qwen 3 Coder 30B)&lt;/td&gt;
&lt;td&gt;Solid for most tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Code generation (complex, multi-file)&lt;/td&gt;
&lt;td&gt;Cloud (Claude Sonnet)&lt;/td&gt;
&lt;td&gt;Better at tracing logic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Short-form writing&lt;/td&gt;
&lt;td&gt;Local (Granite 3.2 8B)&lt;/td&gt;
&lt;td&gt;Good structure, decent tone&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long-form writing (&amp;gt; 1500 words)&lt;/td&gt;
&lt;td&gt;Myself, with local for structure&lt;/td&gt;
&lt;td&gt;Coherence window is the limit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vision (simple)&lt;/td&gt;
&lt;td&gt;Local (Granite Vision 2B)&lt;/td&gt;
&lt;td&gt;"What's in this photo" — fine&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vision (complex)&lt;/td&gt;
&lt;td&gt;Cloud (GPT-4o Vision)&lt;/td&gt;
&lt;td&gt;Architecture diagrams, UI mocks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Current information queries&lt;/td&gt;
&lt;td&gt;Cloud (Perplexity/GPT-4o)&lt;/td&gt;
&lt;td&gt;Local doesn't know today's news&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data extraction&lt;/td&gt;
&lt;td&gt;Local&lt;/td&gt;
&lt;td&gt;Surprisingly good, even at 7B&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Monthly cloud spend now:&lt;/strong&gt; ~$25 (down from $187)&lt;br&gt;
&lt;strong&gt;Monthly local spend:&lt;/strong&gt; ~$12 (electricity)&lt;br&gt;
&lt;strong&gt;Total:&lt;/strong&gt; ~$37/month for 85% local, 15% cloud&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Takeaway
&lt;/h2&gt;

&lt;p&gt;Local AI is incredible for what it does well. I saved over $2,000 in a year. I run my entire content pipeline, my code review, my daily summaries, and my monitoring — all on hardware I own.&lt;/p&gt;

&lt;p&gt;But if you're starting out and someone tells you "you don't need cloud AI at all," they're either lying or they don't do the kind of work that requires deep reasoning, current data, or complex visual understanding.&lt;/p&gt;

&lt;p&gt;The honest answer is hybrid. Local for the 80-85% of tasks where it's good enough (and cheaper and faster and more private). Cloud for the rest. And being smart about which is which.&lt;/p&gt;

&lt;p&gt;The biggest trap isn't the technology. It's the denial. When you've invested time and money into a local setup, you want it to work for everything. You start accepting lower quality outputs because "it's free." But free output that's wrong isn't free — it costs your time to fix it.&lt;/p&gt;

&lt;p&gt;I still use local for the vast majority of my work. I just know when to stop.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Sam Hartley runs a multi-machine AI home lab and writes about the infrastructure that makes local AI actually usable. Yes, he still has a ChatGPT subscription. No, he's not proud of it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;→ &lt;a href="http://www.fiverr.com/s/XLyg" rel="noopener noreferrer"&gt;Custom automation setups on Fiverr&lt;/a&gt;&lt;br&gt;
→ &lt;a href="https://t.me/celebibot_en" rel="noopener noreferrer"&gt;Follow CelebiBots on Telegram&lt;/a&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  ai #llm #selfhosted #ollama #honest #localai #buildinpublic
&lt;/h1&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>selfhosted</category>
      <category>ollama</category>
    </item>
  </channel>
</rss>
