<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ariel Chang</title>
    <description>The latest articles on DEV Community by Ariel Chang (@arielchangdev).</description>
    <link>https://dev.to/arielchangdev</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4167664%2Fb69c8bc2-75de-47cc-9799-e80f9cbb6dd9.jpg</url>
      <title>DEV Community: Ariel Chang</title>
      <link>https://dev.to/arielchangdev</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/arielchangdev"/>
    <language>en</language>
    <item>
      <title>Nine Bugs That Taught Me Reliability: Post-Mortem of a Zero-Cost Hybrid AI System</title>
      <dc:creator>Ariel Chang</dc:creator>
      <pubDate>Wed, 07 Oct 2026 03:51:46 +0000</pubDate>
      <link>https://dev.to/arielchangdev/nine-bugs-that-taught-me-reliability-post-mortem-of-a-zero-cost-hybrid-ai-system-4bce</link>
      <guid>https://dev.to/arielchangdev/nine-bugs-that-taught-me-reliability-post-mortem-of-a-zero-cost-hybrid-ai-system-4bce</guid>
      <description>&lt;p&gt;canonical_url: &lt;a href="https://arielchangdev.github.io/angelina-finance-agent/" rel="noopener noreferrer"&gt;https://arielchangdev.github.io/angelina-finance-agent/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Building a hybrid local+cloud AI system under two hard rules — $0 cost and never break production. The interesting part wasn't the happy path; it was the 9 production bugs and what each taught me about where reliability actually lives.&lt;/p&gt;

&lt;p&gt;A war-stories write-up. The system was fun to design. The &lt;em&gt;interesting&lt;/em&gt; part was the nine ways it quietly broke — and what each one taught me about where reliability actually lives.&lt;/p&gt;

&lt;p&gt;I spent a few months building and operating a small AI agent under two rules that never moved:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;It must cost exactly $0.&lt;/strong&gt; Every component lives inside a provider's always-free tier, with a $1 budget alert sitting on top as a tripwire.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It must never regress the running production service.&lt;/strong&gt; There's a live daily job that real usage depends on. Every change had to ship without breaking what was already working.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Those two constraints turned out to be the whole story. When you can't throw money or a staging cluster at a problem, you're forced to get the boring details right. This post is about the boring details — specifically, the nine production bugs I had to fight, written up so you can skip the ones I didn't.&lt;/p&gt;

&lt;p&gt;If you want the full architecture and trade-off write-up, that's in a separate case study (linked at the bottom). Here I'll keep the system context short and spend the words on the failures.&lt;/p&gt;




&lt;h2&gt;
  
  
  The system, in one paragraph
&lt;/h2&gt;

&lt;p&gt;It's a self-hosted AI agent (Python + FastAPI, a hosted LLM on its free tier, a vector DB for retrieval, SQLite for state). A cron job runs a scheduled daily task, logs results to a shared spreadsheet, and keeps a small knowledge base in sync. Over three iterations, it grew from one home-lab VM into a &lt;strong&gt;hybrid topology&lt;/strong&gt;: a &lt;em&gt;local&lt;/em&gt; VM is the primary and does all the real work; a &lt;em&gt;cloud&lt;/em&gt; always-free VM sits passive and only takes over if the local node goes dark. Coordination happens through a heartbeat written to a shared sheet; data moves through a token-gated HTTPS sync endpoint. That's it. Everything below is about keeping that two-node, zero-budget setup correct.&lt;/p&gt;

&lt;p&gt;Here's roughly how a normal day flows:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;local VM  ──(daily task + heartbeat stamp)──▶  shared sheet  ◀──(23:00 read)──  cloud VM (passive)
   └──────────────── token-gated HTTPS sync (last-write-wins) ────────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the fun part.&lt;/p&gt;




&lt;h2&gt;
  
  
  Bug 1 — The deploy that clobbered its own credentials
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Symptom.&lt;/strong&gt; I pushed a cleaned-up, "ready for GitHub" copy of the code onto the production box. The next daily run died: the LLM call came back &lt;code&gt;400 "API key not valid"&lt;/code&gt;, and the downstream notification returned &lt;code&gt;404&lt;/code&gt;. Nothing in the diff &lt;em&gt;looked&lt;/em&gt; like it touched credentials.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Investigation.&lt;/strong&gt; The code logic was identical. So it wasn't logic. I diffed the deployed files against what had been running and found the clean copy carried &lt;strong&gt;placeholder&lt;/strong&gt; credential values — the ones I'd scrubbed in for the public repo.&lt;/p&gt;

&lt;p&gt;*&lt;em&gt;Root cause. *&lt;/em&gt; Credentials were living &lt;em&gt;in source&lt;/em&gt;. "Deploy the latest code" therefore also meant "overwrite the real keys with placeholders." The deploy did exactly what I told it to; I'd just told it something stupid.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fix.&lt;/strong&gt; Pull every secret out of source and into the environment / a git-ignored &lt;code&gt;.env&lt;/code&gt;. Committed files carry placeholders only; real values never enter the repo.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# committed (safe)
API_KEY=replace-me

# on the box, in a git-ignored .env (never committed)
API_KEY=&amp;lt;the real value&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Takeaway.&lt;/strong&gt; Config is not code. The moment you can deploy code without touching config, "deploy latest" becomes a safe, boring operation instead of a loaded gun. This single separation prevented a whole category of future incidents.&lt;/p&gt;




&lt;h2&gt;
  
  
  Bug 2 — Cron ran, but with none of my environment
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Symptom.&lt;/strong&gt; The app worked perfectly when I ran it by hand. The cron-triggered run behaved as if its configuration didn't exist.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Investigation.&lt;/strong&gt; Classic "works in my shell, not in cron." I dumped the environment the cron process actually saw. It was nearly empty — none of the variables I rely on from &lt;code&gt;.env&lt;/code&gt; were present.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Root cause.&lt;/strong&gt; The cron entry invoked &lt;code&gt;python&lt;/code&gt; directly without sourcing &lt;code&gt;.env&lt;/code&gt; first. Cron runs in a minimal, non-login shell: your &lt;code&gt;.bashrc&lt;/code&gt;, your profile, your exported variables — none of it comes along.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fix.&lt;/strong&gt; Source the environment inside the cron command before launching:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;0 23 &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt;  &lt;span class="nb"&gt;cd&lt;/span&gt; /opt/app &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt; /opt/app/.env &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; /opt/app/venv/bin/python &lt;span class="nt"&gt;-m&lt;/span&gt; app. daily &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; /var/log/app.log 2&amp;gt;&amp;amp;1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Takeaway.&lt;/strong&gt; cron is not your shell. If a job depends on environment, make the job load that environment explicitly. Don't assume anything from an interactive session is inherited.&lt;/p&gt;




&lt;h2&gt;
  
  
  Bug 3 — "Works on my machine" had a version number attached
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Symptom.&lt;/strong&gt; Code that ran clean locally threw a &lt;code&gt;SyntaxError&lt;/code&gt; on the cloud VM — at &lt;em&gt;import&lt;/em&gt; time, before any logic executed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Investigation.&lt;/strong&gt; A &lt;code&gt;SyntaxError&lt;/code&gt; on identical source means the parser differs, which means the interpreter version differs. Local was Python 3.12; the cloud VM was on 3.11.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Root cause.&lt;/strong&gt; An f-string containing a backslash. Python 3.12 relaxed the f-string grammar and accepts it; 3.11 does not. Same characters, different verdict.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fix.&lt;/strong&gt; Rewrite the expression to be valid on the &lt;em&gt;older&lt;/em&gt; interpreter — pull the backslash out of the f-string into a named variable — so it parses everywhere.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# fails to parse on 3.11
&lt;/span&gt;&lt;span class="n"&gt;msg&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;line one&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;line two&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="c1"&gt;# version-safe
&lt;/span&gt;&lt;span class="n"&gt;nl&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;msg&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;line one&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;nl&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;line two&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Takeaway.&lt;/strong&gt; "Works on my machine" always has a version number attached. Either match the runtime across environments, or write to the lowest version you have to support. Runtime parity is part of the contract, not a detail.&lt;/p&gt;




&lt;h2&gt;
  
  
  Bug 4 — SQLite refused my upsert because the index was &lt;em&gt;partial&lt;/em&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Symptom.&lt;/strong&gt; An &lt;code&gt;ON CONFLICT&lt;/code&gt; upsert on sync records failed outright. The error pointed at the conflict target.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Investigation.&lt;/strong&gt; The &lt;code&gt;sync_id&lt;/code&gt; column had a unique index, so an upsert &lt;em&gt;should&lt;/em&gt; have an arbiter to key on. Reading the SQLite docs more carefully: the conflict target of an upsert must map to a specific kind of index.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Root cause.&lt;/strong&gt; The unique index was a &lt;strong&gt;partial&lt;/strong&gt; index (it had a &lt;code&gt;WHERE&lt;/code&gt; clause). SQLite will not use a partial index as an upsert conflict arbiter. It wasn't a bug in my SQL so much as a bug in my assumption that "unique index" and "upsert arbiter" are the same thing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fix.&lt;/strong&gt; Convert the index to a &lt;strong&gt;full&lt;/strong&gt; unique index, plus a self-healing migration that detects the old partial index on startup and repairs existing databases in place.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- won't arbitrate an upsert&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;UNIQUE&lt;/span&gt; &lt;span class="k"&gt;INDEX&lt;/span&gt; &lt;span class="n"&gt;ix_sync&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;records&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sync_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;sync_id&lt;/span&gt; &lt;span class="k"&gt;IS&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- will&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;UNIQUE&lt;/span&gt; &lt;span class="k"&gt;INDEX&lt;/span&gt; &lt;span class="n"&gt;ix_sync&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;records&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sync_id&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Takeaway.&lt;/strong&gt; Upsert arbiters have precise requirements, and they differ by engine. When the database rejects something that "should" work, read the exact constraints before you reach for a workaround.&lt;/p&gt;




&lt;h2&gt;
  
  
  Bug 5 — The sync table that grew forever (namespace bounce)
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Symptom.&lt;/strong&gt; Every daily sync, the conversation table got bigger. Not from new activity — from the &lt;em&gt;same&lt;/em&gt; records multiplying.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Investigation.&lt;/strong&gt; I traced individual records across sync rounds. A record that originated on local showed up on cloud, then came &lt;em&gt;back&lt;/em&gt; to local, then went out to cloud again — each hop with a slightly different ID. The system saw each hop as a brand-new record and inserted it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Root cause.&lt;/strong&gt; On each export, records were re-namespaced (&lt;code&gt;local: N&lt;/code&gt; → &lt;code&gt;cloud: M&lt;/code&gt; → &lt;code&gt;local:P&lt;/code&gt; → …). Because identity changed every round-trip, nothing was ever recognized as "already seen." Classic bi-directional-sync duplication.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fix.&lt;/strong&gt; Two rules: (1) preserve the &lt;strong&gt;original origin id&lt;/strong&gt; on re-export so identity is stable for the life of the record, and (2) when exporting, &lt;strong&gt;skip records that originated on this instance&lt;/strong&gt; — don't ship something back to where it was born.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway.&lt;/strong&gt; In bi-directional sync, identity must be stable and origin-aware. If a record's identity mutates as it travels, you don't have sync — you have a photocopier.&lt;/p&gt;




&lt;h2&gt;
  
  
  Bug 6 — The sync endpoint timed out because it did the heavy work inline
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Symptom.&lt;/strong&gt; Sync requests started timing out, specifically on the small (1GB RAM) VM.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Investigation.&lt;/strong&gt; The round-trip did an exchange &lt;em&gt;and&lt;/em&gt; re-embedded the imported text into vectors synchronously, inside the request handler. Re-embedding a few hundred chunks on a CPU-only, memory-starved box takes real wall-clock time — well past the HTTP timeout. On a beefier node it squeaked by, which is why it hid for a while.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Root cause.&lt;/strong&gt; Expensive, variable-duration work (re-embedding) was sitting directly on the request path. The response couldn't return until the slowest possible operation finished.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fix.&lt;/strong&gt; Return the pre-exchange snapshot &lt;em&gt;immediately&lt;/em&gt;, and move the import + re-embed into a background task. The HTTP round-trip now returns in seconds; the heavy lifting completes eventually, off the wire.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;POST /sync/exchange
  ├─ read peer payload
  ├─ respond NOW with our pre-exchange snapshot   ← bounded
  └─ enqueue import + re-embed as background work  ← eventual
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Takeaway.&lt;/strong&gt; Keep slow, unbounded work off the request path. A response time should be bounded by design; the work behind it can be eventual. If a handler's latency depends on data volume, that's a smell.&lt;/p&gt;




&lt;h2&gt;
  
  
  Bug 7 — Failover that only delivered every &lt;em&gt;other&lt;/em&gt; day
&lt;/h2&gt;

&lt;p&gt;This is my favorite, because the bug was &lt;em&gt;subtle&lt;/em&gt; and the symptom was &lt;em&gt;weird&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Symptom.&lt;/strong&gt; During an extended local outage, the cloud backup delivered the daily task — but only on alternating days. Day 1: delivered. Day 2: skipped. Day 3: delivered. Like a metronome.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Investigation.&lt;/strong&gt; The failover logic reads a heartbeat the local node stamps after each successful run; if local has been silent past a ~25-hour window, cloud declares it offline and takes over. So why the oscillation? I looked at &lt;em&gt;who writes the heartbeat&lt;/em&gt;. When cloud failed over, it was stamping the "local presence" heartbeat too.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Root cause. **&lt;/strong&gt; The failover push also stamped the local-presence signal. So after cloud covered Day 1, the heartbeat looked fresh — "local is alive!" — and on Day 2 cloud dutifully skipped. But local was still down, so by Day 3 the heartbeat was stale again, and cloud failed over again. The presence signal was lying because the wrong actor was writing it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fix.&lt;/strong&gt; Only the local role may stamp the local-presence heartbeat. A failover push does its job and does &lt;strong&gt;not&lt;/strong&gt; touch the local heartbeat. Now an extended outage reads as a continuous outage, and cloud covers every day until local comes back.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway.&lt;/strong&gt; A presence signal must be written &lt;em&gt;only by the entity whose presence it represents&lt;/em&gt;. The moment a second actor can write "I'm here" on someone else's behalf, your liveness detection is corrupted. This generalizes well beyond failover — any time you have a "last seen" field, guard who's allowed to touch it.&lt;/p&gt;




&lt;h2&gt;
  
  
  Bug 8 — A momentary 503 cost me a whole day
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Symptom.&lt;/strong&gt; One day: no output at all. The upstream model provider had a brief hiccup right when the job ran.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Investigation.&lt;/strong&gt; The logs showed a short burst of &lt;code&gt;503&lt;/code&gt;s during the daily window. My retry logic gave up after too few attempts, so a transient outage that lasted a couple of minutes turned into a missed day.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Root cause. **&lt;/strong&gt; Thin retries. A single scheduled attempt with minimal backoff can't ride out even a short 5xx burst, and a once-a-day job has no second chance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fix.&lt;/strong&gt; Harden the retry/backoff: 5 attempts with capped backoff (roughly 20/40/60/90/120s), covering &lt;code&gt;500/502/503/504/429&lt;/code&gt; and network errors.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;attempt 1 → 503 → wait 20s
attempt 2 → 503 → wait 40s
attempt 3 → 200 ✓   (delivered on time)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Takeaway.&lt;/strong&gt; For infrequent, high-stakes jobs, retries aren't optional polish — they're the difference between "delivered" and "silently missed." And the payoff was real: the hardened logic later rode out an actual 3×-503 event and still delivered on schedule. Backoff earns its keep on the day you forgot it was there.&lt;/p&gt;




&lt;h2&gt;
  
  
  Bug 9 — The scheduler fired at the wrong hour because UTC
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Symptom.&lt;/strong&gt; The cloud failover check, meant to run at 23:00 local, was firing at 07:00 local. Eight hours off, consistently.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Investigation.&lt;/strong&gt; Eight hours is a timezone offset, not a bug in the cron expression. I checked the VM's clock: it was on UTC.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Root cause.&lt;/strong&gt; Cloud VMs commonly default to UTC. My &lt;code&gt;23 0 * * *&lt;/code&gt; cron was correct — against the &lt;em&gt;wrong wall clock&lt;/em&gt;. The scheduler was doing exactly what I asked, in a timezone I didn't mean.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fix.&lt;/strong&gt; Set the VM timezone explicitly to the intended local zone so the cron expression lines up with the real-world time I care about.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;timedatectl set-timezone &amp;lt;intended/zone&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Takeaway.&lt;/strong&gt; A scheduler is only as correct as the clock beneath it. Never assume the host's timezone — set it explicitly, and write schedules against a timezone you've verified.&lt;/p&gt;




&lt;h2&gt;
  
  
  Meta-lessons: reliability lives in the boring details
&lt;/h2&gt;

&lt;p&gt;Step back from the nine, and a pattern shows up. None of these were exotic. Not one was in the "AI" part. Every single one lived in plumbing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Separate config from code.&lt;/strong&gt; Then deploying code can't poison production config. (Bug 1)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;cron carries no environment.&lt;/strong&gt; Load it explicitly; inherit nothing. (Bug 2)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pin and match your runtime versions.&lt;/strong&gt; "Works on my machine" has a version number. (Bug 3)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Know your database's exact rules.&lt;/strong&gt; Upsert arbiters, index types — read the fine print instead of assuming. (Bug 4)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Make identity stable and origin-aware in sync. **&lt;/strong&gt; Mutating identity means multiplying records. (Bug 5)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep slow work off the request path. **&lt;/strong&gt; Bounded responses, eventual work. (Bug 6)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Presence signals belong to their owner.&lt;/strong&gt; Only the subject writes "I'm here." (Bug 7)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retry with backoff for anything that matters and runs rarely. **&lt;/strong&gt; (Bug 8)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Set timezones explicitly. **&lt;/strong&gt; Schedulers trust the clock blindly. (Bug 9)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The through-line: the model call was never the hard part. Keeping a free, two-node system correct, secure, and cheap through nine real failures was the hard part — and the fixes were almost all about discipline in the unglamorous layer. Config hygiene, environment, runtime parity, database semantics, identity, request-path discipline, signal ownership, backoff, clocks.&lt;/p&gt;

&lt;p&gt;That's where reliability actually lives. Not in the clever bits. In the boring ones you were tempted to skip.&lt;/p&gt;




&lt;h2&gt;
  
  
  Takeaway
&lt;/h2&gt;

&lt;p&gt;If you're building something small and constrained — a side project, a home lab, a lean production service — resist the urge to treat the plumbing as beneath you. The constraints ($0, don't break prod) didn't make the system worse; they forced the discipline that made it reliable. Nine bugs later, it's run at zero cost, survived a multi-day failover with no duplicate output, and shipped a disciplined release history.&lt;/p&gt;

&lt;p&gt;Boring is a feature.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Code &amp;amp; full case study: &lt;a href="https://github.com/arielchangdev/angelina-finance-agent" rel="noopener noreferrer"&gt;https://github.com/arielchangdev/angelina-finance-agent&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  About me / let's connect
&lt;/h2&gt;

&lt;p&gt;I'm &lt;strong&gt;Ariel Chang&lt;/strong&gt; — I build reliable, low-cost, self-hosted systems and write about the unglamorous engineering that keeps them honest.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Full case study &amp;amp; source:&lt;/strong&gt; &lt;a href="https://github.com/arielchangdev/angelina-finance-agent" rel="noopener noreferrer"&gt;https://github.com/arielchangdev/angelina-finance-agent&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Portfolio page:&lt;/strong&gt; &lt;a href="https://arielchangdev.github.io/angelina-finance-agent/" rel="noopener noreferrer"&gt;https://arielchangdev.github.io/angelina-finance-agent/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LinkedIn:&lt;/strong&gt; &lt;a href="https://www.linkedin.com/in/ariel-chang-690a89160" rel="noopener noreferrer"&gt;https://www.linkedin.com/in/ariel-chang-690a89160&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you're building something under tight constraints, I'd genuinely enjoy comparing notes. Questions and war stories of your own are very welcome in the comments.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>python</category>
      <category>sre</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
