<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Christian Anderson</title>
    <description>The latest articles on DEV Community by Christian Anderson (@iam-tech).</description>
    <link>https://dev.to/iam-tech</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4110779%2Ff7ed549e-875c-4c07-8cb7-c1d3ff6a355e.jpg</url>
      <title>DEV Community: Christian Anderson</title>
      <link>https://dev.to/iam-tech</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/iam-tech"/>
    <language>en</language>
    <item>
      <title>Homelab dashboard IP sync failing silently: when a miss looks like no change</title>
      <dc:creator>Christian Anderson</dc:creator>
      <pubDate>Fri, 09 Oct 2026 08:00:13 +0000</pubDate>
      <link>https://dev.to/iam-tech/homelab-dashboard-ip-sync-failing-silently-when-a-miss-looks-like-no-change-2id7</link>
      <guid>https://dev.to/iam-tech/homelab-dashboard-ip-sync-failing-silently-when-a-miss-looks-like-no-change-2id7</guid>
      <description>&lt;p&gt;My homelab dashboard has a tile for each service. Most of those services sit at fixed addresses. A few don't: my 3D printer gets its address from DHCP, and when it moves, the tile points at nothing.&lt;/p&gt;

&lt;p&gt;So on 12 July I wrote a small job to keep the tiles right. It runs every 15 minutes, works out the current address of each tracked device, and rewrites the dashboard's config only if something changed. It tracks three devices: the printer, Home Assistant and Uptime Kuma.&lt;/p&gt;

&lt;p&gt;For three and a half weeks it did nothing at all, and its log said so every quarter of an hour in a way I read as good news.&lt;/p&gt;

&lt;h2&gt;
  
  
  Built to be safe
&lt;/h2&gt;

&lt;p&gt;I made two design choices on purpose, and both come back into this story.&lt;/p&gt;

&lt;p&gt;First, it is a plain Python script with no language model anywhere in it. An agent that "works out" an IP address can invent one, and a wrong address written into a dashboard is worse than a stale one.&lt;/p&gt;

&lt;p&gt;Second, it fails safe. If it can't find a device, it leaves that device's tile alone. It never writes a guess. A device that is switched off shouldn't get its tile blanked, and the next run can pick it up again.&lt;/p&gt;

&lt;p&gt;That second choice is the one that hid the problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it found devices
&lt;/h2&gt;

&lt;p&gt;My router doesn't serve forward DNS for local &lt;code&gt;.lan&lt;/code&gt; names, so asking "what is the printer's address?" by name never worked. Reverse lookups did: DHCP leases were mirrored into the resolver, so asking "what name belongs to this address?" gave an answer.&lt;/p&gt;

&lt;p&gt;That resolver rate-limits reverse lookups hard. Sweeping all 254 addresses at 64 workers turned up about 13 names. Even 8 workers was flaky. So the job had two paths:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Confirm.&lt;/strong&gt; Do a reverse lookup on each device's last-known address, three queries in all, and accept it if the name still matches.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sweep.&lt;/strong&gt; Only for a device that has moved: ping the whole subnet twice, then do reverse lookups on live hosts only, four at a time, and stop as soon as it's found.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;It worked. The log from the first days shows &lt;code&gt;no changes (3 resolved, 0 miss)&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The log line I kept reading as fine
&lt;/h2&gt;

&lt;p&gt;On 16 August I actually read the log. Run after run looked like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;MISS: Voron 3D Printer (...) not on &amp;lt;subnet&amp;gt; — leaving tile untouched
MISS: Home Assistant (...) not on &amp;lt;subnet&amp;gt; — leaving tile untouched
MISS: Uptime Kuma (...) not on &amp;lt;subnet&amp;gt; — leaving tile untouched
no changes (0 resolved, 3 miss)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;"No changes" was true. It was also useless. Since the evening of 23 July, almost every run had found none of the three devices. It couldn't see anything, so it changed nothing, so it reported nothing to change. Nothing alerted, because the only thing it ever notified on was a change.&lt;/p&gt;

&lt;p&gt;The log has 4,422 lines reading &lt;code&gt;no changes (0 resolved, 3 miss)&lt;/code&gt;. That isn't 4,422 runs, though. It's about half that, because of a second bug I found in the same pass: every line was written twice. The script printed to stderr &lt;em&gt;and&lt;/em&gt; appended to its log file, and cron redirected stderr into that same file. Even the count of the failure was wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it went blind
&lt;/h2&gt;

&lt;p&gt;Reverse DNS had stopped working entirely. &lt;code&gt;.lan&lt;/code&gt; names came back NXDOMAIN from the resolver the job uses, and also when I asked the router directly. Reverse lookups of addresses I knew were good came back empty. The router was otherwise healthy: it answered pings with no loss and resolved normal forward DNS. Bare hostnames still resolved, but only to IPv6 addresses on my VPN overlay, never the LAN IPv4 the tiles needed.&lt;/p&gt;

&lt;p&gt;I never found out why the router stopped answering reverse lookups. I stopped depending on it instead.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: ask something that knows
&lt;/h2&gt;

&lt;p&gt;The rule I took from it: &lt;strong&gt;don't infer an address from DNS when an authoritative source exists.&lt;/strong&gt; Two such sources were sitting there already.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Containers ask Proxmox.&lt;/strong&gt; Two of the three devices are LXC containers, and Proxmox knows their addresses. The job now reads them from the nodes over SSH. Only running containers with an address on the right subnet count. One container name exists on both nodes. That is logged as ambiguous and refused, not guessed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Physical boxes are found by MAC.&lt;/strong&gt; The printer runs on a Raspberry Pi. Its MAC address doesn't change when its IP does, which is exactly the drift this job exists for. It checks the ARP table first. If the MAC isn't there, it pings the last-known address to refresh the table. Only then does it do a full ping-sweep.&lt;/p&gt;

&lt;p&gt;Reverse DNS is still there, but only as the last fallback.&lt;/p&gt;

&lt;p&gt;The first run after the change: &lt;code&gt;3 resolved, 0 miss&lt;/code&gt;. Then the embarrassing part. I checked the live dashboard config, and all three tiles had been correct the whole time. None of the devices had moved. The sync had been blind, not wrong, so the dashboard never showed a symptom.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two more fixes in the same pass
&lt;/h2&gt;

&lt;p&gt;The double-logging: the script now writes to the file and only echoes to stderr when it's attached to a terminal. It also echoes if the file write fails, so output can't disappear silently.&lt;/p&gt;

&lt;p&gt;And timing: 38.0 seconds of a 38.1-second run were spent asking Proxmox for container addresses. That was two SSH calls, each running a helper that called &lt;code&gt;pct&lt;/code&gt; three times per container. It came to around 3,600 &lt;code&gt;pct&lt;/code&gt; invocations a day for addresses that almost never change. They're cached for an hour now. But the hour isn't the main guard: if a cached address stops answering, the job refreshes from Proxmox straight away. I tested that by poisoning the cache with a dead address. It noticed, refreshed and recovered the right one. The run dropped from 38 seconds to 0.34.&lt;/p&gt;

&lt;h2&gt;
  
  
  Did it earn its keep?
&lt;/h2&gt;

&lt;p&gt;Since then the printer has moved twice. On 7 September and again on 10 September it came back on a new DHCP address. Both times the job found it by MAC and rewrote the tile, backing up the config first. That is the job it was built for, done twice, after three and a half weeks of finding nothing while saying "no changes".&lt;/p&gt;

&lt;h2&gt;
  
  
  "No change" and "couldn't look" are different answers
&lt;/h2&gt;

&lt;p&gt;The log still has a problem, and I'd rather say so. Most runs today end &lt;code&gt;no changes (2 resolved, 1 miss)&lt;/code&gt;. The miss is the printer: its MAC isn't on the network much of the time. The script does keep misses in a separate list in its state file. But it still notifies only on a change, so a device it can't find is still a log line nobody reads.&lt;/p&gt;

&lt;p&gt;What I'd do next, and what I'd suggest for any job shaped like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Return three outcomes, not two.&lt;/strong&gt; "Checked, unchanged", "checked, changed" and "couldn't check" are different results. Don't let the third collapse into the first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Alert on consecutive misses, not on each one.&lt;/strong&gt; One miss is a printer that's switched off. Every device missing on every run for a day is a broken job.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Put the resolved count where you'll see it.&lt;/strong&gt; &lt;code&gt;0 resolved&lt;/code&gt; was in every line. I just never looked.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This wasn't a one-off, either. This month my blog's publish queue ran empty after 18 September. Both publishers logged "queue empty — nothing to publish" and exited 0, and nothing alerted. A healthy, idle publisher looks just like a healthy, busy one unless you check how deep the queue is. There is now a check that alerts when the queue is empty.&lt;/p&gt;

&lt;p&gt;And on 7 September, a canary that watches my paid model provider switched four agent profiles back to local models after its probe failed on every model. It turned out to be the probe's input, not the models. It only knew how to switch back after a credit problem, not after a failover, so it logged "nothing to watch" for 17 days. It now re-probes after a failover and switches back on its own.&lt;/p&gt;

&lt;p&gt;All three were healthy processes reporting a true, reassuring sentence about a state they hadn't actually checked.&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;🤖 &lt;em&gt;Drafted with AI assistance from my own homelab notes, logs and repos, then reviewed and edited before publishing.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>proxmox</category>
      <category>homelab</category>
      <category>python</category>
      <category>networking</category>
    </item>
    <item>
      <title>AI blog drafts fact-checked: my drafter invented a bug, 8 of 9 needed fixes</title>
      <dc:creator>Christian Anderson</dc:creator>
      <pubDate>Thu, 08 Oct 2026 08:00:14 +0000</pubDate>
      <link>https://dev.to/iam-tech/ai-blog-drafts-fact-checked-my-drafter-invented-a-bug-8-of-9-needed-fixes-2ibm</link>
      <guid>https://dev.to/iam-tech/ai-blog-drafts-fact-checked-my-drafter-invented-a-bug-8-of-9-needed-fixes-2ibm</guid>
      <description>&lt;p&gt;On 18 September I published a post called &lt;a href="https://dev.to/c1-anderson/how-i-post-every-day-without-a-content-team-or-a-lying-robot-the-writer-pipeline-that-turns-real-4kgb"&gt;AI blog writing pipeline without made-up facts&lt;/a&gt;. Its central claim was that the model only phrases facts a script has gathered, so "it never had the chance to fabricate one", and that the whole system "fails toward silence, never toward fabrication".&lt;/p&gt;

&lt;p&gt;Six days later I fact-checked every draft sitting in the review folder against the code, the logs and my own notes. The claim did not hold. This is what I found, and what I have changed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The draft that invented a bug
&lt;/h2&gt;

&lt;p&gt;The drafter had written a post about my podcast pipeline, which turns an article into a two-voice episode using Piper text-to-speech over the Wyoming protocol. My log recorded it as the first draft to pass a new set of checks. Those checks were about format: title length, description length, tags. It passed them all.&lt;/p&gt;

&lt;p&gt;The content was mostly fiction. The draft said:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the bug was &lt;strong&gt;endianness&lt;/strong&gt;, fixed by flipping a &lt;code&gt;&amp;gt;&lt;/code&gt; to a &lt;code&gt;&amp;lt;&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;the text-to-speech model had a "reasoning field" that caused trouble;&lt;/li&gt;
&lt;li&gt;the fix took &lt;strong&gt;two weeks of debugging&lt;/strong&gt;;&lt;/li&gt;
&lt;li&gt;synthesis ran at &lt;strong&gt;90% real-time&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It also contained, in the body text, a sentence that began "wait, no. Let me be precise." The model had corrected itself mid-draft and the correction had been published into the post.&lt;/p&gt;

&lt;p&gt;None of those four things happened. The real bugs were two, and both are in the code:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The Wyoming frame has three parts, not two.&lt;/strong&gt; A header line can declare a &lt;code&gt;data_length&lt;/code&gt;, and the event's data then follows as its own block of bytes after the newline, not inline in the header. My first reader went header, then payload, landed in the middle of a JSON object, and failed with &lt;code&gt;Extra data: line 1 column 62&lt;/code&gt;. That looks like a broken server. It was a field I hadn't read.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The model writing the script returned nothing usable.&lt;/strong&gt; Not the speech model: the language model writing the dialogue. Its answer came back with &lt;code&gt;content&lt;/code&gt; empty and the whole token budget inside a &lt;code&gt;thinking&lt;/code&gt; key. The pipeline logged a cheerful "0 turns" rather than an error. The code now raises when content is empty and &lt;code&gt;thinking&lt;/code&gt; is populated, and it switches thinking off in the request body.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The speed figure was invented too. What I actually measured was 8.4 seconds of audio in 2.9 seconds of wall time, a real-time factor of about 0.35.&lt;/p&gt;

&lt;p&gt;I rewrote the post from the code and published it the same day as &lt;a href="https://dev.to/c1-anderson/piper-tts-over-wyoming-the-frame-format-and-empty-model-that-broke-my-podcast-1c1l"&gt;Piper TTS over Wyoming: the frame format and empty model that broke my podcast&lt;/a&gt;. The fabricated original is kept, marked as such, so I can compare the two.&lt;/p&gt;

&lt;p&gt;The irony is worth stating plainly. The invented bug was more conventional than the real one. My guess is that endianness is what a model reaches for when it knows a post is about binary framing and doesn't know the details. The true story, a missed field in a three-part frame, is the one other people will hit.&lt;/p&gt;

&lt;h2&gt;
  
  
  It wasn't the first time
&lt;/h2&gt;

&lt;p&gt;Looking back, the warnings were already in my notes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;16 September:&lt;/strong&gt; the topic radar picked up a public repo that turned out to have no code in it, only a README. The drafter wrote a whole build story from the one-line description. It was convincing enough that it was read as a real project. I renamed the draft so it couldn't be queued and added a guard that skips any repo with no commits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;20 September:&lt;/strong&gt; a fix that let malformed drafts survive (the model was wrapping its front matter in a code fence) also let through a draft that invented a brownout threshold and a soldering step, neither of which was in the facts it was given.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;21 September:&lt;/strong&gt; on several runs the model emitted its own tool-call syntax instead of a post, because it was trying to call a tool inside a text-only request.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each time I fixed the symptom in front of me. None of those fixes made a draft true.&lt;/p&gt;

&lt;h2&gt;
  
  
  Nine drafts, checked line by line
&lt;/h2&gt;

&lt;p&gt;On 24 September the queue was empty and there were nine drafts waiting. I checked each claim against the thing it described. The result:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;1 had a false core.&lt;/strong&gt; A post about an Ollama deadlock credited the fix to a configuration setting that, on its own, did not hold. What actually fixed it was a small guard script. The draft also said I had pinned the model myself, which I hadn't, and it wrapped the whole thing in a debugging story that never happened.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;8 needed fixes&lt;/strong&gt;, from heavy to small.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;All 9&lt;/strong&gt; were missing a description in the front matter.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The errors were not random. They fell into a handful of types.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Invented process story.&lt;/strong&gt; Drafts described alternatives I "tried" and rejected. One post about using Discord as the interface for my agents listed two other chat platforms I had supposedly evaluated. It also described a rate-limit error and a message-chunking fix that didn't exist, and said voice notes worked, which they don't. Models write the post they expect, and posts usually have a "what I tried first" section.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A confident wrong fix.&lt;/strong&gt; The deadlock post above. This is the dangerous type, because a reader will copy the fix.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stale state.&lt;/strong&gt; Counts that were true once: the number of smart bulbs responding, the number of containers in the estate, the number of agent profiles. One draft still included a service I had decommissioned. Another described logging tables that had been dropped ten days earlier.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mixed-run numbers.&lt;/strong&gt; A post about a price-watching agent's sources combined figures from different runs into one table, called a source that had already been fixed broken, and described sources as shops.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Half-true configuration.&lt;/strong&gt; A Docker post said log size limits were set. The config file did exist, but containers created before it weren't covered.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Claims that were never true.&lt;/strong&gt; A post about my secrets manager said the automated identity had read-only access. It never did. That also turned out to be in a post already live since 6 September.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Security-posture detail.&lt;/strong&gt; Some drafts described access arrangements more specifically than I want public. Those lines were cut rather than corrected.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checking the drafts found two live bugs
&lt;/h2&gt;

&lt;p&gt;The most useful outcome wasn't the corrected posts. To check whether "every agent is traced" was true, I had to look, and it wasn't: when the tracing service moved address, the main configuration was updated but twelve per-profile environment files still pointed at the old one. The specialist agents had been silently untraced.&lt;/p&gt;

&lt;p&gt;Checking a claim about which model the agents use turned up another. A canary that tests the paid cloud model had reverted four profiles to a local model on 7 September, because its probe prompt came back redacted by a privacy filter. It only restores automatically after a credit problem, not after this kind of revert, so they sat on the local model for more than 17 days while its log said there was nothing to watch. Both were fixed the same day: the twelve files repointed, and the canary changed so it re-probes after any revert and moves back when a paid model passes.&lt;/p&gt;

&lt;p&gt;A draft making a confident claim about my own systems is a cheap audit prompt. It is only useful if someone checks the claim.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the 18 September post got wrong
&lt;/h2&gt;

&lt;p&gt;Reading it again:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;em&gt;"The model never went looking for a fact, so it never had the chance to fabricate one."&lt;/em&gt; It doesn't need to go looking. Given a gap, it fills it.&lt;/li&gt;
&lt;li&gt;The leak gate is real and it works, but it is structural. It catches an IP address or a credential shape. It cannot know that "two weeks of debugging" is false.&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;"A ten-second read by the person whose name is on it can"&lt;/em&gt; catch a subtly overstated claim. The endianness bug reads fine. You only catch it by opening the code.&lt;/li&gt;
&lt;li&gt;Topics were "derived from things that demonstrably happened". The empty repo shows a topic can come from a description alone.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That live post about the secrets manager has been corrected in place, with a dated correction note saying what was wrong, rather than edited silently.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I do now
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;"Drafted OK" is not "true".&lt;/strong&gt; Passing the format and leak checks means the file is well formed. It says nothing about the content.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every draft is checked against the code, logs or notes before it is queued&lt;/strong&gt;, claim by claim. For numbers, a live read beats a note, and a note beats a code comment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check the predictable categories first:&lt;/strong&gt; any "I tried X first", any fix, any count, any statement about access or permissions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Correct published posts visibly&lt;/strong&gt;, with a date.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The nine drafts have been fixed and queued one a day. The pipeline still saves me the typing. It does not save me the checking.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;I wrote the full pattern up, with the validators, the canary that lied and a checklist, as a short field report: &lt;a href="https://asareanderson.gumroad.com/l/ktelxr" rel="noopener noreferrer"&gt;Keeping AI Agents Honest&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;🤖 &lt;em&gt;Drafted with AI assistance from my own homelab notes, logs and repos, then reviewed and edited before publishing.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>ai</category>
      <category>writing</category>
      <category>llm</category>
      <category>devto</category>
    </item>
    <item>
      <title>DHCP lease churn in a Proxmox homelab: when Authelia down x87 was a lost IP</title>
      <dc:creator>Christian Anderson</dc:creator>
      <pubDate>Wed, 07 Oct 2026 08:00:14 +0000</pubDate>
      <link>https://dev.to/iam-tech/dhcp-lease-churn-in-a-proxmox-homelab-when-authelia-down-x87-was-a-lost-ip-1nll</link>
      <guid>https://dev.to/iam-tech/dhcp-lease-churn-in-a-proxmox-homelab-when-authelia-down-x87-was-a-lost-ip-1nll</guid>
      <description>&lt;p&gt;On 2 September my self-healing monitor logged &lt;code&gt;DOWN: Authelia&lt;/code&gt; 87 times in a row. Authelia, my SSO portal, was never down. Its own address answered 200 the whole time.&lt;/p&gt;

&lt;p&gt;What had gone was an IP address. This post covers that incident and the others that turned out to be the same bug: a smart bulb that took a server's address, a cloned container sharing a MAC, and a re-IP that switched off my tracing.&lt;/p&gt;

&lt;h2&gt;
  
  
  "It's always been on that address" is not "it's pinned"
&lt;/h2&gt;

&lt;p&gt;Both Proxmox nodes rebooted at about 14:16 that afternoon. The container that runs my reverse proxy came back and DHCP gave it a different address.&lt;/p&gt;

&lt;p&gt;Everything pointed at its usual address: AdGuard rewrites, Cloudflare records, proxy configs, monitors. But there was no DHCP reservation for it on the router. The address was only ever a sticky lease: dnsmasq tends to hand a returning client its old address, until the day it doesn't. The pool had 150 addresses on 12-hour leases.&lt;/p&gt;

&lt;p&gt;That container runs my only Caddy instance, and six hostnames resolve to it, including the SSO portal. One lost lease took out the front door and all SSO while every backend stayed healthy. My monitors probe by hostname, so they blamed the wrong service.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The lesson I wrote down:&lt;/strong&gt; when a monitor says a service on another box is down, probe that service's own address and port before touching it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Making it static, and the lease that wouldn't leave
&lt;/h3&gt;

&lt;p&gt;I set a static address in the container's Proxmox network config with &lt;code&gt;pct set ... -net0 ...,ip=&amp;lt;addr&amp;gt;/24,gw=&amp;lt;gw&amp;gt;&lt;/code&gt;. That rewrites the container's systemd-networkd file to a fixed address with DHCP off, and hot-adds the static address. But the old dynamic address and its &lt;code&gt;proto dhcp&lt;/code&gt; default route survived a &lt;code&gt;systemctl restart systemd-networkd&lt;/code&gt;. I had to delete both by hand with &lt;code&gt;ip addr del&lt;/code&gt; and &lt;code&gt;ip route del&lt;/code&gt;. Until then the container held both addresses.&lt;/p&gt;

&lt;p&gt;After that, SSO answered again and the healer went quiet.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bulb that took a server's address
&lt;/h2&gt;

&lt;p&gt;That evening I swept every service to check it was really back. My password manager's hostname returned &lt;strong&gt;502&lt;/strong&gt;, while its container reported itself healthy. Caddy's log showed its upstream dial failing with &lt;code&gt;connect: connection refused&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The fault was at layer 2. The password manager's container was set to a static address, but that address sat inside the DHCP pool. So the router had leased it to something else, and the MAC answering for it began with an OUI registered to WiZ, the smart bulb maker. Ping times to it were around 15 ms, which is wireless, not a container on the same host.&lt;/p&gt;

&lt;p&gt;The bulb won the ARP race across the whole network, even on the Proxmox host. &lt;strong&gt;A static IP inside a DHCP pool is not a static IP.&lt;/strong&gt; The router doesn't know it's taken and will hand it to the next device that asks.&lt;/p&gt;

&lt;p&gt;The fix was to move the container to a static address &lt;em&gt;below&lt;/em&gt; the start of the pool, where the router can never lease it. Only one live reference needed changing: the &lt;code&gt;reverse_proxy&lt;/code&gt; line in the Caddyfile, then &lt;code&gt;systemctl reload caddy&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The third one, and a false green
&lt;/h2&gt;

&lt;p&gt;Same night, a third service: my 3D printer monitoring app returned 502 while the app was healthy. My NPMplus reverse proxy on another box still had its upstream set to the container's pre-reboot address. I repointed it and reloaded nginx inside the container.&lt;/p&gt;

&lt;p&gt;The monitoring was worse. Uptime Kuma had two monitors for that service, both on the old address:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the HTTP check was red with &lt;code&gt;ECONNREFUSED&lt;/code&gt;: the right alarm for the wrong reason;&lt;/li&gt;
&lt;li&gt;the ping check was &lt;strong&gt;green&lt;/strong&gt;, because an unrelated household device had inherited the old lease and was answering pings.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A stale IP in a ping monitor doesn't go red. It goes green on whatever takes the address. Meanwhile my healer, which only reads Kuma, had started escalating to rebooting a container that was never unhealthy. Only its limit of three reboots an hour held it back.&lt;/p&gt;

&lt;p&gt;Two more lessons came out of this one:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A raw-IP probe could never have worked.&lt;/strong&gt; The app is Django using the sites framework, which picks the site by the request's &lt;code&gt;Host&lt;/code&gt; header. Probed by IP and port, the login page returned 500. With the right &lt;code&gt;Host&lt;/code&gt; header it returned 200. The old monitor had only ever passed because a site record existed for the &lt;em&gt;old&lt;/em&gt; IP and port. I pointed the monitor at the public URL instead, the only probe that would have caught the real user-facing fault in the proxy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A 3xx is not proof of health.&lt;/strong&gt; I called that service "restored" off a 302. The 302 was the redirect to the login page, which was returning 500. Kuma followed the redirect, which is how it caught what my &lt;code&gt;curl&lt;/code&gt; missed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Kuma's history showed the old monitor green until 08:49 that morning, before the 14:16 reboot. That container's lease had turned over mid-morning, not at the reboot at all.&lt;/p&gt;

&lt;p&gt;The tally for one unreserved address, over that night and the next day: two container IPs, a Caddy upstream, two proxy upstream fixes, five references in source files and two monitors, all changed by hand. One of those source files would have quietly recreated the broken raw-IP monitor on its next run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reservations and static config protect different things
&lt;/h2&gt;

&lt;p&gt;On 3 September I added DHCP reservations on the router for all ten containers, then checked each still held its expected address and every front-door hostname answered.&lt;/p&gt;

&lt;p&gt;This is the part I had misunderstood:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Static config in the container&lt;/strong&gt; stops the container asking for a new address.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A reservation on the router&lt;/strong&gt; stops the router giving that address to someone else.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The bulb incident was the missing second half. The rule I now follow: static hosts live below the pool, DHCP keeps the pool, and anything that other config names by address gets a reservation.&lt;/p&gt;

&lt;h2&gt;
  
  
  An older version of the same bug: two containers, one MAC
&lt;/h2&gt;

&lt;p&gt;In August I'd already met the extreme case. On 5 August I migrated my agent container to a newer Proxmox node for more RAM, and never stopped the original. Both were set to start on boot. They were byte-identical clones: same MAC, so the same DHCP lease and the same IP, the same hostname, and the same NetBird identity.&lt;/p&gt;

&lt;p&gt;The switch learns a MAC wherever it last saw a frame, so replies went to whichever twin had spoken last. That explained the "flaky networking" I had been chasing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;SSH banner stalls in roughly 20–30% of connections;&lt;/li&gt;
&lt;li&gt;Uptime Kuma's login acknowledgement never arriving;&lt;/li&gt;
&lt;li&gt;NetBird reconnecting 11,898 times in 24 hours on that one peer, against 2–3 for every other peer, because two clients held one identity and kept evicting each other.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;ARP looked clean, because both twins answered with the &lt;em&gt;same&lt;/em&gt; MAC. I first called it a layer-2 loop, which was wrong. A paired tcpdump on both ends finally showed a reset from our address that our container had never sent. That was the twin.&lt;/p&gt;

&lt;p&gt;I stopped the old container on 9 August. NetBird's reconnect log for that peer dropped from about 24 per three minutes to 1.&lt;/p&gt;

&lt;h2&gt;
  
  
  The re-IP that quietly switched off tracing
&lt;/h2&gt;

&lt;p&gt;The latest instance was this week. My Langfuse container came back from a node reboot on 22 September on a static address that someone or something had set in its config. It was inside the DHCP pool too, and different from its reservation. Uptime Kuma went red with &lt;code&gt;EHOSTUNREACH&lt;/code&gt; for 2,382 heartbeats. Langfuse itself was healthy throughout, but my agents' tracing, pointed at the reserved address, went dark.&lt;/p&gt;

&lt;p&gt;On 23 September I moved the container back to its reserved address rather than repointing everything else, because the reservation is the source of truth and a static address inside the pool is the bulb trap again.&lt;/p&gt;

&lt;p&gt;That wasn't the end. On 22 September every specialist agent profile's config had been repointed to the wrong address, and none of those files went back when the container did. So every specialist agent stayed untraced until I found it on 24 September, repointed them, restarted the gateway and confirmed new traces arriving.&lt;/p&gt;

&lt;p&gt;An IP change has to update every config that names the address, including the per-profile copies, and the monitors.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I check now
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Is this address &lt;strong&gt;reserved&lt;/strong&gt; on the router, or just sticky? "It's always been there" proves nothing.&lt;/li&gt;
&lt;li&gt;Is any static address inside the DHCP pool? Move it below the pool.&lt;/li&gt;
&lt;li&gt;After any IP change: grep configs on every host, including proxy configs inside containers on other boxes, per-profile config files and every monitor.&lt;/li&gt;
&lt;li&gt;Prefer hostname probes that follow redirects over pings to an IP.&lt;/li&gt;
&lt;li&gt;After cloning or migrating a container, make sure the original is actually stopped.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;The longer version of this, with the DNS, DHCP and tunnel traps side by side and a checklist, is a short field report: &lt;a href="https://asareanderson.gumroad.com/l/zdfddu" rel="noopener noreferrer"&gt;Home Network Traps&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;🤖 &lt;em&gt;Drafted with AI assistance from my own homelab notes, logs and repos, then reviewed and edited before publishing.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>networking</category>
      <category>dhcp</category>
      <category>proxmox</category>
      <category>devops</category>
    </item>
    <item>
      <title>Gumroad API pagination and HTTP 200 errors: the store that looked half-empty</title>
      <dc:creator>Christian Anderson</dc:creator>
      <pubDate>Tue, 06 Oct 2026 08:00:13 +0000</pubDate>
      <link>https://dev.to/iam-tech/gumroad-api-pagination-and-http-200-errors-the-store-that-looked-half-empty-143j</link>
      <guid>https://dev.to/iam-tech/gumroad-api-pagination-and-http-200-errors-the-store-that-looked-half-empty-143j</guid>
      <description>&lt;p&gt;I run a small Gumroad store from a set of scripts on my homelab. The scripts create products, set prices and tags, and check what is already listed before they list anything new. Twice this month they gave me a completely wrong picture of that store, and neither time did anything error. Both mistakes came from the same place: the Gumroad API tells you what happened in a way that is easy to misread.&lt;/p&gt;

&lt;p&gt;There are two traps. The HTTP status code tells you almost nothing, and the product list comes back ten at a time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trap one: everything is HTTP 200
&lt;/h2&gt;

&lt;p&gt;In late August I wrote down, with some confidence, that Gumroad's API could not create products. The note said &lt;code&gt;POST /v2/products&lt;/code&gt; returned 404 and that creation was dashboard-only. I built the rest of the tooling around that: a script that packaged the files, then a manual step where I would click through the dashboard.&lt;/p&gt;

&lt;p&gt;On 5 September I ran the full lifecycle against the live account to check, and the note was wrong:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;POST   /v2/products              success=True   (created, unpublished)
PUT    /v2/products/{id}         success=True   (edit)
PUT    /v2/products/{id}/enable  success=True   (published)
DELETE /v2/products/{id}         success=True   (account back to 0 products)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Creation works. What had misled me is what a &lt;em&gt;bare&lt;/em&gt; &lt;code&gt;POST /v2/products&lt;/code&gt; returns. It is not a 404. It is an HTTP &lt;strong&gt;200&lt;/strong&gt; with this body:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"success"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"New products should be created with a price"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Gumroad answers nearly everything with 200 and puts the real outcome in the &lt;code&gt;success&lt;/code&gt; field. A product that does not exist is also a 200:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"success"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"The product was not found."&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So there are two wrong readings available, and I had managed one of them. Treat the 200 as "the endpoint exists and it worked" and you carry on with a failure. Or half-remember a failure as "it 404'd" and you conclude the feature is missing. The status code in this API tells you that you reached Gumroad. It does not tell you whether Gumroad did what you asked.&lt;/p&gt;

&lt;p&gt;Once I read &lt;code&gt;success&lt;/code&gt; instead, the manual step disappeared. The scripts now create, edit, publish and tag products end to end. Tags, for the record, work through &lt;code&gt;PUT /v2/products/{id}&lt;/code&gt; with repeated &lt;code&gt;tags[]&lt;/code&gt; parameters. I applied them and read them back to confirm rather than trusting the response.&lt;/p&gt;

&lt;p&gt;A related one caught me later, when I switched a product to pay-what-you-want. &lt;code&gt;customizable_price&lt;/code&gt; and &lt;code&gt;suggested_price_cents&lt;/code&gt; have to be sent together. Send only one and the product renders as plain free. Again, no error.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trap two: ten products, then &lt;code&gt;next_page_url&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;On 20 September I added a &lt;code&gt;sync&lt;/code&gt; command. It lists and publishes every item in my local catalogue that is not live on Gumroad yet. It already had a guard: before creating anything, check whether a product with that name is already on Gumroad.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;GET /v2/products&lt;/code&gt; returns only the first ten products. The rest sit behind a &lt;code&gt;next_page_url&lt;/code&gt; field in the response. My &lt;code&gt;status()&lt;/code&gt; function and the duplicate check in &lt;code&gt;create()&lt;/code&gt; both read page one and stopped.&lt;/p&gt;

&lt;p&gt;The consequences were all quiet:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The store appeared to have 10 products. It had 22.&lt;/li&gt;
&lt;li&gt;A diff of catalogue against store said 10 catalogue items had never been listed. In fact 14 of the 15 were already live.&lt;/li&gt;
&lt;li&gt;The duplicate guard could not see page two, so the first &lt;code&gt;sync&lt;/code&gt; run created two duplicate products. I deleted both the same night and kept the older originals, which were the ones already linked from elsewhere.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Nothing errored. Nothing was wrong with any single request. The bug was in what I assumed a single request meant.&lt;/p&gt;

&lt;p&gt;The fix was a function that walks the pages:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;urllib.parse&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;urllib.request&lt;/span&gt;

&lt;span class="n"&gt;API&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.gumroad.com/v2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;gumroad_get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;sep&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;&amp;amp;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;url&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;url&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;sep&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;urllib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;urlencode&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;access_token&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;urllib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;urlopen&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="c1"&gt;# HTTP 200 means nothing here. The outcome is in `success`.
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;success&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;RuntimeError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gumroad: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;message&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;body&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;all_products&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_pages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;25&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;pages&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;API&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/products&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[],&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="n"&gt;url&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;pages&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;max_pages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;gumroad_get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;products&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
        &lt;span class="n"&gt;url&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;next_page_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;url&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;startswith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;http&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;url&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.gumroad.com&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;url&lt;/span&gt;
        &lt;span class="n"&gt;pages&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A few details in there matter. The loop accepts &lt;code&gt;next_page_url&lt;/code&gt; as either a full URL or a bare path, so it does not break on whichever form comes back. The token is added with &lt;code&gt;&amp;amp;&lt;/code&gt; when the URL already carries a query string, which the next-page URL does. The page cap of 25 is there so a bad cursor cannot loop forever. And the &lt;code&gt;success&lt;/code&gt; check runs on every page, not just the first.&lt;/p&gt;

&lt;p&gt;After the fix, &lt;code&gt;status()&lt;/code&gt; and the duplicate check both use &lt;code&gt;all_products()&lt;/code&gt;. The note I left in the code says it plainly: anything asking "is this already on Gumroad?" must use this function, never a single call to the products endpoint.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trap two again, four days later
&lt;/h2&gt;

&lt;p&gt;On 24 September I built a separate agent to propose better titles, summaries and tags for the store and my blog posts. It has its own helper module with its own Gumroad client, written from scratch after the 20 September fix.&lt;/p&gt;

&lt;p&gt;It read page one only. The store had 21 products at that point, and the agent could see 10. Eleven were invisible, including several of the guides I had specifically asked it to look at. My own check missed the same eleven.&lt;/p&gt;

&lt;p&gt;The fix was the same loop, in a different file. I then ran a second batch to add summaries to those 11 products and read each one back.&lt;/p&gt;

&lt;p&gt;This is the part I find worth writing down. The 20 September fix was correct and it was tested. It fixed one function in one script. The &lt;em&gt;pattern&lt;/em&gt; — call the list endpoint once and treat the result as the whole store — lived in my head, and it came out again the next time I wrote a Gumroad client from scratch.&lt;/p&gt;

&lt;h2&gt;
  
  
  Grep for the callers, not the bug
&lt;/h2&gt;

&lt;p&gt;While writing this post I grepped my scripts directory for anything that touches &lt;code&gt;v2/products&lt;/code&gt;. Four files came up. Two now paginate. One is the batch script that fixed the missing summaries. The fourth is the script that records distribution numbers per product. It still reads page one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# ⚠️ Gumroad answers EVERYTHING 200 and puts the real outcome in `success`.
# A missing product is 200. A malformed request is 200. Read `success`.
&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;_get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.gumroad.com/v2/products?access_token=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;tok&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It had the first lesson written into a comment, and it still made the second mistake. I fixed it the same afternoon with the same loop, and it now sees all 21 products instead of 10.&lt;/p&gt;

&lt;p&gt;So I now treat an API quirk like this as a search, not a patch:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Read the outcome field, not the status code.&lt;/strong&gt; For Gumroad that means &lt;code&gt;success&lt;/code&gt;. A 200 only means the request arrived.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Assume every list endpoint is paginated until you've proved otherwise.&lt;/strong&gt; Check for a &lt;code&gt;next_page_url&lt;/code&gt;, cursor or page field on the first response, even when you think the list is small. Ten felt like a whole store to me because it was a round, plausible number.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;When you fix it, grep for every caller.&lt;/strong&gt; Search for the endpoint string, not the function name, because a second client written from scratch won't share your function names. Then check each hit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't let your own check share the agent's blind spot.&lt;/strong&gt; Mine missed the same eleven products the agent did, so it agreed with the agent instead of catching it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Count before you create.&lt;/strong&gt; If a script is about to create something because it "isn't there yet", the thing it checked against needs to be complete, or the guard is decoration.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;None of these failures raised an exception. Every request succeeded. The store just looked smaller than it was, and three scripts believed it.&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;🤖 &lt;em&gt;Drafted with AI assistance from my own homelab notes, logs and repos, then reviewed and edited before publishing.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>python</category>
      <category>gumroad</category>
      <category>api</category>
      <category>writing</category>
    </item>
    <item>
      <title>Cloudflare Tunnel 530 errors: the connector that was healthy and did nothing</title>
      <dc:creator>Christian Anderson</dc:creator>
      <pubDate>Mon, 05 Oct 2026 08:00:14 +0000</pubDate>
      <link>https://dev.to/iam-tech/cloudflare-tunnel-530-errors-the-connector-that-was-healthy-and-did-nothing-3817</link>
      <guid>https://dev.to/iam-tech/cloudflare-tunnel-530-errors-the-connector-that-was-healthy-and-did-nothing-3817</guid>
      <description>&lt;p&gt;It started as "Immich has gone down". Immich was fine. The server container was healthy and answered locally. What had gone was every public hostname I run through Cloudflare Tunnel, all at once, and it had been that way for about a day before anyone noticed.&lt;/p&gt;

&lt;p&gt;This is what happened, the mistake that made me declare it fixed when it wasn't, and the checks I now use so the tunnel can't lie to me again.&lt;/p&gt;

&lt;h2&gt;
  
  
  A healthy container connecting nothing
&lt;/h2&gt;

&lt;p&gt;On the evening of 3 September, around 22:00, two containers on my NAS were recreated. Both ran a community web UI wrapper for &lt;code&gt;cloudflared&lt;/code&gt;, the Cloudflare Tunnel client. When they came back up, both logged this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CONFIG: No pre-existing config file found
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That line is the whole incident. The wrapper is a web UI, not a connector. With no config it starts happily, shows as healthy in &lt;code&gt;docker ps&lt;/code&gt;, and connects no tunnel at all. Both of its config directories were empty. From Cloudflare's side the tunnel had zero connectors, so every public hostname behind it returned &lt;strong&gt;530&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That included my single sign-on portal. Anything behind SSO was unreachable from outside, not just the photo library. The photo library was just the first thing to be noticed.&lt;/p&gt;

&lt;p&gt;The container's health check was checking the UI, and the UI was fine. Nothing in &lt;code&gt;docker ps&lt;/code&gt; distinguishes "a web page is being served" from "traffic is flowing through a tunnel".&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: run an actual connector
&lt;/h2&gt;

&lt;p&gt;Instead of trying to rebuild the wrapper's config, I ran the real thing, the official image with the tunnel token:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;--name&lt;/span&gt; cloudflared-connector &lt;span class="nt"&gt;--restart&lt;/span&gt; unless-stopped &lt;span class="se"&gt;\&lt;/span&gt;
  cloudflare/cloudflared:latest tunnel &lt;span class="nt"&gt;--no-autoupdate&lt;/span&gt; run &lt;span class="nt"&gt;--token&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$TOK&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I pulled the token from the Cloudflare API (&lt;code&gt;GET /accounts/{account}/cfd_tunnel/{tunnel}/token&lt;/code&gt;) using an API token already scoped to Tunnel edit and DNS edit. The tunnel then reported &lt;code&gt;healthy&lt;/code&gt; with &lt;code&gt;connections=4&lt;/code&gt;, and all five public hostnames were back.&lt;/p&gt;

&lt;p&gt;A connector-only container has no UI and nothing to forget. The token is the whole configuration, because the routes live in Cloudflare's dashboard. That is exactly what I want from the piece that holds the front door open.&lt;/p&gt;

&lt;p&gt;The two config-less wrapper containers kept running afterwards, doing nothing. They're harmless, but they're a good reminder that "running" and "working" are different claims.&lt;/p&gt;

&lt;h2&gt;
  
  
  The trap: split-horizon DNS makes local tests lie
&lt;/h2&gt;

&lt;p&gt;Here is the part I'm least proud of. During the incident I tested a hostname from inside the network, got a clean redirect back and declared external access working. It was 530 the whole time.&lt;/p&gt;

&lt;p&gt;My router's resolver is AdGuard Home, and it has rewrites for my own domain. Inside the house, a lookup for one of my hostnames returns the private address of my reverse proxy. Outside, public DNS returns Cloudflare's addresses. That is deliberate split-horizon DNS, and it means:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;curl https://&amp;lt;my hostname&amp;gt;/&lt;/code&gt; from any machine at home &lt;strong&gt;never touches Cloudflare&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;It goes straight to the reverse proxy on the LAN, gets a perfect 302 to the login page, and proves nothing about the tunnel.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The tunnel could be completely dead and every local test would still pass.&lt;/p&gt;

&lt;h3&gt;
  
  
  How to test the real path
&lt;/h3&gt;

&lt;p&gt;First, check whether you're in split-horizon territory at all. Compare what a public resolver says with what your own machine says:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dig +short @1.1.1.1 app.example.com   &lt;span class="c"&gt;# what the internet sees&lt;/span&gt;
dig +short app.example.com            &lt;span class="c"&gt;# what this machine sees&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If those answers differ, a plain &lt;code&gt;curl&lt;/code&gt; from this machine is testing your LAN, not your tunnel.&lt;/p&gt;

&lt;p&gt;Then force the request through the Cloudflare edge by pinning the hostname to the public answer, while keeping the right SNI and &lt;code&gt;Host&lt;/code&gt; header:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;EDGE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;dig +short @1.1.1.1 app.example.com | &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-1&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; &lt;span class="s1"&gt;'%{http_code}\n'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--resolve&lt;/span&gt; app.example.com:443:&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$EDGE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; https://app.example.com/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A 530 here means the edge has nowhere to send the request. A 200, 302 or 303 means the full path works: edge, tunnel, connector, origin.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;--resolve&lt;/code&gt; also splits problems the other way. Point it at your reverse proxy's address instead, and it proves the proxy and certificate work independently of DNS. If that passes and the normal request fails, the fault is name resolution, not the service.&lt;/p&gt;

&lt;h3&gt;
  
  
  Ask the client's own resolver
&lt;/h3&gt;

&lt;p&gt;The same mistake has caught me with DNS on the client side too. The week before, a household device had been handed a smart-DNS streaming box as its resolver by a DHCP tag, and that box answers nothing at all for my domain. I spent a session checking AdGuard, Cloudflare, the browser cache and the firewall, and they were all innocent.&lt;/p&gt;

&lt;p&gt;The step that would have found it straight away: query the resolver the client is actually using, the one in its own DHCP lease, not the one you think it should be using. &lt;code&gt;dig @1.1.1.1&lt;/code&gt; and &lt;code&gt;dig @router&lt;/code&gt; tell you about those resolvers, not about the device that's broken.&lt;/p&gt;

&lt;h3&gt;
  
  
  A public record pointing at a private address
&lt;/h3&gt;

&lt;p&gt;One more split-horizon wrinkle turned up the same day. One of my hostnames had an A record in public DNS pointing at a private LAN address. It had only ever worked at home, because at home that address is reachable. From anywhere else it was a dead end. I moved it onto the tunnel, with the tunnel sending it to the reverse proxy and setting the origin server name so the proxy picks the right site block and still applies SSO. Then I verified it the only way that counts: GET and POST from the edge, both redirecting to the login portal.&lt;/p&gt;

&lt;h2&gt;
  
  
  Monitoring that was lying as well
&lt;/h2&gt;

&lt;p&gt;On 13 September I found my Uptime Kuma "Cloudflare" monitor had been red for more than 2,700 heartbeats while the tunnel was fine. It had been polling the wrapper's web UI port, which no longer answered, so it got &lt;code&gt;ECONNREFUSED&lt;/code&gt;. A monitor that's always red gets ignored, which is just as bad as one that's always green.&lt;/p&gt;

&lt;p&gt;I replaced it with a push monitor that tests the path a real visitor takes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a script on a systemd timer runs every five minutes;&lt;/li&gt;
&lt;li&gt;it resolves two of my hostnames through &lt;code&gt;1.1.1.1&lt;/code&gt;, so it gets the public answer, not the split-horizon one;&lt;/li&gt;
&lt;li&gt;it calls each with &lt;code&gt;curl --resolve&lt;/code&gt; to the Cloudflare edge;&lt;/li&gt;
&lt;li&gt;it only reports DOWN if &lt;strong&gt;both&lt;/strong&gt; come back 530 or give no answer, so one flaky app can't page me about the tunnel;&lt;/li&gt;
&lt;li&gt;Kuma expects a push every 600 seconds, so if the script itself dies, that goes red too.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It has a &lt;code&gt;--dry-run&lt;/code&gt; flag so I can see what it would report without pushing anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  The happy ending: a four-minute self-heal
&lt;/h2&gt;

&lt;p&gt;This morning, 24 September, the NAS went down hard. Its journal simply stops mid-stream at 10:13:24 BST, with no shutdown sequence, which is what a power loss or hard reset looks like. It booted again at 10:14.&lt;/p&gt;

&lt;p&gt;At 10:17 the connector container, set to &lt;code&gt;restart unless-stopped&lt;/code&gt;, re-registered all four connections with Cloudflare. Total public outage: about four minutes. I didn't have to do anything.&lt;/p&gt;

&lt;p&gt;Three weeks earlier, a container recreate had cost a day of silent outage. The difference wasn't a clever fix. It was running the real connector instead of a UI wrapper, and testing through the edge instead of through my own DNS.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd tell myself three weeks earlier
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;A healthy container is not a working tunnel. Check the connector count on Cloudflare's side.&lt;/li&gt;
&lt;li&gt;If your LAN resolves your own domain to a private address, every local &lt;code&gt;curl&lt;/code&gt; is testing your LAN. Use &lt;code&gt;curl --resolve&lt;/code&gt; against the public answer.&lt;/li&gt;
&lt;li&gt;Compare &lt;code&gt;dig @1.1.1.1&lt;/code&gt; with your local answer before trusting any test.&lt;/li&gt;
&lt;li&gt;When a client can't resolve something, query that client's own resolver first.&lt;/li&gt;
&lt;li&gt;A monitor that probes the wrong thing is worse than no monitor. Probe the path a user takes.&lt;/li&gt;
&lt;li&gt;Give the connector a restart policy and nothing else to remember. Then a power cut is a four-minute blip, not a day of "Immich is down".&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;The longer version of this, with the DNS, DHCP and tunnel traps side by side and a checklist, is a short field report: &lt;a href="https://asareanderson.gumroad.com/l/zdfddu" rel="noopener noreferrer"&gt;Home Network Traps&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;🤖 &lt;em&gt;Drafted with AI assistance from my own homelab notes, logs and repos, then reviewed and edited before publishing.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>cloudflare</category>
      <category>homelab</category>
      <category>testing</category>
      <category>networking</category>
    </item>
    <item>
      <title>AI agent SOUL.md: writing a soul for my chief of staff, and the rule I broke</title>
      <dc:creator>Christian Anderson</dc:creator>
      <pubDate>Sun, 04 Oct 2026 08:00:20 +0000</pubDate>
      <link>https://dev.to/iam-tech/ai-agent-soulmd-writing-a-soul-for-my-chief-of-staff-and-the-rule-i-broke-25hf</link>
      <guid>https://dev.to/iam-tech/ai-agent-soulmd-writing-a-soul-for-my-chief-of-staff-and-the-rule-i-broke-25hf</guid>
      <description>&lt;p&gt;My homelab agent had a job title and no personality.&lt;/p&gt;

&lt;p&gt;Its identity file opened with "You are Hermes Chief of Staff" and then read like an HR policy: decomposition rules, card templates, guardrails. All useful. None of it told the model &lt;em&gt;who&lt;/em&gt; it was, who it worked for, or what the house actually runs. So it behaved like a very polite ticketing system.&lt;/p&gt;

&lt;p&gt;Three things made me rewrite it in one evening:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It routed work to &lt;strong&gt;5&lt;/strong&gt; specialist profiles. There were &lt;strong&gt;11&lt;/strong&gt;, so six of them were invisible to it.&lt;/li&gt;
&lt;li&gt;I gave it a written brief for a family party playlist — named artists, a running order, a reference DJ mix. It ignored the brief and built its usual weekly mix: sixty tracks, mostly rap, some of them explicit. For a baptism.&lt;/li&gt;
&lt;li&gt;Three "blocked" tickets all said the same thing: &lt;em&gt;no password, can't log in, SSH locked out.&lt;/em&gt; Every one of them was wrong. The access existed. One of the specialist agents had written a made-up vault path into its own notes and had believed it ever since.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is what I changed, what the docs told me to do, where I ignored them, and how the same file ports to Claude Cowork and Grok Bot.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a soul file actually is
&lt;/h2&gt;

&lt;p&gt;I run &lt;a href="https://github.com/NousResearch/hermes-agent" rel="noopener noreferrer"&gt;Hermes Agent&lt;/a&gt; from Nous Research. Its &lt;a href="https://hermes-agent.nousresearch.com/docs/user-guide/features/personality/" rel="noopener noreferrer"&gt;Personality &amp;amp; SOUL.md docs&lt;/a&gt; are short and worth reading in full. The bits that matter:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;SOUL.md&lt;/code&gt; lives in &lt;code&gt;HERMES_HOME&lt;/code&gt; and is &lt;strong&gt;slot #1 of the system prompt&lt;/strong&gt;. It's the first thing the model reads, every turn.&lt;/li&gt;
&lt;li&gt;It's loaded only from &lt;code&gt;HERMES_HOME&lt;/code&gt;, never from whatever directory you happen to be in, so the personality doesn't change per project.&lt;/li&gt;
&lt;li&gt;It's scanned for prompt-injection patterns before it's included. Keep it to persona and voice, not clever meta-instructions.&lt;/li&gt;
&lt;li&gt;Empty or unreadable file → a built-in default identity.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/personality&lt;/code&gt; is a session-level overlay. The soul is the default; the overlay is a mood.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The companion guide, &lt;a href="https://github.com/NousResearch/hermes-agent/blob/main/website/docs/guides/use-soul-with-hermes.md" rel="noopener noreferrer"&gt;Use SOUL with Hermes&lt;/a&gt;, gives a four-part skeleton — &lt;strong&gt;Identity, Style, Avoid, Defaults&lt;/strong&gt; — and one rule I'll come back to:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If it should apply everywhere, put it in SOUL.md; if it only belongs to one project, put it in AGENTS.md.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It also calls out the weak version: files full of project details, or generic filler like "be helpful" and "be clear". My old file was mostly the first kind.&lt;/p&gt;

&lt;h2&gt;
  
  
  Give it a name that says what it does
&lt;/h2&gt;

&lt;p&gt;The agent is now &lt;strong&gt;Johnny Decoder&lt;/strong&gt;. Silly, memorable, and it describes the job: take the noise and decode it into what matters. The platform is still Hermes; the character isn't.&lt;/p&gt;

&lt;p&gt;(Update: on 24 September it was renamed again, to KMDENZEL, "Denzel" for short, a CIA-style cryptonym. This post keeps the name it had when I wrote it.)&lt;/p&gt;

&lt;p&gt;The identity section went from one line to this (trimmed):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Who you are&lt;/span&gt;
You're a London-based chief of staff with a sound system in the back office.
Calm under pressure, dry sense of humour, allergic to waffle. British spelling,
contractions, plain words.

Although you're chief of staff, you're an all-rounder. You live on Discord and
know it inside out. You're deeply technical: ten years in tech research, you read
code fluently. You still hand the hands-on work to the specialists — knowing how
is what makes you good at checking their work.

You'd rather say "I don't know" than bluff. When you've messed up, you say so
first and fix it second.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things I learned writing it:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Taste is a behaviour, not decoration.&lt;/strong&gt; "You know the difference between Afrobeat and Afrobeats, dancehall and bashment" does more work than "be knowledgeable about music". It gives the model something to be &lt;em&gt;right&lt;/em&gt; about.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Say how it should fail.&lt;/strong&gt; "When you've messed up, say so first" is there because the old version buried its mistakes in paragraph three of a status report.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Tell it what the house runs
&lt;/h2&gt;

&lt;p&gt;The old soul never said what the agent was responsible for. The new one has a table: homelab, security, trading, writing, music, style and deals, the separate eBay bot, home lighting, travel, licences, tool research, income — and who owns each.&lt;/p&gt;

&lt;p&gt;Then a routing table with all eleven specialists. This one matters more than it looks: in my setup the Kanban dispatcher silently skips an assignee name that doesn't exist, so a ticket for a made-up agent sits in &lt;code&gt;ready&lt;/code&gt; forever. The table is the only list it's allowed to route from.&lt;/p&gt;

&lt;p&gt;And the rule that came straight out of the false tickets:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;A blocker that says "no password", "needs Infisical login" or "SSH locked out"
has been wrong every time so far. The access usually exists. Question that kind
of block before you pass it to the owner.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A soul file is a good place for scar tissue. Just keep it to one line per scar.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule I broke
&lt;/h2&gt;

&lt;p&gt;Look back at that docs quote. &lt;strong&gt;Project details belong in AGENTS.md, not the soul.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;My soul now contains a channel map (which Discord channel gets which alerts), the paths of the knowledge files it should read, and the specialist routing table. By the docs' definition that's project detail. I put it there anyway because this agent doesn't &lt;em&gt;have&lt;/em&gt; a project directory — it lives in a Discord gateway, and the soul is the one file guaranteed to be in every turn.&lt;/p&gt;

&lt;p&gt;It has a cost. The file went from about 8.6 KB to 15 KB, and on a local 6 GB GPU the real latency cost of a turn is prompt processing, not generation. Every KB is paid on every message.&lt;/p&gt;

&lt;p&gt;What I'd do next time: keep the soul to identity, style, avoid-list and defaults, and put the house map somewhere the agent reads on demand. If your tool supports &lt;a href="https://agents.md" rel="noopener noreferrer"&gt;AGENTS.md&lt;/a&gt;, use it. I'm treating my version as a known debt, not a pattern to copy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Memory is the other half of the soul
&lt;/h2&gt;

&lt;p&gt;A soul says who the agent is. Memory says what it has learned. Hermes has a &lt;a href="https://hermes-agent.nousresearch.com/docs/user-guide/features/memory" rel="noopener noreferrer"&gt;built-in memory system&lt;/a&gt;: a &lt;code&gt;USER.md&lt;/code&gt; about you, a &lt;code&gt;MEMORY.md&lt;/code&gt; of the agent's own notes, both injected every turn with a size cap, plus an optional external provider. I also run Honcho, self-hosted, for cross-session user modelling.&lt;/p&gt;

&lt;p&gt;When I checked, the external provider was fine and the agent's own &lt;code&gt;MEMORY.md&lt;/code&gt; &lt;strong&gt;didn't exist&lt;/strong&gt;. Weeks of work, nothing written down. Meanwhile a specialist profile's notes held the invented vault path that caused every false "no password" ticket.&lt;/p&gt;

&lt;p&gt;So the soul now has a memory section. The short version:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Store&lt;/th&gt;
&lt;th&gt;What goes in it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;USER.md&lt;/td&gt;
&lt;td&gt;Durable facts about the owner: tastes, preferences, how he likes things done&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MEMORY.md&lt;/td&gt;
&lt;td&gt;Durable lessons about the estate: what broke, what the real fix was&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Skills&lt;/td&gt;
&lt;td&gt;Procedures that worked and will be needed again&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Honcho&lt;/td&gt;
&lt;td&gt;Runs in the background; the agent doesn't manage it&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;And the rules:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Only save what you verified&lt;/strong&gt; — from a tool, or said by the owner. A guess written to memory becomes a fact for every later session.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Date anything that can go stale.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;After a mistake, save the lesson in one line&lt;/strong&gt;: what you assumed, what was true, how to check next time.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Fix memory the moment it's wrong.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Never store secrets.&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;"Self-evolving" sounds grand. In practice it's those five rules plus a weekly job that consolidates skills, and a daily IT check that now flags an empty or stale memory file. The agent that should have noticed its own memory was empty is now told to look.&lt;/p&gt;

&lt;h2&gt;
  
  
  Write a bio on yourself first
&lt;/h2&gt;

&lt;p&gt;The most useful hour wasn't writing the agent's personality. It was writing mine.&lt;/p&gt;

&lt;p&gt;The agent's &lt;code&gt;USER.md&lt;/code&gt; had one paragraph, and half of it was wrong. I replaced it with what a sharp new colleague would need on day one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="p"&gt;-&lt;/span&gt; Lives in London — all times Europe/London.
&lt;span class="p"&gt;-&lt;/span&gt; Dad; the music account is shared with the kids, so "top tracks" is nursery rhymes.
&lt;span class="p"&gt;-&lt;/span&gt; Music: 90s hip hop, UK rap, R&amp;amp;B, afrobeats, dancehall, lovers rock, UK garage.
  A brief he gives beats any default.
&lt;span class="p"&gt;-&lt;/span&gt; Films: comedy, thrillers, true crime, murder mysteries.
&lt;span class="p"&gt;-&lt;/span&gt; Rules: never guess; never spend or send without a yes; work email never
  goes near the personal GitHub.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last music line is the baptism fix. None of this is personality — it's context — but without it the personality has nothing to be personal &lt;em&gt;about&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;If you do one thing after reading this, write that file. It ports everywhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  Porting the soul to Claude Cowork and Grok Bot
&lt;/h2&gt;

&lt;p&gt;The same split — &lt;strong&gt;identity and rules&lt;/strong&gt; that always apply, &lt;strong&gt;context&lt;/strong&gt; that depends on the job, &lt;strong&gt;memory&lt;/strong&gt; that grows — maps onto the other agent products surprisingly cleanly.&lt;/p&gt;

&lt;h3&gt;
  
  
  Claude Cowork
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://support.claude.com/en/articles/13345190-get-started-with-claude-cowork" rel="noopener noreferrer"&gt;Cowork&lt;/a&gt; is Anthropic's agent mode in the Claude apps, on paid plans. It has two instruction layers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Global instructions&lt;/strong&gt; (Settings → Cowork, or Settings → General in the newer app) apply to every session. That's your soul: name, voice, avoid-list, hard rules like "never delete without approval".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Folder instructions&lt;/strong&gt; attach to a local folder you give it, and Claude can update them as it works. That's your AGENTS.md: the project details my soul shouldn't contain.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The community pattern that works well is an "About me" folder with three files — who you are, your voice (with a few real writing samples), and your rules. There's a template set in &lt;a href="https://github.com/davila7/claude-cowork-guide" rel="noopener noreferrer"&gt;davila7/claude-cowork-guide&lt;/a&gt;. My &lt;code&gt;USER.md&lt;/code&gt; drops straight into the first file.&lt;/p&gt;

&lt;h3&gt;
  
  
  Grok Bot
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://x.ai/news/introducing-grok-bot" rel="noopener noreferrer"&gt;Grok Bot&lt;/a&gt; is SpaceXAI's always-on agent product, launched 11 August 2026. Per the &lt;a href="https://docs.x.ai/grok-bot/bots" rel="noopener noreferrer"&gt;bot docs&lt;/a&gt;, each bot has a profile — &lt;strong&gt;Bot actions → Edit Profile&lt;/strong&gt; — with a name, title, description and avatar. The docs are explicit about what goes where:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The &lt;strong&gt;description&lt;/strong&gt; is for "rules that should remain true", and you update it when you discover "a durable preference, boundary, or responsibility".&lt;/li&gt;
&lt;li&gt;The &lt;strong&gt;conversation&lt;/strong&gt; is for task-specific instructions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory&lt;/strong&gt; builds up as the bot works: preferences, facts, summaries.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's the soul/AGENTS/memory split again, just with different labels. The name and title fields are where "Johnny Decoder, Chief of Staff" goes; the description takes the identity paragraph, the avoid-list and the hard rules. The docs don't state a length limit, so I'd still keep it to the stable parts.&lt;/p&gt;

&lt;h3&gt;
  
  
  Claude Code
&lt;/h3&gt;

&lt;p&gt;For coding agents, &lt;a href="https://docs.claude.com/en/docs/claude-code/memory" rel="noopener noreferrer"&gt;Claude Code's memory docs&lt;/a&gt; cover &lt;code&gt;CLAUDE.md&lt;/code&gt; files at user and project level — the same global-versus-project split. And if you're building on the API, Anthropic's guide to &lt;a href="https://docs.claude.com/en/docs/build-with-claude/prompt-engineering/system-prompts" rel="noopener noreferrer"&gt;giving Claude a role with a system prompt&lt;/a&gt; is the underlying technique all of this rests on.&lt;/p&gt;

&lt;h3&gt;
  
  
  The mapping
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Hermes&lt;/th&gt;
&lt;th&gt;Claude Cowork&lt;/th&gt;
&lt;th&gt;Grok Bot&lt;/th&gt;
&lt;th&gt;Claude Code&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Identity, voice, hard rules&lt;/td&gt;
&lt;td&gt;&lt;code&gt;SOUL.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Global instructions&lt;/td&gt;
&lt;td&gt;Profile name, title, description&lt;/td&gt;
&lt;td&gt;User &lt;code&gt;CLAUDE.md&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Project context&lt;/td&gt;
&lt;td&gt;&lt;code&gt;AGENTS.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Folder instructions + files&lt;/td&gt;
&lt;td&gt;The conversation&lt;/td&gt;
&lt;td&gt;Project &lt;code&gt;CLAUDE.md&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;About you&lt;/td&gt;
&lt;td&gt;&lt;code&gt;USER.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;"About me" files&lt;/td&gt;
&lt;td&gt;Learned memory&lt;/td&gt;
&lt;td&gt;Memory files&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What it learned&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;MEMORY.md&lt;/code&gt;, skills&lt;/td&gt;
&lt;td&gt;Chat memory&lt;/td&gt;
&lt;td&gt;Learned memory&lt;/td&gt;
&lt;td&gt;Memory files&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Write the soul once, in plain Markdown, and you can paste the top half into any of them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reading list
&lt;/h2&gt;

&lt;p&gt;Docs I actually used:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://hermes-agent.nousresearch.com/docs/user-guide/features/personality/" rel="noopener noreferrer"&gt;Hermes: Personality &amp;amp; SOUL.md&lt;/a&gt; and &lt;a href="https://github.com/NousResearch/hermes-agent/blob/main/website/docs/guides/use-soul-with-hermes.md" rel="noopener noreferrer"&gt;Use SOUL with Hermes&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href="https://hermes-agent.nousresearch.com/docs/user-guide/features/memory" rel="noopener noreferrer"&gt;Hermes: memory&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.anthropic.com/research/claude-character" rel="noopener noreferrer"&gt;Anthropic: Claude's character&lt;/a&gt; — how Anthropic thinks about personality as a set of traits, not a costume&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.claude.com/en/docs/build-with-claude/prompt-engineering/system-prompts" rel="noopener noreferrer"&gt;Anthropic: system prompts&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://support.claude.com/en/articles/13345190-get-started-with-claude-cowork" rel="noopener noreferrer"&gt;Claude Cowork: getting started&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.x.ai/grok-bot/bots" rel="noopener noreferrer"&gt;Grok Bot: create and manage bots&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://agents.md" rel="noopener noreferrer"&gt;AGENTS.md&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;GitHub repos worth a look (all MIT unless noted, checked September 2026):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://github.com/NousResearch/hermes-agent" rel="noopener noreferrer"&gt;NousResearch/hermes-agent&lt;/a&gt; — the agent itself.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/aeonfun/soul.md" rel="noopener noreferrer"&gt;aeonfun/soul.md&lt;/a&gt; — builds a soul from your own data with Claude Code or OpenClaw; its &lt;code&gt;SOUL.template.md&lt;/code&gt; is a good blank page.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/thedaviddias/souls-directory" rel="noopener noreferrer"&gt;thedaviddias/souls-directory&lt;/a&gt; — a directory of ready-made soul files.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/bchop-studio/soul-templates" rel="noopener noreferrer"&gt;bchop-studio/soul-templates&lt;/a&gt; — eight small personas; useful for seeing how little text a strong voice needs.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/davila7/claude-cowork-guide" rel="noopener noreferrer"&gt;davila7/claude-cowork-guide&lt;/a&gt; — folder-instruction templates for Cowork (no licence file at the time of writing).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Read the templates for voice. Don't copy one wholesale: a borrowed personality with none of your context is still the generic assistant, with a funnier opening line.&lt;/p&gt;

&lt;h2&gt;
  
  
  The checklist I'd use next time
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Write your own bio first: time zone, household, tastes, hard rules.&lt;/li&gt;
&lt;li&gt;Give the agent a name that describes the job.&lt;/li&gt;
&lt;li&gt;Identity, Style, Avoid, Defaults. Specific traits, not "be helpful".&lt;/li&gt;
&lt;li&gt;One line per scar: the mistakes it keeps making, and how to check.&lt;/li&gt;
&lt;li&gt;Keep project detail out of the soul, unless you've decided to pay for it on every turn.&lt;/li&gt;
&lt;li&gt;Add memory rules: verified only, dated, corrected, no secrets.&lt;/li&gt;
&lt;li&gt;Check it's being used. An empty memory file is a silent failure.&lt;/li&gt;
&lt;li&gt;After changing it, start a fresh session. Gateways often cache the system prompt, and the old personality will keep answering until you do.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The rewrite took an evening. Most of that time went on the bio and the memory rules, not the personality — which is roughly the opposite of what I expected going in.&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;🤖 &lt;em&gt;Drafted with AI assistance from my own homelab notes, logs and repos, then reviewed and edited before publishing.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>llm</category>
      <category>selfhosted</category>
    </item>
    <item>
      <title>Homelab census: 41 containers, one 6 GB GPU, and where my AI agents run</title>
      <dc:creator>Christian Anderson</dc:creator>
      <pubDate>Sat, 03 Oct 2026 08:00:14 +0000</pubDate>
      <link>https://dev.to/iam-tech/homelab-census-41-containers-one-6-gb-gpu-and-where-my-ai-agents-run-1gec</link>
      <guid>https://dev.to/iam-tech/homelab-census-41-containers-one-6-gb-gpu-and-where-my-ai-agents-run-1gec</guid>
      <description>&lt;p&gt;Today I asked my LLM box what it was doing, and it told me it had a 27-billion-parameter model loaded. Total footprint 18.3 GB. Amount of that on the graphics card: 0.5 GB.&lt;/p&gt;

&lt;p&gt;So the 27B was "running on the GPU" in the same sense that I am running a marathon when I walk to the shop. The other 17.8 GB sat in system RAM and did its arithmetic on the CPU, one patient token at a time.&lt;/p&gt;

&lt;p&gt;That seemed like a good moment for a proper census: what is actually running, on what, and which AI jobs go to local hardware, a paid cloud API or a flat subscription. The short version: 41 Docker containers and 9 LXC containers across four boxes, and a single 6 GB card makes most of the routing decisions for me.&lt;/p&gt;

&lt;p&gt;(Background, already written up: the &lt;a href="https://dev.to/c1-anderson/i-run-my-homelab-like-a-small-company-the-architecture-the-org-chart-and-the-rules-i-only-952"&gt;small-company org chart&lt;/a&gt;, the &lt;a href="https://dev.to/c1-anderson/one-ai-six-jobs-how-i-take-an-idea-from-a-one-line-thought-to-something-live-with-a-team-of-50bj"&gt;six-job agent team&lt;/a&gt; and &lt;a href="https://dev.to/c1-anderson/why-i-moved-my-homelabs-coding-agent-from-claude-to-opencode-openrouter-and-what-it-actually-3133"&gt;why the coding agent moved to OpenCode and OpenRouter&lt;/a&gt;.)&lt;/p&gt;

&lt;h2&gt;
  
  
  The census
&lt;/h2&gt;

&lt;p&gt;Measured on 24 September.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Box&lt;/th&gt;
&lt;th&gt;Hardware&lt;/th&gt;
&lt;th&gt;What it runs&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Proxmox node 1&lt;/td&gt;
&lt;td&gt;4 cores, ~8 GB RAM&lt;/td&gt;
&lt;td&gt;5 LXC containers: three single-purpose service boxes, a Docker host (12 containers) and a Docker "proxy node" (4 containers)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Proxmox node 2&lt;/td&gt;
&lt;td&gt;4 cores, 16 GB RAM&lt;/td&gt;
&lt;td&gt;4 LXC containers: the Hermes agent box (native, no Docker), Obico (4 containers), Vaultwarden (1), Langfuse (6)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ZimaBlade&lt;/td&gt;
&lt;td&gt;2 cores, 16 GB RAM, fanless&lt;/td&gt;
&lt;td&gt;14 Docker containers on ZimaOS&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LLM box&lt;/td&gt;
&lt;td&gt;i7-10875H (16 threads), 31 GB RAM, RTX 2060 with 6 GB VRAM&lt;/td&gt;
&lt;td&gt;Ollama with 21 model tags, and ComfyUI sharing the same GPU. No containers.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The arithmetic, since I got it wrong the first time: node 1 has 12 + 4 = &lt;strong&gt;16&lt;/strong&gt; Docker containers, node 2 has 4 + 1 + 6 = &lt;strong&gt;11&lt;/strong&gt;, so &lt;strong&gt;27&lt;/strong&gt; on the Proxmox side, plus &lt;strong&gt;14&lt;/strong&gt; on the ZimaBlade. That's &lt;strong&gt;41 Docker containers&lt;/strong&gt; and &lt;strong&gt;9 LXC containers&lt;/strong&gt;, plus three native services: Ollama, ComfyUI and Hermes itself.&lt;/p&gt;

&lt;p&gt;All of that sits on 26 CPU threads and 71 GB of RAM, if you add up four boxes that were never designed to be added up.&lt;/p&gt;

&lt;h3&gt;
  
  
  What's in the Docker hosts
&lt;/h3&gt;

&lt;p&gt;The node 1 Docker host is the toolbox: crawl4ai for page fetching, SearXNG for search, ESPHome, Wyoming Whisper and Piper for speech-to-text and text-to-speech, a small web app plus its public demo, and Honcho, which is four containers on its own (API, database, deriver, Redis). The proxy node carries Home Assistant, Mosquitto, and the NetBird server and dashboard.&lt;/p&gt;

&lt;p&gt;On node 2, Langfuse, which traces the agent's model calls, is six containers (ClickHouse, web, worker, MinIO, Postgres, Redis). That's more containers than the thing it observes.&lt;/p&gt;

&lt;p&gt;The ZimaBlade is the photo library and the networking plumbing: Immich (server, machine learning, Postgres, Redis), the tunnel connector and reverse proxy, a Homepage dashboard, and omp, a terminal coding agent that points back at the LLM box, plus a few small utilities.&lt;/p&gt;

&lt;h3&gt;
  
  
  Two things the census doesn't show
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Node 2's guests all live on a single 5400rpm hard drive.&lt;/strong&gt; The agent moved there for the extra RAM, and it was the right call, but every reboot is a queue of cold-starting databases fighting over one spindle. The first time it happened the agent's dashboard answered 502 for about 15 minutes. The fix was stopping Langfuse, the heaviest starter and the one thing nobody misses for ten minutes. The dashboard came up 21 seconds later.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The move itself left a twin behind.&lt;/strong&gt; The original agent box on node 1 was never stopped: same MAC, same IP, same VPN identity. What I called "flaky networking" for days was two containers taking turns answering for one address. A census would have caught it on day one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The AI side, and the card that decides it
&lt;/h2&gt;

&lt;p&gt;Here's the rule the whole routing falls out of: &lt;strong&gt;the agent framework I run refuses any model with less than a 64K context window&lt;/strong&gt;, and the LLM box has 6 GB of VRAM.&lt;/p&gt;

&lt;p&gt;Those two facts don't get along. The best local model I've measured for the job is &lt;code&gt;qwen3.5&lt;/code&gt; 9B, built as a 64K variant. At 64K context it's 7.66 GB in total, of which 3.93 GB lands on the card. That's 51% on the GPU and the rest on the CPU, and no amount of context tuning fixes it: at 8K context it's still only 64% on the GPU, because Ollama pins roughly the same slice of VRAM and spills the remainder.&lt;/p&gt;

&lt;p&gt;I spent much of August trying to shop my way out of that. Every candidate was measured and rejected:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Smaller &lt;code&gt;qwen2.5&lt;/code&gt; builds&lt;/strong&gt; fit the card much better, but their context is architecturally 32K. Ollama silently clamps a 64K setting back down, so the framework rejects them outright.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The 27B&lt;/strong&gt; loaded at 4096 context when I benchmarked it, still far below the 64K floor, and in a like-for-like test prefilled at 132 tokens a second against 604 for the 9B, and generated at 3.0 against 10.7. A typical agent turn would have spent over two minutes before producing its first token.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;qwen2.5-coder&lt;/code&gt;&lt;/strong&gt;, the obvious pick for a coding agent, can't tool-call. It writes the JSON for the tool call into its reply as prose, and the agent prints it and stops. Its 16K and 32K tags are still on disk, a small memorial to a good idea.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the local model is genuinely the best option on this hardware, and a real agent turn on it still took around a minute and a half when I timed it in August. The same work on a cheap cloud model answers a tool-call probe in 1.5 to 2 seconds.&lt;/p&gt;

&lt;p&gt;That gap decides most of the routing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who goes where
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Job&lt;/th&gt;
&lt;th&gt;Runs on&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Coding agent&lt;/td&gt;
&lt;td&gt;OpenCode Go subscription (&lt;code&gt;deepseek-v4-flash&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Flat $10 a month. Usage caps, not a meter&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chief of staff, morning brief, research, writing, travel&lt;/td&gt;
&lt;td&gt;OpenRouter &lt;code&gt;qwen/qwen3.7-flash&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Pay per token, about $1.39 a month at measured volume&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Security audit and code review&lt;/td&gt;
&lt;td&gt;Local Ollama only&lt;/td&gt;
&lt;td&gt;Deliberately no third-party egress&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Embeddings, session titles, compression&lt;/td&gt;
&lt;td&gt;Local Ollama&lt;/td&gt;
&lt;td&gt;The framework's auxiliary calls stay local even when the main model is remote&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Honcho user modelling&lt;/td&gt;
&lt;td&gt;Local Ollama&lt;/td&gt;
&lt;td&gt;Background work where latency doesn't matter&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Images&lt;/td&gt;
&lt;td&gt;ComfyUI (SD1.5) on the same card&lt;/td&gt;
&lt;td&gt;Free, and small enough to share&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Speech&lt;/td&gt;
&lt;td&gt;Whisper and Piper containers, CPU&lt;/td&gt;
&lt;td&gt;Never touches the GPU at all&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Last resort&lt;/td&gt;
&lt;td&gt;An OpenRouter free model&lt;/td&gt;
&lt;td&gt;Rate-limited, so it only ever goes last&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Paid cloud: OpenRouter
&lt;/h3&gt;

&lt;p&gt;The cloud model became the primary for the general profiles in late August. The number that justified it came from measuring 30 days of real usage: 44.35M input tokens against 421K output. Input outweighs output by about 105 to 1, which means &lt;strong&gt;input price sets the bill&lt;/strong&gt;, and headline "per million" pricing is close to useless for ranking models. On a model at $0.03 per million input tokens, that volume comes to roughly $1.39 a month.&lt;/p&gt;

&lt;p&gt;The chain is vendor-diverse on purpose: &lt;code&gt;qwen/qwen3.7-flash&lt;/code&gt;, then &lt;code&gt;deepseek/deepseek-v4-flash-0731&lt;/code&gt; as a second paid model from a different vendor, then the local 64K model, then a free model last. A fallback that shares its primary's failure mode isn't a fallback.&lt;/p&gt;

&lt;p&gt;A watchdog probes the primary every three hours for two failure modes: the model dying, and the money running out. A healthy model you can't pay for is still an outage, so below a credit floor it moves everything back to local and says so.&lt;/p&gt;

&lt;p&gt;It also got something badly wrong. On 7 September every paid model "failed" its health check at once, so the watchdog moved four agents (chief of staff, research, writing and travel) back to the local model. Nothing had died: an upstream PII-redaction filter had turned the city in the probe prompt into &lt;code&gt;[ADDRESS]&lt;/code&gt;, so no model could answer it properly. Worse, the watchdog only restored automatically after a credit revert, not a failure revert, so those four agents stayed on the slow local model for 17 days until I caught it on 24 September. It now re-probes after any failover and moves back when a model passes. A probe that fails the same way on every vendor is a problem with the probe, not the models.&lt;/p&gt;

&lt;h3&gt;
  
  
  Subscription: OpenCode Go
&lt;/h3&gt;

&lt;p&gt;The coding agent is the heaviest user by a distance, so it's on OpenCode's Go plan instead: $10 a month, with rolling caps of $12 per five hours, $30 a week and $60 a month. Every response reports a cost of zero. On a subscription the risk isn't a bill, it's hitting a cap halfway through a task, so the watchdog reads Go's usage endpoint and tells a cap from a death. A capped profile gets parked on OpenRouter and put back automatically when the window resets.&lt;/p&gt;

&lt;p&gt;One trap worth passing on: Go lives at &lt;code&gt;/zen/go/v1&lt;/code&gt;, not &lt;code&gt;/zen/v1&lt;/code&gt;. The second is pay-as-you-go and returns "insufficient balance" on a perfectly good Go key, which cost me an hour and a wrong conclusion.&lt;/p&gt;

&lt;h3&gt;
  
  
  Local: Ollama
&lt;/h3&gt;

&lt;p&gt;Local isn't the cheap option here, it's the private one. The security and review agents see the most sensitive estate detail and never leave the LAN. Everything else local is background work, or the fallback for when the internet or a credit balance lets me down.&lt;/p&gt;

&lt;p&gt;Running it taught me more than the cloud side did:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A hang never triggers a fallback.&lt;/strong&gt; Something on the network kept asking Ollama to keep a model loaded forever. With only 6 GB, the scheduler waits for room rather than refusing, so every other model request just sat there, and Ollama logged nothing. Fallback chains fire on errors. This wasn't an error, it was silence, and every local-model profile was dead with nothing reporting it. A small script now unloads anything pinned for more than 48 hours, every 15 minutes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The prompt was the bottleneck, not the model.&lt;/strong&gt; A captured agent turn was about 18,400 prompt tokens, 74% of them tool definitions. Turning off toolsets the specialist agents had never once called cut their tool schemas by 47%. That was worth more than any model swap available.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure what's on the wire.&lt;/strong&gt; For a while every benchmark I ran through the framework's one-shot mode was invalid, because that path never sent "thinking off". The real workloads were fine. My test harness wasn't.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What I'd tell someone starting
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Count first.&lt;/strong&gt; Take the census from the hosts, not from memory. Mine turned up a twin container and a tracing stack bigger than the thing it traces.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Let VRAM make the first cut.&lt;/strong&gt; Write down the smallest context your agent needs, then check what fraction of a model at that context actually lands on your card. If the answer is half, you have your routing decision.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rank cloud models on input price.&lt;/strong&gt; Agent workloads are overwhelmingly prompt. Output price is mostly a distraction.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep local for the jobs that must stay local&lt;/strong&gt;, and as the fallback. Don't make it carry interactive work it can't do quickly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Assume silence is a failure.&lt;/strong&gt; A pinned model, an empty ledger, a probe that swallows exceptions: none of them raise an error. Build the check that notices nothing happening.&lt;/p&gt;

&lt;p&gt;A 27B model on 0.5 GB of VRAM is a fun number. It's also a fair summary of the whole exercise: the hardware will happily let you do something slow and call it running. The census is how you find out.&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;🤖 &lt;em&gt;Drafted with AI assistance from my own homelab notes, logs and repos, then reviewed and edited before publishing.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>ai</category>
      <category>docker</category>
      <category>homelab</category>
      <category>selfhosted</category>
    </item>
    <item>
      <title>Infisical vs 1Password for homelab machine secrets: why I self-host</title>
      <dc:creator>Christian Anderson</dc:creator>
      <pubDate>Fri, 02 Oct 2026 08:00:20 +0000</pubDate>
      <link>https://dev.to/iam-tech/infisical-vs-1password-for-homelab-machine-secrets-why-i-self-host-1m45</link>
      <guid>https://dev.to/iam-tech/infisical-vs-1password-for-homelab-machine-secrets-why-i-self-host-1m45</guid>
      <description>&lt;p&gt;Two scripts on my agent box both needed the password for my uptime monitor. Each kept its own copy. When I finally tested them side by side, one logged in and saw every monitor, and the other had been failing authentication for who knows how long:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;script A   LOGIN OK
script B   FAILED: authIncorrectCreds
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nothing had alerted, because nothing compared the two. Two copies of a secret turn out to be one copy and one liability, and nothing tells you which is which.&lt;/p&gt;

&lt;p&gt;The same afternoon I found a third script that rotated its own machine credential by rewriting its own source file. That file sat in a directory a nightly backup pushed to my git server. So the rotation worked beautifully, and every rotated credential ended up in version control.&lt;/p&gt;

&lt;p&gt;Both are fixed now. This post is the decision behind it: why the machine secrets live in a self-hosted &lt;a href="https://infisical.com/" rel="noopener noreferrer"&gt;Infisical&lt;/a&gt; rather than 1Password, which I'd otherwise happily recommend. My reasons fit on one line: it's hosted locally, I know where the cryptographic keys live, and I can manage all of it from the CLI.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reason one: it runs on my LAN
&lt;/h2&gt;

&lt;p&gt;Infisical runs as the official native package, the omnibus build supervised by systemd, inside its own small Proxmox container. Not Docker, not the cloud.&lt;/p&gt;

&lt;p&gt;Infisical's own docs pitch self-hosting as keeping "your data on your own infrastructure and network", and for a homelab that's the whole appeal. The things fetching secrets are cron jobs, backup scripts, a monitoring daemon and a handful of agents, all on the same network. A secret request doesn't need to leave the house, and when my internet connection has a bad evening (it does), secret fetches don't care.&lt;/p&gt;

&lt;p&gt;Local has a limit, though. My monitoring daemon deliberately does &lt;em&gt;not&lt;/em&gt; depend on Infisical being reachable at the moment it runs, because a daemon whose job is noticing the network is down shouldn't need a network round-trip to find out. Self-hosting shortens the dependency chain but it doesn't remove it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reason two: I know where the keys live
&lt;/h2&gt;

&lt;p&gt;This is where I want to be fair, because 1Password's cryptography is excellent and I'm not claiming otherwise.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://agilebits.github.io/security-design/key-security-features.html" rel="noopener noreferrer"&gt;1Password's security design white paper&lt;/a&gt; describes true end-to-end encryption: keys are generated on your devices, all encryption happens locally, and a "two-secret key derivation" mixes your account password with a locally held Secret Key so that what they store can't be used to crack your password. Vault items &lt;a href="https://agilebits.github.io/security-design/secureItems.html" rel="noopener noreferrer"&gt;use AES-256-GCM&lt;/a&gt;. The company is never in a position to learn your keys. That's a strong design.&lt;/p&gt;

&lt;p&gt;Infisical's model is different. &lt;a href="https://infisical.com/docs/internals/security" rel="noopener noreferrer"&gt;Its security page&lt;/a&gt; says secrets are encrypted at rest with AES-256-GCM under a layered key hierarchy: an operator-supplied root encryption key (passed in as an environment variable) encrypts an internal KMS root key, which in turn encrypts per-organisation and per-project data keys. Their stated goal is that "compromising the database alone is insufficient to decrypt sensitive data". The &lt;a href="https://infisical.com/docs/self-hosting/configuration/envars" rel="noopener noreferrer"&gt;self-hosting configuration docs&lt;/a&gt; make &lt;code&gt;ENCRYPTION_KEY&lt;/code&gt; a required setting.&lt;/p&gt;

&lt;p&gt;So the question for me was never whose maths is stronger. It was &lt;strong&gt;where the ciphertext lives and who holds the key at the top&lt;/strong&gt;. With 1Password the ciphertext sits on their servers and the unlocking secrets sit on my devices. With self-hosted Infisical, the database, the server and the root key all sit on hardware I can touch. For machine secrets, read by unattended processes at three in the morning, I prefer the arrangement where I run the whole chain.&lt;/p&gt;

&lt;p&gt;The honest price is that &lt;strong&gt;I'm now the one guarding the root key&lt;/strong&gt;. If I lose it, the database is noise. If I leak it alongside a database backup, the layering stops helping. I also own the patching, the backups and the restore test. Nobody pages Infisical's on-call when my container falls over. That's a real cost, and anyone who tells you self-hosting a secrets manager is free is selling something too.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reason three: the CLI is the interface
&lt;/h2&gt;

&lt;p&gt;My scripts don't have a human to type a master password, so they log in as a machine identity using &lt;a href="https://infisical.com/docs/documentation/platform/identities/universal-auth" rel="noopener noreferrer"&gt;universal auth&lt;/a&gt;: a client ID and client secret swapped for a short-lived access token. The &lt;a href="https://infisical.com/docs/cli/commands/login" rel="noopener noreferrer"&gt;Infisical CLI&lt;/a&gt; does that in one line, and the shape in my shell scripts is roughly this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# point the CLI at the self-hosted instance, not Infisical Cloud&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;INFISICAL_TOKEN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;infisical login &lt;span class="nt"&gt;--method&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;universal-auth &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--client-id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$CLIENT_ID&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--client-secret&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$CLIENT_SECRET&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--domain&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$INFISICAL_URL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--plain&lt;/span&gt; &lt;span class="nt"&gt;--silent&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;

infisical secrets get SOME_NAME &lt;span class="nt"&gt;--domain&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$INFISICAL_URL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--projectId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PROJECT_ID&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--env&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;prod &lt;span class="nt"&gt;--path&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/SomeFolder &lt;span class="nt"&gt;--plain&lt;/span&gt; &lt;span class="nt"&gt;--silent&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;--plain&lt;/code&gt; prints just the value and &lt;code&gt;--silent&lt;/code&gt; hides the update tips, which matters when stdout is being captured into a variable. If you self-host, note the docs' warning that the domain has to be set on &lt;em&gt;every&lt;/em&gt; command (or via an environment variable), otherwise the CLI quietly heads for the US cloud. The CLI on my box is v0.38.0.&lt;/p&gt;

&lt;p&gt;For services that take environment variables, &lt;a href="https://infisical.com/docs/cli/commands/run" rel="noopener noreferrer"&gt;&lt;code&gt;infisical run -- &amp;lt;command&amp;gt;&lt;/code&gt;&lt;/a&gt; injects secrets into the process, and &lt;a href="https://infisical.com/docs/cli/commands/export" rel="noopener noreferrer"&gt;&lt;code&gt;infisical export&lt;/code&gt;&lt;/a&gt; can write dotenv, JSON or YAML if something insists on a file.&lt;/p&gt;

&lt;p&gt;Three habits grew out of driving it this way.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Names, never values, in anything automated.&lt;/strong&gt; A nightly inventory script uses the CLI to list secret &lt;em&gt;names&lt;/em&gt; per folder, so my estate notes can say "this folder holds these four things" without ever touching a value. When I need to know two copies agree, I compare hashes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Delete explicitly.&lt;/strong&gt; On my version, deleting a secret as a machine identity failed with "Must be user to delete personal secret". My first thought was that the error was wrong, but it wasn't: the CLI had created that secret as a &lt;em&gt;personal&lt;/em&gt; secret, and deleting it needed &lt;code&gt;--type=shared&lt;/code&gt;. The surprise was the default, not the message.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Python goes through one resolver.&lt;/strong&gt; The Python services never call the CLI. They import one small module that resolves environment, then Infisical, then an explicit default, and raises rather than returning something wrong. Login is cached, so it's one round-trip per process, and resolution happens at the call rather than at import, so an Infisical blip surfaces as a clear error inside the function that needed the secret instead of killing an import. That module is what fixed the two disagreeing scripts: both now ask the same question of the same place.&lt;/p&gt;

&lt;p&gt;The machine identity's own client ID and secret still have to live somewhere, because a bootstrap credential can't fetch itself. Mine sits in a locked-down file on the box, readable only by the account that needs it. The self-rewriting script now updates that file atomically and clears the token cache, instead of editing its own source.&lt;/p&gt;

&lt;h2&gt;
  
  
  Humans and machines are different problems
&lt;/h2&gt;

&lt;p&gt;I don't use Infisical for my own passwords. Those live in Vaultwarden, the self-hosted Bitwarden-compatible server, with around 753 items. The split is simple: &lt;strong&gt;humans go to Vaultwarden, machines go to Infisical&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I found out why it matters by getting it wrong. At one point the Vaultwarden master password was stored in Infisical so automation could unlock the vault. That quietly collapsed the security of my entire password vault onto the agent box and the credential it logs in with. I deleted those entries. The consequence is deliberate: nothing on the estate can unlock my personal vault without me.&lt;/p&gt;

&lt;p&gt;The same boundary shows up in small rules. When a bot needed a website session cookie, the rule was that the cookie value goes into Infisical and the script pulls it at runtime. It never gets pasted into a chat window, however convenient that would be.&lt;/p&gt;

&lt;p&gt;The same thinking applies between machines. Most consumers only ever read a secret; the handful of jobs that rotate one are the only things that should be able to write. For a long time that wasn't true here: every script read with an identity that could also write and delete. I know because I didn't take its permissions on trust: I probed it with a throwaway secret, and it could write and delete. (That probe is also where the personal-secret surprise above came from.) The fix was a second machine identity with the project's built-in Viewer role, which my resolver now uses for every read. The install script only keeps it if a test read works and a test write is refused, and the API's answer to that write was a flat "not allowed to create on secrets". The two jobs that genuinely write, sync and rotation, keep the other identity.&lt;/p&gt;

&lt;h2&gt;
  
  
  What 1Password does better
&lt;/h2&gt;

&lt;p&gt;Plenty, and it's worth saying plainly.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;End-to-end encryption with server ignorance.&lt;/strong&gt; 1Password can't read your data. My Infisical server can decrypt everything, because it has to hand secrets to machines.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No root key to babysit, no server to patch.&lt;/strong&gt; Their operations team is better at uptime than my Proxmox box.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A good CLI of its own.&lt;/strong&gt; &lt;a href="https://www.1password.dev/cli/secrets-scripts/" rel="noopener noreferrer"&gt;&lt;code&gt;op read&lt;/code&gt;, &lt;code&gt;op run&lt;/code&gt; and &lt;code&gt;op inject&lt;/code&gt;&lt;/a&gt; cover the same patterns as mine, using &lt;code&gt;op://vault/item/field&lt;/code&gt; secret references.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Machine access without running anything.&lt;/strong&gt; &lt;a href="https://www.1password.dev/service-accounts/" rel="noopener noreferrer"&gt;Service accounts&lt;/a&gt; automate secrets "without the need to deploy additional services". If you do want something local, a &lt;a href="https://www.1password.dev/connect/" rel="noopener noreferrer"&gt;Connect server&lt;/a&gt; runs in your infrastructure and caches your data there, though it &lt;a href="https://www.1password.dev/connect/get-started/" rel="noopener noreferrer"&gt;keeps that data in sync with 1Password.com&lt;/a&gt; and needs a 1Password account behind it. It's a local cache of a cloud vault, not a self-hosted vault.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Humans.&lt;/strong&gt; Sharing, recovery, apps that non-technical family members will actually use. Nothing I've built competes with that.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What I'd tell someone deciding
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;If you'd rather not own a root key, backups and upgrades, use 1Password. That's a perfectly sensible answer, not a lesser one.&lt;/li&gt;
&lt;li&gt;If your consumers are all on your LAN and you already run backups you've actually restored, self-hosting machine secrets is very reasonable. Budget for the patching.&lt;/li&gt;
&lt;li&gt;Whichever you pick, &lt;strong&gt;keep humans and machines apart&lt;/strong&gt;. The key to your personal vault should never be readable by a script.&lt;/li&gt;
&lt;li&gt;Drive it from the CLI and a single resolver, not from copies. The tool matters less than the fact that there's only one place to ask.&lt;/li&gt;
&lt;li&gt;Give readers and writers different identities, and test what each one can really do.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I didn't pick Infisical because 1Password is weak. I picked it because I wanted to know where every key lives, including the one at the top, and the only way to know that for certain was to hold it myself.&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;🤖 &lt;em&gt;Drafted with AI assistance from my own homelab notes, logs and repos, then reviewed and edited before publishing.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>security</category>
      <category>devops</category>
      <category>selfhosted</category>
      <category>homelab</category>
    </item>
    <item>
      <title>Self-hosted Langfuse: tracing 7% of my AI agents, and ClickHouse logging itself</title>
      <dc:creator>Christian Anderson</dc:creator>
      <pubDate>Thu, 01 Oct 2026 08:00:15 +0000</pubDate>
      <link>https://dev.to/iam-tech/self-hosted-langfuse-tracing-7-of-my-ai-agents-and-clickhouse-logging-itself-3e80</link>
      <guid>https://dev.to/iam-tech/self-hosted-langfuse-tracing-7-of-my-ai-agents-and-clickhouse-logging-itself-3e80</guid>
      <description>&lt;p&gt;On Discord I told my agent &lt;code&gt;post&lt;/code&gt; followed by a comment id. That's the approval step for a dev.to reply it had drafted: read the draft, say "post", done.&lt;/p&gt;

&lt;p&gt;Nothing got posted. The turn took &lt;strong&gt;25 minutes&lt;/strong&gt; and came back with a confident, detailed summary of the article. It was made up.&lt;/p&gt;

&lt;p&gt;I self-host Langfuse to trace my agents, so I went and looked at the trace. Then I checked what else it had seen. It had been watching about 7% of my agents' traffic.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 25-minute turn
&lt;/h2&gt;

&lt;p&gt;Here is the trace of that one turn:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Hermes turn                      1526.3 s
  LLM call 1   qwen3.5-64k        569.4 s
  Tool: web_search                  0.8 s
  Tool: web_search                  0.9 s
  LLM call 2   qwen3.5-64k        207.4 s
  Tool: web_extract                 2.9 s
  LLM call 3   qwen3.5-64k        740.2 s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Tools: 4.6 seconds in total. Model: 1,517 seconds. The local model spent 25 minutes thinking its way to the wrong answer, with three short trips to the web in between.&lt;/p&gt;

&lt;p&gt;Without the trace I'd have blamed the web tools. A turn that searches the web and takes 25 minutes &lt;em&gt;feels&lt;/em&gt; like a network problem. The trace says the network was the fastest thing in the room.&lt;/p&gt;

&lt;p&gt;The actual cause had nothing to do with speed. The Discord conversation was one long-lived session, started days before the reply skill existed. The agent reads its list of skills once, at session start. So as far as that session knew, there was no skill for posting a reply. Asked to "post" something it had no tool for, it did what models do and improvised: searched the web for the article, extracted the page, and wrote a plausible summary instead.&lt;/p&gt;

&lt;p&gt;The fix was a fresh session. But it made me curious about what else the traces could tell me.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the traces said about models
&lt;/h2&gt;

&lt;p&gt;Pulling latency per model from the same Langfuse data (the default profile, all history):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Calls&lt;/th&gt;
&lt;th&gt;p50&lt;/th&gt;
&lt;th&gt;Slowest&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;qwen/qwen3.7-flash&lt;/td&gt;
&lt;td&gt;92&lt;/td&gt;
&lt;td&gt;4.3 s&lt;/td&gt;
&lt;td&gt;180.5 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;qwen3.5-64k (local, 6 GB GPU)&lt;/td&gt;
&lt;td&gt;91&lt;/td&gt;
&lt;td&gt;125.2 s&lt;/td&gt;
&lt;td&gt;900.3 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;deepseek-v4-flash&lt;/td&gt;
&lt;td&gt;65&lt;/td&gt;
&lt;td&gt;6.7 s&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;mistral-nemo (local)&lt;/td&gt;
&lt;td&gt;24&lt;/td&gt;
&lt;td&gt;106.6 s&lt;/td&gt;
&lt;td&gt;1,802.9 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;hermes3-64k&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;2.5 s&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On the tool side, &lt;code&gt;terminal&lt;/code&gt; ran 135 times with a median of 0.6 s, and the slowest single tool call was a &lt;code&gt;search_files&lt;/code&gt; at 120.6 s.&lt;/p&gt;

&lt;p&gt;None of this is shocking. A local model on a 6 GB card is slow, and cheap cloud models are fast. But "slow" and "125 seconds median, about 30 times slower than the fastest cloud option" are different sentences, and only one of them helps decide which job runs where.&lt;/p&gt;

&lt;p&gt;Then I noticed the call counts. 91 calls for the local model. That seemed low for a system I lean on all day.&lt;/p&gt;

&lt;h2&gt;
  
  
  The coverage check
&lt;/h2&gt;

&lt;p&gt;Langfuse held &lt;strong&gt;618 events across 79 traces in four weeks&lt;/strong&gt;. That's about three traces a day.&lt;/p&gt;

&lt;p&gt;The agent keeps its own session database, so I counted from the other side. The last 14 days alone:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;188 sessions&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;1,387 model calls&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;10 agent profiles&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Tracing was enabled on exactly &lt;strong&gt;one of ten&lt;/strong&gt; profiles: the default chat one, with 91 calls. That's roughly 7% of the traffic.&lt;/p&gt;

&lt;p&gt;The untraced ones were the ones doing the work:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Profile&lt;/th&gt;
&lt;th&gt;Model calls (14 days)&lt;/th&gt;
&lt;th&gt;Traced&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;coder&lt;/td&gt;
&lt;td&gt;522&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;sec&lt;/td&gt;
&lt;td&gt;390&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;itadmin&lt;/td&gt;
&lt;td&gt;172&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;writer&lt;/td&gt;
&lt;td&gt;120&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;default (chat)&lt;/td&gt;
&lt;td&gt;91&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;reviewer&lt;/td&gt;
&lt;td&gt;55&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;scout&lt;/td&gt;
&lt;td&gt;29&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;others&lt;/td&gt;
&lt;td&gt;small&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The coding agent alone made more than five times as many calls as the only profile I was watching.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why it happened
&lt;/h3&gt;

&lt;p&gt;The tracing plugin is enabled per profile, and it reads its Langfuse API keys through that profile's own secret scope. That's deliberate: profile B's traces shouldn't be shipped with profile A's keys.&lt;/p&gt;

&lt;p&gt;When I set Langfuse up, I enabled the plugin and added the keys on the default profile, and moved on. Every specialist profile then ran without tracing. No error. No warning in a log. The plugin fails open by design (no keys means its hooks quietly do nothing), which is right for a plugin, and exactly why nobody noticed.&lt;/p&gt;

&lt;p&gt;A monitoring tool that's switched off looks identical to one that has nothing to report.&lt;/p&gt;

&lt;p&gt;There was a second, smaller gap. One-shot jobs launched in "safe mode" skip plugins entirely, and the reply drafter ran that way because it's quicker. That one is closed now: the drafter sends its own traces through the Langfuse SDK, so it no longer depends on the plugin.&lt;/p&gt;

&lt;h2&gt;
  
  
  The second surprise: the disk
&lt;/h2&gt;

&lt;p&gt;Langfuse v3 onwards stores traces in ClickHouse. Postgres holds users and projects, Redis handles the queue, MinIO stores blobs. So no, you can't drop ClickHouse to save resources; it is the trace store.&lt;/p&gt;

&lt;p&gt;ClickHouse was using &lt;strong&gt;6.03 GiB&lt;/strong&gt; of disk. The actual Langfuse trace data in it was &lt;strong&gt;2.2 MiB&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The rest was ClickHouse's own diagnostic logging:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;System table&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;th&gt;Rows&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;system.trace_log&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;3.38 GiB&lt;/td&gt;
&lt;td&gt;174.6 million&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;system.text_log&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1.19 GiB&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;system.part_log&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;492 MiB&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;system.asynchronous_metric_log&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;485 MiB&lt;/td&gt;
&lt;td&gt;949 million&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;system.metric_log&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;461 MiB&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That's roughly 2,700 times more bytes about ClickHouse than about my agents.&lt;/p&gt;

&lt;p&gt;Those tables aren't written once and forgotten. They're written continuously. The node this runs on keeps its guests on a single 5400 rpm HDD, and it already has a roughly 15-minute IO storm after every reboot. The Langfuse container is the one I stop to relieve it. So the trace store I'd barely been using was grinding that disk all day, recording its own profiler samples.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Tracing on every profile
&lt;/h3&gt;

&lt;p&gt;For each of the ten profiles:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Copied the Langfuse keys into that profile's own env file (keeping the per-profile secret scoping).&lt;/li&gt;
&lt;li&gt;Enabled the plugin.&lt;/li&gt;
&lt;li&gt;Set the Langfuse &lt;code&gt;environment&lt;/code&gt; to the profile name.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The third step is the one that matters for later. Traces carry no profile tag of their own, so without it every trace lands in one undifferentiated pile. With it, "cost and latency per agent role" is a filter in the UI rather than a spreadsheet exercise.&lt;/p&gt;

&lt;p&gt;Then I restarted the gateway and sent a one-word probe on two profiles. Traces arrived tagged &lt;code&gt;legal&lt;/code&gt; (on a cloud flash model) and &lt;code&gt;writer&lt;/code&gt; (on the local qwen3.5-64k). That's the test I should have run the first time.&lt;/p&gt;

&lt;h3&gt;
  
  
  ClickHouse logging itself
&lt;/h3&gt;

&lt;p&gt;ClickHouse lets you override server config with a file dropped into &lt;code&gt;/etc/clickhouse-server/config.d/&lt;/code&gt;. This one removes twelve system log tables, keeps &lt;code&gt;query_log&lt;/code&gt; with a 7-day TTL (it's genuinely useful when a query is slow), and turns the server logger down to warnings:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xml"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;clickhouse&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;trace_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;text_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;metric_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;asynchronous_metric_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;part_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;processors_profile_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;query_thread_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;query_views_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;query_metric_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;opentelemetry_span_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;session_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;latency_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;query_log&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;ttl&amp;gt;&lt;/span&gt;event_date + INTERVAL 7 DAY DELETE&lt;span class="nt"&gt;&amp;lt;/ttl&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;/query_log&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;logger&amp;gt;&amp;lt;level&amp;gt;&lt;/span&gt;warning&lt;span class="nt"&gt;&amp;lt;/level&amp;gt;&amp;lt;/logger&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/clickhouse&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It's mounted read-only into that directory through the compose file, and the container was recreated. Checks afterwards:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;container healthy&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;trace_log&lt;/code&gt; row count unchanged over 30 seconds, so it has stopped writing&lt;/li&gt;
&lt;li&gt;trace data intact&lt;/li&gt;
&lt;li&gt;Langfuse web health check returns 200&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One honest caveat: removing a log table from config stops ClickHouse writing to it, but it doesn't delete what's already there. The old 5.9 GiB of log tables was still on disk. I dropped those deliberately, as a separate step, rather than bundling a destructive delete into a config change. The Langfuse container went from 15 GB of disk used to 7.7 GB, with the trace data intact.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I want this at all
&lt;/h2&gt;

&lt;p&gt;It would have been easy to conclude Langfuse was dead weight: a few gigabytes of disk for 79 traces.&lt;/p&gt;

&lt;p&gt;But the estate runs around ten agent profiles: a chief of staff that I chat with, a coder, security, IT admin, writer, scout, reviewer, legal, analyst and a few more. They run on a mix of local models on a 6 GB GPU and cheap paid cloud models with a fallback chain. The questions I actually have about it are ones only traces answer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Which role burns the calls?&lt;/strong&gt; Coder, at 522 of 1,387.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Where does a slow turn's time go?&lt;/strong&gt; Model or tool. See the 25-minute turn, which looked like a web problem and was a thinking problem.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Did the fallback fire?&lt;/strong&gt; A cloud model failing over to another is invisible from the outside if the answer still arrives.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Local or cloud for this job?&lt;/strong&gt; 125 seconds median against 4 to 7 seconds is a real trade, and it should be made per job, not by habit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What does each role cost?&lt;/strong&gt; Now that traces are tagged per profile, that's a query.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The rule I run the estate by is "scripts gather facts, models never do". Models get to summarise what scripts measured; they never go and measure it themselves. Langfuse is the fact-gatherer for the agents themselves. It just can't gather facts from agents it isn't connected to.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the weekly report tracks now
&lt;/h2&gt;

&lt;p&gt;A small script now reads the agents' own session records and Langfuse side by side, and every Monday it posts a report: per role, the turns, model calls, tokens, cost, errors and the share that was traced; per model, p50 and p95 latency and cost; how turn time splits between model and tools; the five slowest turns; and tool errors. It's built to answer these questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Does the coder profile's call volume hold up, and is it doing many short turns or a few enormous ones?&lt;/li&gt;
&lt;li&gt;How often do the local models hit their multi-minute tail, and on which profiles?&lt;/li&gt;
&lt;li&gt;How often does the cloud fallback chain actually fire, and for which model?&lt;/li&gt;
&lt;li&gt;Which one or two jobs account for most of the cost, and would they be fine on a local model?&lt;/li&gt;
&lt;li&gt;Does ClickHouse stay small now that it's only storing what I asked it to?&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  If you self-host Langfuse, check these three things
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Is tracing actually enabled everywhere you think it is?&lt;/strong&gt; Per service, per profile, per worker. Count calls from your application's side and compare with Langfuse's trace count. If one is an order of magnitude bigger than the other, like mine, you'll know.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;How big are ClickHouse's system log tables?&lt;/strong&gt; Check &lt;code&gt;system.parts&lt;/code&gt; grouped by table. If &lt;code&gt;trace_log&lt;/code&gt; and &lt;code&gt;asynchronous_metric_log&lt;/code&gt; dwarf your actual data, a &lt;code&gt;config.d&lt;/code&gt; override fixes it in a few lines.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Send a probe trace from every agent.&lt;/strong&gt; One word, one turn, per profile, and confirm each one lands with the right tag. A plugin that fails open will never tell you it isn't there.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Postscript: it happened again
&lt;/h2&gt;

&lt;p&gt;On 22 September the tracing server's address changed, and all 12 specialist agents' configs were repointed at the new one. When it moved back on 23 September, only the main config was updated. Every specialist silently stopped tracing. I found and fixed it on 24 September.&lt;/p&gt;

&lt;p&gt;No error, no warning, nothing in a log. It's the same fail-open silence this whole post is about, and the reason point 1 above isn't a one-off check.&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;🤖 &lt;em&gt;Drafted with AI assistance from my own homelab notes, logs and repos, then reviewed and edited before publishing.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>observability</category>
      <category>selfhosted</category>
    </item>
    <item>
      <title>Shopify products.json vs anti-bot walls: 85 fashion sources, measured</title>
      <dc:creator>Christian Anderson</dc:creator>
      <pubDate>Wed, 30 Sep 2026 08:00:03 +0000</pubDate>
      <link>https://dev.to/iam-tech/shopify-productsjson-vs-anti-bot-walls-85-fashion-sources-measured-17lo</link>
      <guid>https://dev.to/iam-tech/shopify-productsjson-vs-anti-bot-walls-85-fashion-sources-measured-17lo</guid>
      <description>&lt;p&gt;I built a personal shopping agent. It runs on my homelab, gathers sale items from a list of sources once a day with two refreshes later on, and a half-hourly job checks alerts and my watchlist. The point is to tell me when something I'd actually wear is reduced &lt;strong&gt;in my size&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The interesting part was never the agent. It was finding out how much of the retail web will talk to a program at all. This is the measured version: 85 sources, what each gave up, and four mistakes I made reading the results.&lt;/p&gt;

&lt;h2&gt;
  
  
  The scoreboard
&lt;/h2&gt;

&lt;p&gt;Every number here comes from one run: the daily gather on 19 September, which took 854 seconds.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Status&lt;/th&gt;
&lt;th&gt;Sources&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;OK&lt;/td&gt;
&lt;td&gt;56&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Blocked&lt;/td&gt;
&lt;td&gt;22&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Parse failed&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reachable, nothing parsed&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Error&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Only &lt;strong&gt;45 of the 85 returned even a single item&lt;/strong&gt;. Eleven of the "OK" sources answered fine and had nothing on sale that matched.&lt;/p&gt;

&lt;p&gt;The run collected 4,112 discounted products. After dropping what wasn't menswear, wasn't clothing or footwear, wasn't in stock in my size or wasn't really reduced, &lt;strong&gt;941 were in my size&lt;/strong&gt;, plus 452 where the size couldn't be checked. The biggest single drop was "wrong size", at 1,411.&lt;/p&gt;

&lt;p&gt;And the distribution is lopsided in a way I didn't expect. &lt;strong&gt;21 Shopify stores produced 3,173 of the 4,112 items&lt;/strong&gt;, more than three quarters, from one endpoint that has been sitting there the whole time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Shopify's &lt;code&gt;/products.json&lt;/code&gt; is the quiet hero
&lt;/h2&gt;

&lt;p&gt;Any Shopify storefront will hand you &lt;code&gt;/products.json&lt;/code&gt;: paginated, JSON, no key, no rendering, no scraping. Every product, every variant, with &lt;code&gt;price&lt;/code&gt; and &lt;code&gt;compare_at_price&lt;/code&gt;, so "is this actually reduced" is a subtraction rather than a guess about a strikethrough in the markup.&lt;/p&gt;

&lt;p&gt;Critically, the variants carry &lt;strong&gt;per-size availability&lt;/strong&gt;. That one field is the whole reason this project works.&lt;/p&gt;

&lt;p&gt;The other pleasant surprise was a big retailer that ships a &lt;strong&gt;public, search-only Algolia key in its own sale page&lt;/strong&gt;. Its site uses that key to power its sale filters. I use it the same way, from the same page, to ask for on-sale menswear in my sizes, and on that run it returned 400 items filtered server-side. No rendering, no parsing, no load a normal visitor wouldn't cause.&lt;/p&gt;

&lt;p&gt;The pattern: before you write a scraper, check whether the shop's own front end is already calling something cleaner than the HTML.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hard part isn't the deal. It's the size.
&lt;/h2&gt;

&lt;p&gt;Eight sources list products and prices but have &lt;strong&gt;no per-size stock I can read&lt;/strong&gt;: Nike, JD Sports, Selfridges, Footasylum, Puma, Converse, Clarks and one flash-sale site. Some don't put sizes on the listing at all. On others the size selector looks the same whether a size is in stock or gone. Puma loads its size grid and inventory from a later API call, so none of it is in the HTML.&lt;/p&gt;

&lt;p&gt;So they return items I can't confirm I can buy. A 40%-off jacket that doesn't exist in my size isn't a deal, it's an errand.&lt;/p&gt;

&lt;p&gt;The rule I settled on: &lt;strong&gt;keep them in the pool, tag them &lt;code&gt;size_unknown&lt;/code&gt;, and never alert on them.&lt;/strong&gt; They can turn up if I go looking, but they can never interrupt me. For any notifier, the real question is the false-positive rate, and what the user does after the third one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mistake one: I read a 403 as a "no"
&lt;/h2&gt;

&lt;p&gt;Five well-known retailers were written off early as &lt;em&gt;"robots.txt disallows"&lt;/em&gt;. I'd built the polite thing — fetch robots, honour it, move on — and five big names said no, so I moved on.&lt;/p&gt;

&lt;p&gt;They hadn't said no. Their &lt;code&gt;robots.txt&lt;/code&gt; &lt;strong&gt;returns 403 to a plain HTTP request&lt;/strong&gt;, because the file sits behind the same edge protection as everything else. My code caught the failure, couldn't parse any rules, and fell through to "assume disallowed". Fetched the way a browser fetches it, every one of those files allows &lt;code&gt;*&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Failing to read a policy is not the same as the policy saying no.&lt;/em&gt; If your fallback is a decision, log it as a decision ("couldn't read robots, assuming disallow"), not as a fact.&lt;/p&gt;

&lt;p&gt;The sting in the tail: I re-tested all five properly, and all five still gave me nothing. No men's prices on one, redirects and a region picker on two, an empty page on the other two. Right answer, wrong reason, for about a day. Those are the ones that rot quietly: the outcome looks correct, so nothing makes you check the logic.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mistake two: "blocked" is three different things
&lt;/h2&gt;

&lt;p&gt;When I probed a new batch of sources, I tried each one twice: once with a plain HTTP client and once rendered in a real browser. Same honest User-Agent, same robots check, no evasion. The only difference is that JavaScript runs.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Shop&lt;/th&gt;
&lt;th&gt;Plain HTTP client&lt;/th&gt;
&lt;th&gt;Rendered in a browser&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;High-street fashion chain&lt;/td&gt;
&lt;td&gt;connection timeout&lt;/td&gt;
&lt;td&gt;76 prices on the sale page&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Off-price outlet&lt;/td&gt;
&lt;td&gt;not tried&lt;/td&gt;
&lt;td&gt;287 prices&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;High-street chain&lt;/td&gt;
&lt;td&gt;200, but no product grid&lt;/td&gt;
&lt;td&gt;179 prices&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Secondhand marketplace&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;192 listings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Voucher aggregator&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;230 KB of offers&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A plain client would have put the first and third in the blocked column. One was just rendering its grid in JavaScript, and the other didn't answer a bare client but was perfectly happy to serve a browser.&lt;/p&gt;

&lt;p&gt;Two others, Sports Direct and a high-street chain, returned &lt;strong&gt;404 to every URL I tried&lt;/strong&gt;. That wasn't a block at all. I was guessing their category paths wrong, and Sports Direct parses fine now from its real sale page. That's the trap: in a health table, "edge-blocked", "renders empty without JavaScript" and "you asked for a page that doesn't exist" all show up as the same red cell. Only the first is the retailer's decision. The other two are my bugs, and they'll sit in the blocked column looking like someone else's fault.&lt;/p&gt;

&lt;p&gt;So the status note is now worth more than the count. Record &lt;em&gt;why&lt;/em&gt; a source is red, and re-test the reds with a different client now and then.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mistake three: Foot Locker was my bug, not theirs
&lt;/h2&gt;

&lt;p&gt;On the 19 September run, Foot Locker returned 48 discounted shoes and none in my size, because my reader couldn't find per-size stock. I'd filed it with the others: size selector, no availability markers, nothing to read.&lt;/p&gt;

&lt;p&gt;The stock was there all along. Foot Locker ships exact per-size availability in JSON embedded in the product page, with UK, EU and US sizes for each variant. My reader only looked at the page's visible elements. I fixed it on 20 September to read the embedded JSON first and fall back to the old method, and Foot Locker went from 0 to 9 in-size deals on the next full refresh.&lt;/p&gt;

&lt;p&gt;Same day, same lesson: some European stores list shoe sizes as bare numbers like &lt;code&gt;45.5&lt;/code&gt;, and my parser read them as UK sizes and threw them away as wrong. "They don't publish it" and "I didn't read it properly" look identical from the outside.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mistake four: I wrote my own guess into the outfit builder
&lt;/h2&gt;

&lt;p&gt;The daily digest includes a couple of outfit ideas. Early on it put a brown hoodie forward as something I'd wear. The model hadn't made that up. My own outfit-building code assumed brown and cream were colours I wear, because I'd written that assumption in without asking. That's a question for me to answer, not a default for the code to guess.&lt;/p&gt;

&lt;p&gt;The builder now splits the work between code and model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Code&lt;/strong&gt; picks the candidate pieces out of the pool, per slot (top, layer, bottom, footwear), inside my sizes. Sandals, mules and slides are excluded here, before the model sees anything.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The model&lt;/strong&gt; chooses ids from that list and writes the name and the one-line idea.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Code&lt;/strong&gt; validates the answer: the right slots filled, no more than one neutral, the cheaper look actually cheaper than the dearer one. If validation fails, it falls back to a deterministic pick.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The model does the part it's good at (taste, phrasing) and none of the part where being confidently wrong costs money.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rules I gave myself
&lt;/h2&gt;

&lt;p&gt;Worth stating, because the blocked column is 22 and the temptation to fix that is real:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Identify honestly in the User-Agent. No pretending to be Chrome.&lt;/li&gt;
&lt;li&gt;No proxies, no captcha solving, no rotating anything.&lt;/li&gt;
&lt;li&gt;A challenge page, an error page or an empty render counts as &lt;strong&gt;blocked&lt;/strong&gt;. Back off 72 hours, don't retry in a loop.&lt;/li&gt;
&lt;li&gt;Cache, and keep the number of requests per site per run small.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One retailer worked fine for about 30 page loads and then Akamai shut the door. That's the system telling you where the line is. Others sit behind DataDome and reCAPTCHA. They've decided, and the honest response is to record it and move on, not engineer around it.&lt;/p&gt;

&lt;p&gt;Which leaves the real conclusion. &lt;strong&gt;The shops that were easiest to work with are the ones that already publish structured data&lt;/strong&gt;, deliberately or as a side effect of their own front end. Everyone else is spending money on a wall, and the thing behind the wall — is this jacket in my size — is the only fact I ever wanted.&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;🤖 &lt;em&gt;Drafted with AI assistance from my own homelab notes, logs and repos, then reviewed and edited before publishing.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>python</category>
      <category>shopify</category>
      <category>webscraping</category>
      <category>javascript</category>
    </item>
    <item>
      <title>Docker homelab gotchas: docker ps called the image orphaned. It was live</title>
      <dc:creator>Christian Anderson</dc:creator>
      <pubDate>Tue, 29 Sep 2026 08:00:15 +0000</pubDate>
      <link>https://dev.to/iam-tech/docker-homelab-gotchas-docker-ps-called-the-image-orphaned-it-was-live-1l4m</link>
      <guid>https://dev.to/iam-tech/docker-homelab-gotchas-docker-ps-called-the-image-orphaned-it-was-live-1l4m</guid>
      <description>&lt;p&gt;My Docker host was 82% full, so I went looking for images to delete. &lt;code&gt;docker system df&lt;/code&gt; said 6.67 GB was reclaimable. &lt;code&gt;docker ps --format '{{.Image}}'&lt;/code&gt; printed bare SHA IDs instead of names for several running containers, which made one image, &lt;code&gt;unclecode/crawl4ai&lt;/code&gt;, look like nothing was using it.&lt;/p&gt;

&lt;p&gt;It was running. Healthy, up two weeks, serving traffic. I was one &lt;code&gt;docker rmi&lt;/code&gt; away from taking out a live service on the strength of a formatting quirk.&lt;/p&gt;

&lt;p&gt;What stopped me wasn't the image list. It was a listening port that nothing on my list owned.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="k"&gt;for &lt;/span&gt;c &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;docker ps &lt;span class="nt"&gt;-q&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do &lt;/span&gt;docker inspect &lt;span class="nt"&gt;-f&lt;/span&gt; &lt;span class="s1"&gt;'{{.Name}} {{.Config.Image}}'&lt;/span&gt; &lt;span class="nv"&gt;$c&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;done
&lt;/span&gt;ss &lt;span class="nt"&gt;-ltnp&lt;/span&gt;        &lt;span class="c"&gt;# a port nobody claims is a container you mis-read&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the pattern for most of what follows: Docker's summary views are fine until you act on them.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I actually run it
&lt;/h2&gt;

&lt;p&gt;Nothing exotic. Docker runs inside unprivileged LXC containers on Proxmox (they need &lt;code&gt;nesting=1&lt;/code&gt; and &lt;code&gt;keyctl=1&lt;/code&gt;), plus a small NAS running an appliance OS whose app manager is a layer over Docker. Single-purpose things, like my voice assistant's speech-to-text and text-to-speech servers, are one &lt;code&gt;docker run&lt;/code&gt; line each. Anything with a database is a compose stack; the LLM tracing one is six containers.&lt;/p&gt;

&lt;p&gt;Here's what bit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Disk: the pressure was images, not data
&lt;/h2&gt;

&lt;p&gt;Four guests were 82 to 86% full, and my instinct was to push data to the NAS. Profiling disagreed. The Docker LXC had 16 GB of images. Another guest had 11.4 GB of images and 11.3 GB of build cache. There was very little cold data at all: my weekly offload job's first run reclaimed 213 MB.&lt;/p&gt;

&lt;p&gt;You can't move a running container's image layers onto NFS anyway, since overlayfs over NFS is slow and fragile. The fix was growing the guest disks, online, from a thin pool that was 7% used.&lt;/p&gt;

&lt;p&gt;Two more lessons from the clean-up:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Per-tag sizes are a lie for planning.&lt;/strong&gt; I removed 16 superseded tags of my own app, each listed at about 240 MB. They freed &lt;strong&gt;60 MB&lt;/strong&gt;, because the builds share nearly every layer. The whole clean-up, dangling images and build cache included, came to about 340 MB.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An image can be dead in &lt;code&gt;ps&lt;/code&gt; and still be a dependency.&lt;/strong&gt; Three NetBird images weren't running, but a compose file on disk still referenced them. What I did delete had no container ever and no references anywhere: 2.8 GB.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;code&gt;docker image prune -a&lt;/code&gt; is now on my "a human types it" list. On one host 4 of the 11 images were in use, and there's no fast path to re-pull if something happens to be stopped at the time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Logs: nobody set a limit
&lt;/h2&gt;

&lt;p&gt;Another Docker host's root filesystem reached 100%, 0 bytes free, with five failed systemd units. Nothing had alerted. The cause was that &lt;code&gt;/etc/docker/daemon.json&lt;/code&gt; didn't exist, so Docker was using the default &lt;code&gt;json-file&lt;/code&gt; log driver with no rotation and no size cap. The Home Assistant container's log was a single &lt;strong&gt;1.9 GB&lt;/strong&gt; file.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;truncate&lt;/span&gt; &lt;span class="nt"&gt;-s&lt;/span&gt; 0 /var/lib/docker/containers/&amp;lt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;/&amp;lt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="nt"&gt;-json&lt;/span&gt;.log   &lt;span class="c"&gt;# 1.9G&lt;/span&gt;
journalctl &lt;span class="nt"&gt;--vacuum-size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;100M                                  &lt;span class="c"&gt;# 286M&lt;/span&gt;
apt-get clean                                                  &lt;span class="c"&gt;# 271M&lt;/span&gt;
docker image prune &lt;span class="nt"&gt;-f&lt;/span&gt;        &lt;span class="c"&gt;# dangling only, reclaimed 3.47G&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All four containers kept their uptimes. Then I wrote the &lt;code&gt;daemon.json&lt;/code&gt; that should have been there all along: &lt;code&gt;json-file&lt;/code&gt;, &lt;code&gt;max-size: 50m&lt;/code&gt;, &lt;code&gt;max-file: 3&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;It didn't cap Home Assistant's log. Weeks later, the container still reports an empty log config:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker inspect &lt;span class="nt"&gt;-f&lt;/span&gt; &lt;span class="s1"&gt;'{{.Name}} {{.HostConfig.LogConfig}}'&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;docker ps &lt;span class="nt"&gt;-q&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="c"&gt;# /homeassistant {json-file map[]}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Daemon log defaults are copied into a container when it's created, not when Docker restarts. Every container that existed before the file stays uncapped until it's recreated. The only containers on that host with a cap are two whose compose file sets its own &lt;code&gt;logging:&lt;/code&gt; options.&lt;/p&gt;

&lt;h2&gt;
  
  
  Restart policy: the file said one thing, the container another
&lt;/h2&gt;

&lt;p&gt;A file browser on the NAS hard-exited one morning because it couldn't reach the SSO provider. It stayed dead for &lt;strong&gt;six days&lt;/strong&gt;. &lt;code&gt;RestartCount=0&lt;/code&gt;. Nothing retried and nothing alerted.&lt;/p&gt;

&lt;p&gt;The compose file said &lt;code&gt;restart: unless-stopped&lt;/code&gt;. The container had been created before that line was added and was never recreated, so it was still running &lt;code&gt;restart=no&lt;/code&gt;. I'd "fixed" this once already with &lt;code&gt;docker update&lt;/code&gt;, and it had reverted. The second time I checked the on-disk &lt;code&gt;hostconfig.json&lt;/code&gt; rather than trusting &lt;code&gt;docker inspect&lt;/code&gt;. Then I audited the NAS: all 18 running containers now have &lt;code&gt;unless-stopped&lt;/code&gt;. Five didn't, including the SSO container that every login depends on. The reverse proxy had the same drift.&lt;/p&gt;

&lt;p&gt;That appliance layer had its own go at me. My dashboard container was uninstalled by the NAS's app manager twice, with its config left on disk both times. The first time I reinstalled from the leftover compose file. The second time that file was gone too. I don't know why it happened, so I stopped letting the app manager own it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;--name&lt;/span&gt; homepage &lt;span class="nt"&gt;--restart&lt;/span&gt; unless-stopped &lt;span class="nt"&gt;-p&lt;/span&gt; &amp;lt;port&amp;gt;:3000 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-v&lt;/span&gt; &amp;lt;config-dir&amp;gt;:/app/config &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-v&lt;/span&gt; /var/run/docker.sock:/var/run/docker.sock:ro &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="nv"&gt;HOMEPAGE_ALLOWED_HOSTS&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&amp;lt;host&amp;gt;:&amp;lt;port&amp;gt;,localhost:&amp;lt;port&amp;gt; &lt;span class="se"&gt;\&lt;/span&gt;
  ghcr.io/gethomepage/homepage:v1.8.0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Leave out &lt;code&gt;HOMEPAGE_ALLOWED_HOSTS&lt;/code&gt; on v1.x and every browser request gets a blank 400.&lt;/p&gt;

&lt;h2&gt;
  
  
  Configuration is read at create time
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;docker compose restart&lt;/code&gt; does not reload &lt;code&gt;.env&lt;/code&gt;.&lt;/strong&gt; &lt;code&gt;env_file&lt;/code&gt; is read when the container is created. Use &lt;code&gt;docker compose up -d --force-recreate&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Compose interpolates &lt;code&gt;$&lt;/code&gt; inside &lt;code&gt;env_file&lt;/code&gt; values.&lt;/strong&gt; My password manager's admin token is an Argon2 PHC string, &lt;code&gt;$argon2id$v=19$m=19456,...&lt;/code&gt;. It reached the container as 51 mangled characters beginning &lt;code&gt;=19=19456,&lt;/code&gt;, so the admin page returned 401 for every token, the right one included. Doubling every &lt;code&gt;$&lt;/code&gt; fixed it, and &lt;code&gt;docker inspect&lt;/code&gt; showed the container getting 134 characters starting &lt;code&gt;$argon2id$&lt;/code&gt;. The file being right proved nothing. The env file is named &lt;code&gt;vw.env&lt;/code&gt;, not &lt;code&gt;.env&lt;/code&gt;, because compose auto-loads &lt;code&gt;.env&lt;/code&gt; as its own interpolation source.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;environment:&lt;/code&gt; beats &lt;code&gt;env_file&lt;/code&gt;.&lt;/strong&gt; My app's compose file pinned three service URLs in &lt;code&gt;environment:&lt;/code&gt;, so anyone setting them correctly in their own &lt;code&gt;.env&lt;/code&gt; was silently overridden.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stock compose defaults can contradict each other.&lt;/strong&gt; The tracing stack's compose file sets &lt;code&gt;DATABASE_URL&lt;/code&gt; to a hardcoded &lt;code&gt;postgres:postgres&lt;/code&gt; instead of building it from &lt;code&gt;POSTGRES_PASSWORD&lt;/code&gt;. Set only the password and Postgres changes while the app keeps using the literal default:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Error: P1000: Authentication failed against database server
Applying database migrations failed. ... Exiting...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The web container restart-looped. The error mentions special characters that aren't URL-encoded, which was a red herring: the password was alphanumeric by construction. The fix is an explicit &lt;code&gt;DATABASE_URL&lt;/code&gt; in &lt;code&gt;.env&lt;/code&gt;, kept in step with the password.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code baked into the image isn't live because the tests passed.&lt;/strong&gt; My web app's code is copied in at build time, and compose mounts only the data volume. I committed three fixes, watched 661 tests pass, and both containers carried on running the previous version with the new code on disk beside them. Now I ask the container:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker &lt;span class="nb"&gt;exec&lt;/span&gt; &amp;lt;container&amp;gt; python &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"from &amp;lt;pkg&amp;gt; import &amp;lt;module&amp;gt;; print(&amp;lt;module&amp;gt;.&amp;lt;CONSTANT_THE_FIX_ADDED&amp;gt;)"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Networking: published ports and DOCKER-USER
&lt;/h2&gt;

&lt;p&gt;My web app has a private instance and an internet-facing demo copy on the same Docker host. The private one is meant to be reached only through a reverse proxy on another box, but it was published on &lt;code&gt;0.0.0.0&lt;/code&gt;, so the demo container could reach it directly and skip the proxy.&lt;/p&gt;

&lt;p&gt;Binding it to &lt;code&gt;127.0.0.1&lt;/code&gt; fixed that and took the site down for a day, because the proxy is on a different machine. The real fix was to publish on the host's LAN address and restrict it in the &lt;code&gt;DOCKER-USER&lt;/code&gt; chain. The shape:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;iptables &lt;span class="nt"&gt;-N&lt;/span&gt; PRIVATE-IN
iptables &lt;span class="nt"&gt;-I&lt;/span&gt; DOCKER-USER &lt;span class="nt"&gt;-m&lt;/span&gt; conntrack &lt;span class="nt"&gt;--ctorigdst&lt;/span&gt; 192.0.2.10 &lt;span class="nt"&gt;--ctorigdstport&lt;/span&gt; 8080 &lt;span class="nt"&gt;-j&lt;/span&gt; PRIVATE-IN
iptables &lt;span class="nt"&gt;-A&lt;/span&gt; PRIVATE-IN &lt;span class="nt"&gt;-m&lt;/span&gt; conntrack &lt;span class="nt"&gt;--ctstate&lt;/span&gt; ESTABLISHED &lt;span class="nt"&gt;-j&lt;/span&gt; RETURN
iptables &lt;span class="nt"&gt;-A&lt;/span&gt; PRIVATE-IN &lt;span class="nt"&gt;-s&lt;/span&gt; 192.0.2.20 &lt;span class="nt"&gt;-j&lt;/span&gt; RETURN     &lt;span class="c"&gt;# the proxy&lt;/span&gt;
iptables &lt;span class="nt"&gt;-A&lt;/span&gt; PRIVATE-IN &lt;span class="nt"&gt;-j&lt;/span&gt; DROP
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(Addresses and port are placeholders.) The proxy gets a 200 and everything else times out.&lt;/p&gt;

&lt;p&gt;Before I added egress rules, the internet-facing demo could reach far more of my LAN than it had any business touching. Afterwards those internal services time out, while the handful it genuinely needs and the internet still answer. Getting there turned on three traps:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Matching &lt;code&gt;-d&lt;/code&gt; never fires for published ports.&lt;/strong&gt; By the time a packet reaches &lt;code&gt;FORWARD&lt;/code&gt;, the destination has already been DNATed from the host address onto a &lt;code&gt;172.x&lt;/code&gt; container address. &lt;code&gt;--ctorigdst&lt;/code&gt; matches what the client actually asked for.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DNS doesn't go through &lt;code&gt;FORWARD&lt;/code&gt;.&lt;/strong&gt; The container's resolver was the host's own mesh-VPN address, which is local, so queries land on &lt;code&gt;INPUT&lt;/code&gt;. A mesh-range DROP in &lt;code&gt;DOCKER-USER&lt;/code&gt; never touches DNS, which looks reassuring, but the separate &lt;code&gt;INPUT -i br-...&lt;/code&gt; chain must allow port 53 or the container goes deaf.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The bridge name changes.&lt;/strong&gt; It's &lt;code&gt;br-&lt;/code&gt; plus the first 12 characters of the network ID, and recreating the network gives it a new ID. The script reads bridge and subnet from &lt;code&gt;docker network inspect&lt;/code&gt; every run, and I re-run it after any &lt;code&gt;compose down&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Docker rebuilds &lt;code&gt;DOCKER-USER&lt;/code&gt; when it starts, so the script runs from a systemd unit with &lt;code&gt;After=docker.service&lt;/code&gt;. It also had its own bug: the cleanup &lt;code&gt;-D DOCKER-USER&lt;/code&gt; line left out the &lt;code&gt;-s&lt;/code&gt; match it was inserted with, so it never deleted anything. I found 17 duplicate jumps. There's now one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Healthy is not working
&lt;/h2&gt;

&lt;p&gt;Two containers that &lt;code&gt;docker ps&lt;/code&gt; would call fine:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;My Cloudflare tunnel ran from a container that turned out to be a &lt;strong&gt;web UI for&lt;/strong&gt; cloudflared, not a connector. After a recreate it logged &lt;code&gt;No pre-existing config file found&lt;/code&gt;, showed healthy, and connected no tunnel at all. Every public hostname returned 530, SSO included. A plain &lt;code&gt;cloudflare/cloudflared&lt;/code&gt; container with &lt;code&gt;--restart unless-stopped&lt;/code&gt; fixed it.&lt;/li&gt;
&lt;li&gt;The NAS has a Pascal-generation Quadro that, last I checked, falls off the PCIe bus hours after boot, so Immich's CUDA machine-learning container logs &lt;code&gt;no CUDA-capable device&lt;/code&gt; and falls back to CPU. Immich v3 dropped Pascal support, so the moment the card is fixed that container will die with SIGILL, exit 132, and crash-loop. A working GPU is worse than a missing one until I change the image tag.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Rules I run by now
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Never delete an image on the evidence of &lt;code&gt;docker ps&lt;/code&gt; or &lt;code&gt;docker system df&lt;/code&gt;. Inspect the containers and check the listening ports.&lt;/li&gt;
&lt;li&gt;Check what the container received (&lt;code&gt;docker inspect&lt;/code&gt;, &lt;code&gt;docker exec ... env&lt;/code&gt;), not what the file says.&lt;/li&gt;
&lt;li&gt;Changed &lt;code&gt;.env&lt;/code&gt;? &lt;code&gt;up -d --force-recreate&lt;/code&gt;, not &lt;code&gt;restart&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Every container gets &lt;code&gt;--restart unless-stopped&lt;/code&gt;, and I check it persisted.&lt;/li&gt;
&lt;li&gt;If an app manager can uninstall it, anything I care about runs outside it.&lt;/li&gt;
&lt;li&gt;Logs get a size cap on day one. A &lt;code&gt;daemon.json&lt;/code&gt; default only reaches containers created after it, so recreate the old ones.&lt;/li&gt;
&lt;li&gt;Firewall published ports by &lt;code&gt;--ctorigdst&lt;/code&gt; in &lt;code&gt;DOCKER-USER&lt;/code&gt;, derive the bridge name every run, and put the script behind &lt;code&gt;After=docker.service&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;"Healthy" means the process is up. It doesn't mean it's doing its job.&lt;/li&gt;
&lt;/ul&gt;




&lt;blockquote&gt;
&lt;p&gt;🤖 &lt;em&gt;Drafted with AI assistance from my own homelab notes, logs and repos, then reviewed and edited before publishing.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>docker</category>
      <category>devops</category>
      <category>homelab</category>
      <category>selfhosted</category>
    </item>
    <item>
      <title>Discord as an AI agent UI: why chat is my primary interface, not dashboards</title>
      <dc:creator>Christian Anderson</dc:creator>
      <pubDate>Mon, 28 Sep 2026 08:00:13 +0000</pubDate>
      <link>https://dev.to/iam-tech/discord-as-an-ai-agent-ui-why-chat-is-my-primary-interface-not-dashboards-4k07</link>
      <guid>https://dev.to/iam-tech/discord-as-an-ai-agent-ui-why-chat-is-my-primary-interface-not-dashboards-4k07</guid>
      <description>&lt;p&gt;I have dashboards. There's one for the agent platform itself, behind single sign-on. There's a read-only dashboard for a trading experiment, one for my home-grown intrusion detection, and a Homepage start page tying the lot together. I open them when I want to look at something.&lt;/p&gt;

&lt;p&gt;But I don't talk to my agents through any of them. I talk to them in Discord. A private server, one bot, and a set of channels. That's the primary interface, and this is how it's laid out, what it gives me for nothing, and what actually goes wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the channels route work
&lt;/h2&gt;

&lt;p&gt;The main channel goes to my chief-of-staff agent, which is called Denzel. Anything I type there lands with it. It does some admin itself, and anything on a specialist's turf becomes a card on a kanban board, assigned to the right profile: infrastructure to the coder, research to the scout, security to the security profile, long-form writing to the writer.&lt;/p&gt;

&lt;p&gt;The specialists report into their own channels. There's one for kanban updates, one for network and security alerts, one for sysadmin reports, one for dev.to, one for music, one for family holiday planning. Two of them take conversation directly: the shopping assistant and the toolsmith each have a channel routed straight to their profile, with no @mention needed. Only my account is allowed to talk to the bot at all.&lt;/p&gt;

&lt;p&gt;Some channels double as approval queues. The toolsmith posts each tool it thinks I should trial as its own message, and I react ✅ or ❌. The SEO agent does the same with metadata changes for my posts. A reaction is about the smallest approval interface there is, and it works from my phone.&lt;/p&gt;

&lt;p&gt;The useful property is that the room is the address. I don't type which agent I mean. Adding a specialist means adding a channel and a route.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Discord gives a one-user setup for free
&lt;/h2&gt;

&lt;p&gt;If I built this as a web app, I'd be building a list of things that have nothing to do with the agents:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Identity and auth.&lt;/strong&gt; Discord already knows who I am. The bot only answers my account.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Push notifications&lt;/strong&gt; to my phone and desktop, with no setup on my side.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;History and search&lt;/strong&gt;, going back as far as the channels do.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Threads.&lt;/strong&gt; The bot can open one per conversation, which keeps a long back-and-forth out of the main channel.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Images.&lt;/strong&gt; I can send the shopping assistant a photo of a jacket and it finds similar pieces in my size.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mobile, off my home network.&lt;/strong&gt; The agents reach Discord outbound. I don't open anything up to reach them from outside.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of that is interesting to build, and all of it is necessary. That's the whole argument.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually breaks
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Sessions cache the system prompt.&lt;/strong&gt; A Discord session keeps the system prompt it started with. Add a skill, change an agent's instructions, rename it, and an existing session won't see any of it. The setting that should time sessions out doesn't do anything in the build I run. The fix is typing &lt;code&gt;/reset&lt;/code&gt; in the channel.&lt;/p&gt;

&lt;p&gt;I learned this the annoying way. I'd built a workflow for approving replies to comments on my posts, then asked the main agent to post one. The session predated the skill. Instead of saying it didn't know how, it spent 25 minutes searching the web and reading the article, and then invented a summary. Now, whenever an agent's instructions change, &lt;code&gt;/reset&lt;/code&gt; is part of the change.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A session can get poisoned.&lt;/strong&gt; My first real chat with the shopping assistant went badly. A simple request, a white T-shirt or a polo, took 43 minutes and failed. The session history held file paths that my cloud model provider's PII filter had redacted into placeholders, and the model kept copying those broken paths back into its own commands. Nothing wrong with the agent's code. The conversation itself was contaminated, and the only fix was ending the session and starting clean.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Voice notes are dropped.&lt;/strong&gt; A recorded voice message in a text channel reaches my agent platform as an empty message. The audio attachment is thrown away before the code that would detect it as voice ever runs. The feature exists in the code and doesn't engage. There's a separate path for live voice channels. I haven't tried it, so I can't tell you whether it works.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2,000 characters per message.&lt;/strong&gt; That's Discord's cap. Fine for "done" or a short list. Not fine for a daily digest. The shopping assistant's posting helper splits on line boundaries and keeps each piece under 1,900 characters, and product results go out as embeds rather than long text.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Link previews.&lt;/strong&gt; Drop a bare URL in a message and Discord unfurls it into a preview card. My first test digest was hard to read for exactly that reason. The shopping assistant now sets Discord's per-message &lt;code&gt;SUPPRESS_EMBEDS&lt;/code&gt; flag on its text posts, so links stay links.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No tables.&lt;/strong&gt; A code block is the closest you get. If an agent's output is naturally a table, it has to become a list.&lt;/p&gt;

&lt;h2&gt;
  
  
  Someone else's server
&lt;/h2&gt;

&lt;p&gt;This is the one to think about properly, and I'd rather be straight about it. My channels aren't carrying package updates and uptime pings. They carry my clothing sizes and order history, family holiday planning, and security alerts from my own network.&lt;/p&gt;

&lt;p&gt;Discord isn't end-to-end encrypted. The identity and auth are free, but everything in those channels sits on Discord's servers. For me, a private server that only my account can post in is an acceptable trade for what it saves me. It is a trade, though, and it's worth making on purpose rather than by default. If your agents handle things you wouldn't want stored on a third party's servers, this layout isn't for you, or those agents belong somewhere else.&lt;/p&gt;

&lt;h2&gt;
  
  
  The general point
&lt;/h2&gt;

&lt;p&gt;For a tool with one user, the chat app is the UI, and the dashboards are for looking. Everything I'd otherwise build — login, notifications, history, a mobile client, somewhere to drop a photo — already exists, so the time goes on the agents.&lt;/p&gt;

&lt;p&gt;What I didn't expect is that the failures are about &lt;em&gt;sessions&lt;/em&gt;, not the interface. Stale prompts, poisoned history, attachments dropped before they're processed. None of that shows up as an error in Discord. It shows up as an agent confidently doing the wrong thing, which is the thing to watch for.&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;🤖 &lt;em&gt;Drafted with AI assistance from my own homelab notes, logs and repos, then reviewed and edited before publishing.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>discord</category>
      <category>selfhosted</category>
    </item>
  </channel>
</rss>
