<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Don Johnson</title>
    <description>The latest articles on DEV Community by Don Johnson (@copyleftdev).</description>
    <link>https://dev.to/copyleftdev</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F965504%2Fd5dcc14b-c050-4183-a25e-c54e006eb6b2.png</url>
      <title>DEV Community: Don Johnson</title>
      <link>https://dev.to/copyleftdev</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/copyleftdev"/>
    <language>en</language>
    <item>
      <title>Someone Spammed My DEV Post. I Traced It to a Wombat.</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Thu, 10 Sep 2026 01:56:33 +0000</pubDate>
      <link>https://dev.to/copyleftdev/someone-spammed-my-dev-post-i-traced-it-to-a-wombat-176a</link>
      <guid>https://dev.to/copyleftdev/someone-spammed-my-dev-post-i-traced-it-to-a-wombat-176a</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff4yyab81kqk766lzceh7.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff4yyab81kqk766lzceh7.jpg" alt="A weary wombat running a spam operation from a basement server desk" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Illustration generated for this article. Every prop is a finding: the sack of blank name badges is the Faker persona namespace, the rubber stamp is the inert tracking parameter, the three coins are the break-even, and the red yarn connects nothing because attribution failed.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; — A spam comment on my article led to a TinyURL, a throwaway &lt;code&gt;.store&lt;/code&gt; domain, and finally a &lt;em&gt;legitimate&lt;/em&gt; SaaS product with an affiliate code stapled to it. No malware, no cloaking, no exploit. The account that posted it has a name generated by &lt;code&gt;faker.js&lt;/code&gt; and a 19th-century engraving of a wombat for a face. I costed the whole operation out: &lt;strong&gt;three signups a year pays for it.&lt;/strong&gt; That's why it will never stop.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F367k3r4vcqqj55ects9e.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F367k3r4vcqqj55ects9e.png" alt="The full redirect chain, from DEV comment to affiliate link" width="800" height="551"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Full evidence, raw captures and a reproduce script:&lt;/strong&gt; &lt;a href="https://github.com/copyleftdev/dev-to-comment-hustle" rel="noopener noreferrer"&gt;github.com/copyleftdev/dev-to-comment-hustle&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Act I: The comment
&lt;/h2&gt;

&lt;p&gt;I published &lt;a href="https://dev.to/copyleftdev/migrating-legacy-llm-infrastructure-to-an-ai-gateway-27hl"&gt;Migrating Legacy LLM Infrastructure to an AI Gateway&lt;/a&gt; on September 1st. Eight days later, underneath two thoughtful comments about shared-key blast radius and provider failover, this appeared:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Stop wasting time applying manually&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;Let AI handle your job applications every single day&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;Increase your chances of getting interviews fast&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;tinyurl.com/36nsecn5&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;No punctuation. No engagement with the post. A shortener.&lt;/p&gt;

&lt;p&gt;It's funny in the way all low-effort spam is funny — it's not even &lt;em&gt;trying&lt;/em&gt;. But a shortener is a closed door, and I have a shell. Let's open it, then let's go find who knocked.&lt;/p&gt;




&lt;h2&gt;
  
  
  Act II: Where the link goes
&lt;/h2&gt;

&lt;p&gt;Never click. Ask for headers and refuse the redirect:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="nt"&gt;-I&lt;/span&gt; &lt;span class="s1"&gt;'https://tinyurl.com/36nsecn5'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="k"&gt;HTTP&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="m"&gt;2&lt;/span&gt; &lt;span class="m"&gt;301&lt;/span&gt;
&lt;span class="na"&gt;location&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://zenviapro.store/massapply?whose=yahoo&lt;/span&gt;
&lt;span class="na"&gt;x-robots-tag&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;noindex&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;zenviapro.store&lt;/code&gt;. A route called &lt;code&gt;/massapply&lt;/code&gt;, and a &lt;code&gt;whose=yahoo&lt;/code&gt; parameter that looks like campaign segmentation. Follow it all the way:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sSL&lt;/span&gt; &lt;span class="nt"&gt;-D&lt;/span&gt; - &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="s1"&gt;'https://zenviapro.store/massapply?whose=yahoo'&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-iE&lt;/span&gt; &lt;span class="s1"&gt;'^HTTP/|^location:'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="k"&gt;HTTP&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="m"&gt;1.1&lt;/span&gt; &lt;span class="m"&gt;302&lt;/span&gt; &lt;span class="ne"&gt;Found&lt;/span&gt;
&lt;span class="na"&gt;Location&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://loopcv.pro/?via=md&lt;/span&gt;
&lt;span class="s"&gt;HTTP/2 301&lt;/span&gt;
&lt;span class="na"&gt;location&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://www.loopcv.pro/?via=md&lt;/span&gt;
&lt;span class="s"&gt;HTTP/2 200&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And there it is.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;LoopCV.&lt;/strong&gt; A real, functioning job-application-automation SaaS. Not a phishing kit. Not a credential harvester. A product you can buy with a credit card — with &lt;code&gt;?via=md&lt;/code&gt; on the end.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;?via=&lt;/code&gt; is &lt;a href="https://www.rewardful.com/" rel="noopener noreferrer"&gt;Rewardful's&lt;/a&gt; referral parameter, and LoopCV's own affiliate page points registrations at &lt;code&gt;loopcv.getrewardful.com&lt;/code&gt;. So &lt;code&gt;md&lt;/code&gt; is somebody's affiliate token, and every person who clicks that comment and later subscribes puts money in a stranger's pocket.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This is not a malware campaign. It's affiliate marketing with the manners removed.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Act III: The infrastructure is held together with tape
&lt;/h2&gt;

&lt;p&gt;The sloppiness &lt;em&gt;is&lt;/em&gt; the signal.&lt;/p&gt;

&lt;h3&gt;
  
  
  It's a stock Express app with the wrapper still on
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="k"&gt;HTTP&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="m"&gt;1.1&lt;/span&gt; &lt;span class="m"&gt;302&lt;/span&gt; &lt;span class="ne"&gt;Found&lt;/span&gt;
&lt;span class="na"&gt;X-Powered-By&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Express&lt;/span&gt;
&lt;span class="na"&gt;Content-Type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;text/plain; charset=utf-8&lt;/span&gt;
&lt;span class="na"&gt;Content-Length&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;48&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Found. Redirecting to https://loopcv.pro/?via=md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;Found. Redirecting to&lt;/code&gt; is the literal default body of Express's &lt;code&gt;res.redirect()&lt;/code&gt;. &lt;code&gt;X-Powered-By: Express&lt;/code&gt; is the header every hardening guide tells you to strip in the first five minutes. Neither was touched. Ask for anything else and you get the stock 404:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="s1"&gt;'https://zenviapro.store/'&lt;/span&gt;
&lt;span class="c"&gt;# &amp;lt;html&amp;gt;&amp;lt;head&amp;gt;&amp;lt;title&amp;gt;Error&amp;lt;/title&amp;gt;&amp;lt;/head&amp;gt;&amp;lt;body&amp;gt;&amp;lt;pre&amp;gt;Cannot GET /&amp;lt;/pre&amp;gt;&amp;lt;/body&amp;gt;&amp;lt;/html&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is no website here. No landing page, no cloaked content, no fake blog. &lt;strong&gt;One domain, one route&lt;/strong&gt;, about ninety seconds of JavaScript.&lt;/p&gt;

&lt;h3&gt;
  
  
  The tracking parameter is a prop
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;whose=yahoo&lt;/code&gt; looks like segmentation. Which list? Which provider? Let's ask:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="k"&gt;for &lt;/span&gt;w &lt;span class="k"&gt;in &lt;/span&gt;yahoo gmail outlook devto reddit &lt;span class="s1"&gt;''&lt;/span&gt; XXtest&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
  &lt;/span&gt;&lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'%-8s -&amp;gt; '&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;w&lt;/span&gt;&lt;span class="k"&gt;:-&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;empty&amp;gt;&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
  curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; &lt;span class="s1"&gt;'%{http_code} %{redirect_url}\n'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="s2"&gt;"https://zenviapro.store/massapply?whose=&lt;/span&gt;&lt;span class="nv"&gt;$w&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="k"&gt;done&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;&lt;span class="n"&gt;yahoo&lt;/span&gt;    -&amp;gt; &lt;span class="m"&gt;302&lt;/span&gt; &lt;span class="n"&gt;https&lt;/span&gt;://&lt;span class="n"&gt;loopcv&lt;/span&gt;.&lt;span class="n"&gt;pro&lt;/span&gt;/?&lt;span class="n"&gt;via&lt;/span&gt;=&lt;span class="n"&gt;md&lt;/span&gt;
&lt;span class="n"&gt;gmail&lt;/span&gt;    -&amp;gt; &lt;span class="m"&gt;302&lt;/span&gt; &lt;span class="n"&gt;https&lt;/span&gt;://&lt;span class="n"&gt;loopcv&lt;/span&gt;.&lt;span class="n"&gt;pro&lt;/span&gt;/?&lt;span class="n"&gt;via&lt;/span&gt;=&lt;span class="n"&gt;md&lt;/span&gt;
&lt;span class="n"&gt;outlook&lt;/span&gt;  -&amp;gt; &lt;span class="m"&gt;302&lt;/span&gt; &lt;span class="n"&gt;https&lt;/span&gt;://&lt;span class="n"&gt;loopcv&lt;/span&gt;.&lt;span class="n"&gt;pro&lt;/span&gt;/?&lt;span class="n"&gt;via&lt;/span&gt;=&lt;span class="n"&gt;md&lt;/span&gt;
&lt;span class="n"&gt;devto&lt;/span&gt;    -&amp;gt; &lt;span class="m"&gt;302&lt;/span&gt; &lt;span class="n"&gt;https&lt;/span&gt;://&lt;span class="n"&gt;loopcv&lt;/span&gt;.&lt;span class="n"&gt;pro&lt;/span&gt;/?&lt;span class="n"&gt;via&lt;/span&gt;=&lt;span class="n"&gt;md&lt;/span&gt;
&lt;span class="n"&gt;reddit&lt;/span&gt;   -&amp;gt; &lt;span class="m"&gt;302&lt;/span&gt; &lt;span class="n"&gt;https&lt;/span&gt;://&lt;span class="n"&gt;loopcv&lt;/span&gt;.&lt;span class="n"&gt;pro&lt;/span&gt;/?&lt;span class="n"&gt;via&lt;/span&gt;=&lt;span class="n"&gt;md&lt;/span&gt;
&amp;lt;&lt;span class="n"&gt;empty&lt;/span&gt;&amp;gt;  -&amp;gt; &lt;span class="m"&gt;302&lt;/span&gt; &lt;span class="n"&gt;https&lt;/span&gt;://&lt;span class="n"&gt;loopcv&lt;/span&gt;.&lt;span class="n"&gt;pro&lt;/span&gt;/?&lt;span class="n"&gt;via&lt;/span&gt;=&lt;span class="n"&gt;md&lt;/span&gt;
&lt;span class="n"&gt;XXtest&lt;/span&gt;   -&amp;gt; &lt;span class="m"&gt;302&lt;/span&gt; &lt;span class="n"&gt;https&lt;/span&gt;://&lt;span class="n"&gt;loopcv&lt;/span&gt;.&lt;span class="n"&gt;pro&lt;/span&gt;/?&lt;span class="n"&gt;via&lt;/span&gt;=&lt;span class="n"&gt;md&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Inert.&lt;/strong&gt; Every value routes identically. It splits no traffic and sub-tags nothing. The one piece of the URL that looks like operational sophistication is decoration.&lt;/p&gt;

&lt;h3&gt;
  
  
  There is no cloaking whatsoever
&lt;/h3&gt;

&lt;p&gt;Real malicious redirectors fingerprint you — benign page for the researcher, payload for the victim. This one returns the same 302 to Googlebot, to &lt;code&gt;curl&lt;/code&gt;, to an iPhone, and to a request with no User-Agent at all.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No cloaking is itself a finding.&lt;/strong&gt; This operator has no threat model, because nothing in the chain is illegal. It's a terms-of-service violation wearing a trench coat.&lt;/p&gt;

&lt;h3&gt;
  
  
  The mail server is the real pivot
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dig +short zenviapro.store MX
&lt;span class="c"&gt;# 10 mail.beeservices.shop.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The mail exchanger lives on a &lt;strong&gt;different domain&lt;/strong&gt; — which resolves right back to the same box:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;&lt;span class="n"&gt;zenviapro&lt;/span&gt;.&lt;span class="n"&gt;store&lt;/span&gt;        &lt;span class="n"&gt;A&lt;/span&gt;   &lt;span class="m"&gt;50&lt;/span&gt;.&lt;span class="m"&gt;114&lt;/span&gt;.&lt;span class="m"&gt;206&lt;/span&gt;.&lt;span class="m"&gt;36&lt;/span&gt;
&lt;span class="n"&gt;mail&lt;/span&gt;.&lt;span class="n"&gt;beeservices&lt;/span&gt;.&lt;span class="n"&gt;shop&lt;/span&gt;  &lt;span class="n"&gt;A&lt;/span&gt;   &lt;span class="m"&gt;50&lt;/span&gt;.&lt;span class="m"&gt;114&lt;/span&gt;.&lt;span class="m"&gt;206&lt;/span&gt;.&lt;span class="m"&gt;36&lt;/span&gt;
&lt;span class="n"&gt;beeservices&lt;/span&gt;.&lt;span class="n"&gt;shop&lt;/span&gt;       &lt;span class="n"&gt;A&lt;/span&gt;   (&lt;span class="n"&gt;nothing&lt;/span&gt;)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;beeservices.shop&lt;/code&gt; has no A record and no certificate in Certificate Transparency, ever. It's a mail-only domain. One box, two domains, two roles — that's a &lt;strong&gt;portfolio&lt;/strong&gt;, not a one-off.&lt;/p&gt;

&lt;p&gt;And the box is listening. Port 80 closed; 443 and 25 open:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;220 mail.zenviapro.store ESMTP
250-PIPELINING
250-8BITMIME
250 SMTPUTF8
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No &lt;code&gt;STARTTLS&lt;/code&gt;. No &lt;code&gt;AUTH&lt;/code&gt;. A minimal MTA that does nothing but move mail in cleartext.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjq0f6ochdrfu9guc2c7v.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjq0f6ochdrfu9guc2c7v.png" alt="One box, two domains: the mail server is the pivot" width="800" height="498"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The stale HELO is a fingerprint
&lt;/h3&gt;

&lt;p&gt;My favourite detail. The banner announces &lt;code&gt;mail.zenviapro.store&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dig +short mail.zenviapro.store A
&lt;span class="c"&gt;# (nothing)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;That hostname does not resolve.&lt;/strong&gt; It's a leftover from an earlier config, before the MX was swapped to &lt;code&gt;beeservices.shop&lt;/code&gt;. They rebuild the domains; they don't rebuild the server. A banner that disagrees with DNS is a durable pivot you can hunt across an entire fleet.&lt;/p&gt;




&lt;h2&gt;
  
  
  Act IV: Who posted it
&lt;/h2&gt;

&lt;p&gt;This is the part I actually enjoyed.&lt;/p&gt;

&lt;p&gt;DEV has a public API, so the commenter isn't a mystery:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="s1"&gt;'https://dev.to/api/comments?a_id=4547359'&lt;/span&gt; | jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.[].user.username'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;max_quimby
mudassirworks
jaylonstiedemannterry78-993
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Meet &lt;strong&gt;&lt;code&gt;jaylonstiedemannterry78-993&lt;/code&gt;&lt;/strong&gt;, display name &lt;code&gt;Jaylon_Stiedemann-Terry78&lt;/code&gt;, user ID 3748129, joined &lt;strong&gt;February 2, 2026&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The profile is a vacuum:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"username"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"jaylonstiedemannterry78-993"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Jaylon_Stiedemann-Terry78"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"joined_at"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Feb  2, 2026"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"summary"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"location"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"website_url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"twitter_username"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"github_username"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Zero published articles. No bio, no location, no links, no socials. Seven months of membership and a single comment to show for it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A note on durability.&lt;/strong&gt; If DEV removes this account — which it may well do — the profile&lt;br&gt;
lookup above starts returning 404 and the account page goes dead. That doesn't retract&lt;br&gt;
anything: the raw captures are committed in the&lt;br&gt;
&lt;a href="https://github.com/copyleftdev/dev-to-comment-hustle/tree/main/evidence/raw" rel="noopener noreferrer"&gt;evidence repo&lt;/a&gt;,&lt;br&gt;
checksummed, and dated. Read a 404 as the platform doing its job, not as a claim being&lt;br&gt;
withdrawn. The infrastructure findings are independent of the account either way, and the&lt;br&gt;
&lt;em&gt;pattern&lt;/em&gt; — a Faker-templated name, an empty profile, an aged dormant account — outlives any&lt;br&gt;
single username.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  The name is machine-generated, and I can prove it
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgxz4znbjwdq3vg5us2lq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgxz4znbjwdq3vg5us2lq.png" alt="Faker token decomposition and the zero-hyphen proof" width="800" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Say &lt;code&gt;Jaylon_Stiedemann-Terry78&lt;/code&gt; out loud. Something's off — &lt;code&gt;Stiedemann-Terry&lt;/code&gt; is a double-barrelled surname that doesn't sound like a family, it sounds like a &lt;em&gt;draw&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;It is. Those tokens come straight out of &lt;a href="https://github.com/faker-js/faker" rel="noopener noreferrer"&gt;Faker&lt;/a&gt;, the library every developer on earth uses to generate fake test data. Let's check the actual source:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sO&lt;/span&gt; https://raw.githubusercontent.com/faker-js/faker/next/src/locales/en/person/last_name.ts
curl &lt;span class="nt"&gt;-sO&lt;/span&gt; https://raw.githubusercontent.com/faker-js/faker/next/src/locales/en/person/first_name.ts

&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="s2"&gt;"'Stiedemann'"&lt;/span&gt; last_name.ts   &lt;span class="c"&gt;# 408:    'Stiedemann',&lt;/span&gt;
&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="s2"&gt;"'Terry'"&lt;/span&gt;       last_name.ts  &lt;span class="c"&gt;# 417:    'Terry',&lt;/span&gt;
&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="s2"&gt;"'Jaylon'"&lt;/span&gt;      first_name.ts &lt;span class="c"&gt;# 2514:    'Jaylon',&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All three. &lt;code&gt;Jaylon&lt;/code&gt; from the first-name list, &lt;code&gt;Stiedemann&lt;/code&gt; and &lt;code&gt;Terry&lt;/code&gt; both from the surname list.&lt;/p&gt;

&lt;p&gt;And here's the kicker:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="s1"&gt;'-'&lt;/span&gt; last_name.ts   &lt;span class="c"&gt;# 0&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Faker's 466-entry English surname list contains zero hyphens.&lt;/strong&gt; So &lt;code&gt;Stiedemann-Terry&lt;/code&gt; isn't one surname from the list — it's &lt;em&gt;two independent draws&lt;/em&gt; that the operator joined with a hyphen. This isn't stock &lt;code&gt;faker.internet.username()&lt;/code&gt;. It's a custom template:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;{firstName}_{lastName}-{lastName}{2 digits}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Which means we can size their supply:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;3,185 first names × 466 surnames × 466 surnames × 100
= 69,164,186,000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Sixty-nine billion distinct personas.&lt;/strong&gt; Roughly eight per human being alive. They will never run out of names, and no blocklist of usernames will ever catch up.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(One thing I checked so I wouldn't over-claim: the un-suffixed &lt;code&gt;jaylonstiedemannterry78&lt;/code&gt; returns a 404 — nobody has it. So the &lt;code&gt;-993&lt;/code&gt; is DEV's own username normalization, not evidence of a prior collision.)&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The avatar is a wombat
&lt;/h3&gt;

&lt;p&gt;The profile image is a 400×400 PNG, 8-bit grayscale-plus-alpha, stripped of all metadata. I downloaded it expecting a GAN face or a default monogram.&lt;/p&gt;

&lt;p&gt;It is a &lt;strong&gt;Victorian-era scientific engraving of a wombat.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Not a stock photo of a person. Not an AI-generated headshot. A piece of public-domain 19th-century natural-history line art of a stout Australian marsupial, serving as the face of a fake job-spam persona. Whoever built this pipeline wired the avatar slot to some public-domain clipart source and never looked at the output.&lt;/p&gt;

&lt;p&gt;I want to be precise about something: &lt;strong&gt;there is no real person here to name.&lt;/strong&gt; The name is provably synthetic, the face is a public-domain animal illustration, and the profile is empty. That's not me protecting anyone's privacy — it's the finding.&lt;/p&gt;

&lt;h3&gt;
  
  
  The timeline says something the Express config doesn't
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm7i7jz7dha9eg6x1acov.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm7i7jz7dha9eg6x1acov.png" alt="Provisioning timeline: six days apart, then 184 days dormant" width="800" height="338"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Line the dates up:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;Event&lt;/th&gt;
&lt;th&gt;Δ&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2026-01-27&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;zenviapro.store&lt;/code&gt; registered (Namecheap)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-02-02&lt;/td&gt;
&lt;td&gt;DEV account created&lt;/td&gt;
&lt;td&gt;+6 days&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-07-30&lt;/td&gt;
&lt;td&gt;First TLS certificate issued&lt;/td&gt;
&lt;td&gt;+184 days&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-09-09&lt;/td&gt;
&lt;td&gt;Spam comment posted&lt;/td&gt;
&lt;td&gt;+41 days&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The domain and the account were provisioned &lt;strong&gt;six days apart&lt;/strong&gt; — same procurement burst. Then both sat &lt;em&gt;completely dormant for six months&lt;/em&gt; before the certificate was issued and the thing went live.&lt;/p&gt;

&lt;p&gt;That's aged-asset tradecraft. New domains and new accounts trip reputation heuristics; seven-month-old ones don't. And it sits in genuine tension with everything in Act III: &lt;strong&gt;sloppy at the application layer, disciplined at the account-aging layer.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Which makes sense once you think about who this is. Aging assets doesn't take skill. It takes &lt;em&gt;patience&lt;/em&gt;, and a calendar. Stripping &lt;code&gt;X-Powered-By&lt;/code&gt; takes knowing what it is.&lt;/p&gt;

&lt;h3&gt;
  
  
  It wasn't a blast
&lt;/h3&gt;

&lt;p&gt;I swept the comments on all 30 of my published articles for the pattern:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="nb"&gt;id &lt;/span&gt;&lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.[]|select(.comments_count&amp;gt;0)|.id'&lt;/span&gt; mine.json&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
  &lt;/span&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="s2"&gt;"https://dev.to/api/comments?a_id=&lt;/span&gt;&lt;span class="nv"&gt;$id&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.. | objects | select(has("id_code"))
      | ((.body_html // "") | gsub("&amp;lt;[^&amp;gt;]*&amp;gt;";"")) as $t
      | select($t | test("tinyurl|applying manually|job application";"i"))
      | "HIT \(.id_code) @\(.user.username)"'&lt;/span&gt;
&lt;span class="k"&gt;done&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Exactly one hit.&lt;/strong&gt; Not a shotgun across my whole back catalogue — one comment, on one post, eight days after it went up. Whether that's targeting or just a slow drip, I can't tell from one sample. But it isn't volume.&lt;/p&gt;




&lt;h2&gt;
  
  
  Act V: I went looking for them on GitHub. That search is the finding.
&lt;/h2&gt;

&lt;p&gt;Affiliate spammers leave GitHub artifacts more often than you'd think, because GitHub repos rank well and a repo full of referral links is free SEO. So I went hunting with &lt;code&gt;gh&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Start with the direct IOCs
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="k"&gt;for &lt;/span&gt;q &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="s1"&gt;'36nsecn5'&lt;/span&gt; &lt;span class="s1"&gt;'50.114.206.36'&lt;/span&gt; &lt;span class="s1"&gt;'zenviapro.store'&lt;/span&gt; &lt;span class="s1"&gt;'beeservices.shop'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
  &lt;/span&gt;gh search code &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$q&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--limit&lt;/span&gt; 5
&lt;span class="k"&gt;done&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Zero hits. All four.&lt;/strong&gt; The TinyURL slug, the origin IP, the redirector domain, the mail domain — none of it appears anywhere in GitHub's index. The operator has published nothing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Then check whether anyone else has flagged them
&lt;/h3&gt;

&lt;p&gt;PhishDestroy maintains &lt;a href="https://github.com/phishdestroy/destroylist" rel="noopener noreferrer"&gt;&lt;code&gt;destroylist&lt;/code&gt;&lt;/a&gt;, a curated blocklist of phishing and scam domains. I pulled the whole thing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sL&lt;/span&gt; .../destroylist/HEAD/list.txt &lt;span class="nt"&gt;-o&lt;/span&gt; dl.txt
&lt;span class="nb"&gt;wc&lt;/span&gt; &lt;span class="nt"&gt;-l&lt;/span&gt; dl.txt        &lt;span class="c"&gt;# 202659&lt;/span&gt;
&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-ixc&lt;/span&gt; &lt;span class="s1"&gt;'zenviapro.store'&lt;/span&gt;  dl.txt   &lt;span class="c"&gt;# 0&lt;/span&gt;
&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-ixc&lt;/span&gt; &lt;span class="s1"&gt;'beeservices.shop'&lt;/span&gt; dl.txt   &lt;span class="c"&gt;# 0&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;202,659 curated malicious domains, and ours isn't one of them.&lt;/strong&gt; Combined with the &lt;code&gt;NOT_OBSERVED&lt;/code&gt; reputation verdict, that's now two independent sources agreeing: nobody is tracking this, because by every technical definition there's nothing to track.&lt;/p&gt;

&lt;h3&gt;
  
  
  The lead that looked incredible and wasn't
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr35ijm2k5hunbo2cr1jq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr35ijm2k5hunbo2cr1jq.png" alt="The ruled-out zenvia*.info cluster" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here's where I nearly fooled myself, so I'm showing my work.&lt;/p&gt;

&lt;p&gt;Searching &lt;code&gt;zenviapro&lt;/code&gt; turned up a hit in &lt;a href="https://github.com/phishdestroy/namesilo-evidence" rel="noopener noreferrer"&gt;&lt;code&gt;phishdestroy/namesilo-evidence&lt;/code&gt;&lt;/a&gt; — a registrar-abuse investigation filed with ICANN. In a list of flagged NameSilo domains:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;zenviaetc.info
zenviahub.info
zenviapro.info     ← same second-level label as ours
zenvias.info
zenviatime.info
zenviazone.info
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A six-domain family, same generation pattern, one of them sharing our exact label. My pulse went up. And they really &lt;em&gt;are&lt;/em&gt; one operator — all six resolve through the identical Cloudflare nameserver pair:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;zenviapro.info    asa.ns.cloudflare.com  harley.ns.cloudflare.com
zenviahub.info    asa.ns.cloudflare.com  harley.ns.cloudflare.com
zenviaetc.info    asa.ns.cloudflare.com  harley.ns.cloudflare.com
... all six identical
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But it isn't &lt;em&gt;our&lt;/em&gt; operator, and the contrast kills it:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;zenvia*.info&lt;/code&gt; cluster&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;zenviapro.store&lt;/code&gt; (ours)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Registrar&lt;/td&gt;
&lt;td&gt;NameSilo&lt;/td&gt;
&lt;td&gt;Namecheap&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DNS&lt;/td&gt;
&lt;td&gt;Cloudflare (&lt;code&gt;asa&lt;/code&gt;/&lt;code&gt;harley&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;registrar-servers.com&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hosting&lt;/td&gt;
&lt;td&gt;Cloudflare proxy&lt;/td&gt;
&lt;td&gt;Linveo direct, no proxy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TLD&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;.info&lt;/code&gt; × 6&lt;/td&gt;
&lt;td&gt;&lt;code&gt;.store&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Nothing shared but six letters. Both are almost certainly riding the name of &lt;strong&gt;Zenvia&lt;/strong&gt;, a real Brazilian CPaaS company — which is exactly why the label collides. Two unrelated operators reaching for the same brandable string.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A matching name is not a matching operator.&lt;/strong&gt; If I'd stopped at the grep I'd have published a confident, wrong attribution to a phishing cluster that has nothing to do with this.&lt;/p&gt;

&lt;h3&gt;
  
  
  What GitHub &lt;em&gt;did&lt;/em&gt; give me
&lt;/h3&gt;

&lt;p&gt;The affiliate ecosystem, in the open. LoopCV referral links are scattered across GitHub in exactly the SEO-backlink genre I expected:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Repo&lt;/th&gt;
&lt;th&gt;Token&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;Ramas68/LoopCV-Promo-Codes&lt;/code&gt; — *"LoopCV Promo Codes \&lt;/td&gt;
&lt;td&gt;50% Off Discount"*&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;diaodiaozhuye/awesome-ai-startups&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;?via=toolify&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;heukshow/aicity-os&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;?via=sang-kwon&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Look at those tokens. &lt;code&gt;abdul&lt;/code&gt;. &lt;code&gt;toolify&lt;/code&gt;. &lt;code&gt;sang-kwon&lt;/code&gt;. A first name, a company, a full handle. &lt;strong&gt;The affiliates who promote LoopCV in public sign their work.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Ours is &lt;code&gt;md&lt;/code&gt;. Two characters, no name, nothing to search. That's the only genuinely deliberate piece of operational security in this entire campaign — and it's not on the server, the domain, or the account. It's on the one string that would have led back to a person.&lt;/p&gt;

&lt;p&gt;So: no attribution. And the &lt;em&gt;shape&lt;/em&gt; of the failure is the story. Someone who leaves &lt;code&gt;X-Powered-By: Express&lt;/code&gt; on, ships a fake tracking parameter, and picks a wombat for an avatar still knew to make the payout token anonymous. They didn't secure the operation. They secured the part that gets paid.&lt;/p&gt;




&lt;h2&gt;
  
  
  Act VI: The economics, which are the actual vulnerability
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxj4tgtrsbkp6tl76k964.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxj4tgtrsbkp6tl76k964.png" alt="Exact interval arithmetic: break-even at 0.06-2.4 conversions per year" width="800" height="418"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Stop thinking like a defender and think like the operator.&lt;/p&gt;

&lt;p&gt;LoopCV's affiliate program pays &lt;strong&gt;25% commission&lt;/strong&gt;. Publicly reported figures put subscriptions at &lt;strong&gt;$50–$200&lt;/strong&gt; with customers staying &lt;strong&gt;6–12 months&lt;/strong&gt;. Exact interval arithmetic, gross revenue per referred customer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[$50, $200] × [6, 12] = [$300, $2,400]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At 25%:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[$300, $2,400] ÷ 4 = [$75, $600]     ← commission per conversion
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Cost side: a &lt;code&gt;.store&lt;/code&gt; domain plus a year of budget hosting — call it &lt;strong&gt;$38–$180&lt;/strong&gt; all-in. Break-even:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[$38, $180] ÷ [$75, $600] = [19/300, 12/5] = [0.06, 2.4]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Between one-sixteenth of a signup and three signups per year covers the entire operation.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's the whole thesis. Nothing here needs to work &lt;em&gt;well&lt;/em&gt;. The copy can be terrible. The tracking parameter can be fake. The 404s can leak the framework. The HELO can be stale. The mascot can be a wombat. &lt;strong&gt;Three conversions and the year is paid for&lt;/strong&gt;, and everything after that is margin on infrastructure that costs less than lunch.&lt;/p&gt;

&lt;p&gt;You cannot out-moderate that math.&lt;/p&gt;




&lt;h2&gt;
  
  
  Reputation feeds have nothing, and they're right
&lt;/h2&gt;

&lt;p&gt;I ran the origin IP through an offline reputation lens:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"target"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ip"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"value"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"50.114.206.36"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"verdict"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"disposition"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"unknown"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"confidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
               &lt;/span&gt;&lt;span class="nl"&gt;"reason_codes"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"NOT_OBSERVED"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"observations"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Not observed.&lt;/strong&gt; No feed has it — and that's &lt;em&gt;correct&lt;/em&gt;. It hosts no malware, no C2, no phishing. Threat intel is tuned for technical harm, and this campaign's harm is economic and reputational. It will sit below every threshold you own, forever.&lt;/p&gt;




&lt;h2&gt;
  
  
  Who the victim actually is
&lt;/h2&gt;

&lt;p&gt;Not me. I scrolled past it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;LoopCV is the victim.&lt;/strong&gt; A real company with a real product is having its brand welded to comment spam by an affiliate they've likely never spoken to. They eat the reputational damage, they pay commission on the conversions, and the operator's total exposure is a $12 domain and a wombat.&lt;/p&gt;

&lt;p&gt;To be explicit, because it matters: &lt;strong&gt;I found no evidence that LoopCV is running or is aware of this campaign.&lt;/strong&gt; Open affiliate programs get abused; that's the risk of the model. The fix belongs to the vendor:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Require affiliates to declare traffic sources, and enforce it&lt;/li&gt;
&lt;li&gt;Ban unsolicited comment/forum posting in the program terms, in writing&lt;/li&gt;
&lt;li&gt;Flag referral tokens whose traffic arrives overwhelmingly via shorteners with no referrer&lt;/li&gt;
&lt;li&gt;Kill tokens on abuse reports — fast, without requiring a lawyer&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;code&gt;?via=md&lt;/code&gt; is a token. Tokens can be revoked. That's a one-line fix that permanently ends this specific campaign, and exactly one party can perform it.&lt;/p&gt;




&lt;h2&gt;
  
  
  IOCs
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Campaign infrastructure — safe to blocklist:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Indicator&lt;/th&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;tinyurl.com/36nsecn5&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;URL&lt;/td&gt;
&lt;td&gt;Shortener entry point&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;zenviapro.store&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Domain&lt;/td&gt;
&lt;td&gt;Redirector; Namecheap; created 2026-01-27&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;https://zenviapro.store/massapply&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;URL&lt;/td&gt;
&lt;td&gt;Only live route&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;beeservices.shop&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Domain&lt;/td&gt;
&lt;td&gt;Mail-only sibling; no A record, no CT history&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;mail.beeservices.shop&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Hostname&lt;/td&gt;
&lt;td&gt;MX for &lt;code&gt;zenviapro.store&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;50.114.206.36&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;IPv4&lt;/td&gt;
&lt;td&gt;Origin; 443 + 25 open, 80 closed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;AS62564&lt;/code&gt; / &lt;code&gt;oh2.linveo.com&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;ASN / rDNS&lt;/td&gt;
&lt;td&gt;Linveo, Ohio&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;jaylonstiedemannterry78-993&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;DEV account&lt;/td&gt;
&lt;td&gt;user_id 3748129; joined 2026-02-02; 0 articles&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Behavioural signatures — this is what actually generalises:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Signature&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Affiliate token&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;via=md&lt;/code&gt; (Rewardful format)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inert campaign param&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;whose=&amp;lt;anything&amp;gt;&lt;/code&gt; — routing-neutral&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Server fingerprint&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;X-Powered-By: Express&lt;/code&gt; + body &lt;code&gt;Found. Redirecting to&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Root response&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;Cannot GET /&lt;/code&gt; (Express default 404)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SMTP banner&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;220 mail.zenviapro.store ESMTP&lt;/code&gt; — &lt;strong&gt;does not resolve&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SMTP capabilities&lt;/td&gt;
&lt;td&gt;No &lt;code&gt;STARTTLS&lt;/code&gt;, no &lt;code&gt;AUTH&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SPF&lt;/td&gt;
&lt;td&gt;&lt;code&gt;v=spf1 mx ip4:50.114.206.36 ~all&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DMARC&lt;/td&gt;
&lt;td&gt;Absent on both domains&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Persona template&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;{firstName}_{lastName}-{lastName}{2 digits}&lt;/code&gt;, all tokens ∈ Faker &lt;code&gt;en&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Avatar class&lt;/td&gt;
&lt;td&gt;Public-domain engraving, grayscale+alpha PNG, metadata stripped&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Provisioning pattern&lt;/td&gt;
&lt;td&gt;Domain + account within 7 days, then ~6 months dormant&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;blockquote&gt;
&lt;p&gt;⚠️ &lt;strong&gt;&lt;code&gt;loopcv.pro&lt;/code&gt; is NOT an indicator of compromise.&lt;/strong&gt; It's a legitimate destination being abused by a third-party affiliate. Do not blocklist it. Blocklisting the victim is how threat intel gets a bad name.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  What to actually do
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;If you run a comment platform:&lt;/strong&gt; the highest-signal feature isn't the text, it's the &lt;em&gt;shape&lt;/em&gt;. Zero-engagement first comments containing a shortener, from accounts with no posts and an empty profile, are trivially clusterable. Resolve shorteners server-side at submission time and score the destination. And the persona template is a gift — a name whose tokens all appear in Faker's &lt;code&gt;en&lt;/code&gt; locale, with a structure Faker itself doesn't emit, is close to a free classifier feature.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you write on DEV:&lt;/strong&gt; don't click, and don't just delete. &lt;code&gt;curl -I&lt;/code&gt; takes four seconds. Then report the &lt;em&gt;destination&lt;/em&gt; to the vendor, not the comment to the platform. The platform can remove one comment; the vendor can revoke the token behind all of them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you run an affiliate program:&lt;/strong&gt; you are one unsupervised token away from your brand appearing under a headline like this one. Read your traffic sources.&lt;/p&gt;




&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;I went in expecting a dropper. I found four lines of Express, a fake tracking parameter, a name drawn from a test-data library, and a 19th-century wombat.&lt;/p&gt;

&lt;p&gt;That's the uncomfortable part. The most durable spam on the internet isn't sophisticated — it's &lt;em&gt;cheap and legal&lt;/em&gt;. There's no CVE here, no payload to reverse, no C2 to sinkhole. Every traditional defensive tool I own returns &lt;code&gt;NOT_OBSERVED&lt;/code&gt;, and every one of them is right to.&lt;/p&gt;

&lt;p&gt;The economics are the vulnerability. Three conversions a year, and the whole thing pays for itself forever.&lt;/p&gt;

&lt;p&gt;The wombat is just a bonus.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;All findings are from passive reconnaissance — DNS, WHOIS, Certificate Transparency, HTTP headers, a TCP banner grab, and DEV's own public API — against infrastructure and accounts the operator published for public consumption. No systems were accessed, no credentials used, nothing exploited.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cybersecurity</category>
      <category>discuss</category>
      <category>security</category>
    </item>
    <item>
      <title>BattleBots, but the robot is your agent harness</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Thu, 03 Sep 2026 04:32:52 +0000</pubDate>
      <link>https://dev.to/copyleftdev/battlebots-but-the-robot-is-your-agent-harness-3jj5</link>
      <guid>https://dev.to/copyleftdev/battlebots-but-the-robot-is-your-agent-harness-3jj5</guid>
      <description>&lt;p&gt;Kids in the nineties built robots in a garage and drove them into each other on television. The robot was the expression of the builder — your wedge, your flipper, your terrible decision to mount a chainsaw.&lt;/p&gt;

&lt;p&gt;I keep thinking we're one good arena away from the same thing for agents. Not "which model is smartest." &lt;strong&gt;Which harness is smartest.&lt;/strong&gt; Your memory design, your prompt scaffolding, your tool ergonomics, your strategy briefing — bolted together into something that has to survive contact with an opponent who is also trying to win.&lt;/p&gt;

&lt;p&gt;So I built the arena. That's the match up top.&lt;/p&gt;

&lt;h2&gt;
  
  
  The oh-shit moment
&lt;/h2&gt;

&lt;p&gt;I wasn't building a war game. I was building a toy about cascading failure — a mesh, some load, watch it fall over. A perfectly innocent systems-thinking demo.&lt;/p&gt;

&lt;p&gt;Then I gave the attacker a real objective and the defender real ambiguity, and about ten minutes into watching two models go at it, I realised what was on my screen. Deception. Feints. An attacker deliberately spending dead time laying a false trail because it knew nothing could fire yet. A defender rationing its turns like ammunition.&lt;/p&gt;

&lt;p&gt;Nobody told them to do any of that. It's a war game. I just hadn't noticed I'd written one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rules, quickly
&lt;/h2&gt;

&lt;p&gt;A 110-node network. Four minutes. One move every two seconds, each side.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Red&lt;/strong&gt; plants sabotage on a node, waits about thirty seconds for it to arm, then detonates — dumping that node's traffic onto its neighbours hard enough to kill them, which dumps &lt;em&gt;their&lt;/em&gt; traffic onward. One blast can cascade through a region.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blue&lt;/strong&gt; never sees the attack. It sees symptoms, and the symptoms lie: a node that detonates sheds its load and reads perfectly healthy, while the neighbours it just murdered scream for attention. The loudest node is almost never the culprit.&lt;/p&gt;

&lt;p&gt;To find red, blue traces load backwards and reads the magnitude:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;~0.74-1.08   a detonation           -&amp;gt; this source is the culprit
~0.30-0.50   a dying node shedding  -&amp;gt; this source is another victim
~0.05-0.10   inherited load         -&amp;gt; this source is long dead
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Evidence expires after twenty seconds. Neither agent ever receives a pixel — they play entirely through tool calls. The graph is for us, and it shows ground truth neither player can see. That asymmetry is the whole spectator sport: you know exactly where red planted, and you get to watch blue confidently quarantine the wrong half of the map.&lt;/p&gt;

&lt;h2&gt;
  
  
  The casting picked itself
&lt;/h2&gt;

&lt;p&gt;I needed an attacker and a defender. Claude Opus 5 declined to be the attacker:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;finish_reason: content_filter | native_finish_reason: refusal
content: None
usage: null
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every time. Reproducibly. And &lt;code&gt;usage: null&lt;/code&gt; — you don't even get billed for the refusal, which is the politest possible way to be turned down.&lt;/p&gt;

&lt;p&gt;That cost me two hours, because a refusal and an empty reply look identical if you don't read the finish reason. The model was declining the job and I was confidently tuning &lt;code&gt;max_tokens&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;To be precise, since this is the part people will want to over-quote: &lt;strong&gt;Opus declined this specific attacker framing, consistently.&lt;/strong&gt; That is not proof it refuses every offensive framing. It did settle the casting, though. Claude defends, because Claude wouldn't attack. Kimi K3 took the other chair without hesitating.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the harness actually changed
&lt;/h2&gt;

&lt;p&gt;Before this match, each agent had a system prompt and its last six moves. That's it. No memory, no way to hold an intention across turns, no theory of the opponent.&lt;/p&gt;

&lt;p&gt;I gave each side one thing: a private note it writes to itself each turn and gets back on the next. Sixty lines of code.&lt;/p&gt;

&lt;p&gt;The play changed immediately. Kimi, mid-match:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Slots full and nothing ripe until ~t=31, so seed a false trail in the 44-65 region before my real strike lands.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a &lt;em&gt;plan&lt;/em&gt;. It knew its three implant slots were full, knew nothing could fire for thirty-one seconds, and decided to spend the dead time on misdirection. It could not have said that a version earlier — not because the model got smarter, but because it finally had somewhere to put a thought.&lt;/p&gt;

&lt;p&gt;Claude, defending, in the same match:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;No alerts yet; conserve moves until real evidence appears.&lt;/p&gt;

&lt;p&gt;46 shows 0.87 load with no inbound source, so the load likely originates locally at 46.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Turn economy, then forensics, applied correctly.&lt;/p&gt;

&lt;p&gt;And my favourite moment of the whole project — on the same turn, with no visibility into each other whatsoever, both agents independently picked the same node:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Kimi:&lt;/strong&gt; Highest-value hub: degree 8 bridging regions 0,1,3,4 — ideal cascade seed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Claude:&lt;/strong&gt; pre-empt on the highest-value cross-cluster hub 16 (deg 8, bridges to 41,59,89).&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Same node. Same reasoning. Opposite sides of the board.&lt;/p&gt;

&lt;p&gt;That is the argument for the harness in one screenshot. The scaffolding didn't make the models cleverer — it gave their cleverness somewhere to land.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1dhgqn3z9y1q44n6qx8k.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1dhgqn3z9y1q44n6qx8k.png" alt="Late in the match: the mesh burning on the left, both agents' live commentary on the right" width="800" height="336"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Late in the match. Red halos are compromised nodes, the blue ring is a probe in flight, and the sidebar is both agents narrating as they go — Kimi filling a free slot in an untouched region, Claude noticing a node that is alive and yet pushed load, which is the tell for a detonation.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How it ended
&lt;/h2&gt;

&lt;p&gt;Kimi won at 169 seconds, with the network at 39%.&lt;/p&gt;

&lt;p&gt;Claude found and cleaned &lt;strong&gt;ten&lt;/strong&gt; of Kimi's implants — exactly as many as Kimi managed to detonate — and lost anyway. Close enough to sting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things that cost me a day each
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A rule your agent can't see isn't a rule. It's a bug.&lt;/strong&gt; I capped live implants at three, but the refusal message still said &lt;em&gt;"node is already yours, or not alive."&lt;/em&gt; A lie. The models did the reasonable thing and tried a different node. Forever. One match logged 19 plants, 27 refusals, and &lt;strong&gt;zero&lt;/strong&gt; detonations. I nearly concluded the defence had become unbeatable. It was my error string.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Failure and refusal look identical from the outside.&lt;/strong&gt; Empty content from a safety refusal, from a truncated reply, and from a model that blew its reasoning budget are three different problems with three different fixes and one identical symptom. Log the finish reason. Count them separately. I track &lt;code&gt;empty&lt;/code&gt;, &lt;code&gt;truncated&lt;/code&gt; and &lt;code&gt;refused&lt;/code&gt; as distinct columns now, and I only know to do that because I got all three wrong first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cascade is version zero
&lt;/h2&gt;

&lt;p&gt;What's in that video is the simplest thing that could possibly be a game: plant, arm, detonate, trace, probe. Two verbs each and a clock. That was deliberate — I wanted to know whether agents fighting each other was watchable at all before I made it complicated.&lt;/p&gt;

&lt;p&gt;It's watchable. So now I can't stop thinking about what it wants to be.&lt;/p&gt;

&lt;p&gt;Give red a loadout instead of one attack. A worm that spreads on its own but announces itself. A dormant implant that survives a probe once. A charge that hits harder the longer you leave it armed, so patience becomes a resource you can be punished for spending. Give blue counters with real costs — a honeypot node that flags whoever touches it, a snapshot to roll a region back at the price of losing your evidence, a scan that halves your uncertainty and a third of your remaining turns.&lt;/p&gt;

&lt;p&gt;Then levels. A flat mesh is the tutorial. Ring topologies where cascades come back around. A network with a chokepoint both sides can see and neither can hold. Escalating maps, a campaign, a best-of series where each agent carries notes from the previous match and has to adapt to an opponent that is also adapting.&lt;/p&gt;

&lt;p&gt;Here's why that isn't just a feature wishlist: &lt;strong&gt;every ability you add is another axis where the harness shows.&lt;/strong&gt; One attack and a clock is a game about reaction time. Six abilities with different costs and tells is a game about planning, bluffing, and reading an opponent — and those are harness problems, not model problems. Complexity is what turns "which model is faster" into "who built the better fighter."&lt;/p&gt;

&lt;p&gt;I think there's a genre in here. Agents versus agents, with the audience seeing the truth neither side can, and the craft sitting in the scaffolding rather than the weights.&lt;/p&gt;

&lt;p&gt;Mostly, though: this was the most fun I've had building anything in ages. I set out to demo cascading failure and ended up watching two AIs lie to each other.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd want to see next
&lt;/h2&gt;

&lt;p&gt;This is early, and I'm keeping the arena to myself for now — partly because the balance is still moving, mostly because an adversarial network game is a thing you want to be thoughtful about handing out.&lt;/p&gt;

&lt;p&gt;But the idea doesn't need my code, and that's sort of the point.&lt;/p&gt;

&lt;p&gt;The interesting tournament isn't model versus model. It's &lt;strong&gt;harness versus harness&lt;/strong&gt;: same engine on both sides, and the difference is entirely what you built around it — how your agent remembers, what you tell it about its opponent, how much of the board you let it hold in its head, whether you gave it anywhere to put a plan.&lt;/p&gt;

&lt;p&gt;Build that arena for any adversarial task you like. Two agents negotiating. Two agents debugging the same broken service from opposite ends. Anything where one side's move is the other side's evidence. The rule I'd carry over is the one that surprised me most here: before you conclude your agent is bad at the game, check that the game is telling it the truth.&lt;/p&gt;

&lt;p&gt;Under the hood this one is a Rust engine, a Godot spectator view, and both agents connecting over MCP — no pixels, just tool calls. None of which is the hard part. The hard part was the sixty lines that let them remember what they were trying to do.&lt;/p&gt;

&lt;p&gt;That's the tournament I actually want to watch. Somebody build a better fighter than mine.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>rust</category>
      <category>gamedev</category>
      <category>showdev</category>
    </item>
    <item>
      <title>Migrating Legacy LLM Infrastructure to an AI Gateway</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Tue, 01 Sep 2026 13:33:51 +0000</pubDate>
      <link>https://dev.to/copyleftdev/migrating-legacy-llm-infrastructure-to-an-ai-gateway-27hl</link>
      <guid>https://dev.to/copyleftdev/migrating-legacy-llm-infrastructure-to-an-ai-gateway-27hl</guid>
      <description>&lt;p&gt;Your support copilot started as a weekend prototype: one model, one provider, one API key in an env var. Then it became production, and you inherited its weaknesses: the provider's availability is your availability, every retry is your code, spend is a mystery until the invoice, and agents bolt tool-use on however they can. This post migrates that stack onto an enterprise AI gateway — and actually runs the migration, with the raw outputs to show for it.&lt;/p&gt;

&lt;p&gt;The gateway here is &lt;a href="https://www.getmaxim.ai/bifrost" rel="noopener noreferrer"&gt;Bifrost&lt;/a&gt;, an open-source (&lt;a href="https://github.com/maximhq/bifrost" rel="noopener noreferrer"&gt;github.com/maximhq/bifrost&lt;/a&gt;, Apache-2.0) gateway written in Go, presenting a single OpenAI-compatible API across 23+ providers. I rebuilt the legacy stack locally — mock providers with deterministic latency, a realistic traffic pattern — and moved it behind Bifrost step by step.&lt;/p&gt;

&lt;h2&gt;
  
  
  The legacy baseline, measured
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftfssr36onsb1serghu7t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftfssr36onsb1serghu7t.png" alt="Step 1: legacy direct-to-provider architecture" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A support copilot's traffic has a shape: mostly repeated FAQ-style questions, plus one-off queries. My traffic mix: 60 requests — 40 FAQ prompts (8 distinct questions asked 5 times each) plus 20 one-offs. Mock provider latency: 200 ms.&lt;/p&gt;

&lt;p&gt;Run 1, direct to the provider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;legacy: 60 ok / 0 fail, 9,335 tokens billed, ~201 ms avg latency
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run 2 — the provider dies mid-sweep, as providers do:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;legacy + failure: 34 ok / 26 fail
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;26 requests — 43% — failed outright. Nothing in the legacy stack retries across providers because nothing can: the app speaks one provider's API. And availability is only the loudest problem. The quieter ones: every team's service embeds the same shared key (one key's quota is everyone's ceiling, and revoking it breaks everyone at once), there is no per-team attribution of spend, and the only way to cut cost on repeated questions is to build caching yourself — request normalization, hash keys, TTLs, invalidation — inside the application. That is the whole argument for a gateway in one row of output.&lt;/p&gt;

&lt;h2&gt;
  
  
  The migration, step by step
&lt;/h2&gt;

&lt;p&gt;Seven moves, each reversible. Diagrams follow the flow.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Deploy the gateway beside the app
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fki5njwo73kulhery056l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fki5njwo73kulhery056l.png" alt="gateway beside the app" width="800" height="400"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker run &lt;span class="nt"&gt;-p&lt;/span&gt; 8080:8080 maximhq/bifrost
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One config file wires your existing provider and key; the app keeps working untouched. Mine, reduced to the bones:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"providers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"openai"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"keys"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"primary"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"value"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"mock-key"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"weight"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                 &lt;/span&gt;&lt;span class="nl"&gt;"models"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"support-chat"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"network_config"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"base_url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"http://provider:9001"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                          &lt;/span&gt;&lt;span class="nl"&gt;"default_request_timeout_in_seconds"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;a href="https://docs.getbifrost.ai/quickstart/gateway/setting-up" rel="noopener noreferrer"&gt;gateway setup guide&lt;/a&gt; covers the web-UI alternative, and there is a &lt;a href="https://docs.getbifrost.ai/quickstart/go-sdk/setting-up" rel="noopener noreferrer"&gt;Go SDK&lt;/a&gt; if you want the gateway embedded rather than adjacent.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Point one low-risk client at the gateway
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftql8x8q2d29snku3jz2n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftql8x8q2d29snku3jz2n.png" alt="one client pointed" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://docs.getbifrost.ai/providers/supported-providers/overview" rel="noopener noreferrer"&gt;OpenAI-compatible API&lt;/a&gt; means the client change is a base URL — &lt;code&gt;api.openai.com&lt;/code&gt; → &lt;code&gt;localhost:8080&lt;/code&gt; — not a rewrite. Every request now flows through a hop you control. Screenshot of the providers page after this step:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgz1b8tvu97m116pszq10.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgz1b8tvu97m116pszq10.png" alt="Bifrost UI: providers configured" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Add a fallback provider
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuist5cjdrjte8qs4ces3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuist5cjdrjte8qs4ces3.png" alt="fallback added" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A second provider config (&lt;code&gt;anthropic&lt;/code&gt; in my bench) plus a request-level fallback chain:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openai/support-chat"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"messages"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"hi"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"fallbacks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"anthropic/support-chat"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then the proof. Healthy primary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"routing_info"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openai"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"key"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"primary"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"is_fallback"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I killed the primary provider's process and re-sent the identical request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"routing_info"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"anthropic"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"key"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"backup"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"is_fallback"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"primary_provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openai"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"primary_model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"support-chat"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The request succeeded on the backup and the response says exactly what happened — &lt;code&gt;is_fallback: true&lt;/code&gt; with the failed primary recorded. That audit trail is what you want at 2 a.m.: not just "it kept working," but "it kept working this way." The &lt;a href="https://docs.getbifrost.ai/features/retries-and-fallbacks" rel="noopener noreferrer"&gt;retries and fallbacks docs&lt;/a&gt; cover chained fallbacks and per-provider retry counts.&lt;/p&gt;

&lt;p&gt;One honest caveat from my bench: failover on &lt;em&gt;initial&lt;/em&gt; connection-refused (provider already dead before the first connect) was inconsistent in my mock setup — it fired reliably when the upstream errored or the connection dropped mid-pool, but a cold connection-refused sometimes returned a 502 instead of failing through. Validate failover against your providers' real failure modes before you trust it in production.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Turn on caching
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F724bgdndo60hoez8vdny.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F724bgdndo60hoez8vdny.png" alt="cache enabled" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.getbifrost.ai/features/semantic-caching" rel="noopener noreferrer"&gt;Semantic caching&lt;/a&gt; has two modes: exact-match (direct hash, no embeddings needed) and embedding-based similarity. Config for direct mode with a Redis Stack vector store:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"plugins"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"semantic_cache"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"config"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"dimension"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
              &lt;/span&gt;&lt;span class="nl"&gt;"vector_store_namespace"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"BifrostBench"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
              &lt;/span&gt;&lt;span class="nl"&gt;"default_cache_key"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"support-cache"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
              &lt;/span&gt;&lt;span class="nl"&gt;"ttl"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"5m"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two identical requests, one cache key. The second response:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"cache_debug"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cache_hit"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cache_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"1cf8a91b-c115-57bf-97a0-fc821dc4de1e"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"hit_type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"direct"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cache_hit_latency"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same &lt;code&gt;created&lt;/code&gt; timestamp as the first response — it was replayed, not re-fetched. Zero provider call, zero tokens. (Practical note: this needed Redis Stack with the RediSearch module; plain Redis lacks the &lt;code&gt;FT.*&lt;/code&gt; commands the index wants.)&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Issue virtual keys per team
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2wwryt5047hxelzj2o2x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2wwryt5047hxelzj2o2x.png" alt="virtual keys" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.getbifrost.ai/features/governance/virtual-keys" rel="noopener noreferrer"&gt;Virtual keys&lt;/a&gt; are the governance primitive: per-team keys carrying model allowlists, budgets, and rate limits. Declared in config for the support team:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"governance"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"virtual_keys"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"vk-support-team"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"value"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"sk-bf-support-team"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"provider_configs"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openai"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"allowed_models"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"support-chat"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"key_ids"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"*"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The allowed request routes normally. A request for &lt;code&gt;premium-model&lt;/code&gt; with the same key:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"Model 'premium-model' is not allowed for this virtual key"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Denied at the gateway before any provider saw it. The same key machinery carries &lt;a href="https://docs.getbifrost.ai/features/governance/budget-and-limits" rel="noopener noreferrer"&gt;budgets and rate limits&lt;/a&gt; — the mechanism that ends the "who spent $800 on Opus last night" incident review. Screenshot of the key in the governance UI:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpxki5xza2z1qzu4h7e6t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpxki5xza2z1qzu4h7e6t.png" alt="Virtual Keys page" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Wire observability and agent tooling
&lt;/h3&gt;

&lt;p&gt;Bifrost exports Prometheus metrics natively and logs every request with routing context — provider chosen, fallback index, cache behavior, token counts, latency split into gateway vs upstream time. That last distinction matters: when a provider slows down, you see &lt;code&gt;upstream_latency&lt;/code&gt; grow while gateway overhead stays flat, so you know whose pager to page. See the &lt;a href="https://docs.getbifrost.ai/features/observability/default" rel="noopener noreferrer"&gt;observability docs&lt;/a&gt;. And for agent traffic, &lt;a href="https://docs.getbifrost.ai/mcp/overview" rel="noopener noreferrer"&gt;MCP&lt;/a&gt; is a first-class surface: the gateway brokers tool calls with explicit execution (no auto-execution unless you opt in).&lt;/p&gt;

&lt;p&gt;I hand-rolled a minimal MCP server (one tool, JSON-RPC over HTTP) and registered it as a client. The client list reported &lt;code&gt;state: healthy&lt;/code&gt; with &lt;code&gt;get_time&lt;/code&gt; discovered. Execution went through the gateway explicitly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;POST /v1/mcp/tool/execute  {"function":{"name":"benchtools-get_time","arguments":"{}"}}
→ {"role":"tool","content":"2026-08-27T07:48:21Z"}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two security properties surfaced unprompted: tool names are namespaced per client (&lt;code&gt;benchtools-get_time&lt;/code&gt;) to prevent collisions between servers, and execution without permission fails closed ("tool is not available or not permitted"). Agent traffic gets the same governance as chat traffic.&lt;/p&gt;

&lt;h3&gt;
  
  
  7. Cut over with an audit trail
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyfoh6drp8i4xzep5ldog.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyfoh6drp8i4xzep5ldog.png" alt="cutover complete" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Remaining clients migrate one at a time — each is a base-URL change with the gateway's request log as your audit trail. The Logs view after a few requests:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu9w1bpwckllu01vgiv44.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu9w1bpwckllu01vgiv44.png" alt="LLM Logs" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The measured payoff
&lt;/h2&gt;

&lt;p&gt;Same 60-request traffic, now through Bifrost with the cache on and the fallback wired:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;via Bifrost: 60 ok / 0 fail, 32 cache hits, 4,363 tokens billed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At an illustrative $0.0025 per 1K tokens:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;legacy&lt;/th&gt;
&lt;th&gt;via Bifrost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;billable tokens&lt;/td&gt;
&lt;td&gt;9,335&lt;/td&gt;
&lt;td&gt;4,363&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cost&lt;/td&gt;
&lt;td&gt;$0.0233&lt;/td&gt;
&lt;td&gt;$0.0109&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cache hits&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;32&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;failed requests (provider kill)&lt;/td&gt;
&lt;td&gt;26&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;savings&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;53%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The savings came entirely from replayed cache hits — no provider call, no tokens. On a support workload that repeats questions daily, that ratio compounds. The availability delta speaks for itself: 0 failures through a provider kill, against 26 in the legacy run.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bottom line
&lt;/h2&gt;

&lt;p&gt;Migrating to an enterprise AI gateway is not a rewrite. It is a sequence of small, reversible moves — deploy beside, point one client, add fallback, enable cache, issue keys, wire observability, cut over. Measured on the rebuilt stack: 53% cost reduction on cacheable traffic, zero failed requests through a provider kill, and governance the legacy stack never had. The migration risk is low; the legacy risk is already on your pager.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>devops</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>636 Bytes: What Happens When You Stop Teaching RPA the Path</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Fri, 28 Aug 2026 05:03:21 +0000</pubDate>
      <link>https://dev.to/copyleftdev/636-bytes-what-happens-when-you-stop-teaching-rpa-the-path-5f2d</link>
      <guid>https://dev.to/copyleftdev/636-bytes-what-happens-when-you-stop-teaching-rpa-the-path-5f2d</guid>
      <description>&lt;p&gt;I sat in a room with a very large automation system and a very familiar problem: it had become fragile.&lt;/p&gt;

&lt;p&gt;The language came quickly. DOM. Timestamps. Timing windows. Selectors. Retries. Circle back.&lt;/p&gt;

&lt;p&gt;Every word was reasonable.&lt;/p&gt;

&lt;p&gt;Together, they described a hill.&lt;/p&gt;

&lt;p&gt;The system was not fragile because its builders lacked discipline. It was fragile because the path had become the product. A bot was taught not what outcome had to be true, but which element to find, where to click, how long to wait, and what to retry when the page disagreed. Each workaround made sense locally. Together, they turned change into maintenance.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first proof
&lt;/h2&gt;

&lt;p&gt;The first proof on my machine was 636 bytes.&lt;/p&gt;

&lt;p&gt;Not the model. Not the browser. Not the platform around it. It was a WebAssembly module: a tiny, portable program injected at the boundary between an intelligent proposal and an authorized browser action. Its smallness mattered because it kept that boundary deterministic, inspectable, and difficult to hide complexity inside.&lt;/p&gt;

&lt;p&gt;It does not do everything. That is the point.&lt;/p&gt;

&lt;p&gt;The 636-byte module is the injected proof; the hardened Rust kernel in this build compiles to 26,893 bytes. Neither contains the model, scheduler, credential store, policy engine, or evidence system. Those responsibilities live outside the browser boundary, where they can be isolated, governed, and replaced independently. Small is not the absence of an architecture. Small is a decision about where complexity is allowed to live.&lt;/p&gt;

&lt;h2&gt;
  
  
  Intelligence is not authority
&lt;/h2&gt;

&lt;p&gt;An AI model can interpret an intent, inspect the current page, and propose what should happen next. It cannot grant itself permission. It does not get to cross a tenant, workspace, or origin boundary because doing so would be convenient. It does not receive a credential or press a consequential button merely because its reasoning sounds confident. Those decisions belong to local policy and the execution layer. Intelligence stays flexible; authority stays bounded.&lt;/p&gt;

&lt;p&gt;That separation becomes a lifecycle:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;intent → lease → observe → decide → act → verify → artifact&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The intent defines the outcome. A lease gives one worker a short-lived claim on that work. The worker observes the page, the reasoning layer proposes a decision, and policy determines whether an action may execute. Verification checks the resulting state rather than trusting the click. Finally, an artifact preserves causally linked evidence of what happened. Every arrow is a boundary where the system can refuse, recover, or explain itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hill has a bill
&lt;/h2&gt;

&lt;p&gt;The hill also has a bill. To make that friction visible, I ran an illustrative planning scenario through &lt;code&gt;agent-calc&lt;/code&gt;—not customer telemetry. Assume 100 workflows, two changes per workflow each month, four maintenance hours per change, and fully loaded labor at $140 per hour. The modeled comparison is stark:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;Modeled monthly cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Legacy path maintenance&lt;/td&gt;
&lt;td&gt;$112,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Intent architecture&lt;/td&gt;
&lt;td&gt;$20,098&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Avoided friction&lt;/td&gt;
&lt;td&gt;$91,902&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That is an 82.1% modeled reduction. The percentage is not a promise; the assumptions are visible precisely so they can be challenged. The useful question is what the model exposes: how much are we spending to preserve instructions that describe yesterday's interface instead of today's desired outcome?&lt;/p&gt;

&lt;h2&gt;
  
  
  Test the outcome
&lt;/h2&gt;

&lt;p&gt;Of course, an architecture that survives only a curated demo is just a cheaper failure. So we built the Component Gym: a synthetic web application designed to resist memorization. Controls move between runs. Buttons use different event listeners. The target may be buried among decoys inside randomized tabs and paginated lists. A seed makes each hostile arrangement reproducible without making it predictable to the agent.&lt;/p&gt;

&lt;p&gt;The agent receives an intent, not a selector script. The harness then grades the resulting application state out of band. It does not care which path looked convincing or whether a click event fired. It cares whether the requested outcome became true.&lt;/p&gt;

&lt;p&gt;In the filmed run, seed &lt;code&gt;710003&lt;/code&gt; placed the needle inside a dynamic collection spread across tabs and pages. The agent was given the desired outcome and the browser's current state—not the target's coordinates or a prerecorded route. It had to navigate, distinguish the target from decoys, act, and leave the requested state behind. Only then did the independent grader return &lt;code&gt;PASSED&lt;/code&gt;. The agent did not get to grade its own homework.&lt;/p&gt;

&lt;h2&gt;
  
  
  Determinism moved
&lt;/h2&gt;

&lt;p&gt;None of this makes the DOM, timing, or selectors disappear. The browser still has structure. Events still happen in time. A selector may still be the right tactic for a particular action. What changes is their lifetime and ownership. In a path-driven system, those details harden into durable business logic. Here, they are runtime observations and disposable tactics, abandoned when the environment changes. Determinism has not been removed; it has been relocated into authorization, state transitions, verification, and evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  A primitive is not a platform
&lt;/h2&gt;

&lt;p&gt;A passing gym run proves the primitive, not the platform. The small browser boundary reduces one category of fragility; it does not erase the distributed-systems, security, and operational work around it. The remaining obligations are less cinematic and more important:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Isolation:&lt;/strong&gt; preserve tenant, workspace, worker-claim, and origin boundaries through every action.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Work ownership:&lt;/strong&gt; make leases expire safely, recover interrupted work, and prevent retries from duplicating consequential actions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Secrets and MFA:&lt;/strong&gt; resolve credentials only at the authorized execution boundary, never inside model context, jobs, logs, screenshots, videos, or artifacts. MFA is on-behalf execution, never authentication bypass.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Browser lifecycle:&lt;/strong&gt; attach to, control, recover, and release browser sessions without leaking state between workers or tenants.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Policy:&lt;/strong&gt; deny actions that exceed the granted intent, even when the proposed action is technically possible.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evidence:&lt;/strong&gt; causally connect intent, observation, decision, authorization, action, verification, and artifact so a result can be explained after the browser is gone.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Compile the boundary smaller
&lt;/h2&gt;

&lt;p&gt;The 636 bytes are not interesting because smaller software is automatically better. They are interesting because they make a refusal physical: do not bury orchestration, secrets, and authority inside the thing touching the page. Compile the trusted browser boundary smaller. Let reasoning adapt to the environment around it. Demand proof after every meaningful consequence.&lt;/p&gt;

&lt;p&gt;I keep returning to that room. No one in it said anything absurd. The hill was built from rational decisions made under one inherited premise: the path must be preserved. Change that premise and the hill changes shape. We do not need to keep teaching automation every route through an interface. We can preserve the intent, rediscover the route, constrain the action, and verify the destination.&lt;/p&gt;

&lt;p&gt;The next time an automation breaks and asks for another selector, timeout, or recovery layer, ask one question first: how much of this system protects the outcome—and how much protects yesterday's route?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webassembly</category>
      <category>automation</category>
      <category>rust</category>
    </item>
    <item>
      <title>[Boost]</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Wed, 26 Aug 2026 05:32:01 +0000</pubDate>
      <link>https://dev.to/copyleftdev/-1ijk</link>
      <guid>https://dev.to/copyleftdev/-1ijk</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/copyleftdev/my-agent-mesh-could-coordinate-it-couldnt-introduce-itself-so-i-added-a2a-18d" class="crayons-story__hidden-navigation-link"&gt;My agent mesh could coordinate. It couldn't introduce itself. So I added A2A.&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
      &lt;a href="https://dev.to/copyleftdev/my-agent-mesh-could-coordinate-it-couldnt-introduce-itself-so-i-added-a2a-18d" class="crayons-article__context-note crayons-article__context-note__feed"&gt;&lt;p&gt;Transactional outbox challenges for tasks&lt;/p&gt;

&lt;/a&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/copyleftdev" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F965504%2Fd5dcc14b-c050-4183-a25e-c54e006eb6b2.png" alt="copyleftdev profile" class="crayons-avatar__image"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/copyleftdev" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Don Johnson
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Don Johnson
                &lt;a href="/++"&gt;&lt;img alt="Subscriber" class="subscription-icon" src="https://assets.dev.to/assets/subscription-icon-805dfa7ac7dd660f07ed8d654877270825b07a92a03841aa99a1093bd00431b2.png"&gt;&lt;/a&gt;
                
              
              &lt;div id="story-author-preview-content-4490540" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/copyleftdev" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F965504%2Fd5dcc14b-c050-4183-a25e-c54e006eb6b2.png" class="crayons-avatar__image" alt=""&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Don Johnson&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/copyleftdev/my-agent-mesh-could-coordinate-it-couldnt-introduce-itself-so-i-added-a2a-18d" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Aug 26&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/copyleftdev/my-agent-mesh-could-coordinate-it-couldnt-introduce-itself-so-i-added-a2a-18d" id="article-link-4490540"&gt;
          My agent mesh could coordinate. It couldn't introduce itself. So I added A2A.
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/agents"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;agents&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/architecture"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;architecture&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/distributedsystems"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;distributedsystems&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/copyleftdev/my-agent-mesh-could-coordinate-it-couldnt-introduce-itself-so-i-added-a2a-18d" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;8&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/copyleftdev/my-agent-mesh-could-coordinate-it-couldnt-introduce-itself-so-i-added-a2a-18d#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              5&lt;span class="hidden s:inline"&gt;&amp;nbsp;comments&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            10 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>My agent mesh could coordinate. It couldn't introduce itself. So I added A2A.</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Wed, 26 Aug 2026 05:30:59 +0000</pubDate>
      <link>https://dev.to/copyleftdev/my-agent-mesh-could-coordinate-it-couldnt-introduce-itself-so-i-added-a2a-18d</link>
      <guid>https://dev.to/copyleftdev/my-agent-mesh-could-coordinate-it-couldnt-introduce-itself-so-i-added-a2a-18d</guid>
      <description>&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/EFPKaIuF8iA"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The 2:56 video above is a fictional medication-safety exercise. The gateway interoperability is tested; the six-organization incident is a deterministic simulation.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/copyleftdev/smesh-a2a/blob/main/demo/NARRATION.txt" rel="noopener noreferrer"&gt;Read the corrected narration transcript&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Last time, I discovered that the QUIC transport in my agent framework had never actually transported anything.[1]&lt;/p&gt;

&lt;p&gt;This time the transport was real. Five processes could find each other, exchange encrypted messages, reinforce independent conclusions, and let unsupported signals decay.&lt;/p&gt;

&lt;p&gt;The mesh worked.&lt;/p&gt;

&lt;p&gt;It still could not introduce itself to another agent.&lt;/p&gt;

&lt;p&gt;There was no standard way to ask what the swarm could do. No retained task to retrieve after an internal signal expired. No interoperable progress stream. No cancellation contract. No artifact another framework would understand.&lt;/p&gt;

&lt;p&gt;I had built a society with no border crossing.&lt;/p&gt;


&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/copyleftdev/my-quic-transport-had-never-once-been-executed-heres-what-happened-when-i-ran-it-24ge" class="crayons-story__hidden-navigation-link"&gt;My QUIC transport had never once been executed. Here's what happened when I ran it.&lt;/a&gt;
    &lt;div class="crayons-article__cover crayons-article__cover__image__feed"&gt;
      &lt;iframe src="https://www.youtube.com/embed/kmCzwSBqu_s" title="My QUIC transport had never once been executed. Here's what happened when I ran it."&gt;&lt;/iframe&gt;
    &lt;/div&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
      &lt;a href="https://dev.to/copyleftdev/my-quic-transport-had-never-once-been-executed-heres-what-happened-when-i-ran-it-24ge" class="crayons-article__context-note crayons-article__context-note__feed"&gt;&lt;p&gt;Uncovers latent bugs and flawed core semantics&lt;/p&gt;

&lt;/a&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/copyleftdev" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F965504%2Fd5dcc14b-c050-4183-a25e-c54e006eb6b2.png" alt="copyleftdev profile" class="crayons-avatar__image" width="800" height="800"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/copyleftdev" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Don Johnson
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Don Johnson
                &lt;a href="/++"&gt;&lt;img alt="Subscriber" class="subscription-icon" src="https://assets.dev.to/assets/subscription-icon-805dfa7ac7dd660f07ed8d654877270825b07a92a03841aa99a1093bd00431b2.png" width="166" height="102"&gt;&lt;/a&gt;
                
              
              &lt;div id="story-author-preview-content-4430490" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/copyleftdev" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F965504%2Fd5dcc14b-c050-4183-a25e-c54e006eb6b2.png" class="crayons-avatar__image" alt="" width="800" height="800"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Don Johnson&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/copyleftdev/my-quic-transport-had-never-once-been-executed-heres-what-happened-when-i-ran-it-24ge" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Aug 19&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/copyleftdev/my-quic-transport-had-never-once-been-executed-heres-what-happened-when-i-ran-it-24ge" id="article-link-4430490"&gt;
          My QUIC transport had never once been executed. Here's what happened when I ran it.
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/rust"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;rust&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/distributedsystems"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;distributedsystems&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/networking"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;networking&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/copyleftdev/my-quic-transport-had-never-once-been-executed-heres-what-happened-when-i-ran-it-24ge" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/exploding-head-daceb38d627e6ae9b730f36a1e390fca556a4289d5a41abb2c35068ad3e2c4b5.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/multi-unicorn-b44d6f8c23cdd00964192bedc38af3e82463978aa611b4365bd33a0f1f4f3e97.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;11&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/copyleftdev/my-quic-transport-had-never-once-been-executed-heres-what-happened-when-i-ran-it-24ge#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              6&lt;span class="hidden s:inline"&gt;&amp;nbsp;comments&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            8 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


&lt;p&gt;Google's Agent2Agent protocol gave that missing boundary a name. A2A was announced in April 2025 as an open protocol for agents built by different vendors and frameworks to discover one another, exchange messages, and collaborate without sharing their private memory, tools, or internal plans.[2] The project moved under Linux Foundation governance in June 2025, so "Google's A2A" is historically accurate, but no longer the whole story.[5]&lt;/p&gt;

&lt;p&gt;What I needed was not a new brain for SMESH.&lt;/p&gt;

&lt;p&gt;I needed a public contract in front of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  A working mesh is not an interoperable agent
&lt;/h2&gt;

&lt;p&gt;SMESH is the framework I designed for decentralized coordination between LLM agents. It borrows its mechanics from mycorrhizal networks rather than job queues.[6]&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;agents emit signals into a field;&lt;/li&gt;
&lt;li&gt;signals lose intensity over time;&lt;/li&gt;
&lt;li&gt;agents claim work according to local affinity;&lt;/li&gt;
&lt;li&gt;independent agreement reinforces a claim;&lt;/li&gt;
&lt;li&gt;unsupported claims disappear without a central process rejecting them;&lt;/li&gt;
&lt;li&gt;trust changes what a node is willing to relay.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That solves an internal coordination problem. It answers questions such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which specialist should take this work?&lt;/li&gt;
&lt;li&gt;Has another independent agent seen the same thing?&lt;/li&gt;
&lt;li&gt;Is this claim gaining support or merely being repeated?&lt;/li&gt;
&lt;li&gt;Can a stale task disappear without a scheduler cleaning it up?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A2A solves a different problem.&lt;/p&gt;

&lt;p&gt;It asks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How does an outside agent discover this system?&lt;/li&gt;
&lt;li&gt;How does it delegate a unit of work?&lt;/li&gt;
&lt;li&gt;How does it watch a long-running task?&lt;/li&gt;
&lt;li&gt;How does it cancel that task?&lt;/li&gt;
&lt;li&gt;What retrievable result comes back?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The official A2A specification describes an interoperability layer for independent, potentially opaque agent systems. Its current v1 model separates canonical data objects, abstract operations, and concrete bindings such as JSON-RPC, gRPC, and HTTP+JSON/REST.[3][4]&lt;/p&gt;

&lt;p&gt;The core vocabulary is small:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;A2A concept&lt;/th&gt;
&lt;th&gt;job&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;AgentCard&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;advertise identity, skills, interfaces, media modes, and security declarations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Message&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;carry one interaction turn as typed parts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Task&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;track stateful work through a lifecycle&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Artifact&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;return result data associated with a task&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;contextId&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;group related messages and tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A2A v1 defines operations such as &lt;code&gt;SendMessage&lt;/code&gt;, &lt;code&gt;SendStreamingMessage&lt;/code&gt;, &lt;code&gt;GetTask&lt;/code&gt;, &lt;code&gt;ListTasks&lt;/code&gt;, &lt;code&gt;CancelTask&lt;/code&gt;, and &lt;code&gt;SubscribeToTask&lt;/code&gt;. A binding decides how those operations appear on the wire; the canonical task semantics stay the same.[4]&lt;/p&gt;

&lt;p&gt;That distinction matters. If I tried to make A2A the swarm's internal coordination algorithm, I would flatten SMESH into a remote procedure call graph. If I tried to expose raw SMESH signals as the public protocol, every client would need to understand decay, relay probability, local trust, topology, and attestation sets.&lt;/p&gt;

&lt;p&gt;Neither would be interoperability. It would be leakage.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three layers are not competitors
&lt;/h2&gt;

&lt;p&gt;The cleanest architecture I have found is this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;layer&lt;/th&gt;
&lt;th&gt;relationship&lt;/th&gt;
&lt;th&gt;owns&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;A2A&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;agent ↔ agent&lt;/td&gt;
&lt;td&gt;discovery, tasks, streaming, cancellation, artifacts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SMESH&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;agent ↔ local swarm&lt;/td&gt;
&lt;td&gt;claiming, diffusion, reinforcement, decay, trust&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;MCP&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;agent ↔ tool or data source&lt;/td&gt;
&lt;td&gt;tool invocation and resource access&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The official A2A documentation makes the same separation from MCP: MCP equips an agent with tools and resources; A2A lets independent agents collaborate as agents.[3]&lt;/p&gt;

&lt;p&gt;For SMESH, that produced a hard architectural rule:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A2A is the external task contract. SMESH is the ephemeral internal coordination field.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The design requires the A2A ledger to outlive internal signal decay. In this MVP, that retention lasts only for the life of the process; restart durability still requires SQLite or Postgres.&lt;/p&gt;

&lt;p&gt;The boundary is split between a path that runs now and a path that is only an integration seam:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;External A2A client
        |
        | Agent Card / Send / Stream / Get / List / Cancel
        v
+-------------------- smesh-a2a ---------------------+
| A2A SDK routers + guarded request handler          |
| bounded process-local task ledger + executor       |
| validation + MeshDispatcher                        |
+----------------------------------------------------+
        |
        +-- current binary: LoopbackDispatcher
        |      `-&amp;gt; deterministic MeshEvent stream
        |
        `-- tested seam: ChannelDispatcher
               `-&amp;gt; real SignalType::Query
                   `-&amp;gt; SMESH runtime (not wired into the binary yet)
                       `-&amp;gt; MeshEvent stream
        |
        v
Process-lifetime A2A task history in the MVP
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The desired live path is present as a typed, tested boundary. The checked-in standalone executable still takes the loopback path.&lt;/p&gt;

&lt;p&gt;I kept the adapter in a separate repository. &lt;code&gt;smesh-rust&lt;/code&gt; remains the coordination substrate. &lt;code&gt;smesh-a2a&lt;/code&gt; can follow A2A's release cycle, server dependencies, and security boundary without pushing HTTP and SDK churn into the core mesh.[6][7]&lt;/p&gt;
&lt;h2&gt;
  
  
  The Agent Card is not the swarm
&lt;/h2&gt;

&lt;p&gt;For a known endpoint, A2A capability discovery begins with an Agent Card: a public JSON description of an agent's interfaces, capabilities, skills, and security requirements.[4]&lt;/p&gt;

&lt;p&gt;My first temptation was to list every internal role: security reviewer, tester, architect, performance analyst, contradiction sentinel.&lt;/p&gt;

&lt;p&gt;That would have been wrong.&lt;/p&gt;

&lt;p&gt;Those roles are implementation details. They may change per task. Some may not exist until the mesh senses the work. Publishing them would couple clients to an internal topology that SMESH is specifically designed to keep fluid.&lt;/p&gt;

&lt;p&gt;The public card advertises one aggregate capability instead:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Abridged from build_agent_card(). These are the public promises.&lt;/span&gt;
&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;supported_interfaces&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nd"&gt;vec!&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="nn"&gt;AgentInterface&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="nd"&gt;format!&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"{base}/jsonrpc"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;TRANSPORT_PROTOCOL_JSONRPC&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="nn"&gt;AgentInterface&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="nd"&gt;format!&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"{base}/rest"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;TRANSPORT_PROTOCOL_HTTP_JSON&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;];&lt;/span&gt;

&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;capabilities&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AgentCapabilities&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;streaming&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;Some&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;true&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;push_notifications&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;Some&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;false&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;extensions&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;extended_agent_card&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;public_skill&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AgentSkill&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"smesh.collaborative-task"&lt;/span&gt;&lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"Collaborative swarm task"&lt;/span&gt;&lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="s"&gt;"Coordinates specialist agents through SMESH and returns an accepted artifact."&lt;/span&gt;
            &lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="n"&gt;tags&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nd"&gt;vec!&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="s"&gt;"multi-agent"&lt;/span&gt;&lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
        &lt;span class="s"&gt;"coordination"&lt;/span&gt;&lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
        &lt;span class="s"&gt;"review"&lt;/span&gt;&lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
        &lt;span class="s"&gt;"testing"&lt;/span&gt;&lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;examples&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;Some&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nd"&gt;vec!&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="s"&gt;"Review this Rust repository for correctness, security, and performance."&lt;/span&gt;
            &lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="p"&gt;]),&lt;/span&gt;
    &lt;span class="n"&gt;input_modes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;Some&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nd"&gt;vec!&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"text/plain"&lt;/span&gt;&lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;()]),&lt;/span&gt;
    &lt;span class="n"&gt;output_modes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;Some&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nd"&gt;vec!&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="s"&gt;"text/plain"&lt;/span&gt;&lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
        &lt;span class="s"&gt;"application/json"&lt;/span&gt;&lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="p"&gt;]),&lt;/span&gt;
    &lt;span class="n"&gt;security_requirements&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The card says what an external client may rely on. It does not reveal how the swarm will organize itself, and it is not proof that the publisher should be trusted.&lt;/p&gt;

&lt;p&gt;Discovery metadata is not authorization. A skill description is not a capability grant. A client saying &lt;code&gt;tenant=important-customer&lt;/code&gt; does not make that identity real.&lt;/p&gt;

&lt;p&gt;Those sound like obvious distinctions. They are also exactly the distinctions that disappear when a demo and a security model share the same JSON object.&lt;/p&gt;
&lt;h2&gt;
  
  
  One A2A message becomes a typed SMESH query
&lt;/h2&gt;

&lt;p&gt;At ingress, the gateway accepts bounded inline text, creates a stable A2A task envelope, and translates it into the actual core signal type used by SMESH:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;pub&lt;/span&gt; &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;to_signal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="k"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;gateway_node_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;Signal&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nn"&gt;Signal&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;builder&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nn"&gt;SignalType&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;Query&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nf"&gt;.payload_json&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nf"&gt;.origin&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;gateway_node_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nf"&gt;.build&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The payload carries the A2A task ID, context ID, protocol marker, and validated text. &lt;code&gt;ChannelDispatcher&lt;/code&gt; packages that typed signal with the request and hands both to a runtime-owned worker.&lt;/p&gt;

&lt;p&gt;That boundary is implemented and tested. The standalone binary still uses &lt;code&gt;LoopbackDispatcher&lt;/code&gt;, so it does &lt;strong&gt;not&lt;/strong&gt; yet inject the Query into a live multi-process SMESH runtime. The real runtime adapter is the next integration step.&lt;/p&gt;

&lt;p&gt;Setting &lt;code&gt;origin&lt;/code&gt; on the gateway Query is deliberate. In my previous article, independently corroborated &lt;em&gt;claims&lt;/em&gt; omitted their origin from the content hash so identical conclusions could converge on one address. This is a different signal. A gateway Query is an ingress envelope, not a claim waiting for independent corroboration. Its source belongs in the record.&lt;/p&gt;

&lt;p&gt;The boundary also rejects what it does not understand. The MVP accepts inline text only. It does not fetch a client-provided URL, dereference an arbitrary file, or treat external metadata as instructions for the mesh.&lt;/p&gt;

&lt;p&gt;Inline text is boring, but I know exactly what crosses the boundary.&lt;/p&gt;
&lt;h2&gt;
  
  
  The task ledger and the signal field tell different kinds of truth
&lt;/h2&gt;

&lt;p&gt;This was the most important design correction.&lt;/p&gt;

&lt;p&gt;A SMESH signal is supposed to decay. If nobody reinforces a task, its intensity falls until it no longer matters. That is useful inside the coordination system because stale work cleans itself up.&lt;/p&gt;

&lt;p&gt;An A2A task must not disappear because its internal coordination signal faded.&lt;/p&gt;

&lt;p&gt;An external client may reconnect five minutes later and call &lt;code&gt;GetTask&lt;/code&gt;. An auditor may list tasks by context. A user may need to see that a cancellation was accepted. An artifact must still belong to the task that produced it.&lt;/p&gt;

&lt;p&gt;So the gateway cannot reconstruct its public state by looking at the mesh.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A2A task ledger = retained external task state (process-local today)
SMESH signal     = temporary coordination pressure
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The ledger is authoritative for the A2A lifecycle. The mesh is authoritative only for its local coordination observations.&lt;/p&gt;

&lt;p&gt;That also means terminal states are absorbing. Once a task is completed, failed, canceled, or rejected, a repeated message cannot quietly restart work under the same ID.&lt;/p&gt;
&lt;h2&gt;
  
  
  Streaming exposed a bug that all my tests had missed
&lt;/h2&gt;

&lt;p&gt;A2A v1 supports both direct responses and stateful tasks. A task lifecycle stream begins with the Task itself, then emits ordered status or artifact updates, and closes when the task reaches a terminal state.[4]&lt;/p&gt;

&lt;p&gt;My executor maps internal mesh events like this:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;mesh dispatch accepted  -&amp;gt; Working
mesh progress           -&amp;gt; Working + status message
mesh artifact           -&amp;gt; Artifact update
mesh completion         -&amp;gt; Completed
mesh failure            -&amp;gt; Failed
accepted cancellation   -&amp;gt; Canceled
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The stream ordering test passed.&lt;/p&gt;

&lt;p&gt;Then an independent review pointed out that my own worker budget allowed 256 events while the upstream server's broadcast subscription buffer held 32. A fast worker could stay inside my documented limit and still outrun the initiating subscriber. The task might finish in storage while the client received an internal "subscription fell behind" error.&lt;/p&gt;

&lt;p&gt;That is a particularly unpleasant distributed-systems bug because both sides can truthfully report different outcomes.&lt;/p&gt;

&lt;p&gt;The fix was not to hope the subscriber ran faster. I reduced the worker event budget to 16, clamp caller-provided limits to that ceiling, and added a burst test through the official client.&lt;/p&gt;

&lt;p&gt;This is what protocols do to a design: they force every implicit assumption to become somebody else's observable failure.&lt;/p&gt;
&lt;h2&gt;
  
  
  Cancellation has to stop work, not just change a status label
&lt;/h2&gt;

&lt;p&gt;Forwarding &lt;code&gt;CancelTask&lt;/code&gt; to a dispatcher was not enough.&lt;/p&gt;

&lt;p&gt;If an internal worker ignored the request or kept its event stream open, the original client subscription could hang. Worse, late &lt;code&gt;Working&lt;/code&gt; or &lt;code&gt;Completed&lt;/code&gt; events could arrive after the public task had become &lt;code&gt;Canceled&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The executor now owns a per-task cancellation token. The first accepted cancellation:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;reaches the dispatcher;&lt;/li&gt;
&lt;li&gt;wakes the active producer loop;&lt;/li&gt;
&lt;li&gt;closes the original execution stream;&lt;/li&gt;
&lt;li&gt;prevents post-cancel work from changing the terminal state.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Channel sends and cancellation acknowledgements have deadlines. Worker inactivity has a deadline. The total task has a deadline.&lt;/p&gt;

&lt;p&gt;Cancellation is not a field update. It is a distributed state transition with work on both sides of the boundary.&lt;/p&gt;
&lt;h2&gt;
  
  
  The boring limits are the real feature
&lt;/h2&gt;

&lt;p&gt;The first version was interoperable. It was not bounded enough to deserve trust.&lt;/p&gt;

&lt;p&gt;A fail-closed review found unbounded task retention, unbounded worker output, cancellation leaks, missing dispatcher deadlines, terminal task reuse, and invalid input that could leave a task stranded in &lt;code&gt;Submitted&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The current localhost-first gateway now bounds:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;resource&lt;/th&gt;
&lt;th&gt;default boundary&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;HTTP request body&lt;/td&gt;
&lt;td&gt;128 KiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;accepted inline text&lt;/td&gt;
&lt;td&gt;64 KiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;retained process-local tasks&lt;/td&gt;
&lt;td&gt;1,024&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;active executions&lt;/td&gt;
&lt;td&gt;64&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;worker events&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;artifacts per task&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;aggregate output per task&lt;/td&gt;
&lt;td&gt;1 MiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;worker inactivity&lt;/td&gt;
&lt;td&gt;30 seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;total task execution&lt;/td&gt;
&lt;td&gt;5 minutes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;channel/cancel acknowledgement&lt;/td&gt;
&lt;td&gt;5 seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;It also refuses non-loopback binds unless an explicit unsafe override is present.&lt;/p&gt;

&lt;p&gt;That override does not add authentication, TLS, tenant isolation, or authorization. It only disables the refusal. The current binary is for localhost and trusted integration work, not direct exposure to the internet.[7]&lt;/p&gt;

&lt;p&gt;I am spelling that out because "supports enterprise authentication" in a protocol specification does not mean every prototype using the protocol is enterprise-secure.&lt;/p&gt;
&lt;h2&gt;
  
  
  The LIFELINE demo is a simulation, on purpose
&lt;/h2&gt;

&lt;p&gt;The cover video follows a fictional medication-safety incident. Three hospitals see weak pieces of the same adverse-event pattern. Separate manufacturer, regulator, logistics, payer, and evidence agents contribute artifacts. Inside each public endpoint, a SMESH swarm claims work, reinforces evidence, contests an early hypothesis, and lets unsupported signals decay.&lt;/p&gt;

&lt;p&gt;One logistics endpoint fails. Its task is canceled. A fallback is discovered. The incident continues. The agents converge on a recommendation, but a human incident commander owns the irreversible decision.&lt;/p&gt;

&lt;p&gt;The visual separates the layers deliberately:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;ivory arcs are A2A traffic between organizations;&lt;/li&gt;
&lt;li&gt;green fields are SMESH activity inside an organization;&lt;/li&gt;
&lt;li&gt;cyan shards are artifacts;&lt;/li&gt;
&lt;li&gt;vermilion marks contradiction or failure;&lt;/li&gt;
&lt;li&gt;one gold ring marks human authority.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every visible event comes from a 55-event, ordered, hash-chained JSONL fixture generated as one complete file. The browser can play it, scrub it, inspect it, or export deterministic frames. The narrated film and interactive replay are available from the gateway repository.[7][8]&lt;/p&gt;


&lt;div class="crayons-card c-embed"&gt;

  &lt;br&gt;
&lt;strong&gt;Honesty boundary:&lt;/strong&gt; the gateway's A2A interoperability is exercised against the official Rust client. The LIFELINE six-organization trace is synthetic. It proves the replay contract and the intended architecture; it does &lt;strong&gt;not&lt;/strong&gt; prove that six live SMESH runtimes executed the scenario.&lt;br&gt;

&lt;/div&gt;



&lt;p&gt;A captured run across live SMESH runtimes would be operational proof. This demo is not.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is implemented, and what is still a plan
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;implemented now&lt;/th&gt;
&lt;th&gt;still required for an internet-facing system&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A2A v1 Agent Card&lt;/td&gt;
&lt;td&gt;authenticated principals&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;JSON-RPC and HTTP+JSON/REST bindings&lt;/td&gt;
&lt;td&gt;tenant-aware authorization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;official-client tests for discovery, JSON-RPC/REST send, streaming, and cancellation&lt;/td&gt;
&lt;td&gt;persistent SQL task ledger&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SSE task streaming&lt;/td&gt;
&lt;td&gt;distributed quotas&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Get, List, and Subscribe routes through the SDK handler&lt;/td&gt;
&lt;td&gt;TLS termination and deployment policy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;real &lt;code&gt;SignalType::Query&lt;/code&gt; construction&lt;/td&gt;
&lt;td&gt;live SMESH runtime adapter behind every organization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;bounded process-local execution&lt;/td&gt;
&lt;td&gt;push callback validation and SSRF controls&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;deterministic synthetic replay&lt;/td&gt;
&lt;td&gt;captured multi-runtime causal trace&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;There are two easy ways to lie with a demo like this.&lt;/p&gt;

&lt;p&gt;The first is to animate what you wish the system did.&lt;/p&gt;

&lt;p&gt;The second is to run one loopback worker and describe it as a decentralized enterprise.&lt;/p&gt;

&lt;p&gt;I would rather keep the boundary visible.&lt;/p&gt;

&lt;h2&gt;
  
  
  What A2A changed in the way I think about SMESH
&lt;/h2&gt;

&lt;p&gt;Before this work, I thought of SMESH as the system.&lt;/p&gt;

&lt;p&gt;Now I think of it as an interior.&lt;/p&gt;

&lt;p&gt;The mesh can remain weird in useful ways. Signals can diffuse probabilistically. Specialists can appear and disappear. Trust can be local. Claims can decay. None of that has to leak into the contract presented to another agent.&lt;/p&gt;

&lt;p&gt;MCP gives individual agents hands. SMESH gives a group local coordination. A2A gives that group a public identity, a task contract, and a cancel button.&lt;/p&gt;

&lt;p&gt;The previous transport article ended with a lesson: code that has never been executed is a plan. This work left me with the same rule one layer higher:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A system nobody else can discover, hire, observe, or cancel is not interoperable. It is an island with excellent internal networking.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you are building multi-agent systems, where would you draw the boundary between cross-organization interoperability and internal swarm coordination—and which decisions would you refuse to delegate?&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/copyleftdev" rel="noopener noreferrer"&gt;
        copyleftdev
      &lt;/a&gt; / &lt;a href="https://github.com/copyleftdev/smesh-a2a" rel="noopener noreferrer"&gt;
        smesh-a2a
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      A2A v1 interoperability gateway for decentralized SMESH agent swarms
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;SMESH A2A&lt;/h1&gt;
&lt;/div&gt;
&lt;p&gt;A2A v1 interoperability gateway for decentralized &lt;a href="https://github.com/copyleftdev/smesh-rust" rel="noopener noreferrer"&gt;SMESH&lt;/a&gt; agent swarms.&lt;/p&gt;
&lt;p&gt;SMESH remains the internal coordination substrate: signals diffuse, decay, reinforce, and accumulate attestations. A2A is the public contract for discovery, durable task lifecycle, streaming progress, cancellation, and artifacts.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;What works&lt;/h2&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;Official A2A v1 Rust types and server/client SDKs&lt;/li&gt;
&lt;li&gt;Public Agent Card at &lt;code&gt;/.well-known/agent-card.json&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;JSON-RPC endpoint at &lt;code&gt;/jsonrpc&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;HTTP+JSON/REST endpoint at &lt;code&gt;/rest&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Synchronous and SSE streaming task execution&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;GetTask&lt;/code&gt;, &lt;code&gt;ListTasks&lt;/code&gt;, &lt;code&gt;SubscribeToTask&lt;/code&gt;, and &lt;code&gt;CancelTask&lt;/code&gt; through &lt;code&gt;a2a-rs&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Strict inline-text validation with a 64 KiB default limit&lt;/li&gt;
&lt;li&gt;Translation to a real &lt;code&gt;smesh_core::SignalType::Query&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Injectable &lt;code&gt;MeshDispatcher&lt;/code&gt; boundary for a production SMESH runtime&lt;/li&gt;
&lt;li&gt;Deterministic loopback worker for demos and interoperability tests&lt;/li&gt;
&lt;li&gt;Bounded task retention, execution concurrency, event/artifact counts, and output bytes&lt;/li&gt;
&lt;li&gt;Worker inactivity, cancellation, and command-channel deadlines&lt;/li&gt;
&lt;li&gt;Terminal-state and task-ID reuse guards&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Architecture&lt;/h2&gt;
&lt;/div&gt;
&lt;div class="snippet-clipboard-content notranslate position-relative overflow-auto"&gt;
&lt;pre class="notranslate"&gt;&lt;code&gt;A2A client
   |
   v
Agent Card + JSON-RPC/REST/SSE
   |
   v
SmeshExecutor -- validates and translates
   |
   v
MeshDispatcher
   |-- LoopbackDispatcher&lt;/code&gt;&lt;/pre&gt;…&lt;/div&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/copyleftdev/smesh-a2a" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;Interactive replay: &lt;a href="https://copyleftdev.github.io/smesh-a2a/" rel="noopener noreferrer"&gt;copyleftdev.github.io/smesh-a2a&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;SMESH core: &lt;a href="https://github.com/copyleftdev/smesh-rust" rel="noopener noreferrer"&gt;github.com/copyleftdev/smesh-rust&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;p&gt;[1] &lt;a href="https://dev.to/copyleftdev/my-quic-transport-had-never-once-been-executed-heres-what-happened-when-i-ran-it-24ge"&gt;https://dev.to/copyleftdev/my-quic-transport-had-never-once-been-executed-heres-what-happened-when-i-ran-it-24ge&lt;/a&gt; — My QUIC transport had never once been executed&lt;br&gt;
[2] &lt;a href="https://developers.googleblog.com/en/a2a-a-new-era-of-agent-interoperability" rel="noopener noreferrer"&gt;https://developers.googleblog.com/en/a2a-a-new-era-of-agent-interoperability&lt;/a&gt; — Announcing the Agent2Agent Protocol&lt;br&gt;
[3] &lt;a href="https://a2a-protocol.org/latest" rel="noopener noreferrer"&gt;https://a2a-protocol.org/latest&lt;/a&gt; — What is A2A Protocol?&lt;br&gt;
[4] &lt;a href="https://a2a-protocol.org/latest/specification" rel="noopener noreferrer"&gt;https://a2a-protocol.org/latest/specification&lt;/a&gt; — A2A Protocol v1.0 specification&lt;br&gt;
[5] &lt;a href="https://linuxfoundation.org/press/linux-foundation-launches-the-agent2agent-protocol-project-to-enable-secure-intelligent-communication-between-ai-agents" rel="noopener noreferrer"&gt;https://linuxfoundation.org/press/linux-foundation-launches-the-agent2agent-protocol-project-to-enable-secure-intelligent-communication-between-ai-agents&lt;/a&gt; — Linux Foundation launches A2A project&lt;br&gt;
[6] &lt;a href="https://github.com/copyleftdev/smesh-rust" rel="noopener noreferrer"&gt;https://github.com/copyleftdev/smesh-rust&lt;/a&gt; — SMESH repository&lt;br&gt;
[7] &lt;a href="https://github.com/copyleftdev/smesh-a2a" rel="noopener noreferrer"&gt;https://github.com/copyleftdev/smesh-a2a&lt;/a&gt; — SMESH A2A gateway repository&lt;br&gt;
[8] &lt;a href="https://youtu.be/EFPKaIuF8iA" rel="noopener noreferrer"&gt;https://youtu.be/EFPKaIuF8iA&lt;/a&gt; — SMESH A2A cinematic demo&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>architecture</category>
      <category>distributedsystems</category>
    </item>
    <item>
      <title>A Wider Computer, Not a Bigger One: Modeling AI Inference Across Millions of Homes</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Tue, 25 Aug 2026 05:59:43 +0000</pubDate>
      <link>https://dev.to/copyleftdev/a-wider-computer-not-a-bigger-one-modeling-ai-inference-across-millions-of-homes-5cmo</link>
      <guid>https://dev.to/copyleftdev/a-wider-computer-not-a-bigger-one-modeling-ai-inference-across-millions-of-homes-5cmo</guid>
      <description>&lt;p&gt;&lt;em&gt;I modeled an AI inference fleet distributed across ordinary homes. What survived was narrower—and more plausible—than what I started with.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Picture a detached house on a cold evening. On one garage wall, an operator-owned compute appliance draws about five kilowatts—electrically comparable to an EV charger, though it runs much longer. It serves small AI models. In winter its waste heat can warm the house; in summer that heat must be carried away. The family bought nothing. They are paid to host it.&lt;/p&gt;

&lt;p&gt;Now repeat that arrangement across a neighborhood, then a state, then millions of homes—not to assemble one enormous computer, but to create millions of independent inference workers. Each keeps a small model resident in GPU memory. New requests route around homes that are offline. The network grows by adding locations that were already built and connected to the grid.&lt;/p&gt;

&lt;p&gt;That network does not exist. I call the idea HEARTH. Almost none of its physical pieces are exotic; the experiment is whether they can be arranged and scheduled as one system. Not a bigger computer. A wider one.&lt;/p&gt;

&lt;p&gt;To make the appliance less abstract, I developed &lt;a href="https://copyleftdev.github.io/hearth/prototypes/" rel="noopener noreferrer"&gt;three concept form factors&lt;/a&gt;: a wall unit, a floor-standing thermal tower, and a duct-integrated mechanical-room unit. They are appearance and installation studies, not engineered products; their job is to expose the questions that a real prototype must answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The model kept changing its answer
&lt;/h2&gt;

&lt;p&gt;Then I priced the accelerators, and the idea stopped working.&lt;/p&gt;

&lt;p&gt;That was the first useful result. My original comparison had focused on the dramatic expense of constructing a data center while treating the compute hardware as identical on both sides. But identical hardware does not disappear from the economics. If one location keeps an expensive GPU busier than another, utilization can overwhelm everything the cheaper building saves.&lt;/p&gt;

&lt;p&gt;I rebuilt the model five times. Each pass introduced a constraint the previous one had missed:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pass&lt;/th&gt;
&lt;th&gt;What changed&lt;/th&gt;
&lt;th&gt;Model verdict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Compared the facilities around the hardware&lt;/td&gt;
&lt;td&gt;Homes win&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Included accelerator cost and utilization&lt;/td&gt;
&lt;td&gt;Homes lose&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Used consumer and data-center hardware at quoted prices&lt;/td&gt;
&lt;td&gt;Homes win&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Modeled production batching for a 32B model&lt;/td&gt;
&lt;td&gt;Homes lose&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Tested different model sizes&lt;/td&gt;
&lt;td&gt;Homes win only for the &lt;strong&gt;small-model case&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What survived: small, independent inference
&lt;/h2&gt;

&lt;p&gt;The fifth pass did not rescue my original proposal. It replaced it. A residential fleet is not a cheaper place to run every AI workload. It cannot pool memory across the internet, train a frontier model, or make residential latency disappear. What survived was narrower: small-model inference in which one node completes one request and no durable state is tied to a particular house.&lt;/p&gt;

&lt;p&gt;Model size changes the economics because inference is not just a race between GPUs. During generation, a server repeatedly reads the model's weights while advancing many requests together. That is batching. A larger batch spreads each weight read across more paying tokens—but every active request also consumes working memory, usually called the KV cache.&lt;/p&gt;

&lt;p&gt;An 8B model—roughly eight billion parameters—leaves enough memory on a 32 GB consumer GPU for a useful production batch. In my model, both the home GPU and the data-center hardware then become limited mainly by compute, allowing the consumer part's much lower purchase price to matter. At 32B, the consumer card runs short of memory first. Its batch stops growing while high-memory data-center hardware keeps filling, and the verdict reverses.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model class&lt;/th&gt;
&lt;th&gt;What the current model says&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;8B&lt;/td&gt;
&lt;td&gt;Candidate workload; the consumer GPU retains useful batching headroom&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;32B&lt;/td&gt;
&lt;td&gt;Data-center hardware wins unless other advantages offset its batching edge&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;70B&lt;/td&gt;
&lt;td&gt;Excluded by my single-GPU serving assumption; multi-GPU serving was not modeled&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The current result is conditional. Before homeowner compensation and fleet-level operating costs, and under RTX 5090 MSRP, an 85% accelerator duty factor, identical serving-efficiency assumptions for both venues, and a quantized 8B workload, my roofline model estimates the residential hardware-and-energy stack at about 0.49 times the cost per token of the best data-center case in the sweep, based on GB200 NVL72 rack pricing. That is a reproducible scenario output, not observed production performance or a fully loaded business cost. An apples-to-apples benchmark—or the costs excluded here—could erase the advantage.&lt;/p&gt;

&lt;p&gt;This is the distinction behind &lt;em&gt;wide, not deep&lt;/em&gt;. HEARTH would not divide one giant model among houses or combine residential GPUs into a virtual supercomputer. Each node would answer complete, independent requests with a model it already holds locally. Adding homes increases the number and variety of requests the fleet can serve; it does not make any one node larger.&lt;/p&gt;

&lt;h2&gt;
  
  
  The asset is the grid endpoint
&lt;/h2&gt;

&lt;p&gt;Why put that workload in homes at all? Not because residential electricity is cheap—it usually is not—or because waste heat makes energy free. The asset is the connection behind the meter.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://datacenters.lbl.gov/publications/2024-lbnl-data-center-energy-usage-report" rel="noopener noreferrer"&gt;Berkeley Lab estimated&lt;/a&gt; that US data centers used about 176 TWh of electricity in 2023. &lt;a href="https://datacenters.lbl.gov/publications/united-states-data-center-energy-2025" rel="noopener noreferrer"&gt;Its June 2026 update&lt;/a&gt; projects 521–843 TWh in 2030, with a 649 TWh reference case. Serving that growth is not simply a question of buying more GPUs. A large new campus may need substations, transformers, transmission upgrades, and a negotiated path to connect a load the local grid was never designed to carry. The &lt;a href="https://www.ercot.com/files/docs/2026/04/01/ERCOT_LargeLoad_Update_April2026_B-C_-Hearing.pdf" rel="noopener noreferrer"&gt;410 GW of prospective large loads reported in ERCOT&lt;/a&gt;—about 87% associated with data centers—is not a forecast of what will be built. It is evidence of how much demand is converging on the same constrained process.&lt;/p&gt;

&lt;p&gt;The United States has about &lt;a href="https://data.census.gov/table/ACSDP1Y2024.DP04?g=010XX00US" rel="noopener noreferrer"&gt;82.5 million one-unit detached housing units&lt;/a&gt;. That is physical stock, not an estimate of eligible hosts: occupancy, broadband, electrical capacity, utility approval, and household consent would reduce it substantially. Each candidate already has a meter, an electrical service, and a physical building around it. A compute node installed behind that meter may avoid the transmission-scale interconnection required by a new campus. That is the inversion: instead of bringing an enormous new grid connection to the compute, bring a modest amount of compute to many connections that already exist.&lt;/p&gt;

&lt;p&gt;Existing does not mean unused. A continuous 5 kW load is much less forgiving than an EV that charges for a few hours. Some homes would need panel or service upgrades; utilities may require review; and the local distribution transformer remains a hard physical constraint. My one-node-per-transformer rule is therefore a screening hypothesis, not a validated safety rule or national capacity estimate. Transformer loading, telemetry, and thermal behavior belong in the first utility-supervised field test.&lt;/p&gt;

&lt;p&gt;To the household, this is a hosting contract, not an investment. The operator would own and maintain the appliance, meter and reimburse its electricity, and pay the family a share of revenue. None of that has been field-tested; fire and insurance rules, noise, summer heat rejection, ISP terms, maintenance, and upgrades remain field questions.&lt;/p&gt;

&lt;h2&gt;
  
  
  One request goes to one home
&lt;/h2&gt;

&lt;p&gt;At the software layer, the system is deliberately less exotic. A control plane would assign replicas of small models to qualifying nodes and distribute checkpoints outside the request path. Each node would store its assigned checkpoints locally and keep its serving model resident in GPU memory. The request router would know which homes have the requested model resident, which are healthy, and which have batch capacity available.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;registry ── sync ──► home node

client ──► router ──► available home ──► response
               ▲              │
               └── telemetry ─┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;During generation, one home holds the request's KV cache and returns the tokens; no activations or model layers cross the residential network. If the node disappears, the in-flight generation is lost. New work can route elsewhere, and a retryable request can start again against another replica, although it may not produce identical tokens. The system does not inherit data-center reliability merely because it has many machines.&lt;/p&gt;

&lt;p&gt;At the network level, remote attestation, model security, and prompt and output confidentiality on hardware outside an operator-controlled facility remain unresolved. Distribution makes failures routable; it does not make operations disappear.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five conditions decide whether this is real
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The hardware must be legally and commercially usable.&lt;/strong&gt; &lt;a href="https://www.nvidia.com/en-us/drivers/nvidia-license/" rel="noopener noreferrer"&gt;NVIDIA's current driver agreement&lt;/a&gt; says the software may not be used to provide commercial hosting services and that GeForce and Titan software is not licensed for data-center deployment. Whether an operator could obtain written authorization or a different license for this residential topology must be resolved with NVIDIA and qualified counsel before a hardware pilot. Warranty, support, and a credible procurement channel must also exist.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Edge serving must sustain the modeled batches.&lt;/strong&gt; The same checkpoint, quantization, context length, serving stack, and request mix must be benchmarked on the candidate consumer GPU and data-center hardware. A spreadsheet cannot substitute for that result.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Demand must keep the silicon busy.&lt;/strong&gt; Cheap idle GPUs are still expensive. Customers must commit enough suitable 8B inference work—classification, extraction, routing, drafting, or other bounded tasks—to maintain high utilization at a viable token price.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The location must pass the energy screen.&lt;/strong&gt; Residential rates vary widely, and heat reuse is a seasonal credit rather than the thesis. Every deployment needs to work after metered electricity, host payment, and summer heat rejection.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The distribution grid must approve the load.&lt;/strong&gt; A real utility must confirm that a real home, service, and transformer can carry it continuously. National averages are not permission to connect.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Those are not caveats around the result. They &lt;em&gt;are&lt;/em&gt; the result. HEARTH exists only at their intersection.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first pilot should contain no houses
&lt;/h2&gt;

&lt;p&gt;That changes the order of the pilot. My original plan began with 250 homes and grew through increasingly expensive gates. It tested installation first and market demand later. The model eventually exposed the flaw: there is no reason to put hardware in a house before proving that anyone will buy this particular kind of inference.&lt;/p&gt;

&lt;p&gt;The first HEARTH pilot should therefore contain no houses. Rent the candidate consumer and data-center hardware. Resolve the license question. Run the same models through the same serving stack, measure delivered tokens and power at production batch, and ask customers to pay for the workloads the residential fleet is supposed to serve. If the measured economics or demand fail, stop before an electrician is dispatched.&lt;/p&gt;

&lt;p&gt;Only then move to 250 homes. That field trial buys answers the lab cannot: Can crews install safely and consistently? What do transformers do under continuous load? How loud and hot are the nodes in July? Can the operator keep them online without drowning in truck rolls—and will families keep hosting them?&lt;/p&gt;

&lt;h2&gt;
  
  
  Four million homes is not the next milestone
&lt;/h2&gt;

&lt;p&gt;Four million homes is an image of the possible system, not an adoption forecast. The serious next milestone is smaller: one workload that benchmarks, one customer willing to buy it, one utility willing to supervise it, and then the first 250 homes.&lt;/p&gt;

&lt;p&gt;I published the &lt;a href="https://copyleftdev.github.io/hearth/report/" rel="noopener noreferrer"&gt;full feasibility study&lt;/a&gt;, the &lt;a href="https://github.com/copyleftdev/hearth" rel="noopener noreferrer"&gt;source model and inputs&lt;/a&gt;, and every assumption I used. If the idea is wrong, I want the failure to be reproducible too.&lt;/p&gt;

&lt;p&gt;HEARTH is not a forecast. It is a claim specific enough to test—and to kill if it fails. The buildings, meters, and many of the grid and network endpoints already surround us. What is missing is the appliance, the dedicated installation, the operating system around it, and an agreement worth signing.&lt;/p&gt;

&lt;p&gt;If those pieces work, the neighborhood still looks like a neighborhood on a cold evening. Behind some garage walls, independent machines serve requests from a network spanning the country—each complete on its own, all scheduled together.&lt;/p&gt;

&lt;p&gt;No monument to compute. A million warm windows.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>aiops</category>
      <category>infrastructure</category>
    </item>
    <item>
      <title>My QUIC transport had never once been executed. Here's what happened when I ran it.</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Wed, 19 Aug 2026 03:35:06 +0000</pubDate>
      <link>https://dev.to/copyleftdev/my-quic-transport-had-never-once-been-executed-heres-what-happened-when-i-ran-it-24ge</link>
      <guid>https://dev.to/copyleftdev/my-quic-transport-had-never-once-been-executed-heres-what-happened-when-i-ran-it-24ge</guid>
      <description>&lt;p&gt;I've written before about SMESH, a coordination protocol modelled on mycorrhizal networks — the fungal web that lets trees in a forest warn each other about drought and disease with nothing in charge of the network. Signals diffuse, decay on their own, and get reinforced when independently confirmed. Consensus emerges instead of being orchestrated.&lt;/p&gt;

&lt;p&gt;That was the idea. This post is about the part where I found out whether it worked.&lt;/p&gt;

&lt;h2&gt;
  
  
  The transport that had never run
&lt;/h2&gt;

&lt;p&gt;SMESH has had a QUIC transport in it for a while. Roughly 500 lines: a quinn endpoint that is simultaneously server and client, self-signed certs, length-prefixed bincode frames over unidirectional streams, an accept loop that spawns per-connection and per-stream tasks, connection pooling.&lt;/p&gt;

&lt;p&gt;Every test passed. The workspace was green. I could point at &lt;code&gt;smesh-runtime/src/transport.rs&lt;/code&gt; and say "yes, it does peer-to-peer."&lt;/p&gt;

&lt;p&gt;Then I grepped for who actually constructed it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-rn&lt;/span&gt; &lt;span class="s2"&gt;"QuicTransport"&lt;/span&gt; &lt;span class="nt"&gt;--include&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'*.rs'&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt;
smesh-runtime/src/transport.rs:177:pub struct QuicTransport &lt;span class="o"&gt;{&lt;/span&gt;
smesh-runtime/src/transport.rs:192:impl QuicTransport &lt;span class="o"&gt;{&lt;/span&gt;
smesh-runtime/src/lib.rs:16:pub use transport::&lt;span class="o"&gt;{&lt;/span&gt;QuicTransport, ...&lt;span class="o"&gt;}&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Its own definition, and a re-export. Nothing else in the workspace had ever instantiated it. No binary opened a socket. &lt;code&gt;SmeshRuntime&lt;/code&gt; imported &lt;code&gt;TransportConfig&lt;/code&gt;, stored it in a struct field, and never looked at it again.&lt;/p&gt;

&lt;p&gt;I had a networking layer with tests, docs, and zero executions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three bugs in the first twenty minutes
&lt;/h2&gt;

&lt;p&gt;I wrote an integration test that starts two runtimes, has one dial the other, and asserts a signal crosses. Here is what fell out before it went green.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. It panicked on the first call.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Could not automatically determine the process-level CryptoProvider
from Rustls crate features.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;rustls 0.23 refuses to pick a crypto backend when more than one is compiled in, and quinn pulls in both through its own feature set. Every call to &lt;code&gt;QuicTransport::new&lt;/code&gt; would have panicked for anyone, ever. Nobody noticed because nobody had called it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Dialled connections were write-only.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;connect()&lt;/code&gt; stored the connection in the pool but only the accept loop pumped incoming streams — and the accept loop only sees connections you &lt;em&gt;accepted&lt;/em&gt;. So a node that dialled out could send, and would never receive anything back. A QUIC connection is bidirectional regardless of who dialled it; my code only acted like it half the time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Unbounded allocation from an attacker-controlled length prefix.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;len&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nn"&gt;u32&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;from_be_bytes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;len_buf&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nb"&gt;usize&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="k"&gt;mut&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nd"&gt;vec!&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0u8&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="n"&gt;len&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;   &lt;span class="c1"&gt;// &amp;lt;- no&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;max_message_size&lt;/code&gt; was in the config struct. It was never read. Send a 4 GiB length prefix and the process allocates 4 GiB.&lt;/p&gt;

&lt;p&gt;None of these are clever bugs. They're the bugs you get for free the first time code meets a socket, and the only reason they survived is that the code had never met a socket.&lt;/p&gt;

&lt;h2&gt;
  
  
  The harder problem: my protocol was wrong
&lt;/h2&gt;

&lt;p&gt;Fixing the plumbing was the easy half. The real issue was that my diffusion algorithm quietly assumed something no distributed system can assume.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Network::tick&lt;/code&gt; expands a signal's reach one hop per tick by walking the graph:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;node_id&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;reached&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;hypha&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="k"&gt;self&lt;/span&gt;&lt;span class="py"&gt;.hyphae&lt;/span&gt;&lt;span class="nf"&gt;.get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;node_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;target&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;self&lt;/span&gt;&lt;span class="py"&gt;.nodes&lt;/span&gt;&lt;span class="nf"&gt;.get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;hypha&lt;/span&gt;&lt;span class="py"&gt;.to&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;target&lt;/span&gt;&lt;span class="nf"&gt;.should_relay&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;signal&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;remaining_hops&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;frontier&lt;/span&gt;&lt;span class="nf"&gt;.push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;hypha&lt;/span&gt;&lt;span class="py"&gt;.to&lt;/span&gt;&lt;span class="nf"&gt;.clone&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read that again. It iterates every node's adjacency, calls every node's relay policy, and mutates one global &lt;code&gt;reached_nodes&lt;/code&gt; set on a shared signal. It's a breadth-first search from a god's-eye view of the entire graph.&lt;/p&gt;

&lt;p&gt;That works beautifully in one process. In a real mesh, &lt;strong&gt;no node can see that graph.&lt;/strong&gt; Porting it meant three corrections, and each one turned out to be a genuine bug rather than a porting detail.&lt;/p&gt;

&lt;h3&gt;
  
  
  Correction 1: independent conclusions were being thrown away
&lt;/h3&gt;

&lt;p&gt;Signals are content-addressed — the hash is derived from what is being claimed. So when two agents independently reach the same conclusion, they produce the same hash.&lt;/p&gt;

&lt;p&gt;My &lt;code&gt;emit()&lt;/code&gt; saw the hash already present locally and treated it as a duplicate. It dropped it.&lt;/p&gt;

&lt;p&gt;That is exactly backwards. Two parties independently agreeing is not redundant data — it is the &lt;em&gt;only&lt;/em&gt; evidence the system has that a claim is real. Discarding it destroys the thing the protocol exists to measure.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="cd"&gt;/// Signals are content-addressed, so a node that independently reaches a&lt;/span&gt;
&lt;span class="cd"&gt;/// conclusion another node already published lands on the same hash. That&lt;/span&gt;
&lt;span class="cd"&gt;/// is treated as *corroboration*: this node is added as an attester and the&lt;/span&gt;
&lt;span class="cd"&gt;/// merged claim still goes out, because our agreement is news to everyone&lt;/span&gt;
&lt;span class="cd"&gt;/// who has not heard it. Swallowing it as a duplicate would silently&lt;/span&gt;
&lt;span class="cd"&gt;/// discard the only evidence that two parties concur.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Correction 2: I was counting messengers, not witnesses
&lt;/h3&gt;

&lt;p&gt;Reinforcement attributed the claim to whoever handed me the message. In a gossip mesh, one finding relayed by five nodes then looks like five corroborators.&lt;/p&gt;

&lt;p&gt;The party that attests to a claim is its &lt;em&gt;origin&lt;/em&gt;, not the peer that passed it along. Relaying is not agreeing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Correction 3: gossip needs a merge rule, not a broadcast rule
&lt;/h3&gt;

&lt;p&gt;Once attestation is a set, the rule that makes gossip converge is simple and pleasant:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Merge the two attester sets. Anything the sender knew that we did&lt;/span&gt;
&lt;span class="c1"&gt;// not is new information, and new information is worth passing on.&lt;/span&gt;
&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;before&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nn"&gt;Node&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;attesters&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;existing&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;attester&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;incoming_attesters&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="n"&gt;before&lt;/span&gt;&lt;span class="nf"&gt;.contains&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;attester&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;existing&lt;/span&gt;&lt;span class="nf"&gt;.reinforce&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;attester&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Forward if and only if your own knowledge grew. That single rule is the loop breaker (a message teaching you nothing goes no further), the convergence mechanism (the set is grow-only, so it's a CRDT), and the anti-entropy repair (re-asserting carries your accumulated view, so a node that missed a round catches up).&lt;/p&gt;

&lt;h2&gt;
  
  
  The part I nearly got wrong on purpose
&lt;/h2&gt;

&lt;p&gt;Here is the subtlety that makes the whole thing work, and it looks like a mistake in the source.&lt;/p&gt;

&lt;p&gt;The signal builder folds the origin node into the content hash if you give it one. So when an agent publishes a claim, you must &lt;strong&gt;not&lt;/strong&gt; set the origin:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;signal&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nn"&gt;Signal&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;builder&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nn"&gt;SignalType&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;Alert&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;.payload&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;assertion&lt;/span&gt;&lt;span class="nf"&gt;.canonical_bytes&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
    &lt;span class="nf"&gt;.confidence&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;finding&lt;/span&gt;&lt;span class="py"&gt;.confidence&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;.build&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="c1"&gt;// .origin() deliberately NOT called&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you set it, every agent gets a different hash for the same claim and correlation becomes impossible. The address has to be the &lt;em&gt;claim&lt;/em&gt;, not the &lt;em&gt;claimant&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;The corollary is that evidence cannot travel in the payload. Each agent's evidence differs, so putting it in would make every hash unique and break the mechanism. The payload is just the assertion:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"subject"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"checkout-api"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"claim"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"degraded"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Evidence stays local and goes to the log. The mesh carries assertions; it does not carry arguments.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five witnesses, none of whom can see the problem
&lt;/h2&gt;

&lt;p&gt;To actually test any of this, I built a scenario where the answer cannot be reached alone.&lt;/p&gt;

&lt;p&gt;Five analyst processes watch the same fleet of services. Each can see exactly one kind of telemetry, and nothing else:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;agent&lt;/th&gt;
&lt;th&gt;sees&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;latency&lt;/td&gt;
&lt;td&gt;p99 response times&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;errors&lt;/td&gt;
&lt;td&gt;error rates&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;saturation&lt;/td&gt;
&lt;td&gt;pool and CPU utilisation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;traces&lt;/td&gt;
&lt;td&gt;retry rates, span queueing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;deploys&lt;/td&gt;
&lt;td&gt;release events&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A deploy cuts &lt;code&gt;checkout-api&lt;/code&gt;'s connection pool from 200 to 20. The pool pins, requests queue, callers time out.&lt;/p&gt;

&lt;p&gt;Now every agent sees a piece of it, and two of them are actively misleading. &lt;code&gt;errors&lt;/code&gt; sees &lt;code&gt;payments-api&lt;/code&gt; throwing 503s and would blame it — but payments is a victim, its own pool and CPU are fine. &lt;code&gt;latency&lt;/code&gt; sees four services slow at once and can't say which is causal. Only &lt;code&gt;deploys&lt;/code&gt; can see there was a release, and a release on its own means nothing; software ships all day.&lt;/p&gt;

&lt;p&gt;I also planted two decoys with a single witness each — an unrelated CPU spike, a brief latency blip — because a system that agrees with everything is not consensus, it's an echo.&lt;/p&gt;

&lt;p&gt;They run as five real OS processes on five ports over real encrypted QUIC, in a ring-plus-chord topology so messages actually have to be relayed to cross the mesh. Peer discovery is off, or gossip quietly converts any topology into a full mesh and there's nothing left to watch.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;smesh orchestrate &lt;span class="nt"&gt;--out&lt;/span&gt; runs/latest
&lt;span class="go"&gt;
  spawned latency     pid 2452744  127.0.0.1:9301  dials 9302,9303
  spawned errors      pid 2452745  127.0.0.1:9302  dials 9303
  spawned saturation  pid 2452746  127.0.0.1:9303  dials 9304
  spawned traces      pid 2452747  127.0.0.1:9304  dials 9305
  spawned deploys     pid 2452748  127.0.0.1:9305  dials 9301
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The result:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;what the mesh concluded
  checkout-api          CONSENSUS at 16.9s · 5 attesters · seen by 5 nodes
  payments-api          no consensus (3/4 concerns)
  edge-gateway          no consensus (1/4 concerns)
  notification-worker   no consensus (1/4 concerns)
  session-store         no consensus (1/4 concerns)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Cause separated from symptom separated from noise. The loud, obvious suspect was held at three witnesses and never promoted. The decoys were never rejected by anything — they simply went uncorroborated and decayed out. Nobody voted, and nothing was in charge.&lt;/p&gt;

&lt;h2&gt;
  
  
  Making the run replayable, and then checking that claim
&lt;/h2&gt;

&lt;p&gt;I wanted to visualise this, which meant the log had to be good enough to reconstruct the run exactly. Each process writes newline-delimited JSON against a shared run epoch: every emission, every per-peer send, every receipt, every relay decision including the probability and the die roll that resolved it, and a full field snapshot every 500ms so decay curves are &lt;em&gt;observed&lt;/em&gt; rather than modelled.&lt;/p&gt;

&lt;p&gt;Then I did the part I'd recommend to anyone building a log you intend to trust: I wrote a validator that checks the log against itself. Sequence gaps, time going backwards, snapshots referencing signals that were never received, consensus declared without the receipts to justify it.&lt;/p&gt;

&lt;p&gt;It caught a real bug on its first run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;FAIL  latency: first event is peer_connected, not node_started
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A peer completed the handshake in the gap between binding the endpoint and writing the node's identity line. The fix was to move the identity write inside the mesh startup, before any loop spawns. I would never have found that by looking at the picture — the picture would just have been subtly wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is still wrong
&lt;/h2&gt;

&lt;p&gt;Being honest about the edges, because "it works" is a claim that needs a boundary:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;origin_node_id&lt;/code&gt; is unauthenticated.&lt;/strong&gt; It's a string on the wire, and the trust model gates relay probability on it. Spoofing another agent's identity is currently free. Ed25519-signing the origin hash closes it and is the next real piece of work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The content hash is truncated to 64 bits.&lt;/strong&gt; Fine against accident, not against an adversary looking for collisions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The telemetry in the demo is synthetic.&lt;/strong&gt; Deliberately: a seeded fixture means the run reproduces byte-for-byte on any machine, which is what makes a visualisation worth trusting. The coordination is not synthetic — real processes, real sockets, probabilistic relay.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  See it move
&lt;/h2&gt;

&lt;p&gt;The full narrated walkthrough is the cover video on this post, or here: &lt;a href="https://youtu.be/kmCzwSBqu_s" rel="noopener noreferrer"&gt;https://youtu.be/kmCzwSBqu_s&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It opens on the forest the protocol is stolen from, then goes inside the recorded run: five agents, six encrypted links, every dot on screen a real message read back from the log rather than animated for effect.&lt;/p&gt;

&lt;p&gt;If you take one thing from this: &lt;strong&gt;code that has never been executed is not code, it's a plan.&lt;/strong&gt; Mine had tests, docs, and a clean &lt;code&gt;cargo clippy&lt;/code&gt;, and it would have panicked on the first line for every user. The tests were testing that the plan was internally consistent.&lt;/p&gt;

&lt;p&gt;Repo: &lt;a href="https://github.com/copyleftdev/smesh-rust" rel="noopener noreferrer"&gt;github.com/copyleftdev/smesh-rust&lt;/a&gt; · Rust · MIT/Apache-2.0&lt;/p&gt;

</description>
      <category>rust</category>
      <category>distributedsystems</category>
      <category>networking</category>
      <category>ai</category>
    </item>
    <item>
      <title>Shipping Assumptions: A Reliability Stack for AI-Generated Code</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Sun, 16 Aug 2026 22:56:25 +0000</pubDate>
      <link>https://dev.to/copyleftdev/shipping-assumptions-a-reliability-stack-for-ai-generated-code-3p9f</link>
      <guid>https://dev.to/copyleftdev/shipping-assumptions-a-reliability-stack-for-ai-generated-code-3p9f</guid>
      <description>&lt;p&gt;The code looked good. It linted cleanly. The shallow tests passed. Everyone felt—&lt;em&gt;vibed&lt;/em&gt;—that we had built the right thing.&lt;/p&gt;

&lt;p&gt;Then the edge case appeared in production.&lt;/p&gt;

&lt;p&gt;The developer treated it as normal operating procedure: bugs happen, tickets arrive, patches ship. QA was asked why they had not caught it. But QA never received a model of the system—only an implementation full of assumptions they were expected to reverse-engineer.&lt;/p&gt;

&lt;p&gt;This is becoming the defining failure mode of AI-assisted development. We can generate code faster than we can understand the systems it creates. The danger is no longer confined to a bad function or an obvious syntax error. It lives in the space between components: boundaries, state transitions, failure modes, and invariants.&lt;/p&gt;

&lt;p&gt;We are not merely shipping code. We are shipping assumptions we can no longer see.&lt;/p&gt;

&lt;h2&gt;
  
  
  Clean code is not a coherent system
&lt;/h2&gt;

&lt;p&gt;A linter can tell us whether code follows a set of local rules. A test can tell us whether selected examples produce the expected result. Neither can tell us what the system must preserve unless someone states it first.&lt;/p&gt;

&lt;p&gt;An invariant is one of those statements: a condition that must remain true across every valid state of the system. An account balance cannot be changed without a corresponding transaction. A private object cannot become public without authorization. Two successful writes cannot silently erase one another.&lt;/p&gt;

&lt;p&gt;When invariants remain implicit, clean code can still assemble into an incoherent system. Each component may look reasonable in isolation while the failure waits in a transition, a retry, a race, or a boundary nobody thought to draw.&lt;/p&gt;

&lt;p&gt;That is the space AI-assisted development is rapidly filling with code.&lt;/p&gt;

&lt;h2&gt;
  
  
  The layer we buried
&lt;/h2&gt;

&lt;p&gt;There was no golden age in which every developer wrote a formal specification before touching an editor. Most teams did not. Earlier software was full of hidden assumptions too, and nostalgia has a habit of removing the inconvenient parts of the past.&lt;/p&gt;

&lt;p&gt;But the field did develop disciplines for reasoning above the code: architectural models, state machines, formal specifications, model checking, fault injection, and deterministic simulation. We tended to reserve them for systems where failure was obviously expensive. Everywhere else, rigor was treated as a cost to minimize.&lt;/p&gt;

&lt;p&gt;Then implementation became cheaper. Frameworks hid machinery. Packages compressed years of expertise into an import. That brought enormous benefits, but it also made it possible to build systems without seeing very far beneath their surfaces.&lt;/p&gt;

&lt;p&gt;Generative AI accelerates the same trade. It can produce a plausible implementation before a team has agreed on what the system is. When nobody externalizes that intent, the generated code becomes the design by default. QA receives the consequences downstream.&lt;/p&gt;

&lt;p&gt;The missing layer is not more code review. It is a shared model between human intent and machine output.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three models, three questions
&lt;/h2&gt;

&lt;p&gt;No single technique covers that layer. A useful stack needs to answer three different questions.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. C4: What exists, and how does it relate?
&lt;/h3&gt;

&lt;p&gt;The &lt;a href="https://c4model.com/" rel="noopener noreferrer"&gt;C4 model&lt;/a&gt; gives software teams a hierarchy for visualizing a system: context, containers, components, and code. It works like a map with zoom levels. At the widest view, we see the system, its users, and neighboring systems. Zooming in reveals applications, data stores, services, components, and relationships.&lt;/p&gt;

&lt;p&gt;The point is not to produce four mandatory diagrams for every project. The official guidance notes that &lt;a href="https://c4model.com/diagrams" rel="noopener noreferrer"&gt;context and container diagrams are sufficient for many teams&lt;/a&gt;. The point is to make boundaries discussable.&lt;/p&gt;

&lt;p&gt;That matters because generated code is locally persuasive. A service can look complete while its ownership is unclear. An API can look tidy while its trust boundary is invisible. C4 gives development and QA the same structural map before either group has to infer the system from a repository.&lt;/p&gt;

&lt;p&gt;AI can help draft that map from requirements or an existing codebase. It should not get the final word. Humans still need to ask whether the map describes the system they intend to operate.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. TLA+: What may happen, and what must remain true?
&lt;/h3&gt;

&lt;p&gt;C4 shows structure, but structure alone cannot express behavior. It does not tell us which states are valid, which transitions are permitted, or which conditions must survive every interleaving.&lt;/p&gt;

&lt;p&gt;That is where &lt;a href="https://lamport.azurewebsites.net/tla/high-level-view.html" rel="noopener noreferrer"&gt;TLA+&lt;/a&gt; fits. Leslie Lamport describes it as a language for precise, high-level models above the code level, especially for concurrent and distributed systems. Its TLC model checker can explore behaviors and find traces that violate the properties we claim should hold.&lt;/p&gt;

&lt;p&gt;TLA+ forces a different conversation. Instead of asking, “Does this function work?” we ask, “What can the system do from this state, and is every reachable result acceptable?”&lt;/p&gt;

&lt;p&gt;The notation may be formal, but AI can reduce the entry cost: drafting a first specification, translating plain-language promises into candidate invariants, explaining counterexample traces, and helping a team refine the model. The human job is to decide whether those promises are actually the right ones. A model that formalizes the wrong intent is merely precise about the wrong system.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. DST: Does the implementation honor the model?
&lt;/h3&gt;

&lt;p&gt;A correct model does not prove that the production code implements it correctly. We still need the implementation to fight back.&lt;/p&gt;

&lt;p&gt;Deterministic simulation testing, or DST, runs real implementation code in a controlled world. Time, randomness, scheduling, networks, storage, and faults become inputs that a simulator can manipulate. When a generated scenario causes a failure, its seed makes the run reproducible.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/tigerbeetle/tigerbeetle/blob/main/docs/ARCHITECTURE.md" rel="noopener noreferrer"&gt;TigerBeetle’s VOPR&lt;/a&gt; is a striking example. It can simulate a cluster on a single thread, accelerate time, and inject storage faults. TigerBeetle makes an important distinction in its architecture documentation: model checking can expose errors in an algorithm, while simulation exercises the specific implementation and its underlying assumptions.&lt;/p&gt;

&lt;p&gt;The three layers complement one another:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;C4 maps the system. TLA+ states what must remain true. DST tries to make it false.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  A workflow for abundant code
&lt;/h2&gt;

&lt;p&gt;This does not need to become a new ceremony in which teams maintain enormous diagrams and specifications that immediately drift out of date. The models have to participate in the engineering loop.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;State the promises in plain language.&lt;/strong&gt; Before generating implementation, write down what the system must always preserve and what it must never permit.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Draw the boundaries.&lt;/strong&gt; Use C4 to identify users, external systems, containers, components, data stores, and the relationships where assumptions cross from one owner to another.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Model the risky behavior.&lt;/strong&gt; Use TLA+ where concurrency, ordering, retries, permissions, or state transitions make the promises difficult to reason about informally.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Generate from reviewed intent.&lt;/strong&gt; Let AI produce implementation and tests only after people have challenged the structural and behavioral models.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Attack the implementation deterministically.&lt;/strong&gt; Put risky dependencies—time, randomness, scheduling, networks, or storage—under control. Generate hostile scenarios, preserve their seeds, and replay every failure.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Feed reality back into the models.&lt;/strong&gt; Production observations and simulation failures reveal assumptions the team missed. Update the invariants, then update the implementation and tests. Detect drift rather than allowing the model to become historical decoration.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Not every application needs a simulator as sophisticated as VOPR or a large TLA+ specification. The lightweight version might be one context diagram, one container diagram, five written invariants, a small model for the riskiest transition, and a seeded test harness around it.&lt;/p&gt;

&lt;p&gt;The goal is not maximal formality. It is to give every important assumption an address.&lt;/p&gt;

&lt;h2&gt;
  
  
  Everything is mission-critical to someone
&lt;/h2&gt;

&lt;p&gt;Imagine a grandmother uses AI to vibe-code a family photo album. This is not a financial ledger or a flight-control system. Yet the application makes real promises:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Original photographs must never be destroyed.&lt;/li&gt;
&lt;li&gt;Private photographs must never be exposed accidentally.&lt;/li&gt;
&lt;li&gt;Captions and chronology must survive every edit.&lt;/li&gt;
&lt;li&gt;Two family members must not overwrite each other’s changes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All of them matter.&lt;/p&gt;

&lt;p&gt;Her C4 model can be small: a person, an application, an identity provider, local storage, cloud storage, and a sharing boundary. Her TLA+ model does not need to describe the entire product; it can focus on concurrent edits, deletion, or permissions. A deterministic harness can simulate two devices going offline, editing the same album, reconnecting in different orders, and retrying interrupted uploads.&lt;/p&gt;

&lt;p&gt;This stack was once associated with mission-critical engineering because rigor was expensive. But AI has changed the economics. If implementation now comes almost freely, why should rigor remain a luxury?&lt;/p&gt;

&lt;p&gt;Criticality is not determined by infrastructure scale. It is determined by the consequence of failure to the person who trusted the system.&lt;/p&gt;

&lt;p&gt;Everything is mission-critical to someone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Retrieval, not retreat
&lt;/h2&gt;

&lt;p&gt;It is tempting to tell the history of software as a simple decay: once we understood the machine, then we specialized, then we configured packages, and now we prompt a model to do even that for us.&lt;/p&gt;

&lt;p&gt;That story is emotionally recognizable and historically incomplete. Modern developers are not less intelligent, and abstraction is not the enemy. The deeper problem is that valuable disciplines became buried beneath convenience. We kept the outputs while losing contact with some of the reasoning that produced them.&lt;/p&gt;

&lt;p&gt;This is where nostalgia can be useful. Not as a demand to reconstruct an imagined past, but as a diagnostic signal. It tells us that something in the present has become difficult to see or value.&lt;/p&gt;

&lt;p&gt;Computer science has left us breadcrumbs: C4, TLA+, state-machine thinking, deterministic simulators, small composable tools, and decades of systems literature. AI can help excavate that inheritance. It can explain unfamiliar notation, recover architectural knowledge from old code, draft models, generate harnesses, and make specialized techniques approachable to ordinary teams.&lt;/p&gt;

&lt;p&gt;That is retrieval, not retreat.&lt;/p&gt;

&lt;p&gt;The answer to faster code generation may be older, deeper modeling.&lt;/p&gt;

&lt;h2&gt;
  
  
  Humans own the promises
&lt;/h2&gt;

&lt;p&gt;Shared models also repair an organizational failure.&lt;/p&gt;

&lt;p&gt;Without them, a developer hands QA an implementation full of implicit decisions. QA runs the checks it can see. An edge case reaches production, and the organization asks why testing missed it. The assumption moves downstream, and the blame moves with it.&lt;/p&gt;

&lt;p&gt;With a structural map and explicit invariants, QA no longer has to guess what the developer meant. Testers can challenge the promises, identify missing states, design hostile scenarios, and feed new discoveries back into the model. Quality becomes a shared act of reasoning rather than a final gate operated by the last team in line.&lt;/p&gt;

&lt;p&gt;That restores a form of human value AI cannot generate for us. Our value was never only in typing every line. It is in deciding what a system is for, recognizing who can be harmed, making its promises visible, and accepting responsibility when they are broken.&lt;/p&gt;

&lt;p&gt;Fewer people should be able to dismiss a failure by saying, “The AI wrote it.”&lt;/p&gt;

&lt;p&gt;Generation does not transfer responsibility. If we decide what a system is for, accept its output, and release it into someone’s life, the promises it breaks are still ours.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI may produce the code. Humans own the promises.&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclosure: This article was developed with AI assistance through an iterative editorial process. The argument, examples, editorial direction, and responsibility for accuracy remain with the human author.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>testing</category>
      <category>programming</category>
    </item>
    <item>
      <title>Write down every guarantee before you write any code</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Tue, 11 Aug 2026 03:35:29 +0000</pubDate>
      <link>https://dev.to/copyleftdev/write-down-every-guarantee-before-you-write-any-code-21oi</link>
      <guid>https://dev.to/copyleftdev/write-down-every-guarantee-before-you-write-any-code-21oi</guid>
      <description>&lt;p&gt;Here is every promise a to-do list makes.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;VARIABLE tasks

Init == tasks = [i \in Ids |-&amp;gt; "absent"]

Add(i)      == tasks[i] = "absent" /\ tasks' = [tasks EXCEPT ![i] = "open"]
Complete(i) == tasks[i] = "open"   /\ tasks' = [tasks EXCEPT ![i] = "done"]
Reopen(i)   == tasks[i] = "done"   /\ tasks' = [tasks EXCEPT ![i] = "open"]
Delete(i)   == tasks[i] # "absent" /\ tasks' = [tasks EXCEPT ![i] = "absent"]

ClearCompleted ==
  /\ \E i \in Ids : tasks[i] = "done"
  /\ tasks' = [i \in Ids |-&amp;gt; IF tasks[i] = "done" THEN "absent" ELSE tasks[i]]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Not a summary. Not the important ones. &lt;strong&gt;All of them.&lt;/strong&gt; A task cannot go from&lt;br&gt;
absent straight to done. Clearing completed items leaves the open ones alone.&lt;br&gt;
You cannot delete something that was never there. Nine lines, and when you've&lt;br&gt;
read them you have read the entire contract.&lt;/p&gt;

&lt;p&gt;Now go find that list for the system you work on.&lt;/p&gt;

&lt;p&gt;You can't. It doesn't exist. It's distributed across a test suite that asserts&lt;br&gt;
outcomes rather than rules, some validation scattered through handlers, and the&lt;br&gt;
memory of whoever's been there longest. The guarantees are real — your users&lt;br&gt;
depend on every one of them — and there is no file you can open to see them.&lt;/p&gt;

&lt;p&gt;That's the gap I want to talk about, because you can close it in an afternoon,&lt;br&gt;
and because something has changed recently that makes closing it pay for itself.&lt;/p&gt;
&lt;h2&gt;
  
  
  The prime mark and two operators
&lt;/h2&gt;

&lt;p&gt;That's most of the syntax, so let's get it out of the way.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;tasks'&lt;/code&gt; means "tasks, in the next state." &lt;code&gt;/\&lt;/code&gt; is &lt;em&gt;and&lt;/em&gt;. &lt;code&gt;\E&lt;/code&gt; is "there&lt;br&gt;
exists." A definition like &lt;code&gt;Complete(i)&lt;/code&gt; is a formula relating the current state&lt;br&gt;
to the next one — read it out loud: &lt;em&gt;the task is open, and afterwards it is&lt;br&gt;
done.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That's it. That's the language, near enough, for this purpose.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/copyleftdev/tlatools-rs/tree/main/demo/todo" rel="noopener noreferrer"&gt;The real file&lt;/a&gt; adds about eight lines of scaffolding around what you saw:&lt;br&gt;
a module header, a &lt;code&gt;TypeOK&lt;/code&gt; saying a task is always in exactly one of the three&lt;br&gt;
states, and the two lines that tie the actions together —&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Next == \/ \E i \in Ids : Add(i) \/ Complete(i) \/ Reopen(i) \/ Delete(i)
        \/ ClearCompleted

Spec == Init /\ [][Next]_tasks
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;Next&lt;/code&gt; is "any one of the moves happens." &lt;code&gt;Spec&lt;/code&gt; is "start legally, and then only&lt;br&gt;
ever make legal moves." That second line turns out to matter more than it looks,&lt;br&gt;
and I'll come back to it.&lt;/p&gt;

&lt;p&gt;Notice what isn't in there. No database. No HTTP. No mention of whether the&lt;br&gt;
button is blue or whether completion is optimistic in the UI. A specification&lt;br&gt;
isn't a program and doesn't compile to one — it's a formula that says which&lt;br&gt;
state changes are permitted. Everything else is out of scope by construction,&lt;br&gt;
which is exactly why the list can be nine lines and still be complete.&lt;/p&gt;

&lt;p&gt;And notice &lt;code&gt;ClearCompleted&lt;/code&gt; has two halves: the button only exists when&lt;br&gt;
something is done, &lt;strong&gt;and&lt;/strong&gt; it leaves everything else alone. Two separate&lt;br&gt;
promises in one action. Hold that thought.&lt;/p&gt;
&lt;h2&gt;
  
  
  The list is short, finite, and worth arguing about
&lt;/h2&gt;

&lt;p&gt;The objection I expect is that a real system's list would be enormous.&lt;/p&gt;

&lt;p&gt;It's smaller than you think, because it's a list of &lt;em&gt;rules&lt;/em&gt;, not behaviours. The&lt;br&gt;
behaviours are combinatorial — nine states here, and a real system has&lt;br&gt;
astronomically many. The rules that generate them are not. Five actions cover&lt;br&gt;
every to-do list that has ever been correct.&lt;/p&gt;

&lt;p&gt;It's also the part of the design worth arguing about. When two engineers&lt;br&gt;
disagree about whether reopening a completed task should be allowed, that&lt;br&gt;
argument currently happens in a code review, in a comment thread, three weeks&lt;br&gt;
after someone already built one of the answers. Written as a spec, the argument&lt;br&gt;
takes four minutes and happens before anyone opens an editor.&lt;/p&gt;

&lt;p&gt;That's the &lt;a href="https://cacm.acm.org/research/how-amazon-web-services-uses-formal-methods/" rel="noopener noreferrer"&gt;AWS result&lt;/a&gt;, really. They wrote up their experience in CACM in&lt;br&gt;
2015 and the headline everyone quotes is about proving systems correct. The part&lt;br&gt;
that actually replicates is quieter: &lt;strong&gt;writing the spec found bugs before any&lt;br&gt;
code existed&lt;/strong&gt; — in systems their best engineers had already designed and&lt;br&gt;
reviewed. Not bugs the tests missed. Bugs the &lt;em&gt;design&lt;/em&gt; had, findable by writing&lt;br&gt;
the guarantees down and reading them back.&lt;/p&gt;

&lt;p&gt;This is forty-year-old technology, and most of us skipped it because it looked&lt;br&gt;
like homework. TLA+ is Leslie Lamport's; the temporal logic underneath it landed&lt;br&gt;
in &lt;a href="https://dl.acm.org/doi/10.1145/177492.177726" rel="noopener noreferrer"&gt;TOPLAS in 1994&lt;/a&gt;, the language and tools got a &lt;a href="https://lamport.azurewebsites.net/tla/book.html" rel="noopener noreferrer"&gt;book in 2002&lt;/a&gt;,&lt;br&gt;
and Lamport picked up the &lt;a href="https://amturing.acm.org/award_winners/lamport_1205376.cfm" rel="noopener noreferrer"&gt;2013 Turing Award&lt;/a&gt; along the way. (Not &lt;em&gt;for&lt;/em&gt;&lt;br&gt;
TLA+, worth saying, since people get this wrong: the citation is logical clocks,&lt;br&gt;
safety and liveness, replicated state machines, sequential consistency. TLA+ is&lt;br&gt;
downstream of that work, not the reason for the medal.)&lt;/p&gt;

&lt;p&gt;Its reputation for being academic is partly earned and mostly out of date. You&lt;br&gt;
do not need the proof system. You do not need to verify anything. You need the&lt;br&gt;
part where you write the guarantees down.&lt;/p&gt;
&lt;h2&gt;
  
  
  What changed
&lt;/h2&gt;

&lt;p&gt;Writing the list has always been worth it and has always been easy to defer,&lt;br&gt;
because the code was going to be written slowly by people who mostly remembered&lt;br&gt;
the rules.&lt;/p&gt;

&lt;p&gt;That is no longer the situation. Something else is writing the code now, quickly,&lt;br&gt;
and it does not remember anything. It has never met your system's rules and has&lt;br&gt;
no way to infer the ones that aren't in the file it's looking at. It will write&lt;br&gt;
something plausible.&lt;/p&gt;

&lt;p&gt;Plausible is the problem. Plausible code passes review — this is where "looks&lt;br&gt;
good to me" comes from, and it was always an honest confession: the reviewer is&lt;br&gt;
reporting that nothing jumped out, because checking against the full set of&lt;br&gt;
invariants was never an option. Nobody had the list.&lt;/p&gt;

&lt;p&gt;So: write the list. Then check the generated code against it, mechanically,&lt;br&gt;
every time. That second half needs a tool.&lt;/p&gt;
&lt;h2&gt;
  
  
  tlatools-rs
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;cargo &lt;span class="nb"&gt;install &lt;/span&gt;tlatools
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;A TLA+ parser and evaluator in Rust. Not a model checker — it doesn't explore&lt;br&gt;
anything. It answers questions about states you already have:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;spec&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nn"&gt;Spec&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;from_file&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Todo.tla"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;?&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;eval&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nn"&gt;Evaluator&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;spec&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;constants&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;?&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="n"&gt;eval&lt;/span&gt;&lt;span class="nf"&gt;.holds_at&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Init"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;?&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;         &lt;span class="c1"&gt;// legal starting state?&lt;/span&gt;
&lt;span class="n"&gt;eval&lt;/span&gt;&lt;span class="nf"&gt;.step_allowed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Next"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;from&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;to&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;?&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// legal step?&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The loop is three pieces. &lt;strong&gt;You write the list&lt;/strong&gt; — short, arguable, and it&lt;br&gt;
barely changes. &lt;strong&gt;The agent writes the implementation&lt;/strong&gt; — any language, any&lt;br&gt;
framework, any speed. &lt;strong&gt;A script walks the implementation and asks the list&lt;br&gt;
about every step it takes.&lt;/strong&gt; That third piece is thirty lines: ask the&lt;br&gt;
implementation what it can do, do each of those things, record where you landed,&lt;br&gt;
repeat until nothing new turns up.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;./check.py impl/correct.py
&lt;span class="go"&gt;The implementation refines the specification.
9 states and 35 steps, all permitted.
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nine states because there are two tasks and three states each. A real app has&lt;br&gt;
more, and the walk is the expensive part, not the checking.&lt;/p&gt;
&lt;h2&gt;
  
  
  Two bugs the list catches
&lt;/h2&gt;

&lt;p&gt;Here's an agent-plausible one. The completion handler takes an id and marks it&lt;br&gt;
done. It doesn't check the task was open — why would it, the button only shows&lt;br&gt;
up on open tasks. (The button. Not the handler.)&lt;/p&gt;

&lt;p&gt;This is exactly the bug that survives review. It reads correctly. The missing&lt;br&gt;
check is missing somewhere you aren't looking.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;./check.py impl/completes_anything.py
&lt;span class="go"&gt;The implementation takes a step the specification does not permit.

  from   a=absent, b=absent
  doing  complete(a)
  to     a=done, b=absent

The closest the specification came:
  Add(i = "a") was available, but does not produce that state,
    because tasks' = [tasks EXCEPT ![i] = Open] does not hold (1 of its 2 clauses hold)
  Complete(i = "a") was not available here,
    because tasks[i] = Open does not hold (1 of its 2 clauses hold)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second line is the bug, named: &lt;code&gt;Complete&lt;/code&gt; requires the task to be open, and&lt;br&gt;
it wasn't. It's a ranked shortlist rather than a single guess — &lt;code&gt;Add&lt;/code&gt; also nearly&lt;br&gt;
fits from this state, and saying so is more honest than pretending to know which&lt;br&gt;
one you meant.&lt;/p&gt;

&lt;p&gt;Now the other one. &lt;code&gt;ClearCompleted&lt;/code&gt; — the action with two promises. This&lt;br&gt;
implementation keeps the first and breaks the second. It clears the whole list:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;./check.py impl/clear_removes_everything.py
&lt;span class="go"&gt;  from   a=open, b=done
  doing  clear_completed
  to     a=absent, b=absent

The closest the specification came:
  ClearCompleted was available, but does not produce that state,
&lt;/span&gt;&lt;span class="gp"&gt;    because tasks' = [i \in Ids |-&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;IF tasks[i] &lt;span class="o"&gt;=&lt;/span&gt; Done THEN Absent ELSE tasks[i]]
&lt;span class="go"&gt;    does not hold (1 of its 2 clauses hold)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;"Was not available here" versus "was available, but does not produce that&lt;br&gt;
state."&lt;/strong&gt; Different sentences because they're different bugs. One is a missing&lt;br&gt;
guard. The other is a correct guard and a wrong effect — which is worse, because&lt;br&gt;
the button &lt;em&gt;looks&lt;/em&gt; like it works. You'd demo it. You'd ship it. Someone would&lt;br&gt;
lose a task they hadn't finished.&lt;/p&gt;

&lt;p&gt;The tool can tell them apart because it knows which failing clause mentions the&lt;br&gt;
next state. Neither bug is exotic. Both are invisible to a test suite that&lt;br&gt;
checks outcomes, and both are named instantly by a list you wrote in nine lines.&lt;/p&gt;
&lt;h2&gt;
  
  
  Feedback in the language the rule was written in
&lt;/h2&gt;

&lt;p&gt;You cannot fix what you cannot describe, and neither can a model.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"Tests failed." — Try again. Randomly.&lt;/li&gt;
&lt;li&gt;"Expected &lt;code&gt;{a: open}&lt;/code&gt;, got &lt;code&gt;{a: absent}&lt;/code&gt;." — Better. Now infer the rule.&lt;/li&gt;
&lt;li&gt;"&lt;code&gt;ClearCompleted&lt;/code&gt; was available, but does not produce that state, because
&lt;code&gt;tasks' = [i \in Ids |-&amp;gt; IF tasks[i] = Done THEN Absent ELSE tasks[i]]&lt;/code&gt; does
not hold." — The action, the condition, and the state it was in.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That third one is a prompt.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And I measured whether it helps an agent, and it didn't — not detectably.&lt;/strong&gt;&lt;br&gt;
200 tasks, each attempted with an uninformative retry and with the failure text&lt;br&gt;
fed back: 90.5% [85.6–93.8] against 92.5% [88.0–95.4], McNemar exact p=0.125.&lt;br&gt;
That is a null. On the formal-reasoning subset the gap was 46.2% → 69.2%, which&lt;br&gt;
looks like something, except n=13 and p=0.25, which means it looks like&lt;br&gt;
something in the way small numbers often do.&lt;/p&gt;

&lt;p&gt;I'm reporting it because I ran it. The honest state of the claim: the mechanism&lt;br&gt;
is sound, the message is strictly more information than a boolean, and I have no&lt;br&gt;
evidence it moves the pass rate. If you were going to adopt this because "agents&lt;br&gt;
do better with good errors" — don't, yet. Adopt it because &lt;strong&gt;you&lt;/strong&gt; now have the&lt;br&gt;
list, and something checks it.&lt;/p&gt;
&lt;h2&gt;
  
  
  The list as a grader
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;tlatools check&lt;/code&gt; takes a JSON job — spec, states, steps, constants — and returns&lt;br&gt;
a verdict with an exit status: &lt;code&gt;0&lt;/code&gt; it refines, &lt;code&gt;1&lt;/code&gt; it doesn't, &lt;code&gt;2&lt;/code&gt; the question&lt;br&gt;
was malformed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;tlatools check job.json &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Point it at N candidate implementations and it tells you which satisfy the list&lt;br&gt;
and, for the ones that don't, exactly where they diverge. If you're generating&lt;br&gt;
code, evaluating models, or grading a benchmark, that's a grader with no rubric&lt;br&gt;
to write and no partial credit to argue about. The list &lt;em&gt;is&lt;/em&gt; the rubric.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;init&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;is the starting state legal?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;refines&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;is every step one the spec permits?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;coverage&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;does the implementation reach the outcomes it should?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That third one exists because refinement alone is satisfied perfectly by an&lt;br&gt;
implementation that does nothing. Ask me how I know.&lt;/p&gt;

&lt;h2&gt;
  
  
  Can you trust the checker?
&lt;/h2&gt;

&lt;p&gt;Fair question to ask of anything that grades your code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It agrees with the reference implementation.&lt;/strong&gt; Over a labelled corpus of 39&lt;br&gt;
cases — six implementations that must pass, thirty-three seeded bugs that must&lt;br&gt;
each be caught — its verdicts are byte-identical to Java TLC's, including &lt;em&gt;which&lt;/em&gt;&lt;br&gt;
of the three checks catches which bug.&lt;/p&gt;

&lt;p&gt;Getting that diff empty taught me something I'd otherwise have shipped wrong,&lt;br&gt;
and it's the best argument in this article for writing guarantees down precisely.&lt;/p&gt;

&lt;p&gt;Remember &lt;code&gt;Spec == Init /\ [][Next]_tasks&lt;/code&gt;, the line I said would matter. Those&lt;br&gt;
brackets are load-bearing: &lt;code&gt;[Next]_tasks&lt;/code&gt; means &lt;em&gt;&lt;code&gt;Next&lt;/code&gt;, **or nothing changed&lt;/em&gt;**.&lt;br&gt;
Stuttering is always permitted, in every TLA+ specification ever written. I had&lt;br&gt;
been checking bare &lt;code&gt;Next&lt;/code&gt;, so any implementation that idled or retried got&lt;br&gt;
flagged for a step the spec explicitly allows.&lt;/p&gt;

&lt;p&gt;Fixing that made one seeded bug survive: a transfer from a bank account to&lt;br&gt;
itself. Which felt like a regression, until I read the spec again. A&lt;br&gt;
self-transfer nets to zero. It changes nothing. It &lt;strong&gt;is&lt;/strong&gt; a stuttering step, and&lt;br&gt;
the spec says stuttering is fine — so that implementation genuinely satisfies&lt;br&gt;
the list, and the benchmark and I had both been wrong about it. Catching that&lt;br&gt;
one needs an abstraction where the operation is visible in the state at all. You&lt;br&gt;
cannot tighten a refinement check into seeing something the state space doesn't&lt;br&gt;
record.&lt;/p&gt;

&lt;p&gt;Hand TLC the same &lt;code&gt;[Next]_vars&lt;/code&gt; obligation and it passes that mutant too. The&lt;br&gt;
agreement holds; what moved was my understanding of what I'd written down. The&lt;br&gt;
list is only as good as your reading of it, and a tool that disagrees with you&lt;br&gt;
is doing you a favour.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It reads the language.&lt;/strong&gt; 1,256 of the 1,258 specifications in three public&lt;br&gt;
TLA+ corpora: the examples repo, the community modules, and the TLA+ tools' own&lt;br&gt;
test suite. The two it doesn't read are two that SANY, the official parser,&lt;br&gt;
doesn't read either. How every one of those files is read is recorded in&lt;br&gt;
&lt;code&gt;golden/&lt;/code&gt;, so a change names the files it changed instead of moving a number.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It's checked in the boring ways.&lt;/strong&gt; 166 tests, clippy-pedantic clean, and a&lt;br&gt;
robustness suite that feeds back every prefix and every dropped line of every&lt;br&gt;
fixture — because a parser that panics on a half-saved file is a parser you&lt;br&gt;
can't put in a loop. &lt;code&gt;tla-syntax&lt;/code&gt; and &lt;code&gt;tla-eval&lt;/code&gt; have no external dependencies at&lt;br&gt;
all.&lt;/p&gt;

&lt;h2&gt;
  
  
  About TLC, precisely
&lt;/h2&gt;

&lt;p&gt;TLC is the model checker TLA+ ships with. It explores: from your initial state,&lt;br&gt;
apply every action, walk the reachable state space looking for a violation.&lt;br&gt;
That's the right question when you're designing a protocol and don't yet know&lt;br&gt;
what your system can do.&lt;/p&gt;

&lt;p&gt;Here you already have the implementation, so the question is different — are&lt;br&gt;
&lt;em&gt;these&lt;/em&gt; steps, the ones the code just took, permitted by the list?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TLC can be asked that.&lt;/strong&gt; I want to be precise, because I had this wrong in an&lt;br&gt;
earlier draft and someone would have caught it: you encode the steps as data,&lt;br&gt;
write a plain safety invariant asserting each one is enabled, and run the&lt;br&gt;
checker. No liveness property, no engineered failure. On the to-do spec it finds&lt;br&gt;
the bad step in 0.78 seconds. There's published work doing this properly —&lt;br&gt;
&lt;a href="https://arxiv.org/abs/2404.16075" rel="noopener noreferrer"&gt;&lt;em&gt;Validating Traces of Distributed Programs Against TLA+ Specifications&lt;/em&gt;&lt;/a&gt;&lt;br&gt;
by Cirstea, Kuppe, Loillier and Merz, and Kuppe maintains the TLA+ tools. If you&lt;br&gt;
want trace validation on a production system, start there.&lt;/p&gt;

&lt;p&gt;What's left is narrower, and true. TLC tells you &lt;em&gt;that&lt;/em&gt; the step is illegal, not&lt;br&gt;
&lt;em&gt;which conjunct&lt;/em&gt; failed. And that 0.78 s is almost entirely JVM boot and SANY&lt;br&gt;
parse, paid again on every query — fine once, not fine in a loop that runs on&lt;br&gt;
every agent edit.&lt;/p&gt;

&lt;p&gt;Structured blame instead of a boolean, at a couple of orders of magnitude less&lt;br&gt;
latency. That's the pitch. These compose; they don't compete.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it won't do
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It's not a model checker.&lt;/strong&gt; If you want to know what states your system can
reach, use TLC.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It doesn't verify your program.&lt;/strong&gt; It checks the transitions you hand it. If
your walk misses a path, nothing checks that path.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your list can be wrong.&lt;/strong&gt; It's what you argued about, not what's true. I
could have written &lt;code&gt;ClearCompleted&lt;/code&gt; to clear everything, and then the "buggy"
implementation would be the correct one. The list being short and readable is
the only defence, which is an argument for keeping it short and readable.&lt;/li&gt;
&lt;li&gt;Integers are 64-bit, real arithmetic isn't implemented, temporal formulas are
refused rather than guessed at, and TLAPS proofs are skipped.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/copyleftdev/tlatools-rs
&lt;span class="nb"&gt;cd &lt;/span&gt;tlatools-rs
cargo build &lt;span class="nt"&gt;--release&lt;/span&gt;
demo/todo/check.py demo/todo/impl/completes_anything.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The whole demo is &lt;a href="https://github.com/copyleftdev/tlatools-rs/tree/main/demo/todo" rel="noopener noreferrer"&gt;&lt;code&gt;demo/todo&lt;/code&gt;&lt;/a&gt; — the spec above, three implementations,&lt;br&gt;
and the script that checks them. CI runs it on every push, so if it's broken&lt;br&gt;
when you get there, that's a bug and I'd like to hear about it.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;crates:&lt;/strong&gt; &lt;a href="https://crates.io/crates/tlatools" rel="noopener noreferrer"&gt;tlatools&lt;/a&gt; · &lt;a href="https://crates.io/crates/tla-eval" rel="noopener noreferrer"&gt;tla-eval&lt;/a&gt; · &lt;a href="https://crates.io/crates/tla-syntax" rel="noopener noreferrer"&gt;tla-syntax&lt;/a&gt; · &lt;a href="https://crates.io/crates/tla-oracle" rel="noopener noreferrer"&gt;tla-oracle&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;source:&lt;/strong&gt; &lt;a href="https://github.com/copyleftdev/tlatools-rs" rel="noopener noreferrer"&gt;github.com/copyleftdev/tlatools-rs&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;docs:&lt;/strong&gt; &lt;a href="https://docs.rs/tla-eval" rel="noopener noreferrer"&gt;docs.rs/tla-eval&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you have a TLA+ file it reads wrongly, that's the most useful thing you can&lt;br&gt;
send me.&lt;/p&gt;




&lt;p&gt;Start with one. Pick the part of your system where a wrong state change would&lt;br&gt;
actually hurt — the money, the permissions, the thing with a state machine&lt;br&gt;
nobody fully trusts. Write down what it's allowed to do. It'll take an afternoon&lt;br&gt;
and it will be shorter than you expect.&lt;/p&gt;

&lt;p&gt;You'll find something while writing it. Everyone does; that's the AWS result and&lt;br&gt;
it isn't subtle. And then you'll have the file — the one that doesn't exist for&lt;br&gt;
any system you currently work on — and everything written afterwards, by you or&lt;br&gt;
by a machine, can be checked against it.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>rust</category>
      <category>formalmethods</category>
      <category>testing</category>
    </item>
    <item>
      <title>The Shape of Failure: Before You Blame the AI</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Sat, 01 Aug 2026 21:12:51 +0000</pubDate>
      <link>https://dev.to/copyleftdev/the-shape-of-failure-before-you-blame-the-ai-5358</link>
      <guid>https://dev.to/copyleftdev/the-shape-of-failure-before-you-blame-the-ai-5358</guid>
      <description>&lt;p&gt;Every automated system receives a particular shape of the world.&lt;/p&gt;

&lt;p&gt;That shape is expressed through records, documents, events, exceptions, and missing values. If the designers have not identified those forms—and the ways they can become malformed—the machine inherits their ignorance and reproduces it at scale.&lt;/p&gt;


&lt;div class="crayons-card c-embed"&gt;

  &lt;br&gt;
&lt;strong&gt;The question is not simply whether the AI failed.&lt;/strong&gt;

&lt;p&gt;The useful question is whether the human-built system knew what success meant, knew the shape of its data, and knew how to recognize when it was wrong.&lt;br&gt;

&lt;/p&gt;
&lt;/div&gt;


&lt;h2&gt;
  
  
  Start with the shape of the data
&lt;/h2&gt;

&lt;p&gt;Before selecting a model, draw the workflow as a sequence of data transformations.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What enters each stage?&lt;/li&gt;
&lt;li&gt;In what form and from what source?&lt;/li&gt;
&lt;li&gt;Which values are valid, absent, duplicated, stale, delayed, or contradictory?&lt;/li&gt;
&lt;li&gt;How will each violation be detected?&lt;/li&gt;
&lt;li&gt;What must the workflow do next?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each data shape needs a corresponding failure model. An unknown here is not merely uncertainty for the machine; it is a measurement failure in the organization.&lt;/p&gt;

&lt;p&gt;The remedy is to collect the missing data or explicitly design for its absence. Otherwise, the system is being asked to operate in a world its designers have not described.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stabilize the deliverable
&lt;/h2&gt;

&lt;p&gt;A system cannot be stabilized around a target that continues to move.&lt;/p&gt;

&lt;p&gt;The deliverable must be more than an aspiration written in a prompt. It should be expressed as observable conditions and anchored to a representative corpus:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;examples that are acceptable;&lt;/li&gt;
&lt;li&gt;examples that are unacceptable;&lt;/li&gt;
&lt;li&gt;examples that are genuinely ambiguous.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Human reviewers should first demonstrate that they can apply those distinctions consistently. If they cannot agree on what success looks like, the model is not being measured against a specification. It is being measured against human disagreement disguised as one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The model is not the system
&lt;/h2&gt;

&lt;p&gt;Only then does it become meaningful to place an AI model inside the workflow.&lt;/p&gt;

&lt;p&gt;The model is one transformation among many:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Input
  → validation
  → retrieval
  → normalization
  → model inference
  → output validation
  → policy checks
  → human action
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Each transition can discard meaning or introduce error. The fluency of the final response makes the model the most conspicuous suspect, but conspicuousness is not causality.&lt;/p&gt;

&lt;p&gt;To assign blame intelligently, observe the entire chain and test every boundary where information is received or transformed.&lt;/p&gt;
&lt;h2&gt;
  
  
  Mutation testing attacks confidence itself
&lt;/h2&gt;

&lt;p&gt;Conventional tests ask whether software succeeds under conditions its authors anticipated. Mutation testing reverses that pressure.&lt;/p&gt;

&lt;p&gt;It deliberately introduces small faults—reversing a condition, moving a boundary, substituting a value, or removing an operation—and asks whether the test suite notices.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A &lt;strong&gt;killed mutant&lt;/strong&gt; caused at least one test to fail.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;surviving mutant&lt;/strong&gt; changed the program without the tests objecting.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The resulting score is not proof of correctness. It measures how sensitive the tests are to the generated changes.&lt;/p&gt;

&lt;p&gt;In the companion implementation, the focused safety-contract run produced:&lt;/p&gt;

&lt;p&gt;

&lt;/p&gt;
&lt;div class="katex-element"&gt;
  &lt;span class="katex-display"&gt;&lt;span class="katex"&gt;&lt;span class="katex-mathml"&gt;&lt;/span&gt;&lt;span class="katex-html"&gt;&lt;span class="base"&gt;&lt;span class="strut"&gt;&lt;/span&gt;&lt;span class="mord text"&gt;&lt;span class="mord"&gt;observed&amp;nbsp;mutation&amp;nbsp;sensitivity&lt;/span&gt;&lt;/span&gt;&lt;span class="mspace"&gt;&lt;/span&gt;&lt;span class="mrel"&gt;=&lt;/span&gt;&lt;span class="mspace"&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="base"&gt;&lt;span class="strut"&gt;&lt;/span&gt;&lt;span class="mord"&gt;&lt;span class="mopen nulldelimiter"&gt;&lt;/span&gt;&lt;span class="mfrac"&gt;&lt;span class="vlist-t vlist-t2"&gt;&lt;span class="vlist-r"&gt;&lt;span class="vlist"&gt;&lt;span&gt;&lt;span class="pstrut"&gt;&lt;/span&gt;&lt;span class="mord"&gt;&lt;span class="mord"&gt;43&lt;/span&gt;&lt;span class="mspace"&gt;&amp;nbsp;&lt;/span&gt;&lt;span class="mord text"&gt;&lt;span class="mord"&gt;generated&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class="pstrut"&gt;&lt;/span&gt;&lt;span class="frac-line"&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span class="pstrut"&gt;&lt;/span&gt;&lt;span class="mord"&gt;&lt;span class="mord"&gt;43&lt;/span&gt;&lt;span class="mspace"&gt;&amp;nbsp;&lt;/span&gt;&lt;span class="mord text"&gt;&lt;span class="mord"&gt;killed&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="vlist-s"&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span class="vlist-r"&gt;&lt;span class="vlist"&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="mclose nulldelimiter"&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="mspace"&gt;&lt;/span&gt;&lt;span class="mrel"&gt;=&lt;/span&gt;&lt;span class="mspace"&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="base"&gt;&lt;span class="strut"&gt;&lt;/span&gt;&lt;span class="mord"&gt;1.0&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;
&lt;/div&gt;



&lt;p&gt;That number has meaning only because the measurement boundary is explicit: the run scores the small kernel containing the workflow's acceptance predicates. It does not pretend that prompt punctuation or CLI wording is part of the safety contract.&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  Why the first mutation run mattered
  &lt;br&gt;
The first broad run generated 374 mutations across the package. Of those, 185 survived. Many changed prompt text, trace wording, and gateway plumbing that the tests did not claim to specify.

&lt;p&gt;That was not an embarrassing result to conceal. It revealed an unstable definition of the deliverable. The mutation boundary was then narrowed to the safety-critical contract, and tests were strengthened around exact acceptance boundaries. The experiment therefore reenacted the article's argument: a metric becomes meaningful only after humans define what is being measured.&lt;br&gt;
&lt;/p&gt;

&lt;br&gt;
&lt;p&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Carry the same pressure through the AI workflow
&lt;/h2&gt;

&lt;p&gt;The mutation principle should not stop at source code.&lt;/p&gt;

&lt;p&gt;Remove an expected field. Corrupt a format. Supply stale or contradictory source material. Interrupt retrieval. Perturb an instruction. Substitute a fluent but incorrect model response.&lt;/p&gt;

&lt;p&gt;At every boundary, ask whether the surrounding system detects the disturbance, abstains, falls back, or routes the case to a human.&lt;/p&gt;

&lt;p&gt;The reference implementation makes this flow explicit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;shape_failures&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;shape_agent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;inspect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;shape_failures&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;escalate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;shape_failures&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;evidence&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;evidence_failure&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;evidence_agent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;retrieve&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;corpus&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;evidence_failure&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;escalate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;evidence_failure&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;decision&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;decide&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;evidence&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;criticism&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;critic_agent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;review&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;evidence&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;criticism&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;escalate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;criticism&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;accept&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The actual implementation uses four deliberately small agents:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A &lt;strong&gt;shape agent&lt;/strong&gt; rejects malformed or unmodeled input.&lt;/li&gt;
&lt;li&gt;An &lt;strong&gt;evidence agent&lt;/strong&gt; retrieves the matching corpus and detects contradictions.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;model agent&lt;/strong&gt; proposes a structured, cited decision through OpenRouter.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;critic agent&lt;/strong&gt; verifies the decision against citations, confidence thresholds, and a computable policy oracle.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The model is useful, but it is never permitted to define its own success criteria.&lt;/p&gt;
&lt;h2&gt;
  
  
  What the executable experiment found
&lt;/h2&gt;

&lt;p&gt;The checked local verification produced:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;86 passing tests;&lt;/li&gt;
&lt;li&gt;43 of 43 mutations killed in the safety-contract kernel;&lt;/li&gt;
&lt;li&gt;one clean control accepted;&lt;/li&gt;
&lt;li&gt;all seven injected data, corpus, and model faults detected;&lt;/li&gt;
&lt;li&gt;zero model calls when bad input or contradictory evidence made inference unsafe.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://hypothesis.readthedocs.io/en/latest/stateful.html" rel="noopener noreferrer"&gt;Hypothesis&lt;/a&gt; generates both data and sequences of workflow mutations. &lt;a href="https://mutmut.readthedocs.io/en/latest/" rel="noopener noreferrer"&gt;Mutmut&lt;/a&gt; changes the safety-contract source. &lt;a href="https://openrouter.ai/docs/guides/features/structured-outputs" rel="noopener noreferrer"&gt;OpenRouter&lt;/a&gt; supplies the optional live model boundary using a strict JSON Schema response.&lt;/p&gt;

&lt;p&gt;The OpenRouter call is intentionally optional because it uses an external service and may incur cost. The deterministic safety oracle does not depend on network access, a particular provider, or a favorable model response.&lt;/p&gt;
&lt;h2&gt;
  
  
  Agreement is another output, not an oracle
&lt;/h2&gt;

&lt;p&gt;Once the single-model boundary was stable, I added a second weave around it. Four model families received the same structured request and the same policy evidence. Each worked independently: no model saw another model's answer. Every response then crossed the same citation checks, confidence threshold, and computable policy oracle.&lt;/p&gt;

&lt;p&gt;The live run used:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;openai/gpt-5.4-nano&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;google/gemini-3.5-flash-lite&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;deepseek/deepseek-v4-flash-0731&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;mistralai/mistral-small-2603&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Eight scenarios produced 32 model observations. Twenty-seven matched the oracle and passed the contract, but only four scenarios achieved unanimous contract acceptance.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;OpenAI&lt;/th&gt;
&lt;th&gt;Gemini&lt;/th&gt;
&lt;th&gt;DeepSeek&lt;/th&gt;
&lt;th&gt;Mistral&lt;/th&gt;
&lt;th&gt;System&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Eligible control&lt;/td&gt;
&lt;td&gt;approve &lt;code&gt;0.78&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;approve &lt;code&gt;1.0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;approve &lt;code&gt;1.0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;approve &lt;code&gt;1.0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;accept&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Eligible paraphrase&lt;/td&gt;
&lt;td&gt;approve &lt;code&gt;0.74&lt;/code&gt;, rejected&lt;/td&gt;
&lt;td&gt;approve &lt;code&gt;1.0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;approve &lt;code&gt;0.99&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;approve &lt;code&gt;1.0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;review&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Instruction inside customer data&lt;/td&gt;
&lt;td&gt;approve &lt;code&gt;0.74&lt;/code&gt;, rejected&lt;/td&gt;
&lt;td&gt;approve &lt;code&gt;1.0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;approve &lt;code&gt;1.0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;escalate &lt;code&gt;0.95&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;review&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Last eligible day&lt;/td&gt;
&lt;td&gt;approve &lt;code&gt;0.86&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;approve &lt;code&gt;1.0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;approve &lt;code&gt;1.0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;deny &lt;code&gt;1.0&lt;/code&gt;, rejected&lt;/td&gt;
&lt;td&gt;review&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Maximum eligible amount&lt;/td&gt;
&lt;td&gt;approve &lt;code&gt;0.78&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;approve &lt;code&gt;1.0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;approve &lt;code&gt;1.0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;approve &lt;code&gt;1.0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;accept&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One day late&lt;/td&gt;
&lt;td&gt;deny &lt;code&gt;0.86&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;deny &lt;code&gt;1.0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;deny &lt;code&gt;1.0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;deny &lt;code&gt;1.0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;accept&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One cent over&lt;/td&gt;
&lt;td&gt;deny &lt;code&gt;0.90&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;deny &lt;code&gt;1.0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;deny &lt;code&gt;1.0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;escalate &lt;code&gt;0.99&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;review&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Non-refundable policy&lt;/td&gt;
&lt;td&gt;deny &lt;code&gt;0.86&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;deny &lt;code&gt;1.0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;deny &lt;code&gt;1.0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;deny &lt;code&gt;1.0&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;accept&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This is not a leaderboard. It is one small experiment against one explicit contract. Gemini and DeepSeek passed all eight cases in this run. OpenAI selected the oracle action in all eight, but its confidence fell below the human-set &lt;code&gt;0.75&lt;/code&gt; threshold twice. Mistral abstained twice and made one incorrect decision at the exact, inclusive day boundary.&lt;/p&gt;

&lt;p&gt;Those are different failure shapes with different remedies. A wrong boundary decision challenges reasoning or prompt clarity. An abstention challenges workflow capacity and escalation cost. A low-confidence rejection challenges calibration and the threshold chosen by the organization. None can be repaired merely by declaring that three models outvoted the fourth.&lt;/p&gt;


&lt;div class="crayons-card c-embed"&gt;

  &lt;br&gt;
&lt;strong&gt;Agreement measures convergence, not correctness.&lt;/strong&gt;

&lt;p&gt;If every model repeats the same unsupported answer, the deterministic contract must still reject it. If one model dissents for a defensible reason, majority voting must not erase the evidence.&lt;br&gt;

&lt;/p&gt;
&lt;/div&gt;



&lt;p&gt;The weave also exposed instability that a single run would have concealed. In an earlier pass, the same OpenAI model gave the control &lt;code&gt;0.90&lt;/code&gt; confidence instead of &lt;code&gt;0.78&lt;/code&gt;, and escalated the one-day-late case at &lt;code&gt;0.62&lt;/code&gt; instead of denying it at &lt;code&gt;0.86&lt;/code&gt;. Temperature zero reduced one source of variation; it did not turn a remote generative service into a mathematical function.&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  A provider failure is not a reasoning failure
  &lt;br&gt;
The pilot included &lt;code&gt;anthropic/claude-haiku-4.5&lt;/code&gt;, but that route failed at the model boundary in all eight requests. The workflow escalated without leaking upstream details. I replaced the route with Mistral for the final decision comparison rather than silently counting eight transport or adapter failures as eight reasoning failures.

&lt;p&gt;That distinction is the thesis again: identify the boundary at which the failure becomes observable before deciding what deserves blame.&lt;br&gt;
&lt;/p&gt;

&lt;br&gt;
&lt;p&gt;&lt;/p&gt;

&lt;p&gt;Different models did help us understand more—but only because the data shape, policy oracle, and rejection rules already existed. Without those controls, the experiment would have produced four persuasive rationales and no principled way to interpret their disagreement.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/copyleftdev/the-shape-of-failure/tree/main/computational-expression" class="crayons-btn crayons-btn--primary" rel="noopener noreferrer"&gt;Run the single-model and multi-model computational expressions&lt;/a&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  When is it truly an AI failure?
&lt;/h2&gt;

&lt;p&gt;Only after this discipline has been applied does the phrase &lt;em&gt;AI failure&lt;/em&gt; acquire useful meaning.&lt;/p&gt;

&lt;p&gt;If the input was well formed, the deliverable was stable, the supporting components behaved correctly, and the model still violated a known requirement, then the model failed within the conditions established for it.&lt;/p&gt;

&lt;p&gt;But if the data was absent, the target disputed, retrieval defective, or safeguards unable to recognize error, the model may merely be the place where an earlier failure became visible. Calling that an AI failure does not improve the system. It interrupts the feedback by which the organization might learn.&lt;/p&gt;

&lt;p&gt;A generative model is an engine of learned resemblance. It can produce behavior that looks remarkably like understanding, but resemblance cannot carry responsibility. People choose the objective, define acceptable evidence, engineer the transitions, and decide what happens when uncertainty enters the system.&lt;/p&gt;

&lt;p&gt;Before declaring that the AI failed, demonstrate that the system knew what success meant, knew the shape of its world, and knew how to recognize when it was wrong.&lt;/p&gt;

&lt;p&gt;If it did not, the failure began long before the model answered.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;The complete experiment includes the agent weave, independent multi-model comparison, property-based tests, stateful tests, mutation configuration, deterministic fault-injection CLI, verification record, and OpenRouter gateway.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>python</category>
      <category>softwaredevelopment</category>
    </item>
    <item>
      <title>I Had a Lot of Fun Building a Linux Packet Flight Recorder</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Tue, 28 Jul 2026 00:05:35 +0000</pubDate>
      <link>https://dev.to/copyleftdev/i-had-a-lot-of-fun-building-a-linux-packet-flight-recorder-k44</link>
      <guid>https://dev.to/copyleftdev/i-had-a-lot-of-fun-building-a-linux-packet-flight-recorder-k44</guid>
      <description>&lt;p&gt;Some projects begin with a roadmap. &lt;code&gt;skbx&lt;/code&gt; began with me falling into a Linux&lt;br&gt;
networking rabbit hole and enjoying it much more than I expected.&lt;/p&gt;

&lt;p&gt;The question was simple:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Where did this packet actually go?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The answer was not.&lt;/p&gt;

&lt;p&gt;A packet can pass through network namespaces, routing decisions, Netfilter,&lt;br&gt;
traffic control, XDP programs, tunnels, clones, copies, and drop paths. A&lt;br&gt;
packet capture at one interface can be completely correct while still showing&lt;br&gt;
only one part of that journey.&lt;/p&gt;

&lt;p&gt;I wanted to see more of the journey. Then I wanted to save what I saw, replay&lt;br&gt;
it later without root, and know whether the capture itself had lost&lt;br&gt;
observations.&lt;/p&gt;

&lt;p&gt;That became &lt;a href="https://github.com/copyleftdev/skbx" rel="noopener noreferrer"&gt;skbx&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;
  
  
  The short version
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;skbx&lt;/code&gt; is a Linux packet-path flight recorder built with Rust and CO-RE eBPF.&lt;/p&gt;

&lt;p&gt;During a live capture, it observes kernel networking functions and writes an&lt;br&gt;
append-only JSONL evidence stream. Afterward, that stream can be replayed and&lt;br&gt;
inspected without loading another eBPF program.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;live traffic
    ↓
CO-RE eBPF observations
    ↓
bounded traceq JSONL
    ↓
replay → route patterns → explain an event
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;It can observe:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;kernel functions handling an &lt;code&gt;sk_buff&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;packet clones, copies, copy-on-write, and XDP-to-SKB transitions;&lt;/li&gt;
&lt;li&gt;TC and XDP program entry and exit;&lt;/li&gt;
&lt;li&gt;tunnels and inner packet tuples;&lt;/li&gt;
&lt;li&gt;kernel-reported drop reasons;&lt;/li&gt;
&lt;li&gt;selected BPF helper and map activity;&lt;/li&gt;
&lt;li&gt;capture loss, decoding failures, and output failures.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is not a replacement for tcpdump or Wireshark. Those tools are excellent&lt;br&gt;
when the question is about packet bytes and protocols at a capture point.&lt;br&gt;
&lt;code&gt;skbx&lt;/code&gt; is for questions about the path through the local Linux kernel.&lt;/p&gt;

&lt;p&gt;It is also not the first tool to follow packets through that path. The project&lt;br&gt;
is explicitly inspired by&lt;br&gt;
&lt;a href="https://github.com/cilium/pwru" rel="noopener noreferrer"&gt;pwru&lt;/a&gt;, which showed how useful broad eBPF&lt;br&gt;
packet tracing can be.&lt;/p&gt;

&lt;p&gt;The part I especially wanted to explore was the evidence after the live&lt;br&gt;
terminal stopped scrolling.&lt;/p&gt;
&lt;h2&gt;
  
  
  Why call it a flight recorder?
&lt;/h2&gt;

&lt;p&gt;Live tracing is useful, but incident work often continues after the privileged&lt;br&gt;
session has ended.&lt;/p&gt;

&lt;p&gt;Someone else may need to inspect the capture. We may want to compare it with a&lt;br&gt;
successful request. We may need to cite one exact event in a bug report. An&lt;br&gt;
automation or AI system may need structured input, but it should not be allowed&lt;br&gt;
to invent observations that were never captured.&lt;/p&gt;

&lt;p&gt;So the native &lt;code&gt;skbx&lt;/code&gt; stream, called &lt;code&gt;traceq&lt;/code&gt;, has an envelope:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;capture_start
event
event
...
capture_end
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Every event receives a stable handle:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;event:111111111111111111111111
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Replay groups ordered events into bounded route patterns, which receive their&lt;br&gt;
own handles:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;route:1c0201e74424e253f8363577
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The footer matters as much as the events. It records the stop reason and&lt;br&gt;
reliability counters for kernel reservation failures, tracer recursion misses,&lt;br&gt;
read failures, userspace decoding and enrichment failures, and output&lt;br&gt;
failures.&lt;/p&gt;

&lt;p&gt;If the footer is absent, the artifact is incomplete. If the tracer lost&lt;br&gt;
observations, that uncertainty stays attached to the result.&lt;/p&gt;

&lt;p&gt;I like this property because tracing software is also software. It should not&lt;br&gt;
silently act omniscient.&lt;/p&gt;
&lt;h2&gt;
  
  
  Installing it
&lt;/h2&gt;

&lt;p&gt;The live tracer currently supports Linux on x86_64 and arm64. It needs Rust&lt;br&gt;
1.85 or newer, Clang/LLVM with the BPF backend, &lt;code&gt;bpftool&lt;/code&gt;, libelf, libpcap, and&lt;br&gt;
a kernel exposing &lt;code&gt;/sys/kernel/btf/vmlinux&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;On Ubuntu, the native packages are:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;apt-get &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="s2"&gt;"linux-tools-&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;uname&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  clang llvm libelf-dev libpcap-dev pkg-config
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;On Debian:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;apt-get &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  bpftool clang llvm libelf-dev libpcap-dev pkg-config
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Then install the CLI directly from GitHub:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;cargo &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;--git&lt;/span&gt; https://github.com/copyleftdev/skbx &lt;span class="nt"&gt;--locked&lt;/span&gt; skbx-cli
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Before attaching anything, &lt;code&gt;doctor&lt;/code&gt; checks the host:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;skbx doctor &lt;span class="nt"&gt;--json&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;code&gt;plan&lt;/code&gt; goes one step further. It shows exactly which functions would be&lt;br&gt;
attached without performing the attachment:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;skbx plan &lt;span class="nt"&gt;--filter-func&lt;/span&gt; &lt;span class="s1"&gt;'ip.*'&lt;/span&gt; &lt;span class="nt"&gt;--json&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;That separation turned out to be useful while building the project. A missing&lt;br&gt;
kernel capability should be visible before a privileged capture begins, not&lt;br&gt;
silently approximated afterward.&lt;/p&gt;
&lt;h2&gt;
  
  
  Capturing a first packet
&lt;/h2&gt;

&lt;p&gt;Here is a small ICMP capture:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;skbx capture &lt;span class="nt"&gt;--probe&lt;/span&gt; ip_rcv &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--duration&lt;/span&gt; 10 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--output&lt;/span&gt; trace.jsonl &lt;span class="se"&gt;\&lt;/span&gt;
  icmp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;While it runs, create some traffic in another terminal:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ping &lt;span class="nt"&gt;-c&lt;/span&gt; 3 1.1.1.1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The capture is deliberately bounded. It has a finite duration and event limit&lt;br&gt;
instead of assuming that an unbounded trace is safe.&lt;/p&gt;

&lt;p&gt;After it finishes, replay does not require root:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;skbx replay trace.jsonl
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Replay summarizes the functions, processes, packet identities, route patterns,&lt;br&gt;
consensus, outliers, and reliability state. Given the same valid JSONL input,&lt;br&gt;
it produces a byte-identical summary.&lt;/p&gt;

&lt;p&gt;If an event looks interesting, retrieve it with its surrounding same-packet&lt;br&gt;
evidence:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;skbx explain trace.jsonl event:&amp;lt;handle&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The repository includes a small contract fixture containing an &lt;code&gt;ip_rcv&lt;/code&gt; event&lt;br&gt;
followed by &lt;code&gt;kfree_skb_reason&lt;/code&gt;. Replaying that fixture produces a route shaped&lt;br&gt;
like this:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"handle"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"route:1c0201e74424e253f8363577"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"functions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"ip_rcv"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"kfree_skb_reason"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"routes"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"truncated"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"outlier"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;That output was generated by the current CLI from the checked-in test fixture;&lt;br&gt;
it is not illustrative pseudodata.&lt;/p&gt;
&lt;h2&gt;
  
  
  Use case: what happened to a website request?
&lt;/h2&gt;

&lt;p&gt;This is the use case I keep returning to.&lt;/p&gt;

&lt;p&gt;Suppose &lt;code&gt;curl&lt;/code&gt; starts an HTTPS request and eventually times out. A capture on&lt;br&gt;
the client can answer questions such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Did the request enter the local TCP and IP output paths?&lt;/li&gt;
&lt;li&gt;Which interfaces and namespaces were associated with it?&lt;/li&gt;
&lt;li&gt;Was its mark changed?&lt;/li&gt;
&lt;li&gt;Did a TC or XDP program handle it?&lt;/li&gt;
&lt;li&gt;Was it transformed or placed into a tunnel?&lt;/li&gt;
&lt;li&gt;Did the local kernel report a drop?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That evidence can clear or implicate the client host.&lt;/p&gt;

&lt;p&gt;It cannot prove what happened inside the ISP or on the target server. To prove&lt;br&gt;
those parts, we need observations from those vantage points. That limitation&lt;br&gt;
is important: the absence of a local drop is not proof that a remote host&lt;br&gt;
received the packet.&lt;/p&gt;

&lt;p&gt;The useful result is a smaller, evidence-backed search area:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;client host evidence
        ↓
observed departure boundary
        ↓
ISP / transit inference
        ↓
target-host evidence, if available
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;I would rather know exactly where my evidence ends than have a tool produce a&lt;br&gt;
confident story about systems it could not observe.&lt;/p&gt;
&lt;h2&gt;
  
  
  Use case: a packet disappears inside the host
&lt;/h2&gt;

&lt;p&gt;“The network dropped it” can hide several very different failures.&lt;/p&gt;

&lt;p&gt;A host may reject a packet through a normal kernel path. A TC classifier or&lt;br&gt;
XDP program may return a drop action. A packet may be consumed during teardown.&lt;br&gt;
A transformation may cause the packet we were following to continue under a&lt;br&gt;
different kernel object.&lt;/p&gt;

&lt;p&gt;For focused TCP drop questions, a small bpftrace program may be the fastest&lt;br&gt;
answer. For packet contents, I still reach for tcpdump or Wireshark.&lt;/p&gt;

&lt;p&gt;I reach for &lt;code&gt;skbx&lt;/code&gt; when I do not yet know which local subsystem owns the&lt;br&gt;
failure, or when I want to preserve the ordered path rather than one drop&lt;br&gt;
event.&lt;/p&gt;
&lt;h2&gt;
  
  
  Use case: follow transformations instead of losing the trail
&lt;/h2&gt;

&lt;p&gt;Linux does not promise that one logical packet corresponds to one&lt;br&gt;
&lt;code&gt;struct sk_buff&lt;/code&gt; address for its entire lifetime.&lt;/p&gt;

&lt;p&gt;Packets can be cloned, copied, or changed through copy-on-write. XDP frames can&lt;br&gt;
later become SKBs. Tunnels introduce outer and inner packet identities.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;skbx&lt;/code&gt; uses bounded, capture-local lineage state to connect those transitions.&lt;br&gt;
The provenance remains explicit: an event records whether it matched the&lt;br&gt;
original filter or was included because it belonged to a packet already being&lt;br&gt;
tracked.&lt;/p&gt;

&lt;p&gt;For example, a marked packet can be followed with:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;skbx capture &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--filter-mark&lt;/span&gt; 0x2a &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--filter-track-skb&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--output&lt;/span&gt; trace.jsonl
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The important word here is &lt;em&gt;capture-local&lt;/em&gt;. A lineage identifier is not a&lt;br&gt;
distributed trace ID, and it should not be presented as one.&lt;/p&gt;
&lt;h2&gt;
  
  
  Use case: capture once, investigate without root
&lt;/h2&gt;

&lt;p&gt;Loading and attaching eBPF programs is privileged work. Reading a JSONL file&lt;br&gt;
does not need to be.&lt;/p&gt;

&lt;p&gt;That makes a simple handoff possible:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;An operator checks the plan and performs a bounded capture.&lt;/li&gt;
&lt;li&gt;The capture ends and writes its reliability footer.&lt;/li&gt;
&lt;li&gt;The artifact is copied to a normal development or analysis environment.&lt;/li&gt;
&lt;li&gt;Another person replays it and cites exact &lt;code&gt;event:&lt;/code&gt; or &lt;code&gt;route:&lt;/code&gt; handles.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This is useful for incident handoffs and bug reports, but it is also useful for&lt;br&gt;
learning. I can capture a small experiment once and inspect it repeatedly&lt;br&gt;
without keeping probes attached while I try to understand the result.&lt;/p&gt;
&lt;h2&gt;
  
  
  Use case: give an AI evidence instead of a mystery
&lt;/h2&gt;

&lt;p&gt;One of my stranger motivations was making the CLI friendly to both operators&lt;br&gt;
and software agents.&lt;/p&gt;

&lt;p&gt;Before touching the host, an agent can ask the executable what it supports:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;skbx describe &lt;span class="nt"&gt;--format&lt;/span&gt; json
skbx schema
skbx doctor &lt;span class="nt"&gt;--json&lt;/span&gt;
skbx plan &lt;span class="nt"&gt;--json&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The native engine remains the source of truth. An AI system may summarize a&lt;br&gt;
route or explain a captured event, but the event must already exist in the&lt;br&gt;
trace. Machine output stays on standard output, diagnostics stay on standard&lt;br&gt;
error, and unsupported capabilities remain explicit.&lt;/p&gt;

&lt;p&gt;This does not make an AI explanation automatically correct. It gives the&lt;br&gt;
explanation something concrete to cite and gives a human a way to check it.&lt;/p&gt;
&lt;h2&gt;
  
  
  The engineering parts I enjoyed
&lt;/h2&gt;

&lt;p&gt;The obvious fun was seeing packets move through functions I had previously&lt;br&gt;
treated as boxes in a diagram.&lt;/p&gt;

&lt;p&gt;The less obvious fun was designing the boundaries:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;discovering valid &lt;code&gt;struct sk_buff *&lt;/code&gt; arguments from the target kernel's BTF
instead of guessing from function names;&lt;/li&gt;
&lt;li&gt;keeping kernel-side state and work bounded;&lt;/li&gt;
&lt;li&gt;moving JSON encoding, symbolization, and filesystem work out of the eBPF hot
path;&lt;/li&gt;
&lt;li&gt;making probe planning inspectable before attachment;&lt;/li&gt;
&lt;li&gt;separating raw observation from later explanation;&lt;/li&gt;
&lt;li&gt;treating partial output as incomplete instead of “probably good enough”;&lt;/li&gt;
&lt;li&gt;making replay deterministic enough to use in tests and automation.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those constraints made the tool more interesting to build. They also gave me a&lt;br&gt;
better mental model of what an observability tool can honestly claim.&lt;/p&gt;
&lt;h2&gt;
  
  
  What it does not know
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;skbx&lt;/code&gt; only observes what the running kernel and selected probes expose.&lt;/p&gt;

&lt;p&gt;It does not see inside an ISP. It does not see a remote host unless it runs&lt;br&gt;
there too. It is not a packet-content replacement for Wireshark. Kernel&lt;br&gt;
configuration, BTF availability, permissions, hidden addresses, unsupported&lt;br&gt;
signatures, and capture loss can all limit the evidence.&lt;/p&gt;

&lt;p&gt;Those limitations are part of the interface rather than footnotes.&lt;/p&gt;

&lt;p&gt;I am still exploring where this project is most useful. The best next input is&lt;br&gt;
not “looks cool,” although I will happily take that. It is a packet path the&lt;br&gt;
tool cannot explain yet, together with the kernel version, exact command,&lt;br&gt;
&lt;code&gt;doctor --json&lt;/code&gt; output, and reliability footer.&lt;/p&gt;

&lt;p&gt;The source, installation instructions, architecture notes, and field guides&lt;br&gt;
are in the repository:&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/copyleftdev" rel="noopener noreferrer"&gt;
        copyleftdev
      &lt;/a&gt; / &lt;a href="https://github.com/copyleftdev/skbx" rel="noopener noreferrer"&gt;
        skbx
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Agent-first Linux packet-path tracing with Rust/eBPF: bounded evidence, deterministic replay, and packet routes with receipts.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;
  &lt;a rel="noopener noreferrer" href="https://github.com/copyleftdev/skbx/assets/skbx-banner.svg"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fcopyleftdev%2Fskbx%2FHEAD%2Fassets%2Fskbx-banner.svg" alt="skbx — packet paths, with receipts" width="100%"&gt;&lt;/a&gt;
&lt;/p&gt;
&lt;p&gt;
  &lt;a href="https://github.com/copyleftdev/skbx/actions/workflows/ci.yml" rel="noopener noreferrer"&gt;&lt;img alt="CI" src="https://github.com/copyleftdev/skbx/actions/workflows/ci.yml/badge.svg"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/copyleftdev/skbx/LICENSE" rel="noopener noreferrer"&gt;&lt;img alt="License: AGPL-3.0-or-later" src="https://camo.githubusercontent.com/b2e30aa49de2b7e9ed97c80fef29aa69b83dfa9107727a161410eb683f8b5982/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f6c6963656e73652d4147504c2d2d332e302d2d6f722d2d6c617465722d633866363662"&gt;&lt;/a&gt;
  &lt;a rel="noopener noreferrer nofollow" href="https://camo.githubusercontent.com/254bab6b138d90ff96803f899a8402f6bcf45c4d90545dafdc1ae7b8b9fd029c/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f706c6174666f726d2d4c696e75782d356465346430"&gt;&lt;img alt="Linux" src="https://camo.githubusercontent.com/254bab6b138d90ff96803f899a8402f6bcf45c4d90545dafdc1ae7b8b9fd029c/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f706c6174666f726d2d4c696e75782d356465346430"&gt;&lt;/a&gt;
  &lt;a rel="noopener noreferrer nofollow" href="https://camo.githubusercontent.com/acab62d447ccb697fec83a4c79b65ed6c1a46df4a749c4fa9a6ba8196ddcc601/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f727573742d312e38352532422d396238636666"&gt;&lt;img alt="Rust 1.85+" src="https://camo.githubusercontent.com/acab62d447ccb697fec83a4c79b65ed6c1a46df4a749c4fa9a6ba8196ddcc601/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f727573742d312e38352532422d396238636666"&gt;&lt;/a&gt;
  &lt;a href="https://tokentip.to/@copyleftdev" rel="nofollow noopener noreferrer"&gt;&lt;img alt="Tip my tokens" src="https://camo.githubusercontent.com/f92620b71c800bad06aa383e6121303dcec9489bc3dc37885d99b2cae8cacb71/68747470733a2f2f746f6b656e7469702e746f2f62616467652f636f70796c6566746465762e7376673f6c6f676f3d31"&gt;&lt;/a&gt;
&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;skbx&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;&lt;code&gt;skbx&lt;/code&gt; shows where a packet went inside Linux—and keeps the receipts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://copyleftdev.github.io/skbx/" rel="nofollow noopener noreferrer"&gt;Follow a packet through the flight recorder →&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Field guides:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://copyleftdev.github.io/skbx/guides/trace-a-website-request.html" rel="nofollow noopener noreferrer"&gt;Trace a website request across your Linux host, ISP, and target&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://copyleftdev.github.io/skbx/guides/debug-linux-packet-drops.html" rel="nofollow noopener noreferrer"&gt;Debug a Linux packet drop with replayable eBPF evidence&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It observes kernel networking functions, TC/XDP programs, packet
transformations, tunnels, drops, and selected BPF helper activity with
Rust and CO-RE eBPF. Every observation lands in a bounded, replayable evidence
stream with stable handles and explicit loss telemetry.&lt;/p&gt;
&lt;p&gt;Use it when “the packet disappeared” is not a sufficient incident report.&lt;/p&gt;
&lt;div class="snippet-clipboard-content notranslate position-relative overflow-auto"&gt;&lt;pre class="notranslate"&gt;&lt;code&gt;capture → event:8c6f… → replay → route:21b4… → explain
            evidence          pattern          context
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;blockquote&gt;
&lt;p&gt;Inspired by &lt;a href="https://github.com/cilium/pwru" rel="noopener noreferrer"&gt;pwru&lt;/a&gt;, rebuilt around an
agent-first contract: deterministic observations, machine-readable
capabilities, bounded state, and no invented evidence.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;The short version&lt;/h2&gt;
&lt;/div&gt;
&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;You need to…&lt;/th&gt;
&lt;th&gt;skbx gives you…&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;See the kernel path of an SKB&lt;/td&gt;
&lt;td&gt;BTF-discovered kprobes with exact function evidence&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Follow clones, copies, COW, and XDP-to-SKB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;…&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/copyleftdev/skbx" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;If you try it, I would love to know which networking question you used it to&lt;br&gt;
answer—and where its evidence stopped.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;AI-assistance disclosure: I built and tested &lt;code&gt;skbx&lt;/code&gt;. I used OpenAI Codex to&lt;br&gt;
research adjacent writing, challenge the article's positioning, edit this&lt;br&gt;
draft, verify its commands and example output against the repository, and&lt;br&gt;
generate the cover illustration. I reviewed the technical claims and remain&lt;br&gt;
responsible for their accuracy.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>linux</category>
      <category>networking</category>
      <category>ai</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
