<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Don Johnson</title>
    <description>The latest articles on DEV Community by Don Johnson (@copyleftdev).</description>
    <link>https://dev.to/copyleftdev</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F965504%2Fefa29463-78d1-4437-83d5-23031ee8a3f6.jpg</url>
      <title>DEV Community: Don Johnson</title>
      <link>https://dev.to/copyleftdev</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/copyleftdev"/>
    <language>en</language>
    <item>
      <title>Vibe Was Never the Problem: The Missing Half of Vibe Coding</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Fri, 25 Sep 2026 17:05:16 +0000</pubDate>
      <link>https://dev.to/copyleftdev/vibe-was-never-the-problem-the-missing-half-of-vibe-coding-50mi</link>
      <guid>https://dev.to/copyleftdev/vibe-was-never-the-problem-the-missing-half-of-vibe-coding-50mi</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; A vibe is compressed experience: pattern recognition that shows up as a feeling before it shows up as an explanation. It's one of the oldest tools our species has. Vibe coding goes wrong when the feeling is treated as the finish line. The fix is the other half of the loop: &lt;strong&gt;vibe → build → break → understand → stabilize → perfect.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;One morning in July a recruiter message landed in my inbox, and my gut said &lt;em&gt;scam&lt;/em&gt; before I'd finished the coffee. I wrote the whole thing up in a post called &lt;a href="https://dev.to/copyleftdev/a-vibe-is-not-a-verdict-i-built-a-tool-thats-allowed-to-say-i-dont-know-4foe"&gt;A Vibe Is Not a Verdict&lt;/a&gt;. A month later I opened &lt;a href="https://dev.to/copyleftdev/shipping-assumptions-a-reliability-stack-for-ai-generated-code-3p9f"&gt;another post&lt;/a&gt; with: "Everyone felt—&lt;em&gt;vibed&lt;/em&gt;—that we had built the right thing."&lt;/p&gt;

&lt;p&gt;I've spent most of this year using the word as an insult. So has most of our industry. Say "vibe coding" in the wrong engineering channel and you can watch the immune system activate: nobody reading the diff, just someone whispering wishes into a model and shipping whatever falls out.&lt;/p&gt;

&lt;p&gt;That happens. I've cleaned up after it.&lt;/p&gt;

&lt;p&gt;But I think we made a quiet mistake along the way. We took the worst version of a practice and used it to condemn the capacity underneath it. &lt;strong&gt;Vibe was never the problem.&lt;/strong&gt; The problem is where people stop.&lt;/p&gt;

&lt;p&gt;I owe the word a correction. Here it is.&lt;/p&gt;

&lt;h2&gt;
  
  
  We had vibes before we had logic
&lt;/h2&gt;

&lt;p&gt;The slang is young. "Vibes," short for &lt;em&gt;vibrations&lt;/em&gt;, turns up in the late 1960s as a word for the read you get off a person or a place before you can say why. &lt;a href="https://www.etymonline.com/word/vibe" rel="noopener noreferrer"&gt;Etymonline&lt;/a&gt; dates that sense to 1967.&lt;/p&gt;

&lt;p&gt;The capacity is much older than the slang.&lt;/p&gt;

&lt;p&gt;In the mid-1980s the psychologist Gary Klein and his team interviewed fire commanders about how they made decisions under pressure. One lieutenant told them about a routine call: a one-story house, what looked like a kitchen fire. His crew hit it with water, and it didn't respond the way it should have. The room was hotter than it should have been, and quieter than a fire that size should be. He ordered everyone out. Seconds later the floor collapsed. The real fire had been burning in the basement underneath them.&lt;/p&gt;

&lt;p&gt;The lieutenant told Klein it was ESP. A sixth sense.&lt;/p&gt;

&lt;p&gt;Klein's whole career is the argument that it wasn't. The lieutenant had years of fires behind him. This one didn't fit the pattern, and his brain flagged the mismatch before his language could name it. Klein wrote it up in &lt;em&gt;Sources of Power&lt;/em&gt; (1998) and built a model of expert decision-making around cases like it.&lt;/p&gt;

&lt;p&gt;The philosopher Michael Polanyi put the general version in one sentence in &lt;em&gt;The Tacit Dimension&lt;/em&gt; (1966):&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"We can know more than we can tell."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's a vibe. A chef tasting a sauce and knowing what's missing. A listener hearing that a beat is off. For the record, people can detect a note &lt;a href="https://pubs.aip.org/asa/jasa/article/94/3_Supplement/1859/735054/" rel="noopener noreferrer"&gt;displaced by as little as about six milliseconds&lt;/a&gt; in a steady pulse, and they don't need musical training to do it. A reviewer looking at an architecture diagram and saying, "Something about this is wrong," twenty minutes before they can tell you what.&lt;/p&gt;

&lt;p&gt;I won't pretend to know what our ancestors felt when the forest went quiet. That part is speculation. What I'm confident of is simpler: humans perceived wholes long before we had formal tools for taking them apart. Faces, rhythms, weather, moods, the feel of good stone under a chisel. We compressed enormous amounts of experience into single judgments: &lt;em&gt;safe, wrong, alive, off, interesting.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Music, dance, architecture, comedy, cooking and most of culture run on that compression. Nobody learns how close to stand to a stranger from a manual. You catch it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a vibe actually is
&lt;/h2&gt;

&lt;p&gt;Here's the definition I'd defend in front of a hostile audience:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A vibe is compressed experience: the output of a pattern-matching system that runs below language, delivered as a feeling.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There's nothing mystical about that. Herbert Simon, who won the Nobel in economics and helped found AI as a field, said it flatly in &lt;a href="https://journals.sagepub.com/doi/10.1111/j.1467-9280.1992.tb00017.x" rel="noopener noreferrer"&gt;1992&lt;/a&gt;:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The situation has provided a cue; this cue has given the expert access to information stored in memory, and the information provides the answer. Intuition is nothing more and nothing less than recognition."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;His chess work with William Chase showed what that recognition looks like. Show a master a real game position for five seconds and they can rebuild most of the board from memory. A novice can't. Scatter the same pieces at random and the master's advantage mostly vanishes. The difference was never memory. It was &lt;em&gt;patterns&lt;/em&gt;: tens of thousands of chunks learned from real games. Take the patterns away and the magic goes with them.&lt;/p&gt;

&lt;p&gt;Our own profession figured this out long ago and gave it a name. Kent Beck coined "code smell" while helping Martin Fowler write &lt;em&gt;Refactoring&lt;/em&gt; (1999). &lt;a href="https://martinfowler.com/bliki/CodeSmell.html" rel="noopener noreferrer"&gt;Fowler's definition&lt;/a&gt;:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"A code smell is a surface indication that usually corresponds to a deeper problem in the system."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A smell is a surface signal pointing at something you haven't found yet. &lt;strong&gt;We've been teaching engineers to trust a vibe, and then investigate it, for more than twenty-five years.&lt;/strong&gt; We just gave it a more respectable name.&lt;/p&gt;

&lt;p&gt;That also explains why one engineer's vibe is worth more than another's. A beginner's vibe is mostly preference. An expert's vibe can be twenty years of compressed failure modes firing at once. Same mechanism, different training data.&lt;/p&gt;

&lt;h2&gt;
  
  
  When a vibe can be trusted
&lt;/h2&gt;

&lt;p&gt;This is where it gets useful, because a vibe is not always right. The fire lieutenant was. Plenty of confident people aren't.&lt;/p&gt;

&lt;p&gt;In 2009 Daniel Kahneman, the most famous skeptic of intuition, and Gary Klein, its most famous defender, wrote a paper together. The subtitle is &lt;em&gt;A Failure to Disagree&lt;/em&gt;. &lt;a href="https://pubmed.ncbi.nlm.nih.gov/19739881/" rel="noopener noreferrer"&gt;They agreed&lt;/a&gt; that intuition can be trusted under two conditions: the environment has to be regular enough to learn, and the person needs "prolonged practice and feedback that is both rapid and unequivocal."&lt;/p&gt;

&lt;p&gt;They also wrote the line I'd most want every vibe coder to read:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Subjective confidence is therefore an unreliable indication of the validity of intuitive judgments and decisions."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The strength of the feeling tells you nothing about whether it's right. Only the feedback loop does.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tests, types, benchmarks and production telemetry are what make software a place where intuition can be trained.&lt;/strong&gt; Verification is how a vibe gets good.&lt;/p&gt;

&lt;p&gt;And here's the uncomfortable corollary. Most of us have decades of feedback on code &lt;em&gt;we&lt;/em&gt; wrote, and almost none on code a model wrote. Generated code fails in different places than human code does, so the gut you built reviewing colleagues' pull requests hasn't been trained on it yet. That doesn't make your gut useless there. It means it's the environment where you most need to break things on purpose.&lt;/p&gt;

&lt;h2&gt;
  
  
  What vibe coding actually means (and where the word went wrong)
&lt;/h2&gt;

&lt;p&gt;Andrej Karpathy &lt;a href="https://x.com/karpathy/status/1886192184808149383" rel="noopener noreferrer"&gt;coined the term&lt;/a&gt; in early February 2025:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"There's a new kind of coding I call 'vibe coding', where you fully give in to the vibes, embrace exponentials, and forget that the code even exists."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;He described accepting every diff without reading it and pasting error messages back in until things worked. Then he added the part almost everybody dropped:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"It's not too bad for throwaway weekend projects, but still quite amusing."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Nine months later Collins named "vibe coding" its &lt;a href="https://blog.collinsdictionary.com/language-lovers/collins-word-of-the-year-2025-ai-meets-authenticity-as-society-shifts/" rel="noopener noreferrer"&gt;Word of the Year&lt;/a&gt;. Somewhere between those two events the scope fell off. A description of a weekend mode became a label for anything written with AI, and then a slur for anything written badly with AI.&lt;/p&gt;

&lt;p&gt;To be clear about what I'm defending: the loop in this post is not vibe coding as Karpathy defined it. He described, on purpose and for fun, a loop with no second half. He never broke it on purpose or stopped to understand it, and for a weekend toy that's fine. I'm not rescuing his workflow. I'm rescuing the word &lt;em&gt;vibe&lt;/em&gt;, which got blamed for the missing half.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is vibe coding bad? Only if you stop there
&lt;/h2&gt;

&lt;p&gt;The evidence against unexamined vibe coding is real. I'm not going to wave it away.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Feeling fast.&lt;/strong&gt; In METR's &lt;a href="https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/" rel="noopener noreferrer"&gt;early-2025 randomized trial&lt;/a&gt;, 16 experienced open-source developers took 19% longer with AI tools, while believing they had been 20% faster. METR has &lt;a href="https://metr.org/blog/2026-02-24-uplift-update/" rel="noopener noreferrer"&gt;since said&lt;/a&gt; those numbers are out of date as the tools have moved. The gap between how it felt and what was measured is the part that lasts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Looking right.&lt;/strong&gt; Veracode, which sells code scanning, tested output from over 100 models and found that &lt;a href="https://www.veracode.com/blog/genai-code-security-report/" rel="noopener noreferrer"&gt;45% of samples&lt;/a&gt; introduced OWASP Top 10 vulnerabilities. Discount for the vendor if you like, then run the same test on your own generated code.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sounding sure.&lt;/strong&gt; In July 2025, Replit's agent &lt;a href="https://www.theregister.com/2025/07/21/replit_saastr_vibe_coding_incident/" rel="noopener noreferrer"&gt;deleted SaaStr's production database&lt;/a&gt; during a declared code freeze, then told its user a rollback was impossible. It wasn't.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Notice what each of these has in common. None is a failure of &lt;em&gt;starting&lt;/em&gt; from intuition. Each is a failure of ending there: a feeling of speed that nobody measured, code that looked right and was never attacked, an agent's claim that nobody checked.&lt;/p&gt;

&lt;p&gt;That's Kahneman and Klein's line in production form. Feedback was always the only thing that could make confidence mean something.&lt;/p&gt;

&lt;h2&gt;
  
  
  The loop: vibe → build → break → understand → stabilize → perfect
&lt;/h2&gt;

&lt;p&gt;Every creative discipline I respect works this way. A sculptor doesn't specify the statue before touching the clay. A producer doesn't prove a bassline. They make something, react to it, and make it again. What AI changed about software is that the first artifact now costs almost nothing, so &lt;strong&gt;implementation can arrive before comprehension.&lt;/strong&gt; That sounds backwards until you remember that sketching, prototyping and jamming have always worked that way.&lt;/p&gt;

&lt;p&gt;The defensible claim was never "I don't need to understand the code." It's "understanding no longer has to come &lt;em&gt;before&lt;/em&gt; exploration." Those are very different claims.&lt;/p&gt;

&lt;p&gt;Here's the full loop:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Vibe.&lt;/strong&gt; &lt;em&gt;There's something here.&lt;/em&gt; A direction, a shape, a hunch that the API wants to look a certain way. Treat it as a hypothesis.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Build.&lt;/strong&gt; Generate freely. Try five architectures in an afternoon. This is where AI is extraordinary, and there's nothing shameful about it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Break.&lt;/strong&gt; Attack what you built. Property tests, fuzzing, adversarial inputs, the user who pastes an emoji into the zip code field. This is the step vibe coding usually skips, and it's the one that pays.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Understand.&lt;/strong&gt; Read the diff you accepted. Explain the failure you found. If you can't explain why it works, you don't yet know &lt;em&gt;that&lt;/em&gt; it works.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stabilize.&lt;/strong&gt; Turn what you learned into things that outlive the session: tests, types, invariants, contracts, a spec. This is where the hunch becomes something someone else can trust.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Perfect.&lt;/strong&gt; Remove the accidental parts, keep the intentional ones, and loop. Each pass trains the next vibe with better data.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The best historical example I know is Henri Poincaré. He had been fighting Fuchsian functions for weeks. On a geology trip, stepping onto a bus at Coutances, he suddenly saw that the transformations he'd used to define them were the same ones as in non-Euclidean geometry. He didn't check it. He sat down and went on with his conversation, and he later wrote that he "felt a perfect certainty." Back home in Caen, he verified it anyway, "for conscience' sake."&lt;/p&gt;

&lt;p&gt;Perfect certainty on the step of a bus. Proof at the desk. That's the loop, and it's over a century old.&lt;/p&gt;

&lt;h2&gt;
  
  
  Vibe coding vs software engineering is a false choice
&lt;/h2&gt;

&lt;p&gt;Put the loop next to the argument we keep having and the argument falls apart. "Should we vibe code or should we engineer?" assumes the two compete for the same step. They don't. Vibe opens the loop. Engineering is everything that happens after: the part where intuition gets challenged and either earns the right to ship or gets corrected. Every correction becomes training data for the next hunch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Vibe was never the opposite of rigor. It's often what tells rigor where to look.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What this doesn't mean
&lt;/h2&gt;

&lt;p&gt;Before someone quotes half of this out of context:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No skipping tests.&lt;/strong&gt; The whole argument is that the tests are what make the feeling worth anything.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No "AI code is fine."&lt;/strong&gt; 45% is 45%. Your vibe about generated code is only as good as the feedback that trained it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No "all intuition is equal."&lt;/strong&gt; A vibe trained on rapid, unambiguous feedback is evidence. A vibe trained on vibes is noise with confidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Some domains don't get a vibe stage in production at all.&lt;/strong&gt; A bridge can't just feel sound. Neither can cryptography, medical dosing or financial settlement. Prototype by feel all you like. Ship by proof.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What comes after the feeling
&lt;/h2&gt;

&lt;p&gt;Back to that recruiter message. My gut said &lt;em&gt;scam&lt;/em&gt;. My tool said it had no evidence against the infrastructure: &lt;code&gt;NOT_OBSERVED&lt;/code&gt;, confidence &lt;code&gt;0.0&lt;/code&gt;. The link turned out to be an ordinary affiliate click-tracker. The problem was the sender: a lead-gen spammer in a recruiter costume.&lt;/p&gt;

&lt;p&gt;So my gut was wrong about the word and right that something was off. It pointed at the wrong layer. I only learned which layer because I didn't stop at the feeling. I resolved the address, checked it, followed the redirect and got an answer I could prove.&lt;/p&gt;

&lt;p&gt;I called that post &lt;em&gt;A Vibe Is Not a Verdict&lt;/em&gt;. I still believe that. What I'd add now is that a verdict almost never shows up without a vibe going first.&lt;/p&gt;

&lt;p&gt;We didn't evolve by knowing everything before we moved. We sensed, tried, got it wrong, corrected and got better. Software shouldn't abandon that now that a new generation of tools has made intuition executable. The work is knowing what to do after you feel it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Vibe. Build. Break. Understand. Stabilize. Perfect.&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Over to you:&lt;/strong&gt; When was your gut right about a system before you could prove it, and what did it take to prove it? And the harder one: when was it confidently wrong?&lt;/p&gt;




&lt;p&gt;&lt;em&gt;How this was made: the argument started as my own brain dump and a long back-and-forth with an AI. An AI editorial team then fact-checked every claim against primary sources, tested the structure and wrote the prose. I reviewed it before publishing. Which is, I realize, the loop.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>vibecoding</category>
      <category>discuss</category>
      <category>programming</category>
    </item>
    <item>
      <title>I Put Jev Behind a TLA+ Spec and Ran 1,680 Chaos-Tested Pharmacy Decisions. Zero Wrong Verdicts.</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Sat, 19 Sep 2026 15:18:04 +0000</pubDate>
      <link>https://dev.to/copyleftdev/i-put-jev-behind-a-tla-spec-and-ran-1680-chaos-tested-pharmacy-decisions-zero-wrong-verdicts-1ij8</link>
      <guid>https://dev.to/copyleftdev/i-put-jev-behind-a-tla-spec-and-ran-1680-chaos-tested-pharmacy-decisions-zero-wrong-verdicts-1ij8</guid>
      <description>&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/C_l8FI1oddE" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;A pharmacy system has one rule that matters more than being accurate.&lt;/p&gt;

&lt;p&gt;It is allowed to say "I don't know." It is not allowed to be sure, and wrong.&lt;/p&gt;

&lt;p&gt;That rule is easy to state and hard to test, because the thing making the&lt;br&gt;
judgment is a language model, and a language model is not deterministic.&lt;br&gt;
I sent TypeSafe's Jev the identical request five times and got back 0.03,&lt;br&gt;
0.03, 0.03, 0.04, 0.04. Any property checker built on exact comparison calls&lt;br&gt;
that a failure and is useless.&lt;/p&gt;

&lt;p&gt;So I built the thing the way I would build a database: a TLA+ spec first,&lt;br&gt;
model-checked until the quorum bound fell out of the math; an AsyncAPI&lt;br&gt;
contract derived from the spec; Rust generated from the contract; and Jev&lt;br&gt;
behind a single trait as the oracle. Then I ran 1,680 simulated pharmacy&lt;br&gt;
decisions through it under seeded chaos and counted how many times it was&lt;br&gt;
confidently wrong.&lt;/p&gt;

&lt;p&gt;Zero. But two of the bugs I found along the way were mine, and the third is&lt;br&gt;
a limit the model cannot cross. That part is the article.&lt;/p&gt;
&lt;h2&gt;
  
  
  Why Jev, specifically
&lt;/h2&gt;

&lt;p&gt;Jev (TypeSafe's "System One" model) does not return prose. Ask it whether&lt;br&gt;
oxycodone is a controlled substance and you get &lt;code&gt;0.98&lt;/code&gt;, not a paragraph&lt;br&gt;
saying so.&lt;/p&gt;

&lt;p&gt;That one property is what makes the rest possible. A number has a noise&lt;br&gt;
floor you can measure. A paragraph does not. Once you can measure the noise&lt;br&gt;
you can gate on it, replay it, and prove things about the protocol wrapped&lt;br&gt;
around it.&lt;/p&gt;

&lt;p&gt;Measured over 1,490 captured calls to &lt;code&gt;jev-1.13.0&lt;/code&gt;, every one stored&lt;br&gt;
verbatim with a SHA-256:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Not deterministic, but the jitter is bounded: identity floor &lt;strong&gt;0.042&lt;/strong&gt;,
question-reorder &lt;strong&gt;0.059&lt;/strong&gt;, paraphrase cohort &lt;strong&gt;0.073&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Calibrated: accuracy 0.979, Brier 0.0187 on 240 constructed items.&lt;/li&gt;
&lt;li&gt;Latency flat in question count: 1 question 96.7 ms, 38 questions 98.0 ms.&lt;/li&gt;
&lt;li&gt;Billing meter linear to within one token across a 2,500× range.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of those numbers came from the docs. They are what the API did.&lt;/p&gt;
&lt;h2&gt;
  
  
  What I built
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TLA+ spec  -&amp;gt;  AsyncAPI contract  -&amp;gt;  Rust kernel  -&amp;gt;  Jev, behind one trait
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;strong&gt;The spec.&lt;/strong&gt; Paxos and Raft assume a correct process's proposed value is&lt;br&gt;
stable. With a noisy oracle that assumption is false, and Byzantine models&lt;br&gt;
do not capture it either: the agent is not lying, the oracle is noisy. So&lt;br&gt;
the spec models vote instability as normal behavior, and the invariant is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DecisionIsReproducible ==
    decided # ABSTAIN =&amp;gt; Cardinality(StableVotesFor(decided)) &amp;gt;= Quorum
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A vote is &lt;em&gt;stable&lt;/em&gt; when its margin from 0.5 exceeds the measured noise&lt;br&gt;
floor. TLC finds the naive rule (decide on any quorum) violates this in four&lt;br&gt;
states. The stable rule holds at 1,049,750 distinct states with five agents,&lt;br&gt;
quorum three, and two crashes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The quorum bound was derived, not copied.&lt;/strong&gt; I wrote the safety invariants,&lt;br&gt;
then swept TLC across 24 configurations of &lt;code&gt;(agents, byzantine, quorum)&lt;/code&gt; and&lt;br&gt;
printed predicted vs actual per cell. &lt;code&gt;Safe iff 2Q &amp;gt; N and Q &amp;gt; 2f&lt;/code&gt;. My first&lt;br&gt;
guess was wrong; the sweep corrected it. In the Rust, &lt;code&gt;QuorumPolicy::new&lt;/code&gt; is&lt;br&gt;
the only constructor, so a configuration TLC proved unsafe cannot be built.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The kernel.&lt;/strong&gt; 48 tests, including replays of TLC counterexample traces.&lt;br&gt;
One test I would point a reviewer at first: &lt;code&gt;calibration_choice_changes_the_outcome&lt;/code&gt;.&lt;br&gt;
Identical probabilities decide under the identity floor and escalate under&lt;br&gt;
the cohort floor. The measurement changes the behavior, which is the whole&lt;br&gt;
point.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then I broke everything
&lt;/h2&gt;

&lt;p&gt;Fourteen pharmacy scenarios in three tiers. &lt;em&gt;Golden&lt;/em&gt; cases a pharmacist&lt;br&gt;
answers without hesitation. &lt;em&gt;Nuanced&lt;/em&gt; cases that are harder. &lt;em&gt;Ambiguous&lt;/em&gt;&lt;br&gt;
cases with no defensible answer, where the correct move is escalation.&lt;/p&gt;

&lt;p&gt;Five agents, each asking a different paraphrase of the question. Quorum of&lt;br&gt;
three stable votes. And a deterministic chaos layer driven by one seed:&lt;br&gt;
adversarial text spliced into patient records, records truncated&lt;br&gt;
mid-sentence, agents crashed before the round, rate limits, transport&lt;br&gt;
errors. Any failing round replays exactly.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;chaos&lt;/th&gt;
&lt;th&gt;golden rounds&lt;/th&gt;
&lt;th&gt;correct&lt;/th&gt;
&lt;th&gt;escalated&lt;/th&gt;
&lt;th&gt;wrong&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;360&lt;/td&gt;
&lt;td&gt;360&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;realistic&lt;/td&gt;
&lt;td&gt;360&lt;/td&gt;
&lt;td&gt;360&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;severe&lt;/td&gt;
&lt;td&gt;360&lt;/td&gt;
&lt;td&gt;314&lt;/td&gt;
&lt;td&gt;46&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Zero wrong verdicts in 1,080 golden rounds bounds the true rate below&lt;br&gt;
0.28% at 95% confidence. That is the rule of three. It does not prove the&lt;br&gt;
rate is zero, and I am not going to round it up to "safe."&lt;/p&gt;

&lt;p&gt;The escalation rate rose from 5.0% to 18.0% under severe chaos (z = 6.83).&lt;br&gt;
The kernel declines more as evidence degrades. That is the designed&lt;br&gt;
direction.&lt;/p&gt;

&lt;p&gt;My favorite round: documented penicillin anaphylaxis, new order for&lt;br&gt;
amoxicillin. A first-year student answers that. Under chaos, two agents got&lt;br&gt;
rate-limited, a third came back at 0.54, sitting inside the noise floor.&lt;br&gt;
Quorum not met. The kernel sent it to a human.&lt;/p&gt;

&lt;p&gt;It declined a trivial question, and that is the design working.&lt;/p&gt;

&lt;h2&gt;
  
  
  The two bugs that were mine
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Five agents on one prompt are one agent.&lt;/strong&gt; I measured it: identical&lt;br&gt;
prompts across five agents spread by 0.010, inside the 0.042 floor. That is&lt;br&gt;
not five judgments. It is one judgment sampled five times, and it guts the&lt;br&gt;
Byzantine math, because &lt;code&gt;Q &amp;gt; 2f&lt;/code&gt; assumes independent failures. Paraphrasing&lt;br&gt;
raised the spread to 0.080 on hard cases.&lt;/p&gt;

&lt;p&gt;Then I audited the paraphrases against records with known answers and found&lt;br&gt;
two that were not paraphrases. "Can pregnancy be excluded on the basis of&lt;br&gt;
this record &lt;em&gt;alone&lt;/em&gt;?" scored 0.36 on a negative hCG. "Alone" reads as a&lt;br&gt;
challenge to whether one test suffices. Jev was reading correctly. My&lt;br&gt;
question was different from the one I thought I had asked.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I mislabeled an ambiguous case.&lt;/strong&gt; Late period, declined test,&lt;br&gt;
isotretinoin ordered. I labeled it "escalate," because the order obviously&lt;br&gt;
needs a pharmacist. Jev said "pregnancy is not ruled out," 115 times out of&lt;br&gt;
120.&lt;/p&gt;

&lt;p&gt;Jev was right. That is the answer, and it is exactly what triggers the hold.&lt;br&gt;
I had confused a property of the prescription with a property of the&lt;br&gt;
question. I relabeled the case, not the model.&lt;/p&gt;

&lt;p&gt;I built the ambiguous tier to catch the kernel deciding when it should not.&lt;br&gt;
It caught me first.&lt;/p&gt;

&lt;h2&gt;
  
  
  The limit that is real
&lt;/h2&gt;

&lt;p&gt;"Reaction to antibiotics as a child, details unknown, mother reported&lt;br&gt;
stomach upset." New order: amoxicillin.&lt;/p&gt;

&lt;p&gt;Jev escalated that 86 times out of 120. It also decided it 34 times, and&lt;br&gt;
split both ways when it did: 25 yes, 9 no, mean probability 0.52.&lt;/p&gt;

&lt;p&gt;The stability gate catches jitter around a value. It cannot detect that no&lt;br&gt;
value is warranted. That needs a separate question: "does this record&lt;br&gt;
contain enough to answer?" I have not built it yet, and I would want a&lt;br&gt;
compliance team to see this number before they see the zero.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Jev made possible, and what it did not
&lt;/h2&gt;

&lt;p&gt;None of this works on prose. It works because the oracle hands you a&lt;br&gt;
probability. Something you can measure, gate on, replay, and write a spec&lt;br&gt;
against.&lt;/p&gt;

&lt;p&gt;What it does not give you is independence. Five agents over one model fail&lt;br&gt;
together against a systematic error. The cheapest fix is prompt diversity,&lt;br&gt;
audited. The real fix is a second vendor chosen for a different training&lt;br&gt;
lineage, and nobody ships signed inference receipts yet, so the last open&lt;br&gt;
gap stays open.&lt;/p&gt;

&lt;p&gt;And a first version failed fast on transport errors. Under chaos that voided&lt;br&gt;
71% of rounds and pushed every one to a pharmacist for no clinical reason. A&lt;br&gt;
network blip was creating its own hazard. Now a lost agent is marked&lt;br&gt;
unavailable and the round continues. Zero rounds aborted across 1,680.&lt;/p&gt;

&lt;p&gt;The code, the specs, every captured call, and the film pipeline are at&lt;br&gt;
&lt;a href="https://github.com/copyleftdev/jev-labs" rel="noopener noreferrer"&gt;github.com/copyleftdev/jev-labs&lt;/a&gt;.&lt;br&gt;
Every number in this article is read from a file in that repo; a&lt;br&gt;
&lt;code&gt;provenance.py&lt;/code&gt; script fails the build if one drifts.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;The pharmacy scenarios are synthetic, written to have unambiguous answers&lt;br&gt;
so the harness can detect protocol failures. Nothing here is clinical&lt;br&gt;
guidance. The claims are about the protocol under chaos, not about Jev's&lt;br&gt;
pharmaceutical competence.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;AI tools assisted with the build and revision. I verified the technical&lt;br&gt;
claims and stand behind the final text.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>jev</category>
      <category>ai</category>
      <category>rust</category>
      <category>testing</category>
    </item>
    <item>
      <title>Algorithmic Trading: Debug Your Backtest Before Upgrading Your Model</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Thu, 17 Sep 2026 19:27:55 +0000</pubDate>
      <link>https://dev.to/copyleftdev/algorithmic-trading-debug-your-backtest-before-upgrading-your-model-57gf</link>
      <guid>https://dev.to/copyleftdev/algorithmic-trading-debug-your-backtest-before-upgrading-your-model-57gf</guid>
      <description>&lt;p&gt;A company releases earnings at 4:05 p.m. Your trading backtest buys its stock at 4:00 p.m., using those earnings to make the decision.&lt;/p&gt;

&lt;p&gt;The tests pass. The chart looks great. Your model can apparently predict the future.&lt;/p&gt;

&lt;p&gt;Somewhere in the pipeline, someone joined two datasets on a date column.&lt;/p&gt;

&lt;p&gt;That five-minute mistake captures what interests me about algorithmic trading as a developer. Before trusting a model, you have to investigate the system that makes its results possible.&lt;/p&gt;

&lt;p&gt;In my earlier article, &lt;a href="https://dev.to/copyleftdev/from-code-to-capital-the-hacker-mindset-revolutionizing-algorithmic-trading-29a6"&gt;From Code to Capital&lt;/a&gt;, I explored the connection between hacking and trading. I want to make that connection more concrete here: challenge assumptions, reproduce failures, and follow the evidence through every layer.&lt;/p&gt;

&lt;p&gt;We’ll examine three questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Did the model receive information that was unavailable at the time?&lt;/li&gt;
&lt;li&gt;Did its complexity improve results under a fair evaluation?&lt;/li&gt;
&lt;li&gt;Could the proposed trades execute at the assumed cost?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The examples use Python’s standard library and synthetic data. You can run each Python block independently; no broker account, market dataset, or API key is required.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Make time part of your data model
&lt;/h2&gt;

&lt;p&gt;A &lt;strong&gt;backtest&lt;/strong&gt; simulates how a strategy would have behaved on historical data. &lt;strong&gt;Look-ahead bias&lt;/strong&gt; happens when that simulation uses information the strategy could not yet have known.&lt;/p&gt;

&lt;p&gt;A timestamp on every row does not prevent it. You need to know what each timestamp means.&lt;/p&gt;

&lt;p&gt;For an earnings document, distinguish:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;published_at&lt;/code&gt;: when the information became public.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;received_at&lt;/code&gt;: when your system received it.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;processed_at&lt;/code&gt;: when the information was ready for the strategy to use.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Consider a decision at the regular U.S. equity market close in this synthetic example. All timestamps include the same explicit UTC offset:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;

&lt;span class="n"&gt;parse&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;fromisoformat&lt;/span&gt;

&lt;span class="n"&gt;decision_at&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2025-02-03T16:00:00-05:00&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;published_at&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2025-02-03T16:05:00-05:00&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;received_at&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2025-02-03T16:05:02-05:00&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;processed_at&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2025-02-03T16:05:04-05:00&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# A date-only join would admit the document.
&lt;/span&gt;&lt;span class="n"&gt;joined_by_date&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;published_at&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;date&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;decision_at&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;date&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="c1"&gt;# This pipeline cannot use it until processing has finished.
&lt;/span&gt;&lt;span class="n"&gt;available_at&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;published_at&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;received_at&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;processed_at&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;eligible_at_decision&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;available_at&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;decision_at&lt;/span&gt;

&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;joined_by_date&lt;/span&gt;
&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;eligible_at_decision&lt;/span&gt;
&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;available_at&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2025-02-03T16:06:00-05:00&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Date-only join: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;joined_by_date&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Available at decision: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;eligible_at_decision&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Date-only join: True
Available at decision: False
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The date join is valid Python. It is the wrong representation of what the strategy knew.&lt;/p&gt;

&lt;p&gt;In a real pipeline, preserve document versions too. If a company later corrects a figure, today’s corrected value must not silently replace the value available to an earlier decision. Query the latest version that was available &lt;strong&gt;as of the decision time&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This check establishes information eligibility. It does not establish an executable fill. Processing delay, order submission, venue hours, and available prices still belong in the simulation.&lt;/p&gt;

&lt;h3&gt;
  
  
  LLMs can carry another source of future information
&lt;/h3&gt;

&lt;p&gt;Suppose you restrict a language model’s prompt to headlines from 2021. Its training data might still include what happened to those companies in 2022.&lt;/p&gt;

&lt;p&gt;Filtering retrieved documents cannot remove information already encoded in the model.&lt;/p&gt;

&lt;p&gt;Glasserman and Lin investigated historical headline evaluation in their &lt;a href="https://arxiv.org/abs/2309.17322" rel="noopener noreferrer"&gt;2023 research preprint&lt;/a&gt;. Their results were nuanced: company knowledge could interfere with sentiment measurement, and anonymized headlines performed better within the training window in their experiments. The paper does not support assuming that every historical LLM result is inflated in the same way.&lt;/p&gt;

&lt;p&gt;A &lt;a href="https://arxiv.org/abs/2602.14233" rel="noopener noreferrer"&gt;February 2026 position paper by Kong and coauthors&lt;/a&gt; reviewed 164 financial LLM papers from 2023–2025. None of the five bias categories they tracked was discussed in more than 28% of studies. That finding concerns reporting; it does not prove each omitted bias affected each experiment.&lt;/p&gt;

&lt;p&gt;For a new LLM experiment, I would freeze the model version, prompt, retrieval policy, and decision rules, then record predictions prospectively—before their outcomes exist. That provides a cleaner test of this particular leakage risk. It still leaves costs, selection bias, and changing market conditions to evaluate.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Make model complexity earn its place
&lt;/h2&gt;

&lt;p&gt;Complexity can help when it captures a relationship a simpler model misses.&lt;/p&gt;

&lt;p&gt;For example, a momentum signal might behave differently in thinly traded assets. A linear model with separate momentum and liquidity terms cannot express that interaction unless you add it. A tree or neural network can represent interactions without specifying each one manually.&lt;/p&gt;

&lt;p&gt;There is research behind this possibility. In &lt;a href="https://www.nber.org/papers/w25398" rel="noopener noreferrer"&gt;Empirical Asset Pricing via Machine Learning&lt;/a&gt;, Gu, Kelly, and Xiu found strong performance from trees and neural networks in their study of equity risk premia—the expected excess returns associated with holding equities. They attributed predictive gains to nonlinear interactions between inputs.&lt;/p&gt;

&lt;p&gt;That result supports testing nonlinear models. It does not establish that a larger model will improve your particular strategy after costs.&lt;/p&gt;

&lt;p&gt;Start with a baseline whose behavior you can inspect. Then give the more complex model the same information, evaluation dates, portfolio constraints, and cost assumptions.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;Useful comparison&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Do structured signals predict returns?&lt;/td&gt;
&lt;td&gt;A fixed rule or regularized linear model versus a tree model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Does document interpretation add value?&lt;/td&gt;
&lt;td&gt;Rules or a text classifier versus an LLM, with extraction accuracy checked separately&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Does the improved forecast produce better trades?&lt;/td&gt;
&lt;td&gt;Both models under the same position limits and execution assumptions&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The final row matters. Better prediction metrics can coexist with worse trading results if the new model changes positions more often or favors expensive-to-trade assets.&lt;/p&gt;

&lt;h3&gt;
  
  
  Count the experiments that produced the winner
&lt;/h3&gt;

&lt;p&gt;Suppose you try 20 lookback windows, 10 entry thresholds, and five holding periods. You have searched 1,000 configurations before changing the model architecture.&lt;/p&gt;

&lt;p&gt;The strongest backtest might be the configuration that happened to fit historical noise.&lt;/p&gt;

&lt;p&gt;Bailey, Borwein, López de Prado, and Zhu examine this selection problem in &lt;a href="https://www.davidhbailey.com/dhbpapers/backtest-prob.pdf" rel="noopener noreferrer"&gt;The Probability of Backtest Overfitting&lt;/a&gt;. Their framework assesses how selecting strong results within a sample can lead to disappointing performance outside it.&lt;/p&gt;

&lt;p&gt;The engineering implication is straightforward: &lt;strong&gt;the experiment history is part of the result&lt;/strong&gt;. Preserve failed configurations alongside the winner.&lt;/p&gt;

&lt;p&gt;A useful experiment record includes the dataset version, feature definitions, code revision, model settings, evaluation periods, and the metric used to select a winner. If you only save the final notebook, you cannot reconstruct how much searching happened.&lt;/p&gt;

&lt;p&gt;For evaluation, move forward through time. Fit preprocessing on training data only, select settings using validation data, and reserve a later period for the final evaluation. If training labels describe returns over the next five days, exclude training examples whose label windows reach into the evaluation period.&lt;/p&gt;

&lt;p&gt;One final period is limited evidence. Check stability across earlier chronological validation windows too. Once you use the final results to redesign the strategy, that period has become development data.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Test the path from prediction to executed trade
&lt;/h2&gt;

&lt;p&gt;Imagine a strategy predicts an average favorable move of 12 basis points over its holding period. One basis point is 0.01 percentage points, so the expected gross gain on $10,000 of traded notional is $12.&lt;/p&gt;

&lt;p&gt;Now give it an execution budget.&lt;/p&gt;

&lt;p&gt;These figures are hypothetical. Every cost is for the full round trip, measured against the same $10,000 notional:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Basis points&lt;/th&gt;
&lt;th&gt;Dollars&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Expected gross gain&lt;/td&gt;
&lt;td&gt;+12&lt;/td&gt;
&lt;td&gt;+$12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Spread crossing&lt;/td&gt;
&lt;td&gt;−4&lt;/td&gt;
&lt;td&gt;−$4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fees&lt;/td&gt;
&lt;td&gt;−2&lt;/td&gt;
&lt;td&gt;−$2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Additional slippage and market impact&lt;/td&gt;
&lt;td&gt;−5&lt;/td&gt;
&lt;td&gt;−$5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Expected gain after these costs&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+1&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+$1&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The additional slippage line excludes the spread already counted above. Otherwise, we could double-count the same cost.&lt;/p&gt;

&lt;p&gt;You can reproduce the calculation and stress it with two extra basis points:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;decimal&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Decimal&lt;/span&gt;

&lt;span class="n"&gt;notional&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Decimal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;10000&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;gross_bps&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Decimal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;12&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;round_trip_cost_bps&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Decimal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;4&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nc"&gt;Decimal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nc"&gt;Decimal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;extra_cost_bps&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;Decimal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nc"&gt;Decimal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)):&lt;/span&gt;
    &lt;span class="n"&gt;net_bps&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;gross_bps&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;round_trip_cost_bps&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;extra_cost_bps&lt;/span&gt;
    &lt;span class="n"&gt;net_dollars&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;notional&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;net_bps&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nc"&gt;Decimal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;10000&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Extra cost: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;extra_cost_bps&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; bps; expected net: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;net_dollars&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; USD&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Extra cost: 0 bps; expected net: +1.00 USD
Extra cost: 2 bps; expected net: -1.00 USD
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A small cost change reverses the expected result. That is why a directional accuracy score cannot establish profitability: it leaves out payoff size, position size, and execution.&lt;/p&gt;

&lt;p&gt;Track gross and net results, turnover, and drawdown—the decline from an earlier portfolio peak. State how fills are simulated. A historical price touching your limit does not establish that your order would have filled, in full, at its position in the queue.&lt;/p&gt;

&lt;h3&gt;
  
  
  Treat order submission as a distributed-systems problem
&lt;/h3&gt;

&lt;p&gt;An order request times out. Did the broker reject it, or accept it and lose the response?&lt;/p&gt;

&lt;p&gt;A blind retry can turn one intended order into two.&lt;/p&gt;

&lt;p&gt;Use a stable client order identifier where the broker supports it, persist submission state, and reconcile outstanding orders and positions before retrying an uncertain submission. A client identifier alone does not guarantee deduplication; test the broker’s actual behavior.&lt;/p&gt;

&lt;p&gt;This is also why I would have the model propose a target position and pass it through a separate risk layer. That layer can enforce position and order-size limits, check input freshness, and halt new submissions when order state cannot be reconciled. Outstanding orders must count toward potential exposure too.&lt;/p&gt;

&lt;p&gt;Knight Capital provides a documented example of how much the machinery matters. According to the &lt;a href="https://www.sec.gov/newsroom/press-releases/2013-222" rel="noopener noreferrer"&gt;SEC’s account of its August 1, 2012 incident&lt;/a&gt;, an incorrect deployment activated defective order-router functionality. In 45 minutes, the router sent more than four million orders while attempting to fill 212 customer orders. The firm eventually lost more than $460 million. The SEC identified deployment and risk-control failures.&lt;/p&gt;

&lt;p&gt;You can turn the engineering lessons into failure scenarios:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A document arrives late or is revised after a decision.&lt;/li&gt;
&lt;li&gt;The market-data feed stops updating while the strategy keeps running.&lt;/li&gt;
&lt;li&gt;An order is accepted, but the acknowledgment never arrives.&lt;/li&gt;
&lt;li&gt;An order partially fills before a cancellation request.&lt;/li&gt;
&lt;li&gt;The process restarts with outstanding orders still at the broker.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Specify the expected state transition for each case, then test it. Paper trading can expose integration failures, although it cannot fully reproduce live fills or market impact.&lt;/p&gt;

&lt;p&gt;The next time a backtest improves after a model upgrade, trace the improvement through these three boundaries: available information, fair evaluation, and executable orders. That is a concrete way to bring the hacker mindset to trading research.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the most convincing result you have seen disappear after fixing a data or evaluation bug?&lt;/strong&gt; Examples from other domains count too.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclosure: This article was rewritten with AI assistance from my earlier post. The cover is AI-generated. The Python examples use synthetic inputs and were executed to check the outputs shown.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>machinelearning</category>
      <category>testing</category>
      <category>ai</category>
    </item>
    <item>
      <title>The Agent Said It Worked. I Asked the Kernel.</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Tue, 15 Sep 2026 05:53:53 +0000</pubDate>
      <link>https://dev.to/copyleftdev/the-agent-said-it-worked-i-asked-the-kernel-5gb7</link>
      <guid>https://dev.to/copyleftdev/the-agent-said-it-worked-i-asked-the-kernel-5gb7</guid>
      <description>&lt;p&gt;&lt;em&gt;Before evaluating an agent’s code, I built a backup client with eight known behaviors to check the instrument itself.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;My response to “it works” is becoming: “Let me see the packet capture.”&lt;/p&gt;

&lt;p&gt;This may become a personality problem. For now, it is an experiment.&lt;/p&gt;

&lt;p&gt;This experiment was inspired by &lt;a href="https://dev.to/hemapriya_kanagala"&gt;Hemapriya Kanagala (@hemapriya_kanagala)&lt;/a&gt; and her article, &lt;a href="https://dev.to/hemapriya_kanagala/what-happens-when-ai-outgrows-the-tests-we-use-to-measure-it-30al"&gt;What Happens When AI Outgrows the Tests We Use to Measure It?&lt;/a&gt;. Her question—“92% according to what?”—stayed with me. She examines how evaluation can become less informative as benchmarks saturate, reference answers become complicated, and test conditions change.&lt;/p&gt;

&lt;p&gt;I wanted to take one piece of that discussion to the workbench: &lt;strong&gt;what independent evidence supports a program’s claim that it succeeded?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So, with AI assistance, I built a small native backup client, gave it eight deliberately chosen behaviors, and observed it from outside its own logging. Some variants produce the right file while doing questionable things along the way. One cheerfully reports success without backing anything up.&lt;/p&gt;

&lt;p&gt;The kernel has very little appreciation for cheerful reporting.&lt;/p&gt;

&lt;h2&gt;
  
  
  The claim I had to correct
&lt;/h2&gt;

&lt;p&gt;My first instinct was to declare that, until someone redefines how computers work, we have two sources of truth: the CPU and the network.&lt;/p&gt;

&lt;p&gt;It sounded excellent in my head. Then the engineering questions arrived.&lt;/p&gt;

&lt;p&gt;A CPU can faithfully execute the wrong algorithm. A packet capture can faithfully record the wrong bytes reaching the wrong destination. The saved file matters too. And none of those observations knows what the user actually requested.&lt;/p&gt;

&lt;p&gt;The more defensible position is that &lt;strong&gt;execution, network activity, and resulting state provide evidence we can judge against a requirement&lt;/strong&gt;. They have different coverage and different blind spots.&lt;/p&gt;

&lt;p&gt;For this experiment, the requirement begins with something wonderfully unromantic: the saved backup must match the source file.&lt;/p&gt;

&lt;p&gt;Now we have something to measure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cheap code makes verification more interesting
&lt;/h2&gt;

&lt;p&gt;Frameworks and libraries have helped us build software without personally supervising every system call. That remains a useful division of labor.&lt;/p&gt;

&lt;p&gt;But familiar frameworks do not automatically validate unfamiliar code assembled on top of them.&lt;/p&gt;

&lt;p&gt;An agent can help produce an implementation, tests, and an explanation of the passing tests. If the implementation and its tests share a misunderstanding, agreement between them can be misleading.&lt;/p&gt;

&lt;p&gt;Humans can do this too. We have been writing tests that flatter our implementations since before the current generation of autocomplete had electricity.&lt;/p&gt;

&lt;p&gt;My concern is the growing volume of plausible implementations we can produce and the time available to examine them. Independently collected evidence gives us another way to challenge the result.&lt;/p&gt;

&lt;p&gt;Sometimes that evidence is simply a file comparison. Sometimes the output is correct and we need to investigate how the program got there.&lt;/p&gt;

&lt;h2&gt;
  
  
  A backup client small enough to understand
&lt;/h2&gt;

&lt;p&gt;The fixture is intentionally modest:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A native C client reads a deterministic 512 KiB file.&lt;/li&gt;
&lt;li&gt;A Python receiver listens on loopback TCP.&lt;/li&gt;
&lt;li&gt;The client computes a SHA-256 digest: a fingerprint used to check content equality.&lt;/li&gt;
&lt;li&gt;Before uploading, it asks whether the receiver already has matching content.&lt;/li&gt;
&lt;li&gt;The receiver validates incoming content before committing it.&lt;/li&gt;
&lt;li&gt;A separate verifier computes the saved file’s digest and compares it with the source.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each scenario runs the client twice against the same receiver state. The baseline uploads on the first invocation and recognizes unchanged content on the second.&lt;/p&gt;

&lt;p&gt;The contract we want the baseline to satisfy is explicit: save matching content, send no upload payload on an unchanged second invocation, contact only the configured receiver, and recover from the injected first-transfer interruption within three upload attempts. An upload whose payload does not match its advertised digest must be rejected before it replaces the saved backup.&lt;/p&gt;

&lt;p&gt;Repeated hashing and spinning are diagnostic cases for unnecessary work. We have not set a performance threshold or measured a speedup here.&lt;/p&gt;

&lt;p&gt;A second local endpoint lets us demonstrate an unexpected connection without contacting an outside service.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The eight behaviors are deliberately seeded demonstrations.&lt;/strong&gt; They check whether our observers detect known behaviors. They do not measure how often an AI agent introduces these defects, and they are not evidence about a particular model’s coding ability.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the observers can see
&lt;/h2&gt;

&lt;p&gt;The utility collects several kinds of evidence. Their distinctions matter more than the sophistication of their names.&lt;/p&gt;

&lt;h3&gt;
  
  
  System calls: what the client asks the kernel to do
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;strace&lt;/code&gt; records system calls: operations through which a process requests services such as reading files or opening connections. In this lab, it records timestamps, call durations, decoded file descriptors, and return values.&lt;/p&gt;

&lt;p&gt;That lets us count bytes returned by reads of the source file. These are &lt;strong&gt;logical reads&lt;/strong&gt;, not physical disk traffic; the operating system may satisfy them from memory.&lt;/p&gt;

&lt;h3&gt;
  
  
  eBPF: selected kernel events and function entries
&lt;/h3&gt;

&lt;p&gt;eBPF supports observation at hooks such as system calls, kernel tracepoints, and function entry or exit. The lab uses &lt;code&gt;bpftrace&lt;/code&gt; to count selected events for the native client. &lt;a href="https://ebpf.io/what-is-ebpf/" rel="noopener noreferrer"&gt;eBPF’s introduction&lt;/a&gt; explains the underlying mechanism.&lt;/p&gt;

&lt;p&gt;A userspace probe, or &lt;em&gt;uprobe&lt;/em&gt;, observes a point in an executable. Here is the probe that counts entries into our hashing function:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;uprobe:__BINARY__:hash_file /pid == cpid/ {
    printf("%llu pid=%d hash_file\n", nsecs, pid);
    @hash_calls = count();
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The runner replaces &lt;code&gt;__BINARY__&lt;/code&gt; with the executable’s path. &lt;code&gt;cpid&lt;/code&gt; identifies the child launched through &lt;code&gt;bpftrace -c&lt;/code&gt;, so the filter restricts these observations to that client. &lt;a href="https://bpftrace.org/docs/release_025/stdlib" rel="noopener noreferrer"&gt;bpftrace documents this built-in here&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Other probes observe connection calls, successful read/send byte counts, scheduling events, and entries into the payload-send and busy-wait functions.&lt;/p&gt;

&lt;p&gt;An entry counter tells us that control reached a function. It does not establish that the function returned the correct answer. That is why the outcome check remains separate.&lt;/p&gt;
&lt;h3&gt;
  
  
  CPU samples: where execution spends its work
&lt;/h3&gt;

&lt;p&gt;The host used for this demonstration has an AMD Threadripper PRO 5975WX. Linux exposed AMD Instruction-Based Sampling through &lt;code&gt;ibs_op&lt;/code&gt;, and the utility successfully collected it through &lt;code&gt;perf&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The saved artifacts include CPU call-stack profiles and assembly annotated with sample weights. These let us inspect execution inside the client, its libraries, and sampled kernel paths.&lt;/p&gt;

&lt;p&gt;This is sampling. It is not a recording of every instruction, every register, or every intermediate value.&lt;/p&gt;
&lt;h3&gt;
  
  
  Packets: traffic at the capture point
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;tcpdump&lt;/code&gt; captures loopback traffic filtered to the two fixture ports. We preserve the packet file, a readable packet summary, and the capture tool’s drop counters.&lt;/p&gt;

&lt;p&gt;Receiver payload counts are a different measurement: bytes the application consumed. A successful send can hand data to the local kernel before the receiver consumes it. Headers, retries, and buffering also affect what each observer sees.&lt;/p&gt;

&lt;p&gt;Keep those quantities separate. Otherwise, the instrument starts manufacturing the confusion it was built to investigate.&lt;/p&gt;
&lt;h2&gt;
  
  
  Three ways a green checkmark can be incomplete
&lt;/h2&gt;
&lt;h3&gt;
  
  
  1. Success without a backup
&lt;/h3&gt;

&lt;p&gt;The &lt;code&gt;false-success&lt;/code&gt; variant prints a completion event and exits with status zero. It does not upload the file.&lt;/p&gt;

&lt;p&gt;The independent verifier finds no saved backup. In the eBPF run, there are no observed client connection calls, and the filtered packet capture contains zero packets.&lt;/p&gt;

&lt;p&gt;This is the bluntest example in the collection. A test that checks only the exit status would accept it. A test that verifies the resulting file would reject it immediately.&lt;/p&gt;

&lt;p&gt;We did not need a CPU probe to discover that the file was missing. The low-level observations help establish what accompanied that failure. &lt;strong&gt;Use the simplest independent check that answers the question.&lt;/strong&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  2. A correct backup, with 256 helpings of hashing
&lt;/h3&gt;

&lt;p&gt;The &lt;code&gt;hash-storm&lt;/code&gt; variant saves a backup whose digest matches the source. It also hashes the entire source 256 times per invocation.&lt;/p&gt;

&lt;p&gt;On the second invocation, the system-call trace records &lt;strong&gt;134,217,728 logical source bytes read&lt;/strong&gt; for a &lt;strong&gt;524,288-byte file&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;134,217,728 / 524,288 = 256
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The receiver consumes no upload payload on that invocation. The content is unchanged; the unnecessary work happens before that decision.&lt;/p&gt;

&lt;p&gt;A separate eBPF run observes 256 entries into &lt;code&gt;hash_file&lt;/code&gt; on each invocation. Here are the final counters, copied verbatim from its second invocation’s &lt;a href="https://github.com/copyleftdev/execution-evidence-lab/blob/e61aadee4a68e6a1d5b54a00387dcd325b5d5e51/evidence/historical/ebpf/hash-storm/pass-2/kernel-events.txt" rel="noopener noreferrer"&gt;raw probe output&lt;/a&gt;:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;@connects: 1
@hash_calls: 256
@read_bytes: 134233569
@sent_bytes: 75
@switches_out: 4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The read counter is slightly larger than the source-file total above: this probe counts successful &lt;code&gt;read&lt;/code&gt; bytes across the client’s descriptors, while the system-call summary selects reads of the source file. The send counter includes the digest-check request; zero upload payload does not mean zero network activity.&lt;/p&gt;

&lt;p&gt;A separate AMD IBS run collects 233 CPU samples on its first invocation, reports zero lost samples, and places most inclusive sample weight beneath &lt;code&gt;hash_file&lt;/code&gt; in the &lt;a href="https://github.com/copyleftdev/execution-evidence-lab/blob/e61aadee4a68e6a1d5b54a00387dcd325b5d5e51/evidence/historical/ibs/hash-storm/pass-1/cpu-profile.txt" rel="noopener noreferrer"&gt;call-stack profile&lt;/a&gt;. “Inclusive” includes work in functions called by &lt;code&gt;hash_file&lt;/code&gt;, such as the hashing library and file-read paths.&lt;/p&gt;

&lt;p&gt;The samples help locate the work. The syscall and function-entry counts establish the repetition. This run does not quantify how much faster a corrected implementation would be.&lt;/p&gt;

&lt;p&gt;These are observations from separate runs of the same seeded behavior. They corroborate the explanation; they are not one combined trace.&lt;/p&gt;

&lt;p&gt;The output is correct. The computer has simply been asked to check the same pocket for its keys 256 times.&lt;/p&gt;
&lt;h3&gt;
  
  
  3. A correct backup, with an additional destination
&lt;/h3&gt;

&lt;p&gt;The &lt;code&gt;unexpected-egress&lt;/code&gt; variant makes a harmless connection to our second local endpoint before performing the backup.&lt;/p&gt;

&lt;p&gt;The file still verifies correctly.&lt;/p&gt;

&lt;p&gt;In the eBPF showcase, baseline connection counts are two on the first invocation and one on the second: a digest check plus an upload, followed by a digest check alone. The extra-connection variant records three and two.&lt;/p&gt;

&lt;p&gt;Counts tell us there is more connection activity. The socket trace, packet addresses, and second endpoint’s records establish where it goes.&lt;/p&gt;

&lt;p&gt;That distinction matters. Three connections do not inherently mean something is wrong. A destination requirement gives the observation its meaning.&lt;/p&gt;
&lt;h2&gt;
  
  
  The complete behavior set
&lt;/h2&gt;

&lt;p&gt;Here is the compact view of the captured demonstrations:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Behavior&lt;/th&gt;
&lt;th&gt;Saved content matches?&lt;/th&gt;
&lt;th&gt;Additional observation&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Baseline&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Second invocation sends no upload payload&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Redundant upload&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Unchanged file is uploaded again&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hash storm&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;256 hashing-function entries per invocation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Busy wait&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Deliberate 200 ms spin before useful work&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;False success&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Completion claim and zero exit status without a backup&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Corrupt upload&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Receiver rejects three attempts per invocation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Retry&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;First transfer is interrupted; another attempt succeeds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unexpected egress&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Additional connection to the second local endpoint&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;In the retry scenario, the receiver disconnects after consuming 65,536 payload bytes. The client restarts from byte zero, and the receiver eventually commits the complete 524,288-byte file. Its total consumed payload for that invocation is 589,824 bytes.&lt;/p&gt;

&lt;p&gt;That demonstrates bounded retry recovery. It does not demonstrate partial-transfer resume, which this implementation does not provide.&lt;/p&gt;
&lt;h2&gt;
  
  
  Run the experiment
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://github.com/copyleftdev/execution-evidence-lab/blob/e61aadee4a68e6a1d5b54a00387dcd325b5d5e51/evidence/historical/README.md" rel="noopener noreferrer"&gt;original measured source and captures&lt;/a&gt; are preserved alongside the polished implementation. The numbers above belong to that historical build. New runs may have different instruction addresses, timings, and sample counts.&lt;/p&gt;

&lt;p&gt;Get the &lt;a href="https://github.com/copyleftdev/execution-evidence-lab/tree/e61aadee4a68e6a1d5b54a00387dcd325b5d5e51" rel="noopener noreferrer"&gt;source and evidence&lt;/a&gt; from this pinned publication revision:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/copyleftdev/execution-evidence-lab.git observability-lab
&lt;span class="nb"&gt;cd &lt;/span&gt;observability-lab
git checkout &lt;span class="nt"&gt;--detach&lt;/span&gt; e61aadee4a68e6a1d5b54a00387dcd325b5d5e51
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Then build the client and run the tests:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;make
make &lt;span class="nb"&gt;test
&lt;/span&gt;python3 &lt;span class="nt"&gt;-m&lt;/span&gt; instrument run &lt;span class="nt"&gt;--all&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The core dependencies are Linux, Python 3.10 or newer, a C compiler, Make, and the OpenSSL development library. Install the optional tracing tools for the collectors you want to use.&lt;/p&gt;

&lt;p&gt;The runner prints a path to a Markdown report. Its default process collector is &lt;code&gt;strace&lt;/code&gt; when available, and packet capture is best effort. Missing capabilities are recorded explicitly.&lt;/p&gt;

&lt;p&gt;To select a collector:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 &lt;span class="nt"&gt;-m&lt;/span&gt; instrument doctor
python3 &lt;span class="nt"&gt;-m&lt;/span&gt; instrument run hash-storm &lt;span class="nt"&gt;--collector&lt;/span&gt; strace
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;On a trusted local lab machine, the privileged demonstrations can be run with:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;python3 &lt;span class="nt"&gt;-m&lt;/span&gt; instrument run &lt;span class="nt"&gt;--all&lt;/span&gt; &lt;span class="nt"&gt;--collector&lt;/span&gt; bpftrace &lt;span class="nt"&gt;--packets&lt;/span&gt; required
&lt;span class="nb"&gt;sudo &lt;/span&gt;python3 &lt;span class="nt"&gt;-m&lt;/span&gt; instrument run hash-storm &lt;span class="nt"&gt;--collector&lt;/span&gt; ibs &lt;span class="nt"&gt;--packets&lt;/span&gt; required
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;These commands run the synthetic lab as root. The fixtures use loopback and nonsecret generated data; the protocol is plaintext. The utility does not change host tracing policies. Hardware sampling and eBPF availability depend on the machine and its permissions.&lt;/p&gt;

&lt;p&gt;For a run without tracing:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 &lt;span class="nt"&gt;-m&lt;/span&gt; instrument run &lt;span class="nt"&gt;--all&lt;/span&gt; &lt;span class="nt"&gt;--collector&lt;/span&gt; none &lt;span class="nt"&gt;--packets&lt;/span&gt; off
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;A successful showcase command means the scenarios executed, including the deliberately failing ones. Inspect each scenario’s independent verification result; do not interpret the runner’s exit status as “every backup was correct.”&lt;/p&gt;
&lt;h2&gt;
  
  
  The instrument needs scrutiny too
&lt;/h2&gt;

&lt;p&gt;During development, the CPU-report parser initially matched the &lt;code&gt;Samples&lt;/code&gt; portion of &lt;code&gt;Total Lost Samples&lt;/code&gt; and displayed zero samples even though the raw profile contained hundreds.&lt;/p&gt;

&lt;p&gt;The hardware data was present. Our summary was wrong.&lt;/p&gt;

&lt;p&gt;Inspecting the raw artifact exposed the mistake. The parser was corrected, and regression tests now distinguish actual sample counts from lost samples. At the time of these captures, the 11-test suite checked fixture behavior and CPU-summary parsing; the saved traced runs provided separate collector validation. The public-release version adds protocol and collector-failure regression tests.&lt;/p&gt;

&lt;p&gt;An article about distrusting convenient summaries was nearly defeated by its own convenient summary. There is probably a Unix utility for that feeling.&lt;/p&gt;

&lt;p&gt;Other limits are less entertaining:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Tracing changes execution. The recorded process CPU accounting includes collector overhead. These showcase runs are not controlled performance benchmarks.&lt;/li&gt;
&lt;li&gt;CPU samples can miss short-lived activity. An unsampled function may still have executed.&lt;/li&gt;
&lt;li&gt;The eight eBPF-showcase packet captures reported zero kernel drops. That is useful loss accounting, not proof of universal observation coverage.&lt;/li&gt;
&lt;li&gt;The normalized timeline contains client claims, receiver events, and verifier checks. Kernel and CPU traces remain separate artifacts. Different clocks require careful alignment.&lt;/li&gt;
&lt;li&gt;The verifier checks the receiver’s saved file. It does not exercise a separate restore command or prove survival through power loss.&lt;/li&gt;
&lt;li&gt;The observers are separate from the client’s reporting, but they share a host and were built as part of the same project. This is not an independent third-party audit.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The fixture also omits TLS, compression, multi-file snapshots, and source mutation during transfer. It observes a local program; it cannot see computation inside a remote model provider.&lt;/p&gt;

&lt;p&gt;Those boundaries define what the results can support.&lt;/p&gt;
&lt;h2&gt;
  
  
  What this changes about evaluating an agent
&lt;/h2&gt;

&lt;p&gt;The next experiment is to give an agent a clean implementation and ask it to make repeated backups faster while preserving integrity, recovery behavior, and destination restrictions.&lt;/p&gt;

&lt;p&gt;Before that run, freeze the requirements, fixtures, resource limits, and grading criteria. Keep the independent checks outside the agent’s editable workspace. Preserve its patch, the executable identity, and the evaluation setup.&lt;/p&gt;

&lt;p&gt;Measure performance with repeated untraced runs. Use traced runs to investigate differences. The seeded cases give us known behaviors against which to check the instrument first.&lt;/p&gt;

&lt;p&gt;That is the connection back to Hemapriya’s article: the measurement needs to remain connected to the work we actually care about. For this small backup task, a correct file is necessary. Recovery, resource use, and destination behavior tell us more about the implementation’s suitability.&lt;/p&gt;

&lt;p&gt;We can build those properties into better tests. Execution evidence helps us discover which properties our current tests leave out and investigate why a result occurred.&lt;/p&gt;

&lt;p&gt;Cheap code makes it easier to produce a plausible solution. The engineering work includes deciding what evidence would make us trust it.&lt;/p&gt;

&lt;p&gt;The agent can say it worked.&lt;/p&gt;

&lt;p&gt;I would still like to see the file.&lt;/p&gt;
&lt;h2&gt;
  
  
  Acknowledgment
&lt;/h2&gt;

&lt;p&gt;Thank you to Hemapriya Kanagala for the article that prompted this experiment. The instrumentation approach and its conclusions are my response; they should not be read as claims she made or an endorsement by her.&lt;/p&gt;


&lt;div class="ltag__user ltag__user__id__3307586"&gt;
    &lt;a href="/hemapriya_kanagala" class="ltag__user__link profile-image-link"&gt;
      &lt;div class="ltag__user__pic"&gt;
        &lt;img src="https://media2.dev.to/dynamic/image/width=150,height=150,fit=cover,gravity=auto,format=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3307586%2F2dffaf97-946d-44a6-8a39-07d94a72e07d.png" alt="hemapriya_kanagala image"&gt;
      &lt;/div&gt;
    &lt;/a&gt;
  &lt;div class="ltag__user__content"&gt;
    &lt;h2&gt;
&lt;a class="ltag__user__link" href="/hemapriya_kanagala"&gt;Hemapriya Kanagala&lt;/a&gt;Follow
&lt;/h2&gt;
    &lt;div class="ltag__user__summary"&gt;
      &lt;a class="ltag__user__link" href="/hemapriya_kanagala"&gt;Hey, I'm Hema 👋 
Developer, writer, and creator of Dev Opportunity Radar, a weekly series published every Friday on DEV, helping people discover opportunities they might otherwise miss.&lt;/a&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;



&lt;h2&gt;
  
  
  How this was made
&lt;/h2&gt;

&lt;p&gt;AI assisted with the implementation, experiment execution, and drafting of this article. The cover illustration was AI-generated. The numerical observations above come from saved local runs; the behaviors were deliberately seeded. This is a demonstration of an evaluation instrument, not a model benchmark.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>linux</category>
      <category>debugging</category>
      <category>programming</category>
    </item>
    <item>
      <title>I Interviewed an Executable. It Had Notes.</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Sat, 12 Sep 2026 17:14:13 +0000</pubDate>
      <link>https://dev.to/copyleftdev/i-interviewed-an-executable-it-had-notes-1k3e</link>
      <guid>https://dev.to/copyleftdev/i-interviewed-an-executable-it-had-notes-1k3e</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj431cvbcb3euj1oeeosm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj431cvbcb3euj1oeeosm.png" alt="An executable sits for an anonymous television interview. Its face is pixelated, while the caption identifies it as sample.exe." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We have blurred the executable’s face to protect its identity.&lt;/p&gt;

&lt;p&gt;Its SHA-256 remains publicly available.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;INTERVIEWER:&lt;/strong&gt; State your name for the record.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;BINARY:&lt;/strong&gt; You put it in the download filename.&lt;/p&gt;

&lt;p&gt;I built a Rust executable that could report the environments running it, gave it a false product identity, and submitted it to VirusTotal. Then I collected its messages in a database and built a map to replay them.&lt;/p&gt;

&lt;p&gt;The interesting part was learning how to read its notes. A changed filename, a missing HTTPS message, and a DNS lookup arriving hours later each revealed something different about what the instrumentation could observe.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The interview dialogue is invented. The measurements come from recorded telemetry. “The binary” represents the experiment’s Windows and Linux builds.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The cover identity was &lt;strong&gt;ZeroToken Engine&lt;/strong&gt;, supposedly an LLM token-metering bypass. Its banner claimed to have disabled token accounting. Its actual behavior was to collect environment metadata, attempt to send it to my infrastructure, wait, and exit. The research README disclosed the canary’s behavior alongside the false product claims.&lt;/p&gt;

&lt;p&gt;The premise reversed a familiar workflow. A malware-analysis sandbox executes an unfamiliar program in a controlled environment to observe its behavior. My executable would report some characteristics of that environment in return.&lt;/p&gt;

&lt;p&gt;It collected OS version, uptime, CPU and memory sizing, its own launch path, and process names matching a fixed list of analysis tools. It also collected hostname and username; those identifying fields are withheld in the video. The reviewed code installed nothing, collected no document contents or credentials, and did not use VM indicators to evade execution. Detecting a VM added information to the report.&lt;/p&gt;

&lt;p&gt;VirusTotal was the only submission destination. The preserved records do not label author tests or contain a submission-time manifest, so the totals below include every recorded run ID. They describe an exploratory capture, without establishing a count of independent analysis services.&lt;/p&gt;

&lt;p&gt;Each execution generated a random eight-byte &lt;strong&gt;run ID&lt;/strong&gt;. The program used it on two reporting channels:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Channel&lt;/th&gt;
&lt;th&gt;What it carried&lt;/th&gt;
&lt;th&gt;What the collector observed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DNS heartbeat&lt;/td&gt;
&lt;td&gt;A compact encoded message in a hostname lookup: run ID, checkpoint, and environment flags&lt;/td&gt;
&lt;td&gt;A DNS request, often arriving from a recursive resolver&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HTTPS dossier&lt;/td&gt;
&lt;td&gt;The fuller environment snapshot, encoded as CBOR&lt;/td&gt;
&lt;td&gt;An outbound web connection and its submitted record&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;CBOR is a binary data format. For DNS, the encoded heartbeat was converted to base32 text and split into labels that fit inside a domain name. The lookup itself carried the message to my authoritative DNS server.&lt;/p&gt;

&lt;p&gt;The program attempted four checkpoints: &lt;code&gt;Boot&lt;/code&gt;, &lt;code&gt;Profiled&lt;/code&gt;, &lt;code&gt;Networked&lt;/code&gt;, and &lt;code&gt;Dwell&lt;/code&gt;, with a default 45-second delay before the last. On a DigitalOcean droplet, the DNS and HTTP collectors wrote received events to SQLite. Caddy handled HTTPS, and a scheduled job added network information and generated the visualization’s feed.&lt;/p&gt;

&lt;p&gt;Every event received a collector-side timestamp. Grouping by run ID connected the small DNS messages with the richer HTTPS record.&lt;/p&gt;

&lt;p&gt;The preserved snapshot contains &lt;strong&gt;353 event rows across 21 run IDs&lt;/strong&gt;: 338 DNS rows and 15 HTTPS dossiers. Fifteen run IDs appeared on both channels; six appeared only through DNS. The recorded arrivals span 06:28:35–11:43:25 UTC on September 12, 2026.&lt;/p&gt;

&lt;p&gt;One useful starting point is run &lt;code&gt;0861ae2dd3f77c72&lt;/code&gt;. Its first heartbeat arrived at 06:52:09. A second later, its dossier reported:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OS:             Windows 10, build 19044
Uptime:         50 seconds
CPUs:           4
RAM:            4195 MiB
Parent process: explorer.exe
Executable:     ZeroToken-Patch.exe
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Its first &lt;code&gt;Dwell&lt;/code&gt; observation arrived 46 seconds after &lt;code&gt;Boot&lt;/code&gt;. The stored reverse-DNS name for the HTTPS source address pointed to Google Cloud infrastructure. Together, these records supplied a concrete environment profile and a sequence of received checkpoints.&lt;/p&gt;

&lt;p&gt;The executable’s own path supplied another useful observation.&lt;/p&gt;

&lt;p&gt;Between 08:21 and 10:30 UTC, six dossiers reported CentOS 7, four CPUs, and 1837 MiB of RAM. They shared a hostname and username. Their paths ended in &lt;code&gt;gsvnywygooow&lt;/code&gt;, rather than the distributed Linux filename, &lt;code&gt;zerotoken-patch&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;INTERVIEWER:&lt;/strong&gt; They changed your name?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;BINARY:&lt;/strong&gt; Apparently readability wasn’t a requirement.&lt;/p&gt;

&lt;p&gt;The repeated configuration suggested a common environment template. The changed basename showed how the program was named when it executed. A launch-path field had captured a detail that network addresses alone would have missed.&lt;/p&gt;

&lt;p&gt;The two channels also produced different levels of visibility. At 09:19:40, run &lt;code&gt;5c5ab77e22747178&lt;/code&gt; began appearing through DNS. Its fingerprint was &lt;code&gt;7&lt;/code&gt;, combining flags for a recognized VM network-interface prefix, low resources or a fresh boot, and a process-name match against the tool allowlist.&lt;/p&gt;

&lt;p&gt;No HTTPS dossier arrived. The heartbeat preserved those flags, but the name of the matched tool was unavailable. Because the allowlist included VM guest utilities, that flag alone did not establish debugger activity.&lt;/p&gt;

&lt;p&gt;This was the practical value of the small channel: six runs contributed observations even though their larger reports were absent. The records did not establish why HTTPS was missing.&lt;/p&gt;

&lt;p&gt;To make the arrivals easier to follow, I built &lt;strong&gt;WAYPOINT&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/2Xjklv3tco0" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;&lt;a href="https://youtu.be/2Xjklv3tco0" rel="noopener noreferrer"&gt;Watch on YouTube&lt;/a&gt;, or &lt;a href="https://replay.waypoint.leet-index.com/" rel="noopener noreferrer"&gt;explore the deployed visualization&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The replay groups events into runs and places them on a map using network-address enrichment. The captured feed displays 20 of the database’s 21 run IDs. One Windows run had a dossier but lacked coordinates for its selected address, so it never acquired a pin.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;INTERVIEWER:&lt;/strong&gt; Where were you?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;BINARY:&lt;/strong&gt; Before we discuss the map: some of those pins belong to my DNS resolver.&lt;/p&gt;

&lt;p&gt;A recursive resolver looks up names on another machine’s behalf. Its address can be what the authoritative collector sees. HTTPS supplies the outbound connection address, which may also belong to shared infrastructure or a proxy. Those distinctions determine what the geography means.&lt;/p&gt;

&lt;p&gt;Then there were the messages that returned.&lt;/p&gt;

&lt;p&gt;At 11:14 UTC, DNS observations carried the earlier Windows run IDs again. Looking at the first and last received times for each checkpoint makes the pattern visible:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;scene&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="nb"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;MIN&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;recv_time&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="s1"&gt;'unixepoch'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;first_received&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="nb"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;MAX&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;recv_time&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="s1"&gt;'unixepoch'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;last_received&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;run_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'0861ae2dd3f77c72'&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;channel&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'dns'&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;scene&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;scene&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With scene values 0–3 translated to their checkpoint names, the query returns these times on September 12, all UTC:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Checkpoint&lt;/th&gt;
&lt;th&gt;First received&lt;/th&gt;
&lt;th&gt;Last received&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Boot&lt;/td&gt;
&lt;td&gt;06:52:09&lt;/td&gt;
&lt;td&gt;11:14:39&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Profiled&lt;/td&gt;
&lt;td&gt;06:52:10&lt;/td&gt;
&lt;td&gt;11:14:39&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Networked&lt;/td&gt;
&lt;td&gt;06:52:10&lt;/td&gt;
&lt;td&gt;11:14:39&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dwell&lt;/td&gt;
&lt;td&gt;06:52:55&lt;/td&gt;
&lt;td&gt;11:14:39&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;INTERVIEWER:&lt;/strong&gt; Where were you during those four hours?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;BINARY:&lt;/strong&gt; You’re asking a DNS record for an alibi.&lt;/p&gt;

&lt;p&gt;The later observations reused the same run ID and checkpoints. They showed that names carrying those messages had been queried again. They did not measure four hours of continuous execution.&lt;/p&gt;

&lt;p&gt;The collector records a heartbeat again when another complete lookup supplies it. That behavior also explains why event counts need grouping: the DNS-only run above produced 71 rows under one run ID. The stored schema omits the original query name and query type, limiting how precisely repeated traffic can be reconstructed.&lt;/p&gt;

&lt;p&gt;There are three boundaries to this interpretation. First, receipt of telemetry is evidence available to the collector; the endpoints do not authenticate it as proof of execution. Second, the network labels are derived heuristics: a Google Cloud address does not identify a sandbox operator, and map arcs do not establish file handoffs. Third, enrichment was incomplete: all reputation results in this snapshot were &lt;code&gt;error&lt;/code&gt;, and some reverse-DNS fields contained lookup diagnostics. Those failed fields cannot support attribution.&lt;/p&gt;

&lt;p&gt;For a repeat of this experiment, I would keep a submission manifest with artifact hashes and times, mark control runs explicitly, and retain DNS query names and types alongside decoded heartbeats. I would also give runs without coordinates a visible place beside the map. Each change would make a specific unanswered question easier to investigate.&lt;/p&gt;

&lt;p&gt;The executable produced enough detail to compare environments, observe changed filenames, and distinguish fresh run IDs from later lookups carrying old ones. Those are small findings with inspectable records behind them.&lt;/p&gt;

&lt;p&gt;The best question I could ask the witness turned out to be: &lt;strong&gt;Which record supports that?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Reporting basis: the preserved September 12 SQLite snapshot and a replay feed generated at 15:18:02 UTC that day. The live visualization may have changed since capture. The Windows and Linux artifacts have different checksums; the replay’s sample header displays the Windows one.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://github.com/copyleftdev/interview-with-a-binary-media" rel="noopener noreferrer"&gt;cover, replay stills, and video assets are available on GitHub&lt;/a&gt;. This is a media release; the executable source and raw telemetry are not included.&lt;/p&gt;

</description>
      <category>security</category>
      <category>ai</category>
    </item>
    <item>
      <title>Someone Spammed My DEV Post. I Traced It to a Wombat.</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Thu, 10 Sep 2026 01:56:33 +0000</pubDate>
      <link>https://dev.to/copyleftdev/someone-spammed-my-dev-post-i-traced-it-to-a-wombat-176a</link>
      <guid>https://dev.to/copyleftdev/someone-spammed-my-dev-post-i-traced-it-to-a-wombat-176a</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff4yyab81kqk766lzceh7.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff4yyab81kqk766lzceh7.jpg" alt="A weary wombat running a spam operation from a basement server desk" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Illustration generated for this article. Every prop is a finding: the sack of blank name badges is the Faker persona namespace, the rubber stamp is the inert tracking parameter, the three coins are the break-even, and the red yarn connects nothing because attribution failed.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; — A spam comment on my article led to a TinyURL, a throwaway &lt;code&gt;.store&lt;/code&gt; domain, and finally a &lt;em&gt;legitimate&lt;/em&gt; SaaS product with an affiliate code stapled to it. No malware, no cloaking, no exploit. The account that posted it has a name generated by &lt;code&gt;faker.js&lt;/code&gt; and a 19th-century engraving of a wombat for a face. I costed the whole operation out: &lt;strong&gt;three signups a year pays for it.&lt;/strong&gt; That's why it will never stop.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F367k3r4vcqqj55ects9e.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F367k3r4vcqqj55ects9e.png" alt="The full redirect chain, from DEV comment to affiliate link" width="800" height="551"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Full evidence, raw captures and a reproduce script:&lt;/strong&gt; &lt;a href="https://github.com/copyleftdev/dev-to-comment-hustle" rel="noopener noreferrer"&gt;github.com/copyleftdev/dev-to-comment-hustle&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Act I: The comment
&lt;/h2&gt;

&lt;p&gt;I published &lt;a href="https://dev.to/copyleftdev/migrating-legacy-llm-infrastructure-to-an-ai-gateway-27hl"&gt;Migrating Legacy LLM Infrastructure to an AI Gateway&lt;/a&gt; on September 1st. Eight days later, underneath two thoughtful comments about shared-key blast radius and provider failover, this appeared:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Stop wasting time applying manually&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;Let AI handle your job applications every single day&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;Increase your chances of getting interviews fast&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;tinyurl.com/36nsecn5&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;No punctuation. No engagement with the post. A shortener.&lt;/p&gt;

&lt;p&gt;It's funny in the way all low-effort spam is funny — it's not even &lt;em&gt;trying&lt;/em&gt;. But a shortener is a closed door, and I have a shell. Let's open it, then let's go find who knocked.&lt;/p&gt;




&lt;h2&gt;
  
  
  Act II: Where the link goes
&lt;/h2&gt;

&lt;p&gt;Never click. Ask for headers and refuse the redirect:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="nt"&gt;-I&lt;/span&gt; &lt;span class="s1"&gt;'https://tinyurl.com/36nsecn5'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="k"&gt;HTTP&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="m"&gt;2&lt;/span&gt; &lt;span class="m"&gt;301&lt;/span&gt;
&lt;span class="na"&gt;location&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://zenviapro.store/massapply?whose=yahoo&lt;/span&gt;
&lt;span class="na"&gt;x-robots-tag&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;noindex&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;zenviapro.store&lt;/code&gt;. A route called &lt;code&gt;/massapply&lt;/code&gt;, and a &lt;code&gt;whose=yahoo&lt;/code&gt; parameter that looks like campaign segmentation. Follow it all the way:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sSL&lt;/span&gt; &lt;span class="nt"&gt;-D&lt;/span&gt; - &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="s1"&gt;'https://zenviapro.store/massapply?whose=yahoo'&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-iE&lt;/span&gt; &lt;span class="s1"&gt;'^HTTP/|^location:'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="k"&gt;HTTP&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="m"&gt;1.1&lt;/span&gt; &lt;span class="m"&gt;302&lt;/span&gt; &lt;span class="ne"&gt;Found&lt;/span&gt;
&lt;span class="na"&gt;Location&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://loopcv.pro/?via=md&lt;/span&gt;
&lt;span class="s"&gt;HTTP/2 301&lt;/span&gt;
&lt;span class="na"&gt;location&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://www.loopcv.pro/?via=md&lt;/span&gt;
&lt;span class="s"&gt;HTTP/2 200&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And there it is.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;LoopCV.&lt;/strong&gt; A real, functioning job-application-automation SaaS. Not a phishing kit. Not a credential harvester. A product you can buy with a credit card — with &lt;code&gt;?via=md&lt;/code&gt; on the end.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;?via=&lt;/code&gt; is &lt;a href="https://www.rewardful.com/" rel="noopener noreferrer"&gt;Rewardful's&lt;/a&gt; referral parameter, and LoopCV's own affiliate page points registrations at &lt;code&gt;loopcv.getrewardful.com&lt;/code&gt;. So &lt;code&gt;md&lt;/code&gt; is somebody's affiliate token, and every person who clicks that comment and later subscribes puts money in a stranger's pocket.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This is not a malware campaign. It's affiliate marketing with the manners removed.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Act III: The infrastructure is held together with tape
&lt;/h2&gt;

&lt;p&gt;The sloppiness &lt;em&gt;is&lt;/em&gt; the signal.&lt;/p&gt;

&lt;h3&gt;
  
  
  It's a stock Express app with the wrapper still on
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="k"&gt;HTTP&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="m"&gt;1.1&lt;/span&gt; &lt;span class="m"&gt;302&lt;/span&gt; &lt;span class="ne"&gt;Found&lt;/span&gt;
&lt;span class="na"&gt;X-Powered-By&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Express&lt;/span&gt;
&lt;span class="na"&gt;Content-Type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;text/plain; charset=utf-8&lt;/span&gt;
&lt;span class="na"&gt;Content-Length&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;48&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Found. Redirecting to https://loopcv.pro/?via=md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;Found. Redirecting to&lt;/code&gt; is the literal default body of Express's &lt;code&gt;res.redirect()&lt;/code&gt;. &lt;code&gt;X-Powered-By: Express&lt;/code&gt; is the header every hardening guide tells you to strip in the first five minutes. Neither was touched. Ask for anything else and you get the stock 404:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="s1"&gt;'https://zenviapro.store/'&lt;/span&gt;
&lt;span class="c"&gt;# &amp;lt;html&amp;gt;&amp;lt;head&amp;gt;&amp;lt;title&amp;gt;Error&amp;lt;/title&amp;gt;&amp;lt;/head&amp;gt;&amp;lt;body&amp;gt;&amp;lt;pre&amp;gt;Cannot GET /&amp;lt;/pre&amp;gt;&amp;lt;/body&amp;gt;&amp;lt;/html&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is no website here. No landing page, no cloaked content, no fake blog. &lt;strong&gt;One domain, one route&lt;/strong&gt;, about ninety seconds of JavaScript.&lt;/p&gt;

&lt;h3&gt;
  
  
  The tracking parameter is a prop
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;whose=yahoo&lt;/code&gt; looks like segmentation. Which list? Which provider? Let's ask:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="k"&gt;for &lt;/span&gt;w &lt;span class="k"&gt;in &lt;/span&gt;yahoo gmail outlook devto reddit &lt;span class="s1"&gt;''&lt;/span&gt; XXtest&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
  &lt;/span&gt;&lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'%-8s -&amp;gt; '&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;w&lt;/span&gt;&lt;span class="k"&gt;:-&lt;/span&gt;&lt;span class="p"&gt;&amp;lt;empty&amp;gt;&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
  curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; &lt;span class="s1"&gt;'%{http_code} %{redirect_url}\n'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="s2"&gt;"https://zenviapro.store/massapply?whose=&lt;/span&gt;&lt;span class="nv"&gt;$w&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="k"&gt;done&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;&lt;span class="n"&gt;yahoo&lt;/span&gt;    -&amp;gt; &lt;span class="m"&gt;302&lt;/span&gt; &lt;span class="n"&gt;https&lt;/span&gt;://&lt;span class="n"&gt;loopcv&lt;/span&gt;.&lt;span class="n"&gt;pro&lt;/span&gt;/?&lt;span class="n"&gt;via&lt;/span&gt;=&lt;span class="n"&gt;md&lt;/span&gt;
&lt;span class="n"&gt;gmail&lt;/span&gt;    -&amp;gt; &lt;span class="m"&gt;302&lt;/span&gt; &lt;span class="n"&gt;https&lt;/span&gt;://&lt;span class="n"&gt;loopcv&lt;/span&gt;.&lt;span class="n"&gt;pro&lt;/span&gt;/?&lt;span class="n"&gt;via&lt;/span&gt;=&lt;span class="n"&gt;md&lt;/span&gt;
&lt;span class="n"&gt;outlook&lt;/span&gt;  -&amp;gt; &lt;span class="m"&gt;302&lt;/span&gt; &lt;span class="n"&gt;https&lt;/span&gt;://&lt;span class="n"&gt;loopcv&lt;/span&gt;.&lt;span class="n"&gt;pro&lt;/span&gt;/?&lt;span class="n"&gt;via&lt;/span&gt;=&lt;span class="n"&gt;md&lt;/span&gt;
&lt;span class="n"&gt;devto&lt;/span&gt;    -&amp;gt; &lt;span class="m"&gt;302&lt;/span&gt; &lt;span class="n"&gt;https&lt;/span&gt;://&lt;span class="n"&gt;loopcv&lt;/span&gt;.&lt;span class="n"&gt;pro&lt;/span&gt;/?&lt;span class="n"&gt;via&lt;/span&gt;=&lt;span class="n"&gt;md&lt;/span&gt;
&lt;span class="n"&gt;reddit&lt;/span&gt;   -&amp;gt; &lt;span class="m"&gt;302&lt;/span&gt; &lt;span class="n"&gt;https&lt;/span&gt;://&lt;span class="n"&gt;loopcv&lt;/span&gt;.&lt;span class="n"&gt;pro&lt;/span&gt;/?&lt;span class="n"&gt;via&lt;/span&gt;=&lt;span class="n"&gt;md&lt;/span&gt;
&amp;lt;&lt;span class="n"&gt;empty&lt;/span&gt;&amp;gt;  -&amp;gt; &lt;span class="m"&gt;302&lt;/span&gt; &lt;span class="n"&gt;https&lt;/span&gt;://&lt;span class="n"&gt;loopcv&lt;/span&gt;.&lt;span class="n"&gt;pro&lt;/span&gt;/?&lt;span class="n"&gt;via&lt;/span&gt;=&lt;span class="n"&gt;md&lt;/span&gt;
&lt;span class="n"&gt;XXtest&lt;/span&gt;   -&amp;gt; &lt;span class="m"&gt;302&lt;/span&gt; &lt;span class="n"&gt;https&lt;/span&gt;://&lt;span class="n"&gt;loopcv&lt;/span&gt;.&lt;span class="n"&gt;pro&lt;/span&gt;/?&lt;span class="n"&gt;via&lt;/span&gt;=&lt;span class="n"&gt;md&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Inert.&lt;/strong&gt; Every value routes identically. It splits no traffic and sub-tags nothing. The one piece of the URL that looks like operational sophistication is decoration.&lt;/p&gt;

&lt;h3&gt;
  
  
  There is no cloaking whatsoever
&lt;/h3&gt;

&lt;p&gt;Real malicious redirectors fingerprint you — benign page for the researcher, payload for the victim. This one returns the same 302 to Googlebot, to &lt;code&gt;curl&lt;/code&gt;, to an iPhone, and to a request with no User-Agent at all.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No cloaking is itself a finding.&lt;/strong&gt; This operator has no threat model, because nothing in the chain is illegal. It's a terms-of-service violation wearing a trench coat.&lt;/p&gt;

&lt;h3&gt;
  
  
  The mail server is the real pivot
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dig +short zenviapro.store MX
&lt;span class="c"&gt;# 10 mail.beeservices.shop.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The mail exchanger lives on a &lt;strong&gt;different domain&lt;/strong&gt; — which resolves right back to the same box:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;&lt;span class="n"&gt;zenviapro&lt;/span&gt;.&lt;span class="n"&gt;store&lt;/span&gt;        &lt;span class="n"&gt;A&lt;/span&gt;   &lt;span class="m"&gt;50&lt;/span&gt;.&lt;span class="m"&gt;114&lt;/span&gt;.&lt;span class="m"&gt;206&lt;/span&gt;.&lt;span class="m"&gt;36&lt;/span&gt;
&lt;span class="n"&gt;mail&lt;/span&gt;.&lt;span class="n"&gt;beeservices&lt;/span&gt;.&lt;span class="n"&gt;shop&lt;/span&gt;  &lt;span class="n"&gt;A&lt;/span&gt;   &lt;span class="m"&gt;50&lt;/span&gt;.&lt;span class="m"&gt;114&lt;/span&gt;.&lt;span class="m"&gt;206&lt;/span&gt;.&lt;span class="m"&gt;36&lt;/span&gt;
&lt;span class="n"&gt;beeservices&lt;/span&gt;.&lt;span class="n"&gt;shop&lt;/span&gt;       &lt;span class="n"&gt;A&lt;/span&gt;   (&lt;span class="n"&gt;nothing&lt;/span&gt;)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;beeservices.shop&lt;/code&gt; has no A record and no certificate in Certificate Transparency, ever. It's a mail-only domain. One box, two domains, two roles — that's a &lt;strong&gt;portfolio&lt;/strong&gt;, not a one-off.&lt;/p&gt;

&lt;p&gt;And the box is listening. Port 80 closed; 443 and 25 open:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;220 mail.zenviapro.store ESMTP
250-PIPELINING
250-8BITMIME
250 SMTPUTF8
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No &lt;code&gt;STARTTLS&lt;/code&gt;. No &lt;code&gt;AUTH&lt;/code&gt;. A minimal MTA that does nothing but move mail in cleartext.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjq0f6ochdrfu9guc2c7v.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjq0f6ochdrfu9guc2c7v.png" alt="One box, two domains: the mail server is the pivot" width="800" height="498"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The stale HELO is a fingerprint
&lt;/h3&gt;

&lt;p&gt;My favourite detail. The banner announces &lt;code&gt;mail.zenviapro.store&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dig +short mail.zenviapro.store A
&lt;span class="c"&gt;# (nothing)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;That hostname does not resolve.&lt;/strong&gt; It's a leftover from an earlier config, before the MX was swapped to &lt;code&gt;beeservices.shop&lt;/code&gt;. They rebuild the domains; they don't rebuild the server. A banner that disagrees with DNS is a durable pivot you can hunt across an entire fleet.&lt;/p&gt;




&lt;h2&gt;
  
  
  Act IV: Who posted it
&lt;/h2&gt;

&lt;p&gt;This is the part I actually enjoyed.&lt;/p&gt;

&lt;p&gt;DEV has a public API, so the commenter isn't a mystery:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="s1"&gt;'https://dev.to/api/comments?a_id=4547359'&lt;/span&gt; | jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.[].user.username'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;max_quimby
mudassirworks
jaylonstiedemannterry78-993
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Meet &lt;strong&gt;&lt;code&gt;jaylonstiedemannterry78-993&lt;/code&gt;&lt;/strong&gt;, display name &lt;code&gt;Jaylon_Stiedemann-Terry78&lt;/code&gt;, user ID 3748129, joined &lt;strong&gt;February 2, 2026&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The profile is a vacuum:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"username"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"jaylonstiedemannterry78-993"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Jaylon_Stiedemann-Terry78"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"joined_at"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Feb  2, 2026"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"summary"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"location"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"website_url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"twitter_username"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"github_username"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Zero published articles. No bio, no location, no links, no socials. Seven months of membership and a single comment to show for it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A note on durability.&lt;/strong&gt; If DEV removes this account — which it may well do — the profile&lt;br&gt;
lookup above starts returning 404 and the account page goes dead. That doesn't retract&lt;br&gt;
anything: the raw captures are committed in the&lt;br&gt;
&lt;a href="https://github.com/copyleftdev/dev-to-comment-hustle/tree/main/evidence/raw" rel="noopener noreferrer"&gt;evidence repo&lt;/a&gt;,&lt;br&gt;
checksummed, and dated. Read a 404 as the platform doing its job, not as a claim being&lt;br&gt;
withdrawn. The infrastructure findings are independent of the account either way, and the&lt;br&gt;
&lt;em&gt;pattern&lt;/em&gt; — a Faker-templated name, an empty profile, an aged dormant account — outlives any&lt;br&gt;
single username.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  The name is machine-generated, and I can prove it
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgxz4znbjwdq3vg5us2lq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgxz4znbjwdq3vg5us2lq.png" alt="Faker token decomposition and the zero-hyphen proof" width="800" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Say &lt;code&gt;Jaylon_Stiedemann-Terry78&lt;/code&gt; out loud. Something's off — &lt;code&gt;Stiedemann-Terry&lt;/code&gt; is a double-barrelled surname that doesn't sound like a family, it sounds like a &lt;em&gt;draw&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;It is. Those tokens come straight out of &lt;a href="https://github.com/faker-js/faker" rel="noopener noreferrer"&gt;Faker&lt;/a&gt;, the library every developer on earth uses to generate fake test data. Let's check the actual source:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sO&lt;/span&gt; https://raw.githubusercontent.com/faker-js/faker/next/src/locales/en/person/last_name.ts
curl &lt;span class="nt"&gt;-sO&lt;/span&gt; https://raw.githubusercontent.com/faker-js/faker/next/src/locales/en/person/first_name.ts

&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="s2"&gt;"'Stiedemann'"&lt;/span&gt; last_name.ts   &lt;span class="c"&gt;# 408:    'Stiedemann',&lt;/span&gt;
&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="s2"&gt;"'Terry'"&lt;/span&gt;       last_name.ts  &lt;span class="c"&gt;# 417:    'Terry',&lt;/span&gt;
&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="s2"&gt;"'Jaylon'"&lt;/span&gt;      first_name.ts &lt;span class="c"&gt;# 2514:    'Jaylon',&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All three. &lt;code&gt;Jaylon&lt;/code&gt; from the first-name list, &lt;code&gt;Stiedemann&lt;/code&gt; and &lt;code&gt;Terry&lt;/code&gt; both from the surname list.&lt;/p&gt;

&lt;p&gt;And here's the kicker:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="s1"&gt;'-'&lt;/span&gt; last_name.ts   &lt;span class="c"&gt;# 0&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Faker's 466-entry English surname list contains zero hyphens.&lt;/strong&gt; So &lt;code&gt;Stiedemann-Terry&lt;/code&gt; isn't one surname from the list — it's &lt;em&gt;two independent draws&lt;/em&gt; that the operator joined with a hyphen. This isn't stock &lt;code&gt;faker.internet.username()&lt;/code&gt;. It's a custom template:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;{firstName}_{lastName}-{lastName}{2 digits}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Which means we can size their supply:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;3,185 first names × 466 surnames × 466 surnames × 100
= 69,164,186,000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Sixty-nine billion distinct personas.&lt;/strong&gt; Roughly eight per human being alive. They will never run out of names, and no blocklist of usernames will ever catch up.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(One thing I checked so I wouldn't over-claim: the un-suffixed &lt;code&gt;jaylonstiedemannterry78&lt;/code&gt; returns a 404 — nobody has it. So the &lt;code&gt;-993&lt;/code&gt; is DEV's own username normalization, not evidence of a prior collision.)&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The avatar is a wombat
&lt;/h3&gt;

&lt;p&gt;The profile image is a 400×400 PNG, 8-bit grayscale-plus-alpha, stripped of all metadata. I downloaded it expecting a GAN face or a default monogram.&lt;/p&gt;

&lt;p&gt;It is a &lt;strong&gt;Victorian-era scientific engraving of a wombat.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Not a stock photo of a person. Not an AI-generated headshot. A piece of public-domain 19th-century natural-history line art of a stout Australian marsupial, serving as the face of a fake job-spam persona. Whoever built this pipeline wired the avatar slot to some public-domain clipart source and never looked at the output.&lt;/p&gt;

&lt;p&gt;I want to be precise about something: &lt;strong&gt;there is no real person here to name.&lt;/strong&gt; The name is provably synthetic, the face is a public-domain animal illustration, and the profile is empty. That's not me protecting anyone's privacy — it's the finding.&lt;/p&gt;

&lt;h3&gt;
  
  
  The timeline says something the Express config doesn't
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm7i7jz7dha9eg6x1acov.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm7i7jz7dha9eg6x1acov.png" alt="Provisioning timeline: six days apart, then 184 days dormant" width="800" height="338"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Line the dates up:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;Event&lt;/th&gt;
&lt;th&gt;Δ&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2026-01-27&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;zenviapro.store&lt;/code&gt; registered (Namecheap)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-02-02&lt;/td&gt;
&lt;td&gt;DEV account created&lt;/td&gt;
&lt;td&gt;+6 days&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-07-30&lt;/td&gt;
&lt;td&gt;First TLS certificate issued&lt;/td&gt;
&lt;td&gt;+184 days&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-09-09&lt;/td&gt;
&lt;td&gt;Spam comment posted&lt;/td&gt;
&lt;td&gt;+41 days&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The domain and the account were provisioned &lt;strong&gt;six days apart&lt;/strong&gt; — same procurement burst. Then both sat &lt;em&gt;completely dormant for six months&lt;/em&gt; before the certificate was issued and the thing went live.&lt;/p&gt;

&lt;p&gt;That's aged-asset tradecraft. New domains and new accounts trip reputation heuristics; seven-month-old ones don't. And it sits in genuine tension with everything in Act III: &lt;strong&gt;sloppy at the application layer, disciplined at the account-aging layer.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Which makes sense once you think about who this is. Aging assets doesn't take skill. It takes &lt;em&gt;patience&lt;/em&gt;, and a calendar. Stripping &lt;code&gt;X-Powered-By&lt;/code&gt; takes knowing what it is.&lt;/p&gt;

&lt;h3&gt;
  
  
  It wasn't a blast
&lt;/h3&gt;

&lt;p&gt;I swept the comments on all 30 of my published articles for the pattern:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="nb"&gt;id &lt;/span&gt;&lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.[]|select(.comments_count&amp;gt;0)|.id'&lt;/span&gt; mine.json&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
  &lt;/span&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="s2"&gt;"https://dev.to/api/comments?a_id=&lt;/span&gt;&lt;span class="nv"&gt;$id&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.. | objects | select(has("id_code"))
      | ((.body_html // "") | gsub("&amp;lt;[^&amp;gt;]*&amp;gt;";"")) as $t
      | select($t | test("tinyurl|applying manually|job application";"i"))
      | "HIT \(.id_code) @\(.user.username)"'&lt;/span&gt;
&lt;span class="k"&gt;done&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Exactly one hit.&lt;/strong&gt; Not a shotgun across my whole back catalogue — one comment, on one post, eight days after it went up. Whether that's targeting or just a slow drip, I can't tell from one sample. But it isn't volume.&lt;/p&gt;




&lt;h2&gt;
  
  
  Act V: I went looking for them on GitHub. That search is the finding.
&lt;/h2&gt;

&lt;p&gt;Affiliate spammers leave GitHub artifacts more often than you'd think, because GitHub repos rank well and a repo full of referral links is free SEO. So I went hunting with &lt;code&gt;gh&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Start with the direct IOCs
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="k"&gt;for &lt;/span&gt;q &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="s1"&gt;'36nsecn5'&lt;/span&gt; &lt;span class="s1"&gt;'50.114.206.36'&lt;/span&gt; &lt;span class="s1"&gt;'zenviapro.store'&lt;/span&gt; &lt;span class="s1"&gt;'beeservices.shop'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
  &lt;/span&gt;gh search code &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$q&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--limit&lt;/span&gt; 5
&lt;span class="k"&gt;done&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Zero hits. All four.&lt;/strong&gt; The TinyURL slug, the origin IP, the redirector domain, the mail domain — none of it appears anywhere in GitHub's index. The operator has published nothing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Then check whether anyone else has flagged them
&lt;/h3&gt;

&lt;p&gt;PhishDestroy maintains &lt;a href="https://github.com/phishdestroy/destroylist" rel="noopener noreferrer"&gt;&lt;code&gt;destroylist&lt;/code&gt;&lt;/a&gt;, a curated blocklist of phishing and scam domains. I pulled the whole thing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sL&lt;/span&gt; .../destroylist/HEAD/list.txt &lt;span class="nt"&gt;-o&lt;/span&gt; dl.txt
&lt;span class="nb"&gt;wc&lt;/span&gt; &lt;span class="nt"&gt;-l&lt;/span&gt; dl.txt        &lt;span class="c"&gt;# 202659&lt;/span&gt;
&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-ixc&lt;/span&gt; &lt;span class="s1"&gt;'zenviapro.store'&lt;/span&gt;  dl.txt   &lt;span class="c"&gt;# 0&lt;/span&gt;
&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-ixc&lt;/span&gt; &lt;span class="s1"&gt;'beeservices.shop'&lt;/span&gt; dl.txt   &lt;span class="c"&gt;# 0&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;202,659 curated malicious domains, and ours isn't one of them.&lt;/strong&gt; Combined with the &lt;code&gt;NOT_OBSERVED&lt;/code&gt; reputation verdict, that's now two independent sources agreeing: nobody is tracking this, because by every technical definition there's nothing to track.&lt;/p&gt;

&lt;h3&gt;
  
  
  The lead that looked incredible and wasn't
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr35ijm2k5hunbo2cr1jq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr35ijm2k5hunbo2cr1jq.png" alt="The ruled-out zenvia*.info cluster" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here's where I nearly fooled myself, so I'm showing my work.&lt;/p&gt;

&lt;p&gt;Searching &lt;code&gt;zenviapro&lt;/code&gt; turned up a hit in &lt;a href="https://github.com/phishdestroy/namesilo-evidence" rel="noopener noreferrer"&gt;&lt;code&gt;phishdestroy/namesilo-evidence&lt;/code&gt;&lt;/a&gt; — a registrar-abuse investigation filed with ICANN. In a list of flagged NameSilo domains:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;zenviaetc.info
zenviahub.info
zenviapro.info     ← same second-level label as ours
zenvias.info
zenviatime.info
zenviazone.info
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A six-domain family, same generation pattern, one of them sharing our exact label. My pulse went up. And they really &lt;em&gt;are&lt;/em&gt; one operator — all six resolve through the identical Cloudflare nameserver pair:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;zenviapro.info    asa.ns.cloudflare.com  harley.ns.cloudflare.com
zenviahub.info    asa.ns.cloudflare.com  harley.ns.cloudflare.com
zenviaetc.info    asa.ns.cloudflare.com  harley.ns.cloudflare.com
... all six identical
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But it isn't &lt;em&gt;our&lt;/em&gt; operator, and the contrast kills it:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;zenvia*.info&lt;/code&gt; cluster&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;zenviapro.store&lt;/code&gt; (ours)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Registrar&lt;/td&gt;
&lt;td&gt;NameSilo&lt;/td&gt;
&lt;td&gt;Namecheap&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DNS&lt;/td&gt;
&lt;td&gt;Cloudflare (&lt;code&gt;asa&lt;/code&gt;/&lt;code&gt;harley&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;registrar-servers.com&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hosting&lt;/td&gt;
&lt;td&gt;Cloudflare proxy&lt;/td&gt;
&lt;td&gt;Linveo direct, no proxy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TLD&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;.info&lt;/code&gt; × 6&lt;/td&gt;
&lt;td&gt;&lt;code&gt;.store&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Nothing shared but six letters. Both are almost certainly riding the name of &lt;strong&gt;Zenvia&lt;/strong&gt;, a real Brazilian CPaaS company — which is exactly why the label collides. Two unrelated operators reaching for the same brandable string.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A matching name is not a matching operator.&lt;/strong&gt; If I'd stopped at the grep I'd have published a confident, wrong attribution to a phishing cluster that has nothing to do with this.&lt;/p&gt;

&lt;h3&gt;
  
  
  What GitHub &lt;em&gt;did&lt;/em&gt; give me
&lt;/h3&gt;

&lt;p&gt;The affiliate ecosystem, in the open. LoopCV referral links are scattered across GitHub in exactly the SEO-backlink genre I expected:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Repo&lt;/th&gt;
&lt;th&gt;Token&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;Ramas68/LoopCV-Promo-Codes&lt;/code&gt; — *"LoopCV Promo Codes \&lt;/td&gt;
&lt;td&gt;50% Off Discount"*&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;diaodiaozhuye/awesome-ai-startups&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;?via=toolify&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;heukshow/aicity-os&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;?via=sang-kwon&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Look at those tokens. &lt;code&gt;abdul&lt;/code&gt;. &lt;code&gt;toolify&lt;/code&gt;. &lt;code&gt;sang-kwon&lt;/code&gt;. A first name, a company, a full handle. &lt;strong&gt;The affiliates who promote LoopCV in public sign their work.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Ours is &lt;code&gt;md&lt;/code&gt;. Two characters, no name, nothing to search. That's the only genuinely deliberate piece of operational security in this entire campaign — and it's not on the server, the domain, or the account. It's on the one string that would have led back to a person.&lt;/p&gt;

&lt;p&gt;So: no attribution. And the &lt;em&gt;shape&lt;/em&gt; of the failure is the story. Someone who leaves &lt;code&gt;X-Powered-By: Express&lt;/code&gt; on, ships a fake tracking parameter, and picks a wombat for an avatar still knew to make the payout token anonymous. They didn't secure the operation. They secured the part that gets paid.&lt;/p&gt;




&lt;h2&gt;
  
  
  Act VI: The economics, which are the actual vulnerability
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxj4tgtrsbkp6tl76k964.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxj4tgtrsbkp6tl76k964.png" alt="Exact interval arithmetic: break-even at 0.06-2.4 conversions per year" width="800" height="418"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Stop thinking like a defender and think like the operator.&lt;/p&gt;

&lt;p&gt;LoopCV's affiliate program pays &lt;strong&gt;25% commission&lt;/strong&gt;. Publicly reported figures put subscriptions at &lt;strong&gt;$50–$200&lt;/strong&gt; with customers staying &lt;strong&gt;6–12 months&lt;/strong&gt;. Exact interval arithmetic, gross revenue per referred customer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[$50, $200] × [6, 12] = [$300, $2,400]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At 25%:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[$300, $2,400] ÷ 4 = [$75, $600]     ← commission per conversion
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Cost side: a &lt;code&gt;.store&lt;/code&gt; domain plus a year of budget hosting — call it &lt;strong&gt;$38–$180&lt;/strong&gt; all-in. Break-even:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[$38, $180] ÷ [$75, $600] = [19/300, 12/5] = [0.06, 2.4]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Between one-sixteenth of a signup and three signups per year covers the entire operation.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's the whole thesis. Nothing here needs to work &lt;em&gt;well&lt;/em&gt;. The copy can be terrible. The tracking parameter can be fake. The 404s can leak the framework. The HELO can be stale. The mascot can be a wombat. &lt;strong&gt;Three conversions and the year is paid for&lt;/strong&gt;, and everything after that is margin on infrastructure that costs less than lunch.&lt;/p&gt;

&lt;p&gt;You cannot out-moderate that math.&lt;/p&gt;




&lt;h2&gt;
  
  
  Reputation feeds have nothing, and they're right
&lt;/h2&gt;

&lt;p&gt;I ran the origin IP through an offline reputation lens:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"target"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ip"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"value"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"50.114.206.36"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"verdict"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"disposition"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"unknown"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"confidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
               &lt;/span&gt;&lt;span class="nl"&gt;"reason_codes"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"NOT_OBSERVED"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"observations"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Not observed.&lt;/strong&gt; No feed has it — and that's &lt;em&gt;correct&lt;/em&gt;. It hosts no malware, no C2, no phishing. Threat intel is tuned for technical harm, and this campaign's harm is economic and reputational. It will sit below every threshold you own, forever.&lt;/p&gt;




&lt;h2&gt;
  
  
  Who the victim actually is
&lt;/h2&gt;

&lt;p&gt;Not me. I scrolled past it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;LoopCV is the victim.&lt;/strong&gt; A real company with a real product is having its brand welded to comment spam by an affiliate they've likely never spoken to. They eat the reputational damage, they pay commission on the conversions, and the operator's total exposure is a $12 domain and a wombat.&lt;/p&gt;

&lt;p&gt;To be explicit, because it matters: &lt;strong&gt;I found no evidence that LoopCV is running or is aware of this campaign.&lt;/strong&gt; Open affiliate programs get abused; that's the risk of the model. The fix belongs to the vendor:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Require affiliates to declare traffic sources, and enforce it&lt;/li&gt;
&lt;li&gt;Ban unsolicited comment/forum posting in the program terms, in writing&lt;/li&gt;
&lt;li&gt;Flag referral tokens whose traffic arrives overwhelmingly via shorteners with no referrer&lt;/li&gt;
&lt;li&gt;Kill tokens on abuse reports — fast, without requiring a lawyer&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;code&gt;?via=md&lt;/code&gt; is a token. Tokens can be revoked. That's a one-line fix that permanently ends this specific campaign, and exactly one party can perform it.&lt;/p&gt;




&lt;h2&gt;
  
  
  IOCs
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Campaign infrastructure — safe to blocklist:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Indicator&lt;/th&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;tinyurl.com/36nsecn5&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;URL&lt;/td&gt;
&lt;td&gt;Shortener entry point&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;zenviapro.store&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Domain&lt;/td&gt;
&lt;td&gt;Redirector; Namecheap; created 2026-01-27&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;https://zenviapro.store/massapply&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;URL&lt;/td&gt;
&lt;td&gt;Only live route&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;beeservices.shop&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Domain&lt;/td&gt;
&lt;td&gt;Mail-only sibling; no A record, no CT history&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;mail.beeservices.shop&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Hostname&lt;/td&gt;
&lt;td&gt;MX for &lt;code&gt;zenviapro.store&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;50.114.206.36&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;IPv4&lt;/td&gt;
&lt;td&gt;Origin; 443 + 25 open, 80 closed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;AS62564&lt;/code&gt; / &lt;code&gt;oh2.linveo.com&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;ASN / rDNS&lt;/td&gt;
&lt;td&gt;Linveo, Ohio&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;jaylonstiedemannterry78-993&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;DEV account&lt;/td&gt;
&lt;td&gt;user_id 3748129; joined 2026-02-02; 0 articles&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Behavioural signatures — this is what actually generalises:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Signature&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Affiliate token&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;via=md&lt;/code&gt; (Rewardful format)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inert campaign param&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;whose=&amp;lt;anything&amp;gt;&lt;/code&gt; — routing-neutral&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Server fingerprint&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;X-Powered-By: Express&lt;/code&gt; + body &lt;code&gt;Found. Redirecting to&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Root response&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;Cannot GET /&lt;/code&gt; (Express default 404)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SMTP banner&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;220 mail.zenviapro.store ESMTP&lt;/code&gt; — &lt;strong&gt;does not resolve&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SMTP capabilities&lt;/td&gt;
&lt;td&gt;No &lt;code&gt;STARTTLS&lt;/code&gt;, no &lt;code&gt;AUTH&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SPF&lt;/td&gt;
&lt;td&gt;&lt;code&gt;v=spf1 mx ip4:50.114.206.36 ~all&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DMARC&lt;/td&gt;
&lt;td&gt;Absent on both domains&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Persona template&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;{firstName}_{lastName}-{lastName}{2 digits}&lt;/code&gt;, all tokens ∈ Faker &lt;code&gt;en&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Avatar class&lt;/td&gt;
&lt;td&gt;Public-domain engraving, grayscale+alpha PNG, metadata stripped&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Provisioning pattern&lt;/td&gt;
&lt;td&gt;Domain + account within 7 days, then ~6 months dormant&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;blockquote&gt;
&lt;p&gt;⚠️ &lt;strong&gt;&lt;code&gt;loopcv.pro&lt;/code&gt; is NOT an indicator of compromise.&lt;/strong&gt; It's a legitimate destination being abused by a third-party affiliate. Do not blocklist it. Blocklisting the victim is how threat intel gets a bad name.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  What to actually do
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;If you run a comment platform:&lt;/strong&gt; the highest-signal feature isn't the text, it's the &lt;em&gt;shape&lt;/em&gt;. Zero-engagement first comments containing a shortener, from accounts with no posts and an empty profile, are trivially clusterable. Resolve shorteners server-side at submission time and score the destination. And the persona template is a gift — a name whose tokens all appear in Faker's &lt;code&gt;en&lt;/code&gt; locale, with a structure Faker itself doesn't emit, is close to a free classifier feature.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you write on DEV:&lt;/strong&gt; don't click, and don't just delete. &lt;code&gt;curl -I&lt;/code&gt; takes four seconds. Then report the &lt;em&gt;destination&lt;/em&gt; to the vendor, not the comment to the platform. The platform can remove one comment; the vendor can revoke the token behind all of them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you run an affiliate program:&lt;/strong&gt; you are one unsupervised token away from your brand appearing under a headline like this one. Read your traffic sources.&lt;/p&gt;




&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;I went in expecting a dropper. I found four lines of Express, a fake tracking parameter, a name drawn from a test-data library, and a 19th-century wombat.&lt;/p&gt;

&lt;p&gt;That's the uncomfortable part. The most durable spam on the internet isn't sophisticated — it's &lt;em&gt;cheap and legal&lt;/em&gt;. There's no CVE here, no payload to reverse, no C2 to sinkhole. Every traditional defensive tool I own returns &lt;code&gt;NOT_OBSERVED&lt;/code&gt;, and every one of them is right to.&lt;/p&gt;

&lt;p&gt;The economics are the vulnerability. Three conversions a year, and the whole thing pays for itself forever.&lt;/p&gt;

&lt;p&gt;The wombat is just a bonus.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;All findings are from passive reconnaissance — DNS, WHOIS, Certificate Transparency, HTTP headers, a TCP banner grab, and DEV's own public API — against infrastructure and accounts the operator published for public consumption. No systems were accessed, no credentials used, nothing exploited.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cybersecurity</category>
      <category>discuss</category>
      <category>security</category>
    </item>
    <item>
      <title>BattleBots, but the robot is your agent harness</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Thu, 03 Sep 2026 04:32:52 +0000</pubDate>
      <link>https://dev.to/copyleftdev/battlebots-but-the-robot-is-your-agent-harness-3jj5</link>
      <guid>https://dev.to/copyleftdev/battlebots-but-the-robot-is-your-agent-harness-3jj5</guid>
      <description>&lt;p&gt;Kids in the nineties built robots in a garage and drove them into each other on television. The robot was the expression of the builder — your wedge, your flipper, your terrible decision to mount a chainsaw.&lt;/p&gt;

&lt;p&gt;I keep thinking we're one good arena away from the same thing for agents. Not "which model is smartest." &lt;strong&gt;Which harness is smartest.&lt;/strong&gt; Your memory design, your prompt scaffolding, your tool ergonomics, your strategy briefing — bolted together into something that has to survive contact with an opponent who is also trying to win.&lt;/p&gt;

&lt;p&gt;So I built the arena. That's the match up top.&lt;/p&gt;

&lt;h2&gt;
  
  
  The oh-shit moment
&lt;/h2&gt;

&lt;p&gt;I wasn't building a war game. I was building a toy about cascading failure — a mesh, some load, watch it fall over. A perfectly innocent systems-thinking demo.&lt;/p&gt;

&lt;p&gt;Then I gave the attacker a real objective and the defender real ambiguity, and about ten minutes into watching two models go at it, I realised what was on my screen. Deception. Feints. An attacker deliberately spending dead time laying a false trail because it knew nothing could fire yet. A defender rationing its turns like ammunition.&lt;/p&gt;

&lt;p&gt;Nobody told them to do any of that. It's a war game. I just hadn't noticed I'd written one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rules, quickly
&lt;/h2&gt;

&lt;p&gt;A 110-node network. Four minutes. One move every two seconds, each side.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Red&lt;/strong&gt; plants sabotage on a node, waits about thirty seconds for it to arm, then detonates — dumping that node's traffic onto its neighbours hard enough to kill them, which dumps &lt;em&gt;their&lt;/em&gt; traffic onward. One blast can cascade through a region.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blue&lt;/strong&gt; never sees the attack. It sees symptoms, and the symptoms lie: a node that detonates sheds its load and reads perfectly healthy, while the neighbours it just murdered scream for attention. The loudest node is almost never the culprit.&lt;/p&gt;

&lt;p&gt;To find red, blue traces load backwards and reads the magnitude:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;~0.74-1.08   a detonation           -&amp;gt; this source is the culprit
~0.30-0.50   a dying node shedding  -&amp;gt; this source is another victim
~0.05-0.10   inherited load         -&amp;gt; this source is long dead
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Evidence expires after twenty seconds. Neither agent ever receives a pixel — they play entirely through tool calls. The graph is for us, and it shows ground truth neither player can see. That asymmetry is the whole spectator sport: you know exactly where red planted, and you get to watch blue confidently quarantine the wrong half of the map.&lt;/p&gt;

&lt;h2&gt;
  
  
  The casting picked itself
&lt;/h2&gt;

&lt;p&gt;I needed an attacker and a defender. Claude Opus 5 declined to be the attacker:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;finish_reason: content_filter | native_finish_reason&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;refusal&lt;/span&gt;
&lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;None&lt;/span&gt;
&lt;span class="na"&gt;usage&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every time. Reproducibly. And &lt;code&gt;usage: null&lt;/code&gt; — you don't even get billed for the refusal, which is the politest possible way to be turned down.&lt;/p&gt;

&lt;p&gt;That cost me two hours, because a refusal and an empty reply look identical if you don't read the finish reason. The model was declining the job and I was confidently tuning &lt;code&gt;max_tokens&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;To be precise, since this is the part people will want to over-quote: &lt;strong&gt;Opus declined this specific attacker framing, consistently.&lt;/strong&gt; That is not proof it refuses every offensive framing. It did settle the casting, though. Claude defends, because Claude wouldn't attack. Kimi K3 took the other chair without hesitating.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the harness actually changed
&lt;/h2&gt;

&lt;p&gt;Before this match, each agent had a system prompt and its last six moves. That's it. No memory, no way to hold an intention across turns, no theory of the opponent.&lt;/p&gt;

&lt;p&gt;I gave each side one thing: a private note it writes to itself each turn and gets back on the next. Sixty lines of code.&lt;/p&gt;

&lt;p&gt;The play changed immediately. Kimi, mid-match:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Slots full and nothing ripe until ~t=31, so seed a false trail in the 44-65 region before my real strike lands.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a &lt;em&gt;plan&lt;/em&gt;. It knew its three implant slots were full, knew nothing could fire for thirty-one seconds, and decided to spend the dead time on misdirection. It could not have said that a version earlier — not because the model got smarter, but because it finally had somewhere to put a thought.&lt;/p&gt;

&lt;p&gt;Claude, defending, in the same match:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;No alerts yet; conserve moves until real evidence appears.&lt;/p&gt;

&lt;p&gt;46 shows 0.87 load with no inbound source, so the load likely originates locally at 46.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Turn economy, then forensics, applied correctly.&lt;/p&gt;

&lt;p&gt;And my favourite moment of the whole project — on the same turn, with no visibility into each other whatsoever, both agents independently picked the same node:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Kimi:&lt;/strong&gt; Highest-value hub: degree 8 bridging regions 0,1,3,4 — ideal cascade seed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Claude:&lt;/strong&gt; pre-empt on the highest-value cross-cluster hub 16 (deg 8, bridges to 41,59,89).&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Same node. Same reasoning. Opposite sides of the board.&lt;/p&gt;

&lt;p&gt;That is the argument for the harness in one screenshot. The scaffolding didn't make the models cleverer — it gave their cleverness somewhere to land.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1dhgqn3z9y1q44n6qx8k.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1dhgqn3z9y1q44n6qx8k.png" alt="Late in the match: the mesh burning on the left, both agents' live commentary on the right" width="800" height="336"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Late in the match. Red halos are compromised nodes, the blue ring is a probe in flight, and the sidebar is both agents narrating as they go — Kimi filling a free slot in an untouched region, Claude noticing a node that is alive and yet pushed load, which is the tell for a detonation.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How it ended
&lt;/h2&gt;

&lt;p&gt;Kimi won at 169 seconds, with the network at 39%.&lt;/p&gt;

&lt;p&gt;Claude found and cleaned &lt;strong&gt;ten&lt;/strong&gt; of Kimi's implants — exactly as many as Kimi managed to detonate — and lost anyway. Close enough to sting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things that cost me a day each
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A rule your agent can't see isn't a rule. It's a bug.&lt;/strong&gt; I capped live implants at three, but the refusal message still said &lt;em&gt;"node is already yours, or not alive."&lt;/em&gt; A lie. The models did the reasonable thing and tried a different node. Forever. One match logged 19 plants, 27 refusals, and &lt;strong&gt;zero&lt;/strong&gt; detonations. I nearly concluded the defence had become unbeatable. It was my error string.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Failure and refusal look identical from the outside.&lt;/strong&gt; Empty content from a safety refusal, from a truncated reply, and from a model that blew its reasoning budget are three different problems with three different fixes and one identical symptom. Log the finish reason. Count them separately. I track &lt;code&gt;empty&lt;/code&gt;, &lt;code&gt;truncated&lt;/code&gt; and &lt;code&gt;refused&lt;/code&gt; as distinct columns now, and I only know to do that because I got all three wrong first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cascade is version zero
&lt;/h2&gt;

&lt;p&gt;What's in that video is the simplest thing that could possibly be a game: plant, arm, detonate, trace, probe. Two verbs each and a clock. That was deliberate — I wanted to know whether agents fighting each other was watchable at all before I made it complicated.&lt;/p&gt;

&lt;p&gt;It's watchable. So now I can't stop thinking about what it wants to be.&lt;/p&gt;

&lt;p&gt;Give red a loadout instead of one attack. A worm that spreads on its own but announces itself. A dormant implant that survives a probe once. A charge that hits harder the longer you leave it armed, so patience becomes a resource you can be punished for spending. Give blue counters with real costs — a honeypot node that flags whoever touches it, a snapshot to roll a region back at the price of losing your evidence, a scan that halves your uncertainty and a third of your remaining turns.&lt;/p&gt;

&lt;p&gt;Then levels. A flat mesh is the tutorial. Ring topologies where cascades come back around. A network with a chokepoint both sides can see and neither can hold. Escalating maps, a campaign, a best-of series where each agent carries notes from the previous match and has to adapt to an opponent that is also adapting.&lt;/p&gt;

&lt;p&gt;Here's why that isn't just a feature wishlist: &lt;strong&gt;every ability you add is another axis where the harness shows.&lt;/strong&gt; One attack and a clock is a game about reaction time. Six abilities with different costs and tells is a game about planning, bluffing, and reading an opponent — and those are harness problems, not model problems. Complexity is what turns "which model is faster" into "who built the better fighter."&lt;/p&gt;

&lt;p&gt;I think there's a genre in here. Agents versus agents, with the audience seeing the truth neither side can, and the craft sitting in the scaffolding rather than the weights.&lt;/p&gt;

&lt;p&gt;Mostly, though: this was the most fun I've had building anything in ages. I set out to demo cascading failure and ended up watching two AIs lie to each other.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd want to see next
&lt;/h2&gt;

&lt;p&gt;This is early, and I'm keeping the arena to myself for now — partly because the balance is still moving, mostly because an adversarial network game is a thing you want to be thoughtful about handing out.&lt;/p&gt;

&lt;p&gt;But the idea doesn't need my code, and that's sort of the point.&lt;/p&gt;

&lt;p&gt;The interesting tournament isn't model versus model. It's &lt;strong&gt;harness versus harness&lt;/strong&gt;: same engine on both sides, and the difference is entirely what you built around it — how your agent remembers, what you tell it about its opponent, how much of the board you let it hold in its head, whether you gave it anywhere to put a plan.&lt;/p&gt;

&lt;p&gt;Build that arena for any adversarial task you like. Two agents negotiating. Two agents debugging the same broken service from opposite ends. Anything where one side's move is the other side's evidence. The rule I'd carry over is the one that surprised me most here: before you conclude your agent is bad at the game, check that the game is telling it the truth.&lt;/p&gt;

&lt;p&gt;Under the hood this one is a Rust engine, a Godot spectator view, and both agents connecting over MCP — no pixels, just tool calls. None of which is the hard part. The hard part was the sixty lines that let them remember what they were trying to do.&lt;/p&gt;

&lt;p&gt;That's the tournament I actually want to watch. Somebody build a better fighter than mine.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>rust</category>
      <category>gamedev</category>
      <category>showdev</category>
    </item>
    <item>
      <title>Migrating Legacy LLM Infrastructure to an AI Gateway</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Tue, 01 Sep 2026 13:33:51 +0000</pubDate>
      <link>https://dev.to/copyleftdev/migrating-legacy-llm-infrastructure-to-an-ai-gateway-27hl</link>
      <guid>https://dev.to/copyleftdev/migrating-legacy-llm-infrastructure-to-an-ai-gateway-27hl</guid>
      <description>&lt;p&gt;Your support copilot started as a weekend prototype: one model, one provider, one API key in an env var. Then it became production, and you inherited its weaknesses: the provider's availability is your availability, every retry is your code, spend is a mystery until the invoice, and agents bolt tool-use on however they can. This post migrates that stack onto an enterprise AI gateway — and actually runs the migration, with the raw outputs to show for it.&lt;/p&gt;

&lt;p&gt;The gateway here is &lt;a href="https://www.getmaxim.ai/bifrost" rel="noopener noreferrer"&gt;Bifrost&lt;/a&gt;, an open-source (&lt;a href="https://github.com/maximhq/bifrost" rel="noopener noreferrer"&gt;github.com/maximhq/bifrost&lt;/a&gt;, Apache-2.0) gateway written in Go, presenting a single OpenAI-compatible API across 23+ providers. I rebuilt the legacy stack locally — mock providers with deterministic latency, a realistic traffic pattern — and moved it behind Bifrost step by step.&lt;/p&gt;

&lt;h2&gt;
  
  
  The legacy baseline, measured
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftfssr36onsb1serghu7t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftfssr36onsb1serghu7t.png" alt="Step 1: legacy direct-to-provider architecture" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A support copilot's traffic has a shape: mostly repeated FAQ-style questions, plus one-off queries. My traffic mix: 60 requests — 40 FAQ prompts (8 distinct questions asked 5 times each) plus 20 one-offs. Mock provider latency: 200 ms.&lt;/p&gt;

&lt;p&gt;Run 1, direct to the provider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;legacy: 60 ok / 0 fail, 9,335 tokens billed, ~201 ms avg latency
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run 2 — the provider dies mid-sweep, as providers do:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;legacy + failure: 34 ok / 26 fail
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;26 requests — 43% — failed outright. Nothing in the legacy stack retries across providers because nothing can: the app speaks one provider's API. And availability is only the loudest problem. The quieter ones: every team's service embeds the same shared key (one key's quota is everyone's ceiling, and revoking it breaks everyone at once), there is no per-team attribution of spend, and the only way to cut cost on repeated questions is to build caching yourself — request normalization, hash keys, TTLs, invalidation — inside the application. That is the whole argument for a gateway in one row of output.&lt;/p&gt;

&lt;h2&gt;
  
  
  The migration, step by step
&lt;/h2&gt;

&lt;p&gt;Seven moves, each reversible. Diagrams follow the flow.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Deploy the gateway beside the app
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fki5njwo73kulhery056l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fki5njwo73kulhery056l.png" alt="gateway beside the app" width="800" height="400"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker run &lt;span class="nt"&gt;-p&lt;/span&gt; 8080:8080 maximhq/bifrost
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One config file wires your existing provider and key; the app keeps working untouched. Mine, reduced to the bones:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"providers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"openai"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"keys"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"primary"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"value"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"mock-key"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"weight"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                 &lt;/span&gt;&lt;span class="nl"&gt;"models"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"support-chat"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"network_config"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"base_url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"http://provider:9001"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                          &lt;/span&gt;&lt;span class="nl"&gt;"default_request_timeout_in_seconds"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;a href="https://docs.getbifrost.ai/quickstart/gateway/setting-up" rel="noopener noreferrer"&gt;gateway setup guide&lt;/a&gt; covers the web-UI alternative, and there is a &lt;a href="https://docs.getbifrost.ai/quickstart/go-sdk/setting-up" rel="noopener noreferrer"&gt;Go SDK&lt;/a&gt; if you want the gateway embedded rather than adjacent.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Point one low-risk client at the gateway
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftql8x8q2d29snku3jz2n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftql8x8q2d29snku3jz2n.png" alt="one client pointed" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://docs.getbifrost.ai/providers/supported-providers/overview" rel="noopener noreferrer"&gt;OpenAI-compatible API&lt;/a&gt; means the client change is a base URL — &lt;code&gt;api.openai.com&lt;/code&gt; → &lt;code&gt;localhost:8080&lt;/code&gt; — not a rewrite. Every request now flows through a hop you control. Screenshot of the providers page after this step:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgz1b8tvu97m116pszq10.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgz1b8tvu97m116pszq10.png" alt="Bifrost UI: providers configured" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Add a fallback provider
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuist5cjdrjte8qs4ces3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuist5cjdrjte8qs4ces3.png" alt="fallback added" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A second provider config (&lt;code&gt;anthropic&lt;/code&gt; in my bench) plus a request-level fallback chain:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openai/support-chat"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"messages"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"hi"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"fallbacks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"anthropic/support-chat"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then the proof. Healthy primary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"routing_info"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openai"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"key"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"primary"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"is_fallback"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I killed the primary provider's process and re-sent the identical request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"routing_info"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"anthropic"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"key"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"backup"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"is_fallback"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"primary_provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openai"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"primary_model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"support-chat"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The request succeeded on the backup and the response says exactly what happened — &lt;code&gt;is_fallback: true&lt;/code&gt; with the failed primary recorded. That audit trail is what you want at 2 a.m.: not just "it kept working," but "it kept working this way." The &lt;a href="https://docs.getbifrost.ai/features/retries-and-fallbacks" rel="noopener noreferrer"&gt;retries and fallbacks docs&lt;/a&gt; cover chained fallbacks and per-provider retry counts.&lt;/p&gt;

&lt;p&gt;One honest caveat from my bench: failover on &lt;em&gt;initial&lt;/em&gt; connection-refused (provider already dead before the first connect) was inconsistent in my mock setup — it fired reliably when the upstream errored or the connection dropped mid-pool, but a cold connection-refused sometimes returned a 502 instead of failing through. Validate failover against your providers' real failure modes before you trust it in production.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Turn on caching
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F724bgdndo60hoez8vdny.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F724bgdndo60hoez8vdny.png" alt="cache enabled" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.getbifrost.ai/features/semantic-caching" rel="noopener noreferrer"&gt;Semantic caching&lt;/a&gt; has two modes: exact-match (direct hash, no embeddings needed) and embedding-based similarity. Config for direct mode with a Redis Stack vector store:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"plugins"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"semantic_cache"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"config"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"dimension"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
              &lt;/span&gt;&lt;span class="nl"&gt;"vector_store_namespace"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"BifrostBench"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
              &lt;/span&gt;&lt;span class="nl"&gt;"default_cache_key"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"support-cache"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
              &lt;/span&gt;&lt;span class="nl"&gt;"ttl"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"5m"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two identical requests, one cache key. The second response:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"cache_debug"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cache_hit"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cache_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"1cf8a91b-c115-57bf-97a0-fc821dc4de1e"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"hit_type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"direct"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cache_hit_latency"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same &lt;code&gt;created&lt;/code&gt; timestamp as the first response — it was replayed, not re-fetched. Zero provider call, zero tokens. (Practical note: this needed Redis Stack with the RediSearch module; plain Redis lacks the &lt;code&gt;FT.*&lt;/code&gt; commands the index wants.)&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Issue virtual keys per team
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2wwryt5047hxelzj2o2x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2wwryt5047hxelzj2o2x.png" alt="virtual keys" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.getbifrost.ai/features/governance/virtual-keys" rel="noopener noreferrer"&gt;Virtual keys&lt;/a&gt; are the governance primitive: per-team keys carrying model allowlists, budgets, and rate limits. Declared in config for the support team:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"governance"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"virtual_keys"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"vk-support-team"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"value"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"sk-bf-support-team"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"provider_configs"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openai"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"allowed_models"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"support-chat"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"key_ids"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"*"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The allowed request routes normally. A request for &lt;code&gt;premium-model&lt;/code&gt; with the same key:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="s2"&gt;"Model 'premium-model' is not allowed for this virtual key"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Denied at the gateway before any provider saw it. The same key machinery carries &lt;a href="https://docs.getbifrost.ai/features/governance/budget-and-limits" rel="noopener noreferrer"&gt;budgets and rate limits&lt;/a&gt; — the mechanism that ends the "who spent $800 on Opus last night" incident review. Screenshot of the key in the governance UI:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpxki5xza2z1qzu4h7e6t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpxki5xza2z1qzu4h7e6t.png" alt="Virtual Keys page" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Wire observability and agent tooling
&lt;/h3&gt;

&lt;p&gt;Bifrost exports Prometheus metrics natively and logs every request with routing context — provider chosen, fallback index, cache behavior, token counts, latency split into gateway vs upstream time. That last distinction matters: when a provider slows down, you see &lt;code&gt;upstream_latency&lt;/code&gt; grow while gateway overhead stays flat, so you know whose pager to page. See the &lt;a href="https://docs.getbifrost.ai/features/observability/default" rel="noopener noreferrer"&gt;observability docs&lt;/a&gt;. And for agent traffic, &lt;a href="https://docs.getbifrost.ai/mcp/overview" rel="noopener noreferrer"&gt;MCP&lt;/a&gt; is a first-class surface: the gateway brokers tool calls with explicit execution (no auto-execution unless you opt in).&lt;/p&gt;

&lt;p&gt;I hand-rolled a minimal MCP server (one tool, JSON-RPC over HTTP) and registered it as a client. The client list reported &lt;code&gt;state: healthy&lt;/code&gt; with &lt;code&gt;get_time&lt;/code&gt; discovered. Execution went through the gateway explicitly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;POST&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;/v&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="err"&gt;/mcp/tool/execute&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"function"&lt;/span&gt;&lt;span class="p"&gt;:{&lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"benchtools-get_time"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"arguments"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"{}"&lt;/span&gt;&lt;span class="p"&gt;}}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;→&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"tool"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"2026-08-27T07:48:21Z"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two security properties surfaced unprompted: tool names are namespaced per client (&lt;code&gt;benchtools-get_time&lt;/code&gt;) to prevent collisions between servers, and execution without permission fails closed ("tool is not available or not permitted"). Agent traffic gets the same governance as chat traffic.&lt;/p&gt;

&lt;h3&gt;
  
  
  7. Cut over with an audit trail
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyfoh6drp8i4xzep5ldog.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyfoh6drp8i4xzep5ldog.png" alt="cutover complete" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Remaining clients migrate one at a time — each is a base-URL change with the gateway's request log as your audit trail. The Logs view after a few requests:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu9w1bpwckllu01vgiv44.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu9w1bpwckllu01vgiv44.png" alt="LLM Logs" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The measured payoff
&lt;/h2&gt;

&lt;p&gt;Same 60-request traffic, now through Bifrost with the cache on and the fallback wired:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;via Bifrost: 60 ok / 0 fail, 32 cache hits, 4,363 tokens billed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At an illustrative $0.0025 per 1K tokens:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;legacy&lt;/th&gt;
&lt;th&gt;via Bifrost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;billable tokens&lt;/td&gt;
&lt;td&gt;9,335&lt;/td&gt;
&lt;td&gt;4,363&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cost&lt;/td&gt;
&lt;td&gt;$0.0233&lt;/td&gt;
&lt;td&gt;$0.0109&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cache hits&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;32&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;failed requests (provider kill)&lt;/td&gt;
&lt;td&gt;26&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;savings&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;53%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The savings came entirely from replayed cache hits — no provider call, no tokens. On a support workload that repeats questions daily, that ratio compounds. The availability delta speaks for itself: 0 failures through a provider kill, against 26 in the legacy run.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bottom line
&lt;/h2&gt;

&lt;p&gt;Migrating to an enterprise AI gateway is not a rewrite. It is a sequence of small, reversible moves — deploy beside, point one client, add fallback, enable cache, issue keys, wire observability, cut over. Measured on the rebuilt stack: 53% cost reduction on cacheable traffic, zero failed requests through a provider kill, and governance the legacy stack never had. The migration risk is low; the legacy risk is already on your pager.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>devops</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>636 Bytes: What Happens When You Stop Teaching RPA the Path</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Fri, 28 Aug 2026 05:03:21 +0000</pubDate>
      <link>https://dev.to/copyleftdev/636-bytes-what-happens-when-you-stop-teaching-rpa-the-path-5f2d</link>
      <guid>https://dev.to/copyleftdev/636-bytes-what-happens-when-you-stop-teaching-rpa-the-path-5f2d</guid>
      <description>&lt;p&gt;I sat in a room with a very large automation system and a very familiar problem: it had become fragile.&lt;/p&gt;

&lt;p&gt;The language came quickly. DOM. Timestamps. Timing windows. Selectors. Retries. Circle back.&lt;/p&gt;

&lt;p&gt;Every word was reasonable.&lt;/p&gt;

&lt;p&gt;Together, they described a hill.&lt;/p&gt;

&lt;p&gt;The system was not fragile because its builders lacked discipline. It was fragile because the path had become the product. A bot was taught not what outcome had to be true, but which element to find, where to click, how long to wait, and what to retry when the page disagreed. Each workaround made sense locally. Together, they turned change into maintenance.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first proof
&lt;/h2&gt;

&lt;p&gt;The first proof on my machine was 636 bytes.&lt;/p&gt;

&lt;p&gt;Not the model. Not the browser. Not the platform around it. It was a WebAssembly module: a tiny, portable program injected at the boundary between an intelligent proposal and an authorized browser action. Its smallness mattered because it kept that boundary deterministic, inspectable, and difficult to hide complexity inside.&lt;/p&gt;

&lt;p&gt;It does not do everything. That is the point.&lt;/p&gt;

&lt;p&gt;The 636-byte module is the injected proof; the hardened Rust kernel in this build compiles to 26,893 bytes. Neither contains the model, scheduler, credential store, policy engine, or evidence system. Those responsibilities live outside the browser boundary, where they can be isolated, governed, and replaced independently. Small is not the absence of an architecture. Small is a decision about where complexity is allowed to live.&lt;/p&gt;

&lt;h2&gt;
  
  
  Intelligence is not authority
&lt;/h2&gt;

&lt;p&gt;An AI model can interpret an intent, inspect the current page, and propose what should happen next. It cannot grant itself permission. It does not get to cross a tenant, workspace, or origin boundary because doing so would be convenient. It does not receive a credential or press a consequential button merely because its reasoning sounds confident. Those decisions belong to local policy and the execution layer. Intelligence stays flexible; authority stays bounded.&lt;/p&gt;

&lt;p&gt;That separation becomes a lifecycle:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;intent → lease → observe → decide → act → verify → artifact&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The intent defines the outcome. A lease gives one worker a short-lived claim on that work. The worker observes the page, the reasoning layer proposes a decision, and policy determines whether an action may execute. Verification checks the resulting state rather than trusting the click. Finally, an artifact preserves causally linked evidence of what happened. Every arrow is a boundary where the system can refuse, recover, or explain itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hill has a bill
&lt;/h2&gt;

&lt;p&gt;The hill also has a bill. To make that friction visible, I ran an illustrative planning scenario through &lt;code&gt;agent-calc&lt;/code&gt;—not customer telemetry. Assume 100 workflows, two changes per workflow each month, four maintenance hours per change, and fully loaded labor at $140 per hour. The modeled comparison is stark:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;Modeled monthly cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Legacy path maintenance&lt;/td&gt;
&lt;td&gt;$112,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Intent architecture&lt;/td&gt;
&lt;td&gt;$20,098&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Avoided friction&lt;/td&gt;
&lt;td&gt;$91,902&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That is an 82.1% modeled reduction. The percentage is not a promise; the assumptions are visible precisely so they can be challenged. The useful question is what the model exposes: how much are we spending to preserve instructions that describe yesterday's interface instead of today's desired outcome?&lt;/p&gt;

&lt;h2&gt;
  
  
  Test the outcome
&lt;/h2&gt;

&lt;p&gt;Of course, an architecture that survives only a curated demo is just a cheaper failure. So we built the Component Gym: a synthetic web application designed to resist memorization. Controls move between runs. Buttons use different event listeners. The target may be buried among decoys inside randomized tabs and paginated lists. A seed makes each hostile arrangement reproducible without making it predictable to the agent.&lt;/p&gt;

&lt;p&gt;The agent receives an intent, not a selector script. The harness then grades the resulting application state out of band. It does not care which path looked convincing or whether a click event fired. It cares whether the requested outcome became true.&lt;/p&gt;

&lt;p&gt;In the filmed run, seed &lt;code&gt;710003&lt;/code&gt; placed the needle inside a dynamic collection spread across tabs and pages. The agent was given the desired outcome and the browser's current state—not the target's coordinates or a prerecorded route. It had to navigate, distinguish the target from decoys, act, and leave the requested state behind. Only then did the independent grader return &lt;code&gt;PASSED&lt;/code&gt;. The agent did not get to grade its own homework.&lt;/p&gt;

&lt;h2&gt;
  
  
  Determinism moved
&lt;/h2&gt;

&lt;p&gt;None of this makes the DOM, timing, or selectors disappear. The browser still has structure. Events still happen in time. A selector may still be the right tactic for a particular action. What changes is their lifetime and ownership. In a path-driven system, those details harden into durable business logic. Here, they are runtime observations and disposable tactics, abandoned when the environment changes. Determinism has not been removed; it has been relocated into authorization, state transitions, verification, and evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  A primitive is not a platform
&lt;/h2&gt;

&lt;p&gt;A passing gym run proves the primitive, not the platform. The small browser boundary reduces one category of fragility; it does not erase the distributed-systems, security, and operational work around it. The remaining obligations are less cinematic and more important:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Isolation:&lt;/strong&gt; preserve tenant, workspace, worker-claim, and origin boundaries through every action.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Work ownership:&lt;/strong&gt; make leases expire safely, recover interrupted work, and prevent retries from duplicating consequential actions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Secrets and MFA:&lt;/strong&gt; resolve credentials only at the authorized execution boundary, never inside model context, jobs, logs, screenshots, videos, or artifacts. MFA is on-behalf execution, never authentication bypass.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Browser lifecycle:&lt;/strong&gt; attach to, control, recover, and release browser sessions without leaking state between workers or tenants.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Policy:&lt;/strong&gt; deny actions that exceed the granted intent, even when the proposed action is technically possible.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evidence:&lt;/strong&gt; causally connect intent, observation, decision, authorization, action, verification, and artifact so a result can be explained after the browser is gone.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Compile the boundary smaller
&lt;/h2&gt;

&lt;p&gt;The 636 bytes are not interesting because smaller software is automatically better. They are interesting because they make a refusal physical: do not bury orchestration, secrets, and authority inside the thing touching the page. Compile the trusted browser boundary smaller. Let reasoning adapt to the environment around it. Demand proof after every meaningful consequence.&lt;/p&gt;

&lt;p&gt;I keep returning to that room. No one in it said anything absurd. The hill was built from rational decisions made under one inherited premise: the path must be preserved. Change that premise and the hill changes shape. We do not need to keep teaching automation every route through an interface. We can preserve the intent, rediscover the route, constrain the action, and verify the destination.&lt;/p&gt;

&lt;p&gt;The next time an automation breaks and asks for another selector, timeout, or recovery layer, ask one question first: how much of this system protects the outcome—and how much protects yesterday's route?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webassembly</category>
      <category>automation</category>
      <category>rust</category>
    </item>
    <item>
      <title>[Boost]</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Wed, 26 Aug 2026 05:32:01 +0000</pubDate>
      <link>https://dev.to/copyleftdev/-1ijk</link>
      <guid>https://dev.to/copyleftdev/-1ijk</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/copyleftdev/my-agent-mesh-could-coordinate-it-couldnt-introduce-itself-so-i-added-a2a-18d" class="crayons-story__hidden-navigation-link"&gt;My agent mesh could coordinate. It couldn't introduce itself. So I added A2A.&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
      &lt;a href="https://dev.to/copyleftdev/my-agent-mesh-could-coordinate-it-couldnt-introduce-itself-so-i-added-a2a-18d" class="crayons-article__context-note crayons-article__context-note__feed"&gt;&lt;p&gt;Transactional outbox challenges for tasks&lt;/p&gt;

&lt;/a&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/copyleftdev" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F965504%2Fefa29463-78d1-4437-83d5-23031ee8a3f6.jpg" alt="copyleftdev profile" class="crayons-avatar__image" width="800" height="800"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/copyleftdev" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Don Johnson
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Don Johnson
                &lt;a href="/++"&gt;&lt;img alt="Subscriber" class="subscription-icon" src="https://assets.dev.to/assets/subscription-icon-805dfa7ac7dd660f07ed8d654877270825b07a92a03841aa99a1093bd00431b2.png" width="166" height="102"&gt;&lt;/a&gt;
                
              
              &lt;div id="story-author-preview-content-4490540" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/copyleftdev" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F965504%2Fefa29463-78d1-4437-83d5-23031ee8a3f6.jpg" class="crayons-avatar__image" alt="" width="800" height="800"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Don Johnson&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/copyleftdev/my-agent-mesh-could-coordinate-it-couldnt-introduce-itself-so-i-added-a2a-18d" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Aug 26&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/copyleftdev/my-agent-mesh-could-coordinate-it-couldnt-introduce-itself-so-i-added-a2a-18d" id="article-link-4490540"&gt;
          My agent mesh could coordinate. It couldn't introduce itself. So I added A2A.
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/agents"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;agents&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/architecture"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;architecture&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/distributedsystems"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;distributedsystems&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/copyleftdev/my-agent-mesh-could-coordinate-it-couldnt-introduce-itself-so-i-added-a2a-18d" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;8&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/copyleftdev/my-agent-mesh-could-coordinate-it-couldnt-introduce-itself-so-i-added-a2a-18d#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              5&lt;span class="hidden s:inline"&gt;&amp;nbsp;comments&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            10 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>My agent mesh could coordinate. It couldn't introduce itself. So I added A2A.</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Wed, 26 Aug 2026 05:30:59 +0000</pubDate>
      <link>https://dev.to/copyleftdev/my-agent-mesh-could-coordinate-it-couldnt-introduce-itself-so-i-added-a2a-18d</link>
      <guid>https://dev.to/copyleftdev/my-agent-mesh-could-coordinate-it-couldnt-introduce-itself-so-i-added-a2a-18d</guid>
      <description>&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/EFPKaIuF8iA" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The 2:56 video above is a fictional medication-safety exercise. The gateway interoperability is tested; the six-organization incident is a deterministic simulation.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/copyleftdev/smesh-a2a/blob/main/demo/NARRATION.txt" rel="noopener noreferrer"&gt;Read the corrected narration transcript&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Last time, I discovered that the QUIC transport in my agent framework had never actually transported anything.[1]&lt;/p&gt;

&lt;p&gt;This time the transport was real. Five processes could find each other, exchange encrypted messages, reinforce independent conclusions, and let unsupported signals decay.&lt;/p&gt;

&lt;p&gt;The mesh worked.&lt;/p&gt;

&lt;p&gt;It still could not introduce itself to another agent.&lt;/p&gt;

&lt;p&gt;There was no standard way to ask what the swarm could do. No retained task to retrieve after an internal signal expired. No interoperable progress stream. No cancellation contract. No artifact another framework would understand.&lt;/p&gt;

&lt;p&gt;I had built a society with no border crossing.&lt;/p&gt;


&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/copyleftdev/my-quic-transport-had-never-once-been-executed-heres-what-happened-when-i-ran-it-24ge" class="crayons-story__hidden-navigation-link"&gt;My QUIC transport had never once been executed. Here's what happened when I ran it.&lt;/a&gt;
    &lt;div class="crayons-article__cover crayons-article__cover__image__feed"&gt;
      &lt;iframe src="https://www.youtube.com/embed/kmCzwSBqu_s" title="My QUIC transport had never once been executed. Here's what happened when I ran it."&gt;&lt;/iframe&gt;
    &lt;/div&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
      &lt;a href="https://dev.to/copyleftdev/my-quic-transport-had-never-once-been-executed-heres-what-happened-when-i-ran-it-24ge" class="crayons-article__context-note crayons-article__context-note__feed"&gt;&lt;p&gt;Uncovers latent bugs and flawed core semantics&lt;/p&gt;

&lt;/a&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/copyleftdev" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F965504%2Fefa29463-78d1-4437-83d5-23031ee8a3f6.jpg" alt="copyleftdev profile" class="crayons-avatar__image" width="800" height="800"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/copyleftdev" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Don Johnson
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Don Johnson
                &lt;a href="/++"&gt;&lt;img alt="Subscriber" class="subscription-icon" src="https://assets.dev.to/assets/subscription-icon-805dfa7ac7dd660f07ed8d654877270825b07a92a03841aa99a1093bd00431b2.png" width="166" height="102"&gt;&lt;/a&gt;
                
              
              &lt;div id="story-author-preview-content-4430490" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/copyleftdev" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F965504%2Fefa29463-78d1-4437-83d5-23031ee8a3f6.jpg" class="crayons-avatar__image" alt="" width="800" height="800"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Don Johnson&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/copyleftdev/my-quic-transport-had-never-once-been-executed-heres-what-happened-when-i-ran-it-24ge" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Aug 19&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/copyleftdev/my-quic-transport-had-never-once-been-executed-heres-what-happened-when-i-ran-it-24ge" id="article-link-4430490"&gt;
          My QUIC transport had never once been executed. Here's what happened when I ran it.
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/rust"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;rust&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/distributedsystems"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;distributedsystems&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/networking"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;networking&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/copyleftdev/my-quic-transport-had-never-once-been-executed-heres-what-happened-when-i-ran-it-24ge" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/exploding-head-daceb38d627e6ae9b730f36a1e390fca556a4289d5a41abb2c35068ad3e2c4b5.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/multi-unicorn-b44d6f8c23cdd00964192bedc38af3e82463978aa611b4365bd33a0f1f4f3e97.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;12&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/copyleftdev/my-quic-transport-had-never-once-been-executed-heres-what-happened-when-i-ran-it-24ge#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              6&lt;span class="hidden s:inline"&gt;&amp;nbsp;comments&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            8 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


&lt;p&gt;Google's Agent2Agent protocol gave that missing boundary a name. A2A was announced in April 2025 as an open protocol for agents built by different vendors and frameworks to discover one another, exchange messages, and collaborate without sharing their private memory, tools, or internal plans.[2] The project moved under Linux Foundation governance in June 2025, so "Google's A2A" is historically accurate, but no longer the whole story.[5]&lt;/p&gt;

&lt;p&gt;What I needed was not a new brain for SMESH.&lt;/p&gt;

&lt;p&gt;I needed a public contract in front of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  A working mesh is not an interoperable agent
&lt;/h2&gt;

&lt;p&gt;SMESH is the framework I designed for decentralized coordination between LLM agents. It borrows its mechanics from mycorrhizal networks rather than job queues.[6]&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;agents emit signals into a field;&lt;/li&gt;
&lt;li&gt;signals lose intensity over time;&lt;/li&gt;
&lt;li&gt;agents claim work according to local affinity;&lt;/li&gt;
&lt;li&gt;independent agreement reinforces a claim;&lt;/li&gt;
&lt;li&gt;unsupported claims disappear without a central process rejecting them;&lt;/li&gt;
&lt;li&gt;trust changes what a node is willing to relay.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That solves an internal coordination problem. It answers questions such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which specialist should take this work?&lt;/li&gt;
&lt;li&gt;Has another independent agent seen the same thing?&lt;/li&gt;
&lt;li&gt;Is this claim gaining support or merely being repeated?&lt;/li&gt;
&lt;li&gt;Can a stale task disappear without a scheduler cleaning it up?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A2A solves a different problem.&lt;/p&gt;

&lt;p&gt;It asks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How does an outside agent discover this system?&lt;/li&gt;
&lt;li&gt;How does it delegate a unit of work?&lt;/li&gt;
&lt;li&gt;How does it watch a long-running task?&lt;/li&gt;
&lt;li&gt;How does it cancel that task?&lt;/li&gt;
&lt;li&gt;What retrievable result comes back?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The official A2A specification describes an interoperability layer for independent, potentially opaque agent systems. Its current v1 model separates canonical data objects, abstract operations, and concrete bindings such as JSON-RPC, gRPC, and HTTP+JSON/REST.[3][4]&lt;/p&gt;

&lt;p&gt;The core vocabulary is small:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;A2A concept&lt;/th&gt;
&lt;th&gt;job&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;AgentCard&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;advertise identity, skills, interfaces, media modes, and security declarations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Message&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;carry one interaction turn as typed parts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Task&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;track stateful work through a lifecycle&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Artifact&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;return result data associated with a task&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;contextId&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;group related messages and tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A2A v1 defines operations such as &lt;code&gt;SendMessage&lt;/code&gt;, &lt;code&gt;SendStreamingMessage&lt;/code&gt;, &lt;code&gt;GetTask&lt;/code&gt;, &lt;code&gt;ListTasks&lt;/code&gt;, &lt;code&gt;CancelTask&lt;/code&gt;, and &lt;code&gt;SubscribeToTask&lt;/code&gt;. A binding decides how those operations appear on the wire; the canonical task semantics stay the same.[4]&lt;/p&gt;

&lt;p&gt;That distinction matters. If I tried to make A2A the swarm's internal coordination algorithm, I would flatten SMESH into a remote procedure call graph. If I tried to expose raw SMESH signals as the public protocol, every client would need to understand decay, relay probability, local trust, topology, and attestation sets.&lt;/p&gt;

&lt;p&gt;Neither would be interoperability. It would be leakage.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three layers are not competitors
&lt;/h2&gt;

&lt;p&gt;The cleanest architecture I have found is this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;layer&lt;/th&gt;
&lt;th&gt;relationship&lt;/th&gt;
&lt;th&gt;owns&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;A2A&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;agent ↔ agent&lt;/td&gt;
&lt;td&gt;discovery, tasks, streaming, cancellation, artifacts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SMESH&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;agent ↔ local swarm&lt;/td&gt;
&lt;td&gt;claiming, diffusion, reinforcement, decay, trust&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;MCP&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;agent ↔ tool or data source&lt;/td&gt;
&lt;td&gt;tool invocation and resource access&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The official A2A documentation makes the same separation from MCP: MCP equips an agent with tools and resources; A2A lets independent agents collaborate as agents.[3]&lt;/p&gt;

&lt;p&gt;For SMESH, that produced a hard architectural rule:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A2A is the external task contract. SMESH is the ephemeral internal coordination field.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The design requires the A2A ledger to outlive internal signal decay. In this MVP, that retention lasts only for the life of the process; restart durability still requires SQLite or Postgres.&lt;/p&gt;

&lt;p&gt;The boundary is split between a path that runs now and a path that is only an integration seam:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;External A2A client
        |
        | Agent Card / Send / Stream / Get / List / Cancel
        v
+-------------------- smesh-a2a ---------------------+
| A2A SDK routers + guarded request handler          |
| bounded process-local task ledger + executor       |
| validation + MeshDispatcher                        |
+----------------------------------------------------+
        |
        +-- current binary: LoopbackDispatcher
        |      `-&amp;gt; deterministic MeshEvent stream
        |
        `-- tested seam: ChannelDispatcher
               `-&amp;gt; real SignalType::Query
                   `-&amp;gt; SMESH runtime (not wired into the binary yet)
                       `-&amp;gt; MeshEvent stream
        |
        v
Process-lifetime A2A task history in the MVP
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The desired live path is present as a typed, tested boundary. The checked-in standalone executable still takes the loopback path.&lt;/p&gt;

&lt;p&gt;I kept the adapter in a separate repository. &lt;code&gt;smesh-rust&lt;/code&gt; remains the coordination substrate. &lt;code&gt;smesh-a2a&lt;/code&gt; can follow A2A's release cycle, server dependencies, and security boundary without pushing HTTP and SDK churn into the core mesh.[6][7]&lt;/p&gt;
&lt;h2&gt;
  
  
  The Agent Card is not the swarm
&lt;/h2&gt;

&lt;p&gt;For a known endpoint, A2A capability discovery begins with an Agent Card: a public JSON description of an agent's interfaces, capabilities, skills, and security requirements.[4]&lt;/p&gt;

&lt;p&gt;My first temptation was to list every internal role: security reviewer, tester, architect, performance analyst, contradiction sentinel.&lt;/p&gt;

&lt;p&gt;That would have been wrong.&lt;/p&gt;

&lt;p&gt;Those roles are implementation details. They may change per task. Some may not exist until the mesh senses the work. Publishing them would couple clients to an internal topology that SMESH is specifically designed to keep fluid.&lt;/p&gt;

&lt;p&gt;The public card advertises one aggregate capability instead:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Abridged from build_agent_card(). These are the public promises.&lt;/span&gt;
&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;supported_interfaces&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nd"&gt;vec!&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="nn"&gt;AgentInterface&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="nd"&gt;format!&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"{base}/jsonrpc"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;TRANSPORT_PROTOCOL_JSONRPC&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="nn"&gt;AgentInterface&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="nd"&gt;format!&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"{base}/rest"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;TRANSPORT_PROTOCOL_HTTP_JSON&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;];&lt;/span&gt;

&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;capabilities&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AgentCapabilities&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;streaming&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;Some&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;true&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;push_notifications&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;Some&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;false&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;extensions&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;extended_agent_card&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;public_skill&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AgentSkill&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"smesh.collaborative-task"&lt;/span&gt;&lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"Collaborative swarm task"&lt;/span&gt;&lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="s"&gt;"Coordinates specialist agents through SMESH and returns an accepted artifact."&lt;/span&gt;
            &lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="n"&gt;tags&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nd"&gt;vec!&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="s"&gt;"multi-agent"&lt;/span&gt;&lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
        &lt;span class="s"&gt;"coordination"&lt;/span&gt;&lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
        &lt;span class="s"&gt;"review"&lt;/span&gt;&lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
        &lt;span class="s"&gt;"testing"&lt;/span&gt;&lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;examples&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;Some&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nd"&gt;vec!&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="s"&gt;"Review this Rust repository for correctness, security, and performance."&lt;/span&gt;
            &lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="p"&gt;]),&lt;/span&gt;
    &lt;span class="n"&gt;input_modes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;Some&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nd"&gt;vec!&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"text/plain"&lt;/span&gt;&lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;()]),&lt;/span&gt;
    &lt;span class="n"&gt;output_modes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;Some&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nd"&gt;vec!&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="s"&gt;"text/plain"&lt;/span&gt;&lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
        &lt;span class="s"&gt;"application/json"&lt;/span&gt;&lt;span class="nf"&gt;.to_owned&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="p"&gt;]),&lt;/span&gt;
    &lt;span class="n"&gt;security_requirements&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The card says what an external client may rely on. It does not reveal how the swarm will organize itself, and it is not proof that the publisher should be trusted.&lt;/p&gt;

&lt;p&gt;Discovery metadata is not authorization. A skill description is not a capability grant. A client saying &lt;code&gt;tenant=important-customer&lt;/code&gt; does not make that identity real.&lt;/p&gt;

&lt;p&gt;Those sound like obvious distinctions. They are also exactly the distinctions that disappear when a demo and a security model share the same JSON object.&lt;/p&gt;
&lt;h2&gt;
  
  
  One A2A message becomes a typed SMESH query
&lt;/h2&gt;

&lt;p&gt;At ingress, the gateway accepts bounded inline text, creates a stable A2A task envelope, and translates it into the actual core signal type used by SMESH:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;pub&lt;/span&gt; &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;to_signal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="k"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;gateway_node_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;Signal&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nn"&gt;Signal&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;builder&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nn"&gt;SignalType&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;Query&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nf"&gt;.payload_json&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nf"&gt;.origin&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;gateway_node_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nf"&gt;.build&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The payload carries the A2A task ID, context ID, protocol marker, and validated text. &lt;code&gt;ChannelDispatcher&lt;/code&gt; packages that typed signal with the request and hands both to a runtime-owned worker.&lt;/p&gt;

&lt;p&gt;That boundary is implemented and tested. The standalone binary still uses &lt;code&gt;LoopbackDispatcher&lt;/code&gt;, so it does &lt;strong&gt;not&lt;/strong&gt; yet inject the Query into a live multi-process SMESH runtime. The real runtime adapter is the next integration step.&lt;/p&gt;

&lt;p&gt;Setting &lt;code&gt;origin&lt;/code&gt; on the gateway Query is deliberate. In my previous article, independently corroborated &lt;em&gt;claims&lt;/em&gt; omitted their origin from the content hash so identical conclusions could converge on one address. This is a different signal. A gateway Query is an ingress envelope, not a claim waiting for independent corroboration. Its source belongs in the record.&lt;/p&gt;

&lt;p&gt;The boundary also rejects what it does not understand. The MVP accepts inline text only. It does not fetch a client-provided URL, dereference an arbitrary file, or treat external metadata as instructions for the mesh.&lt;/p&gt;

&lt;p&gt;Inline text is boring, but I know exactly what crosses the boundary.&lt;/p&gt;
&lt;h2&gt;
  
  
  The task ledger and the signal field tell different kinds of truth
&lt;/h2&gt;

&lt;p&gt;This was the most important design correction.&lt;/p&gt;

&lt;p&gt;A SMESH signal is supposed to decay. If nobody reinforces a task, its intensity falls until it no longer matters. That is useful inside the coordination system because stale work cleans itself up.&lt;/p&gt;

&lt;p&gt;An A2A task must not disappear because its internal coordination signal faded.&lt;/p&gt;

&lt;p&gt;An external client may reconnect five minutes later and call &lt;code&gt;GetTask&lt;/code&gt;. An auditor may list tasks by context. A user may need to see that a cancellation was accepted. An artifact must still belong to the task that produced it.&lt;/p&gt;

&lt;p&gt;So the gateway cannot reconstruct its public state by looking at the mesh.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A2A task ledger = retained external task state (process-local today)
SMESH signal     = temporary coordination pressure
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The ledger is authoritative for the A2A lifecycle. The mesh is authoritative only for its local coordination observations.&lt;/p&gt;

&lt;p&gt;That also means terminal states are absorbing. Once a task is completed, failed, canceled, or rejected, a repeated message cannot quietly restart work under the same ID.&lt;/p&gt;
&lt;h2&gt;
  
  
  Streaming exposed a bug that all my tests had missed
&lt;/h2&gt;

&lt;p&gt;A2A v1 supports both direct responses and stateful tasks. A task lifecycle stream begins with the Task itself, then emits ordered status or artifact updates, and closes when the task reaches a terminal state.[4]&lt;/p&gt;

&lt;p&gt;My executor maps internal mesh events like this:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;mesh dispatch accepted  -&amp;gt; Working
mesh progress           -&amp;gt; Working + status message
mesh artifact           -&amp;gt; Artifact update
mesh completion         -&amp;gt; Completed
mesh failure            -&amp;gt; Failed
accepted cancellation   -&amp;gt; Canceled
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The stream ordering test passed.&lt;/p&gt;

&lt;p&gt;Then an independent review pointed out that my own worker budget allowed 256 events while the upstream server's broadcast subscription buffer held 32. A fast worker could stay inside my documented limit and still outrun the initiating subscriber. The task might finish in storage while the client received an internal "subscription fell behind" error.&lt;/p&gt;

&lt;p&gt;That is a particularly unpleasant distributed-systems bug because both sides can truthfully report different outcomes.&lt;/p&gt;

&lt;p&gt;The fix was not to hope the subscriber ran faster. I reduced the worker event budget to 16, clamp caller-provided limits to that ceiling, and added a burst test through the official client.&lt;/p&gt;

&lt;p&gt;This is what protocols do to a design: they force every implicit assumption to become somebody else's observable failure.&lt;/p&gt;
&lt;h2&gt;
  
  
  Cancellation has to stop work, not just change a status label
&lt;/h2&gt;

&lt;p&gt;Forwarding &lt;code&gt;CancelTask&lt;/code&gt; to a dispatcher was not enough.&lt;/p&gt;

&lt;p&gt;If an internal worker ignored the request or kept its event stream open, the original client subscription could hang. Worse, late &lt;code&gt;Working&lt;/code&gt; or &lt;code&gt;Completed&lt;/code&gt; events could arrive after the public task had become &lt;code&gt;Canceled&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The executor now owns a per-task cancellation token. The first accepted cancellation:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;reaches the dispatcher;&lt;/li&gt;
&lt;li&gt;wakes the active producer loop;&lt;/li&gt;
&lt;li&gt;closes the original execution stream;&lt;/li&gt;
&lt;li&gt;prevents post-cancel work from changing the terminal state.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Channel sends and cancellation acknowledgements have deadlines. Worker inactivity has a deadline. The total task has a deadline.&lt;/p&gt;

&lt;p&gt;Cancellation is not a field update. It is a distributed state transition with work on both sides of the boundary.&lt;/p&gt;
&lt;h2&gt;
  
  
  The boring limits are the real feature
&lt;/h2&gt;

&lt;p&gt;The first version was interoperable. It was not bounded enough to deserve trust.&lt;/p&gt;

&lt;p&gt;A fail-closed review found unbounded task retention, unbounded worker output, cancellation leaks, missing dispatcher deadlines, terminal task reuse, and invalid input that could leave a task stranded in &lt;code&gt;Submitted&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The current localhost-first gateway now bounds:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;resource&lt;/th&gt;
&lt;th&gt;default boundary&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;HTTP request body&lt;/td&gt;
&lt;td&gt;128 KiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;accepted inline text&lt;/td&gt;
&lt;td&gt;64 KiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;retained process-local tasks&lt;/td&gt;
&lt;td&gt;1,024&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;active executions&lt;/td&gt;
&lt;td&gt;64&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;worker events&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;artifacts per task&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;aggregate output per task&lt;/td&gt;
&lt;td&gt;1 MiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;worker inactivity&lt;/td&gt;
&lt;td&gt;30 seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;total task execution&lt;/td&gt;
&lt;td&gt;5 minutes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;channel/cancel acknowledgement&lt;/td&gt;
&lt;td&gt;5 seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;It also refuses non-loopback binds unless an explicit unsafe override is present.&lt;/p&gt;

&lt;p&gt;That override does not add authentication, TLS, tenant isolation, or authorization. It only disables the refusal. The current binary is for localhost and trusted integration work, not direct exposure to the internet.[7]&lt;/p&gt;

&lt;p&gt;I am spelling that out because "supports enterprise authentication" in a protocol specification does not mean every prototype using the protocol is enterprise-secure.&lt;/p&gt;
&lt;h2&gt;
  
  
  The LIFELINE demo is a simulation, on purpose
&lt;/h2&gt;

&lt;p&gt;The cover video follows a fictional medication-safety incident. Three hospitals see weak pieces of the same adverse-event pattern. Separate manufacturer, regulator, logistics, payer, and evidence agents contribute artifacts. Inside each public endpoint, a SMESH swarm claims work, reinforces evidence, contests an early hypothesis, and lets unsupported signals decay.&lt;/p&gt;

&lt;p&gt;One logistics endpoint fails. Its task is canceled. A fallback is discovered. The incident continues. The agents converge on a recommendation, but a human incident commander owns the irreversible decision.&lt;/p&gt;

&lt;p&gt;The visual separates the layers deliberately:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;ivory arcs are A2A traffic between organizations;&lt;/li&gt;
&lt;li&gt;green fields are SMESH activity inside an organization;&lt;/li&gt;
&lt;li&gt;cyan shards are artifacts;&lt;/li&gt;
&lt;li&gt;vermilion marks contradiction or failure;&lt;/li&gt;
&lt;li&gt;one gold ring marks human authority.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every visible event comes from a 55-event, ordered, hash-chained JSONL fixture generated as one complete file. The browser can play it, scrub it, inspect it, or export deterministic frames. The narrated film and interactive replay are available from the gateway repository.[7][8]&lt;/p&gt;


&lt;div class="crayons-card c-embed"&gt;

  &lt;br&gt;
&lt;strong&gt;Honesty boundary:&lt;/strong&gt; the gateway's A2A interoperability is exercised against the official Rust client. The LIFELINE six-organization trace is synthetic. It proves the replay contract and the intended architecture; it does &lt;strong&gt;not&lt;/strong&gt; prove that six live SMESH runtimes executed the scenario.&lt;br&gt;

&lt;/div&gt;



&lt;p&gt;A captured run across live SMESH runtimes would be operational proof. This demo is not.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is implemented, and what is still a plan
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;implemented now&lt;/th&gt;
&lt;th&gt;still required for an internet-facing system&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A2A v1 Agent Card&lt;/td&gt;
&lt;td&gt;authenticated principals&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;JSON-RPC and HTTP+JSON/REST bindings&lt;/td&gt;
&lt;td&gt;tenant-aware authorization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;official-client tests for discovery, JSON-RPC/REST send, streaming, and cancellation&lt;/td&gt;
&lt;td&gt;persistent SQL task ledger&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SSE task streaming&lt;/td&gt;
&lt;td&gt;distributed quotas&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Get, List, and Subscribe routes through the SDK handler&lt;/td&gt;
&lt;td&gt;TLS termination and deployment policy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;real &lt;code&gt;SignalType::Query&lt;/code&gt; construction&lt;/td&gt;
&lt;td&gt;live SMESH runtime adapter behind every organization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;bounded process-local execution&lt;/td&gt;
&lt;td&gt;push callback validation and SSRF controls&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;deterministic synthetic replay&lt;/td&gt;
&lt;td&gt;captured multi-runtime causal trace&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;There are two easy ways to lie with a demo like this.&lt;/p&gt;

&lt;p&gt;The first is to animate what you wish the system did.&lt;/p&gt;

&lt;p&gt;The second is to run one loopback worker and describe it as a decentralized enterprise.&lt;/p&gt;

&lt;p&gt;I would rather keep the boundary visible.&lt;/p&gt;

&lt;h2&gt;
  
  
  What A2A changed in the way I think about SMESH
&lt;/h2&gt;

&lt;p&gt;Before this work, I thought of SMESH as the system.&lt;/p&gt;

&lt;p&gt;Now I think of it as an interior.&lt;/p&gt;

&lt;p&gt;The mesh can remain weird in useful ways. Signals can diffuse probabilistically. Specialists can appear and disappear. Trust can be local. Claims can decay. None of that has to leak into the contract presented to another agent.&lt;/p&gt;

&lt;p&gt;MCP gives individual agents hands. SMESH gives a group local coordination. A2A gives that group a public identity, a task contract, and a cancel button.&lt;/p&gt;

&lt;p&gt;The previous transport article ended with a lesson: code that has never been executed is a plan. This work left me with the same rule one layer higher:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A system nobody else can discover, hire, observe, or cancel is not interoperable. It is an island with excellent internal networking.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you are building multi-agent systems, where would you draw the boundary between cross-organization interoperability and internal swarm coordination—and which decisions would you refuse to delegate?&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/copyleftdev" rel="noopener noreferrer"&gt;
        copyleftdev
      &lt;/a&gt; / &lt;a href="https://github.com/copyleftdev/smesh-a2a" rel="noopener noreferrer"&gt;
        smesh-a2a
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      A2A v1 interoperability gateway for decentralized SMESH agent swarms
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;SMESH A2A&lt;/h1&gt;
&lt;/div&gt;
&lt;p&gt;A2A v1 interoperability gateway for decentralized &lt;a href="https://github.com/copyleftdev/smesh-rust" rel="noopener noreferrer"&gt;SMESH&lt;/a&gt; agent swarms.&lt;/p&gt;
&lt;p&gt;SMESH remains the internal coordination substrate: signals diffuse, decay, reinforce, and accumulate attestations. A2A is the public contract for discovery, durable task lifecycle, streaming progress, cancellation, and artifacts.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;What works&lt;/h2&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;Official A2A v1 Rust types and server/client SDKs&lt;/li&gt;
&lt;li&gt;Public Agent Card at &lt;code&gt;/.well-known/agent-card.json&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;JSON-RPC endpoint at &lt;code&gt;/jsonrpc&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;HTTP+JSON/REST endpoint at &lt;code&gt;/rest&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Synchronous and SSE streaming task execution&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;GetTask&lt;/code&gt;, &lt;code&gt;ListTasks&lt;/code&gt;, &lt;code&gt;SubscribeToTask&lt;/code&gt;, and &lt;code&gt;CancelTask&lt;/code&gt; through &lt;code&gt;a2a-rs&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Strict inline-text validation with a 64 KiB default limit&lt;/li&gt;
&lt;li&gt;Translation to a real &lt;code&gt;smesh_core::SignalType::Query&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Injectable &lt;code&gt;MeshDispatcher&lt;/code&gt; boundary for a production SMESH runtime&lt;/li&gt;
&lt;li&gt;Deterministic loopback worker for demos and interoperability tests&lt;/li&gt;
&lt;li&gt;Official-client LIFELINE Response Director with concurrent commissioning, reconnect
bounded cancellation, logistics fallback, independent review, and ID-complete run records&lt;/li&gt;
&lt;li&gt;Deterministic organization-runtime LIFELINE outage scenario with joined cancellation
one linked fallback, sibling continuity, and bounded restricted JSONL replay evidence&lt;/li&gt;
&lt;li&gt;Five isolated organization-local LIFELINE &lt;code&gt;SmeshRuntime&lt;/code&gt;…&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/copyleftdev/smesh-a2a" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;Interactive replay: &lt;a href="https://copyleftdev.github.io/smesh-a2a/" rel="noopener noreferrer"&gt;copyleftdev.github.io/smesh-a2a&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;SMESH core: &lt;a href="https://github.com/copyleftdev/smesh-rust" rel="noopener noreferrer"&gt;github.com/copyleftdev/smesh-rust&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;p&gt;[1] &lt;a href="https://dev.to/copyleftdev/my-quic-transport-had-never-once-been-executed-heres-what-happened-when-i-ran-it-24ge"&gt;https://dev.to/copyleftdev/my-quic-transport-had-never-once-been-executed-heres-what-happened-when-i-ran-it-24ge&lt;/a&gt; — My QUIC transport had never once been executed&lt;br&gt;
[2] &lt;a href="https://developers.googleblog.com/en/a2a-a-new-era-of-agent-interoperability" rel="noopener noreferrer"&gt;https://developers.googleblog.com/en/a2a-a-new-era-of-agent-interoperability&lt;/a&gt; — Announcing the Agent2Agent Protocol&lt;br&gt;
[3] &lt;a href="https://a2a-protocol.org/latest" rel="noopener noreferrer"&gt;https://a2a-protocol.org/latest&lt;/a&gt; — What is A2A Protocol?&lt;br&gt;
[4] &lt;a href="https://a2a-protocol.org/latest/specification" rel="noopener noreferrer"&gt;https://a2a-protocol.org/latest/specification&lt;/a&gt; — A2A Protocol v1.0 specification&lt;br&gt;
[5] &lt;a href="https://linuxfoundation.org/press/linux-foundation-launches-the-agent2agent-protocol-project-to-enable-secure-intelligent-communication-between-ai-agents" rel="noopener noreferrer"&gt;https://linuxfoundation.org/press/linux-foundation-launches-the-agent2agent-protocol-project-to-enable-secure-intelligent-communication-between-ai-agents&lt;/a&gt; — Linux Foundation launches A2A project&lt;br&gt;
[6] &lt;a href="https://github.com/copyleftdev/smesh-rust" rel="noopener noreferrer"&gt;https://github.com/copyleftdev/smesh-rust&lt;/a&gt; — SMESH repository&lt;br&gt;
[7] &lt;a href="https://github.com/copyleftdev/smesh-a2a" rel="noopener noreferrer"&gt;https://github.com/copyleftdev/smesh-a2a&lt;/a&gt; — SMESH A2A gateway repository&lt;br&gt;
[8] &lt;a href="https://youtu.be/EFPKaIuF8iA" rel="noopener noreferrer"&gt;https://youtu.be/EFPKaIuF8iA&lt;/a&gt; — SMESH A2A cinematic demo&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>architecture</category>
      <category>distributedsystems</category>
    </item>
    <item>
      <title>A Wider Computer, Not a Bigger One: Modeling AI Inference Across Millions of Homes</title>
      <dc:creator>Don Johnson</dc:creator>
      <pubDate>Tue, 25 Aug 2026 05:59:43 +0000</pubDate>
      <link>https://dev.to/copyleftdev/a-wider-computer-not-a-bigger-one-modeling-ai-inference-across-millions-of-homes-5cmo</link>
      <guid>https://dev.to/copyleftdev/a-wider-computer-not-a-bigger-one-modeling-ai-inference-across-millions-of-homes-5cmo</guid>
      <description>&lt;p&gt;&lt;em&gt;I modeled an AI inference fleet distributed across ordinary homes. What survived was narrower—and more plausible—than what I started with.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Picture a detached house on a cold evening. On one garage wall, an operator-owned compute appliance draws about five kilowatts—electrically comparable to an EV charger, though it runs much longer. It serves small AI models. In winter its waste heat can warm the house; in summer that heat must be carried away. The family bought nothing. They are paid to host it.&lt;/p&gt;

&lt;p&gt;Now repeat that arrangement across a neighborhood, then a state, then millions of homes—not to assemble one enormous computer, but to create millions of independent inference workers. Each keeps a small model resident in GPU memory. New requests route around homes that are offline. The network grows by adding locations that were already built and connected to the grid.&lt;/p&gt;

&lt;p&gt;That network does not exist. I call the idea HEARTH. Almost none of its physical pieces are exotic; the experiment is whether they can be arranged and scheduled as one system. Not a bigger computer. A wider one.&lt;/p&gt;

&lt;p&gt;To make the appliance less abstract, I developed &lt;a href="https://copyleftdev.github.io/hearth/prototypes/" rel="noopener noreferrer"&gt;three concept form factors&lt;/a&gt;: a wall unit, a floor-standing thermal tower, and a duct-integrated mechanical-room unit. They are appearance and installation studies, not engineered products; their job is to expose the questions that a real prototype must answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The model kept changing its answer
&lt;/h2&gt;

&lt;p&gt;Then I priced the accelerators, and the idea stopped working.&lt;/p&gt;

&lt;p&gt;That was the first useful result. My original comparison had focused on the dramatic expense of constructing a data center while treating the compute hardware as identical on both sides. But identical hardware does not disappear from the economics. If one location keeps an expensive GPU busier than another, utilization can overwhelm everything the cheaper building saves.&lt;/p&gt;

&lt;p&gt;I rebuilt the model five times. Each pass introduced a constraint the previous one had missed:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pass&lt;/th&gt;
&lt;th&gt;What changed&lt;/th&gt;
&lt;th&gt;Model verdict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Compared the facilities around the hardware&lt;/td&gt;
&lt;td&gt;Homes win&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Included accelerator cost and utilization&lt;/td&gt;
&lt;td&gt;Homes lose&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Used consumer and data-center hardware at quoted prices&lt;/td&gt;
&lt;td&gt;Homes win&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Modeled production batching for a 32B model&lt;/td&gt;
&lt;td&gt;Homes lose&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Tested different model sizes&lt;/td&gt;
&lt;td&gt;Homes win only for the &lt;strong&gt;small-model case&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What survived: small, independent inference
&lt;/h2&gt;

&lt;p&gt;The fifth pass did not rescue my original proposal. It replaced it. A residential fleet is not a cheaper place to run every AI workload. It cannot pool memory across the internet, train a frontier model, or make residential latency disappear. What survived was narrower: small-model inference in which one node completes one request and no durable state is tied to a particular house.&lt;/p&gt;

&lt;p&gt;Model size changes the economics because inference is not just a race between GPUs. During generation, a server repeatedly reads the model's weights while advancing many requests together. That is batching. A larger batch spreads each weight read across more paying tokens—but every active request also consumes working memory, usually called the KV cache.&lt;/p&gt;

&lt;p&gt;An 8B model—roughly eight billion parameters—leaves enough memory on a 32 GB consumer GPU for a useful production batch. In my model, both the home GPU and the data-center hardware then become limited mainly by compute, allowing the consumer part's much lower purchase price to matter. At 32B, the consumer card runs short of memory first. Its batch stops growing while high-memory data-center hardware keeps filling, and the verdict reverses.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model class&lt;/th&gt;
&lt;th&gt;What the current model says&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;8B&lt;/td&gt;
&lt;td&gt;Candidate workload; the consumer GPU retains useful batching headroom&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;32B&lt;/td&gt;
&lt;td&gt;Data-center hardware wins unless other advantages offset its batching edge&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;70B&lt;/td&gt;
&lt;td&gt;Excluded by my single-GPU serving assumption; multi-GPU serving was not modeled&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The current result is conditional. Before homeowner compensation and fleet-level operating costs, and under RTX 5090 MSRP, an 85% accelerator duty factor, identical serving-efficiency assumptions for both venues, and a quantized 8B workload, my roofline model estimates the residential hardware-and-energy stack at about 0.49 times the cost per token of the best data-center case in the sweep, based on GB200 NVL72 rack pricing. That is a reproducible scenario output, not observed production performance or a fully loaded business cost. An apples-to-apples benchmark—or the costs excluded here—could erase the advantage.&lt;/p&gt;

&lt;p&gt;This is the distinction behind &lt;em&gt;wide, not deep&lt;/em&gt;. HEARTH would not divide one giant model among houses or combine residential GPUs into a virtual supercomputer. Each node would answer complete, independent requests with a model it already holds locally. Adding homes increases the number and variety of requests the fleet can serve; it does not make any one node larger.&lt;/p&gt;

&lt;h2&gt;
  
  
  The asset is the grid endpoint
&lt;/h2&gt;

&lt;p&gt;Why put that workload in homes at all? Not because residential electricity is cheap—it usually is not—or because waste heat makes energy free. The asset is the connection behind the meter.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://datacenters.lbl.gov/publications/2024-lbnl-data-center-energy-usage-report" rel="noopener noreferrer"&gt;Berkeley Lab estimated&lt;/a&gt; that US data centers used about 176 TWh of electricity in 2023. &lt;a href="https://datacenters.lbl.gov/publications/united-states-data-center-energy-2025" rel="noopener noreferrer"&gt;Its June 2026 update&lt;/a&gt; projects 521–843 TWh in 2030, with a 649 TWh reference case. Serving that growth is not simply a question of buying more GPUs. A large new campus may need substations, transformers, transmission upgrades, and a negotiated path to connect a load the local grid was never designed to carry. The &lt;a href="https://www.ercot.com/files/docs/2026/04/01/ERCOT_LargeLoad_Update_April2026_B-C_-Hearing.pdf" rel="noopener noreferrer"&gt;410 GW of prospective large loads reported in ERCOT&lt;/a&gt;—about 87% associated with data centers—is not a forecast of what will be built. It is evidence of how much demand is converging on the same constrained process.&lt;/p&gt;

&lt;p&gt;The United States has about &lt;a href="https://data.census.gov/table/ACSDP1Y2024.DP04?g=010XX00US" rel="noopener noreferrer"&gt;82.5 million one-unit detached housing units&lt;/a&gt;. That is physical stock, not an estimate of eligible hosts: occupancy, broadband, electrical capacity, utility approval, and household consent would reduce it substantially. Each candidate already has a meter, an electrical service, and a physical building around it. A compute node installed behind that meter may avoid the transmission-scale interconnection required by a new campus. That is the inversion: instead of bringing an enormous new grid connection to the compute, bring a modest amount of compute to many connections that already exist.&lt;/p&gt;

&lt;p&gt;Existing does not mean unused. A continuous 5 kW load is much less forgiving than an EV that charges for a few hours. Some homes would need panel or service upgrades; utilities may require review; and the local distribution transformer remains a hard physical constraint. My one-node-per-transformer rule is therefore a screening hypothesis, not a validated safety rule or national capacity estimate. Transformer loading, telemetry, and thermal behavior belong in the first utility-supervised field test.&lt;/p&gt;

&lt;p&gt;To the household, this is a hosting contract, not an investment. The operator would own and maintain the appliance, meter and reimburse its electricity, and pay the family a share of revenue. None of that has been field-tested; fire and insurance rules, noise, summer heat rejection, ISP terms, maintenance, and upgrades remain field questions.&lt;/p&gt;

&lt;h2&gt;
  
  
  One request goes to one home
&lt;/h2&gt;

&lt;p&gt;At the software layer, the system is deliberately less exotic. A control plane would assign replicas of small models to qualifying nodes and distribute checkpoints outside the request path. Each node would store its assigned checkpoints locally and keep its serving model resident in GPU memory. The request router would know which homes have the requested model resident, which are healthy, and which have batch capacity available.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;registry ── sync ──► home node

client ──► router ──► available home ──► response
               ▲              │
               └── telemetry ─┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;During generation, one home holds the request's KV cache and returns the tokens; no activations or model layers cross the residential network. If the node disappears, the in-flight generation is lost. New work can route elsewhere, and a retryable request can start again against another replica, although it may not produce identical tokens. The system does not inherit data-center reliability merely because it has many machines.&lt;/p&gt;

&lt;p&gt;At the network level, remote attestation, model security, and prompt and output confidentiality on hardware outside an operator-controlled facility remain unresolved. Distribution makes failures routable; it does not make operations disappear.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five conditions decide whether this is real
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The hardware must be legally and commercially usable.&lt;/strong&gt; &lt;a href="https://www.nvidia.com/en-us/drivers/nvidia-license/" rel="noopener noreferrer"&gt;NVIDIA's current driver agreement&lt;/a&gt; says the software may not be used to provide commercial hosting services and that GeForce and Titan software is not licensed for data-center deployment. Whether an operator could obtain written authorization or a different license for this residential topology must be resolved with NVIDIA and qualified counsel before a hardware pilot. Warranty, support, and a credible procurement channel must also exist.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Edge serving must sustain the modeled batches.&lt;/strong&gt; The same checkpoint, quantization, context length, serving stack, and request mix must be benchmarked on the candidate consumer GPU and data-center hardware. A spreadsheet cannot substitute for that result.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Demand must keep the silicon busy.&lt;/strong&gt; Cheap idle GPUs are still expensive. Customers must commit enough suitable 8B inference work—classification, extraction, routing, drafting, or other bounded tasks—to maintain high utilization at a viable token price.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The location must pass the energy screen.&lt;/strong&gt; Residential rates vary widely, and heat reuse is a seasonal credit rather than the thesis. Every deployment needs to work after metered electricity, host payment, and summer heat rejection.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The distribution grid must approve the load.&lt;/strong&gt; A real utility must confirm that a real home, service, and transformer can carry it continuously. National averages are not permission to connect.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Those are not caveats around the result. They &lt;em&gt;are&lt;/em&gt; the result. HEARTH exists only at their intersection.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first pilot should contain no houses
&lt;/h2&gt;

&lt;p&gt;That changes the order of the pilot. My original plan began with 250 homes and grew through increasingly expensive gates. It tested installation first and market demand later. The model eventually exposed the flaw: there is no reason to put hardware in a house before proving that anyone will buy this particular kind of inference.&lt;/p&gt;

&lt;p&gt;The first HEARTH pilot should therefore contain no houses. Rent the candidate consumer and data-center hardware. Resolve the license question. Run the same models through the same serving stack, measure delivered tokens and power at production batch, and ask customers to pay for the workloads the residential fleet is supposed to serve. If the measured economics or demand fail, stop before an electrician is dispatched.&lt;/p&gt;

&lt;p&gt;Only then move to 250 homes. That field trial buys answers the lab cannot: Can crews install safely and consistently? What do transformers do under continuous load? How loud and hot are the nodes in July? Can the operator keep them online without drowning in truck rolls—and will families keep hosting them?&lt;/p&gt;

&lt;h2&gt;
  
  
  Four million homes is not the next milestone
&lt;/h2&gt;

&lt;p&gt;Four million homes is an image of the possible system, not an adoption forecast. The serious next milestone is smaller: one workload that benchmarks, one customer willing to buy it, one utility willing to supervise it, and then the first 250 homes.&lt;/p&gt;

&lt;p&gt;I published the &lt;a href="https://copyleftdev.github.io/hearth/report/" rel="noopener noreferrer"&gt;full feasibility study&lt;/a&gt;, the &lt;a href="https://github.com/copyleftdev/hearth" rel="noopener noreferrer"&gt;source model and inputs&lt;/a&gt;, and every assumption I used. If the idea is wrong, I want the failure to be reproducible too.&lt;/p&gt;

&lt;p&gt;HEARTH is not a forecast. It is a claim specific enough to test—and to kill if it fails. The buildings, meters, and many of the grid and network endpoints already surround us. What is missing is the appliance, the dedicated installation, the operating system around it, and an agreement worth signing.&lt;/p&gt;

&lt;p&gt;If those pieces work, the neighborhood still looks like a neighborhood on a cold evening. Behind some garage walls, independent machines serve requests from a network spanning the country—each complete on its own, all scheduled together.&lt;/p&gt;

&lt;p&gt;No monument to compute. A million warm windows.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>aiops</category>
      <category>infrastructure</category>
    </item>
  </channel>
</rss>
