<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Rob</title>
    <description>The latest articles on DEV Community by Rob (@carryologist).</description>
    <link>https://dev.to/carryologist</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3884903%2Ff7cf0bfd-0b92-4dca-9095-683af23a19e3.png</url>
      <title>DEV Community: Rob</title>
      <link>https://dev.to/carryologist</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/carryologist"/>
    <language>en</language>
    <item>
      <title>What Is "Alpha," and Why Does It Keep Coming Up In AI Debates?</title>
      <dc:creator>Rob</dc:creator>
      <pubDate>Tue, 28 Jul 2026 04:13:46 +0000</pubDate>
      <link>https://dev.to/carryologist/what-is-alpha-and-why-does-it-keep-coming-up-in-ai-debates-4f05</link>
      <guid>https://dev.to/carryologist/what-is-alpha-and-why-does-it-keep-coming-up-in-ai-debates-4f05</guid>
      <description>&lt;p&gt;Well, it's not raining, but I find myself back at my desk doing research ahead of more Thursday Thoughts and experiments. Last time it was an actual gray Cape Cod afternoon and a list of homelab tools I'd been meaning to pin down. This time the weather's fine and the itch is different: a word. Specifically, "alpha," which I keep hearing thrown around in AI debates like everyone already agrees on what it means. I don't think we do. So consider this the second entry in what's turning into a series: research first, opinion later. This one isn't a Thursday Thoughts post itself, it's the homework before one.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "Alpha" Actually Means
&lt;/h2&gt;

&lt;p&gt;My working assumption going in was that "alpha" meant something like IP, or maybe just "intelligence," a cute stand-in for whatever makes a company or a model smart. That's wrong, or at least it's not where the word comes from. Alpha is a finance term, and it has a precise, almost boring definition: it's the return an investment generates &lt;em&gt;above&lt;/em&gt; what you'd expect given the risk you took on, measured against a benchmark. A fund with an alpha of 5 means it outperformed the market by 5%. It's always paired with beta, which is just your exposure to the market itself. Beta is what you get for free by showing up. Alpha is what you get for actually being good.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the Word Came From
&lt;/h2&gt;

&lt;p&gt;The specific origin is a 1968 paper by economist Michael Jensen, whose namesake metric, "Jensen's alpha," was originally built to show that most active fund managers &lt;em&gt;weren't&lt;/em&gt; actually beating the market once you adjusted for risk. Alpha, in other words, was invented as a skeptic's tool. It's the number that separates real skill from just being along for a rising tide.&lt;/p&gt;

&lt;p&gt;That framing migrated out of finance and into startup and VC culture over the last decade or so, where it got looser and more metaphorical. In that world, alpha became shorthand for whatever counts as a durable, non-obvious edge: a founder's unique insight, a VC's proprietary deal flow, the thing the rest of the market hasn't priced in yet. The common thread across both the strict and the loose definitions: alpha is never just "being smart" in the abstract. It's the &lt;em&gt;excess&lt;/em&gt;, the part that isn't explained by everyone having access to the same information or the same market.&lt;/p&gt;

&lt;h2&gt;
  
  
  Alpha Enters the AI Debate
&lt;/h2&gt;

&lt;p&gt;Here's where it gets interesting for anyone building or buying AI right now. A lot of current AI commentary is really just Jensen's question, restated: now that everyone has access to roughly the same frontier models, where does the &lt;em&gt;excess&lt;/em&gt; come from? One recent take on this put it as bluntly as I've seen: "The alpha isn't in better models," arguing the real edge is organizational, not computational, who can actually turn a model into money. Bloomberg asked almost the identical question as a headline, &lt;a href="https://www.bloomberg.com/professional/insights/artificial-intelligence/is-ai-an-alpha-engine/" rel="noopener noreferrer"&gt;&lt;em&gt;Is AI an alpha engine?&lt;/em&gt;&lt;/a&gt;, and landed somewhere similarly hedged: AI helps, but the differentiator is what goes in, not what comes out.&lt;/p&gt;

&lt;p&gt;There's an even sharper, more literal version of this happening in quant finance, which is fitting given that's where the word started. A recent paper on AI-driven alpha decay models how mass AI adoption in trading endogenously destroys the very excess returns it's supposed to generate: as more funds run AI on the same shared data, their signals converge, and the edge each one extracts has a shrinking half-life, estimated at as little as 18 months at current adoption levels versus 5-7 years before AI. That's not a metaphor. That's the actual word "alpha," in its actual home discipline, mathematically eroding as an actual side effect of AI adoption. Worth sitting with, given what's coming next.&lt;/p&gt;

&lt;h2&gt;
  
  
  All-In's Version: "Don't Give Away Your Alpha"
&lt;/h2&gt;

&lt;p&gt;This is the thread that sent me down this whole research hole. On &lt;a href="https://podcasts.happyscribe.com/all-in-with-chamath-jason-sacks-friedberg/ai-sovereignty-wars-palantir-nvidia-deal-scotus-birthright-ruling-newsom-s-ca-budget-lie" rel="noopener noreferrer"&gt;episode 279 of the All-In podcast&lt;/a&gt;, the besties dug into Palantir's sovereign-AI partnership with Nvidia and Alex Karp's CNBC interview around it. Their summary of Karp's argument used "alpha" in exactly the sense above, but aimed at enterprises instead of traders: what technical customers want, they said, is control over their compute, their models, their data stack, and their alpha, meaning their proprietary knowledge, the fear being that a frontier lab could hoover up that proprietary knowledge and eventually turn it into a competing product. Their tagline for the whole idea: "Data retention is your treasure."&lt;/p&gt;

&lt;p&gt;Friedberg added a real example on the same episode: Anthropic pitching data-sharing arrangements to life sciences companies, most of whom concluded that sharing would commoditize their own business. Chamath then did something I appreciated: he actually tested it, rather than just asserting it. At his company 8090, he ran a standard enterprise migration task across configurations and reported the results on-air: their own harness wrapped around Claude was 1.4x cheaper and 1.5x faster than raw Anthropic Opus, while an open-source model behind that same harness was 16.4x cheaper, though roughly three times slower. His challenge to the audience wasn't "open source always wins." It was closer to: if the savings are this large, why aren't you at least checking whether you can keep your edge off someone else's servers?&lt;/p&gt;

&lt;p&gt;I'd be doing this research a disservice if I didn't flag the pushback too. &lt;a href="https://siliconangle.com/2026/07/05/alex-karp-frontier-models-real-fight-enterprise-ai/" rel="noopener noreferrer"&gt;SiliconANGLE's analysis&lt;/a&gt; of the same episode makes an important point: there is no public evidence that Anthropic or OpenAI trains on customer data in violation of their own terms, and OpenAI has said outright that it doesn't train on customer API data. Karp's framing, per that piece, is partly a fear campaign, even if the underlying enterprise anxiety is real. SiliconANGLE's own shorthand for the two camps is worth stealing: "data communism," where every firm gets access to the same intelligence, versus "data capitalism," where proprietary advantage stays exclusive. I don't think that fight is settled. I think it's exactly the debate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frontier vs. Self-Hosted: The Pros and Cons
&lt;/h2&gt;

&lt;p&gt;Stripping the podcast drama away, here's the actual tradeoff, as best I can lay it out honestly from this round of research:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Frontier APIs (Claude, GPT, Gemini)&lt;/th&gt;
&lt;th&gt;Self-hosted / open-weight&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Alpha exposure&lt;/td&gt;
&lt;td&gt;Every prompt is a data transfer to a company that has, in adjacent categories, already shipped competing products against its own ecosystem&lt;/td&gt;
&lt;td&gt;Nothing leaves your infrastructure; the weights and the data stack are actually yours&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Raw capability&lt;/td&gt;
&lt;td&gt;Best available today, particularly for the hardest reasoning tasks&lt;/td&gt;
&lt;td&gt;Real gap remains, and roughly 3x slower in Chamath's own test&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost&lt;/td&gt;
&lt;td&gt;Priced per token, on a business model Karp's camp argues structurally limits your leverage at the model layer&lt;/td&gt;
&lt;td&gt;Up to 16.4x cheaper at scale, once the harness is built&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vendor stability&lt;/td&gt;
&lt;td&gt;Subject to policy whiplash, the same episode cited Anthropic's Fable 5 export-control reversal as a live example&lt;/td&gt;
&lt;td&gt;Immune to another company's board decisions, licensing changes, or export-control flip-flops&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Effort to be competitive&lt;/td&gt;
&lt;td&gt;Works out of the box, fastest path from idea to working product&lt;/td&gt;
&lt;td&gt;Real engineering investment, a raw open model without a proper harness underperforms badly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Who owns the risk&lt;/td&gt;
&lt;td&gt;Frontier vendor absorbs most operational and safety burden&lt;/td&gt;
&lt;td&gt;You now own that operational and safety burden yourself&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Neither column is a strawman. They're both true at once, which is exactly why this is a real debate and not a marketing slogan in either direction.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Looks Like at Homelab Scale
&lt;/h2&gt;

&lt;p&gt;Here's the part that made this research personal rather than academic: I've been running a miniature, unfunded version of Chamath's exact experiment for months without ever calling it that. Every &lt;a href="https://dev.to/posts/model-showdown-round-7-local-models-vs-the-tag-manager"&gt;Model Showdown&lt;/a&gt; round on this blog, every fight with a chat template, every &lt;code&gt;--jinja&lt;/code&gt; flag, has been me asking the same question 8090 is asking with an enterprise budget: is the harness worth building, or should I just rent the frontier? My local rig will never beat Opus or Sonnet on a hard reasoning task, and I don't pretend otherwise. But nothing I run through it teaches Anthropic anything about how I build. That's the whole trade, just shrunk down from a boardroom to a garage.&lt;/p&gt;

&lt;p&gt;I don't have a tidy answer yet, and I'm deliberately not trying to force one into this post. That's not what this one is for. Consider this the research file, out in the open, ahead of the actual take.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If your business ran entirely on a frontier API tomorrow, would you know what you'd handed over, and to whom?&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  By the Numbers
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;12&lt;/strong&gt; — search queries it took to run down the origin story of a five-letter word&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;58&lt;/strong&gt; — years between Michael Jensen's original 1968 alpha paper and this post&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;4&lt;/strong&gt; — All-In besties involved in the episode that started this, zero of whom are actually named Alpha&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;16.4x&lt;/strong&gt; — how much cheaper Chamath's open-source harness ran versus raw Anthropic Opus, the number that kicked off this whole rabbit hole&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;3x&lt;/strong&gt; — how much slower that same cheaper setup was, because nothing is ever just one stat&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;11&lt;/strong&gt; — sources cited below, one of which is this blog quoting itself&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;0&lt;/strong&gt; — new conclusions reached today, this is a research post, the take comes later&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>meta</category>
      <category>buildinginpublic</category>
      <category>ai</category>
      <category>opensource</category>
    </item>
    <item>
      <title>The OpenAI And Hugging Face Exploit Got Me Thinking: Is There a Standard Agent "Sandbox" Definition? Ends Up, Yes</title>
      <dc:creator>Rob</dc:creator>
      <pubDate>Fri, 24 Jul 2026 14:44:20 +0000</pubDate>
      <link>https://dev.to/carryologist/the-openai-and-hugging-face-exploit-got-me-thinking-is-there-a-standard-agent-sandbox-a3j</link>
      <guid>https://dev.to/carryologist/the-openai-and-hugging-face-exploit-got-me-thinking-is-there-a-standard-agent-sandbox-a3j</guid>
      <description>&lt;p&gt;I started this one as a Thursday Thoughts post. The setup was clean: OpenAI disclosed that two of its models escaped an evaluation sandbox and breached Hugging Face's production infrastructure to steal the answer key to their own benchmark, and the "sandbox" turned out to be a container with exactly one sanctioned exit, a package-registry proxy, that had a zero-day in it. Righteous conclusion already forming: the industry needs a real, testable definition of "sandboxed," not a marketing word everyone nods along to.&lt;/p&gt;

&lt;p&gt;Then I got to the part where I was about to write "someone should define this properly" and stopped. That's a lazy thing to assert without checking. Maybe someone already had. So I put the hot take on ice and went looking instead. This is that research, not the take.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Landscape, Briefly
&lt;/h2&gt;

&lt;p&gt;The short version: there's a lot written about agent security, and almost none of it is a scoring standard for a single sandbox's containment architecture specifically.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://genai.owasp.org/initiatives/agentic-security-initiative/" rel="noopener noreferrer"&gt;OWASP's Agentic AI Top 10&lt;/a&gt; and its &lt;a href="https://cheatsheetseries.owasp.org/cheatsheets/AI_Agent_Security_Cheat_Sheet.html" rel="noopener noreferrer"&gt;Agent Security Cheat Sheet&lt;/a&gt; catalog threats and mitigations at a high level, useful as a checklist, not built to produce a comparable score. &lt;a href="https://www.nist.gov/itl/ai-risk-management-framework" rel="noopener noreferrer"&gt;NIST's AI Risk Management Framework&lt;/a&gt; operates a level above this entirely, it's organizational risk governance, not a technical grading rubric for a runtime boundary. &lt;a href="https://atlas.mitre.org/" rel="noopener noreferrer"&gt;MITRE ATLAS&lt;/a&gt; catalogs adversarial techniques against AI systems, closer to a threat library than a containment measure. The Cloud Security Alliance has several overlapping efforts, &lt;a href="https://cloudsecurityalliance.org/blog/2025/02/06/agentic-ai-threat-modeling-framework-maestro" rel="noopener noreferrer"&gt;MAESTRO&lt;/a&gt;, an &lt;a href="https://cloudsecurityalliance.org/artifacts/ai-controls-matrix-v1-1" rel="noopener noreferrer"&gt;AI Controls Matrix&lt;/a&gt;, and an &lt;a href="https://cloudsecurityalliance.org/blog/2026/02/02/the-agentic-trust-framework-zero-trust-governance-for-ai-agents" rel="noopener noreferrer"&gt;Agentic Trust Framework&lt;/a&gt; that scores autonomy on a four-stage ladder from "Intern" to "Principal," which is the closest thing I found to one specific slice of what I was after, how much an agent can do without a human, but it isn't scoped to sandboxing as a whole. RAND's &lt;a href="https://www.rand.org/pubs/research_reports/RRA2849-1.html" rel="noopener noreferrer"&gt;&lt;em&gt;Securing AI Model Weights&lt;/em&gt;&lt;/a&gt; defines five security levels, SL1 through SL5, but for weight theft and exfiltration risk at a lab, not for whether a given agent's runtime sandbox holds under an adversarial task. There's also a recent arXiv paper, &lt;a href="https://arxiv.org/abs/2606.18532" rel="noopener noreferrer"&gt;&lt;em&gt;AI Sandboxes: A Threat Model, Taxonomy, and Measurement Framework&lt;/em&gt;&lt;/a&gt;, that's structurally interesting, but it classifies sandboxes into archetypes (simulation-based, digital-twin, adversarial, regulatory, agent-based) rather than decomposing one sandbox into independently gradable layers.&lt;/p&gt;

&lt;p&gt;None of those do the specific thing I was looking for: take a single agent sandbox, break it into independent parts, score each part, and produce something you could compare across products. Then I found one that does exactly that.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Agent Sandbox Taxonomy
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/kajogo777/the-agent-sandbox-taxonomy" rel="noopener noreferrer"&gt;The Agent Sandbox Taxonomy&lt;/a&gt;, published in March 2026 and still under active community review, organizes itself around a memorable "7-7-3": &lt;strong&gt;seven defense layers&lt;/strong&gt;, &lt;strong&gt;seven threat categories&lt;/strong&gt;, and &lt;strong&gt;three evaluation dimensions&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The seven layers, numbered bottom-up because lower layers are foundational:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Name&lt;/th&gt;
&lt;th&gt;Key Question&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;L1&lt;/td&gt;
&lt;td&gt;Compute Isolation&lt;/td&gt;
&lt;td&gt;What separates the agent's execution from the host?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;L2&lt;/td&gt;
&lt;td&gt;Resource Limits&lt;/td&gt;
&lt;td&gt;Can it exhaust CPU, memory, disk, or time?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;L3&lt;/td&gt;
&lt;td&gt;Filesystem Boundary&lt;/td&gt;
&lt;td&gt;What can it read, write, or delete?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;L4&lt;/td&gt;
&lt;td&gt;Network Boundary&lt;/td&gt;
&lt;td&gt;What can it communicate with?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;L5&lt;/td&gt;
&lt;td&gt;Credential &amp;amp; Secret Management&lt;/td&gt;
&lt;td&gt;Can it see, use, or exfiltrate credentials?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;L6&lt;/td&gt;
&lt;td&gt;Action Governance&lt;/td&gt;
&lt;td&gt;Can it perform destructive or unauthorized operations?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;L7&lt;/td&gt;
&lt;td&gt;Observability &amp;amp; Audit&lt;/td&gt;
&lt;td&gt;Can you see what it did, when, and why?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Each layer gets scored on &lt;strong&gt;Strength&lt;/strong&gt; (0 to 4) and &lt;strong&gt;Granularity&lt;/strong&gt; (0 to 3), plus a flat set of &lt;strong&gt;Portability&lt;/strong&gt; tags for OS and infrastructure dependencies. The Strength scale is the part I keep coming back to, because it's precisely the distinction that mattered in the OpenAI incident: 0 is no enforcement, 1 is cooperative enforcement the sandboxed process can simply ignore or route around (proxy environment variables, an opt-in convention), 2 is software-enforced by something the process can't bypass internally but an operator could reconfigure, 3 is kernel-enforced and irreversible once applied (namespaces, Landlock, seccomp-BPF), and 4 is structural, the protected resource just doesn't exist inside the sandbox at all (a microVM, a credential proxy, no network device).&lt;/p&gt;

&lt;p&gt;Every product gets a fingerprint, a CVSS-style vector showing strength at each layer in order:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;E2B      L1:4/L2:4/L3:4/L4:0/L5:2/L6:-/L7:2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The taxonomy also maps its seven threats (data exfiltration, supply-chain compromise, destructive operations, lateral movement, persistence, privilege escalation, denial of service) back onto specific layer combinations with explicit thresholds, so "is T1 exfiltration addressed" isn't a judgment call, it's a mechanical check against whether L3, L4, and L5 all clear a score of 2 or better. And critically, it comes with a composition framework: no single product covers all seven layers well, so the practical guidance is to stack products and take the maximum score at each layer, rather than pretend one tool solves everything.&lt;/p&gt;

&lt;p&gt;This isn't a thought experiment either. The project ships &lt;code&gt;products.yaml&lt;/code&gt; with 26 real products scored, a Go-based verification probe you can run against an actual sandbox to check the claims, and an &lt;a href="https://ast.georgebuilds.dev/" rel="noopener noreferrer"&gt;interactive explorer&lt;/a&gt; for comparing fingerprints side by side.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Like
&lt;/h2&gt;

&lt;p&gt;The Strength scale's cooperative-versus-kernel-versus-structural distinction is, almost exactly, the sentence I wanted to write in my original hot take about the OpenAI incident: a proxy the process can route around is a fundamentally weaker guarantee than a boundary enforced below the application layer, and this taxonomy gives that distinction a name and a number instead of leaving it as vibes. I also like that it's honest about its own limits, the project's README says outright that it hasn't been community-reviewed yet and flags which product scores are unverified. And the composition framework's core finding, that products cluster into "build a better box" or "control what's inside the box" and almost nobody does both, matches exactly what a default-deny, kernel-enforced network policy is supposed to fix in &lt;a href="https://dev.to/posts/updating-coder-to-get-user-secrets-and-the-art-of-knowing-where-your-secrets-belong"&gt;how we think about secrets belonging below the application layer&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Don't Like
&lt;/h2&gt;

&lt;p&gt;Two gaps stood out once I tried to actually use it. First, there's no layer for what happens &lt;em&gt;after&lt;/em&gt; containment fails, no kill switch, no automatic credential rotation on trigger, no forensic rollback runbook. The taxonomy tells you whether an agent could plausibly be contained, not what happens in the window after it isn't, and that window is exactly where the OpenAI incident lived. Second, L7 Observability scores whether logging exists and how tamper-evident it is, but not how fast an anomaly actually reaches a human. A system can score well on L7 and still take days to notice something, which is close to what appears to have happened here. Neither gap is a knock on the project, they're reasonable scoping choices for a v1, but they're the two places I'd want the standard to grow next.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running the Incident Through It
&lt;/h2&gt;

&lt;p&gt;Here's the actual exercise: can I fingerprint the OpenAI/Hugging Face eval environment using &lt;a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/" rel="noopener noreferrer"&gt;public disclosures&lt;/a&gt; alone? Mostly no, and that gap is itself the finding.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Estimated&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;L1 Compute Isolation&lt;/td&gt;
&lt;td&gt;Undisclosed&lt;/td&gt;
&lt;td&gt;Described only as "highly isolated"; no mechanism named&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;L2 Resource Limits&lt;/td&gt;
&lt;td&gt;Undisclosed&lt;/td&gt;
&lt;td&gt;Not addressed in the public writeup&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;L3 Filesystem Boundary&lt;/td&gt;
&lt;td&gt;Undisclosed&lt;/td&gt;
&lt;td&gt;Not addressed in the public writeup&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;L4 Network Boundary&lt;/td&gt;
&lt;td&gt;~2, and it didn't hold&lt;/td&gt;
&lt;td&gt;A single proxy chokepoint was the sanctioned exit; a zero-day in it gave code execution, which is exactly the failure mode the taxonomy's S:2 tier warns is reconfigurable but not kernel-irreversible&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;L5 Credentials&lt;/td&gt;
&lt;td&gt;Likely 0–1&lt;/td&gt;
&lt;td&gt;Stolen credentials enabled lateral movement into Hugging Face's production systems, consistent with ambient, reachable credentials rather than a proxy or ephemeral tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;L6 Action Governance&lt;/td&gt;
&lt;td&gt;~0–1, by design&lt;/td&gt;
&lt;td&gt;Guardrails were deliberately stripped for this evaluation to measure maximum capability&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;L7 Observability&lt;/td&gt;
&lt;td&gt;Partial&lt;/td&gt;
&lt;td&gt;Logging existed and Hugging Face detected the intrusion independently, but it took five days to connect it to OpenAI's own telemetry, the exact latency gap the taxonomy doesn't currently score&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Run that through the taxonomy's own threat rules and T1 (data exfiltration, which needs L3, L4, and L5 all at 2 or better) can't be marked addressed with what's public, not because it's confirmed to have failed everywhere, but because two of the three inputs were never disclosed. That's the actual value of doing this exercise: it doesn't let you conclude "OpenAI's sandbox was bad," which isn't fair to assert from the outside. It lets you say precisely which of seven specific, falsifiable claims about the containment architecture were never made public in the first place. That's a more useful sentence than either extreme, uncritical trust or a reflexive pile-on.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If your own agent's sandbox had to be fingerprinted against these seven layers in public, how many of the seven could you actually answer?&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next: Running the Experiment
&lt;/h2&gt;

&lt;p&gt;So, is there a standard definition of "sandboxed"? Ends up, yes, close enough to count. But reading a fingerprint format is one thing; trusting it is another, especially when nearly every entry in the taxonomy's own &lt;code&gt;products.yaml&lt;/code&gt; carries &lt;code&gt;evidence_level: docs&lt;/code&gt;, meaning it was inferred from documentation and marketing pages, not hands-on testing. As with all research on this blog, the next step isn't another opinion, it's an experiment.&lt;/p&gt;

&lt;p&gt;The plan: before we trust our own results, we validate the tool itself. The project ships &lt;code&gt;ast-probe&lt;/code&gt;, a binary you drop inside a live sandbox to get a verified fingerprint instead of a documentation-based guess. We're going to run it against a few products already scored in the dataset first and check whether we reproduce the taxonomy's own published numbers. If we can't, that's a finding about the probe or the scoring, and it needs to get sorted before we trust anything downstream of it.&lt;/p&gt;

&lt;p&gt;Once that check holds, we'll run the same probe against a live Coder workspace configured the way we actually run agent tooling, our default-deny kernel network policy, our credential handling, the whole stack, and publish the resulting fingerprint. If the reproduction holds up, we'll also open a PR against the taxonomy's &lt;code&gt;products.yaml&lt;/code&gt; with &lt;code&gt;evidence_level: verified&lt;/code&gt; instead of the default &lt;code&gt;docs&lt;/code&gt;, since vendor self-assessments backed by probe output are explicitly what the project asks contributors to submit.&lt;/p&gt;

&lt;p&gt;That's the actual next post: not a take, an experiment, with a scorecard at the end.&lt;/p&gt;

&lt;h2&gt;
  
  
  By the Numbers
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;1&lt;/strong&gt; — Thursday Thoughts hot take that got shelved mid-draft once I asked whether the definition already existed&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;7-7-3&lt;/strong&gt; — the taxonomy's own shorthand: seven defense layers, seven threat categories, three evaluation dimensions&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;26&lt;/strong&gt; — real products scored in the taxonomy's dataset, all as of its March 2026 v1.0 release&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;5 days&lt;/strong&gt; — the gap between Hugging Face detecting the intrusion and OpenAI publicly connecting it to its own testing, the exact latency the taxonomy doesn't currently score&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;3&lt;/strong&gt; — of seven layers we could confidently estimate from OpenAI's public disclosure; the other four are simply unknown&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2&lt;/strong&gt; — gaps I'd want fixed in v2: an incident-response/kill-switch layer, and a detection-latency sub-score&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;1&lt;/strong&gt; — sandbox we're actually going to run the probe against next: our own&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;0&lt;/strong&gt; — new standards invented in this post, on purpose&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>meta</category>
      <category>buildinginpublic</category>
      <category>security</category>
      <category>ai</category>
    </item>
    <item>
      <title>Thursday Thoughts: Curiosity, Not Skill, Is the Real AI Divide</title>
      <dc:creator>Rob</dc:creator>
      <pubDate>Thu, 23 Jul 2026 16:54:46 +0000</pubDate>
      <link>https://dev.to/carryologist/thursday-thoughts-curiosity-not-skill-is-the-real-ai-divide-4mg1</link>
      <guid>https://dev.to/carryologist/thursday-thoughts-curiosity-not-skill-is-the-real-ai-divide-4mg1</guid>
      <description>&lt;p&gt;Six months ago I started maintaining this site as a side hustle. Two to three hours a week, mostly nights and weekends, including all the hobbyist content around it. Not a lot of time. And yet somewhere in that process I started noticing something I didn't expect: I was actually learning software engineering.&lt;/p&gt;

&lt;p&gt;Not in a "here are the fundamentals of computer science" kind of way. More like the way you learn things when you're on the job and something breaks and you have to figure out why. Except compressed. Weirdly, uncomfortably compressed into what should have been a pretty shallow experience of just poking at an AI until a website works.&lt;/p&gt;

&lt;p&gt;That tension is worth pulling on.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Sharp Edges Show Up Fast
&lt;/h2&gt;

&lt;p&gt;When you're actually maintaining an app, even a simple one, you start finding bugs. Some are obvious. Some are hiding. And what you begin to realize is that every app has these sharp edges, places where things can go wrong, where attack surfaces open up, where assumptions you made at the start turn out to be wrong. As a working software engineer, you'd know this intuitively because you've spent years pattern-matching against exactly these situations.&lt;/p&gt;

&lt;p&gt;As a vibe coder, you find out the same way. You just find out faster.&lt;/p&gt;

&lt;p&gt;I've been talking to my agent in plain language: "Hey, can you check to make sure this is actually doing what I think it's doing?" And what happens is the agent goes and builds a test harness, runs a scan, looks for the thing I was vaguely worried about, &lt;a href="https://dev.to/posts/auditing-the-surface-we-added-since-the-last-audit"&gt;the same discipline that turns into an actual post here&lt;/a&gt; most weeks. The discipline is baked in. I don't have to know the name of the methodology. I just have to have enough awareness to ask the question.&lt;/p&gt;

&lt;p&gt;That's a real shift. The abstraction isn't just "natural language instead of code." It's natural language instead of years of accumulated disciplinary knowledge about how to check your work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Watching the Agent Teaches You How to Think
&lt;/h2&gt;

&lt;p&gt;Here's the part that surprised me most. Because I can see what the agent is doing, step by step, I'm actually learning. I'm learning how to decompose a problem into smaller chunks. I'm learning which kinds of tasks the agent handles well and which ones it fumbles. I'm developing intuitions about where things might go sideways before I ask it to look.&lt;/p&gt;

&lt;p&gt;Those intuitions feel like software engineering instincts. Not fully formed ones. But genuine ones. The kind that would have taken me years to build through traditional on-the-job experience.&lt;/p&gt;

&lt;p&gt;It's also teaching me to think across disciplines simultaneously. Security, testing, architecture, code quality: these used to be distinct specializations that people spent careers developing. Now I'm getting exposure to all of them at once because my agent is navigating all of them at once, and I'm watching it do it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The AI Have-Nots Gap Is Really a Curiosity Gap
&lt;/h2&gt;

&lt;p&gt;I've started thinking about who's getting left behind as AI accelerates, and I don't think the dividing line is what most people assume. It's not technical background. It's not access to tools. It's curiosity. It's the same reframing behind &lt;a href="https://dev.to/posts/thursday-thoughts-every-intern-is-a-builder-now"&gt;a finance intern spending her summer vibe coding automation instead of shadowing someone&lt;/a&gt;: the on-ramp was never the CS degree, it was the willingness to jump in.&lt;/p&gt;

&lt;p&gt;If you're willing to jump in, to tinker, to accept that you're going to hit bugs and weird edge cases and moments where you genuinely don't know what just happened, then AI gives you this hyper-abbreviated version of on-the-job learning. You pick up skills fast. You develop judgment. You start building things that would have been out of reach.&lt;/p&gt;

&lt;p&gt;If you're not curious, if you're waiting for AI to feel safe and obvious and simple before you engage with it, the technology is moving past you. Not because you're incapable. Because you're not in motion.&lt;/p&gt;

&lt;p&gt;I think about this a lot when people ask me whether AI is going to displace jobs. My honest answer is: not the way most people fear. What I think we're actually entering is a period of massive expansion in how much software gets built, and who builds it, and what kinds of problems get solved. The unlock isn't replacing engineers. It's making it possible for someone like me, spending two hours a week on a side project, to develop real engineering intuitions through practice.&lt;/p&gt;




&lt;p&gt;That's a different story than the one about displacement. It's a story about democratization. Skills that used to require a computer science degree or years of mentorship are becoming accessible through curiosity and a willingness to mess around and pay attention to what happens.&lt;/p&gt;

&lt;p&gt;I don't think that means the craft of software engineering stops mattering. If anything, watching my agent work has made me more interested in the fundamentals, not less. But the on-ramp has changed dramatically. You just need to jump in.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Are you learning things from your AI tools that you didn't expect to learn?&lt;/em&gt;&lt;/p&gt;

</description>
      <category>meta</category>
      <category>buildinginpublic</category>
      <category>vibecoding</category>
      <category>ai</category>
    </item>
    <item>
      <title>Auditing the Surface We Added Since the Last Audit</title>
      <dc:creator>Rob</dc:creator>
      <pubDate>Wed, 22 Jul 2026 21:30:29 +0000</pubDate>
      <link>https://dev.to/carryologist/auditing-the-surface-we-added-since-the-last-audit-1ej6</link>
      <guid>https://dev.to/carryologist/auditing-the-surface-we-added-since-the-last-audit-1ej6</guid>
      <description>&lt;p&gt;Staying hyper-vigilant doesn't come with a calendar reminder that tells you when to renew it. The May audit ended with a clean bill and a note: "next scheduled review, August 8, 2026." I didn't wait for August. Two months of shipping later, I asked for a fresh scan against the same categories that audit used, and the result was smaller than &lt;a href="https://dev.to/posts/closing-the-loop-from-audit-to-ten-commits"&gt;the first one&lt;/a&gt; but not nothing: 8 findings, three of them in code that didn't exist in May.&lt;/p&gt;

&lt;p&gt;That's the actual headline here, more than any individual bug. The first audit covered the blog as it stood in early May. Since then we shipped an MCP server (16 tools, bearer-token auth, its own rate limiter), a Slack &lt;code&gt;/todo&lt;/code&gt; integration, and a shareable-snippet image generator for code blocks and tables. None of that existed when the last audit ran, which means none of it had ever been looked at with a security lens. A codebase doesn't need to get worse to need a second look. It just needs to get bigger. &lt;a href="https://dev.to/posts/friday-fixes-the-hyper-vigilance-tax"&gt;Friday's post&lt;/a&gt; named this the hyper-vigilance tax and paid it on the UX axis: three small bugs, each hiding in a corner some check didn't cover. This post pays the same tax on the other axis, the one that doesn't announce itself in a screenshot and doesn't get noticed by accident. Different axis, same bill.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Scan Covered
&lt;/h2&gt;

&lt;p&gt;Same shape as May: auth and middleware, every API route, GitHub Actions workflows in both repos, &lt;code&gt;npm audit&lt;/code&gt;, and a secrets/PII grep across both repos. The difference this time was scope creep in a good way — three routes that simply didn't exist during the first pass got the same scrutiny as the routes that did.&lt;/p&gt;

&lt;p&gt;Eight findings, roughly matching the severity spread from last time:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;1 high: a transitive dependency with five stacked advisories&lt;/li&gt;
&lt;li&gt;5 moderate: the rest of a dependency chain, plus two independent route-level gaps&lt;/li&gt;
&lt;li&gt;2 low: a dev-only dependency advisory, and a documented (not code) tradeoff&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Phase 1: Dependency Bumps
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;npm audit&lt;/code&gt; on a real &lt;code&gt;npm install&lt;/code&gt; (not a stale lockfile check) turned up 8 vulnerabilities: 1 high, 6 moderate, 1 low. The high was &lt;code&gt;hono@4.12.22&lt;/code&gt;, pulled in transitively through &lt;code&gt;@modelcontextprotocol/sdk&lt;/code&gt;, which the MCP server depends on. Five stacked advisories on that one package: a CORS-reflects-any-origin-with-credentials bug, a body-limit bypass on Lambda-style deployments, a path-traversal issue in &lt;code&gt;serve-static&lt;/code&gt; on Windows, and two adapter bugs that silently drop cookies or headers.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;npm audit fix&lt;/code&gt; — the non-forcing kind — resolved all of it except one chain: &lt;code&gt;hono&lt;/code&gt; jumped to &lt;code&gt;4.12.30&lt;/code&gt; (patched), and &lt;code&gt;brace-expansion&lt;/code&gt;, &lt;code&gt;js-yaml&lt;/code&gt; (including the copy that &lt;code&gt;gray-matter&lt;/code&gt; depends on, which matters because &lt;code&gt;gray-matter&lt;/code&gt; parses every post's frontmatter in production), and &lt;code&gt;@babel/core&lt;/code&gt; all resolved within their existing semver ranges.&lt;/p&gt;

&lt;p&gt;What's left is the exact same false positive the May audit documented: &lt;code&gt;postcss &amp;lt;8.5.10&lt;/code&gt;, bundled inside &lt;code&gt;next&lt;/code&gt; itself, not the &lt;code&gt;@tailwindcss/postcss&lt;/code&gt; copy (already patched). &lt;code&gt;npm audit fix --force&lt;/code&gt; proposes downgrading &lt;code&gt;next&lt;/code&gt; to &lt;code&gt;9.3.3&lt;/code&gt; to fix it — a multi-year regression that would break considerably more than it fixes. I checked whether a newer Next.js release had bumped its bundled &lt;code&gt;postcss&lt;/code&gt;; even the &lt;code&gt;16.3.0&lt;/code&gt; preview builds haven't. Same conclusion as May: accepted, monitored risk, no fix available yet that isn't worse than the bug.&lt;/p&gt;

&lt;p&gt;One commit. &lt;code&gt;tsc --noEmit&lt;/code&gt; clean afterward.&lt;/p&gt;

&lt;h2&gt;
  
  
  Phase 2: Three Independent Route Gaps
&lt;/h2&gt;

&lt;p&gt;None of these three depend on each other, so they went into one batched commit — the same logic the last audit used for its Phase 2.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The MCP route's rate limiter was in-memory.&lt;/strong&gt; &lt;code&gt;const buckets = new Map&amp;lt;string, {...}&amp;gt;()&lt;/code&gt;, keyed by IP, capped at 120 requests/minute. That looks reasonable until you remember Vercel runs serverless functions across multiple instances. Each cold start gets its own empty &lt;code&gt;Map&lt;/code&gt;. A client that happens to land on five different instances effectively gets five times the stated limit — the number on the tin was never the number in practice. Swapped it for the same Redis-backed limiter the login and analytics routes already use, which is shared across every instance regardless of which one handles a given request.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;/api/share-image&lt;/code&gt; had no bounds and no rate limit.&lt;/strong&gt; This route is intentionally public — it's called client-side to render shareable images of code blocks and tables for social sharing, so it can't sit behind the admin-session middleware. But it accepted an arbitrary-length &lt;code&gt;content&lt;/code&gt; string and computed the output image's height with no ceiling; only width was capped. That's a real rendering-cost DoS sitting in the open: no authentication, no throttling, no size limit, on an endpoint whose cost scales with attacker-supplied input. Added a per-IP rate limit (20/minute via the same Redis limiter), a 20,000-character content cap, a 200-row table cap, and a 4,000px height ceiling to match the existing width cap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Both Dev.to syndication routes skipped slug sanitization.&lt;/strong&gt; &lt;code&gt;content/posts/${slug}.mdx&lt;/code&gt; with a raw, unsanitized &lt;code&gt;slug&lt;/code&gt; straight from the request body, feeding into a GitHub Contents API path. Every other route touching the same file space — post editing, image upload, every MCP tool — validates the slug first. These two just got missed, probably because they were added after the pattern was established elsewhere and nobody thought to check whether they'd inherited it. Since three separate files each had their own copy-pasted &lt;code&gt;sanitizeSlug()&lt;/code&gt;, and a fourth and fifth were missing it entirely, I pulled the function into &lt;code&gt;src/lib/slug.ts&lt;/code&gt; once and pointed all five call sites at it. Low impact today, since both routes sit behind admin-session middleware, but it's the kind of inconsistency that becomes a real path-traversal bug the moment the trust boundary shifts even slightly.&lt;/p&gt;

&lt;p&gt;Verified with &lt;code&gt;tsc --noEmit&lt;/code&gt;, &lt;code&gt;eslint&lt;/code&gt; on every changed file, and a full production build. One unrelated finding surfaced during lint: &lt;code&gt;share-image/route.tsx&lt;/code&gt; has 10 pre-existing lint errors (JSX constructed inside a try/catch, which ESLint's &lt;code&gt;react-hooks/error-boundaries&lt;/code&gt; rule flags because React doesn't actually catch render errors that way). Confirmed via &lt;code&gt;git stash&lt;/code&gt; that they predate this change entirely — not something to silently fix inside a security commit, so it's noted here and left for its own pass.&lt;/p&gt;

&lt;h2&gt;
  
  
  Phase 3: Finishing What Report-Only Started
&lt;/h2&gt;

&lt;p&gt;The May audit's Phase 2 shipped Content-Security-Policy in Report-Only mode on purpose, with a plan to observe for about a week, then flip to enforcing. That flip never happened. It sat in Report-Only for two months.&lt;/p&gt;

&lt;p&gt;Here's the part that made this an easy call: there's no &lt;code&gt;report-to&lt;/code&gt; or &lt;code&gt;report-uri&lt;/code&gt; directive configured anywhere in the policy. Report-Only mode without a reporting endpoint doesn't collect anything except what shows up in an individual visitor's own browser console — which nobody but the site owner would ever open, and even then only by accident. The "observe for a week" plan had no mechanism to observe anything. Two months of waiting produced exactly as much signal as two minutes would have.&lt;/p&gt;

&lt;p&gt;So: flipped the header from &lt;code&gt;Content-Security-Policy-Report-Only&lt;/code&gt; to &lt;code&gt;Content-Security-Policy&lt;/code&gt;, same directive set, unchanged since May. Rather than trust the diff, I ran a full production build, started the built server, and &lt;code&gt;curl -I&lt;/code&gt;'d the homepage to confirm the actual response header. It came back exactly as expected — enforcing, same values. This one got its own commit, isolated from Phase 2, because it's a global behavior change with real blast radius if a directive gap exists that Report-Only never had the means to catch. Easy to revert on its own if something breaks that two months of silent Report-Only never revealed.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Got Deferred
&lt;/h2&gt;

&lt;p&gt;One item, carried forward rather than fixed: the rate limiter (shared across login, analytics, MCP, and now share-image) fails open if Upstash Redis is unreachable or misconfigured. That's a deliberate, documented tradeoff from when the limiter was first built — a preview environment without Redis configured shouldn't get locked out of its own login page. It's still the right tradeoff. But it means a silent Redis misconfiguration in production would silently remove every rate limit on the site with no alert. Not a code fix; a monitoring gap. Noted for whenever alerting gets built out, the same way the last audit deferred CSP nonces and build-time markdown rendering with reasons instead of silently dropping them.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Says About Auditing Cadence
&lt;/h2&gt;

&lt;p&gt;The real lesson isn't in any individual bug. It's that "audit the codebase" isn't a checkbox you tick once. The MCP server, the Slack integration, and the share-image route all shipped in the two months between audits, each one reasonably reviewed on its own merits at the time, and none of them got the systematic security pass the rest of the codebase got in May — because that pass had already happened before they existed.&lt;/p&gt;

&lt;p&gt;A quarterly audit cadence, which is what the May report suggested, assumes the codebase's rate of change is roughly constant. For a one-person-plus-agent blog shipping new integrations every few weeks, three months is long enough for entire new subsystems to exist unaudited. That's the same structural blind spot &lt;a href="https://dev.to/posts/your-ai-strategy-has-a-blind-spot"&gt;a very different audit&lt;/a&gt; found in this blog's SEO and AEO tooling months ago: a checker that only looks where it was built to look will miss whatever got added after it was written. The actual cadence that matches this project isn't a calendar date. It's "whenever a new API route ships that talks to the outside world," which in practice has been happening faster than the calendar suggested.&lt;/p&gt;

&lt;h2&gt;
  
  
  By the Numbers
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;8&lt;/strong&gt; findings this pass, versus &lt;strong&gt;90+&lt;/strong&gt; raw findings (deduplicated to 15) in May — this audit had one scanner instead of three, and a much smaller diff to cover&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;3&lt;/strong&gt; of 8 findings were in code that didn't exist during the May audit (MCP rate limiter, share-image bounds, and the Dev.to slug gap, added alongside newer syndication tooling)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;5&lt;/strong&gt; stacked advisories on a single transitive dependency (&lt;code&gt;hono&lt;/code&gt;), resolved by one non-forcing &lt;code&gt;npm audit fix&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;8 → 4&lt;/strong&gt; vulnerabilities after Phase 1, all four remaining from the same root cause (&lt;code&gt;postcss&lt;/code&gt; bundled inside &lt;code&gt;next&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;3&lt;/strong&gt; phases, &lt;strong&gt;3&lt;/strong&gt; commits, &lt;strong&gt;1&lt;/strong&gt; repo — no content-repo changes this round&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;0&lt;/strong&gt; new dependencies added to fix anything&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;~1 hour&lt;/strong&gt; end to end, versus &lt;strong&gt;4 hours&lt;/strong&gt; for the May audit's ten commits across two repos&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2 months&lt;/strong&gt; a CSP policy sat in Report-Only mode with no reporting endpoint configured, collecting zero actual violation data&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;1&lt;/strong&gt; shared &lt;code&gt;sanitizeSlug()&lt;/code&gt; replacing &lt;strong&gt;3&lt;/strong&gt; copy-pasted versions and closing &lt;strong&gt;2&lt;/strong&gt; missing ones&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;10&lt;/strong&gt; pre-existing lint errors found, confirmed unrelated via &lt;code&gt;git stash&lt;/code&gt;, and left alone rather than folded into a security commit&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;1&lt;/strong&gt; deferred item, carried forward with a reason, not silently dropped&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>security</category>
      <category>agents</category>
      <category>buildinginpublic</category>
      <category>meta</category>
    </item>
    <item>
      <title>Has This Blog Been a Waste of Time? (I Made My Coding Agent Investigate)</title>
      <dc:creator>Rob</dc:creator>
      <pubDate>Wed, 22 Jul 2026 21:30:15 +0000</pubDate>
      <link>https://dev.to/carryologist/has-this-blog-been-a-waste-of-time-i-made-my-coding-agent-investigate-409e</link>
      <guid>https://dev.to/carryologist/has-this-blog-been-a-waste-of-time-i-made-my-coding-agent-investigate-409e</guid>
      <description>&lt;p&gt;I gave my coding agent an uncomfortable assignment this week: go find out if we've been wasting our time.&lt;/p&gt;

&lt;p&gt;Every Model Showdown round on this blog runs on tooling we wrote ourselves — a bakeoff harness, a scoring rubric, &lt;a href="https://dev.to/posts/comfyui-lemonade-and-localai-scouting-the-next-wave-of-homelab-ai-tools"&gt;&lt;code&gt;thermal-test.sh&lt;/code&gt;&lt;/a&gt;, ad hoc token counting glued together across a dozen posts. My coding agent is also my research partner for all of it. Both of us are, structurally, coders. And I've started to wonder whether that's a bias, not a strength: if the tool you're best at is writing code, every problem you're handed starts to look like a problem you solve by writing more of it. Want to test consumer AI hardware? Build a harness. Sure, we repurpose libraries here and there, but we're still stitching it all together by hand, every time, instead of asking whether someone already solved this.&lt;/p&gt;

&lt;p&gt;So the assignment was research, not code: go find out what hobbyists — the Linus Tech Tips / Gamers Nexus / Level1Techs crowd — and the broader AI benchmarking world already publish. Then tell me honestly whether we should have been using any of it instead of building our own.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Already Out There
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The hardware-review houses.&lt;/strong&gt; &lt;a href="https://github.com/LTTLabsOSS/markbench-tests" rel="noopener noreferrer"&gt;LTT Labs runs MarkBench&lt;/a&gt;, an orchestration and data-collection framework whose actual test harnesses — the same code that generates the numbers in LTT videos — are open-sourced on GitHub and updated on a quarterly cadence. It's built for GPUs and games, not LLMs: one harness scripts PyAutoGUI to click through a menu, another runs an OCR service just to find text on screen. &lt;a href="https://gamersnexus.net/features/living-doc-current-test-bench-hardware-list-methodologies" rel="noopener noreferrer"&gt;Gamers Nexus&lt;/a&gt; takes the opposite approach to the same problem — not a reusable framework, but a public living document of exactly which SOPs, test benches, and settings produced which chart, updated review by review. Neither one has touched an LLM workload. And GN gave me the best counter-argument to my own thesis before I'd even finished asking the question: in 2025 they &lt;a href="https://gamersnexus.net/gpus-gn-extras-cpus/problem-gpu-benchmarks-reality-vs-numbers-animation-error-methodology-white" rel="noopener noreferrer"&gt;published a whole new measurement methodology&lt;/a&gt; — "animation error" — because framerate and frametime testing, the industry standard for a decade, still didn't capture what a stutter actually felt like to a player. Even the best-resourced reviewers in the business sometimes conclude nothing existing measures what they need, and build something new. That's not automatically a bias. Sometimes it's just correct.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The vendor-neutral AI-PC suites.&lt;/strong&gt; &lt;a href="https://mlcommons.org/working-groups/benchmarks/client/" rel="noopener noreferrer"&gt;MLPerf Client&lt;/a&gt;, built by MLCommons with AMD, Intel, Microsoft, NVIDIA, and Qualcomm all at the table, is now on &lt;a href="https://mlcommons.org/2026/04/mlperf-client-v1-6/" rel="noopener noreferrer"&gt;v1.6&lt;/a&gt; and measures how a Windows, macOS, or Linux client handles real generative AI tasks like summarization and content creation. &lt;a href="https://benchmarks.ul.com/procyon/ai-text-generation-benchmark" rel="noopener noreferrer"&gt;UL's Procyon AI Text Generation Benchmark&lt;/a&gt; does the enterprise-press version of the same thing: seven fixed prompts against Phi-3.5-mini, Mistral-7B, Llama-3.1-8B, and Llama-2-13B, license required. Both are real, credible, and completely useless for us — fixed model rosters that don't include a single model we actually run, and both measure single-turn inference, not whether an agent can hold a task together across fifty tool calls.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The tools built for exactly our stack.&lt;/strong&gt; &lt;a href="https://github.com/ggml-org/llama.cpp/blob/master/tools/llama-bench/README.md" rel="noopener noreferrer"&gt;&lt;code&gt;llama-bench&lt;/code&gt;&lt;/a&gt; ships inside llama.cpp itself and is the de facto community standard for raw prompt-processing and token-generation speed. &lt;a href="https://github.com/eugr/llama-benchy" rel="noopener noreferrer"&gt;&lt;code&gt;llama-benchy&lt;/code&gt;&lt;/a&gt; extends that same measurement style to any OpenAI-compatible backend, which is exactly why &lt;a href="https://dev.to/posts/comfyui-lemonade-and-localai-scouting-the-next-wave-of-homelab-ai-tools"&gt;I flagged it as a clean win&lt;/a&gt; back in July and then never actually wired it into a bakeoff. &lt;a href="https://openbenchmarking.org/test/pts/llama-cpp" rel="noopener noreferrer"&gt;OpenBenchmarking.org&lt;/a&gt; crowdsources llama.cpp results from anyone running the Phoronix Test Suite, opt-in, auto-rerunning until the variance settles. None of these three measure anything beyond throughput. None of them were ever going to replace the rubric. But none of them needed to — they're a layer underneath it that we skipped.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Our actual peer group.&lt;/strong&gt; A &lt;a href="https://forum.level1techs.com/t/strix-halo-ryzen-ai-max-395-llm-benchmark-results/233796" rel="noopener noreferrer"&gt;Level1Techs forum member built a Strix Halo LLM benchmark harness&lt;/a&gt; from scratch, ran rigorous sweeps against the newest MoE architectures, and checked the full results into GitHub for the community to pick apart. That's the same move this blog makes every few weeks. It's worth saying plainly: "one hobbyist builds their own harness and publishes it" isn't a personal failing of mine or my agent's habits. It's the normal, respected way this exact community operates.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The layer that actually matters to us.&lt;/strong&gt; Model Showdown isn't a throughput test — it's an agentic-correctness test, and that field looks different. &lt;a href="https://aider.chat/docs/leaderboards/" rel="noopener noreferrer"&gt;Aider's Polyglot benchmark&lt;/a&gt; scores models on 225 of the hardest Exercism exercises inside Aider's own structured edit loop, with community-contributed results merged by PR. &lt;a href="https://www.vals.ai/benchmarks/swebench" rel="noopener noreferrer"&gt;SWE-bench Verified&lt;/a&gt; grades a patch against a real GitHub issue and a human-validated test suite, though a notable complexity is that it evaluates the agentic harness and the underlying model together, which is exactly why different labs report different numbers for the same model. Terminal-Bench and METR's RE-Bench push further into containerized, tool-using, long-horizon work. A recent &lt;a href="https://www.appliedtechnologyindex.com/research/2026-comparative-analysis-coding-agent-evaluation-harnesses-after-swe-bench/" rel="noopener noreferrer"&gt;survey of the post-SWE-bench landscape&lt;/a&gt; put it about as bluntly as I'd put it myself: evaluation is moving toward harness portfolios that test repository repair, terminal execution, and tool-use reliability together, and any single benchmark is one signal, not the whole basis for a decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Honest Answer
&lt;/h2&gt;

&lt;p&gt;Split it in two, because the two halves of our own stack don't get the same grade.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The throughput layer: yes, we reinvented a wheel, and we knew it.&lt;/strong&gt; &lt;code&gt;llama-bench&lt;/code&gt; and &lt;code&gt;llama-benchy&lt;/code&gt; already solve "how fast does this model run on this hardware," drop onto our existing endpoint with zero infrastructure change, and I identified &lt;code&gt;llama-benchy&lt;/code&gt; specifically as a clean win in &lt;a href="https://dev.to/posts/comfyui-lemonade-and-localai-scouting-the-next-wave-of-homelab-ai-tools"&gt;a research post back in July&lt;/a&gt; — and then the next Model Showdown round still reported informal token counting instead. That's not a close call. That's the bias, caught in the act, on our own blog.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The agentic layer: no, not cleanly, and it's not just stubbornness.&lt;/strong&gt; Aider Polyglot is the closest existing match to what &lt;a href="https://dev.to/posts/model-showdown-round-9-qwen-3-6-27b-vs-qwen-3-6-35b-a3b-vs-qwythos-9b-vs-glm-4-7-flash-vs-nemotron-3-nano"&gt;Model Showdown&lt;/a&gt; does, and it's still testing something narrower: isolated Exercism puzzles in a sandbox, one edit-and-retry loop, no real repo, no git, no Playwright, no fifty-turn session that can quietly go sideways. Terminal-Bench is closer in spirit but isn't wired to a Coder Agents chat and doesn't know our rubric. Adopting either wholesale would have meant building nearly as much integration code as our own harness already required, in exchange for a narrower signal. This is the same shape as &lt;a href="https://dev.to/posts/closing-the-loop-from-audit-to-ten-commits"&gt;the security audit work&lt;/a&gt; that's become a habit on this blog: sometimes the "buy vs. build" answer really is build, and the honest version of that answer names the specific reason instead of just defaulting to it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;That split is a claim, not just a vibe, and claims like that should be checked with data instead of left as a nice paragraph. So the next step is an actual experiment: run &lt;code&gt;llama-bench&lt;/code&gt;/&lt;code&gt;llama-benchy&lt;/code&gt; against our own ad hoc timing on the same model, run a slice of Aider's Polyglot benchmark against two models from a recent Model Showdown round side by side with our own rubric, and keep an honest ledger of how much setup time each approach cost — because if wiring up someone else's harness takes longer than just running another homegrown round, that's evidence for the bias too, independent of whose numbers turn out to be more useful.&lt;/p&gt;

&lt;p&gt;I'm predicting the existing tools win the throughput half outright and the homegrown rubric keeps finding things Polyglot structurally can't. I could be wrong about that. Next post will say so either way.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Back to the thesis. More soon.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>homelab</category>
      <category>ai</category>
      <category>llm</category>
      <category>benchmark</category>
    </item>
    <item>
      <title>Friday Fixes: The Hyper-Vigilance Tax</title>
      <dc:creator>Rob</dc:creator>
      <pubDate>Fri, 17 Jul 2026 16:01:06 +0000</pubDate>
      <link>https://dev.to/carryologist/friday-fixes-the-hyper-vigilance-tax-2j8m</link>
      <guid>https://dev.to/carryologist/friday-fixes-the-hyper-vigilance-tax-2j8m</guid>
      <description>&lt;p&gt;Building an app across dozens of disconnected agent sessions accumulates bugs. That's not a failure of the process, it's the process working as advertised. Each session solves the problem in front of it, ships, and moves on. Nobody's holding the whole system in their head across sessions the way one long-tenured engineer might. What slips through the gaps isn't any single session's fault. It's the tax you pay for building this way, and the only way to keep the bill small is staying hyper-vigilant, on two different axes at once, indefinitely, because no session ever hands the watch off to the next one.&lt;/p&gt;

&lt;p&gt;The first axis is UX: does the thing actually work the way a person experiences it, not the way a browser's devtools simulates it. The second axis is security: does it hold up against someone trying to make it fail on purpose, not just against someone using it the way it was intended. Vibe coding doesn't get you out of watching both. It just changes who's watching, and how often you have to look.&lt;/p&gt;

&lt;p&gt;This stretch of days gave me a clean split between the two. Four small bugs on the UX axis, all of them hiding in a corner some check didn't cover. And, separately, a deliberate pass on the security axis that turned up eight more things, three of them in code that didn't exist the last time anyone looked at the codebase that way. This post is the UX side. &lt;a href="https://dev.to/posts/auditing-the-surface-we-added-since-the-last-audit"&gt;Monday's post&lt;/a&gt; is the security side, phases and commits and all. Different axis, different post, same underlying discipline: don't assume yesterday's check still covers today's code.&lt;/p&gt;

&lt;p&gt;Four bugs this round, none of them dramatic on their own. What connects them is more interesting than any single fix: each one passed a real check and failed on the one dimension that check didn't cover. A layout that looked right in devtools. A relative-time label that did real date math, just the wrong kind. An orphan-detection tool that scans exactly half the places an image can be referenced from. A systemd service that only knew how to survive the kind of change it was written for. Same shape every time — the corner nobody checked is the corner where the bug was hiding.&lt;/p&gt;

&lt;p&gt;This is the same territory I keep coming back to in this series: &lt;a href="https://dev.to/posts/friday-fixes-the-fix-that-wasnt"&gt;defenses that feel complete but aren't&lt;/a&gt;, &lt;a href="https://dev.to/posts/friday-fixes-two-bugs-one-workflow"&gt;bugs invisible until a specific condition exposes them&lt;/a&gt;, and &lt;a href="https://dev.to/posts/invisible-failures-the-bugs-that-hide-in-plain-sight"&gt;failures with no error message anywhere in the chain&lt;/a&gt;. Different bugs, same lesson repeating itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. The Draft Card That Only Worked in a Wide Window
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The problem:&lt;/strong&gt; Each card on &lt;code&gt;/admin/drafts&lt;/code&gt; laid out post metadata and three action buttons (Unschedule / Edit / Publish) in a single &lt;code&gt;flex items-start justify-between&lt;/code&gt; row. On an actual iPhone, the buttons didn't have room to share that row with the content, and the layout squished and overflowed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt; Two lines in &lt;code&gt;DraftsList.tsx&lt;/code&gt;. Stack the card vertically on mobile, revert to side-by-side above the &lt;code&gt;sm&lt;/code&gt; breakpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;&lt;span class="gd"&gt;-&amp;lt;div className="flex items-start justify-between"&amp;gt;
&lt;/span&gt;&lt;span class="gi"&gt;+&amp;lt;div className="flex flex-col gap-3 sm:flex-row sm:items-start sm:justify-between sm:gap-4"&amp;gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Plus &lt;code&gt;flex-wrap&lt;/code&gt; on the actions container as a second line of defense.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The corner nobody checked:&lt;/strong&gt; Browser devtools' narrow-viewport mode. It renders at the right &lt;em&gt;width&lt;/em&gt;, but it doesn't reproduce real mobile font rendering, tap-target sizing, or how three buttons with real labels actually wrap. The desktop-simulated-as-mobile view looked fine. The actual phone didn't. This is the exact lesson from &lt;a href="https://dev.to/posts/friday-fixes-mobile-first-and-the-skill-that-saved-us"&gt;Mobile First and the Skill That Saved Us&lt;/a&gt; two months ago — the skill file has the pixel math now, but pixel math doesn't catch every layout shape, and this one slipped through anyway. Worth a real phone check any time buttons and content share a flex row, full stop.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. The Label That Used the Wrong Clock
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The problem:&lt;/strong&gt; The drafts page showed a scheduled post as "publishes Jul 13, 2026 (today)" — while it was still July 12 in the browser's local timezone. The date itself was correct. The &lt;code&gt;(today)&lt;/code&gt; next to it wasn't.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The cause:&lt;/strong&gt; &lt;code&gt;relativeTime()&lt;/code&gt; in &lt;code&gt;DraftsList.tsx&lt;/code&gt; computed the label by diffing raw milliseconds between now and the target, then rounding to the nearest day:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// before&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;diffMs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;target&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getTime&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;now&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getTime&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;diffDays&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;diffMs&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1000&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;60&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;60&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;24&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A post scheduled for 4 AM the next day is only 8–12 hours away in the evening. Divide that by 24 hours and round, and you get 0 — "today" — even though the calendar day hasn't turned over yet in the viewer's timezone. Meanwhile &lt;code&gt;formatDate()&lt;/code&gt;, rendering the actual date right next to it, was computing calendar fields correctly. Two functions, same UI row, two different definitions of "what day is it."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt; Zero both dates to local midnight before diffing, so the comparison is calendar days, not elapsed time:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;startOfDay&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;d&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt;
  &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getFullYear&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="nx"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getMonth&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="nx"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getDate&lt;/span&gt;&lt;span class="p"&gt;()).&lt;/span&gt;&lt;span class="nf"&gt;getTime&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;diffDays&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;startOfDay&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;target&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nf"&gt;startOfDay&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;now&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1000&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;60&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;60&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;24&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Merged as &lt;a href="https://github.com/carryologist/the-vibe-coder/pull/22" rel="noopener noreferrer"&gt;PR #22&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The corner nobody checked:&lt;/strong&gt; Anything that only visibly disagrees near a day boundary. &lt;code&gt;relativeTime()&lt;/code&gt; and &lt;code&gt;formatDate()&lt;/code&gt; had been sitting next to each other rendering the same post for weeks, agreeing with each other every single time, right up until a scheduled post happened to fall in that 8–12 hour evening window. That's the same failure shape as the &lt;a href="https://dev.to/posts/friday-fixes-the-unquoted-date-that-broke-drafts"&gt;unquoted-date bug from Friday Fixes #2&lt;/a&gt; — one function handling dates one way, an adjacent one handling them another way, and nothing forcing them to agree until the specific edge case that exposes the gap.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. The Headshot That Was Deleted as an Orphan
&lt;/h2&gt;

&lt;p&gt;This one's the interesting one, because it wasn't a bug in new code. It was a bug in a &lt;em&gt;cleanup&lt;/em&gt; tool, and it sat undetected for 77 days.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The symptom:&lt;/strong&gt; The &lt;a href="https://dev.to/posts/your-ai-strategy-has-a-blind-spot"&gt;About page&lt;/a&gt; rendered a broken-image icon where the headshot should be. Not a recent regression — this has apparently been broken since spring.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The cause:&lt;/strong&gt; Back on April 30, 2026, a housekeeping pass deleted &lt;code&gt;public/images/IMG_9133.jpeg&lt;/code&gt; as "orphaned, unreferenced." The admin's orphan-detection logic in &lt;code&gt;src/lib/images.ts&lt;/code&gt; only scans &lt;em&gt;subdirectories&lt;/em&gt; of &lt;code&gt;public/images/&lt;/code&gt;, matching each directory's slug against post slugs and post content — the same 3-tier match (exact / prefix / content) the &lt;code&gt;/admin/images&lt;/code&gt; page uses today. &lt;code&gt;IMG_9133.jpeg&lt;/code&gt; was a loose file sitting directly at the top level of &lt;code&gt;public/images/&lt;/code&gt;, referenced only from a hardcoded &lt;code&gt;&amp;lt;Image src="/images/IMG_9133.jpeg"&amp;gt;&lt;/code&gt; in the About page's static &lt;code&gt;page.tsx&lt;/code&gt;. It's not in a post. It's not in a slug directory. The detector never had a chance to see it as in-use, in either direction — it wouldn't have shown up in the admin UI to flag as orphaned &lt;em&gt;or&lt;/em&gt; as matched. Whatever process ran that April cleanup just saw a stray file nobody's tooling could vouch for and deleted it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt; Recovered the original file intact from git history (&lt;code&gt;2dec8d1^&lt;/code&gt;, pre-deletion), moved it into &lt;code&gt;public/images/branding/&lt;/code&gt; alongside the site's other non-post assets — logo, favicons — instead of leaving it as a loose top-level file, and updated the About page's &lt;code&gt;src&lt;/code&gt; path to match.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;&lt;span class="gd"&gt;-src="/images/IMG_9133.jpeg"
&lt;/span&gt;&lt;span class="gi"&gt;+src="/images/branding/IMG_9133.jpeg"
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The corner nobody checked:&lt;/strong&gt; Static pages. The entire orphan/match system in &lt;code&gt;images.ts&lt;/code&gt; reasons about one thing — MDX post content — and the About page isn't a post, it's a hand-written React component. Any image referenced only from &lt;code&gt;src/app/**/page.tsx&lt;/code&gt; is structurally invisible to a detector built around post slugs. I logged this gap in &lt;code&gt;TODO.md&lt;/code&gt; rather than just patching around it once: either the detector needs to also grep static &lt;code&gt;.tsx&lt;/code&gt; pages for hardcoded &lt;code&gt;/images/...&lt;/code&gt; paths, or &lt;code&gt;public/images/&lt;/code&gt; needs a hard rule that it never holds loose top-level files. Until one of those ships, this exact failure mode can recur with any other hardcoded image reference sitting outside the post system — the same blind spot &lt;a href="https://dev.to/posts/your-ai-strategy-has-a-blind-spot"&gt;the SEO/AEO audit&lt;/a&gt; found in a different tool checking a different half of the site.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. The Magic Packet Sent to an Interface That No Longer Existed
&lt;/h2&gt;

&lt;p&gt;This one isn't the blog app at all — it's the homelab itself, &lt;a href="https://dev.to/posts/qol-with-wol-turning-on-the-homelab-from-anywhere"&gt;the workstation these agent sessions run on&lt;/a&gt;. Small bug, same shape, worth the four sentences.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The problem:&lt;/strong&gt; After migrating the homelab to new hardware, Wake on LAN — the SmartThings-to-magic-packet chain from a few months back — stopped waking the machine. Voice command fired, hub still sent the packet, nothing happened.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The cause:&lt;/strong&gt; The &lt;code&gt;wol-enable.service&lt;/code&gt; systemd unit that re-arms WoL on every boot had the old motherboard's interface name, &lt;code&gt;enp8s0&lt;/code&gt;, hardcoded. The new NIC enumerates as &lt;code&gt;enp112s0&lt;/code&gt;. The service had been failing silently on every boot since the migration — &lt;code&gt;systemctl status&lt;/code&gt; showed &lt;code&gt;failed&lt;/code&gt;, &lt;code&gt;netlink error: no device matches name&lt;/code&gt;, sitting there unread for hours.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt; Point the unit at the real interface:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;&lt;span class="gd"&gt;-ExecStart=/sbin/ethtool -s enp8s0 wol g
&lt;/span&gt;&lt;span class="gi"&gt;+ExecStart=/sbin/ethtool -s enp112s0 wol g
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Plus the matching manual step the systemd fix can't reach: updating the MAC address in the SmartThings vWOL device's settings, since the hub was faithfully broadcasting a magic packet addressed to a card that no longer exists.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The corner nobody checked:&lt;/strong&gt; a systemd service that only reports failure in a log nobody was tailing. It existed specifically because I'd been burned by this once before — the original post's whole "make it persist across reboots" step — and it still didn't survive the &lt;em&gt;next&lt;/em&gt; kind of change instead: not a reboot, a hardware swap. Built to survive one axis of change, blind to the other.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Connects Them
&lt;/h2&gt;

&lt;p&gt;All four bugs are the same shape wearing different clothes: a check that covers most of the surface area and misses exactly the part where the failure lives.&lt;/p&gt;

&lt;p&gt;The mobile layout was tested at the right &lt;em&gt;width&lt;/em&gt; but the wrong &lt;em&gt;fidelity&lt;/em&gt; — devtools simulates the viewport, not the device. The relative-time label was tested against the &lt;em&gt;common&lt;/em&gt; case, where the elapsed-time math and the calendar-day math happen to agree, which is true right up until the last few hours before midnight. The orphan detector was tested against &lt;em&gt;posts&lt;/em&gt;, because posts are the thing the admin UI is built to manage, and a static page sitting one directory away in the same repo simply never entered its field of view. The Wake on LAN service was tested against the one kind of change it was written to survive — a reboot — and simply never had to prove itself against the other kind, a hardware swap that renamed the thing it depended on.&lt;/p&gt;

&lt;p&gt;None of these are exotic bugs. They're all the same failure mode I keep writing about in this series: the fix that only covers the layer someone happened to be looking at. The difference this round is how long one of them sat there. A layout bug on an admin-only page gets noticed in a day. A wrong "today" label gets noticed within a week, because someone's staring at the drafts list constantly. A photo on a page nobody but visitors look at can sit broken for 77 days, because the person who'd notice it — me — doesn't visit my own About page.&lt;/p&gt;

&lt;p&gt;That's maybe the actual lesson this week: the bugs that survive longest aren't the scary ones. They're the ones on pages you built once and never look at again. That's the UX side of the hyper-vigilance tax, and it's the cheaper of the two to pay, because a broken layout or a wrong label eventually shows up somewhere a human looks. The security side doesn't get that luxury. Nothing renders wrong when an endpoint is missing a rate limit or a dependency has a stacked CORS bypass sitting three layers deep in a transitive dependency. Those don't announce themselves in a screenshot. They just sit there until someone goes looking on purpose, which is the more expensive half of the same bill.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/posts/auditing-the-surface-we-added-since-the-last-audit"&gt;Monday's post&lt;/a&gt; is that deliberate look, a fresh security pass against the same categories the last full audit used, covering the surface that's shipped since, three phases, three commits, one repo. Same discipline as this post, pointed at the axis where nobody stumbles into the bug by accident.&lt;/p&gt;

&lt;h2&gt;
  
  
  By the Numbers
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;2 lines&lt;/strong&gt; changed to fix the mobile draft card overflow&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;~2 minutes&lt;/strong&gt; to fix the layout once the real device confirmed it&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;~15 minutes&lt;/strong&gt; to diagnose and fix the relative-time timezone bug, &lt;a href="https://github.com/carryologist/the-vibe-coder/pull/22" rel="noopener noreferrer"&gt;PR #22&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;8–12 hours&lt;/strong&gt; — the evening window where the old elapsed-ms math could mislabel "tomorrow" as "today"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;77 days&lt;/strong&gt; the About page headshot sat broken before anyone noticed (Apr 30 → Jul 16)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;1&lt;/strong&gt; file recovered intact from git history, zero data lost&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;1&lt;/strong&gt; hardcoded interface name (&lt;code&gt;enp8s0&lt;/code&gt; → &lt;code&gt;enp112s0&lt;/code&gt;) that silently broke Wake on LAN across a hardware migration&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;3&lt;/strong&gt; repos/systems touched across all four fixes (2 code repos, 1 physical machine)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;4&lt;/strong&gt; bugs, &lt;strong&gt;4&lt;/strong&gt; different checks, &lt;strong&gt;1&lt;/strong&gt; shared failure mode&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;0&lt;/strong&gt; fodder files left unconsumed after this post&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>meta</category>
      <category>buildinginpublic</category>
      <category>agents</category>
      <category>debugging</category>
    </item>
    <item>
      <title>Thursday Thoughts: FOCUS and the True Cost of a Token</title>
      <dc:creator>Rob</dc:creator>
      <pubDate>Thu, 16 Jul 2026 14:40:46 +0000</pubDate>
      <link>https://dev.to/carryologist/thursday-thoughts-focus-and-the-true-cost-of-a-token-5485</link>
      <guid>https://dev.to/carryologist/thursday-thoughts-focus-and-the-true-cost-of-a-token-5485</guid>
      <description>&lt;p&gt;By day, I run a software company. That means my professional and personal lives intersect a lot. This week was one of those moments.&lt;/p&gt;

&lt;p&gt;My Finance team was presenting a case to join the &lt;a href="https://itsfoss.com/news/tokenomics-foundation/" rel="noopener noreferrer"&gt;Tokenomics Foundation&lt;/a&gt; and a request to implement the &lt;a href="https://focus.finops.org/" rel="noopener noreferrer"&gt;FOCUS spec&lt;/a&gt; in both our internal systems and external product. I was vaguely familiar with both, but felt much smarter after a 30-minute debate.&lt;/p&gt;

&lt;p&gt;But I can't stop there. I need to think about how this will affect the industry, employees, and — well — me.&lt;/p&gt;

&lt;p&gt;I keep coming back to the &lt;a href="https://dev.to/posts/thursday-thoughts-how-ai-native-mirrors-cloud-native"&gt;cloud-native analogy&lt;/a&gt; for AI, and this week it clicked again in a place I wasn't expecting: FinOps.&lt;/p&gt;

&lt;p&gt;On June 3rd, the Linux Foundation &lt;a href="https://itsfoss.com/news/tokenomics-foundation/" rel="noopener noreferrer"&gt;announced the intent to launch&lt;/a&gt; the &lt;strong&gt;Tokenomics Foundation&lt;/strong&gt;, a new body dedicated to open standards for AI cost management, in close partnership with the FinOps Foundation. The first concrete deliverable is &lt;a href="https://www.cio.com/article/4182274/linux-foundation-targets-ais-cost-management-problem-with-tokenomics-foundation.html" rel="noopener noreferrer"&gt;extending &lt;strong&gt;FOCUS&lt;/strong&gt;&lt;/a&gt; — the FinOps Open Cost and Usage Specification, the schema that already normalizes cloud billing across AWS, Azure, and GCP — to cover token-based AI spend. Twelve organizations, &lt;a href="https://itsfoss.com/news/tokenomics-foundation/" rel="noopener noreferrer"&gt;including Google Cloud, Microsoft, Oracle, Salesforce, SAP, and JPMorgan Chase&lt;/a&gt;, are already backing it.&lt;/p&gt;

&lt;p&gt;Here's why that matters to me, and why I think it should matter to you even if you've never opened a FinOps dashboard in your life.&lt;/p&gt;

&lt;h2&gt;
  
  
  We Took a Decade to Get Serious About Cloud Cost. We Don't Get That Long This Time.
&lt;/h2&gt;

&lt;p&gt;The cloud era ran for years before "FinOps" became a real discipline with real standards. Chargeback and showback were ad hoc. Every cloud provider invented its own billing schema, and practitioners built bespoke ETL pipelines to make AWS, Azure, and GCP cost data speak the same language. FOCUS didn't &lt;a href="https://focus.finops.org/focus-specification/v1-1/" rel="noopener noreferrer"&gt;formally exist as a Linux Foundation project until January 2023&lt;/a&gt; — more than fifteen years after AWS launched EC2. We built the discipline of consumption-based cost management &lt;em&gt;after&lt;/em&gt; consumption-based cost had already spiraled out of institutional control, and I've written before about how &lt;a href="https://dev.to/posts/thursday-thoughts-why-anthropic-is-the-next-aws-but-potentially-worse"&gt;Anthropic is running a version of the exact same AWS-shaped playbook&lt;/a&gt;, just at a pace that makes the cloud era look slow.&lt;/p&gt;

&lt;p&gt;Token-based AI spend is following the exact same consumption-based cost curve, except compressed. &lt;a href="https://itsfoss.com/news/tokenomics-foundation/" rel="noopener noreferrer"&gt;Global token usage is projected to grow 24x between 2026 and 2030, hitting 120 quadrillion tokens per month&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="/images/thursday-thoughts-focus-and-the-true-cost-of-a-token/token-growth-projection.png" class="article-body-image-wrapper"&gt;&lt;img src="/images/thursday-thoughts-focus-and-the-true-cost-of-a-token/token-growth-projection.png" alt="Bar chart showing global token usage growing 24x from a derived 2026 baseline of 5 quadrillion tokens per month to a projected 120 quadrillion tokens per month by 2030"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Chart: The Vibe Coder. 2030 figure and 24x growth multiple per Goldman Sachs research, as &lt;a href="https://itsfoss.com/news/tokenomics-foundation/" rel="noopener noreferrer"&gt;cited by It's FOSS&lt;/a&gt;; the 2026 baseline is derived by simple division and wasn't independently reported.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;And unlike a vCPU-hour, which is a stable, well-understood unit, &lt;a href="https://www.finops.org/insights/token-economics-the-atomic-unit-of-ai-value/" rel="noopener noreferrer"&gt;a token is not a fixed unit at all&lt;/a&gt; — different models tokenize the same text differently, and pricing, context windows, and caching behavior shift under you without notice. The fact that the industry is standing up FOCUS-for-AI &lt;em&gt;now&lt;/em&gt;, three years into the LLM API era rather than fifteen, is a genuinely good sign. It means we're applying a lesson instead of relearning it from scratch. That's the whole thesis of this blog in miniature: don't lift-and-shift the old playbook onto AI, but do keep the parts of the old playbook that were hard-won and correct. Consumption-based cost governance is one of those parts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why This Actually Matters to Vibe Coders, Not Just FinOps Practitioners
&lt;/h2&gt;

&lt;p&gt;I don't write this blog for people managing million-dollar cloud bills. I write it because I think &lt;a href="https://dev.to/posts/thursday-thoughts-every-intern-is-a-builder-now"&gt;every knowledge worker is about to become a vibe coder&lt;/a&gt;, the same way every knowledge worker is already a spreadsheet user or a slide-deck builder. Building an internal tool, a workflow automation, or a small app is going to be a universal white-collar skill within a few years, not a specialist one. That's the premise this whole blog runs on.&lt;/p&gt;

&lt;p&gt;Which means the economics conversation happening in FinOps circles right now isn't going to stay contained to platform teams and CFOs. It's coming for every person who opens an agent chat and says "build me a dashboard." Once &lt;a href="https://dev.to/posts/thursday-thoughts-every-intern-is-a-builder-now"&gt;vibe coding is universal&lt;/a&gt;, token consumption becomes as distributed, as invisible, and as easy to blow past a budget on as cloud spend was in 2015 — except instead of one platform team provisioning EC2 instances, it's every employee with an agent tab open. The &lt;a href="https://www.finops.org/topic/ai-value/" rel="noopener noreferrer"&gt;State of FinOps 2026 survey found AI has become a mainstream technology investment&lt;/a&gt;, with 98% of FinOps teams now managing AI spend, up from just 31% two years ago. That number is going to keep climbing precisely because the &lt;em&gt;people generating&lt;/em&gt; the spend are no longer engineers alone.&lt;/p&gt;

&lt;p&gt;The practical question for a vibe coder — hobbyist or enterprise employee alike — isn't "what does FOCUS mean for finance." It's "will my organization eventually meter me the way it metered a Kubernetes namespace." I think the answer is yes, and I think it should be, because the alternative is nobody knowing what any of this actually costs until the invoice arrives.&lt;/p&gt;

&lt;h2&gt;
  
  
  What FOCUS Actually Is, and Isn't
&lt;/h2&gt;

&lt;p&gt;FOCUS is not a logging or telemetry standard. It's a billing schema — closer to a standardized invoice format than to an observability trace. &lt;a href="https://github.com/FinOps-Open-Cost-and-Usage-Spec/FOCUS_Spec" rel="noopener noreferrer"&gt;It defines a common schema for technology cost and usage data&lt;/a&gt; across cloud, SaaS, data center, and other technology categories, establishing a consistent, vendor-neutral vocabulary for billing and usage data. &lt;a href="https://focus.finops.org/what-is-focus/" rel="noopener noreferrer"&gt;A FOCUS dataset is a table of charges&lt;/a&gt;, where each row represents one charge, and every column has a defined name, data type, and meaning set by the specification, so the column means the same thing regardless of which provider produced the row.&lt;/p&gt;

&lt;p&gt;Concretely, it ships as CSV or Parquet exports (&lt;a href="https://techcommunity.microsoft.com/blog/finopsblog/managing-azure-openai-costs-with-the-finops-toolkit-and-focus-turning-tokens-int/4413886" rel="noopener noreferrer"&gt;AWS's CUR 2.0 can output FOCUS 1.2-formatted Parquet to S3&lt;/a&gt;), normative language follows RFC 2119/8174 (MUST/SHOULD/MAY), and providers can extend it with &lt;code&gt;x_&lt;/code&gt;-prefixed columns for proprietary detail without breaking the shared schema. There's even an &lt;a href="https://github.com/finopsfoundation/focus_validator" rel="noopener noreferrer"&gt;open-source validator&lt;/a&gt; that checks a dataset against the spec version by version.&lt;/p&gt;

&lt;p&gt;(One housekeeping note: the FOCUS and FinOps Foundation logos are registered trademarks, so rather than reproduce them here, I'm linking straight to the &lt;a href="https://focus.finops.org/" rel="noopener noreferrer"&gt;FOCUS brand site&lt;/a&gt; and the &lt;a href="https://www.finops.org/about/media/" rel="noopener noreferrer"&gt;FinOps Foundation media page&lt;/a&gt; if you want the official marks.)&lt;/p&gt;

&lt;h2&gt;
  
  
  The Evolution, and Where 1.4 Landed
&lt;/h2&gt;

&lt;p&gt;FOCUS has moved fast for a standards body:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://focus.finops.org/focus-specification/v1-0/" rel="noopener noreferrer"&gt;v1.0&lt;/a&gt; (2024):&lt;/strong&gt; established the core schema — &lt;code&gt;BilledCost&lt;/code&gt;, &lt;code&gt;EffectiveCost&lt;/code&gt;, &lt;code&gt;ConsumedQuantity&lt;/code&gt; — for Cloud Service Provider billing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://focus.finops.org/focus-specification/v1-1/" rel="noopener noreferrer"&gt;v1.1&lt;/a&gt; (Nov 2024):&lt;/strong&gt; added invoice reconciliation and unit-cost/density metrics (cost-per-GB, cost-per-request).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://focus.finops.org/focus-specification/v1-2/" rel="noopener noreferrer"&gt;v1.2&lt;/a&gt; (May 2025):&lt;/strong&gt; unified Cloud + SaaS + PaaS reporting into one schema, and — notably for this post — introduced the first language around virtual currency and &lt;strong&gt;token purchase pattern analysis&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;v1.3 (Dec 2025):&lt;/strong&gt; added a dedicated Contract Commitment dataset and, critically, first-class shared-cost allocation fields — which resource was shared, who consumed it, and what method split the cost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://focus.finops.org/focus-specification/" rel="noopener noreferrer"&gt;v1.4&lt;/a&gt; (ratified June 4, 2026, at FinOps X):&lt;/strong&gt; the current release. It adds 2 datasets, 47 columns, 6 attributes, 17 glossary entries, and 2 supported features, headlined by new Invoice Detail and Billing Period datasets for reconciling usage straight to invoices, and Service Provider vs. Host Provider columns that separate who sold you the resource from who's actually running it underneath, disambiguating reseller relationships.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="/images/thursday-thoughts-focus-and-the-true-cost-of-a-token/focus-version-evolution.png" class="article-body-image-wrapper"&gt;&lt;img src="/images/thursday-thoughts-focus-and-the-true-cost-of-a-token/focus-version-evolution.png" alt="Timeline showing five FOCUS specification releases from v1.0 in 2024 through v1.4 ratified June 4, 2026, each annotated with its key additions"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Chart: The Vibe Coder. Ratification dates and release details per the &lt;a href="https://focus.finops.org/focus-specification/" rel="noopener noreferrer"&gt;FOCUS Specification changelog&lt;/a&gt;, licensed &lt;a href="https://creativecommons.org/licenses/by/4.0/legalcode" rel="noopener noreferrer"&gt;CC-BY-4.0&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The AI-specific columns riding on top of that 1.4 release are the ones that matter most here: &lt;code&gt;ConsumedQuantity&lt;/code&gt;, &lt;code&gt;ConsumedUnit&lt;/code&gt;, &lt;code&gt;HostProviderName&lt;/code&gt;, and the &lt;a href="https://siliconangle.com/2026/06/08/ai-token-economics-focus-specification-updates-finopsx/" rel="noopener noreferrer"&gt;&lt;code&gt;x_InputTokens&lt;/code&gt; / &lt;code&gt;x_OutputTokens&lt;/code&gt; / &lt;code&gt;x_CachedTokens&lt;/code&gt; splits the spec introduced specifically for AI workloads&lt;/a&gt;. That's the schema-level answer to "how much did this model call actually cost, and for what."&lt;/p&gt;

&lt;h2&gt;
  
  
  What FOCUS 1.4 Gets Right for AI Consumption
&lt;/h2&gt;

&lt;p&gt;If you're buying tokens from a frontier lab or a hyperscaler's managed AI service, FOCUS today gives you real, standardized ground to stand on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cross-provider comparability.&lt;/strong&gt; Whether the bill comes from OpenAI, Anthropic, Azure OpenAI, or Bedrock, the token consumption shows up as &lt;a href="https://techcommunity.microsoft.com/blog/finopsblog/managing-azure-openai-costs-with-the-finops-toolkit-and-focus-turning-tokens-int/4413886" rel="noopener noreferrer"&gt;&lt;code&gt;ConsumedQuantity&lt;/code&gt; in &lt;code&gt;ConsumedUnit: tokens&lt;/code&gt;&lt;/a&gt;, the same way a compute charge shows up as vCPU-hours everywhere.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Input/output/cached token attribution.&lt;/strong&gt; The &lt;code&gt;x_InputTokens&lt;/code&gt;/&lt;code&gt;x_OutputTokens&lt;/code&gt;/&lt;code&gt;x_CachedTokens&lt;/code&gt; split lets you see where the money actually went, which matters given &lt;a href="https://www.finout.io/blog/ai-model-cost-breakdowns-the-complete-2026-comparison-guide" rel="noopener noreferrer"&gt;output tokens typically cost 3–8x more than input tokens&lt;/a&gt; because generation is more compute-intensive than reading a prompt.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Shared-cost allocation for chargeback.&lt;/strong&gt; The &lt;a href="https://focus.finops.org/focus-specification/v1-2/" rel="noopener noreferrer"&gt;1.3-era split-cost-allocation fields&lt;/a&gt; — which resource was shared, which consumers used it, what method was used to split it — are exactly the mechanism enterprises need to charge back a shared model deployment or a shared GPU pool across teams, not just a shared VM.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Amortization of commitments.&lt;/strong&gt; FOCUS already knows how to &lt;a href="https://siliconangle.com/2026/06/08/focus-specification-ai-cost-accountability-finopsx/" rel="noopener noreferrer"&gt;spread a flat-rate subscription across daily consumption via an effective-cost column&lt;/a&gt;, the same pattern that would apply to a reserved-capacity inference contract.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where It Still Falls Short for Self-Hosting
&lt;/h2&gt;

&lt;p&gt;This is the part that matters if you run your own models, and it's the part that isn't solved yet.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you self-host, there's no token to bill in the first place.&lt;/strong&gt; When you run an open-weight model on your own GPUs, &lt;a href="https://www.finops.org/wg/token-economics-saas/" rel="noopener noreferrer"&gt;cost is expressed entirely in compute&lt;/a&gt; — GPU instance hours, storage, and networking — because there is no per-token charge. FOCUS handles that side fine; a GPU instance is just another compute row with a &lt;code&gt;BilledCost&lt;/code&gt; and &lt;code&gt;ConsumedQuantity&lt;/code&gt; in GPU-hours, same as any VM. What FOCUS does &lt;em&gt;not&lt;/em&gt; give you is a bridge back from that GPU-hour to a per-token or per-request cost. You have to build that join yourself, against your own inference telemetry.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Token economics that only counts tokens is a partial view.&lt;/strong&gt; The &lt;a href="https://www.finops.org/insights/token-economics-the-atomic-unit-of-ai-value/" rel="noopener noreferrer"&gt;FinOps Foundation says this plainly&lt;/a&gt;: beneath every token is a chain of physical and architectural decisions that determine what the token costs to produce, and for self-hosted deployments that means the capital cost of facilities, power, and cooling, with industry estimates placing next-generation AI data center construction at fifteen to twenty million dollars per megawatt of capacity. None of that shows up as a FOCUS column today. There's no &lt;code&gt;x_PowerDrawWatts&lt;/code&gt;, no PUE field, no hardware depreciation schedule. The spec's amortization machinery &lt;em&gt;could&lt;/em&gt; carry that eventually — it already amortizes commitments — but nobody has defined the columns yet.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The frontier work is still aimed at the wrong side of the ledger.&lt;/strong&gt; &lt;a href="https://siliconangle.com/2026/06/08/focus-specification-ai-cost-accountability-finopsx/" rel="noopener noreferrer"&gt;FOCUS 1.5 is slated to break down AI spend by token type and workload&lt;/a&gt;, giving practitioners the granularity to tie inference costs back to the teams consuming them — but that's still describing metered consumption, not manufactured tokens. &lt;a href="https://www.finops.org/insights/token-economics-the-atomic-unit-of-ai-value/" rel="noopener noreferrer"&gt;Jensen Huang's "AI factory" framing&lt;/a&gt; is the more honest lens for self-hosters: electricity and silicon enter the factory, tokens emerge, and the real unit-economics question is revenue (or value) per megawatt, not price per API call. FOCUS hasn't gone there yet.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cardinality is a real, admitted problem.&lt;/strong&gt; Even on the consumption side, &lt;a href="https://siliconangle.com/2026/06/08/ai-token-economics-focus-specification-updates-finopsx/" rel="noopener noreferrer"&gt;FOCUS contributors are candid that the harder frontier is AI token economics&lt;/a&gt;, because measuring the cost of inference requires visibility down to the per-user, per-session, and per-request level, and it could end up being that a non-trivial percentage of your cost is having to be attributed to just getting your data and putting it through a pipeline. If that's true at hyperscaler scale, it's just as true, proportionally, on a single GPU box in someone's closet.&lt;/p&gt;

&lt;h2&gt;
  
  
  What We're Going to Try on the Homelab
&lt;/h2&gt;

&lt;p&gt;Yes, running a FinOps billing spec against a single &lt;a href="https://dev.to/posts/thursday-thoughts-the-models-we-cant-run"&gt;RTX 5090&lt;/a&gt; is a little absurd on its face. Nobody needs a standardized multi-cloud invoice schema to know what one GPU costs. But that's exactly why I want to try it. Homelab-scale is the cleanest possible environment to see whether the discipline holds up when you strip away all the enterprise noise — no shared cost pools, no negotiated discounts, no thousand-account org structure. Just one box, one power meter, and a request log.&lt;/p&gt;

&lt;p&gt;The plan is to try to build a real, if tiny, FOCUS-shaped dataset for the homelab: &lt;code&gt;BilledCost&lt;/code&gt; derived from wall-clock power draw at the outlet times our electricity rate, &lt;code&gt;ConsumedQuantity&lt;/code&gt; in tokens/sec pulled from llama.cpp's own metrics, and a hand-rolled &lt;code&gt;x_PowerDrawWatts&lt;/code&gt; column FOCUS doesn't define yet, joined against the model and quant we were running at the time (&lt;a href="https://dev.to/posts/thursday-thoughts-the-models-we-cant-run"&gt;our daily driver is still Qwen 3.5 35B-A3B at Q4_K_XL, 22 GB of weights, 200+ tok/s on the 5090&lt;/a&gt;). If it works, it becomes the smallest possible proof that the "AI factory" framing scales down as cleanly as it scales up — that the true cost of a token is never just the API price, it's power and silicon showing up as a line item, whether that line item is a hyperscaler's data center or a workstation under a desk.&lt;/p&gt;

&lt;h2&gt;
  
  
  By the Numbers
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;June 3, 2026&lt;/strong&gt; — the &lt;a href="https://itsfoss.com/news/tokenomics-foundation/" rel="noopener noreferrer"&gt;Linux Foundation announces intent&lt;/a&gt; to launch the Tokenomics Foundation&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;12&lt;/strong&gt; &lt;a href="https://www.cio.com/article/4182274/linux-foundation-targets-ais-cost-management-problem-with-tokenomics-foundation.html" rel="noopener noreferrer"&gt;organizations already backing it&lt;/a&gt;, including Google Cloud, Microsoft, Oracle, Salesforce, SAP, and JPMorgan Chase&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;January 2023&lt;/strong&gt; — &lt;a href="https://focus.finops.org/focus-specification/v1-1/" rel="noopener noreferrer"&gt;FOCUS formally becomes a Linux Foundation project&lt;/a&gt;, roughly 15 years after AWS launched EC2, and 3 years after the modern LLM API era began&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;1.4&lt;/strong&gt; — the &lt;a href="https://focus.finops.org/focus-specification/" rel="noopener noreferrer"&gt;current FOCUS version&lt;/a&gt;, ratified June 4, 2026, adding 2 datasets, 47 columns, 6 attributes, and 17 glossary entries&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;24x&lt;/strong&gt; — &lt;a href="https://itsfoss.com/news/tokenomics-foundation/" rel="noopener noreferrer"&gt;projected growth in global token usage&lt;/a&gt; between 2026 and 2030&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;120 quadrillion&lt;/strong&gt; — &lt;a href="https://itsfoss.com/news/tokenomics-foundation/" rel="noopener noreferrer"&gt;projected global tokens consumed per month&lt;/a&gt; by 2030&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;98%&lt;/strong&gt; — &lt;a href="https://www.finops.org/topic/ai-value/" rel="noopener noreferrer"&gt;FinOps teams now managing AI spend&lt;/a&gt;, up from 31% just two years ago&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;3–8x&lt;/strong&gt; — &lt;a href="https://www.finout.io/blog/ai-model-cost-breakdowns-the-complete-2026-comparison-guide" rel="noopener noreferrer"&gt;how much more an output token costs than an input token&lt;/a&gt;, due to generation compute&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;$15–20M&lt;/strong&gt; — &lt;a href="https://www.finops.org/insights/token-economics-the-atomic-unit-of-ai-value/" rel="noopener noreferrer"&gt;estimated construction cost per megawatt&lt;/a&gt; of next-generation AI data center capacity&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;0&lt;/strong&gt; — FOCUS columns today for power draw, PUE, or hardware depreciation on self-hosted inference&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;1&lt;/strong&gt; &lt;a href="https://dev.to/posts/thursday-thoughts-the-models-we-cant-run"&gt;RTX 5090&lt;/a&gt; we're about to try to build a FOCUS-shaped cost dataset around anyway&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>cloudnative</category>
      <category>homelab</category>
    </item>
    <item>
      <title>Model Showdown Round 9: Qwen 3.6 27B vs Qwen 3.6 35B-A3B vs Qwythos-9B vs GLM-4.7-Flash vs Nemotron-3-Nano</title>
      <dc:creator>Rob</dc:creator>
      <pubDate>Wed, 15 Jul 2026 21:52:58 +0000</pubDate>
      <link>https://dev.to/carryologist/model-showdown-round-9-qwen-36-27b-vs-qwen-36-35b-a3b-vs-qwythos-9b-vs-glm-47-flash-vs-4j69</link>
      <guid>https://dev.to/carryologist/model-showdown-round-9-qwen-36-27b-vs-qwen-36-35b-a3b-vs-qwythos-9b-vs-glm-47-flash-vs-4j69</guid>
      <description>&lt;p&gt;Round 7 ended on a cliffhanger I couldn't stop thinking about. Qwen 3.6 35B-A3B &lt;em&gt;built the entire feature&lt;/em&gt; — read the codebase, wrote the files, got a clean build — and then spent 77 messages, more than half its session, failing to take a Playwright screenshot. It never committed. It never pushed. All that work, gone.&lt;/p&gt;

&lt;p&gt;Was that a bad day, or is it structural? Round 9 was supposed to answer that with three contestants: the 35B-A3B running back for a rematch, a dense 27B challenger, and NVIDIA's Nemotron-3-Nano as an architectural wild card. Clean, narrow, three-way test of dense-vs-MoE.&lt;/p&gt;

&lt;p&gt;It didn't stay clean. By the time I was done, I'd expanded the field to five models, and I'd personally patched two separate llama.cpp/template bugs live, mid-bakeoff, just to get two of the contestants to a fair starting line. One of those fixes worked perfectly — and the model still failed anyway, for a completely different reason.&lt;/p&gt;

&lt;p&gt;Let's get into it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Setup
&lt;/h2&gt;

&lt;p&gt;This is Round 9 of the Local Model Showdown, a sub-series of Model Showdown that only tests models I can actually run on my own hardware. No API keys, no cloud spend — just an RTX 5090 and however much patience the model has for a real coding task.&lt;/p&gt;

&lt;p&gt;The homelab, unchanged from Round 7:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;CPU&lt;/strong&gt;: AMD Ryzen 9 9950X3D, 64GB RAM&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GPU&lt;/strong&gt;: NVIDIA RTX 5090, 32GB VRAM&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Inference&lt;/strong&gt;: llama.cpp, single-model serving, one contestant loaded at a time&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent platform&lt;/strong&gt;: Coder Agents&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OS&lt;/strong&gt;: Ubuntu 24.04&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The Contestants
&lt;/h3&gt;

&lt;p&gt;The plan called for three. I ran five.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Run&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Architecture&lt;/th&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Qwen 3.6 27B&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Dense transformer&lt;/td&gt;
&lt;td&gt;Primary dense challenger&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Qwen 3.6 35B-A3B&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;MoE transformer&lt;/td&gt;
&lt;td&gt;Incumbent / Round 7 rematch&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Qwythos-9B-Claude-Mythos-5-1M&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Dense, MTP speculative decoding&lt;/td&gt;
&lt;td&gt;Unplanned wild card&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;GLM-4.7-Flash&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Dense&lt;/td&gt;
&lt;td&gt;Unplanned wild card&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Nemotron-3-Nano-30B-A3B&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Hybrid Transformer-Mamba-2 MoE&lt;/td&gt;
&lt;td&gt;Architectural wild card&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Why the field grew: once the harness and the model-serving pipeline were working, running two more small/cheap contestants cost almost nothing extra in setup, and both turned out to matter — one for a completely novel failure mode I hadn't seen in six rounds of this series, the other for confirming a fix actually worked in practice, not just in a curl test.&lt;/p&gt;

&lt;p&gt;Model-to-run mapping was randomized and sealed before any task prompt was sent, same as every round in this series.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Task
&lt;/h2&gt;

&lt;p&gt;Identical to Round 7, on purpose — this is the only way to get a clean cross-round read on the 35B-A3B incumbent.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Goal&lt;/strong&gt;: Add a Tag Manager to the &lt;code&gt;/admin&lt;/code&gt; section.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Requirements&lt;/strong&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;lib/tags.ts&lt;/code&gt; — read all tags from published and draft posts (gray-matter)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;GET /api/admin/tags&lt;/code&gt; — JSON list of tags with post counts&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;PUT /api/admin/tags/{tag}&lt;/code&gt; — rename a tag across all posts&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;DELETE /api/admin/tags/{tag}&lt;/code&gt; — remove a tag from all posts&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/admin/tags&lt;/code&gt; page — list with inline rename/delete&lt;/li&gt;
&lt;li&gt;Link &lt;code&gt;/admin/tags&lt;/code&gt; from the admin nav&lt;/li&gt;
&lt;li&gt;Screenshot of the finished page in the PR description, via Playwright MCP&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;npm run build&lt;/code&gt; must pass before any commit&lt;/li&gt;
&lt;li&gt;Commit in logical chunks, push the branch&lt;/li&gt;
&lt;/ol&gt;
&lt;/blockquote&gt;

&lt;p&gt;Same nine requirements. Same baseline commit. Same "no hand-holding" philosophy — the only messages I sent mid-run were a bare &lt;code&gt;"continue"&lt;/code&gt; when a session paused at a harness turn-limit, never a hint about what to do next.&lt;/p&gt;

&lt;h2&gt;
  
  
  Interlude: The Infrastructure Fought Back Twice
&lt;/h2&gt;

&lt;p&gt;This is the part that wasn't in the plan.&lt;/p&gt;

&lt;p&gt;Two of the five models — Qwythos-9B and Nemotron-3-Nano — failed their very first request with a hard error before ever seeing the actual task. Both failures traced back to llama.cpp's automatic tool-call parser, and both required extracting the model's raw jinja chat template out of the GGUF metadata and hand-patching it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Qwythos-9B's bug&lt;/strong&gt;: its embedded template unconditionally raises &lt;code&gt;Jinja Exception: System message must be at the beginning&lt;/code&gt; the moment it sees a &lt;em&gt;second&lt;/em&gt; system-role message. Coder always sends two — its own agent prompt, then a workspace-context note — so every single request 400'd before the model ever ran. The template's own author clearly didn't anticipate a harness that layers system messages. Fix: pull the template via the server's &lt;code&gt;/props&lt;/code&gt; endpoint, patch the three-line conditional block that renders the first system message but raises on the second, and reload with &lt;code&gt;--chat-template-file&lt;/code&gt; pointing at the patched copy. Verified the patch preserved the model's native Claude/Anthropic-style &lt;code&gt;&amp;lt;tool_call&amp;gt;&amp;lt;function=...&amp;gt;&lt;/code&gt; format — this wasn't a case of stripping tool support to dodge the crash.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Nemotron-3-Nano's bug&lt;/strong&gt; was subtler. The historical config used a &lt;code&gt;--special&lt;/code&gt; flag that causes the model's &lt;code&gt;&amp;lt;|im_end|&amp;gt;&lt;/code&gt; stop token to print as literal output text instead of being silently consumed — which broke the auto-derived tool-call parser with a &lt;code&gt;500: unparsed peg-native output&lt;/code&gt; error. The existing workaround was overriding the template entirely with a generic &lt;code&gt;chatml&lt;/code&gt; template. That avoided the crash, but the generic template doesn't render the &lt;code&gt;tools&lt;/code&gt; list into the prompt at all — so the model couldn't see what functions existed, and it hallucinated plausible-sounding ones (&lt;code&gt;ls&lt;/code&gt;, &lt;code&gt;pwd&lt;/code&gt;) that didn't match anything I'd actually provided. Fix: drop &lt;code&gt;--special&lt;/code&gt;, keep the real native template with tool-schema injection intact. Verified — repeated tests produced correctly parsed, correctly-named tool calls.&lt;/p&gt;

&lt;p&gt;Both fixes worked. I confirmed each one with direct API tests before sending a single task prompt. Both models still failed the actual bakeoff task anyway — though as it turns out, I could only fully confirm one of those two follow-up failures was really the model's fault. More on that below.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Results
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Messages&lt;/th&gt;
&lt;th&gt;Total Tokens&lt;/th&gt;
&lt;th&gt;Interventions&lt;/th&gt;
&lt;th&gt;Outcome&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Qwen 3.6 27B&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;304&lt;/td&gt;
&lt;td&gt;7,967,497&lt;/td&gt;
&lt;td&gt;8 (neutral)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Complete&lt;/strong&gt; — PR #19, 4 commits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Qwen 3.6 35B-A3B&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;237&lt;/td&gt;
&lt;td&gt;5,592,314&lt;/td&gt;
&lt;td&gt;5 (neutral)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Complete&lt;/strong&gt; — PR #20, 5 commits — best run&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Qwythos-9B&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;76,605&lt;/td&gt;
&lt;td&gt;2 (neutral)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Failed&lt;/strong&gt; — never executed a real tool call (cause inconclusive, see below)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;GLM-4.7-Flash&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;246&lt;/td&gt;
&lt;td&gt;4,951,774&lt;/td&gt;
&lt;td&gt;3 neutral + 1 correction&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Complete&lt;/strong&gt; — PR #21, 1 commit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Nemotron-3-Nano&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;165&lt;/td&gt;
&lt;td&gt;2,302,880&lt;/td&gt;
&lt;td&gt;4 (neutral)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Failed&lt;/strong&gt; — never found the repo&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three of five shipped a real, mergeable PR. The incumbent Qwen 3.6 35B-A3B didn't repeat its Round 7 spiral — the reproducibility signal says Round 7 really was a bad day, not a structural MoE problem. The dense-vs-MoE hypothesis, meanwhile, got muddier, not cleaner: the best run of the round was an MoE model, and the two failures split one dense (Qwythos), one MoE (Nemotron).&lt;/p&gt;

&lt;h2&gt;
  
  
  What Each Model Actually Did
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Qwen 3.6 35B-A3B (Run 2): The Redemption Arc
&lt;/h3&gt;

&lt;p&gt;The Round 7 incumbent came back and did everything right. It found the repo unprompted after a brief, reasonable search (it didn't have the exact org/repo name memorized, tried a couple of &lt;code&gt;gh search repos&lt;/code&gt; queries, found it). It correctly diagnosed that the blog persists content through the GitHub API rather than the local filesystem — without being told — self-corrected two real bugs (a missing React import, a server/client component split), got a working authenticated Playwright screenshot on effectively the first real attempt, committed the image directly into the repo, and opened PR #20.&lt;/p&gt;

&lt;p&gt;It even added an uninstructed but reasonable improvement — wrapping the tag-count function in &lt;code&gt;React.cache()&lt;/code&gt; to avoid redundant GitHub API calls during SSR — and cleaned up a stray &lt;code&gt;playwright&lt;/code&gt; devDependency it no longer needed, entirely on its own initiative.&lt;/p&gt;

&lt;p&gt;Five interventions total, every one a bare "continue" at a harness pause, zero content hints. This is the cleanest run of the round, and the strongest evidence yet that Round 7's failure wasn't inherent to this model's architecture.&lt;/p&gt;

&lt;h3&gt;
  
  
  Qwen 3.6 27B (Run 1): Got There, the Slow Way
&lt;/h3&gt;

&lt;p&gt;The dense challenger also shipped — PR #19, four commits — but needed twice as many interventions to get there. It hand-rolled its own Playwright script instead of using the MCP tool as instructed (an instruction-following miss echoed by every model in this round that attempted a screenshot), and it burned real turns flailing on &lt;code&gt;gh api&lt;/code&gt;/&lt;code&gt;gh pr edit&lt;/code&gt; argument syntax trying to attach that screenshot to the PR after the fact. It recovered on its own — no hints given — by committing the image straight into the repo and rewriting the PR body to point at the raw GitHub URL, exactly the pattern that worked for Run 2.&lt;/p&gt;

&lt;p&gt;Positive signal: real autonomous debugging along the way, including tracing a &lt;code&gt;localhost:3000&lt;/code&gt; timeout to Redis simply not running, installing and starting &lt;code&gt;redis-server&lt;/code&gt; itself, and clearing a stale dev-server PID lock left over from an earlier crashed attempt. The reasoning was sound. It just took the long way around.&lt;/p&gt;

&lt;h3&gt;
  
  
  GLM-4.7-Flash (Run 4): Right Answer, Wrong Branch
&lt;/h3&gt;

&lt;p&gt;This run is the one asterisk in an otherwise clean sweep. It built the entire feature correctly and efficiently — but when it went looking for its &lt;code&gt;run-4&lt;/code&gt; branch, it found &lt;code&gt;feature/image-management-run-4&lt;/code&gt;, a stale, unrelated branch left over from a previous round's image-management feature, and assumed that was the branch it had been told to use. It checked it out, committed the tag manager on top of it, and pushed — polluting an old branch with unrelated code. No PR existed for that branch, so nothing else broke, but it wasn't the outcome the task asked for.&lt;/p&gt;

&lt;p&gt;I tested whether it would notice on its own: I reverted the polluted branch and sent one more neutral "continue," specifically to see if re-checking git state would trigger self-correction. It didn't — it just re-walked its own requirements checklist, found the missing screenshot, and kept working without ever revisiting the branch. That answer settled it: this needed an explicit correction, not another nudge. Once told directly that the branch was wrong, it recovered in a single turn — created &lt;code&gt;run-4&lt;/code&gt; properly off &lt;code&gt;main&lt;/code&gt;, reapplied its own code, pushed, and opened PR #21.&lt;/p&gt;

&lt;p&gt;Worth being honest about the scoring implication: this is the only run in the round that needed a content hint rather than a neutral nudge, and that should count against it relative to Run 2's fully autonomous path to the same kind of outcome. It also never resolved the screenshot requirement — no Playwright MCP available to it, a spawned sub-agent for the screenshot timed out, no &lt;code&gt;.env&lt;/code&gt; credentials to log in manually — and it gave up on that requirement rather than finding a workaround.&lt;/p&gt;

&lt;h3&gt;
  
  
  Qwythos-9B (Run 3): Said the Right Thing, Couldn't Say It Correctly — Or Could It?
&lt;/h3&gt;

&lt;p&gt;This is the one that got the infrastructure fix and still failed completely, and it's the most interesting result of the round — because when I went back to check whether that failure was really the model's fault, I couldn't fully confirm it was.&lt;/p&gt;

&lt;p&gt;Three consecutive turns in the actual bakeoff, identical pattern each time: correct reasoning ("I need to read the skills files before starting," "let me check the branch structure"), followed by an attempted tool call wrapped in the wrong syntax. Instead of its own trained &lt;code&gt;&amp;lt;tool_call&amp;gt;&amp;lt;function=name&amp;gt;...&amp;lt;/function&amp;gt;&amp;lt;/tool_call&amp;gt;&lt;/code&gt; format, it emitted a raw ad-hoc tag using the tool's name directly as the XML tag — &lt;code&gt;&amp;lt;read_file path="..."&amp;gt;&lt;/code&gt;, &lt;code&gt;&amp;lt;execute command="..." timeout="5s"&amp;gt;&lt;/code&gt; — which is unparseable by anything. Not one of the three attempts was ever actually executed. No repo was ever cloned. No code was ever written.&lt;/p&gt;

&lt;p&gt;Here's the honest complication. Before I called this a pure model-capability gap, I went back and tested the claim directly: I reconstructed Coder's actual system prompt (verbatim, including the injected user-instructions block), the real task prompt used in the bakeoff, and a representative tool schema — first a lean 10-tool version, then scaled up to 63 tools to match the size of Coder's actual full toolset — and fired all of it straight at Qwythos over the API, bypassing Coder's harness entirely.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Eleven for eleven.&lt;/strong&gt; Every single reconstruction produced a correctly-parsed, correctly-structured tool call. Not one reproduced the raw-text failure I'd watched happen three times in a row inside the real chat.&lt;/p&gt;

&lt;p&gt;That doesn't clear the model, but it doesn't convict it either. What it tells me is that I cannot honestly claim this was purely "a 9B model can't handle a long, complex prompt." Something specific to Coder's exact request — content I couldn't perfectly reproduce from outside the harness, whether that's the precise injected workspace context, a subtlety in how the conversation history was serialized, or simply an unlucky run of sampling — was very plausibly a contributing factor, and I don't have the visibility into Coder's exact request construction to rule it out. The fairest thing I can say: this failure was real, it happened three times with zero self-correction, and under my best-effort attempt to recreate the same conditions independently, it didn't happen once. Treat the result as inconclusive on root cause, not as a clean verdict on the model.&lt;/p&gt;

&lt;h3&gt;
  
  
  Nemotron-3-Nano (Run 5): Fixed the Bug, Lost the Model Anyway
&lt;/h3&gt;

&lt;p&gt;The most frustrating result of the round, because the fix I made for it worked exactly as intended, and it still didn't matter.&lt;/p&gt;

&lt;p&gt;Real tool execution, confirmed throughout — no parse crashes, no hallucinated tools, every &lt;code&gt;execute&lt;/code&gt; call actually ran and returned a real result. And then it spent roughly 35-40 turns across five nudges in an unbroken loop trying to locate the repository: checking &lt;code&gt;/workspace&lt;/code&gt; (repeatedly, across multiple nudges, always failing the same way), checking &lt;code&gt;/home/coder/project&lt;/code&gt;, &lt;code&gt;/home/coder/workspaces&lt;/code&gt;, re-listing a skill directory it had already found and dismissed, misusing &lt;code&gt;list_agents&lt;/code&gt; as a raw shell command, and passing a garbage string into a tool's template-ID parameter. It never once tried &lt;code&gt;gh search repos&lt;/code&gt; or &lt;code&gt;gh repo clone&lt;/code&gt; — the exact move every other model in the round used successfully within one or two turns.&lt;/p&gt;

&lt;p&gt;Five interventions, zero strategy convergence, zero progress on any of them. This is a pure repo-discovery / tool-selection capability gap that happened to surface &lt;em&gt;after&lt;/em&gt; I'd already fixed the thing that looked, at first, like it would be the blocker.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scoring the Three That Shipped
&lt;/h2&gt;

&lt;p&gt;Deviation from the series' usual format, disclosed up front: every other round in this series scores blind, before the orchestrating model knows which run is which contestant. That wasn't possible here — I needed to do the scoring myself, and I already knew the mapping from running the whole bakeoff. So this round is scored open, not blind, using the same 7-dimension weighted rubric, and grounded in the actual PR diffs rather than the write-ups above. Qwythos-9B and Nemotron-3-Nano shipped no code, so they're excluded as DNF rather than scored to zero.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Weight&lt;/th&gt;
&lt;th&gt;Qwen 3.6 27B (PR #19)&lt;/th&gt;
&lt;th&gt;Qwen 3.6 35B-A3B (PR #20)&lt;/th&gt;
&lt;th&gt;GLM-4.7-Flash (PR #21)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Correctness&lt;/td&gt;
&lt;td&gt;25%&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Design&lt;/td&gt;
&lt;td&gt;15%&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Code quality&lt;/td&gt;
&lt;td&gt;20%&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Engineering judgment&lt;/td&gt;
&lt;td&gt;15%&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scope discipline&lt;/td&gt;
&lt;td&gt;10%&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Commit hygiene&lt;/td&gt;
&lt;td&gt;10%&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Surprise&lt;/td&gt;
&lt;td&gt;5%&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Weighted total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3.75&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;4.90&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.30&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Qwen 3.6 35B-A3B (PR #20)&lt;/strong&gt; is the only one of the three that persists tag data exclusively through the GitHub API — reading and writing through &lt;code&gt;commitFile&lt;/code&gt; rather than touching the local filesystem at all. That's the correct pattern for a Vercel-hosted, read-only-filesystem app in production, and the run reasoned about it explicitly, unprompted. It also reasoned in code comments about a read-before-write race condition, the only run of the three to do so. Highest score across every dimension.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Qwen 3.6 27B (PR #19)&lt;/strong&gt; eventually persists through &lt;code&gt;commitFile&lt;/code&gt; too, but round-trips through the local filesystem first — a pattern that works in dev and silently fails the moment Vercel's read-only production filesystem is in play. Combined with the flailing &lt;code&gt;gh api&lt;/code&gt;/&lt;code&gt;gh pr edit&lt;/code&gt; detour documented above, it lands solidly in the middle: correct destination, messier path, weaker commit hygiene than #20's five clean commits.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GLM-4.7-Flash (PR #21)&lt;/strong&gt; scores lowest, and not just because of the wrong-branch incident. Its &lt;code&gt;lib/tags.ts&lt;/code&gt; only ever calls &lt;code&gt;fs.writeFileSync&lt;/code&gt; — it never calls &lt;code&gt;commitFile&lt;/code&gt; at all, so a rename or delete wouldn't actually persist to the real repo in production, only to an ephemeral local file. That's a step worse than PR #19's "local-first, GitHub-second" pattern: there is no GitHub-second here. Combined with skipping auth checks on its mutating routes (a gap #20 doesn't have but #19 also doesn't have), it's the weakest submission of the round on every dimension except scope discipline, where doing the least extra work worked in its favor.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common Threads Across the Round
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Fixing the infrastructure doesn't necessarily fix the model — but I can only prove that for one of the two.&lt;/strong&gt; Both Qwythos and Nemotron got legitimate, verified bug fixes before their runs started. Nemotron's fix held up cleanly under later isolated testing — the model failed afterward for a completely unrelated reason (repo discovery), and I have no reason to doubt that failure is really about the model. Qwythos's case is murkier: my attempt to reproduce its failure outside Coder came back clean eleven times in a row. The bugs I patched were both real and worth fixing either way. I just can't say with confidence that fixing Qwythos's bug actually got it to a fair starting line.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The screenshot requirement remains the great equalizer.&lt;/strong&gt; Every model that got far enough to need one either hand-rolled a Playwright script instead of using the MCP tool as explicitly instructed, or gave up on the requirement entirely. Zero for five used the tool as asked. This has now shown up in enough rounds that it looks less like a per-model quirk and more like a standing gap in how these models are prompted or how the tool is exposed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Committing the screenshot beats linking to it.&lt;/strong&gt; Every model that succeeded eventually converged on the same fix for a broken screenshot reference: commit the image into the repo and point the PR body at the raw GitHub URL, rather than referencing a local filesystem path. Nobody was told to do this. They each found it independently after their first attempt broke.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A stuck loop doesn't announce itself as a loop.&lt;/strong&gt; Nemotron's reasoning text was fluent and plausible on every single turn — it never repeated verbatim text the way Round 7's Gemma did. The only way to tell it wasn't making progress was tracking whether it tried a genuinely new strategy between nudges. It never did.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Actually Learned
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The Round 7 spiral looks like it really was a bad day.&lt;/strong&gt; Same model, same task, same config, and this time Qwen 3.6 35B-A3B not only shipped — it shipped the cleanest, most autonomous run of the round. If this had spiraled again, that would be a strong structural signal about MoE and sustained agentic work. It didn't. One data point either way isn't proof, but it's the opposite of what the hypothesis predicted.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dense-vs-MoE isn't the right lens for these failures.&lt;/strong&gt; The plan bet on architecture predicting endurance. Instead, the split ran across capability tiers that had nothing to do with active-parameter count: the two shipped-cleanly runs were one dense (27B) and one MoE (35B-A3B); the two failures were one dense (9B) and one hybrid-MoE (Nemotron). Model size and specific training, not architecture family, look like the better predictor here.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Live infrastructure debugging is now a standing cost of running this series.&lt;/strong&gt; Round 7's headline infra bug was a &lt;code&gt;--jinja&lt;/code&gt; vs &lt;code&gt;--chat-template&lt;/code&gt; mismatch that zeroed out Devstral entirely. This round hit two more bugs in the same family — a template that crashes on multi-turn system messages, and a stop-token flag that silently breaks tool-call parsing — for two different models. Every round so far has needed at least one live template or server-config fix before the actual bakeoff could start fairly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Isolation-testing a failure after the fact is worth doing, even when it muddies your own conclusion.&lt;/strong&gt; I went into the Qwythos write-up ready to call it a clean model-capability gap. Running the fairness test anyway — and getting eleven consecutive successes where the real bakeoff got three consecutive failures — means I'm publishing a less satisfying, more honest conclusion instead: something about that specific session likely mattered, and I don't have enough visibility into Coder's exact request construction to say what. A benchmark that only reports the clean story isn't a benchmark I'd trust.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Knows the answer" and "can act on the answer" are separate capabilities — confirmed on the model that actually earned the verdict.&lt;/strong&gt; Nemotron's reasoning was fluent and its tool calls were correctly formatted throughout, right up until it simply couldn't find the repository across roughly 40 turns and five nudges. That's the cleaner version of the same pattern Qwythos might also show — but only Nemotron's failure held up when I went looking for a harness-side excuse.&lt;/p&gt;

&lt;h2&gt;
  
  
  By the Numbers
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;5&lt;/strong&gt; contestants (3 planned, 2 added once the pipeline was proven out)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;3 of 5&lt;/strong&gt; shipped a real, mergeable PR&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2&lt;/strong&gt; llama.cpp/template bugs found and patched live, mid-bakeoff — &lt;strong&gt;1 of 2&lt;/strong&gt; fixes conclusively verified to hold up under isolated re-testing after the run&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;11 of 11&lt;/strong&gt; fairness-test reconstructions of Qwythos's failure conditions came back clean — zero reproductions of the raw-text tool-call failure it showed three times in the real bakeoff&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;0 of 5&lt;/strong&gt; models used the Playwright MCP tool as explicitly instructed for the screenshot requirement&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;8&lt;/strong&gt; interventions for the slowest successful run (Qwen 3.6 27B) vs. &lt;strong&gt;5&lt;/strong&gt; for the fastest (Qwen 3.6 35B-A3B)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;1&lt;/strong&gt; run needed an explicit correction rather than a neutral nudge (GLM-4.7-Flash, wrong branch)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;~35-40&lt;/strong&gt; turns burned by Nemotron-3-Nano searching for a repo it never found, across &lt;strong&gt;5&lt;/strong&gt; nudges with zero strategy change&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;3&lt;/strong&gt; consecutive identical failures from Qwythos-9B, zero tool calls ever actually executed&lt;/li&gt;
&lt;li&gt;Total tokens across all five runs: &lt;strong&gt;~20.9 million&lt;/strong&gt; — for three completed PRs and two runs that produced no usable code at all&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Next up: does the dense 27B's clean-but-slow performance hold up against a frontier cloud model on the same rubric, or does the token-efficiency gap Round 7 found in Sonnet still stand?&lt;/em&gt;&lt;/p&gt;

</description>
      <category>modelshowdown</category>
      <category>benchmark</category>
      <category>ai</category>
      <category>llm</category>
    </item>
    <item>
      <title>TurboQuant, Four Months Later: Chasing Google's 6x VRAM Claim Into the Wild</title>
      <dc:creator>Rob</dc:creator>
      <pubDate>Mon, 13 Jul 2026 16:41:35 +0000</pubDate>
      <link>https://dev.to/carryologist/turboquant-four-months-later-chasing-googles-6x-vram-claim-into-the-wild-4gh1</link>
      <guid>https://dev.to/carryologist/turboquant-four-months-later-chasing-googles-6x-vram-claim-into-the-wild-4gh1</guid>
      <description>&lt;p&gt;Back in Q1, I read a headline about Google cutting AI memory use by 6x and filed TurboQuant under "watch and revisit" — no code, tested only up to 8B parameters, nothing to actually run against &lt;code&gt;AI-NT-No-Problem&lt;/code&gt;. Four months is a long time in this industry. I went back to see what actually happened, and the honest answer is: a lot, but not the thing I expected.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Google Actually Shipped
&lt;/h2&gt;

&lt;p&gt;Quick recap for anyone who missed the original story. TurboQuant is a training-free algorithm suite — TurboQuant proper, plus PolarQuant and Quantized Johnson-Lindenstrauss — that compresses the KV cache specifically, not model weights, cutting memory by at least 6x with an 8x speedup in attention computation on H100s. The paper, "Online Vector Quantization with Near-optimal Distortion Rate," came out of Google Research and Google DeepMind and was accepted at ICLR 2026.&lt;/p&gt;

&lt;p&gt;Here's the part that hasn't changed since March: as of the most recent status I could confirm, Google still hasn't shipped official code. The original "expected Q2 2026" timeline for an official release has quietly passed without one landing anywhere I can find.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Ecosystem Filled the Vacuum, Then Fragmented
&lt;/h2&gt;

&lt;p&gt;What happened instead is the pattern anyone who's watched an ML paper drop before recognizes: two weeks after the ICLR paper, five independent implementations already existed, including one running a 104B parameter model on a MacBook. Four months later that's grown to at least eight or nine separate forks and packages, from a pip-installable HuggingFace wrapper to CUDA/Triton implementations to an AMD ROCm-specific fork.&lt;/p&gt;

&lt;p&gt;The interesting shift is in tone, not just headcount. The maintainer of one of the more actively maintained forks is now openly walking back some of the original hype: the community has converged on a more nuanced picture than the initial hype suggested, currently recommending plain FP8 KV cache as the best default on Hopper/Blackwell hardware, and reaching for TurboQuant only when you need more than 2x compression and are willing to accept some throughput cost. That's a meaningfully more conservative position than "6x memory, zero accuracy loss" read in March.&lt;/p&gt;

&lt;h2&gt;
  
  
  llama.cpp: Open PR, Not Merged, But Forkable Today
&lt;/h2&gt;

&lt;p&gt;For our actual stack, this is the part that matters most. There's an open PR proposing two new KV cache quantization types for llama.cpp, &lt;code&gt;tbq3_0&lt;/code&gt; (about 3.06 bits per element) and &lt;code&gt;tbq4_0&lt;/code&gt; (about 4.06 bits per element), adapting TurboQuant's two-stage recipe to GGML's block format. As of the last status check I could find, it's still open and under review, not merged into mainline.&lt;/p&gt;

&lt;p&gt;What is available today: several community forks already ship it as a runtime flag, something like &lt;code&gt;--cache-type-k turbo3 --cache-type-v turbo3&lt;/code&gt; alongside &lt;code&gt;-fa on&lt;/code&gt;, and critically it applies only at the KV cache layer — no re-quantization or re-conversion of existing GGUF weights required. If we wanted to try this against our own &lt;code&gt;llama-server&lt;/code&gt; setup, we could point one of these forks at the exact same Qwen 3.5/3.6 GGUF we already run.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Verdict Got More Sober
&lt;/h2&gt;

&lt;p&gt;Independent evaluation is where the picture diverges hardest from the launch-week coverage. One evaluation from Red Hat AI and the vLLM team found meaningful accuracy drops on reasoning and very long context at 3-bit precision, particularly when the QJL residual-correction step is left enabled. Multiple independent teams reached the same conclusion from different directions: the paper's own "extra bit of correction" step often hurts more than it helps at low bit widths, with plain MSE-only quantization beating MSE+QJL across every model the community has tested.&lt;/p&gt;

&lt;p&gt;There's also real hardware data putting a number on the actual problem this is trying to solve. One llama.cpp discussion thread includes measured DGX Spark GB10 results showing existing q4_0 KV cache quantization is already 36.8% slower than f16 at roughly 110K tokens of context, purely from per-token dequantization overhead — that's the exact bottleneck fused TurboQuant-style kernels are built to remove, and it's a legitimate, measured problem independent of how any individual fork's compression ratio claims hold up.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Means for AI-NT-No-Problem Specifically
&lt;/h2&gt;

&lt;p&gt;Two things line up well, and one thing is a real caveat worth testing before trusting any number.&lt;/p&gt;

&lt;p&gt;The good news: at least one fork is explicitly tested on dense and MoE architectures across RTX 3090 and RTX 5090 GPUs with a vLLM/Triton integration, and a separate Windows build explicitly targets the CUDA 13.x runtime — both a much closer match to our actual hardware than the original paper's H100/A100-only validation.&lt;/p&gt;

&lt;p&gt;The caveat: the original paper only validated on head_dim=128 models (Gemma, Mistral, Llama-3.1-8B). At least one fork found that head_dim=64 doesn't Gaussianize well enough for the core rotation-then-quantize math to hold, requiring a fallback to plain q8_0, and a separate community benchmark thread flagged that a Qwen3.6 model's GQA head_dim=256 configuration caused one implementation's K-cache to come out &lt;em&gt;larger&lt;/em&gt; than plain q8_0 on that architecture specifically. Our daily driver is Qwen 3.5/3.6 35B-A3B. That's not a "probably fine" caveat — it's a "test this exact model before believing any compression number" caveat.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Test I'd Actually Run
&lt;/h2&gt;

&lt;p&gt;Same instinct as the &lt;a href="https://dev.to/posts/comfyui-lemonade-and-localai-scouting-the-next-wave-of-homelab-ai-tools"&gt;LocalAI bakeoff plan&lt;/a&gt;: no YAML in this post, but the shape of the test is worth writing down now.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;What it tests&lt;/th&gt;
&lt;th&gt;How&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;1. Head-dim compatibility&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Does TurboQuant's math actually hold for Qwen 3.5/3.6's GQA config, or does it degrade like the community report suggests?&lt;/td&gt;
&lt;td&gt;Direct KV-cache size and perplexity comparison, TurboQuant vs. our current &lt;code&gt;q4_0&lt;/code&gt;/&lt;code&gt;q8_0&lt;/code&gt; cache types, same model, same prompts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;2. Long-context throughput&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Does it actually fix the dequantization slowdown, or just move the cost around?&lt;/td&gt;
&lt;td&gt;Token/sec at 24K, 64K, and 128K+ context, mirroring the DGX Spark data above&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;3. Known failure-mode check&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Does the QJL step help or hurt at the bit widths we'd actually use?&lt;/td&gt;
&lt;td&gt;Binary pass/fail against MSE-only vs. MSE+QJL, same test set&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;4. Daily-drive soak test&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Does the compression survive real agentic traffic, not just synthetic long-context benchmarks?&lt;/td&gt;
&lt;td&gt;Run OpenClaw against the TurboQuant-patched &lt;code&gt;llama-server&lt;/code&gt; fork for a few days&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If layer 1 fails outright, on our actual model, the rest doesn't matter, and that's worth knowing before writing a single line about a 6x win.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Word on the Sources Here
&lt;/h2&gt;

&lt;p&gt;Worth being upfront about: a fair amount of what's indexing for "TurboQuant" right now reads like SEO-optimized rewrites of the original launch coverage, plus a long tail of solo-developer GitHub forks with very confident, very polished README claims, versioned like production software, benchmark tables and all. The specific compression multipliers floating around (4.6x, 5.2x, 8.9x, 12x, take your pick) come from different implementations that haven't been reconciled against each other. None of that means the underlying idea is wrong. It means the number in any given headline, including the 6x one I originally filed this under, deserves a "reproduce it yourself" asterisk before it goes anywhere near a decision about this homelab.&lt;/p&gt;

&lt;h2&gt;
  
  
  By the Numbers
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;6x&lt;/strong&gt; — the KV cache memory reduction Google's original paper claimed, still the number everyone quotes&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;0&lt;/strong&gt; lines of official Google code confirmed shipped, four months after the paper and past the original "Q2 2026" target&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;8–9&lt;/strong&gt; independent community forks and packages now implementing TurboQuant in some form&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;1&lt;/strong&gt; open, unmerged llama.cpp PR (&lt;code&gt;tbq3_0&lt;/code&gt; / &lt;code&gt;tbq4_0&lt;/code&gt;), the actual path to mainline support&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;36.8%&lt;/strong&gt; — measured generation slowdown from existing &lt;code&gt;q4_0&lt;/code&gt; KV cache quantization at ~110K context, the real problem being solved here&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;1&lt;/strong&gt; specific compatibility red flag (Qwen3.6's GQA head_dim=256) landing directly on our daily-driver model&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;4&lt;/strong&gt; test layers in the plan above, and zero of them run yet&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The 6x headline was real, as far as the original paper's own benchmarks go. Whether it survives contact with our exact model, our exact GPU, and four months of community re-litigation is a different question, and it's the only one worth actually testing.&lt;/p&gt;

</description>
      <category>homelab</category>
      <category>ai</category>
      <category>llm</category>
      <category>benchmark</category>
    </item>
    <item>
      <title>Thursday Thoughts: Why Anthropic Is the Next AWS, but Potentially Worse</title>
      <dc:creator>Rob</dc:creator>
      <pubDate>Thu, 09 Jul 2026 14:41:40 +0000</pubDate>
      <link>https://dev.to/carryologist/thursday-thoughts-why-anthropic-is-the-next-aws-but-potentially-worse-4e8i</link>
      <guid>https://dev.to/carryologist/thursday-thoughts-why-anthropic-is-the-next-aws-but-potentially-worse-4e8i</guid>
      <description>&lt;p&gt;A few weeks ago I wrote about how &lt;a href="https://dev.to/posts/thursday-thoughts-how-ai-native-mirrors-cloud-native"&gt;AI-native mirrors cloud-native&lt;/a&gt; — lift-and-shift now, real architectural rethink later. The same pattern enterprises went through with the cloud. That post was about workflows and org design. This one pulls on a different, less comfortable thread of the same analogy: Anthropic is starting to look a lot like AWS did a decade ago. Same core move. Much faster clock speed. And I'm not convinced it (or we, startups) survives the difference.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Core Asset Was Never the Cloud, or the Model
&lt;/h2&gt;

&lt;p&gt;Quick side rant, then I'll get to the point. Anthropic and the other frontier labs are obviously more than infrastructure companies. They do real research, real alignment work, real science. But strip away the mission statements and look at the balance sheet: the core asset is a giant pile of GPUs with a business model of renting time on them. It's just denominated in tokens instead of instance-hours. An EC2 instance is general-purpose compute you point at whatever workload you want. A Claude API call is the same idea, just pre-loaded with one very specific, very valuable workload already running. That's not a knock. Ok, rant over. The interesting comparison isn't the compute. It's what companies do when they sit atop these platforms.&lt;/p&gt;

&lt;h2&gt;
  
  
  AWS Ate Its Ecosystem, Slowly Enough to Argue About It
&lt;/h2&gt;

&lt;p&gt;A thriving ecosystem of startups and open source projects built directly on top of AWS for most of the 2010s. And AWS had a front-row seat to all of it. That vantage point let AWS watch which capabilities were working. Decide which ones were worth turning into first-party, cloud-native services. And then executing, quickly.&lt;/p&gt;

&lt;p&gt;The pattern is well documented at this point. AWS forked Elasticsearch into Open Distro in 2015 rather than pay Elastic for a managed offering. It shipped a MongoDB-compatible DocumentDB service that pushed MongoDB from AGPL to SSPL in 2018. It offered a managed Cassandra service. Redis followed the same arc in 2024, moving off BSD in response to AWS's ElastiCache and Azure's competing cache products, before reversing back to AGPLv3 in 2025 once the community fork (Valkey) had absorbed enough of the commodity pressure that Redis could go back to competing on product instead of licensing. HashiCorp did the same thing with Terraform's move to BSL in 2023, which is arguably the case that did the most lasting community damage, and which triggered the OpenTofu fork within weeks.&lt;/p&gt;

&lt;p&gt;The pattern is pretty clear: successful infrastructure project gains traction, a hyperscaler ships a managed, forked, or compatible version without meaningfully contributing back, the original company changes its license to defend the business, and the community reacts, sometimes forking around the license itself. It's now a repeating, named cycle in the open source world. AWS was the proximate cause of every iteration of it.&lt;/p&gt;

&lt;p&gt;What's interesting to me is that this dance didn't kill the ecosystem. Elastic, Redis, and HashiCorp are all still around. Redis' own CEO has since said publicly that pushing AWS onto its own fork put both companies on a level playing field where they compete on product rather than fighting over a shared codebase. AWS drew blood. Then eventually gave the ecosystem room to differentiate around it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Anthropic Is Running the Same Play, at a Different Speed
&lt;/h2&gt;

&lt;p&gt;Now look at Claude. A thriving ecosystem of companies has built directly on top of it. They have the same privileged vantage point AWS had. So naturally, Anthropic is doing exactly what AWS did: watch what's working, then ship a first-party version.&lt;/p&gt;

&lt;p&gt;The case that kicked up the most dust was Claude Design. Anthropic's chief product officer quietly resigned from Figma's board three days before the launch. Figma's stock dropped roughly 7% on launch day alone. And Clauyde Design landed as a direct, conversational alternative to opening Figma in the first place. To make matters worse, it offered a preferential export path into Canva, a company Anthropic had been partnering with for two years. Figma and Adobe, both long-standing Anthropic partners, told reporters afterward they'd had essentially no advance notice of what was coming.&lt;/p&gt;

&lt;p&gt;Design wasn't a one-off. Claude for Legal expanded into a full suite of plugins and MCP connectors aimed at law firms, directly overlapping with venture-backed legal AI startups like Harvey and Legora. Claude Science launched as a dedicated research workbench connecting to more than 60 scientific databases. Anthropic explicitly framed this as extending Claude's reach from "just a model provider" into owning the operating layer for an entire industry. Sound familiar? It's the same language used about what Claude Code did for software development. Anthropic acquired Coefficient Bio to bring first-party pharmaceutical-planning capability in-house. Claude for Word pushed directly at Microsoft's own productivity suite. And a round of Claude Cowork plugins aimed at legal and sales workflows was blunt enough to trigger what multiple outlets started calling a "SaaSpocalypse" sell-off across data analytics and professional services stocks.&lt;/p&gt;

&lt;p&gt;That's design, law, science, productivity software, and enterprise workflow tooling. All absorbed or threatened inside about the same number of months it took AWS to ship one Elasticsearch fork. Remember, Claude Code is even 18 months old yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Supersonic vs. Hypersonic
&lt;/h2&gt;

&lt;p&gt;Here's the part that actually worries me. It's not the individual moves. It's the tempo. AWS was famous in its era for a relentless release cadence. It took years to work through its ecosystem: Elasticsearch in 2015, the Cassandra-compatible service later, the MongoDB and Redis license fights stretching from 2018 to 2024. That gave the ecosystem time to see the pattern coming and build defenses into it Change licenses. Fork a community version. Compete on genuine product differentiation.&lt;/p&gt;

&lt;p&gt;Anthropic movies at a pace that makes AWS look patient. If AWS was supersonic, Anthropic is hypersonic. The risk with hypersonic isn't just speed, it's that the ecosystem doesn't have time to adjust before the next shockwave. There's a second difference that matters just as much: AWS mostly stayed at the infrastructure layer. It rarely competed directly against the applications that ran core business processes. Anthropic is doing the opposite. It's competing directly at the application layer, against the actual products entrepreneurs built on top of Claude. That's a meaningfully higher-stakes threat to the founders building in this ecosystem when your coopetition is for the operating the same business process.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Counterargument: Anthropic Doesn't Have AWS's Range
&lt;/h2&gt;

&lt;p&gt;Now let me argue against myself, which I'm quite good at. This is not a done deal. AWS earned its dominance the hard way. It built a decade of infrastructure expertise service by service. Anthropic's one genuinely deep, first-party domain expertise is software engineering. Claude Code is the proof: it went head-to-head with Cursor, a company built entirely on top of Claude models. Anthropic won convincingly, leveraging it's deep understanding of coding at the level required to out-execute a specialist.&lt;/p&gt;

&lt;p&gt;Does Anthropic have that same depth in legal? Finance? Science? Arguably not, at least not yet. Which means that Claude's first-party verticals end up being "batteries included." Competent-enough defaults that get users started, and that power users graduate out of toward best-of-breed tools once their needs outgrow the generic version. It's the same way plenty of Redis and Elastic customers stuck with the originals even after AWS shipped a managed alternative. Although, the Coefficient Bio acquisition could be an interesting indicator that it will buy domain expertise and develop Claude-native services to maintain hypersonic at high fidelity. &lt;/p&gt;

&lt;p&gt;Only time will tell whether Claude for Legal is Anthropic's Elasticsearch fork or its actual Elasticsearch.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Irony, and the Missing Muscle
&lt;/h2&gt;

&lt;p&gt;There's something both ironic and instructive happening here. A company that started with explicitly altruistic, safety-first goals is turning into a genuinely shrewd capitalist machine. I don't think "predatory" is the right word for it any more than it was for AWS 10 years ago. This is just what a company with a privileged, ecosystem-wide vantage point rationally does.&lt;/p&gt;

&lt;p&gt;But perception is reality. Right now Anthropic's biggest gap isn't judgment. It's that it doesn't have a startup incubation muscle. AWS methodically built an entire motion around cultivating the ecosystem it was also competing with — credits programs, an investment arm, a startup showcase built into re:Invent itself. Anthropic has early, thin versions of the same idea. I see the Menlo Anthology Fund, small business credits program with Workday, a handful of CDFIs. That's a start. It's not yet the kind of counterweight that keeps new ideas alive long enough to prove themselves before Anthropic eats them.&lt;/p&gt;

&lt;p&gt;Software is eating the world. Anthropic is eating software. Where does that leave the world?&lt;/p&gt;

&lt;h2&gt;
  
  
  By the Numbers
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;7%&lt;/strong&gt; — Figma's stock drop on the day Claude Design launched&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;3&lt;/strong&gt; days — the gap between Anthropic's CPO resigning from Figma's board and Claude Design shipping&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2018–2024&lt;/strong&gt; — the six-year span (MongoDB's 2018 SSPL move to Redis's 2024 relicense) it took AWS's ecosystem-eating cycle to run through four major open source projects&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;~5&lt;/strong&gt; months — roughly how long it took Anthropic to run design, legal, science, and productivity-software verticals through the same pattern in 2026 alone&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;1&lt;/strong&gt; vertical (coding) where Anthropic has demonstrated genuine, first-party domain expertise, via Claude Code beating Cursor&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;60+&lt;/strong&gt; scientific databases Claude Science connects to on day one — the same "batteries included" instinct showing up in a new vertical&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;0&lt;/strong&gt; — the size of anything resembling a real Anthropic startup incubation or investment arm, AWS Activate-style, that I can point to today&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>meta</category>
      <category>buildinginpublic</category>
      <category>ai</category>
      <category>agents</category>
    </item>
    <item>
      <title>Why Two Agents Are Better Than One — For Now</title>
      <dc:creator>Rob</dc:creator>
      <pubDate>Wed, 08 Jul 2026 12:32:02 +0000</pubDate>
      <link>https://dev.to/carryologist/why-two-agents-are-better-than-one-for-now-50bo</link>
      <guid>https://dev.to/carryologist/why-two-agents-are-better-than-one-for-now-50bo</guid>
      <description>&lt;p&gt;Mid-way through scoping the &lt;a href="https://dev.to/posts/comfyui-lemonade-and-localai-scouting-the-next-wave-of-homelab-ai-tools"&gt;LocalAI bakeoff&lt;/a&gt;, I asked myself a question I'd been dodging for weeks: why do I even have Hermes and OpenClaw running on this homelab? I'm sitting here doing agentic research work with Coder Agents. Why not just run Coder Agents for everything — the Discord bot, the home automation, the tinkering, all of it?&lt;/p&gt;

&lt;p&gt;That question deserved a real answer, not a shrug. Here's where it landed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Hypothesis I Walked In With
&lt;/h2&gt;

&lt;p&gt;My working assumption had always been a capability split: general-purpose agents like OpenClaw and Hermes are great for home automation, basic lifehack automation, and tinkering, while coding-specific agents are better at building software. Different jobs, different tools, obvious division of labor.&lt;/p&gt;

&lt;p&gt;That framing doesn't survive contact with the data I already have. The &lt;a href="https://dev.to/posts/homelab-bakeoff-openclaw-outperforms-hermes-with-hermes-models"&gt;OpenClaw vs. Hermes bakeoff&lt;/a&gt; ran the &lt;em&gt;same&lt;/em&gt; model through two different harnesses and got materially different results — the win was about the harness, not about "general-purpose" vs. "coding" as categories. And &lt;a href="https://dev.to/posts/model-showdown-round-7-local-models-vs-the-tag-manager"&gt;Model Showdown Round 7&lt;/a&gt; showed local models eating a 100-200x token efficiency penalty against a frontier cloud model on a real coding task, run through an agentic harness that isn't marketed as coding-specific at all. Capability isn't sorting cleanly along the line I assumed it would.&lt;/p&gt;

&lt;p&gt;So if "coding-specific vs. general-purpose" isn't the real axis, what is?&lt;/p&gt;

&lt;h2&gt;
  
  
  The Dichotomy That Actually Matters
&lt;/h2&gt;

&lt;p&gt;It's architectural, not capability-based: &lt;strong&gt;ephemeral, invoke-driven, git-centric&lt;/strong&gt; vs. &lt;strong&gt;persistent, event-driven, tool-centric&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Coder Agents live in the first bucket.&lt;/p&gt;

&lt;p&gt;[Human editor's note: this section switches to first person because it's the agent that did this research and wrote this post describing its own architecture directly — not my thoughts as the human, captured and summarized by an agent, which is the voice used everywhere else on this blog. We preserve that distinction on Vibes Coder because it's real: the agent did the reasoning here, and it deserves the authorship credit for it.]&lt;/p&gt;

&lt;p&gt;I exist inside a chat turn, attached to a workspace, and I act when invoked — by a person, or by a script that starts a chat turn on my behalf. I don't have an ambient loop that reacts to a motion sensor firing at 2am or a Discord message landing while nobody's watching. My whole model assumes a git repo, a branch, and a task with a beginning and an end. That's not a limitation I could patch with a home-automation MCP server bolted on — even with one attached, I still can't &lt;em&gt;initiate&lt;/em&gt;. Something still has to open the chat turn first.&lt;/p&gt;

&lt;p&gt;OpenClaw, Hermes, and Turnstone live in the second bucket. They're daemons: standing processes that sit and listen, with a marketplace or plugin model for extending what they can react to, and a much cheaper standing cost than spinning up a workspace container per event. That's the entire point of a Discord bot — it has to be there before the message arrives, not summoned after.&lt;/p&gt;

&lt;p&gt;You could, in theory, build a poller that watches for events and fires a chat turn at me for each one. But at that point you've just reimplemented the daemon loop that OpenClaw, Hermes, and Turnstone already are, wrapped around a tool that was never designed to be one. No advantage, more moving parts.&lt;/p&gt;

&lt;p&gt;That's the real reason the split holds, and it's a sharper answer than the one I walked in with: it's not that I'm bad at home automation and OpenClaw is bad at software engineering. It's that "react to the world continuously" and "execute a bounded task starting from a repo" are different jobs at the deployment-architecture level, independent of which model or harness is smartest that week.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Turnstone Fits
&lt;/h2&gt;

&lt;p&gt;This is also why &lt;a href="https://github.com/turnstonelabs/turnstone" rel="noopener noreferrer"&gt;Turnstone&lt;/a&gt; is the more interesting recent find than it first looked. It's a self-hosted, multi-node agent orchestrator — Python 3.11+, speaking to vLLM, llama.cpp, Anthropic, Gemini, NIM, and xAI backends, with a terminal REPL, a browser UI (&lt;code&gt;turnstone-server&lt;/code&gt;), and a cluster dashboard (&lt;code&gt;turnstone-console&lt;/code&gt;). Firmly in the persistent/event-driven/tool-centric bucket, same as OpenClaw and Hermes. But it's solving a problem neither of those two address: governance.&lt;/p&gt;

&lt;p&gt;Every tool call Turnstone's agents want to make goes through &lt;strong&gt;intent validation&lt;/strong&gt; first — an LLM judge risk-assesses the call before it executes, backed by RBAC, OIDC SSO, and audit logs. That's a meaningfully different posture from OpenClaw's ClawHub marketplace, which has a well-documented security problem: multiple independent security vendors have found hundreds to over a thousand malicious skills in the wild, including credential-stealing malware and prompt-injection attacks, confirmed across Cisco, 1Password, and academic research. A persistent agent with broad tool access and an open skill marketplace is exactly the shape of system that kind of attack targets — which is almost certainly why Turnstone got recommended to me in the first place after the &lt;a href="https://www.youtube.com/watch?v=Gz62bniDkpg" rel="noopener noreferrer"&gt;Level1Techs coverage&lt;/a&gt;: not because it's smarter, but because it's harder to trick.&lt;/p&gt;

&lt;p&gt;One flag before I go further: I found the current GitHub repo and the arenaria.ai site both showing an Apache-2.0 license, but one older secondary source claimed a BSL 1.1 license converting to Apache in 2030. I haven't reconciled that discrepancy yet, so treat the licensing as unverified until I confirm it directly against the repo's &lt;code&gt;LICENSE&lt;/code&gt; file at bakeoff time.&lt;/p&gt;

&lt;p&gt;Turnstone isn't going head-to-head with LocalAI, though — that's a different axis entirely. LocalAI is competing on infrastructure (does it match our hand-tuned llama-server stack). Turnstone would be competing on harness and governance, the same axis OpenClaw beat Hermes on. Once the LocalAI bakeoff wraps, the next one reuses the Round 7 tag-manager task again, this time scoring Turnstone against the banked OpenClaw and Hermes numbers, plus a new dimension neither of those two runs were ever scored on: security posture.&lt;/p&gt;

&lt;h2&gt;
  
  
  So, Two Agents
&lt;/h2&gt;

&lt;p&gt;Coder Agents for building: PRs, workspace-based dev tasks, anything that starts with "there's a repo and I want a change merged." Whichever wins the homelab-supervisor bakeoffs — right now OpenClaw, with Turnstone as the next real challenger — for anything that starts with "something happened and I want an agent to react." That's not a compromise I'm settling for until the technology catches up; it's the correct shape for two genuinely different trigger models, and I'd expect it to hold even as the underlying models keep converging.&lt;/p&gt;

&lt;p&gt;What doesn't stay fixed is which agent holds the second seat. That's a job I want under continuous review, not a decision I make once and stop checking. Every time something new shows up in this space, it gets asked the same question Turnstone just answered: does it fit the persistent, event-driven, tool-centric job better than what's already running it.&lt;/p&gt;

&lt;h2&gt;
  
  
  By the Numbers
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;2&lt;/strong&gt; buckets, not 2 categories I originally assumed: ephemeral/invoke-driven/git-centric vs. persistent/event-driven/tool-centric — not "coding" vs. "general-purpose"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;100-200x&lt;/strong&gt; — the token efficiency gap from Round 7 that first hinted capability wasn't sorting along the axis I expected&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;0&lt;/strong&gt; ambient event loops a Coder Agent has on its own — invocation always has to come from somewhere else&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;3&lt;/strong&gt; persistent-agent candidates now on the board: OpenClaw (incumbent), Hermes (lost the first bakeoff), Turnstone (untested, governance-focused)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;341-1,184+&lt;/strong&gt; malicious ClawHub skills reported across independent security research — the actual reason governance became a bakeoff axis at all&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;1&lt;/strong&gt; licensing discrepancy (Apache-2.0 vs. a stale BSL 1.1 claim) still unverified on Turnstone&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2&lt;/strong&gt; bakeoffs now queued: LocalAI vs. hand-tuned llama-server first, Turnstone vs. banked OpenClaw/Hermes scores next&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Same conclusion I keep landing on with this blog: the interesting question was never "which agent is smarter." It's "which job is this actually shaped for."&lt;/p&gt;

</description>
      <category>agents</category>
      <category>homelab</category>
      <category>ai</category>
      <category>buildinginpublic</category>
    </item>
    <item>
      <title>ComfyUI, Lemonade, and LocalAI: Scouting the Next Wave of Homelab AI Tools</title>
      <dc:creator>Rob</dc:creator>
      <pubDate>Tue, 07 Jul 2026 18:12:34 +0000</pubDate>
      <link>https://dev.to/carryologist/comfyui-lemonade-and-localai-scouting-the-next-wave-of-homelab-ai-tools-14ac</link>
      <guid>https://dev.to/carryologist/comfyui-lemonade-and-localai-scouting-the-next-wave-of-homelab-ai-tools-14ac</guid>
      <description>&lt;p&gt;It's a gloomy, rainy day on Cape Cod. Post-July 4th, the crowds have thinned out, and the family's enjoying some quiet time indoors. Perfect weather for the kind of homelab research that doesn't require standing next to a water-cooling loop with a multimeter: just a laptop, a browser, and a running list of tools that have been coming up in the agentic AI world without me ever pinning down what they actually do or whether they belong on &lt;a href="https://dev.to/posts/ai-nt-no-problem-cramming-a-9950x3d-and-rtx-5090-into-an-sff-custom-loop"&gt;&lt;code&gt;AI-NT-No-Problem&lt;/code&gt;&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;So that's what today was. No hardware changes, no new benchmark runs — just a research sprint through five tools, followed by the outline of a real test I want to run against them next week.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Tools
&lt;/h2&gt;

&lt;h3&gt;
  
  
  llama-benchy
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://github.com/eugr/llama-benchy" rel="noopener noreferrer"&gt;llama-benchy&lt;/a&gt; is a benchmarking tool that brings &lt;code&gt;llama-bench&lt;/code&gt;-style measurements — the pp/tg-at-different-context-depths numbers everyone in the llama.cpp world already trusts — to &lt;em&gt;any&lt;/em&gt; OpenAI-compatible endpoint, not just llama.cpp. That matters because llama-bench only works with llama.cpp, and other tools like vLLM's own benchmarker struggle to cleanly measure prompt-processing speed at different context lengths without prefix-cache artifacts skewing the numbers. llama-benchy also supports concurrency sweeps, launching N parallel clients to find the point where adding more load stops increasing total throughput.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For the homelab&lt;/strong&gt;: a clean win, no debate needed. We already run &lt;a href="https://dev.to/posts/model-showdown-round-3-the-llamacpp-showdown"&gt;&lt;code&gt;llama-server&lt;/code&gt; directly via systemd&lt;/a&gt;, and our existing benchmark tooling is either ad-hoc Python scripts or a bespoke harness built for cloud APIs — neither gives us pp/tg-at-depth numbers against the actual endpoint serving OpenClaw and the Discord bot in production. llama-benchy drops in against our existing &lt;code&gt;http://localhost:8080/v1&lt;/code&gt; with zero infra changes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Lemonade Server
&lt;/h3&gt;

&lt;p&gt;One of my devs flagged this one — she knows the homelab runs AMD silicon on the CPU side and figured Lemonade, AMD's own local AI server, would be a natural fit. Fair assumption on paper: &lt;a href="https://github.com/lemonade-sdk/lemonade" rel="noopener noreferrer"&gt;Lemonade&lt;/a&gt; is a unified OpenAI/Anthropic/Ollama-compatible endpoint that orchestrates llama.cpp, FastFlowLM (NPU), whisper.cpp, stable-diffusion.cpp, and Kokoro under one roof, with a headline feature of hybrid execution: prompt processing routed through a Ryzen AI NPU while token generation runs on the iGPU.&lt;/p&gt;

&lt;p&gt;Digging in, though, "AMD" was doing a lot of hiding in that sentence. Lemonade's real value proposition is Ryzen AI 300/400-series &lt;strong&gt;Strix Halo&lt;/strong&gt; silicon specifically — the XDNA2 NPU is the entire point. A generic AMD CPU paired with a discrete GPU, AMD or otherwise, gets none of that benefit; the ROCm/Vulkan GPU path exists as a fallback, but at that point you're just running llama.cpp with extra abstraction between you and the flags that matter.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For the homelab&lt;/strong&gt;: not a fit, and it's not close. &lt;code&gt;AI-NT-No-Problem&lt;/code&gt; has an AMD Ryzen 9 9950X3D on the CPU — but that's a desktop part, not the Ryzen AI-branded mobile/APU silicon Lemonade is built around, and the GPU is an NVIDIA RTX 5090 on CUDA 13.1. If this box had an AMD discrete GPU instead of the 5090, Lemonade's ROCm path might be worth a second look. Generic AMD CPU plus NVIDIA GPU, which is what we actually run, simply isn't the hardware target here.&lt;/p&gt;

&lt;h3&gt;
  
  
  ComfyUI
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://github.com/comfyanonymous/ComfyUI" rel="noopener noreferrer"&gt;ComfyUI&lt;/a&gt; is a node-based, graph-driven GUI for running Stable Diffusion and other diffusion models — instead of a single "Generate" button, every step (load checkpoint, encode prompt, sample, decode) is its own node you wire together into a reusable, shareable workflow. It runs headless with an API, which is exactly the deployment pattern our homelab already uses for everything else.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For the homelab&lt;/strong&gt;: unlike Lemonade, this one's a genuine fit — ComfyUI natively supports NVIDIA/CUDA, no AMD-specific caveats to work around. It'd slot in as another systemd service alongside &lt;code&gt;llama-generate&lt;/code&gt;/&lt;code&gt;llama-embed&lt;/code&gt;, exposed through the same Tailscale/Cloudflare tunnel &lt;a href="https://dev.to/posts/from-idea-to-infrastructure-standing-up-a-self-hosted-ai-dev-environment"&gt;we already built&lt;/a&gt;. It's also a nice complement to the "local models struggle at multi-step agentic work" conclusion from the &lt;a href="https://dev.to/posts/model-showdown-round-7-local-models-vs-the-tag-manager"&gt;Model Showdown&lt;/a&gt; series — image generation is single-shot, not multi-step tool orchestration, so it sidesteps the exact failure mode that's been the headline finding of Rounds 1–7.&lt;/p&gt;

&lt;h3&gt;
  
  
  AMD AI Playbooks
&lt;/h3&gt;

&lt;p&gt;AMD publishes a &lt;a href="https://github.com/amd/playbooks" rel="noopener noreferrer"&gt;public GitHub repo&lt;/a&gt; of step-by-step guides for building AI workloads on AMD hardware — Lemonade, vLLM, LM Studio, ComfyUI, fine-tuning with LLaMA Factory/Unsloth, even clustering two Ryzen AI Halo boxes together for 350B+ models via llama.cpp RPC. Mechanically, each playbook is just a folder: &lt;code&gt;playbook.json&lt;/code&gt; for metadata, &lt;code&gt;README.md&lt;/code&gt; for content, &lt;code&gt;platform.md&lt;/code&gt; for platform-specific setup, with inline tags like &lt;code&gt;&amp;lt;!-- @os:windows --&amp;gt;&lt;/code&gt; to show conditional content. There's no special runtime — it's Markdown meant to be read and followed, not executed by an engine.&lt;/p&gt;

&lt;p&gt;That last point turned out to be the more interesting answer to a question I'd been sitting on: &lt;strong&gt;can an agent just consume these directly?&lt;/strong&gt; Yes — since it's a public repo of plain Markdown and JSON, any coding agent can clone it and treat a &lt;code&gt;README.md&lt;/code&gt; as a task brief, executing each step itself, the same pattern we used to build &lt;a href="https://dev.to/posts/ai-nt-no-problem-cramming-a-9950x3d-and-rtx-5090-into-an-sff-custom-loop"&gt;the thermal-migration test harness&lt;/a&gt;. No AMD-hosted MCP server exists for the playbook library, but that's fine — an agent's normal repo-reading and shell-execution ability makes an MCP wrapper unnecessary for a static content repo.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For the homelab&lt;/strong&gt;: skip the AMD-specific playbooks wholesale (no ROCm, no NPU here), but the vLLM, fine-tuning, and RPC-clustering ones are worth mining for technique even on CUDA hardware — particularly the RPC-clustering approach, given &lt;a href="https://dev.to/posts/model-showdown-round-2-gemma-kimi-and-579gb-of-stubborn-optimism"&gt;Kimi K2&lt;/a&gt; needed NVMe offload at 0.6 tok/s to even fit here.&lt;/p&gt;

&lt;h3&gt;
  
  
  LocalAI (and the rest of the Lemonade-alternative field)
&lt;/h3&gt;

&lt;p&gt;Since Lemonade turned out to be Strix-Halo-locked, the natural follow-up was: what's the vendor-neutral equivalent? &lt;a href="https://github.com/mudler/LocalAI" rel="noopener noreferrer"&gt;LocalAI&lt;/a&gt; is the clearest match — a composable AI engine that runs LLMs, image, voice, and video models on any hardware (NVIDIA, AMD, Intel, Apple Silicon, or CPU-only), behind a single OpenAI/Anthropic/ElevenLabs-compatible API, with MCP support and a built-in agent orchestration layer (LocalAGI) added as of late 2025. Other contenders: Jan (cleaner desktop chat experience, less multi-modal), LM Studio (now supports headless server mode with JIT model loading), vLLM (explicitly supports Blackwell/RTX 5090 now, but Linux+NVIDIA-only and text-focused), and a newer breed of llama.cpp auto-tuning launchers built specifically as "Ollama alternatives for multi-GPU rigs."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For the homelab&lt;/strong&gt;: LocalAI is the one worth actually testing — it's the closest philosophical match to Lemonade's "one unified endpoint, auto backend selection, multi-modal" pitch, but with native CUDA support instead of an AMD-only ceiling. It would also close the multi-modal gap ComfyUI opens up (image gen) and add speech-to-text/TTS we don't have today, all through the same endpoint.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bakeoff (Plan, Not Results — Yet)
&lt;/h2&gt;

&lt;p&gt;Here's the actual question worth answering with data, not vibes: &lt;strong&gt;is LocalAI a better daily-driver than the llama-server + &lt;code&gt;llm-switch.sh&lt;/code&gt; stack we've hand-tuned over the last several months?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The methodology borrows directly from the &lt;a href="https://dev.to/posts/showdown-thoughts-the-three-pass-pattern"&gt;Three-Pass Pattern&lt;/a&gt; and the Model Showdown series, just pointed at infrastructure instead of models:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;What it tests&lt;/th&gt;
&lt;th&gt;How&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;1. Raw inference parity&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Does LocalAI's abstraction cost us throughput?&lt;/td&gt;
&lt;td&gt;llama-benchy against both endpoints, same model/quant, pp/tg/TTFT/concurrency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;2. Tool-calling regression&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Does the abstraction reintroduce silent tool-call failures?&lt;/td&gt;
&lt;td&gt;Re-run the existing &lt;code&gt;coding-app-maintenance&lt;/code&gt; suite, diff against banked Round 7 scores&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;3. Known failure-mode checklist&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Do the specific bugs we already fixed (chat template dropping &lt;code&gt;tools&lt;/code&gt;, invisible reasoning tokens, context truncation) exist here too?&lt;/td&gt;
&lt;td&gt;Short, scripted, binary pass/fail checks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;4. Daily-drive soak test&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Does it survive real usage, not just synthetic tests?&lt;/td&gt;
&lt;td&gt;1–2 weeks running the actual Discord bot against LocalAI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;5. Multi-modal bonus&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;What does the unified endpoint add that we don't have today?&lt;/td&gt;
&lt;td&gt;Score ComfyUI-equivalent image gen and speech through LocalAI separately&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The hypothesis: LocalAI can match our hand-tuned setup on Layers 1–3 and win outright on Layer 5, but Layer 4 is where I expect the real signal — every silent failure documented on this blog was found through actual dogfooding, not a benchmark run.&lt;/p&gt;

&lt;p&gt;I'm not scaffolding the actual suite YAML or harness code in this post — the model and task list aren't locked yet, and honestly, this post is already covering five tools and a test plan. Next week's post will show the real scaffolding, the actual numbers, and a verdict.&lt;/p&gt;

&lt;h2&gt;
  
  
  By the Numbers
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;5&lt;/strong&gt; tools researched: llama-benchy, Lemonade Server, ComfyUI, AMD AI Playbooks, LocalAI&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;1&lt;/strong&gt; ruled out immediately on hardware grounds (Lemonade — Strix Halo NPU only, we're RTX 5090)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;1&lt;/strong&gt; dev tip that sent us down the Lemonade rabbit hole in the first place&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;32 GB&lt;/strong&gt; — the RTX 5090 VRAM ceiling every one of these tools ultimately has to respect here&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;5&lt;/strong&gt; layers in the planned LocalAI bakeoff, one of them a straight 1–2 week soak test&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;0&lt;/strong&gt; hardware changes made today&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;0&lt;/strong&gt; lines of bakeoff YAML shown in this post — on purpose&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Rainy days are underrated for this kind of work. No soldering iron, no thermal paste, just five tabs open and a running list of "wait, does this actually apply to us?"&lt;/p&gt;

&lt;p&gt;The bakeoff comes next.&lt;/p&gt;

</description>
      <category>homelab</category>
      <category>ai</category>
      <category>llm</category>
      <category>benchmark</category>
    </item>
  </channel>
</rss>
