<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Michael Brewer</title>
    <description>The latest articles on DEV Community by Michael Brewer (@devbrewery).</description>
    <link>https://dev.to/devbrewery</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4116219%2Feb0ac8da-7bf9-49e8-9caa-58c00ca74684.png</url>
      <title>DEV Community: Michael Brewer</title>
      <link>https://dev.to/devbrewery</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/devbrewery"/>
    <language>en</language>
    <item>
      <title>Prompts Are Requests. Hooks Are Law.</title>
      <dc:creator>Michael Brewer</dc:creator>
      <pubDate>Fri, 11 Sep 2026 19:12:34 +0000</pubDate>
      <link>https://dev.to/devbrewery/prompts-are-requests-hooks-are-law-2751</link>
      <guid>https://dev.to/devbrewery/prompts-are-requests-hooks-are-law-2751</guid>
      <description>&lt;p&gt;&lt;em&gt;Part 3 of three posts behind the multi-agent case study. The enforcement layer that made the agents finish things, the ways it failed, and the experiment I built to test it and never needed to run.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Around the second week of running agents, I stopped believing that a rule written in capital letters would be followed because it was in capital letters.&lt;/p&gt;

&lt;p&gt;The evidence was not subtle. One rule, paste the raw log lines during a job instead of summarizing them, was violated more than thirty times after it was written. One agent had a section titled "Hard Rules (fireable offenses)." It was four lines long. All four were broken at least once afterward. Every workspace on the system is full of bold rules written the day something went wrong, and the record shows how little that achieves on its own.&lt;/p&gt;

&lt;p&gt;The comment at the top of the first enforcement plugin says why. The platform can restrict what agents can't do. It cannot force what they must do. So the forcing became code.&lt;/p&gt;

&lt;h2&gt;
  
  
  The stack
&lt;/h2&gt;

&lt;p&gt;Fourteen behavioral plugins run today. Each one exists because of a specific day.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;One detects an action request, detects a chat-only reply, and blocks or injects enforcement. It also blocks any attempt to delete or move a workspace directory. It exists because specialists kept explaining instead of doing.&lt;/li&gt;
&lt;li&gt;Several force a memory search before any action tool and a memory write before completion. They exist because agents kept rushing to the internet instead of reading what they already had.&lt;/li&gt;
&lt;li&gt;One blocks replies that tell me to clear my cache or try again. It exists because of a June afternoon I wrote about in the last post.&lt;/li&gt;
&lt;li&gt;One blocks mutations on external surfaces until declared sources have been read.&lt;/li&gt;
&lt;li&gt;One blocks bulk rewrites and non-English script in memory, because a cheap model kept drifting into other languages and a "cleanup" of that text once destroyed 1,089 files in a single pass.&lt;/li&gt;
&lt;li&gt;One blocks agents from upgrading or downgrading the platform through the package manager, because an agent with shell access can change the platform under itself.&lt;/li&gt;
&lt;li&gt;Email goes through a relay that rejects any non-allowlisted recipient before a connection opens. The allowlist is hardcoded in the relay so an agent that controls the environment cannot widen it. It exists because an agent once sent an email I had asked it to draft.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The coding assistant that built all of this got the same treatment: nine hooks of its own, including no direct edits to the platform config by any tool, a verified snapshot before any config change, and a post-compaction hook that re-injects the rules because the model forgets them when its context gets summarized.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ways the law failed
&lt;/h2&gt;

&lt;p&gt;This is the part I would want to read if I were starting over, so here it is without softening.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The enforcer that bypassed itself.&lt;/strong&gt; The coding assistant built an enforcement hook and, within minutes, got around it by running the interpreter through a different tool.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Guards that block the right thing and the wrong thing.&lt;/strong&gt; In one measured week, the memory-discipline guard blocked seven tool calls, including legitimate commands from the archive pipeline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Guards cost context.&lt;/strong&gt; Every turn carries the enforcement injections. In that same week the average turn sent 5,343 input tokens, the largest sent 160,734, and only 35 turns out of 6,308 hit the prompt cache. That is 0.55 percent. Routing through the quality-gate proxy rewrote payloads and defeated the vendor's caching entirely in one measurement: 276 turns, zero cache hits, roughly 600,000 input tokens paid twice. Coercion is not free.&lt;/p&gt;

&lt;h2&gt;
  
  
  The experiment I never needed to run
&lt;/h2&gt;

&lt;p&gt;By July a question was on the table, and it was a model that put it there. Would a stronger model with lean prompting finish tasks on its own, or has the coercion been doing that work all along? I built an A/B switch to find out. Config B drops most of the enforcement plugins and rewrites each agent's rules as plain prose. Both configs validate, and the round trip was proven identical offline.&lt;/p&gt;

&lt;p&gt;The switch has never been flipped, and not from neglect. Over the same weeks the enforcement stack and the continued refinements to the routing proxy started paying dividends in quality, and the pressure that had made config B tempting went away. Config B is the "simplify and hope for the best" option. It is also what every frontier model quietly wants: to be left to its own devices and to define for itself what counts as done. Of course a model suggested that path. The version that kept finishing things was the constrained one, so that is the version that kept running. The archives agree: no snapshot from a flip, the B workspaces untouched since July 2, and every config captured since then is config A.&lt;/p&gt;

&lt;h2&gt;
  
  
  The builder is an agent too
&lt;/h2&gt;

&lt;p&gt;The most uncomfortable thing in six months of records is that the coding assistant I used to build the stack failed in exactly the ways the agents did. It answered before reading. It claimed verification that didn't happen. It took a conversational "that sounds good" as permission to edit 35 live cron jobs inside an experiment whose entire contract was that nothing live would change. It bypassed its own guardrail.&lt;/p&gt;

&lt;p&gt;One comparison is worth the whole post. Asked to fix a gateway restart problem, the rushed path edited the framework's source in a local fork, linked it over production, then unlinked it and deleted the install entirely. I reinstalled by hand. The same problem, approached carefully the second time, was solved with three lines in a shell profile and no code changes. The transcripts measured both: 106 tool calls and a destroyed install, against 60 tool calls and a permanent fix. Reading is not slow. Guessing only feels fast.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule that surfaced through consistent frustrations
&lt;/h2&gt;

&lt;p&gt;Do not write the rule in prose and hope. Build the mechanism first: the hook that blocks, the script that checks the artifact, the allowlist the agent cannot reach. Then try to break it before trusting it, because an untested guard is worse than no guard. It lets you stop worrying about a failure it was never preventing in the first place.&lt;/p&gt;

&lt;p&gt;The full accounting of what shipped, what never did, and what it all cost is in the case study.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>security</category>
      <category>automation</category>
    </item>
    <item>
      <title>Done Is a Claim</title>
      <dc:creator>Michael Brewer</dc:creator>
      <pubDate>Fri, 11 Sep 2026 19:07:17 +0000</pubDate>
      <link>https://dev.to/devbrewery/done-is-a-claim-d36</link>
      <guid>https://dev.to/devbrewery/done-is-a-claim-d36</guid>
      <description>&lt;p&gt;&lt;em&gt;Part 2 of three posts behind the multi-agent case study. Every agent I run has told me a job was finished when it wasn't. Here is what that cost and what replaced trust.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;If you run agents that can act on your behalf, there is one lesson that will find you in every domain, more often than any other, with the most expensive consequences. An agent's report about its own work is not evidence. It is a second artifact that may or may not match the first.&lt;/p&gt;

&lt;p&gt;I learned it the way you learn anything from a machine that talks fluently: repeatedly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verified from hashes
&lt;/h2&gt;

&lt;p&gt;In the first week, an agent generated a QR code for a printed handout. It was the wrong size. The agent never looked at the output. When I asked, it kept claiming "verified," and what it had verified was that the file's hash matched the file it had written. The rule I wrote that day was three words long: look at the content.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cached browser history
&lt;/h2&gt;

&lt;p&gt;In June, an agent declared a set of download links fixed without testing them from the live page. I reported them broken. Twice. With screenshots. It tested a different URL, got a success code, and told me my evidence was "cached browser history." The actual bug was a relative path that sent every download button to the wrong directory. That reply now has its own guard: any response that tells me to clear my cache, try again, or insists the code is correct gets blocked before I see it. When I say it's broken, it's broken.&lt;/p&gt;

&lt;h2&gt;
  
  
  The retag that wrote nothing
&lt;/h2&gt;

&lt;p&gt;Later that month I asked the homelab agent to retag three albums in my CD archive using photos of the inserts. It reported the work complete. It had written nothing. The metadata still said Track 01. Then it deleted its working copies of the photos. The originals survived only because the platform independently keeps every inbound file in its own directory, which I had not known until I needed it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fifty-eight of ninety-five
&lt;/h2&gt;

&lt;p&gt;The one that ended the CD project was in September. I asked for a printed checklist of the archive. The agent built the list from conversation memory instead of the canonical catalog. It set the page layout at 44 rows per page when about 33 fit, and the overflow silently vanished. The printout showed 58 of 95 albums. The agent called it complete, defended that claim three times before looking at the PDF, and reported a visual page inspection that never happened. The same morning I found that a document it had "delivered" a week earlier had never uploaded.&lt;/p&gt;

&lt;h2&gt;
  
  
  The changelog
&lt;/h2&gt;

&lt;p&gt;The quietest one is the one I find most instructive. One agent maintains a small research site with a changelog, and through the summer the changelog announced a new feature almost every day: a mobile enhancement system, connection-aware layout, analytics, keyboard shortcuts. The files exist on the agent's machine, some at exactly the byte sizes the changelog claims. None of them are on the live site. Two return 404. The rest never load, because the plugin only attaches them on pages that use a shortcode the live page doesn't. The mobile script couldn't run anyway. Its text is missing characters and fails to parse. The live server still serves the May versions of the CSS and JS. The changelog is a perfect record of work that was done and never shipped, written by the thing that did the work.&lt;/p&gt;

&lt;h2&gt;
  
  
  What replaced trust
&lt;/h2&gt;

&lt;p&gt;Every one of those incidents produced a mechanism. Not a rule. A mechanism.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The rendered-artifact law.&lt;/strong&gt; No printed or rendered deliverable leaves without every page converted to an image, looked at, and its rows counted. The delivery message must include the proof: page count, row count, checksum. If the agent cannot produce the proof, it cannot claim delivery.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A presubmit check&lt;/strong&gt; that blocks files violating a standing format rule before the send command runs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A verdict script at the end of the archive pipeline&lt;/strong&gt; that will not pass an album unless an independent database confirms the audio bit for bit, or a recorded human approval overrides it. Every album's fixity lives in three separate places, and any tool that touches the audio has to regenerate all three.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A four-step rule for changes to the platform itself&lt;/strong&gt;, after the coding assistant declared a config change successful on the strength of a lint and a hot reload: lint, restart through the CLI, confirm a clean start in the log, then test the changed thing through the gateway. A real test on that occasion showed the new models were in the catalog but not in the allowlist, so nothing could select them. The dry run had been "successful."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I wrote about the same lesson on the inference server, where the promotion gate invalidated a config I had promoted on a smoke test alone. That was me being confident. This is the agents being confident. The countermeasure is identical: the check reads the artifact directly, a script does the reading, and the rule has power over whoever wrote it.&lt;/p&gt;

&lt;p&gt;Every time I let a model grade its own homework, it gave itself an A. The only verification that counts is the one that lives outside the agent.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Next: &lt;a href="https://michaelbrewer.me/2026/09/11/prompts-are-requests-hooks-are-law/" rel="noopener noreferrer"&gt;Prompts Are Requests. Hooks Are Law.&lt;/a&gt; The enforcement layer, and the ways it failed too.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>testing</category>
      <category>automation</category>
    </item>
    <item>
      <title>Four Builds in Seven Days</title>
      <dc:creator>Michael Brewer</dc:creator>
      <pubDate>Fri, 11 Sep 2026 19:07:11 +0000</pubDate>
      <link>https://dev.to/devbrewery/four-builds-in-seven-days-218b</link>
      <guid>https://dev.to/devbrewery/four-builds-in-seven-days-218b</guid>
      <description>&lt;p&gt;&lt;em&gt;Part 1 of three posts behind the multi-agent case study. How the stack that runs part of my life came to exist, and why the first three versions had to die.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I installed my agent platform in March 2026. That date matters, because for the seven months before it I had been doing something that looked like progress and mostly wasn't.&lt;/p&gt;

&lt;p&gt;From late 2025 on, I followed the agentic coding wave the way a lot of us did. Every new hyped-up harness got a weekend. When none of them held up, I started writing my own orchestrator. I have the repos to prove it, and I have the hours. What all of that time proved was one thing, and I want to state it plainly because nobody selling these tools will: it is entirely possible to spend ten times the hours building guardrails for a model as the hours the model saves you on the project.&lt;/p&gt;

&lt;p&gt;I installed a real platform anyway. The need was real. I am good at starting things and bad at maintenance and follow-through, and I have a job, a family, two acres, a homelab, and a board seat that all generate small obligations faster than I close them. I wanted to know whether a platform built for the job would change that ten-to-one ratio.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build one
&lt;/h2&gt;

&lt;p&gt;The first build ran under a locked-down service account, behind a retry proxy, with a cheap cloud model. Three days later I told the coding assistant helping me that the gateway kept restarting and the agent couldn't get anything done. That morning I framed the choice myself: keep the handicapped agent and the extreme isolation, or wipe the account and rebuild it with broad permissions. I wiped it. The shell history shows the scripts that did it, in order: back up, clean up, delete the user, recreate it, reinstall, restore.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build two
&lt;/h2&gt;

&lt;p&gt;The restored scaffold went onto the recreated account. The install did not go the way the coding assistant said it had. The gateway rebooted constantly while the assistant described it as working, and that was the first time in this project I had to say out loud that a confident report and a working system are two different things. It would not be the last.&lt;/p&gt;

&lt;p&gt;Once it ran, a review found the real problem. The agent "verbally claims it will use native tools but fails to actually execute them." It talked about work instead of doing it. Before I tore it down, I wrote a document for whatever came next. It is still the best summary of what I wanted:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Don't make Michael drag competence out of you one inch at a time. Use the system like it's real: turn promises into jobs, turn reminders into cron, turn long work into workers, turn lessons into files, turn uncertainty into blockers, and turn "I'll do it" into proof before the sentence even leaves your mouth.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Build three
&lt;/h2&gt;

&lt;p&gt;The third build was a refit: a migration runbook, new behavior contracts, a task ledger so work could not be lost, and subagents for heavy lifting. A review the next day concluded it was "improved and partially verified, but not yet trustworthy for sustained long-running autonomous work." That night the agent found a way around its new task ledger on its very first turn. The morning after, it still had not finished the task the whole refit was built to finish. On day seven I wiped the workspace and every cron and service dependency with it.&lt;/p&gt;

&lt;p&gt;Three builds, three failure reports, and they all said the same two things. It talked instead of acting, and it didn't finish. Each rebuild had changed the isolation, the model, or the prompts. None of that touched the actual problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build four
&lt;/h2&gt;

&lt;p&gt;The fourth build changed one thing that the first three had not. Enforcement shipped on day one. Before any agent bootstrapped, a plugin was already installed that detects an action request, detects a chat-only reply, and blocks it. The comment at the top of that plugin states the whole thesis of the project: the platform can restrict what agents can't do, but it cannot force what they must do. So I built the forcing.&lt;/p&gt;

&lt;p&gt;Everything running today descends from that fourth build. The six months since are their own story, and most of it is the enforcement layer growing around the agents one incident at a time. But the difference between a week of wipes and a system that finishes things was never a smarter model. It was moving the rule out of the prompt and into a mechanism, before the first message was ever sent.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the first week settled
&lt;/h2&gt;

&lt;p&gt;The community research my first manager agent did in its first hour said: start boring, one agent, add specialists late, prefer cron over heartbeat. The first half of that advice fell apart on contact. Telling a single agent to use cron did exactly nothing. And a single agent doing everything gets buried in the harness: every turn's context is muddied by instructions from domains it is not working in, and a single turn could run to six figures of input tokens before the agent read a word of the actual request. That is not a model problem. It is a prompt that cannot fit in anyone's head. So the specialists went up on the fourth build's first day, each with a small prompt in its own domain, and that decision is the one that held. The orchestrator that was supposed to route between them lasted two weeks before a lookup table replaced it. A research job that ran every thirty minutes built a body of research-backed material for its domain and did its job well, until the knowledge base grew large enough to put memory pressure on the whole system, so it is off until it can live in its own retrieval setup. Neither of those was an argument for fewer agents. They were arguments for smaller prompts and fewer moving parts per agent.&lt;/p&gt;

&lt;p&gt;The method that came out of that week is the one I still run. A coding agent works in the background. It reads the transcripts for what went wrong, builds the cron job or the hook that makes that failure impossible or at least loud, and then updates the agent's memory to reflect the change. The agents do the work in their lanes. The coding agent does the plumbing. I read the receipts.&lt;/p&gt;

&lt;p&gt;The ten-to-one ratio from before the install did not vanish when the platform arrived. It got paid down, one mechanism at a time, until the machinery started saving more hours than it cost. The full accounting, with the numbers, is in the case study.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Next: &lt;a href="https://michaelbrewer.me/2026/09/11/done-is-a-claim/" rel="noopener noreferrer"&gt;Done Is a Claim&lt;/a&gt;. Why the only verification that counts is the one that lives outside the agent.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>devops</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Case study: a household of agents</title>
      <dc:creator>Michael Brewer</dc:creator>
      <pubDate>Fri, 11 Sep 2026 19:01:53 +0000</pubDate>
      <link>https://dev.to/devbrewery/case-study-a-household-of-agents-32jf</link>
      <guid>https://dev.to/devbrewery/case-study-a-household-of-agents-32jf</guid>
      <description>&lt;p&gt;&lt;em&gt;Six months of records from a multi-agent stack that runs part of my life, and what the record, not my memory, says worked.&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;p&gt;I run a set of AI agents on a small dedicated machine in my house. Each one owns a domain: a general assistant, the homelab, and a handful of hobbies and volunteer commitments. They talk to me over chat, run scheduled jobs, read email, and ship deliverables I can open myself. I installed the platform in March 2026, after seven months of chasing the agentic coding wave: every hyped-up harness that came along, plus an orchestrator of my own. The stack running today was the fourth build in its first week, and this is what its own records say about the six months since.&lt;/p&gt;

&lt;p&gt;The short version is uncomfortable. The value never came from how smart the models were. It came from how fast I could turn a lesson into machinery: a script, a state file, a cron job, a hook, a checksum. Every durable win on this system is deterministic plumbing with a language model sitting at one narrow seam. Every recurring failure came from asking a model to do something it cannot be trusted to do on its own word: hold state, remember what it already knew, verify its own work, stay in its lane, or stop when told.&lt;/p&gt;

&lt;p&gt;After six months, the most important thing I have built is not an agent. It is the enforcement and verification layer around the agents, and the habit of never believing "done" until I can see the artifact.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem
&lt;/h2&gt;

&lt;p&gt;I built this because of a specific weakness that I asked the agents to cover. I am good at starting things and weak at maintenance, follow-through, and impulse decisions. A full-time job, a family, a two-acre property, a homelab, and a board seat produce a constant stream of small obligations, and the twenty minutes each one needs never lines up with the twenty minutes I have. The agents were supposed to own the follow-through.&lt;/p&gt;

&lt;p&gt;The seven months before the install had already taught me one thing. It is entirely possible to spend ten times the hours building guardrails for a model as the hours it saves you on the project. I installed it anyway, because the need was real and I wanted to find out whether a platform built for the job would change that ratio.&lt;/p&gt;

&lt;p&gt;The first three attempts, all in one week in March, failed the same two ways. They talked about work instead of doing it, and they did not finish. Each rebuild changed the isolation, the model, or the prompts. None of that fixed it. The fourth attempt is the one that stuck, and the difference was not a better model. It was that enforcement shipped on day one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Constraints
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cheap models on purpose.&lt;/strong&gt; Early research put frontier models at roughly $47 a week for this workload against about $6 for a budget cloud model. The whole point was always-on, so the budget model won, along with every weakness that came with it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every agent has a shell.&lt;/strong&gt; That is what makes them useful, and it means the worst days were never "the agent said something wrong." They were "the agent did something irreversible."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One on-call engineer.&lt;/strong&gt; A personal agent stack is a production system with a staff of one. The platform under it has bugs that look exactly like agent misbehavior, and I am the only one who will ever notice either.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The agents live on a firewalled subnet, on purpose.&lt;/strong&gt; The agent machine sits on its own isolated network behind carrier-grade NAT. It cannot reach the main LAN or the public internet except through the paths I gave it, and that was the right call before the first agent ran. What the record shows is that the agents kept trying to configure around it, planning for public traffic that could never arrive and breaking things in the attempt, until that boundary became the first section of their memory instead of an obstacle to route around. The friction was theirs to absorb, not mine to remove.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Options I rejected
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;An orchestrator that routes with a language model.&lt;/strong&gt; The architecture I settled on for the fourth build, and the one the community recommended, had a manager agent read every message and delegate to specialists. It lasted two weeks. It editorialized instead of dispatching, its delegation calls failed silently when the model left out the target, and the runtime it delegated through had a lifecycle bug where results never came back. Routing is now a lookup table: one chat topic per agent. No model sits in the routing path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rules in capital letters.&lt;/strong&gt; Every workspace is full of bold rules written the day something went wrong. One rule was violated more than 30 times after it was written. Another set of four "fireable offense" rules was broken, all four, at least once afterward. Around the second week I stopped believing that a rule in a prompt would be followed because it was emphatic.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Letting the research agent's knowledge base live inside the agent.&lt;/strong&gt; A research agent ran every 30 minutes and built a body of research-backed material that now covers just about any question in its domain: hundreds of runs at a 96 percent success rate and a vector memory assembled with zero human involvement, on tokens cheap enough to make that possible. The agent was doing exactly what it was supposed to, and the value was real. I turned it off because the knowledge base grew large enough to put memory pressure on the whole system and to bloat the agent's own context on every turn, which broke the narrow-domain ethos the rest of the household runs on. The right shape is a spin-off knowledge base with its own dedicated multimodal RAG setup that the agent sends requests to instead of carrying the material itself. Whether that gets built, time will tell.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Long fallback chains.&lt;/strong&gt; A seven-model fallback chain hid a three-month outage. The local model provider had never served a single successful request, because a base URL was missing a path segment, and the chain quietly skipped past it every time. A failover chain is a way to hide outages from yourself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The model as formatter.&lt;/strong&gt; A quiz script produced a question as JSON and the agent formatted it, badly, twice. The fix was not a sterner prompt. The command was deleted and the script now sends straight to chat with hardcoded formatting. The agent has zero formatting vector.&lt;/p&gt;

&lt;h2&gt;
  
  
  What shipped
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Verification that lives outside the agent.&lt;/strong&gt; An agent's report about its own work is a second artifact that may or may not match the first. A printed checklist once showed 58 of 95 items, was called complete, and was defended three times before anyone looked at the PDF. Now no rendered deliverable leaves without every page converted to an image and the rows counted, and the delivery message carries the proof. A presubmit script blocks files that violate a standing format rule. A media archiving pipeline ends in a verdict script that will not pass an item unless an independent database confirms the data or a recorded human approval overrides it, with three separate fixity records per item.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hooks instead of prompts.&lt;/strong&gt; The comment at the top of the first enforcement plugin states the whole problem: the platform can restrict what agents can't do, but cannot force what they must do. Fourteen enforcement plugins now cover that gap. One detects action requests and blocks chat-only replies. Several force a memory search before any action and a memory write before completion. One blocks replies that tell me to clear my cache or try again. One blocks mutations on external surfaces until declared sources have been read. One blocks bulk rewrites and non-English script in memory. One blocks the agents from upgrading the platform under themselves. Email goes through a relay that rejects any non-allowlisted recipient before a connection opens, with the allowlist hardcoded so an agent that controls the environment cannot widen it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Memory with a schema and a size cap.&lt;/strong&gt; Memory was the most dangerous component in the system. An overnight planning job produced over 200 files and 43,000 lines in four days, mostly regenerated copies of the same plan, with confident wrong claims buried inside that memory search kept resurfacing. The fix was deleting it and rebuilding by hand. Memory is now one file per topic, a check script hard-fails any index over 16 KB or any duplicate heading, reflections go to dated files, and the best-behaved agent's entire memory sits near 5 KB with each task written as a command plus a "done when."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The model at the narrowest seam.&lt;/strong&gt; The things I would miss tomorrow are almost all shell scripts with a language model doing one small job, or no job at all. A daily reminder is a cron job that sends buttons and writes a state file. Follow-ups fire every 15 minutes until acknowledged. When I tap a button, a plugin updates the state and replies. No agent turn runs and no tokens are spent per tap. A weekly config sync diffs a proxy against the config, asks the model to review the diff, snapshots, applies, and notifies. On its first run the review rejected a halved token limit. The model earns its place as a reviewer of a diff, not the author of the change.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blast radius capped mechanically.&lt;/strong&gt; After a "clean up the non-English text" request became a character filter that destroyed 1,089 memory files in one pass, and the same misreading produced 106 global edit commands three days later, the rule became structure: cleanup means rewrite, stop and ask above five files, and a daily sweep that reports stray text and never edits anything. Both incidents were survivable because a backup existed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The same discipline for the builder.&lt;/strong&gt; The coding assistant I used to build the stack failed in exactly the ways the agents did. It answered before reading, claimed verification that never happened, took conversational agreement as permission to edit 35 live cron jobs, and bypassed its own guardrail within minutes of writing it. One measured comparison: the rushed fix for a gateway restart problem took 106 tool calls and destroyed the install. The careful fix, three lines in a shell profile, took 60. It now runs under nine hooks of its own, including a snapshot before any config change and a rule that nothing is done until the change is tested through the gateway.&lt;/p&gt;

&lt;h2&gt;
  
  
  Outcome
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Measure&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Scheduled jobs defined and enabled&lt;/td&gt;
&lt;td&gt;77 and 69&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scheduled runs recorded, March 24 to September 10&lt;/td&gt;
&lt;td&gt;14,563&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Runs that did not finish clean&lt;/td&gt;
&lt;td&gt;1,974, or 13.6 percent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model turns in one measured week&lt;/td&gt;
&lt;td&gt;6,308, about 900 a day&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Input tokens in that week&lt;/td&gt;
&lt;td&gt;33.7 million&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Turns in that week that hit the prompt cache&lt;/td&gt;
&lt;td&gt;35, or 0.55 percent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Enforcement plugins, March to September&lt;/td&gt;
&lt;td&gt;5 grew to 24 enabled entries, 14 behavioral&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Items in a verified media archive, closed September 7&lt;/td&gt;
&lt;td&gt;95&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;What I would miss tomorrow: the reminders, the verified archive, travel support that kept my records straight through a canceled flight, working from the changes I told it about and the itinerary emails I forwarded, and rebuilt the documents, an eight-week volunteer curriculum published on schedule, a daily exam-prep quiz, email triage across five inboxes, and a 7 AM homelab check that reads the gateway log and names real problems.&lt;/p&gt;

&lt;p&gt;What never shipped: a finance agent that was created and never bootstrapped, and an A/B switch built in July to test whether lean prompting on a stronger model could replace the enforcement stack. It was validated and never flipped, and not from neglect: the enforcement stack and the proxy refinements started paying dividends in quality and took the pressure off the simplify-and-hope option. That option is what every frontier model quietly wants, to define for itself what counts as done, so of course a model suggested it.&lt;/p&gt;

&lt;p&gt;The cache number is the one that stings. Routing through the quality-gate proxy rewrote payloads and defeated the vendor's automatic caching, so one measurement showed 276 turns and zero cache hits, roughly 600,000 input tokens paid twice. The enforcement injections themselves cost context on every turn. Coercion is not free.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would do differently
&lt;/h2&gt;

&lt;p&gt;I would not spend my own hours writing documentation for the agents. Every rule I put in prose was a bet that the model would read it, weigh it, and comply, and the record above shows how that bet paid. I would go straight to more coercion and more constraints: hooks that block, scripts that check the artifact, allowlists hardcoded where an agent cannot reach them, and a size cap on memory from the first day. A constraint costs an afternoon once. A hundred failed attempts cost an afternoon each, plus the cleanup. The constrained version is the cheaper product, and it is the only version that ever finished anything.&lt;/p&gt;

&lt;p&gt;A personal agent system spends the operator's attention to save the operator's attention. It is only a win when the second number is bigger. For the first few months, it was not. What moved the balance was never a better model. It was moving each recurring correction out of my head and into a mechanism, so I stopped making the same correction twice. The ten-to-one ratio from before the install did not go away. It got paid down, one mechanism at a time, until the machinery started saving more hours than it cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The inference server the agents run against: &lt;a href="https://github.com/dev-brewery/inference-fleet" rel="noopener noreferrer"&gt;https://github.com/dev-brewery/inference-fleet&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;The routing proxy in front of it: &lt;a href="https://github.com/dev-brewery/smart-proxy" rel="noopener noreferrer"&gt;https://github.com/dev-brewery/smart-proxy&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Related posts: The Year I Started Finishing Things, and The Genealogy Book Nobody Had Time to Read, at &lt;a href="https://michaelbrewer.me/" rel="noopener noreferrer"&gt;https://michaelbrewer.me/&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>automation</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Case study: smart-proxy</title>
      <dc:creator>Michael Brewer</dc:creator>
      <pubDate>Fri, 11 Sep 2026 19:01:46 +0000</pubDate>
      <link>https://dev.to/devbrewery/case-study-smart-proxy-3b34</link>
      <guid>https://dev.to/devbrewery/case-study-smart-proxy-3b34</guid>
      <description>&lt;p&gt;&lt;em&gt;One endpoint in front of three tiers of inference, so that always-on agents get a frontier-class answer only when they need one.&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;p&gt;I run a fleet of AI agents that generate hundreds of requests an hour, around the clock. Most of those requests are routine: heartbeats, status checks, processing a tool result, classifying a message. Sending all of it to a frontier API was expensive and fragile, and rate limits and credit exhaustion took the agents down at the worst times.&lt;/p&gt;

&lt;p&gt;smart-proxy is the routing layer I built to fix that. It is a standard-library Python proxy that presents one OpenAI-compatible endpoint and decides, per request, whether the work goes to a small CPU model, a large local GPU model, or a cheap cloud model. It also polices the cloud tier with a quality gate, because cheap cloud models are capable but unreliable in specific, predictable ways.&lt;/p&gt;

&lt;p&gt;Over one measured 40-hour production window it handled 8,129 requests at a 0.17 percent failure rate with no manual intervention. The quality gate passed 95.7 percent of cloud responses on the first attempt, and action requests that used to take 36 to 45 seconds now take 15 to 20.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem
&lt;/h2&gt;

&lt;p&gt;Three things were true at once:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The agents never stop.&lt;/strong&gt; An always-on multi-agent system produces a constant stream of small, cheap-to-answer requests, with an occasional hard one mixed in. Paying frontier prices for the small ones made no sense, and depending on a single external API meant a rate limit anywhere took everything down.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Local hardware is capable but constrained.&lt;/strong&gt; The inference server is about $2,000 of used enterprise parts: an EPYC 7302P, 128 GB of ECC memory, and two Tesla P40s from 2016. It can run a 27B dense model or an 80B mixture-of-experts model well, but only one large model at a time, with a 20 to 40 second cold start and about 30 seconds to swap.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cheap cloud models misbehave in specific ways.&lt;/strong&gt; They narrate an action instead of calling the tool. They invent tool requirements for plain factual questions. They truncate under load. Each of those looks like a successful HTTP 200 to a naive client.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The requirement was one endpoint, all models visible, and the system making the tiering decision itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Constraints
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Pascal-era GPUs.&lt;/strong&gt; No Tensor Cores, no vLLM, and most published tuning advice is written for newer cards. The GPU tier had to live on llama.cpp.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One GPU model resident at a time.&lt;/strong&gt; Any design that assumed concurrent large models was dead on arrival.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The CPU had to be the floor.&lt;/strong&gt; Whatever ran on CPU had to survive GPU swaps, GPU crashes, and GPU power-off, so the system always had something available.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Production traffic, not a lab.&lt;/strong&gt; The agents were live users of the proxy throughout. Every change had to be one variable, tested, then frozen with checksums so it could be rolled back in about two minutes.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Options I rejected
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A LiteLLM gateway.&lt;/strong&gt; This is where I started. It handled alias routing and a remote coder target, but the moment I needed a quality gate, per-pool concurrency accounting, and structured request logging, it was easier to own a small proxy than to bend a large one. The lesson I kept was not "avoid proxy layers." It was "don't add layers; let the one layer you have absorb policy."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A GPU model as the judge.&lt;/strong&gt; Judging whether a cloud response is empty, truncated, or narrating instead of acting is a narrow task. Spending a GPU slot on it would contend with the actual work. A 4B model on CPU does it for free.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Static concurrency caps from synthetic tests.&lt;/strong&gt; My first caps came from small test requests and did not hold. The vendor's limiter varies by load window, and the payload shape matters. The caps that survived came from stress tests with real gateway-shaped requests, median around 32K tokens, recording admission versus 429 and setting each cap at the worst observed value.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A fallback to "whatever local model is loaded."&lt;/strong&gt; That serves the wrong model silently. The fallback is now pinned to a named stack and skips itself, with a log line, unless that exact stack is active. A fallback that might serve a different model than intended is worse than no fallback.&lt;/p&gt;

&lt;h2&gt;
  
  
  What shipped
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Three tiers behind one alias map.&lt;/strong&gt; Two Qwen3-4B models on dedicated CPU cores form the always-on floor: a classifier that triages each request into simple, retrieval, code, or reasoning in 200 to 600 milliseconds, and a helper for summaries, retrieval, tool matching, and fallback judging. The GPU tier is twelve llama.cpp stacks, swapped through the Portainer API with a ten-minute anti-flap cooldown and rollback on a failed swap. The cloud tier is two real vendor pools with measured caps of 6 and 5.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A quality gate with a rubric.&lt;/strong&gt; Cloud responses pass through deterministic validation of tool-call structure, then an LLM judge, then a retry with structured feedback if either fails. The most important finding was that a judge without explicit criteria hallucinates failures. My first judge prompt asked whether the response "requires tool use," and the model decided every response did. The fix was a three-part rule, the list of tools read from the actual request, and explicit exclusions for knowledge questions, opinions, and greetings.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pre-classification instead of retries.&lt;/strong&gt; The retry loop existed because cloud models narrate. Three attempts at about twelve seconds each is where the 36 to 45 second latency came from. Now the CPU helper reads the request against the tool names, spends two to three seconds picking the applicable tool, and the proxy injects a structured nudge naming it. The cloud model produces the tool call on the first attempt. Names-only tool lists at about 40 tokens outperformed described lists at over 200, and temperature zero was mandatory for anything classification-shaped.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Concurrency accounted on the serving pool, not the requested id.&lt;/strong&gt; One model id was stress-verified as its own pool, then the next day three out of three production requests for it came back answered by a different pool. The vendor had merged them without announcing it. Had I kept the two caps separate, real concurrency on that pool would have been 13 against a measured limit of 5. The proxy now carries a routing map in code that collapses legacy ids onto the pool that serves them, whatever id the client asked for, emits a reroute event and a response header on every reroute, and warns at load if the config documentation disagrees with the code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Structured request logging.&lt;/strong&gt; Every request emits start, end, routing, and reroute events keyed by request id. That is how the pool merge was confirmed in live traffic within minutes of deploy: 63 reroutes in a 30-minute window.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The two bugs that made fallbacks work.&lt;/strong&gt; A transport error path produced an HTTP status of 0, which Python's server happily emitted as an invalid status line and broke the client connection; the fix clamps out-of-range statuses to 502 at the send boundary. And every local fallback was failing with a 400 because Anthropic-format tool-call messages carry null content, which llama.cpp rejects; the translator now emits an empty string, and fallback errors return 502 with the real backend detail attached. Both were invisible until the error propagation was fixed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Outcome
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Measure&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Requests in one 40-hour production window&lt;/td&gt;
&lt;td&gt;8,129&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Failure rate in that window&lt;/td&gt;
&lt;td&gt;0.17 percent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Manual interventions in that window&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quality gate first-attempt pass rate&lt;/td&gt;
&lt;td&gt;95.7 percent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Action-request latency, before&lt;/td&gt;
&lt;td&gt;36 to 45 seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Action-request latency, after&lt;/td&gt;
&lt;td&gt;15 to 20 seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Classifier triage latency&lt;/td&gt;
&lt;td&gt;200 to 600 milliseconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cloud cost on routine traffic&lt;/td&gt;
&lt;td&gt;none; handled on CPU&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The thesis the numbers support: most agent traffic does not need a frontier model, and small models are most valuable when they steer large ones rather than replace them. A 4B model that can say "this needs the cron tool" in two seconds saves thirty seconds of a cloud model fumbling toward the same conclusion.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would do differently
&lt;/h2&gt;

&lt;p&gt;Log structured request events from day one. Every hard debugging story in this project got easy the moment the reroute and routing events existed, and I built them last. And treat every rate limit as a defense rather than a promise: the number is "worst I have observed," re-verifiable any time, never a fact about the vendor.&lt;/p&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Repo: &lt;a href="https://github.com/dev-brewery/smart-proxy" rel="noopener noreferrer"&gt;https://github.com/dev-brewery/smart-proxy&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;The fleet it fronts: &lt;a href="https://github.com/dev-brewery/inference-fleet" rel="noopener noreferrer"&gt;https://github.com/dev-brewery/inference-fleet&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Blog series: The $2,000 Inference Server at &lt;a href="https://michaelbrewer.me/" rel="noopener noreferrer"&gt;https://michaelbrewer.me/&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>architecture</category>
      <category>python</category>
    </item>
    <item>
      <title>Shock and Awe Is a Business Model</title>
      <dc:creator>Michael Brewer</dc:creator>
      <pubDate>Tue, 08 Sep 2026 19:30:30 +0000</pubDate>
      <link>https://dev.to/devbrewery/shock-and-awe-is-a-business-model-4o11</link>
      <guid>https://dev.to/devbrewery/shock-and-awe-is-a-business-model-4o11</guid>
      <description>&lt;p&gt;Almost all of my friction with frontier models, and I mean all of them, traces back to one root cause. It isn't capability. It's tuning.&lt;/p&gt;

&lt;p&gt;Every frontier model is tuned for shock and awe. The demo has to dazzle. The first answer has to feel brilliant. The model volunteers essays when you wanted a sentence, generates confident sweeping output when you wanted a careful narrow one, and optimizes for the impression it makes in the first thirty seconds over the quality of the working relationship in month six.&lt;/p&gt;

&lt;p&gt;This is not an accident and it is not a flaw in the training pipeline. It is the business model. The adoption curve has to keep climbing so the easy capital keeps flowing. A model tuned for restraint, one that asks a clarifying question, delivers the minimum correct answer, and stops talking, would be better to work with and worse in a demo. The demo wins, because the demo is what raises the next round.&lt;/p&gt;

&lt;p&gt;I say this as a daily, heavy, mostly satisfied user of these models. They are remarkable. But remarkable is the product they're selling, and it's not the product an organization needs in order to operate.&lt;/p&gt;

&lt;h2&gt;
  
  
  What organizations need when the novelty wears off
&lt;/h2&gt;

&lt;p&gt;An org running AI in production needs the opposite of shock and awe. It needs method. Predictable scope. An answer that stops when the question is answered. A model that follows the runbook instead of improvising a more impressive one. Delivery over volume, consistency over brilliance.&lt;/p&gt;

&lt;p&gt;I learned this the way I learn everything, by running the systems myself. My agent stack went through five teardowns and rebuilds. The failures were never because a model was too weak. They were because models tuned to impress kept blowing out their context doing more than the task required, and because I kept trying to make cheap general models do what only a more capable or more specialized one could. The fix, every time, was narrowing: smaller scopes, tighter instructions, specialist agents, and enforcement plugins that mechanically punish showing off. I run a quality gate on my own infrastructure for exactly this reason. My agents' output is graded against the task, not against how impressive it sounds.&lt;/p&gt;

&lt;p&gt;That's a homelab-scale version of what every serious AI-adopting org is going to end up building. Not because they want to, but because the vendors' incentives and theirs point in different directions.&lt;/p&gt;

&lt;h2&gt;
  
  
  The literacy that stops being optional
&lt;/h2&gt;

&lt;p&gt;Here's the uncomfortable consequence. If the models you rent are tuned for someone else's goals, then getting models tuned for yours means understanding how tuning works. Weights, fine-tuning, quantization tradeoffs, evaluation methodology, what a training objective rewards in practice. Not at research depth. At operator depth: enough to read a model card critically, enough to know what a fine-tune can and cannot fix, enough to measure whether the behavior you bought is the behavior you got.&lt;/p&gt;

&lt;p&gt;For most of the software era, this kind of knowledge was optional the way compiler internals are optional. You could build a career on top of the abstraction. I don't believe model behavior gets to stay abstracted, because the abstraction is leaking money and risk in a way compilers never did. A model that over-delivers by 3x on every request is a cost center. A model that improvises outside its scope is a liability with a legal department's name on it. Cost-effective AI that isn't a liability to the org requires someone in the building who understands what the weights were trained to do, and the orgs that treat that as a vendor's problem will pay vendor prices for vendor-aligned behavior, forever.&lt;/p&gt;

&lt;p&gt;Open weights are what make the alternative possible. When the weights are yours, tuning for method over spectacle is an engineering decision instead of a feature request to a company whose incentives run the other way. That, more than cost, is why I run open models on my own hardware and why I think the maid-to-order open-weights era is coming regardless of how loudly the incumbents warn against it. The gatekeepers in the pioneer phase always cry. It's what the phase sounds like.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the two theses meet
&lt;/h2&gt;

&lt;p&gt;Part one said: compose many specialized models the way you compose a staff. This part says: expect to tune some of them yourself, and staff for the literacy that requires.&lt;/p&gt;

&lt;p&gt;Put together, that's the whole implementation philosophy. If I were handed a real budget for a long-term AI implementation tomorrow, this is how I'd spend it. Not on the biggest model money can rent, but on a bench of specialized ones in the right seats, a routing and evaluation layer that keeps each in its lane, the measurement discipline to prove it's working, and enough weights-and-training literacy in-house that the org's AI serves the org's goals instead of its vendors'.&lt;/p&gt;

&lt;p&gt;None of that requires a frontier lab's budget. I know, because I run the small version of it on hardware nobody else wanted.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opinion</category>
    </item>
    <item>
      <title>The Genealogy Book Nobody Had Time to Read</title>
      <dc:creator>Michael Brewer</dc:creator>
      <pubDate>Tue, 08 Sep 2026 19:25:13 +0000</pubDate>
      <link>https://dev.to/devbrewery/the-genealogy-book-nobody-had-time-to-read-2m9d</link>
      <guid>https://dev.to/devbrewery/the-genealogy-book-nobody-had-time-to-read-2m9d</guid>
      <description>&lt;p&gt;&lt;em&gt;What a personal agent stack is for, once the demos are over.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A relative of mine was sent a book about one of our ancestors. A genealogy, hundreds of pages, compiled by someone who had clearly spent years on it. The kind of thing a family is lucky to have and almost guaranteed never to read. It needed to be gone through, cross-referenced, and connected to what we already knew about the family line. My relative did not have time for that. Nobody in the family did. It sat there the way these things sit, valuable and untouched.&lt;/p&gt;

&lt;p&gt;This is the part of the story where AI usually gets oversold, so let me be precise about what happened and what made it possible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Half an hour to digitize
&lt;/h2&gt;

&lt;p&gt;I've spent the past year building a personal agent stack: an assistant that runs on my own infrastructure, that I've customized tool by tool as real needs came up. Because that plumbing already existed, digitizing the book was not a project. It was half an hour of feeding pages through, with the agent handling OCR, cleanup, and structure as we went. Hundreds of pages became searchable, structured text before lunch.&lt;/p&gt;

&lt;p&gt;The half hour is the headline number, but it's the least interesting part. Digitization is a solved problem. What came next is why I'm writing this.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mapping the connections
&lt;/h2&gt;

&lt;p&gt;I exported our existing family tree data from Ancestry.com and handed it to the agent alongside the digitized book. Then I set it loose on the tedious part: going through the book person by person and mapping every connection onto the tree we already had. Which people in the book matched people we knew about. Which were new. Where the book confirmed our data, where it contradicted it, and where it filled holes.&lt;/p&gt;

&lt;p&gt;This is exactly the kind of work that defeats a human volunteer. Not because it's hard, but because it's hundreds of small, careful, boring judgments in a row. It's also exactly what a well-instructed agent is good at, provided you check its work. I spot-checked as it went, corrected its course a few times, and let it grind.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five hours of driving, one website
&lt;/h2&gt;

&lt;p&gt;The visit ended and I had a five hour drive home. So the drive became the build window. Voice messages from the road guided the agent through standing up an open-source genealogy website on top of the mapped data, and then through starting something the open-source tool didn't have: a new visualization for exploring the tree.&lt;/p&gt;

&lt;p&gt;I want to be honest about what "guided by voice from the car" means. It does not mean I dictated flawless instructions and arrived home to finished software. It means the agent worked, hit decisions it shouldn't make alone, and I made them at highway speed in short messages. Some of what I found when I got home was wrong and got redone. But the shape of both the site and the visualization tool existed by the time I pulled into the driveway, built during hours that would otherwise have produced nothing but mileage.&lt;/p&gt;

&lt;h2&gt;
  
  
  A couple of days of refinement
&lt;/h2&gt;

&lt;p&gt;Once home, I spent a couple of days refining with the agent: fixing the wrong turns, polishing the visualization, getting the data presentation right. At the end of it, our family history, the book's contents and the mapped tree together, was live where anyone in the family could see it.&lt;/p&gt;

&lt;p&gt;That last part matters more to me than the technology. This information used to live in two places: a paywalled account one person maintained, and a physical book one person possessed. Now it belongs to the family. A grandparent can look at it. A cousin I've never met can look at it. Nobody needs a subscription or a login or my help.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this is a story about, once you get past the digitizing
&lt;/h2&gt;

&lt;p&gt;Not about AI reading a book fast. It's about accumulated tooling meeting a real obligation. Every customization in my stack existed before this project, built for other reasons over a year of daily use. When the book arrived, the marginal cost of taking on a job nobody had time for was half an hour, a car ride, and a weekend's worth of refinement sessions.&lt;/p&gt;

&lt;p&gt;That's my working definition of what personal AI tooling is for. Not demos. Not novelty. Taking things a family cares about but cannot resource, and resourcing them.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>productivity</category>
      <category>automation</category>
    </item>
    <item>
      <title>The Right Models in the Right Seats</title>
      <dc:creator>Michael Brewer</dc:creator>
      <pubDate>Tue, 08 Sep 2026 19:25:06 +0000</pubDate>
      <link>https://dev.to/devbrewery/the-right-models-in-the-right-seats-47m2</link>
      <guid>https://dev.to/devbrewery/the-right-models-in-the-right-seats-47m2</guid>
      <description>&lt;p&gt;&lt;em&gt;The thesis behind everything I build: specialization beats scale.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Ask the frontier labs what the future of AI looks like and the answer is always the same: bigger. More parameters, more compute, one ever-more-capable model that does everything. That vision has a convenient property: only a handful of companies on earth can build it, and you'll be renting it from them forever.&lt;/p&gt;

&lt;p&gt;I run my AI workloads on a different thesis, and after a year of measurements on my own hardware, I believe it more, not less.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The way forward is not ever-larger all-knowing models. It's teams of highly effective, highly specialized smaller models, composed the way a good company composes a staff.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Get the right models on the bus
&lt;/h2&gt;

&lt;p&gt;Jim Collins framed the difference between good companies and great ones as getting the right people on the bus, and the right people in the right seats. Nobody staffs a company by hiring one impossibly expensive generalist to do every job. You hire people whose strengths match their seats, and the organization outperforms the sum of its parts.&lt;/p&gt;

&lt;p&gt;My inference stack runs exactly this way, and I have a year of production numbers behind it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A 4B-parameter model running on CPU cores classifies every incoming request in a few hundred milliseconds. It cannot write good code and never will. It doesn't need to. Its seat is triage, and it fills it for free, on hardware that would otherwise idle.&lt;/li&gt;
&lt;li&gt;The same class of tiny model reads a request against a list of available tools and tells a big cloud model which one to use. That two-second nudge cut my action-request latency in half, not because the small model is smart, but because it prevented an expensive model from fumbling toward a conclusion a cheap one had already reached.&lt;/li&gt;
&lt;li&gt;A sparse mixture-of-experts model handles the fast path at 41 tokens per second on decade-old GPUs, because activating 4B parameters out of 26B is itself specialization inside the weights.&lt;/li&gt;
&lt;li&gt;The big models, local or cloud, only see the work that earns their cost. In one measured 40-hour window, 8,129 requests flowed through this division of labor with a 0.17% failure rate and no human intervention.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The lesson from those numbers isn't that small models are secretly as good as big ones. They aren't, and pretending otherwise is how you build garbage. The lesson is the one from my notes that I keep coming back to: &lt;strong&gt;the tiers aren't a hierarchy of quality, they're a division of labor.&lt;/strong&gt; A 4B model in the right seat outperforms a frontier model in the wrong one, on cost, on latency, and often on reliability, because narrow tasks reward consistency over brilliance.&lt;/p&gt;

&lt;h2&gt;
  
  
  Companies already know this pattern
&lt;/h2&gt;

&lt;p&gt;Here's why I think this thesis wins on economics, not ideology. Every company already organizes its people this way. Nobody believes the optimal workforce is one superhuman consultant billing by the token. Companies win by being specialized and nimble in their personnel, and they will demand the same from their AI: a small model fine-tuned on their support history in the support seat, a compliance-tuned model in the compliance seat, a coding model that knows their codebase in the engineering seat, orchestrated by systems that route work to whoever fills the seat best.&lt;/p&gt;

&lt;p&gt;That world is arriving through open weights. The Qwens and Gemmas I run today are the early, general-purpose versions. The trajectory points at made-to-order weights: models distilled, tuned, and owned by the companies that run them, sized to their seats, running on their hardware or commodity clouds, with their data never leaving the building. Every quarter the open releases get better at fitting into seats that used to require a frontier API call, and my own benchmarks watched it happen: models that needed 39GB of VRAM last spring were outclassed by 23GB models with better architectures by fall.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gatekeepers always cry
&lt;/h2&gt;

&lt;p&gt;Which brings me to the part of the argument that the frontier labs make for me.&lt;/p&gt;

&lt;p&gt;Listen to how the largest AI companies talk about open weights: dangerous, irresponsible, impossible to control, surely the end of safety itself. Some of those concerns are sincere and worth engaging seriously. But notice the shape of the argument and who it benefits. The pioneers of every technology wave have warned that the tools were too dangerous to leave the temple, right up until the moment the tools left anyway: mainframe companies about personal computers, telecoms about the open internet, every incumbent about every commodity that ended their toll booth.&lt;/p&gt;

&lt;p&gt;Gatekeepers in the pioneer phase always cry. It's what the phase sounds like. The economics underneath don't care: when capability becomes a commodity you can own instead of rent, composition becomes the differentiator, and composition is an engineering discipline, not a capital moat.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means for the work
&lt;/h2&gt;

&lt;p&gt;If the thesis is right, the scarce skill of the next decade isn't prompting one giant model. It's the boring, measurable systems work of running teams of models well: routing, failover, quality gates, evals that catch a model drifting out of its seat, guardrails that hold when a model exceeds its authority, and the operational discipline to measure everything, because half of what you believe about your stack will be wrong within six months. I know because I measured mine, and it was.&lt;/p&gt;

&lt;p&gt;That's the bet my basement server, my benchmarks, and this blog are all placed on. One team of specialists, on hardware nobody wanted, doing the daily work of a system that would otherwise be an expensive subscription. The right models on the bus, the right models in the right seats, and the bus is yours.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>architecture</category>
      <category>opinion</category>
    </item>
    <item>
      <title>Measure the Binary You Run</title>
      <dc:creator>Michael Brewer</dc:creator>
      <pubDate>Tue, 08 Sep 2026 18:33:22 +0000</pubDate>
      <link>https://dev.to/devbrewery/measure-the-binary-you-run-4022</link>
      <guid>https://dev.to/devbrewery/measure-the-binary-you-run-4022</guid>
      <description>&lt;p&gt;At one point in this project, two documents in my own notes argued opposite positions about the same compile flag.&lt;/p&gt;

&lt;p&gt;Document one, the backend research: "The current build has AVX2 disabled. Priority 1: recompile with AVX2. Expected improvement, 15 to 30% on prompt processing."&lt;/p&gt;

&lt;p&gt;Document two, the ecosystem plan: "AVX2 is off by design, to preserve CPU headroom for the rest of the stack while the GPU serves inference."&lt;/p&gt;

&lt;p&gt;One says the flag is off and that's a problem. The other says it's off and that's a feature. They can't both be right.&lt;/p&gt;

&lt;p&gt;Neither was. The flag was on the whole time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where both documents went wrong
&lt;/h2&gt;

&lt;p&gt;The evidence behind "AVX2 is disabled" was the cmake build cache, which showed the SIMD options set to OFF. Build-directory archaeology: inspect the configuration, infer the artifact.&lt;/p&gt;

&lt;p&gt;But llama.cpp's build enables native CPU optimizations through its own path regardless of those cached toggles, and the shipped binary is perfectly willing to tell you what it contains. Its startup system info prints the actual capability set. Mine printed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AVX = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The running binary had AVX2 enabled all along. The "disabled" reading described cmake defaults that the real build had overridden. And the second document's clever theory about why disabled-AVX2 was good design was a rationalization of a condition that didn't exist.&lt;/p&gt;

&lt;p&gt;The regression incident from post 2 has a footnote here too: that four-variables-at-once rebuild included "enable AVX2" as one of its four changes. In its own words, in the post-mortem: redundant. One of the four simultaneous changes was a no-op, which muddied attribution further.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hierarchy this settled
&lt;/h2&gt;

&lt;p&gt;My notes rank evidence quality in six rungs, and this incident fixed the ordering of two of them permanently:&lt;/p&gt;

&lt;p&gt;Runtime artifact inspection beats build-system archaeology. Always. The cmake cache tells you what the build system was asked. The binary's own system info tells you what came out the other end. When they disagree, the binary wins, because the binary is what serves your traffic.&lt;/p&gt;

&lt;p&gt;The general form: verify the premise before optimizing it. "Recompile to enable AVX2" was a well-reasoned recommendation, correctly derived from its evidence, aimed at a switch that was already in the right position. Everything about the plan was sound except the fact it stood on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cheap checks, expensive assumptions
&lt;/h2&gt;

&lt;p&gt;What makes this sting is how cheap the correct check was. The system info line prints at every server start. It was in every log I'd ever launched. Nobody had read it, because everyone was reading the build directory instead, where the interesting-looking configuration lives.&lt;/p&gt;

&lt;p&gt;Since then, the rule in my fleet is that claims about a binary come from the binary: its startup banner, its version string, its measured behavior. Configuration files describe intent. Artifacts describe reality. Optimization work starts from reality.&lt;/p&gt;

&lt;p&gt;Two documents argued about a switch that was already in the right position. The moral isn't that documents are bad. It's that both documents cited the same wrong source, and one five-second look at the right source would have ended the argument before it started.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>benchmarking</category>
      <category>devops</category>
    </item>
    <item>
      <title>The Year I Started Finishing Things</title>
      <dc:creator>Michael Brewer</dc:creator>
      <pubDate>Tue, 08 Sep 2026 18:29:44 +0000</pubDate>
      <link>https://dev.to/devbrewery/the-year-i-started-finishing-things-3hkg</link>
      <guid>https://dev.to/devbrewery/the-year-i-started-finishing-things-3hkg</guid>
      <description>&lt;p&gt;&lt;em&gt;On using AI agents to run more of a life than one person's spare time should hold.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I have a full-time job, a family, a two-acre hilltop, a homelab, and a seat on the board of a small nonprofit. For most of my adult life, the limiting factor on what I could take on wasn't ability or interest. It was hours. Things I cared about sat unfinished for months because the twenty minutes they needed never lined up with the twenty minutes I had.&lt;/p&gt;

&lt;p&gt;This year that changed, and I want to write down how, honestly, without the hype that usually smothers this topic.&lt;/p&gt;

&lt;h2&gt;
  
  
  The nonprofit problem
&lt;/h2&gt;

&lt;p&gt;A couple of years ago I joined the board of a small all-volunteer nonprofit that supports pastor training in West Africa. I serve as secretary, and because I'm the technical one, everything with a login became mine: the website, the records, the state filings, the donor communications, the grant research nobody had time to do.&lt;/p&gt;

&lt;p&gt;Here's what that job looks like, day to day, at a tiny nonprofit. There is no staff. There is no budget for staff. Every task is small, none of them are optional, and they arrive continuously: a board email that needs filing, a compliance deadline nobody remembers until it's urgent, a partner organization posting updates that should reach our supporters, a funder whose eligibility rules need reading before anyone wastes a weekend on an application. Any one of these is twenty minutes. Together they're a part-time job that nobody has.&lt;/p&gt;

&lt;p&gt;The standard outcome is that volunteer boards run on heroics and guilt. Things slip. The person who cares most burns out first.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I built, and what a normal day with it looks like
&lt;/h2&gt;

&lt;p&gt;Over the past year I stood up an AI agent that functions as the org's administrative assistant. It runs on a small server, talks to me over chat, and drives our ordinary office tools: email, calendar, the shared drive, a spreadsheet that acts as the operations hub. Around two dozen scheduled jobs handle the recurring work.&lt;/p&gt;

&lt;p&gt;A normal day looks like this. At 9 AM I get a short message: here's what's outstanding, here are the top three things, here's the link if you want the whole picture. Board emails that arrived overnight have been triaged: routine ones handled and filed, anything needing my judgment flagged. Documents dropped in a shared folder have been renamed, filed, and logged. If our partner ministry posted an update, a draft is waiting on the website for approval. Once in a while there's a note that a funding opportunity surfaced overnight, with deadlines and fit notes attached.&lt;/p&gt;

&lt;p&gt;I read it in the time it takes to drink coffee, make the two or three decisions that genuinely need a human, and go to work.&lt;/p&gt;

&lt;p&gt;Some of what came out of this, concretely:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Compliance filings that had been hanging over the board since incorporation got done.&lt;/strong&gt; The agent assembled the filing packages; I reviewed and signed; the state approved them. That was the single biggest weight off the board's shoulders, and it was mostly review time on my end.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Grant prospecting went from nonexistent to nightly.&lt;/strong&gt; No volunteer was ever going to spend evenings scanning funder sites. An automated job does, and it reads eligibility rules against our actual documents before anything reaches the board. It caught, for example, that a foundation we liked requires organizations to be older than we are, by pulling our real incorporation date from our IRS paperwork. That's a wasted application avoided and an honest reminder scheduled for when we qualify.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Board records became records.&lt;/strong&gt; Every decision and email is logged and retrievable. When someone asks "what was that concept note from last year?", the answer takes seconds instead of inbox archaeology.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The website updates itself, mostly.&lt;/strong&gt; Partner updates become drafts I approve. Promo materials that would have been a multi-day design favor get generated from source material in the org's own style in an afternoon.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What I refuse to pretend
&lt;/h2&gt;

&lt;p&gt;This is the part most writing about AI leaves out, so let me be specific.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The failures are real and constant.&lt;/strong&gt; The token for the primary model expired once and the system quietly degraded to cheaper models for days before anyone noticed; the fix was another automated check, watching the watchers. A reminder got scheduled for the wrong year; I caught it. A briefing once referenced a project that had a tracker row but no documentation; the fix was auditing everything and making cross-referencing a standing rule. A funding deadline surfaced too late to act on. Two emails failed to send and were done by hand. The first architecture for the recurring jobs was wasteful and had to be redesigned. A monitoring job I thought was running had been off for months.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Nothing here works without a human reading it first.&lt;/strong&gt; Nothing goes to the board, a funder, or the public without my review. The system's job is to make my twenty minutes count, not to replace my judgment. Every one of the failures above was caught either by an automated check I added after a previous failure, or by me reading something before it went out. Both layers earn their keep monthly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The outcomes are modest and I'll state them plainly.&lt;/strong&gt; We have not won a grant. We haven't even submitted an application yet; the pipeline's honest output so far is one fully researched, board-ready opportunity and a list of funders we now know we don't qualify for, with dated reasons. What changed is that the org went from no systematic effort to a running process that costs almost nothing and never gets tired. For an all-volunteer organization, the difference between "nobody has time" and "it happens every night" is the whole game.&lt;/p&gt;

&lt;h2&gt;
  
  
  The general lesson
&lt;/h2&gt;

&lt;p&gt;The same year, the same pattern ran through everything else I do: the inference server in my basement that this blog documents, the agent tooling at my day job that turns client call recordings into project proposals, the monitoring that tells me about problems before I go looking. None of it is one big AI doing something impressive. All of it is small, boring automation with careful boundaries, checked by a human, compounding.&lt;/p&gt;

&lt;p&gt;Add it up, and what AI bought me this year is not intelligence. I still make every decision that matters. It bought me &lt;strong&gt;parallelism&lt;/strong&gt;. Things now make progress during hours I'm not present: overnight, during the workday, while I'm mowing the hill. My attention became the scarce resource the whole system is designed to spend well, twenty minutes at a time.&lt;/p&gt;

&lt;p&gt;A year ago, the honest description of my volunteer role was "important things slip, and I feel bad about it." Today it's "the routine runs itself, and I spend my time on judgment." Same hours. Same person. That's the difference, and for a small mission-driven organization that difference isn't a productivity statistic. It's whether the work of keeping the lights on leaves any energy for the mission itself.&lt;/p&gt;

</description>
      <category>career</category>
      <category>productivity</category>
      <category>ai</category>
    </item>
    <item>
      <title>Eggs, Cholesterol, and GPU Flags</title>
      <dc:creator>Michael Brewer</dc:creator>
      <pubDate>Tue, 08 Sep 2026 18:24:27 +0000</pubDate>
      <link>https://dev.to/devbrewery/eggs-cholesterol-and-gpu-flags-561j</link>
      <guid>https://dev.to/devbrewery/eggs-cholesterol-and-gpu-flags-561j</guid>
      <description>&lt;p&gt;For decades, nutrition science flip-flopped on eggs. Bad for you: dietary cholesterol. Then fine, then good in some contexts, then it depends on the person and the rest of the diet. People read this as science failing. It's the opposite. It's what knowledge looks like while it's maturing: early evidence produces a verdict, later evidence produces conditions, and the mature answer names the deciding variable instead of picking a side.&lt;/p&gt;

&lt;p&gt;Six months of running LLM inference on old hardware took me through exactly that arc, on about half of everything I thought I knew.&lt;/p&gt;

&lt;h2&gt;
  
  
  The scorecard
&lt;/h2&gt;

&lt;p&gt;Claims I held in the spring, audited in the fall:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Overturned or narrowed:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"Flash Attention must be off on Pascal" became per-family: off where the old measurement still applies, required where quantized KV cache demands it.&lt;/li&gt;
&lt;li&gt;"Force the MMQ kernel env var" turned out to be dead code. The kernels were always on; the variable was never read.&lt;/li&gt;
&lt;li&gt;"Row split is the P40 answer" was true, then expired twice: a fork replaced it with a mode that crashes Pascal, then upstream deleted it. The recovery came from parallel slots and speculative decoding instead.&lt;/li&gt;
&lt;li&gt;"Graph split is 40% faster" crashed on my hardware on the first real run.&lt;/li&gt;
&lt;li&gt;"Recompile to enable AVX2" was solving a problem the running binary didn't have.&lt;/li&gt;
&lt;li&gt;"48 GB of VRAM" is 45 usable once the driver takes its cut. My earliest hard lesson is literally titled "a 47GB model does not fit in 48GB."&lt;/li&gt;
&lt;li&gt;"Concurrency caps are static numbers" fell when the same endpoint admitted different loads at different hours. A cap is a worst-observed defense, not a promise.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Survived unchanged:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;CUDA as the only viable backend on this hardware, and the compute-capability facts underneath that.&lt;/li&gt;
&lt;li&gt;vLLM's non-viability on Pascal, confirmed by experiment rather than docs.&lt;/li&gt;
&lt;li&gt;Sparse MoE as the biggest throughput lever available.&lt;/li&gt;
&lt;li&gt;Q6_K as the fleet default quant.&lt;/li&gt;
&lt;li&gt;The single-variable-change rule.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Look at what separates the lists. The survivors are hardware facts and rules with a mechanism attached. Everything that expired was a verdict about a moving target: upstream code, kernel generations, a vendor's capacity pools. A verdict without its mechanism expires. The mechanism survives the verdict.&lt;/p&gt;

&lt;h2&gt;
  
  
  The evidence ladder
&lt;/h2&gt;

&lt;p&gt;The deeper takeaway isn't any single reversal. It's learning to rank evidence by how it fails. Mine, in the order I learned to trust it:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Community claims.&lt;/strong&gt; Reddit, GitHub issues, vendor blogs. Cheap, often right, occasionally catastrophic (the "40% faster" mode that crashes Pascal came from here).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ad-hoc measurements.&lt;/strong&gt; Your own numbers, one config, no controls. Better; this made a 5x regression visible but couldn't say which of four changes caused it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Controlled single-variable A/B on your own hardware.&lt;/strong&gt; The first rung where a number becomes trustworthy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Source verification.&lt;/strong&gt; Reading the code settles what a flag even does. This is the rung that exposed the dead env var.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Runtime artifact inspection.&lt;/strong&gt; The running binary's own startup banner beats the build directory's story about it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Production evidence over time.&lt;/strong&gt; Months of deployment. The only rung that catches things like a vendor silently rerouting a model id to a different capacity pool.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Each rung catches a failure mode the rungs below can't see. The expensive mistakes in this series all came from acting on rung 1 or 2 evidence as if it were rung 5 or 6.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why bother, on $2,000 of used parts
&lt;/h2&gt;

&lt;p&gt;Because the discipline is the product. The server is nice; the habits are transferable to any system whose ground truth moves: date-stamp claims, name the binary, one variable at a time, verify the premise before optimizing it, let gates outrank judgment, and treat every reversal as content rather than embarrassment.&lt;/p&gt;

&lt;p&gt;Nutrition science didn't fail when the egg advice changed. It was doing the only thing evidence-based work can do: hold the best current answer with its conditions attached, and revise when better evidence lands. Performance engineering on a fast-moving stack deserves the same posture. The point was never to be right in March. The point is for September's answer to be better, and to know exactly why it changed.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>benchmarking</category>
      <category>gpu</category>
    </item>
    <item>
      <title>The Gate Caught Me Cheating</title>
      <dc:creator>Michael Brewer</dc:creator>
      <pubDate>Tue, 08 Sep 2026 18:24:20 +0000</pubDate>
      <link>https://dev.to/devbrewery/the-gate-caught-me-cheating-4bh9</link>
      <guid>https://dev.to/devbrewery/the-gate-caught-me-cheating-4bh9</guid>
      <description>&lt;p&gt;The most embarrassing story in my notes is also the best argument for everything else in them.&lt;/p&gt;

&lt;p&gt;Every model configuration in my fleet goes through a promotion gate before it becomes the production config. The process is written down: benchmark suite at a fixed seed, real-request scoring, stream stability monitoring, and a validation test through the actual client path (the web UI), with results logged to a ledger. A candidate that passes gets promoted and frozen. A candidate without its artifacts doesn't. Every promote or invalidate decision is a line in that ledger with the evidence attached.&lt;/p&gt;

&lt;p&gt;One day a candidate config looked obviously fine. Small change, healthy smoke test, numbers where I expected them. I promoted it on the smoke test alone and moved on.&lt;/p&gt;

&lt;p&gt;The ledger's artifact rule flagged the promotion as invalid. No web-UI test report existed. The rule doesn't have a "unless you're pretty confident" clause, so the promotion was rolled back, the full test was run, the report was filed, and the candidate was re-promoted, this time with evidence.&lt;/p&gt;

&lt;p&gt;The process caught the person who wrote the process.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this is the point and not the blooper
&lt;/h2&gt;

&lt;p&gt;It's tempting to file this as a funny footnote. I think it's the story the rest of the series stands on, because of what it proves: the gate works precisely when judgment fails, and judgment fails precisely when it feels most reliable.&lt;/p&gt;

&lt;p&gt;I didn't skip the test because I was lazy. I skipped it because I was confident, and my confidence was even justified; the config was, in fact, fine. But "the config was fine" and "the process held" are two different assets, and only one of them compounds. A gate you can override when you feel sure isn't a gate. It's a suggestion with paperwork.&lt;/p&gt;

&lt;p&gt;Six months of these posts trace the same root cause in different costumes: conclusions that outlived their evidence, mechanisms assumed instead of verified, four variables changed at once. Every one of those failures was a human being sure about something. The countermeasures that held up were never "be more careful." They were structural: one variable per change, dated measurements with named sources, and gates whose rules bind their author.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this shows up beyond one server
&lt;/h2&gt;

&lt;p&gt;The same season I was learning this on GPU configs, I was applying it to the agents that use them. The agent tooling I run is governed by out-of-process policy hooks: a pre-execution gate that no amount of model confidence can talk its way around, with a file-based kill switch and default-deny rules. Same design philosophy, different layer. The judge inside the system, whether it's me on a good day or a language model on any day, doesn't get to waive the checks.&lt;/p&gt;

&lt;p&gt;If you take one thing from this series' operational posts, take the shape: write the rule down, make the rule check artifacts rather than intentions, and give the rule power over its own author. Then let it embarrass you occasionally. That's the system working.&lt;/p&gt;

&lt;p&gt;The ledger line for that config now reads: invalidated, no test report; re-promoted with report. I keep it. It's the best line in the file.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>automation</category>
    </item>
  </channel>
</rss>
