<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: dev-brewery</title>
    <description>The latest articles on DEV Community by dev-brewery (@devbrewery).</description>
    <link>https://dev.to/devbrewery</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4116219%2Feb0ac8da-7bf9-49e8-9caa-58c00ca74684.png</url>
      <title>DEV Community: dev-brewery</title>
      <link>https://dev.to/devbrewery</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/devbrewery"/>
    <language>en</language>
    <item>
      <title>Shock and Awe Is a Business Model</title>
      <dc:creator>dev-brewery</dc:creator>
      <pubDate>Tue, 08 Sep 2026 19:30:30 +0000</pubDate>
      <link>https://dev.to/devbrewery/shock-and-awe-is-a-business-model-4o11</link>
      <guid>https://dev.to/devbrewery/shock-and-awe-is-a-business-model-4o11</guid>
      <description>&lt;p&gt;Almost all of my friction with frontier models, and I mean all of them, traces back to one root cause. It isn't capability. It's tuning.&lt;/p&gt;

&lt;p&gt;Every frontier model is tuned for shock and awe. The demo has to dazzle. The first answer has to feel brilliant. The model volunteers essays when you wanted a sentence, generates confident sweeping output when you wanted a careful narrow one, and optimizes for the impression it makes in the first thirty seconds over the quality of the working relationship in month six.&lt;/p&gt;

&lt;p&gt;This is not an accident and it is not a flaw in the training pipeline. It is the business model. The adoption curve has to keep climbing so the easy capital keeps flowing. A model tuned for restraint, one that asks a clarifying question, delivers the minimum correct answer, and stops talking, would be better to work with and worse in a demo. The demo wins, because the demo is what raises the next round.&lt;/p&gt;

&lt;p&gt;I say this as a daily, heavy, mostly satisfied user of these models. They are remarkable. But remarkable is the product they're selling, and it's not the product an organization actually needs to operate.&lt;/p&gt;

&lt;h2&gt;
  
  
  What organizations actually need
&lt;/h2&gt;

&lt;p&gt;An org running AI in production needs the opposite of shock and awe. It needs method. Predictable scope. An answer that stops when the question is answered. A model that follows the runbook instead of improvising a more impressive one. Delivery over volume, consistency over brilliance.&lt;/p&gt;

&lt;p&gt;I learned this the way I learn everything, by running the systems myself. My agent stack went through five teardowns and rebuilds. The failures were never because a model was too weak. They were because models tuned to impress kept blowing out their context doing more than the task required, and because I kept trying to make cheap general models do what only a more capable or more specialized one could. The fix, every time, was narrowing: smaller scopes, tighter instructions, specialist agents, and enforcement plugins that mechanically punish showing off. I run a quality gate on my own infrastructure for exactly this reason. My agents' output is graded against the task, not against how impressive it sounds.&lt;/p&gt;

&lt;p&gt;That's a homelab-scale version of what every serious AI-adopting org is going to end up building. Not because they want to, but because the vendors' incentives and theirs point in different directions.&lt;/p&gt;

&lt;h2&gt;
  
  
  The literacy that stops being optional
&lt;/h2&gt;

&lt;p&gt;Here's the uncomfortable consequence. If the models you rent are tuned for someone else's goals, then getting models tuned for yours means understanding how tuning works. Weights, fine-tuning, quantization tradeoffs, evaluation methodology, what a training objective actually rewards. Not at research depth. At operator depth: enough to read a model card critically, enough to know what a fine-tune can and cannot fix, enough to measure whether the behavior you bought is the behavior you got.&lt;/p&gt;

&lt;p&gt;For most of the software era, this kind of knowledge was optional the way compiler internals are optional. You could build a career on top of the abstraction. I don't believe model behavior gets to stay abstracted, because the abstraction is leaking money and risk in a way compilers never did. A model that over-delivers by 3x on every request is a cost center. A model that improvises outside its scope is a liability with a legal department's name on it. Cost-effective AI that isn't a liability to the org requires someone in the building who understands what the weights were trained to do, and the orgs that treat that as a vendor's problem will pay vendor prices for vendor-aligned behavior, forever.&lt;/p&gt;

&lt;p&gt;Open weights are what make the alternative possible. When the weights are yours, tuning for method over spectacle is an engineering decision instead of a feature request to a company whose incentives run the other way. That, more than cost, is why I run open models on my own hardware and why I think the maid-to-order open-weights era is coming regardless of how loudly the incumbents warn against it. The gatekeepers in the pioneer phase always cry. It's what the phase sounds like.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the two theses meet
&lt;/h2&gt;

&lt;p&gt;Part one said: compose many specialized models the way you compose a staff. This part says: expect to tune some of them yourself, and staff for the literacy that requires.&lt;/p&gt;

&lt;p&gt;Put together, that's the whole implementation philosophy. If I were handed a real budget for a long-term AI implementation tomorrow, this is how I'd spend it. Not on the biggest model money can rent, but on a bench of specialized ones in the right seats, a routing and evaluation layer that keeps each in its lane, the measurement discipline to prove it's working, and enough weights-and-training literacy in-house that the org's AI serves the org's goals instead of its vendors'.&lt;/p&gt;

&lt;p&gt;None of that requires a frontier lab's budget. I know, because I run the small version of it on hardware nobody else wanted.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opinion</category>
    </item>
    <item>
      <title>The Genealogy Book Nobody Had Time to Read</title>
      <dc:creator>dev-brewery</dc:creator>
      <pubDate>Tue, 08 Sep 2026 19:25:13 +0000</pubDate>
      <link>https://dev.to/devbrewery/the-genealogy-book-nobody-had-time-to-read-2m9d</link>
      <guid>https://dev.to/devbrewery/the-genealogy-book-nobody-had-time-to-read-2m9d</guid>
      <description>&lt;p&gt;&lt;em&gt;What a personal agent stack is actually for.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A relative of mine was sent a book about one of our ancestors. A genealogy, hundreds of pages, compiled by someone who had clearly spent years on it. The kind of thing a family is lucky to have and almost guaranteed never to read. It needed to be gone through, cross-referenced, and connected to what we already knew about the family line. My relative did not have time for that. Nobody in the family did. It sat there the way these things sit, valuable and untouched.&lt;/p&gt;

&lt;p&gt;This is the part of the story where AI usually gets oversold, so let me be precise about what happened and what made it possible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Half an hour to digitize
&lt;/h2&gt;

&lt;p&gt;I've spent the past year building a personal agent stack: an assistant that runs on my own infrastructure, that I've customized tool by tool as real needs came up. Because that plumbing already existed, digitizing the book was not a project. It was half an hour of feeding pages through, with the agent handling OCR, cleanup, and structure as we went. Hundreds of pages became searchable, structured text before lunch.&lt;/p&gt;

&lt;p&gt;The half hour is the headline number, but it's the least interesting part. Digitization is a solved problem. What came next is why I'm writing this.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mapping the connections
&lt;/h2&gt;

&lt;p&gt;I exported our existing family tree data from Ancestry.com and handed it to the agent alongside the digitized book. Then I set it loose on the tedious part: going through the book person by person and mapping every connection onto the tree we already had. Which people in the book matched people we knew about. Which were new. Where the book confirmed our data, where it contradicted it, and where it filled holes.&lt;/p&gt;

&lt;p&gt;This is exactly the kind of work that defeats a human volunteer. Not because it's hard, but because it's hundreds of small, careful, boring judgments in a row. It's also exactly what a well-instructed agent is good at, provided you check its work. I spot-checked as it went, corrected its course a few times, and let it grind.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five hours of driving, one website
&lt;/h2&gt;

&lt;p&gt;The visit ended and I had a five hour drive home. So the drive became the build window. Voice messages from the road guided the agent through standing up an open-source genealogy website on top of the mapped data, and then through starting something the open-source tool didn't have: a new visualization for exploring the tree.&lt;/p&gt;

&lt;p&gt;I want to be honest about what "guided by voice from the car" means. It does not mean I dictated flawless instructions and arrived home to finished software. It means the agent worked, hit decisions it shouldn't make alone, and I made them at highway speed in short messages. Some of what I found when I got home was wrong and got redone. But the shape of both the site and the visualization tool existed by the time I pulled into the driveway, built during hours that would otherwise have produced nothing but mileage.&lt;/p&gt;

&lt;h2&gt;
  
  
  A couple of days of refinement
&lt;/h2&gt;

&lt;p&gt;Once home, I spent a couple of days refining with the agent: fixing the wrong turns, polishing the visualization, getting the data presentation right. At the end of it, our family history, the book's contents and the mapped tree together, was live where anyone in the family could see it.&lt;/p&gt;

&lt;p&gt;That last part matters more to me than the technology. This information used to live in two places: a paywalled account one person maintained, and a physical book one person possessed. Now it belongs to the family. A grandparent can look at it. A cousin I've never met can look at it. Nobody needs a subscription or a login or my help.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this is actually a story about
&lt;/h2&gt;

&lt;p&gt;Not about AI reading a book fast. It's about accumulated tooling meeting a real obligation. Every customization in my stack existed before this project, built for other reasons over a year of daily use. When the book arrived, the marginal cost of taking on a job nobody had time for was half an hour, a car ride, and a weekend's worth of refinement sessions.&lt;/p&gt;

&lt;p&gt;That's my working definition of what personal AI tooling is for. Not demos. Not novelty. Taking things a family cares about but cannot resource, and resourcing them.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>productivity</category>
      <category>automation</category>
    </item>
    <item>
      <title>The Right Models in the Right Seats</title>
      <dc:creator>dev-brewery</dc:creator>
      <pubDate>Tue, 08 Sep 2026 19:25:06 +0000</pubDate>
      <link>https://dev.to/devbrewery/the-right-models-in-the-right-seats-47m2</link>
      <guid>https://dev.to/devbrewery/the-right-models-in-the-right-seats-47m2</guid>
      <description>&lt;p&gt;&lt;em&gt;The thesis behind everything I build: specialization beats scale.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Ask the frontier labs what the future of AI looks like and the answer is always the same: bigger. More parameters, more compute, one ever-more-capable model that does everything. That vision has a convenient property: only a handful of companies on earth can build it, and you'll be renting it from them forever.&lt;/p&gt;

&lt;p&gt;I run my AI workloads on a different thesis, and after a year of measurements on my own hardware, I believe it more, not less.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The way forward is not ever-larger all-knowing models. It's teams of highly effective, highly specialized smaller models, composed the way a good company composes a staff.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Get the right models on the bus
&lt;/h2&gt;

&lt;p&gt;Jim Collins framed the difference between good companies and great ones as getting the right people on the bus, and the right people in the right seats. Nobody staffs a company by hiring one impossibly expensive generalist to do every job. You hire people whose strengths match their seats, and the organization outperforms the sum of its parts.&lt;/p&gt;

&lt;p&gt;My inference stack runs exactly this way, and I have a year of production numbers behind it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A 4B-parameter model running on CPU cores classifies every incoming request in a few hundred milliseconds. It cannot write good code and never will. It doesn't need to. Its seat is triage, and it fills it for free, on hardware that would otherwise idle.&lt;/li&gt;
&lt;li&gt;The same class of tiny model reads a request against a list of available tools and tells a big cloud model which one to use. That two-second nudge cut my action-request latency in half, not because the small model is smart, but because it prevented an expensive model from fumbling toward a conclusion a cheap one had already reached.&lt;/li&gt;
&lt;li&gt;A sparse mixture-of-experts model handles the fast path at 41 tokens per second on decade-old GPUs, because activating 4B parameters out of 26B is itself specialization inside the weights.&lt;/li&gt;
&lt;li&gt;The big models, local or cloud, only see the work that actually needs them. In one measured 40-hour window, 8,129 requests flowed through this division of labor with a 0.17% failure rate and no human intervention.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The lesson from those numbers isn't that small models are secretly as good as big ones. They aren't, and pretending otherwise is how you build garbage. The lesson is the one from my notes that I keep coming back to: &lt;strong&gt;the tiers aren't a hierarchy of quality, they're a division of labor.&lt;/strong&gt; A 4B model in the right seat outperforms a frontier model in the wrong one, on cost, on latency, and often on reliability, because narrow tasks reward consistency over brilliance.&lt;/p&gt;

&lt;h2&gt;
  
  
  Companies already know this pattern
&lt;/h2&gt;

&lt;p&gt;Here's why I think this thesis wins on economics, not ideology. Every company already organizes its people this way. Nobody believes the optimal workforce is one superhuman consultant billing by the token. Companies win by being specialized and nimble in their personnel, and they will demand the same from their AI: a small model fine-tuned on their support history in the support seat, a compliance-tuned model in the compliance seat, a coding model that knows their codebase in the engineering seat, orchestrated by systems that route work to whoever fills the seat best.&lt;/p&gt;

&lt;p&gt;That world is arriving through open weights. The Qwens and Gemmas I run today are the early, general-purpose versions. The trajectory points at made-to-order weights: models distilled, tuned, and owned by the companies that run them, sized to their seats, running on their hardware or commodity clouds, with their data never leaving the building. Every quarter the open releases get better at fitting into seats that used to require a frontier API call, and my own benchmarks watched it happen: models that needed 39GB of VRAM last spring were outclassed by 23GB models with better architectures by fall.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gatekeepers always cry
&lt;/h2&gt;

&lt;p&gt;Which brings me to the part of the argument that the frontier labs make for me.&lt;/p&gt;

&lt;p&gt;Listen to how the largest AI companies talk about open weights: dangerous, irresponsible, impossible to control, surely the end of safety itself. Some of those concerns are sincere and worth engaging seriously. But notice the shape of the argument and who it benefits. The pioneers of every technology wave have warned that the tools were too dangerous to leave the temple, right up until the moment the tools left anyway: mainframe companies about personal computers, telecoms about the open internet, every incumbent about every commodity that ended their toll booth.&lt;/p&gt;

&lt;p&gt;Gatekeepers in the pioneer phase always cry. It's what the phase sounds like. The economics underneath don't care: when capability becomes a commodity you can own instead of rent, composition becomes the differentiator, and composition is an engineering discipline, not a capital moat.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means for the work
&lt;/h2&gt;

&lt;p&gt;If the thesis is right, the scarce skill of the next decade isn't prompting one giant model. It's the boring, measurable systems work of running teams of models well: routing, failover, quality gates, evals that catch a model drifting out of its seat, guardrails that hold when a model exceeds its authority, and the operational discipline to measure everything, because half of what you believe about your stack will be wrong within six months. I know because I measured mine, and it was.&lt;/p&gt;

&lt;p&gt;That's the bet my basement server, my benchmarks, and this blog are all placed on. One team of specialists, on hardware nobody wanted, doing the daily work of a system that would otherwise be an expensive subscription. The right models on the bus, the right models in the right seats, and the bus is yours.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>architecture</category>
      <category>opinion</category>
    </item>
    <item>
      <title>Measure the Binary You Run</title>
      <dc:creator>dev-brewery</dc:creator>
      <pubDate>Tue, 08 Sep 2026 18:33:22 +0000</pubDate>
      <link>https://dev.to/devbrewery/measure-the-binary-you-run-4022</link>
      <guid>https://dev.to/devbrewery/measure-the-binary-you-run-4022</guid>
      <description>&lt;p&gt;At one point in this project, two documents in my own notes argued opposite positions about the same compile flag.&lt;/p&gt;

&lt;p&gt;Document one, the backend research: "The current build has AVX2 disabled. Priority 1: recompile with AVX2. Expected improvement, 15 to 30% on prompt processing."&lt;/p&gt;

&lt;p&gt;Document two, the ecosystem plan: "AVX2 is off by design, to preserve CPU headroom for the rest of the stack while the GPU serves inference."&lt;/p&gt;

&lt;p&gt;One says the flag is off and that's a problem. The other says it's off and that's a feature. They can't both be right.&lt;/p&gt;

&lt;p&gt;Neither was. The flag was on the whole time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where both documents went wrong
&lt;/h2&gt;

&lt;p&gt;The evidence behind "AVX2 is disabled" was the cmake build cache, which showed the SIMD options set to OFF. Build-directory archaeology: inspect the configuration, infer the artifact.&lt;/p&gt;

&lt;p&gt;But llama.cpp's build enables native CPU optimizations through its own path regardless of those cached toggles, and the shipped binary is perfectly willing to tell you what it contains. Its startup system info prints the actual capability set. Mine printed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AVX = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The running binary had AVX2 enabled all along. The "disabled" reading described cmake defaults that the real build had overridden. And the second document's clever theory about why disabled-AVX2 was good design was a rationalization of a condition that didn't exist.&lt;/p&gt;

&lt;p&gt;The regression incident from post 2 has a footnote here too: that four-variables-at-once rebuild included "enable AVX2" as one of its four changes. In its own words, in the post-mortem: redundant. One of the four simultaneous changes was a no-op, which muddied attribution further.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hierarchy this settled
&lt;/h2&gt;

&lt;p&gt;My notes rank evidence quality in six rungs, and this incident fixed the ordering of two of them permanently:&lt;/p&gt;

&lt;p&gt;Runtime artifact inspection beats build-system archaeology. Always. The cmake cache tells you what the build system was asked. The binary's own system info tells you what you are actually running. When they disagree, the binary wins, because the binary is what serves your traffic.&lt;/p&gt;

&lt;p&gt;The general form: verify the premise before optimizing it. "Recompile to enable AVX2" was a well-reasoned recommendation, correctly derived from its evidence, aimed at a switch that was already in the right position. Everything about the plan was sound except the fact it stood on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cheap checks, expensive assumptions
&lt;/h2&gt;

&lt;p&gt;What makes this sting is how cheap the correct check was. The system info line prints at every server start. It was in every log I'd ever launched. Nobody had read it, because everyone was reading the build directory instead, where the interesting-looking configuration lives.&lt;/p&gt;

&lt;p&gt;Since then, the rule in my fleet is that claims about a binary come from the binary: its startup banner, its version string, its measured behavior. Configuration files describe intent. Artifacts describe reality. Optimization work starts from reality.&lt;/p&gt;

&lt;p&gt;Two documents argued about a switch that was already in the right position. The moral isn't that documents are bad. It's that both documents cited the same wrong source, and one five-second look at the right source would have ended the argument before it started.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>benchmarking</category>
      <category>devops</category>
    </item>
    <item>
      <title>The Year I Started Finishing Things</title>
      <dc:creator>dev-brewery</dc:creator>
      <pubDate>Tue, 08 Sep 2026 18:29:44 +0000</pubDate>
      <link>https://dev.to/devbrewery/the-year-i-started-finishing-things-3hkg</link>
      <guid>https://dev.to/devbrewery/the-year-i-started-finishing-things-3hkg</guid>
      <description>&lt;p&gt;&lt;em&gt;On using AI agents to run more of a life than one person's spare time should hold.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I have a full-time job, a family, a two-acre hilltop, a homelab, and a seat on the board of a small nonprofit. For most of my adult life, the limiting factor on what I could take on wasn't ability or interest. It was hours. Things I cared about sat unfinished for months because the twenty minutes they needed never lined up with the twenty minutes I had.&lt;/p&gt;

&lt;p&gt;This year that changed, and I want to write down how, honestly, without the hype that usually smothers this topic.&lt;/p&gt;

&lt;h2&gt;
  
  
  The nonprofit problem
&lt;/h2&gt;

&lt;p&gt;A couple of years ago I joined the board of a small all-volunteer nonprofit that supports pastor training in West Africa. I serve as secretary, and because I'm the technical one, everything with a login became mine: the website, the records, the state filings, the donor communications, the grant research nobody had time to do.&lt;/p&gt;

&lt;p&gt;Here's what that job actually looks like at a tiny nonprofit. There is no staff. There is no budget for staff. Every task is small, none of them are optional, and they arrive continuously: a board email that needs filing, a compliance deadline nobody remembers until it's urgent, a partner organization posting updates that should reach our supporters, a funder whose eligibility rules need reading before anyone wastes a weekend on an application. Any one of these is twenty minutes. Together they're a part-time job that nobody has.&lt;/p&gt;

&lt;p&gt;The standard outcome is that volunteer boards run on heroics and guilt. Things slip. The person who cares most burns out first.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I actually built
&lt;/h2&gt;

&lt;p&gt;Over the past year I stood up an AI agent that functions as the org's administrative assistant. It runs on a small server, talks to me over chat, and drives our ordinary office tools: email, calendar, the shared drive, a spreadsheet that acts as the operations hub. Around two dozen scheduled jobs handle the recurring work.&lt;/p&gt;

&lt;p&gt;A normal day looks like this. At 9 AM I get a short message: here's what's outstanding, here are the top three things, here's the link if you want the whole picture. Board emails that arrived overnight have been triaged: routine ones handled and filed, anything needing my judgment flagged. Documents dropped in a shared folder have been renamed, filed, and logged. If our partner ministry posted an update, a draft is waiting on the website for approval. Once in a while there's a note that a funding opportunity surfaced overnight, with deadlines and fit notes attached.&lt;/p&gt;

&lt;p&gt;I read it in the time it takes to drink coffee, make the two or three decisions that actually need a human, and go to work.&lt;/p&gt;

&lt;p&gt;Some of what came out of this, concretely:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Compliance filings that had been hanging over the board since incorporation got done.&lt;/strong&gt; The agent assembled the filing packages; I reviewed and signed; the state approved them. That was the single biggest weight off the board's shoulders, and it was mostly review time on my end.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Grant prospecting went from nonexistent to nightly.&lt;/strong&gt; No volunteer was ever going to spend evenings scanning funder sites. An automated job does, and it reads eligibility rules against our actual documents before anything reaches the board. It caught, for example, that a foundation we liked requires organizations to be older than we are, by pulling our real incorporation date from our IRS paperwork. That's a wasted application avoided and an honest reminder scheduled for when we qualify.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Board records became records.&lt;/strong&gt; Every decision and email is logged and retrievable. When someone asks "what was that concept note from last year?", the answer takes seconds instead of inbox archaeology.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The website updates itself, mostly.&lt;/strong&gt; Partner updates become drafts I approve. Promo materials that would have been a multi-day design favor get generated from source material in the org's own style in an afternoon.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What I refuse to pretend
&lt;/h2&gt;

&lt;p&gt;This is the part most writing about AI leaves out, so let me be specific.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The failures are real and constant.&lt;/strong&gt; The token for the primary model expired once and the system quietly degraded to cheaper models for days before anyone noticed; the fix was another automated check, watching the watchers. A reminder got scheduled for the wrong year; I caught it. A briefing once referenced a project that had a tracker row but no documentation; the fix was auditing everything and making cross-referencing a standing rule. A funding deadline surfaced too late to act on. Two emails failed to send and were done by hand. The first architecture for the recurring jobs was wasteful and had to be redesigned. A monitoring job I thought was running had been off for months.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Human oversight is load-bearing.&lt;/strong&gt; Nothing goes to the board, a funder, or the public without my review. The system's job is to make my twenty minutes count, not to replace my judgment. Every one of the failures above was caught either by an automated check I added after a previous failure, or by me reading something before it went out. Both layers earn their keep monthly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The outcomes are modest and I'll state them plainly.&lt;/strong&gt; We have not won a grant. We haven't even submitted an application yet; the pipeline's honest output so far is one fully researched, board-ready opportunity and a list of funders we now know we don't qualify for, with dated reasons. What changed is that the org went from no systematic effort to a running process that costs almost nothing and never gets tired. For an all-volunteer organization, the difference between "nobody has time" and "it happens every night" is the whole game.&lt;/p&gt;

&lt;h2&gt;
  
  
  The general lesson
&lt;/h2&gt;

&lt;p&gt;The same year, the same pattern ran through everything else I do: the inference server in my basement that this blog documents, the agent tooling at my day job that turns client call recordings into project proposals, the monitoring that tells me about problems before I go looking. None of it is one big AI doing something impressive. All of it is small, boring automation with careful boundaries, checked by a human, compounding.&lt;/p&gt;

&lt;p&gt;What AI actually bought me this year is not intelligence. I still make every decision that matters. It bought me &lt;strong&gt;parallelism&lt;/strong&gt;. Things now make progress during hours I'm not present: overnight, during the workday, while I'm mowing the hill. My attention became the scarce resource the whole system is designed to spend well, twenty minutes at a time.&lt;/p&gt;

&lt;p&gt;A year ago, the honest description of my volunteer role was "important things slip, and I feel bad about it." Today it's "the routine runs itself, and I spend my time on judgment." Same hours. Same person. That's the difference, and for a small mission-driven organization that difference isn't a productivity statistic. It's whether the work of keeping the lights on leaves any energy for the mission itself.&lt;/p&gt;

</description>
      <category>career</category>
      <category>productivity</category>
      <category>ai</category>
    </item>
    <item>
      <title>Eggs, Cholesterol, and GPU Flags</title>
      <dc:creator>dev-brewery</dc:creator>
      <pubDate>Tue, 08 Sep 2026 18:24:27 +0000</pubDate>
      <link>https://dev.to/devbrewery/eggs-cholesterol-and-gpu-flags-561j</link>
      <guid>https://dev.to/devbrewery/eggs-cholesterol-and-gpu-flags-561j</guid>
      <description>&lt;p&gt;For decades, nutrition science flip-flopped on eggs. Bad for you: dietary cholesterol. Then fine, then good in some contexts, then it depends on the person and the rest of the diet. People read this as science failing. It's the opposite. It's what knowledge looks like while it's maturing: early evidence produces a verdict, later evidence produces conditions, and the mature answer names the deciding variable instead of picking a side.&lt;/p&gt;

&lt;p&gt;Six months of running LLM inference on old hardware took me through exactly that arc, on about half of everything I thought I knew.&lt;/p&gt;

&lt;h2&gt;
  
  
  The scorecard
&lt;/h2&gt;

&lt;p&gt;Claims I held in the spring, audited in the fall:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Overturned or narrowed:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"Flash Attention must be off on Pascal" became per-family: off where the old measurement still applies, required where quantized KV cache demands it.&lt;/li&gt;
&lt;li&gt;"Force the MMQ kernel env var" turned out to be dead code. The kernels were always on; the variable was never read.&lt;/li&gt;
&lt;li&gt;"Row split is the P40 answer" was true, then expired twice: a fork replaced it with a mode that crashes Pascal, then upstream deleted it. The recovery came from parallel slots and speculative decoding instead.&lt;/li&gt;
&lt;li&gt;"Graph split is 40% faster" crashed on my hardware on the first real run.&lt;/li&gt;
&lt;li&gt;"Recompile to enable AVX2" was solving a problem the running binary didn't have.&lt;/li&gt;
&lt;li&gt;"48 GB of VRAM" is 45 usable once the driver takes its cut. My earliest hard lesson is literally titled "a 47GB model does not fit in 48GB."&lt;/li&gt;
&lt;li&gt;"Concurrency caps are static numbers" fell when the same endpoint admitted different loads at different hours. A cap is a worst-observed defense, not a promise.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Survived unchanged:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;CUDA as the only viable backend on this hardware, and the compute-capability facts underneath that.&lt;/li&gt;
&lt;li&gt;vLLM's non-viability on Pascal, confirmed by experiment rather than docs.&lt;/li&gt;
&lt;li&gt;Sparse MoE as the biggest throughput lever available.&lt;/li&gt;
&lt;li&gt;Q6_K as the fleet default quant.&lt;/li&gt;
&lt;li&gt;The single-variable-change rule.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Look at what separates the lists. The survivors are hardware facts and rules with a mechanism attached. Everything that expired was a verdict about a moving target: upstream code, kernel generations, a vendor's capacity pools. A verdict without its mechanism expires. The mechanism survives the verdict.&lt;/p&gt;

&lt;h2&gt;
  
  
  The evidence ladder
&lt;/h2&gt;

&lt;p&gt;The deeper takeaway isn't any single reversal. It's learning to rank evidence by how it fails. Mine, in the order I learned to trust it:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Community claims.&lt;/strong&gt; Reddit, GitHub issues, vendor blogs. Cheap, often right, occasionally catastrophic (the "40% faster" mode that crashes Pascal came from here).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ad-hoc measurements.&lt;/strong&gt; Your own numbers, one config, no controls. Better; this made a 5x regression visible but couldn't say which of four changes caused it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Controlled single-variable A/B on your own hardware.&lt;/strong&gt; The first rung where a number becomes trustworthy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Source verification.&lt;/strong&gt; Reading the code settles what a flag even does. This is the rung that exposed the dead env var.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Runtime artifact inspection.&lt;/strong&gt; The running binary's own startup banner beats the build directory's story about it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Production evidence over time.&lt;/strong&gt; Months of deployment. The only rung that catches things like a vendor silently rerouting a model id to a different capacity pool.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Each rung catches a failure mode the rungs below can't see. The expensive mistakes in this series all came from acting on rung 1 or 2 evidence as if it were rung 5 or 6.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why bother, on $2,000 of used parts
&lt;/h2&gt;

&lt;p&gt;Because the discipline is the product. The server is nice; the habits are transferable to any system whose ground truth moves: date-stamp claims, name the binary, one variable at a time, verify the premise before optimizing it, let gates outrank judgment, and treat every reversal as content rather than embarrassment.&lt;/p&gt;

&lt;p&gt;Nutrition science didn't fail when the egg advice changed. It was doing the only thing evidence-based work can do: hold the best current answer with its conditions attached, and revise when better evidence lands. Performance engineering on a fast-moving stack deserves the same posture. The point was never to be right in March. The point is for September's answer to be better, and to know exactly why it changed.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>benchmarking</category>
      <category>gpu</category>
    </item>
    <item>
      <title>The Gate Caught Me Cheating</title>
      <dc:creator>dev-brewery</dc:creator>
      <pubDate>Tue, 08 Sep 2026 18:24:20 +0000</pubDate>
      <link>https://dev.to/devbrewery/the-gate-caught-me-cheating-4bh9</link>
      <guid>https://dev.to/devbrewery/the-gate-caught-me-cheating-4bh9</guid>
      <description>&lt;p&gt;The most embarrassing story in my notes is also the best argument for everything else in them.&lt;/p&gt;

&lt;p&gt;Every model configuration in my fleet goes through a promotion gate before it becomes the production config. The process is written down: benchmark suite at a fixed seed, real-request scoring, stream stability monitoring, and a validation test through the actual client path (the web UI), with results logged to a ledger. A candidate that passes gets promoted and frozen. A candidate without its artifacts doesn't. Every promote or invalidate decision is a line in that ledger with the evidence attached.&lt;/p&gt;

&lt;p&gt;One day a candidate config looked obviously fine. Small change, healthy smoke test, numbers where I expected them. I promoted it on the smoke test alone and moved on.&lt;/p&gt;

&lt;p&gt;The ledger's artifact rule flagged the promotion as invalid. No web-UI test report existed. The rule doesn't have a "unless you're pretty confident" clause, so the promotion was rolled back, the full test was run, the report was filed, and the candidate was re-promoted, this time with evidence.&lt;/p&gt;

&lt;p&gt;The process caught the person who wrote the process.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this is the point and not the blooper
&lt;/h2&gt;

&lt;p&gt;It's tempting to file this as a funny footnote. I think it's the load-bearing story of the whole series, because of what it proves: the gate works precisely when judgment fails, and judgment fails precisely when it feels most reliable.&lt;/p&gt;

&lt;p&gt;I didn't skip the test because I was lazy. I skipped it because I was confident, and my confidence was even justified; the config was, in fact, fine. But "the config was fine" and "the process held" are two different assets, and only one of them compounds. A gate you can override when you feel sure isn't a gate. It's a suggestion with paperwork.&lt;/p&gt;

&lt;p&gt;Six months of these posts trace the same root cause in different costumes: conclusions that outlived their evidence, mechanisms assumed instead of verified, four variables changed at once. Every one of those failures was a human being sure about something. The countermeasures that actually worked were never "be more careful." They were structural: one variable per change, dated measurements with named sources, and gates whose rules bind their author.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this shows up beyond one server
&lt;/h2&gt;

&lt;p&gt;The same season I was learning this on GPU configs, I was applying it to the agents that use them. The agent tooling I run is governed by out-of-process policy hooks: a pre-execution gate that no amount of model confidence can talk its way around, with a file-based kill switch and default-deny rules. Same design philosophy, different layer. The judge inside the system, whether it's me on a good day or a language model on any day, doesn't get to waive the checks.&lt;/p&gt;

&lt;p&gt;If you take one thing from this series' operational posts, take the shape: write the rule down, make the rule check artifacts rather than intentions, and give the rule power over its own author. Then let it embarrass you occasionally. That's the system working.&lt;/p&gt;

&lt;p&gt;The ledger line for that config now reads: invalidated, no test report; re-promoted with report. I keep it. It's the best line in the file.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>automation</category>
    </item>
    <item>
      <title>Buying Speed With Architecture</title>
      <dc:creator>dev-brewery</dc:creator>
      <pubDate>Tue, 08 Sep 2026 18:19:04 +0000</pubDate>
      <link>https://dev.to/devbrewery/buying-speed-with-architecture-424d</link>
      <guid>https://dev.to/devbrewery/buying-speed-with-architecture-424d</guid>
      <description>&lt;p&gt;Every post so far has been about flags: flags that died, flags that did nothing, flags that flipped. This one is about the uncomfortable truth on the other side of all that tuning: on fixed hardware, the biggest speed wins in this project didn't come from configuration at all. They came from what the model is and what the server can do with it.&lt;/p&gt;

&lt;p&gt;Ranked by measured impact on my dual Tesla P40s:&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Sparse MoE: 41 tokens/sec on 2016 GPUs
&lt;/h2&gt;

&lt;p&gt;The single largest lever, by a wide margin. Gemma 4 26B-A4B is a mixture-of-experts model: 26B parameters on disk, about 3.8B active per token. My dense 27B models decode at 8.5 to 20 tokens/sec depending on config. The MoE runs 41.&lt;/p&gt;

&lt;p&gt;Nothing about my hardware changed. The model simply does less work per token, by design. Sparse activation bought more throughput than every flag decision in this series combined, and it's not close. If your hardware is old and your workload tolerates the model family, MoE is the first question to ask, not the last.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. MTP speculative decoding: +57% single-stream
&lt;/h2&gt;

&lt;p&gt;The daily-driver stack (27B dense, Q6_K) measured 8.46 tokens/sec single-stream after the row-split removal. Enabling MTP speculative decoding, where the model's own multi-token-prediction layers draft ahead and the main pass verifies, lifted that to about 13.3 tokens/sec in a controlled A/B. Acceptance rates ran 0.38 to 0.63, and output correctness was verified against non-speculative runs.&lt;/p&gt;

&lt;p&gt;Free speed is rare. This is the closest thing to it I found: no quality cost, no extra VRAM of consequence, one flag, 57%.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Parallel slots: aggregate throughput nearly doubles
&lt;/h2&gt;

&lt;p&gt;Same stack, measured at 1, 2, and 4 concurrent slots: 8.46 / 12.8 / 15.0 tokens/sec aggregate. Per-request latency degrades gently until the slots saturate, which means one swapped-in model can serve several concurrent agents acceptably instead of queuing them.&lt;/p&gt;

&lt;p&gt;For an agent workload, aggregate is the number that matters. My traffic is dozens of tool-call round-trips, not one human reading one stream. Four slots turned "one user at a time" into "the whole agent fleet."&lt;/p&gt;

&lt;h2&gt;
  
  
  The stacking effect
&lt;/h2&gt;

&lt;p&gt;These compose. MoE where the family fits, MTP and parallel slots where it doesn't. The result across the fleet: the post-row-split era ended up faster than the row-split era it mourned. The 12-14 tokens/sec I lost to a deleted flag came back as 15 aggregate with better concurrency, and the fast path tripled it.&lt;/p&gt;

&lt;p&gt;There's a budget lesson in that. I spent months on flags worth 10 to 40% each, some of which later turned out to be dead code or expired rules. The architecture-level choices were worth 100 to 400%, and they were sitting in plain sight the whole time: pick a sparser model, turn on the decoding feature the model ships with, let the server batch.&lt;/p&gt;

&lt;p&gt;Tune the flags, but tune them last. On fixed hardware, architecture is the knob with the range.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>llm</category>
      <category>performance</category>
    </item>
    <item>
      <title>The Template Mattered More Than the Quant</title>
      <dc:creator>dev-brewery</dc:creator>
      <pubDate>Tue, 08 Sep 2026 18:18:57 +0000</pubDate>
      <link>https://dev.to/devbrewery/the-template-mattered-more-than-the-quant-2elg</link>
      <guid>https://dev.to/devbrewery/the-template-mattered-more-than-the-quant-2elg</guid>
      <description>&lt;p&gt;Ask anyone tuning local LLMs where quality lives and you'll hear about quantization. Q4 versus Q6 versus Q8, perplexity curves, "never go below Q5 for reasoning." It's the knob everyone debates because it's the knob with numbers attached.&lt;/p&gt;

&lt;p&gt;Here are my measured results on a 32B model, same benchmark suite, fixed seed:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Quantization comparison (20-question MCQ suite, seed 42):&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Q8_0: 15/20 correct, 19.6 tokens/sec, ~39.5 GB&lt;/li&gt;
&lt;li&gt;Q6_K: 15/20 correct, 20.3 tokens/sec, ~29 GB&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A full quantization step moved accuracy not at all. Zero questions. The smaller quant was actually faster (less memory traffic) and left 10 GB more room for KV cache. That one table set my fleet's default: Q6_K as the daily driver, Q8_0 as a stretch when VRAM allows.&lt;/p&gt;

&lt;p&gt;Now the same model, same quant, same seed, changing only the chat template:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Template comparison (scored /10):&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Model's own template: 3-4&lt;/li&gt;
&lt;li&gt;ChatML template: 0&lt;/li&gt;
&lt;li&gt;no_think mode + model template: 3&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The wrong template zeroed the model. Not degraded it, zeroed it. A model that scores reliably with its own template produced nothing scoreable when wrapped in ChatML, the template half the internet's example configs default to.&lt;/p&gt;

&lt;p&gt;Template choice had a larger effect size than any quantization decision I measured. It isn't close.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this surprises people
&lt;/h2&gt;

&lt;p&gt;Quantization is a continuous, well-instrumented knob with academic literature behind it. Templates are a formatting detail buried in a GGUF's metadata or a server flag. One looks like engineering, the other looks like plumbing.&lt;/p&gt;

&lt;p&gt;But think about what each one actually perturbs. Quantization adds small rounding error to every weight, and modern quant schemes are engineered to keep that error away from what matters. A wrong template perturbs the input distribution itself: the model sees token sequences that never appeared in its training in that arrangement, malformed role markers, missing control tokens. From the model's perspective, quantization is a slight blur; the wrong template is a foreign language.&lt;/p&gt;

&lt;p&gt;The failure is also silent. A mis-templated model still generates fluent text. It doesn't crash, it doesn't warn, it just gets steadily and confidently worse at its job. If my benchmark hadn't scored outputs, I could have shipped that config and spent weeks blaming the quant, the sampler, or the model itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The discipline that caught it
&lt;/h2&gt;

&lt;p&gt;Nothing clever caught this. A boring benchmark suite did: fixed seed, fixed question set, scored outputs, run per candidate config before promotion. The suite exists because eyeballing model quality is how you fool yourself; scores at a fixed seed are how the difference between "seems fine" and "scores zero" becomes visible.&lt;/p&gt;

&lt;p&gt;The order of operations this experience installed:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Verify the template first. It's the highest-leverage, least-discussed setting in local inference.&lt;/li&gt;
&lt;li&gt;Then choose the smallest quant that holds your benchmark scores. The savings go to context or speed.&lt;/li&gt;
&lt;li&gt;Distrust any quality comparison, including between whole models, that doesn't control for template. Some fraction of "model A beats model B locally" posts are template bugs wearing a costume.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Check the plumbing before you argue about the engineering. The cheap setting nobody benchmarks moved my scores more than the expensive setting everybody does.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>ai</category>
      <category>programming</category>
    </item>
    <item>
      <title>When the Reason Changes, the Flag Flips</title>
      <dc:creator>dev-brewery</dc:creator>
      <pubDate>Tue, 08 Sep 2026 18:13:31 +0000</pubDate>
      <link>https://dev.to/devbrewery/when-the-reason-changes-the-flag-flips-5b8l</link>
      <guid>https://dev.to/devbrewery/when-the-reason-changes-the-flag-flips-5b8l</guid>
      <description>&lt;p&gt;For most of this project's life, one rule was untouchable: Flash Attention off. Non-negotiable. Pascal GPUs have no Tensor Cores, and the measured result on my Tesla P40s matched the community's: enabling FA ran about 50% slower. The flag &lt;code&gt;-fa off&lt;/code&gt; sat in every start script, and it belonged there.&lt;/p&gt;

&lt;p&gt;Today, several of my most-used stacks run with Flash Attention on. On the same GPUs. And they're right to.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed
&lt;/h2&gt;

&lt;p&gt;Not the hardware. The reason.&lt;/p&gt;

&lt;p&gt;The stacks in question run quantized KV cache (&lt;code&gt;-ctv q8_0&lt;/code&gt;), which halves the memory cost of every token of context. llama.cpp requires Flash Attention for quantized V cache; that's a hard constraint in the code, not a preference. So on those stacks, FA on is the admission price for q8_0 KV.&lt;/p&gt;

&lt;p&gt;And q8_0 KV is what buys the headline capability of my daily-driver stack: a 262k-token context window on a 27B dense model, in about 23 GB of VRAM, on a GPU that was current when Obama was president. Without quantized KV, that context would need roughly twice the memory and would not fit.&lt;/p&gt;

&lt;p&gt;So the deployed matrix today looks like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Dense SSM-hybrid stacks: FA on, because quantized KV requires it, and long context is their job.&lt;/li&gt;
&lt;li&gt;One MoE hybrid: FA on with f16 KV, deployed as measured-best on that family (recorded honestly in my notes as observed-deployed, not a controlled A/B).&lt;/li&gt;
&lt;li&gt;The older families, Gemma 4 and the pre-hybrid Qwens and the 80B MoEs: FA off, because for them the classic rule is still correct.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Same silicon. Three different answers. The deciding variables are model family and KV quantization, not GPU generation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The trade underneath
&lt;/h2&gt;

&lt;p&gt;Nothing here is free, and naming the cost is the point. Quantized KV costs about 16% decode speed, a known measured penalty. My long-context stack pays 16% of its speed for double the context per byte of VRAM. For a latency-sensitive stack, the same trade would be wrong.&lt;/p&gt;

&lt;p&gt;"FA off on Pascal" was a measured fact about one kernel generation and one use case. I mistook it for a law of the hardware. The measurement was never wrong; my generalization of it was.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern
&lt;/h2&gt;

&lt;p&gt;This is the third post in a row that lands on the same shape of lesson, and that's not an accident, it's the thesis of the series:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The deleted flag (post 2): a true rule expired when upstream changed.&lt;/li&gt;
&lt;li&gt;The dead env var (post 3): a true conclusion rode on a false mechanism.&lt;/li&gt;
&lt;li&gt;Flash Attention (this post): a true measurement got over-generalized into a false rule.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A verdict without its mechanism expires. "FA off" was a verdict. "FA is slower on Pascal's kernel generation, unless a feature you need requires it, in which case pay the price knowingly" is a mechanism with its conditions attached. The second form survives upstream changes, new model families, and new requirements. The first form silently goes stale and costs you capabilities you didn't know you'd given up.&lt;/p&gt;

&lt;p&gt;When someone hands you a hardware rule, ask what it's conditioned on. If the answer is "nothing, it's just true," it's probably a verdict that hasn't met its exception yet.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>opensource</category>
      <category>programming</category>
    </item>
    <item>
      <title>The Critical Env Var That Did Nothing</title>
      <dc:creator>dev-brewery</dc:creator>
      <pubDate>Tue, 08 Sep 2026 18:13:25 +0000</pubDate>
      <link>https://dev.to/devbrewery/the-critical-env-var-that-did-nothing-1ohi</link>
      <guid>https://dev.to/devbrewery/the-critical-env-var-that-did-nothing-1ohi</guid>
      <description>&lt;p&gt;Every start script in my fleet carried the same line, under the same banner comment:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# CRITICAL: Forcing MMQ kernels for Pascal GPUs&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;GGML_CUDA_FORCE_MMQ&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;My hardware constraints doc listed it as a requirement. Community guidance for Tesla P40s repeats it everywhere: Pascal has no Tensor Cores, its strength is INT8 matmul through the &lt;code&gt;dp4a&lt;/code&gt; instruction, and MMQ is the kernel path that uses it. Forcing MMQ on is the single most repeated piece of P40 advice on the internet.&lt;/p&gt;

&lt;p&gt;In May 2026, while writing build documentation for a new stack, I actually read the kernel selection code.&lt;/p&gt;

&lt;p&gt;The env var is dead code. On upstream llama.cpp, &lt;code&gt;ggml_cuda_should_use_mmq()&lt;/code&gt; selects MMQ unconditionally on compute capability 6.1. Pascal meets the DP4A minimum and has no FP16 tensor core path, so there is no other kernel the code could choose. The runtime environment variable is never read. Only a cmake option by a similar name exists, and it does something different at build time.&lt;/p&gt;

&lt;p&gt;The kernels I was "forcing" were always on. They could not have been off.&lt;/p&gt;

&lt;h2&gt;
  
  
  The interesting part is why nobody noticed
&lt;/h2&gt;

&lt;p&gt;The advice was harmless. That's exactly what made it invisible. Exporting the variable cost nothing, broke nothing, and the resulting performance was good, so the ritual survived every review. If the variable had hurt performance, someone would have caught it years ago.&lt;/p&gt;

&lt;p&gt;This is the definition of a cargo-cult flag: a config line that travels from tutorial to tutorial because removing it feels riskier than keeping it, and no one's measurement would change either way.&lt;/p&gt;

&lt;p&gt;I want to be precise about what was wrong here, because the distinction matters. The underlying claim was true: INT8 MMQ matmul genuinely is the right kernel path for a P40. Only the mechanism claim was false, the belief that you had to force it. A true conclusion propped up by a false mechanism is still a landmine, because you'll carry the false mechanism into the next decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  What source verification buys you
&lt;/h2&gt;

&lt;p&gt;My notes rank evidence in six rungs, and this incident is why "read the source" outranks "benchmark it yourself." No benchmark would have caught this. A/B testing the env var produces identical numbers on both sides, which is easy to misread as "the flag is so important it's saturated," rather than "the flag is disconnected."&lt;/p&gt;

&lt;p&gt;Only the code answers what a flag actually does:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Community claims told me to set it.&lt;/li&gt;
&lt;li&gt;My own measurements couldn't distinguish it.&lt;/li&gt;
&lt;li&gt;The source told me it was never read.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The fix in my fleet was deliberately conservative. New stacks drop the export and document why. Old stacks keep it, because it's harmless and their configs are frozen with their measurements. And the constraints doc got the stale line flagged for review rather than silently edited, because silently rewriting your own historical record is how you lose the ability to trust it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;Before you optimize a flag, verify the flag is connected to anything. The cheapest possible check, reading the selection logic, settled in minutes what years of repeated community advice never questioned.&lt;/p&gt;

&lt;p&gt;Somewhere in your config, right now, there is a line that does nothing. It's probably the one with "CRITICAL" in the comment.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>debugging</category>
      <category>devops</category>
    </item>
    <item>
      <title>The Flag We Tuned Around Got Deleted</title>
      <dc:creator>dev-brewery</dc:creator>
      <pubDate>Tue, 08 Sep 2026 18:06:52 +0000</pubDate>
      <link>https://dev.to/devbrewery/the-flag-we-tuned-around-got-deleted-31ge</link>
      <guid>https://dev.to/devbrewery/the-flag-we-tuned-around-got-deleted-31ge</guid>
      <description>&lt;p&gt;The single most important llama.cpp flag for my dual Tesla P40 setup was &lt;code&gt;-sm row&lt;/code&gt;. It split every layer's tensors across both GPUs and it was worth nearly double the throughput of the alternative: 12-14 tokens/sec against about 7 for layer split. Every stack I built was tuned around it.&lt;/p&gt;

&lt;p&gt;In July 2026, upstream llama.cpp deleted it. Not deprecated. Deleted.&lt;/p&gt;

&lt;p&gt;This is the story of a performance rule that died twice, and what replaced the throughput it took with it. It's the longest arc in my notes, and it runs in five acts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Act 1: Row wins
&lt;/h2&gt;

&lt;p&gt;February 2026. My first serious model, a 72B, started at 2.7 to 3.6 tokens/sec in what my notes politely call "poor configuration." Working up the offload ladder to full GPU residency and switching to row split produced the first real daily driver: about 10.3 tokens/sec generation, 60 tokens/sec prompt processing.&lt;/p&gt;

&lt;p&gt;The standing rule crystallized: row split, 12-14 tokens/sec. Layer split, about 7. And a rumor from a vendor blog said a newer "graph" split mode was worth another 30-40%.&lt;/p&gt;

&lt;h2&gt;
  
  
  Act 2: The regression incident
&lt;/h2&gt;

&lt;p&gt;March 2026. I rebuilt the same model on a "modernized" fork, and changed four variables at once: split mode, SIMD compile flags, kernel selection method, and compression settings. Prompt processing collapsed from 153 tokens/sec to 29. Five times slower, and with four simultaneous changes, nothing was attributable.&lt;/p&gt;

&lt;p&gt;The cleanup A/B on March 4 isolated everything. Same model, same prompt, one variable at a time:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Original binary, row split: 60 tokens/sec prompt, 10.3 generation. Works.&lt;/li&gt;
&lt;li&gt;New fork, layer split: half the speed.&lt;/li&gt;
&lt;li&gt;New fork, graph split: CUDA crash. &lt;code&gt;ROPE failed: an illegal memory access&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The "40% faster" graph mode does not run on Pascal at all. That's the difference between a community claim and a measurement on your own hardware: one of them can crash.&lt;/p&gt;

&lt;p&gt;The verdict written that day: the original binary is optimal for this hardware, do not switch. And a rule was born that I now treat as non-negotiable: one variable per change. That rule was paid for in lost throughput and a wasted week.&lt;/p&gt;

&lt;h2&gt;
  
  
  Act 3: The fast mode becomes the wrong mode
&lt;/h2&gt;

&lt;p&gt;May 2026. Gemma 4 arrived, and its architecture uses shared KV layers implemented as tensor views. Those crash row split on multi-GPU, a hard assert deep in the CUDA backend, known upstream issue. Every Gemma stack I run is layer split by necessity, knowingly paying the throughput cost. The Qwen architectures have no such bug and kept row.&lt;/p&gt;

&lt;p&gt;So split mode is not a performance dial. It's also a correctness knob, and the answer is per-architecture. "Row is faster" was true and useless without the condition attached.&lt;/p&gt;

&lt;h2&gt;
  
  
  Act 4: Upstream deletes row
&lt;/h2&gt;

&lt;p&gt;July 6, 2026. llama.cpp removed &lt;code&gt;-sm row&lt;/code&gt; entirely. Any stack built from a newer clone has layer split as its only multi-GPU option. The flag I had organized my fleet around no longer exists in the binaries I build.&lt;/p&gt;

&lt;p&gt;Row died twice, in two lineages: first the fork replaced it with a mode that crashes Pascal, then upstream removed it outright.&lt;/p&gt;

&lt;h2&gt;
  
  
  Act 5: The win comes back from somewhere else
&lt;/h2&gt;

&lt;p&gt;Here's the part that justified the whole ordeal. On the new stack, layer split alone measured 8.46 tokens/sec single-stream, right where the old "layer is about 7" rule predicted. But two features that didn't exist in my February binaries changed the math:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;--parallel 4&lt;/code&gt; (four concurrent slots): 8.46 / 12.8 / 15.0 tokens/sec at 1, 2, and 4 slots. Aggregate throughput nearly doubles before per-slot latency degrades.&lt;/li&gt;
&lt;li&gt;MTP speculative decoding: 8.46 to about 13.3 tokens/sec single-stream, a 57% lift, acceptance rates 0.38 to 0.63, output correctness verified.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Net result: the fleet ended up faster than the row-split era without row split. The recovery didn't come from finding a replacement flag. It came from features orthogonal to the one I lost.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this taught me
&lt;/h2&gt;

&lt;p&gt;Date-stamp every performance claim and name the binary it was measured on. A tuning rule is a fact about a specific artifact at a specific moment, not a law of the hardware.&lt;/p&gt;

&lt;p&gt;And when the flag you tuned around disappears, re-measure before assuming regression. The replacement win may live somewhere you weren't looking. Mine did, twice over.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>opensource</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
