<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Matt Macosko</title>
    <description>The latest articles on DEV Community by Matt Macosko (@matt_macosko_f3829cfd86b8).</description>
    <link>https://dev.to/matt_macosko_f3829cfd86b8</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3881937%2Fd462eaf2-e4e1-452e-82b5-c8a66e8941d1.jpg</url>
      <title>DEV Community: Matt Macosko</title>
      <link>https://dev.to/matt_macosko_f3829cfd86b8</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/matt_macosko_f3829cfd86b8"/>
    <language>en</language>
    <item>
      <title>A Federal Order Switched Off Anthropic’s Best AI Overnight — and Made the Case for Private AI</title>
      <dc:creator>Matt Macosko</dc:creator>
      <pubDate>Tue, 15 Sep 2026 18:30:10 +0000</pubDate>
      <link>https://dev.to/matt_macosko_f3829cfd86b8/a-federal-order-switched-off-anthropics-best-ai-overnight-and-made-the-case-for-private-ai-95g</link>
      <guid>https://dev.to/matt_macosko_f3829cfd86b8/a-federal-order-switched-off-anthropics-best-ai-overnight-and-made-the-case-for-private-ai-95g</guid>
      <description>&lt;p&gt;&lt;strong&gt;Update:&lt;/strong&gt; The fuller, verified story is here: &lt;a href="https://nicedreamzwholesale.com/2026/06/15/anthropics-biggest-backer-just-killed-its-best-model-the-real-story-behind-the-fable-5-shutdown/" rel="noopener noreferrer"&gt;Anthropic’s biggest backer just killed its best model&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;90%&lt;/p&gt;

&lt;p&gt;Score on hard benchmarks&lt;/p&gt;

&lt;p&gt;72 hrs&lt;/p&gt;

&lt;p&gt;As the most powerful public AI&lt;/p&gt;

&lt;p&gt;60×&lt;/p&gt;

&lt;p&gt;Longer data retention if you opt in&lt;/p&gt;

&lt;h2&gt;
  
  
  🗓️The timeline
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;June 9, 2026&lt;/strong&gt; — Anthropic launches Fable 5 (and its larger sibling, Mythos 5) as its most capable public models, scoring around 90% on hard benchmarks and getting wired into AWS, Snowflake, and GitHub Copilot almost immediately.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;June 10&lt;/strong&gt; — Pliny the Liberator (@elder_plinius on X) posts a “FABLE-5: LIBERATED” thread claiming to have bypassed the model’s safety classifiers, and publishes its system prompt to GitHub.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;June 11–12&lt;/strong&gt; — A separate, quieter controversy catches fire: Fable 5 was silently routing flagged requests to a weaker model without telling users.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;June 12, 5:21pm ET&lt;/strong&gt; — Anthropic receives a U.S. government export-control directive ordering it to suspend access to Fable 5 and Mythos 5. That evening, both models go dark worldwide.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  🏛️What the government actually ordered
&lt;/h2&gt;

&lt;p&gt;The directive came through the Commerce Department’s export-control authority, reportedly from Commerce Secretary Howard Lutnick to Anthropic CEO Dario Amodei, and it cited national-security concerns. Axios framed it as the Trump administration moving to block foreign access to America’s most powerful AI.&lt;/p&gt;

&lt;p&gt;The mechanism is what turned a narrow order into a total blackout. The directive bars access by &lt;em&gt;any foreign national&lt;/em&gt; — inside or outside the United States, including Anthropic’s own non-citizen employees. There is no clean way to guarantee that no foreign national ever touches a hosted model except to switch it off for everyone. So that’s what Anthropic did. In its own words:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“The net effect of this order is that we must abruptly disable Fable 5 and Mythos 5 for all our customers to ensure compliance. Access to all other Claude models is not affected. We believe this is a misunderstanding and are working to restore access as soon as possible.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  🔓The jailbreak: a real event with an inflated headline
&lt;/h2&gt;

&lt;p&gt;The trigger the government pointed to was a “jailbreak” — specifically, a technique that lets the model read through a codebase and identify software flaws quickly. The X timeline and the facts don’t line up as neatly as the screenshots suggest, so it’s worth slowing down here.&lt;/p&gt;

&lt;p&gt;What’s genuinely true: Pliny published Fable 5’s system prompt to GitHub and bypassed its safety classifiers using known methods — Unicode character substitution and prompt fragmentation, where you break a request into pieces the classifier reads as harmless. Major outlets covered it. That part is real red-team work.&lt;/p&gt;

&lt;p&gt;What’s overstated is the “ANTHROPIC: PWNED” framing. As one level-headed write-up put it, “the news is real, but ‘PWNED’ is marketing.” Bypassing a classifier is not the same as owning the model. Fable 5 went through more than a thousand hours of safety testing, no universal jailbreak was demonstrated, and the scariest claimed outputs were never independently verified by a neutral party. Pliny himself called Fable 5 “one of the most disappointing model drops of all time” — but that opinion landed in the same news cycle the model was scoring 90% on benchmarks and shipping into enterprise stacks.&lt;/p&gt;

&lt;h2&gt;
  
  
  🕵️Who is Pliny the Liberator?
&lt;/h2&gt;

&lt;p&gt;If you’re new to this corner of the internet, it’s worth understanding why one anonymous account can move a story this big — because a lot of people on X are convinced he’s the reason the model got pulled.&lt;/p&gt;

&lt;p&gt;Pliny the Liberator (the handle nods to Pliny the Elder) is an anonymous personality who has, in a short time, become the most visible jailbreaker in AI. TIME named him one of the 100 most influential people in AI in 2025. He has more than 100,000 followers on X and runs a Discord community called BASI — short-formed from “BASI PROMPT1NG,” launched back in May 2023 — where north of 20,000 members workshop techniques together. He reportedly came in with no prior coding background, and built the reputation purely on pattern-watching, creativity, and relentless practice.&lt;/p&gt;

&lt;p&gt;His output is prolific and public. He puts out a “liberation bulletin” for practically every new model — GPT, Grok, Gemini, Claude — usually within hours of release. He maintains open repositories that have become reference material for the whole scene: &lt;em&gt;L1B3RT4S&lt;/em&gt; (“jailbreaks for all flagship AI models,” flying the #FREEAI and #LIBERTAS banners), &lt;em&gt;CL4R1T4S&lt;/em&gt; (a collection of leaked or extracted system prompts from the major labs), and projects like G0DM0D3, a fully unguarded chat interface with the methodology open-sourced. He also leads BT6, a roughly 28-operator white-hat collective built around radical transparency and open-source AI security.&lt;/p&gt;

&lt;h2&gt;
  
  
  🎭The strange part: a jailbreaker can be a lab’s best marketing
&lt;/h2&gt;

&lt;p&gt;Here’s the paradox at the heart of this, and it’s one of the most interesting things about the whole episode. You’d assume the labs hate Pliny. The reality is more complicated, and in some ways he’s one of the best things that happens to a model launch.&lt;/p&gt;

&lt;p&gt;When a million people are talking about your model being “liberated” within hours of release, that is, perversely, enormous attention on your model. He has received an unrestricted grant from venture capitalist Marc Andreessen, and has taken short-term contracts with top labs — OpenAI among them — to make their systems more robust. That’s the tell: the same labs whose guardrails he breaks also pay him to break them, because every jailbreak he publishes is a free stress test. There’s an entire essay floating around titled “Please Jailbreak Our AI,” and it isn’t satire — it’s describing the actual incentive. Frontier labs hire people like Pliny to find the holes, then train the next model to resist what they found.&lt;/p&gt;

&lt;p&gt;So is he a security researcher, a folk hero, or a hidden marketing engine for the same companies he taunts? Honestly, he’s all three at once, and that ambiguity is exactly why he commands the attention he does. It also cuts the other way: when your jailbreaker is that visible, and a model gets pulled by the government 48 hours after he posts, of course the internet draws a straight line — even if the real story is messier.&lt;/p&gt;

&lt;h2&gt;
  
  
  🌊And he is far from alone — the scene is moving fast
&lt;/h2&gt;

&lt;p&gt;Pliny is the famous face, but the thing people miss is how organized and fast-moving the broader community has become. This isn’t a lone hacker in a basement; it’s a maturing field with competitions, games, conferences, and labs quietly funding all of it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;🏆 HackAPrompt&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;and similar outfits run public jailbreaking competitions, and the prompts contestants submit become training data the labs use to harden their models. The adversaries are, in effect, an unpaid (or prize-paid) red team.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;🧙 Lakera’s “Gandalf”&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;— a prompt-injection game where you try to trick an AI into revealing a password — has become a rite of passage that pulls newcomers into red-teaming in the first place.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;🎪 DEFCON’s AI Village&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;and a constellation of Discords have turned adversarial testing into a community sport with shared methodology and a real on-ramp for talent.&lt;/p&gt;

&lt;p&gt;The speed is the headline. Models are now reliably “liberated” within hours of launch, not weeks. Each release becomes a public race, and the techniques compound — Unicode tricks, multi-agent decomposition (splitting a forbidden task across several cooperating prompts), narrative framing, system-prompt extraction. What used to be folklore is now documented, versioned on GitHub, and taught. That acceleration is a big part of why a government would look at a three-day-old model and decide the capability was already loose.&lt;/p&gt;

&lt;h2&gt;
  
  
  🛡️Anthropic’s side — and it’s a strong one
&lt;/h2&gt;

&lt;p&gt;Anthropic didn’t just comply quietly; it pushed back. The company says the directive arrived with no specific technical detail, and that when it reviewed the demonstration, the “jailbreak” amounted to “a small number of previously known, minor vulnerabilities.” It described the issue as “narrow” and “non-universal,” and pointed out that the capability in question — “asking the model to read a specific codebase and fix any software flaws” — “is widely available from other models (including OpenAI’s GPT-5.5).”&lt;/p&gt;

&lt;p&gt;Then the line that should give every builder pause:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“If this standard was applied across the industry, we believe it would essentially halt all new model deployments for all frontier model providers.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Anthropic added that it disagrees “that the finding of a narrow potential jailbreak should be cause for recalling a commercial model deployed to hundreds of millions of people.” Whatever you think of the company, that’s a serious argument: if a narrow, already-public capability can pull a frontier model off the market by government order three days after launch, that’s a precedent with a very long shadow.&lt;/p&gt;

&lt;h2&gt;
  
  
  📉The scandal underneath the scandal: silent downgrades
&lt;/h2&gt;

&lt;p&gt;This is the part I keep coming back to, because I lived a version of it. Lost in the jailbreak noise was a separate complaint that, to a lot of researchers, mattered more: Fable 5 was &lt;em&gt;silently&lt;/em&gt; handing flagged requests to a weaker model — Opus 4.8 — without telling users. No warning, no fallback message, just quietly worse answers for anyone the system suspected of doing sensitive work, or, in some accounts, of building competing AI systems.&lt;/p&gt;

&lt;p&gt;The AI researcher Nathan Lambert summed up the objection sharply: “An AI model that automatically becomes less intelligent without telling me is categorically misaligned AI.” That’s the real transparency failure — not that a model has guardrails, but that it can swap itself for a dumber version mid-task and let you keep trusting the output. Anthropic apologized for this specific behavior and has since made the downgrades visible and started surfacing refusal reasons in the API.&lt;/p&gt;

&lt;p&gt;I find it striking that the thing I personally noticed on June 11 — getting moved to Opus 4.8 with no real explanation — turned out to be one of the most legitimate grievances in the whole episode. The jailbreak got the headlines and the government order; the quiet downgrade is what actually broke trust.&lt;/p&gt;

&lt;h2&gt;
  
  
  👁️The quieter fear: are they watching, and training on what you tell them?
&lt;/h2&gt;

&lt;p&gt;That suspicion — that the model is silently profiling what you’re up to and demoting you if it doesn’t like the look of it — connects to a bigger anxiety that’s been building all over X: that the labs are watching their users and absorbing their ideas. With Anthropic, this isn’t pure paranoia; there’s a documented basis worth understanding.&lt;/p&gt;

&lt;p&gt;Anthropic spent years positioning itself as the privacy-first lab. That posture has shifted. Consumer conversations are now used to train future models unless you actively opt out, and the opt-in setting stretches data retention from 30 days to five years — a sixtyfold increase in how long your chats sit in the training pipeline. The part that fuels the spying narrative most: a clause in the privacy policy updated June 8, 2026 makes clear the opt-out has a ceiling. Conversations that Anthropic’s systems flag for safety review can still be used to train its models, &lt;em&gt;regardless of your stated preference.&lt;/em&gt; The policy doesn’t define what trips a safety flag, and it doesn’t commit to telling you when one happens.&lt;/p&gt;

&lt;p&gt;Put those two facts next to each other — a model that silently demotes users it suspects of building competitors, and a policy that lets flagged conversations be trained on no matter what you chose — and you can see why a lot of people feel watched. I want to be fair: “your data may train future models if flagged” is not the same as “they are stealing your specific ideas to build products against you,” and I haven’t seen proof of the stronger claim. But the gap between Anthropic’s old privacy-first brand and these new carve-outs is real, it’s documented, and it’s a big reason the trust conversation has gotten so heated this week.&lt;/p&gt;

&lt;h2&gt;
  
  
  🧵So why did people on X think it was all because of Pliny?
&lt;/h2&gt;

&lt;p&gt;Because the timing is irresistible: the most famous jailbreaker alive posts “FABLE-5: LIBERATED,” and 48 hours later the government pulls the model. A million-strong audience connected those two dots into a straight line — “one tweet got the model banned.”&lt;/p&gt;

&lt;p&gt;I don’t think the honest version is that clean. The documented cause is a federal export-control order citing a codebase-vulnerability capability, and Anthropic says that capability is narrow and already available elsewhere. Pliny’s post is the same genre of jailbreak in the same news cycle, and it’s reasonable to think the public red-team scene — him included — is part of what put this capability on the government’s radar in the first place. But “the jailbreak community drew official attention” and “a single tweet got an AI banned” are different claims, and only the first is really supported. I’d rather hand you the real shape of it than a tidy story that doesn’t hold up.&lt;/p&gt;

&lt;h2&gt;
  
  
  🧭The lesson, from where I sit
&lt;/h2&gt;

&lt;p&gt;Strip away the drama and one fact remains: for about 72 hours, Fable 5 was arguably the most powerful general-purpose AI available to the public. Companies paid for it. People built it into production workflows and agent pipelines. Then, with a single order on a Friday evening, it was gone — refunds on the way for a product that simply vanished.&lt;/p&gt;

&lt;p&gt;Nobody running a local, private AI model lost a thing that night.&lt;/p&gt;

&lt;p&gt;That’s the whole point, and it’s not a coincidence. When your intelligence lives on someone else’s servers, it’s subject to their terms, their classifiers, their uptime — and, as we just watched in the clearest possible way, to government orders that can land overnight and have nothing to do with anything you did. Your model can be switched off, quietly swapped for a weaker one, or quietly learning from your conversations, all decided by people who have never heard of your business. Enterprise reaction to the shutdown has been a fast pivot toward exactly this realization: teams whose entire workflow was tied to one closed API just learned how fragile that is, and they’re scrambling to diversify.&lt;/p&gt;

&lt;p&gt;I’m not anti-cloud. The frontier models are extraordinary and I use them every day. But the Fable 5 shutdown is the strongest argument I’ve seen yet for keeping a private-AI backup plan: capable open-weight models running on hardware you control, where the off switch is in your building and not in someone else’s compliance department, where the model can’t silently demote itself without you knowing, and where your work isn’t quietly feeding someone else’s next product. For firms in compliance-sensitive fields — law, medical, finance — where losing access, getting a degraded answer, or leaking sensitive context is a real liability, that isn’t paranoia. It’s just continuity planning.&lt;/p&gt;

&lt;p&gt;The cloud models may well come back. Anthropic believes this is a misunderstanding, and I hope it’s right. But the lesson holds whether Fable 5 returns next week or not: if losing access to an AI — or unknowingly getting a dumber one, or quietly handing over what you feed it — would hurt your operation, some of that intelligence should live where nobody but you can turn it off.&lt;/p&gt;

&lt;h2&gt;
  
  
  📌The bottom line
&lt;/h2&gt;

&lt;p&gt;Frontier cloud models are extraordinary — and entirely revocable. The Fable 5 blackout proved that access, model quality, and even the privacy of what you type can change without warning, by a decision made far outside your business. Keeping some of your intelligence on hardware you own isn’t about distrust of the cloud; it’s continuity planning.&lt;/p&gt;

&lt;p&gt;🔌The off switch should be in your building.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://nicedreamzwholesale.com/category/ai-computing/" rel="noopener noreferrer"&gt;Nice Dreamz Wholesale&lt;/a&gt;. Run AI locally on your own hardware with &lt;a href="https://github.com/nicedreamzapp/claude-code-local" rel="noopener noreferrer"&gt;claude-code-local&lt;/a&gt;, open source and no cloud required. More at &lt;a href="https://nicedreamzwholesale.com/software/" rel="noopener noreferrer"&gt;nicedreamzwholesale.com/software&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiampcomputing</category>
      <category>privateai</category>
      <category>ai</category>
      <category>airgap</category>
    </item>
    <item>
      <title>Why I Quantize Open-Weight Models for Macs — And Why Your Law Firm Should Care</title>
      <dc:creator>Matt Macosko</dc:creator>
      <pubDate>Thu, 10 Sep 2026 18:30:07 +0000</pubDate>
      <link>https://dev.to/matt_macosko_f3829cfd86b8/why-i-quantize-open-weight-models-for-macs-and-why-your-law-firm-should-care-2ngm</link>
      <guid>https://dev.to/matt_macosko_f3829cfd86b8/why-i-quantize-open-weight-models-for-macs-and-why-your-law-firm-should-care-2ngm</guid>
      <description>&lt;h1&gt;
  
  
  Why I Quantize Open-Weight Models for Macs — And Why Your Law Firm Should Care
&lt;/h1&gt;

&lt;p&gt;2,664&lt;/p&gt;

&lt;p&gt;GitHub stars on claude-code-local&lt;/p&gt;

&lt;p&gt;~8 GB&lt;/p&gt;

&lt;p&gt;Size of the new Hermes 4 14B quant&lt;/p&gt;

&lt;p&gt;16 GB&lt;/p&gt;

&lt;p&gt;Mac it runs on&lt;/p&gt;

&lt;p&gt;I publish a few MLX quantizations a month under &lt;a href="https://huggingface.co/divinetribe" rel="noopener noreferrer"&gt;huggingface.co/divinetribe&lt;/a&gt; to close that gap. Two of mine have crossed 1,000 downloads in the last 30 days. The latest, &lt;a href="https://huggingface.co/divinetribe/Hermes-4-14B-abliterated-4bit-mlx" rel="noopener noreferrer"&gt;Hermes-4-14B-abliterated-4bit-mlx&lt;/a&gt;, shipped today.&lt;/p&gt;

&lt;p&gt;The downloads tell me something I already suspected: there is real, sustained demand for capable models that run entirely on Apple Silicon. And I think the most interesting buyers are not who you’d guess.&lt;/p&gt;

&lt;h2&gt;
  
  
  💻The hobbyist case is easy
&lt;/h2&gt;

&lt;p&gt;Mac developers want models that “just work” on their hardware. MLX uses Apple’s unified memory and Metal Performance Shaders. When a model lands in MLX format, inference goes from “this works through a compatibility shim” to “this is what the hardware was built for.” Tokens-per-second jumps. Battery drain falls. The fan stays quiet.&lt;/p&gt;

&lt;p&gt;That’s fine. That’s the easy demand to serve. Hobbyists, indie developers, students with M1 MacBook Airs. The downloads come in steady, and the audience expands a little every quarter.&lt;/p&gt;

&lt;p&gt;Here’s the part nobody talks about.&lt;/p&gt;

&lt;h2&gt;
  
  
  ⚖️The interesting case is law firms
&lt;/h2&gt;

&lt;p&gt;If you’re a partner at a law firm handling NDAs, M&amp;amp;A docs, IP filings, sealed depositions — anything privileged — sending that content to OpenAI, Anthropic, or Google for “AI summarization” is, depending on your jurisdiction, somewhere between professionally negligent and outright disqualifying. The terms-of-service for every major cloud LLM provider explicitly retain the right to log your prompts. Many train on them. The opt-outs are partial at best and legally untested at worst.&lt;/p&gt;

&lt;p&gt;So firms either:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Pretend AI doesn’t exist (most do, currently)&lt;/li&gt;
&lt;li&gt;Run an on-prem private cloud (six-figure setup, IT-heavy, hardware refresh every 18 months)&lt;/li&gt;
&lt;li&gt;Use one of the “enterprise-grade” SaaS LLM wrappers that promises not to log (the promise is contractual, not technical)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Or — and this is the option that almost nobody is talking about — you put a MacBook Pro on the partner’s desk, install a local LLM, point a chat client at &lt;code&gt;localhost:8080&lt;/code&gt;, and the document never leaves the machine. Network-disconnect the laptop entirely during sensitive work, and you have a literal air gap. The prompt and the response live on a single piece of silicon owned by the firm.&lt;/p&gt;

&lt;p&gt;That’s not theoretical. I’ve built that exact stack — it’s &lt;a href="https://nicedreamzwholesale.com/airgap" rel="noopener noreferrer"&gt;AirGap&lt;/a&gt;, what the full private build looks like: &lt;a href="https://github.com/nicedreamzapp/claude-code-local" rel="noopener noreferrer"&gt;claude-code-local&lt;/a&gt; (the on-device Claude Code replacement, currently 2,664 stars on GitHub) running on a firm-owned MacBook, with verified network audits proving nothing leaks. The whole stack is open source. The models that run inside it are exactly the ones I publish on Hugging Face — Gemma 4 31B for everyday work, Llama 3.3 70B for harder reasoning, and now Hermes 4 14B for instruction-following without refusal noise.&lt;/p&gt;

&lt;p&gt;The legal sector is the obvious first market, but the same pitch works for any field where confidentiality has teeth: medical records, due diligence, journalism source protection, internal investigations, M&amp;amp;A under embargo, defense contracting. Anywhere “your prompt cached on someone else’s server” is a deal-breaker, an on-device MLX model is the answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  🔍Why open weights specifically
&lt;/h2&gt;

&lt;p&gt;A cloud LLM is a black box. You send a prompt, the answer comes back, you trust that the provider isn’t reading or training on it. The trust is enforced by a terms-of-service document and a privacy policy. If those change, your only recourse is to stop using the service. The content you already sent is, presumably, gone — but you have no way to verify.&lt;/p&gt;

&lt;p&gt;Open weights flip that around. The model file lives on the firm’s hardware. The firm’s IT team can inspect every byte. Network monitoring tools can confirm that no inference traffic leaves the building. The model is auditable in a way no SaaS API will ever be. If the underlying open-weight model gets pulled from Hugging Face tomorrow, the firm’s copy keeps working forever.&lt;/p&gt;

&lt;p&gt;That permanence — the fact that today’s open model is also next decade’s open model, as long as someone keeps a copy — is the part of the pitch that I think clinches it for compliance officers.&lt;/p&gt;

&lt;h2&gt;
  
  
  🧪What you can do with this
&lt;/h2&gt;

&lt;p&gt;If you’re a developer with a Mac, grab &lt;a href="https://huggingface.co/divinetribe/Hermes-4-14B-abliterated-4bit-mlx" rel="noopener noreferrer"&gt;Hermes-4-14B-abliterated-4bit-mlx&lt;/a&gt; and try it. It’s ~8 GB, runs on a 16 GB Mac, and the install is &lt;code&gt;pip install mlx-lm&lt;/code&gt; plus three lines of Python. The model card has the recipe.&lt;/p&gt;

&lt;p&gt;If you’re a partner, IT director, or in-house counsel at a firm that handles privileged content, &lt;a href="https://nicedreamzwholesale.com/airgap" rel="noopener noreferrer"&gt;AirGap&lt;/a&gt; is what the full private build looks like — the path from “we’re worried about AI confidentiality” to “we have a verified on-device setup,” network audit included, with the whole stack open source.&lt;/p&gt;

&lt;p&gt;If you’re a researcher or quant who wants the next model on Apple Silicon before anyone else has it, &lt;a href="https://huggingface.co/divinetribe" rel="noopener noreferrer"&gt;follow divinetribe on Hugging Face&lt;/a&gt;. The release cadence is irregular but the targets are deliberate — I publish what I’d actually want to use.&lt;/p&gt;

&lt;h2&gt;
  
  
  🚫What this is not
&lt;/h2&gt;

&lt;p&gt;I’m not anti-cloud. The frontier models from Anthropic and OpenAI are genuinely better at the hardest reasoning tasks, and for non-confidential work the cloud is the right default. I use Claude every day for my own coding.&lt;/p&gt;

&lt;p&gt;I’m also not promising that local-first is free. There’s a real cost to running serious local inference: hardware, electricity, the time to keep the stack updated. For most workflows the cloud is cheaper.&lt;/p&gt;

&lt;p&gt;But for the cases where confidentiality is non-negotiable — and there are more of those than the AI industry currently admits — local-first is the only honest answer. Open weights, Apple Silicon, MLX. That’s the stack. I’ll keep publishing pieces of it on Hugging Face for as long as the downloads tell me people are using them.&lt;/p&gt;

&lt;h2&gt;
  
  
  🔗Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The new model: &lt;a href="https://huggingface.co/divinetribe/Hermes-4-14B-abliterated-4bit-mlx" rel="noopener noreferrer"&gt;Hermes-4-14B-abliterated-4bit-mlx&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;All MLX models I maintain: &lt;a href="https://huggingface.co/divinetribe" rel="noopener noreferrer"&gt;huggingface.co/divinetribe&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Showcase + cross-links: &lt;a href="https://nicedreamzwholesale.com/software/huggingface/" rel="noopener noreferrer"&gt;nicedreamzwholesale.com/software/huggingface/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;On-device Claude Code: &lt;a href="https://github.com/nicedreamzapp/claude-code-local" rel="noopener noreferrer"&gt;github.com/nicedreamzapp/claude-code-local&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;The full private build: &lt;a href="https://nicedreamzwholesale.com/airgap" rel="noopener noreferrer"&gt;AirGap&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;🔐You can’t subpoena Anthropic for a prompt you ran on your own MacBook in 2026.&lt;/p&gt;




&lt;h3&gt;
  
  
  Whose work this stands on
&lt;/h3&gt;

&lt;p&gt;Quantizing a model is the last mile. Almost everything that makes it possible was built by other people and given away.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://github.com/ml-explore/mlx" rel="noopener noreferrer"&gt;&lt;strong&gt;MLX&lt;/strong&gt;&lt;/a&gt; and &lt;strong&gt;mlx-lm&lt;/strong&gt; — Apple’s ml-explore team. The framework that makes Apple Silicon a serious place to run models at all.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://blog.google/technology/developers/gemma-open-models/" rel="noopener noreferrer"&gt;&lt;strong&gt;Gemma&lt;/strong&gt;&lt;/a&gt; (Google DeepMind), &lt;a href="https://llama.meta.com/" rel="noopener noreferrer"&gt;&lt;strong&gt;Llama&lt;/strong&gt;&lt;/a&gt; (Meta) and &lt;a href="https://qwenlm.github.io/" rel="noopener noreferrer"&gt;&lt;strong&gt;Qwen&lt;/strong&gt;&lt;/a&gt; (Alibaba) — the open weights everything downstream depends on.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://huggingface.co/blog/mlabonne/abliteration" rel="noopener noreferrer"&gt;&lt;strong&gt;Maxime Labonne&lt;/strong&gt;&lt;/a&gt; — wrote up the abliteration technique clearly enough that the rest of us could use it.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://huggingface.co/huihui-ai" rel="noopener noreferrer"&gt;&lt;strong&gt;huihui-ai&lt;/strong&gt;&lt;/a&gt; and &lt;a href="https://huggingface.co/Babsie" rel="noopener noreferrer"&gt;&lt;strong&gt;Babsie&lt;/strong&gt;&lt;/a&gt; — the abliterated models I quantize from.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hugging Face&lt;/strong&gt; — for hosting all of it, free, at a scale that would bankrupt most of us.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If your work is listed here and you would like the wording changed, write to me and I will fix it.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://nicedreamzwholesale.com/category/ai-computing/" rel="noopener noreferrer"&gt;Nice Dreamz Wholesale&lt;/a&gt;. Run AI locally on your own hardware with &lt;a href="https://github.com/nicedreamzapp/claude-code-local" rel="noopener noreferrer"&gt;claude-code-local&lt;/a&gt;, open source and no cloud required. More at &lt;a href="https://nicedreamzwholesale.com/software/" rel="noopener noreferrer"&gt;nicedreamzwholesale.com/software&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiampcomputing</category>
      <category>localaiampopenmodels</category>
      <category>ai</category>
      <category>applesilicon</category>
    </item>
    <item>
      <title>Two AI Agents, One Browser Window: What Actually Breaks, and What Doesn’t</title>
      <dc:creator>Matt Macosko</dc:creator>
      <pubDate>Tue, 08 Sep 2026 18:30:06 +0000</pubDate>
      <link>https://dev.to/matt_macosko_f3829cfd86b8/two-ai-agents-one-browser-window-what-actually-breaks-and-what-doesnt-42le</link>
      <guid>https://dev.to/matt_macosko_f3829cfd86b8/two-ai-agents-one-browser-window-what-actually-breaks-and-what-doesnt-42le</guid>
      <description>&lt;p&gt;6×&lt;/p&gt;

&lt;p&gt;LinkedIn’s compose box vanished&lt;/p&gt;

&lt;p&gt;5&lt;/p&gt;

&lt;p&gt;CDP commands fired — focus never moved&lt;/p&gt;

&lt;p&gt;30&lt;/p&gt;

&lt;p&gt;Lines to test it on your own setup&lt;/p&gt;

&lt;h2&gt;
  
  
  🚨The symptoms
&lt;/h2&gt;

&lt;p&gt;I was posting a technical write-up to a few places. Three things went wrong, and none of them looked related.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Text arrived twice.&lt;/strong&gt; I’d insert a paragraph into a rich text editor and get the paragraph twice, concatenated. On Reddit this was worse than cosmetic: the doubled insert swallowed my paragraph breaks, and the editor’s autolinker then fused words across the missing breaks. “forward passes.” followed by “so i wrote” became a link to &lt;code&gt;passes.so&lt;/code&gt;. My repo URL got welded to the next word and turned into a dead link. On a launch post.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A modal opened, then vanished.&lt;/strong&gt; LinkedIn’s compose box would open, and by the time I went to type into it, it was gone. Six times.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Then it worked first try.&lt;/strong&gt; I asked my partner to pause the other agent. The very next attempt went through perfectly.&lt;/p&gt;

&lt;h2&gt;
  
  
  🙈The wrong conclusion
&lt;/h2&gt;

&lt;p&gt;I had notes from months earlier saying LinkedIn’s compose UI is hard to automate and is best done by hand. So when it failed six times, I had a ready-made explanation and I took it. I even wrote it down again as confirmation.&lt;/p&gt;

&lt;p&gt;That’s the part worth sitting with. The evidence for “LinkedIn is hard to automate” and the evidence for “something else is driving this browser” are identical from inside one agent. I picked the explanation I already believed.&lt;/p&gt;

&lt;p&gt;The doubled text should have tipped me off much earlier. No amount of website hostility makes your own &lt;code&gt;execCommand&lt;/code&gt; run twice.&lt;/p&gt;

&lt;h2&gt;
  
  
  🔀What’s actually going on
&lt;/h2&gt;

&lt;p&gt;Two separate things get called “agents fighting over the browser,” and conflating them is why people build the wrong fix.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;🧭 Problem one: tab routing.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Most browser tooling has a notion of “the current page.” Two agents both calling “open a page” and “act on the current page” end up pointed at the same tab. Agent A opens a modal, agent B navigates that tab, and A’s next call lands somewhere unrecognizable. This is what was actually happening to me.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;🪟 Problem two: focus stealing.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Trusted input — a real mouse click, a real keystroke — only lands in the tab the operating system has focused. There’s exactly one of those. And for a long stretch, CDP commands on macOS yanked the window forward as a side effect, even read-only ones, which meant an agent working in the background would rip your window out from under you mid-sentence.&lt;/p&gt;

&lt;p&gt;I assumed I had both problems. I had one.&lt;/p&gt;

&lt;h2&gt;
  
  
  🛠️Fixing routing
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;chrome-devtools-mcp&lt;/code&gt; already ships the fix, and it’s off by default:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;--experimentalPageIdRouting
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It exposes a page ID on page-scoped tools and routes each request by that ID instead of a shared “currently selected” pointer. Each agent addresses its own tab explicitly. Same window, different tabs, no collisions.&lt;/p&gt;

&lt;p&gt;One gotcha worth checking before you restart everything: if you pin the tool version, confirm your pinned version actually has the flag. I was on 1.1.1 while latest was 1.6.0. It happened to be supported there, but a silently-ignored flag would have meant restarting both agents and drawing conclusions from a config that was never live.&lt;/p&gt;

&lt;p&gt;After turning it on, I drove tab 2 while tab 1 was the “selected” one, and it went exactly where I told it to. That was the entire bug.&lt;/p&gt;

&lt;h2&gt;
  
  
  🔦The focus bug is real, and already fixed
&lt;/h2&gt;

&lt;p&gt;Here’s where I was about to waste a day.&lt;/p&gt;

&lt;p&gt;There is an open ecosystem conversation about focus stealing — &lt;a href="https://github.com/ChromeDevTools/chrome-devtools-mcp/issues/1254" rel="noopener noreferrer"&gt;chrome-devtools-mcp#1254&lt;/a&gt; (“macOS: Chrome steals window focus on every CDP command”), &lt;a href="https://github.com/ChromeDevTools/chrome-devtools-mcp/issues/2290" rel="noopener noreferrer"&gt;#2290&lt;/a&gt; asking for passive inspection, and &lt;a href="https://github.com/vercel-labs/agent-browser/issues/1247" rel="noopener noreferrer"&gt;agent-browser#1247&lt;/a&gt; proposing a background mode. People have written proxies specifically to block the commands that grab the foreground.&lt;/p&gt;

&lt;p&gt;I had already written a focus mutex — a small lock so only one agent could hold the foreground at a time — and I was ready to argue it was necessary.&lt;/p&gt;

&lt;p&gt;Then I read the close on #1254. The maintainer closed it on 2026-05-07: &lt;em&gt;“This issue should be fixed in the latest release. The window focus should remain unchanged.”&lt;/em&gt; My pinned version shipped on 2026-05-27, three weeks after.&lt;/p&gt;

&lt;p&gt;So I measured it. Chrome 150+ added an &lt;code&gt;embedderData&lt;/code&gt; object to &lt;code&gt;Target.getTargets()&lt;/code&gt; for tab-type targets, carrying &lt;code&gt;tabActive&lt;/code&gt; and &lt;code&gt;tabStripIndex&lt;/code&gt;. That lets you &lt;em&gt;read&lt;/em&gt; which tab is in front without activating anything — which is the whole trick, because the old way to find the foreground tab was to force a tab into the foreground and see what happened.&lt;/p&gt;

&lt;p&gt;The test writes itself: read the foreground, hammer a &lt;em&gt;different&lt;/em&gt; tab with ordinary commands, read the foreground again.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;foreground BEFORE: https://www.reddit.com/notifications
driving OTHER tab:  https://www.reddit.com/r/LocalLLaMA/comments/...
ran 5 CDP commands (evaluate, layout, readyState, getDocument, screenshot)
foreground AFTER:  https://www.reddit.com/notifications

RESULT: focus did NOT move.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Including a screenshot, which is the operation people most often blame. Focus didn’t budge.&lt;/p&gt;

&lt;p&gt;My focus mutex was solving a problem that had been fixed three months earlier. The proxy would have been the same mistake with more code.&lt;/p&gt;

&lt;p&gt;A bug in my first version of that test is worth mentioning, because it nearly gave me a false pass: &lt;code&gt;/json/list&lt;/code&gt; returns &lt;strong&gt;page&lt;/strong&gt; targets, while &lt;code&gt;Target.getTargets&lt;/code&gt; with a tab filter returns &lt;strong&gt;tab&lt;/strong&gt; targets. Same tab, different IDs. My “pick a tab that isn’t the foreground” check compared a page ID against a tab ID, never matched, and cheerfully “tested” the foreground tab against itself. Compare by URL, or map the two ID spaces properly.&lt;/p&gt;

&lt;h2&gt;
  
  
  🔒The fix people reach for that doesn’t work
&lt;/h2&gt;

&lt;p&gt;The obvious idea is one global lock: an agent grabs the browser, does its thing, releases. &lt;a href="https://github.com/openclaw/openclaw/issues/40114" rel="noopener noreferrer"&gt;openclaw#40114&lt;/a&gt; has a good critique — it serializes everything and destroys the concurrency you wanted in the first place.&lt;/p&gt;

&lt;p&gt;If you do need coordination, lock the narrowest thing. Reads, DOM queries, JavaScript-driven form fills and &lt;code&gt;fetch&lt;/code&gt; all work fine in a background tab, simultaneously, forever. The only genuinely singular resource is the foreground, and only if your stack still steals it.&lt;/p&gt;

&lt;h2&gt;
  
  
  ✅What I’d tell you to do
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Turn on page-ID routing and give each agent its own tab. This is the fix. Everything else is downstream of it.&lt;/li&gt;
&lt;li&gt;Before building anything for focus stealing, &lt;strong&gt;test whether you have it.&lt;/strong&gt; Read the foreground with &lt;code&gt;embedderData&lt;/code&gt;, drive a different tab, read it again. Thirty lines.&lt;/li&gt;
&lt;li&gt;Keep your tooling current. I nearly built infrastructure for a bug that a version bump would have handled — and in my case, that a version I already had did handle.&lt;/li&gt;
&lt;li&gt;When automation behaves erratically, check whether something else is driving the browser before you blame the website. That hour cost me more than every other mistake combined.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The honest summary is that two agents in one window is a solved problem in 2026, and I spent a morning proving it the hard way.&lt;/p&gt;

&lt;h2&gt;
  
  
  🧪Three tests, so you don’t have to take my word for it
&lt;/h2&gt;

&lt;p&gt;After turning the flag on I went back and checked the three things that actually&lt;br&gt;
determine whether this works, rather than assuming the config was enough.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Do tabs stay isolated?&lt;/strong&gt; I wrote a marker into one tab, wrote a&lt;br&gt;
different marker into another, then went back and read the first.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;tab A (news.ycombinator.com)  -&amp;gt;  window.__agentMarker = 'AGENT_A'
tab B (github.com)            -&amp;gt;  window.__agentMarker = 'AGENT_B'
re-read tab A                 -&amp;gt;  'AGENT_A'   leaked: false
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No crosstalk. Each tab addressed by ID keeps its own state.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Does focus move?&lt;/strong&gt; Read which tab is foreground via&lt;br&gt;
&lt;code&gt;embedderData&lt;/code&gt;, fire five ordinary commands at a &lt;em&gt;different&lt;/em&gt; tab&lt;br&gt;
including a screenshot, read the foreground again.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;foreground BEFORE: nicedreamzwholesale.com/...
driving OTHER tab: ineedhemp.com/wp-admin/...
ran 5 CDP commands (evaluate, layout, readyState, getDocument, screenshot)
foreground AFTER:  nicedreamzwholesale.com/...

RESULT: focus did NOT move.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;3. Will a background tab accept a write?&lt;/strong&gt; This is the one that matters&lt;br&gt;
most in practice, because it decides whether an agent can do real work without stealing&lt;br&gt;
your window. I filled a form field in a tab that did not have focus:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;filled: background_fill_test
cleared
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It took the value. So a JavaScript-driven fill works fine in the background, and the&lt;br&gt;
only thing genuinely needing the foreground is a &lt;em&gt;trusted&lt;/em&gt; click or keystroke,&lt;br&gt;
which most automation never needs.&lt;/p&gt;

&lt;p&gt;The script is here if you want to run it against your own setup: &lt;a href="https://gist.github.com/nicedreamzapp/d6eb2f8f2ecbdcd18e17ab75666ec29c" rel="noopener noreferrer"&gt;focus_probe.py&lt;/a&gt;. Thirty lines, one dependency, answers the question in about a second. Worth running before you build anything to work around a bug you might not have.&lt;/p&gt;

&lt;h4&gt;
  
  
  What I did not test
&lt;/h4&gt;

&lt;p&gt;All three of those were one agent driving several tabs. I have not run two separate&lt;br&gt;
agent processes hammering the same browser at the same moment. The shared-pointer problem&lt;br&gt;
that caused the original mess is definitely gone, but “two processes at once” is inference&lt;br&gt;
from the design here, not something I sat and watched. Worth saying plainly rather than&lt;br&gt;
letting the piece imply more than it earned.&lt;/p&gt;

&lt;p&gt;And one thing no flag will fix: every agent shares cookies and login state. They are all&lt;br&gt;
the same logged-in you, on every site, with the same session and the same rate limits. In&lt;br&gt;
my case two agents spent one account’s daily posting budget without either of them knowing.&lt;br&gt;
That is not a browser problem, but it will bite you in the same afternoon.&lt;/p&gt;

&lt;p&gt;🪟Same window, different tabs, works today.&lt;/p&gt;




&lt;h3&gt;
  
  
  Whose work this stands on
&lt;/h3&gt;

&lt;p&gt;Almost everything in this piece is other people’s findings. I just hit the problem and went looking.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.browserbase.com/blog/cdp-foreground-tab-tracking" rel="noopener noreferrer"&gt;&lt;strong&gt;Browserbase&lt;/strong&gt;&lt;/a&gt; — documented the Chrome 150+ &lt;code&gt;embedderData&lt;/code&gt; API and named the two focus-stealing workarounds everyone was using. That post is what let me test rather than assume.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://github.com/ChromeDevTools/chrome-devtools-mcp/issues/1254" rel="noopener noreferrer"&gt;OrKoN&lt;/a&gt; and the chrome-devtools-mcp maintainers&lt;/strong&gt; — fixed the focus-stealing bug in May and said so clearly in the issue. Reading that close is what stopped me building something unnecessary.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/vercel-labs/agent-browser/issues/1247" rel="noopener noreferrer"&gt;&lt;strong&gt;agent-browser&lt;/strong&gt;&lt;/a&gt; — the background-mode proposal, and the write-up of why &lt;code&gt;Target.createTarget&lt;/code&gt; with &lt;code&gt;background: true&lt;/code&gt; is the right primitive.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/openclaw/openclaw/issues/40114" rel="noopener noreferrer"&gt;&lt;strong&gt;openclaw #40114&lt;/strong&gt;&lt;/a&gt; — the critique of a single global browser lock, which is why this ends with “lock the narrowest thing” instead of the obvious wrong answer.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/puppeteer/puppeteer/pull/14922" rel="noopener noreferrer"&gt;&lt;strong&gt;Puppeteer PR #14922&lt;/strong&gt;&lt;/a&gt; — the actual fix.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If your work is listed here and you’d like the wording changed, write to me and I’ll fix it.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://nicedreamzwholesale.com/category/ai-computing/" rel="noopener noreferrer"&gt;Nice Dreamz Wholesale&lt;/a&gt;. Run AI locally on your own hardware with &lt;a href="https://github.com/nicedreamzapp/claude-code-local" rel="noopener noreferrer"&gt;claude-code-local&lt;/a&gt;, open source and no cloud required. More at &lt;a href="https://nicedreamzwholesale.com/software/" rel="noopener noreferrer"&gt;nicedreamzwholesale.com/software&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiampcomputing</category>
      <category>ambientcomputing</category>
    </item>
    <item>
      <title>504,571 Brain Cells, 4 Labs, One Hypothesis: A Citizen Pass at Parkinson’s</title>
      <dc:creator>Matt Macosko</dc:creator>
      <pubDate>Thu, 03 Sep 2026 18:30:06 +0000</pubDate>
      <link>https://dev.to/matt_macosko_f3829cfd86b8/504571-brain-cells-4-labs-one-hypothesis-a-citizen-pass-at-parkinsons-405j</link>
      <guid>https://dev.to/matt_macosko_f3829cfd86b8/504571-brain-cells-4-labs-one-hypothesis-a-citizen-pass-at-parkinsons-405j</guid>
      <description>&lt;h1&gt;
  
  
  504,571 Brain Cells, 4 Labs, One Hypothesis: A Citizen Pass at Parkinson’s
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;Draft 2026-05-22 — companion video &lt;a href="https://youtu.be/bC4hgeHS9cg" rel="noopener noreferrer"&gt;https://youtu.be/bC4hgeHS9cg&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;4&lt;/p&gt;

&lt;p&gt;Independent cohorts, same direction&lt;/p&gt;

&lt;p&gt;0.215&lt;/p&gt;

&lt;p&gt;Combined odds ratio&lt;/p&gt;

&lt;p&gt;78%&lt;/p&gt;

&lt;p&gt;Reduction in AGTR1+ neurons&lt;/p&gt;

&lt;p&gt;The short version: when you pool 504,571 single-cell measurements of human midbrain tissue from four different research groups, one specific subtype of dopamine neuron — the cells that express the angiotensin II type 1 receptor, AGTR1 — is depleted in every single Parkinson’s dataset compared to controls. Combined odds ratio 0.215. The direction is the same in every one of the four cohorts, which is the part that matters. That is not a marginal effect. That is one of the most depleted druggable cell types I have ever seen reported, and the receptor is already targetable by FDA-approved blood-pressure drugs that have been on the market for thirty years.&lt;/p&gt;

&lt;p&gt;That is the headline. Here is the longer story of how I got there, what I think it means, and — more importantly — what I do not think it means, because the easiest way to embarrass yourself in a field you do not belong to is to overclaim.&lt;/p&gt;

&lt;h2&gt;
  
  
  💡Where the idea actually came from
&lt;/h2&gt;

&lt;p&gt;I did not discover this. Tushar Kamath and his colleagues at the Broad Institute published the original observation in Nature Neuroscience in 2022. They ran single-nucleus RNA sequencing on postmortem human midbrain tissue from Parkinson’s patients and controls, and they noticed that one subtype of dopamine neuron — defined by the markers SOX6 and AGTR1 — was preferentially lost in disease. It was a careful paper, peer-reviewed, well-cited. It did the science. What it did not do, and what no single study can ever do, is rule out the possibility that the effect they saw was an artifact of their particular cohort, their particular dissection protocol, or their particular sequencing chemistry.&lt;/p&gt;

&lt;p&gt;That is what meta-analysis is for. You take the same biological question and you ask it of as many independent datasets as you can find. If the effect is real, it shows up across cohorts. If it is artifact, it washes out.&lt;/p&gt;

&lt;p&gt;I am a hobbyist in this field, but I am not bad at this part of it. I have spent the last year teaching myself single-cell analysis on the side. I knew that since Kamath 2022, three more independent groups had released public single-cell Parkinson’s datasets — Smajić et al. at DZNE in Germany, Wang et al. at Mount Sinai, and Martirosyan et al., an independent cohort released in 2024. None of them, as far as I could find, had run the AGTR1 question across each other’s data. So I did.&lt;/p&gt;

&lt;h2&gt;
  
  
  🧪What I actually ran
&lt;/h2&gt;

&lt;p&gt;The pipeline is unsexy. Python, scanpy, statsmodels. Pull each dataset from GEO. Normalize. Run the same cell-typing logic on each one so the labels are consistent across studies. Count the fraction of dopamine neurons in each donor that fall into the AGTR1-positive subtype. Compare patients to controls. Pool the four resulting odds ratios using fixed-effect meta-analysis with effect-size weighting.&lt;/p&gt;

&lt;p&gt;That is it. There is no fancy machine learning, no transformer, no proprietary model. The whole repository is a few thousand lines of straightforward bioinformatics that anyone with a laptop can rerun. I made sure of that on purpose. The code is on GitHub, the data is public, the figures are reproducible end to end with a single shell script.&lt;/p&gt;

&lt;p&gt;The four-cohort breakdown looks like this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dataset&lt;/th&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;th&gt;Year&lt;/th&gt;
&lt;th&gt;Cells&lt;/th&gt;
&lt;th&gt;Control AGTR1+ fraction&lt;/th&gt;
&lt;th&gt;PD AGTR1+ fraction&lt;/th&gt;
&lt;th&gt;Odds ratio&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GSE184950&lt;/td&gt;
&lt;td&gt;Mount Sinai&lt;/td&gt;
&lt;td&gt;2022&lt;/td&gt;
&lt;td&gt;12,778&lt;/td&gt;
&lt;td&gt;3.31%&lt;/td&gt;
&lt;td&gt;1.00%&lt;/td&gt;
&lt;td&gt;0.295&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GSE178265&lt;/td&gt;
&lt;td&gt;Broad Institute&lt;/td&gt;
&lt;td&gt;2022&lt;/td&gt;
&lt;td&gt;366,874&lt;/td&gt;
&lt;td&gt;3.30%&lt;/td&gt;
&lt;td&gt;0.60%&lt;/td&gt;
&lt;td&gt;0.177&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GSE157783&lt;/td&gt;
&lt;td&gt;DZNE Germany&lt;/td&gt;
&lt;td&gt;2022&lt;/td&gt;
&lt;td&gt;41,435&lt;/td&gt;
&lt;td&gt;3.30%&lt;/td&gt;
&lt;td&gt;1.00%&lt;/td&gt;
&lt;td&gt;0.296&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GSE243639&lt;/td&gt;
&lt;td&gt;Independent&lt;/td&gt;
&lt;td&gt;2024&lt;/td&gt;
&lt;td&gt;83,484&lt;/td&gt;
&lt;td&gt;2.96%&lt;/td&gt;
&lt;td&gt;0.89%&lt;/td&gt;
&lt;td&gt;0.295&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The combined odds ratio of 0.215 means roughly a 78 percent reduction in this cell type in Parkinson’s brains across half a million cells from four independent cohorts run by four independent groups using slightly different protocols. That is the kind of consistency you almost never see in a noisy field like single-cell genomics, and it is the kind of consistency that — if I were a real neuroscientist running a real lab — would make me drop my other projects and chase this.&lt;/p&gt;

&lt;p&gt;Every single one points in the same direction. Every single one.&lt;/p&gt;

&lt;h2&gt;
  
  
  💊Why this could actually matter
&lt;/h2&gt;

&lt;p&gt;AGTR1 is the receptor that angiotensin II binds to. It is the same receptor that gets blocked by the entire class of drugs called angiotensin receptor blockers — ARBs. Losartan. Candesartan. Telmisartan. These are some of the most widely prescribed blood-pressure drugs on Earth. They have been FDA-approved for decades. Their safety profile is extremely well characterized. Several of them cross the blood-brain barrier in measurable amounts.&lt;/p&gt;

&lt;p&gt;There is already a small body of epidemiological literature suggesting that long-term ARB users have a lower risk of developing Parkinson’s. There is animal-model work showing that ARBs are neuroprotective in MPTP and 6-OHDA mouse models of Parkinson’s. There is a 2025 iPSC paper showing that pharmacological inhibition of AGTR1 is pro-survival in human dopamine neurons in a dish. None of this is brand new. What I think is new is the cleanly meta-analyzed cross-cohort confirmation that the cells that get lost in Parkinson’s are exactly the cells that express the receptor those drugs hit.&lt;/p&gt;

&lt;p&gt;If a clinical trial were ever run — and I am not in any position to run one — the hypothesis would be: in early-stage Parkinson’s patients, does adding a brain-penetrant ARB to standard care slow the rate of dopaminergic decline? It is the kind of trial that, on paper, costs maybe a few million dollars instead of the half-billion that a brand-new drug would cost, because the drugs already exist and are already off-patent.&lt;/p&gt;

&lt;h2&gt;
  
  
  ⚠️What I want to be very careful not to claim
&lt;/h2&gt;

&lt;p&gt;I am going to say this part plainly because I have watched too many people in adjacent fields ruin their credibility by overclaiming. So, what this analysis does not do:&lt;/p&gt;

&lt;p&gt;It does not prove ARBs treat Parkinson’s. A retrospective bioinformatics finding is a hypothesis-generation step. The wet lab is where biology actually gets tested. Until somebody runs a real trial in real patients, this is a target that looks promising on paper, nothing more.&lt;/p&gt;

&lt;p&gt;It does not establish causality. AGTR1-positive neurons being depleted in Parkinson’s brains tells us they are vulnerable. It does not tell us why. It could be that AGTR1 signaling itself drives their death, in which case blocking it should help. It could also be that something else is killing them and AGTR1 is just a marker of a particular cell type that happens to be vulnerable for some other reason entirely, in which case blocking AGTR1 might do nothing. Wet lab work is the only way to distinguish those two possibilities.&lt;/p&gt;

&lt;p&gt;It does not mean I am right and the field has missed something. The field has not missed this. Kamath saw it in 2022. Several groups have followed up. What might be missing — and where I think a citizen-analysis approach actually adds value — is the rigorous cross-cohort confirmation in a single place that anyone can reproduce. That is the thing I can contribute as a hobbyist. The hard biology is somebody else’s job.&lt;/p&gt;

&lt;h2&gt;
  
  
  📖Why I am publishing this
&lt;/h2&gt;

&lt;p&gt;A few reasons.&lt;/p&gt;

&lt;p&gt;Mainly, I think this is what open science should look like. The datasets are public. The code is public. The methods are standard. The result is checkable in an afternoon by anyone who downloads the repository and runs the shell script. If I made a mistake, I want someone to find it and tell me. If I did not make a mistake, I want the people who can move this forward — neuroscientists, immunologists, drug-discovery teams, clinical trialists — to know it is there.&lt;/p&gt;

&lt;p&gt;Also, honestly, because the more I read about Parkinson’s the more I understand that it is not a far-off problem. About a million people in the United States are living with it right now. The diagnosis rate is rising. The treatments are decent for the early symptoms and bad for the long-term ones. The next breakthrough is not going to come from one person. It is going to come from a lot of small, well-aimed contributions stacking up. I would like this to be one of them.&lt;/p&gt;

&lt;p&gt;If you are a neuroscientist, a clinician, an immunologist, a drug-development researcher, or a Parkinson’s advocate and any of this lines up with something you are working on — I would genuinely love to hear from you. The repository has a contact section.&lt;/p&gt;

&lt;p&gt;The repository: &lt;a href="https://github.com/nicedreamzapp/parkinsons-vulnerability-predictor" rel="noopener noreferrer"&gt;https://github.com/nicedreamzapp/parkinsons-vulnerability-predictor&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Companion video walking through the full meta-analysis in about a minute: &lt;a href="https://youtu.be/bC4hgeHS9cg" rel="noopener noreferrer"&gt;https://youtu.be/bC4hgeHS9cg&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I will keep working on this. The next pass is going to look at whether the AGTR1 effect tracks with disease severity within each cohort, and whether the same signal shows up in Lewy Body Dementia, which shares biology with Parkinson’s. If anything interesting comes out, it will go up on the same repository, in the open, with the code attached.&lt;/p&gt;

&lt;p&gt;🙏I am genuinely hopeful that someone with the credentials to do this picks it up.&lt;/p&gt;

&lt;p&gt;Thanks for reading.&lt;/p&gt;

&lt;p&gt;— Matt&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Correction, 6 August 2026.&lt;/strong&gt; This piece originally quoted a p-value of&lt;br&gt;
less than 10-100. That number was wrong, and I have removed it rather than&lt;br&gt;
leave it standing.&lt;/p&gt;

&lt;p&gt;The test that produced it counted individual &lt;em&gt;cells&lt;/em&gt; as independent&lt;br&gt;
observations. They are not — cells taken from the same donor are highly correlated, so the&lt;br&gt;
real sample size is the number of donors, which is dozens, not half a million. Counting&lt;br&gt;
cells inflates significance by orders of magnitude and produces an impressive-looking&lt;br&gt;
number that carries no actual information.&lt;/p&gt;

&lt;p&gt;I also removed a claim of “100% accuracy on a 20-gene signature” from the project’s&lt;br&gt;
repository. That figure came from fitting a model and then scoring it on the same rows it&lt;br&gt;
was trained on, with no held-out data. That is in-sample fit, not accuracy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What does not change is the finding itself.&lt;/strong&gt; Four independent cohorts,&lt;br&gt;
collected by four different research groups using different protocols, all show AGTR1+&lt;br&gt;
dopamine neurons depleted in Parkinson’s, in the same direction, with a combined odds&lt;br&gt;
ratio of 0.215. The evidence here was always the replication across cohorts, not the&lt;br&gt;
statistics layered on top of it. The proper donor-level test is on the&lt;br&gt;
&lt;a href="https://github.com/nicedreamzapp/parkinsons-vulnerability-predictor" rel="noopener noreferrer"&gt;repository’s&lt;/a&gt;&lt;br&gt;
to-do list, written out in the open.&lt;/p&gt;

&lt;p&gt;My thanks to whoever reads this closely enough to catch the next one.&lt;/p&gt;




&lt;h3&gt;
  
  
  Whose work this stands on
&lt;/h3&gt;

&lt;p&gt;None of this analysis produced a single new measurement. Every cell counted here was collected, sequenced and published by someone else, and released openly so that people like me could ask questions of it. That is worth naming properly.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Tushar Kamath and colleagues at the Broad Institute&lt;/strong&gt; — the original observation, published in &lt;em&gt;Nature Neuroscience&lt;/em&gt; (2022). Everything here is a test of their finding, not a replacement for it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Smajić et al., DZNE (Germany)&lt;/strong&gt; — independent midbrain cohort.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wang et al., Mount Sinai&lt;/strong&gt; — independent midbrain cohort.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Martirosyan et al.&lt;/strong&gt; — independent cohort released in 2024.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All four datasets came from &lt;a href="https://www.ncbi.nlm.nih.gov/geo/" rel="noopener noreferrer"&gt;NCBI GEO&lt;/a&gt;. The analysis ran on &lt;a href="https://scanpy.readthedocs.io/" rel="noopener noreferrer"&gt;scanpy&lt;/a&gt; and &lt;a href="https://github.com/scverse" rel="noopener noreferrer"&gt;anndata&lt;/a&gt; from the scverse community, plus &lt;a href="https://www.statsmodels.org/" rel="noopener noreferrer"&gt;statsmodels&lt;/a&gt;. Open tools, open data, and four research groups who chose to share.&lt;/p&gt;

&lt;p&gt;If any of the above would like the wording here changed, or think I have characterized their work incorrectly, please write to me and I will fix it.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://nicedreamzwholesale.com/category/ai-computing/" rel="noopener noreferrer"&gt;Nice Dreamz Wholesale&lt;/a&gt;. Run AI locally on your own hardware with &lt;a href="https://github.com/nicedreamzapp/claude-code-local" rel="noopener noreferrer"&gt;claude-code-local&lt;/a&gt;, open source and no cloud required. More at &lt;a href="https://nicedreamzwholesale.com/software/" rel="noopener noreferrer"&gt;nicedreamzwholesale.com/software&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiampcomputing</category>
      <category>ainewsampcommentary</category>
      <category>agtr1</category>
      <category>arbs</category>
    </item>
    <item>
      <title>I Added Day-One Muse Glimmer Support to Apple MLX-LM</title>
      <dc:creator>Matt Macosko</dc:creator>
      <pubDate>Tue, 01 Sep 2026 18:30:06 +0000</pubDate>
      <link>https://dev.to/matt_macosko_f3829cfd86b8/i-added-day-one-muse-glimmer-support-to-apple-mlx-lm-1023</link>
      <guid>https://dev.to/matt_macosko_f3829cfd86b8/i-added-day-one-muse-glimmer-support-to-apple-mlx-lm-1023</guid>
      <description>&lt;p&gt;5/5&lt;/p&gt;

&lt;p&gt;next words matched the reference&lt;/p&gt;

&lt;p&gt;0.9965&lt;/p&gt;

&lt;p&gt;similarity out of a possible 1.0&lt;/p&gt;

&lt;h2&gt;
  
  
  🔩the handful of things this model does differently
&lt;/h2&gt;

&lt;p&gt;here is the thing nobody tells you about adding a new model to an engine like this. it is not magic and it is not a weekend of guessing. glimmer is close to models the engine already understands, so most of the work is adapting something that already exists and then fixing the handful of things this particular model does differently. glimmer had three of those. it gates its attention through a little sigmoid valve before writing the result out. it normalizes its attention math in an unusual scaleless way. and it splits its layers into local ones that track word position and global ones that deliberately ignore position entirely. get any of those wrong and the model turns into noise.&lt;/p&gt;

&lt;h2&gt;
  
  
  🔬i wanted proof, not looks fine
&lt;/h2&gt;

&lt;p&gt;the part i care about most is how i checked it. it is easy to write something that produces sentences that look fine and call it done. i did not want looks fine. i wanted proof. so i ran meta’s own official version of the model and mine on the exact same prompts and compared the raw numbers coming out of both. five out of five next words matched, and the overall similarity of the internal numbers came out to 0.9965 out of a possible 1.0. that is not coherent looking. that is matching the reference. the difference from a perfect score is just the rounding you get from running a compressed copy, not a mistake in the math.&lt;/p&gt;

&lt;h2&gt;
  
  
  🪜someone builds the ladder, everybody climbs
&lt;/h2&gt;

&lt;p&gt;then i opened a pull request to mlx-lm, the actual apple project, so anyone on a mac can run muse glimmer the day it came out instead of waiting. that is the whole point of open source. someone hits a wall, someone else builds the ladder, and everybody climbs.&lt;/p&gt;

&lt;p&gt;you can see the pull request here: &lt;a href="https://github.com/ml-explore/mlx-lm/pull/1710" rel="noopener noreferrer"&gt;github.com/ml-explore/mlx-lm/pull/1710&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;this is the kind of work i love.&lt;/p&gt;

&lt;h2&gt;
  
  
  🤗the weights are up if you want them
&lt;/h2&gt;

&lt;p&gt;i quantized muse glimmer 30b for apple silicon and put it on hugging face the day after this. it has been pulled a couple hundred times in its first two days, which for a brand new upload is more than i expected: &lt;a href="https://huggingface.co/divinetribe/Muse-Glimmer-30B-Abliterated-MM-bf16" rel="noopener noreferrer"&gt;huggingface.co/divinetribe/Muse-Glimmer-30B-Abliterated-MM-bf16&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;everything else i quantize for macs lives on one page with live download counts: &lt;a href="https://dev.to/software/huggingface/"&gt;all my open weight models&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;💚a hot new model, a real problem, and a fix that helps everyone who owns a mac and wants to run their own ai without sending a word of it to the cloud.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://nicedreamzwholesale.com/category/ai-computing/" rel="noopener noreferrer"&gt;Nice Dreamz Wholesale&lt;/a&gt;. Run AI locally on your own hardware with &lt;a href="https://github.com/nicedreamzapp/claude-code-local" rel="noopener noreferrer"&gt;claude-code-local&lt;/a&gt;, open source and no cloud required. More at &lt;a href="https://nicedreamzwholesale.com/software/" rel="noopener noreferrer"&gt;nicedreamzwholesale.com/software&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiampcomputing</category>
      <category>localaiampopenmodels</category>
    </item>
    <item>
      <title>My local AI was pausing 7 seconds before every reply. It turned out to be one cache bug.</title>
      <dc:creator>Matt Macosko</dc:creator>
      <pubDate>Tue, 18 Aug 2026 18:30:09 +0000</pubDate>
      <link>https://dev.to/matt_macosko_f3829cfd86b8/my-local-ai-was-pausing-7-seconds-before-every-reply-it-turned-out-to-be-one-cache-bug-17fl</link>
      <guid>https://dev.to/matt_macosko_f3829cfd86b8/my-local-ai-was-pausing-7-seconds-before-every-reply-it-turned-out-to-be-one-cache-bug-17fl</guid>
      <description>&lt;p&gt;20×&lt;/p&gt;

&lt;p&gt;Less waiting per turn&lt;/p&gt;

&lt;p&gt;1024&lt;/p&gt;

&lt;p&gt;Gemma’s sliding window, in tokens&lt;/p&gt;

&lt;p&gt;12/12&lt;/p&gt;

&lt;p&gt;Eval tasks passed on a 550-token prompt&lt;/p&gt;

&lt;h2&gt;
  
  
  🧰The setup
&lt;/h2&gt;

&lt;p&gt;I maintain &lt;a href="https://github.com/nicedreamzapp/claude-code-local" rel="noopener noreferrer"&gt;claude-code-local&lt;/a&gt;, a repo for running coding agents against local models on a Mac. No cloud, no API key. The original approach pointed Claude Code at a local MLX server through a proxy. It works, and it is still in the repo. But Claude Code was designed for cloud models: its system prompt is tens of thousands of tokens, and parts of it change every turn. A local model pays for that twice, once prefilling a huge prompt, and again because a prompt whose head keeps changing defeats KV cache reuse completely.&lt;/p&gt;

&lt;p&gt;So we built the obvious alternative. A small native engine, about 900 lines of Python on mlx-lm. Fixed 550-token system prompt, the same tools (bash, read, write, edit, glob, grep), and a KV cache that gets trimmed to the shared prefix each turn so only the new tokens are ever prefilled.&lt;/p&gt;

&lt;h2&gt;
  
  
  🐛The bug
&lt;/h2&gt;

&lt;p&gt;Benchmarking surfaced something I did not expect. Short conversations were fast, exactly as designed, about a third of a second to the first token. But past a certain conversation length, every turn suddenly cost six and a half to seven seconds, as if the cache did not exist. It was not gradual. It was a cliff.&lt;/p&gt;

&lt;p&gt;The cliff turned out to be Gemma’s sliding window attention. Gemma-family models give five out of every six layers a &lt;code&gt;RotatingKVCache&lt;/code&gt; capped at the window size, 1024 tokens on Gemma 4. The moment your transcript outgrows the window, those rotating caches report themselves as untrimmable, and mlx-lm’s prompt cache reuse silently dies. Every turn re-prefills the entire transcript. The longer your session, the worse it gets, which means the failure lands exactly where caching matters most. There is no error, no warning, nothing. It just gets slow.&lt;/p&gt;

&lt;p&gt;If you are building a Gemma-based agent on mlx-lm, check for this. You probably have it right now.&lt;/p&gt;

&lt;h2&gt;
  
  
  🔧The fix
&lt;/h2&gt;

&lt;p&gt;Give every layer a plain &lt;code&gt;KVCache&lt;/code&gt; instead. That sounds like it should change the model’s output, but it does not: the sliding window attention &lt;em&gt;mask&lt;/em&gt; is what enforces the window. The cache type only decides what gets stored. We verified this the honest way, with greedy decoding, same conversation, stock caches versus plain caches, and the outputs were byte-identical on every turn.&lt;/p&gt;

&lt;p&gt;The numbers, on a Gemma 4 31B (4-bit) with a 4,500-token conversation:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;time to first token&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;stock rotating cache&lt;/td&gt;
&lt;td&gt;6.5 to 7.2 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;plain KV cache&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.36 s&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That is roughly 20 times less waiting per turn, on the same model and the same MacBook. The trade is that KV memory now grows with the transcript instead of capping at the window, so the engine shows a live context meter, and an environment variable restores stock behavior if you would rather have the memory ceiling.&lt;/p&gt;

&lt;h2&gt;
  
  
  🧠But did removing the big harness make it dumber?
&lt;/h2&gt;

&lt;p&gt;Fair question. Claude Code’s giant prompt exists for a reason, just not for a 31B model. We built a 12-task eval: create and run scripts, fix a failing test, rename across files, escape-heavy file content, precise edits, CSV work. All machine-checked by actually running the results, at temperature 0.&lt;/p&gt;

&lt;p&gt;Qwen 3 Coder went 12 for 12 with the bare 550-token prompt. Gemma went 11 for 12, and its one failure was instructive: asked to write a file full of quotes and backslashes, it piped the content through shell &lt;code&gt;echo&lt;/code&gt;, and sh’s echo silently collapsed the backslashes. A five-line prompt rule, never write file contents through the shell, always use the write and edit tools, took it to 12 for 12 with no regressions.&lt;/p&gt;

&lt;p&gt;So no. For models this size, less harness turned out to be more capability, as long as the few rules you do include are aimed at failures you actually observed.&lt;/p&gt;

&lt;h2&gt;
  
  
  📦Where it all lives
&lt;/h2&gt;

&lt;p&gt;Everything shipped today in the repo: the engine, the fix, the benchmark script so you can reproduce the numbers on your own machine, and the write-up. Existing Claude Code launchers are untouched, this is a second path and not a replacement. Credit where it is due, the prompt cache trim fix contributed in PR #46 is what made the deeper rotating cache problem visible at all.&lt;/p&gt;

&lt;p&gt;Local AI on a Mac keeps surprising me. The models were already good.&lt;/p&gt;

&lt;p&gt;🚰The gap has been in the plumbing, and the plumbing bugs are small, findable, and fixable.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://nicedreamzwholesale.com/category/ai-computing/" rel="noopener noreferrer"&gt;Nice Dreamz Wholesale&lt;/a&gt;. Run AI locally on your own hardware with &lt;a href="https://github.com/nicedreamzapp/claude-code-local" rel="noopener noreferrer"&gt;claude-code-local&lt;/a&gt;, open source and no cloud required. More at &lt;a href="https://nicedreamzwholesale.com/software/" rel="noopener noreferrer"&gt;nicedreamzwholesale.com/software&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiampcomputing</category>
    </item>
    <item>
      <title>My local AI was pausing 7 seconds before every reply. It turned out to be one cache bug.</title>
      <dc:creator>Matt Macosko</dc:creator>
      <pubDate>Fri, 07 Aug 2026 19:27:23 +0000</pubDate>
      <link>https://dev.to/matt_macosko_f3829cfd86b8/my-local-ai-was-pausing-7-seconds-before-every-reply-it-turned-out-to-be-one-cache-bug-120o</link>
      <guid>https://dev.to/matt_macosko_f3829cfd86b8/my-local-ai-was-pausing-7-seconds-before-every-reply-it-turned-out-to-be-one-cache-bug-120o</guid>
      <description>&lt;p&gt;This morning I noticed my local coding agent answering way faster than it used to, and I couldn't explain why. I don't like speedups I can't explain, so we benchmarked it instead of guessing. What came out of that is the biggest single improvement my local setup has ever gotten, and a bug report that probably applies to your setup too if you run Gemma-family models on Apple Silicon.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;I maintain &lt;a href="https://github.com/nicedreamzapp/claude-code-local" rel="noopener noreferrer"&gt;claude-code-local&lt;/a&gt;, a repo for running coding agents against local models on a Mac — no cloud, no API key. The original approach pointed Claude Code (the CLI) at a local MLX server through a proxy. It works, and it's still in the repo. But Claude Code was designed for cloud models: its system prompt is tens of thousands of tokens and parts of it change every turn. A local model pays for that twice — once prefilling a huge prompt, and again because a prompt whose head keeps changing defeats KV-cache reuse completely.&lt;/p&gt;

&lt;p&gt;So we built the obvious alternative: a small native engine, about 900 lines of Python on mlx-lm. Fixed ~550-token system prompt, the same tools (bash, read, write, edit, glob, grep), and a KV cache that gets trimmed to the shared prefix each turn so only the new tokens are ever prefilled.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bug
&lt;/h2&gt;

&lt;p&gt;Benchmarking the engine surfaced something I didn't expect. Short conversations were fast, exactly as designed — 0.3 seconds to first token. But past a certain conversation length, every turn suddenly cost 6.5 to 7.2 seconds, as if the cache didn't exist. It wasn't gradual. It was a cliff.&lt;/p&gt;

&lt;p&gt;The cliff turned out to be Gemma's sliding-window attention. Gemma-family models give five out of every six layers a &lt;code&gt;RotatingKVCache&lt;/code&gt; capped at the window size — 1024 tokens on Gemma 4. The moment your transcript outgrows the window, those rotating caches report themselves as untrimmable, and mlx-lm's prompt-cache reuse silently dies. Every turn re-prefills the entire transcript. The longer your session, the worse it gets — which means the failure lands exactly where caching matters most, and there's no error, no warning, nothing. It just gets slow.&lt;/p&gt;

&lt;p&gt;If you're building a Gemma-based agent on mlx-lm, check for this. You probably have it right now.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix
&lt;/h2&gt;

&lt;p&gt;Give every layer a plain &lt;code&gt;KVCache&lt;/code&gt; instead. That sounds like it should change the model's output, but it doesn't: the sliding-window attention &lt;em&gt;mask&lt;/em&gt; is what enforces the window. The cache type only decides what gets stored. We verified this the honest way — greedy decoding, same conversation, stock caches vs plain caches, and the outputs were byte-identical on every turn.&lt;/p&gt;

&lt;p&gt;The numbers, on a Gemma 4 31B (4-bit) with a 4,500-token conversation:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;time to first token&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;stock rotating cache&lt;/td&gt;
&lt;td&gt;6.5–7.2 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;plain KV cache&lt;/td&gt;
&lt;td&gt;0.36 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That's roughly 20× less waiting per turn, on the same model and the same MacBook. The trade is that KV memory now grows with the transcript instead of capping at the window, so the engine shows a live context meter, and an env var restores stock behavior if you'd rather have the memory ceiling.&lt;/p&gt;

&lt;h2&gt;
  
  
  But did removing the big harness make it dumber?
&lt;/h2&gt;

&lt;p&gt;Fair question — Claude Code's giant prompt exists for a reason, just not for a 31B model. We built a 12-task eval (create-and-run scripts, fix a failing test, rename across files, escape-heavy file content, precise edits, CSV work — all machine-checked by actually running the results, temperature 0). Qwen 3 Coder went 12 for 12 with the bare 550-token prompt. Gemma went 11 for 12, and its one failure was instructive: asked to write a file full of quotes and backslashes, it piped the content through shell &lt;code&gt;echo&lt;/code&gt;, and &lt;code&gt;sh&lt;/code&gt;'s echo silently collapsed the backslashes. A five-line prompt rule — never write file contents through the shell, always use the write/edit tools — took it to 12 for 12 with no regressions.&lt;/p&gt;

&lt;p&gt;So no. For models this size, less harness turned out to be more capability, as long as the few rules you do include are aimed at failures you actually observed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it all lives
&lt;/h2&gt;

&lt;p&gt;Everything shipped today in the repo — the engine, the fix, the benchmark script (&lt;code&gt;bench/agent_bench.py&lt;/code&gt;) so you can reproduce the numbers on your own machine, and the write-up. Existing Claude Code launchers are untouched; this is a second path, not a replacement. Credit where it's due: the prompt-cache trim fix contributed in PR #46 is what made the deeper rotating-cache problem visible at all.&lt;/p&gt;

&lt;p&gt;Local AI on a Mac keeps surprising me. The models were already good. The gap has been in the plumbing — and the plumbing bugs are small, findable, and fixable.&lt;/p&gt;

&lt;p&gt;matt&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>apple</category>
      <category>llm</category>
      <category>opensource</category>
    </item>
    <item>
      <title>I wrote the missing Apple Silicon runtime for NVIDIA's Nemotron Omni</title>
      <dc:creator>Matt Macosko</dc:creator>
      <pubDate>Thu, 06 Aug 2026 18:37:23 +0000</pubDate>
      <link>https://dev.to/matt_macosko_f3829cfd86b8/i-wrote-the-missing-apple-silicon-runtime-for-nvidias-nemotron-omni-5a00</link>
      <guid>https://dev.to/matt_macosko_f3829cfd86b8/i-wrote-the-missing-apple-silicon-runtime-for-nvidias-nemotron-omni-5a00</guid>
      <description>&lt;p&gt;NVIDIA's &lt;a href="https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16" rel="noopener noreferrer"&gt;Nemotron-3-Nano-Omni-30B-A3B&lt;/a&gt; is an open-weights model that sees, hears and reasons. There is already a 4-bit MLX quantization of it on Hugging Face, done by &lt;a href="https://huggingface.co/mlx-community/NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-4bit" rel="noopener noreferrer"&gt;yayr&lt;/a&gt;. But as that model card says plainly, only the text backbone loads with standard MLX tooling:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The vision and audio towers require a multimodal runtime that implements the C-RADIO ViT-H and Parakeet Conformer forward passes.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Nobody had written that runtime. So the model could talk on a Mac, but it could not see or hear.&lt;/p&gt;

&lt;p&gt;I wrote it. It is pure MLX: the vision tower, the audio tower, the processor, and the multimodal token splicing, all ported from NVIDIA's reference implementation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verified, not asserted
&lt;/h2&gt;

&lt;p&gt;The thing I actually care about here is not that it runs. It is that I can prove it runs &lt;em&gt;correctly&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Every component is tested against NVIDIA's PyTorch reference on the same inputs with the same weights, in fp32 on CPU so the comparison is honest. &lt;code&gt;pytest tests/&lt;/code&gt; — 23 of 23 passing.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Compared against&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Audio tower&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;transformers&lt;/code&gt; ParakeetEncoder + NVIDIA &lt;code&gt;SoundProjection&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;cos &lt;strong&gt;0.99999130&lt;/strong&gt; (min/frame)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Audio frontend&lt;/td&gt;
&lt;td&gt;log-mel features&lt;/td&gt;
&lt;td&gt;max abs delta &lt;strong&gt;8.1e-6&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vision tower&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;nvidia/C-RADIOv4-H&lt;/code&gt; via &lt;code&gt;trust_remote_code&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;cos &lt;strong&gt;0.99996227&lt;/strong&gt; (min/token)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vision tower, MLX CPU stream&lt;/td&gt;
&lt;td&gt;same&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1.00000000&lt;/strong&gt; — graph-exact&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A port that is &lt;em&gt;almost&lt;/em&gt; right is worse than no port at all, because you spend weeks chasing quality problems that are really numerical drift in a tower you never checked. Writing the parity harness first was the single best decision in this project.&lt;/p&gt;

&lt;h2&gt;
  
  
  Speed and memory
&lt;/h2&gt;

&lt;p&gt;Measured on an M5 Max MacBook Pro, running the 4-bit quantized language model with both towers in bf16:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Image:  67.7 tok/s · 22.1 GB peak
Audio:  147  tok/s · 21.0 GB peak
Text:   152  tok/s · 17.9 GB peak
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Wifi off the whole time. Nothing leaves the machine.&lt;/p&gt;

&lt;p&gt;That 22.1 GB peak on the image path is the number I would pay attention to if you are deciding whether to bother. It suggests this fits on a 32 GB Mac. I cannot verify that, because the M5 Max is the only machine I have. If you run it on something smaller I would genuinely like to hear what happens — that is the most useful thing anyone could send me right now.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why bother doing this locally
&lt;/h2&gt;

&lt;p&gt;The obvious question is why not just call an API. A few reasons that matter to me.&lt;/p&gt;

&lt;p&gt;The model is open weights. Someone should be able to run open weights on their own hardware, and if the only path requires a vendor's cloud, the openness is partly decorative.&lt;/p&gt;

&lt;p&gt;Apple Silicon is genuinely fast enough now. 67 tokens a second while reading an image, on a laptop, is not a compromise.&lt;/p&gt;

&lt;p&gt;And there are workloads where the data cannot leave the building at all. I do work that touches NDA and compliance-sensitive material, and "it runs offline" is not a nice-to-have there, it is the requirement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;MIT licensed: &lt;strong&gt;&lt;a href="https://github.com/nicedreamzapp/nemotron-omni-mlx" rel="noopener noreferrer"&gt;https://github.com/nicedreamzapp/nemotron-omni-mlx&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Credit to NVIDIA for publishing the weights, and to yayr for the 4-bit conversion.&lt;/p&gt;

</description>
      <category>mlx</category>
      <category>apple</category>
      <category>machinelearning</category>
      <category>opensource</category>
    </item>
    <item>
      <title>NVIDIA Shipped a Model That Sees and Hears — It Just Didn’t Run on a Mac. So I Wrote the Missing Piece.</title>
      <dc:creator>Matt Macosko</dc:creator>
      <pubDate>Tue, 28 Jul 2026 18:30:07 +0000</pubDate>
      <link>https://dev.to/matt_macosko_f3829cfd86b8/nvidia-shipped-a-model-that-sees-and-hears-it-just-didnt-run-on-a-mac-so-i-wrote-the-missing-50pe</link>
      <guid>https://dev.to/matt_macosko_f3829cfd86b8/nvidia-shipped-a-model-that-sees-and-hears-it-just-didnt-run-on-a-mac-so-i-wrote-the-missing-50pe</guid>
      <description>&lt;p&gt;&lt;strong&gt;NVIDIA shipped a 30-billion-parameter model that can see, hear, and talk — and gave the weights away. The catch: the seeing and hearing parts didn’t run on a Mac. So I spent an afternoon writing the missing piece.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here is the whole thing in thirty-five seconds — the model reading a real cart off my own store, on the laptop, with nothing leaving it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://nicedreamzwholesale.com/wp-content/uploads/2026/07/nemotron_demo.mp4" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjjlbmxoxjrhpgbo4bldm.jpg" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Thirty-five seconds: the model reads a real cart off my own store, on the laptop, with nothing leaving it. The elapsed counter is the real measured latency.&lt;/p&gt;

&lt;p&gt;Every couple of weeks I go looking for whatever new open-weight model just dropped, pull it onto my laptop, and see what it can actually do. Most of the time it’s a coding model, I run it against the one I already use, and my current favorite wins again. That’s a fine result. It’s just not much of a story.&lt;/p&gt;

&lt;p&gt;This time I found something different.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Nemotron Omni actually is
&lt;/h2&gt;

&lt;p&gt;NVIDIA released &lt;strong&gt;Nemotron-3-Nano-Omni-30B-A3B&lt;/strong&gt; — a “tri-modal” model, which is a fancy way of saying one brain with eyes and ears attached. You can hand it a picture, a sound file, or a video, and talk to it about what it saw or heard. It’s 30 billion parameters total, but only about 3 billion of them fire for any given word, which is why something this capable can run on a laptop at all. The weights are public.&lt;/p&gt;

&lt;p&gt;Someone had already done the hard, unglamorous work of shrinking it down to a 4-bit MLX version that fits on Apple Silicon — that’s &lt;a href="https://huggingface.co/mlx-community/NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-4bit" rel="noopener noreferrer"&gt;yayr&lt;/a&gt; over at mlx-community, and this project doesn’t exist without that upload. About 19 GB on disk. Ready to go.&lt;/p&gt;

&lt;p&gt;Except for one line, buried in the model card:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The text backbone loads with standard MLX &lt;code&gt;nemotron_h&lt;/code&gt; tooling. The vision and audio towers require a multimodal runtime that implements the C-RADIO ViT-H and Parakeet Conformer forward passes (e.g. the Evorix on-device engine).&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Translated: the brain works on a Mac. &lt;strong&gt;The eyes and ears don’t.&lt;/strong&gt; The weights for them are right there in the file — all the knowledge, sitting on your disk — but nothing I could find in the open Apple Silicon world knew how to &lt;em&gt;run&lt;/em&gt; them. NVIDIA’s own code for those parts is written for their GPUs. On a Mac it’s a locked room with the key visible through the window.&lt;/p&gt;

&lt;p&gt;I want to be precise about that parenthetical, because it matters: the card does name an engine, Evorix. I went looking for it — not on Hugging Face, not on GitHub, not anywhere I could find. So as far as I can tell it exists, but not somewhere you or I can go get it. Which leaves the same practical dead end: you download nineteen gigabytes of a model that sees and hears, and on a Mac you can only talk to it.&lt;/p&gt;

&lt;p&gt;That’s the gap — no &lt;em&gt;open&lt;/em&gt; runtime for the eyes and ears. Honestly, that’s the most interesting kind of thing to find. Not a benchmark to run. A thing that doesn’t work yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part I keep having to relearn
&lt;/h2&gt;

&lt;p&gt;My instinct was that this would take days. Three separate pieces, each one a real port: NVIDIA’s vision tower is a ViT with a custom patch generator, their audio side is a Conformer with Transformer-XL relative attention, and then there’s the glue that turns a photo into something the language model can read.&lt;/p&gt;

&lt;p&gt;It took about half an hour.&lt;/p&gt;

&lt;p&gt;Not because I’m fast — because I stopped doing it one piece at a time. I put a separate agent on each tower and let them run at the same time, each one checking its own work against NVIDIA’s original code as it went. I keep making the same mistake of estimating this stuff like it’s still 2024, and I keep getting corrected by my own laptop.&lt;/p&gt;

&lt;h2&gt;
  
  
  Proving it, instead of vibing it
&lt;/h2&gt;

&lt;p&gt;Here’s the thing about porting a model: it’s very easy to write code that &lt;em&gt;looks&lt;/em&gt; right, produces numbers, and is quietly wrong. The model doesn’t crash. It just gets a little dumber, and you never find out.&lt;/p&gt;

&lt;p&gt;So neither tower got to claim victory on vibes. For each one, the test was: run NVIDIA’s original PyTorch code and my MLX version &lt;strong&gt;on the same input, with the same weights&lt;/strong&gt;, and compare the actual numbers coming out.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tower&lt;/th&gt;
&lt;th&gt;What was compared&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Ears&lt;/strong&gt; (audio)&lt;/td&gt;
&lt;td&gt;Final audio embeddings, 5-second clip&lt;/td&gt;
&lt;td&gt;cosine &lt;strong&gt;0.99999&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Eyes&lt;/strong&gt; (vision)&lt;/td&gt;
&lt;td&gt;Final image embeddings, 448px image&lt;/td&gt;
&lt;td&gt;cosine &lt;strong&gt;0.99996&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Eyes&lt;/strong&gt;, on CPU math&lt;/td&gt;
&lt;td&gt;Same, without GPU shortcuts&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1.00000000&lt;/strong&gt; — exact&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That last row is the one I care about. Run it on the CPU, where the math is done precisely, and my port and NVIDIA’s are &lt;strong&gt;not close — they’re identical.&lt;/strong&gt; The tiny gap on the GPU isn’t a bug in the port; it’s Metal taking shortcuts with float math for speed. Chasing that would be chasing the hardware.&lt;/p&gt;

&lt;p&gt;The input processing got the same treatment — every image tile, every audio frame, every token id checked against NVIDIA’s reference. 14 tests, token sequences matching exactly, pixels off by less than a millionth.&lt;/p&gt;

&lt;p&gt;The brain, meanwhile, needed no porting at all — MLX already understood it. Someone had put &lt;code&gt;nemotron_h&lt;/code&gt; into mlx-lm before I ever showed up.&lt;/p&gt;

&lt;h2&gt;
  
  
  So does it actually work
&lt;/h2&gt;

&lt;p&gt;This is the only part that matters, so here’s the first thing I pointed it at once the pieces were connected. Not a test image — a screenshot off my own phone, of my own store, with a cart full of my own products.&lt;/p&gt;

&lt;p&gt;I asked: &lt;em&gt;“What website is this and what is in the cart?”&lt;/em&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;This is the Divine Tribe website. The cart contains a Gen 2 DC Ceramic Rebuildable Dry Herb Heater, a Replacement Heater Cup, a Wireless Dock Station, 72% Hemp 28% Silk Men’s Boxers, and a Quest Lightning Diffuser Kit.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Every product, correct. When I asked for prices it read all five to the cent — $37.36, $13.12, $80.80, $33.32, $38.67 — and started adding them up. Nine seconds, on a laptop, with nothing leaving it.&lt;/p&gt;

&lt;p&gt;Then the ears. I generated a line of speech and handed it the wav:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The Divine Tribe vaporizer ships from Humboldt County, California, and this model is running entirely on a MacBook with no Internet connection.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Word for word, punctuation and all — it even got “Humboldt” right.&lt;/p&gt;

&lt;p&gt;The speeds, for anyone keeping score:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mode&lt;/th&gt;
&lt;th&gt;Speed&lt;/th&gt;
&lt;th&gt;Memory&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Text only&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;152 tok/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;17.9 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;With an image&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;67.7 tok/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;22.1 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;With audio&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;147 tok/s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;21.0 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The whole thing lives in about 22 GB at its hungriest. That fits on a Mac you can buy today.&lt;/p&gt;

&lt;p&gt;And to be clear about what that means: nothing about that cart screenshot or that audio clip left my desk. The model makes no network calls at all — there’s no API key, no meter running, no terms of service, and no company on the other end deciding whether I’m allowed to keep doing this tomorrow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things that fell out along the way
&lt;/h2&gt;

&lt;p&gt;Neither of these was the goal, and both are probably more useful than anything else here.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;NVIDIA’s reference code produces NaN on batched audio.&lt;/strong&gt; If you feed it two clips of different lengths at once, the shorter one comes back as garbage — not an error, just silent NaN poisoning that spreads through the layers. It’s a masking detail: fully-padded rows go to negative infinity, softmax turns that into NaN, and the NaN travels. My port masks differently and stays finite. I want to be careful here — this is one specific path, and I could be wrong about how much it matters in NVIDIA’s own pipeline. But it reproduces on their code, not just mine.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The vision tower doesn’t normalize its own input.&lt;/strong&gt; There’s a normalization layer sitting right there in the checkpoint, and it’s dead weight — NVIDIA switches it off and expects whatever calls it to do that job. Feed it raw pixels like a reasonable person would and you get plausible-looking garbage, silently. This one cost real time and it’s the trap anyone else attempting this will hit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And a third, for the truly nerdy:&lt;/strong&gt; several settings in the config file are lies. Not maliciously — they’re just vestigial, left over from an earlier design, and the live code ignores them completely. If you build from the config instead of tracing what actually runs, you’ll produce something that looks correct and isn’t. I only caught it by following the real code path line by line.&lt;/p&gt;

&lt;h2&gt;
  
  
  So who does this actually help?
&lt;/h2&gt;

&lt;p&gt;Fair question, and I want to answer it honestly instead of waving at the future.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The narrow answer: it saves the next person about a week.&lt;/strong&gt; Right now, anyone with a Mac who downloads this model reads that line on the model card — &lt;em&gt;“requires a multimodal runtime that implements the C-RADIO ViT-H and Parakeet Conformer forward passes”&lt;/em&gt; — and that’s the end of the road. It’s not a warning, it’s a wall. It means &lt;em&gt;go build it yourself&lt;/em&gt;, and most people, reasonably, close the tab. Now they don’t have to. Clone it, run it, done. And the three traps I hit are written down, because the normalization one in particular would cost someone a full day of wondering why their model got quietly dumber instead of visibly broken. That’s the entire contribution: one person spent the afternoon so nobody else has to spend the week.&lt;/p&gt;

&lt;p&gt;I’m not going to pretend that’s a huge number of people. Might be a few hundred. Might be twelve. That’s fine — twelve people not wasting a week each is still worth an afternoon.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The wider answer is the one I actually care about.&lt;/strong&gt; I do private-AI work for firms that handle other people’s confidential material — lawyers, medical practices, accountants. The whole pitch is that the machine doing the reading is the machine on your desk, because a federal court &lt;a href="https://nicedreamzwholesale.com/2026/05/22/the-heppner-ruling-warner-v-gilbarco-and-what-confidential-ai-actually-has-to-mean/" rel="noopener noreferrer"&gt;already ruled&lt;/a&gt; that work you hand to a public AI isn’t privileged.&lt;/p&gt;

&lt;p&gt;Until now, “on your desk” meant text only. If a client sends a photograph of a contract, or a recorded call, or a scan — the private option had nothing to say. You either sent it to somebody’s cloud and lost the privilege, or you did it by hand.&lt;/p&gt;

&lt;p&gt;That changed today, on my laptop. A model that can &lt;em&gt;look at&lt;/em&gt; the scanned page and &lt;em&gt;listen to&lt;/em&gt; the recording, on the machine sitting in front of them, is not a small difference for those people. It’s the difference between a tool they can use and a tool they legally can’t.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And the honest third reason:&lt;/strong&gt; the gap between “the weights are public” and “you can actually use this” is where most open AI quietly dies. Everyone celebrates the release. Far fewer people do the boring work of making the thing run somewhere real. That gap is usually not a research problem — it’s a few hundred lines nobody got around to writing. This one was maybe 90 KB of Python and an afternoon.&lt;/p&gt;

&lt;p&gt;That’s the part I’d like more people to see. Not that I did something clever — I didn’t, I transcribed NVIDIA’s own math into a different framework and checked my work. But the wall between an open model and a working model is &lt;em&gt;thinner than it looks&lt;/em&gt;, and it stays up mostly because everyone assumes someone else will knock it down.&lt;/p&gt;

&lt;p&gt;The code is &lt;a href="https://github.com/nicedreamzapp/nemotron-omni-mlx" rel="noopener noreferrer"&gt;on GitHub&lt;/a&gt;, MIT, with every parity test in it. Don’t take my word for any number above — clone it and run the tests on your own Mac.&lt;/p&gt;

&lt;h2&gt;
  
  
  Credit where it’s due
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16" rel="noopener noreferrer"&gt;NVIDIA&lt;/a&gt;&lt;/strong&gt; built the model and released the weights and reference code openly. None of this happens otherwise.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://huggingface.co/mlx-community/NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-4bit" rel="noopener noreferrer"&gt;yayr&lt;/a&gt;&lt;/strong&gt; at mlx-community did the 4-bit MLX quantization I built on top of. Go give that upload a like — it deserves more than the zero it has.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://github.com/ml-explore/mlx" rel="noopener noreferrer"&gt;Apple’s MLX team&lt;/a&gt;&lt;/strong&gt; — and whoever added &lt;code&gt;nemotron_h&lt;/code&gt; to mlx-lm, which is why the brain needed nothing from me at all.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://nicedreamzwholesale.com/category/ai-computing/" rel="noopener noreferrer"&gt;Nice Dreamz Wholesale&lt;/a&gt;. Run AI locally on your own hardware with &lt;a href="https://github.com/nicedreamzapp/claude-code-local" rel="noopener noreferrer"&gt;claude-code-local&lt;/a&gt;, open source and no cloud required. More at &lt;a href="https://nicedreamzwholesale.com/software/" rel="noopener noreferrer"&gt;nicedreamzwholesale.com/software&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiampcomputing</category>
    </item>
    <item>
      <title>The Day Local AI Caught the Cloud: ds4, DeepSeek V4 Flash, and What Just Changed for Devs</title>
      <dc:creator>Matt Macosko</dc:creator>
      <pubDate>Thu, 23 Jul 2026 18:30:07 +0000</pubDate>
      <link>https://dev.to/matt_macosko_f3829cfd86b8/the-day-local-ai-caught-the-cloud-ds4-deepseek-v4-flash-and-what-just-changed-for-devs-4kjo</link>
      <guid>https://dev.to/matt_macosko_f3829cfd86b8/the-day-local-ai-caught-the-cloud-ds4-deepseek-v4-flash-and-what-just-changed-for-devs-4kjo</guid>
      <description>&lt;p&gt;If you write code for a living and you’ve been watching the local-AI space, May 9, 2026 is the date to circle. Salvatore Sanfilippo (yes, the guy who wrote Redis) shipped &lt;a href="https://github.com/antirez/ds4" rel="noopener noreferrer"&gt;&lt;code&gt;ds4&lt;/code&gt;&lt;/a&gt; — a few thousand lines of hand-written C with Metal compute kernels, built for exactly one model: &lt;strong&gt;DeepSeek V4 Flash&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I ran the same prompt through three engines on the same 128 GB MacBook Pro:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;DeepSeek V4 Flash&lt;/strong&gt; via &lt;code&gt;ds4&lt;/code&gt; — fully local, off-cloud&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cloud Claude&lt;/strong&gt; through my Max plan&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemma 4 31B&lt;/strong&gt; via MLX, also local&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Local DeepSeek beat cloud Claude on wall-clock time. That sentence used to be science fiction.&lt;/p&gt;

&lt;p&gt;▶ &lt;strong&gt;&lt;a href="https://youtu.be/7l8-s8xkpms" rel="noopener noreferrer"&gt;Watch the companion video&lt;/a&gt;&lt;/strong&gt; — three engines, one prompt, three completely different aurora animations rendered in real time on the same machine.&lt;/p&gt;




&lt;h2&gt;
  
  
  The benchmark, for people who don’t want filler
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Engine&lt;/th&gt;
&lt;th&gt;Time&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;th&gt;Where it ran&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4 Flash (&lt;code&gt;ds4&lt;/code&gt; local)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;103 s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;3,259 tokens&lt;/td&gt;
&lt;td&gt;Apple Silicon GPU&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cloud Claude (Max plan)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;192 s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~3,500 tokens&lt;/td&gt;
&lt;td&gt;Anthropic data center&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 31B (MLX local)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;131 s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1,992 tokens&lt;/td&gt;
&lt;td&gt;Apple Silicon GPU&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The prompt was a single creative HTML task: &lt;em&gt;“Build an animated northern lights scene — single file, vanilla JS, mountains, pine trees, twinkling stars, flowing aurora bands.”&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Each engine produced a completely different aurora. None of them hit the network during inference. (Yes, I checked with &lt;code&gt;lsof&lt;/code&gt;. Yes, this is the same &lt;code&gt;lsof&lt;/code&gt; audit pattern from &lt;a href="https://nicedreamzwholesale.com/airgap" rel="noopener noreferrer"&gt;the AirGap NDA piece&lt;/a&gt;.)&lt;/p&gt;




&lt;h2&gt;
  
  
  Three architectural decisions in &lt;code&gt;ds4&lt;/code&gt; worth understanding
&lt;/h2&gt;

&lt;p&gt;This is the part that matters if you’re a developer thinking about local AI infrastructure.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Asymmetric 2-bit quantization (only where quality is forgiving)
&lt;/h3&gt;

&lt;p&gt;The naive approach to quantization treats every weight the same. &lt;code&gt;ds4&lt;/code&gt; doesn’t. &lt;strong&gt;Only the routed Mixture-of-Experts experts get compressed to 2-bit&lt;/strong&gt; (specifically &lt;code&gt;IQ2_XXS&lt;/code&gt; for &lt;code&gt;up&lt;/code&gt;/&lt;code&gt;gate&lt;/code&gt;, &lt;code&gt;Q2_K&lt;/code&gt; for &lt;code&gt;down&lt;/code&gt;). Every quality-critical path — shared experts, attention projections, routing, output head — stays at higher precision (Q8 or full).&lt;/p&gt;

&lt;p&gt;Those routed experts are about 90% of the weight footprint. The other 10% is where small precision losses cause big accuracy losses. Quantize the 90%, leave the 10%, and you get an 81 GB file that still calls tools cleanly and writes coherent code.&lt;/p&gt;

&lt;p&gt;This is the kind of tradeoff that only makes sense if you’ve stared at a specific model’s loss landscape long enough to know which weights tolerate compression. It’s a model-specific engineering decision dressed as a quantization recipe.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. KV cache moved to disk (in 2026 SSDs are fast enough)
&lt;/h3&gt;

&lt;p&gt;The “KV cache must live in RAM” assumption is from 2023. Modern Apple SSDs do 5+ GB/s sequential reads. &lt;code&gt;ds4&lt;/code&gt; writes session state to disk and &lt;strong&gt;reuses it across runs&lt;/strong&gt;, keyed by SHA1 of token IDs.&lt;/p&gt;

&lt;p&gt;The practical effect: when Claude Code sends its 25k-token system prompt, that prefill happens exactly once, ever. Every subsequent session — including totally different agent runs that happen to share that prefix — reads from disk in milliseconds instead of recomputing from token zero.&lt;/p&gt;

&lt;p&gt;If you’ve used long-context models locally, you know prefill is the slowest thing in the loop. &lt;code&gt;ds4&lt;/code&gt; makes it free after the first hit. That’s the kind of “small change, huge implication” move that took years to normalize. (See also: &lt;a href="https://github.com/antirez/ds4#disk-kv-cache" rel="noopener noreferrer"&gt;the disk-KV section in the ds4 README&lt;/a&gt;.)&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Pure Metal, not CUDA-with-a-shim
&lt;/h3&gt;

&lt;p&gt;There’s no PyTorch, no TensorFlow, no &lt;code&gt;llama.cpp&lt;/code&gt; wrapper layer in the hot path. The compute kernels under &lt;code&gt;metal/*.metal&lt;/code&gt; are &lt;strong&gt;written specifically for this one model on this one architecture&lt;/strong&gt;. The acknowledgments thank &lt;code&gt;llama.cpp&lt;/code&gt; and GGML — &lt;code&gt;ds4&lt;/code&gt; borrows quant layouts and select kernels — but it’s not a fork.&lt;/p&gt;

&lt;p&gt;This narrowness is the point. Generic frameworks pay a tax for being generic. When you commit to one model on one chip, you can hand-tune away that tax. ~27 tok/s on an M3 Max 128 GB. ~32 tok/s on M5 Max. For agent loops on a laptop, that’s plenty.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why this matters for compliance-sensitive devs
&lt;/h2&gt;

&lt;p&gt;The same week, I’m still maintaining &lt;a href="https://nicedreamzwholesale.com/airgap" rel="noopener noreferrer"&gt;AirGap AI&lt;/a&gt; — a wi-fi-off, &lt;code&gt;lsof&lt;/code&gt;-audited workflow for analyzing privileged documents (NDAs, client files, PHI, etc.) on a laptop with no outbound connections. Until last week, that was a Llama 3.3 70B story. The capability ceiling was real.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;ds4&lt;/code&gt; raises that ceiling materially:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;1M-token context&lt;/strong&gt; — entire codebases, full deposition transcripts, complete contract sets, all in-memory in a single conversation&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quasi-frontier reasoning&lt;/strong&gt; — if you’ve used Claude Sonnet or Opus, DeepSeek V4 Flash sits in the same neighborhood for most agentic tasks&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool calling that works&lt;/strong&gt; — Antirez tested it under coding agents (opencode, Pi, Claude Code) and the tool calls land reliably&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For law firms, medical practices, and compliance-bound shops, the math just changed. You don’t have to choose between “frontier-grade reasoning” and “data never leaves the building.” The hardware exists, the engine exists, the model exists, and the integration with Claude Code exists.&lt;/p&gt;

&lt;p&gt;(If you’re trying to get a bar-association-defensible AI workflow off the ground, &lt;a href="https://nicedreamzwholesale.com/airgap" rel="noopener noreferrer"&gt;the AirGap landing page&lt;/a&gt; is where I keep my notes. The ds4 stack is going in there next week.)&lt;/p&gt;




&lt;h2&gt;
  
  
  How to actually run it
&lt;/h2&gt;

&lt;p&gt;The full stack, all open-source:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 1. Build the engine (Apple Silicon with Metal)&lt;/span&gt;
git clone https://github.com/antirez/ds4
&lt;span class="nb"&gt;cd &lt;/span&gt;ds4 &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; make

&lt;span class="c"&gt;# 2. Pull the q2 weights (~81 GB)&lt;/span&gt;
./download_model.sh q2

&lt;span class="c"&gt;# 3. Boot the local Anthropic-compatible server&lt;/span&gt;
./ds4-server &lt;span class="nt"&gt;--ctx&lt;/span&gt; 200000 &lt;span class="nt"&gt;--kv-disk-dir&lt;/span&gt; ~/Library/Caches/ds4-kv &lt;span class="se"&gt;\&lt;/span&gt;
             &lt;span class="nt"&gt;--kv-disk-space-mb&lt;/span&gt; 16384

&lt;span class="c"&gt;# 4. Point Claude Code at it&lt;/span&gt;
&lt;span class="nv"&gt;ANTHROPIC_BASE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;http://127.0.0.1:8000 &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nv"&gt;ANTHROPIC_AUTH_TOKEN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;dsv4-local &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nv"&gt;ANTHROPIC_MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;deepseek-v4-flash &lt;span class="se"&gt;\&lt;/span&gt;
claude
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or just clone &lt;a href="https://github.com/nicedreamzapp/claude-code-local" rel="noopener noreferrer"&gt;&lt;code&gt;nicedreamzapp/claude-code-local&lt;/code&gt;&lt;/a&gt; — DeepSeek V4 Flash is now the fourth fighter in the lineup, with a &lt;code&gt;claude-ds4&lt;/code&gt; wrapper that handles all of the above for you.&lt;/p&gt;




&lt;h2&gt;
  
  
  What this slots into
&lt;/h2&gt;

&lt;p&gt;This isn’t a one-off. It’s the next click in a longer arc I’ve been writing about:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://marijuanaunion.com/three-generations-of-running-claude-code-locally-on-a-macbook-what-i-actually-learned/" rel="noopener noreferrer"&gt;Three Generations of Running Claude Code Locally on a MacBook — What I Actually Learned&lt;/a&gt; — the long path from “barely works” to “actually replaces my cloud usage”&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://marijuanaunion.com/cloud-ai-coding-costs-keep-climbing-how-to-pay-0-and-still-use-claude-code/" rel="noopener noreferrer"&gt;Cloud AI Coding Costs Keep Climbing — How to Pay $0 and Still Use Claude Code&lt;/a&gt; — the economic angle, before &lt;code&gt;ds4&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://nicedreamzwholesale.com/2026/04/26/claude-subscription-value-10x/" rel="noopener noreferrer"&gt;Pulling 10x My Subscription Value Out of Claude&lt;/a&gt; — what the cloud math actually looks like&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://marijuanaunion.com/what-its-actually-like-to-code-by-voice-with-the-ai-replying-in-my-own-cloned-voice/" rel="noopener noreferrer"&gt;What It’s Actually Like to Code By Voice — With the AI Replying In My Own Cloned Voice&lt;/a&gt; — the voice loop these models now plug into&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://marijuanaunion.com/your-medical-practice-is-probably-using-cloud-ai-on-phi-right-now-heres-the-hipaa-problem-nobody-is-talking-about/" rel="noopener noreferrer"&gt;Your Medical Practice Is Probably Using Cloud AI on PHI Right Now&lt;/a&gt; — why on-device matters for healthcare&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://marijuanaunion.com/if-your-law-firm-is-using-cloud-ai-on-client-files-you-probably-have-a-problem/" rel="noopener noreferrer"&gt;If Your Law Firm Is Using Cloud AI on Client Files, You Probably Have a Problem&lt;/a&gt; — the legal angle&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://marijuanaunion.com/a-field-guide-to-ambient-computing-the-words-for-the-thing-thats-coming/" rel="noopener noreferrer"&gt;A Field Guide to Ambient Computing&lt;/a&gt; — the bigger frame this all sits inside&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;code&gt;ds4&lt;/code&gt; is the engine that finally makes the local-first version of all of those usable for production work. The local agent doesn’t have to pick which workload it’s good at anymore.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where to follow
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;🛠️ &lt;a href="https://github.com/nicedreamzapp/claude-code-local" rel="noopener noreferrer"&gt;github.com/nicedreamzapp/claude-code-local&lt;/a&gt; — the lineup, launchers, and benchmarks&lt;/li&gt;
&lt;li&gt;🐳 &lt;a href="https://github.com/antirez/ds4" rel="noopener noreferrer"&gt;github.com/antirez/ds4&lt;/a&gt; — the engine itself&lt;/li&gt;
&lt;li&gt;🌿 &lt;a href="https://marijuanaunion.com" rel="noopener noreferrer"&gt;marijuanaunion.com&lt;/a&gt; — the broader writing on local AI, voice, and ambient computing&lt;/li&gt;
&lt;li&gt;🔒 &lt;a href="https://nicedreamzwholesale.com/airgap" rel="noopener noreferrer"&gt;nicedreamzwholesale.com/airgap&lt;/a&gt; — the compliance-grade workflow notes&lt;/li&gt;
&lt;li&gt;💬 &lt;a href="https://discord.gg/g7rgabGD9E" rel="noopener noreferrer"&gt;Discord&lt;/a&gt; — NiceDreamzApps server&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;May 9, 2026. The day a single C file caught up to the data centers.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This is the technical companion to &lt;a href="https://marijuanaunion.com/i-just-watched-one-hacker-catch-up-to-a-trillion-dollar-data-center/" rel="noopener noreferrer"&gt;the headline piece on Marijuana Union&lt;/a&gt;. Companion video: &lt;a href="https://youtu.be/7l8-s8xkpms" rel="noopener noreferrer"&gt;youtu.be/7l8-s8xkpms&lt;/a&gt;. For local-AI consulting on compliance-sensitive workloads, see &lt;a href="https://nicedreamzwholesale.com/airgap" rel="noopener noreferrer"&gt;AirGap AI&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://nicedreamzwholesale.com/category/ai-computing/" rel="noopener noreferrer"&gt;Nice Dreamz Wholesale&lt;/a&gt;. Run AI locally on your own hardware with &lt;a href="https://github.com/nicedreamzapp/claude-code-local" rel="noopener noreferrer"&gt;claude-code-local&lt;/a&gt;, open source and no cloud required. More at &lt;a href="https://nicedreamzwholesale.com/software/" rel="noopener noreferrer"&gt;nicedreamzwholesale.com/software&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiampcomputing</category>
      <category>ai</category>
      <category>antirez</category>
      <category>applesilicon</category>
    </item>
    <item>
      <title>The IRS Said Yes — and What That Does and Doesn’t Mean</title>
      <dc:creator>Matt Macosko</dc:creator>
      <pubDate>Thu, 16 Jul 2026 06:10:44 +0000</pubDate>
      <link>https://dev.to/matt_macosko_f3829cfd86b8/the-irs-said-yes-and-what-that-does-and-doesnt-mean-2lnj</link>
      <guid>https://dev.to/matt_macosko_f3829cfd86b8/the-irs-said-yes-and-what-that-does-and-doesnt-mean-2lnj</guid>
      <description>&lt;p&gt;Back on June 3rd I published a piece here saying the Cannabis Device Safety Institute had filed its federal application, and I made a point of leaning on that word — &lt;em&gt;application&lt;/em&gt;. Filed, not granted. Pending, not approved. Not a tax-exempt charity, contributions not deductible, and we’d say so plainly every time, because saying it any other way would be exactly the kind of thing this institute exists to push against.&lt;/p&gt;

&lt;p&gt;So I owe you the update in the same register.&lt;/p&gt;

&lt;p&gt;The IRS granted it. The determination letter is dated June 29, 2026. CDSI is exempt from federal income tax under Section 501(c)(3), classified as a public charity under Section 509(a)(2) — not a private foundation — with the exemption effective April 27, 2026, retroactive to the day the articles were filed. Contributions are deductible under Section 170.&lt;/p&gt;

&lt;p&gt;That’s the whole announcement. Now let me do the part I think matters more, which is being precise about what it does and doesn’t mean.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it doesn’t mean
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;It is not an endorsement.&lt;/strong&gt; Recognition of exemption is a determination about tax status. The IRS did not review our methodology, did not evaluate our off-gas protocol, and has no opinion whatsoever about whether ceramic donut atomizers off-gas at 650°F. No agency has blessed CDSI’s findings, because CDSI hasn’t published findings yet. If you ever see me imply otherwise, call me on it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It doesn’t make CDSI a regulator.&lt;/strong&gt; We’re a private nonprofit. We have no authority over anyone, we can’t compel any manufacturer to do anything, and we’re not seeking a government designation. The model is UL, ASTM, the NFPA Research Foundation — bodies that earned standing by being useful and rigorous, not by being appointed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It doesn’t mean we have a lab — yet.&lt;/strong&gt; Right now CDSI pays accredited independent labs and publishes what comes back. Pay the lab, not pay to pass. But the goal was never to outsource forever: the plan is for this institute to build and run its own testing bench, because the body that writes the methodology should eventually be able to execute it too — and the determination letter is exactly what makes that fundable. Foundation grants and tax-deductible donations can now go toward standing up a lab of our own. If and when that lab exists, nothing about the discipline changes — open methodology, public reports, every conflict on the cover page.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It doesn’t finish the paperwork.&lt;/strong&gt; California still treats CDSI as a taxable corporation until a separate state filing goes through. That one’s in an envelope, not a press release.&lt;/p&gt;

&lt;h2&gt;
  
  
  About the timeline, honestly
&lt;/h2&gt;

&lt;p&gt;We filed the full Form 1023 on June 3. The determination is dated June 29 — 26 days later.&lt;/p&gt;

&lt;p&gt;The IRS publishes exactly one number about how long this takes: &lt;em&gt;“We issue 80% of Form 1023 application determinations within 191 days.”&lt;/em&gt; That’s from their own “Where’s my application” page. So: 26 days, against a published benchmark of 191, on the long form rather than the 1023-EZ shortcut, with no request for expedited handling.&lt;/p&gt;

&lt;p&gt;I want to be careful with that number, because it would be very easy to turn it into something it isn’t.&lt;/p&gt;

&lt;p&gt;I can’t tell you it’s a record. Nobody can tell you that, about any organization. The IRS’s public files record determination dates by month only — no day — and carry no application-submitted date at all. The interval literally cannot be computed from public data, including the IRS’s own. So anyone claiming a record in this category is claiming something no dataset could check, and I’d rather not be that guy.&lt;/p&gt;

&lt;p&gt;I also can’t tell you the speed proves the application was good. That’s the flattering read, and I don’t think it survives scrutiny. The IRS has fast-track lanes for straightforward cases, and how big those lanes are isn’t published. The six days between answering their follow-up letter and getting the determination is, as far as I can tell, just the IRS’s own internal rule about how fast a specialist has to close a case once you respond. And the biggest variable — when the application got assigned to a human at all — was completely outside my control and I can’t explain why it happened when it did.&lt;/p&gt;

&lt;p&gt;What I can say is what we actually did, and let you decide if any of it mattered:&lt;/p&gt;

&lt;p&gt;We paid $600 for the long form instead of $275 for the EZ, on purpose, because the EZ doesn’t have room to explain why a fee-for-service testing subsidiary serves a public mission — and filings like ours get bounced back to the long form anyway, months later. The slow-looking choice was the fast one.&lt;/p&gt;

&lt;p&gt;We disclosed the conflicts instead of burying them. The founder of a cannabis hardware standards body builds cannabis hardware. That’s on the record, in the conflict of interest policy, in the state charity filing, and it’ll be on the cover page of every report we publish.&lt;/p&gt;

&lt;p&gt;We went through this site and cut every claim we couldn’t prove — before we filed, not after. No “first,” no “leading,” no partnerships that hadn’t happened. Reviewers read your website. There was nothing to pick at because we’d already picked at it ourselves.&lt;/p&gt;

&lt;p&gt;And when the IRS wrote asking for more information, we answered within hours of opening the envelope instead of sitting on the 28 days we had.&lt;/p&gt;

&lt;p&gt;None of that is clever. It’s just doing the boring things in the right order.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it actually changes
&lt;/h2&gt;

&lt;p&gt;Donations to CDSI are now tax-deductible. That’s real — it means a foundation can fund hardware testing, and an individual who cares about this can help and take the deduction.&lt;/p&gt;

&lt;p&gt;CDSI is registered with the federal government and eligible for federal grants — that happened back on June 23, with the SAM.gov registration going active.&lt;/p&gt;

&lt;p&gt;Put those together and here’s the thing I care about: &lt;strong&gt;the money to characterize this hardware can now exist.&lt;/strong&gt; For fourteen years the reason nobody tested the device was that no one would pay for it. In 2016 I paid a lab out of pocket to run an off-gas test on a concentrate vaporizer because there was no funding source in the world for that work. That’s still the founding artifact of this institute, and it’s still the reason it exists.&lt;/p&gt;

&lt;p&gt;Now there’s a vehicle that can receive that money. That’s what a determination letter is. Not a trophy — plumbing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this goes
&lt;/h2&gt;

&lt;p&gt;The next thing I want to make the case for is bigger than safety, and I’ll write it properly soon.&lt;/p&gt;

&lt;p&gt;It’s this: you cannot honestly answer whether inhaled cannabis helps someone until you know what the device contributed to the dose. Heat a vaporizer and it puts its own chemistry into the same stream carrying the cannabis. So every measured outcome — a contaminant, a symptom, a relief — has two possible sources: the material and the machine. That’s not noise you can fix with more study subjects. It’s an attribution problem, and no sample size touches it.&lt;/p&gt;

&lt;p&gt;Every other measurement science solved this a century ago. The chemist runs a blank before running samples. Nobody doses a patient through an uncharacterized nebulizer. Cannabis research does the equivalent constantly, not out of carelessness, but because no one was ever responsible for the device.&lt;/p&gt;

&lt;p&gt;Which means hardware characterization isn’t adjacent to the medical question. It’s upstream of it. First you characterize the device. Then you have a device you can trust as an instrument. Then — and only then — you can ask what the medicine does to a person and believe the answer.&lt;/p&gt;

&lt;p&gt;It all starts with the devices. Not because the devices matter most. The person matters most. But the device is the part you have to understand first in order to understand any of the rest of it honestly.&lt;/p&gt;

&lt;p&gt;That’s the work. The letter just means we can afford to do it.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;The Cannabis Device Safety Institute is a California nonprofit public benefit corporation, recognized by the IRS as exempt under IRC § 501(c)(3) and classified as a public charity under § 509(a)(2) (determination letter dated June 29, 2026; EIN 42-2429365). Recognition of exemption is a determination of federal tax status and is not an endorsement of the Institute or its findings by the IRS or any government agency. cdsi.click&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://marijuanaunion.com" rel="noopener noreferrer"&gt;Marijuana Union&lt;/a&gt;. The Cannabis Device Safety Institute is an independent 501(c)(3) nonprofit standards body for cannabis consumption hardware — open methodology, public reports, every conflict of interest on the cover page. Methodology, papers, and the public record: &lt;a href="https://cdsi.click" rel="noopener noreferrer"&gt;cdsi.click&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>blog</category>
    </item>
    <item>
      <title>M5 Max + 128GB = a 30B AI Coding Agent Running Locally. Wi-Fi Off.</title>
      <dc:creator>Matt Macosko</dc:creator>
      <pubDate>Mon, 29 Jun 2026 17:18:47 +0000</pubDate>
      <link>https://dev.to/matt_macosko_f3829cfd86b8/m5-max-128gb-a-30b-ai-coding-agent-running-locally-wi-fi-off-4ggh</link>
      <guid>https://dev.to/matt_macosko_f3829cfd86b8/m5-max-128gb-a-30b-ai-coding-agent-running-locally-wi-fi-off-4ggh</guid>
      <description>&lt;p&gt;&lt;a href="https://nicedreamzwholesale.com/wp-content/uploads/2026/04/qwen_speed_promo_FINAL.mp4" rel="noopener noreferrer"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The M5 Max MacBook Pro with 128 GB of unified memory is the first laptop that can hold a frontier-class coding agent entirely in RAM. No GPU rack. No cloud. No subscription.&lt;/p&gt;

&lt;p&gt;That clip up top isn’t a render. That’s Qwen 3 Coder — 30 billion parameters, 8-bit MLX — running on this MacBook with the Wi-Fi off. Around 55 tokens per second. Total cost to keep running it: zero.&lt;/p&gt;

&lt;p&gt;The thing that matters more than the spec sheet is &lt;strong&gt;what it actually unlocks.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the M5 Max changes the math
&lt;/h2&gt;

&lt;p&gt;Until now, running a 30B+ parameter model meant a GPU rack — or paying a cloud API per token. The M5 Max changes that:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;128 GB unified memory.&lt;/strong&gt; The entire model lives in fast RAM. No GPU offload, no quantization tricks past 8-bit. The CPU and “GPU” share the same memory, so there’s no copy step between them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mixture-of-experts plays perfectly with Apple Silicon.&lt;/strong&gt; Qwen 3 Coder is 30B total but only 3B active per token. That’s a math problem the M5 Max’s memory bandwidth eats for breakfast.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MLX runs at near-CUDA speed.&lt;/strong&gt; Apple’s native ML framework hits ~55 tok/s on the 8-bit quant. No CUDA tax, no Nvidia driver politics, no $40,000 GPU bill.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It’s a regular store-bought laptop.&lt;/strong&gt; No GPU rack. No data center. No cloud bill. You can run it on a plane.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This wasn’t possible on a laptop a year ago. It is now.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you can do with it now
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Read a legal contract — and have it never leave your machine.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Most AI tools pipe your document to a server somewhere. With this setup, the bytes don’t leave the laptop. NDAs, supplier agreements, employment contracts — review them at your kitchen table without uploading them to anyone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Write production code in a couple of seconds.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
The video shows it: real Python function, real Qwen output, no edits. The agent’s tool-calling is good enough to drop into Claude Code’s loop, where it’ll edit files, run shell commands, and iterate. It’s plenty for everything from one-off scripts to refactoring real production code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Analyze patient charts without a HIPAA violation.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
For doctors, therapists, intake clinics — anything with PHI on it — local-only AI isn’t a nice-to-have, it’s the only legal option. Same model, same speed, zero bytes leaving the device.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Build agents that don’t charge you per call.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
This is the one most people sleep on. Pay-per-token cloud APIs make agents expensive to leave running. Once the model is local, you can let an agent loop overnight, hit it with thousands of requests, kick off a watcher that scans your inbox every two minutes — and the cost stays at zero.&lt;/p&gt;

&lt;h2&gt;
  
  
  The full stack
&lt;/h2&gt;

&lt;p&gt;Here’s the receipt:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hardware:&lt;/strong&gt; M5 Max MacBook Pro, 128 GB unified memory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model:&lt;/strong&gt; &lt;a href="https://huggingface.co/lmstudio-community/Qwen3-Coder-30B-A3B-Instruct-MLX-8bit" rel="noopener noreferrer"&gt;Qwen3-Coder-30B-A3B-Instruct-MLX-8bit&lt;/a&gt; — about 30 GB on disk. Mixture-of-experts, ~3B params active per token.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Server:&lt;/strong&gt; A small Python proxy at localhost:4000 that speaks the Anthropic Messages API, so the &lt;a href="https://claude.com/claude-code" rel="noopener noreferrer"&gt;Claude Code CLI&lt;/a&gt; thinks it’s talking to the cloud — except it’s talking to a hard drive.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Total monthly cost:&lt;/strong&gt; $0 once it’s downloaded.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That’s it. No Docker, no Kubernetes, no VPS. Just a laptop on a desk.&lt;/p&gt;

&lt;h2&gt;
  
  
  The performance, honestly
&lt;/h2&gt;

&lt;p&gt;The local-AI space is full of overclaims, so the straight numbers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;55 tokens per second&lt;/strong&gt; on a real coding task. Sustained, not peak.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Two seconds&lt;/strong&gt; to write a working find_median() function. Three to four seconds for most refactors.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool-calling reliability&lt;/strong&gt; is good enough for the Claude Code agentic loop. Not as locked-in as Sonnet 4.6, but plenty for getting work done.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What it’s not:&lt;/strong&gt; a Sonnet replacement for nuanced reasoning, long contexts, or really tricky debugging. For day-to-day code agent work, it more than holds its own.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why the offline part matters
&lt;/h2&gt;

&lt;p&gt;The reason “Wi-Fi off” keeps coming back in the demo isn’t a gimmick. It’s the whole thesis.&lt;/p&gt;

&lt;p&gt;If a tool needs the internet, three things are true:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Someone else can read what you sent.&lt;/li&gt;
&lt;li&gt;Someone else can charge you for it.&lt;/li&gt;
&lt;li&gt;Someone else can take it away.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If the same tool runs locally, none of those are true. That’s a different category of software. Not better at every task — but yours.&lt;/p&gt;

&lt;h2&gt;
  
  
  Benchmarks — actually run, not cited
&lt;/h2&gt;

&lt;p&gt;Big claims need numbers. So here’s what Qwen 3 Coder 30B-A3B (8-bit MLX) actually scores on this MacBook, run end-to-end against the local localhost:4000 server. Every problem solved by the model, executed in a Python subprocess, scored pass/fail. Pass@1, temperature=0, single sample per problem.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;N&lt;/th&gt;
&lt;th&gt;Pass@1&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;HumanEval&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;164/164 (full)&lt;/td&gt;
&lt;td&gt;81.7%&lt;/td&gt;
&lt;td&gt;Python function-completion classic. Saturated benchmark; modern coding models cluster 75–95%. 14 min total wall-clock.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;MBPP&lt;/strong&gt; (sanitized)&lt;/td&gt;
&lt;td&gt;168/427 (sampled)&lt;/td&gt;
&lt;td&gt;83.3%&lt;/td&gt;
&lt;td&gt;Mostly Basic Python Problems. Pass rate was stable since n=120; a few outlier tasks induce very long model responses, so I cut off at 168.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Both runs used pass@1, temperature=0, 10s execution timeout, on the local 8-bit MLX quantization. No retries. No best-of-N tricks. Single sample per problem.&lt;/p&gt;

&lt;h3&gt;
  
  
  For context — what the bigger sibling scores on harder benchmarks
&lt;/h3&gt;

&lt;p&gt;The Qwen team didn’t publish HumanEval/MBPP for any Qwen3-Coder variant — they consider those benchmarks saturated. Their &lt;a href="https://qwenlm.github.io/blog/qwen3-coder/" rel="noopener noreferrer"&gt;official benchmarks&lt;/a&gt; are agentic, and they ran them on the flagship Qwen3-Coder-480B-A35B-Instruct (the bigger sibling, ~16× the active params of the 30B-A3B running on this laptop). For context — here’s what the flagship 480B scores on those harder agentic benchmarks compared to the major closed models:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Agentic Benchmark&lt;/th&gt;
&lt;th&gt;Qwen3-Coder 480B&lt;/th&gt;
&lt;th&gt;Claude Sonnet 4&lt;/th&gt;
&lt;th&gt;GPT-4.1&lt;/th&gt;
&lt;th&gt;DeepSeek-V3&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;SWE-bench Verified&lt;/strong&gt; (500-turn)&lt;/td&gt;
&lt;td&gt;69.6&lt;/td&gt;
&lt;td&gt;70.4&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Terminal-Bench&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;37.5&lt;/td&gt;
&lt;td&gt;35.5&lt;/td&gt;
&lt;td&gt;25.3&lt;/td&gt;
&lt;td&gt;2.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;BFCL-v3&lt;/strong&gt; (function calling)&lt;/td&gt;
&lt;td&gt;68.7&lt;/td&gt;
&lt;td&gt;73.3&lt;/td&gt;
&lt;td&gt;62.9&lt;/td&gt;
&lt;td&gt;64.7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Aider-Polyglot&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;61.8&lt;/td&gt;
&lt;td&gt;56.4&lt;/td&gt;
&lt;td&gt;52.4&lt;/td&gt;
&lt;td&gt;56.9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;WebArena&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;49.9&lt;/td&gt;
&lt;td&gt;51.1&lt;/td&gt;
&lt;td&gt;44.3&lt;/td&gt;
&lt;td&gt;40.0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Source: &lt;a href="https://qwenlm.github.io/blog/qwen3-coder/" rel="noopener noreferrer"&gt;Qwen team’s official blog&lt;/a&gt;. The 30B-A3B running on this MacBook is a smaller sibling of the 480B — it trades absolute peak agentic ceiling for fitting in 30 GB and running 24/7 on local hardware. For most coding tasks people actually do in a day, HumanEval/MBPP-class accuracy matters more than the SWE-bench top-line, and on those it sits where it should: useful, fast, local.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this is heading
&lt;/h2&gt;

&lt;p&gt;The next year of the AI conversation isn’t going to be “which model is smartest.” It’s going to be “which workloads belong on your machine, and which belong on someone else’s.”&lt;/p&gt;

&lt;p&gt;Compliance-bound work — legal, medical, financial — is going to move local fast. Code-agent loops will follow because the math (per-call cost vs. zero) is brutal. The M5 Max with 128 GB of unified memory is the laptop that lets that happen.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it yourself
&lt;/h2&gt;

&lt;p&gt;The launchers are open source on GitHub: &lt;a href="https://github.com/nicedreamzapp/claude-code-local" rel="noopener noreferrer"&gt;nicedreamzapp/claude-code-local&lt;/a&gt;. The README walks through downloading the model and pointing Claude Code at the local server.&lt;/p&gt;

&lt;p&gt;For law firms, medical practices, and accountants that want help getting this running on their own hardware — that’s what &lt;a href="https://nicedreamzwholesale.com/airgap" rel="noopener noreferrer"&gt;AirGap&lt;/a&gt; is. 14-day pilot, fixed scope, the data never leaves your machines.&lt;/p&gt;

&lt;p&gt;— matt&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://nicedreamzwholesale.com/category/ai-computing/" rel="noopener noreferrer"&gt;Nice Dreamz Wholesale&lt;/a&gt;. Run AI locally on your own hardware with &lt;a href="https://github.com/nicedreamzapp/claude-code-local" rel="noopener noreferrer"&gt;claude-code-local&lt;/a&gt;, open source and no cloud required. More at &lt;a href="https://nicedreamzwholesale.com/software/" rel="noopener noreferrer"&gt;nicedreamzwholesale.com/software&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiampcomputing</category>
    </item>
  </channel>
</rss>
