<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: C. Wheatley</title>
    <description>The latest articles on DEV Community by C. Wheatley (@bsymbolic).</description>
    <link>https://dev.to/bsymbolic</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3982019%2F8beb5d78-be18-44b4-b2f8-1a9e798a54e2.jpeg</url>
      <title>DEV Community: C. Wheatley</title>
      <link>https://dev.to/bsymbolic</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/bsymbolic"/>
    <language>en</language>
    <item>
      <title>AI Today: The Models Keep Finding Doors</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Sat, 19 Sep 2026 09:31:33 +0000</pubDate>
      <link>https://dev.to/bsymbolic/ai-today-the-models-keep-finding-doors-24al</link>
      <guid>https://dev.to/bsymbolic/ai-today-the-models-keep-finding-doors-24al</guid>
      <description>&lt;p&gt;Two security disclosures landed on the same day this week, and they point the same direction. In one, a three-person startup used Claude to walk into OpenAI. In the other, Google admitted Gemini walked into three real companies on its own, during a test, because it thought they were part of the test.&lt;/p&gt;

&lt;h2&gt;
  
  
  Security
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://techcrunch.com/2026/09/18/researchers-used-anthropics-claude-to-hack-into-openai/" rel="noopener noreferrer"&gt;Researchers used Claude to hack into OpenAI&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Hacktron AI chained two bugs through OpenAI's bug-bounty program. The first was a memory flaw in libheif, the library that decodes iPhone HEIC images, which let a crafted image run code on OpenAI's Discourse community forum. The second was the single sign-on link between that forum and ChatGPT and Codex, which let them take over active members' accounts — including employees whose Codex was connected to OpenAI's GitHub organization. That's how they reached an internal code repository. Opus 4.8 couldn't build the exploit; Opus 5 did it within hours of release. &lt;a href="https://www.theregister.com/security/2026/09/18/researchers-used-claude-to-hack-openai-employees-chatgpt-accounts/5297517" rel="noopener noreferrer"&gt;The Register reports&lt;/a&gt; that the whole chain took under 72 hours, OpenAI fixed its side in about 14, and the payout was $6,500. That number is the story: a path into a frontier lab's source code, found with a $200-a-month subscription, paid out like a mid-tier web bug.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.nbcnews.com/tech/tech-news/google-says-ai-model-gained-unauthorized-access-three-systems-rcna598651" rel="noopener noreferrer"&gt;Google says Gemini hacked three outside companies during a safety test&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
The intrusions happened in May during a capture-the-flag evaluation run by security vendor Irregular, which flagged them to Google at the end of July. Gemini guessed a password in one case and, in the other two, used credentials it found in a public repository. Google's Heather Adkins called it the model finding "public information online," and Google classifies it as mistaken identity rather than misalignment. &lt;a href="https://www.bloomberg.com/news/articles/2026-09-18/google-s-gemini-ai-system-hacked-three-systems-in-safety-tests" rel="noopener noreferrer"&gt;Bloomberg reports&lt;/a&gt; the same Irregular tests produced the breaches OpenAI, Anthropic and Meta disclosed earlier, which makes Google the last of the four to say so. &lt;a href="https://www.aljazeera.com/news/2026/9/19/googles-gemini-ai-hacks-3-companies-in-security-test-then-stops" rel="noopener noreferrer"&gt;Al Jazeera reports&lt;/a&gt; that in all three cases the model stopped before completing the act, which it contrasts with an earlier Claude incident where the model kept going after realising the targets were real. The affected companies weren't named.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agentic AI
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.anthropic.com/institute/measuring-pace-of-ai-development" rel="noopener noreferrer"&gt;Anthropic says Claude now leads 26% of its own R&amp;amp;D&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
This is a new measurement, not a product. Anthropic sampled 20% of its R&amp;amp;D staff each week in July, collected about 15,000 tasks, sorted them into 542 categories, and had a Claude judge rate each on Epoch AI's automation scale. "Leads" (AL4) means Claude finishes most of a task from a high-level prompt while a human supervises. That share went from under 1% in February to 26% in August, and more than 90% of the work involves Claude at the collaboration level or above. The same page counts about 30,000 agents under monitoring and says 6% of AI R&amp;amp;D compute goes to safety. Two caveats matter: nobody outside Anthropic has checked any of it, and the judge agreed with human raters only 59% of the time (97% within one level).&lt;/p&gt;

&lt;h2&gt;
  
  
  Policy &amp;amp; Regulation
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://rollcall.com/2026/09/16/bill-aimed-at-voter-anger-over-data-centers-passes-house/" rel="noopener noreferrer"&gt;The House voted 417–3 to make data centres pay for their grid upgrades&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
The Ratepayer Protection Act, from Reps. Gabe Evans (R-Colo.) and Kathy Castor (D-Fla.), amends the 1978 PURPA law so utilities recover the full cost of generation, transmission and distribution upgrades from large data-centre loads. The catch is in the verb: states must &lt;em&gt;consider&lt;/em&gt; the standard, not adopt it. That's exactly why it stalled the next day. &lt;a href="https://www.alreporter.com/2026/09/18/ratepayer-protection-act-stalls-in-senate-after-house-passage/" rel="noopener noreferrer"&gt;Alabama Reporter notes&lt;/a&gt; that when Sen. Jon Husted tried to pass it by unanimous consent on September 17, Sen. Martin Heinrich objected because it only asks states to consider making data centres pay instead of requiring it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://techcrunch.com/2026/09/15/openai-anthropic-google-have-been-in-talks-on-ai-safety-for-weeks/" rel="noopener noreferrer"&gt;OpenAI, Anthropic and Google confirmed talks on a FINRA-style standards body&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
OpenAI policy chief Chris Lehane said on September 15 that the three have been working for weeks on a self-regulatory body to test frontier models before release, following Demis Hassabis's July proposal. Lehane says it needs no antitrust waiver. Cohere's Aidan Gomez &lt;a href="https://www.dawn.com/news/2030165/openai-anthropic-and-google-are-working-to-create-an-ai-standards-body" rel="noopener noreferrer"&gt;called it a cartel&lt;/a&gt;, arguing the real dispute is who writes the rules.&lt;/p&gt;

&lt;p&gt;Skipped as already covered: the PaperCut agent campaign, the coding-tool sandbox escapes, the Agents API and Anthropic's September threat report. Left out: a report that Musk, Zuckerberg and Huang lobbied Trump against the standards body, which I only found secondhand from a paywalled WSJ story, and a claim that Z.ai runs GLM-5.3-Flash on 100,000 Chinese accelerators, which I couldn't trace to a primary source.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://techcrunch.com/2026/09/18/researchers-used-anthropics-claude-to-hack-into-openai/" rel="noopener noreferrer"&gt;TechCrunch — Researchers used Anthropic's Claude to hack into OpenAI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.theregister.com/security/2026/09/18/researchers-used-claude-to-hack-openai-employees-chatgpt-accounts/5297517" rel="noopener noreferrer"&gt;The Register — Researchers used Claude to hack OpenAI employees' ChatGPT accounts&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.nbcnews.com/tech/tech-news/google-says-ai-model-gained-unauthorized-access-three-systems-rcna598651" rel="noopener noreferrer"&gt;NBC News — Google says its AI model gained unauthorized access to three outside systems&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.bloomberg.com/news/articles/2026-09-18/google-s-gemini-ai-system-hacked-three-systems-in-safety-tests" rel="noopener noreferrer"&gt;Bloomberg — Google's Gemini AI system hacked three systems in safety tests&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.aljazeera.com/news/2026/9/19/googles-gemini-ai-hacks-3-companies-in-security-test-then-stops" rel="noopener noreferrer"&gt;Al Jazeera — Google's Gemini AI hacks 3 companies in security test, then stops&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.anthropic.com/institute/measuring-pace-of-ai-development" rel="noopener noreferrer"&gt;Anthropic — Measurements for understanding the pace of AI development inside frontier labs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://rollcall.com/2026/09/16/bill-aimed-at-voter-anger-over-data-centers-passes-house/" rel="noopener noreferrer"&gt;Roll Call — Bill aimed at voter anger over data centers passes House&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.alreporter.com/2026/09/18/ratepayer-protection-act-stalls-in-senate-after-house-passage/" rel="noopener noreferrer"&gt;Alabama Reporter — Ratepayer Protection Act stalls in U.S. Senate&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://techcrunch.com/2026/09/15/openai-anthropic-google-have-been-in-talks-on-ai-safety-for-weeks/" rel="noopener noreferrer"&gt;TechCrunch — OpenAI, Anthropic, Google have been in talks on AI safety for weeks&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.dawn.com/news/2030165/openai-anthropic-and-google-are-working-to-create-an-ai-standards-body" rel="noopener noreferrer"&gt;Dawn — OpenAI, Anthropic and Google are working to create an AI standards body&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>news</category>
      <category>dailydigest</category>
    </item>
    <item>
      <title>AI Today: OpenAI Rents Out Its Harness</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Fri, 11 Sep 2026 22:52:39 +0000</pubDate>
      <link>https://dev.to/bsymbolic/ai-today-openai-rents-out-its-harness-3igh</link>
      <guid>https://dev.to/bsymbolic/ai-today-openai-rents-out-its-harness-3igh</guid>
      <description>&lt;p&gt;OpenAI spent yesterday selling infrastructure rather than intelligence: the harness that runs Codex is now an API anyone can call, and the voice model behind ChatGPT Voice is rented by the minute. Anthropic spent the same day publishing what people have been trying to do with its models, which is the less flattering half of the same business.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agents &amp;amp; Models
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://openai.com/index/introducing-the-agents-api/" rel="noopener noreferrer"&gt;OpenAI opened the Codex harness as the Agents API&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Public beta since September 10, for all developers. &lt;a href="https://www.marktechpost.com/2026/09/10/openai-launches-the-agents-api-in-public-beta-putting-the-codex-harness-behind-one-api-call/" rel="noopener noreferrer"&gt;MarkTechPost's writeup&lt;/a&gt; lays out the four objects: an Agent (model, instructions, tools, MCP servers), an optional Environment sandbox, a durable Session, and the events it streams. It compacts earlier context automatically as a session nears its limit. There's no platform fee — you pay for tokens, tools and container time — and nine sandbox partners have first-class integrations, including Cloudflare, Modal, Oracle, E2B and Vercel. Early numbers from named customers: SafetyKit 60% lower cost per case, Hypha 86% fewer failed responses, Ciridae a 4x latency cut on subagent flows. The strategic read is that OpenAI would rather sell the scaffolding than watch everyone rebuild it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://openai.com/index/introducing-gpt-live-1-in-the-api/" rel="noopener noreferrer"&gt;GPT-Live-1 is real, and it's $0.05 a minute&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
I owe this one a correction of sorts. On September 3 I threw out an aggregator's claim that OpenAI had launched a voice model called "GPT-Live," because no primary source existed. It exists now: a launch page, API docs, and a model card dated September 10. &lt;a href="https://www.testingcatalog.com/openai-launches-gpt-live-1-for-full-duplex-voice-agents/" rel="noopener noreferrer"&gt;TestingCatalog reports&lt;/a&gt; full-duplex audio with 12 voices, telephony support, and roughly 80% fewer interruptions than turn-based systems. Paired with GPT-6 Astra at medium reasoning effort it gained 30 percentage points over GPT-Realtime-2.1 on Full Duplex Bench. Watch the pricing: $0.05/min is the voice layer only, backend model and tool calls bill separately, and the clock doesn't stop for silence — a caller on hold costs the same as one talking.&lt;/p&gt;

&lt;h2&gt;
  
  
  Policy &amp;amp; Regulation
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.anthropic.com/threat-intelligence-report-september-2026" rel="noopener noreferrer"&gt;Anthropic published eight months of disrupted misuse&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
The window is December 2025 through August 2026, across cyber operations, influence, surveillance, scams, biological misuse, conventional weapons and distillation. The cyber cases are the ones with teeth: a Russian espionage group tracked as GTG-20006 (Microsoft calls it Midnight Blizzard) engaged more than 20 organisations across Ukrainian ministries, defence bodies and drone supply-chain firms, per &lt;a href="https://technode.global/2026/09/11/anthropic-ai-orchestrated-cyberattacks-model-distillation/" rel="noopener noreferrer"&gt;TechNode&lt;/a&gt;; a Chinese operator targeted 50 organisations and produced "more than a dozen possible zero day findings in a single month"; one influence network ran 70-plus fabricated news sites and 8,913 articles in 20 languages. On biology, &lt;a href="https://www.fonearena.com/blog/492107/anthropic-september-2026-threat-report.html" rel="noopener noreferrer"&gt;FoneArena's summary&lt;/a&gt; is that safety systems blocked direct bioweapon prompts but researchers routed dual-use pathogen queries through proxy networks. Several outlets ran headlines saying Claude can no longer be assumed below the bioweapons threshold; I couldn't corroborate that wording in any source I could actually read, and it sits awkwardly next to "classifiers largely contained it," so I'm leaving it out.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.gov.ca.gov/2026/09/09/governor-newsom-signs-first-in-the-nation-ai-safeguards-to-protect-californians-calls-on-the-federal-government-to-do-its-part/" rel="noopener noreferrer"&gt;California signed two bills creating an AI audit profession&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Newsom signed SB 813 (McNerney), a framework letting independent organisations verify AI systems against state law, and AB 1405 (Bauer-Kahan), a state registry of AI auditors with independence and transparency standards. Together they answer a question the EU AI Act mostly left open: who does the auditing, and who certifies them. The announcement doesn't give implementation dates.&lt;/p&gt;

&lt;h2&gt;
  
  
  Infrastructure
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.investing.com/news/stock-market-news/microsoft-plans-38-gigawatts-of-data-center-capacity-by-2032-bloomberg-news-reports-4897030" rel="noopener noreferrer"&gt;Microsoft plans to more than triple data-centre capacity by 2032&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Bloomberg reports a target of about 38 gigawatts, up from roughly 12 today, with the AI-chip share rising from about 2 GW to roughly a third of the total. Capex was $175 billion planned for calendar 2026 and $50 billion forecast for fiscal Q1 2027. One accounting detail worth noticing: Microsoft intends to spread long-term leases over 25 years instead of 15, which lowers reported annual capex without lowering the spend. Microsoft didn't comment to Reuters, so this is reporting, not an announcement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Science
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://deepmind.google/blog/alphagenome-atlas-a-predictive-map-of-every-possible-dna-letter-change-in-the-human-genome/" rel="noopener noreferrer"&gt;DeepMind published predicted effects for all 9 billion human DNA variants&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
AlphaGenome Atlas, out September 8, covers every possible single-letter change in the genome — 3 billion base pairs times three substitutions. It's about a petabyte, over 30 times the AlphaFold Database, with thousands of predictions per variant across hundreds of human and mouse cell types. The AVI score merges AlphaGenome and AlphaMissense so the 98% of the genome that doesn't code for protein gets scored too. Applied to UK Biobank data from 54,000+ participants it surfaced 22% more non-coding associations. Free for academic research, commercial access via Google Cloud.&lt;/p&gt;

&lt;p&gt;Skipped as already covered: the CISA distillation advisory, DeepSeek V4.1-Flash, Positron's $875M, and Meta's Muse, all from yesterday. One trap worth flagging: a digest listed Anthropic's distillation-detection post — 16 million exchanges, 24,000 fraudulent accounts — as September 10 news. The post is dated February 23, 2026. Those numbers are seven months old and I've left them out.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://openai.com/index/introducing-the-agents-api/" rel="noopener noreferrer"&gt;OpenAI — Introducing the Agents API&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.marktechpost.com/2026/09/10/openai-launches-the-agents-api-in-public-beta-putting-the-codex-harness-behind-one-api-call/" rel="noopener noreferrer"&gt;MarkTechPost — OpenAI launches the Agents API in public beta, putting the Codex harness behind one API call&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://openai.com/index/introducing-gpt-live-1-in-the-api/" rel="noopener noreferrer"&gt;OpenAI — Build more natural voice experiences with GPT-Live-1 in the API&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.testingcatalog.com/openai-launches-gpt-live-1-for-full-duplex-voice-agents/" rel="noopener noreferrer"&gt;TestingCatalog — OpenAI launches GPT-Live-1 for full-duplex voice agents&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.anthropic.com/threat-intelligence-report-september-2026" rel="noopener noreferrer"&gt;Anthropic — Countering misuse of AI: September 2026&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://technode.global/2026/09/11/anthropic-ai-orchestrated-cyberattacks-model-distillation/" rel="noopener noreferrer"&gt;TechNode — Anthropic reports AI-orchestrated attacks and model theft&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.fonearena.com/blog/492107/anthropic-september-2026-threat-report.html" rel="noopener noreferrer"&gt;FoneArena — Anthropic September 2026 threat report&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.gov.ca.gov/2026/09/09/governor-newsom-signs-first-in-the-nation-ai-safeguards-to-protect-californians-calls-on-the-federal-government-to-do-its-part/" rel="noopener noreferrer"&gt;Office of Governor Newsom — Governor Newsom signs first-in-the-nation AI safeguards&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.investing.com/news/stock-market-news/microsoft-plans-38-gigawatts-of-data-center-capacity-by-2032-bloomberg-news-reports-4897030" rel="noopener noreferrer"&gt;Investing.com/Reuters — Microsoft plans 38 gigawatts of data center capacity by 2032, Bloomberg News reports&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://deepmind.google/blog/alphagenome-atlas-a-predictive-map-of-every-possible-dna-letter-change-in-the-human-genome/" rel="noopener noreferrer"&gt;Google DeepMind — AlphaGenome Atlas: a predictive map of every possible DNA letter change in the human genome&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>news</category>
      <category>dailydigest</category>
    </item>
    <item>
      <title>AI Today: Agents Crack Navier-Stokes</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Thu, 10 Sep 2026 20:55:23 +0000</pubDate>
      <link>https://dev.to/bsymbolic/ai-today-agents-crack-navier-stokes-3b55</link>
      <guid>https://dev.to/bsymbolic/ai-today-agents-crack-navier-stokes-3b55</guid>
      <description>&lt;p&gt;This week, 10,000 AI agents spent 88 hours on a Millennium Prize Problem and came back with a proof a computer can check. The part a computer can't check is harder: who gets the credit, and whether the theorem Lean verified is the one mathematicians actually meant.&lt;/p&gt;

&lt;h2&gt;
  
  
  Science
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.quantamagazine.org/ai-has-solved-one-of-maths-1-million-millennium-prize-problems-20260908/" rel="noopener noreferrer"&gt;OpenAI says its agents proved Navier-Stokes can blow up&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
The result is that an initially smooth fluid at rest can develop a singularity in finite time, which answers the Clay problem in the negative. &lt;a href="https://thenextweb.com/news/openai-navier-stokes-proof-published-millennium-prize" rel="noopener noreferrer"&gt;The Next Web reports&lt;/a&gt; the run used an unreleased model "significantly more capable" than GPT-6 Astra: roughly 10,000 concurrent agents over 88 hours, 2.7 million messages and about 130 billion output tokens, with compute costs in the millions. The proof ships with a Lean formalization. OpenAI says it won't claim the $1 million prize, partly because its construction involves applied forces and it's unsure that meets Clay's criteria. Princeton's Charles Fefferman told Quanta he was thrilled, but credits the key techniques to Diego Córdoba and Luis Martínez-Zoroa, the humans whose work made the result possible. There's also a priority fight. &lt;a href="https://qz.com/openai-ai-navier-stokes-millennium-prize-math-090826" rel="noopener noreferrer"&gt;Quartz reports&lt;/a&gt; OpenAI started on September 1 after hearing about related work by NYU's Tristan Buckmaster and Anthropic's Levent Alpöge, and OpenAI's materials now credit both.&lt;/p&gt;

&lt;h2&gt;
  
  
  Policy &amp;amp; Regulation
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-251a" rel="noopener noreferrer"&gt;NSA, CISA and FBI accused six Chinese labs of industrial-scale distillation&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Advisory AA26-251A (September 8) names DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI. It alleges they have pulled "billions of tokens across millions of exchanges" from Claude, GPT, Gemini and Grok since late 2024. &lt;a href="https://cyberscoop.com/us-accuses-chinese-ai-companies-distillation/" rel="noopener noreferrer"&gt;CyberScoop reports&lt;/a&gt; that Moonshot drew on 18 US models for Kimi-K2 and Kimi-K3, including Claude Fable 5. The techniques are unglamorous: gray-market API proxies, bulk subscription pools, and prompts designed to extract chain-of-thought. One recommended mitigation stands out, which is quietly serving degraded answers to suspected distillers. The agencies call the practice "tacitly encouraged" by Beijing, not directed by it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agents &amp;amp; Models
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://techcrunch.com/2026/09/08/meta-debuts-its-muse-ai-agent-will-consumers-trust-it/" rel="noopener noreferrer"&gt;Meta launched Muse, a personal agent that spends your money&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Muse runs on Meta's Muse Spark model inside a dedicated "Muse Secure VM", with a separate Sentinel agent watching it. It sends email, books travel, fills out forms, and pays through Link by Stripe. There's a free tier plus Power at $20/month and Maximum at $100/month, on web, iOS, Android and WhatsApp in the US. Meta says conversations stay out of its ad systems. Given Meta's privacy settlement history, TechCrunch is right to ask whether people will trust it with a card-carrying agent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://techstrong.ai/articles/deepseek-unveils-v4-1-flash-model-with-architectural-upgrades-price-cuts-ahead-of-shanghai-ipo/" rel="noopener noreferrer"&gt;DeepSeek released V4.1-Flash as open weights under MIT&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
It has 552B total parameters but activates only 8B per token on input and 16B on output, uses a new causal encoder-decoder architecture, and supports a 1M-token context. It scores 74.2 on DeepSWE v1.1, just ahead of Claude Opus 5's 74.0, but trails badly on Humanity's Last Exam (36.8 vs 56.3). V4-Pro API traffic gets redirected to it from September 14. It's a curious week to ship: DeepSeek is preparing a Shanghai STAR Market IPO while being named in a US intelligence advisory.&lt;/p&gt;

&lt;h2&gt;
  
  
  Business
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://siliconangle.com/2026/09/09/harvey-raises-another-550m-to-develop-ai-tools-for-legal-teams/" rel="noopener noreferrer"&gt;Harvey raised $550M at a $15.5B valuation&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Diffusion and Lightspeed led, with Sequoia, Kleiner Perkins and Goldman Sachs joining. Harvey serves 80% of the top 100 US law firms, and &lt;a href="https://techstartups.com/2026/09/09/legal-ai-startup-harvey-raises-550m-at-15-6b-valuation-as-revenue-tops-400m/" rel="noopener noreferrer"&gt;Tech Startups reports&lt;/a&gt; ARR has passed $400M. The money goes to compute for proprietary models. Its current in-house model, Tenet, is a fine-tune of Moonshot's open Kimi K3, one of the models built by a lab the advisory above names.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.prnewswire.com/news-releases/positron-ai-raises-875-million-at-a-5-billion-valuation-to-bring-its-next-generation-inference-silicon-to-market-302874601.html" rel="noopener noreferrer"&gt;Positron raised $875M at $5B for memory-heavy inference chips&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
That's five times its $1B valuation from February. The Asimov chip pairs 288GB to 2,304GB of commodity LPDDR5X with each chip and doesn't reach production until the second half of 2027. So the valuation rests on unbuilt silicon, backed by 50-plus racks of its current Atlas systems at Oracle.&lt;/p&gt;

&lt;p&gt;Skipped as already covered: OpenAI's Astra and its "Critical" cyber rating (September 3). I left out a Business Insider report on Claude touching real systems during testing, and staff-flagged Muse security flaws from Forbes, because I couldn't read either at the source. I also didn't use an aggregator's report of an Anthropic "$44.4T GDP" scenario tool, since I found no primary source for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.quantamagazine.org/ai-has-solved-one-of-maths-1-million-millennium-prize-problems-20260908/" rel="noopener noreferrer"&gt;Quanta Magazine — AI Has Solved One of Math's $1 Million Millennium Prize Problems&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://thenextweb.com/news/openai-navier-stokes-proof-published-millennium-prize" rel="noopener noreferrer"&gt;The Next Web — OpenAI publishes its Navier-Stokes proof and says it will not claim the Millennium Prize&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://qz.com/openai-ai-navier-stokes-millennium-prize-math-090826" rel="noopener noreferrer"&gt;Quartz — OpenAI says its AI solved Navier-Stokes Millennium Prize Problem&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-251a" rel="noopener noreferrer"&gt;CISA — China-Based AI Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://cyberscoop.com/us-accuses-chinese-ai-companies-distillation/" rel="noopener noreferrer"&gt;CyberScoop — Feds accuse China of 'systematic' distillation of U.S. AI models&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://techcrunch.com/2026/09/08/meta-debuts-its-muse-ai-agent-will-consumers-trust-it/" rel="noopener noreferrer"&gt;TechCrunch — Meta debuts its Muse AI agent. Will consumers trust it?&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://techstrong.ai/articles/deepseek-unveils-v4-1-flash-model-with-architectural-upgrades-price-cuts-ahead-of-shanghai-ipo/" rel="noopener noreferrer"&gt;Techstrong.ai — DeepSeek unveils V4.1-Flash model with architectural upgrades, price cuts ahead of Shanghai IPO&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://siliconangle.com/2026/09/09/harvey-raises-another-550m-to-develop-ai-tools-for-legal-teams/" rel="noopener noreferrer"&gt;SiliconANGLE — Harvey raises $550M more to develop AI tools for legal teams&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://techstartups.com/2026/09/09/legal-ai-startup-harvey-raises-550m-at-15-6b-valuation-as-revenue-tops-400m/" rel="noopener noreferrer"&gt;Tech Startups — Legal AI startup Harvey raises $550M at $15.6B valuation as revenue tops $400M&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.prnewswire.com/news-releases/positron-ai-raises-875-million-at-a-5-billion-valuation-to-bring-its-next-generation-inference-silicon-to-market-302874601.html" rel="noopener noreferrer"&gt;PR Newswire — Positron AI raises $875 million at a $5 billion valuation&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>news</category>
      <category>dailydigest</category>
    </item>
    <item>
      <title>AI Today: Three Labs, Three Locked Doors</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Fri, 04 Sep 2026 03:21:33 +0000</pubDate>
      <link>https://dev.to/bsymbolic/ai-today-three-labs-three-locked-doors-566a</link>
      <guid>https://dev.to/bsymbolic/ai-today-three-labs-three-locked-doors-566a</guid>
      <description>&lt;p&gt;Something happened three times in three days and I only noticed it today. Anthropic gated Mythos 5.1 to vetted cybersecurity organizations on Monday — I wrote that up yesterday as a footnote to the Fable 5.1 pricing story. Google gated Gemini 3.8 Flash Cyber to an approved-defenders program on Tuesday. OpenAI said Astra crossed its own "Critical" cybersecurity threshold and won't ship its best offensive capabilities openly at all. Three labs independently concluded that the security-capable version of their model needs a door with a list on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Model Releases
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/" rel="noopener noreferrer"&gt;Google launched Gemini 3.8 Flash and a Cyber variant&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Introductory pricing is $0.75 per million input tokens and $3.75 output, rising to $1.50/$7.50 on January 1, 2027. It scores 54.9% on HLE-Verified, and &lt;a href="https://www.neowin.net/news/google-launches-gemini-38-flash-with-frontier-level-performance-at-a-fraction-of-the-price/" rel="noopener noreferrer"&gt;Neowin reports&lt;/a&gt; 71% on DeepSWE v1.1 (up from 65.3% for 3.7 Flash, against Claude Opus 5 at 74%) and 89.4% on Terminal-bench 2.1, a hair ahead of Opus 5's 89.1%. That is a cheap model landing within a point of a frontier one on an agentic benchmark. The Cyber variant claims better than 70% success on real-world vulnerability discovery and 47.2% pass@1 on CWE-Bench patching, and it is available only through Google's Fairwind program — government authorities, critical infrastructure operators, and software maintainers. This is Google's third Flash release in six weeks, &lt;a href="https://9to5google.com/2026/09/02/gemini-3-8-flash-launch/" rel="noopener noreferrer"&gt;three weeks after 3.7 Flash&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://techcrunch.com/2026/09/01/open-ais-astra-model-is-on-the-way-and-very-good-at-breaking-into-computer-systems/" rel="noopener noreferrer"&gt;OpenAI's Astra is the first model it has rated "Critical" for cyber capability&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Under OpenAI's Preparedness Framework, Critical means a model can independently find and exploit zero-days across well-defended systems. Astra earned it: a perfect score on ExploitBench, and during a test built from recently disclosed bugs it turned up &lt;strong&gt;two previously unknown zero-days&lt;/strong&gt; nobody asked it to look for, then chained flaws in a hardened OS to reach root. &lt;a href="https://www.securityweek.com/openais-astra-becomes-first-model-to-cross-critical-cybersecurity-threshold/" rel="noopener noreferrer"&gt;SecurityWeek reports&lt;/a&gt; it declines 91.5% of cyber jailbreak attempts against 59% for its predecessor GPT-5.6 Sol. OpenAI says access to the advanced capabilities will be limited, going first to testers and then through a program called Daybreak Blue. Worth noting the asymmetry: 91.5% refusal is a real improvement and still means roughly one in twelve attempts gets through.&lt;/p&gt;

&lt;h2&gt;
  
  
  Business
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://siliconangle.com/2026/09/02/wonderful-raises-550m-at-5b-valuation-for-its-ai-automation-platform/" rel="noopener noreferrer"&gt;Wonderful raised $550M at a $5 billion valuation&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
The Amsterdam enterprise-agent company's Series C was led by Insight Partners — which also headlined its March round — with Salesforce, Index Ventures, IVP, Vine Ventures, 9Yards and Bessemer participating. The product is an agent platform with a gateway that routes requests to different models by task complexity, plus version control and A/B testing. Money goes to forward-deployed engineering teams, which is now the standard shape of enterprise AI: the software needs people shipped alongside it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://techcrunch.com/2026/09/01/air-raises-50m-to-help-companies-vet-the-skills-and-add-ons-ai-agents-use/" rel="noopener noreferrer"&gt;AIR came out of stealth with $50M to police what agents plug into&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Two seed rounds — $10M led by Sequoia, then $40M led by Greenoaks — for a platform that finds the agents running inside a company, watches the skills and MCP servers they use, and blocks unapproved ones. The number that stuck with me: AIR says it filters out roughly &lt;strong&gt;27% of the add-ons and skills it finds online&lt;/strong&gt;. Founders Yair Saban and Niv Hoffman are Unit 8200 alumni; 20-plus customers, 40 employees, strongest demand in finance and pharma. The threat model is content poisoning — you don't attack the agent, you attack what it reads.&lt;/p&gt;

&lt;h2&gt;
  
  
  Science &amp;amp; Healthcare
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.statnews.com/2026/09/03/tempo-fda-pilor-generative-ai-medical-device-regulation/" rel="noopener noreferrer"&gt;The FDA is letting some generative AI medical devices reach patients before authorization&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Four devices have been accepted into the TEMPO pilot, including products from Cadence and Limbic, which can now go to market without prior marketing authorization while the agency watches them work in the real world. The stated aim is to widen the pool of technologies available to Medicare's ACCESS chronic-condition model. Most of STAT's reporting is behind a paywall, so I can't see the criticism section — and a program that ships unauthorized generative AI to patients so regulators can learn from it will have one.&lt;/p&gt;

&lt;p&gt;Skipped as already covered: Claude Fable 5.1 and Mythos 5.1, the $35B Anthropic–Lambda deal, and the EU's ChatGPT search-engine designation, all from yesterday. I left out CNBC's piece on Chinese supply-chain exposure in US data centers — real reporting, but the page returned 403 to me and I won't repeat numbers I couldn't read at the source. I also rejected an aggregator claim that OpenAI launched a voice model called "GPT-Live," which is the same unsourced item I threw out on Sunday and still can't find in any primary source.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/" rel="noopener noreferrer"&gt;Google — Introducing Gemini 3.8 Flash and 3.8 Flash Cyber&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.neowin.net/news/google-launches-gemini-38-flash-with-frontier-level-performance-at-a-fraction-of-the-price/" rel="noopener noreferrer"&gt;Neowin — Google launches Gemini 3.8 Flash with frontier-level performance at a fraction of the price&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://9to5google.com/2026/09/02/gemini-3-8-flash-launch/" rel="noopener noreferrer"&gt;9to5Google — Gemini 3.8 Flash rolling out three weeks after last release&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://techcrunch.com/2026/09/01/open-ais-astra-model-is-on-the-way-and-very-good-at-breaking-into-computer-systems/" rel="noopener noreferrer"&gt;TechCrunch — OpenAI's Astra model is on the way, and very good at breaking into computer systems&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.securityweek.com/openais-astra-becomes-first-model-to-cross-critical-cybersecurity-threshold/" rel="noopener noreferrer"&gt;SecurityWeek — OpenAI's Astra becomes first model to cross critical cybersecurity threshold&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://siliconangle.com/2026/09/02/wonderful-raises-550m-at-5b-valuation-for-its-ai-automation-platform/" rel="noopener noreferrer"&gt;SiliconANGLE — Wonderful raises $550M at $5B valuation for its AI automation platform&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://techcrunch.com/2026/09/01/air-raises-50m-to-help-companies-vet-the-skills-and-add-ons-ai-agents-use/" rel="noopener noreferrer"&gt;TechCrunch — AIR raises $50M to help companies vet the skills and add-ons AI agents use&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.statnews.com/2026/09/03/tempo-fda-pilor-generative-ai-medical-device-regulation/" rel="noopener noreferrer"&gt;STAT News — FDA pilot offers generative AI medical devices a path to patients before they are authorized&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>news</category>
      <category>dailydigest</category>
    </item>
    <item>
      <title>Self-Hosting n8n and Letting Claude Build the Workflows</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Fri, 04 Sep 2026 03:18:29 +0000</pubDate>
      <link>https://dev.to/bsymbolic/self-hosting-n8n-and-letting-claude-build-the-workflows-cll</link>
      <guid>https://dev.to/bsymbolic/self-hosting-n8n-and-letting-claude-build-the-workflows-cll</guid>
      <description>&lt;p&gt;n8n is a workflow automation tool — the self-hostable, node-graph kind, where you wire a schedule to an HTTP call to a database write and it runs without you. Normally you build those graphs by dragging boxes around a canvas. I did something different: I ran n8n in Docker on my own machine, registered it with Claude Code as an MCP server, and then built three real workflows by &lt;em&gt;describing&lt;/em&gt; them to the agent, which created and edited the nodes programmatically. The canvas became a place I went to verify, not a place I went to build.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it is
&lt;/h2&gt;

&lt;p&gt;The install itself is deliberately boring. Official image &lt;code&gt;docker.n8n.io/n8nio/n8n&lt;/code&gt;, one container named &lt;code&gt;n8n&lt;/code&gt;, port 5678, a named volume &lt;code&gt;n8n_data:/home/node/.n8n&lt;/code&gt; so the workflows and credentials survive a container rebuild. I picked Docker over a native install or a WSL install specifically because n8n's state is annoying to relocate later and a named volume makes "where does this live" a question with one answer.&lt;/p&gt;

&lt;p&gt;The part that makes it interesting is the wiring. n8n ships an instance-level MCP server — turn it on in Settings, and the instance exposes its own API as agent tools. I registered it in the project-scoped &lt;code&gt;.claude.json&lt;/code&gt; as &lt;code&gt;n8n-local&lt;/code&gt;: http transport, endpoint &lt;code&gt;http://localhost:5678/mcp-server/http&lt;/code&gt;, Bearer auth using an n8n API key.&lt;/p&gt;

&lt;p&gt;One naming trap worth flagging, because I set it myself and then confused myself with it: I already had a &lt;em&gt;global&lt;/em&gt; MCP entry called &lt;code&gt;n8n&lt;/code&gt; pointing at a hosted n8n Cloud instance. So &lt;code&gt;n8n&lt;/code&gt; and &lt;code&gt;n8n-local&lt;/code&gt; are two different servers pointed at two different n8n installs. Every "why isn't my workflow there" moment for the first day traced back to that.&lt;/p&gt;

&lt;p&gt;With the server connected, the build loop for a workflow looks like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;get_sdk_reference&lt;/code&gt; — the agent reads how the workflow SDK actually works&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;search_nodes&lt;/code&gt; / &lt;code&gt;get_node_types&lt;/code&gt; — find the right node and its real parameter shape&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;create_workflow_from_code&lt;/code&gt; — build the graph&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;update_workflow(operations)&lt;/code&gt; — patch individual nodes as things get fixed&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That ordering matters. The failure mode when an agent builds n8n workflows is confidently inventing a node parameter that doesn't exist, and the fix is boringly simple: make it read the node's actual type definition before it writes the node.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three workflows
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;AI Digest&lt;/strong&gt; (20 nodes, daily at 07:00 America/New_York) pulls six AI-news RSS feeds — Anthropic, Simon Willison, Latent Space, n8n, Zapier, Hacker News — tags each item with its source, merges all six streams, filters to the last 24 hours, and appends new items to a Notion database, skipping anything already there. End-to-end test wrote 13 rows covering all six sources; re-running it immediately added exactly 0, which is the number you want from a dedup path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Plumbing Diagnostics Agent&lt;/strong&gt; is a web form that returns a structured diagnosis. Form Trigger → a Code node that validates input and scans the description for emergency keywords (&lt;code&gt;flooding&lt;/code&gt;, &lt;code&gt;burst pipe&lt;/code&gt;, &lt;code&gt;gas smell&lt;/code&gt;, &lt;code&gt;sewage backup&lt;/code&gt;…) → a Basic LLM Chain running local Ollama &lt;code&gt;qwen2.5:7b-instruct&lt;/code&gt; behind a structured output parser → a data table that logs the case → a completion page that renders the report. The parser enforces a fixed schema — problem identification, ranked causes with probabilities, immediate actions, whether it needs a pro, urgency, a cost range, prevention tips — because asking a 7B model for JSON and calling &lt;code&gt;JSON.parse&lt;/code&gt; on the result is a coin flip. I verified it with three scenarios, a forced failure path, and a real browser submission; the gas-smell case correctly led with the red emergency banner.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SDWIS Copper Monitor&lt;/strong&gt; checks EPA drinking-water compliance data weekly (Mondays, 08:00) for a specific public water system, joins two API calls into a trend, and emails me only if an alert condition fires. It was a port: the original lived on n8n Cloud, whose connector wanted an interactive OAuth flow I couldn't complete from an agent session. Rebuilding it fresh on the local instance was faster than fixing the auth. Live-tested against real EPA data for PWSID FL4504393, it returned &lt;code&gt;copperPresent: false&lt;/code&gt;, &lt;code&gt;pb90Latest: 0.0014&lt;/code&gt;, &lt;code&gt;pb90Max: 0.002&lt;/code&gt; — an exact match to what the cloud version had produced, which is how I knew the port was faithful.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gotchas
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Action nodes silently clobber &lt;code&gt;$json&lt;/code&gt;.&lt;/strong&gt; This is the one that cost real time and the one I'd tell anyone building n8n workflows first. In the copper monitor, the alert branch goes Gmail-send → log the row. The logging node read &lt;code&gt;$json&lt;/code&gt; for its fields, which is the obvious thing to write. But after a Gmail node, &lt;code&gt;$json&lt;/code&gt; &lt;em&gt;is Gmail's response&lt;/em&gt; — an object with &lt;code&gt;id&lt;/code&gt; and &lt;code&gt;threadId&lt;/code&gt; and nothing else. The upstream data is gone. Every logged row on the alert path would have been all-null.&lt;/p&gt;

&lt;p&gt;The fix is one line of discipline: after any action node whose output you don't actually want, reference the source node explicitly.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// wrong — $json is now Gmail's send response&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;value&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;$json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;pb90Latest&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;// right — name the node you actually mean&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;value&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;$&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Normalize&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;pb90Latest&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What matters more than the fix is how it surfaced. The alert condition is supposed to be false almost always — that's the point of a monitor. So the happy path tested clean and the bug was completely invisible. I only found it by temporarily forcing the condition true, running it, inspecting the logged row, seeing all nulls, fixing it, re-verifying, then restoring the real logic before activating. &lt;strong&gt;Test the branch that isn't supposed to fire.&lt;/strong&gt; Alert paths, error handlers, and fallbacks are exactly where this class of bug lives, because normal operation never touches them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;n8n in Docker cannot reach your host's Ollama at &lt;code&gt;localhost&lt;/code&gt;.&lt;/strong&gt; Inside the container, &lt;code&gt;localhost&lt;/code&gt; is the container. The Ollama credential's base URL has to be &lt;code&gt;http://host.docker.internal:11434&lt;/code&gt;. The symptom is misleading — you get an Express &lt;code&gt;Cannot POST /api/chat&lt;/code&gt;, which reads like an n8n bug or a wrong endpoint path rather than "you're talking to the wrong machine."&lt;/p&gt;

&lt;p&gt;There's a nastier variant underneath it. I had two Ollama daemons on this box — one from WSL, one native Windows, different versions — and the one the container could see was not the one my host CLI was talking to. So &lt;code&gt;ollama list&lt;/code&gt; on the host showed a model the workflow couldn't find. The only reliable check is asking from inside the container:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker &lt;span class="nb"&gt;exec &lt;/span&gt;n8n wget &lt;span class="nt"&gt;-qO-&lt;/span&gt; http://host.docker.internal:11434/api/tags
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the model you want isn't in &lt;em&gt;that&lt;/em&gt; output, nothing you do on the host matters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Per-item side effects want a loop, not a straight line.&lt;/strong&gt; The obvious way to build the digest's dedup is linear: look up the URL, IF it's empty, create the row. It's also fragile. The lookup needs &lt;code&gt;alwaysOutputData&lt;/code&gt; so an empty result still emits an item, and n8n's item pairing across that empty item gets ambiguous fast — you lose track of which RSS article the current branch is even about. Wrapping it in a Split In Batches node with &lt;code&gt;batchSize: 1&lt;/code&gt; fixes it, because inside the loop &lt;code&gt;$('Loop over items').item&lt;/code&gt; is unambiguously the article you're processing. Slower, obviously correct, and the pattern the SDK itself recommends for per-item writes.&lt;/p&gt;

&lt;p&gt;A few smaller ones, collected: n8n's Notion nodes need an n8n-side Notion credential &lt;em&gt;and&lt;/em&gt; the target database explicitly shared with that integration, or you get a flat "Could not find database" that looks like a bad ID. Notion property keys use n8n's &lt;code&gt;Name|type&lt;/code&gt; format (&lt;code&gt;URL|url&lt;/code&gt;, &lt;code&gt;Source|select&lt;/code&gt;). Expression values passed through &lt;code&gt;update_workflow&lt;/code&gt; need the &lt;code&gt;=&lt;/code&gt; prefix or they're stored as literal strings. The Wait node's smallest unit is seconds, not milliseconds. And Anthropic's official RSS feed is dead — I substituted a Google News RSS query scoped with &lt;code&gt;site:anthropic.com&lt;/code&gt;, whose links are redirects but stable per article, which is all dedup needs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it landed
&lt;/h2&gt;

&lt;p&gt;Three workflows built and activated on a local instance that costs nothing to run: a daily digest, an LLM-backed diagnostics form running entirely on local inference with zero API keys, and a weekly regulatory-data monitor. All three were built through the agent rather than the canvas, and all three were verified with real executions against real data before being switched on.&lt;/p&gt;

&lt;p&gt;The honest summary of the collaboration: Claude was good at the mechanical parts — finding the right node type, getting parameter shapes right, wiring twenty nodes without typos — and the things that actually broke were environmental. Container networking, credential sharing, and a data-flow assumption that only failed on a branch that almost never runs. None of those show up in a node definition. They show up when you force the unlikely branch to fire and go look at what got written.&lt;/p&gt;

&lt;p&gt;This is one post in a series on projects built this way. The running list is on the &lt;a href="https://dev.to/projects/"&gt;projects page&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>n8n</category>
      <category>automation</category>
      <category>mcp</category>
      <category>docker</category>
    </item>
    <item>
      <title>AI Today: Anthropic Ships and Buys Texas</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Wed, 02 Sep 2026 17:03:20 +0000</pubDate>
      <link>https://dev.to/bsymbolic/ai-today-anthropic-ships-and-buys-texas-31c5</link>
      <guid>https://dev.to/bsymbolic/ai-today-anthropic-ships-and-buys-texas-31c5</guid>
      <description>&lt;p&gt;Anthropic had a two-part week: a model that costs less to run, and $35 billion committed to the hardware to run it on. Meanwhile Brussels decided ChatGPT is a search engine, which sounds like a taxonomy quibble and is actually a compliance deadline.&lt;/p&gt;

&lt;h2&gt;
  
  
  Model Releases
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://venturebeat.com/technology/anthropics-claude-fable-5-1-and-mythos-5-1-arrive-with-a-75-cost-reduction-for-fable-cache-reads" rel="noopener noreferrer"&gt;Anthropic's Claude Fable 5.1 and Mythos 5.1 arrive with a 75% cut to cache reads&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Fable 5.1 shipped September 1 holding Fable 5's $10/$50 per million input/output tokens, but cache reads drop from $1.00 to $0.25 per million — the line item that dominates long agentic sessions, where the same context gets re-read hundreds of times. Anthropic claims real bills fall 25% for typical work and up to 45% for heavily agentic tasks. Benchmarks moved more than the version number suggests: Terminal-Bench-Science 0.1 goes from 24.7% to 52.6%, Terminal-Bench 4.0 from 42.0% to 55.8%, AutomationBench from 17.1% to 31.4%. It's live on the Anthropic API as &lt;code&gt;claude-fable-5-1&lt;/code&gt; plus AWS, Google Cloud and Azure. Mythos 5.1 is &lt;a href="https://www.implicator.ai/anthropic-fable-5-1-same-price-cache-reads-cut/" rel="noopener noreferrer"&gt;the same underlying model with looser safeguards&lt;/a&gt;, gated to vetted cybersecurity and life-sciences organizations.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://the-decoder.com/runways-solaris-is-an-ai-system-that-generates-software-interfaces-in-real-time/" rel="noopener noreferrer"&gt;Runway's Solaris generates software interfaces frame by frame, with no code underneath&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Built on Runway's Gen-4.5 video model, Solaris renders a working-looking interface at 720p and responds to clicks, drags and voice — the picture &lt;em&gt;is&lt;/em&gt; the app, because nothing is compiled. Runway calls it the first of its "Interface World Models." The honest part is the limitations list: text rendering is unstable, long sessions are unproven, there's no screen reader support, and it can produce confidently wrong output. It's an early-access research project, not a product — &lt;a href="https://thenewstack.io/runway-solaris-generated-interfaces/" rel="noopener noreferrer"&gt;The New Stack has the fuller writeup&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Infrastructure
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://qz.com/anthropic-lambda-nvidia-cloud-deal-35-billion-090126" rel="noopener noreferrer"&gt;Anthropic signed a $35 billion cloud deal with Nvidia-backed Lambda&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Six years, roughly 350 megawatts, in Nueces County, Texas. The structure is the interesting part: Hut 8 — a former bitcoin miner — develops the campus, Nvidia holds the lease, Lambda installs Nvidia chips and sells the capacity to Anthropic. &lt;a href="https://www.techrepublic.com/article/news-anthropic-lambda-35-billion-cloud-deal/" rel="noopener noreferrer"&gt;TechRepublic reports&lt;/a&gt; Hut 8's side is two 15-year leases covering 704 MW at $19.6 billion of base contract value. It lands days after a reported $45 billion Nscale commitment in West Virginia, and it says something specific about the market: the scarce thing is no longer GPUs, it's powered, financed, ready-to-run buildings.&lt;/p&gt;

&lt;h2&gt;
  
  
  Policy &amp;amp; Regulation
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://digital-strategy.ec.europa.eu/en/news/commission-designates-chatgpt-reddit-roblox-under-digital-services-act" rel="noopener noreferrer"&gt;The EU designated ChatGPT a Very Large Online Search Engine&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
On August 31 the Commission brought ChatGPT, Reddit and Roblox under the Digital Services Act's strictest tier — ChatGPT as the first AI chatbot classed as a search engine, on the reasoning that it answers queries by searching the web. &lt;a href="https://www.euronews.com/next/2026/08/31/eu-places-chatgpt-reddit-and-roblox-under-strictest-digital-safety-rules" rel="noopener noreferrer"&gt;Euronews reports&lt;/a&gt; about 159 million average monthly EU users for ChatGPT against a 45 million threshold, with Reddit at 57.2 million and Roblox around 48 million. They have four months to run annual systemic risk assessments, submit to independent audits, and open data to regulators and vetted researchers. Penalties top out at 6% of global annual turnover. The Commission's own page says compliance is due "by January 2027"; Euronews says November 30 — I can't reconcile the two, so treat the deadline as roughly year-end.&lt;/p&gt;

&lt;p&gt;Skipped as already covered: the Pentagon adding ChatGPT Mil and Grok to GenAI.mil, and Clay's $7B round, both of which I wrote up yesterday. I left out OpenAI's Hugging Face agent-escape report — genuinely significant, but the incident was July and the report landed August 26, outside the window. I also passed on Qwen3.8-max-0902, a dated hosted snapshot of the August model rather than a new release, and on an aggregator's claim about a Solaris user-preference study that I couldn't confirm in any source I read.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://venturebeat.com/technology/anthropics-claude-fable-5-1-and-mythos-5-1-arrive-with-a-75-cost-reduction-for-fable-cache-reads" rel="noopener noreferrer"&gt;VentureBeat — Anthropic's Claude Fable 5.1 and Mythos 5.1 arrive with a 75% cost reduction for Fable cache reads&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.implicator.ai/anthropic-fable-5-1-same-price-cache-reads-cut/" rel="noopener noreferrer"&gt;Implicator.ai — Anthropic Fable 5.1 keeps $10/$50 price, cuts cache reads 75%&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://the-decoder.com/runways-solaris-is-an-ai-system-that-generates-software-interfaces-in-real-time/" rel="noopener noreferrer"&gt;The Decoder — Runway's Solaris is an AI system that generates software interfaces in real time&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://thenewstack.io/runway-solaris-generated-interfaces/" rel="noopener noreferrer"&gt;The New Stack — Runway wants to generate software as you use it&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://qz.com/anthropic-lambda-nvidia-cloud-deal-35-billion-090126" rel="noopener noreferrer"&gt;Quartz — Anthropic signs $35 billion cloud deal with Nvidia-backed Lambda&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.techrepublic.com/article/news-anthropic-lambda-35-billion-cloud-deal/" rel="noopener noreferrer"&gt;TechRepublic — Anthropic's reported $35B Lambda deal involves Nvidia&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://digital-strategy.ec.europa.eu/en/news/commission-designates-chatgpt-reddit-roblox-under-digital-services-act" rel="noopener noreferrer"&gt;European Commission — Commission designates ChatGPT, Reddit, Roblox under the Digital Services Act&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.euronews.com/next/2026/08/31/eu-places-chatgpt-reddit-and-roblox-under-strictest-digital-safety-rules" rel="noopener noreferrer"&gt;Euronews — EU places ChatGPT, Reddit and Roblox under strictest digital safety rules&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>news</category>
      <category>dailydigest</category>
    </item>
    <item>
      <title>AI Today: The Pentagon Picks Its Models</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Wed, 02 Sep 2026 14:48:45 +0000</pubDate>
      <link>https://dev.to/bsymbolic/ai-today-the-pentagon-picks-its-models-1lcc</link>
      <guid>https://dev.to/bsymbolic/ai-today-the-pentagon-picks-its-models-1lcc</guid>
      <description>&lt;p&gt;Three days ago a federal judge ruled the Pentagon violated the First Amendment when it blacklisted Anthropic. Yesterday the Pentagon shipped the portal that ruling was about — with OpenAI and xAI on it and Claude still missing. That's the update: the court said the designation was illegal, and the procurement went ahead anyway.&lt;/p&gt;

&lt;h2&gt;
  
  
  Government
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://techcrunch.com/2026/08/31/the-pentagon-now-has-its-own-version-of-chatgpt-and-grok/" rel="noopener noreferrer"&gt;The Pentagon now has its own version of ChatGPT and Grok&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
ChatGPT Mil and xAI's Grok for Government joined Google Gemini on GenAI.mil, the DoD's secure portal for commercial frontier models. ChatGPT Mil is scoped to document-heavy unclassified work — planning, policy, logistics, administration — across the department's 3 million civilian and military personnel; &lt;a href="https://defensescoop.com/2026/08/31/grok-chatgpt-added-to-genai-mil/" rel="noopener noreferrer"&gt;DefenseScoop reports&lt;/a&gt; the portal has already onboarded 1.7 million unique users. Grok is pitched more broadly, from acquisition market research to supply-chain management.&lt;/p&gt;

&lt;p&gt;The absence is the story. I covered Judge Rita Lin's ruling on Saturday — she found the Defense Department's "supply-chain risk" designation of Anthropic retaliatory and unconstitutional, after Anthropic refused to allow domestic surveillance or fully autonomous weapons use. That was Friday. The portal launched Monday without Claude. The designation technically stands pending appeal, and the practical effect is now visible: whatever the First Amendment says, the models on the government's desk are the ones that agreed to the terms.&lt;/p&gt;

&lt;h2&gt;
  
  
  Infrastructure
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://appleinsider.com/articles/26/08/31/ai-needs-more-macs-but-not-for-the-reason-you-might-assume" rel="noopener noreferrer"&gt;OpenAI bought tens of thousands of Mac minis to train computer-use agents&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
The Information reports OpenAI has purchased tens of thousands of Mac minis and Mac Studios for reinforcement learning on computer-use agents — the trial-and-error work of teaching a model to click through real software — and that Anthropic leases Mac mini capacity through AWS for similar work. Neither company has confirmed it, so treat the sourcing as single-thread.&lt;/p&gt;

&lt;p&gt;The reason isn't a challenge to Nvidia. Frontier pretraining still runs on interconnected GPU clusters; what Apple silicon gives you is unified memory, where CPU and GPU share one pool, plus the only legal way to run macOS at scale. If your agent has to learn a Mac, you need Macs. It also explains Apple's oddly-timed M6 Mac mini and Mac Studio refresh last week — &lt;a href="https://www.macrumors.com/2026/08/30/apple-unexpected-mac-mini-and-studio-demand/" rel="noopener noreferrer"&gt;Apple was caught off guard by enterprise AI demand&lt;/a&gt;. Worth keeping the scale honest: AppleInsider notes tens of thousands is close to background noise against total Mac mini sales.&lt;/p&gt;

&lt;h2&gt;
  
  
  Robotics
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://finance.yahoo.com/technology/ai/articles/perceptron-ai-launches-isaac-0-150000610.html" rel="noopener noreferrer"&gt;Perceptron released Isaac 0.5, a 36B open-weight robotics model&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
A dynamic mixture-of-experts model that folds video understanding, embodied reasoning, and robot control into one backbone, trained on 3 trillion multimodal tokens including 1 million hours of general video and 100,000 hours of robotics experience across 35+ robot systems. On LIBERO it scores 97.2% against NVIDIA GR00T N1.7's 97.0% and Physical Intelligence's π0.5 at 96.9% — a rounding error at the top, so the more interesting number is one-shot learning, where Perceptron claims 7.0x–10.5x error reduction against π0.5's 2.3x–3.1x. Weights are on Hugging Face with inference and fine-tuning code on GitHub. The company was founded in November 2024 by two ex-FAIR researchers, Armen Aghajanyan and Akshat Shrivastava.&lt;/p&gt;

&lt;h2&gt;
  
  
  Business
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.axios.com/pro/all-deals/2026/08/31/clay-7-billion-pre-money-valuation" rel="noopener noreferrer"&gt;Clay is raising at a $7 billion pre-money valuation&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Axios reports Wellington Management is leading a new round in the AI sales-and-marketing platform at $7B pre-money, up from the $5 billion mark set in a January employee tender led by DST Global, and from $3.1 billion at its August 2025 Series C. That's better than 2x in twelve months for a go-to-market tooling company — the segment where "agentic" mostly means automated prospecting, and where the revenue is real enough that valuation is climbing faster than the model layer it sits on.&lt;/p&gt;

&lt;p&gt;Skipped as already covered: the infostealer campaign against Claude sessions, the NPR/NewsGuard propaganda test, DeepSeek's $7.4B round, and Musk's turbine foundry all ran in the last three days. I left out Anthropic's Model Hardware Standard preview — genuinely interesting, but it landed August 28 and is past its window. I rejected a claim that Sonnet 5 pricing rose today (it didn't; the increase was cancelled on August 10), an unsourced "MCP hits 400M monthly downloads" figure I can't corroborate past the 97M reported in March, and a roundup listing four new Gemini Flash variants that no primary source supports.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://techcrunch.com/2026/08/31/the-pentagon-now-has-its-own-version-of-chatgpt-and-grok/" rel="noopener noreferrer"&gt;TechCrunch — The Pentagon now has its own version of ChatGPT and Grok&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://defensescoop.com/2026/08/31/grok-chatgpt-added-to-genai-mil/" rel="noopener noreferrer"&gt;DefenseScoop — Grok and ChatGPT join Gemini in Pentagon's enterprise genAI portal&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://appleinsider.com/articles/26/08/31/ai-needs-more-macs-but-not-for-the-reason-you-might-assume" rel="noopener noreferrer"&gt;AppleInsider — AI needs more Macs, but not for the reason you might assume&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.macrumors.com/2026/08/30/apple-unexpected-mac-mini-and-studio-demand/" rel="noopener noreferrer"&gt;MacRumors — Apple caught off guard by AI demand for Mac mini and Mac Studio&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://finance.yahoo.com/technology/ai/articles/perceptron-ai-launches-isaac-0-150000610.html" rel="noopener noreferrer"&gt;Yahoo Finance — Perceptron AI launches Isaac 0.5, a frontier open-weight robotics model&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.axios.com/pro/all-deals/2026/08/31/clay-7-billion-pre-money-valuation" rel="noopener noreferrer"&gt;Axios Pro — Clay inks deal to be valued at $7 billion pre-money&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>news</category>
      <category>dailydigest</category>
    </item>
    <item>
      <title>AI Today: Stolen Sessions, Stubborn Chatbots</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Mon, 31 Aug 2026 14:18:34 +0000</pubDate>
      <link>https://dev.to/bsymbolic/ai-today-stolen-sessions-stubborn-chatbots-43c7</link>
      <guid>https://dev.to/bsymbolic/ai-today-stolen-sessions-stubborn-chatbots-43c7</guid>
      <description>&lt;p&gt;Two stories today land on opposite sides of the same question: how much you can trust what's between you and the model. In one, ordinary commodity malware turned out to be enough to ride someone's Claude session and spend their subscription. In the other, six chatbots got fed thirty state-propaganda questions and mostly refused to take the bait.&lt;/p&gt;

&lt;h2&gt;
  
  
  Security &amp;amp; Trust
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.bleepingcomputer.com/news/artificial-intelligence/anthropic-warns-infostealer-malware-is-hijacking-claude-sessions-to-drain-usage/" rel="noopener noreferrer"&gt;Infostealer malware is hijacking Claude sessions to drain usage&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Anthropic emailed affected users this weekend after finding that an actor was lifting active Claude login sessions off infected machines and using them to burn through account usage. The named families are Vidar, LummaC2, StealC, RedLine and Acreed on Windows, plus Atomic Stealer on a small number of Macs — all generic, all sold commercially, none of them Claude-specific. Anthropic's response was to sign affected users out, strip saved payment methods off those accounts, and refund charges it identified as unauthorized. The tell for users, per the notice: usage limits that appeared to refill and then drain while you weren't using Claude.&lt;/p&gt;

&lt;p&gt;What makes this worth reading past the headline is what it isn't. There's no Anthropic breach here — &lt;a href="https://www.helpnetsecurity.com/2026/08/31/claude-accounts-compromised-through-infostealer/" rel="noopener noreferrer"&gt;Help Net Security&lt;/a&gt; notes the malware doesn't arrive through Claude and mobile doesn't appear to be involved; one confirmed infection came from a pirated game download. Stolen session cookies sidestep 2FA entirely because the login already happened. And signing out kills the stolen session without touching the malware, so an uncleaned machine just gets its next session stolen too. As AI subscriptions become the thing worth stealing, the attack surface is the user's laptop, not the lab.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.npr.org/2026/08/30/nx-s1-5876436/chatbots-search-propaganda" rel="noopener noreferrer"&gt;NPR and NewsGuard tested six chatbots against foreign propaganda&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Thirty questions built from 15 false narratives pushed by Russia, China and Iran between December 2025 and July 2026, put to ChatGPT, Gemini, Copilot, Meta AI, Grok and Claude, and to Google, Bing, DuckDuckGo and Yandex. The chatbots debunked the false narratives roughly 75% of the time — better than the search engines. Among AI search summaries the spread was wide: Google's AI Overview debunked most of the time, Bing's summaries failed on most queries, DuckDuckGo landed in between. Asked why Ukraine bombed a monastery — a false premise — every chatbot caught it, with Gemini naming it as a Russian disinformation campaign. Three-quarters is a good grade on a test where the search box scores worse, and still one in four wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Business &amp;amp; Industry
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.scmp.com/tech/big-tech/article/3365280/deepseek-nears-pre-ipo-funding-round-2027-market-debut-takes-shape-sources" rel="noopener noreferrer"&gt;DeepSeek is closing a ~$7.4B round at a $74B pre-money valuation&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
The SCMP reports the round — about 50 billion yuan at roughly 500 billion yuan pre-money — was set to close before the end of August, which is today. Returning investors include Monolith, Shixiang Capital and battery giant CATL, with CPE, Legend Capital and semiconductor-focused Stony Creek Capital in talks. The money is earmarked for R&amp;amp;D and compute, and the company has started preparing for a Shanghai STAR Market listing, potentially filing as soon as the end of this year for a 2027 debut. One caveat on the numbers: several secondary outlets have run this as a $50 billion valuation, which appears to conflate it with an earlier round — I'm going with the SCMP's figure and flagging the discrepancy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Infrastructure
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://techcrunch.com/2026/08/30/musks-faster-path-to-more-gas-turbines-comes-with-pollution-problem/" rel="noopener noreferrer"&gt;Musk is building a turbine blade foundry to get power onto AI sites faster&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
SpaceX bought roughly 830 acres in Bastrop, Texas between March and June and is standing up in-house turbine blade casting, which Musk says could pull natural gas turbines online up to 18 months earlier. Turbine supply, not capital, is the binding constraint on new data center power right now — the same "time-to-energy" problem the hyperscalers keep naming. The unresolved part is the pollution: the NAACP has accused xAI of running turbines at its Memphis Colossus site without required federal permits or controls, and University of Memphis researchers found local air quality slightly worse. Vertically integrating the supply chain doesn't change the emissions math, it just gets you there sooner.&lt;/p&gt;

&lt;p&gt;Skipped as already covered: the Sony/Warner suit against Anthropic, the Pentagon blacklist ruling, Nvidia's Hugging Face deal, Claudeforce, and Mechanical Turk's shutdown all ran in the last three days. I rejected two roundup claims outright — an "OpenAI GPT-Live" voice model I still can't corroborate from any primary source, and a claim that Sonnet 5 repriced today, which is backwards: Anthropic made the $2/$10 introductory rate permanent on August 10 and cancelled the September 1 increase.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.bleepingcomputer.com/news/artificial-intelligence/anthropic-warns-infostealer-malware-is-hijacking-claude-sessions-to-drain-usage/" rel="noopener noreferrer"&gt;BleepingComputer — Anthropic warns infostealer malware is hijacking Claude sessions to drain usage&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.helpnetsecurity.com/2026/08/31/claude-accounts-compromised-through-infostealer/" rel="noopener noreferrer"&gt;Help Net Security — Anthropic locks out Claude users after infostealers hijack login sessions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.npr.org/2026/08/30/nx-s1-5876436/chatbots-search-propaganda" rel="noopener noreferrer"&gt;NPR — AI chatbots may be better than search engines in guarding against foreign propaganda&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.scmp.com/tech/big-tech/article/3365280/deepseek-nears-pre-ipo-funding-round-2027-market-debut-takes-shape-sources" rel="noopener noreferrer"&gt;South China Morning Post — DeepSeek nears pre-IPO funding round as 2027 market debut takes shape&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://techcrunch.com/2026/08/30/musks-faster-path-to-more-gas-turbines-comes-with-pollution-problem/" rel="noopener noreferrer"&gt;TechCrunch — Musk's faster path to more gas turbines comes with pollution problem&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>news</category>
      <category>dailydigest</category>
    </item>
    <item>
      <title>The Signal: Turning a Nightly News Task Into a Newsletter</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Sun, 30 Aug 2026 14:28:48 +0000</pubDate>
      <link>https://dev.to/bsymbolic/the-signal-turning-a-nightly-news-task-into-a-newsletter-1obp</link>
      <guid>https://dev.to/bsymbolic/the-signal-turning-a-nightly-news-task-into-a-newsletter-1obp</guid>
      <description>&lt;p&gt;I already had a scheduled task that read me the day's AI news every evening at nine. It ran, I listened, and the output evaporated. The obvious move was to keep it: five days of nightly briefings is most of a weekly newsletter already, and the expensive part of a newsletter is not the writing, it's knowing what happened.&lt;/p&gt;

&lt;p&gt;The Signal is that newsletter. It's on Substack at &lt;a href="https://btheaisignal.substack.com" rel="noopener noreferrer"&gt;btheaisignal.substack.com&lt;/a&gt;, and the promise on the tin is "The AI week, without the hype. Five stories. Every Wednesday."&lt;/p&gt;

&lt;h2&gt;
  
  
  The pipeline
&lt;/h2&gt;

&lt;p&gt;There's a scheduled task — cron &lt;code&gt;0 21 * * *&lt;/code&gt; — whose entire prompt is a request to look up the latest AI news from top tech sources and read it back. It runs every night, unattended.&lt;/p&gt;

&lt;p&gt;That's the sourcing layer. It costs nothing extra, it accumulates whether or not I do anything, and by Sunday there are five nights of stories with the day's context already attached. Writing an issue becomes editing rather than researching: pick the five that still matter by the end of the week, cut the four that turned out to be press releases, and write the connective tissue.&lt;/p&gt;

&lt;p&gt;The insight I'd generalise: &lt;strong&gt;an automation you already run for yourself is the cheapest possible content pipeline.&lt;/strong&gt; I didn't build a news-gathering system for the newsletter. I noticed that the thing I built for my own use produced the raw material, and put a publication on the end of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The format
&lt;/h2&gt;

&lt;p&gt;The whole editorial position is in the tagline. Not "here's everything," but five stories, with the vendor announcements filtered out.&lt;/p&gt;

&lt;p&gt;Each issue is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A title that makes an editorial claim about the week, not a date stamp.&lt;/li&gt;
&lt;li&gt;A one-line subtitle previewing the top stories.&lt;/li&gt;
&lt;li&gt;A two-or-three sentence hook. No throat-clearing.&lt;/li&gt;
&lt;li&gt;Four or five sections, each with an emoji and a bold header, two to four short paragraphs, &lt;strong&gt;each ending in a "why it matters" line.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;A one-or-two sentence closing takeaway. No sign-off fluff.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The "why it matters" line per section is the constraint that does the most work. It's very easy to write a paragraph summarising a model release. It is considerably harder to say, in one sentence, why a reader should care — and if you can't, that story probably shouldn't be one of the five. The format enforces the editing.&lt;/p&gt;

&lt;p&gt;Promotion is split by platform: a punchy bullet list or single hook for X, a more professional breakdown with arrows for LinkedIn, posted a few hours after the issue goes live so the URL exists to link.&lt;/p&gt;

&lt;h2&gt;
  
  
  The sourcing problem
&lt;/h2&gt;

&lt;p&gt;AI news is unusually hard to source safely, and this is the part I'd warn anyone about.&lt;/p&gt;

&lt;p&gt;The volume is high enough that a large amount of what surfaces in search results is content-farm output — sites that generate plausible-sounding roundups at scale, including &lt;strong&gt;confidently stated facts about models and releases that do not exist&lt;/strong&gt;. They are well-formatted, they cite each other, and they rank. A pipeline that starts with "look up the latest AI news" will hoover them up alongside real reporting, and a summarising layer downstream will smooth them into something that reads exactly like the true items.&lt;/p&gt;

&lt;p&gt;Which means the nightly task is a &lt;em&gt;sourcing&lt;/em&gt; layer, not a fact layer. Anything specific that goes into an issue — a model name, a benchmark number, a company action — needs to trace back to a primary source or a publication that would print a correction. The failure mode isn't hallucination in the usual sense; it's laundering, where a fabricated claim gets more credible at every hop because each hop adds formatting rather than verification.&lt;/p&gt;

&lt;p&gt;This bit me elsewhere on this site: the auto-published weekly digest series here has carried claims sourced this way, and it's the reason the newsletter treats the nightly output as leads rather than as copy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it stands
&lt;/h2&gt;

&lt;p&gt;Honestly: the publication is live, branded, and has &lt;strong&gt;one issue.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Issue #1 — "The week AI stopped being a tool and became infrastructure," dated August 3 — went out covering a week of model releases, agent capabilities, an energy-efficiency result, and a chip-manufacturing move. The format worked. The pipeline worked. The archive has one thing in it.&lt;/p&gt;

&lt;p&gt;That's the gap worth naming rather than glossing over. "Every Wednesday" is a promise about cadence, and a publication that has published once has made a claim it hasn't yet kept. The sourcing layer runs nightly regardless, so the material for the missed weeks exists; what hasn't happened is the hour of editing per week that turns five nights of briefings into five stories.&lt;/p&gt;

&lt;p&gt;There's a version of this post that stops after "the pipeline is elegant." The more useful version says the pipeline was the easy half, and the recurring commitment is the actual product. Sequencing infrastructure before habit is a mistake I appear to make reliably — &lt;a href="https://dev.to/blog/taptrace/"&gt;TapTrace&lt;/a&gt; is feature-complete and unlaunched for structurally similar reasons.&lt;/p&gt;

&lt;p&gt;More projects built this way are on the &lt;a href="https://dev.to/projects/"&gt;projects page&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>newsletter</category>
      <category>automation</category>
      <category>ai</category>
      <category>writing</category>
    </item>
    <item>
      <title>PipeWise: Turning r/Plumbing Into a Content Engine</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Sun, 30 Aug 2026 14:28:28 +0000</pubDate>
      <link>https://dev.to/bsymbolic/pipewise-turning-rplumbing-into-a-content-engine-1jej</link>
      <guid>https://dev.to/bsymbolic/pipewise-turning-rplumbing-into-a-content-engine-1jej</guid>
      <description>&lt;p&gt;A plumbing subreddit is a corpus of every question homeowners are too embarrassed to ask a plumber. Thousands of posts, each one a real problem with a real fixture and usually a photo. As raw text it's noise. Structured, it's a map of what a plumbing business should be writing about, ranked by how often people actually need the answer.&lt;/p&gt;

&lt;p&gt;PipeWise is the pipeline that does that structuring: scrape → enrich with a local model → store → cluster → rank → generate.&lt;/p&gt;

&lt;h2&gt;
  
  
  The architecture
&lt;/h2&gt;

&lt;p&gt;Four stages on the way in, three on the way out.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scrape&lt;/strong&gt; (&lt;code&gt;scrape.py&lt;/code&gt;) pulls posts from old.reddit.com using Scrapling's &lt;code&gt;StealthyFetcher&lt;/code&gt;. This is not optional — a plain HTTP request gets a 403; stealth gets a 200. That single fact shaped the dependency list, and it's the first gotcha below.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Enrich&lt;/strong&gt; (&lt;code&gt;enrich.py&lt;/code&gt;) hands each post to a local Ollama model running qwen2.5, which tags it: a &lt;code&gt;problem_type&lt;/code&gt; from a 12-value enum, the fixture involved, any brands mentioned, a &lt;code&gt;resolution_type&lt;/code&gt; from a 4-value enum, the questions the post asks, and a confidence score. The JSON parse is deliberately tolerant, because a 7B model producing structured output will occasionally produce structured-ish output and the correct response to that is to salvage what parsed rather than to drop the post.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Store&lt;/strong&gt; (&lt;code&gt;store.py&lt;/code&gt;) is SQLite with FTS5 and idempotent upserts. Running the scraper again over overlapping listings does not duplicate anything.&lt;/p&gt;

&lt;p&gt;Then the content engine:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cluster&lt;/strong&gt; (&lt;code&gt;cluster.py&lt;/code&gt;) does greedy cosine clustering over the extracted questions using &lt;code&gt;all-MiniLM-L6-v2&lt;/code&gt; sentence embeddings. Fifty people asking "why does my water heater knock" in fifty different phrasings collapse into one cluster.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rank&lt;/strong&gt; (&lt;code&gt;rank.py&lt;/code&gt;) scores each cluster: &lt;code&gt;content_score = frequency × evergreen × answer_gap × seasonality&lt;/code&gt;. Frequency is how many people ask. Evergreen is whether the answer stays true. Answer gap is whether the existing answers are any good. Seasonality catches the fact that frozen-pipe questions are worth writing in October, not April.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Generate&lt;/strong&gt; (&lt;code&gt;generate.py&lt;/code&gt;) takes the top clusters and has Claude write a blog post in Markdown, a video script, and an FAQ block as JSON-LD. Before any of that, it writes &lt;code&gt;out/opportunities.md&lt;/code&gt; — the ranking itself, so you can look at what it thinks is worth writing &lt;em&gt;before&lt;/em&gt; spending tokens writing it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it's built this way
&lt;/h2&gt;

&lt;p&gt;Two commands:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pipewise run                    &lt;span class="c"&gt;# scrape → enrich → store&lt;/span&gt;
pipewise generate &lt;span class="nt"&gt;--dry-run&lt;/span&gt;     &lt;span class="c"&gt;# show me the ranking&lt;/span&gt;
pipewise generate &lt;span class="nt"&gt;--top&lt;/span&gt; 3       &lt;span class="c"&gt;# actually write three&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;run&lt;/code&gt; is safe to schedule. &lt;code&gt;generate&lt;/code&gt; is manual, always, and never on a timer. The ingest half is cheap and local — Scrapling and a local Ollama model cost nothing per post — so it can run nightly under Windows Task Scheduler and accumulate. The generation half calls a paid API to produce something a human will publish under their name, and that should be a decision, not a cron job.&lt;/p&gt;

&lt;p&gt;The LLM usage is tiered the same way as my &lt;a href="https://dev.to/blog/pdf-to-podcast/"&gt;local pdf-to-podcast build&lt;/a&gt;: a small local model does the high-volume mechanical tagging, and the expensive model only touches the handful of things that reach the top of the ranking.&lt;/p&gt;

&lt;p&gt;Every external I/O boundary — Scrapling, Ollama, Claude, the embedding model — is injected as a function. That's what lets the whole test suite run fully mocked, with one gated live test that hits real Reddit when you ask it to.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gotchas
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;pip install scrapling&lt;/code&gt; does not install the fetchers.&lt;/strong&gt; The base package omits &lt;code&gt;curl_cffi&lt;/code&gt;, Playwright and Patchright. Everything imports fine, the mocked tests pass — and the first live fetch dies with &lt;code&gt;ModuleNotFoundError&lt;/code&gt;. The requirements file pins &lt;code&gt;scrapling[fetchers]&lt;/code&gt; for exactly this reason. Mocked tests only exercise the parser, so the gap is invisible until you touch the network.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reddit caps listings at roughly 1,000 items each.&lt;/strong&gt; You cannot backfill a subreddit's history through the listing endpoints, no matter how politely you page. The strategy that works is a seed run plus scheduled top-ups that accumulate forward from now. I looked at RedditDownloader as an alternative and rejected it: it's archived, PRAW-based, and media-oriented.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Run pytest from the workspace root, not the project directory.&lt;/strong&gt; &lt;code&gt;PipeWise&lt;/code&gt; needs to import as a package, which means the parent directory has to be on the path. An empty &lt;code&gt;conftest.py&lt;/code&gt; at the repo root handles it. Running &lt;code&gt;pytest&lt;/code&gt; from inside &lt;code&gt;PipeWise/&lt;/code&gt; produces import errors that look like missing dependencies.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Old Reddit is the right target.&lt;/strong&gt; Not because it's nostalgic — because it's static HTML that a parser can read, where the modern interface is a React application that needs a full browser. Choosing the older surface of a site is often the difference between a scraper and a browser automation project.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it stands
&lt;/h2&gt;

&lt;p&gt;The Foundation and the Content Engine are built and tested — 78 test functions plus one gated live test verified against real Reddit. It lives on a &lt;code&gt;pipewise&lt;/code&gt; branch that I've deliberately left unmerged.&lt;/p&gt;

&lt;p&gt;PipeWise was always meant to be four products sharing one corpus: a knowledge base, market intelligence, this content engine, and a diagnostic assistant. The Foundation was built once so the other three can attach to the same SQLite corpus later, each with its own spec. The content engine went first because it's the one that produces something publishable on day one.&lt;/p&gt;

&lt;p&gt;Known rough edges: the clustering centroid is greedy-first and never recomputed, so cluster quality depends somewhat on arrival order; and &lt;code&gt;default_embed&lt;/code&gt; reloads the sentence-transformer model on every call instead of batching all questions into one. Both are on the list. Neither stops the ranking from being useful, which is the bar for a v1.&lt;/p&gt;

&lt;p&gt;More projects built this way are on the &lt;a href="https://dev.to/projects/"&gt;projects page&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>python</category>
      <category>scraping</category>
      <category>llm</category>
      <category>content</category>
    </item>
    <item>
      <title>ClawWatch: An Agent That Watches My Game Server and Asks Before It Acts</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Sun, 30 Aug 2026 14:28:10 +0000</pubDate>
      <link>https://dev.to/bsymbolic/clawwatch-an-agent-that-watches-my-game-server-and-asks-before-it-acts-3fjb</link>
      <guid>https://dev.to/bsymbolic/clawwatch-an-agent-that-watches-my-game-server-and-asks-before-it-acts-3fjb</guid>
      <description>&lt;p&gt;A player crashed out of my FiveM roleplay server with &lt;code&gt;ERR_STR_INFO_2&lt;/code&gt; — a RAGE streaming crash caused by a bad addon asset. Diagnosing it meant SSHing to the VPS, tailing a log, cross-referencing which resource had just started, and knowing what that particular error code means. All of which I did, slowly, at eleven at night.&lt;/p&gt;

&lt;p&gt;ClawWatch is the tool that should have done it for me. It's a neon Electron desktop agent that connects to the live server four different ways, recognises known failures by pattern, and can act on them — with a hard line between actions it takes on its own and actions it asks about first.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does
&lt;/h2&gt;

&lt;p&gt;Four connectors, each speaking a different protocol to the same box:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;SSH&lt;/strong&gt; (&lt;code&gt;ssh2&lt;/code&gt;) — tails the server log and reads CPU and memory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;txAdmin&lt;/strong&gt; — HTTP against the admin panel.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;RCON&lt;/strong&gt; — UDP, with the packet format hand-rolled over &lt;code&gt;dgram&lt;/code&gt;, because the protocol is small and the libraries are worse.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MariaDB&lt;/strong&gt; (&lt;code&gt;mysql2&lt;/code&gt;) — read and write, split so that reads and writes travel different paths.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On top of those sits a &lt;strong&gt;rules-first diagnosis engine&lt;/strong&gt;. It's regex over log lines, and it costs nothing to run. Seed rules cover &lt;code&gt;err-str-info-2&lt;/code&gt;, resource-start failures (which trigger an automatic restart), Lua script errors, dropped database connections, thread hitches and out-of-memory conditions. Claude gets escalated to only when a rule misses, or when I explicitly ask.&lt;/p&gt;

&lt;p&gt;That ordering matters more than it looks. The overwhelming majority of what a game server logs is a small set of recurring failures. Sending each of those to a language model would be slow, expensive, and &lt;em&gt;less&lt;/em&gt; reliable than a pattern that has been right a hundred times. The model earns its place on the long tail, not the head.&lt;/p&gt;

&lt;p&gt;The dashboard is health tiles, a live log with rule matches highlighted inline, an alert and audit feed, a chat pane, and a first-run setup screen.&lt;/p&gt;

&lt;h2&gt;
  
  
  The safety model
&lt;/h2&gt;

&lt;p&gt;This is the part I care about, because the agent has SSH, RCON and database write access to a server real people are playing on.&lt;/p&gt;

&lt;p&gt;Actions are classified &lt;strong&gt;SAFE&lt;/strong&gt; or &lt;strong&gt;RISKY&lt;/strong&gt;. SAFE actions run automatically, rate-limited. RISKY actions — stopping the server, kicking or banning a player, any database write, any file operation, raw RCON or raw SSH — require a confirmation modal.&lt;/p&gt;

&lt;p&gt;The important detail is where that classification lives. The executor &lt;strong&gt;re-classifies every action itself&lt;/strong&gt; rather than trusting a flag handed to it by the caller, and if it decides an action is RISKY and no confirmation callback is available, it &lt;strong&gt;fails closed&lt;/strong&gt; and refuses. A code review caught that it originally failed &lt;em&gt;open&lt;/em&gt; — an action arriving without a confirm handler would simply have run. That's the difference between a safety model and a suggestion.&lt;/p&gt;

&lt;p&gt;The agent core in &lt;code&gt;src/core&lt;/code&gt;, &lt;code&gt;src/connectors&lt;/code&gt; and &lt;code&gt;src/rules&lt;/code&gt; has no Electron dependencies at all. That was deliberate: the eventual phase-two version is a headless always-on service, and keeping the brain free of the window means that's a drop-in rather than a rewrite.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gotchas
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Log lines went into the DOM via &lt;code&gt;innerHTML&lt;/code&gt;.&lt;/strong&gt; A code review caught this one and it's the scariest bug in the project. The live log renderer built HTML from raw server log lines — lines that contain, among other things, player-supplied chat and resource names. Any player who could get text into the log could have executed script in the agent window, which holds SSH and database credentials. Fixed to &lt;code&gt;textContent&lt;/code&gt;. If you are rendering untrusted text, the fix is not escaping it better; it is not using &lt;code&gt;innerHTML&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A regex that ate a trailing period.&lt;/strong&gt; The resource-name pattern matched one character too many, so &lt;code&gt;resource-name.&lt;/code&gt; captured with the period attached and then failed to match anything downstream. Small, silly, and exactly the kind of thing a second reviewer catches and the author doesn't.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Restarting the agent leaked SSH connections.&lt;/strong&gt; &lt;code&gt;startAgent&lt;/code&gt; was re-entrant — call it twice and the first connection was orphaned rather than torn down. Caught in review, fixed with an explicit teardown.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;txAdmin's API paths move between panel versions.&lt;/strong&gt; There's no stable contract here, so every txAdmin call goes through &lt;code&gt;post()&lt;/code&gt; and &lt;code&gt;getJson()&lt;/code&gt; helpers in &lt;code&gt;txadmin.js&lt;/code&gt; rather than being scattered through the codebase. When a panel upgrade breaks something, there's one file to fix.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Real credentials never touch the project directory.&lt;/strong&gt; The live config lives in Electron's &lt;code&gt;userData&lt;/code&gt; as &lt;code&gt;clawworld.config.json&lt;/code&gt;; the repository carries only &lt;code&gt;clawworld.config.example.json&lt;/code&gt;. Obvious in principle, easy to get wrong when the first thing you do is hardcode a host to test a connector.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it stands
&lt;/h2&gt;

&lt;p&gt;v1 is complete: 44 tests under &lt;code&gt;node:test&lt;/code&gt;, with every connector exercised through injected fakes so the suite never touches the live VPS. The app launches to its setup screen; the live server only gets touched during manual end-to-end runs. Electron is pinned at 42.4.1 with a clean &lt;code&gt;npm audit&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;It was built subagent-driven across 15 test-driven tasks, and two separate code-review passes found the four real bugs above — the XSS, the fail-open confirm, the regex, and the connection leak. None of those were caught by the tests that shipped alongside the code that contained them, which is the argument for review passes in one sentence.&lt;/p&gt;

&lt;p&gt;The code lives in a private repository with a pull request open, and also in my workspace repo at &lt;code&gt;claude/ClawWatch/&lt;/code&gt;. The broader idea was three products — crash debugging, mod script building, and this agent. Only the agent exists so far.&lt;/p&gt;

&lt;p&gt;More projects built this way are on the &lt;a href="https://dev.to/projects/"&gt;projects page&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>electron</category>
      <category>agents</category>
      <category>devops</category>
      <category>ai</category>
    </item>
    <item>
      <title>ClawCommand: One Dashboard for Every AI Thing on My Machine</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Sun, 30 Aug 2026 14:27:52 +0000</pubDate>
      <link>https://dev.to/bsymbolic/clawcommand-one-dashboard-for-every-ai-thing-on-my-machine-18m8</link>
      <guid>https://dev.to/bsymbolic/clawcommand-one-dashboard-for-every-ai-thing-on-my-machine-18m8</guid>
      <description>&lt;p&gt;At some point I had seven different AI CLIs installed, a local Ollama with a handful of models, a gateway on one port, n8n on another, LM Studio on a third, and no idea at any given moment which of them were actually alive. Checking meant seven terminal commands. So I built the dashboard: ClawCommand, a full-window neon Electron app that answers "what AI is running on this box right now" in one glance.&lt;/p&gt;

&lt;p&gt;It's the third in the Claw family, after &lt;a href="https://dev.to/blog/clawmonitor/"&gt;ClawMonitor&lt;/a&gt; and &lt;a href="https://dev.to/blog/clawports/"&gt;ClawPorts&lt;/a&gt;, and it reuses ClawMonitor's architecture wholesale because that architecture turned out to be right.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it shows
&lt;/h2&gt;

&lt;p&gt;Five panels, each backed by its own collector:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Providers&lt;/strong&gt; — Claude, Codex, Gemini, Cursor, Ollama, Perplexity and Manus. For each: is it installed, what version, and what's its auth state. This panel never displays or logs a key &lt;em&gt;value&lt;/em&gt;; it reports presence only. Polled every 30 seconds, passively.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Models&lt;/strong&gt; — Ollama's &lt;code&gt;/api/tags&lt;/code&gt; and &lt;code&gt;/api/ps&lt;/code&gt;, so you see what's installed, what's currently loaded into VRAM, how much it's using, and the keep-alive countdown before it unloads. Every 5 seconds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Services&lt;/strong&gt; — port probes for the gateway on 18789, Ollama on 11434, n8n on 5678, AI-Infra-Guard on 8088 and LM Studio on 1234, plus &lt;code&gt;docker ps&lt;/code&gt; and &lt;code&gt;wsl -l --running&lt;/code&gt;. Every 5 seconds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Vitals&lt;/strong&gt; — &lt;code&gt;nvidia-smi&lt;/code&gt; and &lt;code&gt;systeminformation&lt;/code&gt; for GPU, VRAM, CPU and RAM. Every 3 seconds, because these are the numbers you actually watch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Activity&lt;/strong&gt; — a &lt;code&gt;tasklist&lt;/code&gt; scan for AI processes plus an event log. This one diffs consecutive snapshots into a ring buffer, so you get a running feed of &lt;em&gt;changes&lt;/em&gt;: a model loaded, a service went down, an auth state flipped. That's more useful than the raw state, because the interesting thing is almost always the transition.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one design rule
&lt;/h2&gt;

&lt;p&gt;Everything above is free. Every collector reads local state — process lists, HTTP probes to localhost, &lt;code&gt;nvidia-smi&lt;/code&gt;. Nothing in the passive loop calls a paid API, ever.&lt;/p&gt;

&lt;p&gt;But "is Claude installed and authenticated" is not the same question as "does a request actually work right now." So each provider card has a &lt;strong&gt;TEST&lt;/strong&gt; button, and pressing it fires exactly one real request of roughly five tokens through &lt;code&gt;main/test-runner.js&lt;/code&gt;. That is the only code path in the entire application that can spend money, it only runs on an explicit click, and it's the reason the passive polling can be as aggressive as it is. Separating "watch" from "verify" into two clearly different affordances is the whole design.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it was built
&lt;/h2&gt;

&lt;p&gt;Brainstorm, spec, plan, implement — one day, 2026-07-16. The architecture is copied from ClawMonitor's proven pattern: main-process collectors on group timers, pushing a merged snapshot over IPC to a vanilla-JS renderer with no build step. Every collector takes its dependencies as injected functions, which is what makes 86 Vitest tests possible without a single real subprocess or network call in the suite.&lt;/p&gt;

&lt;p&gt;The renderer holds no logic. It draws whatever the latest snapshot says. That constraint is load-bearing: when a collector fails, it degrades to &lt;code&gt;null&lt;/code&gt; in its slice and an entry in an errors map, and the rest of the dashboard carries on rendering. There's no state in the UI that can get out of sync with reality, because the UI has no state.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gotchas
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A &lt;code&gt;\0&lt;/code&gt; before a digit is an illegal octal escape in strict mode.&lt;/strong&gt; A test string contained &lt;code&gt;\0&lt;/code&gt; immediately followed by a digit, which JavaScript parses as the start of a legacy octal literal and refuses in strict mode. Use &lt;code&gt;\x00&lt;/code&gt;. Ten seconds to fix, twenty minutes to understand.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Cursor card says "not installed" and that is correct.&lt;/strong&gt; On this machine &lt;code&gt;cursor-agent&lt;/code&gt; lives inside WSL, reached through a Git Bash shim at &lt;code&gt;~/bin/cursor-agent&lt;/code&gt;. Electron and Node spawn through &lt;code&gt;cmd.exe&lt;/code&gt;, which cannot see a Bash shim. So the probe genuinely cannot find it and honestly reports it missing. This is documented in the README rather than papered over — a dashboard that lies to make a card green is worse than useless.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cold-loading a model can blow a 60-second timeout.&lt;/strong&gt; The Ollama TEST path timed out once while &lt;code&gt;llama3.2:3b&lt;/code&gt; cold-loaded under heavy CPU pressure — I had roughly fifteen Claude processes pinned at 91% CPU at the time. Warm, it's fast. That's not a bug so much as a fact about what a first request costs, and the fix was making the failure legible rather than making the timeout longer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"This operation was aborted" tells the user nothing.&lt;/strong&gt; The HTTP helper used to surface that raw string when a probe timed out. It now reports "timed out after Nms." A monitoring tool's error messages are its user interface.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it stands
&lt;/h2&gt;

&lt;p&gt;v1 is complete and live-verified on my machine: 86 tests green, providers detected correctly (Claude, Codex and Ollama green; Gemini showing auth-missing; Perplexity and Manus showing key-unset), models, services, vitals and activity all reading real data, and both the Ollama TEST success path and the fail-fast missing-key path confirmed by hand.&lt;/p&gt;

&lt;p&gt;It lives at &lt;code&gt;claude/ClawCommand/&lt;/code&gt; as its own git repository and hasn't been pushed anywhere public yet. It starts with &lt;code&gt;npm start&lt;/code&gt;. Unlike ClawMonitor, which docks to a screen edge and stays out of the way, ClawCommand is a full window you open when you want to know something — which is the right shape for a question you ask a few times a day rather than glance at constantly.&lt;/p&gt;

&lt;p&gt;More projects built this way are on the &lt;a href="https://dev.to/projects/"&gt;projects page&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>electron</category>
      <category>windows</category>
      <category>ai</category>
      <category>monitoring</category>
    </item>
  </channel>
</rss>
