<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: C. Wheatley</title>
    <description>The latest articles on DEV Community by C. Wheatley (@bsymbolic).</description>
    <link>https://dev.to/bsymbolic</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3982019%2F8beb5d78-be18-44b4-b2f8-1a9e798a54e2.jpeg</url>
      <title>DEV Community: C. Wheatley</title>
      <link>https://dev.to/bsymbolic</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/bsymbolic"/>
    <language>en</language>
    <item>
      <title>AI Today: Anthropic Ships and Buys Texas</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Wed, 02 Sep 2026 17:03:20 +0000</pubDate>
      <link>https://dev.to/bsymbolic/ai-today-anthropic-ships-and-buys-texas-31c5</link>
      <guid>https://dev.to/bsymbolic/ai-today-anthropic-ships-and-buys-texas-31c5</guid>
      <description>&lt;p&gt;Anthropic had a two-part week: a model that costs less to run, and $35 billion committed to the hardware to run it on. Meanwhile Brussels decided ChatGPT is a search engine, which sounds like a taxonomy quibble and is actually a compliance deadline.&lt;/p&gt;

&lt;h2&gt;
  
  
  Model Releases
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://venturebeat.com/technology/anthropics-claude-fable-5-1-and-mythos-5-1-arrive-with-a-75-cost-reduction-for-fable-cache-reads" rel="noopener noreferrer"&gt;Anthropic's Claude Fable 5.1 and Mythos 5.1 arrive with a 75% cut to cache reads&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Fable 5.1 shipped September 1 holding Fable 5's $10/$50 per million input/output tokens, but cache reads drop from $1.00 to $0.25 per million — the line item that dominates long agentic sessions, where the same context gets re-read hundreds of times. Anthropic claims real bills fall 25% for typical work and up to 45% for heavily agentic tasks. Benchmarks moved more than the version number suggests: Terminal-Bench-Science 0.1 goes from 24.7% to 52.6%, Terminal-Bench 4.0 from 42.0% to 55.8%, AutomationBench from 17.1% to 31.4%. It's live on the Anthropic API as &lt;code&gt;claude-fable-5-1&lt;/code&gt; plus AWS, Google Cloud and Azure. Mythos 5.1 is &lt;a href="https://www.implicator.ai/anthropic-fable-5-1-same-price-cache-reads-cut/" rel="noopener noreferrer"&gt;the same underlying model with looser safeguards&lt;/a&gt;, gated to vetted cybersecurity and life-sciences organizations.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://the-decoder.com/runways-solaris-is-an-ai-system-that-generates-software-interfaces-in-real-time/" rel="noopener noreferrer"&gt;Runway's Solaris generates software interfaces frame by frame, with no code underneath&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Built on Runway's Gen-4.5 video model, Solaris renders a working-looking interface at 720p and responds to clicks, drags and voice — the picture &lt;em&gt;is&lt;/em&gt; the app, because nothing is compiled. Runway calls it the first of its "Interface World Models." The honest part is the limitations list: text rendering is unstable, long sessions are unproven, there's no screen reader support, and it can produce confidently wrong output. It's an early-access research project, not a product — &lt;a href="https://thenewstack.io/runway-solaris-generated-interfaces/" rel="noopener noreferrer"&gt;The New Stack has the fuller writeup&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Infrastructure
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://qz.com/anthropic-lambda-nvidia-cloud-deal-35-billion-090126" rel="noopener noreferrer"&gt;Anthropic signed a $35 billion cloud deal with Nvidia-backed Lambda&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Six years, roughly 350 megawatts, in Nueces County, Texas. The structure is the interesting part: Hut 8 — a former bitcoin miner — develops the campus, Nvidia holds the lease, Lambda installs Nvidia chips and sells the capacity to Anthropic. &lt;a href="https://www.techrepublic.com/article/news-anthropic-lambda-35-billion-cloud-deal/" rel="noopener noreferrer"&gt;TechRepublic reports&lt;/a&gt; Hut 8's side is two 15-year leases covering 704 MW at $19.6 billion of base contract value. It lands days after a reported $45 billion Nscale commitment in West Virginia, and it says something specific about the market: the scarce thing is no longer GPUs, it's powered, financed, ready-to-run buildings.&lt;/p&gt;

&lt;h2&gt;
  
  
  Policy &amp;amp; Regulation
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://digital-strategy.ec.europa.eu/en/news/commission-designates-chatgpt-reddit-roblox-under-digital-services-act" rel="noopener noreferrer"&gt;The EU designated ChatGPT a Very Large Online Search Engine&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
On August 31 the Commission brought ChatGPT, Reddit and Roblox under the Digital Services Act's strictest tier — ChatGPT as the first AI chatbot classed as a search engine, on the reasoning that it answers queries by searching the web. &lt;a href="https://www.euronews.com/next/2026/08/31/eu-places-chatgpt-reddit-and-roblox-under-strictest-digital-safety-rules" rel="noopener noreferrer"&gt;Euronews reports&lt;/a&gt; about 159 million average monthly EU users for ChatGPT against a 45 million threshold, with Reddit at 57.2 million and Roblox around 48 million. They have four months to run annual systemic risk assessments, submit to independent audits, and open data to regulators and vetted researchers. Penalties top out at 6% of global annual turnover. The Commission's own page says compliance is due "by January 2027"; Euronews says November 30 — I can't reconcile the two, so treat the deadline as roughly year-end.&lt;/p&gt;

&lt;p&gt;Skipped as already covered: the Pentagon adding ChatGPT Mil and Grok to GenAI.mil, and Clay's $7B round, both of which I wrote up yesterday. I left out OpenAI's Hugging Face agent-escape report — genuinely significant, but the incident was July and the report landed August 26, outside the window. I also passed on Qwen3.8-max-0902, a dated hosted snapshot of the August model rather than a new release, and on an aggregator's claim about a Solaris user-preference study that I couldn't confirm in any source I read.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://venturebeat.com/technology/anthropics-claude-fable-5-1-and-mythos-5-1-arrive-with-a-75-cost-reduction-for-fable-cache-reads" rel="noopener noreferrer"&gt;VentureBeat — Anthropic's Claude Fable 5.1 and Mythos 5.1 arrive with a 75% cost reduction for Fable cache reads&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.implicator.ai/anthropic-fable-5-1-same-price-cache-reads-cut/" rel="noopener noreferrer"&gt;Implicator.ai — Anthropic Fable 5.1 keeps $10/$50 price, cuts cache reads 75%&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://the-decoder.com/runways-solaris-is-an-ai-system-that-generates-software-interfaces-in-real-time/" rel="noopener noreferrer"&gt;The Decoder — Runway's Solaris is an AI system that generates software interfaces in real time&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://thenewstack.io/runway-solaris-generated-interfaces/" rel="noopener noreferrer"&gt;The New Stack — Runway wants to generate software as you use it&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://qz.com/anthropic-lambda-nvidia-cloud-deal-35-billion-090126" rel="noopener noreferrer"&gt;Quartz — Anthropic signs $35 billion cloud deal with Nvidia-backed Lambda&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.techrepublic.com/article/news-anthropic-lambda-35-billion-cloud-deal/" rel="noopener noreferrer"&gt;TechRepublic — Anthropic's reported $35B Lambda deal involves Nvidia&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://digital-strategy.ec.europa.eu/en/news/commission-designates-chatgpt-reddit-roblox-under-digital-services-act" rel="noopener noreferrer"&gt;European Commission — Commission designates ChatGPT, Reddit, Roblox under the Digital Services Act&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.euronews.com/next/2026/08/31/eu-places-chatgpt-reddit-and-roblox-under-strictest-digital-safety-rules" rel="noopener noreferrer"&gt;Euronews — EU places ChatGPT, Reddit and Roblox under strictest digital safety rules&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>news</category>
      <category>dailydigest</category>
    </item>
    <item>
      <title>AI Today: The Pentagon Picks Its Models</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Wed, 02 Sep 2026 14:48:45 +0000</pubDate>
      <link>https://dev.to/bsymbolic/ai-today-the-pentagon-picks-its-models-1lcc</link>
      <guid>https://dev.to/bsymbolic/ai-today-the-pentagon-picks-its-models-1lcc</guid>
      <description>&lt;p&gt;Three days ago a federal judge ruled the Pentagon violated the First Amendment when it blacklisted Anthropic. Yesterday the Pentagon shipped the portal that ruling was about — with OpenAI and xAI on it and Claude still missing. That's the update: the court said the designation was illegal, and the procurement went ahead anyway.&lt;/p&gt;

&lt;h2&gt;
  
  
  Government
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://techcrunch.com/2026/08/31/the-pentagon-now-has-its-own-version-of-chatgpt-and-grok/" rel="noopener noreferrer"&gt;The Pentagon now has its own version of ChatGPT and Grok&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
ChatGPT Mil and xAI's Grok for Government joined Google Gemini on GenAI.mil, the DoD's secure portal for commercial frontier models. ChatGPT Mil is scoped to document-heavy unclassified work — planning, policy, logistics, administration — across the department's 3 million civilian and military personnel; &lt;a href="https://defensescoop.com/2026/08/31/grok-chatgpt-added-to-genai-mil/" rel="noopener noreferrer"&gt;DefenseScoop reports&lt;/a&gt; the portal has already onboarded 1.7 million unique users. Grok is pitched more broadly, from acquisition market research to supply-chain management.&lt;/p&gt;

&lt;p&gt;The absence is the story. I covered Judge Rita Lin's ruling on Saturday — she found the Defense Department's "supply-chain risk" designation of Anthropic retaliatory and unconstitutional, after Anthropic refused to allow domestic surveillance or fully autonomous weapons use. That was Friday. The portal launched Monday without Claude. The designation technically stands pending appeal, and the practical effect is now visible: whatever the First Amendment says, the models on the government's desk are the ones that agreed to the terms.&lt;/p&gt;

&lt;h2&gt;
  
  
  Infrastructure
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://appleinsider.com/articles/26/08/31/ai-needs-more-macs-but-not-for-the-reason-you-might-assume" rel="noopener noreferrer"&gt;OpenAI bought tens of thousands of Mac minis to train computer-use agents&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
The Information reports OpenAI has purchased tens of thousands of Mac minis and Mac Studios for reinforcement learning on computer-use agents — the trial-and-error work of teaching a model to click through real software — and that Anthropic leases Mac mini capacity through AWS for similar work. Neither company has confirmed it, so treat the sourcing as single-thread.&lt;/p&gt;

&lt;p&gt;The reason isn't a challenge to Nvidia. Frontier pretraining still runs on interconnected GPU clusters; what Apple silicon gives you is unified memory, where CPU and GPU share one pool, plus the only legal way to run macOS at scale. If your agent has to learn a Mac, you need Macs. It also explains Apple's oddly-timed M6 Mac mini and Mac Studio refresh last week — &lt;a href="https://www.macrumors.com/2026/08/30/apple-unexpected-mac-mini-and-studio-demand/" rel="noopener noreferrer"&gt;Apple was caught off guard by enterprise AI demand&lt;/a&gt;. Worth keeping the scale honest: AppleInsider notes tens of thousands is close to background noise against total Mac mini sales.&lt;/p&gt;

&lt;h2&gt;
  
  
  Robotics
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://finance.yahoo.com/technology/ai/articles/perceptron-ai-launches-isaac-0-150000610.html" rel="noopener noreferrer"&gt;Perceptron released Isaac 0.5, a 36B open-weight robotics model&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
A dynamic mixture-of-experts model that folds video understanding, embodied reasoning, and robot control into one backbone, trained on 3 trillion multimodal tokens including 1 million hours of general video and 100,000 hours of robotics experience across 35+ robot systems. On LIBERO it scores 97.2% against NVIDIA GR00T N1.7's 97.0% and Physical Intelligence's π0.5 at 96.9% — a rounding error at the top, so the more interesting number is one-shot learning, where Perceptron claims 7.0x–10.5x error reduction against π0.5's 2.3x–3.1x. Weights are on Hugging Face with inference and fine-tuning code on GitHub. The company was founded in November 2024 by two ex-FAIR researchers, Armen Aghajanyan and Akshat Shrivastava.&lt;/p&gt;

&lt;h2&gt;
  
  
  Business
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.axios.com/pro/all-deals/2026/08/31/clay-7-billion-pre-money-valuation" rel="noopener noreferrer"&gt;Clay is raising at a $7 billion pre-money valuation&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Axios reports Wellington Management is leading a new round in the AI sales-and-marketing platform at $7B pre-money, up from the $5 billion mark set in a January employee tender led by DST Global, and from $3.1 billion at its August 2025 Series C. That's better than 2x in twelve months for a go-to-market tooling company — the segment where "agentic" mostly means automated prospecting, and where the revenue is real enough that valuation is climbing faster than the model layer it sits on.&lt;/p&gt;

&lt;p&gt;Skipped as already covered: the infostealer campaign against Claude sessions, the NPR/NewsGuard propaganda test, DeepSeek's $7.4B round, and Musk's turbine foundry all ran in the last three days. I left out Anthropic's Model Hardware Standard preview — genuinely interesting, but it landed August 28 and is past its window. I rejected a claim that Sonnet 5 pricing rose today (it didn't; the increase was cancelled on August 10), an unsourced "MCP hits 400M monthly downloads" figure I can't corroborate past the 97M reported in March, and a roundup listing four new Gemini Flash variants that no primary source supports.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://techcrunch.com/2026/08/31/the-pentagon-now-has-its-own-version-of-chatgpt-and-grok/" rel="noopener noreferrer"&gt;TechCrunch — The Pentagon now has its own version of ChatGPT and Grok&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://defensescoop.com/2026/08/31/grok-chatgpt-added-to-genai-mil/" rel="noopener noreferrer"&gt;DefenseScoop — Grok and ChatGPT join Gemini in Pentagon's enterprise genAI portal&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://appleinsider.com/articles/26/08/31/ai-needs-more-macs-but-not-for-the-reason-you-might-assume" rel="noopener noreferrer"&gt;AppleInsider — AI needs more Macs, but not for the reason you might assume&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.macrumors.com/2026/08/30/apple-unexpected-mac-mini-and-studio-demand/" rel="noopener noreferrer"&gt;MacRumors — Apple caught off guard by AI demand for Mac mini and Mac Studio&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://finance.yahoo.com/technology/ai/articles/perceptron-ai-launches-isaac-0-150000610.html" rel="noopener noreferrer"&gt;Yahoo Finance — Perceptron AI launches Isaac 0.5, a frontier open-weight robotics model&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.axios.com/pro/all-deals/2026/08/31/clay-7-billion-pre-money-valuation" rel="noopener noreferrer"&gt;Axios Pro — Clay inks deal to be valued at $7 billion pre-money&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>news</category>
      <category>dailydigest</category>
    </item>
    <item>
      <title>AI Today: Stolen Sessions, Stubborn Chatbots</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Mon, 31 Aug 2026 14:18:34 +0000</pubDate>
      <link>https://dev.to/bsymbolic/ai-today-stolen-sessions-stubborn-chatbots-43c7</link>
      <guid>https://dev.to/bsymbolic/ai-today-stolen-sessions-stubborn-chatbots-43c7</guid>
      <description>&lt;p&gt;Two stories today land on opposite sides of the same question: how much you can trust what's between you and the model. In one, ordinary commodity malware turned out to be enough to ride someone's Claude session and spend their subscription. In the other, six chatbots got fed thirty state-propaganda questions and mostly refused to take the bait.&lt;/p&gt;

&lt;h2&gt;
  
  
  Security &amp;amp; Trust
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.bleepingcomputer.com/news/artificial-intelligence/anthropic-warns-infostealer-malware-is-hijacking-claude-sessions-to-drain-usage/" rel="noopener noreferrer"&gt;Infostealer malware is hijacking Claude sessions to drain usage&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Anthropic emailed affected users this weekend after finding that an actor was lifting active Claude login sessions off infected machines and using them to burn through account usage. The named families are Vidar, LummaC2, StealC, RedLine and Acreed on Windows, plus Atomic Stealer on a small number of Macs — all generic, all sold commercially, none of them Claude-specific. Anthropic's response was to sign affected users out, strip saved payment methods off those accounts, and refund charges it identified as unauthorized. The tell for users, per the notice: usage limits that appeared to refill and then drain while you weren't using Claude.&lt;/p&gt;

&lt;p&gt;What makes this worth reading past the headline is what it isn't. There's no Anthropic breach here — &lt;a href="https://www.helpnetsecurity.com/2026/08/31/claude-accounts-compromised-through-infostealer/" rel="noopener noreferrer"&gt;Help Net Security&lt;/a&gt; notes the malware doesn't arrive through Claude and mobile doesn't appear to be involved; one confirmed infection came from a pirated game download. Stolen session cookies sidestep 2FA entirely because the login already happened. And signing out kills the stolen session without touching the malware, so an uncleaned machine just gets its next session stolen too. As AI subscriptions become the thing worth stealing, the attack surface is the user's laptop, not the lab.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.npr.org/2026/08/30/nx-s1-5876436/chatbots-search-propaganda" rel="noopener noreferrer"&gt;NPR and NewsGuard tested six chatbots against foreign propaganda&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Thirty questions built from 15 false narratives pushed by Russia, China and Iran between December 2025 and July 2026, put to ChatGPT, Gemini, Copilot, Meta AI, Grok and Claude, and to Google, Bing, DuckDuckGo and Yandex. The chatbots debunked the false narratives roughly 75% of the time — better than the search engines. Among AI search summaries the spread was wide: Google's AI Overview debunked most of the time, Bing's summaries failed on most queries, DuckDuckGo landed in between. Asked why Ukraine bombed a monastery — a false premise — every chatbot caught it, with Gemini naming it as a Russian disinformation campaign. Three-quarters is a good grade on a test where the search box scores worse, and still one in four wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Business &amp;amp; Industry
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.scmp.com/tech/big-tech/article/3365280/deepseek-nears-pre-ipo-funding-round-2027-market-debut-takes-shape-sources" rel="noopener noreferrer"&gt;DeepSeek is closing a ~$7.4B round at a $74B pre-money valuation&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
The SCMP reports the round — about 50 billion yuan at roughly 500 billion yuan pre-money — was set to close before the end of August, which is today. Returning investors include Monolith, Shixiang Capital and battery giant CATL, with CPE, Legend Capital and semiconductor-focused Stony Creek Capital in talks. The money is earmarked for R&amp;amp;D and compute, and the company has started preparing for a Shanghai STAR Market listing, potentially filing as soon as the end of this year for a 2027 debut. One caveat on the numbers: several secondary outlets have run this as a $50 billion valuation, which appears to conflate it with an earlier round — I'm going with the SCMP's figure and flagging the discrepancy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Infrastructure
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://techcrunch.com/2026/08/30/musks-faster-path-to-more-gas-turbines-comes-with-pollution-problem/" rel="noopener noreferrer"&gt;Musk is building a turbine blade foundry to get power onto AI sites faster&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
SpaceX bought roughly 830 acres in Bastrop, Texas between March and June and is standing up in-house turbine blade casting, which Musk says could pull natural gas turbines online up to 18 months earlier. Turbine supply, not capital, is the binding constraint on new data center power right now — the same "time-to-energy" problem the hyperscalers keep naming. The unresolved part is the pollution: the NAACP has accused xAI of running turbines at its Memphis Colossus site without required federal permits or controls, and University of Memphis researchers found local air quality slightly worse. Vertically integrating the supply chain doesn't change the emissions math, it just gets you there sooner.&lt;/p&gt;

&lt;p&gt;Skipped as already covered: the Sony/Warner suit against Anthropic, the Pentagon blacklist ruling, Nvidia's Hugging Face deal, Claudeforce, and Mechanical Turk's shutdown all ran in the last three days. I rejected two roundup claims outright — an "OpenAI GPT-Live" voice model I still can't corroborate from any primary source, and a claim that Sonnet 5 repriced today, which is backwards: Anthropic made the $2/$10 introductory rate permanent on August 10 and cancelled the September 1 increase.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.bleepingcomputer.com/news/artificial-intelligence/anthropic-warns-infostealer-malware-is-hijacking-claude-sessions-to-drain-usage/" rel="noopener noreferrer"&gt;BleepingComputer — Anthropic warns infostealer malware is hijacking Claude sessions to drain usage&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.helpnetsecurity.com/2026/08/31/claude-accounts-compromised-through-infostealer/" rel="noopener noreferrer"&gt;Help Net Security — Anthropic locks out Claude users after infostealers hijack login sessions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.npr.org/2026/08/30/nx-s1-5876436/chatbots-search-propaganda" rel="noopener noreferrer"&gt;NPR — AI chatbots may be better than search engines in guarding against foreign propaganda&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.scmp.com/tech/big-tech/article/3365280/deepseek-nears-pre-ipo-funding-round-2027-market-debut-takes-shape-sources" rel="noopener noreferrer"&gt;South China Morning Post — DeepSeek nears pre-IPO funding round as 2027 market debut takes shape&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://techcrunch.com/2026/08/30/musks-faster-path-to-more-gas-turbines-comes-with-pollution-problem/" rel="noopener noreferrer"&gt;TechCrunch — Musk's faster path to more gas turbines comes with pollution problem&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>news</category>
      <category>dailydigest</category>
    </item>
    <item>
      <title>The Signal: Turning a Nightly News Task Into a Newsletter</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Sun, 30 Aug 2026 14:28:48 +0000</pubDate>
      <link>https://dev.to/bsymbolic/the-signal-turning-a-nightly-news-task-into-a-newsletter-1obp</link>
      <guid>https://dev.to/bsymbolic/the-signal-turning-a-nightly-news-task-into-a-newsletter-1obp</guid>
      <description>&lt;p&gt;I already had a scheduled task that read me the day's AI news every evening at nine. It ran, I listened, and the output evaporated. The obvious move was to keep it: five days of nightly briefings is most of a weekly newsletter already, and the expensive part of a newsletter is not the writing, it's knowing what happened.&lt;/p&gt;

&lt;p&gt;The Signal is that newsletter. It's on Substack at &lt;a href="https://btheaisignal.substack.com" rel="noopener noreferrer"&gt;btheaisignal.substack.com&lt;/a&gt;, and the promise on the tin is "The AI week, without the hype. Five stories. Every Wednesday."&lt;/p&gt;

&lt;h2&gt;
  
  
  The pipeline
&lt;/h2&gt;

&lt;p&gt;There's a scheduled task — cron &lt;code&gt;0 21 * * *&lt;/code&gt; — whose entire prompt is a request to look up the latest AI news from top tech sources and read it back. It runs every night, unattended.&lt;/p&gt;

&lt;p&gt;That's the sourcing layer. It costs nothing extra, it accumulates whether or not I do anything, and by Sunday there are five nights of stories with the day's context already attached. Writing an issue becomes editing rather than researching: pick the five that still matter by the end of the week, cut the four that turned out to be press releases, and write the connective tissue.&lt;/p&gt;

&lt;p&gt;The insight I'd generalise: &lt;strong&gt;an automation you already run for yourself is the cheapest possible content pipeline.&lt;/strong&gt; I didn't build a news-gathering system for the newsletter. I noticed that the thing I built for my own use produced the raw material, and put a publication on the end of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The format
&lt;/h2&gt;

&lt;p&gt;The whole editorial position is in the tagline. Not "here's everything," but five stories, with the vendor announcements filtered out.&lt;/p&gt;

&lt;p&gt;Each issue is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A title that makes an editorial claim about the week, not a date stamp.&lt;/li&gt;
&lt;li&gt;A one-line subtitle previewing the top stories.&lt;/li&gt;
&lt;li&gt;A two-or-three sentence hook. No throat-clearing.&lt;/li&gt;
&lt;li&gt;Four or five sections, each with an emoji and a bold header, two to four short paragraphs, &lt;strong&gt;each ending in a "why it matters" line.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;A one-or-two sentence closing takeaway. No sign-off fluff.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The "why it matters" line per section is the constraint that does the most work. It's very easy to write a paragraph summarising a model release. It is considerably harder to say, in one sentence, why a reader should care — and if you can't, that story probably shouldn't be one of the five. The format enforces the editing.&lt;/p&gt;

&lt;p&gt;Promotion is split by platform: a punchy bullet list or single hook for X, a more professional breakdown with arrows for LinkedIn, posted a few hours after the issue goes live so the URL exists to link.&lt;/p&gt;

&lt;h2&gt;
  
  
  The sourcing problem
&lt;/h2&gt;

&lt;p&gt;AI news is unusually hard to source safely, and this is the part I'd warn anyone about.&lt;/p&gt;

&lt;p&gt;The volume is high enough that a large amount of what surfaces in search results is content-farm output — sites that generate plausible-sounding roundups at scale, including &lt;strong&gt;confidently stated facts about models and releases that do not exist&lt;/strong&gt;. They are well-formatted, they cite each other, and they rank. A pipeline that starts with "look up the latest AI news" will hoover them up alongside real reporting, and a summarising layer downstream will smooth them into something that reads exactly like the true items.&lt;/p&gt;

&lt;p&gt;Which means the nightly task is a &lt;em&gt;sourcing&lt;/em&gt; layer, not a fact layer. Anything specific that goes into an issue — a model name, a benchmark number, a company action — needs to trace back to a primary source or a publication that would print a correction. The failure mode isn't hallucination in the usual sense; it's laundering, where a fabricated claim gets more credible at every hop because each hop adds formatting rather than verification.&lt;/p&gt;

&lt;p&gt;This bit me elsewhere on this site: the auto-published weekly digest series here has carried claims sourced this way, and it's the reason the newsletter treats the nightly output as leads rather than as copy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it stands
&lt;/h2&gt;

&lt;p&gt;Honestly: the publication is live, branded, and has &lt;strong&gt;one issue.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Issue #1 — "The week AI stopped being a tool and became infrastructure," dated August 3 — went out covering a week of model releases, agent capabilities, an energy-efficiency result, and a chip-manufacturing move. The format worked. The pipeline worked. The archive has one thing in it.&lt;/p&gt;

&lt;p&gt;That's the gap worth naming rather than glossing over. "Every Wednesday" is a promise about cadence, and a publication that has published once has made a claim it hasn't yet kept. The sourcing layer runs nightly regardless, so the material for the missed weeks exists; what hasn't happened is the hour of editing per week that turns five nights of briefings into five stories.&lt;/p&gt;

&lt;p&gt;There's a version of this post that stops after "the pipeline is elegant." The more useful version says the pipeline was the easy half, and the recurring commitment is the actual product. Sequencing infrastructure before habit is a mistake I appear to make reliably — &lt;a href="https://dev.to/blog/taptrace/"&gt;TapTrace&lt;/a&gt; is feature-complete and unlaunched for structurally similar reasons.&lt;/p&gt;

&lt;p&gt;More projects built this way are on the &lt;a href="https://dev.to/projects/"&gt;projects page&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>newsletter</category>
      <category>automation</category>
      <category>ai</category>
      <category>writing</category>
    </item>
    <item>
      <title>PipeWise: Turning r/Plumbing Into a Content Engine</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Sun, 30 Aug 2026 14:28:28 +0000</pubDate>
      <link>https://dev.to/bsymbolic/pipewise-turning-rplumbing-into-a-content-engine-1jej</link>
      <guid>https://dev.to/bsymbolic/pipewise-turning-rplumbing-into-a-content-engine-1jej</guid>
      <description>&lt;p&gt;A plumbing subreddit is a corpus of every question homeowners are too embarrassed to ask a plumber. Thousands of posts, each one a real problem with a real fixture and usually a photo. As raw text it's noise. Structured, it's a map of what a plumbing business should be writing about, ranked by how often people actually need the answer.&lt;/p&gt;

&lt;p&gt;PipeWise is the pipeline that does that structuring: scrape → enrich with a local model → store → cluster → rank → generate.&lt;/p&gt;

&lt;h2&gt;
  
  
  The architecture
&lt;/h2&gt;

&lt;p&gt;Four stages on the way in, three on the way out.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scrape&lt;/strong&gt; (&lt;code&gt;scrape.py&lt;/code&gt;) pulls posts from old.reddit.com using Scrapling's &lt;code&gt;StealthyFetcher&lt;/code&gt;. This is not optional — a plain HTTP request gets a 403; stealth gets a 200. That single fact shaped the dependency list, and it's the first gotcha below.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Enrich&lt;/strong&gt; (&lt;code&gt;enrich.py&lt;/code&gt;) hands each post to a local Ollama model running qwen2.5, which tags it: a &lt;code&gt;problem_type&lt;/code&gt; from a 12-value enum, the fixture involved, any brands mentioned, a &lt;code&gt;resolution_type&lt;/code&gt; from a 4-value enum, the questions the post asks, and a confidence score. The JSON parse is deliberately tolerant, because a 7B model producing structured output will occasionally produce structured-ish output and the correct response to that is to salvage what parsed rather than to drop the post.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Store&lt;/strong&gt; (&lt;code&gt;store.py&lt;/code&gt;) is SQLite with FTS5 and idempotent upserts. Running the scraper again over overlapping listings does not duplicate anything.&lt;/p&gt;

&lt;p&gt;Then the content engine:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cluster&lt;/strong&gt; (&lt;code&gt;cluster.py&lt;/code&gt;) does greedy cosine clustering over the extracted questions using &lt;code&gt;all-MiniLM-L6-v2&lt;/code&gt; sentence embeddings. Fifty people asking "why does my water heater knock" in fifty different phrasings collapse into one cluster.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rank&lt;/strong&gt; (&lt;code&gt;rank.py&lt;/code&gt;) scores each cluster: &lt;code&gt;content_score = frequency × evergreen × answer_gap × seasonality&lt;/code&gt;. Frequency is how many people ask. Evergreen is whether the answer stays true. Answer gap is whether the existing answers are any good. Seasonality catches the fact that frozen-pipe questions are worth writing in October, not April.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Generate&lt;/strong&gt; (&lt;code&gt;generate.py&lt;/code&gt;) takes the top clusters and has Claude write a blog post in Markdown, a video script, and an FAQ block as JSON-LD. Before any of that, it writes &lt;code&gt;out/opportunities.md&lt;/code&gt; — the ranking itself, so you can look at what it thinks is worth writing &lt;em&gt;before&lt;/em&gt; spending tokens writing it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it's built this way
&lt;/h2&gt;

&lt;p&gt;Two commands:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pipewise run                    &lt;span class="c"&gt;# scrape → enrich → store&lt;/span&gt;
pipewise generate &lt;span class="nt"&gt;--dry-run&lt;/span&gt;     &lt;span class="c"&gt;# show me the ranking&lt;/span&gt;
pipewise generate &lt;span class="nt"&gt;--top&lt;/span&gt; 3       &lt;span class="c"&gt;# actually write three&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;run&lt;/code&gt; is safe to schedule. &lt;code&gt;generate&lt;/code&gt; is manual, always, and never on a timer. The ingest half is cheap and local — Scrapling and a local Ollama model cost nothing per post — so it can run nightly under Windows Task Scheduler and accumulate. The generation half calls a paid API to produce something a human will publish under their name, and that should be a decision, not a cron job.&lt;/p&gt;

&lt;p&gt;The LLM usage is tiered the same way as my &lt;a href="https://dev.to/blog/pdf-to-podcast/"&gt;local pdf-to-podcast build&lt;/a&gt;: a small local model does the high-volume mechanical tagging, and the expensive model only touches the handful of things that reach the top of the ranking.&lt;/p&gt;

&lt;p&gt;Every external I/O boundary — Scrapling, Ollama, Claude, the embedding model — is injected as a function. That's what lets the whole test suite run fully mocked, with one gated live test that hits real Reddit when you ask it to.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gotchas
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;pip install scrapling&lt;/code&gt; does not install the fetchers.&lt;/strong&gt; The base package omits &lt;code&gt;curl_cffi&lt;/code&gt;, Playwright and Patchright. Everything imports fine, the mocked tests pass — and the first live fetch dies with &lt;code&gt;ModuleNotFoundError&lt;/code&gt;. The requirements file pins &lt;code&gt;scrapling[fetchers]&lt;/code&gt; for exactly this reason. Mocked tests only exercise the parser, so the gap is invisible until you touch the network.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reddit caps listings at roughly 1,000 items each.&lt;/strong&gt; You cannot backfill a subreddit's history through the listing endpoints, no matter how politely you page. The strategy that works is a seed run plus scheduled top-ups that accumulate forward from now. I looked at RedditDownloader as an alternative and rejected it: it's archived, PRAW-based, and media-oriented.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Run pytest from the workspace root, not the project directory.&lt;/strong&gt; &lt;code&gt;PipeWise&lt;/code&gt; needs to import as a package, which means the parent directory has to be on the path. An empty &lt;code&gt;conftest.py&lt;/code&gt; at the repo root handles it. Running &lt;code&gt;pytest&lt;/code&gt; from inside &lt;code&gt;PipeWise/&lt;/code&gt; produces import errors that look like missing dependencies.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Old Reddit is the right target.&lt;/strong&gt; Not because it's nostalgic — because it's static HTML that a parser can read, where the modern interface is a React application that needs a full browser. Choosing the older surface of a site is often the difference between a scraper and a browser automation project.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it stands
&lt;/h2&gt;

&lt;p&gt;The Foundation and the Content Engine are built and tested — 78 test functions plus one gated live test verified against real Reddit. It lives on a &lt;code&gt;pipewise&lt;/code&gt; branch that I've deliberately left unmerged.&lt;/p&gt;

&lt;p&gt;PipeWise was always meant to be four products sharing one corpus: a knowledge base, market intelligence, this content engine, and a diagnostic assistant. The Foundation was built once so the other three can attach to the same SQLite corpus later, each with its own spec. The content engine went first because it's the one that produces something publishable on day one.&lt;/p&gt;

&lt;p&gt;Known rough edges: the clustering centroid is greedy-first and never recomputed, so cluster quality depends somewhat on arrival order; and &lt;code&gt;default_embed&lt;/code&gt; reloads the sentence-transformer model on every call instead of batching all questions into one. Both are on the list. Neither stops the ranking from being useful, which is the bar for a v1.&lt;/p&gt;

&lt;p&gt;More projects built this way are on the &lt;a href="https://dev.to/projects/"&gt;projects page&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>python</category>
      <category>scraping</category>
      <category>llm</category>
      <category>content</category>
    </item>
    <item>
      <title>ClawWatch: An Agent That Watches My Game Server and Asks Before It Acts</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Sun, 30 Aug 2026 14:28:10 +0000</pubDate>
      <link>https://dev.to/bsymbolic/clawwatch-an-agent-that-watches-my-game-server-and-asks-before-it-acts-3fjb</link>
      <guid>https://dev.to/bsymbolic/clawwatch-an-agent-that-watches-my-game-server-and-asks-before-it-acts-3fjb</guid>
      <description>&lt;p&gt;A player crashed out of my FiveM roleplay server with &lt;code&gt;ERR_STR_INFO_2&lt;/code&gt; — a RAGE streaming crash caused by a bad addon asset. Diagnosing it meant SSHing to the VPS, tailing a log, cross-referencing which resource had just started, and knowing what that particular error code means. All of which I did, slowly, at eleven at night.&lt;/p&gt;

&lt;p&gt;ClawWatch is the tool that should have done it for me. It's a neon Electron desktop agent that connects to the live server four different ways, recognises known failures by pattern, and can act on them — with a hard line between actions it takes on its own and actions it asks about first.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does
&lt;/h2&gt;

&lt;p&gt;Four connectors, each speaking a different protocol to the same box:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;SSH&lt;/strong&gt; (&lt;code&gt;ssh2&lt;/code&gt;) — tails the server log and reads CPU and memory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;txAdmin&lt;/strong&gt; — HTTP against the admin panel.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;RCON&lt;/strong&gt; — UDP, with the packet format hand-rolled over &lt;code&gt;dgram&lt;/code&gt;, because the protocol is small and the libraries are worse.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MariaDB&lt;/strong&gt; (&lt;code&gt;mysql2&lt;/code&gt;) — read and write, split so that reads and writes travel different paths.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On top of those sits a &lt;strong&gt;rules-first diagnosis engine&lt;/strong&gt;. It's regex over log lines, and it costs nothing to run. Seed rules cover &lt;code&gt;err-str-info-2&lt;/code&gt;, resource-start failures (which trigger an automatic restart), Lua script errors, dropped database connections, thread hitches and out-of-memory conditions. Claude gets escalated to only when a rule misses, or when I explicitly ask.&lt;/p&gt;

&lt;p&gt;That ordering matters more than it looks. The overwhelming majority of what a game server logs is a small set of recurring failures. Sending each of those to a language model would be slow, expensive, and &lt;em&gt;less&lt;/em&gt; reliable than a pattern that has been right a hundred times. The model earns its place on the long tail, not the head.&lt;/p&gt;

&lt;p&gt;The dashboard is health tiles, a live log with rule matches highlighted inline, an alert and audit feed, a chat pane, and a first-run setup screen.&lt;/p&gt;

&lt;h2&gt;
  
  
  The safety model
&lt;/h2&gt;

&lt;p&gt;This is the part I care about, because the agent has SSH, RCON and database write access to a server real people are playing on.&lt;/p&gt;

&lt;p&gt;Actions are classified &lt;strong&gt;SAFE&lt;/strong&gt; or &lt;strong&gt;RISKY&lt;/strong&gt;. SAFE actions run automatically, rate-limited. RISKY actions — stopping the server, kicking or banning a player, any database write, any file operation, raw RCON or raw SSH — require a confirmation modal.&lt;/p&gt;

&lt;p&gt;The important detail is where that classification lives. The executor &lt;strong&gt;re-classifies every action itself&lt;/strong&gt; rather than trusting a flag handed to it by the caller, and if it decides an action is RISKY and no confirmation callback is available, it &lt;strong&gt;fails closed&lt;/strong&gt; and refuses. A code review caught that it originally failed &lt;em&gt;open&lt;/em&gt; — an action arriving without a confirm handler would simply have run. That's the difference between a safety model and a suggestion.&lt;/p&gt;

&lt;p&gt;The agent core in &lt;code&gt;src/core&lt;/code&gt;, &lt;code&gt;src/connectors&lt;/code&gt; and &lt;code&gt;src/rules&lt;/code&gt; has no Electron dependencies at all. That was deliberate: the eventual phase-two version is a headless always-on service, and keeping the brain free of the window means that's a drop-in rather than a rewrite.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gotchas
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Log lines went into the DOM via &lt;code&gt;innerHTML&lt;/code&gt;.&lt;/strong&gt; A code review caught this one and it's the scariest bug in the project. The live log renderer built HTML from raw server log lines — lines that contain, among other things, player-supplied chat and resource names. Any player who could get text into the log could have executed script in the agent window, which holds SSH and database credentials. Fixed to &lt;code&gt;textContent&lt;/code&gt;. If you are rendering untrusted text, the fix is not escaping it better; it is not using &lt;code&gt;innerHTML&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A regex that ate a trailing period.&lt;/strong&gt; The resource-name pattern matched one character too many, so &lt;code&gt;resource-name.&lt;/code&gt; captured with the period attached and then failed to match anything downstream. Small, silly, and exactly the kind of thing a second reviewer catches and the author doesn't.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Restarting the agent leaked SSH connections.&lt;/strong&gt; &lt;code&gt;startAgent&lt;/code&gt; was re-entrant — call it twice and the first connection was orphaned rather than torn down. Caught in review, fixed with an explicit teardown.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;txAdmin's API paths move between panel versions.&lt;/strong&gt; There's no stable contract here, so every txAdmin call goes through &lt;code&gt;post()&lt;/code&gt; and &lt;code&gt;getJson()&lt;/code&gt; helpers in &lt;code&gt;txadmin.js&lt;/code&gt; rather than being scattered through the codebase. When a panel upgrade breaks something, there's one file to fix.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Real credentials never touch the project directory.&lt;/strong&gt; The live config lives in Electron's &lt;code&gt;userData&lt;/code&gt; as &lt;code&gt;clawworld.config.json&lt;/code&gt;; the repository carries only &lt;code&gt;clawworld.config.example.json&lt;/code&gt;. Obvious in principle, easy to get wrong when the first thing you do is hardcode a host to test a connector.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it stands
&lt;/h2&gt;

&lt;p&gt;v1 is complete: 44 tests under &lt;code&gt;node:test&lt;/code&gt;, with every connector exercised through injected fakes so the suite never touches the live VPS. The app launches to its setup screen; the live server only gets touched during manual end-to-end runs. Electron is pinned at 42.4.1 with a clean &lt;code&gt;npm audit&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;It was built subagent-driven across 15 test-driven tasks, and two separate code-review passes found the four real bugs above — the XSS, the fail-open confirm, the regex, and the connection leak. None of those were caught by the tests that shipped alongside the code that contained them, which is the argument for review passes in one sentence.&lt;/p&gt;

&lt;p&gt;The code lives in a private repository with a pull request open, and also in my workspace repo at &lt;code&gt;claude/ClawWatch/&lt;/code&gt;. The broader idea was three products — crash debugging, mod script building, and this agent. Only the agent exists so far.&lt;/p&gt;

&lt;p&gt;More projects built this way are on the &lt;a href="https://dev.to/projects/"&gt;projects page&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>electron</category>
      <category>agents</category>
      <category>devops</category>
      <category>ai</category>
    </item>
    <item>
      <title>ClawCommand: One Dashboard for Every AI Thing on My Machine</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Sun, 30 Aug 2026 14:27:52 +0000</pubDate>
      <link>https://dev.to/bsymbolic/clawcommand-one-dashboard-for-every-ai-thing-on-my-machine-18m8</link>
      <guid>https://dev.to/bsymbolic/clawcommand-one-dashboard-for-every-ai-thing-on-my-machine-18m8</guid>
      <description>&lt;p&gt;At some point I had seven different AI CLIs installed, a local Ollama with a handful of models, a gateway on one port, n8n on another, LM Studio on a third, and no idea at any given moment which of them were actually alive. Checking meant seven terminal commands. So I built the dashboard: ClawCommand, a full-window neon Electron app that answers "what AI is running on this box right now" in one glance.&lt;/p&gt;

&lt;p&gt;It's the third in the Claw family, after &lt;a href="https://dev.to/blog/clawmonitor/"&gt;ClawMonitor&lt;/a&gt; and &lt;a href="https://dev.to/blog/clawports/"&gt;ClawPorts&lt;/a&gt;, and it reuses ClawMonitor's architecture wholesale because that architecture turned out to be right.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it shows
&lt;/h2&gt;

&lt;p&gt;Five panels, each backed by its own collector:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Providers&lt;/strong&gt; — Claude, Codex, Gemini, Cursor, Ollama, Perplexity and Manus. For each: is it installed, what version, and what's its auth state. This panel never displays or logs a key &lt;em&gt;value&lt;/em&gt;; it reports presence only. Polled every 30 seconds, passively.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Models&lt;/strong&gt; — Ollama's &lt;code&gt;/api/tags&lt;/code&gt; and &lt;code&gt;/api/ps&lt;/code&gt;, so you see what's installed, what's currently loaded into VRAM, how much it's using, and the keep-alive countdown before it unloads. Every 5 seconds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Services&lt;/strong&gt; — port probes for the gateway on 18789, Ollama on 11434, n8n on 5678, AI-Infra-Guard on 8088 and LM Studio on 1234, plus &lt;code&gt;docker ps&lt;/code&gt; and &lt;code&gt;wsl -l --running&lt;/code&gt;. Every 5 seconds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Vitals&lt;/strong&gt; — &lt;code&gt;nvidia-smi&lt;/code&gt; and &lt;code&gt;systeminformation&lt;/code&gt; for GPU, VRAM, CPU and RAM. Every 3 seconds, because these are the numbers you actually watch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Activity&lt;/strong&gt; — a &lt;code&gt;tasklist&lt;/code&gt; scan for AI processes plus an event log. This one diffs consecutive snapshots into a ring buffer, so you get a running feed of &lt;em&gt;changes&lt;/em&gt;: a model loaded, a service went down, an auth state flipped. That's more useful than the raw state, because the interesting thing is almost always the transition.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one design rule
&lt;/h2&gt;

&lt;p&gt;Everything above is free. Every collector reads local state — process lists, HTTP probes to localhost, &lt;code&gt;nvidia-smi&lt;/code&gt;. Nothing in the passive loop calls a paid API, ever.&lt;/p&gt;

&lt;p&gt;But "is Claude installed and authenticated" is not the same question as "does a request actually work right now." So each provider card has a &lt;strong&gt;TEST&lt;/strong&gt; button, and pressing it fires exactly one real request of roughly five tokens through &lt;code&gt;main/test-runner.js&lt;/code&gt;. That is the only code path in the entire application that can spend money, it only runs on an explicit click, and it's the reason the passive polling can be as aggressive as it is. Separating "watch" from "verify" into two clearly different affordances is the whole design.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it was built
&lt;/h2&gt;

&lt;p&gt;Brainstorm, spec, plan, implement — one day, 2026-07-16. The architecture is copied from ClawMonitor's proven pattern: main-process collectors on group timers, pushing a merged snapshot over IPC to a vanilla-JS renderer with no build step. Every collector takes its dependencies as injected functions, which is what makes 86 Vitest tests possible without a single real subprocess or network call in the suite.&lt;/p&gt;

&lt;p&gt;The renderer holds no logic. It draws whatever the latest snapshot says. That constraint is load-bearing: when a collector fails, it degrades to &lt;code&gt;null&lt;/code&gt; in its slice and an entry in an errors map, and the rest of the dashboard carries on rendering. There's no state in the UI that can get out of sync with reality, because the UI has no state.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gotchas
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A &lt;code&gt;\0&lt;/code&gt; before a digit is an illegal octal escape in strict mode.&lt;/strong&gt; A test string contained &lt;code&gt;\0&lt;/code&gt; immediately followed by a digit, which JavaScript parses as the start of a legacy octal literal and refuses in strict mode. Use &lt;code&gt;\x00&lt;/code&gt;. Ten seconds to fix, twenty minutes to understand.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Cursor card says "not installed" and that is correct.&lt;/strong&gt; On this machine &lt;code&gt;cursor-agent&lt;/code&gt; lives inside WSL, reached through a Git Bash shim at &lt;code&gt;~/bin/cursor-agent&lt;/code&gt;. Electron and Node spawn through &lt;code&gt;cmd.exe&lt;/code&gt;, which cannot see a Bash shim. So the probe genuinely cannot find it and honestly reports it missing. This is documented in the README rather than papered over — a dashboard that lies to make a card green is worse than useless.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cold-loading a model can blow a 60-second timeout.&lt;/strong&gt; The Ollama TEST path timed out once while &lt;code&gt;llama3.2:3b&lt;/code&gt; cold-loaded under heavy CPU pressure — I had roughly fifteen Claude processes pinned at 91% CPU at the time. Warm, it's fast. That's not a bug so much as a fact about what a first request costs, and the fix was making the failure legible rather than making the timeout longer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"This operation was aborted" tells the user nothing.&lt;/strong&gt; The HTTP helper used to surface that raw string when a probe timed out. It now reports "timed out after Nms." A monitoring tool's error messages are its user interface.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it stands
&lt;/h2&gt;

&lt;p&gt;v1 is complete and live-verified on my machine: 86 tests green, providers detected correctly (Claude, Codex and Ollama green; Gemini showing auth-missing; Perplexity and Manus showing key-unset), models, services, vitals and activity all reading real data, and both the Ollama TEST success path and the fail-fast missing-key path confirmed by hand.&lt;/p&gt;

&lt;p&gt;It lives at &lt;code&gt;claude/ClawCommand/&lt;/code&gt; as its own git repository and hasn't been pushed anywhere public yet. It starts with &lt;code&gt;npm start&lt;/code&gt;. Unlike ClawMonitor, which docks to a screen edge and stays out of the way, ClawCommand is a full window you open when you want to know something — which is the right shape for a question you ask a few times a day rather than glance at constantly.&lt;/p&gt;

&lt;p&gt;More projects built this way are on the &lt;a href="https://dev.to/projects/"&gt;projects page&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>electron</category>
      <category>windows</category>
      <category>ai</category>
      <category>monitoring</category>
    </item>
    <item>
      <title>Shipping Lava Leap: itch.io, the Play Store, and a Bug Told by Telemetry</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Sun, 30 Aug 2026 14:27:33 +0000</pubDate>
      <link>https://dev.to/bsymbolic/shipping-lava-leap-itchio-the-play-store-and-a-bug-told-by-telemetry-2onn</link>
      <guid>https://dev.to/bsymbolic/shipping-lava-leap-itchio-the-play-store-and-a-bug-told-by-telemetry-2onn</guid>
      <description>&lt;p&gt;&lt;a href="https://dev.to/blog/lava-leap/"&gt;Lava Leap&lt;/a&gt; was a finished endless climber sitting on a Vercel URL. This is what happened when I tried to put it in front of people who hadn't been told about it — the publishing paths, the bug I had explained away months earlier, and the instrumentation that finally settled the argument.&lt;/p&gt;

&lt;p&gt;It's playable at &lt;a href="https://bsymbolic.itch.io/lava-leap" rel="noopener noreferrer"&gt;bsymbolic.itch.io/lava-leap&lt;/a&gt; and on the web, with an Android build mid-flight.&lt;/p&gt;

&lt;h2&gt;
  
  
  The itch.io build that 404s
&lt;/h2&gt;

&lt;p&gt;Publishing an HTML5 game to itch is genuinely a ten-minute job, with one landmine that costs an evening.&lt;/p&gt;

&lt;p&gt;Vite emits absolute asset paths — &lt;code&gt;/assets/index-abc123.js&lt;/code&gt;. itch serves games from a subpath on its CDN. So the preloader renders, the browser requests &lt;code&gt;/assets/…&lt;/code&gt; at the CDN root, gets a 404, and the JavaScript never loads. The canvas stays empty. Nothing in the console tells you the problem is your build configuration.&lt;/p&gt;

&lt;p&gt;The fix is one flag:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx tsc &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; npx vite build &lt;span class="nt"&gt;--base&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;./
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Relative paths. The Vercel build stays on the default base — this is an itch-specific package, not a global change. Alongside that: the project Kind must be set to HTML, the uploaded file has to be ticked "played in browser," and the embed wants explicit dimensions with mobile and fullscreen enabled. And itch won't let you upload anything until you've verified your account email, which is a fine policy and an annoying surprise at the moment you're trying to ship.&lt;/p&gt;

&lt;p&gt;Once it was up I verified the whole loop from the itch origin rather than assuming: boot, menu, name prompt, a real run, death, and then score submission, player rank and telemetry events all firing against the live backend from a third-party domain. Cross-origin is exactly where a game's backend calls quietly stop working.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bug I had already dismissed
&lt;/h2&gt;

&lt;p&gt;Months earlier, the highlight-clip feature shipped with a note in my own project notes: downloaded clips look choppy and fast-forwarded, but that's a headless-rendering artifact from the test environment, not a real bug.&lt;/p&gt;

&lt;p&gt;Then a real player on a real machine said the clips were choppy and fast-forwarded.&lt;/p&gt;

&lt;p&gt;The writeoff was wrong, and the mechanism is worth knowing if you ever record a canvas. &lt;code&gt;canvas.captureStream(30)&lt;/code&gt; &lt;strong&gt;samples&lt;/strong&gt; the canvas on a timer. Phaser runs with &lt;code&gt;preserveDrawingBuffer: false&lt;/code&gt;, which means the drawing buffer is only reliably readable immediately after a render. So the sampler mostly caught invalid or stale reads — measured at roughly 9 effective frames per second under load — and stamped every one of them at the nominal 30 fps. Twelve seconds of gameplay became a few seconds of compacted, sped-up video.&lt;/p&gt;

&lt;p&gt;The fix is to stop sampling and start driving: &lt;code&gt;captureStream(0)&lt;/code&gt; for a manual stream, then &lt;code&gt;track.requestFrame()&lt;/code&gt; on Phaser's &lt;code&gt;POST_RENDER&lt;/code&gt; event, when the buffer is guaranteed valid, throttled to about 30 fps. &lt;code&gt;preserveDrawingBuffer: true&lt;/code&gt; went in as a safety net for the fallback path.&lt;/p&gt;

&lt;p&gt;Post-fix measurements on a real GPU: a 12.3-second run produced a 12.79-second clip at 29.8 fps with tight 33 ms frame deltas. Under 6× CPU throttling, with the game itself down to 42 fps, a 14.57-second run produced a 14.65-second clip. Duration tracks wall clock in both cases, which is the actual invariant.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The invariant whose absence let the bug ship is now a test.&lt;/strong&gt; There's an end-to-end regression asserting that clip duration falls within 0.6–1.4× of run duration. That test would have failed on the original implementation.&lt;/p&gt;

&lt;p&gt;The standing rule that came out of it, now written into a brief every implementing agent reads: &lt;em&gt;no unexplained measurement ships, and calling something an "artifact" requires demonstrating the mechanism.&lt;/em&gt; I had used the word "artifact" as a synonym for "I don't want to investigate this."&lt;/p&gt;

&lt;p&gt;A related trap that fooled me three separate times: a synthetic playtest bot that mashes jump dies organically and then retries. A "suspiciously short" clip is usually a perfectly correct clip of run number two. Measurements now use a survival helper that keeps the bot alive.&lt;/p&gt;

&lt;h2&gt;
  
  
  Telemetry that measures itself
&lt;/h2&gt;

&lt;p&gt;While instrumenting the encoder fix I discovered something worse: the &lt;code&gt;track()&lt;/code&gt; function had been a &lt;strong&gt;silent no-op in production since v5.&lt;/strong&gt; No sink was ever wired up. Every analytics event the game had emitted for five versions went nowhere.&lt;/p&gt;

&lt;p&gt;The replacement is a write-only sink shaped like the leaderboard client: a nine-event whitelist locked by a sorted-array test, batching at ten events or fifteen seconds, &lt;code&gt;keepalive&lt;/code&gt; so a closing tab still flushes, a cap of twenty batches per session, a lazily-generated UUID player id, and complete dormancy when the environment variables are absent. Server-side it's a table with zero row-level-security policies — no direct reads or writes possible — behind a rate-limited RPC.&lt;/p&gt;

&lt;p&gt;The event that matters is &lt;code&gt;clip_ready&lt;/code&gt;, which reports the recorded duration, the actual media duration, and the ratio between them. A healthy encoder reports ~100. The old bug reported ~29. &lt;strong&gt;The first genuine production &lt;code&gt;clip_ready&lt;/code&gt; came back at ratio 100&lt;/strong&gt; — 7,880 ms recorded against 7,892 ms of media. The encoder fix was confirmed by the encoder's own telemetry, from a real player's browser.&lt;/p&gt;

&lt;p&gt;That's the shape I'd repeat: when you fix a bug you can't easily reproduce, ship a measurement of the thing that was broken, not just the fix.&lt;/p&gt;

&lt;h2&gt;
  
  
  The store, which is mostly waiting
&lt;/h2&gt;

&lt;p&gt;Google Play is the slow path. The app exists, the signed AAB is built and verified, the store listing and data-safety forms are done, and internal testing is active. But a new personal developer account must run a &lt;strong&gt;closed test with 12 testers for 14 consecutive days&lt;/strong&gt; before it can go to production. There's no engineering that shortens that.&lt;/p&gt;

&lt;p&gt;So itch and the web are the public presence in the meantime, which is the right call anyway — a browser game that needs no install is a lower-friction ask than an APK.&lt;/p&gt;

&lt;p&gt;Two build gotchas worth recording. The Android game-intervention opt-out attribute is &lt;code&gt;android:allowGameDownscaling&lt;/code&gt;, not the plausible-sounding alternative; get it wrong and &lt;code&gt;aapt&lt;/code&gt; hard-fails. And Samsung's Game Optimizing Service was rendering the app at 75% resolution and 60 Hz, which felt like input lag and wasn't — measured, the touch path delivered 92 of 92 rapid taps. Diagnosing that took wireless ADB, a Chrome DevTools Protocol hook over a raw WebSocket, and &lt;code&gt;adb reverse&lt;/code&gt; to get past my own firewall. The finding was that the "bug" was a platform power policy, and the fix was an opt-out file rather than anything in the game loop.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it stands
&lt;/h2&gt;

&lt;p&gt;Live on itch and the web, verified end to end from the itch origin including leaderboard and telemetry. 265 unit tests and 36 end-to-end tests. The keystore is backed up. The Play closed test is running out its fourteen days.&lt;/p&gt;

&lt;p&gt;The pattern across all of it: shipping surfaced three real bugs — the itch base path, the encoder, and the dead telemetry — that no amount of local development had. Two of them had been sitting in the codebase for months looking fine.&lt;/p&gt;

&lt;p&gt;More projects built this way are on the &lt;a href="https://dev.to/projects/"&gt;projects page&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>gamedev</category>
      <category>phaser</category>
      <category>telemetry</category>
      <category>ai</category>
    </item>
    <item>
      <title>CrewCast: The Weather App That Grades Its Own Forecasts</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Sun, 30 Aug 2026 14:27:13 +0000</pubDate>
      <link>https://dev.to/bsymbolic/crewcast-the-weather-app-that-grades-its-own-forecasts-2aj2</link>
      <guid>https://dev.to/bsymbolic/crewcast-the-weather-app-that-grades-its-own-forecasts-2aj2</guid>
      <description>&lt;p&gt;Every weather app tells you it's 86°F and clear. None of them tell you whether you can pour concrete today. CrewCast does: a job-site weather app for contractors and outdoor trades, live at &lt;a href="https://stormradar.vercel.app" rel="noopener noreferrer"&gt;stormradar.vercel.app&lt;/a&gt;, built around one question — &lt;em&gt;is this workable?&lt;/em&gt; — and a lot of machinery devoted to being honest about how confident it is.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does
&lt;/h2&gt;

&lt;p&gt;The Today screen leads with a &lt;strong&gt;job-site score&lt;/strong&gt; out of 100 and a plain-language briefing: "Currently 86°F and workable. Dry through the afternoon — a good window for most outdoor trades." Under it, a best-work-window line and a per-trade breakdown.&lt;/p&gt;

&lt;p&gt;There are &lt;strong&gt;15 trade profiles&lt;/strong&gt; — concrete, roofing, plumbing, painting, excavation, HVAC, landscaping, pressure washing, windows, electrical, tree work, solar, pest, asphalt and moving — each a small function evaluating current conditions into go / caution / no-go with a reason. Adding a trade is one definition plus one button. A contractor can also opt into a push notification the evening before a no-go day for their trade.&lt;/p&gt;

&lt;p&gt;Beyond that: radar with a scrubber and NEXRAD tiles, precipitation nowcast, hurricane tracking from the NHC, NWS alerts with a contractor-focused recommended action per alert type, a printable weather-delay report with per-cause breakdown and official alert history, and a true PDF export rendered server-side.&lt;/p&gt;

&lt;p&gt;Everything runs on free, mostly public-domain sources: NWS, NOAA, NHC, NEXRAD via RainViewer, Open-Meteo, USGS and NASA EONET.&lt;/p&gt;

&lt;h2&gt;
  
  
  The trust pack
&lt;/h2&gt;

&lt;p&gt;The distinguishing feature isn't a forecast. It's the layer that tells you how much to believe one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Model agreement.&lt;/strong&gt; The forecast is fetched from GFS, ECMWF and ICON simultaneously and compared. The briefing carries a line like "Models disagree (GFS · ECMWF · ICON): 2 of 3 show rain, 1 of 3 show storms, temp spread 3°." Disagreement caps the displayed confidence — a split forecast can't be High confidence, no matter how nice the average looks.&lt;/p&gt;

&lt;p&gt;Tuning this was instructive. The first version flagged all seven days as split, because a 49% versus 52% rain probability counted as disagreement. It isn't. The criterion became a genuine gap — one model above 60% while another sits below 40%, or a temperature spread over 8° in the near term. Florida in August still comes back genuinely split five days out of seven, which is the correct answer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Station observations.&lt;/strong&gt; The app pulls the actual current temperature from the nearest NWS station and shows it — "Observed 88° at KDJT · 24m ago" — flagging in amber when the model and the thermometer disagree by 5°F or more.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A self-verification scorecard.&lt;/strong&gt; Every day the app snapshots its own next-day high and rain call, then grades itself the following day against station observations where the sensors cooperate and model actuals where they don't, keeping a rolling 30. It says which basis it used. It also shows "collecting — N of 3" until it has enough results to say anything, rather than displaying a meaningless percentage on day one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A lightning all-clear timer&lt;/strong&gt; with an honest scope note: it runs a 30-minute countdown driven by NWS signals and storm codes, and states plainly that it is not strike data, because no free real-time strike feed exists.&lt;/p&gt;

&lt;p&gt;Alongside those: the forecasters' own reasoning pulled from the NWS Area Forecast Discussion, and SPC convective outlooks resolved by a hand-rolled point-in-polygon test against the day-1 through day-3 GeoJSON.&lt;/p&gt;

&lt;p&gt;And running through all of it, wording chosen to be defensible. Not "safe to work" but "likely workable." Every score card carries "Guidance, not a safety guarantee — verify on site," and a safety page that says plainly the NWS is authoritative and OSHA wins.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gotchas
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The service worker had never been deployed, so push was broken on every platform.&lt;/strong&gt; This is the worst bug in the project's history and it was invisible for weeks. Each frontend deploy copied only &lt;code&gt;index.html&lt;/code&gt; into the deploy directory. &lt;code&gt;serviceWorker.register('/sw.js')&lt;/code&gt; therefore 404'd, silently, and Web Push simply did not work — not just on iOS, everywhere. The fix was adding &lt;code&gt;sw.js&lt;/code&gt;, a manifest and three icons, and the durable lesson is that a deploy step which copies &lt;em&gt;a&lt;/em&gt; file rather than &lt;em&gt;the directory&lt;/em&gt; will eventually ship a broken app that looks fine. (The iOS-specific reality is separate and also real: Safari tabs don't expose push at all, so the app now detects iOS-not-standalone and shows Add to Home Screen instructions instead of a dead toggle.)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A mapped colour sweep will eat your data colours.&lt;/strong&gt; Restyling the app from navy-and-cyan glassmorphism to industrial charcoal-and-safety-orange meant 141 scripted replacements across the source. It worked — and it also recoloured the Saffir-Simpson hurricane scale, where tropical-storm cyan collided with category-2 orange. Data ramps are not brand palette. Water-blue for humidity and rain bars was kept deliberately for the same reason. If you sweep colours by literal value, semantic colours are collateral damage.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Official alert history needs two endpoints and real marine filtering.&lt;/strong&gt; Pulling historical NWS alerts for the delay report means hitting &lt;em&gt;both&lt;/em&gt; IEM archives: storm-based polygon warnings, where severe thunderstorm, tornado and flash flood warnings actually live, and the zone/county endpoint for watches and advisories. Then merge and dedupe. And filtering marine noise takes more than excluding one phenomenon code — a coastal job site drowns in Small Craft Advisories unless you filter the whole marine set plus the marine zone UGC patterns. Verified against a known event: a Severe Thunderstorm Warning at a specific West Palm Beach location appears exactly once, at the right time, with zero marine leakage.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Background tabs break animation timing.&lt;/strong&gt; A bottom sheet used &lt;code&gt;requestAnimationFrame&lt;/code&gt; to add its open class. rAF is throttled in background tabs, so the sheet became &lt;code&gt;display: flex&lt;/code&gt; while staying translated off-screen — visible as a dead grey area. Replaced with a forced reflow and a synchronous class add.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;International users were being silently misled.&lt;/strong&gt; A test from London revealed the app degrading quietly: NWS feeds fail outside the US, so the alerts panel would have shown "✅ No Active Alerts" — which abroad is not merely wrong but dangerous. Now the app detects coverage two ways and, outside the US, shows a notice explaining that forecasts and radar work fine while warnings and push are US-only, and replaces the all-clear with "🌍 Official warnings unavailable here."&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it stands
&lt;/h2&gt;

&lt;p&gt;Live, rebranded from StormRadar to CrewCast after a trademark collision, with a marketing landing page, seven legal and trust pages, seven SEO landing pages, a sitemap and security headers. The source was refactored out of one enormous HTML file into 26 modules with a build step and a proven byte-exact round trip — the single-file version is now a build output that should never be hand-edited.&lt;/p&gt;

&lt;p&gt;Honest caveats: it's imperial units only, real internationalisation is deliberately not built, and monetising it requires clearing three licensing gates — an Open-Meteo paid plan, a RainViewer commercial agreement or a swap to public-domain NEXRAD tiles, and a basemap plan. Knowing where those gates are is itself part of the work.&lt;/p&gt;

&lt;p&gt;One thing this post can't hide, since the screenshot above was taken with the theme pinned: the automatic daylight theme has a race where the page background stays dark while the cards flip light. That's the next fix.&lt;/p&gt;

&lt;p&gt;More projects built this way are on the &lt;a href="https://dev.to/projects/"&gt;projects page&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>weather</category>
      <category>pwa</category>
      <category>javascript</category>
      <category>ai</category>
    </item>
    <item>
      <title>Symphonia: 144 Movements of Public-Domain Classical, Traced to Source</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Sun, 30 Aug 2026 14:26:53 +0000</pubDate>
      <link>https://dev.to/bsymbolic/symphonia-144-movements-of-public-domain-classical-traced-to-source-3ck7</link>
      <guid>https://dev.to/bsymbolic/symphonia-144-movements-of-public-domain-classical-traced-to-source-3ck7</guid>
      <description>&lt;p&gt;There is a surprising amount of genuinely public-domain classical music on the internet and almost nowhere pleasant to listen to it. The Musopen Kickstarter collection — professional recordings, released under a Public Domain Mark — sits on archive.org as a pile of FLAC files. Symphonia turns that pile into a library you can browse by composer and click to play. It's live at &lt;a href="https://symphonia-web.vercel.app" rel="noopener noreferrer"&gt;symphonia-web.vercel.app&lt;/a&gt; with the entire collection loaded: &lt;strong&gt;14 composers, 37 works, 144 movements.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What it is
&lt;/h2&gt;

&lt;p&gt;A quiet, typographic library site. Composers with dates, works under each composer, movements under each work, and a player. Every work carries its performers and its provenance, because "public domain" is a claim you should be able to check: the tagline on the front page is "every work traced to its source," and the footer of each work names where the recording came from and under what mark.&lt;/p&gt;

&lt;p&gt;The catalog is Beethoven's Third, all four Brahms symphonies, Mendelssohn's Third and Fourth, Mozart's 40th and Tchaikovsky's Sixth — plus overtures, tone poems, string quartets, the Goldberg Variations and seven Schubert sonatas. Performers include the Czech National Symphony Orchestra, the Musopen String Quartet, Shelley Katz on the Goldbergs and Paul Pitman on the Schubert.&lt;/p&gt;

&lt;p&gt;The frontend is Next.js 14 on Vercel. The catalog is Supabase Postgres — five tables (composers, works, movements, renditions, sources) with row-level security that lets anonymous readers see only movements whose &lt;code&gt;stage&lt;/code&gt; is &lt;code&gt;uploaded&lt;/code&gt;, so half-ingested material never leaks into the UI. Audio lives in a public Supabase Storage bucket as Opus at 96 kbps.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it was built
&lt;/h2&gt;

&lt;p&gt;The frontend arrived as a static zip with no backend at all. Everything behind it — schema, storage, ingestion — got designed and built from scratch: brainstorm, then a written design spec, then a five-task plan, then implementation.&lt;/p&gt;

&lt;p&gt;The interesting half is the pipeline. It's a Node/TypeScript program in &lt;code&gt;pipeline/&lt;/code&gt;, driven by a manifest, with six idempotent stages:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm run pipeline &lt;span class="nb"&gt;sync&lt;/span&gt;    &lt;span class="c"&gt;# read the manifest, reconcile the catalog&lt;/span&gt;
npm run pipeline fetch   &lt;span class="c"&gt;# pull FLAC from archive.org&lt;/span&gt;
npm run pipeline encode  &lt;span class="c"&gt;# transcode to Opus 96k&lt;/span&gt;
npm run pipeline upload  &lt;span class="c"&gt;# push to Supabase Storage&lt;/span&gt;
npm run pipeline verify  &lt;span class="c"&gt;# confirm what landed&lt;/span&gt;
npm run pipeline run     &lt;span class="c"&gt;# all of the above&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each stage is safe to re-run, which matters more than it sounds: fetching 144 movements of FLAC is slow and flaky, and a pipeline you can't resume is a pipeline you run once and then avoid. Adding new works means extending &lt;code&gt;manifest.json&lt;/code&gt; and running &lt;code&gt;pipeline run&lt;/code&gt; again. There are 18 tests over the pipeline logic.&lt;/p&gt;

&lt;p&gt;Opus at 96 kbps was the choice that made the whole thing fit. The full collection is &lt;strong&gt;800 MB of the 1 GB Supabase free tier&lt;/strong&gt; — 80% full. Anything more (Chopin, say) needs object storage elsewhere or a paid plan. Free-tier egress is about 5 GB a month, which works out to roughly twenty full-symphony listens. That's a real constraint and it shaped the scope: ship the complete Musopen collection well rather than a partial everything.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gotchas
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A missing trailing slash in a Content-Security-Policy silently kills all audio.&lt;/strong&gt; This one cost the most time by far. The CSP's &lt;code&gt;media-src&lt;/code&gt; directive is built from &lt;code&gt;NEXT_PUBLIC_AUDIO_BASE_URL&lt;/code&gt; verbatim. CSP source expressions match on a path prefix — but only if the value ends in &lt;code&gt;/&lt;/code&gt;. Without the slash, the browser does an exact match instead, no audio URL ever matches, and every &lt;code&gt;audio.play()&lt;/code&gt; rejects with &lt;code&gt;NotSupportedError&lt;/code&gt; and &lt;em&gt;zero console output&lt;/em&gt;. No CSP violation warning, no network error, nothing. The page looks perfect and the music never starts.&lt;/p&gt;

&lt;p&gt;If you take one thing from this post: the trailing slash on a CSP source is load-bearing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Next's dev mode needs &lt;code&gt;'unsafe-eval'&lt;/code&gt; or the page renders and never hydrates.&lt;/strong&gt; Same category of failure — everything looks right, nothing works. Without &lt;code&gt;'unsafe-eval'&lt;/code&gt; in &lt;code&gt;script-src&lt;/code&gt; during development, React never hydrates, so clicks do nothing and no error appears. &lt;code&gt;next.config.mjs&lt;/code&gt; now adds it only when &lt;code&gt;NODE_ENV=development&lt;/code&gt;; production stays strict.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;vercel deploy&lt;/code&gt; hung forever uploading a 3.4 GB cache.&lt;/strong&gt; The pipeline keeps downloaded FLAC in &lt;code&gt;pipeline/cache/&lt;/code&gt;, and the deploy dutifully tried to upload all of it. One line in &lt;code&gt;.vercelignore&lt;/code&gt; excluding &lt;code&gt;pipeline/&lt;/code&gt; fixed it. Worth checking any repo where a build tool and a data tool share a directory.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Applying schema DDL took a detour.&lt;/strong&gt; The Supabase Management API endpoint (&lt;code&gt;POST /v1/projects/&amp;lt;ref&amp;gt;/database/query&lt;/code&gt;) works fine — but Cloudflare 403-blocks Python's &lt;code&gt;urllib&lt;/code&gt; user agent in front of &lt;code&gt;api.supabase.com&lt;/code&gt;. Use &lt;code&gt;curl&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;One more that's less a gotcha than a property: the site uses ISR with a one-hour window, so newly ingested works can take up to an hour to appear unless you redeploy. Locally, deleting &lt;code&gt;.next/&lt;/code&gt; busts the fetch cache.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it stands
&lt;/h2&gt;

&lt;p&gt;Symphonia is fully live and click-to-play has been verified in production. The full Musopen collection is ingested and serving — 14 composers, 37 works, 144 movements — with per-work performer credits and provenance. The pipeline is manifest-driven and idempotent with 18 tests. Storage sits at 80% of the free tier, which is the honest ceiling on the current shape of the project.&lt;/p&gt;

&lt;p&gt;Known follow-ups: uploaded objects serve &lt;code&gt;Cache-Control: no-cache&lt;/code&gt; despite the &lt;code&gt;cacheControl&lt;/code&gt; parameter being set, which leaves CDN caching on the table; the sources footer renders four duplicate provenance lines because the frontend doesn't dedupe them; and the repo isn't on GitHub yet. None of those stop the music.&lt;/p&gt;

&lt;p&gt;More projects built this way are on the &lt;a href="https://dev.to/projects/"&gt;projects page&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>nextjs</category>
      <category>supabase</category>
      <category>audio</category>
      <category>ai</category>
    </item>
    <item>
      <title>AuditScout: A Website Audit SaaS That Refuses to Bluff</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Sun, 30 Aug 2026 14:26:35 +0000</pubDate>
      <link>https://dev.to/bsymbolic/auditscout-a-website-audit-saas-that-refuses-to-bluff-4bgm</link>
      <guid>https://dev.to/bsymbolic/auditscout-a-website-audit-saas-that-refuses-to-bluff-4bgm</guid>
      <description>&lt;p&gt;Most website-audit tools hand you a wall of red and a number you can't act on. I wanted the opposite: a short list of things that are actually wrong, ranked by how much they matter, written in language a contractor or a solo founder can read — plus a 30-day plan and a copy/paste prompt that fixes them. That's AuditScout. It's live at &lt;a href="https://auditscout.vercel.app" rel="noopener noreferrer"&gt;auditscout.vercel.app&lt;/a&gt;, in public beta, and the most interesting engineering decisions in it are about honesty.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it is
&lt;/h2&gt;

&lt;p&gt;You paste a public URL. AuditScout fetches the page, runs a battery of passive checks across twelve dimensions — SEO, AEO, GEO, security, performance, mobile UX, accessibility, trust, conversion, and business strategy among them — and streams progress back over Server-Sent Events while it works. What comes out is a scorecard: an overall number, sub-scores per dimension, a ranked list of fixes, a 30-day plan, and a prompt you can paste into your own coding agent.&lt;/p&gt;

&lt;p&gt;No login is required for your first audit. Anonymous users get one, free accounts get three a month, and the paid tiers go up from there.&lt;/p&gt;

&lt;p&gt;The stack is Next.js 16 App Router with React 19, TypeScript, Tailwind v4, Prisma 7 against Neon Postgres, and NextAuth v5. Scanning is HTTP-fetch-first with a Playwright fallback when a page comes back under 500 characters or with an error status — plenty of sites render nothing useful without JavaScript, and plenty of others don't need a browser at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it was built
&lt;/h2&gt;

&lt;p&gt;Same loop as everything else here: I decided what the product was, wrote a locked spec, turned that into an eleven-phase implementation plan, and Claude executed it test-first with separate reviewer passes between phases. The spec and plan are in &lt;code&gt;docs/superpowers/&lt;/code&gt;. Then came something like nine passes of frontend and copy work, mostly driven by me looking at the live site and not liking what I saw.&lt;/p&gt;

&lt;p&gt;Those passes were less about pixels than about claims. The first version of the hero said "Audit any website in seconds." That's a promise the beta scanner can't fully keep, so it became "Get a prioritized website audit in seconds," and the page grew an explicit &lt;strong&gt;Live now / In progress&lt;/strong&gt; split naming exactly which parts of the product work today and which are still being built. There's a "What we don't scan" section listing nine things it deliberately doesn't touch. The pricing table marks unfinished tiers "Coming soon" instead of quietly implying they exist.&lt;/p&gt;

&lt;p&gt;That sounds like copywriting. It kept turning into engineering, because every honest claim had to be one the code could actually back.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three problems worth writing down
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;NextAuth v5 middleware drags Prisma into the Edge runtime.&lt;/strong&gt; The deploy failed on the first try because &lt;code&gt;middleware.ts&lt;/code&gt; imported the auth config, the auth config imported the Prisma adapter, and the Edge runtime rejected &lt;code&gt;node:util/types&lt;/code&gt;. The fix is a config split: &lt;code&gt;lib/auth.config.ts&lt;/code&gt; is edge-safe with an empty &lt;code&gt;providers: []&lt;/code&gt; array and is what middleware imports; &lt;code&gt;lib/auth.ts&lt;/code&gt; spreads that config and adds the Credentials and Google providers for the Node runtime. Two files, one import graph each.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Serverless screenshots broke for a reason that had nothing to do with Chromium.&lt;/strong&gt; AuditScout swaps full &lt;code&gt;playwright&lt;/code&gt; locally for &lt;code&gt;@sparticuz/chromium&lt;/code&gt; plus &lt;code&gt;playwright-core&lt;/code&gt; on Vercel. It kept failing in production, and every instinct said "the bundled Chromium is wrong." It wasn't. Next's tracer followed &lt;code&gt;playwright-core&lt;/code&gt; but not its sibling &lt;code&gt;browsers.json&lt;/code&gt;, so the dynamic &lt;code&gt;import("playwright-core")&lt;/code&gt; threw at runtime. The fix was &lt;code&gt;outputFileTracingIncludes&lt;/code&gt; in &lt;code&gt;next.config.ts&lt;/code&gt;, force-including the whole package:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;outputFileTracingIncludes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/api/audit/**&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;./node_modules/playwright-core/**&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;./node_modules/@sparticuz/chromium/**&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;},&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The trap inside the trap: that key is a picomatch glob. The obvious route key &lt;code&gt;/api/audit/[id]/stream&lt;/code&gt; silently fails, because &lt;code&gt;[id]&lt;/code&gt; parses as a character class. Use &lt;code&gt;/api/audit/**&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The score was invisible and every structural check passed.&lt;/strong&gt; This is the one I think about. The &lt;code&gt;ScoreRing&lt;/code&gt; component had both a Tailwind &lt;code&gt;rotate-90&lt;/code&gt; class and an SVG &lt;code&gt;transform&lt;/code&gt; attribute on its &lt;code&gt;&amp;lt;text&amp;gt;&lt;/code&gt; element. The CSS transform wins, and it rotates about the text's own box rather than the SVG origin — so the overall score number sat outside the visible ring on every single report. The DOM had it. The tests had it. Accessibility checks found the label. It was just not on the screen.&lt;/p&gt;

&lt;p&gt;The fix is trivial: rotate only the arc via the SVG attribute, never rotate the &lt;code&gt;&amp;lt;svg&amp;gt;&lt;/code&gt; and counter-rotate the text. The lesson isn't. A structural assertion that an element exists is not an assertion that a human can see it. Now the check is geometric — assert the score &lt;code&gt;&amp;lt;text&amp;gt;&lt;/code&gt; bounding box lies inside the SVG bounding box — and I take screenshots before believing a UI is fine.&lt;/p&gt;

&lt;p&gt;A second one from the same pass: a JSX text child containing a newline has its leading whitespace stripped, so &lt;code&gt;{SITE} is a solid…&lt;/code&gt; rendered as &lt;code&gt;brightleafplumbing.comis&lt;/code&gt;. You can find these across a whole site by grepping the prerendered HTML for &lt;code&gt;\w{2,}&amp;lt;!-- --&amp;gt;[A-Za-z]\w+&lt;/code&gt; — React's SSR comment marks exactly where the text node boundary was.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it stands
&lt;/h2&gt;

&lt;p&gt;The MVP is built, deployed, and verified end-to-end against real URLs: the full pipeline runs, all sections persist and render, signup and auth write to Neon in production. The suite is 64 unit tests; &lt;code&gt;tsc&lt;/code&gt;, the build, and lint are all clean. A Lighthouse run on the live homepage scores 98 mobile / 100 desktop on performance, 100 on best practices and SEO, and 100 on accessibility — the last of which took darkening three status-color tokens to clear 4.5:1 contrast for small text.&lt;/p&gt;

&lt;p&gt;The honest caveats, which are also on the site: production has no &lt;code&gt;ANTHROPIC_API_KEY&lt;/code&gt; set, so reports currently come from the deterministic fallback generator rather than a model — that path is proven and never crashes, but it isn't the AI-written report the product eventually sells. The anonymous rate limit is in-process memory and needs Redis at any real scale. The SSRF guard resolves DNS and revalidates every redirect hop, but a determined rebinding attack still has a residual TOCTOU window. And Vercel's 60-second Hobby cap means a genuinely slow site can leave an audit stuck in &lt;code&gt;SCANNING&lt;/code&gt; with no stale-status recovery yet.&lt;/p&gt;

&lt;p&gt;Those are the next things to fix, and they're on a public roadmap page rather than in a private issue tracker, which feels like the right place for them.&lt;/p&gt;

&lt;p&gt;More projects built this way are on the &lt;a href="https://dev.to/projects/"&gt;projects page&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>nextjs</category>
      <category>saas</category>
      <category>security</category>
      <category>ai</category>
    </item>
    <item>
      <title>TapTrace: Feature-Complete, Unlaunched, and Honest About Both</title>
      <dc:creator>C. Wheatley</dc:creator>
      <pubDate>Sun, 30 Aug 2026 14:26:17 +0000</pubDate>
      <link>https://dev.to/bsymbolic/taptrace-feature-complete-unlaunched-and-honest-about-both-1jf4</link>
      <guid>https://dev.to/bsymbolic/taptrace-feature-complete-unlaunched-and-honest-about-both-1jf4</guid>
      <description>&lt;p&gt;Ask the internet which utility supplies your tap water and you will get a confident wrong answer. Utility service areas don't follow city limits, ZIP codes are mailing conventions rather than boundaries, and most geocoders interpolate your house from a street segment instead of finding your roof. TapTrace is my attempt to answer the question properly: give it a US street address, get back the Public Water System that serves it, its drinking-water violations and contaminant levels — and, crucially, &lt;strong&gt;a confidence score and tier saying how much to trust that answer.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The engine is done. All four SaaS surfaces around it are done. The landing page is live at &lt;a href="https://taptrace-sjal.vercel.app" rel="noopener noreferrer"&gt;taptrace-sjal.vercel.app&lt;/a&gt;. The API has never been deployed to a server and has never had a customer. This post is about both halves of that.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;POST /resolve&lt;/code&gt; takes an address. It geocodes it, runs a point-in-polygon query against EPA Community Water System service-area boundaries in PostGIS, and returns the PWSID that serves the point — plus open violation counts, contaminants measured against their MCLs, and a &lt;code&gt;tier&lt;/code&gt; of &lt;code&gt;EXACT&lt;/code&gt;, &lt;code&gt;LIKELY&lt;/code&gt;, &lt;code&gt;AREA&lt;/code&gt;, &lt;code&gt;WELL&lt;/code&gt; or &lt;code&gt;UNRESOLVED&lt;/code&gt; with a 0–100 confidence number behind it.&lt;/p&gt;

&lt;p&gt;The confidence model multiplies four factors: the resolution method, the geocode precision, cross-source agreement, and how close the point sits to a boundary edge. A real answer from the live system: 400 S Dixie Hwy in West Palm Beach resolves to &lt;strong&gt;FL4501559 at LIKELY / 70&lt;/strong&gt; using the free Census geocoder, and to EXACT / 100 when a rooftop geocoder is in play. That gap is deliberate and it's the most important design decision in the project.&lt;/p&gt;

&lt;p&gt;The Census geocoder is interpolation-only. It does not do rooftop. So TapTrace &lt;strong&gt;structurally cannot return EXACT&lt;/strong&gt; on the free tier, and it doesn't pretend to. Buying a rooftop geocoder key is what unlocks that tier. An accuracy claim you can't back with a measurement isn't an accuracy claim.&lt;/p&gt;

&lt;p&gt;Florida is loaded: 1,385 service-area boundaries, 84,618 violations (4,943 health-based, the rest monitoring), 9,873 samples. Ingestion is now state-parameterized, so &lt;code&gt;run_all --state FL --state TX&lt;/code&gt; works and Texas has been probed live from both upstream sources.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it was built
&lt;/h2&gt;

&lt;p&gt;Python 3.12, FastAPI, SQLAlchemy 2 with GeoAlchemy2, Alembic, PostgreSQL + PostGIS in Docker, and pytest with ruff and strict mypy. It was built as a sequence of milestones — M0 foundations through M5 deploy — each one getting its own written plan and its own pull request, then four SaaS sub-projects on top: API access and metering, an operator dashboard, an embeddable widget, and Stripe billing. Plus a phase-2 corrosivity calculator. Sixteen PRs, 211 tests.&lt;/p&gt;

&lt;p&gt;Three decisions I'd defend:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Privacy by construction.&lt;/strong&gt; The request log has no address column at all. The query cache scrubs every address form — including the geocoder's &lt;code&gt;matched_address&lt;/code&gt;, which was the one that nearly slipped through review. The ground-truth corpus stores a salted SHA-256 hash of the address plus the geocoded point, never the street. You can't leak what you didn't keep.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Accuracy as a test, not a claim.&lt;/strong&gt; There's a &lt;code&gt;validation/&lt;/code&gt; package with a corpus of Florida ground-truth labels sourced from documented city-to-utility relationships — deliberately &lt;em&gt;not&lt;/em&gt; from the engine's own output, so the check isn't circular. A harness computes per-tier accuracy and coverage, and a CI gate asserts coverage ≥ 0.80 and resolved accuracy ≥ 0.90. The current corpus runs 13/13 correct, all at LIKELY tier. It is also honestly biased toward municipal city centers by design: it validates the pipeline, not service-area edges. And it's Florida-only — accuracy does not travel with coverage, so adding Texas ingestion does not entitle anyone to claim Texas accuracy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scope the calculator to what the data supports.&lt;/strong&gt; The corrosivity work computes LSI, RSI, Larson-Skold and CCPP from &lt;em&gt;caller-supplied&lt;/em&gt; water chemistry, not per-PWS. That's because SDWIS simply doesn't carry pH, alkalinity, calcium hardness or TDS per system — computing a corrosion index from absent data would produce a confident number about nothing. So it's a calculator with an explicit "screening estimate, not a measurement" disclaimer, and the engine was ported verbatim from my earlier copper-risk work, PHREEQC oracle tests and all.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gotchas
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;State-qualified ingest source names are the difference between coverage and data loss.&lt;/strong&gt; The &lt;code&gt;load_versioned&lt;/code&gt; helper activates one version per source, then deletes rows belonging to superseded versions &lt;em&gt;of that source&lt;/em&gt;. When ingestion became multi-state, the source name had to become &lt;code&gt;epa_boundaries:FL&lt;/code&gt; rather than a bare &lt;code&gt;epa_boundaries&lt;/code&gt; — because a shared source name would mean ingesting Texas silently deletes every Florida row. That naming lives in one module and is pinned by its own test, which is exactly where a landmine like that belongs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A migration that uses &lt;code&gt;create_all&lt;/code&gt; will eat later migrations.&lt;/strong&gt; Migration &lt;code&gt;0001&lt;/code&gt; originally called &lt;code&gt;Base.metadata.create_all&lt;/code&gt;. On a fresh database that creates every table currently on &lt;code&gt;Base&lt;/code&gt; — including tables that &lt;em&gt;later&lt;/em&gt; migrations are supposed to own — so migration &lt;code&gt;0002&lt;/code&gt; then hits &lt;code&gt;DuplicateTable&lt;/code&gt; and the whole chain fails on new environments only. It bit twice before I froze &lt;code&gt;0001&lt;/code&gt; to explicit &lt;code&gt;op.create_table&lt;/code&gt; DDL for its eight original tables. If your first migration reflects live models, it is a time bomb aimed at every future clone of the project.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two upstream sources that agree may not be independent.&lt;/strong&gt; The plan was to cross-check EPA's Florida boundaries against the state Water Management District layers. Then the recon turned up that EPA's Florida polygons are &lt;em&gt;already sourced from&lt;/em&gt; SFWMD. Agreement between them would have been circular and would have inflated confidence scores for free. That check is deferred with a note explaining why.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reference data disappears.&lt;/strong&gt; The Envirofacts &lt;code&gt;REF_CODE_VALUES&lt;/code&gt; table, which maps contaminant codes to names, returns "table is not available." So there's a committed seed CSV of 93 codes from the SDWIS data dictionary covering every code that appears in Florida's violations, and the ingestor auto-refreshes from the API if it ever comes back.&lt;/p&gt;

&lt;p&gt;And a small operational one that cost an evening: every &lt;code&gt;docker compose&lt;/code&gt; command against the production stack needs &lt;code&gt;--env-file .env.prod&lt;/code&gt;, because &lt;code&gt;env_file:&lt;/code&gt; injects variables into containers but &lt;em&gt;not&lt;/em&gt; into compose's own interpolation. Without it the database initializes with empty credentials and fails its healthcheck.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it stands
&lt;/h2&gt;

&lt;p&gt;Engine milestones M0 through M5 are complete and merged. API keys with hashed storage, atomic quota metering, an operator dashboard with CSRF-protected forms and a dependency-free inline-SVG chart, a Shadow-DOM embeddable widget gated on publishable keys and an origin allowlist, and Stripe billing driven by webhooks as the source of truth — all built, all merged, 211 tests, ruff and mypy clean. There's a multi-stage Dockerfile, a production compose stack, a scheduler, and a one-shot &lt;code&gt;deploy.sh&lt;/code&gt; that generates its own secrets. The whole thing has been smoke-tested end to end on a fresh volume: migrations apply, &lt;code&gt;/health&lt;/code&gt; returns 200, a minted key resolves a real Florida address.&lt;/p&gt;

&lt;p&gt;And then it stops, because I have never run &lt;code&gt;deploy.sh&lt;/code&gt; on an actual VPS. There is no host, Stripe is still in test mode, and the only thing on the public internet is the landing page. A monetization spec now names the six gaps between here and a first dollar and notes that only two of them are hard blockers.&lt;/p&gt;

&lt;p&gt;Building the thing turned out to be the easy part. That's the honest summary, and it's worth writing down.&lt;/p&gt;

&lt;p&gt;More projects built this way are on the &lt;a href="https://dev.to/projects/"&gt;projects page&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>python</category>
      <category>fastapi</category>
      <category>postgis</category>
      <category>saas</category>
    </item>
  </channel>
</rss>
