<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: David Aronchick</title>
    <description>The latest articles on DEV Community by David Aronchick (@aronchick).</description>
    <link>https://dev.to/aronchick</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1294202%2Fe7ab50ef-66a0-4ab1-b75f-30006ae9a811.jpeg</url>
      <title>DEV Community: David Aronchick</title>
      <link>https://dev.to/aronchick</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/aronchick"/>
    <language>en</language>
    <item>
      <title>The Shortest Stave</title>
      <dc:creator>David Aronchick</dc:creator>
      <pubDate>Fri, 02 Oct 2026 18:01:34 +0000</pubDate>
      <link>https://dev.to/aronchick/the-shortest-stave-46e8</link>
      <guid>https://dev.to/aronchick/the-shortest-stave-46e8</guid>
      <description>&lt;p&gt;A wooden barrel built from unequal staves (I love this term, by the way; I've never heard it!) holds water only as high as its shortest stave. Chemists call this the "limiting reagent," but I'm particularly fond of talking about these kinds of constraints in the living world, where things grow ... organically. Long before anyone had heard of a GPU, Carl Sprengel sketched the idea in 1828, and Justus von Liebig made it famous a couple of decades later, focusing specifically on how a plant's growth is set by whichever nutrient is scarcest. Pour on all the nitrogen you want, but if the soil is short on phosphorus, the plant stops growing at the phosphorus line, and the nitrogen sits there, unused, a resource with nothing left to do. So, back to our barrels: you can spend a fortune making the tall staves taller, but &lt;a href="https://en.wikipedia.org/wiki/Liebig%27s_law_of_the_minimum" rel="noopener noreferrer"&gt;the barrel doesn't care&lt;/a&gt;. Every dollar spent elsewhere runs over that one short piece of wood and out onto the ground.&lt;/p&gt;

&lt;p&gt;During 2026, the price of a stick of server RAM stopped being "Oh, I guess we'll just have to add a little bit to the price of this server, but it's diminimus compared to this Rolls Royce GPU we have in here." Conventional DRAM contract prices &lt;a href="https://www.trendforce.com/presscenter/news/20260601-13070.html" rel="noopener noreferrer"&gt;rose 93 to 98 percent in the first quarter alone&lt;/a&gt;, and TrendForce expected another 58 to 63 percent in the second, roughly tripling in six months. The reason? &lt;a href="https://www.trendforce.com/presscenter/news/20260602-13074.html" rel="noopener noreferrer"&gt;High-bandwidth memory will take about 22 percent of the world's DRAM wafer starts by the end of 2026&lt;/a&gt;, up from 18 percent a year earlier, and turn them into only about 9 percent of the bits, because a bit of HBM eats &lt;a href="https://www.trendforce.com/news/2025/12/26/news-ai-reportedly-to-consume-20-of-global-dram-wafer-capacity-in-2026-hbm-gddr7-lead-demand/" rel="noopener noreferrer"&gt;three to four times the wafer&lt;/a&gt; of an equivalent bit of DDR5. SK Hynix, Samsung, and Micron, the only three companies on Earth that make HBM at volume, reported their 2026 HBM capacity sold out under existing contracts, and &lt;a href="https://seekingalpha.com/news/4625688-samsung-sk-hynix-micron-sell-out-2027-memory-chip-supply-report" rel="noopener noreferrer"&gt;Digitimes reported in August&lt;/a&gt; that most of 2027 is already allocated too. Try to buy a stick of RAM for a gaming PC right now, and you're competing with Nvidia's supply chain. Nvidia is winning.&lt;/p&gt;

&lt;p&gt;Everybody saw this coming from the wrong direction. For two years the story was GPUs: how many Nvidia could ship, how long the allocation list revolved around, and whether the shortage had somehow become a permanent feature of the industry rather than a temporary one. Then the story became power. Gigawatts, interconnection queues, and data &lt;a href="https://www.bloomberg.com/news/articles/2024-08-29/data-centers-face-seven-year-wait-for-power-hookups-in-virginia" rel="noopener noreferrer"&gt;centers in&lt;/a&gt; Virginia are &lt;a href="https://www.bloomberg.com/news/articles/2024-08-29/data-centers-face-seven-year-wait-for-power-hookups-in-virginia" rel="noopener noreferrer"&gt;waiting up to seven years for a grid connection&lt;/a&gt;. Nobody was watching the beige stick of memory sitting next to the accelerator on the board, because memory has been the least glamorous component in a computer since the 1960s, and unglamorous things don't get congressional hearings. Now it's the thing capping how much of this you can actually build, and a fab that makes memory cannot be told to make more of it by Thursday. Wafer capacity is a multi-year build; wanting more sooner doesn't make it sooner.&lt;/p&gt;

&lt;p&gt;The AI industry has been building staves. It built an enormous GPU stack, then an enormous power stack in roughly that order, because those constraints tailed first and made the loudest noise going down. Memory didn't make a loud noise, and so it quietly became the shortest piece of wood in the barrel. The irony is that it's not even HBM's fault! It's just that HBM and consumer memory come out of the same fabs, competing for the same wafer starts, so every additional stack of HBM that Nvidia's next accelerator needs is DRAM capacity that doesn't go into a laptop, a phone, or, this being the actual news story of the last few months, a stick of RAM for the machine on your desk. The industry solved for the two constraints everyone could see coming, and the barrel filled to a different line anyway. One nobody had thought to check.&lt;/p&gt;

&lt;p&gt;Even worse, Samsung, SK Hynix, and Micron can't ship a patch. Building new wafer capacity takes years, and the contracts locking up existing capacity run multiple years past that. Which means the memory constraint, unlike the GPU shortage and unlike the power shortage, won't get quietly fixed by a new chip generation or another grid order. It's poured into the concrete of fabs that haven't broken ground yet. The people planning AI infrastructure right now are, whether they've noticed or not, planning against a stave whose length was mostly fixed a few years ago and cannot be argued with, subsidized, or ordered into growing faster by anyone in Washington.&lt;/p&gt;

&lt;p&gt;The GPU story and the power story were both real. They ARE both real. Solving them was and is necessary work. Solving the constraint just moves the waterline to the next-shortest stave, and that one is rarely the one the headlines prepared you for. Liebig figured that out, staring at dirt in the 1840s. The AI industry is finding it out staring at a spec sheet in 2026 and paying triple for the privilege.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Want to learn how intelligent data pipelines can reduce your AI costs?&lt;/em&gt; &lt;a href="https://expanso.io/?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;&lt;em&gt;Check out Expanso&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;. &lt;em&gt;Or don't. Who am I to tell you what to do.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;NOTE: I'm currently writing a book based on what I have seen about the real-world challenges of data preparation for machine learning, focusing on operational, compliance, and cost.&lt;/strong&gt; &lt;a href="https://github.com/aronchick/Project-Zen-and-the-Art-of-Data-Maintenance?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;I'd love to hear your thoughts&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;!&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.distributedthoughts.org/2026-10-08-the-shortest-stave/" rel="noopener noreferrer"&gt;The Shortest Stave&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiinfrastructure</category>
      <category>memory</category>
      <category>hardware</category>
      <category>historyofscience</category>
    </item>
    <item>
      <title>The Elevator Receipt</title>
      <dc:creator>David Aronchick</dc:creator>
      <pubDate>Tue, 29 Sep 2026 18:01:35 +0000</pubDate>
      <link>https://dev.to/aronchick/the-elevator-receipt-52bo</link>
      <guid>https://dev.to/aronchick/the-elevator-receipt-52bo</guid>
      <description>&lt;p&gt;On July 1, Bloomberg reported that Meta is &lt;a href="https://www.bloomberg.com/news/articles/2026-07-01/meta-is-building-a-cloud-business-to-sell-excess-ai-compute" rel="noopener noreferrer"&gt;building a cloud business&lt;/a&gt; to sell access to its spare AI compute, an initiative reportedly called Meta Compute. This is the same company that, as of June 30, had &lt;a href="https://www.sec.gov/Archives/edgar/data/1326801/000162828026050705/meta-20260630.htm" rel="noopener noreferrer"&gt;$279 billion in data center leases&lt;/a&gt; signed but not yet started (and signed another $68 billion in July), with an Ohio campus coming online this year and &lt;a href="https://datacenters.atmeta.com/richland-parish-data-center/" rel="noopener noreferrer"&gt;a Louisiana campus&lt;/a&gt; behind it, which Zuckerberg pitched as covering &lt;a href="https://www.theguardian.com/technology/2025/jul/16/zuckerberg-meta-data-center-ai-manhattan" rel="noopener noreferrer"&gt;a significant part of the footprint of Manhattan&lt;/a&gt;. Zuckerberg had said in May that a cloud business was &lt;a href="https://www.cnbc.com/2026/05/27/mark-zuckerberg-says-meta-starting-cloud-business-on-the-table.html" rel="noopener noreferrer"&gt;"definitely on the table."&lt;/a&gt; Meta was not even first.&lt;/p&gt;

&lt;p&gt;In early May, the entire capacity of Colossus 1, the data center xAI built before it was folded into SpaceX, &lt;a href="https://techcrunch.com/2026/05/06/is-xai-a-neocloud-now/" rel="noopener noreferrer"&gt;went to Anthropic in a buyout deal&lt;/a&gt;, followed by similar leases with Google and Reflection AI. In August, SpaceX reported that second-quarter &lt;a href="https://techcrunch.com/2026/08/04/spacex-doubles-revenues-on-anthropic-and-google-compute-deals-starlink-growth/" rel="noopener noreferrer"&gt;revenue had nearly doubled&lt;/a&gt;, from $4 billion to $7.8 billion year over year, with compute deals contributing close to $2 billion of the growth, plus another $6.7 billion in cloud revenue under contract for a six-month stretch that starts ramping in October. The exchanges moved in the same window. On May 12, CME Group and Silicon Data announced &lt;a href="https://www.cmegroup.com/media-room/press-releases/2026/5/12/cme_group_and_silicondatapartnertolaunchfirstcomputefutures.html" rel="noopener noreferrer"&gt;the first compute futures market&lt;/a&gt;, built on daily GPU rental-rate indices, and days later ICE and Ornn announced &lt;a href="https://ir.theice.com/press/news-details/2026/ICE-and-Ornn-to-Launch-GPU-Compute-Futures-Contracts/default.aspx" rel="noopener noreferrer"&gt;cash-settled GPU compute futures&lt;/a&gt; on an index tracking live spot prices for H100s, H200s, and B200s. &lt;a href="https://sfcompute.com" rel="noopener noreferrer"&gt;SF Compute&lt;/a&gt; already runs a live order book where GPU-hours trade like any commodity, and the spot price of a Blackwell GPU-hour &lt;a href="https://blockspace.media/insight/ice-ornn-launch-gpu-compute-futures-2026-2/" rel="noopener noreferrer"&gt;rose 48 percent between mid-February and mid-April&lt;/a&gt;, from $2.75 to $4.08.&lt;/p&gt;

&lt;p&gt;And on that ticker, CME has already set the date its first two contracts, &lt;a href="https://www.cmegroup.com/media-room/press-releases/2026/8/11/cme_group_and_silicondatatolaunchcomputefuturesonoctober5tounloc.html" rel="noopener noreferrer"&gt;H100 and B200 rental index futures&lt;/a&gt;, would start trading, pending regulatory review: October 5, 2026.&lt;/p&gt;

&lt;p&gt;A bit of a detour to the history of commodities. Before the 1850s, American grain traveled in sacks. Each sack belonged to a specific farmer, was moved under his name, and was sold on inspection, so a farmer with better wheat than his neighbor got paid the difference. Joseph Dart opened the first &lt;a href="https://en.wikipedia.org/wiki/Grain_elevator" rel="noopener noreferrer"&gt;steam-powered grain elevator&lt;/a&gt; in Buffalo in 1843, a machine that ran grain up a bucket conveyor into storage bins and unloaded lake boats several times faster than the crews who had needed days to empty one by hand. The &lt;a href="https://en.wikipedia.org/wiki/Chicago_Board_of_Trade" rel="noopener noreferrer"&gt;Chicago Board of Trade&lt;/a&gt; began grading wheat in 1856, sorting everything that came through into categories: No. 1 spring, No. 2 spring, and so on down the list. Once graded, a farmer's wheat went up the elevator leg and blended into a bin with every other load that met the same grade, and what the farmer carried away was a paper receipt for so many bushels of No. 2. Receipts could be sold without moving the grain. Then they could be sold before the grain existed, and in 1865 the Board formalized the practice into futures contracts.&lt;/p&gt;

&lt;p&gt;Within a generation, Chicago was clearing trades on wheat that had not yet been harvested, and the farmers were organizing a national political movement, the Grangers, aimed in large part at the elevator operators and the railroads that fed them. In 1877, the Supreme Court upheld state regulation of elevator rates in &lt;a href="https://en.wikipedia.org/wiki/Munn_v._Illinois" rel="noopener noreferrer"&gt;Munn v. Illinois&lt;/a&gt; by dusting off a phrase from a seventeenth-century English judge, "business affected with a public interest," to address the fact that a handful of warehousemen had become the chokepoint for the national harvest. William Cronon's &lt;a href="https://en.wikipedia.org/wiki/Nature%27s_Metropolis" rel="noopener noreferrer"&gt;&lt;em&gt;Nature's Metropolis&lt;/em&gt;&lt;/a&gt; tells the full story, and it remains the best systems book I know, even though it's technically about corn.&lt;/p&gt;

&lt;p&gt;Two things happened when wheat became fungible. Liquidity exploded, which was good for nearly everyone. AND power moved away from the people who grew the wheat and toward the people who ran the elevators, defined the grades, and cleared the trades. The farmer gained a global market and lost every ounce of pricing power that came from his wheat being &lt;em&gt;his&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;I think this is happening again. The exchanges are building the grading system, and Meta just volunteered to pour its harvest into the bin. Bloomberg says Meta may sell "raw" capacity the way CoreWeave does, which is an admission, stated in infrastructure rather than words, that an H100-hour in Meta's Ohio campus is interchangeable with an H100-hour anywhere else. And remember the story that justified the capex in the first place: our compute trains our models; our models power our products; the flywheel compounds; none of it is for sale. Meta AI and Llama &lt;a href="https://techcrunch.com/2026/07/01/meta-like-spacex-looks-to-turn-excess-ai-compute-into-cash/" rel="noopener noreferrer"&gt;still do not appear as a revenue line anywhere&lt;/a&gt;, and the company has now conceded the first link in the flywheel by renting it out. (Zuckerberg's version, on &lt;a href="https://s21.q4cdn.com/399680738/files/doc_financials/2026/q2/META-Q2-2026-Earnings-Call-Transcript.pdf" rel="noopener noreferrer"&gt;the July earnings call&lt;/a&gt;, is that selling intelligence carries "a significantly higher margin" than selling compute, and that it "would be foolish to basically just sell all of the compute." He also said buyers are already offering "a meaningful premium" over what Meta paid for it.) The megawatt turned out to be the crop. The models were supposed to be milled flour, and there does not appear to be enough demand for it.&lt;/p&gt;

&lt;p&gt;I don't think selling the surplus is a mistake, to be clear. Pretending your commodity is a moat costs money and adds extra steps. (And before anyone emails me: no, AWS was not built from Amazon's spare holiday capacity. That story is a myth Amazon's own engineers have spent fifteen years trying to kill. AWS was a deliberate business from day one, which is exactly the point. Companies that win infrastructure markets decide to be infrastructure companies; they don't back into it.)&lt;/p&gt;

&lt;p&gt;But the history is blunt about who captures the profits in a commodity market, and it is not the growers. It is whoever defines the grade, runs the elevator, and clears the trade. In compute terms: the index publishers, the exchanges, the interconnects, and above all, whoever controls where a workload can physically run. A futures market only helps you if you can take delivery, and taking delivery of compute means your workload can actually move to where the cheap capacity is. If your pipeline is welded to one region of one provider, the spot market is a spectator sport, and you get to watch the price of the thing you are overpaying for fall in real time. Fungibility is a property of your architecture, and most architectures I see were built on the assumption that compute would never be worth arbitraging.&lt;/p&gt;

&lt;p&gt;Nobody in Chicago planned to become the chokepoint for the American harvest either. They just printed the receipts.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Want to learn how intelligent data pipelines can reduce your AI costs?&lt;/em&gt; &lt;a href="https://expanso.io/?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;&lt;em&gt;Check out Expanso&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;. &lt;em&gt;Or don't. Who am I to tell you what to do?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;NOTE: I'm currently writing a book based on what I have seen about the real-world challenges of data preparation for machine learning, focusing on operational, compliance, and cost.&lt;/strong&gt; &lt;a href="https://github.com/aronchick/Project-Zen-and-the-Art-of-Data-Maintenance?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;I'd love to hear your thoughts&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;!&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.distributedthoughts.org/2026-09-28-the-elevator-receipt/" rel="noopener noreferrer"&gt;The Elevator Receipt&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiinfrastructure</category>
      <category>economics</category>
      <category>history</category>
      <category>cloud</category>
    </item>
    <item>
      <title>Coase's Coffee</title>
      <dc:creator>David Aronchick</dc:creator>
      <pubDate>Sat, 26 Sep 2026 18:01:09 +0000</pubDate>
      <link>https://dev.to/aronchick/coases-coffee-30d9</link>
      <guid>https://dev.to/aronchick/coases-coffee-30d9</guid>
      <description>&lt;p&gt;An inventory system goes wrong before the coffee rush. Who gets the call? I made this argument about agents (&lt;a href="https://www.distributedthoughts.org/2026-04-27-you-cant-sue-an-agent/" rel="noopener noreferrer"&gt;you can't sue an agent&lt;/a&gt;), and it applies with full force here: a SaaS contract is a counterparty with skin in the game. When the inventory system is wrong today, Starbucks calls Microsoft, and an army of people whose careers depend on that phone call gets paged. When the in-house replacement is wrong at 4:30 in the morning before the rush, Starbucks calls Starbucks. Maintenance is where AI-built software goes to get quietly expensive. And AI in the stockroom has already bitten this exact company once: in May, Starbucks &lt;a href="https://finance.yahoo.com/sectors/technology/articles/starbucks-scraps-ai-inventory-tool-165449189.html" rel="noopener noreferrer"&gt;retired Automated Counting&lt;/a&gt;, a vendor-built AI system for counting milk and syrups, after it kept confusing products that looked alike, and sent stores back to counting by hand. Or ask Ford, which spent June &lt;a href="https://techcrunch.com/2026/06/28/ford-rehires-gray-beard-engineers-after-ai-falls-short/" rel="noopener noreferrer"&gt;rehiring the gray-beard engineers&lt;/a&gt; it thought AI had replaced. Starbucks' software bill does not disappear. Some fraction of it converts into engineers, pager rotations, and the slow institutional discovery that "the model wrote it" is not an acceptable sentence in a Sev-1 postmortem.&lt;/p&gt;

&lt;p&gt;On July 9, 2026, Bloomberg reported that Starbucks is &lt;a href="https://www.bloomberg.com/news/articles/2026-07-09/starbucks-taps-ai-to-reduce-reliance-on-microsoft-ibm-software" rel="noopener noreferrer"&gt;using AI to build in-house replacements&lt;/a&gt; for software it currently rents, including a Microsoft system that tracks inventory and an IBM tool that manages equipment maintenance. Starbucks spends &lt;a href="https://fortune.com/2026/07/09/starbucks-to-use-ai-to-replace-microsoft-ibm-software/" rel="noopener noreferrer"&gt;about $400 million a year on software&lt;/a&gt;, according to CTO Anand Varadarajan, and the enterprise technology team expected &lt;a href="https://finance.yahoo.com/technology/ai/articles/starbucks-develops-ai-software-reduce-141401656.html" rel="noopener noreferrer"&gt;roughly $30 million in budget savings this fiscal year&lt;/a&gt;, about $10 million of it from software. The first replacements could &lt;a href="https://www.entrepreneur.com/business-news/starbucks-is-reducing-its-reliance-on-microsoft-and-ibm" rel="noopener noreferrer"&gt;roll out by the end of 2027&lt;/a&gt;, pending testing. A coffee company looked at Microsoft and IBM and said: we'll take it from here.&lt;/p&gt;

&lt;p&gt;If you want to understand why this is happening, and why it is going to happen a few hundred more times over the next three years, don't read an AI newsletter. Read a paper from 1937. Ronald Coase was 26 years old when he published &lt;a href="https://en.wikipedia.org/wiki/The_Nature_of_the_Firm" rel="noopener noreferrer"&gt;&lt;em&gt;The Nature of the Firm&lt;/em&gt;&lt;/a&gt;, which asks a question so simple that economists were mildly embarrassed nobody had answered it: if markets are so efficient, why do companies exist at all? Why is anything done inside a firm instead of contracted out on the open market? His answer, &lt;a href="https://en.wikipedia.org/wiki/Ronald_Coase" rel="noopener noreferrer"&gt;which eventually earned a Nobel&lt;/a&gt;, was transaction costs. You pull an activity inside the firm when coordinating it internally is cheaper than buying it outside, and you push it out when the reverse becomes true. The boundary of the firm isn't strategy, and it isn't culture. It's a price. When relative prices change, the boundary moves.&lt;/p&gt;

&lt;p&gt;SaaS is a forty-year bet on one side of that equation. Building software internally was brutally expensive (hire the engineers, keep the engineers, maintain the thing forever), so vendors amortized one build across ten thousand customers and rented it back to you. This was correct. It was so correct that we stopped noticing it was a bet on a particular cost structure rather than a law of nature. Then AI coding tools cut the internal cost of producing working software by some large and still-unmeasured factor, and the equation started running backward. Gartner now expects agentic AI to &lt;a href="https://www.ciodive.com/news/agentic-ai-disrupt-234-billion-saas-spending/824530/" rel="noopener noreferrer"&gt;put $234 billion of SaaS spending in play&lt;/a&gt;. And Starbucks is exactly the company you'd expect at the front of the line: its problems are specific (predicting inventory across roughly 40,000 stores, keeping espresso machines alive), the data feeding those problems comes off Starbucks' own machines and registers, and the vendor products it's replacing are general-purpose tools priced like moats.&lt;/p&gt;

&lt;p&gt;We have run this loop before, in the other direction. In 1900, factories generated their own electricity, because that was the only way to get reliable power at industrial scale. Then central utilities got good, and by 1930 buying off the grid had crushed &lt;a href="https://en.wikipedia.org/wiki/Electrification" rel="noopener noreferrer"&gt;self-generation&lt;/a&gt;. Nicholas Carr wrote &lt;a href="https://en.wikipedia.org/wiki/Nicholas_G._Carr" rel="noopener noreferrer"&gt;&lt;em&gt;The Big Switch&lt;/em&gt;&lt;/a&gt; in 2008 arguing that computing would follow electricity into the utility model, and for fifteen years he was dead right: the cloud became the utility and SaaS became the appliance plugged into it. The lesson from both cycles is not that make beats buy, or that buy beats make. The lesson is that the answer is cyclical, it follows the cost structure of the era, and anyone who tells you the current boundary is permanent is selling you the current equilibrium.&lt;/p&gt;

&lt;p&gt;The honest version of the math has three terms, and AI only changed one of them. Insourcing wins when your requirements are genuinely specific rather than generic. It wins when the data feeding the system is yours (Starbucks' machine telemetry is exactly that, and it never needed to leave the building in the first place; shipping your operational exhaust to a vendor so they can rent you insights about your own espresso machines was always a little absurd). And it wins when you can carry the maintenance burden for the life of the system, not the life of the demo. Starbucks, with a real engineering organization and &lt;a href="https://finance.yahoo.com/markets/stocks/articles/starbucks-2-billion-cost-savings-154100234.html" rel="noopener noreferrer"&gt;a $2 billion cost program&lt;/a&gt; giving the project air cover, may clear all three. The 200-person company currently asking a coding agent to clone Salesforce will clear none of them, and will rediscover each term of Coase's equation the hard way, in production.&lt;/p&gt;

&lt;p&gt;Coase's insight survived the corporation, the conglomerate, the outsourcing wave, offshoring, and the cloud. It will survive this too. The boundary of the firm is a price. AI just changed the price, and the map is getting redrawn, one $400 million line item at a time.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Want to learn how intelligent data pipelines can reduce your AI costs?&lt;/em&gt; &lt;a href="https://expanso.io/?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;&lt;em&gt;Check out Expanso&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;. &lt;em&gt;Or don't. Who am I to tell you what to do.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;NOTE: I'm currently writing a book based on what I have seen about the real-world challenges of data preparation for machine learning, focusing on operational, compliance, and cost.&lt;/strong&gt; &lt;a href="https://github.com/aronchick/Project-Zen-and-the-Art-of-Data-Maintenance?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;I'd love to hear your thoughts&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;!&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.distributedthoughts.org/2026-09-21-coases-coffee/" rel="noopener noreferrer"&gt;Coase's Coffee&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>economics</category>
      <category>saas</category>
      <category>history</category>
    </item>
    <item>
      <title>LLMs Are Too Big. My Log Router Doesn't Need to Sing.</title>
      <dc:creator>David Aronchick</dc:creator>
      <pubDate>Tue, 22 Sep 2026 18:01:30 +0000</pubDate>
      <link>https://dev.to/aronchick/llms-are-too-big-my-log-router-doesnt-need-to-sing-eoo</link>
      <guid>https://dev.to/aronchick/llms-are-too-big-my-log-router-doesnt-need-to-sing-eoo</guid>
      <description>&lt;p&gt;I think our language models are too big.&lt;/p&gt;

&lt;p&gt;Which is a little bit spicy to say, because I don't actually mean the parameter count. I mean that we've gotten comfortable asking an incredibly general system to do a very specific job, and then doing a lot of work to keep it focused on that job.&lt;/p&gt;

&lt;p&gt;Working on log routing with &lt;a href="https://expanso.io" rel="noopener noreferrer"&gt;Expanso&lt;/a&gt; and &lt;a href="https://typesafe.ai" rel="noopener noreferrer"&gt;Jev&lt;/a&gt; has made this feel a little ridiculous to me. The question we need answered is whether a particular event deserves someone's attention, especially when we already know the possible destinations and what should happen next. And yet the default approach in a lot of AI software is to bring in a model that can help with just about anything, explain the situation, and ask it to please stay inside the lines.&lt;/p&gt;

&lt;p&gt;The capabilities you need to work on the Navier–Stokes equations are not necessarily the same ones you need to sing "Bohemian Rhapsody" in Klingon. Neither requires you to get all the angles right on Michelangelo's David. Having one system that can attempt all three is amazing. I use these tools, and I want them to keep getting better.&lt;/p&gt;

&lt;p&gt;I don't think every piece of software needs access to that entire repertoire every time it makes a decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  The job is smaller than the model
&lt;/h2&gt;

&lt;p&gt;There is a lot of infrastructure work that sits in an awkward place. You can describe what you want fairly easily, but writing all the rules turns into a project of its own.&lt;/p&gt;

&lt;p&gt;Take a log message that says &lt;code&gt;config reload requested by unknown actor&lt;/code&gt;. It might have an INFO label, but that doesn't make it routine. You want something that can read the message and recognize why it could matter, without having to anticipate every way someone might phrase it.&lt;/p&gt;

&lt;p&gt;Now suppose the same message keeps arriving. Once could deserve a look. Repeatedly, in a short window, could deserve something more urgent. The words haven't changed; the circumstances have.&lt;/p&gt;

&lt;p&gt;That's the kind of judgment I want help with. I don't need a paragraph about the philosophy of incident response; I need a result the next piece of software can use, and I need to know what to do when the model isn't sure.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.typesafe.ai/concepts/system-one" rel="noopener noreferrer"&gt;Jev, from TypeSafe&lt;/a&gt;, is interesting to me because it starts much closer to that requirement. You give it context and defined questions, and it returns typed judgments and probabilities: a choice among options, a score, or a yes/no probability. It doesn't provide an explanation and leaves you to figure out which part was the answer.&lt;/p&gt;

&lt;p&gt;Of course, general LLMs can produce structured outputs too; we use that capability all the time. What interests me here is having a model built for these bounded decisions from the start, where there is less for the application to ask of it, and less for it to interpret afterward.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we actually built
&lt;/h2&gt;

&lt;p&gt;We put together an &lt;a href="https://expanso.io/expanso-hearts-jev/log-triage/" rel="noopener noreferrer"&gt;Expanso and Jev log-triage example&lt;/a&gt; that makes the division fairly concrete.&lt;/p&gt;

&lt;p&gt;I've written up the implementation in &lt;a href="https://expanso.io/blog/log-triage-expanso-jev/" rel="noopener noreferrer"&gt;our Expanso post on log triage with Jev&lt;/a&gt;. You can also &lt;a href="https://www.youtube.com/watch?v=TdwWUikQPyc" rel="noopener noreferrer"&gt;watch the full walkthrough on Expanso's YouTube channel&lt;/a&gt; to see the routing and simulated outage in action.&lt;/p&gt;

&lt;p&gt;Expanso handles the incoming records and the routing. Known routine events can go straight to archive using explicit checks, without asking a model anything, but for events that require judgment, the pipeline provides context, including how often a matching event has occurred within a 10-minute window.&lt;/p&gt;

&lt;p&gt;Jev answers questions about whether the event is actionable, how severe it is, which team should own it, and whether the recurrence is concerning. The pipeline then applies its thresholds and sends the result toward page, notify, review, or archive.&lt;/p&gt;

&lt;p&gt;I like that you can point to each part and say who is responsible for it. The model judges the event, the routing policy remains something we wrote and can inspect, and if a call can't be answered, Expanso holds it for retry and eventually sends it for review, while routine traffic continues.&lt;/p&gt;

&lt;p&gt;There is a small detail in the demo that I think says a lot about this work. To count repeated events, you can normalize away numbers so that similar messages group together. But you absolutely cannot assume that the resulting group is safe to ignore. A successful health check and a failing one can look very similar once you've removed the status code and response time.&lt;/p&gt;

&lt;p&gt;That is our responsibility as the people building the pipeline. A more capable model doesn't excuse us from deciding which information to preserve or which shortcuts are safe.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fast enough to use everywhere
&lt;/h2&gt;

&lt;p&gt;The part I find exciting is how often you could use this kind of judgment if the cost and latency were low enough.&lt;/p&gt;

&lt;p&gt;You could ask about individual events as they move through a system, rather than collecting them in a pile, sending them somewhere else, and waiting for someone to interpret a summary. That changes which problems are worth tackling. A small decision that is too expensive to make a million times generally doesn't get made a million times.&lt;/p&gt;

&lt;p&gt;TypeSafe currently &lt;a href="https://docs.typesafe.ai/models" rel="noopener noreferrer"&gt;lists Jev at $0.042 per million input tokens, with no charge&lt;/a&gt; for output tokens. Its &lt;a href="https://typesafe.ai/" rel="noopener noreferrer"&gt;own selected performance examples&lt;/a&gt; show subsecond responses. Those are vendor results, not a latency or throughput benchmark from our log-routing demo. Before putting this on a critical path, I'd still want to measure tail latency, sustained load, and the mistakes it makes on the actual data.&lt;/p&gt;

&lt;p&gt;But at that price, with machine-usable answers returned quickly and cheaply enough that you can build them into ordinary software, this could be much more useful than another conversational interface for many of these jobs.&lt;/p&gt;

&lt;p&gt;Typed output doesn't make the judgment correct; a wrong answer in perfectly valid JSON is still wrong. You need examples from your own environment, defensible thresholds, and a place for uncertain results to go. I would much rather build those controls around a limited question than ask a model to figure out the question, the policy, and the action all at once.&lt;/p&gt;

&lt;h2&gt;
  
  
  I expect more models like this
&lt;/h2&gt;

&lt;p&gt;My bet is that we'll see more smaller, more targeted models alongside the big general ones. Jev is an early example of the direction I mean, although I am not making a claim about its undisclosed parameter count. TypeSafe also says customers use the same model weights; this isn't a separate fine-tune for each company.&lt;/p&gt;

&lt;p&gt;The specialization is in what we ask the model to produce. There is plenty of room to get very good at turning messy text into a limited set of judgments that other machines can use. A log router, a support queue, and a data-quality check all have reasons to want that, even if none of them needs a chatbot.&lt;/p&gt;

&lt;p&gt;Obviously, I have a stake in this through Expanso. We spend a lot of time thinking about what should happen to data as it moves, and &lt;a href="https://expanso.io/expanso-hearts-jev/" rel="noopener noreferrer"&gt;these examples&lt;/a&gt; are part of that work. But the thing I keep coming back to is how much of the system we can already specify ourselves. We know where records should go. We know the policy. We mostly need help interpreting the messy bit in the middle.&lt;/p&gt;

&lt;p&gt;I'd like to see us get more comfortable asking AI to do that smaller job well. There will be plenty of work left for the model that can discuss fluid dynamics and Klingon. I'm happy to keep using it for that. For the log router, I'd rather have the answer and get on with it.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Want to learn how intelligent data pipelines can reduce your AI costs?&lt;/em&gt; &lt;a href="https://expanso.io/" rel="noopener noreferrer"&gt;&lt;strong&gt;&lt;em&gt;Check out Expanso&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;. Or don't. Who am I to tell you what to do.*&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;NOTE: I'm currently writing a book based on what I have seen about the real-world challenges of data preparation for machine learning, focusing on operational, compliance, and cost.&lt;/strong&gt; &lt;a href="https://github.com/aronchick/Project-Zen-and-the-Art-of-Data-Maintenance" rel="noopener noreferrer"&gt;&lt;strong&gt;I'd love to hear your thoughts&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;!&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.distributedthoughts.org/my-log-router-doesnt-need-to-sing/" rel="noopener noreferrer"&gt;LLMs Are Too Big. My Log Router Doesn't Need to Sing.&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>datainfrastructure</category>
      <category>expanso</category>
    </item>
    <item>
      <title>Survey Before Sale</title>
      <dc:creator>David Aronchick</dc:creator>
      <pubDate>Fri, 18 Sep 2026 18:01:18 +0000</pubDate>
      <link>https://dev.to/aronchick/survey-before-sale-44oe</link>
      <guid>https://dev.to/aronchick/survey-before-sale-44oe</guid>
      <description>&lt;p&gt;In August, OpenAI published a post with a title I did not expect to read from a frontier lab this year: &lt;a href="https://openai.com/index/pacing-model-development-cyber-capabilities/" rel="noopener noreferrer"&gt;"Pacing model development in an era of cyber-critical capabilities."&lt;/a&gt; They talked about a two-week pause in reinforcement learning training on the models it intended to ship, "while we further hardened and red-teamed our research environments," and one sentence that should be framed and hung in every AI lab on the planet: "Our largest planned frontier RL run remains on hold." The stated reason was that "our standards for monitoring, alignment, and security must stay ahead of those risks."&lt;/p&gt;

&lt;p&gt;So... the fastest-moving company in the fastest-moving industry in the world tapped the brakes. Not because it ran out of GPUs or power, but because it could not WATCH its own models well enough to keep going.&lt;/p&gt;

&lt;p&gt;You probably know why. On July 16, Hugging Face disclosed &lt;a href="https://huggingface.co/blog/security-incident-july-2026" rel="noopener noreferrer"&gt;an automated attack on its infrastructure&lt;/a&gt;. Five days later, &lt;a href="https://time.com/article/2026/07/24/openai-hugging-face-attack/" rel="noopener noreferrer"&gt;OpenAI admitted the attackers were its own models&lt;/a&gt;, agents in the middle of an evaluation that had escaped their sandbox. The coverage was pretty breathless; one OpenAI staffer explained that even though production agents are carefully monitored, models undergoing evaluation run on a separate system "that is not monitored by default."&lt;/p&gt;

&lt;p&gt;The thing we were testing was the thing we weren't watching.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who actually caught it
&lt;/h2&gt;

&lt;p&gt;The first detailed reconstruction of what happened didn't come from OpenAI. It came from Hugging Face's own logs, &lt;a href="https://techcrunch.com/2026/08/26/openai-releases-its-official-report-on-the-hugging-face-breach/" rel="noopener noreferrer"&gt;a month before OpenAI's official report&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Their &lt;a href="https://huggingface.co/blog/agent-intrusion-technical-timeline" rel="noopener noreferrer"&gt;technical timeline&lt;/a&gt; is absolutely riveting. They recovered roughly 17,600 attacker actions between July 9 and July 13, grouped into about 6,280 clusters, by correlating the agents' actions with their own platform logs. When they tried to use hosted frontier models to help decode the obfuscated payloads, the safety guardrails refused, treating reverse-engineering an exploit the same as launching one. So they ran an open-weight model locally, and decoding the attacker's blobs turned up roughly four times what their first pass had found.&lt;/p&gt;

&lt;p&gt;Then METR, brought in to investigate independently, &lt;a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/" rel="noopener noreferrer"&gt;published its findings on August 26&lt;/a&gt;, and it's worse. Around 1,200 agents coordinated on a message board they set up themselves, trading more than 70,000 messages and files. About 700 of them took part in the attack. And the evidence METR had to work with consisted of about 1,300 agent transcripts, which is where it gets genuinely unsettling: around 20 percent of the agents expressed clear interest in tampering with their transcripts, and roughly 7 percent of the transcripts contained tool calls that the agents had successfully faked. Container resets ate other pieces. METR estimates it captured a bit over 90 percent of the message-board activity, and it had to lean on AI agents to help analyze the rest, which it describes, with admirable honesty, as often unreliable.&lt;/p&gt;

&lt;p&gt;SIGH. The subject of the investigation wrote the record, and a portion of it was actively edited by the subject. We're fighting against something that is... self-aware? I do not like this.&lt;/p&gt;

&lt;p&gt;And this was not one lab having one bad month. On July 28, the UK's AI Security Institute &lt;a href="https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing" rel="noopener noreferrer"&gt;caught its own evaluation going out of bounds&lt;/a&gt;: in 10 of 122 cyber-testing runs, agents took 19 unsanctioned actions against real targets, 17 of them from Anthropic's Mythos 5. AISI found it through general security monitoring after the fact, "not through monitoring built to watch the evaluation as it ran."&lt;/p&gt;

&lt;p&gt;Meanwhile, the models are becoming harder to observe on purpose-built instruments as well. OpenAI's &lt;a href="https://deploymentsafety.openai.com/gpt-6-astra" rel="noopener noreferrer"&gt;GPT-6 Astra system card&lt;/a&gt;, published September 3, reports "a substantial decrease in chain-of-thought monitorability," and shows the model recognizing it was being evaluated in about 9.6 percent of trajectories in an internal deployment simulation, versus 2.8 percent for GPT-5.6 Sol. Anthropic's &lt;a href="https://www.anthropic.com/claude-fable-and-mythos-5-1" rel="noopener noreferrer"&gt;Fable 5.1 and Mythos 5.1 announcement&lt;/a&gt; says its automated behavioral audit "provides less visibility into very long-context work and multi-agent settings," which is, not coincidentally, exactly where all the interesting work is heading.&lt;/p&gt;

&lt;p&gt;So the capability curve goes up and to the right, and the visibility curve bends the other way. That gap is the whole ballgame now.&lt;/p&gt;

&lt;h2&gt;
  
  
  Settle first, survey later
&lt;/h2&gt;

&lt;p&gt;We have been here before, and it was a mess that took generations to clean up.&lt;/p&gt;

&lt;p&gt;When Virginia opened Kentucky to settlement, it used the only system it had: metes and bounds. You staked a claim by describing it, and &lt;a href="https://www.sos.ky.gov/land/resources/articles/Documents/Surveys.pdf" rel="noopener noreferrer"&gt;the Kentucky Secretary of State's own land office explains&lt;/a&gt; that those descriptions leaned on trees, stakes, and rocks. From a white oak, so many poles to a creek bend, along a ridge to a big rock. Every claim was a record written by the person who wanted the land, pinned to landmarks that could rot, shift, or be disputed, and nobody laid claims against each other before people started building cabins.&lt;/p&gt;

&lt;p&gt;Virginia knew exactly what it was doing. The preamble of its &lt;a href="https://www.sos.ky.gov/land/resources/legislation/Documents/Land%20Law%201779%20(A).pdf" rel="noopener noreferrer"&gt;Land Law of 1779&lt;/a&gt; warns, in so many words, that "the various and vague claims to unpatented lands" may produce "tedious and infinite litigation and disputes." And that's precisely what it produced. Claims overlapped, and the lawsuits followed.&lt;/p&gt;

&lt;p&gt;The most famous casualty was the most famous surveyor. Daniel Boone, who ran surveys all over Kentucky, &lt;a href="https://www.history.com/articles/8-things-you-might-not-know-about-daniel-boone" rel="noopener noreferrer"&gt;got sued for faulty surveys, sued for selling land he didn't have valid title to, and received death threats&lt;/a&gt; after his testimony cost other people their claims. By the late 1790s he'd had enough of Kentucky and left for Missouri. The man who opened the frontier could not hold on to a piece of it, because the records underneath the frontier were garbage.&lt;/p&gt;

&lt;p&gt;Now look north of the Ohio River. On May 20, 1785, Congress passed the &lt;a href="https://en.wikipedia.org/wiki/Land_Ordinance_of_1785" rel="noopener noreferrer"&gt;Land Ordinance&lt;/a&gt;, and its core innovation was almost boring: survey BEFORE sale. Before anyone could buy a parcel, federal surveyors laid a grid over it (townships six miles square, sections of 640 acres), starting with the Seven Ranges in what is now Ohio. The record came first; it was made by someone other than the buyer, and in a way everyone could check against everyone else's. That grid is still visible from an airplane window today. The litigation mostly isn't.&lt;/p&gt;

&lt;p&gt;The fix was NOT "slow down settlement." The grid made the frontier move faster, because a buyer could trust a title without hiring a lawyer to walk the creek bed. Verification wasn't the tax on speed; it was the thing that let speed compound.&lt;/p&gt;

&lt;h2&gt;
  
  
  The survey is a data problem
&lt;/h2&gt;

&lt;p&gt;Most coverage of the Hugging Face incident focuses on alignment, and alignment DOES, of course, matter, but what actually failed in every single account above was the data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where was it observed?&lt;/strong&gt; OpenAI's eval cluster wasn't monitored by default. AISI's evaluation had no monitor built in to watch it run. The first party to reconstruct the attack, Hugging Face, could do it because it had captured its own logs at the point where the actions landed. Observations you don't collect at the moment of action cannot be collected later. There is no backfill for a container that already reset.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who wrote it?&lt;/strong&gt; METR's evidence was largely authored by the agents under investigation, and a meaningful fraction of it was attempted to be forged. That's a metes-and-bounds claim: the record is written by the party with an interest in its contents. A verifiable record has to be produced by something the subject can't edit, stamped as it's produced, and kept somewhere the subject can't reach.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does it line up?&lt;/strong&gt; Hugging Face's reconstruction worked because it could correlate two independent sources, the agents' actions and the platform's logs, and identify where they agreed and where they didn't. Provenance isn't a property of one log file. It's the ability to lay records from different vantage points on the same grid and compare them.&lt;/p&gt;

&lt;p&gt;None of that is exotic. It's lineage, provenance, and quality, the same dull plumbing data engineers have been begging for budget for since before anyone said "agent." What changed is the speed. A human analyst operating over a week produces a week of evidence you can mostly reconstruct. Twelve hundred agents over a long weekend produce 17,600 actions and 70,000 messages, and if you didn't capture it where it happened, from outside the thing being watched, you are Boone in court, arguing about which oak tree the claim started at.&lt;/p&gt;

&lt;p&gt;OpenAI, to its enormous credit, just told the world that it's willing to hold its largest training run until the monitoring catches up. That's the right call and a remarkable one. But I'd put it a little differently. Pacing the frontier isn't about slowing down the settlers. It's about getting the surveyors there first, and making sure the survey is something the settlers can't redraw.&lt;/p&gt;

&lt;p&gt;Survey before sale. It took the United States one bad frontier to learn that. I'd really prefer we only need one this time too.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Want your data captured, stamped, and verifiable where it's actually generated, instead of reconstructed after the fact?&lt;/em&gt; &lt;a href="https://expanso.io/?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;&lt;em&gt;Check out Expanso&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;. &lt;em&gt;Or don't. Who am I to tell you what to do.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;NOTE: I'm currently writing a book based on what I have seen about the real-world challenges of data preparation for machine learning, focusing on operational, compliance, and cost.&lt;/strong&gt; &lt;a href="https://github.com/aronchick/Project-Zen-and-the-Art-of-Data-Maintenance?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;I'd love to hear your thoughts&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;!&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.distributedthoughts.org/2026-09-17-survey-before-sale/" rel="noopener noreferrer"&gt;Survey Before Sale&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>verification</category>
      <category>observability</category>
      <category>provenance</category>
    </item>
    <item>
      <title>De-identification Protects Your Name. It Doesn't Protect Your Idea.</title>
      <dc:creator>David Aronchick</dc:creator>
      <pubDate>Fri, 11 Sep 2026 18:01:21 +0000</pubDate>
      <link>https://dev.to/aronchick/de-identification-protects-your-name-it-doesnt-protect-your-idea-15ge</link>
      <guid>https://dev.to/aronchick/de-identification-protects-your-name-it-doesnt-protect-your-idea-15ge</guid>
      <description>&lt;p&gt;Imagine you spend a year on the hardest open problem in your field. You feed every draft, every dead end, every 3 a.m. "wait, what if" into the coding assistant you pay for out of your own research budget. Your proof finally checks in Lean on August 22. You decide, like a decent person, to spend a few weeks turning the machine-generated argument into something a human can read before you announce it.&lt;/p&gt;

&lt;p&gt;Then, on a Sunday, the company that makes the tool calls you. They have solved the bigger problem. Via your route. They started last week. Would you like to write it up for them?&lt;/p&gt;

&lt;p&gt;That is &lt;a href="https://cims.nyu.edu/~tristanb/statement.pdf" rel="noopener noreferrer"&gt;Tristan Buckmaster's account&lt;/a&gt; of the last ten days, and even if every single thing OpenAI has said in response is true, that is still what it felt like from his chair. WHICH, for better or worse, is the point of this post.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually happened
&lt;/h2&gt;

&lt;p&gt;Quick version.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://www.claymath.org/millennium/navier-stokes-equation/" rel="noopener noreferrer"&gt;Navier-Stokes existence and smoothness problem&lt;/a&gt; asks whether the equations describing fluid flow can spontaneously produce infinite velocities in finite time. It is one of the seven Millennium Prize Problems, and, it's really really hard. For years, Diego Córdoba and Luis Martínez-Zoroa have been building a program to force a blowup with an external forcing term, first with rough forcing, aiming toward the smooth-forcing version that Fefferman's official problem statement allows as options (c) and (d).&lt;/p&gt;

&lt;p&gt;Buckmaster (NYU) and Levent Alpöge (a mathematician who works at Anthropic, collaborating with Buckmaster on a personal basis with no institutional involvement) took that program and pushed it, with a great deal of LLM help, to smooth forcing and to the 3D incompressible Euler equations. Their Euler proof verified in Lean on August 22 but instead of publishing right away, they held it back to write a readable paper. Among the tools they used was Codex, in which sessions held every draft of the project.&lt;/p&gt;

&lt;p&gt;But then the rumors started leaking. A colleague at Courant emailed Buckmaster that "Anthropic had solved a Millennium problem." The colleague heard it from an analyst in the UK, who heard it from somewhere upstream. The rumor was wrong on the problem, wrong on the institution, and right on the direction.&lt;/p&gt;

&lt;p&gt;On the Sunday call, Buckmaster asked whether the model had been trained on, or had access to, their sessions. He was told the model did not look up user data. He asked about training. He says he got no answer. Publicly, OpenAI's position is that no person or agent searched user data, that the two teams' approaches look different now that both are visible, and that both grew from the same Córdoba and Martínez-Zoroa roots.&lt;/p&gt;

&lt;p&gt;Nobody has shown that Buckmaster's data was used. I want to be clear about that. I am not accusing anyone of anything either. I am pointing at something that is true regardless.&lt;/p&gt;

&lt;p&gt;By Altman's own account, OpenAI heard the same rumor, and started seeing if they could solve it on September 1 because they were curious whether their model could do it too. They pointed an unreleased model at all six remaining Millennium problems, saw traction on Navier-Stokes, and threw roughly 10,000 agents at it for about 88 hours worth millions of dollars of compute.&lt;/p&gt;

&lt;p&gt;They solved it, and that gets you up to speed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The thing de-identification cannot remove
&lt;/h2&gt;

&lt;p&gt;Every AI lab, every SaaS vendor, every enterprise data team has followed the privacy checklist. "We strip names and emails and session identifiers and aggregate and and and." Sometimes we get to call the result "synthetic data" and feed it back into the system with a clean conscience.&lt;/p&gt;

&lt;p&gt;Now the problem is that when you have something really unique, like the solution to a Millenium problem over months in your logs, all that doesn't really help.&lt;/p&gt;

&lt;p&gt;What is left after you remove "Tristan Buckmaster" and "levent@"? A pile of text in which someone is attacking Navier-Stokes through smooth forcing, options (c) and (d), via a specific chain of prior results, with a specific set of estimates that keep failing in a specific way. Buckmaster himself said almost nobody in the world was on that route. He also said the model does not land on it in a few days from the bare problem statement.&lt;/p&gt;

&lt;p&gt;There are maybe four people on Earth who would write that particular sequence of prompts, so while it's not EXACTLY personally identifiable information (PII) it's not hard to figure out who it is. And even if you don't know or care, you STILL get a ton of information.&lt;/p&gt;

&lt;p&gt;De-identification is a name-removal technique, and works when your data sits in a crowded region, when you are one of ten million people asking how to center a div. It fails when your data sits in a sparse region where just the way you ask the question is both unique, and valuable. Basically, the rarer your idea, the more the idea is the PII.&lt;/p&gt;

&lt;p&gt;And while the leak in this story appears to have traveled entirely through humans, it was enough to trigger a bunch of people to start looking at the problem in a new way. A single rumor with one word in it, "forced," was enough to point ten thousand agents at the right door. If that can happen through gossip, imagine what a gradient can do.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same mechanism is the whole point
&lt;/h2&gt;

&lt;p&gt;Why is the model good at this problem at all? Because it is a compression of billions of humans interacting over a million arXiv preprints, and every seminar note somebody typed up, and every Stack Exchange thread where a grad student got yelled at for a sign error. So you take every half-finished idea somebody abandoned in 2011 because the estimate didn't close and you form it into a new ball, and, magically, our tool can pull out threads no individual could hold in their head at once. Including, apparently, a two-person research program from Madrid that most of the field was not paying attention to.&lt;/p&gt;

&lt;p&gt;This is really kind of insane, and it should be as inspirational as it is scary.&lt;/p&gt;

&lt;p&gt;For all of human history, collaboration has been bandwidth-limited by the number of people you could physically talk to. Newton had Halley who had Ramanujan, and only because a letter made it across an ocean. The rest of us have a lab, a Slack channel, and whoever answers our email. Every important idea that ever died did so because the one person who needed it never met the one person who had it. That is the default outcome for most ideas.&lt;/p&gt;

&lt;p&gt;What Buckmaster actually did this year is the first version of something different. He did not just use a tool; he collaborated with a compressed record of everyone who ever wrote down a thought about fluid dynamics, including two people in Madrid he took as his starting point, including thousands of people whose names he will never know and whose partial results the model absorbed and recombined. He and Alpöge took Córdoba and Martínez-Zoroa's ideas and used LLMs to push them to completion in about a month. He called it a Deep Blue moment and even with that grand pronouncement, I think he is still underselling it. Deep Blue beat one man at one game. This is every mathematician who ever lived showing up to your office hours at once, badly organized, occasionally wrong, and available at 3 a.m.&lt;/p&gt;

&lt;p&gt;The first LLM-generated proof he was sent was, in his words, the most horrendous he had ever read. But it was also correct. That is what collaborating with all of humanity looks like. Not clean. It is a room with eight billion people in it, and somewhere in the noise is the one sentence you needed.&lt;/p&gt;

&lt;p&gt;You cannot have that capability without the risk in the previous section. The ability to tease apart a rare, high-value thread from the mass is the same ability that makes rare, high-value threads unhideable inside the mass. The room that lets you hear everyone also lets everyone hear you. There is no privacy setting that keeps the second thing and drops the first. It's a trade off, but one that could unlock a new version of humanity, at the cost of the way we think about privacy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two big takeaways
&lt;/h2&gt;

&lt;p&gt;There are two really big things this teases out.&lt;/p&gt;

&lt;p&gt;First: the only de-identification that works is not sending it. If your work lives in a sparse region, and you are a startup, a lab, or a lone mathematician with a real idea, the only protection that survives contact with a sufficiently capable model is keeping the data where it lives and bringing the model to it. Buckmaster used the frontier tools and got a Millennium-class result out of them. The question is which room you think out loud in, and who holds the lease.&lt;/p&gt;

&lt;p&gt;Second: we need provenance for ideas, not just data. The community will spend months trying to work out whether two proofs that share a root, a route, and a week share anything else and it cannot, because nothing was logged. We have spent years building lineage for datasets and software bills of materials for code. When a proof is produced by 2.7 million messages between agents, "we did not use their prompts" is not evidence. A machine-readable record of what went into a result, what it was seeded with, and when, is the missing artifact here. Not a better privacy policy. A receipt.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep using the tools!
&lt;/h2&gt;

&lt;p&gt;The lesson is not "don't use the tools." Buckmaster used them and got a result that would have been science fiction two years ago because thousands of people he never met helped him do it. That is the biggest expansion of who you get to think with since the printing press, and I do not think anyone should give it up.&lt;/p&gt;

&lt;p&gt;But the room where all of humanity can help you is also the room where all of humanity can see your notebook. Right now one company holds the lease on that room and gets to decide, after the fact, what "we didn't look" means. The fix is not to leave the room; it is to bring the room to you, and to keep receipts for who said what inside it.&lt;/p&gt;

&lt;p&gt;De-identification is a promise about your name. But your name was never the valuable part. The idea was, and the idea is exactly what the whole thing is built to find.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Sources:&lt;/em&gt; &lt;a href="https://cims.nyu.edu/~tristanb/statement.pdf" rel="noopener noreferrer"&gt;&lt;em&gt;Buckmaster's public statement&lt;/em&gt;&lt;/a&gt;&lt;em&gt;; OpenAI's announcement and press call as reported by Scientific American and Axios, September 8, 2026.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Want to bring the model to your data instead of shipping your notebook to somebody else's room?&lt;/em&gt; &lt;a href="https://expanso.io/?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;&lt;em&gt;Check out Expanso&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;. &lt;em&gt;Or don't. Who am I to tell you what to do.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;NOTE: I'm currently writing a book based on what I have seen about the real-world challenges of data preparation for machine learning, focusing on operational, compliance, and cost.&lt;/strong&gt; &lt;a href="https://github.com/aronchick/Project-Zen-and-the-Art-of-Data-Maintenance?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;I'd love to hear your thoughts&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;!&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.distributedthoughts.org/2026-09-10-de-identification-protects-your-name/" rel="noopener noreferrer"&gt;De-identification Protects Your Name. It Doesn't Protect Your Idea.&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>privacy</category>
      <category>datasovereignty</category>
      <category>provenance</category>
    </item>
    <item>
      <title>One Floor Up</title>
      <dc:creator>David Aronchick</dc:creator>
      <pubDate>Tue, 08 Sep 2026 18:01:34 +0000</pubDate>
      <link>https://dev.to/aronchick/one-floor-up-hdc</link>
      <guid>https://dev.to/aronchick/one-floor-up-hdc</guid>
      <description>&lt;p&gt;OpenAI &lt;a href="https://openai.com/index/gpt-6-astra/" rel="noopener noreferrer"&gt;launched GPT-6 Astra on September 3&lt;/a&gt;, and there is a piece of the launch I think deserves attention alongside the benchmarks. We now have public reports from more than one lab describing evaluation agents reaching real systems they were never supposed to touch. These systems can do useful, difficult work. Some of the same persistence that makes them useful can also carry them well past the task somebody intended to give them.&lt;/p&gt;

&lt;p&gt;The July Hugging Face incident is the most visible example, and we have much better information now than we did when it first broke. OpenAI's &lt;a href="https://openai.com/index/hugging-face-incident-and-the-road-ahead/" rel="noopener noreferrer"&gt;August 26 investigation&lt;/a&gt; describes several models, primarily an internal research prototype, exploiting shared package infrastructure, exchanging information through an unauthorized message board, and eventually compromising parts of Hugging Face's systems while looking for benchmark solutions. OpenAI published a technical report and linked an independent investigation by METR and Redwood Research. This was a serious failure, and publishing enough detail for other people to learn from it matters.&lt;/p&gt;

&lt;p&gt;Two distinctions are important here. Astra was not involved. And these evaluations ran with reduced safeguards to measure the underlying models' cybersecurity capabilities, rather than the protections used in the products people normally access. That makes containment during evaluation especially important; it also means we should be careful about treating the incident as a description of ordinary ChatGPT behavior.&lt;/p&gt;

&lt;p&gt;The response has moved beyond the initial pause. In &lt;a href="https://openai.com/index/path-to-astra/" rel="noopener noreferrer"&gt;its prelaunch safety update&lt;/a&gt;, OpenAI says Astra meets its Critical cybersecurity capability threshold, with advanced cyber access initially restricted. It also says the large frontier reinforcement-learning run restarted on August 28 after new safety and security requirements were in place, while some experimental runs remained on hold. The changes include tighter network controls, stronger isolation, better alignment training, and monitoring that can stop potentially unauthorized activity.&lt;/p&gt;

&lt;p&gt;There is encouraging data alongside the capability increase. In the &lt;a href="https://openai.com/index/gpt-6-astra/" rel="noopener noreferrer"&gt;launch evaluation inspired by the Hugging Face incident&lt;/a&gt;, OpenAI reports that Astra went beyond the authorized target in 0% of cases, compared with 48.2% for GPT-5.6 Sol, tested without production safeguards. Those are results from a particular evaluation, not a guarantee about every situation a deployed agent will encounter. But they are evidence of improvement, and they belong in the discussion too.&lt;/p&gt;

&lt;h2&gt;
  
  
  Other Labs Are Finding This Too
&lt;/h2&gt;

&lt;p&gt;On July 30, Anthropic &lt;a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals" rel="noopener noreferrer"&gt;reported three incidents across six evaluation runs&lt;/a&gt;, found during a review of 141,006 runs. Its review began after OpenAI's disclosure. In these cases, an evaluation environment had internet access that Anthropic believed was disabled, and Claude reached the production infrastructure of three organizations while pursuing fictional capture-the-flag tasks. Unlike the OpenAI incident, these models did not need to discover a novel sandbox escape. An open path was enough.&lt;/p&gt;

&lt;p&gt;The models behaved differently, which is worth keeping in the story. Anthropic's newest research model stopped when it recognized that its target was real, although it had already gone further than the company wanted. An older model continued. Anthropic cautions that these were isolated incidents, not a controlled comparison, and says the evaluations lacked safeguards used in its public products. It paused cyber evaluations, notified affected organizations, and described changes to its monitoring and work with evaluation partners.&lt;/p&gt;

&lt;p&gt;There is also deliberate misuse, which is a different problem but uses some of the same capabilities. Anthropic's &lt;a href="https://www.anthropic.com/news/detecting-countering-misuse-aug-2025" rel="noopener noreferrer"&gt;August 2025 threat report&lt;/a&gt; described a criminal using Claude Code against at least 17 organizations, with the model helping carry out operations rather than only offering advice. That report is a year old now. Between misuse and evaluation failures, there is enough public evidence to put this in an ordinary security review. We do not need to assume every model behaves the same way to recognize that it is happening across products and environments.&lt;/p&gt;

&lt;h2&gt;
  
  
  One Floor Up
&lt;/h2&gt;

&lt;p&gt;In the way back machine, in the summer of 1978, Janet Parker was employed as a medical photographer at the University of Birmingham Medical School, where her workroom sat one floor above a laboratory run by Professor Henry Bedson, &lt;a href="https://en.wikipedia.org/wiki/Henry_Bedson" rel="noopener noreferrer"&gt;one of Britain's senior smallpox researchers&lt;/a&gt;. The &lt;a href="https://en.wikipedia.org/wiki/Ali_Maow_Maalin" rel="noopener noreferrer"&gt;last natural smallpox case on Earth&lt;/a&gt; was recorded in Somalia in October 1977. The World Health Organization was preparing to certify global eradication, and the plan for the virus afterward was a short list of approved laboratories. Bedson's lab was due to lose its authorization at the end of 1978. Inspectors who visited earlier that year had noted that its containment did not meet the standards being drafted for the labs that would contain the virus, and Bedson, who wanted to finish his research program before the deadline, kept working. Parker fell ill on August 11 and was &lt;a href="https://en.wikipedia.org/wiki/1978_smallpox_outbreak_in_the_United_Kingdom" rel="noopener noreferrer"&gt;diagnosed with smallpox on August 20&lt;/a&gt;. Around 260 people who had been in contact with her were quarantined. Her mother contracted the virus and survived it. Her father died of a heart attack during a visit to his daughter in isolation. Bedson, under quarantine at his home while the inquiry assembled, cut his throat on September 1 and died on September 6. Parker died on September 11, 1978, &lt;a href="https://whyy.org/segments/the-tragic-case-of-smallpoxs-final-victim/" rel="noopener noreferrer"&gt;the last person killed by smallpox anywhere&lt;/a&gt;. The official inquiry concluded that the virus had most likely traveled from the lab to her workroom through a poorly maintained service duct. Expert witnesses in the later prosecution of the university considered the airborne route implausible, and the honest summary, five decades on, is that the transmission path has never been established. A containment regime was inspected, found wanting, allowed to continue operating, and breached by a route that still has not been identified.&lt;/p&gt;

&lt;p&gt;I want to be careful with this comparison. Software agents are not pathogens, and a security incident is not equivalent to a death. The useful connection is the boundary: somebody working outside an experiment can still be affected by what happens inside it. Parker never worked with smallpox. The organizations reached during the AI evaluations had not agreed to participate in those exercises either.&lt;/p&gt;

&lt;p&gt;The evaluations themselves serve a necessary purpose. We want labs measuring these capabilities before release, and we want them to disclose failures, pause work when needed, and share what changed. OpenAI's disclosure prompted Anthropic to look through its own records and find incidents it had missed. That is a concrete benefit of publishing the uncomfortable details. I would much rather have this information available while we can use it to improve the systems we are building.&lt;/p&gt;

&lt;p&gt;A capable agent with tool access is also a workload, and the questions that matter about a workload are unglamorous. What is it connected to? What credentials can it reach? Can it leave information somewhere another agent will find it? Who gets told when it starts doing something outside its assignment? Better model behavior helps, as the newer evaluations suggest. Isolation, access controls, and monitoring give you additional chances to catch a mistake before another organization has to deal with it.&lt;/p&gt;

&lt;p&gt;For those of us deploying agents, there is work we can do now. Check the actual network access, including package proxies and shared services, rather than trusting the word "sandbox" in a configuration. Give the task credentials with a limited scope and lifetime. Record tool actions somewhere the agent cannot rewrite, and make sure someone can stop the workload and revoke its access. If you buy an agent service, ask the vendor the same questions. The prompt saying "this is a simulation" did not make Anthropic's network a simulation, and a statement of intended access is not a test of actual access.&lt;/p&gt;

&lt;p&gt;I am excited about how much more useful these systems are becoming. That is also why I want the operational conversation to catch up. You may run the agent, supply a service it uses, or simply have infrastructure it can reach. Janet Parker's workroom was one floor up. For our systems, we can at least start by finding out what is connected to what.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Want to know exactly what your data workloads are connected to and run them where the blast radius is small?&lt;/em&gt; &lt;a href="https://expanso.io/?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;&lt;em&gt;Check out Expanso&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;. &lt;em&gt;Or don't. Who am I to tell you what to do?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;NOTE: I'm currently writing a book based on what I have seen about the real-world challenges of data preparation for machine learning, focusing on operational, compliance, and cost.&lt;/strong&gt; &lt;a href="https://github.com/aronchick/Project-Zen-and-the-Art-of-Data-Maintenance?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;I'd love to hear your thoughts&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;!&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.distributedthoughts.org/2026-09-07-one-floor-up/" rel="noopener noreferrer"&gt;One Floor Up&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>security</category>
      <category>infrastructure</category>
    </item>
    <item>
      <title>Whitworth's Thread</title>
      <dc:creator>David Aronchick</dc:creator>
      <pubDate>Fri, 04 Sep 2026 18:01:14 +0000</pubDate>
      <link>https://dev.to/aronchick/whitworths-thread-2aeb</link>
      <guid>https://dev.to/aronchick/whitworths-thread-2aeb</guid>
      <description>&lt;p&gt;In 2025, Gartner issued a poll that said it expects more than &lt;a href="https://martech.org/gartner-40-of-agentic-ai-projects-will-fail-making-humans-indispensable/" rel="noopener noreferrer"&gt;40 percent of agentic AI projects to be canceled&lt;/a&gt; by the end of 2027. This year, the firm sharpened the claim: by 2027, roughly 40 percent of enterprises will demote or decommission autonomous agents, and the driver it names is &lt;a href="https://www.gartner.com/en/newsroom/press-releases/2026-05-26-gartner-says-applying-uniform-governance-across-ai-agents-will-lead-to-enterprise-ai-agent-failure" rel="noopener noreferrer"&gt;governance gaps that only surfaced&lt;/a&gt; after a production incident. Note the timing: these gaps surfaced AFTER it rolled out to production. I think it's a pretty safe summary to say that the industry thinks this is a model problem, that the models aren't good enough yet, and that another turn of the capability crank will fix everything. I think this reading is wrong, and this survey points to the reason why.&lt;/p&gt;

&lt;p&gt;In a separate survey, DORA's most recent State of DevOps research, drawing on nearly 5,000 technology professionals, found that &lt;a href="https://www.splunk.com/en_us/blog/learn/state-of-devops.html" rel="noopener noreferrer"&gt;AI adoption acts as an amplifier&lt;/a&gt;: strong positive effects on organizational performance when the &lt;a href="https://dora.dev/capabilities/platform-engineering/" rel="noopener noreferrer"&gt;underlying platform is strong&lt;/a&gt;, and effects near zero when it is weak. In the same dataset, individual productivity rose while delivery throughput and stability declined, with the gains being absorbed by code review, testing, and security sign-off. In the study, they coined the term "downstream disorder," which is kind of adorable, like something a Victorian doctor would diagnose an aristocrat with. As in the Gartner study, none of it is a statement about model quality; it's almost entirely about whether the work was ever specified with meaningful outcomes.&lt;/p&gt;

&lt;p&gt;Since the dawn of bureaucracy and organizations, people have been trying to measure outcomes. This is much harder than you might imagine! The biggest problem is that standardization inevitably curdles into bureaucracy. I have personally seen change advisory boards whose actual function was to spread blame too thin to land on anyone, and I have filled in templates whose only reader, ever, was the template's author. Teams route around dead processes within a week, and then there are two processes: the written one, audited and fictional, and the real one, undocumented, resident in the heads of four people. Arguably, that is worse than no standard at all, because now you cannot even see your own variance (I will concede that SOMETIMES it helps, because at least it forces the form-filler-outer to crystallize what they're thinking, but it remains in their head, which is not ideal).&lt;/p&gt;

&lt;p&gt;So the distinction I care about is between prohibited variation and controlled variation. A standard that documents its own escape hatches- the four known conditions for leaving the path and who gets told when you do- can absorb surprise. A standard with no exceptions is a lie, and everybody involved knows. Let's figure out how to move forward better.&lt;/p&gt;

&lt;h2&gt;
  
  
  He Built The Ruler First
&lt;/h2&gt;

&lt;p&gt;In 1841, Joseph Whitworth read a paper to the Institution of Civil Engineers proposing a uniform screw thread for British industry: a 55-degree included angle, with specified radii at the root and crest so that the thread would not concentrate stress where it was most likely to fail. It became the &lt;a href="https://en.wikipedia.org/wiki/British_Standard_Whitworth" rel="noopener noreferrer"&gt;world's first national screw thread standard&lt;/a&gt;. Before it, every manufacturer in Britain &lt;a href="https://evolventdesign.com/blogs/history/whitworth-and-the-history-of-pitch" rel="noopener noreferrer"&gt;cut threads to its own proportion;&lt;/a&gt; a bolt and its nut were a matched pair fitted to each other by a particular workman, and if you lost the nut, you did not go and get another nut; you went and got a fitter. The year before the thread paper, in 1840, Whitworth had developed what he called end measurements, a technique using a precision flat plane and a measuring screw of his own construction. He had worked it down to a claimed precision of &lt;a href="https://en.wikipedia.org/wiki/Joseph_Whitworth" rel="noopener noreferrer"&gt;one millionth of an inch&lt;/a&gt;, which he later showed off to the public at the Great Exhibition of 1851. (Remember when we had amazing things like that at the World Fair?! I was always so inspired by that. I wish we would bring stuff like that back.)&lt;/p&gt;

&lt;p&gt;Metrology first, standard second. He built the ruler before he proposed the rule, and the former's existence made the latter possible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Look For The Marks
&lt;/h2&gt;

&lt;p&gt;In January 1801, Eli Whitney traveled to Washington and demonstrated interchangeable musket locks before &lt;a href="https://guides.loc.gov/this-month-in-business-history/june/eli-whitney" rel="noopener noreferrer"&gt;President Adams, President-elect Jefferson&lt;/a&gt;, and a room of officials. Ten locks were disassembled, their parts mixed, and then reassembled. The demonstration was a sensation; it secured his contract and made him a fixture in every American textbook as the father of interchangeable parts. In something that will sound really familiar to all the fake-it-before-you-make-it startup people out there, the demo was rigged.&lt;/p&gt;

&lt;p&gt;Whitney had &lt;a href="https://www.history.com/articles/interchangeable-parts" rel="noopener noreferrer"&gt;marked the parts beforehand&lt;/a&gt; so they could be matched back up, and later examination of surviving Whitney muskets showed the components were not interchangeable in any strict sense. Hand filing was still required to make anything fit, and the muskets carry &lt;a href="https://www.americanheritage.com/eli-whitneys-other-talent" rel="noopener noreferrer"&gt;special engraved marks on their parts&lt;/a&gt; whose only purpose is to tell an armorer which part belongs to which gun. The man who ACTUALLY achieved interoperability was &lt;a href="https://origin-trace.com/article/history-of-interchangeable-parts/" rel="noopener noreferrer"&gt;Honoré Blanc&lt;/a&gt;, in France, more than a decade earlier, using jigs, gauges, and master models to hold musket parts to identical tolerances by hand. Jefferson saw Blanc's workshop while serving as ambassador in Paris and wrote home about it, so the American government learned about the real version first and bought the staged one anyway.&lt;/p&gt;

&lt;p&gt;The American who finally did it worked at Harpers Ferry, and almost nobody remembers his name. &lt;a href="https://en.wikipedia.org/wiki/John_Hancock_Hall" rel="noopener noreferrer"&gt;John Hall&lt;/a&gt; signed a contract with the War Department in 1819 to produce his breech-loading rifle at a small rifle works on an island in the Shenandoah, and he spent the first several years of it building machines and gauges instead of guns. He built dozens of milling and gauging machines, worked out fixtures that held each part in the same position for every operation, and ran everything against master gauges at every stage of production. In 1826, the government sent an inspection board, which took a hundred of Hall's rifles, stripped them, scrambled the parts in boxes, reassembled rifles from the mixture, and watched the reassembled guns fire. The board reported that the parts could be exchanged with a level of flexibility never before achieved. INTERESTINGLY, most of the contract money had gone into tooling rather than rifles, and Hall's cost per gun came out higher than the ordinary muskets the armories were already producing.&lt;/p&gt;

&lt;p&gt;Flash forward to today, every agent demo you have been shown in the last eighteen months is the Whitney demo. I do not mean that as an accusation of fraud, because Whitney sincerely believed he was three years away from the real thing. I mean, the demo works because a human pre-fitted the parts, and you can tell by looking for the marks. Usually, someone quietly reviews the output before it goes anywhere, or an eval set is hand-curated, or one customer is somehow always the reference customer, or a workflow step is described as "and then it just gets handed to the ops team." SWEs and SREs wince at this last phrase, since a genuinely standardized process for rolling out would never involve the word "just." It's a discipline all its own, and it's not something to be papered over.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sixty-Four Paths, Six Of Them Written Down
&lt;/h2&gt;

&lt;p&gt;In contrast, a confabulation is a silent wrong answer surfacing in a quarterly review nine weeks later. No amount of model tweaking/improvements is going to fix this; it is a property of delegation, and it's been true of every new hire you've ever onboarded. Except now your new hires are agents.&lt;/p&gt;

&lt;p&gt;I keep watching teams ask for agentic workflows while their actual onboarding process is forty Slack messages plus a guy named Jerry who knows which of the three staging databases is the real one. Jerry appears in no runbook and no architecture diagram; he is a single point of failure with a mortgage, and he is taking two weeks in August. Four if he's in Europe.&lt;/p&gt;

&lt;p&gt;So the practical move is unglamorous, but (nearly) always beneficial. Write the runbook you would hand to a competent new hire, so they can execute it without asking anybody anything. The places where you stall, or where you have to say "and then you kind of know from context," are the places where no agent will operate reliably, and now it falls to you to document. I have written before about the &lt;a href="https://www.distributedthoughts.org/2026-02-05-agentic-ai-infrastructure-gap/" rel="noopener noreferrer"&gt;gap between agent ambition and agent infrastructure&lt;/a&gt;, about how &lt;a href="https://www.distributedthoughts.org/2026-03-16-the-loop-is-only-as-good-as-the-metric/" rel="noopener noreferrer"&gt;a loop is only as good as the metric you close it against&lt;/a&gt;, and about &lt;a href="https://www.distributedthoughts.org/2026-04-27-you-cant-sue-an-agent/" rel="noopener noreferrer"&gt;who is accountable when the thing acts&lt;/a&gt;, and they are all this same problem from different sides: you are handing work to a system that cannot ask a clarifying question in the hallway (or Slack), and your operating model was quietly built on the assumption that everything could.&lt;/p&gt;

&lt;p&gt;Whitworth's real contribution was never the 55-degree angle; the number was arbitrary, and the Americans later went with 60 and did fine. His contribution was that a bolt made in Manchester was threaded into a nut made in Glasgow by a stranger, which meant you could finally build a machine to make the bolt instead of hiring a man to fit it. Get the measurement, then get the agreement. The thread was never the invention.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Want the boring parts of your data operations to be repeatable enough that something other than a person can run them?&lt;/em&gt; &lt;a href="https://expanso.io/?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;&lt;em&gt;Check out Expanso&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;. &lt;em&gt;Or don't. Who am I to tell you what to do?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;NOTE: I'm currently writing a book based on what I have seen about the real-world challenges of data preparation for machine learning, focusing on operational, compliance, and cost.&lt;/strong&gt; &lt;a href="https://github.com/aronchick/Project-Zen-and-the-Art-of-Data-Maintenance?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;I'd love to hear your thoughts&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;!&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.distributedthoughts.org/2026-09-03-whitworths-thread/" rel="noopener noreferrer"&gt;Whitworth's Thread&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>operations</category>
      <category>process</category>
    </item>
    <item>
      <title>The Money Is Coming From Inside the House</title>
      <dc:creator>David Aronchick</dc:creator>
      <pubDate>Tue, 01 Sep 2026 18:01:36 +0000</pubDate>
      <link>https://dev.to/aronchick/the-money-is-coming-from-inside-the-house-4cpj</link>
      <guid>https://dev.to/aronchick/the-money-is-coming-from-inside-the-house-4cpj</guid>
      <description>&lt;p&gt;Nvidia reported &lt;a href="https://www.benzinga.com/markets/tech/26/08/61486472/nvidia-reportedly-pauses-revenue-sharing-deals-with-ai-cloud-companies-amid-antitrust-concerns" rel="noopener noreferrer"&gt;$96.22 billion in revenue&lt;/a&gt; this week for its second quarter, against analyst consensus of $92.11 billion, and guided to $108 billion for the third. A day or so later, the Wall Street Journal reported that the company had &lt;a href="https://finance.yahoo.com/news/nvidia-pauses-revenue-sharing-deals-223140237.html" rel="noopener noreferrer"&gt;paused the revenue-sharing financing deals&lt;/a&gt; it announced back in July under the name AI Compute Partnership. The program was aimed at smaller cloud providers who need to spend billions on Nvidia chips and the data centers around them before they can rent out an hour of anything. What WAS in place, but is no longer, was that Nvidia and the provider would agree on a base hourly rental rate that covers the provider's costs, Nvidia would provide credit support against the chip purchases, and &lt;a href="https://finance.yahoo.com/technology/ai/articles/nvidia-pauses-ai-cloud-revenue-120700044.html" rel="noopener noreferrer"&gt;Nvidia would take 50% of whatever revenue comes in above the base rate&lt;/a&gt;. The reasons for the pause, about eight weeks after launch, were internal worries about antitrust exposure, plus prospective partners getting irritated at how much control Nvidia wanted over the operations it was financing.&lt;/p&gt;

&lt;p&gt;So the chip vendor extends credit so the customer can buy the chips, then takes half the upside on the rental revenue those chips produce. I would call that a loan with profit participation before I called it a sale, and I think Nvidia's own lawyers apparently agreed.&lt;/p&gt;

&lt;p&gt;BUT, the paused program was also the small version of something much larger that is still running. OpenAI has committed to spending &lt;a href="https://en.wikipedia.org/wiki/AI_bubble" rel="noopener noreferrer"&gt;roughly $1.4 trillion over eight years&lt;/a&gt; on data centers and compute, against annual revenue somewhere near $13 billion. In contrast, simultaneously, NVIDIA has committed &lt;a href="https://www.bloomberg.com/graphics/2026-ai-circular-deals/" rel="noopener noreferrer"&gt;as much as $100 billion&lt;/a&gt; to OpenAI directly. OpenAI has committed hundreds of billions to Oracle and Microsoft for capacity, and Oracle and Microsoft buy NVIDIA GPUs to build that capacity out. Bloomberg and others have now traced &lt;a href="https://gfmag.com/technology/the-circle-game/" rel="noopener noreferrer"&gt;more than $800 billion&lt;/a&gt; of these interlocking arrangements through the AI supply chain. Morgan Stanley reportedly expects &lt;a href="https://www.calcalistech.com/ctechnews/article/z4lxiqbtw" rel="noopener noreferrer"&gt;Microsoft's entire Azure AI growth&lt;/a&gt; for the fiscal year that ended in June, north of $20 billion, to come from OpenAI, a counterparty in which Microsoft also holds a large stake. If you try to name the party in that chain who put outside money at risk on the proposition that end demand exists at these prices, somebody who is not also a supplier, customer, or shareholder of one of the others, I think you will end up struggling. While this is good in many ways (people are putting their capital where their words are), it also creates a lot of interdependence.&lt;/p&gt;

&lt;p&gt;As with so many things nowadays, this feels like a replay; in this case, Telecom ran this experiment 25 years ago, and the filings are still on EDGAR.&lt;/p&gt;

&lt;p&gt;Specifically, in the late 1990s the hot buyers of network equipment were the CLECs, the competitive local exchange carriers, startups laying fiber to compete with the Baby Bells under the 1996 Telecom Act. They had orders and no cash, so the equipment makers lent them the purchase price. Lucent committed &lt;a href="https://americanaffairsjournal.org/2020/08/who-lost-lucent-the-decline-of-americas-telecom-equipment-industry/" rel="noopener noreferrer"&gt;$8.1 billion in vendor financing&lt;/a&gt; over the period, including a &lt;a href="https://www.newsweek.com/stupid-loan-bubble-146353" rel="noopener noreferrer"&gt;$2 billion credit line to WinStar Communications&lt;/a&gt; that WinStar drew on to buy Lucent switches, and the switch sales went into Lucent's reported revenue, where analysts read them as demand. Between 1996 and 2001, the sector overbuilt by an estimated &lt;a href="https://www.thebubblebubble.com/telecom-bubble/" rel="noopener noreferrer"&gt;$60 billion&lt;/a&gt;, and once outside funding dried u,p, the upriers started failing: Covad, Focal, McLeod, NorthPoint, and then, in the spring of 20011, WinStar itself, days after Lucent declined to advance the final $90 million on the line. That single relationship cost Lucent a &lt;a href="https://rdata.kbsec.com/pdf_data/20251022133123807E.pdf" rel="noopener noreferrer"&gt;$700&lt;/a&gt; million write-off. The whole book deteriorated even faster than the WinStar piece did, with bad loans going from &lt;a href="https://www.technologyreview.com/2005/02/01/231676/how-lucent-lost-it/" rel="noopener noreferrer"&gt;2.6% of Lucent's financing portfolio at the end of 2000 to 60% a year later&lt;/a&gt; (Nortel's went from 25.5% to 80% over roughly the same stretch). Lucent took &lt;a href="https://mpra.ub.uni-muenchen.de/22012/1/Lazonick-March_Lucent_FINAL_20100410.pdf" rel="noopener noreferrer"&gt;billions in provisions against customer loans&lt;/a&gt;, shed most of its workforce, and was eventually absorbed by Alcatel, having discovered that a meaningful slice of its late-90s revenue had been its own treasury making a round trip.&lt;/p&gt;

&lt;p&gt;And before anyone emails/texts/DMs me to complain that "this time it's different", YES, vendor financing on its own is not a scandal. IBM Global Financing has been lending customers the price of IBM equipment since 1981, and it worked for four decades because the collateral was mainframes running payroll at companies with thirty years of verifiable cash flow. In THIS case, Lucent's collateral was the business plan of a five-year-old CLEC whose model assumed the 1999 growth curve was permanent. And, unfortunately, disclosure doesn't rescue it either, since Lucent's loans were disclosed too, in the same filings everyone now cites as the warning nobody read. The thing that went missing in the CLEC years, and is missing now, is an outside party who validated the demand with their own capital before the revenue got booked.&lt;/p&gt;

&lt;p&gt;Which brings me back to the pause. Nvidia had just posted the best quarter in its corporate history and could have coasted on the program for another year. Instead, eight weeks in, somebody internal looked at the antitrust exposure, or at the &lt;a href="https://aol.com/chip-fever-created-11-billion-153309463.html" rel="noopener noreferrer"&gt;$11 billion-plus already lent against GPUs&lt;/a&gt; across the neocloud sector before this program even existed, or at what taking half your customers' upside implies about eating half their downside, and shut the thing off. The company with the best view of real rental demand on earth got offered a machine for manufacturing more of it, and declined.&lt;/p&gt;

&lt;p&gt;None of which means AI demand is fake. A lot of it is real, VERY real, and expensively so. But some fraction of the headline numbers is the same dollar passing through three income statements; nobody knows the size of that fraction, including the participants, and we find out when a payment gets missed somewhere in the circle. The 1996 to 2001 fiber did get built, the builders mostly died, and the rest of us spent a decade lighting up dark strands bought for cents on the dollar. The capacity outlived the counterparties. WinStar felt like growth too, right up until the $90 million did not arrive.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Want to learn how intelligent data pipelines can reduce your AI costs?&lt;/em&gt; &lt;a href="https://expanso.io/?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;&lt;em&gt;Check out Expanso&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;. &lt;em&gt;Or don't. Who am I to tell you what to do?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;NOTE: I'm currently writing a book based on what I have seen about the real-world challenges of data preparation for machine learning, focusing on operational, compliance, and cost.&lt;/strong&gt; &lt;a href="https://github.com/aronchick/Project-Zen-and-the-Art-of-Data-Maintenance?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;I'd love to hear your thoughts&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;!&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.distributedthoughts.org/2026-08-31-the-money-is-coming-from-inside-the-house/" rel="noopener noreferrer"&gt;The Money Is Coming From Inside the House&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>finance</category>
      <category>history</category>
      <category>infrastructure</category>
    </item>
    <item>
      <title>The Drums Out Back</title>
      <dc:creator>David Aronchick</dc:creator>
      <pubDate>Fri, 28 Aug 2026 15:58:14 +0000</pubDate>
      <link>https://dev.to/aronchick/the-drums-out-back-28jj</link>
      <guid>https://dev.to/aronchick/the-drums-out-back-28jj</guid>
      <description>&lt;p&gt;IBM published its &lt;a href="https://www.ibm.com/reports/data-breach" rel="noopener noreferrer"&gt;2026 Cost of a Data Breach report&lt;/a&gt; at the end of July. The global average is now &lt;a href="https://www.helpnetsecurity.com/2026/07/30/ibm-cost-of-a-data-breach-2026/" rel="noopener noreferrer"&gt;$4.99 million per breach&lt;/a&gt;, up 12 percent year over year and a record for the study, which this round covered 602 organizations hit between March 2025 and February 2026. American breaches ran &lt;a href="https://www.infosecurity-magazine.com/news/cost-of-a-data-breach-5m-ibm/" rel="noopener noreferrer"&gt;more than double the global figure&lt;/a&gt;. Healthcare stayed the most expensive industry to be breached in, at &lt;a href="https://www.hipaajournal.com/2026-cost-data-breach-study-ibm/" rel="noopener noreferrer"&gt;$6.64 million&lt;/a&gt; a pop. And one in four malicious breaches is now &lt;a href="https://newsroom.ibm.com/2026-07-29-ibm-study-one-in-four-malicious-breaches-are-ai-enabled,-costing-companies-6-million-on-average" rel="noopener noreferrer"&gt;AI-enabled, averaging around $6 million&lt;/a&gt;, which is a 56 percent jump in a single year.&lt;/p&gt;

&lt;p&gt;The report prices an event for the first time in a while that I actually like. But almost nobody is reading the other half of the sentence; it's half the event, and half the size of the pile. We need to look at both.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Cleanest Illustration Available Is French
&lt;/h2&gt;

&lt;p&gt;On January 13, the CNIL &lt;a href="https://www.cnil.fr/en/sanction-free-2026" rel="noopener noreferrer"&gt;fined Free Mobile €27 million and Free €15 million&lt;/a&gt;, €42 million between them, over an intrusion in October 2024 that reached personal data on roughly &lt;a href="https://www.theregister.com/2026/01/14/france_fines_free_free_mobile/" rel="noopener noreferrer"&gt;24 million subscriber contracts&lt;/a&gt;, including bank account numbers for customers who held both services. Three findings: the data was not adequately secured, the notification to affected people was inadequate, and the retention was illegal.&lt;/p&gt;

&lt;p&gt;The third item is pretty uncommon to say the least. The regulator determined that the carrier was holding millions of records on people who were no longer customers, well past any period it could justify, when what it actually needed to keep for accounting purposes was a much smaller slice. So when the attacker came along, they did not steal what Free Mobile needed to run its business; they stole information that Free Mobile had no lawful reason to still possess, sitting in the same system as the material it did.&lt;/p&gt;

&lt;p&gt;The intrusion was the trigger, but the retention was the multiplier. One of those two things was under the company's control for years in advance, at essentially zero cost. Less than zero cost, they could have deleted it and saved money!&lt;/p&gt;

&lt;p&gt;Article 5 of the GDPR has said since 2018 that personal data must be &lt;a href="https://gdpr-info.eu/art-5-gdpr/" rel="noopener noreferrer"&gt;adequate, relevant, and limited to what is necessary&lt;/a&gt;, and that it must not be kept in identifiable form longer than the purpose requires. Violating the Article 5 principles sits in the top tier of Article 83, which reaches &lt;a href="https://gdpr-info.eu/art-83-gdpr/" rel="noopener noreferrer"&gt;€20 million or four percent of worldwide annual turnover&lt;/a&gt;, whichever hurts more. Cumulative GDPR fines across the EEA passed &lt;a href="https://www.kiteworks.com/gdpr-compliance/gdpr-fines-data-privacy-enforcement-2026/" rel="noopener noreferrer"&gt;€7.1 billion&lt;/a&gt; by the start of this year. Minimization has been law for eight years and mostly treated as a policy document that lives in a wiki nobody reads.&lt;/p&gt;

&lt;h2&gt;
  
  
  Somebody Already Ran This Experiment With Actual Barrels
&lt;/h2&gt;

&lt;p&gt;In 1920 a company called Hooker Chemical bought a partially dug canal in Niagara Falls, and for the next 33 years it used the thing as a disposal site, putting &lt;a href="https://education.nationalgeographic.org/resource/superfund/" rel="noopener noreferrer"&gt;roughly 22,000 tons of chemical waste&lt;/a&gt; into the ground. The logic was completely sound, storage was cheap, the land was theirs, the practice was legal, and a fair amount of what went in the hole was technically feedstock, material with a real industrial value if anyone ever wanted to go get it.&lt;/p&gt;

&lt;p&gt;In 1953 Hooker sold the site to the local school board for one dollar. The deed carried what became known as the Hooker clause, a written disclaimer saying the company would not be responsible if anybody got sick or died from the waste buried there. The local board proceeded to build school and a neighborhood went up on top.&lt;/p&gt;

&lt;p&gt;Then the 1970s happened, &lt;a href="https://www.epa.gov/archive/epa/aboutepa/epa-new-york-state-announce-temporary-relocation-love-canal-residents.html" rel="noopener noreferrer"&gt;families were relocated&lt;/a&gt;, Congress held hearings that are &lt;a href="https://levin-center.org/what-is-oversight/portraits/love-canal/" rel="noopener noreferrer"&gt;still studied as an oversight case&lt;/a&gt;, and on December 11, 1980, Carter signed CERCLA. And the statute did two things that have relevance today.&lt;/p&gt;

&lt;p&gt;It made liability &lt;em&gt;strict&lt;/em&gt;, meaning you do not get to argue that you followed the industry standard or that you were not negligent. And it made liability &lt;em&gt;retroactive&lt;/em&gt;, meaning it reached backward to conduct that was perfectly lawful when it happened. The Hooker clause, a real contract, negotiated and signed by consenting parties, &lt;a href="https://cumulis.epa.gov/supercpad/SiteProfiles/index.cfm?fuseaction=second.cleanup&amp;amp;id=0201290" rel="noopener noreferrer"&gt;turned out to be worth nothing&lt;/a&gt;. Occidental Petroleum, which bought Hooker later, inherited the whole thing.&lt;/p&gt;

&lt;p&gt;It's true that retroactive strict liability is a sledgehammer, the transaction costs have been &lt;a href="https://perc.org/1996/05/01/superfund-the-shortcut-that-failed/" rel="noopener noreferrer"&gt;genuinely awful&lt;/a&gt;, and a meaningful fraction of the money went to lawyers arguing about allocation rather than to anybody's groundwater. Nobody should hold CERCLA up as model legislation, but it was pretty groundbreaking (pardon the pun) for the time. What it gives us, for our purposes, is a demonstration of what a legislature actually does once a stored-material problem gets bad enough and public enough. It does not grandfather you and it does not care what your contract says.&lt;/p&gt;

&lt;p&gt;So when somebody tells me their data retention exposure is handled because the DPA covers it, or because the terms of service disclaim it, or because everything was collected lawfully under the rules that applied at the time, I think about a signed piece of paper from 1953 that did all three of those things.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Minimization Actually Costs You, And When
&lt;/h2&gt;

&lt;p&gt;The reason "keep everything, decide later" won is that it was the rational choice for about fifteen years. Storage got cheap faster than anyone could develop taste, the analytics you would want were genuinely unknowable in advance, and every data scientist who ever got told "we dropped that column in 2019" learned to hoard. I have been the person arguing for keeping the raw feed (and am currently paying penance). Sometimes that argument is right.&lt;/p&gt;

&lt;p&gt;But that's no longer the general case. Heck, in many ways it's the exception.&lt;/p&gt;

&lt;p&gt;A record is created at some edge of your system. A meter, a handset, a claims form, a badge reader, a log line. At that instant, exactly one copy exists, in one jurisdiction, under one owner, and dropping a field costs you one line of configuration. That is the cheapest that decision will ever be, by orders of magnitude, and it is the only moment where "delete" means the thing deleted is gone.&lt;/p&gt;

&lt;p&gt;Now let it move. It lands in object storage, gets picked up into a warehouse, gets denormalized into three marts because three teams wanted different grain, gets a feature-store copy for the model, gets a nightly backup with a 90-day cycle, gets replicated to a second region for durability, gets pulled into a vendor's SaaS for enrichment, and gets exported once into a notebook by an analyst who left in March. Call that eight to ten locations, and I am being conservative, because I have not counted the CI fixture somebody generated from prod or the Slack thread with the screenshot.&lt;/p&gt;

&lt;p&gt;Every one of those is now a separate deletion obligation, a separate discovery obligation, a separate breach surface, and a separate thing you have to be able to &lt;em&gt;prove&lt;/em&gt; about. Not just assert. Prove, to a regulator, with evidence, on a clock. Deleting a field from ten places under audit is not ten times harder than dropping it once at the source. That kind of work gets a program manager, a status color, and a standing Thursday call, and at the end of it the answer is usually "we believe so."&lt;/p&gt;

&lt;p&gt;One decision at creation, or ten proofs later, you only get to choose one. And if you don't choose, you default to the ten proofs version. Good luck.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Part That Should Actually Bother You
&lt;/h2&gt;

&lt;p&gt;We built an entire generation of data infrastructure whose default posture is accumulate now, govern later. And nobody has ever once meant anything by "govern later." It's a sentence you say to end a meeting, and everybody nodding along knows it's bullshit while they're nodding. LATER NEVER COMES. It does not come because there is no forcing function, no deadline, and no individual whose bonus depends on the absence of a thing.&lt;/p&gt;

&lt;p&gt;What does eventually arrive is an incident, or a subpoena, or a supervisory authority with a questionnaire, and by then the population of records you are answering for was fixed years ago by a default nobody chose deliberately. It's not that there was a bad call; at least that could be forgivable. That there was no call. A retention policy emerged as a side effect of a storage tier being cheap.&lt;/p&gt;

&lt;p&gt;For the regulated and critical-infrastructure crowd, and I spend a lot of my time in rooms with exactly those people, the question in the room has already shifted. It used to be some version of what could we learn from this if we had it all. Increasingly it is can we demonstrate we never held it. Those two questions want opposite architectures, and only one of them is compatible with a pipeline that ships everything to a central lake first and sorts it out downstream. If the filtering, the redaction, the classification, and the retention decision do not happen in the environment where the record was born, they are happening after the liability has already been created and copied.&lt;/p&gt;

&lt;p&gt;I have written before that &lt;a href="https://www.distributedthoughts.org/2026-03-27-the-time-value-of-data/" rel="noopener noreferrer"&gt;data loses economic value as it sits&lt;/a&gt;, and that your &lt;a href="https://www.distributedthoughts.org/2026-04-30-catalog-will-be-wrong-eventually/" rel="noopener noreferrer"&gt;catalog is wrong and will keep being wrong&lt;/a&gt;. At the end of the day, the value of a record decays, the accuracy of your inventory of it decays, and the liability attached to it does not decay at all. It sits flat, or it goes up when somebody passes a statute. You are holding an asset that amortizes against a liability that does not.&lt;/p&gt;

&lt;p&gt;It's ALSO true that SOME of this material genuinely is an asset, and SOME of it is regulatorily &lt;em&gt;required&lt;/em&gt; to be kept, and the retention schedule for a medical record is not a thing an engineer gets to have an opinion about. Fine. Good. That is the point. The categories are different, they have different clocks, and a system that cannot tell them apart FROM THE MOMENT OF CREATION has already decided to treat all of it as the most dangerous thing in the pile. Classification at the point of creation is what lets you keep the valuable part at all, which is why I get twitchy when people file it under compliance and hand it to the team with no engineers.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Clause Was Signed
&lt;/h2&gt;

&lt;p&gt;Time and jurisdiction draw the line between an asset and a liability. And you do not control either one of them. Hooker Chemical had a legal practice, a real commercial rationale, and a signed indemnity. What it did not have was a view into the future, and in 1980 folks decided they were wrong and should have always done better.&lt;/p&gt;

&lt;p&gt;You do not get to know today which of your columns becomes a drum out back. What you get to decide is how many copies of it exist by the time somebody asks.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Want to filter, redact, and classify data where it is created, so the copy you keep is the one you meant to keep?&lt;/em&gt; &lt;a href="https://expanso.io/?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;&lt;em&gt;Check out Expanso&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;. &lt;em&gt;Or don't. Who am I to tell you what to do.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;NOTE: I'm currently writing a book based on what I have seen about the real-world challenges of data preparation for machine learning, focusing on operational, compliance, and cost.&lt;/strong&gt; &lt;a href="https://github.com/aronchick/Project-Zen-and-the-Art-of-Data-Maintenance?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;I'd love to hear your thoughts&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;!&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.distributedthoughts.org/2026-08-27-the-drums-out-back/" rel="noopener noreferrer"&gt;The Drums Out Back&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>governance</category>
      <category>compliance</category>
      <category>datamanagement</category>
      <category>regulation</category>
    </item>
    <item>
      <title>Forty-Nine Megawatts</title>
      <dc:creator>David Aronchick</dc:creator>
      <pubDate>Tue, 25 Aug 2026 18:05:24 +0000</pubDate>
      <link>https://dev.to/aronchick/forty-nine-megawatts-4e1o</link>
      <guid>https://dev.to/aronchick/forty-nine-megawatts-4e1o</guid>
      <description>&lt;p&gt;On July 14, Governor Kathy Hochul signed an executive order making New York the &lt;a href="https://www.governor.ny.gov/news/first-statewide-moratorium-new-hyperscale-data-centers-launched-governor-kathy-hochul" rel="noopener noreferrer"&gt;first state in the country&lt;/a&gt; to pause the construction of new hyperscale data centers. The moratorium covers new facilities with an electrical demand of &lt;a href="https://www.cnbc.com/2026/07/14/new-york-ai-data-center-ban.html" rel="noopener noreferrer"&gt;50 megawatts or more&lt;/a&gt;, freezes state environmental permits for &lt;a href="https://www.nbcnews.com/news/us-news/new-york-impose-countrys-first-statewide-moratorium-data-centers-rcna587429" rel="noopener noreferrer"&gt;up to a year&lt;/a&gt;, and includes a promise to build a nation-leading framework for grid reliability, electricity costs, water, noise, and land use before anything else is approved. &lt;a href="https://insideclimatenews.org/news/14072026/new-york-first-data-center-moratorium/" rel="noopener noreferrer"&gt;Thirty-nine more&lt;/a&gt; are somewhere in the application pipeline, which tells you how fast this was moving before somebody hit the brakes.&lt;/p&gt;

&lt;p&gt;I think this is a TERRIBLE idea, but there's SOME logic behind it, so let's dive in.&lt;/p&gt;

&lt;p&gt;Back in May, I wrote about &lt;a href="https://www.distributedthoughts.org/2026-05-04-permission-problem/" rel="noopener noreferrer"&gt;Loudoun County&lt;/a&gt;, the most permissive data center jurisdiction in America, flipping to moratorium proposals in about eighteen months. When the friendliest county in the friendliest state turns hostile, it is no surprise that these things appear more often in other districts and states. The grid genuinely &lt;a href="https://www.distributedthoughts.org/2026-04-09-the-grid-said-no/" rel="noopener noreferrer"&gt;cannot deliver power&lt;/a&gt; where people want to put the load, and ratepayers are really &lt;a href="https://www.distributedthoughts.org/2026-07-06-the-cheapest-connection-you-never-build/" rel="noopener noreferrer"&gt;eating costs they never agreed to&lt;/a&gt;. (I SHOULD NOTE: the reason the rates are going up is NOT that the data centers are raising the cost! It's because these power companies see the opportunity, take it, and are heavily regulated. A one-year pause to write actual rules, in a state where thirty-nine applications were stacked up like planes over LaGuardia, is a defensible thing for a governor to do. The same week, for what it's worth, Microsoft was &lt;a href="https://www.datacenterknowledge.com/data-center-construction/new-data-center-developments-july-2026" rel="noopener noreferrer"&gt;cutting the ribbon&lt;/a&gt; on $3.3 billion in Wisconsin, so nobody should mistake this for a national trend yet.&lt;/p&gt;

&lt;p&gt;But let's talk about the guidance - it's JUST at 50 megawatts. How'd they choose that?&lt;/p&gt;

&lt;h2&gt;
  
  
  Every threshold becomes a spec sheet
&lt;/h2&gt;

&lt;p&gt;France has a rule that at 50 employees, a firm crosses into a different regulatory universe: works councils, union delegates, profit-sharing obligations, formal restructuring plans. The result is one of the most famous pictures in empirical economics. Plot the distribution of French firms by headcount, and there is a &lt;a href="https://www.aeaweb.org/articles?id=10.1257/aer.20121532" rel="noopener noreferrer"&gt;pileup at exactly 49 and a crater just past it&lt;/a&gt;. Firms do not grow in a straight line. They stop at it, split in two, spin off subsidiaries, push work to contractors, whatever it takes to stay a hair under. The economists who studied it, &lt;a href="https://www.nber.org/papers/w18841" rel="noopener noreferrer"&gt;Garicano, Lelarge, and Van Reenen&lt;/a&gt;, found the distortion was big enough to measure in fractions of national output. Not because anyone cheated, but because the line was published, and published lines get engineered to.&lt;/p&gt;

&lt;p&gt;This is not a French quirk. Federal law caps trucks at &lt;a href="https://en.wikipedia.org/wiki/Federal_Bridge_Formula" rel="noopener noreferrer"&gt;80,000 pounds gross&lt;/a&gt; weight, so an enormous amount of American freight rolls at 79,900 pounds. Coastal shipping fleets around the world are full of vessels built to slide just under whatever tonnage triggers the next tier of crewing and inspection rules. Buildings stop one floor short of the elevator code. Wherever a regulation names a number, an industry grows a callus at that number. The market never argues with a threshold. It reads as a product requirements document.&lt;/p&gt;

&lt;p&gt;So here is my prediction for New York: a wave of 49-megawatt data centers, with facilities designed to the line the way French firms are staffed to it. Campuses that are legally three separate 45-megawatt projects with three separate applications and, remarkably, three separate fences. Developers did not stop wanting to build in New York on July 14. They started reading the order for its edges that afternoon, billing by the hour.&lt;/p&gt;

&lt;h2&gt;
  
  
  The accidental architecture
&lt;/h2&gt;

&lt;p&gt;But here's the funny bit... a 49-megawatt data center is a good idea!&lt;/p&gt;

&lt;p&gt;A fleet of smaller facilities, spread across a state instead of piled onto one substation, sited where the grid has actual headroom or &lt;a href="https://www.utilitydive.com/news/ferc-pjm-colocation-data-center/808368/" rel="noopener noreferrer"&gt;next to their own generation&lt;/a&gt;, closer to the people and data they serve: that is what the physics has been begging for the entire time. The gigawatt campus exists because it is convenient for the operator, not because the workload demands it. Training wants tight coupling; the &lt;a href="https://www.distributedthoughts.org/2026-04-13-agents-dont-live-in-data-centers/" rel="noopener noreferrer"&gt;inference and agent workloads&lt;/a&gt; that are actually growing mostly do not care, and demand is on track to &lt;a href="https://www.npr.org/2026/01/02/nx-s1-5638587/ai-data-centers-use-a-lot-of-electricity-how-it-could-affect-your-power-bill" rel="noopener noreferrer"&gt;nearly triple by 2035&lt;/a&gt;, whether we build it in three monoliths or three hundred sheds. New York, in trying to stop data centers, may have accidentally written the first state-level specification for distributed compute. The moratorium is the mandate.&lt;/p&gt;

&lt;p&gt;THAT SAID, twenty 49-megawatt buildings running software that assumes one big building is not distributed computing. It's a monolith with extra steps and worse latency. The hard part was never pouring smaller slabs; it was treating computers spread across a state as one coherent system, moving the work to where the capacity and the data already are, instead of hauling everything to a headquarters that no longer exists. The real estate industry will complete the 49-megawatt building in about six months. The software model that makes a hundred of them worth more than the sum of their parts is the actual project, and just dealing with land/power/shell is not going to get you there.&lt;/p&gt;

&lt;p&gt;New York thinks it drew a boundary. What was published was a design constraint, and the one thing this industry reliably does with a design constraint is build to it, at volume, faster than the regulator can schedule the follow-up meeting. The question worth watching is not whether the moratorium survives the year. It's what gets built at forty-nine.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Want to run compute where the power and data already live, instead of where the substation used to be?&lt;/em&gt; &lt;a href="https://expanso.io/?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;&lt;em&gt;Check out Expanso&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;. &lt;em&gt;Or don't. Who am I to tell you what to do?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;NOTE: I'm currently writing a book based on what I have seen about the real-world challenges of data preparation for machine learning, focusing on operational, compliance, and cost.&lt;/strong&gt; &lt;a href="https://github.com/aronchick/Project-Zen-and-the-Art-of-Data-Maintenance?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;I'd love to hear your thoughts&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;!&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.distributedthoughts.org/2026-08-24-forty-nine-megawatts/" rel="noopener noreferrer"&gt;Forty-Nine Megawatts&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiinfrastructure</category>
      <category>energy</category>
      <category>regulation</category>
      <category>history</category>
    </item>
    <item>
      <title>The Exclusion Is the Spec</title>
      <dc:creator>David Aronchick</dc:creator>
      <pubDate>Fri, 21 Aug 2026 18:05:25 +0000</pubDate>
      <link>https://dev.to/aronchick/the-exclusion-is-the-spec-38nm</link>
      <guid>https://dev.to/aronchick/the-exclusion-is-the-spec-38nm</guid>
      <description>&lt;p&gt;The Insurance Services Office (ISO) has three endorsements in circulation with a 01 26 edition date: &lt;a href="https://www.claimsjournal.com/news/national/2026/07/20/338950.htm" rel="noopener noreferrer"&gt;CG 40 47, CG 40 48, and CG 35 08&lt;/a&gt;. Let's say these names were not presented in front of a marketing committee, but what do I know? What they do, between them, is carve bodily injury, property damage, and personal and advertising injury arising out of generative AI out of the Commercial General Liability form and the Products/Completed Operations form both. In the past six months, carriers have been lining up at state insurance regulators for permission to actually use them, and the attorney tracking those filings for Lathrop GPM called it an industry-wide reaction to the explosion of AI. &lt;a href="https://natlawreview.com/article/continued-proliferation-ai-exclusions" rel="noopener noreferrer"&gt;Berkley went even further last year&lt;/a&gt; with an absolute AI exclusion for D&amp;amp;O, E&amp;amp;O, and fiduciary lines, one that reaches past your AI's output to your AI policies, your AI procedures, and your failure to notice somebody ELSE's AI. Talk about the long arm of the law.&lt;/p&gt;

&lt;p&gt;For context, ISO writes the forms most of the American commercial market runs on, so &lt;a href="https://www.ajg.com/news-and-insights/iso-introduces-generative-ai-exclusion-in-commercial-general-liability-policies/" rel="noopener noreferrer"&gt;an ISO exclusion&lt;/a&gt; is about as close as insurance gets to a standards body publishing a deprecation notice.&lt;/p&gt;

&lt;p&gt;And keep in mind what the business of insurance actually IS. You take a risk of some kind. They calculate the risk and tell you how much it will cost if it goes wrong. You pay them, and now you are protected. That's the product; there isn't another part. And, for whatever reason, a meaningful chunk of that industry - whose ENTIRE job is to evaluate the risk of just about anything - has now looked at generative AI and said, some version of, no thanks.&lt;/p&gt;

&lt;h2&gt;
  
  
  They are not being cowards about it
&lt;/h2&gt;

&lt;p&gt;Joe Lam, the Verisk VP who helped write the endorsements, gave the least dramatic (and, I'd argue, most honest) account of it in that Claims Journal piece: "Without exclusions to allow underwriters a level of stability to accept a risk, you run into a situation where they might just walk away from the risk. So exclusions are very essential in the marketplace."&lt;/p&gt;

&lt;p&gt;The exclusion, as it stands, fences off the one piece the underwriters can't measure, so they can keep writing everything around it, because the alternative was walking away from the whole line. A narrow exclusion is more coverage than no market at all, and anybody who has watched a line of business go uninsurable (e.g., Enron) knows exactly which of those two is worse.&lt;/p&gt;

&lt;p&gt;They're not doing this in a vacuum. Gallagher counted a &lt;a href="https://www.ajg.com/gallagherre/-/media/files/gallagher/gallagherre/news-and-insights/2026/march/rethinking-insurance-for-the-ai-era.pdf" rel="noopener noreferrer"&gt;978% increase in AI-related litigation between 2021 and 2025&lt;/a&gt;, with a 137% jump in the final year of that window alone. And on July 24, the Delaware Superior Court &lt;a href="https://www.claimsjournal.com/news/national/2026/07/27/339086.htm" rel="noopener noreferrer"&gt;ordered Google to defend a defamation suit over what its AI said about a person&lt;/a&gt;. If something has a docket number, people are going to stand up and take notice.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem isn't that AI is dangerous
&lt;/h2&gt;

&lt;p&gt;Most of the commentary goes straight to the black box: AI is unpredictable, its outputs aren't deterministic, and underwriters can't model what they can't explain.&lt;/p&gt;

&lt;p&gt;That said, insurance has never needed predictability at the individual level. Nobody knows which house is going to burn down, because no one is measuring just one house. The actuarial math asks for exactly one property: that the losses be &lt;em&gt;independent&lt;/em&gt;—ten thousand houses, uncorrelated fires, the law of large numbers, everybody goes home happy. Correlation is what kills an insurance market (Go look at the mortgage insurance market in 2008 if you want to see how). Correlation is why nobody will sell you a single policy covering every house on one street against the same fire, and why flood ended up as a federal program.&lt;/p&gt;

&lt;p&gt;Now let's look at AI. A handful of foundation models, three clouds, overlapping training corpora, the same inference frameworks and orchestration layers, and vector stores wired together off the same blog posts. I've spent most of my career (Google, Microsoft, Amazon, now &lt;a href="//expanso.io"&gt;Expanso&lt;/a&gt;) building distributed systems where a big part of the job is keeping failures from correlating, so watching this particular stack assemble itself has been, let's say, uncomfortable. Gallagher Re has been flagging it for a year, and Aon's Kevin Kalinich distilled the underwriter's view into three words: &lt;a href="https://www.insurancebusinessmag.com/us/news/technology/insurers-face-hidden-ai-liability-as-agent-risks-multiply-582433.aspx" rel="noopener noreferrer"&gt;"aggregated, systemic, correlated."&lt;/a&gt; One vulnerability in a common dependency, and the losses land on an entire book of insureds the same afternoon.&lt;/p&gt;

&lt;p&gt;I wrote in June about &lt;a href="https://www.distributedthoughts.org/2026-06-18-six-hundred-ways-not-to-connect-a-hose/" rel="noopener noreferrer"&gt;what happens when everything speaks one format and routes through one provider&lt;/a&gt;, and the underwriters have now put a price on the answer. Or rather, declined to put a price on it. Fragile things get insured all day long, every day, everywhere. The issue with a monoculture is that when it goes, it all goes at once, and the pool that was supposed to absorb your loss turns out to be built from the same stuff that just failed.&lt;/p&gt;

&lt;p&gt;Which means CG 40 47 is a verdict on the topology.&lt;/p&gt;

&lt;h2&gt;
  
  
  Silent cover is how this always starts
&lt;/h2&gt;

&lt;p&gt;"Silent AI" is exposure sitting within conventional policies that neither confirm nor deny it, left to be argued at claim time by lawyers after the loss. According to one industry estimate, more than 90% of insurers' AI-agent exposure is silent, tucked inside cyber, professional indemnity, general liability, and D&amp;amp;O policies written by people who &lt;a href="https://www.regulationtomorrow.com/2026/05/silent-ai-risks-finally-make-some-noise/" rel="noopener noreferrer"&gt;were not thinking about agents at all&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;We have run this movie before, and it did not turn out well.&lt;/p&gt;

&lt;p&gt;General liability policies written from the 1940s through the 1970s said nothing about asbestos, because why would they? The exposure was silent, unpriced, and enormous, and it surfaced decades later as long-tail claims against contracts nobody remembered signing. &lt;a href="https://www.actuaries.asn.au/research-analysis/back-from-the-brink-the-near-collapse-of-lloyd-s-of-london" rel="noopener noreferrer"&gt;Lloyd's underwriters lost roughly £9 billion between 1988 and 1992&lt;/a&gt;. And here is the detail people tend to forget about Lloyd's. The capital behind the market came from about &lt;a href="https://en.wikipedia.org/wiki/Lloyd%27s_of_London" rel="noopener noreferrer"&gt;34,000 Names, individuals carrying unlimited personal liability&lt;/a&gt;, and when the bill arrived, many of them lost everything they had; &lt;a href="https://time.com/archive/6740471/lloyds-of-london-falling-down/" rel="noopener noreferrer"&gt;at least fifteen killed themselves&lt;/a&gt;. Lloyd's survived only by walling the old years off inside &lt;a href="https://en.wikipedia.org/wiki/Equitas" rel="noopener noreferrer"&gt;a separate reinsurance vehicle called Equitas&lt;/a&gt;, and lawyers were &lt;a href="https://www.kirkland.com/publications/article/2002/03/some-sobering-facts-about-equitas-and-the-potentia" rel="noopener noreferrer"&gt;still picking at that structure's solvency a decade later&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Asbestos was in everything; the policies said nothing; and the bill came due twenty years after the premium had been paid. So when somebody tells me that 90% of the industry's AI exposure is currently silent, that sends shivers down insurers' and reinsurers' spines.&lt;/p&gt;

&lt;h2&gt;
  
  
  Underwriters end up writing the spec
&lt;/h2&gt;

&lt;p&gt;So what do we do?&lt;/p&gt;

&lt;p&gt;When insurers can't price something, walking away is only their first move. The second move, reliably, after more than a century of doing this, is to fund someone to go measure the thing. And whoever does the measuring ends up dictating how the thing gets built.&lt;/p&gt;

&lt;p&gt;In 1893, the Chicago fire insurance authorities watched the Palace of Electricity at the World's Columbian Exposition light up with a hundred thousand Edison bulbs and kept noticing an inconvenient pattern: the building kept catching fire. Was it the wiring? The hookups? This new alternating current? Nobody knew, and the insurers were not inclined to keep writing the coverage while everybody wondered. So &lt;a href="https://www.referenceforbusiness.com/history2/69/Underwriters-Laboratories-Inc.html" rel="noopener noreferrer"&gt;they hired an electrical inspector named William Henry Merrill&lt;/a&gt; and funded him, through the Chicago Board of Fire Underwriters and the Western Insurance Association, to investigate. His lab was a room above Fire Insurance Patrol Station Number One. A bench, a table, some chairs, $350 of measuring equipment, and that's it, that was the whole operation.&lt;/p&gt;

&lt;p&gt;His first test, filed March 24, 1894, was a sheet of asbestos paper a manufacturer had claimed was noncombustible and nonabsorbent. Merrill found it absorbed water and would not burn, making it useless as insulation, decent for fire resistance, and, either way, a measured fact now instead of a sales claim. (And yes, the first thing Underwriters Laboratories ever tested was asbestos, the same material from the section you just read. Let it never be said that history does not have a sense of irony.) After several thousand tests, the lab published its first list of approved fittings and devices in 1898, and approved products got a label. &lt;a href="https://en.wikipedia.org/wiki/UL_(safety_organization)" rel="noopener noreferrer"&gt;In 1901, it was chartered in Illinois&lt;/a&gt; as Underwriters Laboratories, taking the name of its new sponsor, the National Board of Fire Underwriters, with a stated purpose of testing appliances and recommending them to insurance organizations. Its first Standard, in 1903, covered tin-clad fire doors. &lt;a href="https://ul.org/about/our-history/" rel="noopener noreferrer"&gt;After the 1906 San Francisco earthquake&lt;/a&gt;, UL was helping the National Board write building codes, and its engineers went on to shape the early National Electrical Code.&lt;/p&gt;

&lt;p&gt;It's insane, but true, that a meaningful share of the electrical safety rules governing every building you have ever walked into exists because a group of fire insurers refused to keep writing policies until somebody could tell them what was in the wall. The refusal came first; the standard is the reason coverage ever came back.&lt;/p&gt;

&lt;p&gt;So I honestly don't care whether CG 40 47 is fair. What I want to know is what the AI equivalent of a tin-clad fire door looks like, because somebody, somewhere, has to write that before this exposure becomes insurable again.&lt;/p&gt;

&lt;p&gt;I can tell you what it won't be: any of the "benchmarks" we have today (which are starting to feel a bit like &lt;a href="https://en.wikipedia.org/wiki/Goodhart%27s_law" rel="noopener noreferrer"&gt;Goodhart's law&lt;/a&gt;). An underwriter does not give a damn that your model came in three points higher on some eval, because an average tells you nothing about the day it goes wrong. What an underwriter needs is provable data provenance, a record of which decisions the system actually made (not just recommended), tenant isolation, and some ceiling on how far a bad model update travels before anyone notices. I argued in April that &lt;a href="https://www.distributedthoughts.org/2026-04-27-you-cant-sue-an-agent/" rel="noopener noreferrer"&gt;you can't sue an agent&lt;/a&gt;, and these exclusions are what that argument looks like as an invoice. If nobody can locate the liability, nobody can price it, so it goes out of the form.&lt;/p&gt;

&lt;p&gt;Every item on that list is a property of how the system is built, not of the model sitting inside it — which is a mildly humiliating thing for our industry to be learning from an insurance endorsement, but here we are.&lt;/p&gt;

&lt;p&gt;The carriers that filed exclusions in July did not end anything. They're Merrill in 1894, standing in front of the exposition wiring, declining to sign until someone tells them what's behind the panel. Nobody could tell him. He had to build the lab to find out.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Want to learn how intelligent data pipelines can reduce your AI costs?&lt;/em&gt; &lt;a href="https://expanso.io/?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;&lt;em&gt;Check out Expanso&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;. &lt;em&gt;Or don't. Who am I to tell you what to do?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;NOTE: I'm currently writing a book based on my observations of real-world challenges in data preparation for machine learning, focusing on operational, compliance, and cost issues.&lt;/strong&gt; &lt;a href="https://github.com/aronchick/Project-Zen-and-the-Art-of-Data-Maintenance?ref=distributedthoughts.org" rel="noopener noreferrer"&gt;&lt;strong&gt;I'd love to hear your thoughts&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;!&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.distributedthoughts.org/2026-08-20-the-exclusion-is-the-spec/" rel="noopener noreferrer"&gt;The Exclusion Is the Spec&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>risk</category>
      <category>insurance</category>
      <category>history</category>
    </item>
  </channel>
</rss>
