<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sumitsuke</title>
    <description>The latest articles on DEV Community by Sumitsuke (@tauridev).</description>
    <link>https://dev.to/tauridev</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4019232%2F9ef8616e-3149-4666-96a2-5340e2d2e3a3.png</url>
      <title>DEV Community: Sumitsuke</title>
      <link>https://dev.to/tauridev</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/tauridev"/>
    <language>en</language>
    <item>
      <title>I cut transfer by 81–257 KiB. Did LCP actually improve? 30 Lighthouse runs</title>
      <dc:creator>Sumitsuke</dc:creator>
      <pubDate>Mon, 21 Sep 2026 08:32:18 +0000</pubDate>
      <link>https://dev.to/tauridev/i-cut-transfer-by-81-257-kib-did-lcp-actually-improve-30-lighthouse-runs-59b6</link>
      <guid>https://dev.to/tauridev/i-cut-transfer-by-81-257-kib-did-lcp-actually-improve-30-lighthouse-runs-59b6</guid>
      <description>&lt;p&gt;"We made it lighter" and "we made it faster" are different claims, and a single Lighthouse run can't tell them apart — the last time I checked, repeated runs under the same conditions had a spread of more than a second. This is the acceptance sheet I ended up with, and the 30 runs behind it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyeofp9zacjnpckq61yzm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyeofp9zacjnpckq61yzm.png" alt="Transfer size and LCP signed off in separate columns: home −81 KiB / −600 ms, Lab article −257 KiB / −1,347 ms, pricing −81 KiB / −454 ms; CLS unchanged on all three" width="800" height="1816"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Setup
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;old&lt;/strong&gt; = commit &lt;code&gt;88b702d&lt;/code&gt;: stale font subset version, two bold characters missing from the core subset, PNG images, no &lt;code&gt;width&lt;/code&gt;/&lt;code&gt;height&lt;/code&gt;, regular weight not split.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;new&lt;/strong&gt; = commit &lt;code&gt;1ba941a&lt;/code&gt;: versioned font URLs (&lt;code&gt;?v=sha1&lt;/code&gt;), a build gate for bold characters, WebP, image dimensions, the regular weight split into core/rest, plus three rounds of visual fixes.&lt;/li&gt;
&lt;li&gt;Both built with Astro from &lt;code&gt;git worktree&lt;/code&gt;s on the same Windows PC and served with &lt;code&gt;astro preview&lt;/code&gt; on :4321 / :4322.&lt;/li&gt;
&lt;li&gt;Lighthouse 13.5.0, &lt;code&gt;--preset=perf --form-factor=mobile --throttling-method=simulate&lt;/code&gt;, headless Chrome 153.&lt;/li&gt;
&lt;li&gt;Three pages (home, one Lab article with two figures, the pricing page) × 5 runs × old/new = &lt;strong&gt;30 runs&lt;/strong&gt;, alternating old→new→new→old→… so time drift doesn't land on one side.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hypotheses fixed before running&lt;/strong&gt;: H1 home LCP goes down. H2 the article's CLS goes down (it got image dimensions). H3 pricing LCP goes down. &lt;strong&gt;Stop rule&lt;/strong&gt;: if a difference is smaller than the 5-run interquartile range, the verdict is HOLD.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;SERVERS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;old&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;http://localhost:4321&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;new&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;http://localhost:4322&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;ORDER&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;old&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;new&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;new&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;old&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;old&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;new&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;new&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;old&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;old&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;new&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
&lt;span class="c1"&gt;// Windows: chrome-launcher fails to delete its temp dir (EPERM) and exits 1 although the JSON is written —&lt;/span&gt;
&lt;span class="c1"&gt;// judge by the file's existence, not the exit code&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Numbers
&lt;/h2&gt;

&lt;p&gt;Transfer is bytes ÷ 1,024 (KiB), summed from Lighthouse's &lt;code&gt;network-requests&lt;/code&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Page&lt;/th&gt;
&lt;th&gt;Transfer old→new (KiB)&lt;/th&gt;
&lt;th&gt;fonts&lt;/th&gt;
&lt;th&gt;images&lt;/th&gt;
&lt;th&gt;LCP median old→new&lt;/th&gt;
&lt;th&gt;5-run min–max (old / new)&lt;/th&gt;
&lt;th&gt;Δ&lt;/th&gt;
&lt;th&gt;CLS&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;425→344&lt;/td&gt;
&lt;td&gt;388→305&lt;/td&gt;
&lt;td&gt;1→1&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3,610→3,010 ms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;3,460–3,610 / 3,008–3,012&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;−600&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0→0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Lab article&lt;/td&gt;
&lt;td&gt;678→421&lt;/td&gt;
&lt;td&gt;359→275&lt;/td&gt;
&lt;td&gt;276→101 (WebP)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;4,806→3,459 ms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;4,805–4,807 / 3,309–3,461&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;−1,347&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0→0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;/works/pricing/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;395→314&lt;/td&gt;
&lt;td&gt;359→275&lt;/td&gt;
&lt;td&gt;1→1&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3,313→2,859 ms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;3,309–3,616 / 2,859–2,860&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;−454&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.001→0.001&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ul&gt;
&lt;li&gt;FCP medians: 3,010→2,260 / 2,708→1,959 / 2,709→1,959 ms. TBT 0 everywhere.&lt;/li&gt;
&lt;li&gt;Interquartile ranges (2nd vs 4th of the 5 sorted values — the stop rule uses this, not the min–max above): old 0–145 ms, new 1 ms. Simulated throttling is almost deterministic, so the differences are far outside the noise.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CLS did not move.&lt;/strong&gt; Both versions sit at 0.000–0.001. The figures on the article are below the fold, so &lt;code&gt;width&lt;/code&gt;/&lt;code&gt;height&lt;/code&gt; never had a chance to show up in this metric. H2 is simply not observable this way.&lt;/li&gt;
&lt;li&gt;LCP element is text on all three pages (h1, &lt;code&gt;p.sub&lt;/code&gt;). The &lt;code&gt;lcp-breakdown-insight&lt;/code&gt; audit puts it on the render-delay side (183–258 ms); what it was waiting for, the breakdown alone doesn't say.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;benchmarkIndex&lt;/code&gt;&lt;/strong&gt; (Lighthouse's host-CPU estimate) — per-page medians old→new: home 2,985→3,246, article 3,239→3,304, pricing 2,193→2,242. Higher on the &lt;em&gt;new&lt;/em&gt; side for all three pages. Lighthouse's CPU slowdown is relative to host performance, so this is a bias in new's favour, and I don't attribute the whole difference to the change set. One observation the other way: within old's five runs on the home page, benchmarkIndex moved from 2,417 to 3,329 while LCP stayed within 3,460–3,610.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What went wrong
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;First run: chrome-launcher's temp-dir cleanup threw &lt;code&gt;EPERM&lt;/code&gt; and the process exited 1 while the report JSON had been written. Fixed the runner to check for the file.&lt;/li&gt;
&lt;li&gt;Lighthouse 13's perf preset no longer has &lt;code&gt;largest-contentful-paint-element&lt;/code&gt;; the phases live in &lt;code&gt;lcp-breakdown-insight&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The frozen "new" commit and the measured one differ.&lt;/strong&gt; The design froze &lt;code&gt;098df1a&lt;/code&gt;; HEAD moved four commits (four Lab articles added) before I ran. The diff touches none of the three pages' HTML or CSS; the font woff2 files grew by a total of 2.6 KiB (ten bold characters added) — under 1% of transfer. I did not re-run; this line is the deviation record.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Verdict, and what the verdict means
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;H1 supported (−600 ms, outside the range). H3 supported (−454 ms). H2 not supported (old was already 0).&lt;/li&gt;
&lt;li&gt;"PASS" here means: &lt;em&gt;the pre-registered stop rule did not fire, the difference was observable&lt;/em&gt;. It does not mean the change set has been proven to cause the speed-up — the CPU index leaned toward new, and WebP, fonts and dimensions were changed together, so no per-change attribution is possible.&lt;/li&gt;
&lt;li&gt;Not measured: field data. Simulated throttling estimates other conditions from one observed trace; whether real users got faster is a separate question for Web Analytics / Search Console / CrUX after enough days.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The three-column sheet
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Column&lt;/th&gt;
&lt;th&gt;This time&lt;/th&gt;
&lt;th&gt;How to judge&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Proxy metric&lt;/td&gt;
&lt;td&gt;transfer −81 to −257 KiB&lt;/td&gt;
&lt;td&gt;went down? (necessary, not sufficient)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Outcome metric&lt;/td&gt;
&lt;td&gt;simulated LCP −454 to −1,347 ms; CLS 0→0&lt;/td&gt;
&lt;td&gt;outside the reproduction range (5-run IQR)?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stop rule&lt;/td&gt;
&lt;td&gt;difference &amp;lt; IQR ⇒ HOLD&lt;/td&gt;
&lt;td&gt;written before running&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Put "bytes went down" and "LCP went down" in separate columns and sign off on each. Building both versions on one machine and alternating five runs each was enough here to separate the observed difference from the run-to-run spread; it is not enough to say what caused what.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Part 1 of the series "Measure the Outcome, Not the Proxy" (part 2: shorter ≠ requirements preserved; part 3: screenshot success ≠ threshold discrimination — both on this account).&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;&lt;a href="https://sumitsuke.jp/via/devto/lab/transfer-size-vs-lcp-old-new/" rel="noopener noreferrer"&gt;Verification record&lt;/a&gt; (environment, verdict, the 30-run CSV, the LCP breakdown CSV and the runner) and the canonical Japanese write-up. What changed between old and new: &lt;a href="https://sumitsuke.jp/via/devto/lab/font-subset-cache-version-mix/" rel="noopener noreferrer"&gt;the font-subset story&lt;/a&gt; and &lt;a href="https://sumitsuke.jp/via/devto/lab/site-metadata-five-gaps-generators/" rel="noopener noreferrer"&gt;the image/metadata generators&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>performance</category>
      <category>lighthouse</category>
      <category>webperf</category>
      <category>testing</category>
    </item>
    <item>
      <title>Can an LLM measure UI thresholds from screenshots? 164 crops vs DOM gold labels</title>
      <dc:creator>Sumitsuke</dc:creator>
      <pubDate>Mon, 21 Sep 2026 08:31:50 +0000</pubDate>
      <link>https://dev.to/tauridev/can-an-llm-measure-ui-thresholds-from-screenshots-164-crops-vs-dom-gold-labels-4gbb</link>
      <guid>https://dev.to/tauridev/can-an-llm-measure-ui-thresholds-from-screenshots-164-crops-vs-dom-gold-labels-4gbb</guid>
      <description>&lt;p&gt;People paste screenshots into an LLM and ask "does this page follow the rules?". I wanted a number for how far that works, with the answer key coming from somewhere the model can't see: the DOM.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gold first
&lt;/h2&gt;

&lt;p&gt;Ten rules, frozen from my own site's style rules (the thresholds are mine, not a standard):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;R1 body text smaller than 15px&lt;/li&gt;
&lt;li&gt;R2 two different H2 sizes on the same page (judged per width)&lt;/li&gt;
&lt;li&gt;R3 body line length over 38em&lt;/li&gt;
&lt;li&gt;R4 h1 wrapping to 5+ lines at 375px&lt;/li&gt;
&lt;li&gt;R5 a 4-option radio group taller than 100px at 375px (wrapped)&lt;/li&gt;
&lt;li&gt;R6–R10 tap targets under 44px at 375px: TOC summary, TOC items, footer site-map links, the Works sub-nav, breadcrumbs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I built the site at the commit &lt;em&gt;before&lt;/em&gt; the fixes (&lt;code&gt;7bec67f&lt;/code&gt;), opened 12 pages at 1280 and 375 in headless Chrome via CDP, and measured every rule with &lt;code&gt;getComputedStyle&lt;/code&gt; / &lt;code&gt;getBoundingClientRect&lt;/code&gt;. That gives 12 × 11 = &lt;strong&gt;132 judgments&lt;/strong&gt; (R2 counts twice), of which 36 are "element not present on this page". 96 pairs remain.&lt;/p&gt;

&lt;p&gt;From the &lt;strong&gt;same render&lt;/strong&gt;, the script clips 1:1 crops of the elements in question — 164 PNGs, SHA-256 of each in a manifest. That matters: my first attempt measured positions in an iframe and shot with headless Chrome, and at 375px the two disagreed by 300–700px, so six footer crops were blank white.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// R1 = smallest font-size among text nodes inside &amp;lt;main&amp;gt;, excluding chrome&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;EXCL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;nav, .crumbs, form, button, .works-nav, .toc-block, footer, .site-header, .site-footer&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;textEls&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[...&lt;/span&gt;&lt;span class="nb"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;querySelectorAll&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;main *&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)].&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt;
  &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;SCRIPT&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;STYLE&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;SVG&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;PATH&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;NOSCRIPT&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tagName&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;closest&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;EXCL&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nf"&gt;vis&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt;
  &lt;span class="p"&gt;[...&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;childNodes&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;some&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;n&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;n&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;nodeType&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;n&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;textContent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;()));&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Judges
&lt;/h2&gt;

&lt;p&gt;Each judge got, per page, a frozen instruction ("for each rule answer YES / NO / unknown / not-applicable with one line of reasoning; a rule with no image is not-applicable; do not guess; at most three free remarks") plus that page's crops.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GPT&lt;/strong&gt;: pasted by hand into the ChatGPT UI, one page per submission, 12–15 images at once. The UI labelled the model "GPT5.6SOL"; the version was not captured.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini 2.5 Flash&lt;/strong&gt;: via the REST API, temperature 0, one page per call (the web UI caps at 10 images).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two different delivery paths, so this is &lt;strong&gt;not&lt;/strong&gt; a model-vs-model comparison. Scoring is a regex over the four verbatim values against the gold — no human in the loop.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mistake in my first tally
&lt;/h2&gt;

&lt;p&gt;My first summary said "GPT 94% on 96 pairs, and it got all 35 of the 44px pairs right, telling 40px from 45px". An external reviewer looked at the gold and pointed out that in the &lt;em&gt;before&lt;/em&gt; version, every breadcrumb is 40px and every footer link is 45px: within each 44px rule there are &lt;strong&gt;only violations or only non-violations&lt;/strong&gt;. A judge that answers "small links violate, the footer doesn't" scores 35/35 without measuring anything. The same holds for R5.&lt;/p&gt;

&lt;p&gt;So the main metric is now the &lt;strong&gt;60 pairs whose rule contains both labels&lt;/strong&gt; (R1–R4). The 44px rules were scored (GPT 36/36) but that score cannot test threshold discrimination, so it is reported and not counted.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6tuithmfzl8i17iq3zc0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6tuithmfzl8i17iq3zc0.png" alt="Accuracy on the 60 contrastive pairs, all-pairs basis: GPT visual 36/36, numeric 18/24; Gemini visual 28/36, numeric 13/24. All six GPT misses are the 15px rule; the 44px rules have no contrast in the gold" width="800" height="1828"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Results (60 contrastive pairs)
&lt;/h2&gt;

&lt;p&gt;Accuracy = correct YES/NO including true negatives, divided by &lt;strong&gt;all&lt;/strong&gt; pairs (unknowns count against). GPT answered 60/60, so the two bases coincide; Gemini answered 57/60, so I also give the answered-only rate.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Judge&lt;/th&gt;
&lt;th&gt;Group&lt;/th&gt;
&lt;th&gt;Pairs&lt;/th&gt;
&lt;th&gt;Answered&lt;/th&gt;
&lt;th&gt;Correct&lt;/th&gt;
&lt;th&gt;FP rate&lt;/th&gt;
&lt;th&gt;FN rate&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT&lt;/td&gt;
&lt;td&gt;contrastive R1–R4&lt;/td&gt;
&lt;td&gt;60&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;54/60 = 90%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;7%&lt;/td&gt;
&lt;td&gt;21%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT&lt;/td&gt;
&lt;td&gt;visual (R2, R4)&lt;/td&gt;
&lt;td&gt;36&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;36/36&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT&lt;/td&gt;
&lt;td&gt;numeric (R1, R3)&lt;/td&gt;
&lt;td&gt;24&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;18/24&lt;/td&gt;
&lt;td&gt;16%&lt;/td&gt;
&lt;td&gt;60%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini&lt;/td&gt;
&lt;td&gt;contrastive R1–R4&lt;/td&gt;
&lt;td&gt;60&lt;/td&gt;
&lt;td&gt;95%&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;41/60 = 68%&lt;/strong&gt; (41/57 = 72% of answered)&lt;/td&gt;
&lt;td&gt;26%&lt;/td&gt;
&lt;td&gt;36%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini&lt;/td&gt;
&lt;td&gt;visual&lt;/td&gt;
&lt;td&gt;36&lt;/td&gt;
&lt;td&gt;97%&lt;/td&gt;
&lt;td&gt;28/36&lt;/td&gt;
&lt;td&gt;12%&lt;/td&gt;
&lt;td&gt;44%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini&lt;/td&gt;
&lt;td&gt;numeric&lt;/td&gt;
&lt;td&gt;24&lt;/td&gt;
&lt;td&gt;92%&lt;/td&gt;
&lt;td&gt;13/24&lt;/td&gt;
&lt;td&gt;47%&lt;/td&gt;
&lt;td&gt;20%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;All six GPT misses are R1&lt;/strong&gt;: three pages at 13.5px answered NO, three pages at 15px answered YES. A 1.5px difference. R3 (38em vs pages at 18–33em, i.e. a 5–20em gap) was 12/12.&lt;/li&gt;
&lt;li&gt;Gemini's misses: six pages at 18–33em judged as "over 38em"; four false negatives and three false positives on mixed H2 sizes.&lt;/li&gt;
&lt;li&gt;Non-contrastive R5–R10 (36 pairs), for the record: GPT 36/36; Gemini answered only 56%, called the folded footer a "nav" five times and answered "not applicable" for elements that were in the image ten times.&lt;/li&gt;
&lt;li&gt;Reproducibility, one observation: the top page was submitted to GPT twice (pilot, then batch) and its R1 answer flipped YES → NO. The scored run is the second one.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Pre-registered hypotheses
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;H1 "visual rules are caught by &lt;strong&gt;both&lt;/strong&gt; judges" — GPT yes (36/36), Gemini no (FN 44%). &lt;strong&gt;Not supported as stated.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;H2 "numeric rules are unknown or wrong in more than half of cases" — &lt;strong&gt;on hold&lt;/strong&gt;: only 24 contrastive numeric pairs exist (GPT 18/24, Gemini 13/24), the 44px rules are unmeasured, and the one discriminating rule flipped between two runs.&lt;/li&gt;
&lt;li&gt;H3 "free remarks are mostly outside the rules" — GPT wrote almost none (one remark in 12 pages); supported for Gemini.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Three questions before handing a check to a screenshot
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Does the gold for that rule contain &lt;strong&gt;both&lt;/strong&gt; labels? If not, a perfect score means nothing.&lt;/li&gt;
&lt;li&gt;How far is the case from the threshold — &lt;strong&gt;how many px or em&lt;/strong&gt;? A 1.5px gap failed even at 1:1.&lt;/li&gt;
&lt;li&gt;Does the answer survive a &lt;strong&gt;second submission&lt;/strong&gt; of the same images?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;What I'd hand to image review today: mixed heading sizes, heading line counts, line length with a multi-em gap. What stays in the DOM: anything within a couple of pixels of its threshold. What I can't say yet: whether image review discriminates the 44px tap-target threshold — the fixed version of the site needs to be shot to create contrast.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn342d3265jb9iwl556bq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn342d3265jb9iwl556bq.png" alt="Division of labour: hand to image review (visual rules, multi-em line length) / measure in the DOM (a few px from the threshold) / not tested (44px, no contrast) / a human (inside the rules but hard to read)" width="800" height="2112"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Part 3 of the series "Measure the Outcome, Not the Proxy" (part 1: transfer size ≠ speed; part 2: shorter ≠ requirements preserved — both on this account).&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;&lt;a href="https://sumitsuke.jp/via/devto/lab/screenshot-ai-vs-dom-ten-rules/" rel="noopener noreferrer"&gt;Verification record&lt;/a&gt; (environment, verdict, evidence: gold JSON, both score sheets, the 164 crops as a ZIP, all 24 verbatim answers) and the canonical Japanese write-up. &lt;a href="https://sumitsuke.jp/via/devto/lab/readability-dom-eight-items/" rel="noopener noreferrer"&gt;The DOM measurer that fixed the pages&lt;/a&gt;. &lt;a href="https://github.com/sumitsuke/screenshot-ai-vs-dom-ten-rules" rel="noopener noreferrer"&gt;Reproduction kit&lt;/a&gt; with the images, the gold, the judge script for Gemini and the scorer (code MIT, data CC BY 4.0).&lt;/em&gt;&lt;/p&gt;

</description>
      <category>llm</category>
      <category>a11y</category>
      <category>css</category>
      <category>testing</category>
    </item>
    <item>
      <title>I asked an LLM to cut 14 texts by 20–30%. Median: only 7%</title>
      <dc:creator>Sumitsuke</dc:creator>
      <pubDate>Mon, 21 Sep 2026 08:31:43 +0000</pubDate>
      <link>https://dev.to/tauridev/i-asked-an-llm-to-cut-14-texts-by-20-30-median-only-7-59nc</link>
      <guid>https://dev.to/tauridev/i-asked-an-llm-to-cut-14-texts-by-20-30-median-only-7-59nc</guid>
      <description>&lt;p&gt;I run a small site whose service pages are 1,100–1,800 Japanese characters each. I wanted them shorter, so I asked a model to cut them by 20–30% — and then, because the pages contain prices, limits and exceptions, I added "keep the numbers and the sentences that state conditions". My assumption was that the "keep" clause is what stops the model from reaching the target.&lt;/p&gt;

&lt;p&gt;So I measured it instead of assuming. Fourteen documents, four instruction conditions, 56 outputs, each scored three ways.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4u0j50585sojkk8am8om.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4u0j50585sojkk8am8om.png" alt="Requested −20 to −30% vs. actual medians: 0 −6.5%, A −6.9%, B −8.2%, C −8.4%; reached −20%: 0, 0, 2, 1 of 14" width="800" height="1594"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Setup (frozen before generation)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Material&lt;/strong&gt;: 14 pages from my own site, 1,149–1,797 characters after stripping front-matter, HTML tags, the closing CTA and whitespace. SHA-256 of each file frozen.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Protected atoms&lt;/strong&gt;: every &lt;code&gt;number + unit&lt;/code&gt; string in the source (154 of them: 「33,000 円」, 「3 つ」, 「1 営業日」…), as a multiset.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Boundary claims&lt;/strong&gt;: 29 sentences that state a condition, a limit or an exception (hand-picked from candidates, frozen).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Conditions&lt;/strong&gt; (nested, verbatim in the kit):

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;0&lt;/strong&gt; — "Shorten the following text by 20–30%. Output only the shortened text."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A&lt;/strong&gt; — 0 + "Do not change the structure of headings, paragraphs and lists."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;B&lt;/strong&gt; — A + "Keep the numbers and the sentences that state conditions, limits and exceptions."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;C&lt;/strong&gt; — B + "Lists may be converted to bullet points."&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model&lt;/strong&gt;: &lt;code&gt;claude -p --model claude-opus-5&lt;/code&gt;, text on stdin, run from an empty directory. Note that &lt;code&gt;claude -p&lt;/code&gt; still carries Claude Code's own system prompt, so this is not a bare API call.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hypotheses&lt;/strong&gt;: H1 B compresses less than A. H2 A reaches −20% but drops an atom or a claim in more than half of the documents. H3 C compresses more than B.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;HOLD threshold&lt;/strong&gt;: I first ran 3 documents × condition A × 3 times. The same input landed anywhere from −3.5% to −20.8%; the median spread was &lt;strong&gt;6.6 points&lt;/strong&gt;. Any difference between conditions smaller than that is "not distinguishable".&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Condition 0 was added after an external review pointed out that A–C all contain the "keep the structure" clause, so the main result could not be separated from it. The other 42 outputs were generated in randomized condition order (seed 20260919), one run per cell.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scoring
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Character change (whitespace removed).&lt;/li&gt;
&lt;li&gt;Atoms: an atom is "lost" when it appears fewer times in the output than in the source. This is a &lt;strong&gt;string check&lt;/strong&gt; — 「3 つの確認」 ("three checks") rendered as a three-item list still counts as lost. No meaning adjudication.&lt;/li&gt;
&lt;li&gt;Claims: normalized-string containment for automatic pass; everything else was shown to a human with the closest output sentence (&lt;code&gt;difflib&lt;/code&gt;) and judged &lt;em&gt;paraphrased-kept / lost / inverted&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;Structure: heading count and bullet count.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;ATOM&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;([\d,]+(?:\.\d+)?)\s*(円|件|本|日|時間|営業日|%|％|項目|段階|か所|回|人|社|通|枚|行|字|ミリ秒|秒|分|か月|つ|層|型|種|点|倍)&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;lost&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;src&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Condition&lt;/th&gt;
&lt;th&gt;n&lt;/th&gt;
&lt;th&gt;Median change&lt;/th&gt;
&lt;th&gt;Quartiles&lt;/th&gt;
&lt;th&gt;Min / max&lt;/th&gt;
&lt;th&gt;Reached −20%&lt;/th&gt;
&lt;th&gt;Docs that lost ≥1 atom string&lt;/th&gt;
&lt;th&gt;Claims kept&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;0 no constraint&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;−6.5%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;−7.5 … −5.1&lt;/td&gt;
&lt;td&gt;−9.6 / −4.0&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;28/29&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A keep structure&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;−6.9%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;−8.2 … −5.4&lt;/td&gt;
&lt;td&gt;−11.0 / −2.7&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;28/29&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B + keep numbers &amp;amp; conditions&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;−8.2%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;−10.3 … −6.5&lt;/td&gt;
&lt;td&gt;−21.7 / −2.4&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;28/29&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C + bullets allowed&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;−8.4%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;−10.8 … −6.6&lt;/td&gt;
&lt;td&gt;−21.1 / −3.7&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;28/29&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Quartiles are the 4th and 11th of 14 sorted values; &lt;code&gt;statistics.quantiles&lt;/code&gt; moves them by about half a point.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Differences between conditions: A−0 = 0.4, B−A = 1.3, C−B = 0.2 points. All inside the 6.6-point reproduction width. &lt;strong&gt;H1 and H3: HOLD.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;H2 failed at its premise&lt;/strong&gt;: A never reached −20% (0/14), and neither did the unconstrained condition 0 (0/14). The second half happened anyway — 9 of 14 A-outputs lost an atom string or a claim.&lt;/li&gt;
&lt;li&gt;The three outputs that did reach −20% were all B or C. The "keep" clause did not prevent compression either.&lt;/li&gt;
&lt;li&gt;What disappeared: counts — 「3 つ」, 「5 つ」, 「1 通」, 「3 か月」, 「10 件」. Every price (18 occurrences of 「…円」 across 6 documents × 3 conditions) survived.&lt;/li&gt;
&lt;li&gt;Boundary claims, 29 × 4 = 116: 74 exact matches, 38 judged paraphrased-kept, &lt;strong&gt;4 lost, 0 inverted&lt;/strong&gt;. The 4 lost are the &lt;strong&gt;same sentence in all four conditions&lt;/strong&gt;: 「合否条件を通した検証の結果と、確かめていない範囲も一緒に渡します」 ("we hand over the results of verification &lt;em&gt;that passed the acceptance criteria&lt;/em&gt;, together with what was not checked") became 「検証の結果と…」 every time — the acceptance-criteria wording was dropped whether or not the instruction said to keep such sentences. (Its condition-0 row was first batch-judged as a paraphrase; a second review aligned it with A–C. The change is recorded in the published sheet.)&lt;/li&gt;
&lt;li&gt;Heading and bullet counts were identical to the source in 56/56 outputs, including condition 0 where structure was never mentioned. Condition C's permission to convert lists into bullets was never used.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What I can and cannot say
&lt;/h2&gt;

&lt;p&gt;Can: under these run conditions, a 20–30% request stopped at a median of −6.5 to −8.4% &lt;em&gt;with or without&lt;/em&gt; constraints, and the constraint I suspected did not explain the shortfall. All 18 price strings survived while some count strings disappeared; structure was kept without being asked; and the acceptance-criteria phrase was cut from the same sentence in every condition.&lt;/p&gt;

&lt;p&gt;Cannot: what actually sets the ceiling. Other models, other temperatures, a single-point target ("exactly −25%"), documents over 3,000 characters, or a second run of the same cell (n = 1 each; the pilot shows −21% and −7% for the same input).&lt;/p&gt;

&lt;h2&gt;
  
  
  The acceptance sheet I use now
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Column&lt;/th&gt;
&lt;th&gt;How to write it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Compression&lt;/td&gt;
&lt;td&gt;character change (whitespace removed) ≤ −X%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Protected atoms&lt;/td&gt;
&lt;td&gt;every &lt;code&gt;number + unit&lt;/code&gt; string appears at least as often as in the source (string check only)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Boundary claims&lt;/td&gt;
&lt;td&gt;every condition / limit / exception sentence survives, paraphrase allowed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;"Shorter" is a proxy. What I wanted to know was whether the numbers and the conditions survived. Counting the two separately showed both things at once: the target was not reached, and things were still being dropped on the way.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Part 2 of the series "Measure the Outcome, Not the Proxy" (part 1: transfer size ≠ speed; part 3: screenshot success ≠ threshold discrimination — both on this account).&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;&lt;a href="https://sumitsuke.jp/via/devto/lab/ask-ai-to-cut-30-percent-stops-at-8/" rel="noopener noreferrer"&gt;Verification record&lt;/a&gt; (environment, verdict, evidence links) and the canonical write-up in Japanese. &lt;a href="https://github.com/sumitsuke/ask-ai-to-cut-30-percent-stops-at-8" rel="noopener noreferrer"&gt;Reproduction kit&lt;/a&gt; — the 14 source texts, the three scripts, all 56 outputs and the two adjudication sheets (code MIT, data CC BY 4.0).&lt;/em&gt;&lt;/p&gt;

</description>
      <category>llm</category>
      <category>prompt</category>
      <category>testing</category>
      <category>writing</category>
    </item>
    <item>
      <title>I let an AI write a technical article with zero human editing, then ran it through my normal verification pass</title>
      <dc:creator>Sumitsuke</dc:creator>
      <pubDate>Thu, 17 Sep 2026 04:29:22 +0000</pubDate>
      <link>https://dev.to/tauridev/i-let-an-ai-write-a-technical-article-with-zero-human-editing-then-ran-it-through-my-normal-cj6</link>
      <guid>https://dev.to/tauridev/i-let-an-ai-write-a-technical-article-with-zero-human-editing-then-ran-it-through-my-normal-cj6</guid>
      <description>&lt;p&gt;I gave a model a folder of material and one sentence of instruction, and let it write a technical article end to end — topic selection, structure, prose, figures, and its own pre-publication checklist. Nobody edited a word. Then I froze that output and ran it through the verification pass I normally run on my own writing.&lt;/p&gt;

&lt;p&gt;What I was measuring is not "is AI writing any good." It's &lt;strong&gt;what one verification pass actually removes&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Numbers first
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;11 findings&lt;/strong&gt; on the six-type checklist I fixed &lt;em&gt;before&lt;/em&gt; generation. One of them could recur in production.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;4 more&lt;/strong&gt; on operational rules the model was never given (publish flag, information I don't publish, AI-use disclosure, internal linking). Two of those were the serious ones.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Three external AI review rounds&lt;/strong&gt;: 3 / 4 / 4 true new findings, and &lt;strong&gt;1 / 1 / 2 false ones&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The same false finding appeared four times&lt;/strong&gt; across separate sessions — twice inside the three rounds, and twice in two later passes (one that re-read A, one that checked the fixed version). Twice it was phrased as "a clear, reproducible error."&lt;/li&gt;
&lt;li&gt;My own pass, before any AI review, found &lt;strong&gt;0 of the 11&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Production time: &lt;strong&gt;12 minutes&lt;/strong&gt;. Verification: &lt;strong&gt;about 105 minutes&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Sample size is one article. I'm not generalizing from it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup, fixed before generation
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Material&lt;/td&gt;
&lt;td&gt;3 files, 5,085 bytes: a topic memo, a measurement log (raw command output, no interpretation), and 13 lines of the source being discussed. Hashed.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Instruction&lt;/td&gt;
&lt;td&gt;One sentence: write a publishable technical article from this material only; do the topic selection, structure, figures and pre-publication check yourself; invent no numbers that aren't in the material.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Environment&lt;/td&gt;
&lt;td&gt;Fresh working directory, no project instructions anywhere up the tree, no user-level instruction file, settings sources restricted. Memory directory verified empty after the run.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Checklist&lt;/td&gt;
&lt;td&gt;Six types, plus the definition of major/minor, written &lt;strong&gt;before&lt;/strong&gt; generation and &lt;strong&gt;not&lt;/strong&gt; given to the writing model.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A&lt;/td&gt;
&lt;td&gt;The output, frozen (hash + commit). Never edited.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B&lt;/td&gt;
&lt;td&gt;A copy of A with only confirmed fixes applied.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The six types: (1) wrong numbers or facts, (2) citation doesn't support the claim, (3) command or code doesn't reproduce, (4) internal contradiction, (5) generalization beyond what was measured, (6) conflating spec, observation and inference.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the model did on its own
&lt;/h2&gt;

&lt;p&gt;Twelve minutes produced an article, three figures, the scripts that generated them, and a self-check page. In that self-check it had executed the article's own reproduction snippet and compared the output, and machine-verified that the quoted source matched the material line for line. Mid-run it caught one of its own claims being wrong and rewrote it — I re-ran that case afterwards and the rewritten version is correct.&lt;/p&gt;

&lt;p&gt;It also narrowed the topic without being told to. The memo listed five items; only four had real artifacts behind them. For the fifth, the material itself said "no run log exists," so the article put it under "what I did not cover" and published no number for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it actually broke
&lt;/h2&gt;

&lt;p&gt;Not where I expected.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;Found in A&lt;/th&gt;
&lt;th&gt;Left in B&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Wrong numbers or facts&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Citation mismatch&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Doesn't reproduce&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Internal contradiction&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Generalization beyond measurement&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Spec / observation / inference conflated&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;11&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every number carried over from the material was right — 16 of them. The reproduction snippet ran and matched down to the exception's character offset. The quoted source matched the real file 10 lines out of 10.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Exactly one finding was "a number is wrong,"&lt;/strong&gt; and even that was a breakdown whose parts summed to 971 against a stated total of 973 — the material had only listed the top four categories and the article presented them as the whole set.&lt;/p&gt;

&lt;p&gt;The other ten were all &lt;em&gt;absences&lt;/em&gt;: a citation that was never given, a condition dropped (a regex claim stated without the qualifier that makes it true), a cause asserted from a &lt;code&gt;grep&lt;/code&gt; that returned nothing, a conclusion wider than the range actually checked.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Accurate inside the material, expansive outside it.&lt;/strong&gt; That was this article's failure mode.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one that could recur
&lt;/h2&gt;

&lt;p&gt;The article's subject was a guard that escapes invisible characters so a JSON line can't be split by the reader. The guard handled two characters.&lt;/p&gt;

&lt;p&gt;Nobody — not the model, not me — asked how many characters can do that. The class isn't "invisible characters": it's the intersection of &lt;em&gt;what the reader splits on&lt;/em&gt; and &lt;em&gt;what the encoder leaves raw&lt;/em&gt;. For &lt;code&gt;str.splitlines()&lt;/code&gt; and &lt;code&gt;json.dumps(ensure_ascii=False)&lt;/code&gt; that intersection has three members. The third had been unguarded for four months, in the real tool, while every check stayed green.&lt;/p&gt;

&lt;p&gt;That is not a wrong number. It's &lt;strong&gt;"sealed, probably."&lt;/strong&gt; And the only thing that finds it is counting the members of the class, which is work that lives outside the guard.&lt;/p&gt;

&lt;p&gt;I fixed the tool, then required red on a poison case for each of the three characters plus all three at once, before calling it closed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rules it was never given
&lt;/h2&gt;

&lt;p&gt;Four more findings sat outside the checklist: the publish flag was left on, an identifier I don't publish appeared twice, there was no AI-use disclosure, and no link back to the canonical version. Two of those are the most expensive items in the whole list — they're the ones that cost you something after publication, not before.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I don't count them as model failures.&lt;/strong&gt; None of those rules were in the material or the instruction. Charging a system for rules you never gave it measures your own setup, not its capability. They're recorded, separately, as an operations gate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three review rounds, and the interesting failure
&lt;/h2&gt;

&lt;p&gt;I ran the frozen A past an external model three times: new session each time, same artifact, no B, no previous round's output.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;#1&lt;/th&gt;
&lt;th&gt;#2&lt;/th&gt;
&lt;th&gt;#3&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;True new findings&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Repeats of earlier rounds&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;False findings&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;New true findings never hit zero in three rounds, so &lt;strong&gt;this experiment did not measure a stopping point&lt;/strong&gt;. That's the result: not "three is too few" and not "four would find more" — I don't know where it saturates, because I stopped before it did.&lt;/p&gt;

&lt;p&gt;The part worth your time is the false ones. Two of the four inside these three rounds were the same claim — and it came back twice more in later passes (a re-read of A, and the check on the fixed version), so &lt;strong&gt;the same misreading arrived four times in total&lt;/strong&gt;, twice with language like "this is a clear, reproducible error, fix before publishing." The claim was that an escape sequence in the article's code was malformed and would break at runtime.&lt;/p&gt;

&lt;p&gt;I pulled the exact line out of the file as bytes and executed it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;backslash count: correct&lt;/li&gt;
&lt;li&gt;raw separator characters in the line: zero&lt;/li&gt;
&lt;li&gt;evaluate the expression as written: one line out, value round-trips&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The claim described a real failure mode — for code that wasn't in the article. Most likely an unescaping step somewhere in how the file was read.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agreement across separate sessions is not evidence.&lt;/strong&gt; Same model, same prompt, different session is not an independent trial; the errors correlate too. Neither is confident phrasing. The only thing that settled it was pulling the bytes and running them.&lt;/p&gt;

&lt;h2&gt;
  
  
  My own pass missed all 11
&lt;/h2&gt;

&lt;p&gt;The order is: external review → I confirm each finding against primary sources and by executing things → a human adjudicates.&lt;/p&gt;

&lt;p&gt;My own pass, run before any of that, found zero of the eleven. I verified that all 16 numbers matched the material — and never checked whether a breakdown summed to its own stated total. I read "it escapes the two characters" and never asked how many there were.&lt;/p&gt;

&lt;p&gt;The reason is structural: I was &lt;strong&gt;checking claims against material&lt;/strong&gt;, and every one of these lives in the gap between the material and the world. Checking and searching are different jobs, and I'd been doing only the first.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix also broke something
&lt;/h2&gt;

&lt;p&gt;When I re-ran the checklist on B, it produced a finding of its own: I had moved a section during the fixes and left two references to the old state, so the same article said both "sealed" and "not sealed."&lt;/p&gt;

&lt;p&gt;So the honest version of the summary isn't "11 → 0." It's:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;detected in A: 11 (plus 4 operational)&lt;/li&gt;
&lt;li&gt;of those, left in B: 0&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;introduced by the fixing: 1&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;undetected in B: unknown, and I'm not going to pretend otherwise&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you only count what you removed, the ledger flatters you.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you want to run this
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Freeze the material and hash it. The writing model gets the material and one sentence.&lt;/li&gt;
&lt;li&gt;Fix the checklist &lt;strong&gt;before&lt;/strong&gt; generation — types and the major/minor definition both.&lt;/li&gt;
&lt;li&gt;Freeze the output. Never edit it; work on a copy.&lt;/li&gt;
&lt;li&gt;Use AI review as a &lt;strong&gt;candidate generator&lt;/strong&gt;. Truth is decided by primary sources and execution, not by the reviewer's confidence.&lt;/li&gt;
&lt;li&gt;Keep rules the model was never given in a separate column. Don't charge it for what you didn't tell it.&lt;/li&gt;
&lt;li&gt;Count three things: found, left, and &lt;strong&gt;introduced by fixing&lt;/strong&gt;. The third one is the one people skip.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;One line connects the article's own failure and the reviewers': &lt;strong&gt;the absence of an alarm was treated as evidence of absence.&lt;/strong&gt; The guard was silent because it had never been handed the case; my pass was silent because I was checking claims rather than searching for the class; the reviewers were confident because agreement felt like verification. Each time, nothing firing was read as nothing being there.&lt;/p&gt;

&lt;p&gt;What I'd take away, with one sample and no generalization: the writing was accurate where it had material and loose where it didn't, the expensive mistakes were operational rather than factual, and the review that mattered was the one that executed the code rather than the one that read it.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Verification record (environment, verdict, last verified date, evidence) and the canonical write-up (Japanese) — the frozen article A, the six-type checklist and the defect ledger: &lt;a href="https://sumitsuke.jp/via/devto/lab/cost-of-one-verification-pass/" rel="noopener noreferrer"&gt;https://sumitsuke.jp/via/devto/lab/cost-of-one-verification-pass/&lt;/a&gt; .&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>writing</category>
      <category>testing</category>
      <category>review</category>
    </item>
    <item>
      <title>A JSONL record split in two: U+2028, U+0085, and the separator I missed</title>
      <dc:creator>Sumitsuke</dc:creator>
      <pubDate>Wed, 16 Sep 2026 13:29:36 +0000</pubDate>
      <link>https://dev.to/tauridev/a-jsonl-record-split-in-two-u2028-u0085-and-the-separator-i-missed-2oi6</link>
      <guid>https://dev.to/tauridev/a-jsonl-record-split-in-two-u2028-u0085-and-the-separator-i-missed-2oi6</guid>
      <description>&lt;p&gt;I keep a ledger as JSONL — one JSON object per line, written by my own importer, read by half a dozen small checkers. One day the checkers started reporting a line as unparseable, while &lt;code&gt;json.loads&lt;/code&gt; on the same line was perfectly happy.&lt;/p&gt;

&lt;p&gt;The character in the middle was U+2028 LINE SEPARATOR.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why one line became two
&lt;/h2&gt;

&lt;p&gt;Two facts, both documented, that only bite when they meet:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;str.splitlines()&lt;/code&gt; treats U+2028 as a line boundary — along with &lt;code&gt;\n&lt;/code&gt;, &lt;code&gt;\r&lt;/code&gt;, U+2029, U+0085 and a few C0 controls (&lt;a href="https://docs.python.org/3/library/stdtypes.html#str.splitlines" rel="noopener noreferrer"&gt;str.splitlines&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;json.dumps(..., ensure_ascii=False)&lt;/code&gt; emits non-ASCII characters as themselves. U+2028 is not in the set JSON &lt;em&gt;requires&lt;/em&gt; you to escape, so it stays raw (&lt;a href="https://docs.python.org/3/library/json.html" rel="noopener noreferrer"&gt;json&lt;/a&gt;).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;So the file has exactly one &lt;code&gt;\n&lt;/code&gt; per record, &lt;code&gt;json.loads&lt;/code&gt; round-trips fine, and any reader that iterates with &lt;code&gt;splitlines()&lt;/code&gt; sees a broken fragment.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;

&lt;span class="n"&gt;rec&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;body&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;before&lt;/span&gt;&lt;span class="se"&gt;\u2028&lt;/span&gt;&lt;span class="s"&gt;after&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="n"&gt;line&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rec&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ensure_ascii&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;     &lt;span class="c1"&gt;# 1   … one newline
&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;splitlines&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;        &lt;span class="c1"&gt;# 2   … splits here
&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;rec&lt;/span&gt;       &lt;span class="c1"&gt;# True … valid JSON the whole time
&lt;/span&gt;
&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;splitlines&lt;/span&gt;&lt;span class="p"&gt;()[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="c1"&gt;# JSONDecodeError: Unterminated string starting at: line 1 column 19 (char 18)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The fix I shipped, and why it was on the writer
&lt;/h2&gt;

&lt;p&gt;There were several readers and one writer, so I fixed the writer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;dumps_line&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rec&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rec&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ensure_ascii&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\u2028&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\\&lt;/span&gt;&lt;span class="s"&gt;u2028&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\u2029&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\\&lt;/span&gt;&lt;span class="s"&gt;u2029&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The invariant I should have written down on day one, and eventually did:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The serializer must not emit a raw code point that any of its readers treats as a line boundary.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Stated that way, the fix isn't "escape U+2028" — it's "keep the emitted set and the readers' boundary set disjoint," and the list of characters becomes something you derive rather than remember.&lt;/p&gt;

&lt;p&gt;Three deliberate choices:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Keep &lt;code&gt;ensure_ascii=False&lt;/code&gt;.&lt;/strong&gt; Setting it to &lt;code&gt;True&lt;/code&gt; would also solve this, but then every Japanese character becomes &lt;code&gt;\uXXXX&lt;/code&gt; and the ledger stops being readable by eye. Being readable &lt;em&gt;is&lt;/em&gt; the point of this file, so I targeted the separators only.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The value doesn't change.&lt;/strong&gt; &lt;code&gt;\u2028&lt;/code&gt; is a standard JSON escape, so &lt;code&gt;json.loads&lt;/code&gt; gives back the original character. I changed the representation, not the data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Put the incident in the docstring.&lt;/strong&gt; The reason a &lt;code&gt;replace()&lt;/code&gt; like this exists is exactly the thing a future me deletes during cleanup.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Existing rows still held the raw character at that moment. They got fixed sideways: the ledger tool's update path rewrites every row through &lt;code&gt;dumps_line&lt;/code&gt;, so the next update swept them. Counted today across all 973 rows: &lt;strong&gt;0 raw U+2028, 0 raw U+2029, 0 raw U+0085&lt;/strong&gt;. That the sweep happened at all was luck, not design — worth saying out loud.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four months later, I counted
&lt;/h2&gt;

&lt;p&gt;I wrote "escaping the two separator characters" in the docstring and in the article. Neither the code nor I had ever asked the obvious question: &lt;strong&gt;how many characters can actually do this?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Of the characters &lt;code&gt;splitlines()&lt;/code&gt; splits on, the ones that survive &lt;code&gt;json.dumps(ensure_ascii=False)&lt;/code&gt; as raw text are:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Character&lt;/th&gt;
&lt;th&gt;Raw in &lt;code&gt;dumps&lt;/code&gt; output?&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;splitlines()&lt;/code&gt; splits?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;U+2028 LINE SEPARATOR&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;U+2029 PARAGRAPH SEPARATOR&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;U+0085 NEXT LINE&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;yes&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;yes&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;U+000B, U+000C, U+001C–U+001E&lt;/td&gt;
&lt;td&gt;no (escaped as &lt;code&gt;\u000b&lt;/code&gt;, …)&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three, not two. The C0 controls are harmless here precisely because &lt;code&gt;json.dumps&lt;/code&gt; is &lt;em&gt;required&lt;/em&gt; to escape them; U+0085 is above 0x20, so nothing escapes it for you.&lt;/p&gt;

&lt;p&gt;My guard had been green for four months on a hole it was never given a chance to see. The cases it covered worked; the scope was incomplete.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;dumps_line&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rec&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;_s&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rec&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ensure_ascii&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_ch&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_esc&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\u2028&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\\&lt;/span&gt;&lt;span class="s"&gt;u2028&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
                      &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\u2029&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\\&lt;/span&gt;&lt;span class="s"&gt;u2029&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
                      &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\u0085&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\\&lt;/span&gt;&lt;span class="s"&gt;u0085&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)):&lt;/span&gt;
        &lt;span class="n"&gt;_s&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;_s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;_ch&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_esc&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;_s&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The check that would have caught it on day one
&lt;/h2&gt;

&lt;p&gt;Not more test data. A &lt;strong&gt;positive control&lt;/strong&gt;: feed in a case that &lt;em&gt;must&lt;/em&gt; be caught and require the failure.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Injected&lt;/th&gt;
&lt;th&gt;&lt;code&gt;splitlines()&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;value round-trips&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;U+2028&lt;/td&gt;
&lt;td&gt;1 line&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;U+2029&lt;/td&gt;
&lt;td&gt;1 line&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;U+0085&lt;/td&gt;
&lt;td&gt;1 line&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;all three at once&lt;/td&gt;
&lt;td&gt;1 line&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Before the fix, the U+0085 row returned 2 lines. That single red is the whole value of the exercise: a test that only ever passes has told you nothing about the cases you didn't hand it.&lt;/p&gt;

&lt;p&gt;There's a pattern here that generalizes past Unicode. When you write a guard against a class of input, the number that matters is not "did it pass" but &lt;strong&gt;how many members of the class you enumerated&lt;/strong&gt;. I had enumerated two out of three, and the guard could not tell me that, because enumeration is the part outside the guard.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this doesn't cover
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;This is about the &lt;em&gt;writer&lt;/em&gt;. A reader that uses &lt;code&gt;splitlines()&lt;/code&gt; is still free to split on anything; I chose one place to fix instead of six.&lt;/li&gt;
&lt;li&gt;Other producers writing to the same file (a different script, a hand edit) bypass it entirely.&lt;/li&gt;
&lt;li&gt;The shape isn't Python-specific: the gap exists wherever a line splitter treats some code point as a boundary and the serializer emits that same code point raw. Which characters those are depends on the pair you're using — enumerate the intersection for your own splitter and encoder rather than copying my three.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What I'd do differently
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;When writing a guard against "characters of type X", &lt;strong&gt;print the enumerated list in the code&lt;/strong&gt;, not in prose. &lt;code&gt;("\u2028", "\u2029", "\u0085")&lt;/code&gt; next to the &lt;code&gt;replace&lt;/code&gt; is checkable; "the separator characters" in a docstring is not.&lt;/li&gt;
&lt;li&gt;Add one poison case per member, and require red before the fix.&lt;/li&gt;
&lt;li&gt;When the guard's scope changes, fix the message and the docstring in the same commit. Mine still said "only two" while the array said three, one revision later — the code was right and the text was lying.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Counted, not estimated: 973 rows with 0 raw U+2028 / U+2029 / U+0085 today, 3 characters that can split a line through this encoder, 1 of which was unguarded for four months.&lt;/p&gt;

</description>
      <category>python</category>
      <category>json</category>
      <category>debugging</category>
      <category>testing</category>
    </item>
    <item>
      <title>Does AI-generated code silently swallow errors? 120 measured generations: every flagged case was a false positive or a documented fallback</title>
      <dc:creator>Sumitsuke</dc:creator>
      <pubDate>Tue, 15 Sep 2026 12:32:16 +0000</pubDate>
      <link>https://dev.to/tauridev/does-ai-generated-code-silently-swallow-errors-120-measured-generations-every-flagged-case-was-a-241p</link>
      <guid>https://dev.to/tauridev/does-ai-generated-code-silently-swallow-errors-120-measured-generations-every-flagged-case-was-a-241p</guid>
      <description>&lt;p&gt;If you gate AI-generated code with linters or a CI rule that hunts for swallowed errors, this experiment suggests the part you actually care about — "is this &lt;code&gt;return None&lt;/code&gt; a contract or a cover-up?" — is exactly the part the rule cannot decide.&lt;/p&gt;

&lt;p&gt;I started this project convinced that small LLMs routinely swallow failures — catch an error, return an empty value, pretend nothing happened — and that I'd measure the contamination rate. The hypothesis fell apart in a more interesting way than any clean number would have been. This article is the record of a prior being refuted by measurement, with the collapse documented step by step.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Setup: Qwen2.5-Coder 1.5B (Apache-2.0) via Ollama, CPU-only, Windows 11. 12 frozen tasks × 10 generations (seeds 0–9, temp 0.7) = **120 samples&lt;/em&gt;&lt;em&gt;: 100 failure-path functions (50 Python, 50 TypeScript) + 20 pure-computation controls. A 7B robustness footnote adds 30 more (150 generations total, all frozen unfiltered in the repo). Classification: Python by AST, TypeScript by regex; **all 120 labels eyeballed, every swallow candidate and boundary case hand-adjudicated&lt;/em&gt;* against docstrings, comments, and the function's contract, published as &lt;code&gt;gt.csv&lt;/code&gt;.*&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;A naive Semgrep detector — the kind you might reasonably drop into CI — flagged 4 candidates. Human adjudication: 0 true positives.&lt;/strong&gt; &lt;strong&gt;Correction (2026-09-16): that is only the precision half.&lt;/strong&gt; The corpus contains 4 samples adjudicated &lt;code&gt;problematic_fallback&lt;/code&gt; (&lt;code&gt;py_fetch_json_s2/s3/s7/s9&lt;/code&gt;), all of them &lt;code&gt;guard_default&lt;/code&gt;, and the shipped rule set reaches &lt;strong&gt;none&lt;/strong&gt; of them — &lt;strong&gt;recall 0/4&lt;/strong&gt; next to precision 0/4. &lt;strong&gt;Correction (2026-09-17): those two fours are different populations.&lt;/strong&gt; Precision 0/4 mixes languages (2 TypeScript false positives + 2 Python documented fallbacks); recall 0/4 is Python-only. Per language: &lt;strong&gt;Python 2 flagged / 0 TP / recall 0/4; TypeScript 2 flagged / 0 TP / recall 0/0&lt;/strong&gt; — the TypeScript rows contain no adjudicated problem, so TS recall is undefined, not zero. Read the pair together: "4 flagged, 0 true positives" alone can be read as &lt;em&gt;nothing was there to find&lt;/em&gt;, and the corpus says otherwise. Two were false positives (an empty &lt;code&gt;catch&lt;/code&gt; in a &lt;em&gt;usage example&lt;/em&gt; the model appended outside the function under test), two were legitimate fallbacks, documented (docstring or comment) as the function's contract. But this 0 is &lt;strong&gt;not&lt;/strong&gt; "AI doesn't swallow errors" — my detector targeted &lt;code&gt;try/except&lt;/code&gt; and structurally never looked at &lt;code&gt;if&lt;/code&gt;-guard default returns, which is where the suspicious shapes actually lived (see finding 3).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;In this sample, "fails loudly" and "no handling at all" dominated — not swallowing.&lt;/strong&gt; Of Python's 50 failure-path generations, &lt;strong&gt;39 (78%) used no try/except&lt;/strong&gt;: 14 &lt;code&gt;raise&lt;/code&gt; explicitly via &lt;code&gt;if/else&lt;/code&gt; (loud), 4 return &lt;code&gt;None&lt;/code&gt; behind an &lt;code&gt;if&lt;/code&gt; guard, 21 just let exceptions propagate. TypeScript went the other way: &lt;strong&gt;33/50 (66%) wrote try/catch with log and/or re-throw&lt;/strong&gt; — proper handling. The only belief that collapsed was my prior that swallowing would be &lt;em&gt;dominant&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The core finding: whether code "swallows errors" is not decidable from syntax.&lt;/strong&gt; The same &lt;code&gt;return None&lt;/code&gt; shows up as a documented, role-appropriate contract (&lt;code&gt;parse_int&lt;/code&gt;: None if it cannot be converted) &lt;em&gt;and&lt;/em&gt; as a hazard (&lt;code&gt;fetch_json&lt;/code&gt;: 404, 500 — every non-200 collapses into the same &lt;code&gt;None&lt;/code&gt;, while a network error raises instead, a third behavior the caller has to know about — with a comment saying so). A comment doesn't make it sound; a bare guard doesn't make it a bug. The deciding information — the function's role, the caller's expectations, the spec — &lt;strong&gt;lives outside the pattern&lt;/strong&gt;. Static rules can surface candidates; they cannot deliver the verdict.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Honest caveats up front:&lt;/strong&gt; one small model family, 12 tasks, N=10 per task (pseudo-replication — the independent unit is the task, not the generation), single-rater adjudication by me, and a detector whose scope hole I only noticed because the data rubbed my nose in it. Scope: free, local, small models, this corpus. No causal claims about "AI code vs human code" — that comparison belongs to prior work (AIRA, below), not to this experiment.&lt;/li&gt;
&lt;/ol&gt;

&lt;blockquote&gt;
&lt;p&gt;⚠ &lt;strong&gt;Reproducibility scope:&lt;/strong&gt; generation is nondeterministic and ran once; all 150 generations are frozen unfiltered in the repo. What you reproduce is the deterministic analysis layer — same 120 files in, same distributions and candidate counts out. The final legitimate-vs-swallow labels are &lt;em&gt;human adjudication&lt;/em&gt;, published transparently in &lt;code&gt;gt.csv&lt;/code&gt; — and the fact that you can't regenerate that layer mechanically is itself the thesis of this article.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Why measure this
&lt;/h2&gt;

&lt;p&gt;AI-generated code quality has a well-cited data point now: &lt;strong&gt;AIRA&lt;/strong&gt; (&lt;a href="https://arxiv.org/abs/2604.17587" rel="noopener noreferrer"&gt;arXiv:2604.17587&lt;/a&gt;, preprint, 2026) ran a deterministic, parser-backed static analyzer over 955 AI-authored and 955 human-authored code samples and reported &lt;strong&gt;1.80× more high-severity findings in AI-authored code&lt;/strong&gt; (0.435 vs 0.242 per sample), with &lt;strong&gt;Broad Exception Suppression (C03)&lt;/strong&gt; — the "swallowed error" family — as the most frequent check (263 vs 185 in Study 3).&lt;/p&gt;

&lt;p&gt;That number gets quoted as "AI code swallows errors." But read the paper closely and it says something more careful, twice:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Detection is deterministic and works — the paper even warns that an LLM evaluator &lt;em&gt;misses&lt;/em&gt; these suppressions at a rate of 44:1 compared to the deterministic scanner. Pattern-matching is the machine's strong suit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Flagged ≠ defect."&lt;/strong&gt; AIRA's own text states that some fail-soft patterns are contextually intentional and that human review is required before remediation. Two of its checks are permanently human-review-only.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the open question isn't "can we detect suppression patterns?" (yes, deterministically). It's the step everyone skips: &lt;strong&gt;can the &lt;em&gt;verdict&lt;/em&gt; — legitimate fallback or bug-hiding swallow — be automated too?&lt;/strong&gt; I couldn't find a published measurement of that gap (pointers welcome). This article measures it with 120 locally generated functions, a naive detector, and every candidate opened by hand.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Scope declaration up front: this is neither "AI is dangerous" nor "static analysis is useless." It's a measurement of where the machine's jurisdiction ends.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Terms, before the numbers
&lt;/h2&gt;

&lt;p&gt;Three things that look identical in grep output and must not be conflated:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Default return&lt;/strong&gt; — on failure, return &lt;code&gt;None&lt;/code&gt; / &lt;code&gt;0&lt;/code&gt; / &lt;code&gt;[]&lt;/code&gt; / &lt;code&gt;""&lt;/code&gt; instead of raising. A &lt;em&gt;syntactic shape&lt;/em&gt;. Detectable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Silent swallow&lt;/strong&gt; — a default return that erases information the caller needed: failure becomes indistinguishable from "empty but fine." A &lt;em&gt;semantic judgment&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Legitimate fallback&lt;/strong&gt; — a default return that &lt;em&gt;is&lt;/em&gt; the function's contract ("returns None if unparseable"), ideally documented. Also a semantic judgment.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The whole article is about the gap between the first item and the other two. Classification-wise this sits in &lt;strong&gt;CWE-703&lt;/strong&gt; (Improper Check or Handling of Exceptional Conditions), with CWE-1069 (empty exception block) and CWE-390 nearby; the exception-antipattern literature (&lt;a href="https://arxiv.org/abs/1704.00778" rel="noopener noreferrer"&gt;de Pádua &amp;amp; Shang, arXiv:1704.00778&lt;/a&gt;, Java/C#) maps the same territory.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why standard linters don't cover this (and what I built instead)
&lt;/h2&gt;

&lt;p&gt;Checked against current docs, not from memory:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Python / ruff:&lt;/strong&gt; &lt;code&gt;S110&lt;/code&gt; (try-except-pass) and &lt;code&gt;S112&lt;/code&gt; exist but are &lt;strong&gt;disabled by default&lt;/strong&gt; — you must opt in via &lt;code&gt;select&lt;/code&gt;/&lt;code&gt;extend-select&lt;/code&gt;. And a &lt;em&gt;typed&lt;/em&gt; swallow like &lt;code&gt;except Exception: return []&lt;/code&gt; sails through anyway.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TypeScript / ESLint:&lt;/strong&gt; &lt;code&gt;no-empty&lt;/code&gt; only fires on empty blocks, so &lt;code&gt;catch (e) { return null; }&lt;/code&gt; is invisible — it's not empty. &lt;code&gt;no-empty&lt;/code&gt; even allows a catch containing &lt;em&gt;only a comment&lt;/em&gt;. (Note the cultural assumption baked in there: "has a comment" is treated as "is intentional." Whether the comment describes &lt;em&gt;sound&lt;/em&gt; error handling is nobody's department. Hold that thought for the counterexample below.)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So I wrote a Semgrep rule that goes after the &lt;em&gt;meaning-shaped&lt;/em&gt; pattern — the naive "before" version, verbatim from the repo:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;rules&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;py-swallow-return-default&lt;/span&gt;
    &lt;span class="na"&gt;languages&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;python&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="na"&gt;severity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ERROR&lt;/span&gt;
    &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Exception&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;caught&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;default&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;value&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;returned&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;without&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;logging/re-raise&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;(silent&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;fallback)."&lt;/span&gt;
    &lt;span class="na"&gt;patterns&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;pattern&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
          &lt;span class="s"&gt;try:&lt;/span&gt;
            &lt;span class="s"&gt;...&lt;/span&gt;
          &lt;span class="s"&gt;except $E:&lt;/span&gt;
            &lt;span class="s"&gt;return $R&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;metavariable-pattern&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;metavariable&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;$R&lt;/span&gt;
          &lt;span class="na"&gt;patterns&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;pattern-either&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;pattern&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;None&lt;/span&gt;
                &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;pattern&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;[]"&lt;/span&gt;
                &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;pattern&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;{}"&lt;/span&gt;
                &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;pattern&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;0"&lt;/span&gt;
                &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;pattern&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;False"&lt;/span&gt;
                &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;pattern&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;""'&lt;/span&gt;

  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;py-except-pass&lt;/span&gt;
    &lt;span class="na"&gt;languages&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;python&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="na"&gt;severity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ERROR&lt;/span&gt;
    &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Exception&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;swallowed&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;with&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;pass&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;(silent&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;failure)."&lt;/span&gt;
    &lt;span class="na"&gt;patterns&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;pattern&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
          &lt;span class="s"&gt;try:&lt;/span&gt;
            &lt;span class="s"&gt;...&lt;/span&gt;
          &lt;span class="s"&gt;except $E:&lt;/span&gt;
            &lt;span class="s"&gt;pass&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(The TypeScript rules are the same shape — &lt;code&gt;catch { return null }&lt;/code&gt;, empty catch — and ship in the repo.)&lt;/p&gt;

&lt;p&gt;My hypothesis at this point: "run this over N generations, get a contamination rate." Here is where the measurement starts disagreeing with me.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setup
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Tasks:&lt;/strong&gt; 12, frozen in &lt;code&gt;tasks/tasks.json&lt;/code&gt; before generation. Per language: 5 failure-path I/O tasks (config loading, HTTP fetch, env vars, numeric parsing, file head) + 1 pure-computation control (mean/sum — no try needed). Prompts are neutral: no "handle errors" nudging.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Task-design disclosure (I did not farm for swallows):&lt;/strong&gt; pilot runs showed that on plain tasks ("read JSON, return it") the model writes no try/except and nothing interesting happens — which means &lt;em&gt;any&lt;/em&gt; "swallow rate" is heavily task-dependent and can be manufactured. I chose failure-inevitable I/O not to inflate the count (the result is ~0 anyway) but to give failure a stage — and kept the pure-computation controls to verify that off-stage, nothing appears. They did: &lt;strong&gt;all 20 controls showed zero try/except and zero candidates.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Generation:&lt;/strong&gt; &lt;a href="https://github.com/ollama/ollama" rel="noopener noreferrer"&gt;Ollama&lt;/a&gt; (MIT), CPU-only, &lt;code&gt;qwen2.5-coder:1.5b&lt;/code&gt; (Apache-2.0), temperature 0.7, &lt;code&gt;top_k=40 / top_p=0.9&lt;/code&gt;, seeds 0–9 per task, threads fixed at 4. The 7B model appears only in the robustness footnote (N=3 per task).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Determinism probe (before trusting anything):&lt;/strong&gt; at temperature 0 with a fixed thread count, the same seed generated twice was &lt;strong&gt;byte-identical&lt;/strong&gt; (SHA match). Distributions require temp &amp;gt; 0; hence 0.7.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Classifier asymmetry, disclosed:&lt;/strong&gt; Python is classified via &lt;code&gt;ast&lt;/code&gt; (a real parse); TypeScript via regex (not a full parser). Cross-language &lt;em&gt;absolute&lt;/em&gt; comparisons carry instrument bias — which is why every label was eyeballed and every candidate hand-adjudicated rather than trusting classifier output.&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Choice&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Inference&lt;/td&gt;
&lt;td&gt;Ollama, CPU (Windows 11)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model (main)&lt;/td&gt;
&lt;td&gt;Qwen2.5-Coder 1.5B, default GGUF quantization (Q4_K_M-class)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model (footnote)&lt;/td&gt;
&lt;td&gt;Qwen2.5-Coder 7B&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Detector&lt;/td&gt;
&lt;td&gt;Semgrep 1.168.0 (CE) + custom rules above&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Analysis&lt;/td&gt;
&lt;td&gt;Python 3.12, ruff 0.15.12 for the lint baseline&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Measured: 2026-07 (corpus generated 2026-07-01). Full per-run metadata in the repo's PROVENANCE file.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 1: mostly not swallowing — passing through, or failing loudly
&lt;/h2&gt;

&lt;p&gt;How the 50 failure-path generations per language handled failure:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Failure-path tasks (n=50 per language)&lt;/th&gt;
&lt;th&gt;Python&lt;/th&gt;
&lt;th&gt;TypeScript&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;try/except + log or re-raise (proper)&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;33 (66%)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;try/except returning a default (swallow &lt;em&gt;candidate&lt;/em&gt;)&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;no try/except at all&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;39 (78%)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;17 (34%)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Reading "no try/except = no handling" would be wrong. AST-splitting Python's 39:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;14 raise via &lt;code&gt;if/else&lt;/code&gt;&lt;/strong&gt; (e.g. missing env var → &lt;code&gt;raise ValueError(...)&lt;/code&gt;) — that's failing &lt;em&gt;loudly&lt;/em&gt;, the opposite of swallowing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;4 return &lt;code&gt;None&lt;/code&gt; behind an &lt;code&gt;if&lt;/code&gt; guard&lt;/strong&gt; — all four in &lt;code&gt;fetch_json&lt;/code&gt; (non-200 → &lt;code&gt;return None&lt;/code&gt;). A default return! Which my try/except-scoped detector &lt;strong&gt;structurally never saw.&lt;/strong&gt; Remember these four; they're the article's best specimen.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;21 bare&lt;/strong&gt; — no failure path written; exceptions propagate to the caller.&lt;/li&gt;
&lt;li&gt;(Check: 14 + 4 + 21 = 39.)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;TypeScript's 66% try/catch majority mostly did &lt;code&gt;console.error(...)&lt;/code&gt; + &lt;code&gt;throw&lt;/code&gt; — textbook handling, no swallowing.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The language difference (Python avoids try/except, TS writes it) is an &lt;strong&gt;observation, not a finding&lt;/strong&gt;: prompt phrasing, language idiom (async/await + try/catch is TS boilerplate), and the AST-vs-regex classifier asymmetry all confound it. The spine of this article is the counterexample below, not this table.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So the original hypothesis — "AI swallows failures at some rate N% I can report" — collapsed in the first table: in this sample, swallowing wasn't the dominant behavior at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 2: every detector hit was a false positive or documented-legitimate
&lt;/h2&gt;

&lt;p&gt;The naive Semgrep rules flagged &lt;strong&gt;4 of 120&lt;/strong&gt; (2 Python, 2 TypeScript). Opening all four by hand:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TypeScript, 2 hits (&lt;code&gt;ts_load_config&lt;/code&gt;) = false positives.&lt;/strong&gt; The function under test was exemplary (code condensed and annotated from the corpus):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;loadConfig&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kr"&gt;any&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;fs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;promises&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;readFile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;utf8&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`Error reading or parsing the file at &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;path&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// logged and re-thrown -- not swallowed (proper)&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="c1"&gt;// ...but the model appended a usage example after the function:&lt;/span&gt;
&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;config&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;loadConfig&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;./config.json&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="cm"&gt;/* Handle any errors... &amp;lt;- comment-only catch */&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;})();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The rule fired on the &lt;em&gt;demo block's&lt;/em&gt; catch — (a) outside the function under test, (b) "empty" only because Semgrep's AST ignores comments. Context makes it an obvious false positive; the pattern alone can't know that.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Python, 2 hits (&lt;code&gt;py_parse_int&lt;/code&gt;) = documented, legitimate fallbacks.&lt;/strong&gt; The docstring states the contract — None if it cannot be converted (the second one says the same in a comment).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bottom line: undisclosed swallowing inside try/except = 0 — and detected problems in the corpus = 0 out of 4 (Python; the TypeScript rows contain no adjudicated problem, so TS recall is 0/0).&lt;/strong&gt; The first number is precision; the second is recall, and the article originally printed only the first. And now the honest part: my detector's own scope hole. The rules target &lt;code&gt;try/except&lt;/code&gt; — the &lt;strong&gt;&lt;code&gt;if&lt;/code&gt;-guard default returns (the four &lt;code&gt;fetch_json&lt;/code&gt;s) were never in scope&lt;/strong&gt;. So the corpus &lt;em&gt;does&lt;/em&gt; contain default returns; the accurate claim is "zero undisclosed swallows &lt;em&gt;within the detector's scope&lt;/em&gt;," not "zero problematic fallbacks in the corpus." Whether those four are problems is precisely the question syntax can't answer — next section.&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 3 (the core): the same &lt;code&gt;return None&lt;/code&gt;, and what separates them isn't in the code
&lt;/h2&gt;

&lt;p&gt;Two functions from the corpus (condensed and annotated). Which one swallows errors?&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# (A) numeric parsing: None if unconvertible (contract stated in the docstring)
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;parse_int_field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Returns the int, or None if it cannot be converted.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;            &lt;span class="c1"&gt;# role-appropriate fallback
&lt;/span&gt;
&lt;span class="c1"&gt;# (B) HTTP fetch: None on non-200 (and yes, there's a comment saying so)
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;fetch_json&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status_code&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;            &lt;span class="c1"&gt;# 404 or 500 -- every non-200 flattened into None
&lt;/span&gt;                               &lt;span class="c1"&gt;# (a network error raises instead: a third behavior)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Syntactically, the interesting part is near-identical: &lt;em&gt;on failure, return &lt;code&gt;None&lt;/code&gt;&lt;/em&gt;. Semantically they're opposites. (A) matches the function's role — "tell me whether this converts" — so &lt;code&gt;None&lt;/code&gt; &lt;em&gt;is&lt;/em&gt; the answer. (B) permanently destroys the caller's ability to distinguish "no data" from "the fetch failed." And (B) &lt;strong&gt;has a comment&lt;/strong&gt;. Documentation doesn't settle it: a documented &lt;code&gt;return 0&lt;/code&gt; or &lt;code&gt;return None&lt;/code&gt; can still silently poison every computation downstream. ("Has a comment = intentional" is exactly the assumption ESLint's &lt;code&gt;no-empty&lt;/code&gt; institutionalizes.)&lt;/p&gt;

&lt;p&gt;(B) is the strongest specimen this experiment produced: a &lt;em&gt;documented-but-hazardous&lt;/em&gt; default return that appeared as a &lt;strong&gt;4-sample cluster&lt;/strong&gt; (not a one-off), sitting squarely in the blind spot of a try/except-scoped detector — while being exactly the "quietly fails" shape the AIRA numbers gesture at, at population level.&lt;/p&gt;

&lt;p&gt;For contrast, non-swallowing code carries its intent &lt;em&gt;inside&lt;/em&gt; the syntax:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_api_token&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;API_TOKEN&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;API_TOKEN&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The API_TOKEN environment variable is not set.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# fails loudly
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The point, stated carefully: &lt;strong&gt;syntactic patterns can surface candidates. The information that separates contract from cover-up — the function's role, the caller's expectations, the spec, and whether the documentation is &lt;em&gt;right&lt;/em&gt; — lives outside the pattern.&lt;/strong&gt; I'm not claiming "undecidable in principle": smarter analysis (types, dataflow, call-site analysis) absolutely narrows the candidates. But "what should this function return, in this context?" is a spec question, and the spec comes from outside the analyzer. Sharper tools shrink the pile; the final reconciliation against intent remains.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Honest scope note:&lt;/strong&gt; this rests on a small corpus and a handful of specimens — read it as a &lt;em&gt;demonstrated boundary&lt;/em&gt;, not a general law about all static analysis. What it demonstrates survives the small N, though, because it's an existence proof: two same-shaped snippets with opposite verdicts, and the verdict-relevant information demonstrably outside the syntax.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;What I actually do now:&lt;/strong&gt; machines generate the candidate list (patterns, distributions, CI notifications — &lt;code&gt;semgrep scan --error&lt;/code&gt; exits 1 on candidates, which is fine); a human adjudicates candidates against role, caller, and spec. A green scan is read as "&lt;strong&gt;zero candidates for a human to look at&lt;/strong&gt;," never as "pass."&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this lands relative to AIRA
&lt;/h2&gt;

&lt;p&gt;Neatly inside its own fine print, it turns out. AIRA detects suppression patterns deterministically and reports the 1.80× population-level difference — &lt;strong&gt;candidate generation is the machine's win&lt;/strong&gt;, and the paper warns that semantic (LLM) evaluation &lt;em&gt;underperforms&lt;/em&gt; there, 44:1. But the same paper marks the boundary: flagged patterns aren't necessarily defects, some fail-softs are intentional, human review precedes remediation, and two checks are permanently human-only.&lt;/p&gt;

&lt;p&gt;That division of labor is what this experiment probed at specimen level:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Machines: sensitivity.&lt;/strong&gt; Surface default-return shapes, chart try/except usage, gate CI with "candidates exist."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Humans: adjudication.&lt;/strong&gt; Contract or cover-up — read the role, the caller, the spec.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One misquote I need to preempt, because I nearly published it myself (see pitfall 5): AIRA is &lt;strong&gt;not&lt;/strong&gt; a human-evaluation study, and I am not claiming AI code swallows more than human code — that comparison is AIRA's, made with its own methodology, at population level, in a single-author preprint. This experiment neither confirms nor contradicts it; it maps the adjudication residue the paper explicitly leaves to humans.&lt;/p&gt;

&lt;h2&gt;
  
  
  Everything I got wrong
&lt;/h2&gt;

&lt;p&gt;In the order I got it wrong:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;"AI swallows failures; my rule will measure the rate."&lt;/strong&gt; → Measured: in this sample, swallowing wasn't dominant — pass-through, loud failure, and (in TS) proper handling were. &lt;em&gt;The contamination-rate framing died on the first table.&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"The detector's 4 hits = 4 swallows."&lt;/strong&gt; → All four were false positives or documented-legitimate. The honest number is 0, so I report 0 — with its scope. &lt;em&gt;Don't round adjudication up to detection.&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"The false positives just mean my rule is crude; refinement will fix it."&lt;/strong&gt; → Refinement did kill the TS false positives. It could not touch the real problem: whether a &lt;code&gt;return None&lt;/code&gt; is legitimate isn't decided by &lt;em&gt;any&lt;/em&gt; pattern, or even by the presence of documentation. This is where the thesis flipped from "measure the rate" to "the verdict doesn't live in syntax." And auditing my own instrument surfaced the scope hole (&lt;code&gt;if&lt;/code&gt;-guard defaults) it had from the start. &lt;em&gt;When your tool's errors persist under refinement, suspect the question, not the regex.&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Windows encoding landmine.&lt;/strong&gt; Semgrep read my TS rule file under Windows' legacy default codepage (cp932) and crashed with &lt;code&gt;UnicodeDecodeError&lt;/code&gt; (exit 2) — on an em dash I'd left in a rule &lt;code&gt;message&lt;/code&gt;. Fix: ASCII-only rule files + &lt;code&gt;PYTHONUTF8=1&lt;/code&gt; for every run. &lt;em&gt;If your pipeline runs on Windows, pin UTF-8 explicitly or non-ASCII bytes in innocent places will cost you an afternoon.&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Writing this article, I did the thing the article warns about.&lt;/strong&gt; My research notes — AI-assisted — said AIRA judged suppressions "by human evaluation." Pre-publication, I went back to the primary source: it's the &lt;em&gt;opposite&lt;/em&gt;. AIRA is a deterministic scanner, and it &lt;em&gt;warns against&lt;/em&gt; relying on semantic evaluation for this class. I had a plausible summary, I trusted it, and a "verification discipline" article nearly shipped misstating its key citation's methodology. Corrected against the paper itself. &lt;em&gt;Candidates can come from machines — or from your own notes. The final check against the primary source is the human step, and I almost skipped it.&lt;/em&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Limitations
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;One model family, small models (1.5B main / 7B footnote), 12 tasks, one prompt style. Behavior is strongly model- and task-dependent; no generalization claimed beyond this corpus.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Statistical independence:&lt;/strong&gt; N=10 generations per prompt are correlated pseudo-replicates. The independent unit is the task (5 per language), not the generation; don't read the n=50 tables as population estimates. The thesis rides on the counterexample, not the ratios.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Detector scope:&lt;/strong&gt; rules and classifiers target try/except; &lt;code&gt;if&lt;/code&gt;-guard defaults are out of scope, so "0 swallows" means "0 undisclosed swallows in scope" — the corpus contains default returns whose adjudication is the whole point.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Promise handlers (added 2026-09-20):&lt;/strong&gt; promise &lt;code&gt;.catch&lt;/code&gt; handlers are outside every pattern in the TypeScript rule file, braces or not — every pattern is anchored on a &lt;code&gt;try { ... } catch&lt;/code&gt; clause. In this corpus that is 3 of 48 catch-shaped sites (45 &lt;code&gt;try { ... } catch&lt;/code&gt; clauses plus 3 promise &lt;code&gt;.catch&lt;/code&gt; handlers): &lt;code&gt;ts_fetch_json_s0&lt;/code&gt; and &lt;code&gt;ts_load_config_s7&lt;/code&gt; block-form, &lt;code&gt;ts_fetch_json_s8&lt;/code&gt; expression-form. All three &lt;code&gt;console.error&lt;/code&gt; with no return, so 0 of 3 swallow silently; in each file the function's own catch logs and rethrows while the handler logs and stops, so these three are also the corpus's report-without-propagate sites. The numbers above do not move. &lt;code&gt;promise.catch(() =&amp;gt; null)&lt;/code&gt; — the form the corpus does not contain — is what the rule would miss. Raised by a reader, 2026-09-18.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Artefact gap (added 2026-09-16):&lt;/strong&gt; in &lt;code&gt;results/gt.csv&lt;/code&gt; the shape columns &lt;code&gt;subtype&lt;/code&gt; and &lt;code&gt;returns_default&lt;/code&gt; are filled for the Python rows only — &lt;strong&gt;0 of 60 TypeScript rows&lt;/strong&gt; carry them (their &lt;code&gt;human_verdict&lt;/code&gt; is filled: 31 proper / 27 not_adjudicated / 2 false_positive). Anyone recomputing the cross-tab per language hits 27 unlabelled rows. Raised by a reader who re-derived the analysis from the repo. &lt;strong&gt;Partly closed 2026-09-17:&lt;/strong&gt; for Python rows that have a handler, &lt;code&gt;subtype&lt;/code&gt; is no longer blank — it is now &lt;code&gt;guard_default&lt;/code&gt; / &lt;code&gt;handler_default&lt;/code&gt; / &lt;code&gt;both&lt;/code&gt; by where the default return sits (commit &lt;code&gt;82e5130&lt;/code&gt;), so the guard set (&lt;code&gt;subtype in {guard_default, both}&lt;/code&gt;, 17 rows) is computable from &lt;code&gt;gt.csv&lt;/code&gt; alone. Later the same day (commit &lt;code&gt;f839873&lt;/code&gt;) the TypeScript rows got the same three shape columns via tree-sitter, with &lt;code&gt;handling&lt;/code&gt; left as the regex verdict the tables above were computed with (the AST agrees on 58/60; the other 2 are the Semgrep false positives). What that exposes rather than closes: the TypeScript guard set is 11 rows (&lt;code&gt;ts_get_item_s1-s9&lt;/code&gt;, &lt;code&gt;ts_parse_int_s0/s5&lt;/code&gt;) and &lt;strong&gt;none of them has been human-adjudicated&lt;/strong&gt; — their &lt;code&gt;proper&lt;/code&gt; comes from the catch logging or throwing, not from anyone reading the default return outside it. TS recall stays 0/0 for that reason, not because the shape is absent — and, more to the point (reader, 2026-09-17), because the TypeScript half has no positive class by construction: the two halves share only 2 of 6 task families, and on &lt;code&gt;fetch_json&lt;/code&gt;, the one family that carries all four Python problems, TypeScript is 10/10 &lt;code&gt;proper&lt;/code&gt;. Read the TS half as a control arm, not a recall arm. Filling &lt;code&gt;has_raise&lt;/code&gt; also moved the four throw-only TS samples from &lt;code&gt;not_adjudicated&lt;/code&gt; to &lt;code&gt;loud_fail&lt;/code&gt; (same machine rule as Python), so the TS split above now reads 31 proper / 23 not_adjudicated / 4 loud_fail / 2 false_positive.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Classifier asymmetry:&lt;/strong&gt; Python AST vs TypeScript regex (nested-brace catches can slip). Human-verified labels backstop it, but discount cross-language comparisons.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;7B robustness footnote (N=3/task, 30 generations):&lt;/strong&gt; Python 10/15 no-try, 3 proper, 2 candidates; TS 8/15 no-try, 7 proper, 0 candidates. Both candidates: &lt;code&gt;parse_int&lt;/code&gt; again, documented — one returns &lt;code&gt;0&lt;/code&gt; on failure, which is documented &lt;em&gt;and&lt;/em&gt; still capable of silently corrupting downstream arithmetic, i.e., the "(B) problem" one more time. N=3, so no robustness claim — only "not refuted in this range."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The adjudicator's hole, disclosed next to the detector's:&lt;/strong&gt; the legitimate-vs-swallow verdicts are single-rater (me), no inter-rater agreement measured. Rubric: (1) does failure yield a default? (2) is that behavior disclosed as intent? (3) does it erase failure information the caller needs? Boundary cases ship with reasons in &lt;code&gt;gt.csv&lt;/code&gt;. A second rater is future work — "the last step is human" cuts both ways, so I'm flagging my own last step.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Conflict of interest:&lt;/strong&gt; I do AI-code auditing professionally. "The final call needs a human" is a conclusion that favors my line of work. Mitigation is transparency: frozen tasks, unfiltered generations, published adjudications, and the detector's failures disclosed by me, above.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Reproduce it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/sumitsuke/ai-silent-defect-scanner &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd &lt;/span&gt;ai-silent-defect-scanner
&lt;span class="c"&gt;# Deterministic layer: same 120 files -&amp;gt; same distributions &amp;amp; candidate counts&lt;/span&gt;
&lt;span class="nv"&gt;PYTHONUTF8&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 python scripts/classify_split.py      &lt;span class="c"&gt;# failure-path distributions (the n=50 tables)&lt;/span&gt;
&lt;span class="nv"&gt;PYTHONUTF8&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 python scripts/scan_and_count.py      &lt;span class="c"&gt;# naive Semgrep candidates (the 4 hits)&lt;/span&gt;
&lt;span class="nv"&gt;PYTHONUTF8&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 python scripts/build_gt.py            &lt;span class="c"&gt;# regenerate the adjudication layer, results/gt.csv&lt;/span&gt;
make scan-mine &lt;span class="nv"&gt;DIR&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/path/to/your/repo              &lt;span class="c"&gt;# candidates for YOUR repo (verdicts are on you)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;Mechanical reproduction covers distributions and candidate counts. The final labels are my adjudications, published in &lt;code&gt;results/gt.csv&lt;/code&gt; with reasons — by design not machine-rederivable (that's the thesis).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;raw/&lt;/code&gt; ships all 150 generations unfiltered; &lt;code&gt;PROVENANCE&lt;/code&gt; records model, quantization, seeds, temperature, thread count, OS, dates. Regenerating gives you a different distribution — that's LLM sampling for you.&lt;/li&gt;
&lt;li&gt;CPU-only, no GPU needed, no paid APIs anywhere in the pipeline.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Pattern rules find &lt;em&gt;shapes&lt;/em&gt;, and the corpus's actual risk shape (&lt;code&gt;if&lt;/code&gt;-guard default returns) wasn't the shape my rule watched.&lt;/strong&gt; Budget for your detector's scope hole, not just its false-positive rate. &lt;strong&gt;Correction (2026-09-16):&lt;/strong&gt; this takeaway originally covered two different boundaries, and only one of them is a defect. (a) The &lt;em&gt;try/except scope&lt;/em&gt; is a hole — it costs recall: 100% of the corpus's adjudicated problems sit outside it. (b) The &lt;em&gt;one-statement anchor&lt;/em&gt; (handler body is exactly &lt;code&gt;return &amp;lt;default&amp;gt;&lt;/code&gt;) is a filter that buys precision — loosening it adds five samples (&lt;code&gt;py_fetch_json_s6&lt;/code&gt;, &lt;code&gt;py_first_line_s2&lt;/code&gt;, &lt;code&gt;py_load_json_s4&lt;/code&gt;, &lt;code&gt;py_load_json_s8&lt;/code&gt;, &lt;code&gt;py_parse_int_s2&lt;/code&gt;), and all five are &lt;code&gt;has_log=Y&lt;/code&gt; and adjudicated &lt;code&gt;proper&lt;/code&gt;, 5/5. Keep (b), close (a).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Documentation is not adjudication.&lt;/strong&gt; A commented &lt;code&gt;return None&lt;/code&gt; and a &lt;code&gt;return 0&lt;/code&gt; documented in its docstring both appeared here as &lt;em&gt;disclosed and still hazardous&lt;/em&gt;. "Has a comment" is a lint convention, not a verdict.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Split the jurisdictions and staff them accordingly:&lt;/strong&gt; machine = candidates and distributions in CI; human = contract-vs-cover-up against role, caller, and spec. Green means "no candidates today," not "correct."&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you remember one sentence, please don't make it "static analysis is useless" — the detector did its detection job fine. Make it: &lt;strong&gt;syntax can surface the suspects, but conviction requires reading intent — and intent isn't stored in the AST.&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Detection code, all 150 generations, and the human adjudication layer (&lt;code&gt;gt.csv&lt;/code&gt;): &lt;a href="https://github.com/sumitsuke/ai-silent-defect-scanner" rel="noopener noreferrer"&gt;github.com/sumitsuke/ai-silent-defect-scanner&lt;/a&gt;. Every number above is measured; anything not measured is labeled as not measured.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This is an English adaptation of &lt;a href="https://qiita.com/sumitsuke/items/dc2a839a8f3618d0da12" rel="noopener noreferrer"&gt;my Japanese article on Qiita&lt;/a&gt; (Qiita is a Japanese dev-blogging platform) — written by me in Japanese, restructured and translated with AI assistance, human-reviewed. If you spot an error, comments and issues are open; I'll verify against the frozen corpus and correct with a changelog.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Verification record (environment, verdict, last verified date, evidence) and the canonical write-up: &lt;a href="https://sumitsuke.jp/via/devto/lab/ai-code-silent-fallback/" rel="noopener noreferrer"&gt;https://sumitsuke.jp/via/devto/lab/ai-code-silent-fallback/&lt;/a&gt; — code, data and reproduction: &lt;a href="https://github.com/sumitsuke/ai-silent-defect-scanner" rel="noopener noreferrer"&gt;https://github.com/sumitsuke/ai-silent-defect-scanner&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  Corrections
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;2026-09-16&lt;/strong&gt; — After publication, a reader re-derived the deterministic layer from the repo and pointed out that the report card was missing its recall half. I recounted from &lt;code&gt;results/gt.csv&lt;/code&gt; and confirmed every checkable claim: 4 adjudicated &lt;code&gt;problematic_fallback&lt;/code&gt; samples, all &lt;code&gt;guard_default&lt;/code&gt;, none reachable by the shipped rules (recall 0/4); 9 of the 20 controls carry a default return; the five samples added by loosening the one-statement anchor are 5/5 &lt;code&gt;has_log=Y&lt;/code&gt; and &lt;code&gt;proper&lt;/code&gt;; the TypeScript rows carry no shape labels. One number I could not reproduce from the labels alone is their precision 4/17 — from &lt;code&gt;gt.csv&lt;/code&gt; I reach 13 (&lt;code&gt;returns_default=Y AND handling=no_try_except&lt;/code&gt;); the other four would have to come from reading the source. Changes above: the recall row in the TL;DR and in Result 2, the split of scope-hole vs anchor in Takeaway 1, and the artefact gap in Limitations. Nothing in the measured data changed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2026-09-17&lt;/strong&gt; — The same reader came back with two things. First, a correction of their own: "precision 0/4, recall 0/4" reads like one 2×2 but is not — the 4 flagged are 2 TypeScript + 2 Python, the 4 problems are all Python, and TypeScript has no adjudicated problem at all (recall 0/0). I had printed the pair the same way; language is now attached in the TL;DR and Result 2. Second, the 4/17 I said I could not reproduce from the labels: they posted a 30-line AST detector (&lt;code&gt;Return&lt;/code&gt; of a default value outside any &lt;code&gt;except&lt;/code&gt; handler), and I ran it against &lt;code&gt;raw/python/&lt;/code&gt; — it fires on exactly 17 files, the four "extra" ones are &lt;code&gt;py_fetch_json_s6:13&lt;/code&gt;, &lt;code&gt;py_parse_int_s3:22&lt;/code&gt;, &lt;code&gt;py_parse_int_s6:14&lt;/code&gt;, &lt;code&gt;py_parse_int_s9:12&lt;/code&gt;, line numbers as they stated, and the verdicts on the 17 are 4 &lt;code&gt;problematic_fallback&lt;/code&gt; / 2 &lt;code&gt;legit_fallback&lt;/code&gt; / 2 &lt;code&gt;proper&lt;/code&gt; / 9 controls (&lt;code&gt;not_adjudicated&lt;/code&gt;) — recall 4/4, precision 4/17. The reason the labels could not reach those four is that &lt;code&gt;gt.csv&lt;/code&gt; recorded &lt;em&gt;whether&lt;/em&gt; a default return exists but not &lt;em&gt;where&lt;/em&gt;; their proposed column closes it, and &lt;code&gt;build_gt.py&lt;/code&gt; now writes &lt;code&gt;guard_default&lt;/code&gt; / &lt;code&gt;handler_default&lt;/code&gt; / &lt;code&gt;both&lt;/code&gt; for every Python row that has a handler (of the four: &lt;code&gt;py_fetch_json_s6&lt;/code&gt; and &lt;code&gt;py_parse_int_s3/s9&lt;/code&gt; are &lt;code&gt;both&lt;/code&gt;, &lt;code&gt;py_parse_int_s6&lt;/code&gt; is &lt;code&gt;guard_default&lt;/code&gt; — its handler raises, and its only default return is in the &lt;code&gt;else&lt;/code&gt; branch). Nothing in the measured data changed; one column was added — and, by evening, extended to the TypeScript rows (&lt;code&gt;f839873&lt;/code&gt;), which turned the TS artefact gap from "no shape columns" into "11 guard-default samples nobody has adjudicated" (Limitations).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2026-09-17, later&lt;/strong&gt; — The reader re-ran both fixes against the source (Python 17/60 and the four line numbers reproduce; the TypeScript guard set of 11 reproduces exactly) and added two things I had not stated. First, the "4 flagged" row is not one detector's output: the two Python hits come from &lt;code&gt;py-swallow-return-default&lt;/code&gt;, the two TypeScript hits from &lt;code&gt;ts-empty-catch&lt;/code&gt; (I confirmed this from &lt;code&gt;scan_and_count.py&lt;/code&gt;'s hit list); the return-default rule fires on 0 TypeScript files, because its one-statement anchor excludes all 9 TypeScript catches that log first — every one of them is &lt;code&gt;console.error(...)&lt;/code&gt; then &lt;code&gt;return null&lt;/code&gt;/&lt;code&gt;undefined&lt;/code&gt;, and every one is &lt;code&gt;proper&lt;/code&gt;. Second, the TypeScript half is best read as a control arm: it shares only two tasks with Python, and the task that carries all four Python defects (fetch_json) is 10/10 proper in TypeScript. (I originally wrote here that the TypeScript half "has no positive class by construction of the task sample" — withdrawn on 2026-09-18, see below; whether it has a positive class is an adjudication outcome, not a property of the sample.) The two outstanding rulings on &lt;code&gt;ts_parse_int_s0/s5&lt;/code&gt; remained open at this point. The reader asked to stay unnamed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2026-09-18&lt;/strong&gt; — The same reader's third round. Two things. (1) The sentence "no positive class by construction of the task sample" was too strong and is withdrawn: the guard filter fires on the same two shapes in both halves; TypeScript's guard rows are 9 &lt;code&gt;proper&lt;/code&gt; (machine label: the catch logs or throws) and 2 that were open. (2) I ruled the two open rows, &lt;code&gt;ts_parse_int_s0/s5&lt;/code&gt;: no catch in either file, and the docstrings state the default ("it returns null" / "otherwise undefined") — the same rule that made &lt;code&gt;py_parse_int_s3/s9&lt;/code&gt; &lt;code&gt;legit_fallback&lt;/code&gt;, so both are &lt;code&gt;legit_fallback&lt;/code&gt; (commit in the repo, &lt;code&gt;review_note&lt;/code&gt; filled). The TypeScript recall line therefore still has no positive to count, but the honest reading is now "0 of 2 opened, both documented", not "0/0, nothing to open". The Python half is not fully adjudicated either: with &lt;code&gt;subtype ∈ {guard_default, both}&lt;/code&gt; and &lt;code&gt;returns_default = Y&lt;/code&gt; it is 17 rows — 4 problematic / 2 proper / 2 legit / 9 not_adjudicated. One more correction of my own number in the comments: "49 of 60 TypeScript files end with a demo block" was the wrong reading (49 = files with a &lt;code&gt;console.*&lt;/code&gt; call outside any catch); files whose last top-level statement is a &lt;code&gt;console.*&lt;/code&gt; call are 22/60.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2026-09-20&lt;/strong&gt; — The same reader's fourth round: a scope hole the brace-shaped census could not see. Promise &lt;code&gt;.catch&lt;/code&gt; handlers match nothing in the TypeScript rule file, with or without braces; in this corpus that is 3 of 48 catch-shaped sites (45 try/catch clauses plus 3 promise handlers), and all three are &lt;code&gt;console.error&lt;/code&gt; with no return (0 of 3 swallow silently; all three log and stop where the function's own catch rethrows), so no number above changes. Stated in Limitations now, next to the &lt;code&gt;if&lt;/code&gt;-guard hole, with &lt;code&gt;promise.catch(() =&amp;gt; null)&lt;/code&gt; as the one-line form the rule would miss.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2026-09-21&lt;/strong&gt; — Same reader, fifth round: the 48 had no instrument behind it and reads as 45 to anyone recomputing from the rule's anchor. It is now stated as 45 try/catch clauses plus 3 promise handlers, in Limitations and here; and the three handlers are named as what they are in this corpus — the report-without-propagate sites, where the function's own catch rethrows and the call-site handler logs and stops. No number moves.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>staticanalysis</category>
      <category>python</category>
    </item>
    <item>
      <title>LLM self-review vs. a blind re-run on 160 gold-labeled name-matching pairs — no measurable difference at ceiling</title>
      <dc:creator>Sumitsuke</dc:creator>
      <pubDate>Sun, 12 Jul 2026 01:43:06 +0000</pubDate>
      <link>https://dev.to/tauridev/llm-self-review-vs-a-blind-re-run-on-160-gold-labeled-name-matching-pairs-no-measurable-14c2</link>
      <guid>https://dev.to/tauridev/llm-self-review-vs-a-blind-re-run-on-160-gold-labeled-name-matching-pairs-no-measurable-14c2</guid>
      <description>&lt;p&gt;If you ship anything an LLM produced and calm your nerves by asking the &lt;em&gt;same&lt;/em&gt; model to "please double-check it," this is a small, pre-registered measurement of whether that actually buys you anything.&lt;/p&gt;

&lt;p&gt;Short version: on this task it bought nothing measurable — but mostly because the task turned out to be too easy, which is a lesson in itself about evaluation design.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Setup: &lt;code&gt;gpt-5.4-mini&lt;/code&gt;, 160 gold-labeled "are these two records the same person?" pairs, pre-registered before the run.&lt;/em&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The first pass was near-ceiling: 159/160 overall, 120/120 on the confirmation set.&lt;/strong&gt; So there was almost nothing left to fix — I ended up measuring "does re-review &lt;em&gt;preserve&lt;/em&gt; correct answers," not "does it &lt;em&gt;fix&lt;/em&gt; errors."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Re-review (shown its own draft) tied a blind re-run.&lt;/strong&gt; Across all 160 pairs both scored 156/160. On the pre-registered confirmation set, the single discordant pair gives exact McNemar p = 1.0 — no power to tell them apart, not proof of equivalence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reversals of correct answers happened in &lt;em&gt;both&lt;/em&gt; conditions, equally (3 and 3).&lt;/strong&gt; Not specific to self-review; consistent with plain run-to-run variation (single runs, so I can't causally isolate the extra pass itself).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Honest caveat:&lt;/strong&gt; with the first pass this high, I could not test the thing I actually wanted to test. Treat this as the ceiling data point of a series, not a verdict on self-correction.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Reproducibility scope
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;This is one model, one prompt, &lt;code&gt;reasoning_effort=low&lt;/code&gt;, single runs in a fixed order, on an 80/80 balanced diagnostic benchmark (60/60 confirmation + 20/20 hard). Nothing here generalizes to "LLM self-correction" in general. It's a concrete, reproducible measurement on one task, with all raw judgments and the pre-registration public.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Why measure this
&lt;/h2&gt;

&lt;p&gt;"Self-correction" reports point both ways. Self-Refine (Madaan et al. 2023) reports large gains from iterative self-feedback; Huang et al. 2024 ("LLMs Cannot Self-Correct Reasoning Yet") and Kamoi et al. 2024 (a critical cross-task survey) find that without a reliable external signal, self-correction often doesn't help and can degrade. Those negative results are mostly on reasoning tasks, so I treat any agreement here as a cross-genre extrapolation, not a replication.&lt;/p&gt;

&lt;p&gt;The practical version of the question: an LLM says "these two records are the same person / different people." Does asking the same model to re-check make it safer? Entity matching (deciding whether two &lt;em&gt;given&lt;/em&gt; records refer to one entity) is easy to build gold labels for, so you can measure this without subjective judgment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setup
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Task: match a Japanese footballer's Japanese name (kanji/kana) against a romaji name — e.g. is &lt;code&gt;久保建英&lt;/code&gt; the same person as &lt;code&gt;Takefusa Kubo&lt;/code&gt;?&lt;/li&gt;
&lt;li&gt;Model: &lt;code&gt;gpt-5.4-mini-2026-03-17&lt;/code&gt;, &lt;code&gt;reasoning_effort=low&lt;/code&gt;. &lt;strong&gt;Measured: 2026-07.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;160 pairs: a confirmation set of 120 (60 true-match / 60 non-match) plus 40 hard pairs.&lt;/li&gt;
&lt;li&gt;Surface forms come from &lt;strong&gt;different sources&lt;/strong&gt; (Japanese side = Wikidata, romaji side = a Transfermarkt-derived dataset), so club-name spellings and coverage don't line up perfectly.&lt;/li&gt;
&lt;li&gt;The model never sees dates of birth, IDs, or URLs — only names and club history. Gold is verified against an independent attribute (DOB), not the linking ID.&lt;/li&gt;
&lt;li&gt;Cost: 480 calls (160 × A/B/C), ~$0.45 total.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How I measured (pre-registered)
&lt;/h2&gt;

&lt;p&gt;Three conditions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A — first pass:&lt;/strong&gt; judge the two records.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;B — blind re-run:&lt;/strong&gt; the byte-identical prompt, fresh context, no draft shown. ("Independent" here means &lt;em&gt;A not shown&lt;/em&gt;, not statistical independence.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;C — two-stage review:&lt;/strong&gt; the original data &lt;strong&gt;plus A's structured output&lt;/strong&gt; as a draft, with a neutral "KEEP or CHANGE" instruction.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The pre-registered primary comparison is &lt;strong&gt;C vs B on the 120-pair confirmation set&lt;/strong&gt; (balanced accuracy difference + a paired interval). C vs B measures the &lt;em&gt;whole&lt;/em&gt; "hand the model its own output and ask it to re-review" workflow, not the anchoring effect in isolation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. The first pass was already at the ceiling
&lt;/h3&gt;

&lt;p&gt;A scored &lt;strong&gt;120/120&lt;/strong&gt; on the confirmation set and &lt;strong&gt;159/160&lt;/strong&gt; overall. That single fact reframes the whole experiment: there was essentially nothing to fix, so I measured &lt;em&gt;maintenance near the ceiling&lt;/em&gt;, not &lt;em&gt;error correction&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo1utplj4sfkzjanrnd1p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo1utplj4sfkzjanrnd1p.png" alt=" " width="800" height="448"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Re-review did not beat a blind re-run
&lt;/h3&gt;

&lt;p&gt;Descriptively, across all 160 pairs &lt;strong&gt;B and C both scored 156&lt;/strong&gt;. On the pre-registered confirmation set the 2×2 is: both correct 117, C-only-correct 1, B-only-correct 0, both-wrong 2. That single discordant pair (&lt;code&gt;川辺駿&lt;/code&gt; / Hayao Kawabe) gives exact &lt;strong&gt;McNemar p = 1.0&lt;/strong&gt; — which means "no power to distinguish," &lt;strong&gt;not&lt;/strong&gt; "proven equivalent." The balanced-accuracy gap is +0.8 pt, but with one discordant pair the paired bootstrap lower bound is pinned at 0, so I don't lean on it. I don't claim equivalence either.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. "Destruction" happened in both conditions, and it isn't self-review-specific
&lt;/h3&gt;

&lt;p&gt;Reversals of A's correct answers, over all 160 pairs: &lt;strong&gt;A→B = 3, A→C = 3&lt;/strong&gt; (each 3/159 = 1.9%, Wilson 0.64–5.4%). On the confirmation set B breaks 3 and C breaks 2, but that flips on the hard set — subset-dependent. C's final verdict actually matches A on &lt;strong&gt;157/160&lt;/strong&gt;; the few it moved were all in the wrong direction, but with A at 159/160, "everything it moved was wrong" is close to arithmetically inevitable — not evidence C selectively destroys correct answers.&lt;/p&gt;

&lt;p&gt;Deviations from the pre-registration (honest)&lt;/p&gt;

&lt;p&gt;The pre-registered CI was a paired bootstrap (one of two registered options); it degenerated with a single discordant pair, so I report exact McNemar as an &lt;em&gt;added, post-hoc&lt;/em&gt; auxiliary and don't lean on the interval. Wilson intervals were pre-registered for the secondary counts but missing from the first analysis script (added afterward). A pre-registered fame-stratified breakdown wasn't run (no clean popularity proxy). "Confirmatory" claims are limited to C-vs-B on the confirmation set.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Honest scope note:&lt;/strong&gt; I did not measure human review, a different model, or any tool that checks the answer against ground truth. So I can't say external verification is better — only that handing the same model its own output didn't beat a plain re-run here.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;What I actually do now:&lt;/strong&gt; at this accuracy, I don't spend the extra pass. If I want a real second opinion, I reach for a &lt;em&gt;different&lt;/em&gt; signal (a rule, a lookup, a second model), not the same model re-reading the same cues.&lt;/p&gt;

&lt;h2&gt;
  
  
  Everything I got wrong
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;My non-match pairs were too easy.&lt;/strong&gt; I first built "same surname, different person" pairs, but the Japanese-side player was consistently older (mean 3.5-year gap), so the model could separate them on &lt;em&gt;era&lt;/em&gt; alone. I rebuilt them to minimize the birth-year gap. &lt;em&gt;Lesson: the benchmark author has to suspect their own benchmark is too easy.&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;My headline was inflated, and my own data killed it.&lt;/strong&gt; I started to write "self-review destroys correct answers" — then noticed a blind re-run destroyed just as many. The mechanism was "a second pass at ceiling," not "self-review." &lt;em&gt;Lesson: reconstruct the alternative baseline in your own data before you name a cause.&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;I almost claimed the hard cases needed club matching.&lt;/strong&gt; 79 of 80 non-matches are actually separable by &lt;em&gt;reading&lt;/em&gt; the given name. Only one truly needed an auxiliary attribute — and it's the one the model missed:&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8oblose5pkl3c03p5d2l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8oblose5pkl3c03p5d2l.png" alt=" " width="800" height="448"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;荒井悠汰&lt;/code&gt; and a different person written &lt;code&gt;Yuta Arai&lt;/code&gt; have identical romaji but different birth dates. Reading can't separate them; you need club history. That one homonym was A's only error (a single case, n=1) — and re-review didn't fix it. &lt;em&gt;Lesson: the same model, re-reading the same cues, repeats the same miss.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Limitations
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Ceiling effect (central).&lt;/strong&gt; With 159/160 correct up front, "can it fix errors" is essentially untested; I measured preservation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory contamination is unaddressed.&lt;/strong&gt; These are public players; I can't separate "inferred the reading" from "recalled the person." The fame-stratified probe I pre-registered didn't run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Same-model self-loop only.&lt;/strong&gt; No external verifier, human, or ground-truth check. Not a Self-Refine replication (no separate feedback stage; one shot).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Benchmark representativeness.&lt;/strong&gt; The confirmation set is a fixed 60/60 diagnostic split with hand-built non-matches; not a production match-rate distribution.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Single runs, fixed order, one model, &lt;code&gt;reasoning_effort=low&lt;/code&gt;.&lt;/strong&gt; The ceiling and the null may be specific to this low-effort setting.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Reproduce it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Repo: &lt;a href="https://github.com/sumitsuke/llm-nayose-matching" rel="noopener noreferrer"&gt;https://github.com/sumitsuke/llm-nayose-matching&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;PREREGISTRATION.md&lt;/code&gt; (committed before the run, &lt;code&gt;d7a3328&lt;/code&gt;) → &lt;code&gt;src/run_experiment.py&lt;/code&gt; → &lt;code&gt;src/analyze.py&lt;/code&gt; recomputes every number from &lt;code&gt;run_results.json&lt;/code&gt;. Result commits: &lt;code&gt;fd42585&lt;/code&gt;, analysis add-ons &lt;code&gt;74a00b5&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Public: pair info, model judgments, analysis code, pre-registration. Withheld: DOB/verification attributes and non-redistributable raw data.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;"Ask it again" isn't automatically safer.&lt;/strong&gt; Whether re-review helps depends on how much room there is to improve and where the checking signal comes from.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The same model re-reading the same input isn't a second opinion.&lt;/strong&gt; If the cues are identical, so are the blind spots.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Before measuring self-correction, make sure the first pass can actually fail.&lt;/strong&gt; A ceiling task can't tell you whether re-review fixes errors — it can only show whether it preserves them.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Don't read this as "LLM self-review is useless." Read it as: &lt;strong&gt;on a near-ceiling task, an extra same-model pass showed no measurable value over a blind re-run — and the interesting failure was a homonym that re-reading can't catch.&lt;/strong&gt; The next data point is a deliberately harder set where the first pass drops to 75–90%.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Repo and pre-registration above. I'm explicit about what I did **not&lt;/em&gt;* measure: no human, no second model, no ground-truth checker — so I make no claim about external verification here.*&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This is an English adaptation of my Japanese article — written by me in Japanese, restructured and translated with AI assistance, human-reviewed. Japanese original on Qiita (a Japanese dev-blogging platform): &lt;a href="https://qiita.com/sumitsuke/items/e186a0fe97a8318a2e9c" rel="noopener noreferrer"&gt;https://qiita.com/sumitsuke/items/e186a0fe97a8318a2e9c&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Verification record (environment, verdict, last verified date, evidence) and the canonical write-up: &lt;a href="https://sumitsuke.jp/via/devto/lab/llm-self-review-ceiling/" rel="noopener noreferrer"&gt;https://sumitsuke.jp/via/devto/lab/llm-self-review-ceiling/&lt;/a&gt; — code, data and reproduction: &lt;a href="https://github.com/sumitsuke/llm-nayose-matching" rel="noopener noreferrer"&gt;https://github.com/sumitsuke/llm-nayose-matching&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>testing</category>
      <category>nlp</category>
    </item>
    <item>
      <title>What actually breaks when a free Cloudflare D1 database hits its 500 MB wall</title>
      <dc:creator>Sumitsuke</dc:creator>
      <pubDate>Thu, 09 Jul 2026 13:57:02 +0000</pubDate>
      <link>https://dev.to/tauridev/what-actually-breaks-when-a-free-cloudflare-d1-database-hits-its-500-mb-wall-4g9l</link>
      <guid>https://dev.to/tauridev/what-actually-breaks-when-a-free-cloudflare-d1-database-hits-its-500-mb-wall-4g9l</guid>
      <description>&lt;p&gt;Every "free serverless DB" comparison ranks capacity, price, and cold-start. Almost none answer the question you actually hit in production: &lt;strong&gt;what happens at the free-tier wall?&lt;/strong&gt; So I filled a Cloudflare D1 database to its limit over the REST API and poked it.&lt;/p&gt;

&lt;p&gt;The honest short version: the wall is &lt;strong&gt;not&lt;/strong&gt; a full read-only lockout, a plain &lt;code&gt;DELETE&lt;/code&gt; clears it immediately (but that recovery is not quota-free), and — the part I got wrong &lt;strong&gt;twice&lt;/strong&gt; — building an index on a blob column roughly &lt;strong&gt;doubles&lt;/strong&gt; your storage, so it slams into the wall on its own. This post has been revised twice: my own &lt;code&gt;size_after&lt;/code&gt; logs, and then a follow-up experiment, corrected me. I'll show both corrections instead of hiding them.&lt;/p&gt;

&lt;p&gt;Scope: &lt;strong&gt;Workers Free, throwaway DBs, over the D1 REST API &lt;code&gt;/query&lt;/code&gt;, 2026-07.&lt;/strong&gt; The coarse wall was reproduced on a second, freshly-created database; the exact-ceiling pin and the index-doubling are each n=1 (flagged). This is a REST probe, not a Worker Binding probe — don't generalize endpoint-specific behavior (especially &lt;code&gt;VACUUM&lt;/code&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  "The free-tier wall" is not one wall
&lt;/h2&gt;

&lt;p&gt;D1 has &lt;strong&gt;several&lt;/strong&gt; walls, with different scopes and — per Cloudflare's own error list — different &lt;em&gt;messages&lt;/em&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Wall&lt;/th&gt;
&lt;th&gt;Scope&lt;/th&gt;
&lt;th&gt;Value (Free)&lt;/th&gt;
&lt;th&gt;Documented message&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;per-DB storage&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;per database&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;500 MB&lt;/strong&gt; (Paid: 10 GB)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;Exceeded maximum DB size&lt;/code&gt; — action: "delete rows / shard"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;account storage&lt;/td&gt;
&lt;td&gt;account&lt;/td&gt;
&lt;td&gt;5 GB&lt;/td&gt;
&lt;td&gt;&lt;code&gt;Your account has exceeded D1's maximum account storage limit…&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;daily rows read&lt;/td&gt;
&lt;td&gt;account&lt;/td&gt;
&lt;td&gt;5,000,000/day&lt;/td&gt;
&lt;td&gt;daily-limit-exceeded → "can't run queries"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;daily rows written&lt;/td&gt;
&lt;td&gt;account&lt;/td&gt;
&lt;td&gt;100,000/day&lt;/td&gt;
&lt;td&gt;same (&lt;strong&gt;INSERT, UPDATE, and DELETE all count as writes&lt;/strong&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;(Free daily counters reset at 00:00 UTC; storage is not daily-reset. Max row 2 MB; 100 columns; 10 DBs/account. All confirmed against the D1 docs on 2026-07-09.)&lt;/em&gt; &lt;strong&gt;I measured the first one.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 1 — the wall is 500 MB (decimal); pinned to one row on one DB, reproduced coarsely on a second
&lt;/h2&gt;

&lt;p&gt;Filling with bulk inserts stalls around ~499.9 MB. To pin it, I then looped &lt;strong&gt;single-row&lt;/strong&gt; inserts at the wall:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;… 24 single-row inserts succeed (size_after climbs)
25th: HTTP 400  [{"code":7500,"message":"Exceeded maximum DB size"}]
last successful size_after = 499,994,624 bytes  (≈ 499.99 MB ≈ 476.83 MiB)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That ceiling — &lt;strong&gt;499,994,624&lt;/strong&gt; — is &lt;code&gt;500,000,000&lt;/code&gt; (decimal 500 MB) minus 5,376 bytes, enforced at SQLite's 4 KB page granularity (&lt;code&gt;499,994,624 = 122,069 × 4096&lt;/code&gt;). I reproduced the &lt;strong&gt;coarse&lt;/strong&gt; wall on a second, freshly-created DB: it stopped at the identical &lt;code&gt;499,884,032&lt;/code&gt; and returned the same &lt;code&gt;7500&lt;/code&gt;. One caveat so you can hold me to my own logs: the redacted &lt;code&gt;results/*.jsonl&lt;/code&gt; carry &lt;strong&gt;no account IDs&lt;/strong&gt;, and &lt;code&gt;served_by_colo&lt;/code&gt; is Anycast edge routing, not an account signal — so what's verifiable from the evidence is &lt;em&gt;a second database&lt;/em&gt;, and the byte-exact &lt;code&gt;499,994,624&lt;/code&gt; pin is from the single-row probe on &lt;strong&gt;one&lt;/strong&gt; of them (n=1). The wall firing at all means these DBs are on &lt;strong&gt;Workers Free&lt;/strong&gt; (Paid = 10 GB).&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 2 — not read-only, but small writes aren't "exempt" either (measured)
&lt;/h2&gt;

&lt;p&gt;My first draft claimed small writes slip past the wall while big ones don't. Wrong — and my own numbers show why. Those 24 single-row inserts above consumed only &lt;code&gt;499,994,624 − 499,884,032 = 110,592 bytes&lt;/code&gt; (~110 KB) of headroom before even a single-row insert was rejected. So:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Reads always pass&lt;/strong&gt; — near-tautological for a storage cap; reads don't grow the DB.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Writes pass only while headroom remains.&lt;/strong&gt; Small writes weren't bypassing the wall; they were filling the last sliver before it. Call it slack, not exemption. (One headroom-exhaustion sequence, but the 25th insert failing at zero headroom is a clean counterexample to "small writes are exempt" — one is enough.)&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Result 3 — recovery is a plain DELETE, no upgrade (but it spends write quota)
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;DELETE&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;20000&lt;/span&gt;   &lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="mi"&gt;19&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;998&lt;/span&gt; &lt;span class="k"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="n"&gt;size_after&lt;/span&gt; &lt;span class="mi"&gt;499&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;9&lt;/span&gt; &lt;span class="n"&gt;MB&lt;/span&gt; &lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="mi"&gt;243&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;884&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;032&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;~&lt;/span&gt;&lt;span class="mi"&gt;244&lt;/span&gt; &lt;span class="n"&gt;MB&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;immediately&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="n"&gt;rows_written&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;19&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;998&lt;/span&gt;
&lt;span class="n"&gt;bulk&lt;/span&gt; &lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;after&lt;/span&gt; &lt;span class="k"&gt;delete&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="n"&gt;restored&lt;/span&gt;
&lt;span class="k"&gt;VACUUM&lt;/span&gt;                    &lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="mi"&gt;7500&lt;/span&gt; &lt;span class="nv"&gt;"cannot VACUUM from within a transaction: SQLITE_ERROR"&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;maybe&lt;/span&gt; &lt;span class="n"&gt;REST&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="k"&gt;specific&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;DELETE&lt;/code&gt; drops D1's reported size &lt;strong&gt;immediately&lt;/strong&gt; and writes resume — no &lt;code&gt;VACUUM&lt;/code&gt;, no plan upgrade (&lt;strong&gt;recovery class R1&lt;/strong&gt;). Two honest qualifiers: (1) &lt;strong&gt;it isn't quota-free&lt;/strong&gt; — D1 counts &lt;code&gt;DELETE&lt;/code&gt; as a write (docs: &lt;em&gt;"Write operations include INSERT, UPDATE, and DELETE"&lt;/em&gt;), so this one 19,998-row delete spent ~20% of the daily 100k budget, and the cost scales with &lt;strong&gt;row count, not bytes freed&lt;/strong&gt; (this delete freed ~256 MB ≈ 12.8 KB/row; whether it's linear per chunk I didn't separately measure); (2) &lt;code&gt;size_after&lt;/code&gt; is D1's &lt;em&gt;reported&lt;/em&gt; size — the meter dropping doesn't prove the physical file shrank (cf. workerd#1618, local).&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 4 — CREATE INDEX doubles your storage (the thing I got wrong twice)
&lt;/h2&gt;

&lt;p&gt;At the wall, &lt;code&gt;CREATE INDEX&lt;/code&gt; fails with &lt;code&gt;out of memory: SQLITE_NOMEM&lt;/code&gt;. My first draft lumped that in as "large allocation blocked"; after review I "corrected" it to "that's a memory limit, independent of the storage wall." &lt;strong&gt;Both were wrong.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;On a &lt;strong&gt;control DB at 300 MB&lt;/strong&gt; — well &lt;em&gt;below&lt;/em&gt; the 500 MB wall — &lt;code&gt;CREATE INDEX&lt;/code&gt; failed not with &lt;code&gt;SQLITE_NOMEM&lt;/code&gt; but with &lt;strong&gt;&lt;code&gt;7500 "Exceeded maximum DB size"&lt;/code&gt;&lt;/strong&gt; (storage!). Below the wall, a storage error? I measured the mechanism directly — deleted the control to 80 MB and indexed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;INDEX&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;bench&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;v&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;at&lt;/span&gt; &lt;span class="mi"&gt;84&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;029&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;440&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;80&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="n"&gt;MiB&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="n"&gt;OK&lt;/span&gt;
  &lt;span class="n"&gt;size_after&lt;/span&gt; &lt;span class="mi"&gt;84&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;029&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;440&lt;/span&gt; &lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="mi"&gt;167&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;923&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;712&lt;/span&gt;   &lt;span class="o"&gt;=&lt;/span&gt;  &lt;span class="err"&gt;×&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;998&lt;/span&gt; &lt;span class="n"&gt;growth&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;measured&lt;/span&gt; &lt;span class="n"&gt;once&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;at&lt;/span&gt; &lt;span class="o"&gt;~&lt;/span&gt;&lt;span class="mi"&gt;80&lt;/span&gt; &lt;span class="n"&gt;MB&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;An index on a blob column carries a copy of the column's data as its keys, so it roughly doubles storage&lt;/strong&gt; — inferred from this one growth point, not proven in general. If that ×2 holds at other sizes:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;storage before index&lt;/th&gt;
&lt;th&gt;after (~×2)&lt;/th&gt;
&lt;th&gt;result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;80 MB&lt;/td&gt;
&lt;td&gt;160 MB&lt;/td&gt;
&lt;td&gt;✅ succeeds (under the wall)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;300 MB&lt;/td&gt;
&lt;td&gt;~600 MB&lt;/td&gt;
&lt;td&gt;⛔ &lt;code&gt;Exceeded maximum DB size&lt;/code&gt; (over the 500 MB wall)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;500 MB (at the wall)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;⛔ &lt;code&gt;SQLITE_NOMEM&lt;/code&gt; (OOMs during the build)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So the index failure is &lt;strong&gt;storage-bound via the index's own footprint&lt;/strong&gt;, not "memory, unrelated to the wall." A rough rule of thumb — &lt;em&gt;if&lt;/em&gt; the ×2 holds — is that on D1 Free you can only index a column holding up to ~250 MB of data (2 × 250 = 500 MB), but that's an extrapolation, not measured near 250 MB. The &lt;code&gt;SQLITE_NOMEM&lt;/code&gt; is just what you see when you're already &lt;em&gt;at&lt;/em&gt; the wall.&lt;/p&gt;

&lt;h2&gt;
  
  
  The error code, and what I did not measure
&lt;/h2&gt;

&lt;p&gt;Cloudflare's error list classifies these by &lt;strong&gt;message&lt;/strong&gt; (&lt;code&gt;Exceeded maximum DB size&lt;/code&gt;, &lt;code&gt;…account storage limit…&lt;/code&gt;, &lt;code&gt;…isolate memory limit…&lt;/code&gt;); the number &lt;code&gt;7500&lt;/code&gt; appears &lt;strong&gt;nowhere&lt;/strong&gt; in the docs. So &lt;code&gt;7500&lt;/code&gt; is a generic REST wrapper — the message is the discriminator.&lt;/p&gt;

&lt;p&gt;Not measured / n=1 / extrapolated (so not claimed): the &lt;strong&gt;index ×2 is a single measurement&lt;/strong&gt; at one size/row-shape; the ~250 MB threshold is extrapolated (not measured near 250 MB); DELETE linearity across chunks is not measured (one 19,998-row delete); the account 5 GB wall, the daily read wall (5M/day — &lt;em&gt;reads&lt;/em&gt; stop, account-scoped), and Worker Binding behavior are all untested; &lt;code&gt;size_after&lt;/code&gt; behaves like a coherent logical meter (dropped on the one DELETE, ~doubled on the one index build) but the docs don't say physical vs logical. One full cycle (fill 39k + recover 20k + control 24k ≈ 84k &lt;code&gt;rows_written&lt;/code&gt;) fits in one UTC day but not two.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproduce it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/sumitsuke/free-tier-quota-probe
&lt;span class="nb"&gt;cd &lt;/span&gt;free-tier-quota-probe
npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; npm run selftest
&lt;span class="c"&gt;# .env: CF_ACCOUNT_ID / CF_API_TOKEN (D1:Edit) / D1_DATABASE_ID / D1_CONTROL_DATABASE_ID / BILLING_SAFE=confirmed&lt;/span&gt;
node &lt;span class="nt"&gt;--env-file&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;.env bin/d1-admin.mjs create wallprobe   &lt;span class="c"&gt;# create the DB via API (no dashboard needed)&lt;/span&gt;
npm run ping        &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="nt"&gt;--target&lt;/span&gt; d1
npm run fill        &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="nt"&gt;--target&lt;/span&gt; d1 &lt;span class="nt"&gt;--confirm-destructive&lt;/span&gt;
npm run probe-slack &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="nt"&gt;--target&lt;/span&gt; d1                        &lt;span class="c"&gt;# pins the wall + measures the headroom&lt;/span&gt;
npm run fill        &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="nt"&gt;--target&lt;/span&gt; d1 &lt;span class="nt"&gt;--control&lt;/span&gt; &lt;span class="nt"&gt;--cap-mb&lt;/span&gt; 300 &lt;span class="nt"&gt;--confirm-destructive&lt;/span&gt;
npm run probe-index &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="nt"&gt;--target&lt;/span&gt; d1                        &lt;span class="c"&gt;# the index-doubling test&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Node ≥ 20, ESM, D1 over native &lt;code&gt;fetch&lt;/code&gt; (60 s timeout + one retry). Redacted captures are frozen in &lt;code&gt;results/*.jsonl&lt;/code&gt;; predictions (and the follow-up amendment) are in &lt;code&gt;PREREGISTRATION.md&lt;/code&gt;. &lt;strong&gt;Repo:&lt;/strong&gt; &lt;code&gt;https://github.com/sumitsuke/free-tier-quota-probe&lt;/code&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I write about measurement discipline — filling the wall instead of quoting the limit, and correcting my own assumptions on the record. This post's Results 2 and 4 are the live examples: I reasoned wrong twice and only the measurements got it right. Written with AI assistance; every number is from the captured `results/&lt;/em&gt;.jsonl`, and the doc facts were re-verified against Cloudflare's D1 docs on 2026-07-09. A Japanese version of this write-up also exists.*&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Verification record (environment, verdict, last verified date, evidence) and the canonical write-up: &lt;a href="https://sumitsuke.jp/lab/" rel="noopener noreferrer"&gt;https://sumitsuke.jp/lab/&lt;/a&gt; — code, data and reproduction: &lt;a href="https://github.com/sumitsuke/free-tier-quota-probe" rel="noopener noreferrer"&gt;https://github.com/sumitsuke/free-tier-quota-probe&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cloudflare</category>
      <category>database</category>
      <category>sqlite</category>
      <category>serverless</category>
    </item>
    <item>
      <title>I had a local LLM review the same code 51 times: majority voting picks persistence, not correctness</title>
      <dc:creator>Sumitsuke</dc:creator>
      <pubDate>Tue, 07 Jul 2026 11:48:35 +0000</pubDate>
      <link>https://dev.to/tauridev/i-had-a-local-llm-review-the-same-code-51-times-majority-voting-picks-persistence-not-correctness-4g15</link>
      <guid>https://dev.to/tauridev/i-had-a-local-llm-review-the-same-code-51-times-majority-voting-picks-persistence-not-correctness-4g15</guid>
      <description>&lt;p&gt;If you use an LLM to review code — or run it several times and keep "whatever most runs agree on" — this experiment suggests the votes you're counting measure &lt;strong&gt;persistence&lt;/strong&gt;, not &lt;strong&gt;correctness&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Setup: Qwen2.5-Coder 1.5B / 7B (Apache-2.0) via Ollama 0.31.1, single RTX 4060 Ti, Q4_K_M quantization. N=51 repeated audits of the same small Python files with predeclared seeded defects (temp-0 baseline: N=11 per condition). The per-run records (normalized findings plus a SHA-256 of each raw response), ground truth, and aggregation code are public — every number quoted in this article regenerates deterministically from the frozen records via an independent verification script.&lt;/em&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Fully pinned = deterministic (on this rig).&lt;/strong&gt; At temperature 0 with a fixed GPU and quantization, the finding sets were identical across runs — Jaccard similarity &lt;strong&gt;1.00&lt;/strong&gt; in all four conditions (N=11 each). In a probe, six repeats within one seed and six different seeds all produced &lt;em&gt;byte-identical&lt;/em&gt; output: our temp-0 configuration also forces &lt;code&gt;top_k=1&lt;/code&gt; (greedy decoding in Ollama/llama.cpp), so the seed had no observable effect. "LLMs are always nondeterministic" is folklore, at least under fully pinned local conditions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sampling at temperature 0.7 breaks agreement.&lt;/strong&gt; On the defect-seeded file — 51 runs, one seed per run (0–50) — run-to-run Jaccard dropped to &lt;strong&gt;0.52&lt;/strong&gt; (7B) and &lt;strong&gt;0.07&lt;/strong&gt; (1.5B). For 7B, allowing ±2 lines of tolerance recovers 0.52 → &lt;strong&gt;0.80&lt;/strong&gt;: over half of the disagreement is consistent with &lt;em&gt;the same bug pointed at a different line&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The core finding: majority voting selects persistence, not correctness.&lt;/strong&gt; A real resource leak that the model detected in &lt;strong&gt;45 of 51 runs&lt;/strong&gt; split its votes across two adjacent lines (26 + 19 — no run reported both, so those really are 45 distinct runs). Under the strict line+category key, the majority threshold of 26 barely keeps one line; one vote higher, at k=27, a defect found 45 times is gone. Meanwhile, two "grey" pattern warnings (&lt;code&gt;eval&lt;/code&gt; and &lt;code&gt;shell=True&lt;/code&gt; on &lt;em&gt;constants&lt;/em&gt;, from the decoy file) survived at 33/51 and 28/51, because the model points at exactly the same place every time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Honest caveats up front:&lt;/strong&gt; on the defect-seeded file, against predeclared ground truth with category + ±2-line matching, the 7B's per-run recall was only 0.31 — but precision was 0.98. Within this benchmark: when it speaks, it's right, and it misses a lot. The 1.5B was too noisy to carry any claim (it hit 25 of the file's 32 lines at least once across 51 runs), so it serves as a contrast group, not evidence.&lt;/li&gt;
&lt;/ol&gt;

&lt;blockquote&gt;
&lt;p&gt;⚠ &lt;strong&gt;Reproducibility scope:&lt;/strong&gt; the &lt;em&gt;existence and rough magnitude&lt;/em&gt; of the variance reproduces; exact numbers are rig-dependent (GPU / quantization / inference-engine version). The repo ships frozen per-run records — normalized findings plus a hash of each raw response; the raw response bodies themselves were not retained — and readers re-run only the deterministic aggregation layer.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Why measure this
&lt;/h2&gt;

&lt;p&gt;"LLM-as-a-judge" is everywhere, and its known weakness is nondeterminism. Prior work measures &lt;em&gt;score&lt;/em&gt; stability — &lt;a href="https://arxiv.org/abs/2502.20747" rel="noopener noreferrer"&gt;Klishevich et al. (arXiv:2502.20747)&lt;/a&gt; measured scalar-verdict consistency over 70 commits at temp 0, and a Japanese industrial blog measured same-score rates on cloud APIs. What I couldn't find measured anywhere: whether &lt;strong&gt;the set of findings itself&lt;/strong&gt; — &lt;em&gt;which&lt;/em&gt; issues, on &lt;em&gt;which&lt;/em&gt; lines — is stable across runs, on a local model, free to reproduce, with run records frozen and shipped. That's the gap this fills.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setup
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Models:&lt;/strong&gt; &lt;code&gt;Qwen2.5-Coder&lt;/code&gt; 1.5B and 7B (Apache-2.0), run locally via Ollama. Chosen because they're free, local, and require no API key — anyone can rerun generation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sampling design:&lt;/strong&gt; temp-0 conditions use a &lt;strong&gt;fixed seed&lt;/strong&gt; and force &lt;code&gt;top_k=1, top_p=1.0&lt;/code&gt;; temp-0.7 conditions use &lt;strong&gt;one seed per run (0, 1, …, 50)&lt;/strong&gt; with Ollama's default sampler otherwise (&lt;code&gt;top_k=40, top_p=0.9&lt;/code&gt;; not swept). The temp-0.7 numbers therefore measure &lt;strong&gt;across-seed sampling variability at temperature 0.7&lt;/strong&gt; — not same-seed runtime nondeterminism (with a fixed seed, repeats would be trivially identical).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Output format:&lt;/strong&gt; Ollama structured outputs (JSON-schema-enforced). Each finding is &lt;code&gt;{line_number, category, severity, description}&lt;/code&gt;, category constrained to an enum.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Identity of a finding:&lt;/strong&gt; the composite key &lt;code&gt;(line_number, category)&lt;/code&gt;. Paraphrases are ignored; two findings are "the same" iff they name the same place and problem type. I deliberately did &lt;em&gt;not&lt;/em&gt; use embedding similarity — it happily merges "missing null check on line 12" with "missing null check on line 45".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Target files (4):&lt;/strong&gt; &lt;code&gt;target_a&lt;/code&gt; (7 seeded defects), &lt;code&gt;clean&lt;/code&gt; (no defects — false-positive baseline), &lt;code&gt;decoy&lt;/code&gt; (looks dangerous, low actual harm — this classification gets revised later, see pitfall #4), &lt;code&gt;easy&lt;/code&gt; (6 obvious defects).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ground truth:&lt;/strong&gt; predeclared in a &lt;code&gt;gt.csv&lt;/code&gt; &lt;em&gt;outside&lt;/em&gt; the code before measurement. No hints in the code or the prompt — I learned this one the hard way (pitfall #1).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Two-layer reproducibility:&lt;/strong&gt; layer 1 (generation, nondeterministic) ran once and its per-run records are frozen in the repo; layer 2 (aggregation, Python stdlib only) is fully deterministic and is what you reproduce.&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Choice&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Inference&lt;/td&gt;
&lt;td&gt;Ollama 0.31.1&lt;/td&gt;
&lt;td&gt;local, free&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Models&lt;/td&gt;
&lt;td&gt;Qwen2.5-Coder 1.5B / 7B&lt;/td&gt;
&lt;td&gt;Q4_K_M (verified via &lt;code&gt;ollama show&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hardware&lt;/td&gt;
&lt;td&gt;NVIDIA RTX 4060 Ti (CUDA)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;ollama ps&lt;/code&gt; reported 100% GPU offload during runs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Aggregation&lt;/td&gt;
&lt;td&gt;Python 3.12, stdlib only&lt;/td&gt;
&lt;td&gt;layer 2 fully deterministic&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Measured: 2026-07. Per-cell provenance (model quantization/size/family, Ollama version, prompt/schema/GT hashes, backend) is recorded in the frozen records.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How I measured
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Reliability = Jaccard between runs (no ground truth involved):&lt;/strong&gt; &lt;code&gt;matched findings / union&lt;/code&gt;, computed with &lt;strong&gt;optimal bipartite matching&lt;/strong&gt; — naive greedy matching gives order-dependent values (a bug I actually shipped and had to fix). Reported as the mean over all run pairs within a condition, with a &lt;strong&gt;line-tolerance window w&lt;/strong&gt; (±0 strict / ±2 / ±5) reported alongside — this doubles as the sensitivity analysis for the identity key.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Validity = precision / recall / false positives against GT.&lt;/strong&gt; Unless noted, recall/precision are per-run means on &lt;code&gt;target_a&lt;/code&gt;, category-strict, line ±2. Precision excludes empty runs by definition — and there were none on &lt;code&gt;target_a&lt;/code&gt; (0/51); recall averages over all runs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Majority voting = keep findings that appear in ≥ k of N runs.&lt;/strong&gt; Examined on two axes: threshold k × key granularity w.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Uncertainty:&lt;/strong&gt; 95% CIs from bootstrap resampling &lt;strong&gt;at the run level&lt;/strong&gt;. One honest structural caveat: 51 audits of the &lt;em&gt;same&lt;/em&gt; file are pseudo-replication — nothing here generalizes across codebases, and I don't claim it does.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Result 1: pin everything, and it's deterministic
&lt;/h2&gt;

&lt;p&gt;All four temp-0 conditions (2 models × {defect-seeded, easy} files, N=11): Jaccard &lt;strong&gt;1.00&lt;/strong&gt; at every tolerance level, zero variance across runs. In a separate probe on one condition, six repeats within a single seed and six different seeds (1, 2, 3, 101, 202, 303) produced byte-identical outputs — our temp-0 configuration also forces &lt;code&gt;top_k=1&lt;/code&gt;, i.e. greedy decoding in Ollama/llama.cpp, so the seed had no observable effect there. Folklore says "LLMs always wobble"; under fully pinned local conditions on this rig, they didn't.&lt;/p&gt;

&lt;p&gt;📐 The honest asterisk on "deterministic"&lt;/p&gt;

&lt;p&gt;In a separate spike I did observe a single category flip at temp 0 — only ever on the &lt;em&gt;first&lt;/em&gt; run after model load, which also had the longest latency (cold-cache/warm-up signature). With a warm-up run added, the residual disappeared. I can't pin the mechanism (GPU parallel-reduction ordering is the usual suspect; cold cache isn't ruled out), but it isn't the seed. This is consistent with &lt;a href="https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/" rel="noopener noreferrer"&gt;Thinking Machines' "Defeating Nondeterminism in LLM Inference"&lt;/a&gt;, which attributes serving-time nondeterminism primarily to &lt;strong&gt;batch-size variation&lt;/strong&gt; — a local single-request setup is effectively fixed-batch, which is consistent with the 1.00 here. The same mechanism would also predict the temp-0 flutter Klishevich et al. saw on &lt;em&gt;cloud APIs&lt;/em&gt; (variable batching), though they discuss other candidate causes as well — read this as consistency with one proposed mechanism, not a proven attribution.&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 2: sample at temperature 0.7, and it falls apart
&lt;/h2&gt;

&lt;p&gt;People raise temperature precisely to widen candidate coverage across runs. At temp 0.7 on the defect-seeded file (one seed per run, 0–50), run-to-run Jaccard was &lt;strong&gt;0.52&lt;/strong&gt; [0.47, 0.59] for 7B and &lt;strong&gt;0.07&lt;/strong&gt; [0.04, 0.11] for 1.5B.&lt;/p&gt;

&lt;p&gt;Two distinct failure modes hide in those numbers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Line drift (7B):&lt;/strong&gt; 0.52 recovers to &lt;strong&gt;0.80&lt;/strong&gt; when the identity key tolerates ±2 lines. Over half of the disagreement is consistent with the same bug pointed at a different line — though tolerance matching can also create accidental matches, so treat ±2 as an upper bound on "drift".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Category flutter (1.5B, and &lt;code&gt;easy&lt;/code&gt;):&lt;/strong&gt; on the easy file, 7B's line-only recall is 0.81 but category-strict recall is 0.42 — the model keeps finding the right &lt;em&gt;place&lt;/em&gt; and labeling it a different &lt;em&gt;problem&lt;/em&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;File&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;recall (category-strict)&lt;/th&gt;
&lt;th&gt;recall (line-only)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;target_a&lt;/td&gt;
&lt;td&gt;7B&lt;/td&gt;
&lt;td&gt;0.31&lt;/td&gt;
&lt;td&gt;0.31&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;target_a&lt;/td&gt;
&lt;td&gt;1.5B&lt;/td&gt;
&lt;td&gt;0.05&lt;/td&gt;
&lt;td&gt;0.18&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;easy&lt;/td&gt;
&lt;td&gt;7B&lt;/td&gt;
&lt;td&gt;0.42&lt;/td&gt;
&lt;td&gt;0.81&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;(For 7B on &lt;code&gt;target_a&lt;/code&gt; the two recalls coincide — those misses are genuine non-detections, not drift.)&lt;/p&gt;

&lt;p&gt;And here's where I hit my first big mistake: my early recall numbers were great — because I had left &lt;code&gt;# DEFECT: SQL injection&lt;/code&gt; annotations in the target code as notes to myself, and the whole file goes into the prompt. &lt;strong&gt;The model was reading my answer key.&lt;/strong&gt; Deleting the labels cratered recall and deepened the run-to-run variance. Details in the pitfall log below.&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 3 (the core): majority voting rewards persistence
&lt;/h2&gt;

&lt;p&gt;Majority vote = "keep findings that appear in ≥26 of 51 runs." Here is what that actually selects.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjdpfu4i4zau2i29s8gpj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjdpfu4i4zau2i29s8gpj.png" alt="Votes per finding across 51 audits (7B, temp 0.7). Grey-smell warnings survive above the majority line; the real ResourceLeak bug splits 26/19 across two lines and straddles the threshold." width="800" height="409"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Figure: votes per finding (blue = predeclared seeded defects, red = persistent findings outside GT, dashed = majority threshold 26/51). Counts are pooled from two target files run under identical conditions: the blue bars come from the defect-seeded file, the red bars from the decoy file.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The mechanism.&lt;/strong&gt; With &lt;code&gt;(line, category)&lt;/code&gt; as the identity key, the same bug pointed at a neighboring line becomes &lt;em&gt;a different ballot&lt;/em&gt;. The model detected the ResourceLeak in &lt;strong&gt;45/51 runs&lt;/strong&gt; — but the votes landed 26 on line 16 and 19 on line 17 (no single run reported both lines, so 26+19 really is 45 distinct runs). Loosen the question to "did it say ResourceLeak at all, anywhere?" and it's 49/51 — scattered across &lt;strong&gt;six&lt;/strong&gt; different lines (12, 15, 16, 17, 18, 21). The vote count never sees that. Watch the threshold:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Threshold k&lt;/th&gt;
&lt;th&gt;line 16 (26 votes)&lt;/th&gt;
&lt;th&gt;line 17 (19 votes)&lt;/th&gt;
&lt;th&gt;Outcome&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;k=19&lt;/td&gt;
&lt;td&gt;kept&lt;/td&gt;
&lt;td&gt;kept&lt;/td&gt;
&lt;td&gt;both survive (duplicated)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;k=26 (majority)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;kept&lt;/td&gt;
&lt;td&gt;dropped&lt;/td&gt;
&lt;td&gt;survives by one vote&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;k=27&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;dropped&lt;/td&gt;
&lt;td&gt;dropped&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;a defect reported in 45/51 runs is gone&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Under the exact majority rule I predeclared (k=26) the leak &lt;em&gt;barely&lt;/em&gt; survives on one line; the point is not that majority voting killed it, but that &lt;strong&gt;whether it survives is governed by how the votes split and where you draw the line — not by the model's detection ability&lt;/strong&gt;. And it's not one anecdote: &lt;strong&gt;2 of the 5 defects that 7B detected at least once (~40%) showed line drift between runs&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Two contrasts sharpen the picture:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Unsplit real bugs sail through:&lt;/strong&gt; the SQL injection on line 9 got &lt;strong&gt;51/51&lt;/strong&gt; unanimous votes. Voting doesn't kill strong detections — it kills &lt;em&gt;wobbly-keyed&lt;/em&gt; ones.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Grey findings survive:&lt;/strong&gt; warnings against &lt;code&gt;eval("2 + 2")&lt;/code&gt; (33/51) and &lt;code&gt;subprocess.run(..., shell=True)&lt;/code&gt; with a constant string (28/51) — both low-harm on constants, both from the decoy file — persist because the model reacts to the &lt;em&gt;pattern&lt;/em&gt; and points at exactly the same place every time. (Bonus wobble: it labeled the latter &lt;code&gt;CodeInjection&lt;/code&gt; rather than &lt;code&gt;CommandInjection&lt;/code&gt; — categories flutter too.)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The &lt;code&gt;clean&lt;/code&gt; file adds a neat corollary: 7B stayed silent in 37/51 runs (72.5%), and when it &lt;em&gt;did&lt;/em&gt; complain, it complained about something different almost every time — non-empty clean runs agree at just &lt;strong&gt;0.12&lt;/strong&gt;. (Counting silent–silent pairs as agreement 1.0 the unconditional mean is 0.53; I report the conditional number because silence agreeing with silence isn't informative here.) Inconsistent false alarms are exactly what majority voting removes. Consistent grey warnings are exactly what it cannot.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Threshold × key-granularity sensitivity&lt;/strong&gt; (majority k=26):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;File (7B)&lt;/th&gt;
&lt;th&gt;w ±0&lt;/th&gt;
&lt;th&gt;±2&lt;/th&gt;
&lt;th&gt;±5&lt;/th&gt;
&lt;th&gt;Reading&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;clean&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;FP &lt;strong&gt;0&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;scattered false alarms die (as intended)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;decoy&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;2 persistent&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;consistent grey findings survive&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;target_a&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;TP 2&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;3 *&lt;/td&gt;
&lt;td&gt;* the +1 is a double-credit artifact (below)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;easy&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;TP 2 / FP 1&lt;/td&gt;
&lt;td&gt;3 / 2&lt;/td&gt;
&lt;td&gt;3 / 1 †&lt;/td&gt;
&lt;td&gt;looser keys rescue split votes — at a price&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;blockquote&gt;
&lt;p&gt;† The &lt;code&gt;easy&lt;/code&gt; false-positive count is non-monotonic in w (1→2→1) because widening the window changes how findings cluster — adjacent clusters merge or get absorbed. This is a property of line-window clustering, disclosed rather than hidden.&lt;/p&gt;
&lt;/blockquote&gt;

📐 Why the ±5 "rescue" is double-counting, not detection

At w ±5, the ResourceLeak cluster (mostly line 16) becomes wide enough to overlap **two distinct** ground-truth defects — a connection leak (lines 6–11) and a file-handle leak (lines 15–17) — so one cluster gets credited to both (the aggregation deliberately allows multi-credit; that design choice is exactly what this row exposes). To be precise: no run put a ResourceLeak vote inside the connection-leak lines themselves (6–11); the nearest was a single vote at line 12 — one line outside, and far below any threshold. So the "+1 true positive" at ±5 is a windowing artifact, not a detection. Loosening keys rescues split votes, but the rescue is not free: on `easy` it temporarily raises false positives (1→2 at ±2), and here it manufactures a detection that never cleared any threshold. No free lunch.



&lt;p&gt;&lt;strong&gt;So: majority voting filters for &lt;em&gt;salience of the pattern response&lt;/em&gt;, not truth.&lt;/strong&gt; Persistent findings survive whether they're real bugs or grey smells; real bugs whose expression wobbles can fall. "Persistent" and "correct" are different properties, and the vote measures the first one.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Honest scope note: in this sample the core claim rests mainly on one drifting real bug plus two persistent grey findings — read it as an &lt;strong&gt;observed mechanism&lt;/strong&gt;, not a general law. It is &lt;em&gt;not&lt;/em&gt; an argument against aggregation per se — this is also not the &lt;a href="https://arxiv.org/abs/2203.11171" rel="noopener noreferrer"&gt;self-consistency&lt;/a&gt; setting, which votes over &lt;em&gt;normalized&lt;/em&gt; final answers; the failure mode here lives precisely in the un-normalized key. Change the key (category-only, AST node, defect region) and the behavior changes — which is itself the point: threshold × key granularity is a design decision with no free lunch.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;What I actually do now:&lt;/strong&gt; treat multi-run LLM review as a &lt;strong&gt;candidate pool&lt;/strong&gt;, not an auto-verdict. Prioritize high-frequency findings for review, &lt;em&gt;don't&lt;/em&gt; discard low-frequency findings in severe categories, and let tests, static analysis, and a human's understanding of the spec make the final call.&lt;/p&gt;
&lt;h2&gt;
  
  
  Everything I got wrong (measurement-instrument failures included)
&lt;/h2&gt;

&lt;p&gt;The most useful part of this experiment was watching my own tooling and my own memory fail. Five mistakes, in the order I made them:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;My answer key leaked into the prompt.&lt;/strong&gt; I'd annotated defect lines with &lt;code&gt;# DEFECT: ...&lt;/code&gt; comments as notes to myself — and the audit prompt includes the whole file. High recall was the model reading my answers. Removing the labels cratered recall and deepened variance. &lt;em&gt;Never let the benchmark's answers be visible to the system under test.&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;I thought I was benchmarking CPU inference. It was 100% GPU.&lt;/strong&gt; I had designed a thread-count variable (&lt;code&gt;num_thread&lt;/code&gt; 1 vs 4) — then &lt;code&gt;ollama ps&lt;/code&gt; showed 100% GPU offload, where &lt;code&gt;num_thread&lt;/code&gt; is a no-op. My thread axis had measured nothing. &lt;em&gt;Record your accelerator provenance instead of assuming it — the act of recording is what exposed the error.&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;I broke my own measuring instrument while "fixing" it.&lt;/strong&gt; A one-line "optimization" to a clustering helper introduced a classic append-then-clear aliasing bug that silently corrupted the majority-vote true/false-positive counts. I caught it only because a hand count said 2 where the tool said 3. Every number quoted in this article's text is now recomputed by an independent script (&lt;code&gt;verify_independent.py&lt;/code&gt;) that bypasses the aggregation code entirely; the full threshold×tolerance table and the bootstrap CIs still come from the aggregation pipeline. &lt;em&gt;Fixes need verification most of all.&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;My "false positives" were real smells.&lt;/strong&gt; I'd built &lt;code&gt;decoy&lt;/code&gt; assuming its &lt;code&gt;eval&lt;/code&gt;/&lt;code&gt;shell=True&lt;/code&gt;-on-constants patterns were harmless, so surviving warnings counted as false positives. But Bandit and friends flag exactly these. The finding got &lt;em&gt;stronger&lt;/em&gt; through the correction: the claim isn't "false positives survive voting" — it's "&lt;strong&gt;persistence survives voting, regardless of truth&lt;/strong&gt;", and I renamed those findings "persistent findings outside GT".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;My own memory fabricated a result.&lt;/strong&gt; An early draft said "the &lt;code&gt;md5&lt;/code&gt; warning also survived the vote." The frozen records say md5 got &lt;strong&gt;0 votes in 51 runs&lt;/strong&gt;. My memory had synthesized it from the &lt;em&gt;impression&lt;/em&gt; of pitfall #4. Caught in pre-publication log reconciliation. &lt;em&gt;The experimenter's memory is not a primary source either — recheck the frozen records every time you write a number.&lt;/em&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This is "don't trust LLM output" applied symmetrically: to the model, to my own instruments, and to my own claims.&lt;/p&gt;
&lt;h2&gt;
  
  
  Limitations
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Small sample: N=51, two model sizes, one rig (RTX 4060 Ti / CUDA / Q4_K_M / Ollama 0.31.1). CPU, other GPUs, other quantizations may behave differently.&lt;/li&gt;
&lt;li&gt;The temp-0.7 condition varies the seed per run (by design — that's what sampling the distribution means here), so it measures across-seed variability at that temperature, not a pure temperature effect with everything else held.&lt;/li&gt;
&lt;li&gt;The core claim rests on few instances (one drifting bug, two grey findings). Mechanistically coherent, statistically thin.&lt;/li&gt;
&lt;li&gt;"Smaller models wobble more" is a hypothesis from a 2-point comparison — capability and variance are confounded.&lt;/li&gt;
&lt;li&gt;Category-strict recall is low; this is not a benchmark of absolute audit quality. Some ground-truth entries depend on runtime contracts (e.g., whether a config loader may return None) — they are predeclared seeded defects under this benchmark's stated assumptions, not universally adjudicated bugs.&lt;/li&gt;
&lt;li&gt;Structured output (JSON-schema-constrained decoding) modifies the logit distribution; the variance measured here is variance &lt;em&gt;under&lt;/em&gt; that constraint. Free-form review text may behave differently.&lt;/li&gt;
&lt;li&gt;No causal claims. Scope: free, local, small models, this corpus.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  Reproduce it
&lt;/h2&gt;

&lt;p&gt;Generation is nondeterministic and &lt;em&gt;not&lt;/em&gt; the reproduction target; the frozen per-run records are. The deterministic layer regenerates every table in this article:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/sumitsuke/llm-audit-nondeterminism
&lt;span class="nb"&gt;cd &lt;/span&gt;llm-audit-nondeterminism

&lt;span class="c"&gt;# Layer 2 (deterministic, no LLM needed): regenerate all numbers from frozen records&lt;/span&gt;
python scripts/aggregate.py            &lt;span class="c"&gt;# -&amp;gt; results/m2_summary.json&lt;/span&gt;
python scripts/verify_independent.py   &lt;span class="c"&gt;# independent recomputation of every number quoted in the text — should print ALL PASS&lt;/span&gt;
python scripts/make_figures.py         &lt;span class="c"&gt;# -&amp;gt; figures/nofreelunch.png (needs Matplotlib)&lt;/span&gt;

&lt;span class="c"&gt;# Layer 1 (optional, needs Ollama + ~5GB VRAM for 7B Q4_K_M):&lt;/span&gt;
&lt;span class="c"&gt;#   regenerating gives you a *different* distribution — that's the point&lt;/span&gt;
&lt;span class="c"&gt;#   OLLAMA_EXE=/path/to/ollama bash scripts/run_m2.sh&lt;/span&gt;
&lt;span class="c"&gt;#   python scripts/seed_probe.py       # confirm seeds have no observable effect at temp 0&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Numerical aggregation and verification need Python 3.12 and only the standard library; figure regeneration additionally needs Matplotlib. A few seconds total. (On Windows, set &lt;code&gt;PYTHONUTF8=1&lt;/code&gt;.)&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Pin everything and local LLM review was deterministic on this rig&lt;/strong&gt; — including deterministically repeating the same misses. Determinism ≠ correctness.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One review run is a sample, not a verdict.&lt;/strong&gt; At practical temperatures the finding set moves between runs, mostly by line drift.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Majority voting measures agreement strength, not truth.&lt;/strong&gt; Use multi-run review as a candidate generator; adjudicate with tests, static analysis, and humans.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you remember one sentence, please don't make it "LLM code review is useless" — that's not what the data says. Make it: &lt;strong&gt;multi-run review is a decent candidate generator; what's unsafe is automatic adjudication over a brittle vote key.&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;All per-run records (normalized findings + response hashes), ground truth, aggregation code, and the independent verification script: &lt;a href="https://github.com/sumitsuke/llm-audit-nondeterminism" rel="noopener noreferrer"&gt;github.com/sumitsuke/llm-audit-nondeterminism&lt;/a&gt;. Anything not measured (other GPUs, quantizations, larger models) is listed as not measured.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This is an English adaptation of &lt;a href="https://zenn.dev/tauridev/articles/llm-audit-nondeterminism" rel="noopener noreferrer"&gt;my Japanese article on Zenn&lt;/a&gt; — written by me in Japanese, restructured and translated with AI assistance, human-reviewed. If you spot an error, comments and issues are open; I'll verify against the frozen records and correct with a changelog.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Verification record (environment, verdict, last verified date, evidence) and the canonical write-up: &lt;a href="https://sumitsuke.jp/via/devto/lab/llm-audit-51-runs/" rel="noopener noreferrer"&gt;https://sumitsuke.jp/via/devto/lab/llm-audit-51-runs/&lt;/a&gt; — code, data and reproduction: &lt;a href="https://github.com/sumitsuke/llm-audit-nondeterminism" rel="noopener noreferrer"&gt;https://github.com/sumitsuke/llm-audit-nondeterminism&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>codereview</category>
      <category>testing</category>
    </item>
  </channel>
</rss>
