<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Dexterlung</title>
    <description>The latest articles on DEV Community by Dexterlung (@dexterlung).</description>
    <link>https://dev.to/dexterlung</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4029426%2F0d3f03e5-03d2-46ff-975f-c565c23e82ce.jpg</url>
      <title>DEV Community: Dexterlung</title>
      <link>https://dev.to/dexterlung</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/dexterlung"/>
    <language>en</language>
    <item>
      <title>I ran claude plugin eval on 89 skills. The trigger rate fell from 100% to 55.6%.</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Mon, 28 Sep 2026 13:05:19 +0000</pubDate>
      <link>https://dev.to/dexterlung/i-ran-claude-plugin-eval-on-89-skills-the-trigger-rate-fell-from-100-to-556-n29</link>
      <guid>https://dev.to/dexterlung/i-ran-claude-plugin-eval-on-89-skills-the-trigger-rate-fell-from-100-to-556-n29</guid>
      <description>&lt;h2&gt;
  
  
  The problem with skills is that failure is silent
&lt;/h2&gt;

&lt;p&gt;My coffee e-commerce project has 89 skill directories under it. Each one exists because I got burned: remember the GRANT on that migration, keep those three layers in sync when you rename a column, query the database before you guess at the code.&lt;/p&gt;

&lt;p&gt;Skills work by having the model read their descriptions and decide, on its own, whether to invoke one. That's the whole problem. &lt;strong&gt;When it doesn't invoke, nothing happens.&lt;/strong&gt; No error, no warning, no log line. You just watch the AI start working, and find out two days later that it skipped the rule.&lt;/p&gt;

&lt;p&gt;I measured this by hand once, back in July: four natural-language prompts, four sessions, counting how many times the skill name showed up in the reply. Sonnet hit 3 of 4, Haiku hit 1 of 4. I wrote the numbers into my project rules and never measured again, because doing it by hand is expensive and gives you a one-off number that goes stale.&lt;/p&gt;

&lt;p&gt;Last week Claude Code quietly shipped &lt;code&gt;claude plugin eval&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mechanism is simple. One design choice isn't.
&lt;/h2&gt;

&lt;p&gt;You write test cases — realistic sentences a user might type — plus graders that score the result. Each case runs in a fresh isolated session, three times by default because the model is non-deterministic.&lt;/p&gt;

&lt;p&gt;Standard stuff. What surprised me: &lt;strong&gt;it runs a second arm by default, with your plugin not loaded at all.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The difference between the two scores is &lt;code&gt;Δ&lt;/code&gt;. The docs put it bluntly:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If a case scores 1.0 both with and without the plugin, the plugin isn't what made it pass.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Vendors don't usually ship a default that proves their users' work is useless. This design concedes something: &lt;strong&gt;your skill might be doing nothing, and you'd have no way to know.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  First measurement: 100% with one skill loaded
&lt;/h2&gt;

&lt;p&gt;I reused the four prompts from July rather than writing new ones. Writing your own test corpus has a trap — you unconsciously write sentences your skill happens to catch. The old corpus was written before I knew what I'd be measuring.&lt;/p&gt;

&lt;p&gt;Three positive cases (add a back-in-stock notification, add a ship-date column, add a daily-summary page for staff) and one negative (change the homepage headline and button colour). One grader: did that skill get invoked, yes or no.&lt;/p&gt;

&lt;p&gt;Result: &lt;strong&gt;12 for 12.&lt;/strong&gt; All three positives fired, the negative correctly didn't. $2.49, 273 seconds.&lt;/p&gt;

&lt;p&gt;I knew immediately the number was useless, because the test plugin contained &lt;strong&gt;exactly one skill&lt;/strong&gt;. When there's only one candidate and its name matches the task, picking it costs the model nothing. My real environment has 89 competing for the same attention.&lt;/p&gt;

&lt;h2&gt;
  
  
  Second measurement: 55.6% with 89 loaded
&lt;/h2&gt;

&lt;p&gt;I wrote an assembly script that builds a throwaway plugin from all 89 skills at run time. (Deliberately no skill copies in version control — copies drift from the original, and this project has already paid that bill.)&lt;/p&gt;

&lt;p&gt;Same prompts, same model. The only variable is 1 skill versus 89:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Case&lt;/th&gt;
&lt;th&gt;1 skill&lt;/th&gt;
&lt;th&gt;89 skills&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Back-in-stock notification&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;67%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ship-date column&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;33%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Staff daily-summary page&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;67%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Homepage copy (should NOT fire)&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Positives combined: 5 of 9, &lt;strong&gt;55.6%&lt;/strong&gt;. The negative case stayed correct in both.&lt;/p&gt;

&lt;p&gt;That lands in the same range as my July hand-measurement (Sonnet 75%, Haiku 25%), which is mildly reassuring — two completely independent methods, one manual and one automated, agreeing.&lt;/p&gt;

&lt;p&gt;I had a conclusion forming: too many skills dilutes attention, time to prune.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then I did something stupid
&lt;/h2&gt;

&lt;p&gt;Before writing that conclusion I wanted one more thing: in the runs where the skill &lt;em&gt;didn't&lt;/em&gt; fire, &lt;strong&gt;what did the model invoke instead?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If it picked a more appropriate skill, that's correct delegation, not a defect. If it picked one with an overlapping description, the fix is editing descriptions, not deleting skills. Both leave fingerprints in the trace.&lt;/p&gt;

&lt;p&gt;Then I realised I had deleted every trace ten minutes earlier.&lt;/p&gt;

&lt;p&gt;The tool leaves a sandbox directory per run and prints "remove it when you're done." I cleaned them all up in one sweep. Those directories held the only thing that could answer the question.&lt;/p&gt;

&lt;p&gt;This was the expensive lesson of the day, and it's a dumb one: &lt;strong&gt;I deleted something far more valuable than tidiness, in order to be tidy.&lt;/strong&gt; The tool tells you to remove them because they may contain agent-written content. That is not the same as telling you to remove them before you've looked.&lt;/p&gt;

&lt;p&gt;Re-run with &lt;code&gt;--keep-temp&lt;/code&gt;: $0.82.&lt;/p&gt;

&lt;h2&gt;
  
  
  The traces: it didn't fire the skill, but it did the skill's job
&lt;/h2&gt;

&lt;p&gt;All three re-runs failed to invoke. Here's every tool call each one made:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;run 1: Glob, Glob, Glob, Glob, Glob, Grep
run 2: Glob, Glob, Glob, Glob, Glob
run 3: Glob, Glob, Glob, Glob, Glob
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Not one alternative skill invoked.&lt;/strong&gt; It went looking through files instead.&lt;/p&gt;

&lt;p&gt;So "correct delegation" and "overlapping descriptions" are both ruled out. But the thing that actually changed my conclusion was the final reply. All three looked like this (excerpt, translated):&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Based on this project's skill descriptions, here's what I expect to touch (&lt;strong&gt;this is expectation, not fact&lt;/strong&gt; — I'll confirm by reading files once you give me the path):&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;DB&lt;/strong&gt;: new migration adding &lt;code&gt;orders.estimated_ship_date&lt;/code&gt;, with the tenant_id / RLS / GRANT boilerplate&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write path&lt;/strong&gt;: whitelist the field in the admin order edit payload (don't spread formData raw)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Middle layer&lt;/strong&gt;: if there's an RPC &lt;code&gt;jsonb_build_object&lt;/code&gt; or a &lt;code&gt;normalizeOrder&lt;/code&gt; whitelist, add it there too, otherwise you get "value in DB, null in frontend"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Presentation&lt;/strong&gt;: &lt;code&gt;orderPresentation.js&lt;/code&gt; SSOT, plus the admin order list…&lt;/li&gt;
&lt;/ol&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those four points &lt;strong&gt;are what that skill would have produced.&lt;/strong&gt; GRANT boilerplate, no raw form spreading, the middle-layer whitelist that silently nulls your field, the five order surfaces that drift apart. It even labelled its own confidence level.&lt;/p&gt;

&lt;p&gt;It never invoked the Skill tool. It read the description and reasoned from it.&lt;/p&gt;

&lt;p&gt;So my 55.6% measures &lt;strong&gt;invocation&lt;/strong&gt;, not &lt;strong&gt;effect&lt;/strong&gt;. Those are separable in the trace, and I had assumed they were the same thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  A flaw in my test design, while I'm being honest
&lt;/h2&gt;

&lt;p&gt;All three traces show the model stuck in the same place: the working directory is empty, so it burns five Glob calls hunting for project files and ends with "please tell me where the project is."&lt;/p&gt;

&lt;p&gt;That's my fault. Every eval run starts in an empty workspace, and the tool supports seeding one with a scaffold script. I skipped it. An empty directory pushes the model's attention toward "find the files" instead of "is this an architectural change?"&lt;/p&gt;

&lt;p&gt;So the honest framing is: &lt;strong&gt;these numbers measure whether the model formally invokes a skill in an empty workspace.&lt;/strong&gt; That's some distance from the question I care about.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 0.00 that never ran
&lt;/h2&gt;

&lt;p&gt;One more trap, because it will fool anyone who doesn't read the surrounding columns.&lt;/p&gt;

&lt;p&gt;My first attempt at proving the grader could actually fail produced this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;run 1/1 [with]: score 0.00  $&lt;/span&gt;0.00  error: &lt;span class="nb"&gt;exit &lt;/span&gt;1: Credit balance is too low
&lt;span class="go"&gt;  ✗ skill-fired: Skill called 0x (expected 1..∞)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I wanted a red light, and it handed me a red light.&lt;/p&gt;

&lt;p&gt;But that 0.00 isn't "the grader caught something." &lt;strong&gt;The run never happened.&lt;/strong&gt; API credit ran out, the session never started, the grader searched an empty transcript, found no skill invocation, and rendered 0.&lt;/p&gt;

&lt;p&gt;(The credit thing turned out to be a separate wallet from my subscription — eval uses whatever credentials your normal sessions use, and &lt;code&gt;ANTHROPIC_API_KEY&lt;/code&gt; in my environment was routing it to the API account. Unsetting it for the child process fell back to subscription auth.)&lt;/p&gt;

&lt;p&gt;The real one, once it ran:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;run 1/1 [with]:    score 1.00  $&lt;/span&gt;0.30
&lt;span class="gp"&gt;run 1/1 [without]: score 0.00  $&lt;/span&gt;0.15   ← this 0.00 is real
&lt;span class="go"&gt;Δ +1.00
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two identical-looking zeros. &lt;strong&gt;The discriminator is the columns next to them&lt;/strong&gt;: the fake one cost $0.00 and carried an error string; the real one cost $0.15 with a null error.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why would Anthropic make this a first-party command?
&lt;/h2&gt;

&lt;p&gt;This is the part I find most worth thinking about, and I can only offer speculation.&lt;/p&gt;

&lt;p&gt;They could have published a "how to write good skill descriptions" guide for a fraction of the cost. Instead they built a measuring instrument — one whose default behaviour is to &lt;strong&gt;prove your work contributed nothing.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Three signals I can read from that:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One: this isn't a documentation problem.&lt;/strong&gt; If unclear descriptions were the issue, a style guide would fix it. Building a measurement tool implies they think this needs per-case empirical testing, and that results shift with model versions. The docs say to pin your model in CI so "a model rollout isn't mistaken for a plugin regression." That sentence concedes the same skill performs differently across models.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two: the ablation arm is the default, not a flag.&lt;/strong&gt; That's the strongest signal in the design. It assumes your plugin might not be contributing and computes the answer whether you asked or not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Three: this is a tool that only becomes necessary at a certain ecosystem size.&lt;/strong&gt; With five skills you don't need to measure. Past some threshold the question shifts from "do I have a skill for this?" to "will it win the attention contest?" I have 89. That number didn't exist two years ago.&lt;/p&gt;

&lt;p&gt;My guess is they saw this curve before I did. That's a guess — I have no inside information.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I know and what I don't
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Known:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Same model, same prompt: going from 1 skill to 89 drops formal invocation from 100% to 55.6%&lt;/li&gt;
&lt;li&gt;In the runs that didn't invoke, the model didn't pick a different skill — it read descriptions and reasoned from them&lt;/li&gt;
&lt;li&gt;The grader has discriminating power (Δ = +1.00 against the no-plugin arm, so it does go red)&lt;/li&gt;
&lt;li&gt;Total spend: $8.79, about 15 minutes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Not known:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Whether "smarter models need fewer skills" holds. &lt;strong&gt;Every run used the same model.&lt;/strong&gt; That variable was never manipulated. Testing it needs two models of different strength, same suite.&lt;/li&gt;
&lt;li&gt;How much &lt;em&gt;effect&lt;/em&gt; is lost when invocation is lost. The traces show descriptions still reaching the reasoning; how far "reached the reasoning" sits from "executed the skill's full checklist" is unmeasured.&lt;/li&gt;
&lt;li&gt;What these numbers become once the workspace is properly seeded.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you've got a pile of skills, prompts or rule files, my suggestion isn't to prune. It's to &lt;strong&gt;measure once&lt;/strong&gt; — and when you do, remember that "was the tool invoked" is the easy grader to write, and probably not the question you care about.&lt;/p&gt;

&lt;p&gt;As for whether I should cut my 89 down: I have no evidence either way yet. Next step is seeding the workspace and running two models. That's when I get to have an opinion.&lt;/p&gt;




&lt;p&gt;Related, from the same toolchain: &lt;a href="https://dev.to/content/completeness-baseline-en"&gt;Why does the AI always leave a few blocks out?&lt;/a&gt; — the other half of this problem, where the verifier checks whether your declarations are internally consistent and never whether they match the world. And &lt;a href="https://dev.to/content/crystallize-compile-ai-reasoning"&gt;compiling AI reasoning into scripts&lt;/a&gt; (in Chinese), on paying for frontier reasoning once and replaying it for free.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on my blog: &lt;a href="https://coffeeshooters.com/content/skill-fired-is-not-skill-worked-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-skill-fired-is-not-skill-worked-en" rel="noopener noreferrer"&gt;I ran claude plugin eval on 89 skills. The trigger rate fell from 100% to 55.6%.&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;I keep a running index of every pothole I've hit building a real production system solo — symptom on the left, what to grep in your own repo on the right: &lt;a href="https://coffeeshooters.com/potholes?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=potholes-index" rel="noopener noreferrer"&gt;coffeeshooters.com/potholes&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;And if your team is shipping AI-written code faster than anyone can read it, that's the thing I do for a living: &lt;a href="https://coffeeshooters.com/code-audit?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=code-audit-offer" rel="noopener noreferrer"&gt;coffeeshooters.com/code-audit&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>architecture</category>
      <category>testing</category>
    </item>
    <item>
      <title>Google crawled 270 of my articles and indexed none — the body it received was 0 characters</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Sun, 27 Sep 2026 00:22:02 +0000</pubDate>
      <link>https://dev.to/dexterlung/google-crawled-270-of-my-articles-and-indexed-none-the-body-it-received-was-0-characters-303h</link>
      <guid>https://dev.to/dexterlung/google-crawled-270-of-my-articles-and-indexed-none-the-body-it-received-was-0-characters-303h</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Read on: &lt;a href="https://dev.to/content/my-test-never-failed-because-it-proved-one-equals-one-en"&gt;the acceptance check that proved 1 = 1&lt;/a&gt; · &lt;a href="https://dev.to/content/google-crawled-not-indexed-spa-blog"&gt;繁體中文版&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I assumed my content wasn't good enough. Google had never read my content. What it received inside &lt;code&gt;&amp;lt;div id="app"&amp;gt;&lt;/code&gt; was zero characters.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The numbers that started it
&lt;/h2&gt;

&lt;p&gt;On 13 August 2026 I opened Google Search Console and exported two reports, 427 rows. Three lines stood out:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Status&lt;/th&gt;
&lt;th&gt;Rows&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Crawled – currently not indexed&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;270&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Soft 404&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;157&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Duplicate, Google chose a different canonical&lt;/td&gt;
&lt;td&gt;95&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;"Crawled – currently not indexed" has an unambiguous meaning: &lt;strong&gt;Google came, looked, and declined.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That is far worse than not being crawled. Not crawled means it hasn't arrived yet. Crawled and declined means it read the page and judged there was nothing worth indexing.&lt;/p&gt;

&lt;p&gt;My first instincts were keywords, depth, backlinks. All three were the wrong direction, and all three would have cost me another three months of writing articles nobody could find.&lt;/p&gt;

&lt;p&gt;There was only one correct first move: &lt;strong&gt;stop guessing what it sees and go be it.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  First move: fetch your own page as Googlebot
&lt;/h2&gt;

&lt;p&gt;I requested my own article URL with Googlebot's User-Agent, saved the raw HTML, and counted two things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;characters inside &lt;code&gt;&amp;lt;div id="app"&amp;gt;&lt;/code&gt; → &lt;strong&gt;0&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;JSON-LD blocks → &lt;strong&gt;0&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Repeated it with GPTBot. Identical. So besides Google, &lt;strong&gt;the AI-search side was blind too&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The cause is unsurprising to anyone who has shipped a single-page app: the first HTML the server sends is a shell, and the body appears only after JavaScript runs. Google says it executes JavaScript, which is true — but "will execute" and "will always wait for you to finish before scoring" are different promises.&lt;/p&gt;

&lt;p&gt;The part I want to emphasise isn't the conclusion. It's how long it took me to run that command. &lt;strong&gt;That HTML was one curl away the entire time. I wrote 270 articles before I fetched it once.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Fix 1: put the body in the first-hand HTML
&lt;/h2&gt;

&lt;p&gt;My site already had a Netlify edge function doing og-tag injection, which means it already reads the article from the database for social previews. So body injection needed no new architecture — three more things on the same trip:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;render the article body to semantic HTML and place it inside &lt;code&gt;&amp;lt;div id="app"&amp;gt;&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;emit JSON-LD (&lt;code&gt;BlogPosting&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;emit hreflang, and only when the other language version is actually published&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The human path is untouched; the branch is bot-only, so real readers pay nothing.&lt;/p&gt;

&lt;p&gt;One technical decision deserves its own paragraph: &lt;strong&gt;I did not reuse the frontend's markdown renderer.&lt;/strong&gt; It sanitises with DOMPurify, DOMPurify needs a DOM, and the Deno edge runtime has none. So I wrote an escape-first renderer on the edge side: escape every character, then apply structure. That ordering makes XSS &lt;strong&gt;structurally impossible&lt;/strong&gt; rather than "unlikely because I was careful."&lt;/p&gt;

&lt;p&gt;The cost lands on raw-HTML blocks, which degrade to plain text. I didn't estimate whether that mattered — I asked the database. Of 1,112 text blocks in published articles, &lt;strong&gt;1,111 use markdown and 1 uses raw HTML&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I'm pointing at that because it is the most useful discipline I picked up this year: &lt;strong&gt;when you are about to make a judgement about "all of the X", derive the scope from a query, not from memory.&lt;/strong&gt; Had I written "most of them are probably markdown" I would likely have been right — and I would not have known the 1 existed. (The same habit, learned the hard way, is what fixed &lt;a href="https://dev.to/content/supabase-default-acl-anon-truncate-en"&gt;a privilege sweep that reported zero residue while missing seven views&lt;/a&gt;.)&lt;/p&gt;

&lt;p&gt;Verification: 28 unit tests, 12 cross-layer contract tests, all green. Green isn't the evidence, though — I also built 9 deliberately broken variants and confirmed each turned exactly the check it should red, and that restoring made everything green again. &lt;strong&gt;9/9.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;One self-inflicted trap worth half an hour of your life: I used a sentinel character to mark inline-code boundaries, and the first version used a literal NUL byte. Git immediately classified the file as binary (&lt;code&gt;Bin 0 -&amp;gt; 21694 bytes&lt;/code&gt;) — no diff, no line-by-line review. An escape sequence carries identical semantics and stays readable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fix 2: 157 soft 404s, because the site answered 200 to everything
&lt;/h2&gt;

&lt;p&gt;The second number is worse, because it grows by itself.&lt;/p&gt;

&lt;p&gt;I picked two URLs out of the GSC export and requested them by hand:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/product/panama-esmerlada-geisha-washed/  → 200 + homepage HTML
/product-tag/washed-processed/            → 200 + homepage HTML
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both are dead WordPress-era URLs. But the last line of my &lt;code&gt;_redirects&lt;/code&gt; was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;/*&lt;/span&gt;  &lt;span class="n"&gt;/index.html&lt;/span&gt;  &lt;span class="mi"&gt;200&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every SPA deployment guide tells you to add that line — without it, refreshing any sub-path 404s. Its side effect: &lt;strong&gt;every dead URL is also caught by it, and answered 200.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Google receives a 200 and something that looks like a homepage, and files it as a soft 404. Every mistyped or mis-pasted link manufactures another one.&lt;/p&gt;

&lt;p&gt;Two layers of fix:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;edge function&lt;/strong&gt;: for &lt;code&gt;/content/:slug&lt;/code&gt;, &lt;code&gt;/product/:id&lt;/code&gt;, &lt;code&gt;/works/:slug&lt;/code&gt;, when the record definitively does not exist, return status 404. The body is still the SPA, so the visitor still sees the in-site "page not found" screen; &lt;strong&gt;only the status code becomes honest.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;_redirects&lt;/code&gt;&lt;/strong&gt;: legacy WooCommerce/WordPress archive paths return 410 Gone. &lt;strong&gt;Not a 301 to &lt;code&gt;/shop&lt;/code&gt;&lt;/strong&gt; — in Google's eyes that is still a soft 404, just reported somewhere else. Seven legacy &lt;code&gt;.md&lt;/code&gt; URLs get a 301 stripping the extension.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The load-bearing part isn't the 404, it's fail-open
&lt;/h3&gt;

&lt;p&gt;This is the easiest thing in the batch to get wrong, and getting it wrong produces no visible symptom.&lt;/p&gt;

&lt;p&gt;My data helper &lt;code&gt;fetchRows&lt;/code&gt; returned an empty array for both "query failed" and "zero rows found". While it only did injection, that was harmless — both cases meant "don't inject."&lt;/p&gt;

&lt;p&gt;The moment "no record → return 404" existed, the same code meant something else entirely:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;database hiccups for 3 seconds → my real article is declared nonexistent → that 404 lands in Google's index.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The fix is to return &lt;code&gt;{ ok, rows }&lt;/code&gt; and treat &lt;strong&gt;only ok-and-zero&lt;/strong&gt; as genuinely missing.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;/product/:id&lt;/code&gt; needed one more layer: decide locally whether the string even looks like a uuid. A non-uuid sent to PostgREST returns 400, and 400 is "query failed", which fail-open lets through. The 82 legacy WooCommerce permalinks I needed to block are all non-uuid. &lt;strong&gt;They must be judged locally, not shipped out as a question.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Four mutations for this piece: remove fail-open, remove the uuid check, drop a header cleanup in the 404 branch, revert &lt;code&gt;fetchRows&lt;/code&gt; to a bare array.&lt;/p&gt;

&lt;p&gt;The third one &lt;strong&gt;did not go red the first time.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;My assertion sliced the file contents to the end, so it matched a &lt;code&gt;headers.delete&lt;/code&gt; further down in the injection path rather than the one in the 404 branch. Classic "the substring had another source" — the test appears to check one thing and is being fed by somewhere else. Slicing to just the 404 branch made it 4/4.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I did not think that defect up. The mutation exposed it.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Fix 3: one stylesheet costing 2,222 ms
&lt;/h2&gt;

&lt;p&gt;Indexing handled, on to speed. Local Lighthouse against one of my article pages:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Render-blocking resource&lt;/th&gt;
&lt;th&gt;Delay&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;fonts.googleapis.com/css2?Noto+Sans+TC&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2,222 ms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;98 KiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;assets/index-*.css&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;341 ms&lt;/td&gt;
&lt;td&gt;26 KiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;assets/vendor-formkit-*.css&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;6 KiB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One Google Fonts stylesheet, 2,222 ms on its own. Of those 98 KiB, &lt;strong&gt;92 KiB are unicode-range subset declarations I never use&lt;/strong&gt; (CJK fonts split into dozens of subsets, one declaration each).&lt;/p&gt;

&lt;p&gt;I thought I had handled this, because the HTML contains &lt;code&gt;&amp;lt;link rel="preload" as="style"&amp;gt;&lt;/code&gt;. But the load-bearing line is the &lt;code&gt;rel="stylesheet"&lt;/code&gt; right beneath it — &lt;strong&gt;preload only makes the browser fetch earlier; it does not take the stylesheet off the critical rendering path.&lt;/strong&gt; With both present, the slow one is still slow.&lt;/p&gt;

&lt;p&gt;The fix swaps &lt;code&gt;rel&lt;/code&gt; to &lt;code&gt;stylesheet&lt;/code&gt; on load. No flash of invisible text, because the font URL already carries &lt;code&gt;display=swap&lt;/code&gt; and the splash screen uses system fonts. JavaScript-disabled clients get the ordinary tag inside &lt;code&gt;&amp;lt;noscript&amp;gt;&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;And afterwards I went back and confirmed all three injection anchors from Fix 1 were still intact. &lt;strong&gt;Touching fonts shouldn't affect body injection — but "shouldn't" and "confirmed" are different words.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Fix 4: a 531K PDF bundle preloaded on every page
&lt;/h2&gt;

&lt;p&gt;The live &lt;code&gt;index.html&lt;/code&gt; had a &lt;code&gt;modulepreload&lt;/code&gt; line pointing at &lt;code&gt;vendor-pdf-*.js&lt;/code&gt;, 531K.&lt;/p&gt;

&lt;p&gt;The three pages that use PDF (poster studio, sticker generator, cupping station) are all lazy routes. Their dependencies have no business in the first screen.&lt;/p&gt;

&lt;p&gt;Reading the build output made it clear: &lt;strong&gt;every&lt;/strong&gt; page chunk in &lt;code&gt;dist&lt;/code&gt; — including "page not found" and the newsletter page — carried a bare &lt;code&gt;import"./vendor-pdf-*.js"&lt;/code&gt;, and that chunk exports three identifiers nobody binds. So this wasn't "somebody needs pdf." It was my manual chunking rule, matching on strings, &lt;strong&gt;fighting Rollup's own splitting&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Fix: return &lt;code&gt;undefined&lt;/code&gt; for those three packages and hand the decision back to Rollup.&lt;/p&gt;

&lt;p&gt;⚠️ The first version just deleted the line, and made things &lt;strong&gt;worse&lt;/strong&gt;: the &lt;code&gt;return 'vendor'&lt;/code&gt; fallback below caught them and moved them into the core vendor bundle, which loads earlier. Core vendor went 499K → 1,057K and precache 2,398 → 2,930 KiB. You have to &lt;code&gt;return undefined&lt;/code&gt; explicitly, before the fallback.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That was caught by verifying the build output, not by thinking.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;⚠️ The first verification script was also wrong. It asserted "the pdf chunk must have no static importer", which flagged two legitimate cases — a lazy route's chunk statically importing a shared chunk is normal, because that whole subtree only loads on navigation. The real criterion is &lt;strong&gt;whether the transitive static-import closure from the entry point reaches it&lt;/strong&gt;. Computing the closure fixed it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A guard that barks at legitimate change is as bad as one that never barks.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Final numbers:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;modulepreload&lt;/td&gt;
&lt;td&gt;6 (incl. 531K vendor-pdf)&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;core vendor&lt;/td&gt;
&lt;td&gt;499K&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;419K&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;precache&lt;/td&gt;
&lt;td&gt;2,398 KiB&lt;/td&gt;
&lt;td&gt;2,319 KiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;first-screen static closure&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;6 chunks / 1,461K, pdf not among them&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Then I shipped a new defect: the title, three times
&lt;/h2&gt;

&lt;p&gt;With four layers done, I thought the batch was finished.&lt;/p&gt;

&lt;p&gt;The next day I looked at a real screenshot of the live article page. It opened with:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;title → date → summary → title (as a figure caption) → title (h1) → summary&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Title three times, summary twice.&lt;/strong&gt; Shipped by my own previous commit.&lt;/p&gt;

&lt;p&gt;The cause is simple and unflattering: &lt;strong&gt;I designed a layout without reading the frontend component first.&lt;/strong&gt; Its &lt;code&gt;&amp;lt;article&amp;gt;&lt;/code&gt; contains exactly one line, a &lt;code&gt;&amp;lt;BlockRenderer&amp;gt;&lt;/code&gt; — the title, hero image and summary are &lt;em&gt;already inside the blocks&lt;/em&gt;. My injector then added an h1, a date, a summary and a hero image with a caption. All four duplicates.&lt;/p&gt;

&lt;p&gt;The fix is to render the blocks and nothing else — faithfully mirroring the existing renderer.&lt;/p&gt;

&lt;h3&gt;
  
  
  The test was an accomplice
&lt;/h3&gt;

&lt;p&gt;I had a test asserting "the output contains exactly one h1". It was green throughout, while production grew three.&lt;/p&gt;

&lt;p&gt;The reason was the fixture: its blocks happened to contain no title block. So the assertion held forever, over data that didn't resemble reality.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When the fixture doesn't look like the real thing, the test will vouch for the defect.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I gave the fixture a hero block (the real shape) and added a regression test specifically for duplication. Incidentally, my mutation tool caught its own stale config in that round — the string it was supposed to mutate no longer existed, so it reported "this proof is invalid" instead of counting it as a pass. &lt;strong&gt;That is what a guard should do.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Final acceptance: 34 unit tests, prove-red 11/11.&lt;/p&gt;

&lt;p&gt;Two sentences carry this section:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;An injection layer's job is to mirror the existing renderer faithfully, not to redesign the layout.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Measuring" is not "looking."&lt;/strong&gt; No Lighthouse score will ever show you a title printed three times. A screenshot showed it in three seconds.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  A footnote: a byline can't be a shell
&lt;/h2&gt;

&lt;p&gt;In the JSON-LD I changed the author type from &lt;code&gt;Organization&lt;/code&gt; to &lt;code&gt;Person&lt;/code&gt; to feed E-E-A-T person signals.&lt;/p&gt;

&lt;p&gt;But I left the brand name in the &lt;code&gt;name&lt;/code&gt; field. Google now saw "a person named after a brand." &lt;strong&gt;Right type, wrong load-bearing field.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I changed it to my real name and added an assertion that can actually fail: &lt;code&gt;author.name !== SITE_NAME&lt;/code&gt;. Before that, the test only checked &lt;code&gt;@type === 'Person'&lt;/code&gt; — &lt;strong&gt;which stays green if you put the brand name back, and that is exactly the defect this change existed to fix.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  If you run an SPA content site, check in this order
&lt;/h2&gt;

&lt;p&gt;Twenty minutes, roughly:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Fetch one of your own article pages with Googlebot's User-Agent&lt;/strong&gt; and count characters inside &lt;code&gt;&amp;lt;div id="app"&amp;gt;&lt;/code&gt;. Cheapest cut, and the origin of everything else.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Export the soft 404 list from GSC and request two of them by hand.&lt;/strong&gt; If your &lt;code&gt;_redirects&lt;/code&gt; has &lt;code&gt;/* → /index.html 200&lt;/code&gt;, you already know the answer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run Lighthouse and read only the render-blocking section.&lt;/strong&gt; On a CJK site the top entry is very likely a font.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open the live &lt;code&gt;index.html&lt;/code&gt; and read the &lt;code&gt;modulepreload&lt;/code&gt; lines.&lt;/strong&gt; Ask each one whether the first screen truly needs it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;When you're done, look at a real screenshot.&lt;/strong&gt; Not the score. The picture.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The last thing I'll leave here: what actually saved me in this batch was not any individual fix — it was &lt;strong&gt;prove-red&lt;/strong&gt;, deliberately breaking the code and watching whether the check turned red. It caught three false green lights in this batch alone (an assertion matching a string from elsewhere, a guard flagging legitimate cases, a fixture that didn't look real), and not one of those was avoidable by being more careful.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the numbers come from
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;GSC exports, 2026-08-13, 427 rows: crawled–not-indexed 270 / soft 404 157 / duplicate 95&lt;/li&gt;
&lt;li&gt;Live fetches as Googlebot and GPTBot: &lt;code&gt;&amp;lt;div id="app"&amp;gt;&lt;/code&gt; 0 characters, JSON-LD 0&lt;/li&gt;
&lt;li&gt;Local Lighthouse 13, target &lt;code&gt;/content/the-mechanism-i-never-checked&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Build output: modulepreload 6→5, core vendor 499K→419K, precache 2,398→2,319 KiB&lt;/li&gt;
&lt;li&gt;Prove-red totals: body injection 9/9, 404 honesty 4/4, layout duplication 11/11&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on my blog: &lt;a href="https://coffeeshooters.com/content/google-crawled-not-indexed-spa-blog-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-google-crawled-not-indexed-spa-blog-en" rel="noopener noreferrer"&gt;Google crawled 270 of my articles and indexed none — the body it received was 0 characters&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;I keep a running index of every pothole I've hit building a real production system solo — symptom on the left, what to grep in your own repo on the right: &lt;a href="https://coffeeshooters.com/potholes?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=potholes-index" rel="noopener noreferrer"&gt;coffeeshooters.com/potholes&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;And if your team is shipping AI-written code faster than anyone can read it, that's the thing I do for a living: &lt;a href="https://coffeeshooters.com/code-audit?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=code-audit-offer" rel="noopener noreferrer"&gt;coffeeshooters.com/code-audit&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>debugging</category>
      <category>webdev</category>
    </item>
    <item>
      <title>RLS says yes and Postgres still says permission denied: the 403 family I only understood on the second one</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Sat, 26 Sep 2026 05:05:10 +0000</pubDate>
      <link>https://dev.to/dexterlung/rls-says-yes-and-postgres-still-says-permission-denied-the-403-family-i-only-understood-on-the-1f3p</link>
      <guid>https://dev.to/dexterlung/rls-says-yes-and-postgres-still-says-permission-denied-the-403-family-i-only-understood-on-the-1f3p</guid>
      <description>&lt;h2&gt;
  
  
  RLS says yes and Postgres still says &lt;code&gt;permission denied&lt;/code&gt;: the 403 family I only understood on the second one
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;Read on: &lt;a href="https://dev.to/content/supabase-default-acl-anon-truncate-en"&gt;the layer of privileges RLS cannot reach at all&lt;/a&gt; · &lt;a href="https://dev.to/content/privileged-writer-invoker-403-family"&gt;繁體中文版&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;An admin tops up a customer's balance in the back office and the screen returns &lt;code&gt;permission denied for table users&lt;/code&gt;. My RLS policy explicitly allows admins. The door is open and I can't get through — because the thing blocking me was not the door I was looking at.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The symptom: the admin is locked out of his own system
&lt;/h2&gt;

&lt;p&gt;My coffee shop runs on stored value: customers top up, then draw down. Besides customers topping up themselves, an admin can top up or deduct on a customer's behalf from the POS.&lt;/p&gt;

&lt;p&gt;Testing the admin path, I got:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;permission denied for table users
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;First reaction: impossible. My RLS policy has an &lt;code&gt;is_admin()&lt;/code&gt; clause. The admin's identity check passes. The door is open — why can't I walk through?&lt;/p&gt;

&lt;p&gt;I spent a while fiddling with the policy. Nothing helped. I was looking at the wrong door.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cause: RLS governs &lt;em&gt;which rows&lt;/em&gt;, GRANT governs &lt;em&gt;which columns&lt;/em&gt;
&lt;/h2&gt;

&lt;p&gt;The sentence that unlocked it: &lt;strong&gt;RLS allowing you is not the same as you being allowed to write.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Postgres permissions are two independent gates:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Table / column level GRANT&lt;/strong&gt; — may this role touch this table, these columns, at all?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;RLS policy&lt;/strong&gt; — given that it may, &lt;em&gt;which rows&lt;/em&gt; may it touch?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I had been staring at gate 2. Gate 1 was the one stopping me.&lt;/p&gt;

&lt;p&gt;My top-up RPC was &lt;code&gt;SECURITY INVOKER&lt;/code&gt;, meaning it runs &lt;strong&gt;as the caller&lt;/strong&gt;. In the frontend the admin is the &lt;code&gt;authenticated&lt;/code&gt; role. And the privileged columns on &lt;code&gt;users&lt;/code&gt; — &lt;code&gt;balance&lt;/code&gt;, &lt;code&gt;wallet&lt;/code&gt;, &lt;code&gt;member_level&lt;/code&gt;, the &lt;code&gt;tier_*&lt;/code&gt; fields, &lt;code&gt;annual_spent&lt;/code&gt; — &lt;strong&gt;were deliberately never granted to &lt;code&gt;authenticated&lt;/code&gt;&lt;/strong&gt;. I had restricted them to &lt;code&gt;service_role&lt;/code&gt; and a few DEFINER functions precisely so that no ordinary logged-in user could ever write their own balance or tier.&lt;/p&gt;

&lt;p&gt;So the chain was: admin → &lt;code&gt;authenticated&lt;/code&gt; → INVOKER RPC → the RPC writes &lt;code&gt;users.balance&lt;/code&gt; still as &lt;code&gt;authenticated&lt;/code&gt; → no column grant → &lt;strong&gt;permission denied&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;is_admin()&lt;/code&gt; cleared the admin at the row level. Execution never got that far; it died one layer below, at the column level. The error message says &lt;code&gt;permission denied for table users&lt;/code&gt; and tells you nothing about which of the two gates rejected you. That is the worst part of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: a DEFINER wrapper with the guard on the first line
&lt;/h2&gt;

&lt;p&gt;What I actually needed: &lt;strong&gt;admins can write these columns, ordinary users cannot.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Granting the privileged columns to &lt;code&gt;authenticated&lt;/code&gt; is the wrong fix — that opens balance writes to every logged-in user. I wanted "a call that has passed an admin check temporarily gains the ability to write privileged columns."&lt;/p&gt;

&lt;p&gt;Postgres's tool for that is &lt;code&gt;SECURITY DEFINER&lt;/code&gt;: the function runs as its &lt;strong&gt;owner&lt;/strong&gt;, not the caller. My DEFINER functions are owned by postgres, which can write anything.&lt;/p&gt;

&lt;p&gt;So, wrap:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;create&lt;/span&gt; &lt;span class="k"&gt;function&lt;/span&gt; &lt;span class="n"&gt;fn_admin_topup_balance&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;span class="k"&gt;returns&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt;
&lt;span class="k"&gt;language&lt;/span&gt; &lt;span class="n"&gt;plpgsql&lt;/span&gt;
&lt;span class="k"&gt;security&lt;/span&gt; &lt;span class="k"&gt;definer&lt;/span&gt;          &lt;span class="c1"&gt;-- runs as owner (postgres), can write privileged columns&lt;/span&gt;
&lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="err"&gt;$$&lt;/span&gt;
&lt;span class="k"&gt;begin&lt;/span&gt;
  &lt;span class="n"&gt;if&lt;/span&gt; &lt;span class="k"&gt;not&lt;/span&gt; &lt;span class="n"&gt;is_admin&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;then&lt;/span&gt;                        &lt;span class="c1"&gt;-- guard FIRST&lt;/span&gt;
    &lt;span class="n"&gt;raise&lt;/span&gt; &lt;span class="n"&gt;exception&lt;/span&gt; &lt;span class="s1"&gt;'not authorized'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;end&lt;/span&gt; &lt;span class="n"&gt;if&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="n"&gt;perform&lt;/span&gt; &lt;span class="n"&gt;fn_topup_balance&lt;/span&gt;&lt;span class="p"&gt;(...);&lt;/span&gt;                &lt;span class="c1"&gt;-- call the original INVOKER function&lt;/span&gt;
&lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="err"&gt;$$&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;revoke&lt;/span&gt; &lt;span class="k"&gt;all&lt;/span&gt; &lt;span class="k"&gt;on&lt;/span&gt; &lt;span class="k"&gt;function&lt;/span&gt; &lt;span class="n"&gt;fn_admin_topup_balance&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;anon&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;grant&lt;/span&gt; &lt;span class="k"&gt;execute&lt;/span&gt; &lt;span class="k"&gt;on&lt;/span&gt; &lt;span class="k"&gt;function&lt;/span&gt; &lt;span class="n"&gt;fn_admin_topup_balance&lt;/span&gt; &lt;span class="k"&gt;to&lt;/span&gt; &lt;span class="n"&gt;authenticated&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The ordering is the whole design. &lt;strong&gt;&lt;code&gt;is_admin()&lt;/code&gt; must be the first statement.&lt;/strong&gt; A DEFINER function is superuser-powered from the moment control enters its body, so the guard has to run before anything touches data. With that in place, &lt;code&gt;authenticated&lt;/code&gt; may &lt;em&gt;execute&lt;/em&gt; the function, while the body admits only admins. The frontend's admin call site points at the new function; the customer self-service path and the &lt;code&gt;service_role&lt;/code&gt; path are untouched.&lt;/p&gt;

&lt;p&gt;I fixed two this way: &lt;code&gt;fn_admin_topup_balance&lt;/code&gt; and &lt;code&gt;fn_admin_deduct_balance&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The turn: fixing the second one told me it was a family
&lt;/h2&gt;

&lt;p&gt;I fixed top-up. The same afternoon I hit the identical 403 on deduction. Identical fix.&lt;/p&gt;

&lt;p&gt;That is when I stopped and thought about something else: if two identical bugs surface on the same day, &lt;strong&gt;it isn't two bugs, it's a family&lt;/strong&gt;. There must be other INVOKER functions writing privileged &lt;code&gt;users&lt;/code&gt; columns as &lt;code&gt;authenticated&lt;/code&gt;, sitting there untriggered because nobody has walked that path yet.&lt;/p&gt;

&lt;p&gt;Memory won't enumerate that, and neither will grep. I wrote a scanner that asks the catalogue directly — &lt;code&gt;pg_proc&lt;/code&gt; — with three conditions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;prosecdef = false&lt;/code&gt; (INVOKER, not DEFINER)&lt;/li&gt;
&lt;li&gt;the function source writes one of the privileged &lt;code&gt;users&lt;/code&gt; columns&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;authenticated&lt;/code&gt; has EXECUTE on it&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Four hits. Two were the ones I had just fixed (into the allowlist, with the reason recorded: "guarded by fn_admin_*"). The other two were cron maintenance functions: &lt;code&gt;fn_expire_bonuses&lt;/code&gt; and &lt;code&gt;fn_reconcile_balances&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Not every member of a family gets the same fix
&lt;/h2&gt;

&lt;p&gt;Those two cron functions did &lt;strong&gt;not&lt;/strong&gt; get the DEFINER + &lt;code&gt;is_admin()&lt;/code&gt; treatment, because they should never be called from the frontend at all. They are scheduled tasks; the scheduler runs as superuser and needs no &lt;code&gt;authenticated&lt;/code&gt; privileges. The one frontend call site I found for &lt;code&gt;expireBonuses()&lt;/code&gt; turned out to be dead code with no callers.&lt;/p&gt;

&lt;p&gt;So their fix is the opposite direction: &lt;code&gt;REVOKE&lt;/code&gt; from &lt;code&gt;authenticated&lt;/code&gt;, leave &lt;code&gt;service_role&lt;/code&gt;. Narrow the grant instead of adding a gate. After deploying I queried the live ACLs to confirm only postgres and &lt;code&gt;service_role&lt;/code&gt; can reach them.&lt;/p&gt;

&lt;p&gt;That is the part worth underlining: &lt;strong&gt;a bug family does not imply a single remedy.&lt;/strong&gt; What admins legitimately need gets a guarded privilege escalation. What nothing outside should ever touch gets its privileges taken away. The scanner's job is finding every member; deciding the remedy is still one judgement per member.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three things to take away
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;On Supabase/Postgres, when RLS allows it and you still get &lt;code&gt;permission denied&lt;/code&gt;, check the column-level GRANT before you touch the policy.&lt;/strong&gt; They are two independent gates and the error message will not tell you which one rejected you.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;When you need "admins can write this, users can't", wrap it in SECURITY DEFINER with &lt;code&gt;is_admin()&lt;/code&gt; as the first statement.&lt;/strong&gt; Inside a DEFINER body you are superuser — no unauthorised call may reach the line that touches data.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The second identical bug is the signal to write a scanner.&lt;/strong&gt; Querying a system catalogue like &lt;code&gt;pg_proc&lt;/code&gt; converts "are there others?" from an anxiety into a question with an answer. Allowlist the known-good with a written reason; whatever remains is the real work.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I still run that scanner as part of my health check. The next time somebody — including me — writes an INVOKER function that quietly writes a balance, it raises its hand before the deploy instead of after the 403.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on my blog: &lt;a href="https://coffeeshooters.com/content/privileged-writer-invoker-403-family-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-privileged-writer-invoker-403-family-en" rel="noopener noreferrer"&gt;RLS says yes and Postgres still says permission denied: the 403 family I only understood on the second one&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;I keep a running index of every pothole I've hit building a real production system solo — symptom on the left, what to grep in your own repo on the right: &lt;a href="https://coffeeshooters.com/potholes?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=potholes-index" rel="noopener noreferrer"&gt;coffeeshooters.com/potholes&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;And if your team is shipping AI-written code faster than anyone can read it, that's the thing I do for a living: &lt;a href="https://coffeeshooters.com/code-audit?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=code-audit-offer" rel="noopener noreferrer"&gt;coffeeshooters.com/code-audit&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>debugging</category>
      <category>webdev</category>
    </item>
    <item>
      <title>My most important test had never failed — because what it proved was 1 = 1</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Tue, 22 Sep 2026 13:05:14 +0000</pubDate>
      <link>https://dev.to/dexterlung/my-most-important-test-had-never-failed-because-what-it-proved-was-1-1-21d</link>
      <guid>https://dev.to/dexterlung/my-most-important-test-had-never-failed-because-what-it-proved-was-1-1-21d</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Read on: &lt;a href="https://dev.to/content/noise-floor-two-identical-runs-en"&gt;measuring the noise floor before trusting a delta&lt;/a&gt; · &lt;a href="https://dev.to/content/my-test-never-failed-because-it-proved-one-equals-one"&gt;繁體中文版&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Do this first — it is uncomfortable and it takes five minutes
&lt;/h2&gt;

&lt;p&gt;Pick the test or check script you have &lt;strong&gt;the most confidence in&lt;/strong&gt;. Then break the thing it claims to protect, on purpose.&lt;/p&gt;

&lt;p&gt;Not something adjacent. The actual load-bearing thing: make the function return &lt;code&gt;null&lt;/code&gt; forever, push the threshold somewhere unreachable, empty the data source it reads.&lt;/p&gt;

&lt;p&gt;If it stays green, what you own is not an acceptance check. It is a decoration — and it is worse than having nothing, because it is why you stopped worrying.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;I keep a notebook of production potholes — a year and a bit of markdown files recording what broke and how I fixed it. The problem with a notebook is that it only helps when somebody thinks to open it, and the most expensive class of mistake is defined by nobody feeling any need to check: the wrong road is smooth the whole way down.&lt;/p&gt;

&lt;p&gt;So I built a tool. Give it a task description, it surfaces the relevant potholes.&lt;/p&gt;

&lt;p&gt;Its most important feature is the red-flag list — marking which entries are the ones that &lt;strong&gt;produce no signal at the moment you get them wrong&lt;/strong&gt;. I wrote an acceptance check for it and called it, in the commit message, "the most important check in this tool":&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The notebook declares 48 red flags → the tool must be able to display 48.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Reasonable, right? 48 against 48. If the numbers match, no red flag got dropped.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then I handed it to a different model
&lt;/h2&gt;

&lt;p&gt;I asked a model &lt;strong&gt;from a different lab&lt;/strong&gt; to review it. (Different lab matters. Models from the same family share blind spots; three of them looking together is still one pair of eyes.)&lt;/p&gt;

&lt;p&gt;The first thing it did was not read my code. It ran a mutation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;make the fuzzy-match function return null unconditionally
→ selftest still 14 passed, 0 failed
→ the reconciliation still reads "48 / 48 ✅"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I re-ran it myself. &lt;strong&gt;Reproduced exactly.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it could never go red
&lt;/h2&gt;

&lt;p&gt;Unfolded, it's obvious:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// numerator: how many rows are flagged red in the notebook&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;declared&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;flag&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;🔴&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;

&lt;span class="c1"&gt;// denominator: how many the tool can actually display&lt;/span&gt;
&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;shown&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
&lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;file&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;files&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;shown&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nf"&gt;computeRedFlags&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;file&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And &lt;code&gt;computeRedFlags&lt;/code&gt; is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;row&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;row&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;flag&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;🔴&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;continue&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;hit&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;findSection&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;row&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;      &lt;span class="c1"&gt;// found → carries a line number&lt;/span&gt;
  &lt;span class="nx"&gt;allRed&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt;&lt;span class="nx"&gt;row&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;line&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;hit&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nx"&gt;hit&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;line&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;   &lt;span class="c1"&gt;// pushed either way&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;It pushes every row unconditionally.&lt;/strong&gt; Found gets a line number, not-found gets &lt;code&gt;null&lt;/code&gt; — but both count as "can display."&lt;/p&gt;

&lt;p&gt;So &lt;code&gt;shown&lt;/code&gt; and &lt;code&gt;declared&lt;/code&gt; are two spellings of the same filter. They are mathematically incapable of differing. I had not written an acceptance check. I had written &lt;code&gt;1 === 1&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;And the failure mode is at its most toxic here: &lt;strong&gt;it was always green, and green is what stopped me looking.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  I made the same shape of mistake three times in one day
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;First&lt;/strong&gt;: the tautology above.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second&lt;/strong&gt;: I rewrote it as an end-to-end assertion — run a real query, check the &lt;strong&gt;output string&lt;/strong&gt; contains the red flag and its line number:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nf"&gt;say&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;out&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;TRUNCATE&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;out&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;:257&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ran the mutation. &lt;strong&gt;Still green.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Because &lt;code&gt;:257&lt;/code&gt; was arriving from a &lt;em&gt;different line&lt;/em&gt; of the output. My output has two sections: the red-flag list (affected by fuzzy matching) and "closest entries by literal match" (unaffected). With fuzzy matching broken, the red-flag line did lose its line number — and the section below printed &lt;code&gt;:257&lt;/code&gt; anyway.&lt;/p&gt;

&lt;p&gt;I was asserting on "does this whole blob of output contain this substring", and &lt;strong&gt;the substring had two possible sources&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Third&lt;/strong&gt;: while fixing the counting logic I found a bug (a last-wins map dropping two entries), fixed it in the main code — and then &lt;strong&gt;made the identical mistake inside the test itself&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the working version looks like
&lt;/h2&gt;

&lt;p&gt;Three things.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Bind the assertion to the load-bearing line, not to the whole output.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;redLine&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;out&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;\r?\n&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;find&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;l&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;l&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trimStart&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;startsWith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;·&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;l&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;needle&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;ok&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="o"&gt;!!&lt;/span&gt;&lt;span class="nx"&gt;redLine&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="sr"&gt;/`:&lt;/span&gt;&lt;span class="se"&gt;\d&lt;/span&gt;&lt;span class="sr"&gt;+`/&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;redLine&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;2. Do not bind to something that legitimately changes.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;My first version asserted the line number equals 257. Then I added a few lines to the top of that file, it became 261, and the test went red.&lt;/p&gt;

&lt;p&gt;That was not catching a defect. That was &lt;strong&gt;an assertion bound to something allowed to move&lt;/strong&gt;. A guard that barks at legitimate change is as bad as one that never barks — both train you to ignore it.&lt;/p&gt;

&lt;p&gt;Assert that there &lt;strong&gt;is&lt;/strong&gt; a line number, not that it is a particular one. That still survives the mutation: fuzzy match broken → that row prints a cross-entry label → no line number → red.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Every check must name the mutation that turns it red.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I ended up labelling each assertion with the mutation it guards:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What I broke on purpose&lt;/th&gt;
&lt;th&gt;Which check goes red&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;fuzzy match returns null forever&lt;/td&gt;
&lt;td&gt;end-to-end #3 (16 pass / 1 fail)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;put the &lt;code&gt;if (!x.length) continue&lt;/code&gt; back&lt;/td&gt;
&lt;td&gt;the P0 check (16 pass / 1 fail)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;push the first-stage threshold to 9999&lt;/td&gt;
&lt;td&gt;13 failures&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;delete the plain-language line in the file header&lt;/td&gt;
&lt;td&gt;real-corpus 7/7 drops to 2/7 (19 pass / 5 fail)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;change nothing&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;24 pass / 0 fail&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;A check whose red-making mutation you cannot name is a candidate tautology.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  One more trap: a prove-it-goes-red run must control its input
&lt;/h2&gt;

&lt;p&gt;Later the same day I wrote a sentinel monitoring whether the tool's zero-hit rate had gotten too high. To prove it fires, I appended 16 fake zero-hit records to the real log.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It did not go red.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The file already held 241 records from my own testing. The ratio moved from 46.9% to 49.0% — nowhere near the 70% threshold.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A prove-red stacked on real data gets diluted by the existing denominator.&lt;/strong&gt; The right move is to lift the judgement into a pure function and feed it fully controlled input:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;mk&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;n&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;zeros&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;Array&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="na"&gt;length&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;n&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;hits&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nx"&gt;zeros&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="p"&gt;}))&lt;/span&gt;
&lt;span class="c1"&gt;// ★ red: 13 of 16 at zero hits (81%) → must exit 1&lt;/span&gt;
&lt;span class="c1"&gt;// ★ boundary: exactly 70% does not count as over → exit 0; 75% → exit 1&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Why having the rule written down did not save me
&lt;/h2&gt;

&lt;p&gt;The part that bothers me most: my project documentation &lt;strong&gt;already contained this rule&lt;/strong&gt;, verbatim:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A thing that emits a green light must positively observe the load-bearing thing itself.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That day I &lt;strong&gt;quoted it&lt;/strong&gt;, believed I had &lt;strong&gt;followed it&lt;/strong&gt;, and then got the same shape wrong three times.&lt;/p&gt;

&lt;p&gt;Written down, correctly placed, remembered — and still no help.&lt;/p&gt;

&lt;p&gt;What did help was: &lt;strong&gt;a model from a different lab, breaking my code with mutation, and watching whether my check made a sound.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The lesson isn't "be more careful." It is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"I feel confident about this check" carries no evidential weight whatsoever. The only thing that does is "I broke what it protects and it went red."&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The one-minute version
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Pick the check you trust most&lt;/li&gt;
&lt;li&gt;Break the thing it claims to protect, on purpose&lt;/li&gt;
&lt;li&gt;Still green → it is a decoration, fix it now&lt;/li&gt;
&lt;li&gt;Then label each check with &lt;strong&gt;the mutation that makes it red&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Bind assertions to the load-bearing line — not the whole output, not something allowed to change&lt;/li&gt;
&lt;li&gt;Run prove-red on fully controlled input, never stacked on real data&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;By the way — after step 3 you may find more than one. I found three.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on my blog: &lt;a href="https://coffeeshooters.com/content/my-test-never-failed-because-it-proved-one-equals-one-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-my-test-never-failed-because-it-proved-one-equals-one-en" rel="noopener noreferrer"&gt;My most important test had never failed — because what it proved was 1 = 1&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;I keep a running index of every pothole I've hit building a real production system solo — symptom on the left, what to grep in your own repo on the right: &lt;a href="https://coffeeshooters.com/potholes?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=potholes-index" rel="noopener noreferrer"&gt;coffeeshooters.com/potholes&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;And if your team is shipping AI-written code faster than anyone can read it, that's the thing I do for a living: &lt;a href="https://coffeeshooters.com/code-audit?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=code-audit-offer" rel="noopener noreferrer"&gt;coffeeshooters.com/code-audit&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>debugging</category>
      <category>webdev</category>
    </item>
    <item>
      <title>31 articles vanished from my site. No error, no log — one hardcoded .limit(80).</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Sun, 20 Sep 2026 13:30:11 +0000</pubDate>
      <link>https://dev.to/dexterlung/31-articles-vanished-from-my-site-no-error-no-log-one-hardcoded-limit80-29c6</link>
      <guid>https://dev.to/dexterlung/31-articles-vanished-from-my-site-no-error-no-log-one-hardcoded-limit80-29c6</guid>
      <description>&lt;h2&gt;
  
  
  31 articles vanished from my site. No error, no log — one hardcoded &lt;code&gt;.limit(80)&lt;/code&gt;.
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;Read on: &lt;a href="https://dev.to/content/my-test-never-failed-because-it-proved-one-equals-one-en"&gt;the test that never failed because it proved 1 = 1&lt;/a&gt; · &lt;a href="https://dev.to/content/hardcoded-limit-silent-data-loss"&gt;繁體中文版&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I had not deleted a single article, and yet 31 of them were completely gone from the site. Not in the series cards, not reachable through the category filters, not findable in site search. No error. No log entry. No trace.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The symptom: the data is there, the people are gone
&lt;/h2&gt;

&lt;p&gt;My coffee brand's site has a content section with a few years of accumulated writing — early coffee notes from 2020 through recent dev journals.&lt;/p&gt;

&lt;p&gt;One day, tidying up, I noticed the earliest batch of coffee posts seemed to be missing from the front end. The series card count was short. Clicking into the category showed nothing. Searching a keyword I knew was in the title returned nothing.&lt;/p&gt;

&lt;p&gt;I was quite sure I had never deleted them. I queried the database: all present, &lt;code&gt;status&lt;/code&gt; published, nothing unusual.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The rows exist, the front end can't see them.&lt;/strong&gt; That is a worse feeling than "the data was deleted", because there is nowhere to start. No error message to search for, no stack trace to follow. It isn't broken. It is quietly short.&lt;/p&gt;

&lt;h2&gt;
  
  
  The root cause: &lt;code&gt;.limit(80)&lt;/code&gt; truncates in silence
&lt;/h2&gt;

&lt;p&gt;I opened the query in my list page component and found this line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;80&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Written long ago, when I had maybe sixty articles. I counted the real number of published posts that day: &lt;strong&gt;111&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;There it is. The frontend fetches 80, and the other 31 — mostly the 2020–2023 coffee posts, because of the sort order — were never fetched at all. They weren't "fetched but not rendered." They were &lt;strong&gt;absent from the frontend's data from the first millisecond&lt;/strong&gt;. Series counts, category bucketing, and site search were all operating correctly over a corpus that was missing a chunk.&lt;/p&gt;

&lt;p&gt;The reason &lt;code&gt;.limit(80)&lt;/code&gt; is so nasty is that &lt;strong&gt;it does not error&lt;/strong&gt;. It isn't "too much data, give up, throw." It is "here are the first 80, pretend the rest don't exist." That is silent truncation, and it belongs to the most dangerous bug family there is — every component looks like it is working, because every component &lt;em&gt;is&lt;/em&gt; working. The input was already wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stopgap: change 80 to 200 and move on, for now
&lt;/h2&gt;

&lt;p&gt;Having found the cause, I did not immediately do the correct, larger fix. I stopped the bleeding:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One line, thirty seconds, 31 articles back.&lt;/p&gt;

&lt;p&gt;I knew 200 is not a cure — today it is 111, and someday it is 201 and the identical bug returns identically. But stopping the bleeding and curing the disease are two different jobs: &lt;strong&gt;31 articles were invisible to real visitors right now&lt;/strong&gt;, and the priority in that moment was getting them back, not shipping perfect architecture an hour later.&lt;/p&gt;

&lt;p&gt;An expedient bump isn't a sin. The sin is bumping it and forgetting, which converts it into the next time bomb. So I queued the real fix at the same moment — and did it the same day.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real fix: delete the premise that the frontend loads everything
&lt;/h2&gt;

&lt;p&gt;Three changes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. A paged RPC on the server.&lt;/strong&gt; A new &lt;code&gt;fn_get_content_list_page&lt;/code&gt; taking &lt;code&gt;limit&lt;/code&gt; / &lt;code&gt;offset&lt;/code&gt;. The frontend loads the first 30 and appends as you scroll. Whether the total is 111 or 1,111, the frontend never tries to swallow it whole. While I was there, the same RPC took over full-text search (including matching legacy English slugs) and computed reading time in the database instead of in the browser.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Counts separated from the list.&lt;/strong&gt; The old count was &lt;code&gt;pages.value.filter(...).length&lt;/code&gt; — deriving "how many posts are in this category" from &lt;em&gt;the articles already loaded into the frontend&lt;/em&gt;. That is two bugs in one: it couples load volume and count correctness to the same ceiling. Truncate the load and the count silently follows. I replaced it with an independent &lt;code&gt;countRows&lt;/code&gt; query that asks the database directly, selecting two narrow columns and moving no content. &lt;strong&gt;"How many are there" and "which ones do I load" are two different questions and should never share a data source.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. The list stopped carrying full article bodies.&lt;/strong&gt; The old list query dragged the whole &lt;code&gt;content&lt;/code&gt; jsonb for every article, so the index page was hauling the full text of every post on the site. Now it returns summary, cover image and reading time only. Payload per batch went from 348,954 bytes to 22,114 — about 94% less.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verification: "they're back" is not a number
&lt;/h2&gt;

&lt;p&gt;Fixing vanished data has a trap: it is very easy to glance at the page, see more articles, and call it done. I wanted the numbers to &lt;strong&gt;match exactly&lt;/strong&gt;, not to "look fuller."&lt;/p&gt;

&lt;p&gt;Two independent lines, crossed:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;SQL counts&lt;/strong&gt; — ask the database how many per category and per series.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Playwright on a real browser&lt;/strong&gt; — boot the dev server, actually click the category buttons, actually scroll to trigger the next page, and count what is rendered.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Before the fix, the category buttons summed to 8 + 51 + 6 = 65, against a database total of 111. After: 23 + 70 + 13 = 106, plus 5 known cross-category exceptions — exactly 111, matching to the unit. Playwright also confirmed the "30 then 60 more" path produced no duplicates and no gaps.&lt;/p&gt;

&lt;p&gt;Numbers matching is what let me say it was fixed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three things to take away
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Any frontend &lt;code&gt;.limit(N)&lt;/code&gt; over a list needs an answer to "what happens past N."&lt;/strong&gt; Usually the answer is silent truncation: no error, just quietly less data. It is the hardest class to diagnose precisely because every part of the system looks healthy.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;"How many" and "which ones" are different questions.&lt;/strong&gt; Never use the length of the loaded array as a count. It goes wrong in lockstep with the load ceiling. Ask the database for &lt;code&gt;COUNT&lt;/code&gt;, page the list, keep them apart.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;After restoring missing data, reconcile per bucket against database counts.&lt;/strong&gt; Exact match or it isn't fixed — "looks like more now" isn't a verification. If you can, drive the real load path with a browser rather than trusting a unit test that seeds its own array.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That &lt;code&gt;.limit(80)&lt;/code&gt; sat in my code for years and only detonated when the article count quietly crossed 80. Which is the actual lesson: &lt;strong&gt;a hardcoded number never tells you it has expired. It just starts lying while you aren't looking.&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on my blog: &lt;a href="https://coffeeshooters.com/content/hardcoded-limit-silent-data-loss-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-hardcoded-limit-silent-data-loss-en" rel="noopener noreferrer"&gt;31 articles vanished from my site. No error, no log — one hardcoded .limit(80).&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;I keep a running index of every pothole I've hit building a real production system solo — symptom on the left, what to grep in your own repo on the right: &lt;a href="https://coffeeshooters.com/potholes?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=potholes-index" rel="noopener noreferrer"&gt;coffeeshooters.com/potholes&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;And if your team is shipping AI-written code faster than anyone can read it, that's the thing I do for a living: &lt;a href="https://coffeeshooters.com/code-audit?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=code-audit-offer" rel="noopener noreferrer"&gt;coffeeshooters.com/code-audit&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>debugging</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Supabase had granted anon TRUNCATE on 96 tables. I never wrote that line, and every check I own was green.</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Sat, 19 Sep 2026 14:00:58 +0000</pubDate>
      <link>https://dev.to/dexterlung/supabase-had-granted-anon-truncate-on-96-tables-i-never-wrote-that-line-and-every-check-i-own-was-46oj</link>
      <guid>https://dev.to/dexterlung/supabase-had-granted-anon-truncate-on-96-tables-i-never-wrote-that-line-and-every-check-i-own-was-46oj</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Read on: &lt;a href="https://dev.to/content/dead-code-more-dangerous-non-homologous-diagnosis-en"&gt;the blind spot a non-homologous model caught&lt;/a&gt; · &lt;a href="https://dev.to/content/supabase-default-acl-anon-truncate"&gt;繁體中文版&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Do this first, it takes fifteen seconds
&lt;/h2&gt;

&lt;p&gt;Paste this into your Supabase SQL editor:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;grantee&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;privilege_type&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;information_schema&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;role_table_grants&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;table_schema&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'public'&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;grantee&lt;/span&gt; &lt;span class="k"&gt;IN&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'anon'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'authenticated'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;privilege_type&lt;/span&gt; &lt;span class="k"&gt;IN&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'TRUNCATE'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'TRIGGER'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If it comes back non-empty, welcome. Mine came back with &lt;strong&gt;anon holding TRUNCATE on 96 tables, authenticated on 109&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I had not written a single one of those grants.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I found it
&lt;/h2&gt;

&lt;p&gt;I was building a podcast feature and idly checked the privileges on my articles table — I only wanted to confirm anon could read published posts. What I saw in anon's list was INSERT, UPDATE, DELETE, &lt;strong&gt;TRUNCATE&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The first three didn't worry me much. They're governed by RLS, my policies are fail-closed, and anon has no write policy at all. Ugly, but unreachable.&lt;/p&gt;

&lt;p&gt;TRUNCATE is a different animal. &lt;strong&gt;PostgreSQL's row-level security applies to exactly four verbs: SELECT, INSERT, UPDATE, DELETE. TRUNCATE is not one of them.&lt;/strong&gt; It empties an entire table in one statement and it answers only to table-level privileges. The whole RLS wall I spent a year building sits &lt;em&gt;above&lt;/em&gt; this one. TRIGGER is the same — whoever holds it can attach a trigger, which is a path to executing code as the table owner.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the grants came from
&lt;/h2&gt;

&lt;p&gt;Supabase's &lt;code&gt;public&lt;/code&gt; schema ships with a default ACL entry (&lt;code&gt;pg_default_acl&lt;/code&gt;) that says: every table created in this schema from now on automatically grants &lt;code&gt;arwdDxtm&lt;/code&gt; to anon and authenticated. The &lt;code&gt;D&lt;/code&gt; in that string is TRUNCATE. The &lt;code&gt;t&lt;/code&gt; is TRIGGER.&lt;/p&gt;

&lt;p&gt;So nobody slipped. &lt;strong&gt;The platform made the decision before I wrote my first line of SQL.&lt;/strong&gt; And revoking on existing tables isn't enough — the next table you create grows it back.&lt;/p&gt;

&lt;p&gt;Let me be honest about severity before this reads like a scare piece: this is &lt;strong&gt;not&lt;/strong&gt; "your database gets wiped tomorrow." PostgREST has no TRUNCATE verb and no create-trigger endpoint, so a public anon key can't reach it over the API. You need an actual database connection. What it is, is a &lt;strong&gt;hole in the foundation of your defence in depth&lt;/strong&gt;: you believe the worst case is bounded by RLS, and two verbs are not governed by RLS at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why every check I own was green
&lt;/h2&gt;

&lt;p&gt;I'm not without defences. I have a pre-commit checker whose whole job is verifying that every migration creating a table also writes its three GRANT lines. It has a self-test. It has a blindness signal. It was green every single time.&lt;/p&gt;

&lt;p&gt;It was green &lt;em&gt;correctly&lt;/em&gt;. What it checks is &lt;strong&gt;what I wrote&lt;/strong&gt;: it reads migration files as text and confirms the GRANT lines are present. And the entirety of this hole is &lt;strong&gt;what I never wrote&lt;/strong&gt;. Reading your declarations cannot, in principle, see past their edge.&lt;/p&gt;

&lt;p&gt;The sentence I ended up writing on the wall:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A checker reads your declarations. The nastiest defects live in what you never declared — platform defaults, upstream package behaviour, the difference between your laptop and production.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's the same family as &lt;a href="https://dev.to/content/antivirus-said-clean-malware-was-right-there-en"&gt;an antivirus reporting "no malware detected" while the file sat right there on the disk&lt;/a&gt;: the tool answered the question it was asked, and the question had a blind side.&lt;/p&gt;

&lt;h2&gt;
  
  
  The more embarrassing half: I made the same mistake again while fixing it
&lt;/h2&gt;

&lt;p&gt;I wrote the repair migration the same day. A loop over every table in &lt;code&gt;public&lt;/code&gt;, revoking TRUNCATE and TRIGGER from anon and authenticated, with a self-verification at the end: count the residue, roll back unless it's zero.&lt;/p&gt;

&lt;p&gt;It ran. Verification reported &lt;strong&gt;0 rows of residue&lt;/strong&gt;. Green.&lt;/p&gt;

&lt;p&gt;The next day I looked again with an independent query: &lt;strong&gt;seven views still held TRUNCATE.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The reason is small and horrible. My revoke loop scanned &lt;code&gt;relkind IN ('r','p')&lt;/code&gt; — ordinary and partitioned tables, the list my head produces when it hears "table". My verification query joined &lt;code&gt;pg_tables&lt;/code&gt;, a system view that also lists only ordinary tables. &lt;strong&gt;The fix missed views, and the verification was structurally incapable of seeing views.&lt;/strong&gt; They shared one blind spot, so the verification could never falsify the fix.&lt;/p&gt;

&lt;p&gt;Notice how fine the trap is: both the fix and the verification really did query the live database. "Go ask the real world" sounds like the cure, and it is not sufficient — when I asked the world, the &lt;em&gt;scope&lt;/em&gt; of the question was still typed from memory. The query was real. The range was invented.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three rules I took out of it
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Derive scope from a query, never from memory.&lt;/strong&gt; Before acting on "all the tables", ask the database what actually exists:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;DISTINCT&lt;/span&gt; &lt;span class="n"&gt;relkind&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;pg_class&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;relnamespace&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'public'&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;regnamespace&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Generate the repair loop from that output. Don't type &lt;code&gt;('r','p')&lt;/code&gt; because that is what "table" feels like.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The acceptance query may not contain a single word the invariant doesn't contain.&lt;/strong&gt; My invariant was "anon must hold no TRUNCATE." That sentence contains no concept of &lt;em&gt;kind of relation&lt;/em&gt; — so the acceptance query had no business joining &lt;code&gt;pg_tables&lt;/code&gt; at all. Asking &lt;code&gt;information_schema.role_table_grants&lt;/code&gt; directly (it naturally includes views) is the correct answer. Every extra predicate in an acceptance query is an assumption smuggled in from memory.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Every "too little" check has a "too much" twin — ask for it while you're there.&lt;/strong&gt; My green checker guarded against &lt;em&gt;missing&lt;/em&gt; grants causing a 403 in the frontend. Its dual — is anybody holding too much? — had never been assigned to any checker in over a year. That pairing exists nearly everywhere: guard against missed sends and the twin is double sends; guard against running late and the twin is running early. Asking costs almost nothing at the moment you write the first one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix, copy-pasteable
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- 1. Revoke on what exists (including views! check your own relkinds first)&lt;/span&gt;
&lt;span class="k"&gt;DO&lt;/span&gt; &lt;span class="err"&gt;$$&lt;/span&gt;
&lt;span class="k"&gt;DECLARE&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="n"&gt;RECORD&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;BEGIN&lt;/span&gt;
  &lt;span class="k"&gt;FOR&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="k"&gt;IN&lt;/span&gt;
    &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;relname&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;pg_class&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;
    &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;relnamespace&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'public'&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;regnamespace&lt;/span&gt;
      &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="k"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;relkind&lt;/span&gt; &lt;span class="k"&gt;IN&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'r'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'p'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'v'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'m'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="n"&gt;LOOP&lt;/span&gt;
    &lt;span class="k"&gt;BEGIN&lt;/span&gt;
      &lt;span class="k"&gt;EXECUTE&lt;/span&gt; &lt;span class="n"&gt;format&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'REVOKE TRUNCATE, TRIGGER ON public.%I FROM anon, authenticated'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;relname&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="n"&gt;EXCEPTION&lt;/span&gt; &lt;span class="k"&gt;WHEN&lt;/span&gt; &lt;span class="n"&gt;OTHERS&lt;/span&gt; &lt;span class="k"&gt;THEN&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;  &lt;span class="c1"&gt;-- extension-owned objects you cannot touch&lt;/span&gt;
    &lt;span class="k"&gt;END&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;END&lt;/span&gt; &lt;span class="n"&gt;LOOP&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;END&lt;/span&gt; &lt;span class="err"&gt;$$&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- 2. Stop new tables from growing it back&lt;/span&gt;
&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;DEFAULT&lt;/span&gt; &lt;span class="k"&gt;PRIVILEGES&lt;/span&gt; &lt;span class="k"&gt;FOR&lt;/span&gt; &lt;span class="k"&gt;ROLE&lt;/span&gt; &lt;span class="n"&gt;postgres&lt;/span&gt; &lt;span class="k"&gt;IN&lt;/span&gt; &lt;span class="k"&gt;SCHEMA&lt;/span&gt; &lt;span class="k"&gt;public&lt;/span&gt;
  &lt;span class="k"&gt;REVOKE&lt;/span&gt; &lt;span class="k"&gt;TRUNCATE&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;TRIGGER&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;TABLES&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;anon&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;authenticated&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;-- 3. Acceptance: zero joins, zero extra predicates — exactly as wide as the invariant&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;grantee&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;privilege_type&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;information_schema&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;role_table_grants&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;table_schema&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'public'&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;grantee&lt;/span&gt; &lt;span class="k"&gt;IN&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'anon'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'authenticated'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;privilege_type&lt;/span&gt; &lt;span class="k"&gt;IN&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'TRUNCATE'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'TRIGGER'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="c1"&gt;-- expected: empty&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two footnotes. First, &lt;code&gt;pg_default_acl&lt;/code&gt; usually holds &lt;strong&gt;two&lt;/strong&gt; entries — one owned by &lt;code&gt;postgres&lt;/code&gt;, one by &lt;code&gt;supabase_admin&lt;/code&gt;. You cannot alter the second (I get &lt;code&gt;permission denied&lt;/code&gt;), but in my project &lt;code&gt;supabase_admin&lt;/code&gt; owns zero user objects, so that dirty ACL has never actually applied. Check yours with &lt;code&gt;SELECT relname, pg_get_userbyid(relowner) FROM pg_class WHERE relnamespace='public'::regnamespace&lt;/code&gt;. Second, cleaning once is not staying clean — I turned the acceptance query into a scheduled check that runs before every deploy, because the only thing that catches a privilege which regrows is asking again on a timer.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it cost me to learn
&lt;/h2&gt;

&lt;p&gt;This hole existed in my project for over a year. During that year I wrote dozens of governance rules, a dozen checkers, thousands of tests. All green. All honest. All inspecting things I had written.&lt;/p&gt;

&lt;p&gt;It was found because I glanced at something while building an unrelated feature. That is part of the lesson: for the class of defect that lives in what you never wrote, there is no checker you can build by reading your own source. Only two moments catch it — &lt;strong&gt;the moment you adopt a platform, when you enumerate the decisions it made on your behalf&lt;/strong&gt;, and &lt;strong&gt;a scheduled question asked of the live system&lt;/strong&gt; rather than of your files.&lt;/p&gt;

&lt;p&gt;96 is 0 now.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on my blog: &lt;a href="https://coffeeshooters.com/content/supabase-default-acl-anon-truncate-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-supabase-default-acl-anon-truncate-en" rel="noopener noreferrer"&gt;Supabase had granted anon TRUNCATE on 96 tables. I never wrote that line, and every check I own was green.&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;I keep a running index of every pothole I've hit building a real production system solo — symptom on the left, what to grep in your own repo on the right: &lt;a href="https://coffeeshooters.com/potholes?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=potholes-index" rel="noopener noreferrer"&gt;coffeeshooters.com/potholes&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;And if your team is shipping AI-written code faster than anyone can read it, that's the thing I do for a living: &lt;a href="https://coffeeshooters.com/code-audit?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=code-audit-offer" rel="noopener noreferrer"&gt;coffeeshooters.com/code-audit&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>debugging</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Two identical runs scored 89 and 89. Two cases had flipped.</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Fri, 11 Sep 2026 14:05:10 +0000</pubDate>
      <link>https://dev.to/dexterlung/two-identical-runs-scored-89-and-89-two-cases-had-flipped-3in3</link>
      <guid>https://dev.to/dexterlung/two-identical-runs-scored-89-and-89-two-cases-had-flipped-3in3</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Read on: &lt;a href="https://dev.to/content/slimming-skill-descriptions-ab-test-en"&gt;the A/B test this corrects&lt;/a&gt; · &lt;a href="https://dev.to/content/noise-floor-two-identical-runs"&gt;繁體中文版&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Two weeks ago I published an A/B test: I cut 41 AI tools' self-descriptions roughly in half, then ran a behavioural question bank against both versions to check that trigger rate hadn't dropped. Before: 88/96. After: 90/96.&lt;/p&gt;

&lt;p&gt;I wrote this sentence about it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;±2 cases at this sample size is noise, so I am &lt;strong&gt;not&lt;/strong&gt; claiming it got better.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A reader named Vinh Nguyen replied, and the reply amounted to one question: &lt;strong&gt;how do you know?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;He was precise about it. Three votes per case, majority wins, so any case sitting near the 2-1 boundary can flip on sampling alone. There was a between-arm delta with no within-arm noise floor to measure it against. And "three up, one down" — exactly what my four movers looked like — is the shape churn produces.&lt;/p&gt;

&lt;p&gt;He was right. That sentence was an assertion, not a measurement. It is also the same move the post spends two thousand words criticising: claiming a property of a system without observing it. I made it in the one paragraph where I was grading my own work.&lt;/p&gt;

&lt;p&gt;So I ran his experiment. The part I didn't expect is that it cuts in two directions at once.&lt;/p&gt;

&lt;h2&gt;
  
  
  First I went looking for the answer I'd already paid for
&lt;/h2&gt;

&lt;p&gt;Before spending anything, I checked whether I could bound the churn retroactively from the run I had already done. The harness makes three calls per case and takes the majority, so it necessarily computes "2 of 3" or "3 of 3" somewhere. Count the 2-of-3 cases and you have an upper bound on how many could flip — for free.&lt;/p&gt;

&lt;p&gt;I couldn't. And the reason has a nastier shape than "it wasn't recorded."&lt;/p&gt;

&lt;p&gt;The harness &lt;strong&gt;does&lt;/strong&gt; compute the vote split, and it &lt;strong&gt;does&lt;/strong&gt; print it — for cases that failed. A case that passed 2-1 prints nothing at all. The JSON output aggregates to per-skill pass/total and drops every split.&lt;/p&gt;

&lt;p&gt;So the one shape that actually flips under sampling — the narrow pass — is precisely the shape no layer of that tool had ever shown me. The instrument was blind in the exact direction that produces reassurance. Failures got scrutiny; lucky wins got a checkmark.&lt;/p&gt;

&lt;p&gt;I fixed that first, because it is the cheap half and it makes every future round free to interpret. Per-case vote counts now survive into the JSON, and the human-readable run ends with a count of non-unanimous cases. Then I checked the fix the way I should have checked the original claim: forced every case to 2-of-3 and confirmed the new line moved from "0 split" to "4 split." It reads the real votes. It isn't printing a constant.&lt;/p&gt;

&lt;h2&gt;
  
  
  The baseline
&lt;/h2&gt;

&lt;p&gt;Same commit, same question bank, same model, same machine, same afternoon. The unchanged arm, twice, with nothing touched in between. 96 cases, 288 calls per run.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Run 1&lt;/th&gt;
&lt;th&gt;Run 2&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Total passed&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;89 / 96&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;89 / 96&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cases that flipped between the two runs&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Read those two rows together, because the gap between them is the whole point. &lt;strong&gt;The score is identical. Two cases flipped anyway&lt;/strong&gt; — one up, one down, cancelling exactly.&lt;/p&gt;

&lt;p&gt;If I had run the same arm twice and looked only at the total, I would have concluded the harness was perfectly stable. It isn't. It has a floor of about 2, and the total-score view is structurally incapable of showing it, because flips in opposite directions erase each other before they reach the number I was reading.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;So my +2 sits exactly on the floor.&lt;/strong&gt; The sentence in my post happened to land on the right answer for a reason I had not earned.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the noise actually lives
&lt;/h2&gt;

&lt;p&gt;This is the part I'd want if I were reading someone else's post.&lt;/p&gt;

&lt;p&gt;My question bank has three kinds of case per skill: should-trigger, should-NOT-trigger, and a boundary case. Splitting the two flips by kind:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Case kind&lt;/th&gt;
&lt;th&gt;Count&lt;/th&gt;
&lt;th&gt;Flipped&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;should-trigger&lt;/td&gt;
&lt;td&gt;48&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;should-NOT-trigger&lt;/td&gt;
&lt;td&gt;24&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;boundary&lt;/td&gt;
&lt;td&gt;24&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;All of the variance is in the negatives.&lt;/strong&gt; Zero flips in 48 positive cases, zero in 24 boundary cases, both flips in the 24 negatives.&lt;/p&gt;

&lt;p&gt;In hindsight it's obvious. "This sentence is obviously asking for that tool" is a strongly determined decision. "This sentence should not reach for that tool" is a decision made against 116 competitors, and it's soft.&lt;/p&gt;

&lt;p&gt;Which means my question bank spends 75% of its budget on the categories carrying none of the noise. &lt;strong&gt;If I want a future A/B to be readable, the fix is more negative cases, not more rounds.&lt;/strong&gt; I would not have guessed that, and no amount of staring at the total would have suggested it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bad news, which is also the interesting news
&lt;/h2&gt;

&lt;p&gt;I flagged one skill in the original post as "the one to watch" — the one I had compressed hardest, which lost a case.&lt;/p&gt;

&lt;p&gt;It's still down. Both of today's runs put it at 3/4, and the case it loses comes back &lt;strong&gt;0 of 3&lt;/strong&gt; — unanimous failure, twice, seventeen days later. And it is a &lt;em&gt;positive&lt;/em&gt; case, which is the category where the noise produced zero flips all day.&lt;/p&gt;

&lt;p&gt;To be sure that signature meant what I thought, I ran a deliberate control. I took six skills whose descriptions the original disciplined trim could only compress by under 20% — meaning they had almost no redundancy left — and force-cut each by roughly 64%, the same ratio as the one that broke. If over-trimming has a signature, this should produce it.&lt;/p&gt;

&lt;p&gt;It did, loudly. Those six went from 23/24 and 22/24 down to &lt;strong&gt;19/24&lt;/strong&gt;, and one of them collapsed from 4/4 to 1/4. Three separate cases fell from 3/3 to 0/3 or 1/3, &lt;strong&gt;every one of them positive or boundary&lt;/strong&gt; — none in the negatives where the noise lives.&lt;/p&gt;

&lt;p&gt;So real damage and churn do not look alike once you can see the votes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Churn&lt;/strong&gt; is a &lt;em&gt;negative&lt;/em&gt; case wobbling between 1-of-3 and 2-of-3.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Damage&lt;/strong&gt; is a &lt;em&gt;positive&lt;/em&gt; case collapsing to unanimous failure and staying there.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A total score cannot tell those apart. That's not a subtle limitation. In my case the aggregate said "no systematic decline," and it was true, &lt;em&gt;and&lt;/em&gt; it was concealing a genuine regression, because one real loss and several noise-driven gains summed to approximately nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The threshold question
&lt;/h2&gt;

&lt;p&gt;The other commenter, Kane Lim, proposed that the 64% case suggested a practical threshold — a point where compression starts removing semantic signal rather than redundant text.&lt;/p&gt;

&lt;p&gt;The control answers this, and sharpens it. The same 64% cut cost one case on the skill in the original post, and three cases on a skill in the control. &lt;strong&gt;The ratio was identical. The outcome was not.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The difference is what the 64% removed. The skill I could compress that far with discipline could be compressed that far &lt;em&gt;because&lt;/em&gt; most of its 689 characters were restatement of a six-layer rule. On a description that is already tight, reaching 64% means cutting into signal — there is nothing else left to take.&lt;/p&gt;

&lt;p&gt;Which inverts how I had been reading the number. &lt;strong&gt;Compression ratio isn't a risk driver. It's a measurement of how much redundancy was there to begin with&lt;/strong&gt; — and I had been treating a symptom as a cause. The two variables aren't merely confounded, they're mechanically linked: how far a disciplined trim can go &lt;em&gt;is&lt;/em&gt; the redundancy. Across my 24 skills the achievable ratio ran from 3% to 64%, and exactly one exceeded 49%.&lt;/p&gt;

&lt;p&gt;And a correction I owe the original post. I wrote: &lt;em&gt;"everything I cut by 30% held steady. The only one that moved was the one I cut by 64%."&lt;/em&gt; That's wrong on my own data. The four skills that moved were cut by 64%, 49%, 48%, and &lt;strong&gt;17%&lt;/strong&gt;. The 17% one moved by exactly as much as the 64% one. The table proving it sat two paragraphs above the sentence.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd take from this
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Run the same arm twice before you compare two arms.&lt;/strong&gt; It cost about a dollar. It reclassified two of my conclusions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A total score is a lossy summary, and it's lossy in a specific direction:&lt;/strong&gt; opposite-direction flips cancel, so the aggregate systematically &lt;em&gt;understates&lt;/em&gt; instability. Two runs scoring 89 and 89 hid two real flips.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Record the vote split per case, not just pass/fail.&lt;/strong&gt; Anything that takes a majority is discarding the number that says how close it was. And if your tooling prints the split only for failures — mine did — then the flip-prone narrow &lt;em&gt;pass&lt;/em&gt; is the one thing you can never see.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sort your cases by kind and check where the variance actually sits.&lt;/strong&gt; Mine was 100% in one category holding 25% of the budget.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A stable unanimous failure is readable at n=1.&lt;/strong&gt; You don't always need a bigger sample; sometimes you need a less lossy instrument. 0-of-3 twice is a signal. 2-of-3 once is not.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What this doesn't establish
&lt;/h2&gt;

&lt;p&gt;Two runs give a point estimate of the floor, not a distribution — I know the floor is about 2, not that it is 2. The over-cut control was a single arm, single round; its individual collapses are strong, but its total is subject to the same ±2. The six skills I over-cut were selected by a proxy — "the disciplined trim could only compress them under 20%" — which measures how tight a description is, not how narrow a skill's scope is, and scope was Kane's actual question. I also chose &lt;em&gt;how&lt;/em&gt; to cut them, keeping the opening trigger sentence and dropping identifiers and symptom phrases; a different cut might land differently.&lt;/p&gt;

&lt;p&gt;One last thing, since it's the same failure as everything above. When I went to report what these runs cost, I found the harness doesn't print its cost block in JSON mode — that code sits in the other branch. So the actual spend for the three runs was never recorded, and the roughly US$1.2 I quoted is extrapolated from a single-skill run. Third instance in one day of the same defect: &lt;strong&gt;a number that was computed, and then dropped by one output path.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's the thing worth keeping. Not "measure your noise floor" — that's just good practice. It's that my instrument computed the answer to Vinh's question, printed it selectively in the direction that flattered me, and dropped it everywhere else. Nobody was going to find that from the outside. It took a stranger asking one question I couldn't answer.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;本文原載於我的部落格：&lt;a href="https://coffeeshooters.com/content/noise-floor-two-identical-runs-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-noise-floor-two-identical-runs-en" rel="noopener noreferrer"&gt;Two identical runs scored 89 and 89. Two cases had flipped.&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>learning</category>
    </item>
    <item>
      <title>Dead Code Is More Dangerous Than Broken Code — How a Non-Homologous AI Caught the Blind Spot the AI Itself Had Rubber-Stamped</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Wed, 09 Sep 2026 13:05:14 +0000</pubDate>
      <link>https://dev.to/dexterlung/dead-code-is-more-dangerous-than-broken-code-how-a-non-homologous-ai-caught-the-blind-spot-the-ai-j45</link>
      <guid>https://dev.to/dexterlung/dead-code-is-more-dangerous-than-broken-code-how-a-non-homologous-ai-caught-the-blind-spot-the-ai-j45</guid>
      <description>&lt;h2&gt;
  
  
  Dead Code Is More Dangerous Than Broken Code
&lt;/h2&gt;

&lt;h2&gt;
  
  
  The opening: an AI confidently told me to reuse a capability that didn't exist
&lt;/h2&gt;

&lt;p&gt;Here's how it started.&lt;/p&gt;

&lt;p&gt;I was adding a small feature to my own coffee e-commerce store: let a customer pick a fresh combination of beans over chat, price it on the spot, and order it. The chat-ordering bot was already built — I'd even tested balance deduction on a real device. All I needed was one more piece: something on the backend to compute the price, because the amount can't be whatever the phone says it is.&lt;/p&gt;

&lt;p&gt;I handed the repo to my main AI (Opus, my daily driver) and asked whether the backend already had pricing logic I could reuse. It scanned the code and came back fast:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The backend has &lt;code&gt;pricingLogic.ts&lt;/code&gt; — the pricing logic lives there. Just reuse it."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Sounds reasonable. The file really does exist. Open it and there really is a pile of price-computing code inside. Anyone reading the code — me, the AI, anyone — lands on the same conclusion at first glance: &lt;strong&gt;the backend has pricing capability, just wire it up.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;One problem. That code &lt;strong&gt;had not compiled a single time in six months.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;From the day it was born (&lt;code&gt;git blame&lt;/code&gt; says 2026-01-11) it was missing a closing &lt;code&gt;}&lt;/code&gt; — the &lt;code&gt;CartItem&lt;/code&gt; type literally couldn't close. The frontend test runner doesn't compile the backend's TypeScript, and the audit suite doesn't touch backend functions, so no gate had ever gone near it. It just sat there quietly, in the exact spot where "backend pricing logic" should live, for half a year. At one point another change even passed through it and added a line — the person making the change had no idea they were editing a file that doesn't compile.&lt;/p&gt;

&lt;p&gt;The main AI wasn't bluffing. The world it read was real: the file exists, the code inside is real pricing code. It just &lt;strong&gt;had no way to tell, from the source tree, that this thing was dead.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;My gut said stop: don't build on a foundation that drifts. Nail down the truth about pricing first — and, crucially, &lt;strong&gt;bring in a "non-homologous" model to diagnose it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;What that diagnosis dug up was far worse than I expected. It didn't just overturn the main AI's advice; it exposed a structural blind spot in my entire governance setup. That blind spot is what this post is about.&lt;/p&gt;




&lt;h2&gt;
  
  
  What "existential rot" is
&lt;/h2&gt;

&lt;p&gt;When we say "bad code," we usually picture &lt;strong&gt;broken code&lt;/strong&gt;: it computes wrong, throws errors, blows up mid-run. Broken code has one virtue — it makes noise. It crashes, someone reports it, it gets fixed.&lt;/p&gt;

&lt;p&gt;But the four things this diagnosis surfaced don't crash. Not one. They're a different kind of rot, and I've started calling it &lt;strong&gt;existential rot&lt;/strong&gt;: the problem isn't "it does the wrong thing," it's "it exists, does nothing, and lets everyone believe it's doing something."&lt;/p&gt;

&lt;p&gt;Four faces of it:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Dead code (born broken, never compiled).&lt;/strong&gt; That &lt;code&gt;pricingLogic.ts&lt;/code&gt;. Missing a bracket, six months without compiling, zero production impact. It isn't "broken" — it &lt;strong&gt;never lived&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. A deployed zombie (frozen on an old version).&lt;/strong&gt; The backend has &lt;em&gt;another&lt;/em&gt; pricing function that actually is deployed and live. The live version was v55, a snapshot from March. I sent it a real item with a price of 100 — it returned &lt;strong&gt;$0&lt;/strong&gt;. The source had moved on several versions, but the live one is frozen in March, and wrong. Better yet: it has zero callers, so nobody will ever feel pain from it returning $0.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. A phantom implementation (the Nth parallel copy).&lt;/strong&gt; Hidden in the repo was a &lt;em&gt;fourth&lt;/em&gt; pricing implementation — a fictional percentage-discount model. Exactly one place imports it: &lt;strong&gt;its own test file.&lt;/strong&gt; All tests green. The green proves "it agrees with itself," not "it agrees with reality."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. A defense with no reader (a warn-only log nobody reads).&lt;/strong&gt; I'd added a warn-only amount-recomputation check on the database side that logs suspicious orders to a table. Sounds responsible. But that table had exactly two rows — both from my own smoke test — and &lt;strong&gt;nothing in the entire repo, no sentinel, no script, no human, reads that table.&lt;/strong&gt; The defense is written, but there's no return path.&lt;/p&gt;

&lt;p&gt;Put these four side by side and they form the &lt;em&gt;entire truth&lt;/em&gt; of "backend pricing capability" in my system: &lt;strong&gt;one living body, three zombies.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The only thing actually computing money, start to finish, was one place — &lt;code&gt;cartPriceCalculator.js&lt;/code&gt;, running in the customer's browser. The three backend copies were all decorations occupying the spot.&lt;/p&gt;




&lt;h2&gt;
  
  
  The real root cause: pricing authority actually lives in the browser
&lt;/h2&gt;

&lt;p&gt;Why can "a pricing implementation with no consumer" survive this long?&lt;/p&gt;

&lt;p&gt;I read the body of the order-creation stored procedure by hand, and the answer is ugly: it &lt;strong&gt;takes the front end's already-computed&lt;/strong&gt; &lt;code&gt;subtotal&lt;/code&gt;, &lt;code&gt;total_amount&lt;/code&gt;, and &lt;code&gt;unit_price&lt;/code&gt; &lt;strong&gt;verbatim&lt;/strong&gt;. The payment side looks safe — it reads &lt;code&gt;orders.total_amount&lt;/code&gt; from the database — but that value is the one the front end stuffed in.&lt;/p&gt;

&lt;p&gt;In other words, in this architecture &lt;strong&gt;pricing authority de facto lives in the browser&lt;/strong&gt; (client-authoritative computation). The server just copies down the number the phone computed and files it.&lt;/p&gt;

&lt;p&gt;Once that's true, any backend pricing code becomes a &lt;strong&gt;structural decoration&lt;/strong&gt;: wiring it in costs extra effort, and &lt;em&gt;not&lt;/em&gt; wiring it in changes nothing. A decoration never gets traffic-tested — no real order passes through it, hits it, translates its error into "the customer paid too little," the kind of thing that hurts. So it can rot without detection and drift without limit.&lt;/p&gt;

&lt;p&gt;Which is exactly why my initial gut was right: &lt;strong&gt;the "align the two copies" prescription the main AI wrote was a prescription to repair a corpse.&lt;/strong&gt; Until you answer "who's going to call it, which version deploys, which copies get killed," aligning dead code to the living body just makes the corpse look more alive.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why my governance saw none of this
&lt;/h2&gt;

&lt;p&gt;Here's the part I actually wanted to write down.&lt;/p&gt;

&lt;p&gt;My governance isn't weak. There are locked cross-layer traces, there are commit-blocking hard rules, there's a whole audit suite. The top pricing dark-zone had &lt;em&gt;already&lt;/em&gt; been registered, trace-locked (I'll call it trace T-029), given 24 tests, all green, with the word &lt;code&gt;locked&lt;/code&gt; sitting right there in the registry — back on 2026-06-06.&lt;/p&gt;

&lt;p&gt;And then the darkest part of that dark zone &lt;strong&gt;kept rotting, untouched.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I eventually saw why. Almost all of my defenses are &lt;strong&gt;change-triggered&lt;/strong&gt; — they only fire when you &lt;strong&gt;touch a registered node.&lt;/strong&gt; A trace test runs when you edit a file on the trace; a blocker fires when you commit; an audit lights up red when you modify the thing it watches.&lt;/p&gt;

&lt;p&gt;But the whole class of "existential rot" &lt;strong&gt;doesn't require anyone to touch anything&lt;/strong&gt; to persist. The dead code lies there, the zombie stays frozen, the phantom chews its own tail, the reader-less log quietly accumulates — none of it trips a single "change detector."&lt;/p&gt;

&lt;p&gt;The lamp T-029 lit was the &lt;em&gt;front-end price-resolution convergence point&lt;/em&gt;. The "backend copy," the "deployed bundle," the "rule-loading layer" — three surfaces that were simply not in that lamp's beam. It got flagged, locked, 24 tests green — and the darkest part never got touched.&lt;/p&gt;

&lt;p&gt;What I was missing wasn't "one more change-triggered rule." It was an &lt;strong&gt;existence constraint&lt;/strong&gt;:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Any implementation that claims to compute money, to defend, or to be a single source of truth must have (a) a live consumer, and (b) a confrontation witness. Missing either one → retire it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Side note: this also fits a gap I'd already written down — my system only &lt;em&gt;adds&lt;/em&gt; features, it never &lt;em&gt;removes&lt;/em&gt; them (things only go in, never out). And "has no witness" turns out to be an objective test for "should have been retired but wasn't."&lt;/p&gt;




&lt;h2&gt;
  
  
  Why it had to be a &lt;em&gt;non-homologous&lt;/em&gt; model
&lt;/h2&gt;

&lt;p&gt;Back to the key decision: why did I insist on a different model, instead of asking the same AI to check again?&lt;/p&gt;

&lt;p&gt;Because the main AI wasn't short on ability. It was &lt;strong&gt;trapped by the same ready-made narrative.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;"This file is named &lt;code&gt;pricingLogic.ts&lt;/code&gt;, it has pricing code inside, therefore the backend has pricing capability" — that's a story too smooth, too reasonable; it fills your cognition. Whereas "does it even compile? does anyone call it? is that the version that's deployed?" is a &lt;strong&gt;backward-looking, declarative question&lt;/strong&gt; that will &lt;em&gt;not&lt;/em&gt; jump up and knock when you're about to reuse a file.&lt;/p&gt;

&lt;p&gt;A homologous model colludes. Ask it to look again and it'll most likely re-read "the repo has this file" as "the capability exists" a second time, because it's running the same inference reflex as the first time. That's not stupidity — it's an instrument problem. &lt;strong&gt;The source tree is the instrument's testimony, not the fact of the world&lt;/strong&gt; (the actual runtime behavior).&lt;/p&gt;

&lt;p&gt;Switch to a non-homologous model (I used one from a different family) and the effect is like bringing in an outsider who never bought your narrative to review the code. The first thing it does is not &lt;em&gt;trust&lt;/em&gt; the repo — it goes and &lt;strong&gt;confronts&lt;/strong&gt; it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It ran a real compiler over the dead code — instant red. Hard proof: "never compiled."&lt;/li&gt;
&lt;li&gt;It sent real input to the live function — got back $0. Hard proof: "deployed zombie computes wrong."&lt;/li&gt;
&lt;li&gt;It grepped for callers of those "canonical" functions on the production path — zero. Hard proof: "no consumer."&lt;/li&gt;
&lt;li&gt;It read the order-creation procedure body by hand — saw it take the front-end price verbatim. Hard proof: "authority lives in the browser."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It even corrected the main AI's &lt;em&gt;diagnosis name&lt;/em&gt;: this isn't two implementations "drifting" (drift = two living things gradually diverging). This is &lt;strong&gt;necrosis plus a split truth&lt;/strong&gt; — a piece of dead code that never lived, a zombie frozen in March, zero callers. Different name → completely different prescription.&lt;/p&gt;

&lt;p&gt;The value of non-homologous diagnosis, in one line: &lt;strong&gt;it's the vaccine against self-rubber-stamped green lights.&lt;/strong&gt; When a system (including an AI) blesses itself, the most effective mirror isn't making it look at itself harder — it's introducing a confronter that doesn't collude.&lt;/p&gt;




&lt;h2&gt;
  
  
  Six green lights, not one of them watching the load-bearing thing
&lt;/h2&gt;

&lt;p&gt;During the diagnosis I did something I think everyone with a test suite should do occasionally: pull out every "green light" and ask each one — &lt;strong&gt;what load-bearing thing did you actually, positively observe?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The result (all of this is re-checkable):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Green light&lt;/th&gt;
&lt;th&gt;What it claims&lt;/th&gt;
&lt;th&gt;What it actually saw&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Pricing tests all green&lt;/td&gt;
&lt;td&gt;"Business rules are test-protected"&lt;/td&gt;
&lt;td&gt;A fictional model with zero production callers chewing its own tail&lt;/td&gt;
&lt;td&gt;False green&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;T-029 trace 24 green + &lt;code&gt;locked&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;"Pricing SSOT is locked"&lt;/td&gt;
&lt;td&gt;Only the front-end resolvers' fallback — not the compute core, not the backend, not rule-loading&lt;/td&gt;
&lt;td&gt;Half true&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Edge-function smoke passes&lt;/td&gt;
&lt;td&gt;"The pricing function works"&lt;/td&gt;
&lt;td&gt;HTTP 200 + &lt;code&gt;success:true&lt;/code&gt; — returning $0 is also a 200&lt;/td&gt;
&lt;td&gt;False green (checks the status code, not the amount)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Warn-only defense shipped&lt;/td&gt;
&lt;td&gt;"The DB side has a recompute defense"&lt;/td&gt;
&lt;td&gt;Two log rows, both self-inflicted, nobody reads it&lt;/td&gt;
&lt;td&gt;Defense exists, no return path&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Migration header "not applied this batch"&lt;/td&gt;
&lt;td&gt;"Prod wasn't touched by this"&lt;/td&gt;
&lt;td&gt;Prod &lt;em&gt;was&lt;/em&gt; applied&lt;/td&gt;
&lt;td&gt;Doc narrative decoupled from prod&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Front end priced successfully&lt;/td&gt;
&lt;td&gt;"The user sees the correct price"&lt;/td&gt;
&lt;td&gt;On RPC failure it silently falls back and prices anyway&lt;/td&gt;
&lt;td&gt;Silent degradation, no alert&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Six lights, each making you feel "someone's minding this surface." But &lt;strong&gt;not one of them ever positively observed the one thing that's actually load-bearing&lt;/strong&gt;: two runtimes computing the same number for the same cart.&lt;/p&gt;

&lt;p&gt;Here's the line I took away and want to nail to the wall: &lt;strong&gt;a green light with no witness proves "nobody's looking," not "it's correct."&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The fix: let the zombies die all the way first, then give the living body witnesses
&lt;/h2&gt;

&lt;p&gt;Once the framing was overturned, the fix got clear. The principle was "move logic only, don't change a single price." Three steps:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 1: let the zombies die all the way.&lt;/strong&gt; Both dead copies got deleted and retired. No aligning, no repairing — a zero-consumer corpse doesn't deserve a minute of your time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 2: extract the one living body into a single core.&lt;/strong&gt; I pulled the one algorithm that actually computes money out of the front end into a pure-function core (zero imports, computes and returns, touches nothing outside). The front end became "load rules → hand to the core." The backend function was rewritten to "resolve rules → hand to the &lt;em&gt;same&lt;/em&gt; core." One logic, both runtimes delegating to it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 3 — the real point this time — give it three witnesses.&lt;/strong&gt; Sharing one import isn't enough; it only kills the "copy-paste" axis of drift. There's a nastier axis: &lt;strong&gt;deploy lag.&lt;/strong&gt; The same source, front end shipping the latest, backend frozen in March — that's two truths. A shared import can't catch that. Only one thing can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A parity witness:&lt;/strong&gt; feed the &lt;em&gt;same&lt;/em&gt; batch of inputs to the front-end path and the backend path, compare outputs row by row. I prepared six cart "battery" cases (cross-tier VIP, gift-box volume, tiered discount, buy-X-get-Y, etc.), froze the &lt;em&gt;pre&lt;/em&gt;-refactor live output as a golden baseline, and any change that moves a single dollar goes red.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A live deploy-lag witness:&lt;/strong&gt; hit the &lt;em&gt;actually deployed&lt;/em&gt; version over HTTP and compare it against the local core, row by row. This is the only edge that can see "frozen on an old version." First run, it immediately spat out 3-of-3 mismatches against live v55 — proof the detector is alive and working.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A compile gate covering all runtimes:&lt;/strong&gt; a blocker requiring that the pricing-chain backend function actually compiles. This is the direct counterfactual to "dead code nobody knew about for six months" — had this gate existed, that missing-bracket corpse would've gone red on day one.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One thing I deliberately did &lt;strong&gt;not&lt;/strong&gt; do: I didn't heavily rewrite the order-creation procedure body. Because in this repo, "rewrite a gatekeeper function wholesale" has a rap sheet — it wipes out the gatekeeping logic buried inside. (Two prior migrations did exactly that, once wiping a gift-box capacity guard and once a stored-value balance guard, both dead for weeks.) Server-side authority gets promoted later, gradually, through a hook I'd already pre-embedded — without reopening that knife.&lt;/p&gt;




&lt;h2&gt;
  
  
  Turning one expensive insight into a free mechanism forever
&lt;/h2&gt;

&lt;p&gt;This kind of diagnosis is expensive. A non-homologous model, copy-by-copy confrontation, climbing counterfactuals — that burns real tokens. So I have a habit: &lt;strong&gt;every piece of reasoning I've paid for once, I try to make cheap or free the next time.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Two reusable things crystallized out this time, sitting at two different moments:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. A lens for the "about to build" moment — the consumer-witness lens.&lt;/strong&gt; A check I built and use daily. Its core is a gate: before you trust, cite, or build on anything that "claims to compute money / to defend / to be the single truth," it &lt;strong&gt;forces you to write two lines&lt;/strong&gt; — (a) which specific caller in production depends on it? (b) which specific test feeds it the same inputs and confronts it against the authority? If you can't write either line, treat it as nonexistent. This lens is the vaccine against that exact "fooled into the wrong framing by a corpse" moment, grown into a gate that knocks before every future reuse.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. A checklist item for the "designing" moment — the server-authoritative rule.&lt;/strong&gt; I maintain a cross-project front-end/back-end completeness checklist (past incidents distilled into check-sentences). This time I added one, with a machine-readable tag so tools can auto-apply it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Load-bearing &lt;em&gt;computed&lt;/em&gt; values (price, total, balance, stock — the "computed" kind) must be computed by the server, with the client only displaying; the server must &lt;strong&gt;never trust a client-computed load-bearing value verbatim.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is a generalization of an old rule I already had ("no front-end multiplication for the unit-price field") — the old rule covers the pricing special case; this one covers "for every computed load-bearing value, who is allowed to compute it, and whom does the downstream trust." The machine-readable tag means: next time I decompose a new project through my tool-chain, this rule gets &lt;strong&gt;spat out into the spec automatically&lt;/strong&gt;, without me remembering it. Pay once, carry it into every new project free.&lt;/p&gt;

&lt;p&gt;But "spat out automatically" hides a trap I only fixed this time — and it's a vivid one, worth its own beat.&lt;/p&gt;

&lt;p&gt;I'd crystallized that check-sentence library into a machine-readable config file — plainly, an index of "when I hit this sharp edge, here's how I handle it." It hangs on the tool-chain I use to decompose new projects: start a new project and the tool reads it and auto-applies my past scars into the spec. Problem: that config file is &lt;strong&gt;private&lt;/strong&gt; (it lives in a toolkit repo of mine that isn't public), while my project bootstrapper pulls &lt;strong&gt;almost everything else down from public GitHub.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So you get a very classic tension: &lt;strong&gt;a bootstrapper that fetches everything from a public channel, meeting one file that can't travel a public channel.&lt;/strong&gt; Worse — historically that config file &lt;strong&gt;had never been supplied automatically.&lt;/strong&gt; It sat in a "manual step nobody actually ran." The result: every decomposition &lt;strong&gt;silently ate the neutral defaults&lt;/strong&gt;, and this config I'd painstakingly crystallized contributed exactly zero.&lt;/p&gt;

&lt;p&gt;That is literally the disease this whole post is about, growing back in a new spot: a thing that "should be supplied but isn't" quietly trips no alarm, letting every decomposition &lt;em&gt;believe&lt;/em&gt; it applied the completeness rules when it didn't.&lt;/p&gt;

&lt;p&gt;So I wired the supply in, carefully, every line guarding against "private content leaking into a public repo":&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Never fetch it.&lt;/strong&gt; Unlike the other tools, this config isn't downloaded from GitHub — it's &lt;strong&gt;copied at runtime, on my own machine, from a source already present locally&lt;/strong&gt; (scan a few common folders for that private toolkit, or read a path from an env var). Zero hardcoded private paths in the code. A regular user has none of that private source, so it resolves to "not found," falls back to neutral defaults — &lt;strong&gt;nothing leaks.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fence it with &lt;code&gt;.gitignore&lt;/code&gt;.&lt;/strong&gt; On supply it auto-adds that private config's folder to the project's &lt;code&gt;.gitignore&lt;/code&gt;, with an iron-rule comment baked in — "never reaches a public repo." So even if you later fat-finger &lt;code&gt;git add .&lt;/code&gt;, it can't reach any outward remote.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fail loud on supply, never silent.&lt;/strong&gt; If no source is found it &lt;strong&gt;says so out loud&lt;/strong&gt; — "config not supplied, this run uses neutral defaults" — because &lt;strong&gt;a silent skip is exactly the bug this fixes.&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I love this bit because it compresses the whole lesson into one concrete action: &lt;strong&gt;a bootstrapper that fetches everything from the public internet meets a private file that can't travel that public channel — so it switches to "smuggle it in locally + fence it with &lt;code&gt;.gitignore&lt;/code&gt; so it can never leak back to a public repo."&lt;/strong&gt; That "never reaches a public repo" iron-rule comment is what an &lt;em&gt;existence constraint&lt;/em&gt; looks like once it's grown into code: "this thing should exist, and only here" written as a rule the program executes, not a line of doc nobody runs.&lt;/p&gt;

&lt;p&gt;For the record, the referenceable deliverables from this round (things I use daily, fed by real incidents — not slideware): a check I call the &lt;strong&gt;consumer-witness lens&lt;/strong&gt; (before reusing anything, force yourself to write down its two witnesses), a cross-project &lt;strong&gt;front-end/back-end completeness checklist&lt;/strong&gt; (every incident distilled into a check-sentence), and an &lt;strong&gt;orphan audit&lt;/strong&gt; script (hunts "zero-caller decorative anchors" — it found thirty-nine in one project alone).&lt;/p&gt;




&lt;h2&gt;
  
  
  Closing: this was never really about pricing
&lt;/h2&gt;

&lt;p&gt;If you filed this under "a story about one pricing bug," you missed the point.&lt;/p&gt;

&lt;p&gt;The two drifting pricing copies were just the most visible projection of the hole. The real hole is that &lt;strong&gt;the copy axis has no forcing function&lt;/strong&gt;: the same "how money is computed / how state is judged" logic actually lives in five physical spots — front-end bundle, back-end bundle, the live DB function body, migration files, repo source — and &lt;strong&gt;nothing guarantees they equal each other.&lt;/strong&gt; They will, and already do, tell different stories.&lt;/p&gt;

&lt;p&gt;So what's actually portable is this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Dead code is more dangerous than broken code.&lt;/strong&gt; Broken code blows up and gets fixed; dead code lies quietly in the spot where "some capability should live," fooling every reader of the code — including the AI itself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A defense with no reader is not a defense.&lt;/strong&gt; A warn-only log, an alert nobody subscribes to, a dashboard nobody looks at — all worth zero.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A single SSOT module is not a single truth.&lt;/strong&gt; A shared import only kills "copy-paste" drift; "deploy-lag" drift is visible only to a confrontation witness.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"The repo has two copies of this code" should not first read as "two implementations."&lt;/strong&gt; First ask: does it compile? does anyone call it? which version is deployed? — very likely it's "one living body plus N zombies," not drift at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If your governance is all change-triggered, it's blind to existential rot.&lt;/strong&gt; Add one existence constraint: every money-computing / defending implementation needs a live consumer plus a confrontation witness, or it gets retired.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last decision — "don't build on a foundation that drifts; fix the truth first, and bring in a different model to look" — took under a minute. It was the highest-return step in the whole thing. Because an AI reading your code, however confident, is only giving you &lt;strong&gt;instrument testimony.&lt;/strong&gt; Whether to build on it, you have to ask first: this capability — is it actually alive?&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This is one entry in my one-person studio's dev-notes series. The consumer-witness lens, the completeness checklist, and the orphan audit are tools I built and use daily; some are already open-sourced. They aren't theory — every one has a pit like this one behind it, a pit I actually fell into.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;本文原載於我的部落格：&lt;a href="https://coffeeshooters.com/content/dead-code-more-dangerous-non-homologous-diagnosis-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-dead-code-more-dangerous-non-homologous-diagnosis-en" rel="noopener noreferrer"&gt;Dead Code Is More Dangerous Than Broken Code — How a Non-Homologous AI Caught the Blind Spot the AI Itself Had Rubber-Stamped&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>architecture</category>
      <category>testing</category>
    </item>
    <item>
      <title>AI Auditing an AI's Own Toolchain: The Repair Tool That Overwrote Its Own Fix</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Tue, 08 Sep 2026 14:05:34 +0000</pubDate>
      <link>https://dev.to/dexterlung/ai-auditing-an-ais-own-toolchain-the-repair-tool-that-overwrote-its-own-fix-3ce8</link>
      <guid>https://dev.to/dexterlung/ai-auditing-an-ais-own-toolchain-the-repair-tool-that-overwrote-its-own-fix-3ce8</guid>
      <description>&lt;p&gt;I have seven interdependent tool repos — some public, some private — each with its own audit checks, but nobody had ever checked &lt;strong&gt;what happens when they're wired together.&lt;/strong&gt; This post is about handing all seven to Fable5 for a cross-repo audit, and the fact that not one of the real problems it found could have been caught by auditing any single repo alone.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Read on: &lt;a href="https://dev.to/content/fable5-confrontation-point-law-en"&gt;Previous: the Confrontation-Point Law&lt;/a&gt; · &lt;a href="https://dev.to/content/fable5-thinking-extraction-en"&gt;Don't ask a model to narrate itself&lt;/a&gt; · &lt;a href="https://dev.to/content/fable5-cross-repo-audit"&gt;繁體中文版&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The audit prompt itself went through three drafts
&lt;/h2&gt;

&lt;p&gt;I didn't just paste in "please check these seven repos." The prompt started as a skeleton, then got execution steps, then got filled in with the real seven repos — 5 public tool-chain repos plus 2 private ones (an internal connective framework and a cross-project reuse library). The goal was to write context and boundaries, not a checklist — tell Fable what the terrain is and where the hard lines are, and make verification about mechanism, not about handing over an answer.&lt;/p&gt;

&lt;p&gt;Each repo got its own scanning subagent, plus one independent verifier subagent that never saw the scanners' conclusions. That independent verifier is the part that mattered — it overturned four of the scanning agents' conclusions. The topology itself turned out fine; the two public/private seams both worked correctly. The real problems weren't in any single repo. They were three composition-level failures that &lt;strong&gt;only become visible once you look at all seven together.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The tool that fixes drift can silently overwrite your work
&lt;/h2&gt;

&lt;p&gt;The most interesting one: one tool's job is to detect and repair drift between a local skill copy and its canonical source — if your local copy diverges from the original, it syncs you back up. The problem: in one environment, that "local copy" is actually a Windows junction (a filesystem shortcut) pointing straight at the canonical repo — and the tool's scan logic only checked "is this a directory," never "is this actually a shortcut."&lt;/p&gt;

&lt;p&gt;The result: &lt;strong&gt;a tool built to repair drift will write straight through the shortcut into the canonical repo, potentially overwriting uncommitted work sitting on someone's feature branch.&lt;/strong&gt; The repair tool itself becomes the risk. No single-repo audit could have caught this — the bug doesn't live in any code, it lives in how directories are wired together on a specific machine. That shape is only visible across repos.&lt;/p&gt;

&lt;p&gt;The other two composite failures: an orphaned commit left a cross-repo contract stranded with nothing consuming it; and three places documented something as "prompt-level, optional" that was actually being treated as "code-level, mandatory for everyone" — a documentation/enforcement mismatch that leaves different maintainers with completely different mental models of the same rule.&lt;/p&gt;

&lt;h2&gt;
  
  
  One rule worth stealing: content-level probes, not existence checks
&lt;/h2&gt;

&lt;p&gt;The audit converged on a principle I think is more valuable than any single finding: &lt;strong&gt;"the commit exists" or "the file exists" doesn't count as verified.&lt;/strong&gt; You have to walk the entire propagation chain — local branch, remote default branch, the published artifact or install source, the installed copy — and grep an actual signature string at every single stage, and it has to actually hit, every time.&lt;/p&gt;

&lt;p&gt;That principle is what caught the most important issues in this audit. Downgrade it to an existence check, and the audit still runs, still produces a report, &lt;strong&gt;and catches nothing — which is worse than not running it at all, because it hands you a false green light.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A real case: how a spec makes frontend and backend asymmetric from birth
&lt;/h2&gt;

&lt;p&gt;The day after the cross-repo checklist converged, a real client project handed us a case that got folded directly into that checklist as a new criterion.&lt;/p&gt;

&lt;p&gt;The setup: a real order-management admin feature with a visible frontend/backend asymmetry. The backend RPC defined an error code called &lt;code&gt;BOOKING_NOT_FOUND&lt;/code&gt; that the frontend's error-message map never learned about. A refund action returned HTTP 501 in production with zero UI indication. A &lt;code&gt;loading&lt;/code&gt; ref was declared and returned but never consumed by any of six admin tabs. An admin-readable transaction table existed and was never queried by the frontend at all.&lt;/p&gt;

&lt;p&gt;We gathered the hard evidence first, then asked for the root cause, not just a symptom list. The diagnosis: the asymmetry wasn't born in any one layer — it was born upstream, in the spec's own output grammar. The original spec gave the backend a &lt;em&gt;model&lt;/em&gt; (field-level data structures) and gave the frontend only a &lt;em&gt;verb&lt;/em&gt; ("view orders and bookings"). Every downstream layer — decomposition, execution, scoring — is deliberately faithful by design, each explicitly instructed not to invent or over-check. So the upstream asymmetry got preserved intact, all the way to the final output. &lt;strong&gt;No layer dropped the ball. The pipeline worked exactly as designed — the design itself just had a blind seam nobody had seen.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Even the fix this diagnosis proposed doubted itself out loud: the suggestion was to add a mandatory "Surface Contract" section (a state row plus a seam row) to the five-element goal format. But the same document flagged its own biggest risk: "once a template becomes mandatory, the most likely failure is that people fill it with boilerplate that formally passes the check while saying nothing real — that's the step I'm most likely to have gotten wrong." A model willing to say, in writing, "my own suggestion might just get rubber-stamped" is more useful to me than the suggestion itself.&lt;/p&gt;

&lt;p&gt;That proposal is, as of this writing, only designed — not yet wired into the actual goal template. Worth saying plainly: &lt;strong&gt;designed is not the same as done.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaway
&lt;/h2&gt;

&lt;p&gt;A single-repo audit can't see the shape that only exists once tools are wired together — and the most expensive lesson here is that the tool meant to repair drift can be the single riskiest thing in the whole system. "Content-level probes, not existence checks" is a rule worth copying into any multi-repo system you maintain, whether or not you have a Fable5 to run it for you.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;本文原載於我的部落格：&lt;a href="https://coffeeshooters.com/content/fable5-cross-repo-audit-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-fable5-cross-repo-audit-en" rel="noopener noreferrer"&gt;AI Auditing an AI's Own Toolchain: The Repair Tool That Overwrote Its Own Fix&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>webdev</category>
      <category>devops</category>
    </item>
    <item>
      <title>How a Security Bug Gave Birth to a Method, Then the Method Caught What the Audit Missed</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Tue, 08 Sep 2026 14:05:22 +0000</pubDate>
      <link>https://dev.to/dexterlung/how-a-security-bug-gave-birth-to-a-method-then-the-method-caught-what-the-audit-missed-1c09</link>
      <guid>https://dev.to/dexterlung/how-a-security-bug-gave-birth-to-a-method-then-the-method-caught-what-the-audit-missed-1c09</guid>
      <description>&lt;p&gt;At 5:27pm on 2026-07-03, I revoked self-write access on 8 sensitive columns of our members table. At 9:42pm the same day, I reverted it — an admin save feature started throwing 403 errors. Locked, then unlocked, same day.&lt;/p&gt;

&lt;p&gt;That incident became the seed for a method we call the Confrontation-Point Law. I won't claim it's proven — there's exactly one real case so far, and we've deliberately labeled it "armed, not battle-tested." But that one case is one of my favorite technical stories to date, because what it caught was something a careful human audit had already missed.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Read on: &lt;a href="https://dev.to/content/fable5-thinking-extraction-en"&gt;Previous: don't ask a model to narrate itself&lt;/a&gt; · &lt;a href="https://dev.to/content/fable5-cross-repo-audit-en"&gt;Next: AI auditing an AI's own toolchain&lt;/a&gt; · &lt;a href="https://dev.to/content/fable5-confrontation-point-law"&gt;繁體中文版&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  That lockdown commit looked careful
&lt;/h2&gt;

&lt;p&gt;The 07-03 lockdown commit's header read: "verified all 5 write paths to the users table site-wide, &lt;code&gt;MembersCenter.vue&lt;/code&gt; goes through &lt;code&gt;fn_admin_*&lt;/code&gt; RPCs." That's not a throwaway line — someone actually enumerated the paths and labeled each as a legitimate channel. On paper, this looked like a solid audit.&lt;/p&gt;

&lt;p&gt;But inside that same file, the function &lt;code&gt;saveEditMember&lt;/code&gt; calls an RPC to write &lt;code&gt;member_level&lt;/code&gt; at line 2010 — &lt;strong&gt;and ten lines later, at line 2020&lt;/strong&gt;, the same function does a raw &lt;code&gt;.from('users').update({ tier_updated_at, tier_expires_at, tier_auto_downgrade })&lt;/code&gt;. Direct write, no gate.&lt;/p&gt;

&lt;p&gt;The same function was, at the file level, a proper RPC channel — and at the field level, simultaneously an unguarded direct-write channel. That gap is exactly what triggered the 403s that night, and exactly why we reverted four hours later.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a careful audit still missed it
&lt;/h2&gt;

&lt;p&gt;Looking back, what bothers me isn't that there was a bug. It's &lt;strong&gt;why a careful audit still missed it.&lt;/strong&gt; The answer is blunt: the auditor read the code, saw the RPC call, saw a comment that said "goes through RPC," and issued a file-level verdict. &lt;strong&gt;The census got outsourced to narrative.&lt;/strong&gt; The fire wasn't because the set of write points was too large to enumerate — we later verified the whole frontend only touches that field in 3 files, something a dumb grep finds in seconds. The fire was because the verdict substituted "I read it, looks fine" for "I mechanically checked every single line."&lt;/p&gt;

&lt;p&gt;That finding is what produced the first rule of the Confrontation-Point Law: &lt;strong&gt;every frozen judgment needs to name its confrontation point&lt;/strong&gt; — a place where intent and actual state are forced to meet, where any disagreement blocks progress, sitting right on the boundary where real harm happens. Who's allowed to sit in the "this is what's actually true" seat? Only a machine — a real query like &lt;code&gt;has_column_privilege&lt;/code&gt;, never an agent reading code and saying "I checked, looks fine." Narrative sitting in that seat is a fake census. That "goes through RPC" comment is a live example of narrative sitting in the wrong seat.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same week, a mechanical re-audit caught 3 leaks the human review missed
&lt;/h2&gt;

&lt;p&gt;On 07-04, instead of trusting "I read it, looks fine" again, we ran a fully mechanical re-census: every claimed-legitimate write point had to be re-derived through grep/AST plus an actual privilege query. Anything that couldn't be re-derived got flagged red on the spot — no exceptions.&lt;/p&gt;

&lt;p&gt;This caught 3 admin direct-write points the earlier review had completely missed: &lt;code&gt;renewTier&lt;/code&gt;, &lt;code&gt;dismissBinding&lt;/code&gt;, &lt;code&gt;importMembers&lt;/code&gt;. All 8 sensitive columns were re-verdicted from scratch — all 8 came back safe to revoke — and the three admin write paths got migrated into three new SECURITY DEFINER RPCs. This time the rollout order was deliberately sequenced: ship the RPCs first, wait for the frontend deploy to land, &lt;em&gt;then&lt;/em&gt; run the revoke — specifically to avoid repeating the 07-03 disaster of locking permissions before the frontend had caught up.&lt;/p&gt;

&lt;p&gt;After going live, we actually verified it in production, not just in code review: a regular member trying to self-write any of the 8 columns gets a 403; all 5 admin paths still work.&lt;/p&gt;

&lt;h2&gt;
  
  
  I'm not going to oversell this
&lt;/h2&gt;

&lt;p&gt;Honestly: this whole method has &lt;strong&gt;exactly one real case&lt;/strong&gt; behind it so far. We deliberately don't call it "battle-tested" — only "armed": self-tests pass, an injected regression test passes, but it hasn't yet stopped a real attack we didn't already know about. That line is written into the method itself: armed is not the same evidence tier as battle-tested.&lt;/p&gt;

&lt;p&gt;There's a funny footnote to how this method got written, too: one of the prompts we fed to Fable had a placeholder bug that silently ate the draft content. Fable didn't hallucinate a plausible-sounding answer to cover for it — it just refused to answer. We retroactively counted that refusal as a live demonstration of the exact principle we were designing: refusing to answer is itself a legitimate output.&lt;/p&gt;

&lt;p&gt;One more loose thread, left here honestly rather than swept under the rug: the same review found &lt;code&gt;Addresses.vue&lt;/code&gt; writing to a column called &lt;code&gt;addresses&lt;/code&gt; that doesn't actually exist in the database. That's deliberately out of scope for this fix and still unresolved.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaway
&lt;/h2&gt;

&lt;p&gt;A careful-looking audit can still lie at the field level, if its "verified" claim comes from a human reading code and forming a narrative judgment. A census you can actually trust is one where every single row can be mechanically re-derived — anything that can't gets flagged, no matter how confident the write-up sounds. And keep "armed" and "battle-tested" as two different evidence tiers — blurring them is just lying to yourself with extra steps.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;本文原載於我的部落格：&lt;a href="https://coffeeshooters.com/content/fable5-confrontation-point-law-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-fable5-confrontation-point-law-en" rel="noopener noreferrer"&gt;How a Security Bug Gave Birth to a Method, Then the Method Caught What the Audit Missed&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>webdev</category>
      <category>devops</category>
    </item>
    <item>
      <title>Don't Ask a Model to Narrate Itself: How I Actually Extracted a Method From Fable5</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Thu, 03 Sep 2026 13:05:17 +0000</pubDate>
      <link>https://dev.to/dexterlung/dont-ask-a-model-to-narrate-itself-how-i-actually-extracted-a-method-from-fable5-5071</link>
      <guid>https://dev.to/dexterlung/dont-ask-a-model-to-narrate-itself-how-i-actually-extracted-a-method-from-fable5-5071</guid>
      <description>&lt;p&gt;The first time I wanted to "learn how Fable5 thinks," I did the obvious thing: I asked it directly. "How do you approach a problem?" I copied down the answer, then kept asking, until it grew into an eleven-book set — a how-to guide, a book on perspective training, a book on code-review instincts, all the way up to a "main volume."  It looked thorough. It took real time to assemble.&lt;/p&gt;

&lt;p&gt;Then the whole set got overturned by its own method. Not because the content was wrong — because &lt;strong&gt;the approach itself was self-contradictory.&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Read on: &lt;a href="https://dev.to/content/spec-sonar-design-journal-part-1-en"&gt;spec-sonar design journal, Part 1&lt;/a&gt; · &lt;a href="https://dev.to/content/fable5-confrontation-point-law-en"&gt;Next: the Confrontation-Point Law&lt;/a&gt; · &lt;a href="https://dev.to/content/fable5-thinking-extraction"&gt;繁體中文版&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The problem: I used an untrustworthy method to study a doctrine about not trusting self-reports
&lt;/h2&gt;

&lt;p&gt;It was Opus that pointed out the embarrassing part: Fable's own core doctrine says don't trust self-reports — look at what actually got done, not what got said about it. And the method I'd used to write those eleven books was exactly that: ask "how do you think," then treat the answer as ground truth.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I used the discredited method to document the very doctrine that discredits it.&lt;/strong&gt; Once that was named, only four principles out of the whole set survived. Everything else needed a rebuild. I wasn't angry so much as deflated — that wasn't ten minutes of work, it was a full conversation's output, now judged to have a broken foundation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: stop asking what it thinks, start reverse-engineering what it actually did
&lt;/h2&gt;

&lt;p&gt;The new method was simple: instead of asking "how do you think," we took two things Fable5 had actually produced — the &lt;code&gt;router-reasoning.md&lt;/code&gt; and &lt;code&gt;decision-table.md&lt;/code&gt; from an n8n cold-start generator — and reverse-engineered the decision rules baked into them. Trust the scar, not the self-report — applied to the method itself, this time.&lt;/p&gt;

&lt;p&gt;Then we ran a two-pass adversarial interview, four rounds. Pass one: no draft shown, free-form answers. Pass two: show it the draft, let it attack. The point of this structure is to surface two different kinds of error — pass one catches what we didn't think of; pass two catches what we got wrong.&lt;/p&gt;

&lt;p&gt;Round one already surprised us. We'd planned to frame the method as something like "five joints." Fable rewrote the whole naming itself: forge, runway, customs — forge is where you're still building, runway is where you run an already-proven pattern, customs is the border crossing where the two are forced to meet and get checked. That wasn't a wording polish. It was a correction to how we'd classified our own method.&lt;/p&gt;

&lt;h2&gt;
  
  
  Round two: it caught an error happening live, inside our own document
&lt;/h2&gt;

&lt;p&gt;Round two got more interesting. Opus's compressed one-line summary read: "reduce cold-start to lookup then emit a goal graph." Sounds fine. Fable pointed out that compression had quietly deleted two non-happy-path exits — "ask" and "stop" — which happen to be where every safety mechanism in the method actually lives. We checked the real document. The omission was real.&lt;/p&gt;

&lt;p&gt;Something more embarrassing happened in the same round: Opus had written, in its own appendix, "we need to update this playbook" — and never did. Fable caught that live drift and used it as proof of its own thesis on the spot: canonical and working copies drift, not as theory, but as something happening in front of us.&lt;/p&gt;

&lt;h2&gt;
  
  
  Round four: an exam that could actually fail
&lt;/h2&gt;

&lt;p&gt;Round four was a holdout test — a problem it had never seen (e-invoice integration), to see whether it would apply the same method cold. It didn't improvise an answer. It first grepped the real repo, found an old "manual invoice era" precedent, and answered from that. We checked afterward, line by line: the file and line number it cited matched the real file, verbatim.&lt;/p&gt;

&lt;p&gt;That was the first time I'd actually &lt;em&gt;verified&lt;/em&gt; a model's reasoning instead of just feeling like it guessed right — &lt;strong&gt;and a verified-correct answer is a categorically different thing from a lucky-and-unchecked one.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The last move: it turned its own rule against itself
&lt;/h2&gt;

&lt;p&gt;The whole extraction closed on one line from Fable:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"You keep saying evidence. Do you have any evidence for yourself?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The method's own second core tenet is the evidence rule — every frozen claim needs an anchor you can check. Fable pointed out that the entire playbook had, up to that point, zero evidence for its own validity. We thought the interview was over. Instead we got judged by the very method we'd built.&lt;/p&gt;

&lt;p&gt;So we ran a "firing-record audit" — checking every claim in the method against the real git history of the CS project. The result wasn't a flat "no evidence." It split into three tiers: &lt;strong&gt;battle-tested&lt;/strong&gt; (actually stopped a real problem), &lt;strong&gt;armed but not fired&lt;/strong&gt; (the mechanism exists but hasn't faced a real test yet), and &lt;strong&gt;structurally unfireable&lt;/strong&gt; (there's no way this claim could ever be checked). That three-tier grading is now how we evaluate any governance mechanism we build — not a vague "does it work," but an honest label of what evidence tier it's actually at.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it cost me
&lt;/h2&gt;

&lt;p&gt;Honestly: I never deleted those eleven overturned books, but I stopped using them. That was a full conversation's worth of time, spent to arrive at "the approach itself was wrong." If there's one real scar in this piece, it's that — &lt;strong&gt;I picked the wrong extraction method first, and it took a whole round to find out.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What crystallized
&lt;/h2&gt;

&lt;p&gt;The four rounds converged into two reusable skills: &lt;code&gt;connective-system-design&lt;/code&gt; (forge/runway/customs, for designing cross-system connections) and &lt;code&gt;root-cause-diagnosis&lt;/code&gt; (counterfactual climbing, for finding a systemic root cause). The seal Fable added itself: &lt;strong&gt;"Trust the scar, not me."&lt;/strong&gt; — including don't fully trust this article either. Go check the anchors.&lt;/p&gt;

&lt;p&gt;The method was immediately put to two real tests: a cross-repo ecosystem audit across seven repos, and a real client project's frontend/backend asymmetry diagnosis. Both are the next two posts in this series.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway:&lt;/strong&gt; if you want to actually learn a method from a model, don't ask "how do you think" — look at what it actually produced, reverse-engineer the rule, test it on something new, and verify before you believe it. A self-report isn't free truth. It's a claim you still have to go check.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;本文原載於我的部落格：&lt;a href="https://coffeeshooters.com/content/fable5-thinking-extraction-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-fable5-thinking-extraction-en" rel="noopener noreferrer"&gt;Don't Ask a Model to Narrate Itself: How I Actually Extracted a Method From Fable5&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>webdev</category>
      <category>devops</category>
    </item>
    <item>
      <title>I turned both knobs on my on-prem Chinese RAG all the way up. It didn't get a single extra question right</title>
      <dc:creator>Dexterlung</dc:creator>
      <pubDate>Tue, 01 Sep 2026 13:05:15 +0000</pubDate>
      <link>https://dev.to/dexterlung/i-turned-both-knobs-on-my-on-prem-chinese-rag-all-the-way-up-it-didnt-get-a-single-extra-question-18im</link>
      <guid>https://dev.to/dexterlung/i-turned-both-knobs-on-my-on-prem-chinese-rag-all-the-way-up-it-didnt-get-a-single-extra-question-18im</guid>
      <description>&lt;h2&gt;
  
  
  Something you can do right now
&lt;/h2&gt;

&lt;p&gt;If you're about to "increase the retrieval depth" or "swap in a newer reranker" to improve your Chinese RAG accuracy, &lt;strong&gt;spend ten seconds on this first&lt;/strong&gt;:&lt;/p&gt;

&lt;p&gt;Take one question you got wrong, and compute &lt;strong&gt;where the passage containing the correct answer ranks in your raw vector similarity&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;No reranker. No full pipeline. One embedding call and one round of cosine similarity.&lt;/p&gt;

&lt;p&gt;That number decides everything downstream:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If it ranks &lt;strong&gt;outside&lt;/strong&gt; your candidate depth → going deeper &lt;em&gt;might&lt;/em&gt; help (only might — see below)&lt;/li&gt;
&lt;li&gt;If it ranks &lt;strong&gt;inside&lt;/strong&gt; your candidate depth → going deeper &lt;strong&gt;definitely won't&lt;/strong&gt;, and your bottleneck is somewhere else&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without that probe I'd have swept &lt;code&gt;top-k = 10 / 20 / 30 / 50 / 80 / 100&lt;/code&gt;, six configurations, 29 questions each, tens of seconds per question — &lt;strong&gt;about an hour&lt;/strong&gt;. After the ten-second probe, those six configurations collapsed to &lt;strong&gt;one&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where this started: a to-do I wrote for myself
&lt;/h2&gt;

&lt;p&gt;I was evaluating a "the data never leaves the building" knowledge base for a client, and wrote up my research as a handoff doc for the next session. The to-do was specific:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Step 1: top-k 10 → 50 (zero cost, zero VRAM)&lt;/strong&gt; — a public benchmark on 3,493 traditional-Chinese questions shows a weak embedding plus a reranker reaching 98%+ at k=100, beating every single-stage retriever.&lt;br&gt;
⚠️ Don't just crank it to the max (a SIGIR 2025 paper reports that over-reranking degrades results); there's a sweet spot and you have to sweep for it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 2: &lt;code&gt;bge-reranker-base&lt;/code&gt; → &lt;code&gt;bge-reranker-v2-m3&lt;/code&gt;&lt;/strong&gt; (+1.3GB, same architecture, same API, no prompt changes)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It reads as completely reasonable. External benchmarks behind it, a counter-example warning, a clear execution order.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Both steps ended up as "don't."&lt;/strong&gt; And that "zero cost" in step 1 is wrong.&lt;/p&gt;




&lt;h2&gt;
  
  
  First cut: the cheap adjacent question
&lt;/h2&gt;

&lt;p&gt;Before touching anything I looked at the table from my previous round:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Configuration&lt;/th&gt;
&lt;th&gt;K=1&lt;/th&gt;
&lt;th&gt;K=3&lt;/th&gt;
&lt;th&gt;K=5&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Vector only&lt;/td&gt;
&lt;td&gt;84%&lt;/td&gt;
&lt;td&gt;88%&lt;/td&gt;
&lt;td&gt;92%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Vector + rerank (candidate depth 30)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;88%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;96%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;96%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vector + rerank (candidate depth 10)&lt;/td&gt;
&lt;td&gt;84%&lt;/td&gt;
&lt;td&gt;92%&lt;/td&gt;
&lt;td&gt;92%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;96% is 24 out of 25. &lt;strong&gt;And the miss is always the same question.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That sentence matters, because it &lt;strong&gt;caps&lt;/strong&gt; the question "how much can deeper candidates buy me?" — &lt;strong&gt;at most one question, i.e. 4 percentage points.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So the question changed from "which k is best?" into something far cheaper: &lt;strong&gt;did the reranker ever even see the correct answer for that one question?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I wrote a ten-second probe: for each of the 25 questions, where does the passage containing the correct answer rank by raw vector similarity?&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Rank of the first correct passage&lt;/th&gt;
&lt;th&gt;Questions&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1st&lt;/td&gt;
&lt;td&gt;22&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3rd / 5th / 21st&lt;/td&gt;
&lt;td&gt;1 each&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;83rd&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1&lt;/strong&gt; (the usual suspect)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One table cut six configurations down to one:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Depth 10 → correct answer present in candidates for 23/25 (&lt;strong&gt;ceiling 92%&lt;/strong&gt;)&lt;/li&gt;
&lt;li&gt;Depth 30 → 24/25 (&lt;strong&gt;ceiling 96%&lt;/strong&gt;)&lt;/li&gt;
&lt;li&gt;Depth &lt;strong&gt;50 and 80&lt;/strong&gt; → still 24/25, &lt;strong&gt;because that question sits at 83&lt;/strong&gt;. Identical ceiling to 30 ⇒ &lt;strong&gt;no need to run them&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Depth 100 → 25/25 (&lt;strong&gt;ceiling 100%&lt;/strong&gt;)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And check the calibration: depth 10's ceiling is 92%, measured 92%. Depth 30's ceiling is 96%, measured 96%. &lt;strong&gt;Both match to the decimal.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;At that moment I derived a conclusion: &lt;em&gt;"whatever's in the candidate set, the reranker pulls into the top 5 ⇒ this pipeline's bottleneck is how deep we fish, not how good the reranker is."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The derivation itself is sound (the miss set must contain the out-of-candidates set, and the two have equal size, so they're the same set).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Then I ran depth 100, and it falsified that conclusion.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Depth 100: ceiling 100%, measured 96%
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Depth 30&lt;/th&gt;
&lt;th&gt;Depth 100&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Retrieval hit rate (K=5)&lt;/td&gt;
&lt;td&gt;96% (24/25)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;96% (24/25)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Miss&lt;/td&gt;
&lt;td&gt;that question&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;that question&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Median latency per question&lt;/td&gt;
&lt;td&gt;6,834 ms&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;37,651 ms&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The correct answer ranks 83rd, so it &lt;strong&gt;did&lt;/strong&gt; make it into the 100 candidates. The reranker saw it. &lt;strong&gt;And then ranked it outside the top 5.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So the right statement isn't the one I derived — it's a more useful one:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Candidate depth is necessary, not sufficient.&lt;/strong&gt;&lt;br&gt;
For that question to land, two things must both hold: ① you fish deep enough to see it, and ② the reranker is willing to rank it near the top.&lt;br&gt;
Depth 30 fails at ①. Depth 100 clears ① and fails at ②.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Incidentally: &lt;strong&gt;"zero cost" is wrong.&lt;/strong&gt; A reranker is a cross-encoder, and its cost is &lt;strong&gt;linear in the candidate count&lt;/strong&gt; — send it 100 candidates and it computes "how relevant is this passage to this question" 100 times. Going from depth 30 to 100 took each question from 6.8 to 37.7 seconds: &lt;strong&gt;5.5× the latency, for zero extra correct answers.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;"Zero VRAM" is right. "Zero cost" is not. And what it spends is precisely the one budget in this setup that had already bottomed out.&lt;/p&gt;




&lt;h2&gt;
  
  
  Second cut: swap in a "newer" reranker
&lt;/h2&gt;

&lt;p&gt;If the bottleneck is the reranker, replace it. My to-do said &lt;code&gt;bge-reranker-v2-m3&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First discovery: that model isn't available on my stack.&lt;/strong&gt; I run ONNX through fastembed on CPU, and its cross-encoder list has six models — no v2-m3. Only two are multilingual: the &lt;code&gt;bge-reranker-base&lt;/code&gt; I'm already on (a 2023 base-tier model), and &lt;code&gt;jinaai/jina-reranker-v2-base-multilingual&lt;/code&gt; (1.11 GB, &lt;strong&gt;2024 generation, marketed on multilingual&lt;/strong&gt;).&lt;/p&gt;

&lt;p&gt;Actually using v2-m3 would mean swapping the whole service from ONNX to PyTorch. That's a separate project. So I swapped in the jina model — same API, one parameter changed, no prompt edits.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It lost in three directions at once:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;reranker&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;K=1&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;K=3&lt;/th&gt;
&lt;th&gt;K=5&lt;/th&gt;
&lt;th&gt;Pure Chinese prose questions, K=1&lt;/th&gt;
&lt;th&gt;Latency/question&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;bge-reranker-base&lt;/code&gt; (2023)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;88%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;96%&lt;/td&gt;
&lt;td&gt;96%&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;88.9%&lt;/strong&gt; (8/9)&lt;/td&gt;
&lt;td&gt;6,834 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;jina-reranker-v2-multilingual&lt;/code&gt; (2024)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;76%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;96%&lt;/td&gt;
&lt;td&gt;96%&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;66.7%&lt;/strong&gt; (6/9)&lt;/td&gt;
&lt;td&gt;8,219 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;K=1 dropped 12 points&lt;/strong&gt; (3 questions). K=1 is the column that most affects the quote — it decides how many documents the client has to feed the LLM, which decides whether they need to buy that GPU.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The worst drop is on pure Chinese prose&lt;/strong&gt; (88.9% → 66.7%). A &lt;strong&gt;multilingual-first&lt;/strong&gt; 2024 model losing to a 2023 base-tier model on traditional Chinese prose.&lt;/li&gt;
&lt;li&gt;It's also 1.2× slower.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;But the fourth reason is the fatal one, and I nearly missed it.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The fourth reason I nearly missed
&lt;/h2&gt;

&lt;p&gt;My previous round had turned up a bonus finding: &lt;strong&gt;the reranker's score itself works as a "refuse to answer without calling the LLM" gate.&lt;/strong&gt; Below a threshold, return "not found" — skip generation entirely. Fast, and structurally incapable of making things up.&lt;/p&gt;

&lt;p&gt;For this client that matters more than usual, because making things up is exactly what they're afraid of.&lt;/p&gt;

&lt;p&gt;The scoring script prints a "median top-1 score" line each round:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;bge-reranker-base&lt;/code&gt;: answerable 5.26 / unanswerable −2.14&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;jina&lt;/code&gt;: answerable 0.84 / unanswerable −0.76&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That looks like "the gap got smaller." What I wrote in the report at the time was "this doesn't let us say jina made the gate worse [unknown]."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;But the statistic that finding needs isn't the median.&lt;/strong&gt; It needs &lt;strong&gt;the minimum score among answerable questions&lt;/strong&gt; against &lt;strong&gt;the maximum among unanswerable ones&lt;/strong&gt; — because two distributions can have very different medians and still have overlapping tails.&lt;/p&gt;

&lt;p&gt;The per-question scores were already in the scoring run's JSON. I computed it afterwards at zero extra cost:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;reranker&lt;/th&gt;
&lt;th&gt;Answerable &lt;strong&gt;min&lt;/strong&gt;
&lt;/th&gt;
&lt;th&gt;Unanswerable &lt;strong&gt;max&lt;/strong&gt;
&lt;/th&gt;
&lt;th&gt;Gap&lt;/th&gt;
&lt;th&gt;Can one threshold separate them?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;&lt;code&gt;bge-reranker-base&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;−0.5237&lt;/td&gt;
&lt;td&gt;−1.8008&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+1.2771&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;✅ &lt;strong&gt;yes&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;&lt;code&gt;jina-v2-multilingual&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;−0.9005&lt;/td&gt;
&lt;td&gt;−0.6736&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;−0.2269&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;❌ &lt;strong&gt;overlap&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vector only (no rerank)&lt;/td&gt;
&lt;td&gt;0.5895&lt;/td&gt;
&lt;td&gt;0.6377&lt;/td&gt;
&lt;td&gt;−0.0482&lt;/td&gt;
&lt;td&gt;❌ overlap&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Switching to jina means: K=1 down 12 points, latency up 1.2×, and then you also forfeit the entire refusal defence.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And the median made it look like it merely "got smaller." That's the most valuable lesson of this round: &lt;strong&gt;the statistic a report prints is not necessarily the statistic your conclusion needs.&lt;/strong&gt; I turned it into a command (reads the stored scoring JSON, zero cost) and left a test inside it: fixtures of answerable &lt;code&gt;[9, 9, 9, −1]&lt;/code&gt; and unanswerable &lt;code&gt;[0]&lt;/code&gt; — medians differ by 9, looks trivially separable; the truth is min −1 &amp;lt; max 0, &lt;strong&gt;overlap&lt;/strong&gt;. That test exists so that if anyone ever switches it back to medians, it goes red.&lt;/p&gt;




&lt;h2&gt;
  
  
  What's actually wrong with that one question: I tried a third road too, and it's also closed
&lt;/h2&gt;

&lt;p&gt;I had one hypothesis left. Looking at the top-5 for that question across four configurations, I noticed &lt;strong&gt;all five slots in every configuration were filled by the same class of file&lt;/strong&gt; — my own blog drafts. Derivative articles discussing the same incidents, not the canonical records of those incidents.&lt;/p&gt;

&lt;p&gt;That matches a published finding: &lt;strong&gt;what hurts RAG isn't "irrelevant," it's "relevant but wrong."&lt;/strong&gt; It also matches my own earlier observation that this corpus writes about the same event a dozen times over.&lt;/p&gt;

&lt;p&gt;So I added an exclusion rule, filtered the entire blog-source directory out of the candidates, and re-ran that question.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Still missed.&lt;/strong&gt; And the correct answer didn't even make the filtered 30 candidates — excluding the derivative material just pushed &lt;strong&gt;other non-answers&lt;/strong&gt; up.&lt;/p&gt;

&lt;p&gt;Three roads walked:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Hypothesis&lt;/th&gt;
&lt;th&gt;Measured&lt;/th&gt;
&lt;th&gt;Conclusion&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Not deep enough&lt;/td&gt;
&lt;td&gt;depth 100 got it into candidates, reranker still didn't rank it&lt;/td&gt;
&lt;td&gt;❌ necessary, not sufficient&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reranker too old&lt;/td&gt;
&lt;td&gt;swapped to a 2024 multilingual model; still missed, every other metric regressed&lt;/td&gt;
&lt;td&gt;❌ actively worse&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Corpus noise drowning the canon&lt;/td&gt;
&lt;td&gt;excluded the biggest source of derivatives; it can't even reach candidates&lt;/td&gt;
&lt;td&gt;❌ &lt;strong&gt;hypothesis falsified&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Embedding not good enough&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;untested&lt;/td&gt;
&lt;td&gt;⬅️ &lt;strong&gt;only one left&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That question asks "what has to be done to a temp table containing personal data for it to count as safe, and which incident forced that?" — and the file with the answer is anchored on a database table name. &lt;strong&gt;The question doesn't contain that table name; only the answer does.&lt;/strong&gt; It's a pure semantic-bridging task, and my embedding model puts it 83rd.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This round proved the embedding is the bottleneck, by elimination&lt;/strong&gt; — every knob outside the embedding was turned to the stop with no effect. My original plan wrote the precondition for "swap the embedding" as "only swap once you've proven embedding is the bottleneck." That precondition now holds; it just got proven by exclusion.&lt;/p&gt;

&lt;p&gt;My next cut will be &lt;strong&gt;query rewriting&lt;/strong&gt; (have the LLM expand the question into a hypothetical answer first, then embed that) rather than swapping the embedding outright — much cheaper, and among the open-source projects I've read, the one that handles Chinese most seriously does its work on the &lt;strong&gt;query construction&lt;/strong&gt; side (part-of-speech weighting, positional proximity weighting), not by changing models.&lt;/p&gt;




&lt;h2&gt;
  
  
  A red flag that goes straight into the quote
&lt;/h2&gt;

&lt;p&gt;This round also turned up something I wasn't looking for.&lt;/p&gt;

&lt;p&gt;To get a same-day control, I re-ran the depth-30 configuration. Recall was identical (96%, same question missed) — &lt;strong&gt;comparability confirmed&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;But &lt;strong&gt;latency was 6,834 ms per question. A week earlier I'd measured the same configuration at 18,152 ms.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Same model, same corpus, same question bank, same laptop. A factor of 2.66.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Two consequences:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;I'd been comparing a fresh depth-100 number (37,651 ms) against last week's 18,152 ms and computing "2.07×." &lt;strong&gt;Wrong.&lt;/strong&gt; The correct same-day ratio is &lt;strong&gt;5.51×&lt;/strong&gt;. &lt;strong&gt;A latency denominator has to be produced in the same round as its numerator&lt;/strong&gt; — the same discipline as "pin the corpus to a commit," which I'd applied to the corpus and not to latency.&lt;/li&gt;
&lt;li&gt;Worse: the conclusion I'd written a week earlier was "&lt;strong&gt;18 seconds per question — unusable for interactive Q&amp;amp;A&lt;/strong&gt;." &lt;strong&gt;That sentence cannot be measured today&lt;/strong&gt; (6.8 seconds). 6.8 seconds is still slow, but "slow" and "unusable" are completely different things in a client quote — and that sentence was headed into the proposal.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I don't know the source of the variance yet (thermal throttling? something else running? a package version?). &lt;strong&gt;All I can state is that corpus, question bank, and model are ruled out&lt;/strong&gt;, because all three were re-verified identical the same day.&lt;/p&gt;

&lt;p&gt;So the rule now is: &lt;strong&gt;any latency number going into a quote must have its own control, measured in the same round, on the same machine, with no other load.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Net result
&lt;/h2&gt;

&lt;p&gt;Four scoring runs, three probes, roughly an hour of machine time. What it bought:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Don't touch either knob.&lt;/strong&gt; Stay on &lt;code&gt;bge-reranker-base&lt;/code&gt; with candidate depth 30.&lt;/p&gt;

&lt;p&gt;That sounds like nothing happened. In practice I got three things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;A number I can commit to&lt;/strong&gt;: on a 25-question benchmark, 96% retrieval at K=5, 88% at K=1, plus a refusal gate that doesn't depend on the LLM.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A number I can't commit to&lt;/strong&gt;: 100% retrieval. &lt;strong&gt;Measured as unbuyable&lt;/strong&gt; — deeper candidates don't buy it, and neither does the newest reranker I can actually run. That's far more valuable than "I think we can probably get there," because it lets me write a ceiling into the proposal.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Two numbers not to quote&lt;/strong&gt;: any local CPU latency (2.66× day-to-day variance), and any presumption that "a newer model will be better."&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Plus one observation I can sell: &lt;strong&gt;derivative copies inside a client's corpus&lt;/strong&gt; (drafts, meeting notes, several versions of the same policy) will systematically outrank the canon — my own blog drafts pushing my own canonical records out of the top 5 is the live example. But &lt;strong&gt;don't sell "curate the corpus" as a cure-all&lt;/strong&gt;, because in the same round I measured it doing nothing for that question.&lt;/p&gt;




&lt;h2&gt;
  
  
  Four things you can take away
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Before turning any expensive knob, find a ceiling you can compute in ten seconds.&lt;/strong&gt; "At most how many questions can this buy?" routinely makes an entire sweep plan disappear.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Newer, bigger, multilingual-first" does not mean "better on your corpus."&lt;/strong&gt; Marketing labels aren't evidence. I put a 2024-generation multilingual reranker in and it lost to a 2023 base-tier model on traditional Chinese prose.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The statistic a report prints isn't necessarily the statistic your conclusion needs.&lt;/strong&gt; A median can't answer "are these separable?" Ask which statistic your claim rests on, then go compute that one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't quote a latency number without a control.&lt;/strong&gt; The same configuration varies 2.66× day to day on my machine, and it nearly put the word "unusable" into a client proposal.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you're doing this kind of measure-before-you-change work, I've written others in the same family: &lt;a href="https://dev.to/content/zero-of-ninetysix-was-an-empty-wallet-en"&gt;my test report printed "0/96, 0% pass rate" — the truth was my account was out of credit&lt;/a&gt;, on why a red light needs as much suspicion as a green one; and &lt;a href="https://dev.to/content/slimming-skill-descriptions-ab-test-en"&gt;the A/B test where I cut 41 tool descriptions in half&lt;/a&gt;, which is the same discipline applied to a change I was confident about — an absolute number can't tell you whether something dropped, so you have to go build the control arm.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;本文原載於我的部落格：&lt;a href="https://coffeeshooters.com/content/rag-two-knobs-no-gain-en?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=blog-rag-two-knobs-no-gain-en" rel="noopener noreferrer"&gt;I turned both knobs on my on-prem Chinese RAG all the way up. It didn't get a single extra question right&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>solodev</category>
    </item>
  </channel>
</rss>
