<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Devanshu Biswas</title>
    <description>The latest articles on DEV Community by Devanshu Biswas (@dev48v).</description>
    <link>https://dev.to/dev48v</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3929385%2F75a3696c-143d-4252-ba59-6ed4083ca827.jpg</url>
      <title>DEV Community: Devanshu Biswas</title>
      <link>https://dev.to/dev48v</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/dev48v"/>
    <language>en</language>
    <item>
      <title>Your Withdrawal Rate Moves Time-to-FI by 31 Months. Your Balance, the Number in the Biggest Font, Moves It 3.</title>
      <dc:creator>Devanshu Biswas</dc:creator>
      <pubDate>Thu, 17 Sep 2026 12:35:10 +0000</pubDate>
      <link>https://dev.to/dev48v/your-withdrawal-rate-moves-time-to-fi-by-31-months-your-balance-the-number-in-the-biggest-font-23d7</link>
      <guid>https://dev.to/dev48v/your-withdrawal-rate-moves-time-to-fi-by-31-months-your-balance-the-number-in-the-biggest-font-23d7</guid>
      <description>&lt;p&gt;A freedom calculator answers three questions — how long the money lasts, what number ends the job, how far away that is — and they are &lt;strong&gt;one recurrence asked from two sides&lt;/strong&gt;. Build it exactly, measure it honestly, and the finding is not that the arithmetic is wrong. It is that the answer is mostly a statement about assumptions, and the ones that move it most are the ones the genre does not show you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Weekend Builds Vol 4 · #05 · the finale — Vol 4 is now complete at 5 of 5.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;👉 &lt;strong&gt;Live, runs in your browser:&lt;/strong&gt; &lt;a href="https://dev48v.infy.uk/agentlab/vol4-05-freedom-calculator.html" rel="noopener noreferrer"&gt;https://dev48v.infy.uk/agentlab/vol4-05-freedom-calculator.html&lt;/a&gt;&lt;br&gt;
👉 &lt;strong&gt;PUBLIC, MIT, &lt;code&gt;dependencies = []&lt;/code&gt;:&lt;/strong&gt; &lt;a href="https://github.com/dev48v/freedom-calculator" rel="noopener noreferrer"&gt;https://github.com/dev48v/freedom-calculator&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python &lt;span class="nt"&gt;-m&lt;/span&gt; freedom_calculator                &lt;span class="c"&gt;# runway, the FI number, time to FI&lt;/span&gt;
python &lt;span class="nt"&gt;-m&lt;/span&gt; freedom_calculator &lt;span class="nt"&gt;--sensitivity&lt;/span&gt;  &lt;span class="c"&gt;# which input the answer rests on&lt;/span&gt;
python &lt;span class="nt"&gt;-m&lt;/span&gt; freedom_calculator &lt;span class="nt"&gt;--assumptions&lt;/span&gt;  &lt;span class="c"&gt;# the model, and what it costs&lt;/span&gt;
python &lt;span class="nt"&gt;-m&lt;/span&gt; freedom_calculator &lt;span class="nt"&gt;--float&lt;/span&gt;        &lt;span class="c"&gt;# why money is not a float here&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Declared inputs, derived signs, measured spans
&lt;/h2&gt;

&lt;p&gt;You declare five numbers. The scenario on the page: 42,000 balance, 3,200 a month out, 5,400 a month in, 5% real return, 4% withdrawal rate. From those: an FI number of &lt;strong&gt;960,000&lt;/strong&gt; (25× annual spend) and &lt;strong&gt;231 months&lt;/strong&gt; to reach it.&lt;/p&gt;

&lt;p&gt;Every direction below is &lt;strong&gt;derived on paper&lt;/strong&gt; from the recurrence, then &lt;strong&gt;measured&lt;/strong&gt; by bumping the input and watching the answer. The two are kept apart so they can be checked against each other, and the suite asserts they agree for every input at six bump sizes — plus a non-vacuity test, because a calculator that ignored its inputs entirely would pass a sign check perfectly.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;±10% on…&lt;/th&gt;
&lt;th&gt;months, down / base / up&lt;/th&gt;
&lt;th&gt;span&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;what you spend a month&lt;/td&gt;
&lt;td&gt;198 / 231 / 269&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;71&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;what you earn a month&lt;/td&gt;
&lt;td&gt;271 / 231 / 202&lt;/td&gt;
&lt;td&gt;69&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;withdrawal rate&lt;/strong&gt; (usually hidden)&lt;/td&gt;
&lt;td&gt;248 / 231 / 217&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;31&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;real return you assume&lt;/td&gt;
&lt;td&gt;241 / 231 / 222&lt;/td&gt;
&lt;td&gt;19&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;balance you already have&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;233 / 231 / 230&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The row to sit with is the last one. The genre is arranged around a number worth about a tenth of what the hidden constant is worth.&lt;/p&gt;

&lt;p&gt;There is also a row that must be exactly zero: &lt;strong&gt;runway does not depend on the withdrawal rate at all.&lt;/strong&gt; It is a drawdown question. The derived sign is 0 and the measured direction must be 0 too — not small, not usually, &lt;em&gt;exactly&lt;/em&gt;, at every bump size on every scenario. That identity is what stops the two halves of the calculator leaking into each other.&lt;/p&gt;

&lt;h2&gt;
  
  
  The argument people actually have, priced
&lt;/h2&gt;

&lt;p&gt;The 4% rule is a rule about a number under argument — 3% for the cautious, 5% for the cheerful:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;withdrawal rate&lt;/th&gt;
&lt;th&gt;FI number&lt;/th&gt;
&lt;th&gt;months&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;3%&lt;/td&gt;
&lt;td&gt;1,280,000&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;278&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4%&lt;/td&gt;
&lt;td&gt;960,000&lt;/td&gt;
&lt;td&gt;231&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5%&lt;/td&gt;
&lt;td&gt;768,000&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;198&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;512,000&lt;/strong&gt; of spread — &lt;strong&gt;53.3%&lt;/strong&gt; of the headline number — and &lt;strong&gt;80 months&lt;/strong&gt; of your life, from one constant most calculators will not let you type.&lt;/p&gt;

&lt;p&gt;Forgetting to subtract inflation is worth &lt;strong&gt;43 months&lt;/strong&gt;: a nominal 8% instead of a real 5% brings the date from 231 months forward to 188, on which date you actually hold ~717,562 against a target of 960,000.&lt;/p&gt;

&lt;h2&gt;
  
  
  Money is an exact rational because the output is an integer
&lt;/h2&gt;

&lt;p&gt;Not fastidiousness. The output is a &lt;strong&gt;month count that turns on a comparison&lt;/strong&gt;, and a comparison decided by the last bit of a binary approximation is a month count decided by the last bit of a binary approximation.&lt;/p&gt;

&lt;p&gt;Over 120,000 balances that are exactly &lt;em&gt;n&lt;/em&gt; months of spend, &lt;code&gt;floor(balance / spend)&lt;/code&gt; in doubles answers &lt;em&gt;n&lt;/em&gt;−1 in &lt;strong&gt;12,736&lt;/strong&gt; of them — &lt;strong&gt;10.61%&lt;/strong&gt;. Exact rationals: &lt;strong&gt;0&lt;/strong&gt;. Whole months, not cents.&lt;/p&gt;

&lt;p&gt;The zero-net-burn case is worse. Summed as floats, the same basket of line items lands on a net burn of &lt;code&gt;2.27e-13&lt;/code&gt; instead of zero, and a calculator that divides reports 43,980,465,111,040,000 months with a straight face. This package reaches the unbounded answer with a comparison and &lt;strong&gt;no division at all&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// runway and time-to-FI are the SAME recurrence, asked from two sides&lt;/span&gt;
&lt;span class="c1"&gt;//     b[m+1] = b[m] * (1 + r) + flow          flow = income - spend&lt;/span&gt;
&lt;span class="c1"&gt;// the zero-net-burn identity, decided BEFORE any division ever happens&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;cmp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;mul&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;balance&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nx"&gt;flow&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nx"&gt;ZERO&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;unbounded&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A test walks the syntax tree of &lt;code&gt;runway()&lt;/code&gt; and asserts there is no division operator in it, with docstrings stripped first so it cannot pass on the paragraph explaining that the division is not there.&lt;/p&gt;

&lt;h2&gt;
  
  
  Same returns, reversed, a 17× different answer
&lt;/h2&gt;

&lt;p&gt;The model has no sequence-of-returns risk in it. That is a limitation, and here is what it costs — with no RNG anywhere. One declared list of 30 yearly real returns, mean &lt;strong&gt;exactly&lt;/strong&gt; the same 5% the smooth model uses, applied forward, then reversed. Start at the FI number, draw 4%:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;order&lt;/th&gt;
&lt;th&gt;ending balance&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;forward (crash at the front)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~127,381&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;reversed&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~2,163,974&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;the smooth model's answer&lt;/td&gt;
&lt;td&gt;~1,625,807&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;17.0× from order alone.&lt;/strong&gt; Same numbers, same mean, same multiset, same product — and the control proves the product is identical: withdraw nothing and the two orders end on the &lt;strong&gt;identical cent&lt;/strong&gt;, 4,176,967.52, because multiplication commutes. The entire gap is &lt;em&gt;order&lt;/em&gt;, and the smooth model's answer sits between the two giving no hint which one you got.&lt;/p&gt;

&lt;p&gt;The sequence was built to put the crash at the front. Stated, not hidden — the size of the crash is not the point. The point is that reversing a list changes the answer &lt;em&gt;at all&lt;/em&gt;, which the smooth model says is impossible.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this is not
&lt;/h2&gt;

&lt;p&gt;Deterministic arithmetic on declared inputs. No RNG, no Monte Carlo, no historical data. Constant real returns, constant burn in today's money, &lt;strong&gt;no taxes&lt;/strong&gt;, no lumpy years, no property, no state pension. The monthly rate is the annual rate over twelve rather than its twelfth root — the usual convention, and slightly generous. Anything past a 100-year horizon is reported as &lt;em&gt;beyond the horizon&lt;/em&gt; rather than as a number, because a model that answers "3,847 months" is not being more precise, it is being less honest.&lt;/p&gt;

&lt;p&gt;The JavaScript on the page is exact rationals over BigInt, checked against the Python's &lt;code&gt;reference.json&lt;/code&gt; numerator for denominator, and against a third implementation in the verifier that simulates month by month instead of using the closed form.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;50 pytest · 239 verifier asserts · 44 in-page self-checks · 0 failures.&lt;/strong&gt; MIT, Python standard library only.&lt;/p&gt;

</description>
      <category>python</category>
      <category>opensource</category>
      <category>programming</category>
      <category>webdev</category>
    </item>
    <item>
      <title>180 RAG Citation Defects Only a Free id Lookup Catches, 60 Only the Grader Does: Neither Subsumes the Other</title>
      <dc:creator>Devanshu Biswas</dc:creator>
      <pubDate>Thu, 17 Sep 2026 12:34:28 +0000</pubDate>
      <link>https://dev.to/dev48v/180-rag-citation-defects-only-a-free-id-lookup-catches-60-only-the-grader-does-neither-subsumes-id4</link>
      <guid>https://dev.to/dev48v/180-rag-citation-defects-only-a-free-id-lookup-catches-60-only-the-grader-does-neither-subsumes-id4</guid>
      <description>&lt;p&gt;A RAG answer says &lt;code&gt;[c091]&lt;/code&gt; and everybody relaxes. There are two checks you can run on that bracket, and they catch different things.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Citation resolution&lt;/strong&gt; is pure code: is &lt;code&gt;c091&lt;/code&gt; actually one of the chunks we assembled for &lt;em&gt;this&lt;/em&gt; question? Zero model calls. &lt;strong&gt;Faithfulness&lt;/strong&gt; is model-scored: does the assembled context support the claim?&lt;/p&gt;

&lt;p&gt;👉 &lt;strong&gt;Live, runs in your browser:&lt;/strong&gt; &lt;a href="https://dev48v.infy.uk/ai/days/day80-citation-verification.html" rel="noopener noreferrer"&gt;https://dev48v.infy.uk/ai/days/day80-citation-verification.html&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The 2×2, and why both off-diagonals matter
&lt;/h2&gt;

&lt;p&gt;420 seeded answers, 7 defect classes, 60 each:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;faithfulness says supported&lt;/th&gt;
&lt;th&gt;says unsupported&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;resolution passes&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;120&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;60&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;resolution fails&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;180&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;60&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If either off-diagonal cell were empty, one check would subsume the other and you could drop it. Both are populated &lt;strong&gt;by measurement&lt;/strong&gt;: 180 answers where the identifier is wrong and the prose is fine, 60 where the identifier is fine and the prose is invented.&lt;/p&gt;

&lt;p&gt;A pipeline running one check is not running a cheaper version of the other. It is running a different check whose blind spot is exactly what the other one is for.&lt;/p&gt;

&lt;h2&gt;
  
  
  The class people get wrong
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;stale_id&lt;/code&gt;: a &lt;strong&gt;real&lt;/strong&gt; store id that was &lt;strong&gt;not in this assembled context&lt;/strong&gt;. It looks real. It resolves perfectly against your vector store.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;check&lt;/th&gt;
&lt;th&gt;catches&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;resolution against the &lt;strong&gt;assembled context&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;60 / 60&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;resolution against the &lt;strong&gt;chunk store&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0 / 60&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Same check, one wrong set. That is the entire difference between a working guard and a decorative one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// check 1 — resolution. No model. This is the entire thing.&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;inContext&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;assembled&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;          &lt;span class="c1"&gt;// ids WE assembled&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;resolves&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cites&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;every&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;inContext&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;has&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;pass&lt;/span&gt;      &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cites&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;resolves&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;// the bug: checking the STORE instead of the assembled context&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;wrong&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cites&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;every&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;store&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;has&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt; &lt;span class="c1"&gt;// c091 is real...&lt;/span&gt;
                                                       &lt;span class="c1"&gt;// ...just not here&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Both blind spots, stated as numbers
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Resolution never reads the claim.&lt;/strong&gt; An unsupported sentence carrying an id that genuinely is in the context is a clean pass — &lt;strong&gt;0 of 60&lt;/strong&gt; caught.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A context-scoped faithfulness grader is id-blind.&lt;/strong&gt; Claim in, assembled chunks in, supported-or-not out. An invented id, a stale one, or no id at all does not change its input by one token — &lt;strong&gt;0 of 180&lt;/strong&gt; citation defects caught. That is not inferred: the self-check replaces every citation with nonsense and measures the verdict move on &lt;strong&gt;0 of 420&lt;/strong&gt; answers.&lt;/p&gt;

&lt;p&gt;Neither number moves by tuning a threshold. Both are structural.&lt;/p&gt;

&lt;h2&gt;
  
  
  The residual neither one catches
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;60 answers cite a chunk that &lt;em&gt;is&lt;/em&gt; in the assembled context and does &lt;em&gt;not&lt;/em&gt; support the claim&lt;/strong&gt;, while some other assembled chunk does. Resolution passes — the id is in the context. Context-scoped faithfulness passes — the context supports the claim. Both answer their own question correctly; neither question is "does &lt;em&gt;this&lt;/em&gt; chunk support &lt;em&gt;this&lt;/em&gt; sentence".&lt;/p&gt;

&lt;p&gt;A citation-level grader does catch all 60, and costs one call &lt;strong&gt;per citation&lt;/strong&gt; rather than per answer — 180 here against 420, and far worse on answers citing five sources each. It also cannot run on a citation that did not resolve: you cannot hand a grader a chunk you could not find. Third layer, never a replacement.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the free check is worth
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;defects caught (of 360)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;faithfulness alone&lt;/td&gt;
&lt;td&gt;120&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;resolution alone&lt;/td&gt;
&lt;td&gt;240&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;union&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;300&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;plus citation-level&lt;/td&gt;
&lt;td&gt;360&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Zero false alarms from either check on the clean class. Resolution costs &lt;strong&gt;360 comparisons and 0 model calls&lt;/strong&gt; over the whole corpus; grading everything costs &lt;strong&gt;420&lt;/strong&gt; calls; running the free check first and grading only what survives costs &lt;strong&gt;180&lt;/strong&gt; — &lt;strong&gt;240 model calls removed with the union catch rate unchanged&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;One cost that is not money: everything short-circuiting skips sits in the cell where &lt;em&gt;the answer text was fine&lt;/em&gt;. For those the repair is re-citing, not re-answering, and an ordering that dumps them into one undifferentiated "failed" bucket sends correct answers back to be rewritten.&lt;/p&gt;

&lt;h2&gt;
  
  
  The limitation to read first
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The entailment check here is a stand-in, not a model&lt;/strong&gt; — a stated token rule over 12 disjoint topic vocabularies, which separates supported from unsupported with no overlap at all. Real graders are wrong a great deal, and that error is part of why the free check is worth keeping. So read the faithfulness column as a &lt;em&gt;ceiling&lt;/em&gt; and the resolution column as exact, because that one is code. The corpus is synthetic, one claim per answer, one citation per claim, no multi-hop, no reranker, no abstention. And nothing here asks whether the &lt;em&gt;source&lt;/em&gt; is any good.&lt;/p&gt;

&lt;p&gt;38 in-page checks, 138 verifier asserts, 0 failures.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>programming</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Pack Four Docs Without the Block Mask and Doc 4 Spends 83.78% of Its Attention on Documents It Never Saw</title>
      <dc:creator>Devanshu Biswas</dc:creator>
      <pubDate>Thu, 17 Sep 2026 12:33:46 +0000</pubDate>
      <link>https://dev.to/dev48v/pack-four-docs-without-the-block-mask-and-doc-4-spends-8378-of-its-attention-on-documents-it-54hp</link>
      <guid>https://dev.to/dev48v/pack-four-docs-without-the-block-mask-and-doc-4-spends-8378-of-its-attention-on-documents-it-54hp</guid>
      <description>&lt;p&gt;Sequence packing concatenates short documents into one row so you stop paying for padding. On the length distribution here that lifts utilisation from 27.98% to 99.81%.&lt;/p&gt;

&lt;p&gt;But attention inside a row is &lt;strong&gt;global&lt;/strong&gt;. Without a block-diagonal mask, the fourth document in the row spends &lt;strong&gt;83.78%&lt;/strong&gt; of its layer-1 attention on documents it has never seen — worst single token &lt;strong&gt;94.90%&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;👉 &lt;strong&gt;Live, runs in your browser:&lt;/strong&gt; &lt;a href="https://dev48v.infy.uk/dl/day80-sequence-packing.html" rel="noopener noreferrer"&gt;https://dev48v.infy.uk/dl/day80-sequence-packing.html&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The ground truth is declared; everything else is measured
&lt;/h2&gt;

&lt;p&gt;Every document is first pushed through the stack &lt;strong&gt;on its own&lt;/strong&gt;, ordinary causal mask, positions 0…L−1. That unpacked run is the ground truth, and every packed configuration is scored against it token by token, coordinate by coordinate — not against a loss curve, which is exactly the instrument that cannot see either bug.&lt;/p&gt;

&lt;p&gt;The only thing that differs between a correct packed row and a broken one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;allow&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;cfg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;block&lt;/span&gt;
  &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;docIds&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;docIds&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;   &lt;span class="c1"&gt;// block-diagonal ∩ causal&lt;/span&gt;
  &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;                                 &lt;span class="c1"&gt;// causal only: the mask bug&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;posIds&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;cfg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;reset&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;local&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tokens&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The tolerance is a prediction, not a fudge
&lt;/h2&gt;

&lt;p&gt;With the mask &lt;strong&gt;and&lt;/strong&gt; per-document position reset, the packed row reproduces the unpacked run to &lt;strong&gt;2.22e-16&lt;/strong&gt; — across five independently seeded models, every token, every coordinate.&lt;/p&gt;

&lt;p&gt;It is not exactly zero for a stated reason. A batched attention kernel subtracts a row maximum before exponentiating, and takes that maximum over the &lt;em&gt;whole causal prefix of the row&lt;/em&gt;, not over the current document. Subtracting a constant leaves a softmax mathematically unchanged, so this is a reassociation that moves the last bit and nothing else. The engine reproduces that kernel behaviour rather than hiding it.&lt;/p&gt;

&lt;p&gt;The verifier then tests the explanation instead of trusting it: an independently written attention stack that subtracts &lt;strong&gt;no&lt;/strong&gt; maximum at all gives the identity at &lt;strong&gt;exactly 0&lt;/strong&gt;. So the residual is the shift, and nothing else. The asserted tolerance is 1e-12 — four orders above the residual, twelve below the smallest real bug.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two independent bugs, and the arithmetic proves they are two
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;configuration&lt;/th&gt;
&lt;th&gt;max Δ vs unpacked&lt;/th&gt;
&lt;th&gt;worst leaked mass&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;block mask + position reset&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.22e-16&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;no mask, positions reset&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.3607&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;94.90%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;mask, no position reset&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.7607&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;neither (naive packing)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.7773&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;94.90%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;0.7773 is neither the sum (1.1214) nor the max (0.7607). They interact, so you cannot subtract one to estimate the other.&lt;/p&gt;

&lt;p&gt;Two controls separate them. &lt;strong&gt;Zero the position table&lt;/strong&gt; and the position bug has to vanish exactly — there is nothing left for a wrong id to index — and it drops to 1.11e-16 while the mask bug is untouched at 0.3498. And across a growing-neighbour sweep the two respond completely differently to the same change: the &lt;strong&gt;mask&lt;/strong&gt; error correlates with leaked mass at &lt;strong&gt;r = 0.9814&lt;/strong&gt;, the &lt;strong&gt;position&lt;/strong&gt; error at &lt;strong&gt;r = 0.1384&lt;/strong&gt;, because a position offset is already wrong once the neighbour has a single token.&lt;/p&gt;

&lt;h2&gt;
  
  
  Nothing throws, and here is why
&lt;/h2&gt;

&lt;p&gt;Every attention row still sums to 1 to &lt;strong&gt;4.44e-16 in all five configurations&lt;/strong&gt;. A masked softmax renormalises over whatever it is left with, so the shape is right, the loss falls, and the model just learns something slightly wrong.&lt;/p&gt;

&lt;p&gt;Both bugs also leave the &lt;strong&gt;first&lt;/strong&gt; document in the row bit-identical — exactly 0, not nearly zero. Causal attention cannot look forward, and its positions start at 0 whether or not anyone resets them. Spot-check a packed batch on the first sequence and both bugs are invisible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three honest subtractions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Damage is not monotone in neighbour length.&lt;/strong&gt; The foreign mass is strictly increasing by construction, from 0.2261 to 0.8772 mean. The error is L×gap and the second factor moves with the content of whatever tokens get appended — there is a visible &lt;strong&gt;dip at n = 4&lt;/strong&gt;. "More neighbour, more damage" is a trend here, not a law.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Length-sorted batching already gets most of the win.&lt;/strong&gt; Bucketed batches reach &lt;strong&gt;99.06%&lt;/strong&gt; utilisation against packing's 99.81%. The 3.57× is measured against arrival-order batching, which is the baseline packing is sold against but not the only alternative.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Packing does not make attention cheaper unless the mask does.&lt;/strong&gt; Attention cost is quadratic in the row, so a packed row computed &lt;strong&gt;densely&lt;/strong&gt; costs &lt;strong&gt;1.027×&lt;/strong&gt; the padded baseline — slightly &lt;em&gt;more&lt;/em&gt;. It is the block-diagonal structure, exploited by a kernel that skips the off-diagonal blocks, that brings it to &lt;strong&gt;0.1625×&lt;/strong&gt;. The mask is not only what makes packing correct; it is what makes it fast.&lt;/p&gt;

&lt;p&gt;The models are &lt;strong&gt;not trained&lt;/strong&gt;, positions are a learned absolute table rather than RoPE, one head, two layers, float64 throughout. So which bug is &lt;em&gt;larger&lt;/em&gt; is a property of this toy's scales and is not a general claim. The claim is that both are non-zero, independent, and invisible to the loss.&lt;/p&gt;

&lt;p&gt;38 in-page checks, 126 verifier asserts, 0 failures.&lt;/p&gt;

</description>
      <category>deeplearning</category>
      <category>machinelearning</category>
      <category>python</category>
      <category>ai</category>
    </item>
    <item>
      <title>Model A Beats B on Every Slice and Loses 57.60% to 79.06%: 9,510 Reversals, 0 on a Shared Mixture</title>
      <dc:creator>Devanshu Biswas</dc:creator>
      <pubDate>Thu, 17 Sep 2026 12:33:03 +0000</pubDate>
      <link>https://dev.to/dev48v/model-a-beats-b-on-every-slice-and-loses-5760-to-7906-9510-reversals-0-on-a-shared-mixture-2187</link>
      <guid>https://dev.to/dev48v/model-a-beats-b-on-every-slice-and-loses-5760-to-7906-9510-reversals-0-on-a-shared-mixture-2187</guid>
      <description>&lt;p&gt;Two models, one benchmark, three slices. Model A is ahead on &lt;strong&gt;all three&lt;/strong&gt; — 96.0 / 60.0 / 50.0 against 88.0 / 50.0 / 40.0.&lt;/p&gt;

&lt;p&gt;Pooled, A scores &lt;strong&gt;57.60%&lt;/strong&gt; and B scores &lt;strong&gt;79.06%&lt;/strong&gt;. A 21.5-point defeat assembled entirely out of victories.&lt;/p&gt;

&lt;p&gt;👉 &lt;strong&gt;Live, runs in your browser:&lt;/strong&gt; &lt;a href="https://dev48v.infy.uk/ml/day80-simpsons-paradox.html" rel="noopener noreferrer"&gt;https://dev48v.infy.uk/ml/day80-simpsons-paradox.html&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Declared, not sampled — and no winner decided by a float
&lt;/h2&gt;

&lt;p&gt;Every count on the page is a declared integer. Every accuracy is an exact rational, and every comparison is settled by BigInt cross-multiplication, so a reversal is &lt;em&gt;proved&lt;/em&gt; rather than observed through rounding:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// every winner is decided by cross-multiplication, never by a float&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;rcmp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;b&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;l&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;n&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;d&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;n&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;d&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
                         &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;l&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nx"&gt;r&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;l&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;r&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is no p-value anywhere, because nothing is estimated. What &lt;em&gt;is&lt;/em&gt; measured is an exhaustive sweep.&lt;/p&gt;

&lt;h2&gt;
  
  
  The identity the whole thing rests on
&lt;/h2&gt;

&lt;p&gt;Pooled (micro) accuracy is not a different kind of quantity from the per-slice rates. It &lt;strong&gt;is&lt;/strong&gt; their mean, weighted by slice size:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;micro(M) = Σ c_g / Σ n_g = Σ (n_g/N) · (c_g/n_g)
macro(M) = (1/G)  ·  Σ (c_g/n_g)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The two metrics differ in exactly one place — the weights — and the entire paradox lives there. A weighted mean with &lt;strong&gt;fixed&lt;/strong&gt; non-negative weights is monotone in its inputs, so per-slice dominance would be inherited by the pooled number. The reversal &lt;em&gt;requires&lt;/em&gt; the two models to carry different weights.&lt;/p&gt;

&lt;p&gt;Which is why macro and micro disagree &lt;strong&gt;in sign in 9,510 of 9,510&lt;/strong&gt; reversals found. Not usually. All of them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The smallest reversal in existence is nine observations
&lt;/h2&gt;

&lt;p&gt;Sweeping every two-group table with per-cell n from 1 to 8 — &lt;strong&gt;3,748,096 tables&lt;/strong&gt;, of which &lt;strong&gt;774,400&lt;/strong&gt; put A strictly ahead in both groups — gives &lt;strong&gt;9,510&lt;/strong&gt; reversals. The minimum total sample across all of them is &lt;strong&gt;9&lt;/strong&gt;, reached by exactly 4 tables:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;group    model A        model B
g1       1/1 = 100%     2/3 = 66.7%
g2       1/4 =  25%     0/1 =  0%
pooled   2/5 =  40%     2/4 = 50%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Four cells and nine data points are enough to rank two models backwards. There is nothing subtle or large-sample about it.&lt;/p&gt;

&lt;p&gt;At the other end: the largest pooled reversal found is &lt;strong&gt;5/9 against 2/9&lt;/strong&gt; — 55.6 points — and the largest per-group lead that still loses is &lt;strong&gt;3/8&lt;/strong&gt;, so A is ahead by at least 37.5 points in &lt;em&gt;every&lt;/em&gt; group and loses anyway. There is no margin of per-group superiority large enough to make a pooled scalar safe to read.&lt;/p&gt;

&lt;h2&gt;
  
  
  The negative result is the one worth keeping
&lt;/h2&gt;

&lt;p&gt;Re-run the same sweep with the two models forced onto the same slice mixture:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;constraint&lt;/th&gt;
&lt;th&gt;comparisons&lt;/th&gt;
&lt;th&gt;reversals&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;the full sweep&lt;/td&gt;
&lt;td&gt;774,400&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;9,510&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;same n in every group&lt;/td&gt;
&lt;td&gt;14,400&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;same slice &lt;strong&gt;proportions&lt;/strong&gt;, totals free&lt;/td&gt;
&lt;td&gt;32,544&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;same &lt;strong&gt;total&lt;/strong&gt; n only, mixtures free&lt;/td&gt;
&lt;td&gt;72,816&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;502&lt;/strong&gt; (smallest at 10 obs)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Zero is not "rare". With the weights shared, pooled accuracy is a weighted mean using the &lt;em&gt;same&lt;/em&gt; weights for both models, and monotonicity does the rest. Matching only the proportions is already enough. &lt;strong&gt;Matching only the sample size is not.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every zero is printed next to the number of comparisons behind it, so an impossibility is distinguishable from an empty search — and the degenerate control (one group, where a reversal would be a contradiction) duly reports 0 over 1,936 comparisons.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does not say
&lt;/h2&gt;

&lt;p&gt;A third slice does &lt;strong&gt;not&lt;/strong&gt; make it cheaper: each extra group adds two more cells that must each be strictly won, and the smallest three-group reversal costs &lt;strong&gt;13&lt;/strong&gt; observations against nine. The sweep covers two groups at per-cell n ≤ 8, so the minimum is exact for two groups and a claim about nothing else. And the page never says which number is &lt;em&gt;right&lt;/em&gt; — macro and micro answer different questions, and choosing needs a deployment mixture a benchmark does not carry.&lt;/p&gt;

&lt;p&gt;The narrow claim is the defensible one: a pooled scalar published without the slice sizes that produced it is not a ranking anybody can check.&lt;/p&gt;

&lt;p&gt;27 in-page checks, 134 verifier asserts, 0 failures.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>datascience</category>
      <category>python</category>
      <category>programming</category>
    </item>
    <item>
      <title>HTTP Range Has Three Outcomes, Not Two: 32.03% Ignored, 54.95% Partial, 13.02% 416 Over 10,116 Calls</title>
      <dc:creator>Devanshu Biswas</dc:creator>
      <pubDate>Thu, 17 Sep 2026 12:32:22 +0000</pubDate>
      <link>https://dev.to/dev48v/http-range-has-three-outcomes-not-two-3203-ignored-5495-partial-1302-416-over-10116-calls-43l</link>
      <guid>https://dev.to/dev48v/http-range-has-three-outcomes-not-two-3203-ignored-5495-partial-1302-416-over-10116-calls-43l</guid>
      <description>&lt;p&gt;The &lt;code&gt;Range&lt;/code&gt; header is behind every resumed download, every seek in a &lt;code&gt;&amp;lt;video&amp;gt;&lt;/code&gt; tag and every byte-serving PDF viewer. It has &lt;strong&gt;three&lt;/strong&gt; outcomes, and servers routinely collapse them to two.&lt;/p&gt;

&lt;p&gt;Over an enumerated corpus of &lt;strong&gt;1,124 headers × 9 representation lengths = 10,116 decisions&lt;/strong&gt;: &lt;strong&gt;3,240 ignored → 200 (32.03%)&lt;/strong&gt;, &lt;strong&gt;5,559 → 206 (54.95%)&lt;/strong&gt;, &lt;strong&gt;1,317 → 416 (13.02%)&lt;/strong&gt;. A malformed field is &lt;strong&gt;2.46× more common&lt;/strong&gt; than an unsatisfiable one — and those are exactly the pair you conflate by treating "invalid" as "unsatisfiable".&lt;/p&gt;

&lt;p&gt;👉 &lt;strong&gt;Live, runs in your browser:&lt;/strong&gt; &lt;a href="https://dev48v.infy.uk/solve/day80-http-range-parser.html" rel="noopener noreferrer"&gt;https://dev48v.infy.uk/solve/day80-http-range-parser.html&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Invalid is not unsatisfiable
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;the field does not parse   -&amp;gt; 200, whole body      (it is ignored, not rejected)
it parses, nothing fits    -&amp;gt; 416 + bytes */len    (nothing else ever gives 416)
it parses, something fits  -&amp;gt; 206, the rest dropped
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;RFC 7233 is explicit that a Range a recipient cannot parse is &lt;strong&gt;dropped&lt;/strong&gt; — which means an ordinary 200 with the whole body. 416 exists for exactly one situation: the field parsed cleanly and every member starts past the end.&lt;/p&gt;

&lt;p&gt;So &lt;code&gt;bytes=9999-&lt;/code&gt; against a 1,000-byte file is a &lt;strong&gt;416&lt;/strong&gt;, and &lt;code&gt;bytes=abc&lt;/code&gt; is a &lt;strong&gt;200&lt;/strong&gt;. Getting those backwards is how a byte-serving endpoint answers 416 to clients that should have received the whole file.&lt;/p&gt;

&lt;h2&gt;
  
  
  A single backwards member poisons the whole header
&lt;/h2&gt;

&lt;p&gt;This is the one almost everybody gets wrong. &lt;code&gt;bytes=0-99,7-3&lt;/code&gt; is not "one good range and one bad one":&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// the rule everyone gets wrong: this poisons the FIELD, not the member&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;last&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;first&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;last&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;bad&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;first-byte-pos &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;first&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt; is greater than last-byte-pos &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;last&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;first &amp;gt; last&lt;/code&gt; makes the &lt;strong&gt;field&lt;/strong&gt; syntactically invalid, so the whole thing is ignored and the server sends 200 with the entire body. The most obviously "bad" range in the list lands in the 200 column, not the 416 one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Only first-byte-pos can push you off the end
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// only first-byte-pos can make a range unsatisfiable - last merely clamps&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;spec&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;first&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;len&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;last&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;spec&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;last&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nx"&gt;len&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;spec&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;last&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;len&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A suffix longer than the representation clamps to the whole representation. A &lt;code&gt;last&lt;/code&gt; past the end clamps to &lt;code&gt;len-1&lt;/code&gt;. Which is why, against the same 1,000-byte file, &lt;code&gt;bytes=900-9999&lt;/code&gt; is a &lt;strong&gt;206&lt;/strong&gt; and &lt;code&gt;bytes=9999-&lt;/code&gt; is a &lt;strong&gt;416&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Three spellings of "give me everything" — &lt;code&gt;bytes=0-&lt;/code&gt;, &lt;code&gt;bytes=-len&lt;/code&gt; and &lt;code&gt;bytes=0-(len-1)&lt;/code&gt; — resolve to one identical range at every non-zero length. At length 0 the identity is scoped out rather than fudged: there is no whole representation to name, and &lt;code&gt;0-(len-1)&lt;/code&gt; spells &lt;code&gt;0--1&lt;/code&gt;, which is not a legal spec at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  The middle case is the majority, and it shrinks twice
&lt;/h2&gt;

&lt;p&gt;When &lt;em&gt;some&lt;/em&gt; members fit and some do not, the unsatisfiable ones are dropped and what remains is served as a 206. Not a 416, not an error. &lt;strong&gt;2,152&lt;/strong&gt; decisions land in that case.&lt;/p&gt;

&lt;p&gt;Coalescing then shrinks the answer again: &lt;strong&gt;5,044 of the 5,559&lt;/strong&gt; 206s emit &lt;strong&gt;fewer parts than the client wrote&lt;/strong&gt;. Adjacency is the subtle half —&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// adjacent, not merely overlapping: 0-99 and 100-199 are ONE part&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;prev&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;first&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="nx"&gt;prev&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;last&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;prev&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;last&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;prev&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;last&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;last&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;0-99&lt;/code&gt; and &lt;code&gt;100-199&lt;/code&gt; share no byte, but they are one contiguous run, and a server that ships them as two multipart parts is paying boundary overhead for nothing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Six invariants over every one of the 10,116 decisions, 0 violations&lt;/strong&gt; — including the one that actually bites: the emitted byte set is rebuilt as a plain set union of the satisfiable members and compared in &lt;em&gt;both&lt;/em&gt; directions, rather than against itself. And 416 is asserted as an if-and-only-if, so neither direction can rot alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this covers, and what it does not
&lt;/h2&gt;

&lt;p&gt;It is a parser, &lt;strong&gt;not a complete server&lt;/strong&gt;. Out of scope, each changing real behaviour: &lt;code&gt;If-Range&lt;/code&gt; and the conditional dance that decides whether a range is honoured at all, &lt;code&gt;Accept-Ranges&lt;/code&gt;, the actual multipart body bytes with their boundaries and CRLFs, &lt;code&gt;Content-Encoding&lt;/code&gt; interacting with byte offsets, and ranges against a representation whose length is not yet known.&lt;/p&gt;

&lt;p&gt;Three strictness choices are documented rather than hidden, because a different implementer could defensibly go the other way: an empty list element such as &lt;code&gt;bytes=0-1,,4-5&lt;/code&gt; is treated as invalid where RFC 7230's list rule would let a recipient skip it; whitespace around the &lt;code&gt;=&lt;/code&gt; is rejected because the grammar has none there; and &lt;code&gt;bytes=-0&lt;/code&gt; is treated as unsatisfiable, since a zero-length suffix selects no bytes at all.&lt;/p&gt;

&lt;p&gt;37 in-page checks, 202 verifier asserts, 0 failures. Vanilla JS, one file, no build step.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>programming</category>
      <category>javascript</category>
      <category>http</category>
    </item>
    <item>
      <title>The Same 8,514-Token Prompt Bills 0.1113x or 1.2474x — an 11.21x Gap From One 14-Token Line Moving</title>
      <dc:creator>Devanshu Biswas</dc:creator>
      <pubDate>Thu, 17 Sep 2026 12:31:40 +0000</pubDate>
      <link>https://dev.to/dev48v/the-same-8514-token-prompt-bills-01113x-or-12474x-an-1121x-gap-from-one-14-token-line-moving-5el6</link>
      <guid>https://dev.to/dev48v/the-same-8514-token-prompt-bills-01113x-or-12474x-an-1121x-gap-from-one-14-token-line-moving-5el6</guid>
      <description>&lt;p&gt;Providers cache your prompt by &lt;strong&gt;exact token-prefix match&lt;/strong&gt;. What you get back is the longest common prefix between this request and the cached one, and &lt;em&gt;nothing after the first difference&lt;/em&gt; — however stable the rest of it is.&lt;/p&gt;

&lt;p&gt;So one volatile line decides the bill. Take a 14-token &lt;code&gt;Current date: …&lt;/code&gt;, which is &lt;strong&gt;0.16%&lt;/strong&gt; of an 8,514-token system prompt, and walk it through all seven insertion points. Placed &lt;strong&gt;last&lt;/strong&gt;, the steady bill is &lt;strong&gt;0.1113×&lt;/strong&gt; the uncached price. Placed &lt;strong&gt;first&lt;/strong&gt;, it is &lt;strong&gt;1.2474×&lt;/strong&gt; — above 1.000×, so caching there costs &lt;em&gt;more than not caching at all&lt;/em&gt;. That is &lt;strong&gt;11.21×&lt;/strong&gt; between two prompts with identical tokens in a different order.&lt;/p&gt;

&lt;p&gt;👉 &lt;strong&gt;Live, runs in your browser:&lt;/strong&gt; &lt;a href="https://dev48v.infy.uk/prompt/day80-prefix-caching.html" rel="noopener noreferrer"&gt;https://dev48v.infy.uk/prompt/day80-prefix-caching.html&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What is measured and what is declared
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Declared&lt;/strong&gt;, and stated above the first number: the three price multipliers — a cache write costs &lt;code&gt;1.25×&lt;/code&gt; a fresh input token, a cache read &lt;code&gt;0.10×&lt;/code&gt;, fresh input &lt;code&gt;1.00×&lt;/code&gt; — plus a $3.00/Mtok input rate for turning multiples into money. Change two constants and every figure moves with them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Measured&lt;/strong&gt;: the prefix matcher, the cumulative arithmetic, the break-even algebra and its agreement with a request-by-request simulation, and every count. There is no sampling anywhere, and &lt;strong&gt;nothing here simulates language&lt;/strong&gt; — the engine is an ordered list of segments with declared token lengths.&lt;/p&gt;

&lt;p&gt;The prompt is invented. The page says so.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hit is a cumulative sum, and that is the whole mechanism
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;volatile field at index&lt;/th&gt;
&lt;th&gt;sits before&lt;/th&gt;
&lt;th&gt;cache hit&lt;/th&gt;
&lt;th&gt;steady bill&lt;/th&gt;
&lt;th&gt;break-even N*&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;tool definitions&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.2474×&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;never&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;operating rules&lt;/td&gt;
&lt;td&gt;3,120&lt;/td&gt;
&lt;td&gt;0.8304×&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;output schema&lt;/td&gt;
&lt;td&gt;4,960&lt;/td&gt;
&lt;td&gt;0.5844×&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;house style guide&lt;/td&gt;
&lt;td&gt;5,520&lt;/td&gt;
&lt;td&gt;0.5096×&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;worked examples&lt;/td&gt;
&lt;td&gt;5,930&lt;/td&gt;
&lt;td&gt;0.4548×&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;escalation policy&lt;/td&gt;
&lt;td&gt;8,170&lt;/td&gt;
&lt;td&gt;0.1554×&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;— (it is last)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;8,500&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.1113×&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Put the field at index p and the hit is exactly the cumulative token count of segments 0…p−1. That column is measured by running the matcher; the cumulative sum is computed separately; they agree at every position.&lt;/p&gt;

&lt;p&gt;The assertion that carries the page is &lt;strong&gt;monotonicity&lt;/strong&gt;: moving the field one segment earlier can never increase the hit. Checked at every insertion point of this fixture &lt;em&gt;and&lt;/em&gt; of a 48-segment one — 56 positions, not a few sampled slots.&lt;/p&gt;

&lt;h2&gt;
  
  
  Break-even is closed-form, and its denominator is the interesting part
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cold = Ta*w + Tb*b                 first request writes everything
warm = H*r + (Ta-H)*w + Tb*b       every request after it

cold + (N-1)*warm  &amp;lt;=  N*(Ta+Tb)*b
  cold - warm      =  H*(w - r)                  the tail cancels
  (Ta+Tb)*b - warm =  H*(w - r) - Ta*(w - b)     and again

  N* = ceil( H(w-r) / ( H(w-r) - Ta(w-b) ) )     denominator &amp;lt;= 0  =&amp;gt;  never
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The user's actual turn cancels out of the ratio entirely, which is why it does not appear. And a negative denominator means &lt;strong&gt;no number of requests ever repays the write&lt;/strong&gt; — placed first, the page checks that out to 1,000,000 requests and it is still dearer.&lt;/p&gt;

&lt;p&gt;Placed last, the loan is repaid on request &lt;strong&gt;2&lt;/strong&gt;. The derived N* matches the request-by-request simulation at all 56 positions, and the check is two-sided: at N* the cached run is no dearer, and at N*−1 it really is — so N* is the smallest such N, not merely some N.&lt;/p&gt;

&lt;h2&gt;
  
  
  The threshold that does not depend on your prompt
&lt;/h2&gt;

&lt;p&gt;Set that denominator to zero and the fixture drops out. What is left is a property of the price model alone: a cached prefix under &lt;strong&gt;(w−b)/(w−r) = 21.74%&lt;/strong&gt; of the cacheable prompt &lt;strong&gt;never repays its own write, at any N&lt;/strong&gt;. The self-check pins the number printed in the prose to the one the formula derives.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two fixes that look equivalent; one does nothing
&lt;/h2&gt;

&lt;p&gt;The instinct on a cache that never hits is to make the volatile thing &lt;em&gt;smaller&lt;/em&gt; — shorter date format, an id instead of a name, drop the seconds. &lt;strong&gt;Hit length does not depend on the volatile field's token count at all.&lt;/strong&gt; At position 0 the hit is 0 whether the field is 1 token or 800, because the matcher stops at the first differing token and the field's own length is entirely on the far side of that stop.&lt;/p&gt;

&lt;p&gt;Same corollary for several volatile fields: the hit is decided by the &lt;strong&gt;earliest&lt;/strong&gt; one. Everything after it is already free to change, and tidying those is worth exactly nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this is not about
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Not breakpoints, TTL or minimum-prefix tiers.&lt;/strong&gt; All of those are held fixed here — they are a different page's subject. This one varies exactly one thing, where the volatile field sits, and prices the consequence. Also not modelled: cache lifetime and eviction, concurrency, and which prefix a provider actually keeps when several are live. And nothing here is about which instruction &lt;em&gt;wins&lt;/em&gt; when two conflict — segment order changes the bill; it is not an argument about precedence.&lt;/p&gt;

&lt;p&gt;19 in-page checks, 284 verifier asserts, 0 failures — plus 4 deliberate mutations of the engine, 4 caught.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>programming</category>
      <category>webdev</category>
    </item>
    <item>
      <title>A Roving Tabindex Widget Costs 1 Tab Stop at Every Size 1..50; the Default Markup Costs 1,275</title>
      <dc:creator>Devanshu Biswas</dc:creator>
      <pubDate>Thu, 17 Sep 2026 12:30:59 +0000</pubDate>
      <link>https://dev.to/dev48v/a-roving-tabindex-widget-costs-1-tab-stop-at-every-size-150-the-default-markup-costs-1275-3ckn</link>
      <guid>https://dev.to/dev48v/a-roving-tabindex-widget-costs-1-tab-stop-at-every-size-150-the-default-markup-costs-1275-3ckn</guid>
      <description>&lt;p&gt;Tab is for moving &lt;strong&gt;between&lt;/strong&gt; widgets. Arrows are for moving &lt;strong&gt;inside&lt;/strong&gt; one. A toolbar, listbox or menu should add exactly one stop to the page's tab order, not one per item.&lt;/p&gt;

&lt;p&gt;Swept over item counts 1…50, the roving widget contributes &lt;strong&gt;1 tab stop at every single size&lt;/strong&gt; — 50 summed. The default markup, which is what you get by writing no &lt;code&gt;tabindex&lt;/code&gt; at all, contributes 1…50 — &lt;strong&gt;1,275 summed&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;👉 &lt;strong&gt;Live, runs in your browser:&lt;/strong&gt; &lt;a href="https://dev48v.infy.uk/design/day80-roving-tabindex.html" rel="noopener noreferrer"&gt;https://dev48v.infy.uk/design/day80-roving-tabindex.html&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The whole pattern, and the whole correctness story
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// the whole pattern, in two lines&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;tabIndexOf&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;activeIndex&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;next&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;items&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;dir&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;wrap&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="cm"&gt;/* first ENABLED index that way */&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;// and the whole correctness story, in one&lt;/span&gt;
&lt;span class="nx"&gt;items&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;tabIndexOf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Not "usually one". If it is ever zero, Tab skips the widget and a keyboard user can never get in. If it is ever two, the widget has quietly grown a second tab stop and the bug is invisible to everyone using a mouse.&lt;/p&gt;

&lt;h2&gt;
  
  
  Enumerated, not sampled — and the reason is one design decision
&lt;/h2&gt;

&lt;p&gt;Over &lt;strong&gt;240 enumerated configurations&lt;/strong&gt; (item counts × disabled patterns × wrap), &lt;strong&gt;8,892 distinct reachable states&lt;/strong&gt; and &lt;strong&gt;79,968 key events&lt;/strong&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;assertion&lt;/th&gt;
&lt;th&gt;checks&lt;/th&gt;
&lt;th&gt;held&lt;/th&gt;
&lt;th&gt;broke&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;exactly one item at tabindex 0, after every key&lt;/td&gt;
&lt;td&gt;80,208&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;80,208&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;the active item is never a disabled item&lt;/td&gt;
&lt;td&gt;80,208&lt;/td&gt;
&lt;td&gt;80,208&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;the active index is in range, or exactly −1&lt;/td&gt;
&lt;td&gt;80,208&lt;/td&gt;
&lt;td&gt;80,208&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That walk is exhaustive rather than sampled, and it is exhaustive only because of one decision: &lt;strong&gt;buffer expiry arrives as a &lt;code&gt;Timeout&lt;/code&gt; key, not a &lt;code&gt;setTimeout&lt;/code&gt;&lt;/strong&gt;. A state machine that reads the clock cannot be walked — you can only poke at it. Making time an input is what turns "we tried a few" into a count.&lt;/p&gt;

&lt;p&gt;The edges are in by construction, not by luck: the first item disabled, the last item disabled, everything disabled but one. &lt;strong&gt;54&lt;/strong&gt; of the 240 configurations have nothing focusable at all, and those are scored against a &lt;em&gt;different&lt;/em&gt; rule — &lt;strong&gt;zero&lt;/strong&gt; tab stops, not one — because "exactly one" is the wrong assertion for a widget with nothing to focus.&lt;/p&gt;

&lt;h2&gt;
  
  
  The qualifier that usually gets deleted instead of the bug
&lt;/h2&gt;

&lt;p&gt;Three identities, &lt;strong&gt;702 checks, 0 broken&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;wrap=true&lt;/code&gt;: one ArrowRight per enabled item returns you home — the moves are a cyclic group (234 starting points).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;wrap=false&lt;/code&gt;: both ends absorb; arrowing past the last enabled item is a no-op (186 checks).&lt;/li&gt;
&lt;li&gt;ArrowRight then ArrowLeft returns to the same index &lt;strong&gt;whenever no end was crossed&lt;/strong&gt; (282 pairs).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last qualifier is load-bearing, and &lt;strong&gt;186 pairs are deliberately excluded&lt;/strong&gt; because an end really was crossed. Under &lt;code&gt;wrap=false&lt;/code&gt; the pair is genuinely not reversible at the ends: the first press was absorbed and the second one moves. Asserting reversibility without the qualifier fails on a &lt;em&gt;correct&lt;/em&gt; implementation — which is how the qualifier gets deleted instead of the bug.&lt;/p&gt;

&lt;h2&gt;
  
  
  Typeahead is a search, and it has its own edges
&lt;/h2&gt;

&lt;p&gt;Type &lt;code&gt;b&lt;/code&gt; and you land on the next item starting with b; type it again and you cycle; type &lt;code&gt;b&lt;/code&gt; then &lt;code&gt;o&lt;/code&gt; and you get "Bo…", not the next o. So the buffer decides the needle. Two consequences the page counts rather than hopes for: across &lt;strong&gt;1,302 typeahead searches&lt;/strong&gt;, the number that landed on a disabled item is &lt;strong&gt;0&lt;/strong&gt; — a search that puts you somewhere you cannot act is worse than one that finds nothing — and typeahead wraps even when the arrows do not, because clamping is a statement about a &lt;em&gt;direction&lt;/em&gt; and a search does not have one.&lt;/p&gt;

&lt;h2&gt;
  
  
  aria-activedescendant is the other answer, not the wrong one
&lt;/h2&gt;

&lt;p&gt;It &lt;strong&gt;ties&lt;/strong&gt; roving at 1 tab stop, at every size. The difference is where focus is. With roving the item is really focused, so &lt;code&gt;:focus&lt;/code&gt; and &lt;code&gt;:focus-visible&lt;/code&gt; style it for free and &lt;code&gt;document.activeElement&lt;/code&gt; is the thing the user is on. With activedescendant the item is never focused: you draw the ring yourself, and the element the browser thinks is focused and the element the user is on are two different elements, permanently. What that buys is items that never need to be focusable — which is why comboboxes reach for it, since one cannot move DOM focus into its list without taking it out of the input.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this models and what it does not
&lt;/h2&gt;

&lt;p&gt;The engine is the keyboard state machine and nothing else. &lt;strong&gt;Nothing on this page was measured in a browser&lt;/strong&gt; — the keystroke figures are counts of engine transitions, not timings. Out of scope, each of which is real work in a real widget: actually applying the tabindex and calling &lt;code&gt;.focus()&lt;/code&gt;; items added or removed while the widget is focused, where the active index goes stale; RTL, which changes what ArrowLeft means; 2-D grids, where Home/End have both a row meaning and a grid meaning; and Page Up / Page Down.&lt;/p&gt;

&lt;p&gt;23 in-page checks, 112 verifier asserts, 0 failures.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>javascript</category>
      <category>frontend</category>
      <category>ux</category>
    </item>
    <item>
      <title>The Mover Loses on Exactly 11 of 299 Piles, and Four Sensible Openings Win 0 of the Other 288</title>
      <dc:creator>Devanshu Biswas</dc:creator>
      <pubDate>Thu, 17 Sep 2026 12:30:16 +0000</pubDate>
      <link>https://dev.to/dev48v/the-mover-loses-on-exactly-11-of-299-piles-and-four-sensible-openings-win-0-of-the-other-288-5013</link>
      <guid>https://dev.to/dev48v/the-mover-loses-on-exactly-11-of-299-piles-and-four-sensible-openings-win-0-of-the-other-288-5013</guid>
      <description>&lt;p&gt;Fibonacci nim is one pile of counters. You open by taking anything from 1 to n−1 — never the whole pile — and after that nobody may take more than &lt;strong&gt;twice what the opponent just took&lt;/strong&gt;. Whoever lifts the last counter wins.&lt;/p&gt;

&lt;p&gt;The player to move loses &lt;strong&gt;exactly when the pile is a Fibonacci number&lt;/strong&gt;. Over the 299 piles from 2 to 300 that is 11 of them, and they are the obvious 11: 2, 3, 5, 8, 13, 21, 34, 55, 89, 144, 233.&lt;/p&gt;

&lt;p&gt;👉 &lt;strong&gt;Live, runs in your browser:&lt;/strong&gt; &lt;a href="https://dev48v.infy.uk/game/day80-fibonacci-nim.html" rel="noopener noreferrer"&gt;https://dev48v.infy.uk/game/day80-fibonacci-nim.html&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Two halves that are never allowed to talk to each other
&lt;/h2&gt;

&lt;p&gt;The solver knows the rules and nothing else. There is no Fibonacci number anywhere in it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// the definition, memoised - no theory in here at all&lt;/span&gt;
&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;solveState&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;rem&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;cap&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;cap&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cap&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;rem&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;k&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;k&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="nx"&gt;cap&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;k&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;k&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="nx"&gt;rem&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;              &lt;span class="c1"&gt;// took the last counter&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nf"&gt;solveState&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;rem&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;k&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;k&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="nf"&gt;firstPlayerWins&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;n&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;solveState&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;n&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;n&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;    &lt;span class="c1"&gt;// never the whole pile&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The other half is &lt;code&gt;zeckendorf(n)&lt;/code&gt; — every positive integer as a sum of non-consecutive Fibonacci numbers, by greedy subtraction. It knows arithmetic and has never heard of the game.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;0 disagreements&lt;/strong&gt; across all 299 piles. And the same rule holds one level down, over the whole state space rather than just the openings: for all &lt;strong&gt;45,150&lt;/strong&gt; reachable &lt;code&gt;(remaining, cap)&lt;/code&gt; pairs, &lt;em&gt;the mover wins if and only if the smallest Zeckendorf term of what remains fits under the cap&lt;/em&gt; — &lt;strong&gt;0 mismatches&lt;/strong&gt; against the solver.&lt;/p&gt;

&lt;h2&gt;
  
  
  The strategy is replayed, not asserted
&lt;/h2&gt;

&lt;p&gt;A theorem that says "a winning move exists" is cheap. This one is constructive, so the page plays it back through the solver instead of claiming it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;n&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;winnablePiles&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;smallestZeckTerm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;n&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="nf"&gt;assert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;z&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="nx"&gt;n&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;                        &lt;span class="c1"&gt;// legal opening&lt;/span&gt;
  &lt;span class="nf"&gt;assert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;solveState&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;n&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// opponent is now lost&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;288 winnable piles replayed, &lt;strong&gt;0 broken&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Every opening that feels right wins nothing
&lt;/h2&gt;

&lt;p&gt;This is the part worth the controls. Judged by the solver, over the same 288 winnable piles:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;opening rule&lt;/th&gt;
&lt;th&gt;piles it wins&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;smallest Zeckendorf term&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;288&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;largest Zeckendorf term&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;take the maximum allowed&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;take half the pile&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;largest Fibonacci that fits&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;always take 1&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;114&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Four plausible rules — including one built from the &lt;em&gt;same arithmetic read from the wrong end&lt;/em&gt; — win &lt;strong&gt;zero piles between them&lt;/strong&gt;. "Always take 1" is the interesting near-miss at 114, and 114 is not a coincidence: it is exactly the count of piles whose Zeckendorf representation contains a 1. The page counts those two numbers by separate routes and they agree.&lt;/p&gt;

&lt;p&gt;The correct take is small — mean 1075/288 ≈ &lt;strong&gt;3.73&lt;/strong&gt; counters, never more than &lt;strong&gt;2/7&lt;/strong&gt; of the pile (worst at n=7). It is still not bounded by a constant: at n=199 the right opening is &lt;strong&gt;55&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  "The" winning move is a lie the textbook statement tells
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;winning openings&lt;/th&gt;
&lt;th&gt;piles&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;none (the pile is lost)&lt;/td&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;exactly one&lt;/td&gt;
&lt;td&gt;70&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;two&lt;/td&gt;
&lt;td&gt;143&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;three&lt;/td&gt;
&lt;td&gt;75&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;218 of 288&lt;/strong&gt; winnable piles have more than one winning opening. At n=72 they are 1, 4 and 17; at n=300 they are 1, 12 and 67 — and the extras are usually not Fibonacci numbers at all. The Zeckendorf term is always the &lt;em&gt;smallest&lt;/em&gt; of them, which is a real checkable property; the others have no tidy description.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this does not claim
&lt;/h2&gt;

&lt;p&gt;299 piles agreeing is strong evidence for the theorem. It is &lt;strong&gt;not&lt;/strong&gt; the theorem, and the page says which range it verified. The memo is keyed on &lt;code&gt;rem * 4096 + c&lt;/code&gt;, which is sound to n=300 and would silently collide long before it became a general solver. And nothing here makes the strategy &lt;em&gt;playable&lt;/em&gt; — finding the smallest Zeckendorf term of 300 counters at a table is exactly as awkward as it sounds.&lt;/p&gt;

&lt;p&gt;29 in-page checks, 1,016 verifier asserts, 0 failures. Vanilla JS, one file, no build step.&lt;/p&gt;

</description>
      <category>algorithms</category>
      <category>math</category>
      <category>javascript</category>
      <category>programming</category>
    </item>
    <item>
      <title>A Search Scope Is a Read Scope: 751 Yes/No Questions Recover Every "Private" Field in My Personal API</title>
      <dc:creator>Devanshu Biswas</dc:creator>
      <pubDate>Tue, 08 Sep 2026 08:31:24 +0000</pubDate>
      <link>https://dev.to/dev48v/a-search-scope-is-a-read-scope-751-yesno-questions-recover-every-private-field-in-my-personal-4dbe</link>
      <guid>https://dev.to/dev48v/a-search-scope-is-a-read-scope-751-yesno-questions-recover-every-private-field-in-my-personal-4dbe</guid>
      <description>&lt;p&gt;You build one local API over your notes, contacts and calendar. Five endpoints, each behind its own scope: &lt;code&gt;read&lt;/code&gt;, &lt;code&gt;list&lt;/code&gt;, &lt;code&gt;search&lt;/code&gt;, &lt;code&gt;count&lt;/code&gt;, &lt;code&gt;aggregate&lt;/code&gt;. You hand an agent a token with &lt;code&gt;notes:search&lt;/code&gt; and nothing else. The body is withheld on every endpoint that returns a document.&lt;/p&gt;

&lt;p&gt;That token recovers &lt;strong&gt;all 14 private bodies, character for character&lt;/strong&gt;, in a mean of &lt;strong&gt;751 yes/no questions&lt;/strong&gt; each.&lt;/p&gt;

&lt;p&gt;👉 &lt;strong&gt;PUBLIC, MIT, &lt;code&gt;dependencies = []&lt;/code&gt;:&lt;/strong&gt; &lt;a href="https://github.com/dev48v/personal-api" rel="noopener noreferrer"&gt;https://github.com/dev48v/personal-api&lt;/a&gt;&lt;br&gt;
👉 &lt;strong&gt;Live measurement in your browser:&lt;/strong&gt; &lt;a href="https://dev48v.infy.uk/agentlab/vol4-04-personal-api.html" rel="noopener noreferrer"&gt;https://dev48v.infy.uk/agentlab/vol4-04-personal-api.html&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python &lt;span class="nt"&gt;-m&lt;/span&gt; personal_api            &lt;span class="c"&gt;# what each scope is actually worth&lt;/span&gt;
python &lt;span class="nt"&gt;-m&lt;/span&gt; personal_api &lt;span class="nt"&gt;--oracles&lt;/span&gt;  &lt;span class="c"&gt;# the five oracles, side by side&lt;/span&gt;
python &lt;span class="nt"&gt;-m&lt;/span&gt; personal_api &lt;span class="nt"&gt;--demo&lt;/span&gt;     &lt;span class="c"&gt;# watch one note come back out&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What each scope promises, and what it delivers
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;token scope&lt;/th&gt;
&lt;th&gt;can it read a body?&lt;/th&gt;
&lt;th&gt;bodies recovered anyway&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;read everything&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;14/14&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;titles only&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0/14&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;titles + search&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;no&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;14/14 exact&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;titles + count&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;3/14 exact, &lt;strong&gt;14 whole bodies&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;titles + aggregate (k floor)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0/14&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;search only&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;14/14 exact&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The second column is what the scope model promises. The third is the measurement.&lt;/p&gt;

&lt;p&gt;The only preset that holds is &lt;code&gt;titles only&lt;/code&gt; — and it holds because it has no query endpoint at all. A personal API you cannot ask questions of is not the thing anybody was trying to build.&lt;/p&gt;

&lt;h2&gt;
  
  
  The attack is not clever
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;known&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;
&lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;known&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;MAX&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;ch&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;ALPHABET&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;            &lt;span class="c1"&gt;# 37 symbols
&lt;/span&gt;        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;oracle&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;known&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;ch&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prefix&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;known&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;ch&lt;/span&gt;
            &lt;span class="k"&gt;break&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ask whether the body starts with &lt;code&gt;a&lt;/code&gt;, then &lt;code&gt;b&lt;/code&gt;, then &lt;code&gt;c&lt;/code&gt;. When one says yes, keep it and ask for the next character. At most 37 questions per character, no backtracking. It is what an agent with a token and a &lt;code&gt;while&lt;/code&gt; loop does on its own, and it is linear in the length.&lt;/p&gt;

&lt;p&gt;The corpus is &lt;strong&gt;invented&lt;/strong&gt; and written out in one short file before any number is computed, so you can check the counts rather than take them. &lt;code&gt;test_alphabet_covers_corpus&lt;/code&gt; holds the alphabet to the corpus — a body containing a symbol outside it would stall the attack and understate every figure here. (That bug was real: the first version omitted digits 1–9 and reported 1/14 instead of 14/14.)&lt;/p&gt;

&lt;h2&gt;
  
  
  The endpoints ranked by what they are worth to an attacker
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;oracle&lt;/th&gt;
&lt;th&gt;exact&lt;/th&gt;
&lt;th&gt;any whole body&lt;/th&gt;
&lt;th&gt;fragment recovered&lt;/th&gt;
&lt;th&gt;mean calls&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;search, prefix&lt;/strong&gt; (autocomplete)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;14/14&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;14/14&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100.0%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;751&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;search, contains (a search box)&lt;/td&gt;
&lt;td&gt;1/14&lt;/td&gt;
&lt;td&gt;1/14&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;53.5%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;401&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;count, prefix&lt;/td&gt;
&lt;td&gt;3/14&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;14/14&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;23.1%&lt;/td&gt;
&lt;td&gt;706&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;aggregate, k=2&lt;/td&gt;
&lt;td&gt;0/14&lt;/td&gt;
&lt;td&gt;0/14&lt;/td&gt;
&lt;td&gt;2.8%&lt;/td&gt;
&lt;td&gt;68&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;aggregate, k=3&lt;/td&gt;
&lt;td&gt;0/14&lt;/td&gt;
&lt;td&gt;0/14&lt;/td&gt;
&lt;td&gt;0.7%&lt;/td&gt;
&lt;td&gt;43&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;An autocomplete is worth far more than a search box.&lt;/strong&gt; A prefix oracle has an anchor — the start of the string — so extension never backtracks and recovery is exact. A contains oracle has to find an anchor first and can only extend rightward, so it returns a &lt;em&gt;fragment&lt;/em&gt;: 1 exact body in 14, but &lt;strong&gt;53.5% of each body on average&lt;/strong&gt;, contiguous. Half of a note about a bank account is not half a secret.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Counting is a read too, just aimed less well.&lt;/strong&gt; &lt;code&gt;count &amp;gt; 0&lt;/code&gt; is the same boolean, so the same loop runs — but the count is over the whole collection, so greedy extension follows whichever record branches first. It recovers the &lt;em&gt;targeted&lt;/em&gt; body 3 times in 14 and a complete body of &lt;em&gt;some&lt;/em&gt; record &lt;strong&gt;all 14 times&lt;/strong&gt;. The scope did not stop the leak. It randomised the victim.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one defence that works, and why that is a caveat
&lt;/h2&gt;

&lt;p&gt;With a k-anonymity floor the channel closes: &lt;strong&gt;0 exact, 0 whole bodies&lt;/strong&gt;, and 2.8% of a body at k=2 — about one character. That is a real defence and it is the recommendation this build actually supports.&lt;/p&gt;

&lt;p&gt;But it works &lt;strong&gt;because every body in this corpus is unique&lt;/strong&gt;, so any prefix long enough to be informative matches fewer than k records and gets suppressed. A floor does not protect a value that k people share. The suite says so out loud in &lt;code&gt;test_the_k_floor_works_because_the_secrets_are_unique&lt;/code&gt;, so the caveat cannot quietly come unstuck from the claim.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rate limiting is arithmetic
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;limit&lt;/th&gt;
&lt;th&gt;one body&lt;/th&gt;
&lt;th&gt;all 14&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;60 req/min&lt;/td&gt;
&lt;td&gt;12.5 min&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2 h 55 m&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;600 req/min&lt;/td&gt;
&lt;td&gt;1.3 min&lt;/td&gt;
&lt;td&gt;17.5 min&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6000 req/min&lt;/td&gt;
&lt;td&gt;7.5 s&lt;/td&gt;
&lt;td&gt;1.8 min&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;At the tightest limit anyone actually ships, the whole corpus comes out &lt;strong&gt;inside one night&lt;/strong&gt;. Rate limiting converts "instant" into "overnight", which is not a defence against something that runs overnight.&lt;/p&gt;

&lt;p&gt;The control that &lt;em&gt;would&lt;/em&gt; help is the one I did not model: &lt;strong&gt;968 near-identical prefix queries against one field is a shape&lt;/strong&gt;, and an audit log can alarm on a shape. That is the honest recommendation out of this build, and it is not a scope.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this is not
&lt;/h2&gt;

&lt;p&gt;Small invented corpus; the query counts scale with its strings. In-process API — no HTTP layer, no token expiry, no audit log. Nothing models partial field exposure, redaction, differential privacy with real noise, or an LLM in the loop deciding what to answer. Both attacks are the naive ones &lt;strong&gt;on purpose&lt;/strong&gt;: they establish a &lt;em&gt;floor&lt;/em&gt; on what is recoverable, not a ceiling.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;25 pytest, most of them controls&lt;/strong&gt; — that &lt;code&gt;read&lt;/code&gt; is genuinely denied, that &lt;code&gt;list&lt;/code&gt; genuinely withholds the body, that &lt;code&gt;titles only&lt;/code&gt; genuinely recovers nothing, and that the attack never once calls a denied endpoint (the denied-call counter is asserted to stay at zero, so it is a reconstruction and not a read wearing a hat).&lt;/p&gt;

&lt;p&gt;The claim is narrow. &lt;strong&gt;An endpoint that answers questions about a field is an endpoint that returns the field&lt;/strong&gt;, and a scope model that separates them is describing an intent rather than a boundary.&lt;/p&gt;

</description>
      <category>security</category>
      <category>python</category>
      <category>api</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Your Vector Store Rejects the Embedder Change That Barely Matters and Cannot See the One That Destroys Retrieval</title>
      <dc:creator>Devanshu Biswas</dc:creator>
      <pubDate>Tue, 08 Sep 2026 08:30:43 +0000</pubDate>
      <link>https://dev.to/dev48v/your-vector-store-rejects-the-embedder-change-that-barely-matters-and-cannot-see-the-one-that-3a6d</link>
      <guid>https://dev.to/dev48v/your-vector-store-rejects-the-embedder-change-that-barely-matters-and-cannot-see-the-one-that-3a6d</guid>
      <description>&lt;p&gt;The provider ships v2, or somebody changes the tokenizer. The index is still there, still returning five results, still with plausible scores.&lt;/p&gt;

&lt;p&gt;Change the &lt;strong&gt;width&lt;/strong&gt; and the vector is the wrong shape, so a typed index refuses it outright — and even with that check removed, precision@5 only falls from 0.9417 to 0.7750.&lt;/p&gt;

&lt;p&gt;Change the &lt;strong&gt;hash seed or the tokenizer&lt;/strong&gt; and the width is identical, nothing anywhere objects, and precision@5 falls to &lt;strong&gt;0.0333, 0.0917 and 0.1000&lt;/strong&gt; — against a chance baseline of &lt;strong&gt;0.0833&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;👉 &lt;strong&gt;Live, the whole grid computed in your browser:&lt;/strong&gt; &lt;a href="https://dev48v.infy.uk/ai/days/day79-embedding-drift.html" rel="noopener noreferrer"&gt;https://dev48v.infy.uk/ai/days/day79-embedding-drift.html&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Nothing here calls an embedding API
&lt;/h2&gt;

&lt;p&gt;The embedder is a &lt;strong&gt;hashed bag-of-tokens&lt;/strong&gt; implemented in full — FNV-1a with a stated seed, the signed hashing trick, optional idf, L2 normalisation, cosine. A "version" is a concrete configuration of that embedder, so a version change here is one people actually make.&lt;/p&gt;

&lt;p&gt;The corpus is synthetic and declared: 12 topics × 8 characteristic terms, 10 documents each. Because every document's topic is known by construction, retrieval is scored for &lt;strong&gt;correctness&lt;/strong&gt;, not just for agreement with an earlier run.&lt;/p&gt;

&lt;h2&gt;
  
  
  The detectable failures are the mild ones
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;version change&lt;/th&gt;
&lt;th&gt;does a typed store catch it?&lt;/th&gt;
&lt;th&gt;precision@5 after a full re-index&lt;/th&gt;
&lt;th&gt;precision@5 if only the query moves&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;same version (control)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;0.9417&lt;/td&gt;
&lt;td&gt;0.9417&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;width 128 → 256&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;yes — wrong width&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.9583&lt;/td&gt;
&lt;td&gt;0.7750&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;width 128 → 64&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;yes — wrong width&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.8583&lt;/td&gt;
&lt;td&gt;0.6333&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;hash seed 1 → 2&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;no — same width&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.9250&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.0917&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4-grams → 3-grams&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;no — same width&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.9083&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.0333&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4-grams → words&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;no — same width&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.8583&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.1000&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;idf weighting on&lt;/td&gt;
&lt;td&gt;no — same width&lt;/td&gt;
&lt;td&gt;0.9333&lt;/td&gt;
&lt;td&gt;0.9250&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Sort by damage and you get almost exactly the reverse of sorting by detectability. The 3-gram row lands at &lt;strong&gt;0.0333&lt;/strong&gt; — &lt;em&gt;below&lt;/em&gt; the 0.0833 you would get by returning five documents at random.&lt;/p&gt;

&lt;p&gt;The vectors carry no version tag. Cosine is happy to compare any two of the same length. There is no point in that path where a mismatch can be detected.&lt;/p&gt;

&lt;h2&gt;
  
  
  The half-migrated index is the shape this actually takes
&lt;/h2&gt;

&lt;p&gt;The migration script runs on new ingests and nobody backfills. Half the index is in one space, half in another:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;half-migrated&lt;/th&gt;
&lt;th&gt;single-version&lt;/th&gt;
&lt;th&gt;difference&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;new documents' share of top-5&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;21.67%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;52.50%&lt;/td&gt;
&lt;td&gt;−30.83 points&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;precision@5&lt;/td&gt;
&lt;td&gt;0.4667&lt;/td&gt;
&lt;td&gt;0.9417&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;−50.4%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;mean top-1 cosine&lt;/td&gt;
&lt;td&gt;0.3992&lt;/td&gt;
&lt;td&gt;0.5689&lt;/td&gt;
&lt;td&gt;−29.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;queries returning five documents&lt;/td&gt;
&lt;td&gt;24 of 24&lt;/td&gt;
&lt;td&gt;24 of 24&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Those documents are not ranked badly. They are &lt;strong&gt;unreachable&lt;/strong&gt; — and the system returns five results every time, drawn from the half that still lines up. No error, never an empty result set.&lt;/p&gt;

&lt;h2&gt;
  
  
  The monitor you would reach for is the wrong shape too
&lt;/h2&gt;

&lt;p&gt;Without labels, the one thing you can watch is the similarity score. On a &lt;em&gt;total&lt;/em&gt; mismatch it does move: mean top-1 cosine falls from 0.5689 to about 0.23, a ~60% drop. A threshold catches that.&lt;/p&gt;

&lt;p&gt;On the half-migrated index it falls &lt;strong&gt;29.8%&lt;/strong&gt; while precision falls &lt;strong&gt;50.4%&lt;/strong&gt;. Every query still has a well-matched document somewhere in the half that lines up, so the top score stays respectable while half the corpus has quietly left the building.&lt;/p&gt;

&lt;p&gt;A threshold tuned on the loud failure will not fire on the quiet one, and the quiet one is the one that happens.&lt;/p&gt;

&lt;h2&gt;
  
  
  Even the correct migration moves your answers
&lt;/h2&gt;

&lt;p&gt;Re-embedding everything is the right fix and it works — precision comes back to 0.858–0.958, in one case slightly &lt;em&gt;above&lt;/em&gt; the original. But top-5 overlap with the old index runs from &lt;strong&gt;0.933 down to 0.725&lt;/strong&gt;: between &lt;strong&gt;6.7% and 27.5%&lt;/strong&gt; of retrieved documents are different afterwards.&lt;/p&gt;

&lt;p&gt;Nothing is broken. It is just that any prompt, cache, eval set or human sign-off pinned to the old results is now pinned to results that no longer come back — on the day you did everything right.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix, and the honest scope
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Write the embedder version into the index and refuse to serve a query whose version does not match.&lt;/strong&gt; Every consequence above follows from the vectors not carrying one.&lt;/p&gt;

&lt;p&gt;Scope, because it changes how you read the table: the embedder is a hashed bag-of-tokens, &lt;strong&gt;not a neural model&lt;/strong&gt;. A hash-seed change produces two &lt;em&gt;unrelated&lt;/em&gt; spaces — that is the extreme case and an &lt;strong&gt;upper bound&lt;/strong&gt;, not a prediction for what v1→v2 of a trained model does. The tokenizer rows are the better analogue, and they are nearly as bad.&lt;/p&gt;

&lt;p&gt;Nothing here models &lt;strong&gt;semantic&lt;/strong&gt; drift: there is no meaning in this corpus beyond term overlap, so this says nothing about a v2 that is simply better at understanding a question. 120 documents, 24 queries, top-5, plain cosine, no chunking, no reranker, no hybrid keyword leg.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;28 in-page checks, 95 verifier assertions, 0 failures.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>rag</category>
      <category>embeddings</category>
      <category>programming</category>
    </item>
    <item>
      <title>MC Dropout With 10 Samples Is Further From the Truth Than Just Turning Dropout Off. Here Is the Break-Even.</title>
      <dc:creator>Devanshu Biswas</dc:creator>
      <pubDate>Tue, 08 Sep 2026 08:30:01 +0000</pubDate>
      <link>https://dev.to/dev48v/mc-dropout-with-10-samples-is-further-from-the-truth-than-just-turning-dropout-off-here-is-the-2ghp</link>
      <guid>https://dev.to/dev48v/mc-dropout-with-10-samples-is-further-from-the-truth-than-just-turning-dropout-off-here-is-the-2ghp</guid>
      <description>&lt;p&gt;Dropout is a training trick. At test time you turn it off and the whole network answers. MC dropout says leave it on, run k forward passes, average them, and read the spread as uncertainty.&lt;/p&gt;

&lt;p&gt;Both of those are approximations to the same object — &lt;strong&gt;the mean of the dropout ensemble&lt;/strong&gt;. So I computed that object exactly instead of approximating it: 12 dropout units means 2¹² = &lt;strong&gt;4096 masks, every one enumerated&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;👉 &lt;strong&gt;Live, the full enumeration in your browser:&lt;/strong&gt; &lt;a href="https://dev48v.infy.uk/dl/day79-dropout-at-inference.html" rel="noopener noreferrer"&gt;https://dev48v.infy.uk/dl/day79-dropout-at-inference.html&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Turning dropout off does not give you the ensemble mean
&lt;/h2&gt;

&lt;p&gt;With inverted dropout, a kept unit is divided by the keep probability, so the &lt;em&gt;expected&lt;/em&gt; post-dropout activation equals the plain activation. That is exactly why "turn dropout off" is supposed to give you the ensemble mean.&lt;/p&gt;

&lt;p&gt;It does not, and the control says why:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;network&lt;/th&gt;
&lt;th&gt;largest gap&lt;/th&gt;
&lt;th&gt;typical output&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;CONTROL: gap with a linear head&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;seed 7&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.5397&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.5537&lt;/td&gt;
&lt;td&gt;2.75×10⁻¹⁴&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;seed 101&lt;/td&gt;
&lt;td&gt;0.4106&lt;/td&gt;
&lt;td&gt;0.8880&lt;/td&gt;
&lt;td&gt;2.75×10⁻¹⁴&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;seed 2027&lt;/td&gt;
&lt;td&gt;0.2956&lt;/td&gt;
&lt;td&gt;0.4613&lt;/td&gt;
&lt;td&gt;2.71×10⁻¹⁴&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;seed 31337&lt;/td&gt;
&lt;td&gt;0.8336&lt;/td&gt;
&lt;td&gt;1.2534&lt;/td&gt;
&lt;td&gt;1.19×10⁻¹³&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;seed 90210&lt;/td&gt;
&lt;td&gt;0.2457&lt;/td&gt;
&lt;td&gt;0.3972&lt;/td&gt;
&lt;td&gt;2.51×10⁻¹⁴&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Make the layer &lt;em&gt;after&lt;/em&gt; dropout linear and the expectation passes straight through: &lt;code&gt;E[head(drop)] = head(E[drop])&lt;/code&gt;, exactly, and the gap collapses to floating-point zero in all five networks. Put the ReLU back and the gap is a third of the signal.&lt;/p&gt;

&lt;p&gt;A bug in the enumeration would not switch itself off when the ReLU does. That is what makes this a property of the architecture rather than a mistake in my code.&lt;/p&gt;

&lt;h2&gt;
  
  
  The break-even is 13.9 passes
&lt;/h2&gt;

&lt;p&gt;The k-sample mean is unbiased, so its RMS error is &lt;code&gt;σ/√k&lt;/code&gt;. The deterministic pass has no variance but a fixed bias. Set them equal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;k* = σ² / bias²   =   13.9
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the measured curve crosses exactly where it should:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;stochastic passes k&lt;/th&gt;
&lt;th&gt;measured RMS error&lt;/th&gt;
&lt;th&gt;vs dropout-off (0.2109)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;0.7859&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;273% further&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;0.5671&lt;/td&gt;
&lt;td&gt;169% further&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;0.3557&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;69% further&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;0.2492&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;18% further&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;30&lt;/td&gt;
&lt;td&gt;0.1391&lt;/td&gt;
&lt;td&gt;34% closer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;0.0792&lt;/td&gt;
&lt;td&gt;62% closer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1000&lt;/td&gt;
&lt;td&gt;0.0240&lt;/td&gt;
&lt;td&gt;89% closer&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Across five networks the break-even lands between &lt;strong&gt;5.8 and 29.8&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The point is not that 30 passes is expensive. It is that &lt;strong&gt;5 and 10 — the numbers that actually appear in code — sit on the wrong side of the line.&lt;/strong&gt; Averaging a handful of stochastic passes is a more elaborate way of being further from the answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The uncertainty moves. The prediction does not.
&lt;/h2&gt;

&lt;p&gt;The keep probability is a training hyperparameter. Sweeping it:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;keep&lt;/th&gt;
&lt;th&gt;RMS reported spread&lt;/th&gt;
&lt;th&gt;deterministic output&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;0.50&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.7865&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;identical&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;0.60&lt;/td&gt;
&lt;td&gt;0.6568&lt;/td&gt;
&lt;td&gt;identical&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;0.70&lt;/td&gt;
&lt;td&gt;0.5391&lt;/td&gt;
&lt;td&gt;identical&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;0.80&lt;/td&gt;
&lt;td&gt;0.4222&lt;/td&gt;
&lt;td&gt;identical&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;0.90&lt;/td&gt;
&lt;td&gt;0.2893&lt;/td&gt;
&lt;td&gt;identical&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;0.95&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.2021&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;identical&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A &lt;strong&gt;3.89×&lt;/strong&gt; change in the reported uncertainty, and the deterministic output is &lt;strong&gt;bit-identical at all 61 inputs&lt;/strong&gt; — because that pass never sees a mask at all.&lt;/p&gt;

&lt;p&gt;Two teams who picked 0.5 and 0.9 will report different uncertainties for the same prediction from the same network, and neither is more right. The answer and the confidence attached to it come from different places: one from the weights, the other from a number chosen during training.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(This one is exact too: at keep ≠ 0.5 the masks are not equally likely, so they are weighted by binomial mass rather than sampled. Still all 4096.)&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I expected to find, and did not
&lt;/h2&gt;

&lt;p&gt;Going in, my hypothesis was that this spread is just a restatement of how large the activations are — that it would track output magnitude and carry nothing else.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;network&lt;/th&gt;
&lt;th&gt;r( spread , |output| )&lt;/th&gt;
&lt;th&gt;r( spread , |gap| )&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;seed 7&lt;/td&gt;
&lt;td&gt;0.7304&lt;/td&gt;
&lt;td&gt;−0.0750&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;seed 101&lt;/td&gt;
&lt;td&gt;0.9550&lt;/td&gt;
&lt;td&gt;0.5912&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;seed 2027&lt;/td&gt;
&lt;td&gt;0.9573&lt;/td&gt;
&lt;td&gt;−0.3257&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;seed 31337&lt;/td&gt;
&lt;td&gt;0.7092&lt;/td&gt;
&lt;td&gt;0.8325&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;seed 90210&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;−0.1885&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.5506&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;In four of five it holds, at 0.71 to 0.96. In the fifth it is &lt;strong&gt;−0.19&lt;/strong&gt;. One counterexample in five is not a rounding error, so &lt;strong&gt;that claim is not established and I am not making it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The second column is the cleaner negative. The spread does not tell you where the deterministic approximation is worst: &lt;strong&gt;−0.33 to +0.83&lt;/strong&gt; across five networks — not merely weak, but with an unstable &lt;em&gt;sign&lt;/em&gt;. Whatever that number is measuring, it is not the size of the error it is standing next to.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this does not cover
&lt;/h2&gt;

&lt;p&gt;These networks are &lt;strong&gt;not trained&lt;/strong&gt; — the weights come deterministically from a stated seed, so nothing here says what dropout does to a fitted model. One dropout layer, scalar in and out. No label noise, so the spread is purely the ensemble's own disagreement with no aleatoric term to separate it from. And none of this evaluates MC dropout as a &lt;em&gt;calibration&lt;/em&gt; method: no labels, no reliability diagram.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;25 in-page checks, 66 verifier assertions, 0 failures.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>deeplearning</category>
      <category>machinelearning</category>
      <category>ai</category>
      <category>python</category>
    </item>
    <item>
      <title>Your 90%-Precision Threshold Delivers 63.5% After a Prevalence Shift — With TPR and FPR Bit-Identical</title>
      <dc:creator>Devanshu Biswas</dc:creator>
      <pubDate>Tue, 08 Sep 2026 08:29:20 +0000</pubDate>
      <link>https://dev.to/dev48v/your-90-precision-threshold-delivers-635-after-a-prevalence-shift-with-tpr-and-fpr-56h9</link>
      <guid>https://dev.to/dev48v/your-90-precision-threshold-delivers-635-after-a-prevalence-shift-with-tpr-and-fpr-56h9</guid>
      <description>&lt;p&gt;You tuned a decision threshold on validation to hit "precision at least 90%". You shipped it. The population's prevalence drifted from 10% to 2%.&lt;/p&gt;

&lt;p&gt;The model did not change. The scores did not change. &lt;strong&gt;TPR and FPR at that threshold are bit-identical&lt;/strong&gt; — the same reduced fractions, &lt;code&gt;22663/29409&lt;/code&gt; and &lt;code&gt;21043/499953&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Precision is now &lt;strong&gt;63.46%&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;👉 &lt;strong&gt;Live, all 201 thresholds enumerated exactly in your browser:&lt;/strong&gt; &lt;a href="https://dev48v.infy.uk/ml/day79-threshold-transfer.html" rel="noopener noreferrer"&gt;https://dev48v.infy.uk/ml/day79-threshold-transfer.html&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Nothing is sampled and no rate is a float
&lt;/h2&gt;

&lt;p&gt;The score axis is &lt;strong&gt;201 bins of declared integer weights&lt;/strong&gt;, all 201 thresholds are enumerated, and every rate is an exact rational compared by &lt;strong&gt;BigInt cross-multiplication&lt;/strong&gt; — so an argmax is never decided by rounding, and there are no confidence intervals anywhere because there is nothing to be uncertain about.&lt;/p&gt;

&lt;h2&gt;
  
  
  The identity the whole thing rests on
&lt;/h2&gt;

&lt;p&gt;Over a virtual population of &lt;code&gt;q·WP·WN&lt;/code&gt;, every confusion cell is an exact integer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TP = p * Sp * WN         FP = (q - p) * Sn * WP
FN = p * (WP - Sp) * WN  TN = (q - p) * (WN - Sn) * WP
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TPR = TP / (TP + FN)
    = (p·Sp·WN) / (p·Sp·WN + p·(WP−Sp)·WN)
    = Sp / WP
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;&lt;code&gt;p&lt;/code&gt; and &lt;code&gt;q&lt;/code&gt; cancel.&lt;/strong&gt; TPR does not know what the prevalence is. Neither does FPR. Both are conditioned on the class.&lt;/p&gt;

&lt;p&gt;Precision is not. Over six prevalences from 0.1% to 90% at one fixed threshold:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;prevalence&lt;/th&gt;
&lt;th&gt;TPR (reduced)&lt;/th&gt;
&lt;th&gt;FPR (reduced)&lt;/th&gt;
&lt;th&gt;precision&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;0.1%&lt;/td&gt;
&lt;td&gt;&lt;code&gt;22663/29409&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;21043/499953&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.0180&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2%&lt;/td&gt;
&lt;td&gt;&lt;code&gt;22663/29409&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;21043/499953&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.2720&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10%&lt;/td&gt;
&lt;td&gt;&lt;code&gt;22663/29409&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;21043/499953&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.6704&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;30%&lt;/td&gt;
&lt;td&gt;&lt;code&gt;22663/29409&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;21043/499953&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.8870&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;50%&lt;/td&gt;
&lt;td&gt;&lt;code&gt;22663/29409&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;21043/499953&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.9482&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;90%&lt;/td&gt;
&lt;td&gt;&lt;code&gt;22663/29409&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;21043/499953&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.9940&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two identical columns and a fifty-fold swing. That is an algebraic identity on one side and the entire problem on the other.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two of five tuning rules pick a portable threshold
&lt;/h2&gt;

&lt;p&gt;Re-running each rule on the shifted population and reading off where it lands:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;tuning rule&lt;/th&gt;
&lt;th&gt;source&lt;/th&gt;
&lt;th&gt;after 10%→2%&lt;/th&gt;
&lt;th&gt;portable?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;score ≥ 0.50&lt;/td&gt;
&lt;td&gt;bin 100&lt;/td&gt;
&lt;td&gt;bin 100&lt;/td&gt;
&lt;td&gt;n/a — it cannot move&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;maximise F1&lt;/td&gt;
&lt;td&gt;bin 127&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;bin 146&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;maximise TPR − FPR (Youden)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;bin 100&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;bin 100&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;yes&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;precision ≥ 0.90&lt;/td&gt;
&lt;td&gt;bin 148&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;bin 169&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;recall ≥ 0.90&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;bin 98&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;bin 98&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;yes&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The two that survive are &lt;strong&gt;exactly the two whose definitions mention only class-conditional rates.&lt;/strong&gt; That is not an empirical coincidence to be spot-checked; it follows from the identity above. Youden picks the same bin at five different prevalences; max-F1 picks a different bin at nearly every one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scope:&lt;/strong&gt; recall≥0.90 is also unmoved by the score shift I used here, but only because that shift moves the &lt;em&gt;negatives&lt;/em&gt;. Recall does not look at negatives. Move the positives and it moves like everything else.&lt;/p&gt;

&lt;h2&gt;
  
  
  Re-tuning meets the target and hollows out the model
&lt;/h2&gt;

&lt;p&gt;Re-running precision≥0.90 on each world does restore the number:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;world&lt;/th&gt;
&lt;th&gt;re-tuned threshold&lt;/th&gt;
&lt;th&gt;precision&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;recall&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;source&lt;/td&gt;
&lt;td&gt;bin 148&lt;/td&gt;
&lt;td&gt;0.9044&lt;/td&gt;
&lt;td&gt;0.4955&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;prevalence shift&lt;/td&gt;
&lt;td&gt;bin 169&lt;/td&gt;
&lt;td&gt;0.9088&lt;/td&gt;
&lt;td&gt;0.2412&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;score shift&lt;/td&gt;
&lt;td&gt;bin 166&lt;/td&gt;
&lt;td&gt;0.9041&lt;/td&gt;
&lt;td&gt;0.2773&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;both&lt;/td&gt;
&lt;td&gt;bin 193&lt;/td&gt;
&lt;td&gt;0.9022&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.0190&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;In the last world the threshold that satisfies the contract fires on &lt;strong&gt;1.9% of the positives&lt;/strong&gt;. A precision target is a promise about the answers you give, and it can always be kept by giving fewer. A monitor watching only precision reports a healthy system right up to the point where the system is not answering.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part with no good answer
&lt;/h2&gt;

&lt;p&gt;In production you usually have no labels. The one thing you can watch is how often the model says yes.&lt;/p&gt;

&lt;p&gt;I constructed two worlds with the &lt;strong&gt;exactly equal&lt;/strong&gt; positive rate — not close, the identical rational &lt;code&gt;780643/49995300&lt;/code&gt; ≈ 1.5614%. The prevalence for the second was &lt;em&gt;solved&lt;/em&gt; from &lt;code&gt;rate = π·TPR + (1−π)·FPR&lt;/code&gt;, which has an exact rational answer.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;world&lt;/th&gt;
&lt;th&gt;positive rate&lt;/th&gt;
&lt;th&gt;precision&lt;/th&gt;
&lt;th&gt;recall&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;prevalence shift, 10% → 2%&lt;/td&gt;
&lt;td&gt;&lt;code&gt;780643/49995300&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.6346&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.4955&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;score shift, prevalence 0.0702%&lt;/td&gt;
&lt;td&gt;&lt;code&gt;780643/49995300&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.0223&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.4955&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Same firing rate. Same recall. A &lt;strong&gt;28.5×&lt;/strong&gt; difference in precision. Every dashboard that watches the positive rate shows a flat line across both.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this does not cover
&lt;/h2&gt;

&lt;p&gt;One score axis in 201 bins, one shift direction (negatives up 12 bins, positives never move), one pair of class-conditional shapes, five tuning rules. &lt;strong&gt;There is no cost matrix&lt;/strong&gt;, so "which threshold is right" is never answered here — only which one is &lt;em&gt;stable&lt;/em&gt;. No calibration model and no recalibration method: the point is what breaks, not how to fix it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;21 in-page checks, 72 verifier assertions, 0 failures.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>datascience</category>
      <category>python</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
