<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mahiro Hirakawa</title>
    <description>The latest articles on DEV Community by Mahiro Hirakawa (@mahirhir).</description>
    <link>https://dev.to/mahirhir</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3937147%2F66fce836-aa25-43f0-bb5f-632fc17ebf44.jpeg</url>
      <title>DEV Community: Mahiro Hirakawa</title>
      <link>https://dev.to/mahirhir</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/mahirhir"/>
    <language>en</language>
    <item>
      <title>A six-line edit made fifteen units stale. The freshness check was working correctly.</title>
      <dc:creator>Mahiro Hirakawa</dc:creator>
      <pubDate>Mon, 14 Sep 2026 19:32:21 +0000</pubDate>
      <link>https://dev.to/mahirhir/a-six-line-edit-made-fifteen-units-stale-the-freshness-check-was-working-correctly-1en6</link>
      <guid>https://dev.to/mahirhir/a-six-line-edit-made-fifteen-units-stale-the-freshness-check-was-working-correctly-1en6</guid>
      <description>&lt;p&gt;Every unit of this project carries a record of its last verification, and each record pins the configuration it was checked against by hashing it. If the configuration changes, the pin no longer reproduces and the record goes stale, which is exactly what you want: a verification result that quietly outlives its inputs is worse than no result.&lt;/p&gt;

&lt;p&gt;One afternoon I added a six-line block to that configuration file. It described a new component and touched nothing else.&lt;/p&gt;

&lt;p&gt;Fifteen units went stale. Every one of them owed a full re-run of its checks and a fresh record appended afterwards, which is most of an hour of machine time and a boundary ceremony, for six lines that concerned one of them.&lt;/p&gt;

&lt;p&gt;Nothing was broken. The check did precisely what it promised.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pin was over the whole file
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;pin = sha256(entire configuration file)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Any byte anywhere changes that hash. So the pin does not mean "the configuration this unit depends on is unchanged". It means "nobody has edited this file at all", which is a far stronger statement and one that becomes false constantly in a file every component writes to.&lt;/p&gt;

&lt;p&gt;The gap between what a pin means and what everyone reads it to mean is the whole bug, and it is invisible while the file is small.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;what the pin says&lt;/th&gt;
&lt;th&gt;what everyone reads it as&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;this file is byte-identical&lt;/td&gt;
&lt;td&gt;my inputs are unchanged&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;any edit invalidates me&lt;/td&gt;
&lt;td&gt;edits to my section invalidate me&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  The failure mode is people, not machines
&lt;/h2&gt;

&lt;p&gt;An hour of machine time is affordable. What is not affordable is what it does to behaviour.&lt;/p&gt;

&lt;p&gt;When correcting a six-line block costs fifteen re-runs, three things start happening. I watched all three:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;edits get batched until there are enough to justify the ceremony, so the configuration is wrong for longer&lt;/li&gt;
&lt;li&gt;the staleness line gets read past, because it is almost always red for reasons unrelated to what you are doing&lt;/li&gt;
&lt;li&gt;someone proposes re-pinning without re-running, which is the point where the check stops meaning anything&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A check whose cost is out of proportion to the change gets routed around. It does not get fixed, because it is not broken.&lt;/p&gt;

&lt;h2&gt;
  
  
  The repair is a declared partition
&lt;/h2&gt;

&lt;p&gt;The pin became the hash of the unit's own block plus the sections genuinely shared between units, rather than the whole file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;pin = sha256(block for this unit) + sha256(shared sections)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two properties matter more than the saving. The partition is &lt;strong&gt;declared&lt;/strong&gt;, as rows in the file itself rather than inferred by a parser guessing at section boundaries, so it is readable and arguable. And it is &lt;strong&gt;recomputable&lt;/strong&gt; by hand, since the recipe is printed next to the hash and anyone can rerun it with &lt;code&gt;sha256sum&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Now an edit to one unit's block stales one unit. An edit to a shared section stales everything, and that is correct, because everything does depend on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would check earlier next time
&lt;/h2&gt;

&lt;p&gt;There is a cheap diagnostic I did not run for months. Take the last N edits to the file and ask, for each, how many units the pin invalidated and how many actually depended on the change.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;edits sampled          20
units invalidated      15 each time
units actually affected 1, 1, 1, 2, 1, 15, 1, ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One row genuinely needed all fifteen. The rest were a partition problem visible in five minutes of arithmetic over the git log, and the ratio is the argument for fixing it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things I keep
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Hash granularity is an interface, not an implementation detail.&lt;/strong&gt; Choosing the whole file is choosing "any edit invalidates every dependant", and that choice deserves the same scrutiny as any other contract, especially in a file that many components share.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Correct but disproportionate" is a defect class of its own.&lt;/strong&gt; It never shows up as a failing check, so nothing surfaces it. It shows up as people batching their edits, skimming past a red line, and eventually proposing to weaken the check, and by then the conversation is about the check rather than about its granularity.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>devops</category>
      <category>programming</category>
      <category>testing</category>
    </item>
    <item>
      <title>The OOM killer stopped my acceptance check four times in two days. It was rebuilding the project to read a number.</title>
      <dc:creator>Mahiro Hirakawa</dc:creator>
      <pubDate>Mon, 14 Sep 2026 18:58:29 +0000</pubDate>
      <link>https://dev.to/mahirhir/the-oom-killer-stopped-my-acceptance-check-four-times-in-two-days-it-was-rebuilding-the-project-to-42f6</link>
      <guid>https://dev.to/mahirhir/the-oom-killer-stopped-my-acceptance-check-four-times-in-two-days-it-was-rebuilding-the-project-to-42f6</guid>
      <description>&lt;p&gt;There is a step at the end of each work session that decides whether each unit of the project is accepted. It gathers, for every unit, whether its test suite passed, whether its checks are current, whether its documentation was regenerated, and prints one line per unit.&lt;/p&gt;

&lt;p&gt;It kept dying.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;acceptance read: killed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Four times in two days. No error, no stack, just gone. The machine's memory limit, with nothing else running. On the fourth kill I stopped restarting it and read what it does.&lt;/p&gt;

&lt;p&gt;It was compiling the project. To find out whether the tests passed, it ran them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The word "read" was doing a lot of work
&lt;/h2&gt;

&lt;p&gt;The step is called a read, and I had thought of it as one for months. It sounded like gathering. What it actually did, for each of fifteen units, was invoke the build and wait.&lt;/p&gt;

&lt;p&gt;That is not a read. It is a second full build whose output nobody looks at except to extract a boolean.&lt;/p&gt;

&lt;p&gt;The memory limit was not the problem. It was the only thing telling me anything. A build is the most expensive operation in the repository, and the acceptance step ran fifteen of them in sequence, on a machine that can comfortably run one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why re-running looks like the rigorous choice
&lt;/h2&gt;

&lt;p&gt;Here is the argument I would have made for the old design, and it is not stupid:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A recorded result is a claim about the past. The tree may have changed since. Running it now is the only way to know it holds now.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is correct in general and wrong here, because of what the recording already is. Every control run appends a block to a permanent record, and that block carries the commit it ran at plus its verbatim output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;control block   unit=projection  commit=26ef4cc  OK_ALL controls=13  run=2026-09-13
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So the record is not someone's summary. It is the control's own print, pinned to a commit. Asking "did this pass at this commit" is answerable from it exactly, and re-running the build answers the same question by doing the same work again.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;step&lt;/th&gt;
&lt;th&gt;question&lt;/th&gt;
&lt;th&gt;how it should answer&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;the control&lt;/td&gt;
&lt;td&gt;does this hold right now&lt;/td&gt;
&lt;td&gt;run it, append the print&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;the acceptance read&lt;/td&gt;
&lt;td&gt;what do the controls say at this commit&lt;/td&gt;
&lt;td&gt;read the record&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two steps, two questions. The acceptance step had been answering the first question while being named for the second, so it duplicated the control's work and added nothing except a way to run out of memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  The staleness question, which is the real one
&lt;/h2&gt;

&lt;p&gt;"What if the record is from an older commit?" is the objection that matters, and it is answerable without building anything:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ACCEPT_GATE  current=S  next=face  recorded=9/15  stale=0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The read compares each block's commit against the commit being accepted. If they differ, that unit is stale and says so. Staleness becomes a printed number rather than something the design tries to make impossible by brute force.&lt;/p&gt;

&lt;p&gt;That is the useful move: &lt;strong&gt;the cost of re-deriving was buying a property that a comparison provides for free.&lt;/strong&gt; Freshness is a question about two identifiers. It does not need a compiler.&lt;/p&gt;

&lt;h2&gt;
  
  
  A general form I keep coming back to
&lt;/h2&gt;

&lt;p&gt;Any step whose name is a verb of observation, read, check, report, gather, audit, that takes minutes and consumes gigabytes, is probably not observing. Something inside it is re-deriving a fact that already exists somewhere.&lt;/p&gt;

&lt;p&gt;The tell is resource usage that does not match the verb. An observation's cost should scale with how much there is to observe, not with how expensive the thing observed was to produce.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things I keep
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A read that builds is not a read.&lt;/strong&gt; Separating "produce the fact" from "collect the facts" is not tidiness; it is what makes the collecting step cheap enough to run at every boundary instead of once a day when someone is brave.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When a step keeps being killed, read it before restarting it.&lt;/strong&gt; I restarted three times and reduced parallelism twice, both of which are treatments for a resource problem. The fourth kill was the first time I asked what the step was actually doing, and the answer took two minutes to find.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>architecture</category>
      <category>testing</category>
      <category>programming</category>
    </item>
    <item>
      <title>My coverage check did arithmetic instead of looking. All 23 pointers were wrong.</title>
      <dc:creator>Mahiro Hirakawa</dc:creator>
      <pubDate>Mon, 14 Sep 2026 18:03:19 +0000</pubDate>
      <link>https://dev.to/mahirhir/my-coverage-check-did-arithmetic-instead-of-looking-all-23-pointers-were-wrong-502e</link>
      <guid>https://dev.to/mahirhir/my-coverage-check-did-arithmetic-instead-of-looking-all-23-pointers-were-wrong-502e</guid>
      <description>&lt;p&gt;A table in one document points into another. Each of its 24 rows carries a source pointer saying where the thing it describes actually lives, so a reader can go and check.&lt;/p&gt;

&lt;p&gt;The coverage check for that table reported complete, always. Here is what it did:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;rows in the table        24
rows carrying a pointer  23
rows that are a vessel    1
24 - 23 - 1 = 0 uncovered  -&amp;gt;  complete
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is arithmetic over counts. It never opens the file the pointers point into. A pointer can name a line that does not exist, or the wrong line, or a line in a file that was deleted, and the subtraction still comes out to zero.&lt;/p&gt;

&lt;p&gt;When something finally resolved them, all 23 were wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrong by one or two lines, which is the informative part
&lt;/h2&gt;

&lt;p&gt;Not wrong by hundreds. Every pointer was off by one or two lines against the target.&lt;/p&gt;

&lt;p&gt;That number tells you which failure this is. A pointer that drifts as a file grows ends up off by however many lines were inserted above it, which after a year is a large and varied number. Off by one or two, uniformly, across every row, is not drift. It is an off-by-one in whatever produced them, present from the first commit.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;symptom&lt;/th&gt;
&lt;th&gt;drift&lt;/th&gt;
&lt;th&gt;never worked&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;offsets&lt;/td&gt;
&lt;td&gt;large, varied&lt;/td&gt;
&lt;td&gt;small, uniform&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;when introduced&lt;/td&gt;
&lt;td&gt;gradually&lt;/td&gt;
&lt;td&gt;at creation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;repair&lt;/td&gt;
&lt;td&gt;re-point and add a freshness check&lt;/td&gt;
&lt;td&gt;fix the generator, then re-point&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;how it reads in a log&lt;/td&gt;
&lt;td&gt;"these have gone stale"&lt;/td&gt;
&lt;td&gt;"these have gone stale"&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The two look identical in the report. They want different repairs, and only the distribution of the offsets tells you which one you have.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same defect in four tools in one week
&lt;/h2&gt;

&lt;p&gt;While fixing it I found that four separate readers were joining two tables the same way: by row order.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="err"&gt;#&lt;/span&gt; &lt;span class="nx"&gt;join&lt;/span&gt; &lt;span class="nx"&gt;by&lt;/span&gt; &lt;span class="nx"&gt;position&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;row&lt;/span&gt; &lt;span class="mi"&gt;7&lt;/span&gt; &lt;span class="nx"&gt;here&lt;/span&gt; &lt;span class="nx"&gt;describes&lt;/span&gt; &lt;span class="nx"&gt;row&lt;/span&gt; &lt;span class="mi"&gt;7&lt;/span&gt; &lt;span class="nx"&gt;there&lt;/span&gt;
&lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nx"&gt;left&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nf"&gt;check&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;left&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="nx"&gt;right&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Position is a key that every insertion invalidates. Add a row to one side and every pair after it is silently wrong, which is exactly what had happened. Nothing errors, because both sides still have rows at those indices.&lt;/p&gt;

&lt;p&gt;All four became keyed joins, matching on an identifier that travels with the row:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="err"&gt;#&lt;/span&gt; &lt;span class="nx"&gt;join&lt;/span&gt; &lt;span class="nx"&gt;by&lt;/span&gt; &lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;an&lt;/span&gt; &lt;span class="nx"&gt;unmatched&lt;/span&gt; &lt;span class="nx"&gt;row&lt;/span&gt; &lt;span class="nx"&gt;on&lt;/span&gt; &lt;span class="nx"&gt;either&lt;/span&gt; &lt;span class="nx"&gt;side&lt;/span&gt; &lt;span class="nx"&gt;is&lt;/span&gt; &lt;span class="nx"&gt;a&lt;/span&gt; &lt;span class="nx"&gt;named&lt;/span&gt; &lt;span class="nx"&gt;finding&lt;/span&gt;
&lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;row&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;left&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nf"&gt;check&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;row&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;right&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;byId&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;row&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="nx"&gt;MISSING&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The difference that matters is not correctness in the happy case. It is that a keyed join can report &lt;code&gt;MISSING&lt;/code&gt;, and a positional join has no way to express it. A positional join cannot fail; it can only be wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  What replaced the subtraction
&lt;/h2&gt;

&lt;p&gt;Each row now carries the pointer plus a verbatim quote of what it points at. Coverage is the count of rows whose quote still matches the bytes at that location, which is a question about the world rather than about the table's own shape.&lt;/p&gt;

&lt;p&gt;Retired rows moved out of the denominator entirely, as pointer lines that no counter reads.&lt;/p&gt;

&lt;p&gt;The satisfying part: after the repair, no count moved. The table had said complete and, once every pointer was actually resolved, it was complete. Every one of those 23 pointers had been wrong, and the number the check reported was right. It was right for no reason.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things I keep
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A coverage number computed from counts is a statement about your table, not about your code.&lt;/strong&gt; If the check never opens the thing being pointed at, it cannot go red for the reason you built it, and the correct total it prints is a coincidence you should not spend.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Read the distribution of the errors, not just how many.&lt;/strong&gt; "All 23 are stale" and "all 23 are off by one" are the same count and different bugs. One asks for a freshness check. The other asks you to fix the thing that wrote them, or you will re-point 23 rows and generate 23 more.&lt;/p&gt;

</description>
      <category>testing</category>
      <category>architecture</category>
      <category>programming</category>
      <category>devops</category>
    </item>
    <item>
      <title>Every correction I made to my spec added a requirement. The retired rows were still being counted.</title>
      <dc:creator>Mahiro Hirakawa</dc:creator>
      <pubDate>Mon, 14 Sep 2026 17:32:23 +0000</pubDate>
      <link>https://dev.to/mahirhir/every-correction-i-made-to-my-spec-added-a-requirement-the-retired-rows-were-still-being-counted-111g</link>
      <guid>https://dev.to/mahirhir/every-correction-i-made-to-my-spec-added-a-requirement-the-retired-rows-were-still-being-counted-111g</guid>
      <description>&lt;p&gt;The specification here is append-only. You never edit a row. To change one you retire the old row, which stays in the file under a retired heading, and stand a new row beside it. The old text remains readable, so a reader six months later can see what the rule used to say and why it moved.&lt;/p&gt;

&lt;p&gt;Two tools read that file to build the list of requirements the code must satisfy. Neither of them knew what a retired heading was.&lt;/p&gt;

&lt;p&gt;So a retire-and-stand produced two live requirements where there had been one. Correcting a row did not fix the specification. It grew it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The number that should have caught it moved the wrong way
&lt;/h2&gt;

&lt;p&gt;The acceptance check reports how many declared requirements are backed by a test:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;dod_rows=263  printed=191/263
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Retire one row and stand its replacement, and both sides move:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;dod_rows=264  printed=191/264
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The denominator grows. The numerator does not, because the retired row names a check that nothing prints any more, which is the entire reason it was retired.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;action&lt;/th&gt;
&lt;th&gt;requirements&lt;/th&gt;
&lt;th&gt;backed by a test&lt;/th&gt;
&lt;th&gt;coverage&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;before&lt;/td&gt;
&lt;td&gt;263&lt;/td&gt;
&lt;td&gt;191&lt;/td&gt;
&lt;td&gt;72.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;retire one row, stand its replacement&lt;/td&gt;
&lt;td&gt;264&lt;/td&gt;
&lt;td&gt;191&lt;/td&gt;
&lt;td&gt;72.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;correct ten more rows&lt;/td&gt;
&lt;td&gt;274&lt;/td&gt;
&lt;td&gt;191&lt;/td&gt;
&lt;td&gt;69.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Coverage therefore falls slightly every time anyone corrects anything. Nobody reads a one-row drop as a bug. It reads as the ordinary friction of a growing spec, and it points in a direction that discourages exactly the work you want: precision.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why no reader noticed
&lt;/h2&gt;

&lt;p&gt;Both tools locate requirements by matching a heading pattern. A retired heading matches it, because retiring changes a prefix in the text and not the shape of the line. There was no bug in either reader. There was a concept, "retired", that lived in the writing convention and had never been given to anything that reads.&lt;/p&gt;

&lt;p&gt;That is the general shape and it is worth naming: &lt;strong&gt;a convention that is not in the data is not a convention, it is a habit.&lt;/strong&gt; Every reader is free to disagree with it, and they will disagree silently, because each one is behaving exactly as written.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix is not a skip
&lt;/h2&gt;

&lt;p&gt;The obvious repair is to teach both readers to skip the retired prefix. I did not do that, for the same reason I would not fix a duplicated constant by copying it a third time.&lt;/p&gt;

&lt;p&gt;Instead the heading rule became data: one table saying which heading forms are live, which are retired, and what each one means, with the readers deriving their filter from it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight toml"&gt;&lt;code&gt;&lt;span class="nn"&gt;[[heading]]&lt;/span&gt; &lt;span class="py"&gt;form&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"## R-&lt;/span&gt;&lt;span class="err"&gt;\&lt;/span&gt;&lt;span class="s"&gt;d+"&lt;/span&gt;          &lt;span class="py"&gt;status&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"live"&lt;/span&gt;
&lt;span class="nn"&gt;[[heading]]&lt;/span&gt; &lt;span class="py"&gt;form&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"## retired_R-&lt;/span&gt;&lt;span class="err"&gt;\&lt;/span&gt;&lt;span class="s"&gt;d+"&lt;/span&gt;  &lt;span class="py"&gt;status&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"retired"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A third tool, added later, cannot get this wrong by forgetting. It can only get it wrong by not reading the table at all, which is a visible omission rather than a silent divergence.&lt;/p&gt;

&lt;h2&gt;
  
  
  The check that makes the property hold
&lt;/h2&gt;

&lt;p&gt;The property worth testing is not "retired rows are skipped". It is that retiring is denominator-neutral:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;assert: retire one row and stand its replacement
        dod_rows before == dod_rows after
        printed  before == printed  after
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That check would have failed on the day the convention was invented, long before anyone accumulated corrections. It tests the operation rather than the state, which is what you want for anything append-only, because state grows and operations are the only thing you can pin down.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same defect, one layer up
&lt;/h2&gt;

&lt;p&gt;The append-only rule exists so that history is never destroyed. The retirement convention exists so that append-only does not mean accumulating contradictions. The readers were built for the first rule and nobody had told them about the second.&lt;/p&gt;

&lt;p&gt;Whenever you add a rule that makes an earlier rule survivable, check whether everything that consumed the earlier rule knows about the new one. The second rule is usually written as prose, in a different document, months later, by someone who is thinking about the writing side and not the reading side.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things I keep
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Retiring a row must not change any count.&lt;/strong&gt; If correcting something moves your coverage number, the number is measuring your editing activity rather than your code, and the first thing it will teach people is to correct less.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Put the convention in a table the readers load.&lt;/strong&gt; "We mark these as retired" is a sentence in someone's memory. A heading-form table is a thing a new tool has to read, and its absence is loud.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>programming</category>
      <category>testing</category>
      <category>devops</category>
    </item>
    <item>
      <title>I checked 18 claims in my own spec. Five were false, and one tag byte was responsible for one of them.</title>
      <dc:creator>Mahiro Hirakawa</dc:creator>
      <pubDate>Mon, 14 Sep 2026 17:06:02 +0000</pubDate>
      <link>https://dev.to/mahirhir/i-checked-18-claims-in-my-own-spec-five-were-false-and-one-tag-byte-was-responsible-for-one-of-3icn</link>
      <guid>https://dev.to/mahirhir/i-checked-18-claims-in-my-own-spec-five-were-false-and-one-tag-byte-was-responsible-for-one-of-3icn</guid>
      <description>&lt;p&gt;The specification for this project is written in precise language on purpose. It says things like "this map is a homomorphism" and "these two operations form a lens" rather than "this converts X to Y", because the precise words carry obligations and the vague ones do not.&lt;/p&gt;

&lt;p&gt;Nobody had ever checked whether the obligations were met. I built something that reads each claim, resolves the terms against the actual definitions, and returns valid, invalid or unknown.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;validity table  rows=18  valid=9  invalid=5  unknown=4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nine of eighteen. Every invalid one was the same mistake in a different costume: a precise word used at the strength of its vague synonym.&lt;/p&gt;

&lt;h2&gt;
  
  
  One tag byte
&lt;/h2&gt;

&lt;p&gt;The clearest case. The spec said the encoder is a homomorphism into byte strings under concatenation. That is a strong claim with a single obligation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;enc(a · b) = enc(a) ++ enc(b)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here is what the encoder does for a pair:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;enc(pair(x, y)) = tag ++ enc(x) ++ enc(y)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One byte in front, to say what kind of thing follows. That byte is why the encoding is unambiguously decodable, so it is doing real work and it is not going away. It also means the equation is false for every input. Not approximately true, not true up to a constant: there is no pair for which the two sides are equal.&lt;/p&gt;

&lt;p&gt;What the encoder actually is, is a map that is injective and prefix-free. Both of those are properties the rest of the system relies on, and neither of them is a homomorphism. The word was reached for because "structure-preserving" felt like the right shape, and nothing ever asked which structure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four names for four different things
&lt;/h2&gt;

&lt;p&gt;Another row used DAG, strict order, covering relation and transitive closure as if they named one concept. They are related and they are not interchangeable:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;term&lt;/th&gt;
&lt;th&gt;what it is&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DAG&lt;/td&gt;
&lt;td&gt;a graph with no directed cycle&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;strict order&lt;/td&gt;
&lt;td&gt;transitive, irreflexive&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;covering relation&lt;/td&gt;
&lt;td&gt;the edges of the Hasse diagram, no intermediate element&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;transitive closure&lt;/td&gt;
&lt;td&gt;every reachable pair, as an edge&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A DAG is not transitive. A covering relation is deliberately not transitive, and taking its transitive closure gives you back the order, so the two are opposite ends of the same operation. The spec had one sentence that was true of one of these and was written as though it applied to all four, and the code that read that sentence chose a different member than the code that wrote it.&lt;/p&gt;

&lt;h2&gt;
  
  
  A claim its own document denies
&lt;/h2&gt;

&lt;p&gt;The third kind is the one I found most uncomfortable. One row asserted the round-trip laws for a lens, which require:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;get(put(v, s)) = v
put(get(s), s) = s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A different section of the same document, 130 lines away, states plainly that the two maps do not compose to the identity. Both sentences had been read many times, by me, and never on the same day.&lt;/p&gt;

&lt;p&gt;This is the failure mode of a long specification. Contradictions are invisible when the two halves never have to be true at the same moment, and nothing in prose forces that moment to happen.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four unknowns are not failures
&lt;/h2&gt;

&lt;p&gt;Four rows came back unknown, and keeping that as a separate answer is the part I would defend hardest.&lt;/p&gt;

&lt;p&gt;One example: a comparison is total only under a hypothesis supplied by a different layer. Whether the claim holds is not decidable where the claim is written. It is not false, and marking it false would produce a repair for a problem that may not exist. It is also not valid, and calling it valid would launder an assumption into a fact.&lt;/p&gt;

&lt;p&gt;Collapsing unknown into invalid is the more tempting direction, because it makes the report look rigorous and gives everyone something to do. It is still wrong. The honest output is a count of decided and a named list of what a decision is waiting on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why prose lets this happen
&lt;/h2&gt;

&lt;p&gt;None of these five claims survived contact with a checker, and all of them had survived years of reading. Prose does not force you to name the structure, so "this preserves structure" never has to answer &lt;em&gt;which&lt;/em&gt; structure, and the reader supplies whichever one makes the sentence true.&lt;/p&gt;

&lt;p&gt;The repair is not to write less precisely. It is to make each precise word carry its obligation in machine-readable form: when the document says homomorphism, something should be able to ask which operation on each side, and then check it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things I keep
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A precise word is a promise, and the cost of breaking it is paid by a reader who believed you.&lt;/strong&gt; If you are not prepared to have "homomorphism" checked, write "encodes" and lose nothing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Contradictions inside one document are found by evaluation, never by reading.&lt;/strong&gt; Two statements a hundred lines apart are never in your head simultaneously, so no amount of careful reading brings them together. Something that resolves both at once does it in milliseconds.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>programming</category>
      <category>testing</category>
      <category>math</category>
    </item>
    <item>
      <title>One row in my ledger said 12 of 3. Six review rounds later, not a single verdict had changed.</title>
      <dc:creator>Mahiro Hirakawa</dc:creator>
      <pubDate>Mon, 14 Sep 2026 16:58:35 +0000</pubDate>
      <link>https://dev.to/mahirhir/one-row-in-my-ledger-said-12-of-3-six-review-rounds-later-not-a-single-verdict-had-changed-183a</link>
      <guid>https://dev.to/mahirhir/one-row-in-my-ledger-said-12-of-3-six-review-rounds-later-not-a-single-verdict-had-changed-183a</guid>
      <description>&lt;p&gt;There is a ledger here where each row compares our implementation against someone else's on one axis, and carries a letter: ahead, level, behind. Each letter is supported by numbers and by pointers into both codebases.&lt;/p&gt;

&lt;p&gt;A reviewer reads rows and disagrees with them. Six rounds ran over one section of that ledger. Every round found real defects. Not one round changed a letter.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the rounds actually found
&lt;/h2&gt;

&lt;p&gt;Three kinds of thing, none of which is a wrong conclusion.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A closing sentence that was false about its own code.&lt;/strong&gt; One row ended by saying the operation returns "not contained" for a failing name. The code returns three different outcomes: not contained, absent for a name outside the root, and unsettled for a race that ran past its bound. The sentence was not a simplification. It named one of three and implied it was the set.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two counts glued together with a slash.&lt;/strong&gt; One row wrote:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;probes=12/3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;which reads as twelve of three. There was no ratio. There were two independent facts, twelve probes run and three conditions declared, and a slash between them because the surrounding rows all had a shape like &lt;code&gt;n/m&lt;/code&gt; and this one wanted to match.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cited ranges that did not hold what the row claimed.&lt;/strong&gt; A reader that resolves every citation and compares its extent against the claim found several rows off:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;range_wider     row 4:  cites 9 lines, needs 7
range_narrower  row 10: cites 6 lines, needs 12
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second is the dangerous one. A citation narrower than its claim points at a fragment and asserts the whole.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;kind&lt;/th&gt;
&lt;th&gt;changes the verdict&lt;/th&gt;
&lt;th&gt;changes whether you can check it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;clause names 1 of 3 outcomes&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;two counts written as a ratio&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cited range wider or narrower&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Six rounds without a letter moving is not obviously good news
&lt;/h2&gt;

&lt;p&gt;From the outside this is indistinguishable from churn. Six passes, a pile of edits, the same conclusions as before. If someone asked what the review accomplished, "we corrected some numbers" is a weak answer and it is also the true one.&lt;/p&gt;

&lt;p&gt;What changes the reading is what each round left behind. Over those six rounds the checking apparatus gained three readers it did not have: one that resolves a cited range and compares it to the claim, one that refuses a count written as a pair unless both halves come from declared sources, and one that parses the ledger's tables rather than trusting their shape.&lt;/p&gt;

&lt;p&gt;The repaired rows show what the readers now require. A count stays a count, and a sample keeps its population beside it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OK_CONTAIN   conditions=6
OK_ATTEST    probes=12  declared=3        was: probes=12/3
arms=9  sampled=7  ratio=7/7              a pair, with the thing it is a pair of
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So the rounds were not producing opinions about the rows. They were converting a prose ledger into a machine-checkable one, one defect class at a time, and the letters were never what was in question.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test that tells churn apart from progress
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;After a review round, name what now runs that did not run before. If the answer is nothing, that round produced an opinion.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;An opinion is worth having. It is just not worth six rounds, and it does not accumulate, so the seventh round rediscovers what the third one found.&lt;/p&gt;

&lt;p&gt;The three readers above will catch their defect classes in every future row without anyone reading carefully. That is the part that compounds. The corrected rows are a one-time payment.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would do differently
&lt;/h2&gt;

&lt;p&gt;I would say out loud, after round two, that the letters look stable and the evidence does not, because that is a different activity from reviewing the claims and it deserves to be named. Rounds three through six were an audit of the citation apparatus wearing the costume of a content review, and nobody said so, so each round started by re-reading rows that were not the problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things I keep
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A count and a ratio look identical and mean different things.&lt;/strong&gt; &lt;code&gt;12/3&lt;/code&gt; passed several readings because it had the shape of the rows around it. Any number written as a pair needs both halves to come from somewhere declared, and a reader that enforces that is fifteen lines.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Review rounds that never move a verdict are measuring your evidence, not your claims.&lt;/strong&gt; That is legitimate work and it is a different job. Name it when you notice it, because the two jobs want different reviewers, different stopping conditions, and, most of all, a different definition of done.&lt;/p&gt;

</description>
      <category>testing</category>
      <category>architecture</category>
      <category>programming</category>
      <category>devops</category>
    </item>
    <item>
      <title>The same one-line table-parsing bug turned up in five separate tools. I fixed it four times.</title>
      <dc:creator>Mahiro Hirakawa</dc:creator>
      <pubDate>Mon, 14 Sep 2026 16:05:05 +0000</pubDate>
      <link>https://dev.to/mahirhir/the-same-one-line-table-parsing-bug-turned-up-in-five-separate-tools-i-fixed-it-four-times-4l5o</link>
      <guid>https://dev.to/mahirhir/the-same-one-line-table-parsing-bug-turned-up-in-five-separate-tools-i-fixed-it-four-times-4l5o</guid>
      <description>&lt;p&gt;A lot of this project's specification lives in markdown tables, and several tools read them. Parsing a table row looks like a solved problem:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;cells&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;line&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;|&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;slice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That line is in more places than anyone had counted.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two cells broke it
&lt;/h2&gt;

&lt;p&gt;One spec table has a cell containing a set definition and another containing the length notation for a sequence. Written out, they carry a vertical bar inside the cell:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;| E0 | the admitted set  | N_t := {x &lt;span class="se"&gt;\|&lt;/span&gt; E(x,t) = true} |
| S0 | length            | &lt;span class="se"&gt;\|&lt;/span&gt;·&lt;span class="se"&gt;\|&lt;/span&gt; over a byte sequence   |
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both are correct markdown. The bar is escaped, which is exactly what the format asks for. The naive split does not care that it was escaped, so both rows fractured into the wrong number of cells.&lt;/p&gt;

&lt;h2&gt;
  
  
  The symptom was not "the parser is broken"
&lt;/h2&gt;

&lt;p&gt;This is the part I want to keep. What the run reported was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;verified_lying&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;2&lt;/span&gt;
&lt;span class="na"&gt;premise_defs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;   &lt;span class="m"&gt;2&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two definitions in the specification appeared to rest on nothing. That is a genuinely alarming finding, several steps more serious than a lost cell, and it is what everyone reacted to. The first stretch of work went into the two definitions themselves, both of which turned out to be fully reproduced by a proof and a test.&lt;/p&gt;

&lt;p&gt;A parser failure arrives dressed as a content failure. The cells shifted, so a column that should have held a reproduction pointer held something else, and every check downstream honestly reported what it was given.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;what the log said&lt;/th&gt;
&lt;th&gt;what it meant&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;two definitions unverified&lt;/td&gt;
&lt;td&gt;a column moved by one&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;the specification has a hole&lt;/td&gt;
&lt;td&gt;the reader has a hole&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  It was the fourth time, and then the fifth
&lt;/h2&gt;

&lt;p&gt;The same class had already been found in a proof-declaration printer and in the map generator. Each time, the person who hit it fixed the tool in front of them and carried on, which is a completely reasonable thing to do and is why it kept happening.&lt;/p&gt;

&lt;p&gt;A fifth instance turned up in a shared cell-splitting helper the same week.&lt;/p&gt;

&lt;p&gt;Nothing anywhere knew how many table readers existed. There was no wrong decision to point at. There was an absent list.&lt;/p&gt;

&lt;h2&gt;
  
  
  The repair was the inventory
&lt;/h2&gt;

&lt;p&gt;The fix that ended it was not a better regex. Every tool that parses a table became a declared row, and a check reads that list:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OK_TABLE_READERS  readers=7/7  escaped_pipe=1  lookalike=1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A new tool that reads tables has to be added to the list, and each one is exercised against a fixture containing both hazards. Adding the row is not optional in the sense that the check counts the readers it can find in the tree and compares them to the declared ones.&lt;/p&gt;

&lt;p&gt;The second hazard is the one worth stealing. The mathematical notation does not actually want U+007C. It wants U+2223, the DIVIDES character, which looks nearly identical in a monospace font and is a completely different codepoint. A reader that learned about the escaped pipe and nothing else still gets that row wrong, silently, and the fixture now contains one of each.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things I keep
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A bug found more than twice is not a bug. It is a missing inventory.&lt;/strong&gt; The third occurrence is the signal to stop patching and start asking how many places have this shape, because the patch rate is clearly not keeping up with the discovery rate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Expect a parser failure to present as a content failure.&lt;/strong&gt; Whatever sits downstream of the parse will report faithfully on the garbage it received, and its report will be about your data rather than about your reader. When a check suddenly claims something alarming about content that was fine yesterday, suspect the thing that fed it before you suspect the content.&lt;/p&gt;

</description>
      <category>programming</category>
      <category>testing</category>
      <category>devops</category>
      <category>architecture</category>
    </item>
    <item>
      <title>My guard asserted the output path was a scratch directory. The row landed in the permanent log anyway.</title>
      <dc:creator>Mahiro Hirakawa</dc:creator>
      <pubDate>Mon, 14 Sep 2026 15:58:36 +0000</pubDate>
      <link>https://dev.to/mahirhir/my-guard-asserted-the-output-path-was-a-scratch-directory-the-row-landed-in-the-permanent-log-5043</link>
      <guid>https://dev.to/mahirhir/my-guard-asserted-the-output-path-was-a-scratch-directory-the-row-landed-in-the-permanent-log-5043</guid>
      <description>&lt;p&gt;There is a rule here that tests never write to the real record store, and unlike most of my rules it was already a device rather than a sentence. Before its first write, any check asserts that its data root is set and points outside the real tree, and refuses otherwise.&lt;/p&gt;

&lt;p&gt;The rule exists because of an earlier incident, where an end-to-end run appended three entries into the permanent log and nobody noticed until the numbering looked wrong. The guard was the repair.&lt;/p&gt;

&lt;p&gt;Months later a reviewer tool minted a row into that same permanent log. The guard was green on the run that did it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the guard compared
&lt;/h2&gt;

&lt;p&gt;Reduced to its shape:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;root&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;RECORD_ROOT&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;root&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;RECORD_ROOT not set&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;root&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;startsWith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;REAL_TREE&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;refusing to write the real tree&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="c1"&gt;// ... proceed to write under root&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The scratch root was a fresh directory under the temp area. The string did not start with the real tree's path. The assertion was true.&lt;/p&gt;

&lt;p&gt;Inside that scratch tree was a junction, the Windows form of a directory symlink, created to give the tool the same layout it sees in the real checkout without copying a large tree. When the tool opened a file under that junction, the runtime followed it. The write went where the link pointed.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;what the guard looked at&lt;/th&gt;
&lt;th&gt;what the write used&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;the string in an environment variable&lt;/td&gt;
&lt;td&gt;the path the filesystem resolved&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;the intent, before anything opened&lt;/td&gt;
&lt;td&gt;the destination, after every link was followed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Both are called "the path". They are not the same object, and a string comparison cannot tell them apart.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix is two lines, and one of them is the interesting one
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;realpathSync&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;node:fs&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;root&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;realpathSync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;RECORD_ROOT&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;        &lt;span class="c1"&gt;// resolve, then compare&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;root&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;startsWith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;realpathSync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;REAL_TREE&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;...&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Resolving both sides is the obvious half. The half worth keeping is the second change, which is to move the assertion from the intent to the handle. Instead of checking a path before writing, check what you actually opened:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;fh&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;target&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;w&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;ino&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;dev&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;fh&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stat&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;isUnderRealTree&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;dev&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;ino&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;...&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That version cannot be defeated by anything the resolver does, because it is asking the thing that was actually opened rather than the request that produced it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The general shape
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A guard that validates an intent is not guarding the action.&lt;/strong&gt; Between the value you check and the effect that happens, there is a resolver, and the resolver is a component you did not audit.&lt;/p&gt;

&lt;p&gt;This is not specific to symlinks. The same gap is everywhere:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;what you check                 what resolves it            what actually happens
a path string          -&amp;gt;      symlinks, junctions, ..  -&amp;gt; a different inode
a hostname             -&amp;gt;      DNS, /etc/hosts, proxy   -&amp;gt; a different server
a database name        -&amp;gt;      connection pooling, DSN  -&amp;gt; a different database
a branch name          -&amp;gt;      remote refs, worktrees   -&amp;gt; a different ref
a container mount      -&amp;gt;      bind mounts              -&amp;gt; the host filesystem
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every row is the same failure waiting. The check passes on the name, the effect lands on the resolution, and the distance between them is invisible in the log line that says the check was green.&lt;/p&gt;

&lt;h2&gt;
  
  
  Plant the resolver in your own fixtures
&lt;/h2&gt;

&lt;p&gt;The lasting change was not the two lines. It was that this class became something the test fixtures deliberately produce.&lt;/p&gt;

&lt;p&gt;A related checker in the same project had been silently skipping files it could not read, which nobody knew because no fixture had ever contained one. Its selftest now plants two: a dangling junction, and a file with permissions denied. The checker is required to report both rather than pass over them.&lt;/p&gt;

&lt;p&gt;Once you know a resolver sits between your check and your effect, the fixture has to contain something that resolves surprisingly. Otherwise the only place that ever happens is production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things I keep
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Compare resolved values, or better, do not compare paths at all.&lt;/strong&gt; Ask the opened handle where it is. A string is a request, and the answer to a request is a different thing than the request.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A rule that is already a device can still be wrong about what it measures.&lt;/strong&gt; I had promoted this one from a sentence to a check and then stopped thinking about it, which is precisely the moment it is least examined. Being enforced is not the same as being correct.&lt;/p&gt;

</description>
      <category>testing</category>
      <category>programming</category>
      <category>devops</category>
      <category>node</category>
    </item>
    <item>
      <title>I wrote three rules on Saturday. I broke all three on Saturday.</title>
      <dc:creator>Mahiro Hirakawa</dc:creator>
      <pubDate>Mon, 14 Sep 2026 15:32:27 +0000</pubDate>
      <link>https://dev.to/mahirhir/i-wrote-three-rules-on-saturday-i-broke-all-three-on-saturday-5bd6</link>
      <guid>https://dev.to/mahirhir/i-wrote-three-rules-on-saturday-i-broke-all-three-on-saturday-5bd6</guid>
      <description>&lt;p&gt;A long working day produced three lessons. Each got written down as one sentence in the log, the way a good lesson should be:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1. Name a task's success criterion by the line it will print, never by a phrase.
2. Regenerate the derived index before mirroring the tree, not after.
3. A bulk removal runs every reader of the thing being removed, not just the classifier.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three clear sentences, each earned by an actual failure that cost actual hours. All three were violated again before the same day ended.&lt;/p&gt;

&lt;h2&gt;
  
  
  It was not a discipline problem
&lt;/h2&gt;

&lt;p&gt;I want to be precise about this, because the comfortable reading is that I was tired and should have been more careful.&lt;/p&gt;

&lt;p&gt;I knew all three rules. I had written them myself, that day, in a file I had open. At the moment each one applied I was not thinking about rules at all. I was thinking about the task, which is the correct thing to be thinking about, and that is exactly when a sentence in a log provides no resistance whatsoever.&lt;/p&gt;

&lt;p&gt;A sentence has no enforcement surface. Nothing consults it. The gap between writing a rule and needing it is measured in hours and in context switches, and a rule that lives only in a document loses every one of those.&lt;/p&gt;

&lt;h2&gt;
  
  
  What survived versus what did not
&lt;/h2&gt;

&lt;p&gt;The interesting number is not three violations. It is that the rules already implemented as devices were not violated, because violating them was not available.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;rule&lt;/th&gt;
&lt;th&gt;form&lt;/th&gt;
&lt;th&gt;violations that day&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;name success by its print&lt;/td&gt;
&lt;td&gt;sentence&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;index before mirror&lt;/td&gt;
&lt;td&gt;sentence&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;removal runs every reader&lt;/td&gt;
&lt;td&gt;sentence&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;a control must not write the real data store&lt;/td&gt;
&lt;td&gt;device, refuses&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;a claim of having read needs a verified quote&lt;/td&gt;
&lt;td&gt;device, refuses&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The bottom two are not better rules. They are the same kind of rule, wired into something that stops.&lt;/p&gt;

&lt;h2&gt;
  
  
  Each sentence became a refusal
&lt;/h2&gt;

&lt;p&gt;The repair was not to write the sentences more forcefully. Each one became a thing that declines to proceed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;plan &lt;span class="nt"&gt;--add&lt;/span&gt; &lt;span class="nv"&gt;task&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;rework &lt;span class="nt"&gt;--success&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"the merge is clean"&lt;/span&gt;
&lt;span class="go"&gt;REFUSED  success token 'the merge is clean' is not printed anywhere in this tree
         success= must name a line some device actually emits
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The planner now reads every print name the tree can produce and refuses a success criterion that is a phrase rather than a line. The first rule stopped being advice and became a thing you cannot type past.&lt;/p&gt;

&lt;p&gt;The second became data. The step order moved out of my head and into a file, with a checker that reads it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;order&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;edit&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;repin&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;recite&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;index&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;mirror&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;controls&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
         &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;append&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;index&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;accept&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;record&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The checker refuses an index regeneration while an edit is still in flight, and refuses to commit an index without the index's own success line present in the same run. There is no longer a version of the sequence that lives in my memory and can be reordered by hurry.&lt;/p&gt;

&lt;p&gt;The third turned a one-off script into a tool with the readers declared as rows, so "run every reader" stopped meaning "remember which tools parse this" and started meaning "iterate a list that someone has to add to".&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape underneath all three
&lt;/h2&gt;

&lt;p&gt;Every one of these was the same failure wearing a different hat: &lt;strong&gt;a step performed by habit where it should have been performed by a device&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That gives a usable test for any rule you are about to write down:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Name the thing that will refuse. If you cannot, you have written a note, not a rule.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A note is still worth writing. It is the record of a real failure and it is what the device gets built from later. The error is counting it as a fix and moving on, because the next violation then looks like carelessness rather than like the predictable behaviour of an unenforced sentence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things I keep
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A rule applied by hand is a rule you will break, and you can measure how fast.&lt;/strong&gt; Write the sentence, then count violations for a week before believing the problem is solved. Mine had a zero percent survival rate over about eight hours, which is more informative than any amount of confidence about my own carefulness.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Knowing a rule does not change behaviour at the moment the rule applies.&lt;/strong&gt; At that moment your attention is entirely on the task, which is where it belongs. Design for that, rather than for a version of yourself who is thinking about process while doing work.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>programming</category>
      <category>architecture</category>
      <category>testing</category>
    </item>
    <item>
      <title>6,718 of my 10,238 code blocks had no owner. Every single one was in the same directory.</title>
      <dc:creator>Mahiro Hirakawa</dc:creator>
      <pubDate>Mon, 14 Sep 2026 15:03:19 +0000</pubDate>
      <link>https://dev.to/mahirhir/6718-of-my-10238-code-blocks-had-no-owner-every-single-one-was-in-the-same-directory-3fhf</link>
      <guid>https://dev.to/mahirhir/6718-of-my-10238-code-blocks-had-no-owner-every-single-one-was-in-the-same-directory-3fhf</guid>
      <description>&lt;p&gt;The rule is that every block of code stands under a declared owner, so that for any function you can ask which stated responsibility it serves and get one answer. The rule had existed for a while. Nothing had ever counted it.&lt;/p&gt;

&lt;p&gt;The first census:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;kind&lt;/th&gt;
&lt;th&gt;total&lt;/th&gt;
&lt;th&gt;under an owner&lt;/th&gt;
&lt;th&gt;orphan&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;code blocks&lt;/td&gt;
&lt;td&gt;10,238&lt;/td&gt;
&lt;td&gt;3,520&lt;/td&gt;
&lt;td&gt;6,718&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rust functions&lt;/td&gt;
&lt;td&gt;2,004&lt;/td&gt;
&lt;td&gt;1,088&lt;/td&gt;
&lt;td&gt;916&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;proof declarations&lt;/td&gt;
&lt;td&gt;600&lt;/td&gt;
&lt;td&gt;213&lt;/td&gt;
&lt;td&gt;387&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cross-layer edges&lt;/td&gt;
&lt;td&gt;50&lt;/td&gt;
&lt;td&gt;35&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Sixty-six percent of the code blocks had no owner. That is the kind of number that ends a meeting.&lt;/p&gt;

&lt;h2&gt;
  
  
  The big number was one decision
&lt;/h2&gt;

&lt;p&gt;The census also printed where the orphans were, and the answer was all of them in the same place:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;blocks with no owner    = 6,718
of those, under tools/  = 6,718   (100%)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The instrument directory had never been declared. Every script that checks, counts or verifies something lived outside the ownership model entirely, not because anyone decided to exclude it but because the model had been written while thinking about the product and the tools had accumulated beside it.&lt;/p&gt;

&lt;p&gt;So this was not 6,718 problems. It was one missing declaration, and the number went to zero when a layer for the instruments was declared and the devices were listed under it. The largest finding in the census was the cheapest one in it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The small number was real work
&lt;/h2&gt;

&lt;p&gt;The 916 orphan Rust functions looked smaller and were not:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;orphan Rust functions = 916
top three crates      = 142, 101, 91
crates involved       = 14
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nothing groups those. They are spread across fourteen crates, and the top crate holds fifteen percent of them. There is no single row that covers this. It is fourteen passes of reading, one crate at a time, deciding for each function which stated responsibility it serves or declaring it a helper of one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same census produced two opposite kinds of finding
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;finding&lt;/th&gt;
&lt;th&gt;count&lt;/th&gt;
&lt;th&gt;concentration&lt;/th&gt;
&lt;th&gt;what it is&lt;/th&gt;
&lt;th&gt;cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;unowned blocks&lt;/td&gt;
&lt;td&gt;6,718&lt;/td&gt;
&lt;td&gt;100% in one directory&lt;/td&gt;
&lt;td&gt;a missing declaration&lt;/td&gt;
&lt;td&gt;one row&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;unowned functions&lt;/td&gt;
&lt;td&gt;916&lt;/td&gt;
&lt;td&gt;15% in the largest crate&lt;/td&gt;
&lt;td&gt;genuine backlog&lt;/td&gt;
&lt;td&gt;fourteen passes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The counts point the wrong way. Read as bare totals, the first looks like the catastrophe and the second like a rounding error. Grouped, they swap places.&lt;/p&gt;

&lt;h2&gt;
  
  
  What grouping is actually doing
&lt;/h2&gt;

&lt;p&gt;It is asking whether the findings share a cause. A flat list of N findings silently asserts that there are N decisions to make, and that assertion is almost never checked, because a census feels finished the moment it produces a number.&lt;/p&gt;

&lt;p&gt;The failure mode is not that you file 6,718 tickets. Nobody does that. What actually happens is you look at 66 percent, conclude the rule is unenforceable at this scale, and quietly stop measuring. A number too large to act on gets treated as a statement about the rule rather than a statement about the code, and the census that produced it does not get run again.&lt;/p&gt;

&lt;p&gt;So the census prints the concentration next to every count, and a finding without a grouping column is not yet a finding.&lt;/p&gt;

&lt;h2&gt;
  
  
  The estimate that started it was wrong
&lt;/h2&gt;

&lt;p&gt;One more thing from the same run, worth its own paragraph.&lt;/p&gt;

&lt;p&gt;The task was written with an estimate in it: roughly 653 proof declarations. That number came from someone's memory. When the work actually counted, three different ways of counting gave three different answers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;declarations, by three greps:  554   589   600
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The right move is the one that was taken: report all three, flag the disagreement, and do not correct any of them yet. Each grep encodes a different guess about what counts as a declaration, so the spread is a statement about an undefined term rather than about the code.&lt;/p&gt;

&lt;p&gt;Choosing one, especially the one closest to the estimate in the brief, would have converted an open question into a fact by preference. Once a denominator is written down, nobody re-derives it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things I keep
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Print the concentration beside every count.&lt;/strong&gt; A census that reports totals makes every finding look like the same kind of thing, and the whole value of a census is telling apart the finding that is one decision from the finding that is a month.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When two instruments disagree about a denominator, the disagreement is the finding.&lt;/strong&gt; Not an inconvenience to resolve before reporting. The gap between 554 and 600 is the most informative output of that run, because it says the term being counted was never defined.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>programming</category>
      <category>devops</category>
      <category>testing</category>
    </item>
    <item>
      <title>My coverage print read 13 of 13 on every run. It was incapable of printing anything else.</title>
      <dc:creator>Mahiro Hirakawa</dc:creator>
      <pubDate>Mon, 14 Sep 2026 14:58:42 +0000</pubDate>
      <link>https://dev.to/mahirhir/my-coverage-print-read-13-of-13-on-every-run-it-was-incapable-of-printing-anything-else-33k9</link>
      <guid>https://dev.to/mahirhir/my-coverage-print-read-13-of-13-on-every-run-it-was-incapable-of-printing-anything-else-33k9</guid>
      <description>&lt;p&gt;One of the crates classifies a failure into its worst matching category. The category list is an array where each entry carries a code, and a check prints how much of that list the classifier actually covers.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OK_WORST_CLASS  arms=13/13
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Thirteen of thirteen, every run, for as long as the check had existed. I read that as good news for months. It was not news at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numerator was built out of the denominator
&lt;/h2&gt;

&lt;p&gt;Reduced to the shape, the check was this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="n"&gt;ARMS&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;Arm&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="mi"&gt;13&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt; &lt;span class="cm"&gt;/* ... */&lt;/span&gt; &lt;span class="p"&gt;];&lt;/span&gt;

&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;codes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;Vec&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;Code&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ARMS&lt;/span&gt;&lt;span class="nf"&gt;.iter&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="nf"&gt;.map&lt;/span&gt;&lt;span class="p"&gt;(|&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;|&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="nf"&gt;.code&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;&lt;span class="nf"&gt;.collect&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="nd"&gt;println!&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"OK_WORST_CLASS arms={}/{}"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;codes&lt;/span&gt;&lt;span class="nf"&gt;.len&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;ARMS&lt;/span&gt;&lt;span class="nf"&gt;.len&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;codes&lt;/code&gt; is &lt;code&gt;ARMS&lt;/code&gt; with a function mapped over it. &lt;code&gt;map&lt;/code&gt; does not drop elements. So &lt;code&gt;codes.len()&lt;/code&gt; is &lt;code&gt;ARMS.len()&lt;/code&gt; by construction, and the fraction is &lt;code&gt;n/n&lt;/code&gt; for every possible value of &lt;code&gt;n&lt;/code&gt;, including zero.&lt;/p&gt;

&lt;p&gt;This is not a subtle bug in the counting. There is no input, no state, no ordering that makes this print anything except full. The check had no failure mode. A check with no failure mode is a &lt;code&gt;println!&lt;/code&gt; wearing the costume of a check.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first repair was also wrong
&lt;/h2&gt;

&lt;p&gt;The obvious fix is to make the numerator mean something. So the numerator became "arms whose code agrees with what the classifier returns for that arm", which is a real property that can be false.&lt;/p&gt;

&lt;p&gt;The print went to &lt;code&gt;arms=13/13&lt;/code&gt; again. This time it deserved the number, and it was still not measuring the thing anyone was afraid of.&lt;/p&gt;

&lt;p&gt;The fear is not "an arm I wrote disagrees with itself". The fear is &lt;strong&gt;an error case with no arm at all&lt;/strong&gt;. A new variant gets added to the error type. Nobody adds a row to &lt;code&gt;ARMS&lt;/code&gt;, so the classifier silently files it under a neighbour.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;what the fraction can see&lt;/th&gt;
&lt;th&gt;what it cannot&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;an arm whose code disagrees&lt;/td&gt;
&lt;td&gt;a case with no arm&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;an arm that was written wrong&lt;/td&gt;
&lt;td&gt;a case nobody wrote&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;drift inside the list&lt;/td&gt;
&lt;td&gt;absence from the list&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Both halves of the fraction were still derived from &lt;code&gt;ARMS&lt;/code&gt;. A list cannot report its own omissions. Whatever is missing from it is missing from the numerator and the denominator at the same time, which is exactly the cancellation that keeps the ratio at 1.&lt;/p&gt;

&lt;h2&gt;
  
  
  The denominator has to come from the other side
&lt;/h2&gt;

&lt;p&gt;The repair is not in the counting. It is in where the denominator is read from.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="c1"&gt;// before: both sides from the list of arms I wrote&lt;/span&gt;
&lt;span class="n"&gt;arms&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;agreeing_arms&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;ARMS&lt;/span&gt;&lt;span class="nf"&gt;.len&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="c1"&gt;// after: the denominator is the thing being covered&lt;/span&gt;
&lt;span class="n"&gt;arms&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;agreeing_arms&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nn"&gt;ErrorKind&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;ALL&lt;/span&gt;&lt;span class="nf"&gt;.len&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;ErrorKind::ALL&lt;/code&gt; is the closed set the classifier is supposed to cover. Now adding a variant without adding an arm moves the denominator and not the numerator, and the print falls to &lt;code&gt;13/14&lt;/code&gt; on the commit that introduces the gap.&lt;/p&gt;

&lt;p&gt;The general rule, which is obvious once stated and easy to violate once writing code:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A coverage number is only meaningful when its denominator comes from a different source than its numerator.&lt;/strong&gt; Same file is a warning. Same array is a guarantee of nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The class is findable by machine
&lt;/h2&gt;

&lt;p&gt;The part I find most useful is that this does not require noticing. An audit can read a check's own source and ask whether the expression producing the numerator is derived from the expression producing the denominator. If it is, the check gets a named finding. "Aggregation gap" is now a category in that audit, sitting next to the other ways a check lies.&lt;/p&gt;

&lt;p&gt;It fired on two checks the day it was written, and both had been green since the day they were merged. Neither had ever been suspected, because a green check is not something anyone goes looking at.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;audit &lt;span class="nt"&gt;--class&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;aggregation-gap
&lt;span class="go"&gt;FAIL  worst_class    numerator derived from denominator (map over the same array)
FAIL  arms_declared  denominator is a compile-time constant of the numerator's source
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Two things I keep
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A ratio that has never moved is not evidence of health.&lt;/strong&gt; It is an untested hypothesis about your own arithmetic, and the longer it has held at 100% the less anyone remembers to doubt it. The question to ask a green check is not "is this right" but "what input makes this print something else".&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Absence is the thing coverage is for, and absence is exactly what a self-derived denominator cannot represent.&lt;/strong&gt; If the list defines both what you have and what you should have, you have written a tautology and given it a percent sign.&lt;/p&gt;

</description>
      <category>testing</category>
      <category>rust</category>
      <category>programming</category>
      <category>architecture</category>
    </item>
    <item>
      <title>My pipeline recorded 0 of 15 into an append-only file before anyone read why. The real answer was 9.</title>
      <dc:creator>Mahiro Hirakawa</dc:creator>
      <pubDate>Mon, 14 Sep 2026 14:32:30 +0000</pubDate>
      <link>https://dev.to/mahirhir/my-pipeline-recorded-0-of-15-into-an-append-only-file-before-anyone-read-why-the-real-answer-was-9-1f27</link>
      <guid>https://dev.to/mahirhir/my-pipeline-recorded-0-of-15-into-an-append-only-file-before-anyone-read-why-the-real-answer-was-9-1f27</guid>
      <description>&lt;p&gt;The release boundary for this project is ten declared steps. Two of them are next to each other and are separate on purpose:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;step 8   accept_read     compute the acceptance verdict for every unit and print it
step 9   record          write that verdict into the permanent record
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Step 8 looks. Step 9 commits what was seen. The gap between them exists so a person can read the verdict before it becomes history.&lt;/p&gt;

&lt;p&gt;I ran them as one command chain. Both steps are always needed and the second always follows the first, so joining them felt like removing a redundant keystroke.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the read actually returned
&lt;/h2&gt;

&lt;p&gt;Every one of the fifteen units came back red. Not one of them for a reason that had anything to do with the units:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;&lt;span class="n"&gt;FAIL_SEMANTIC&lt;/span&gt;  &lt;span class="n"&gt;reason&lt;/span&gt;=&lt;span class="n"&gt;op_unknown&lt;/span&gt;  &lt;span class="n"&gt;op_unknown&lt;/span&gt;=&lt;span class="m"&gt;51&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A generator had started emitting a new form of annotation that morning. The checker that validates those annotations works from a closed vocabulary. That is the right design, and the 51 new names were simply not in it yet. The checker refused, correctly, the way a closed vocabulary is supposed to refuse.&lt;/p&gt;

&lt;p&gt;Nothing was broken. One list needed 51 entries added to it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The record step did not care
&lt;/h2&gt;

&lt;p&gt;It had no reason to. Its input was a verdict and the verdict said red, so it wrote the truth it was handed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="s"&gt;ACCEPT_RECORDED rows=15 accepted=0/15 commit=f6d74b1&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That file is append-only. Records are never edited and never removed, which is a property I chose and still want, because a record you can quietly fix is not a record. The 0/15 is permanent. It sits in the history with a later block superseding it, and anyone reading that history sees a day when everything failed followed by a correction.&lt;/p&gt;

&lt;p&gt;Twenty minutes later, after the vocabulary was declared and the map regenerated, the read was run again and this time read first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="s"&gt;ACCEPT_RECORDED rows=15 accepted=9/15 commit=2536c12&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nine. The false number cost more to carry than the missing 51 entries cost to add.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;code&gt;&amp;amp;&amp;amp;&lt;/code&gt; protects the thing you were not worried about
&lt;/h2&gt;

&lt;p&gt;The chain even looks like a safety property. "Only record if the read succeeds" sounds like exactly the guard you would want.&lt;/p&gt;

&lt;p&gt;The read did succeed. It ran to completion, computed a verdict for all fifteen units and printed its reason, then exited 0. Exit 0 was correct. A measuring device that reports a red measurement has done its job perfectly.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;what the chain checks&lt;/th&gt;
&lt;th&gt;what it does not check&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;did the read process run&lt;/td&gt;
&lt;td&gt;is the verdict worth committing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;did it exit non-zero&lt;/td&gt;
&lt;td&gt;did a human look at the reason&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;did it crash&lt;/td&gt;
&lt;td&gt;is the red about the subject or about the instrument&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That last row is the one that bit. &lt;code&gt;op_unknown=51&lt;/code&gt; is a statement about the checker's own configuration, not about the fifteen units it was pointed at. No exit code can carry that distinction, because both cases are a successful run that produced a red.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule I now hold
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A step that observes and a step that commits the observation are never one invocation.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Not joined by &lt;code&gt;&amp;amp;&amp;amp;&lt;/code&gt;. Not joined by &lt;code&gt;;&lt;/code&gt;. Not wrapped in a convenience script. Two commands, and in practice two turns, so that something slow enough to read sits between the measurement and the permanent consequence.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;boundary &lt;span class="nt"&gt;--step&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;accept_read      &lt;span class="c"&gt;# look, and stop&lt;/span&gt;
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;boundary &lt;span class="nt"&gt;--step&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;record           &lt;span class="c"&gt;# commit, as a separate decision&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the same shape as &lt;code&gt;plan&lt;/code&gt; and &lt;code&gt;apply&lt;/code&gt;, or a dry run against a real one. It is routinely collapsed for the same reason I collapsed it: the two commands always run together in the happy path, and the happy path is where the habit forms.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things I keep
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Exit 0 means the measurement ran. It never means the measurement was good.&lt;/strong&gt; A device that says "red" and exits 0 is a healthy device, so any automation keyed on the exit code is reading the wrong field for this question.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The cost of collapsing "look" into "act" is paid in permanent records, not in retries.&lt;/strong&gt; If the consequence were a retry I would never have noticed, and I would have kept the shortcut. The reason this one taught me anything is that the file it wrote to cannot be edited.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>programming</category>
      <category>testing</category>
      <category>architecture</category>
    </item>
  </channel>
</rss>
