<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Self-Correcting Systems</title>
    <description>The latest articles on DEV Community by Self-Correcting Systems (@kenielzep97).</description>
    <link>https://dev.to/kenielzep97</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3948231%2F09b103a6-4a7f-4016-a261-debf92d5b015.jpeg</url>
      <title>DEV Community: Self-Correcting Systems</title>
      <link>https://dev.to/kenielzep97</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kenielzep97"/>
    <language>en</language>
    <item>
      <title>My Health Check Watched the Wrong File</title>
      <dc:creator>Self-Correcting Systems</dc:creator>
      <pubDate>Sun, 04 Oct 2026 18:05:19 +0000</pubDate>
      <link>https://dev.to/kenielzep97/my-health-check-watched-the-wrong-file-p7h</link>
      <guid>https://dev.to/kenielzep97/my-health-check-watched-the-wrong-file-p7h</guid>
      <description>&lt;p&gt;I wrote a rule for my own agents:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Liveness is not usefulness. Artifact age beats PID.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A job that is loaded, looks healthy to the supervisor, and is producing nothing is the failure worth catching. Anyone can detect exit 127.&lt;/p&gt;

&lt;p&gt;Then I applied the rule to my own machine and got the wrong answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I saw
&lt;/h2&gt;

&lt;p&gt;Three listener jobs. &lt;code&gt;launchctl&lt;/code&gt; lists them. &lt;code&gt;pgrep -f sentinel_listen.py&lt;/code&gt; returns a PID. Their log files:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;sentinel_listen.log    8103 bytes   mtime 2026-08-11
courtside_listen.log   8127 bytes   mtime 2026-08-11
eye_listen.log         8149 bytes   mtime 2026-08-11
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Last written August 11. By my own rule, that is three dead jobs wearing a costume. I had the sentence written.&lt;/p&gt;

&lt;p&gt;A word about what these are, so nobody has to guess. They are small Telegram pollers, 148 to 243 lines each. They wait for a message and answer it. They are not the interesting part of anything I build, and the failure below is not a failure of theirs — it is a failure of how I checked them. I am using them as the specimen precisely because they are simple. If a health check cannot get a 148-line polling loop right, it will not get anything harder right either.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the script actually does
&lt;/h2&gt;

&lt;p&gt;Before publishing that, I opened the thing that writes the file. There are three &lt;code&gt;print()&lt;/code&gt; calls in it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;scripts&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;sentinel_listen&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;py&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;120&lt;/span&gt;    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SENTINEL LISTEN: no rail configured ...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;scripts&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;sentinel_listen&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;py&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;128&lt;/span&gt;    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SENTINEL LISTEN: awake, polling every 20s ...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;scripts&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;sentinel_listen&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;py&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;142&lt;/span&gt;    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SENTINEL LISTEN error: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;type&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="n"&gt;__name__&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Startup, banner, error. Courtside and eye are the same shape at :144/:152/:175 and :198/:201/:236.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The successful path prints nothing.&lt;/strong&gt; There is no "still polling" heartbeat. On the normal path it calls &lt;code&gt;getUpdates&lt;/code&gt; with a 25-second long-poll, writes a small state file, sleeps 20, and says nothing at all. So a quiet pass is the poll plus 20 seconds, not 20. A caught exception sleeps 30 and then takes that same sleep on the way round. The interval is not uniform across passes.&lt;/p&gt;

&lt;p&gt;So silence in that log does not mean the job stopped. The file is only written on startup and on a caught exception, and I was reading its silence as death.&lt;/p&gt;

&lt;p&gt;I cannot go further than that, and the reason matters. These three run without &lt;code&gt;-u&lt;/code&gt;, so stdout is block-buffered. An exception printed one minute ago could still be sitting in a buffer instead of on disk. &lt;strong&gt;An unchanged file does not prove no errors occurred.&lt;/strong&gt; It proves the file has not been modified. Those are different statements, and the second one is the only one I can support.&lt;/p&gt;

&lt;p&gt;The 8103 bytes are an old error storm — one banner and eighty caught exceptions: 76 &lt;code&gt;URLError&lt;/code&gt;, 3 &lt;code&gt;timeout&lt;/code&gt;, 1 &lt;code&gt;HTTPError&lt;/code&gt;. That is what the file records. What it has failed to record since is a question the file cannot answer about itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The file that was actually moving
&lt;/h2&gt;

&lt;p&gt;While I was checking whether these jobs were dead, they were writing. I watched all three state files advance. Sentinel, for example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;stat&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; &lt;span class="s1"&gt;'%Sm %N'&lt;/span&gt; &lt;span class="nt"&gt;-t&lt;/span&gt; &lt;span class="s1"&gt;'%H:%M:%S'&lt;/span&gt; scripts/sentinel_listen_state.json
  11:13:03   sentinel_listen_state.json
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;stat&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; &lt;span class="s1"&gt;'%Sm %N'&lt;/span&gt; &lt;span class="nt"&gt;-t&lt;/span&gt; &lt;span class="s1"&gt;'%H:%M:%S'&lt;/span&gt; scripts/sentinel_listen_state.json
  11:15:19   sentinel_listen_state.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The log had not been touched since August 11 and the state file was seconds old. Both belong to the same listener job, but only one of them was designed to move on ordinary loop iterations — and they live in different directories. I was watching &lt;code&gt;agent_outputs/&lt;/code&gt;. The living file was in &lt;code&gt;scripts/&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;I picked the artifact that only moves on startup and on failure. The last startup line in that file is from August 11. The process itself came back on September 18, and at the September 22 check, the banner emitted by that restart had not reached the log file on disk.&lt;/p&gt;

&lt;h2&gt;
  
  
  The better signal is also not proof
&lt;/h2&gt;

&lt;p&gt;Here is where I have to stop myself a second time, because "check the state file instead" is the obvious lesson and it is not quite right either.&lt;/p&gt;

&lt;p&gt;This is the end of courtside's loop:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;                                           &lt;span class="c1"&gt;# :153
&lt;/span&gt;    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;                                              &lt;span class="c1"&gt;# :154
&lt;/span&gt;        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;upd&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;updates&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;result&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[]):&lt;/span&gt;         &lt;span class="c1"&gt;# :156  (body elided)
&lt;/span&gt;            &lt;span class="bp"&gt;...&lt;/span&gt;
            &lt;span class="n"&gt;STATE_PATH&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;  &lt;span class="c1"&gt;# :173  only after a handled message
&lt;/span&gt;    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;                          &lt;span class="c1"&gt;# :174
&lt;/span&gt;        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;COURTSIDE LISTEN error: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;type&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="n"&gt;__name__&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;STATE_PATH&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;          &lt;span class="c1"&gt;# :177  every pass, either way
&lt;/span&gt;    &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;POLL_SECONDS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                          &lt;span class="c1"&gt;# :178
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There are two writes. The one at &lt;code&gt;:173&lt;/code&gt; happens inside the &lt;code&gt;try&lt;/code&gt;, after a message is handled. The one at &lt;code&gt;:177&lt;/code&gt; is at loop level, &lt;strong&gt;outside&lt;/strong&gt; the &lt;code&gt;except&lt;/code&gt;, so it runs on every pass — including a pass that just caught an exception and slept.&lt;/p&gt;

&lt;p&gt;That second write is the one that keeps courtside's timestamp fresh, and eye is built the same way. So for those two, a fresh state mtime proves the loop is turning. It does not prove a poll succeeded, and it cannot tell a good pass from a failed one.&lt;/p&gt;

&lt;p&gt;Sentinel is not the same. It has no loop-level write — only &lt;code&gt;:140&lt;/code&gt;, inside the &lt;code&gt;try&lt;/code&gt;, after &lt;code&gt;getUpdates&lt;/code&gt; returns. So sentinel's timestamp is a stronger signal than the other two. It shows that the API call returned parseable JSON and the loop reached its write. It does not separately record API-level success, whether messages arrived, or whether a reply was delivered. An empty successful poll is normal. Three jobs I had been treating as one row are three different write sites.&lt;/p&gt;

&lt;p&gt;And the contents are not one story either. I checked all three:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;sentinel    offset 0            history 0      13 bytes
courtside   offset 51848716     history 4     545 bytes
eye         offset 63652330     history 8    2988 bytes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I had looked at sentinel, seen &lt;code&gt;{"offset": 0}&lt;/code&gt;, and was about to write that all three had never consumed a message. Two of them hold a non-zero saved cursor and a non-empty history. Three jobs that look identical from &lt;code&gt;launchctl&lt;/code&gt; are in three different states, and the only reason I know that is that I opened all three instead of one.&lt;/p&gt;

&lt;p&gt;Even sentinel's zero is smaller than it looks. It is the saved cursor right now. It is not a history of everything that job has ever done.&lt;/p&gt;

&lt;h2&gt;
  
  
  What each thing actually establishes
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Observation&lt;/th&gt;
&lt;th&gt;Supports&lt;/th&gt;
&lt;th&gt;Does not establish&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;pgrep&lt;/code&gt; returns a PID&lt;/td&gt;
&lt;td&gt;a process exists at this moment&lt;/td&gt;
&lt;td&gt;that it polls or delivers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;log file unchanged&lt;/td&gt;
&lt;td&gt;the file was not modified&lt;/td&gt;
&lt;td&gt;that no exception occurred&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;state mtime advances — courtside, eye&lt;/td&gt;
&lt;td&gt;the loop reached its unconditional write at &lt;code&gt;:177&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;that a poll succeeded; it writes after a caught error too&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;state mtime advances — sentinel&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;getUpdates&lt;/code&gt; returned and parsed, reaching &lt;code&gt;:140&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;API-level success, that anything arrived, or delivery&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;saved offset / history&lt;/td&gt;
&lt;td&gt;the file currently holds those values&lt;/td&gt;
&lt;td&gt;lifetime behavior, or delivery&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;em&gt;a successful-poll receipt&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;I do not emit one&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;em&gt;a delivered-result receipt&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;I do not emit one&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The last two rows are the point. Every signal I have is incidental — something the program happens to leave behind while doing its job. None of them was written to record that the job succeeded. Those are the two I would have to build, and until they exist I am inferring operation from debris.&lt;/p&gt;

&lt;p&gt;My rule was an improvement on checking the PID. It was not the end of the ladder. "Artifact age beats PID" needs a harder second question: &lt;strong&gt;which operation am I checking, what evidence does that operation emit, and under what conditions does absence of that evidence mean failure?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A newer file is not automatically a better signal. You have to know what causes it to change. A quiet listener with no incoming work needs a different expectation than a queue with messages waiting, and a timestamp cannot tell those apart on its own.&lt;/p&gt;

&lt;h2&gt;
  
  
  The buffering part, which is real and separate
&lt;/h2&gt;

&lt;p&gt;There is a second thing going on with those files, and it is worth separating from the write semantics. Two properties compound. The successful path barely writes to stdout at all, so the stream produces very little. And because stdout is block-buffered, the few lines it does produce may not reach disk promptly. Either one alone would be survivable. Together they make log silence ambiguous rather than informative.&lt;/p&gt;

&lt;p&gt;I had a live example of the second half. These processes restarted on 2026-09-18 and the startup path prints a banner at &lt;code&gt;:128&lt;/code&gt;. On September 22 the file still contained exactly one banner, with an mtime of August 11. The September restart ran the print, and the line was nowhere on disk. Either it was waiting in a buffer or it was gone. The result at the end says which.&lt;/p&gt;

&lt;p&gt;These three plists run &lt;code&gt;/usr/bin/python3 script.py&lt;/code&gt;. Four sibling jobs on the same machine run &lt;code&gt;/usr/bin/python3 -u script.py&lt;/code&gt;. When Python's stdout is non-interactive, as it is when these jobs write to a file, it is normally block-buffered. The exact size is a runtime detail rather than a guarantee: Python commonly uses the file's block size and falls back to &lt;code&gt;io.DEFAULT_BUFFER_SIZE&lt;/code&gt;, which on this interpreter reports 8192. &lt;code&gt;-u&lt;/code&gt; and &lt;code&gt;PYTHONUNBUFFERED=1&lt;/code&gt; disable that buffering for stdout and stderr.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;finance    -u   2026-09-22 19:50   52479 bytes
health     -u   2026-09-22 09:22   55759 bytes
kairos     -u   2026-09-22 09:30   55029 bytes
organizer  -u   2026-09-22 09:21   61883 bytes

sentinel   no   2026-08-11 16:52    8103 bytes
courtside  no   2026-08-11 16:52    8127 bytes
eye        no   2026-08-11 16:57    8149 bytes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On September 22, four out of four with the flag had logs that moved that day. Three out of three without it were frozen. That is a seven-job cohort on one machine, not a law of nature, and I am reporting it as what it is.&lt;/p&gt;

&lt;p&gt;Both things are true at once: those programs have almost nothing to say, &lt;strong&gt;and&lt;/strong&gt; without &lt;code&gt;-u&lt;/code&gt;, the little they do say is not guaranteed to reach disk promptly.&lt;/p&gt;

&lt;h2&gt;
  
  
  A prediction, dated before I know
&lt;/h2&gt;

&lt;p&gt;Frozen 2026-09-17. The ledger entry also carried my buffering explanation and an &lt;code&gt;ESTABLISHED&lt;/code&gt; socket condition; the socket clause was removed by an amendment on 2026-09-17 and the script-name test added on 2026-09-21. This is a summary of the observable part that remains:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Those three logs will still carry mtime 2026-08-11 on &lt;strong&gt;2026-10-01&lt;/strong&gt;, while the processes are still alive.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Wrong &lt;strong&gt;on 2026-10-01&lt;/strong&gt; if any of the three logs has an mtime on or after 2026-09-18 while its plist still lacks &lt;code&gt;-u&lt;/code&gt; and &lt;code&gt;PYTHONUNBUFFERED&lt;/code&gt;, or if any of the three listeners is absent &lt;strong&gt;when tested by script name on that date&lt;/strong&gt;. An intermediate restart is not a falsifier; the prediction is about what is true on the resolution date.&lt;/p&gt;

&lt;p&gt;Not by PID. My ledger recorded 86150 / 86156 / 86157. The machine rebooted on 2026-09-18 and they came back as 859 / 860 / 853. On September 21, &lt;code&gt;ps -p 86150,86156,86157&lt;/code&gt; returned no rows. Resolving against those numbers would have written WRONG against a prediction that was still holding at the time.&lt;/p&gt;

&lt;p&gt;That is an &lt;strong&gt;observation&lt;/strong&gt;, and it is the only part I am committing to. The buffering account above is my explanation for it, and a correct prediction would not prove the explanation. If the October 1 check fails, I will say so here.&lt;/p&gt;

&lt;p&gt;I am not adding &lt;code&gt;-u&lt;/code&gt; until then. Changing the launch configuration now would invalidate the observation window I froze for the prediction.&lt;/p&gt;

&lt;h2&gt;
  
  
  October 1: it failed
&lt;/h2&gt;

&lt;p&gt;By September 25, the logs had changed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;              frozen               Sep 29 snapshot (mtime Sep 25)   Oct 3 reading
sentinel   8103 B  2026-08-11    16276 B  2026-09-25 06:32    65095 B
courtside  8127 B  2026-08-11    16315 B  2026-09-25 20:39    65322 B
eye        8149 B  2026-08-11    16333 B  2026-09-25 21:27    57130 B
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The plists still run &lt;code&gt;/usr/bin/python3 script.py&lt;/code&gt;, no &lt;code&gt;-u&lt;/code&gt;, no &lt;code&gt;PYTHONUNBUFFERED&lt;/code&gt;, and their modification times are July 5 and July 11, before the prediction was made. So the launch configuration did not change across the window, and the September 25 log changes satisfy the frozen falsifier retrospectively, even though I failed to run the resolver on October 1. That first falsifier alone makes this &lt;strong&gt;WRONG&lt;/strong&gt;. The process condition was never observed on October 1 itself. Checked later, on October 2 and again on October 3, all three were alive by script name, the same PIDs that started 2026-09-18 06:28:01. Eye's last change was October 2 at 20:20; its Oct 3 number is a reading, not a third move.&lt;/p&gt;

&lt;p&gt;The new content is what the write sites said to expect: caught-error lines. As of October 3, sentinel's log holds 673 error lines. 594 are &lt;code&gt;URLError&lt;/code&gt;, and 459 of those are this one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SENTINEL LISTEN error: URLError: &amp;lt;urlopen error [Errno 8] nodename nor servname provided, or not known&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is a name-resolution failure. The error alone does not tell me whether the name was invalid or whether resolving a valid name failed. Between the frozen sizes and the September 29 snapshot, the files had grown by 8,173, 8,188 and 8,184 bytes. Together with the caught errors and the buffered configuration, those increases are consistent with accumulated output being flushed. I did not trace the writes, so I cannot identify the exact flush count or trigger from these snapshots, and a consistent result does not prove the explanation.&lt;/p&gt;

&lt;p&gt;The bet was the part that lost. I bet the logs would sit still until October 1. Enough new error output accumulated that the logs had moved before October 1.&lt;/p&gt;

&lt;p&gt;The banner question has an answer too. The sentinel log now has two startup banners. The second is line 82, directly after the 81 lines that were there on August 11, and the only startup these processes have had since August 11 was September 18. That banner is the September 18 startup line. &lt;strong&gt;It was not lost. It eventually reached disk.&lt;/strong&gt; That is consistent with the buffering explanation above. I had a second dated prediction that said the opposite, that the restart's output was destroyed and the file would still be 8103 bytes with one banner on October 5. As of October 3 the file is 65,095 bytes with two banners. It is already not that file, so Monday can only confirm a result that is already visible, unless the file shrinks. Under my own rule it stays pending until then.&lt;/p&gt;

&lt;p&gt;One more thing the date did not do. Nobody ran the check on October 1. The resolver is a script. An interim run on September 29 had already reported failing criteria, but no run was recorded on October 1, and the first recorded run after the deadline was October 2 at 23:35 UTC. Writing the falsifier down before the result stops me from moving it afterward. It does not make anyone show up on the day. If you bet on something, put the check on a scheduler, not on your memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to run on yours
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 1. does the process exist&lt;/span&gt;
pgrep &lt;span class="nt"&gt;-fl&lt;/span&gt; your_job.py

&lt;span class="c"&gt;# 2. what does the log say, and WHEN does your program write to it&lt;/span&gt;
&lt;span class="nb"&gt;stat&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; &lt;span class="s1"&gt;'%Sm %z %N'&lt;/span&gt; &lt;span class="nt"&gt;-t&lt;/span&gt; &lt;span class="s1"&gt;'%Y-%m-%d'&lt;/span&gt; /path/to/your_job.log
&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-nE&lt;/span&gt; &lt;span class="s1"&gt;'print\(|log\.|logger\.'&lt;/span&gt; /path/to/your_job.py

&lt;span class="c"&gt;# 3. -u as its own argument, not a substring match&lt;/span&gt;
plutil &lt;span class="nt"&gt;-extract&lt;/span&gt; ProgramArguments json &lt;span class="nt"&gt;-o&lt;/span&gt; - ~/Library/LaunchAgents/com.you.yourjob.plist

&lt;span class="c"&gt;# 4. if you made a dated bet about it, schedule the check&lt;/span&gt;
&lt;span class="c"&gt;#    (a falsifier nobody runs on the day is a note, not a test)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These are macOS commands. On Linux, &lt;code&gt;stat -c '%y %s %n'&lt;/code&gt; and your service manager's unit file do the same jobs.&lt;/p&gt;

&lt;p&gt;Step 2 is the one I skipped. Read the write sites before you interpret the file. If the only &lt;code&gt;print()&lt;/code&gt; is inside an &lt;code&gt;except&lt;/code&gt; block, an old timestamp is not evidence of failure. It is also not evidence of health. It is evidence the file did not change.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this does not establish
&lt;/h2&gt;

&lt;p&gt;It does not prove these three jobs are healthy. On September 22, state mtimes showed courtside and eye reaching their loop-level write, sentinel getting a parseable response from &lt;code&gt;getUpdates&lt;/code&gt;, and two of them holding a non-zero cursor. The October error logs do not re-prove that. What still holds is that I do not have a signal tied to a poll actually succeeding. That is the thing I have to go build.&lt;/p&gt;

&lt;p&gt;It does not prove every quiet daemon is fine. It proves that for these three, on this machine, the file I chose could not answer the question I was asking it.&lt;/p&gt;

&lt;p&gt;It does not prove &lt;code&gt;-u&lt;/code&gt; is the only fix. &lt;code&gt;PYTHONUNBUFFERED=1&lt;/code&gt; would do the same thing.&lt;/p&gt;

&lt;p&gt;And it does not prove the buffering story. The October result fits it, including three aggregate growth deltas clustered around 8 KiB, but fitting is not proof. The prediction tested the observation, and the observation was wrong.&lt;/p&gt;

&lt;p&gt;The rule needed replacing, not defending. "Artifact age beats PID" got me off the PID and onto a file with no per-poll success record in it. The version I can actually use is longer and worse as a slogan: &lt;strong&gt;name the operation, name the evidence it emits, and name the conditions under which its absence is a failure.&lt;/strong&gt; If I cannot write that third part down, I do not have a health check. I have a file I like looking at.&lt;/p&gt;

</description>
      <category>python</category>
      <category>devops</category>
      <category>launchd</category>
      <category>opensource</category>
    </item>
    <item>
      <title>The Data Was Public. The Agent Path Wasn't. So His Mock Became My Documentation.</title>
      <dc:creator>Self-Correcting Systems</dc:creator>
      <pubDate>Wed, 30 Sep 2026 04:36:02 +0000</pubDate>
      <link>https://dev.to/kenielzep97/the-data-was-public-the-agent-path-wasnt-so-his-mock-became-my-documentation-413a</link>
      <guid>https://dev.to/kenielzep97/the-data-was-public-the-agent-path-wasnt-so-his-mock-became-my-documentation-413a</guid>
      <description>&lt;p&gt;Two commands against the same data. Run them yourself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="nt"&gt;--get&lt;/span&gt; &lt;span class="s1"&gt;'https://u58x3mt0.api.sanity.io/v2025-08-15/data/query/production'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
       &lt;span class="nt"&gt;--data-urlencode&lt;/span&gt; &lt;span class="s1"&gt;'query=count(*)'&lt;/span&gt;

&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"query"&lt;/span&gt;:&lt;span class="s2"&gt;"count(*)"&lt;/span&gt;,&lt;span class="s2"&gt;"result"&lt;/span&gt;:120,&lt;span class="s2"&gt;"syncTags"&lt;/span&gt;:[&lt;span class="s2"&gt;"s1:dmd/mg"&lt;/span&gt;&lt;span class="o"&gt;]&lt;/span&gt;,&lt;span class="s2"&gt;"ms"&lt;/span&gt;:11&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;120 documents, no key, no account. Now the path my agent actually uses:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; &lt;span class="s1"&gt;'%{http_code}\n'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="s1"&gt;'https://api.sanity.io/v1/context/organizations/&amp;lt;ORG&amp;gt;/mcp/self-correcting-systems'&lt;/span&gt;

401
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same underlying dataset. One route is publicly queryable. The route my agent takes goes through&lt;br&gt;
an authenticated Context MCP endpoint, which needs an organization API token — not a project&lt;br&gt;
token, which Sanity explicitly does not accept there. I had not provisioned Pouya any access to my&lt;br&gt;
organization, so cloning the repo gave him no way to authenticate the path the agent actually uses.&lt;/p&gt;

&lt;p&gt;I am being careful about "same" here: an MCP endpoint can be configured with its own sources and a&lt;br&gt;
&lt;code&gt;groqFilter&lt;/code&gt;, so it is not guaranteed to expose the same 120 documents the public route does. Same&lt;br&gt;
source, different access path, different interface contract. That distinction turns out to be the&lt;br&gt;
bug.&lt;/p&gt;

&lt;p&gt;That gap is the whole story, and I did not know it was there until the contributor who fell into&lt;br&gt;
it told me.&lt;/p&gt;
&lt;h2&gt;
  
  
  What happened
&lt;/h2&gt;

&lt;p&gt;A developer named Pouya read a post of mine, decided my diagnosis was wrong, cloned the repo and&lt;br&gt;
opened a pull request. First outside contribution the project has ever had.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PR #1   https://github.com/keniel13-ui/ask-the-record/pull/1
        PouyaZX4:fix/groq-schema-routing
opened  2026-09-22T12:33:52Z
merged  2026-09-24T00:53:58Z   -&amp;gt;  36.3 hours
        3 commits, 1 file, +25 / -12, merge 4a2940f
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It is merged, so it will not appear in the default pull request list, which shows open ones only.&lt;br&gt;
A reviewer told me the repo had no PRs at all for exactly that reason, which is a small instance of&lt;br&gt;
the same mistake this whole post is about.&lt;/p&gt;

&lt;p&gt;His diagnosis was reasonable. My agent queries a Sanity dataset with GROQ, and he thought the&lt;br&gt;
failures came from the model guessing at document types and inventing query syntax. His fix: put&lt;br&gt;
the real schema and worked examples into the tool description, so the model is told the shape&lt;br&gt;
instead of guessing it.&lt;/p&gt;

&lt;p&gt;Sound idea. I would have tried the same thing.&lt;/p&gt;

&lt;p&gt;Then I ran his examples.&lt;/p&gt;
&lt;h2&gt;
  
  
  The two queries
&lt;/h2&gt;

&lt;p&gt;Both from the pull request description. Here is the one for a claim, against the live dataset:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;*[_type == "claim" &amp;amp;&amp;amp; _id == "claim-ledger-population"][0].{status, expiryStatus}

-&amp;gt; HTTP 400  "attribute expected"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A stray dot between &lt;code&gt;[0]&lt;/code&gt; and &lt;code&gt;{&lt;/code&gt;. Remove it and it works:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;*[_type == "claim" &amp;amp;&amp;amp; _id == "claim-ledger-population"][0]{status, expiryStatus}

-&amp;gt; {"status": "standing", "expiryStatus": "no_expiry_set"}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the one for a patch record:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xpath"&gt;&lt;code&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nt"&gt;_type&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;==&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;"patch"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;&amp;amp;&amp;amp;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;findingRef._ref&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;==&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;"finding-b1"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="m"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;null&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Null for two separate reasons. There is no &lt;code&gt;findingRef&lt;/code&gt; field on my patch documents — it is&lt;br&gt;
&lt;code&gt;findings&lt;/code&gt;, and it is an array. And the ids are case sensitive: &lt;code&gt;finding-B1&lt;/code&gt;, not &lt;code&gt;finding-b1&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Written correctly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xpath"&gt;&lt;code&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nt"&gt;_type&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;==&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;"patch"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;&amp;amp;&amp;amp;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;"finding-B1"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;in&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;findings&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nt"&gt;_ref&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="m"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]{&lt;/span&gt;&lt;span class="nt"&gt;sha&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;inMain&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s"&gt;"sha"&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;"dd1a654"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;"inMain"&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;false&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Six field names that do not exist
&lt;/h2&gt;

&lt;p&gt;The patch did not just carry two broken examples. It wrote a schema into the tool description,&lt;br&gt;
and the schema did not match my data:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;the patch told the model&lt;/th&gt;
&lt;th&gt;what the database actually has&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;claim.subject&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;nothing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;claim.statement&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;text&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;finding.finders&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;foundBy&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;finding.verifiedReceipts&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;commentId&lt;/code&gt; or &lt;code&gt;commentIds&lt;/code&gt;, depending on the record&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;patch.findingRef&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;findings[]&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;patch.commitHash&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;sha&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Check it yourself, no key required:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="nt"&gt;--get&lt;/span&gt; &lt;span class="s1"&gt;'https://u58x3mt0.api.sanity.io/v2025-08-15/data/query/production'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--data-urlencode&lt;/span&gt; &lt;span class="s1"&gt;'query=*[_id=="finding-B1"][0]'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the record his example referenced. An earlier draft of this post ran the same command&lt;br&gt;
against &lt;code&gt;finding-B8&lt;/code&gt; instead, because B8 carries &lt;code&gt;commentIds&lt;/code&gt; and B1 carries the singular&lt;br&gt;
&lt;code&gt;commentId&lt;/code&gt;, which made my table read cleaner. That is picking the sample that proves the point,&lt;br&gt;
in a post about not doing that, and a reviewer caught it before it shipped.&lt;/p&gt;

&lt;p&gt;So a patch written to stop a model from inventing schema would have installed a mismatched schema&lt;br&gt;
into the exact place the model reads it from. Worse than the bug it fixed, because a guess made at&lt;br&gt;
runtime is a guess, while a guess sitting in the tool description arrives with the authority of&lt;br&gt;
documentation.&lt;/p&gt;
&lt;h2&gt;
  
  
  Why the names were wrong, which is not what I assumed
&lt;/h2&gt;

&lt;p&gt;I wrote a draft of this post that said I did not know how the patch was produced and was not going&lt;br&gt;
to guess. That was the right call, because when I asked, the answer was nothing I would have&lt;br&gt;
guessed. In his words, published with his permission:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The dataset itself is public of course, but I wasn't able to replicate the full MCP setup on my&lt;br&gt;
end at the time I tested it so I was hitting that 401 on the Context MCP endpoint. So I just&lt;br&gt;
tested the query generation logic against an offline mock schema on LM Studio instead, which is&lt;br&gt;
where those field names came from.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Nothing was hallucinated. He hit the 401 at the top of this post.&lt;/p&gt;

&lt;p&gt;He could read every one of my 120 documents in a browser. He could not run my agent, because the&lt;br&gt;
agent does not take the public route — it goes through an org-scoped MCP endpoint needing an&lt;br&gt;
organization credential I had not provided and was not going to publish with the repository.&lt;/p&gt;

&lt;p&gt;So he did the careful thing. He built an offline mock of the schema and tested the query&lt;br&gt;
generation logic against that. And when he built the mock he chose better names than mine:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;When creating that schema, I used clean, self-describing semantic names (commitHash instead of&lt;br&gt;
sha, verifiedReceipts instead of commentIds, finders instead of foundBy) because explicit naming&lt;br&gt;
makes it way easier for the SLMs to understand what fields represent and prevents confusion.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;He is right about that, by the way. &lt;code&gt;commitHash&lt;/code&gt; is a better name than &lt;code&gt;sha&lt;/code&gt;. &lt;code&gt;foundBy&lt;/code&gt; is worse&lt;br&gt;
than &lt;code&gt;finders&lt;/code&gt;. Every wrong name in that table is the name a competent person would pick if they&lt;br&gt;
were designing the schema rather than reading it.&lt;/p&gt;

&lt;p&gt;One of them is worse than a naming difference. &lt;code&gt;verifiedReceipts&lt;/code&gt; does not just rename&lt;br&gt;
&lt;code&gt;commentIds&lt;/code&gt; — it asserts something the data cannot support. A comment id says a comment exists at&lt;br&gt;
that address. It says nothing about whether anyone verified it. That distinction is the entire&lt;br&gt;
subject of the record it was describing.&lt;/p&gt;
&lt;h2&gt;
  
  
  The defect is mine
&lt;/h2&gt;

&lt;p&gt;The mock is not the bug. The mock was a reasonable response to a 401.&lt;/p&gt;

&lt;p&gt;The bug is that &lt;strong&gt;my repository advertised a reproduction path that is not the one the system&lt;br&gt;
uses.&lt;/strong&gt; The README says "Public dataset" and gives the URL at the top of this post. That is true,&lt;br&gt;
and it is what I put in an earlier version of this very post — a &lt;code&gt;curl&lt;/code&gt;, no key, check it&lt;br&gt;
yourself. It reproduces the &lt;em&gt;data&lt;/em&gt;. It does not reproduce the &lt;em&gt;agent&lt;/em&gt;, which reaches the same&lt;br&gt;
records through an endpoint that returns 401 to everyone who is not me.&lt;/p&gt;

&lt;p&gt;A contributor following my README lands in a place where every record is visible and nothing is&lt;br&gt;
runnable. The only way forward is to build a stand-in. And a stand-in built at that boundary does&lt;br&gt;
not stay at the boundary — his mock's field names travelled from a local LM Studio test into the&lt;br&gt;
tool description, which is where my harness explicitly tells the model what the schema is.&lt;/p&gt;

&lt;p&gt;That is the whole difference in one line. A runtime guess is visibly a guess. Put the same guess in&lt;br&gt;
the tool description and I have promoted it into authoritative guidance.&lt;/p&gt;

&lt;p&gt;I have shipped the same class of error. I described my own validator to someone as checking&lt;br&gt;
whether a cited record matched what was retrieved. It does not. I was describing the system I&lt;br&gt;
meant to build, from memory, instead of opening the file.&lt;/p&gt;
&lt;h2&gt;
  
  
  What actually resolved it
&lt;/h2&gt;

&lt;p&gt;Not an argument. Two commands.&lt;/p&gt;

&lt;p&gt;I pulled his branch, ran both examples against the live dataset, and sent him exactly what came&lt;br&gt;
back: the 400, the null, and the corrected queries with their real output. No opinion about his&lt;br&gt;
approach, no debate about the diagnosis.&lt;/p&gt;

&lt;p&gt;He pushed a revision in about two hours. Every field name corrected, both queries rewritten:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;claim&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="err"&gt;claim-ledger-population&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;standing&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;/&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;no_expiry_set&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;patch&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="err"&gt;for&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;finding-B&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="w"&gt;         &lt;/span&gt;&lt;span class="err"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;dd&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="err"&gt;a&lt;/span&gt;&lt;span class="mi"&gt;654&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;/&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;finding&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;finding-B&lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="w"&gt;            &lt;/span&gt;&lt;span class="err"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;commentIds&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="err"&gt;'&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="err"&gt;ee&lt;/span&gt;&lt;span class="mi"&gt;98&lt;/span&gt;&lt;span class="err"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="err"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;status&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;unbuilt&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All three still return exactly that today. Note the third is &lt;strong&gt;B8&lt;/strong&gt;, not the B1 I told you to curl&lt;br&gt;
above — B1 carries the singular &lt;code&gt;commentId&lt;/code&gt; &lt;code&gt;"3eanf"&lt;/code&gt; and status &lt;code&gt;implemented&lt;/code&gt;. Those are two&lt;br&gt;
different records and I have mixed them up once already while writing this.&lt;/p&gt;

&lt;p&gt;Still one stale reference in a routing rule, so I sent that too, with the null it produced. He&lt;br&gt;
fixed it that morning. I re-ran the three patterns, checked the file still parsed and still had no&lt;br&gt;
third-party imports, and merged it.&lt;/p&gt;

&lt;p&gt;Three commits, thirty six hours, between two people who have never met. And one detail I like:&lt;br&gt;
the pull request body still shows &lt;code&gt;[0].{status, expiryStatus}&lt;/code&gt;, the broken stray-dot form, after&lt;br&gt;
the merged code had moved to the corrected one. Documentation can keep a false schema alive after&lt;br&gt;
the executable stops using it, which is the same failure as this whole post, one layer up.&lt;/p&gt;
&lt;h2&gt;
  
  
  The thing worth taking
&lt;/h2&gt;

&lt;p&gt;If you maintain anything an outsider might contribute to, run this check:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can someone who clones your repo actually execute the path your system takes, or only the path&lt;br&gt;
your README documents?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For me those were different, and the gap was invisible from the inside because I hold the token.&lt;br&gt;
Everything worked on my machine for a reason I never had to think about.&lt;/p&gt;

&lt;p&gt;When they differ, a contributor's only option is a mock. They will build a good one — Pouya's&lt;br&gt;
names were better than mine. And then the mock's assumptions become your documentation, because&lt;br&gt;
the tool description is documentation, and the model does not know it was written against a&lt;br&gt;
fixture.&lt;/p&gt;

&lt;p&gt;There is a third option, and it is better than either of the two I first wrote down. I said the&lt;br&gt;
fix was to make the real path reachable, or to say in the README that it is not. Both are weak,&lt;br&gt;
because both still leave the contributor inventing the contract.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If outsiders cannot execute a privileged dependency, ship them a reproducible contract for it.&lt;/strong&gt;&lt;br&gt;
A fixture generated from the real schema, checked into the repo, lets someone test against the same&lt;br&gt;
interface without ever receiving my organization credential. The pipeline becomes&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;production schema -&amp;gt; generated contract fixture -&amp;gt; contributor harness
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;instead of&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;README -&amp;gt; unreachable MCP -&amp;gt; contributor invents a substitute schema
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The problem was never that Pouya mocked the boundary. It is that my repository gave him no&lt;br&gt;
canonical boundary to mock, so he had to design one — and a well-designed guess is still a guess.&lt;/p&gt;

&lt;p&gt;That generalises past Sanity and past agents. Anything an outsider cannot run — a private API, a&lt;br&gt;
payment sandbox, an internal queue, an OAuth service — has this shape. If you do not own the&lt;br&gt;
stand-in, your contributors will build one, and theirs will encode their assumptions instead of&lt;br&gt;
yours.&lt;/p&gt;




&lt;p&gt;Thanks to Pouya for the patch, for taking two rounds of corrections without once arguing the&lt;br&gt;
diagnosis, and for answering the question about where those names came from when he could easily&lt;br&gt;
have let me publish a guess instead. The merge is &lt;code&gt;4a2940f&lt;/code&gt;. Every query above can be run by&lt;br&gt;
anyone against the public dataset — which, as it turns out, is exactly the point.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>testing</category>
      <category>python</category>
    </item>
    <item>
      <title>I Replaced a Gate That Accepted Everyone With a Gate That Accepted No One. My Tests Couldn't Tell the Difference.</title>
      <dc:creator>Self-Correcting Systems</dc:creator>
      <pubDate>Mon, 28 Sep 2026 13:58:33 +0000</pubDate>
      <link>https://dev.to/kenielzep97/i-replaced-a-gate-that-accepted-everyone-with-a-gate-that-accepted-no-one-my-tests-couldnt-tell-2n37</link>
      <guid>https://dev.to/kenielzep97/i-replaced-a-gate-that-accepted-everyone-with-a-gate-that-accepted-no-one-my-tests-couldnt-tell-2n37</guid>
      <description>&lt;p&gt;Run this in a terminal, then run it again under &lt;code&gt;script(1)&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s1"&gt;'import sys; print(sys.stdin.isatty(), sys.stdout.isatty())'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If your code decides whether a human is present by calling &lt;code&gt;isatty()&lt;/code&gt;, and then reads that&lt;br&gt;
human's confirmation from &lt;code&gt;/dev/tty&lt;/code&gt;, those are two different questions, and there are&lt;br&gt;
processes that answer them differently. This is the story of finding that out three times in&lt;br&gt;
one file, each time because the previous fix was wrong in a way my tests could not see.&lt;/p&gt;

&lt;p&gt;Here is the probe. Standard library only, writes no files and mutates no state; full source at the end.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;script &lt;span class="nt"&gt;-q&lt;/span&gt; /dev/null python3 gate_probe.py &lt;span class="nt"&gt;--redirected&lt;/span&gt;
&lt;span class="go"&gt;
  isatty(stdin) and isatty(stdout) : False
  /dev/tty openable                : True
  detail                           : opened
&lt;/span&gt;&lt;span class="gp"&gt;  open("/dev/tty", "r+") -&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;UnsupportedOperation: File or stream is not seekable. &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;errno&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;None&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="go"&gt;
  THE TWO PREDICATES DISAGREE for this process.
  A gate on isatty and a confirmation read on /dev/tty will not agree here.
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That process has an openable controlling terminal even though its redirected stdin and stdout&lt;br&gt;
are not TTYs. For three days that was the shape of my gate.&lt;/p&gt;

&lt;p&gt;One precision, because a reader will check. Default &lt;code&gt;pytest&lt;/code&gt; capture replaces &lt;strong&gt;both&lt;/strong&gt; streams,&lt;br&gt;
which I had wrong in an earlier version of this paragraph. On pytest 9.0.2, launched from a real&lt;br&gt;
terminal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;default capture   stdin=DontReadFromInput (isatty False)   stdout=EncodedFile (isatty False)
pytest -s         stdin=TextIOWrapper     (isatty True)    stdout=TextIOWrapper (isatty True)

/dev/tty under default capture: OPENS
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So plain &lt;code&gt;pytest&lt;/code&gt; from a terminal &lt;strong&gt;is&lt;/strong&gt; the disagreement, on its own — no wrapper needed. My gate&lt;br&gt;
required &lt;code&gt;stdin.isatty() and stdout.isatty()&lt;/code&gt;, false under default capture for two reasons at&lt;br&gt;
once, while the confirmation path could open the terminal the whole time. Under &lt;code&gt;pytest -s&lt;/code&gt; the&lt;br&gt;
two predicates agree, so don't look for a mismatch there.&lt;/p&gt;
&lt;h2&gt;
  
  
  The system, briefly
&lt;/h2&gt;

&lt;p&gt;An agent that can run three frozen commands against an exported copy of a repository. It has&lt;br&gt;
never been authorized for normal execution — the policy file forbids it — and I temporarily&lt;br&gt;
flipped that flag during the test run described below, which is how the rest of this post&lt;br&gt;
exists.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"import sys; sys.path.insert(0,'.')
&lt;/span&gt;&lt;span class="gp"&gt;import stage2a_preflight as PF;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;print(PF.check_policy())"&lt;/span&gt;
&lt;span class="go"&gt;
&lt;/span&gt;&lt;span class="gp"&gt;(False, 'policy forbids execute_shell;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;owner has not authorized Stage 2a&lt;span class="s1"&gt;')
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Before it runs anything, a human is supposed to be shown what will happen and type two&lt;br&gt;
identifiers back.&lt;/p&gt;
&lt;h2&gt;
  
  
  v0 — the gate was a parameter
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;job&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;approval&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;typed_job_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;typed_procedure_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;note&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;confirm_approval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;approval&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;typed_job_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;typed_procedure_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;REFUSED_APPROVAL_MISSING&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;note&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Read it and it looks like a human gate. &lt;code&gt;typed_job_id&lt;/code&gt; is a parameter, and a parameter is&lt;br&gt;
something any caller supplies. A test fixture is a caller with two correct strings.&lt;/p&gt;

&lt;p&gt;I flipped the policy flag that permits execution and ran the suite. Afterward:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;ls&lt;/span&gt; &lt;span class="nt"&gt;-lT&lt;/span&gt; state/attempts | &lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="s1"&gt;'NR&amp;gt;1 {print $6, $7, $8, $10}'&lt;/span&gt;
&lt;span class="go"&gt;Sep 24 16:36:41 TEST-JOB-1-1790282200
Sep 24 16:36:42 TEST-JOB-1-1790282202
Sep 24 16:36:44 TEST-JOB-1-1790282204
Sep 24 16:36:45 TEST-JOB-1-1790282205
Sep 24 16:37:42 TEST-JOB-1-1790282262
Sep 24 16:37:43 TEST-JOB-1-1790282263
Sep 24 16:37:45 TEST-JOB-1-1790282265
Sep 24 16:37:46 TEST-JOB-1-1790282266
Sep 24 16:36:43 TEST-JOB-2-1790282203
Sep 24 16:37:44 TEST-JOB-2-1790282264
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Sorted by name, not time. Read the clock column and it is &lt;strong&gt;two runs of five, about a minute&lt;br&gt;
apart&lt;/strong&gt; — 16:36:41–45 and 16:37:42–46 — eight for one fixture job and two for a second. Each&lt;br&gt;
directory holds a 266,240-byte tar of my own repository, all ten identical in size, extracted&lt;br&gt;
into a 27-file tree.&lt;/p&gt;

&lt;p&gt;The gate opened exactly as written.&lt;/p&gt;
&lt;h3&gt;
  
  
  What those ten do and do not prove
&lt;/h3&gt;

&lt;p&gt;They contain a third entry, &lt;code&gt;runtime-tmp&lt;/code&gt;, and it is empty. Its existence is not proof the&lt;br&gt;
three commands started, because of where it is created:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;frozen_env&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;attempt_root&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;addendum v1 B1. An allowlist, not a filtered copy. HOME absent. No PYTHON*.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;tmp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;attempt_root&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;runtime-tmp&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;tmp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;mkdir&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;parents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;exist_ok&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;run_frozen_commands&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;export_root&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;attempt_root&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;env&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;frozen_env&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;attempt_root&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;entry&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;PF&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;FROZEN_COMMANDS&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;frozen_env()&lt;/code&gt; is the first statement of &lt;code&gt;run_frozen_commands()&lt;/code&gt;, and the &lt;code&gt;for&lt;/code&gt; loop below it is&lt;br&gt;
where commands actually launch. So the directory is created before anything runs. So &lt;code&gt;runtime-tmp&lt;/code&gt; proves the runner was &lt;strong&gt;entered&lt;/strong&gt;. Whether any command&lt;br&gt;
started is not established by these directories, and I am not going to round that up.&lt;/p&gt;

&lt;p&gt;What the ledger says is stranger. The ten exports left &lt;strong&gt;no receipt at all&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"
&lt;/span&gt;&lt;span class="go"&gt;import json
for l in open('state/receipts.jsonl'):
    r = json.loads(l)
    print(r['final_state'], r['job_id'], len(r['commands'] or []), r['written_at'])"

REFUSED_NO_HUMAN_PRESENT tg_912616161 0 2026-09-24T21:29:28.666303+00:00
REFUSED_NO_HUMAN_PRESENT tg_912616161 0 2026-09-24T21:30:40.369098+00:00
REFUSED_NO_HUMAN_PRESENT TEST-JOB-1 0 2026-09-28T01:31:55.420602+00:00
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three rows, zero commands. &lt;strong&gt;None of them is the export run.&lt;/strong&gt; Rows 1 and 2 land fifty-three&lt;br&gt;
minutes after it; row 3 is from this repair session, below.&lt;/p&gt;

&lt;p&gt;There is also a quarantine file from that day, &lt;code&gt;receipts_TEST_POLLUTION_QUARANTINED_2026-09-24&lt;/code&gt;,&lt;br&gt;
and the export run is not in that either — its ten rows are all &lt;code&gt;REFUSED_COMMAND_BOUNDARY&lt;/code&gt;&lt;br&gt;
written between 15:02 and 16:06 UTC, while the exports are 20:36-20:37 UTC. A different event,&lt;br&gt;
four to five hours earlier.&lt;/p&gt;

&lt;p&gt;So: ten real exports in production state, and no receipt for them in either ledger. Where that&lt;br&gt;
receipt went I cannot establish. &lt;code&gt;execute()&lt;/code&gt; writes one in every terminal case, so either it was&lt;br&gt;
written to a redirected path and discarded with a temp directory, or something raised before the&lt;br&gt;
write. I did not preserve the test configuration from that run, and although the directory is a git&lt;br&gt;
repo, those files are not tracked in it, so absence here is absence — not evidence of a&lt;br&gt;
destination.&lt;/p&gt;

&lt;p&gt;Those rows read &lt;code&gt;REFUSED_NO_HUMAN_PRESENT&lt;/code&gt;, which is the same overclaim I rename a function for&lt;br&gt;
further down. That value is now &lt;code&gt;REFUSED_CONTROLLING_TERMINAL_UNAVAILABLE&lt;/code&gt; and the receipt schema&lt;br&gt;
went &lt;code&gt;v1&lt;/code&gt; -&amp;gt; &lt;code&gt;v2&lt;/code&gt; to say so. The three rows above were &lt;strong&gt;not&lt;/strong&gt; rewritten — an append-only ledger&lt;br&gt;
edited to match new vocabulary is not an audit trail — so the file holds both values, and&lt;br&gt;
&lt;code&gt;schema&lt;/code&gt; is what tells a reader which vocabulary a row was written under.&lt;/p&gt;
&lt;h3&gt;
  
  
  The thing that was actually protecting me
&lt;/h3&gt;

&lt;p&gt;Not the typed confirmation. A boolean in a config file that I had left set to &lt;code&gt;false&lt;/code&gt;.&lt;/p&gt;
&lt;h2&gt;
  
  
  v1 — the fix could not open a terminal
&lt;/h2&gt;

&lt;p&gt;Move the read off the parameter list and onto the controlling terminal. There is no argument&lt;br&gt;
to fill:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;read_confirmation&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;approval&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/dev/tty&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;r+&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;tty&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;      &lt;span class="c1"&gt;# &amp;lt;- this line
&lt;/span&gt;        &lt;span class="bp"&gt;...&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On this macOS machine, &lt;code&gt;open("/dev/tty", "r+")&lt;/code&gt; fails on both CPython 3.9.6 and 3.13.9. The&lt;br&gt;
update-mode I/O stack buffers through &lt;code&gt;BufferedRandom&lt;/code&gt;, which wants a seekable raw stream; this&lt;br&gt;
terminal is not one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;script &lt;span class="nt"&gt;-q&lt;/span&gt; /dev/null python3 gate_probe.py &amp;lt;/dev/null
&lt;span class="go"&gt;
  isatty(stdin) and isatty(stdout) : True
  /dev/tty openable                : True
  detail                           : opened
&lt;/span&gt;&lt;span class="gp"&gt;  open("/dev/tty", "r+") -&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;UnsupportedOperation: File or stream is not seekable. &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;errno&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;None&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The last line is the probe opening the same terminal it just opened successfully, in the mode&lt;br&gt;
I had used. Reproduced identically on CPython 3.9.6 and 3.13.9. And then the detail that turned a broken&lt;br&gt;
line into a broken control:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"import io; print(io.UnsupportedOperation.__mro__)"&lt;/span&gt;
&lt;span class="gp"&gt;(&amp;lt;class 'io.UnsupportedOperation'&amp;gt;&lt;/span&gt;, &amp;lt;class &lt;span class="s1"&gt;'OSError'&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;, &amp;lt;class &lt;span class="s1"&gt;'ValueError'&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;, ...&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;io.UnsupportedOperation&lt;/code&gt; &lt;strong&gt;subclasses &lt;code&gt;OSError&lt;/code&gt;&lt;/strong&gt;, and carries &lt;code&gt;errno=None&lt;/code&gt;. My handler was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;OSError&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;REFUSED_NO_HUMAN_PRESENT&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;no controlling terminal&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So a human sitting at a real terminal was told they were not there, and the receipt recorded&lt;br&gt;
the wrong cause. It failed closed, which is the good direction to fail — but the gate now&lt;br&gt;
refused the only caller it was built for.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;In v1, every test that exercised approval replaced &lt;code&gt;read_confirmation&lt;/code&gt; by name.&lt;/strong&gt; Not one&lt;br&gt;
exercised the real tty read. v0 had the earlier version of the same blind spot: there was no&lt;br&gt;
&lt;code&gt;read_confirmation&lt;/code&gt; to replace, because its tests supplied the ids directly as arguments.&lt;br&gt;
Different bypass, same hole — neither broken version had a human-interaction boundary under test,&lt;br&gt;
so the suite was green for both. It was never testing the gate; it was testing the bypass.&lt;/p&gt;
&lt;h2&gt;
  
  
  v2 — and the third defect, which the third test found
&lt;/h2&gt;

&lt;p&gt;Separate handles, and an &lt;code&gt;except&lt;/code&gt; narrow enough to mean something:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;CTTY_UNAVAILABLE_ERRNOS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;frozenset&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="n"&gt;errno&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ENXIO&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;errno&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ENODEV&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;errno&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ENOTTY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;errno&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ENOENT&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;open_controlling_terminal&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/dev/tty&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;w&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;inp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/dev/tty&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;BaseException&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;close&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;inp&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then a test I wrote for this failed, and I nearly patched the test. That would have been the&lt;br&gt;
fourth version of the same mistake. What it had found:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;require_human&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;        &lt;span class="n"&gt;asked&lt;/span&gt;  &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;is sys.stdin a tty?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="nf"&gt;read_confirmation&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;    &lt;span class="n"&gt;asked&lt;/span&gt;  &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;can I read /dev/tty?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Two different predicates for one control, and the stale one ran first.&lt;/strong&gt; They are not&lt;br&gt;
ordered — one asks whether two particular streams are terminal devices, the other asks for the&lt;br&gt;
process's controlling terminal. In the captured &lt;code&gt;pytest&lt;/code&gt; configuration I reproduced, &lt;code&gt;isatty&lt;/code&gt;&lt;br&gt;
was false while &lt;code&gt;/dev/tty&lt;/code&gt; stayed openable, so the stale check refused before the confirmation&lt;br&gt;
path could run. Fail-closed, but the real control was shadowed by a different predicate.&lt;/p&gt;

&lt;p&gt;I first wrote that this shadowing was &lt;em&gt;why&lt;/em&gt; the &lt;code&gt;r+&lt;/code&gt; version survived the suite. That is a&lt;br&gt;
wrong-reason claim in a post about wrong-reason claims, and the probe disproves it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;no controlling terminal  -&amp;gt;  OSError errno=6 (ENXIO)
terminal attached        -&amp;gt;  UnsupportedOperation (errno=None)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;r+&lt;/code&gt; fails &lt;strong&gt;two different ways&lt;/strong&gt;, and &lt;code&gt;except OSError&lt;/code&gt; swallows both. Delete the shadow and&lt;br&gt;
run again: with no terminal you get a real ENXIO, which is the verdict every refusal test&lt;br&gt;
expects; at a terminal you get &lt;code&gt;UnsupportedOperation&lt;/code&gt; relabelled into the same refusal. Either&lt;br&gt;
way, green.&lt;/p&gt;

&lt;p&gt;So the shadowing is a second independent reason nothing could have caught it, not the cause.&lt;br&gt;
The causes are the two already named: every test patched &lt;code&gt;read_confirmation&lt;/code&gt;, and the &lt;code&gt;except&lt;/code&gt;&lt;br&gt;
was wide enough to swallow a defect. What the shadow did do is stop the real path from being&lt;br&gt;
exercised by that integration route — which is why the fix needed a test that calls it&lt;br&gt;
directly.&lt;/p&gt;

&lt;p&gt;So there is now one predicate, in one function, used by both:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;require_controlling_terminal&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;inp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;open_controlling_terminal&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;OSError&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;controlling_terminal_unavailable&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;close&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;inp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;close&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;controlling terminal present and openable at /dev/tty&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And &lt;code&gt;controlling_terminal_unavailable()&lt;/code&gt; re-raises anything outside the set, so a programming&lt;br&gt;
fault can no longer be reported as an absent human. The refusal carries the evidence rather&lt;br&gt;
than the conclusion:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;controlling terminal unavailable: /dev/tty open failed with errno 6 (ENXIO): ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two renames went with this, both the same correction. &lt;code&gt;require_human()&lt;/code&gt; asserted something no&lt;br&gt;
check in that file can establish — it cannot prove a person is present, only that confirmation&lt;br&gt;
is obtainable from a terminal, and naming it after the stronger claim is what licensed a second&lt;br&gt;
predicate to grow beside it. And the errno bucket was called &lt;code&gt;NO_CTTY_ERRNOS&lt;/code&gt; while containing&lt;br&gt;
&lt;code&gt;ENOENT&lt;/code&gt; (&lt;code&gt;/dev/tty&lt;/code&gt; does not exist here) and &lt;code&gt;ENOTTY&lt;/code&gt; (that fd is not a terminal). Neither&lt;br&gt;
literally means "this process has no controlling terminal." A set named for its strongest member&lt;br&gt;
is the same overclaim one level down, so it is &lt;code&gt;CTTY_UNAVAILABLE_ERRNOS&lt;/code&gt; now. The bucket is a&lt;br&gt;
decision about what to refuse on; the errno is the fact.&lt;/p&gt;
&lt;h2&gt;
  
  
  Passing tests prove nothing, so I broke it on purpose
&lt;/h2&gt;

&lt;p&gt;Four mutations against the gate, each run as the full 131-test suite:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;mutation&lt;/th&gt;
&lt;th&gt;caught by&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;restore &lt;code&gt;open("/dev/tty", "r+")&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;4 tests&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;treat any &lt;code&gt;OSError&lt;/code&gt; as an absent human&lt;/td&gt;
&lt;td&gt;3 tests&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;restore the &lt;code&gt;isatty&lt;/code&gt; predicate&lt;/td&gt;
&lt;td&gt;9 tests and subtests&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;drop the explicit &lt;code&gt;isinstance(io.UnsupportedOperation)&lt;/code&gt; guard&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;nothing. 131 passed&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The fourth row is the useful one, and I drew the wrong conclusion from it. &lt;strong&gt;Correction, added&lt;br&gt;
2026-09-28 after this post was published.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I originally wrote that the line was redundant, since &lt;code&gt;errno&lt;/code&gt; is already &lt;code&gt;None&lt;/code&gt; for&lt;br&gt;
&lt;code&gt;UnsupportedOperation&lt;/code&gt; so the errno test re-raises it anyway, and I labelled it in the source as&lt;br&gt;
documentation rather than a control. A reader named &lt;a href="https://dev.to/howcani_howcani_77e786a89"&gt;howcani&lt;/a&gt;&lt;br&gt;
pointed out why that inference does not hold:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A mutation score is taken over mutants times fixtures. With one fixture environment,&lt;br&gt;
everything environment sensitive is unkillable by construction, and zero killed is then a&lt;br&gt;
statement about the fixture matrix rather than about the guard. &lt;strong&gt;Two zeros look identical on&lt;br&gt;
the report.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;They were right, and the cell my fixtures never sampled is one line of Python:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;python3 &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"import io; print(io.UnsupportedOperation('msg').errno, io.UnsupportedOperation(6,'x').errno)"&lt;/span&gt;
None 6
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;UnsupportedOperation&lt;/code&gt; carries &lt;code&gt;errno=None&lt;/code&gt; only when built from a message, and that is the form&lt;br&gt;
CPython raises for &lt;code&gt;open("/dev/tty","r+")&lt;/code&gt; — the single shape every test produced. Built with two&lt;br&gt;
arguments it carries an errno. Hand it &lt;code&gt;ENXIO&lt;/code&gt; and the two paths separate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;guarded    -&amp;gt; re-raised, a defect surfaces as a defect
unguarded  -&amp;gt; "controlling terminal unavailable: errno 6 (ENXIO)"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That second line is the relabel-a-defect-as-an-absent-human bug rebuilt exactly. &lt;strong&gt;So the null&lt;br&gt;
result described my fixtures rather than the code, and "a line the suite cannot distinguish from&lt;br&gt;
its own absence" was a claim about the suite that I stated as a claim about the line.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And then a second correction on top of the first, because the line has &lt;strong&gt;two answers and needs&lt;br&gt;
both stated.&lt;/strong&gt; I tried "documentation", then "load-bearing", then "defence in depth, constructible&lt;br&gt;
but unreached" — and the third was the same mistake as the first two, since it averaged two&lt;br&gt;
populations into one adjective. An independent reviewer caught that, which is the identical defect&lt;br&gt;
as two zeros looking identical on a report.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Population A — what the I/O stack actually raises.&lt;/strong&gt; The message form, &lt;code&gt;errno=None&lt;/code&gt;. Verified&lt;br&gt;
against &lt;code&gt;open("/dev/tty","r+")&lt;/code&gt;, a read on a write handle, and &lt;code&gt;seek&lt;/code&gt; and &lt;code&gt;tell&lt;/code&gt; on a pipe: all&lt;br&gt;
message form, all &lt;code&gt;None&lt;/code&gt;. The errno test below already re-raises those, so deleting the line&lt;br&gt;
changes nothing on that path. &lt;strong&gt;For A the line is documentation&lt;/strong&gt;, exactly as first written.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Population B — the inherited &lt;code&gt;OSError&lt;/code&gt; constructor.&lt;/strong&gt; &lt;code&gt;io.UnsupportedOperation(ENXIO, "x")&lt;/code&gt;&lt;br&gt;
carries &lt;code&gt;errno=6&lt;/code&gt;, and deleting the line relabels it as an absent terminal — the original bug&lt;br&gt;
rebuilt. &lt;strong&gt;For B the line is the only check.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So: redundant against A, sole guard against B, and B arrives only by construction — the test that&lt;br&gt;
kills the mutant &lt;em&gt;builds&lt;/em&gt; the exception rather than obtaining it from I/O, which is itself the&lt;br&gt;
measure of how narrow the guard is. There is no single word for that, and reaching for one is what&lt;br&gt;
produced three wrong labels in a row.&lt;/p&gt;

&lt;p&gt;There is now a test that kills that mutant, and the source comment no longer calls it&lt;br&gt;
documentation.&lt;/p&gt;

&lt;p&gt;The same reader's proposed instrument closes the general case: make the environment a fixture&lt;br&gt;
parameter instead of a constant. Five forced cells across three axes — controlling terminal&lt;br&gt;
present, fd 0 a tty, fd 1 a tty — with the gate's verdict asserted per cell:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;mutant&lt;/th&gt;
&lt;th&gt;killed by which cells&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;isatty&lt;/code&gt; on both streams&lt;/td&gt;
&lt;td&gt;neither_tty, stdin_only, stdout_only&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;isatty&lt;/code&gt; on stdin only&lt;/td&gt;
&lt;td&gt;neither_tty, &lt;strong&gt;stdout_only&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;isatty&lt;/code&gt; on stdout only&lt;/td&gt;
&lt;td&gt;neither_tty, &lt;strong&gt;stdin_only&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gate always returns True&lt;/td&gt;
&lt;td&gt;no_ctty&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The two mixed cells, where one standard stream is a terminal and the other is not, are the only&lt;br&gt;
thing separating the stdin-only mutant from the stdout-only one. Without them both die to the&lt;br&gt;
same single cell and read as the same result. That is the concrete version of two zeros looking&lt;br&gt;
identical, and my suite had neither mixed cell until this correction.&lt;/p&gt;

&lt;p&gt;Population, since it is the whole point of this post: &lt;strong&gt;5 of 131 tests run under a pty&lt;/strong&gt;, and one&lt;br&gt;
more opens &lt;code&gt;/dev/tty&lt;/code&gt; in-process, for six that exercise a terminal-backed path. A pseudo-terminal&lt;br&gt;
is deliberately not evidence of a physical terminal or a person, which is why that phrasing is not&lt;br&gt;
"six tests prove a human."&lt;/p&gt;

&lt;p&gt;The count in order, because it is the argument and a reader should be able to check the ordering:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;version&lt;/th&gt;
&lt;th&gt;what the gate was&lt;/th&gt;
&lt;th&gt;tests opening a terminal&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;v0&lt;/td&gt;
&lt;td&gt;a function parameter&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;v1&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;/dev/tty&lt;/code&gt; with &lt;code&gt;r+&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;v2&lt;/td&gt;
&lt;td&gt;separate handles&lt;/td&gt;
&lt;td&gt;2 — &lt;em&gt;written as part of this repair&lt;/em&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;v3&lt;/td&gt;
&lt;td&gt;one predicate&lt;/td&gt;
&lt;td&gt;5 under a pty, +1 in-process&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The two terminal tests did not exist while &lt;code&gt;r+&lt;/code&gt; was in the tree. They were written to fix it, and&lt;br&gt;
they are what caught it — restore &lt;code&gt;r+&lt;/code&gt; today and four tests go red, which is the first row of the&lt;br&gt;
mutation table above. Zero is the number that carried v0 and v1, and zero is why both shipped&lt;br&gt;
green.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 &lt;span class="nt"&gt;-m&lt;/span&gt; pytest &lt;span class="nt"&gt;-q&lt;/span&gt; tests/
&lt;span class="go"&gt;131 passed, 25 subtests passed
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Population, since it is the whole point of this post: &lt;strong&gt;5 of 131 tests run under a pty&lt;/strong&gt;, and one&lt;br&gt;
more opens &lt;code&gt;/dev/tty&lt;/code&gt; in-process, for six that exercise a terminal-backed path. A pseudo-terminal&lt;br&gt;
is deliberately not evidence of a physical terminal or a person, which is why that phrasing is not&lt;br&gt;
"six tests prove a human."&lt;/p&gt;

&lt;p&gt;The count in order, because it is the argument and a reader should be able to check the ordering:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;version&lt;/th&gt;
&lt;th&gt;what the gate was&lt;/th&gt;
&lt;th&gt;tests opening a terminal&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;v0&lt;/td&gt;
&lt;td&gt;a function parameter&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;v1&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;/dev/tty&lt;/code&gt; with &lt;code&gt;r+&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;v2&lt;/td&gt;
&lt;td&gt;separate handles&lt;/td&gt;
&lt;td&gt;2 — &lt;em&gt;written as part of this repair&lt;/em&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;v3&lt;/td&gt;
&lt;td&gt;one predicate&lt;/td&gt;
&lt;td&gt;5 under a pty, +1 in-process&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The two terminal tests did not exist while &lt;code&gt;r+&lt;/code&gt; was in the tree. They were written to fix it, and&lt;br&gt;
they are what caught it — restore &lt;code&gt;r+&lt;/code&gt; today and four tests go red, which is the first row of the&lt;br&gt;
mutation table above. Zero is the number that carried v0 and v1, and zero is why both shipped&lt;br&gt;
green.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 &lt;span class="nt"&gt;-m&lt;/span&gt; pytest &lt;span class="nt"&gt;-q&lt;/span&gt; tests/
&lt;span class="go"&gt;131 passed, 25 subtests passed
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The honest claim about what this technique is worth
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;sudoers(5)&lt;/code&gt;, verbatim from this machine:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;requiretty   If set, sudo will only run when the user is logged in to a real tty.
             When this flag is set, sudo can only be run from a login session and
             not via other means such as cron(8) or cgi-bin scripts.
             This flag is off by default.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Off by default, and note what the description is about: cron and CGI, not adversaries. Which&lt;br&gt;
is the correct amount of credit to give this. It is not a security boundary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;python3 &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s1"&gt;'import sys; print(sys.stdin.isatty())'&lt;/span&gt;
&lt;span class="go"&gt;False
&lt;/span&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;""&lt;/span&gt; | script &lt;span class="nt"&gt;-q&lt;/span&gt; /dev/null python3 &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s1"&gt;'import sys; print(sys.stdin.isatty())'&lt;/span&gt;
&lt;span class="go"&gt;True
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;script(1)&lt;/code&gt; allocates a pseudo-terminal. Code running with my user's permissions and access to&lt;br&gt;
PTY facilities can do that, including the assistant I write most of this code with. Its default&lt;br&gt;
tool calls fail the check — I have watched them fail, and one of my own tests is that wrapper&lt;br&gt;
passing on purpose:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_a_pty_wrapper_satisfies_the_gate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Stated as a limit, not a defect. This is what the post claims.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So the claim is narrow: &lt;strong&gt;this converts an accidental bypass into a more explicit one.&lt;/strong&gt; My&lt;br&gt;
current noninteractive automation path no longer gets past it merely by supplying function&lt;br&gt;
arguments. An environment that already controls a PTY can still satisfy the terminal path&lt;br&gt;
without writing anything new, so this is not identity and not authentication. It is worth&lt;br&gt;
having, and it is not the same as being safe.&lt;/p&gt;
&lt;h2&gt;
  
  
  The second defect, which came back in three days
&lt;/h2&gt;

&lt;p&gt;The ten exports landed in live state because my tests wrote to production paths. I fixed that&lt;br&gt;
on the 24th by redirecting three module constants in the &lt;code&gt;setUp&lt;/code&gt; of each class that needed it.&lt;/p&gt;

&lt;p&gt;Row 3 of that ledger is the 27th local — &lt;code&gt;01:31 UTC on the 28th&lt;/code&gt;, which is why the timestamp&lt;br&gt;
you read above looks like a different day. A test I wrote &lt;em&gt;during this repair&lt;/em&gt; reached&lt;br&gt;
&lt;code&gt;write_receipt()&lt;/code&gt; without a redirect and appended to the real ledger. Its &lt;code&gt;refusal_detail&lt;/code&gt; is&lt;br&gt;
the old &lt;code&gt;isatty&lt;/code&gt; message, which is how I know which version wrote it.&lt;/p&gt;

&lt;p&gt;The lesson is not "remember to redirect." Isolation was &lt;strong&gt;opt-in&lt;/strong&gt;, so correctness depended&lt;br&gt;
on every future author choosing to comply, and I was the future author who did not. It is now&lt;br&gt;
opt-out:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@pytest.fixture&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;autouse&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;isolate_agent_state&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;node&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_closest_marker&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;live_state&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;yield&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt;
    &lt;span class="p"&gt;...&lt;/span&gt;&lt;span class="n"&gt;redirect&lt;/span&gt; &lt;span class="n"&gt;RECEIPTS&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;NONCES&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ATTEMPTS&lt;/span&gt; &lt;span class="n"&gt;to&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="n"&gt;temp&lt;/span&gt; &lt;span class="nb"&gt;dir&lt;/span&gt;&lt;span class="bp"&gt;...&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A test must now declare &lt;code&gt;@pytest.mark.live_state&lt;/code&gt; to touch real state, and that declaration&lt;br&gt;
is visible in the test source. Verified load-bearing by flipping &lt;code&gt;autouse=False&lt;/code&gt;: four&lt;br&gt;
failures, from tests that exist only to prove the fixture fires.&lt;/p&gt;

&lt;p&gt;If a rule in your project is enforced by someone remembering it, it is a request, not a&lt;br&gt;
control. Mine took three days to prove that.&lt;/p&gt;
&lt;h2&gt;
  
  
  The probe
&lt;/h2&gt;

&lt;p&gt;The tool at the top. Four modes; the one that matters is &lt;code&gt;--redirected&lt;/code&gt;, which reproduces the&lt;br&gt;
disagreement:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 gate_probe.py                                    &lt;span class="c"&gt;# pipeline / CI shape&lt;/span&gt;
script &lt;span class="nt"&gt;-q&lt;/span&gt; /dev/null python3 gate_probe.py &amp;lt;/dev/null     &lt;span class="c"&gt;# terminal shape&lt;/span&gt;
script &lt;span class="nt"&gt;-q&lt;/span&gt; /dev/null python3 gate_probe.py &lt;span class="nt"&gt;--redirected&lt;/span&gt;   &lt;span class="c"&gt;# ctty intact, stdio redirected&lt;/span&gt;
script &lt;span class="nt"&gt;-q&lt;/span&gt; /dev/null python3 gate_probe.py &lt;span class="nt"&gt;--detach&lt;/span&gt;       &lt;span class="c"&gt;# no controlling terminal at all&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And a cheap smell check, labelled as what it is — a string search, not proof:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-rn&lt;/span&gt; &lt;span class="s1"&gt;'open(["'&lt;/span&gt;&lt;span class="s2"&gt;"]/dev/tty' tests/
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A bare &lt;code&gt;grep -rc '/dev/tty' tests/&lt;/code&gt; is worse than useless: it prints one count per file rather&lt;br&gt;
than a single number, and it counts docstrings. On my own suite 9 of its 15 hits are prose&lt;br&gt;
&lt;em&gt;about&lt;/em&gt; the terminal. It can only err in the flattering direction, which is the exact failure&lt;br&gt;
mode this post is about. A zero would not prove much either — a test can reach &lt;code&gt;/dev/tty&lt;/code&gt;&lt;br&gt;
through production code without the string appearing in &lt;code&gt;tests/&lt;/code&gt; at all.&lt;/p&gt;

&lt;p&gt;The evidence that actually settles it is above: a test that exercises the real path without&lt;br&gt;
patching it, plus a mutation showing the suite goes red when that path breaks.&lt;/p&gt;

&lt;p&gt;I cannot show you that number for my own broken versions. The directory &lt;em&gt;is&lt;/em&gt; a git repo with&lt;br&gt;
history, but the files in question are not in it — &lt;code&gt;git ls-files --error-unmatch&lt;br&gt;
stage2a_executor.py&lt;/code&gt; returns nothing, and the same for its tests. An earlier version of this&lt;br&gt;
sentence said "that tree is not under version control," which is wrong in a way a reader could&lt;br&gt;
catch by running &lt;code&gt;git log&lt;/code&gt; in it. What I can show is structural and needs no count: &lt;strong&gt;v1's approval tests replaced&lt;br&gt;
&lt;code&gt;read_confirmation&lt;/code&gt; by name, and v0 had no terminal boundary to exercise at all.&lt;/strong&gt; A suite that&lt;br&gt;
never opens &lt;code&gt;/dev/tty&lt;/code&gt; cannot report anything about a gate that reads from it, whatever its&lt;br&gt;
total.&lt;/p&gt;

&lt;h3&gt;
  
  
  gate_probe.py, in full
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;#!/usr/bin/env python3
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Which question is your human-in-the-loop gate actually asking?

Run it four ways and compare the two verdict columns:

    python3 gate_probe.py                                  # a pipeline / CI shape
    script -q /dev/null python3 gate_probe.py &amp;lt;/dev/null    # a terminal shape
    script -q /dev/null python3 gate_probe.py --redirected   # ctty intact, stdio redirected
    script -q /dev/null python3 gate_probe.py --detach       # no controlling terminal at all

`isatty` and `/dev/tty` are different predicates, and a gate built on the first while it
reads from the second will disagree with itself. The row that matters is the one where the
two columns differ: that is a process your gate classifies one way and your confirmation
code classifies the other.

Standard library only. Writes no files and mutates no state; it prints to stdout.
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;errno&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;io&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;sys&lt;/span&gt;

&lt;span class="c1"&gt;# Named for what the errnos establish, not for the strongest one in the set. ENOENT means
# /dev/tty does not exist here; ENOTTY means that fd is not a terminal. Neither literally
# asserts "this process has no controlling terminal."
&lt;/span&gt;&lt;span class="n"&gt;CTTY_UNAVAILABLE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;errno&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ENXIO&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;errno&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ENODEV&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;errno&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ENOTTY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;errno&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ENOENT&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;probe&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Return (isatty_verdict, dev_tty_verdict, detail).&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;stdio&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stdin&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isatty&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stdout&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isatty&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/dev/tty&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;w&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;inp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/dev/tty&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;close&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="n"&gt;inp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;close&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;stdio&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;opened&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="n"&gt;io&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;UnsupportedOperation&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="c1"&gt;# Not an unavailable terminal. Reported separately because it subclasses OSError with
&lt;/span&gt;        &lt;span class="c1"&gt;# errno None, so an `except OSError` upstream will call this "no human present."
&lt;/span&gt;        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;stdio&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;UnsupportedOperation: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; (errno=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;errno&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;) -- A BUG, NOT AN ABSENCE&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;OSError&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;kind&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;controlling terminal unavailable&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;errno&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;CTTY_UNAVAILABLE&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;UNEXPECTED&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;errno&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;errorcode&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;errno&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;stdio&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;kind&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: errno &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;errno&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; (&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;also_show_the_broken_open&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;The mode that looks correct and is not. A tty is not seekable; text update mode
    wants it to be. Reproduced on CPython 3.9.6 and 3.13.9.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/dev/tty&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;r+&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;close&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;open(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/dev/tty&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;r+&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;) -&amp;gt; opened&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;OSError&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;open(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/dev/tty&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;r+&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;) -&amp;gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;type&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="n"&gt;__name__&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; (errno=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;exc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;errno&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;)&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;me&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;abspath&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;__file__&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--detach&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;argv&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="c1"&gt;# setsid() leaves the session, so the child loses the controlling terminal. Only
&lt;/span&gt;        &lt;span class="c1"&gt;# differs from a plain run if the PARENT had one -- run it under script(1) to see it.
&lt;/span&gt;        &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;executable&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;me&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;preexec_fn&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;setsid&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                           &lt;span class="n"&gt;capture_output&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stdout&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stdout&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stderr&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--redirected&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;argv&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="c1"&gt;# The case that produces the disagreement, and the shape of `pytest` launched from
&lt;/span&gt;        &lt;span class="c1"&gt;# a terminal: stdio replaced, session intact. Run this one under script(1).
&lt;/span&gt;        &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;executable&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;me&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;stdin&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;DEVNULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                           &lt;span class="n"&gt;capture_output&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stdout&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stdout&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stderr&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;

    &lt;span class="n"&gt;stdio&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ctty&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;detail&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;probe&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;  isatty(stdin) and isatty(stdout) : &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;stdio&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;  /dev/tty openable                : &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;ctty&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;  detail                           : &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;detail&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;  &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;also_show_the_broken_open&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;stdio&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;ctty&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;  THE TWO PREDICATES DISAGREE for this process.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;  A gate on isatty and a confirmation read on /dev/tty will not agree here.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;


&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;SystemExit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The question
&lt;/h2&gt;

&lt;p&gt;Not "can a test satisfy your gate" — I had that as the ending for two drafts and my own suite&lt;br&gt;
disproves it. &lt;code&gt;test_a_pty_wrapper_satisfies_the_gate&lt;/code&gt; exists on purpose. A test that can drive the&lt;br&gt;
real boundary deliberately is what good integration coverage looks like.&lt;/p&gt;

&lt;p&gt;The question is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can ordinary automation satisfy your human-in-the-loop gate without crossing the interaction&lt;br&gt;
boundary you intended?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If it can pass merely by supplying the right values into the same function call, then the gate has&lt;br&gt;
established knowledge of those values, not human involvement. v0 established that a caller knew&lt;br&gt;
two strings. It never established anyone read them.&lt;/p&gt;

&lt;p&gt;And the follow-up that took me three versions to reach: &lt;strong&gt;does any test exercise the real&lt;br&gt;
boundary, or do they all replace it?&lt;/strong&gt; Mine all replaced it — which is how one suite certified a&lt;br&gt;
gate that accepted everyone and a gate that accepted no one, and reported both as correct.&lt;/p&gt;

&lt;p&gt;Mine held for exactly as long as the flag next to it was set to &lt;code&gt;false&lt;/code&gt;.&lt;/p&gt;

</description>
      <category>python</category>
      <category>testing</category>
      <category>security</category>
      <category>ai</category>
    </item>
    <item>
      <title>An Agent That Counts My Receipts, Not My Claims</title>
      <dc:creator>Self-Correcting Systems</dc:creator>
      <pubDate>Mon, 21 Sep 2026 14:26:01 +0000</pubDate>
      <link>https://dev.to/kenielzep97/an-agent-that-counts-my-receipts-not-my-claims-a3h</link>
      <guid>https://dev.to/kenielzep97/an-agent-that-counts-my-receipts-not-my-claims-a3h</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/sanity-2026-09-16"&gt;Sanity Challenge, Path One: Ship an Agent That Queries Real Content&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;My ledger says three outside engineers found the same defect. Ask the agent how many locatable receipts support that row and it says one.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Finding B8 has 1 locatable comment&lt;/strong&gt; (&lt;code&gt;3ee98&lt;/code&gt; by Pushpendra). Although &lt;strong&gt;three outside finders&lt;br&gt;
are recorded&lt;/strong&gt;, the receipts for Vinh Nguyen and quashudev were not found in the comment trees&lt;br&gt;
searched.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It does not turn three recorded names into three verified receipts. The dataset stores finders and&lt;br&gt;
receipts as different things, and the agent is not permitted to add them together.&lt;/p&gt;

&lt;p&gt;It does that over a public Sanity record of &lt;strong&gt;120 documents&lt;/strong&gt;: 74 articles, 14 findings, 10 people,&lt;br&gt;
3 patches, and 19 claims. The &lt;strong&gt;claim&lt;/strong&gt; documents separately store &lt;code&gt;asOf&lt;/code&gt;, &lt;code&gt;status&lt;/code&gt; and&lt;br&gt;
&lt;code&gt;expiryStatus&lt;/code&gt;, rather than asking the model to reconstruct those distinctions from prose.&lt;/p&gt;

&lt;p&gt;Three things it will tell you about its own author:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the fix for finding B1 is &lt;strong&gt;not merged into &lt;code&gt;origin/main&lt;/code&gt;&lt;/strong&gt; — it reports &lt;code&gt;inMain: false&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;of 14 findings, &lt;strong&gt;12 were raised by the 8 outside engineers&lt;/strong&gt; and 2 are internal (one raised by my coding agent,
recorded as &lt;code&gt;Ka'el&lt;/code&gt;, and one by an internal audit). &lt;strong&gt;Two are implemented; twelve are not&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;it cannot see what happened to that branch after the record was built, and says so&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://ask-the-record.vercel.app" rel="noopener noreferrer"&gt;https://ask-the-record.vercel.app&lt;/a&gt;&lt;/strong&gt; — no login, no user-supplied key, five buttons.&lt;/p&gt;

&lt;p&gt;Measured answers on this deployment ran &lt;strong&gt;13.8–33.8 seconds&lt;/strong&gt; across the five questions — the raw&lt;br&gt;
timings are in &lt;code&gt;evidence/web-latency-2026-09-21.json&lt;/code&gt;. Evidence-bearing&lt;br&gt;
answers must survive a retrieval check: a substantive verdict returned with no successful read is&lt;br&gt;
marked a contract violation. &lt;code&gt;INSUFFICIENT_EVIDENCE&lt;/code&gt; is allowed to abstain without pretending a&lt;br&gt;
source was retrieved. Retrieval or auth failures return an error instead of an answer. &lt;strong&gt;Every&lt;br&gt;
answer that renders cleanly is required to carry five fields:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ANSWER          the claim, stated plainly
SOURCES         what it read, and any source URL carried in the record
EVIDENCE DATE   the asOf recorded in the document — never today's date
VERDICT         STANDING · RETRACTED · SUPERSEDED · UNBUILT · EXPIRED
                NO_EXPIRY_SET · INSUFFICIENT_EVIDENCE
UNCERTAINTY     what this answer cannot establish. Never empty.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Press the Kubernetes button and it returns &lt;strong&gt;INSUFFICIENT_EVIDENCE&lt;/strong&gt; and no number, because the&lt;br&gt;
Knowledge Base evidence available to that question does not establish a count. It is not claiming&lt;br&gt;
the dataset could never yield one. On the graded run, the agent abstained rather than supplying a&lt;br&gt;
number; the validator does not independently establish that every abstention contains no unsupported claim.&lt;/p&gt;

&lt;p&gt;The page also checks its own output. If an answer is missing a required field, has empty&lt;br&gt;
uncertainty, or fails the citation requirement for the instrument that produced it — a URL for&lt;br&gt;
dataset answers, an entry path &lt;strong&gt;and&lt;/strong&gt; the Knowledge Base id for Knowledge Base answers — the page&lt;br&gt;
shows the violation instead of rendering the answer cleanly. It checks citation &lt;strong&gt;presence and&lt;br&gt;
syntax&lt;/strong&gt;. It does not resolve URLs, and it does not yet prove that every emitted citation belongs&lt;br&gt;
to the evidence actually retrieved.&lt;/p&gt;

&lt;p&gt;Keys live server-side. The browser never receives one.&lt;/p&gt;
&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/keniel13-ui/ask-the-record" rel="noopener noreferrer"&gt;https://github.com/keniel13-ui/ask-the-record&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;studio/     the schema — person · article · finding · patch · claim
harness/    a CLI version of the same agent, and the contract it was built against
web/        the serverless function and page behind the demo
evidence/   breaker verdicts and unedited transcripts
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Python harness and the serverless agent use only the Python standard library. The Studio uses&lt;br&gt;
Sanity's normal React and TypeScript dependencies.&lt;/p&gt;
&lt;h2&gt;
  
  
  How I Used Sanity
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The schema makes the distinctions representable; the endpoint instructions require the agent to&lt;br&gt;
keep them separate; the harness independently enforces retrieval, required output fields, citation&lt;br&gt;
syntax and named-object guards around the answer.&lt;/strong&gt; It does not yet re-verify the returned values against the&lt;br&gt;
document field by field — the instructions carry that, not the validator.&lt;/p&gt;

&lt;p&gt;One honest wrinkle in my own design: the response &lt;code&gt;VERDICT&lt;/code&gt; is a single enum that mixes claim state,&lt;br&gt;
finding state, expiry metadata and retrieval outcome. That is a presentation-level classification.&lt;br&gt;
The underlying &lt;code&gt;status&lt;/code&gt; and &lt;code&gt;expiryStatus&lt;/code&gt; stay separate in the dataset, and the graded claim lookup&lt;br&gt;
reported both separately in its answer text — but the validator does not yet guarantee that&lt;br&gt;
separation for every future answer, and a stricter contract would split the output the way the&lt;br&gt;
schema does.&lt;/p&gt;

&lt;p&gt;Three schema choices do the work:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. &lt;code&gt;status&lt;/code&gt; and &lt;code&gt;expiryStatus&lt;/code&gt; are separate.&lt;/strong&gt; A claim can be &lt;code&gt;standing&lt;/code&gt; &lt;strong&gt;and&lt;/strong&gt; &lt;code&gt;no_expiry_set&lt;/code&gt;&lt;br&gt;
at the same time. &lt;code&gt;no_expiry_set&lt;/code&gt; means nobody dated it — not "permanently true," not "expired."&lt;br&gt;
Collapsing those two facts into one is how a record quietly rots, so the schema represents them as&lt;br&gt;
separate fields instead of forcing one value to stand for both.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. A finding has &lt;code&gt;commentOn&lt;/code&gt; and &lt;code&gt;writtenUpIn&lt;/code&gt;.&lt;/strong&gt; Where someone raised a defect and where I wrote&lt;br&gt;
it up are different articles. pm25coder found B1 in the comments of one article; it was written up&lt;br&gt;
in another. One field would have glued the wrong attribution to the wrong post.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. &lt;code&gt;patch.inMain&lt;/code&gt; is a boolean.&lt;/strong&gt; Public is not merged. All three recorded patches are pushed and&lt;br&gt;
none is in &lt;code&gt;main&lt;/code&gt;. Asked about B1, the agent reports &lt;code&gt;inMain: false&lt;/code&gt; rather than inferring merge&lt;br&gt;
state from the branch being public.&lt;/p&gt;
&lt;h3&gt;
  
  
  Two instruments, because one couldn't do it
&lt;/h3&gt;

&lt;p&gt;I built this on a Knowledge Base first. It handled the four prose-routed frozen questions in the&lt;br&gt;
graded run. But I could not reliably recover a specific claim's status and expiry from the indexed&lt;br&gt;
representation, and the measurement shows why:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;all 24 Knowledge Base entries  = 190,503 characters
'claim-ledger-population'      = 1 occurrence — a label in a Sources list
its values                     = present, but not bound to that identifier
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The values existed in the indexed prose but were not bound to &lt;code&gt;claim-ledger-population&lt;/code&gt;. A prompt&lt;br&gt;
cannot make that binding reliable when retrieval did not preserve it, so I routed that question&lt;br&gt;
shape to GROQ — which is what Sanity recommends when the schema already says where to look.&lt;/p&gt;

&lt;p&gt;I did not grade my own work here. &lt;strong&gt;Independent review means separate, owner-assigned AI reviewer&lt;br&gt;
sessions that did not implement the candidate&lt;/strong&gt; — not outside human testers. The first live graded&lt;br&gt;
run they carried out exposed the wrong-object failure on exactly that question. Synthetic breaker&lt;br&gt;
controls then found more: failed retrieval could reach the model, unsupported answers could exit&lt;br&gt;
zero, parts of the output contract were not enforced, and the contracted tool-discovery step was&lt;br&gt;
missing. The &lt;code&gt;VERDICT.md&lt;/code&gt; in &lt;code&gt;evidence/&lt;/code&gt; is a&lt;br&gt;
BLOCK and it stays there.&lt;/p&gt;

&lt;p&gt;The fix was a second Context endpoint serving GROQ over the same dataset:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight cypher"&gt;&lt;code&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="ss"&gt;[&lt;/span&gt;&lt;span class="n"&gt;_id&lt;/span&gt;&lt;span class="o"&gt;==&lt;/span&gt;&lt;span class="s2"&gt;"claim-ledger-population"&lt;/span&gt;&lt;span class="ss"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="ss"&gt;]&lt;/span&gt;
&lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="s2"&gt;"standing"&lt;/span&gt; &lt;span class="err"&gt;·&lt;/span&gt; &lt;span class="n"&gt;expiryStatus&lt;/span&gt; &lt;span class="s2"&gt;"no_expiry_set"&lt;/span&gt; &lt;span class="err"&gt;·&lt;/span&gt; &lt;span class="n"&gt;asOf&lt;/span&gt; &lt;span class="s2"&gt;"2026-09-11"&lt;/span&gt;
  &lt;span class="n"&gt;sourceUrl&lt;/span&gt; &lt;span class="s2"&gt;"https://dev.to/kenielzep97/my-harness-used-one-label-for-three-different-failures-2gc3"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One document, its fields, and its stored &lt;code&gt;sourceUrl&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;So the agent routes by question shape. &lt;strong&gt;Four of the five frozen questions — B1 attribution, B1&lt;br&gt;
merge state, B8 receipts and Kubernetes — stay on the Knowledge Base. Only the exact claim-field&lt;br&gt;
lookup routes to GROQ.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The router itself is blunter than that description: any question containing a &lt;code&gt;claim-*&lt;/code&gt; token goes&lt;br&gt;
to the dataset endpoint, which is broader than the rule I specified. The breaker preserved that as an open&lt;br&gt;
remainder rather than letting me describe the abstraction as if it were the code.&lt;/p&gt;

&lt;p&gt;By default, a Sanity Context endpoint's mode is derived from its configured sources: an endpoint&lt;br&gt;
with a dataset source serves GROQ mode, and Knowledge Base sources on that endpoint are ignored. So&lt;br&gt;
two endpoints, exactly one instrument per question.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;self-correcting-systems       initial_context, knowledge_base_read
self-correcting-systems-data  initial_context, groq_query, schema_explorer, array_field_reader
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Custom instructions on each endpoint carry the rules — report receipts not claims, keep status and&lt;br&gt;
expiry apart, bound anything about mutable state to the snapshot date. They arrive through&lt;br&gt;
&lt;code&gt;initial_context&lt;/code&gt;. The harness fetches that context before letting the model read a Knowledge Base&lt;br&gt;
entry or query the dataset, so the endpoint's instructions are in context before any evidence is&lt;br&gt;
retrieved.&lt;/p&gt;
&lt;h3&gt;
  
  
  What it does not establish
&lt;/h3&gt;

&lt;p&gt;The current corrected CLI harness held all five frozen questions on &lt;strong&gt;one independent v8 graded run&lt;/strong&gt;&lt;br&gt;
with &lt;code&gt;gemini-3.6-flash&lt;/code&gt;. That is not general reliability. The earlier BLOCK used &lt;code&gt;gemini-2.5-flash&lt;/code&gt;;&lt;br&gt;
I am not claiming the same model failed and then passed.&lt;/p&gt;

&lt;p&gt;One more limitation is preserved in the breaker verdict rather than smoothed over. The graded v6&lt;br&gt;
candidate enforced the Knowledge Base citation rule incorrectly — it accepted an entry path &lt;strong&gt;or&lt;/strong&gt;&lt;br&gt;
the Knowledge Base id where the contract required &lt;strong&gt;both&lt;/strong&gt;. The five live answers happened to supply&lt;br&gt;
both, so their content held, but the guard was weaker than specified. Version 7 corrected that to&lt;br&gt;
an AND and received its own independent run. A later review found that an unparseable &lt;code&gt;VERDICT&lt;/code&gt;&lt;br&gt;
could bypass the verdict-dependent checks. Version 8 rejects those malformed values and received&lt;br&gt;
a new independent run: five of five, exit 0. Its CLI hash matches &lt;code&gt;harness/ask.py&lt;/code&gt; in the repository.&lt;br&gt;
The v6 and v7 transcripts remain evidence for their respective versions.&lt;/p&gt;

&lt;p&gt;All four graded runs — the original BLOCK, v6, v7 and v8 — are in &lt;code&gt;evidence/&lt;/code&gt;, unedited.&lt;br&gt;
No earlier transcript was replaced to conceal its result.&lt;/p&gt;
&lt;h2&gt;
  
  
  Agent Session / Evidence Trail
&lt;/h2&gt;

&lt;p&gt;I kept the failed run instead of replacing it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://github.com/keniel13-ui/ask-the-record/blob/main/evidence/harness-breaker-2026-09-20-kairos/VERDICT.md" rel="noopener noreferrer"&gt;the original BLOCK&lt;/a&gt; and its &lt;a href="https://github.com/keniel13-ui/ask-the-record/blob/main/evidence/harness-breaker-2026-09-20-kairos/live-network-transcript.txt" rel="noopener noreferrer"&gt;unedited transcript&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/keniel13-ui/ask-the-record/tree/main/evidence/harness-graded-v6-2026-09-20-aethar" rel="noopener noreferrer"&gt;v6&lt;/a&gt; — five content-holding answers, weaker citation guard&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/keniel13-ui/ask-the-record/tree/main/evidence/harness-graded-v7-2026-09-20-aethar" rel="noopener noreferrer"&gt;v7&lt;/a&gt; — the AND fix, five of five, exit 0&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/keniel13-ui/ask-the-record/tree/main/evidence/harness-graded-v8-2026-09-21-aethar" rel="noopener noreferrer"&gt;v8&lt;/a&gt; — after a review seat found that an unparseable &lt;code&gt;VERDICT&lt;/code&gt; skipped every check gated on it; five of five on the repaired file&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/keniel13-ui/ask-the-record/blob/main/evidence/web-latency-2026-09-21.json" rel="noopener noreferrer"&gt;the latency receipt&lt;/a&gt; — five anonymous live requests and their measured timings&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The historical verdicts remain unedited. &lt;code&gt;GET /api/ask&lt;/code&gt; reports the sha256 of the deployed function,&lt;br&gt;
so you can check which source is actually serving rather than trusting that a push landed.&lt;/p&gt;
&lt;h2&gt;
  
  
  Sanity Project Details
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="err"&gt;Project&lt;/span&gt; &lt;span class="err"&gt;ID&lt;/span&gt;       &lt;span class="err"&gt;u58x3mt0&lt;/span&gt;
&lt;span class="err"&gt;Dataset&lt;/span&gt;          &lt;span class="err"&gt;production&lt;/span&gt; &lt;span class="err"&gt;(public)&lt;/span&gt;
&lt;span class="err"&gt;Knowledge&lt;/span&gt; &lt;span class="err"&gt;Base&lt;/span&gt;   &lt;span class="err"&gt;kbjnxAgyAimV&lt;/span&gt;
&lt;span class="err"&gt;Studio&lt;/span&gt;           &lt;span class="err"&gt;https://self-correcting-systems.sanity.studio/&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Query the record yourself, no credentials:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;https://u58x3mt0.api.sanity.io/v2025-08-15/data/query/production?query=count(*)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;*[_type=="finding" &amp;amp;&amp;amp; status=="unbuilt"]{code, title, "who": foundBy[]-&amp;gt;handle}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No login is required for the demo or the dataset.&lt;/p&gt;




&lt;p&gt;I'd rather ship an agent that says &lt;em&gt;two of fourteen are implemented, one of three B8 receipts is&lt;br&gt;
locatable, and this record cannot see past its snapshot&lt;/em&gt; than one that turns every stored assertion&lt;br&gt;
into certainty.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>sanitychallenge</category>
      <category>sanity</category>
      <category>ai</category>
    </item>
    <item>
      <title>My Harness Used One Label for Three Different Failures.</title>
      <dc:creator>Self-Correcting Systems</dc:creator>
      <pubDate>Mon, 14 Sep 2026 02:14:53 +0000</pubDate>
      <link>https://dev.to/kenielzep97/my-harness-used-one-label-for-three-different-failures-2gc3</link>
      <guid>https://dev.to/kenielzep97/my-harness-used-one-label-for-three-different-failures-2gc3</guid>
      <description>&lt;p&gt;Three fixtures, three separate calls into the same reducer. Here is the complete&lt;br&gt;
&lt;code&gt;failure_reasons&lt;/code&gt; each one returned, unedited:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;unreadable&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;arriving&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;args&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="err"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"SANDBOX_EVENT_CARDINALITY_INVALID"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                              &lt;/span&gt;&lt;span class="s2"&gt;"EXEC_ARGUMENTS_MISMATCH"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                              &lt;/span&gt;&lt;span class="s2"&gt;"TOOL_RESPONSE_CARDINALITY_INVALID"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;usable&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;args,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;different&lt;/span&gt;&lt;span class="w"&gt;   &lt;/span&gt;&lt;span class="err"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"SANDBOX_EVENT_CARDINALITY_INVALID"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                              &lt;/span&gt;&lt;span class="s2"&gt;"EXEC_ARGUMENTS_MISMATCH"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                              &lt;/span&gt;&lt;span class="s2"&gt;"TOOL_RESPONSE_CARDINALITY_INVALID"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;our&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;comparison&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;threw&lt;/span&gt;&lt;span class="w"&gt;     &lt;/span&gt;&lt;span class="err"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"SANDBOX_EVENT_CARDINALITY_INVALID"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                              &lt;/span&gt;&lt;span class="s2"&gt;"EXEC_ARGUMENTS_MISMATCH"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                              &lt;/span&gt;&lt;span class="s2"&gt;"TOOL_RESPONSE_CARDINALITY_INVALID"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These are minimal fixtures with no sandbox event and no tool response, so the first and last&lt;br&gt;
codes fire in all three and are expected. I am printing them anyway. A post about a receipt that&lt;br&gt;
hides which party failed has no business showing you a cleaned-up receipt.&lt;/p&gt;

&lt;p&gt;The middle line is the one that matters, and across three genuinely different failures it never&lt;br&gt;
changes.&lt;/p&gt;

&lt;p&gt;Fixture one sends arguments the parser rejects.&lt;/p&gt;

&lt;p&gt;Fixture two sends a usable call that disagrees with what I froze.&lt;/p&gt;

&lt;p&gt;Fixture three sends an object &lt;strong&gt;my own canonicalizer&lt;/strong&gt; rejects, so the comparison never completes.&lt;/p&gt;

&lt;p&gt;Constructed inputs, so none of this establishes who caused a failure in production. But one name&lt;br&gt;
covers all three, and that name says &lt;em&gt;argument mismatch&lt;/em&gt; even when nothing was compared. A failure&lt;br&gt;
in the checking stage reads as a deviation in the thing being checked.&lt;/p&gt;
&lt;h2&gt;
  
  
  The reason it reads that way
&lt;/h2&gt;

&lt;p&gt;pm25coder put it in one line, in the comments of &lt;a href="https://dev.to/kenielzep97/compare-against-the-schema-they-shipped-not-the-one-you-expected-3mb8"&gt;the schema-comparison&lt;br&gt;
piece&lt;/a&gt;&lt;br&gt;
(&lt;a href="https://dev.to/pm25coder/comment/3ei6m"&gt;permalink to the comment&lt;/a&gt;):&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;EXEC_ARGUMENTS_MISMATCH&lt;/code&gt; misreads as a verdict on the model precisely because its name carries no&lt;br&gt;
subject.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Three different observations arrive under one name: arguments that were rejected, a comparison that&lt;br&gt;
completed and found a difference, and a comparison that never finished.&lt;/p&gt;

&lt;p&gt;The label is not wrong that something happened. It is silent about which of the three, and a reader&lt;br&gt;
fills that in.&lt;/p&gt;
&lt;h2&gt;
  
  
  What the code actually did
&lt;/h2&gt;

&lt;p&gt;One &lt;code&gt;try&lt;/code&gt; was wrapping three different jobs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;call&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;actual&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;parseStrictJson&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;call&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;function&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;arguments&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;      &lt;span class="c1"&gt;// can the arriving args be read&lt;/span&gt;
      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;typeof&lt;/span&gt; &lt;span class="nx"&gt;actual&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;command&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;string&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;Buffer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;byteLength&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;actual&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;command&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;utf8&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;256&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;failures&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;EXEC_COMMAND_OVERSIZE&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nf"&gt;canonicalJsonBytes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;actual&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                        &lt;span class="c1"&gt;// does the comparison work&lt;/span&gt;
                    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;equals&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;canonicalJsonBytes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;prepared&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;expectedExecArguments&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;failures&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;EXEC_ARGUMENTS_MISMATCH&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;                    &lt;span class="c1"&gt;// do they differ&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;failures&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;EXEC_ARGUMENTS_MISMATCH&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;                      &lt;span class="c1"&gt;// ...everything lands here&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;parseStrictJson&lt;/code&gt; throws when what arrived is unreadable. &lt;code&gt;canonicalJsonBytes&lt;/code&gt; throws when the&lt;br&gt;
comparison itself cannot run. Both fell into the same catch, and the catch named the arguments.&lt;/p&gt;

&lt;p&gt;Full file if you want to read around it:&lt;br&gt;
&lt;a href="https://github.com/keniel13-ui/self-correcting-integration-maintainer/blob/dd1a654/scripts/judgment/live.mjs" rel="noopener noreferrer"&gt;&lt;code&gt;scripts/judgment/live.mjs&lt;/code&gt;&lt;/a&gt;.&lt;br&gt;
The third catch site was in&lt;br&gt;
&lt;a href="https://github.com/keniel13-ui/self-correcting-integration-maintainer/blob/dd1a654/scripts/pr2/reducer.mjs" rel="noopener noreferrer"&gt;&lt;code&gt;scripts/pr2/reducer.mjs&lt;/code&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;
  
  
  The fix
&lt;/h2&gt;

&lt;p&gt;Split the parse from the compare, and give each stage its own catch:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;actual&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;actual&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;parseStrictJson&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;call&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;function&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;arguments&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;failures&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;EXEC_ARGUMENTS_INVALID&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;actual&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="kc"&gt;undefined&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;typeof&lt;/span&gt; &lt;span class="nx"&gt;actual&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;command&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;string&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;Buffer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;byteLength&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;actual&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;command&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;utf8&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;256&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;failures&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;EXEC_COMMAND_OVERSIZE&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nf"&gt;canonicalJsonBytes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;actual&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;equals&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;canonicalJsonBytes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;prepared&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;expectedExecArguments&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;failures&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;EXEC_ARGUMENTS_MISMATCH&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;failures&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;EXEC_COMPARATOR_ERROR&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same three fixtures, same unedited arrays, only the middle line moves:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;unreadable&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;arriving&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;args&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="err"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"SANDBOX_EVENT_CARDINALITY_INVALID"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                              &lt;/span&gt;&lt;span class="s2"&gt;"EXEC_ARGUMENTS_INVALID"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                              &lt;/span&gt;&lt;span class="s2"&gt;"TOOL_RESPONSE_CARDINALITY_INVALID"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;usable&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;args,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;different&lt;/span&gt;&lt;span class="w"&gt;   &lt;/span&gt;&lt;span class="err"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"SANDBOX_EVENT_CARDINALITY_INVALID"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                              &lt;/span&gt;&lt;span class="s2"&gt;"EXEC_ARGUMENTS_MISMATCH"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                              &lt;/span&gt;&lt;span class="s2"&gt;"TOOL_RESPONSE_CARDINALITY_INVALID"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;our&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;comparison&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;threw&lt;/span&gt;&lt;span class="w"&gt;     &lt;/span&gt;&lt;span class="err"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"SANDBOX_EVENT_CARDINALITY_INVALID"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                              &lt;/span&gt;&lt;span class="s2"&gt;"EXEC_COMPARATOR_ERROR"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                              &lt;/span&gt;&lt;span class="s2"&gt;"TOOL_RESPONSE_CARDINALITY_INVALID"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The change spans three source files, plus the tests:&lt;br&gt;
&lt;a href="https://github.com/keniel13-ui/self-correcting-integration-maintainer/commit/dd1a654" rel="noopener noreferrer"&gt;&lt;code&gt;dd1a654&lt;/code&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;
  
  
  The part that was not a rename
&lt;/h2&gt;

&lt;p&gt;pm25coder called it "a rename, not a redesign" — his words, verbatim, in&lt;br&gt;
&lt;a href="https://dev.to/pm25coder/comment/3eanf"&gt;3eanf&lt;/a&gt; — and from outside the repo that is exactly what it&lt;br&gt;
looks like. Inside, two things were waiting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;There are two whitelists, not one.&lt;/strong&gt; Failure reasons are filtered through an ordered array before&lt;br&gt;
they reach the output. In &lt;code&gt;live.mjs&lt;/code&gt;, a reason that is not in the array is silently dropped. Register&lt;br&gt;
the new name in the code and not in the array, and that failure stops appearing. Other failures in the same run still show, so the run&lt;br&gt;
does not go green by itself, but the one you just added becomes invisible.&lt;/p&gt;

&lt;p&gt;There are two registries for this namespace and &lt;strong&gt;they fail differently.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;live.mjs&lt;/code&gt; holds 26 codes and emits output by filtering against the list — &lt;code&gt;FAILURE_ORDER.filter(...)&lt;/code&gt;&lt;br&gt;
at lines 169 and 572, no guard. An unknown reason is silently omitted. &lt;code&gt;scripts/pr2/constants.mjs&lt;/code&gt;&lt;br&gt;
holds 20, a strict subset, and its consumer refuses to guess:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;FAILURE_ORDER&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;reason&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;TypeError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`unknown failure reason: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;reason&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So the narrower list is the strict one, and the list covering more surface is the one that fails&lt;br&gt;
quietly. That is backwards, and it is a worse defect than the naming problem this patch fixes: a&lt;br&gt;
filter that discards unrecognised codes is designed to fail invisibly.&lt;/p&gt;

&lt;p&gt;The throwing version is the correct behaviour. The silent filter should be replaced by it, and the&lt;br&gt;
two lists should be one registry. I did not do that here, because bundling a registry refactor into a&lt;br&gt;
naming fix would make both harder to review and would put a behaviour change in a commit that&lt;br&gt;
claims to be about labels. It is on the list as its own change.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And there was a third catch site&lt;/strong&gt; doing the same collapse in a different module, which I only&lt;br&gt;
found by grepping for every place that name was added rather than trusting the two I knew about.&lt;/p&gt;
&lt;h2&gt;
  
  
  The tests, including the one that is supposed to pass
&lt;/h2&gt;

&lt;p&gt;Six of them. That they pass is not the interesting part. Run them against the parent commit and five&lt;br&gt;
of the six fail. Run against the patch, all six pass. The one that passes both ways is there on purpose:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One correction to the commit message before you click it.&lt;/strong&gt; It says three of the six tests fail&lt;br&gt;
on the parent. That line is stale: it was written when the file had four tests, and two more were&lt;br&gt;
added in the amendment. The real number is five, and the ablation below is the thing to trust. I am&lt;br&gt;
not force-pushing a rewrite of a public SHA to tidy a sentence, so the contradiction stays visible&lt;br&gt;
and this paragraph is the correction.&lt;/p&gt;

&lt;p&gt;All six are in one file, if you want to run the ablation yourself:&lt;br&gt;
&lt;a href="https://github.com/keniel13-ui/self-correcting-integration-maintainer/blob/dd1a654/test/exec-comparator-error.test.mjs" rel="noopener noreferrer"&gt;&lt;code&gt;test/exec-comparator-error.test.mjs&lt;/code&gt;&lt;/a&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;arriving args unusable       -&amp;gt; EXEC_ARGUMENTS_INVALID     fails on parent
usable args, differ          -&amp;gt; EXEC_ARGUMENTS_MISMATCH    passes on both   &amp;lt;- control
comparison cannot complete   -&amp;gt; EXEC_COMPARATOR_ERROR      fails on parent
pr2 reducer, unusable args   -&amp;gt; EXEC_ARGUMENTS_INVALID     fails on parent
prepared transport, digest   -&amp;gt; EXEC_COMPARATOR_ERROR      fails on parent
pr2 reducer, digest read     -&amp;gt; EXEC_COMPARATOR_ERROR      fails on parent
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The control is doing something specific. Splitting one catch into two created two new boundaries&lt;br&gt;
that the mismatch path now has to survive. If either boundary swallowed a case it should have passed&lt;br&gt;
through, a suite that only asserted the two new names would still be green, because the case it ate&lt;br&gt;
would simply never be asserted. The control fails the moment the original path stops producing the&lt;br&gt;
original name.&lt;/p&gt;

&lt;p&gt;The comparator failures are induced, not waited for. The arriving call parses cleanly and &lt;strong&gt;my own&lt;br&gt;
expected object&lt;/strong&gt; carries a &lt;code&gt;BigInt&lt;/code&gt; the canonicalizer refuses.&lt;/p&gt;

&lt;p&gt;That refusal is deliberate, not fragile. The serializer accepts a closed set — null, boolean, safe&lt;br&gt;
integer, NFC string, array, plain object — and rejects everything else, because its output is&lt;br&gt;
hashed and a canonical form cannot have alternatives. So a &lt;code&gt;BigInt&lt;/code&gt; in an expected object is an&lt;br&gt;
invalid internal type, and the test is contrived at the type level.&lt;/p&gt;

&lt;p&gt;Be precise about what that buys: it proves the catch fires when the comparison cannot complete on&lt;br&gt;
otherwise-valid input. It does not prove a spontaneous bug in the canonicalizer, and I am not&lt;br&gt;
claiming one. The two digest-read tests induce a&lt;br&gt;
different internal failure using a throwing getter, and each one asserts the intended read was&lt;br&gt;
actually reached and that the error survives the final filter.&lt;/p&gt;
&lt;h2&gt;
  
  
  Two things I got wrong on the way, both caught by someone else
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;My first repair ran the collapse backwards.&lt;/strong&gt; I moved the whole catch to&lt;br&gt;
&lt;code&gt;EXEC_COMPARATOR_ERROR&lt;/code&gt;, which meant unreadable arriving args were now reported as &lt;em&gt;my&lt;/em&gt; machinery&lt;br&gt;
failing. Same defect, opposite direction.&lt;/p&gt;

&lt;p&gt;The reason it survived my own review is worth more than the bug. I wrote the implementation, then&lt;br&gt;
wrote a test asserting what the implementation did. An assertion written against unverified output&lt;br&gt;
cannot fail, because it was derived from the thing it is supposed to check. That workflow&lt;br&gt;
guarantees you codify your own bugs, and it produced a test that defended the error.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Then I claimed one of the three catches was dead code.&lt;/strong&gt; I had constructed a bad expected object,&lt;br&gt;
watched it get caught by an earlier gate, and concluded nothing could reach that catch. One input&lt;br&gt;
class, generalized to all inputs. It is reachable — a failing read on the manifest digest lands&lt;br&gt;
there, because that read sits inside the &lt;code&gt;try&lt;/code&gt; while the one I tested sits outside it.&lt;/p&gt;

&lt;p&gt;Neither was caught by me. The backwards repair was caught before commit. The dead-code claim went&lt;br&gt;
into the first commit and came out in an amendment, so it was caught before push.&lt;/p&gt;
&lt;h2&gt;
  
  
  What this does not do
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;It does not put the contract in the name.&lt;/strong&gt; pm25coder's fuller point was that the label should&lt;br&gt;
say what it was compared against, something closer to &lt;code&gt;args_mismatch_under_contract=&amp;lt;id&amp;gt;&lt;/code&gt;, so the&lt;br&gt;
name states the authority rather than leaving a reader to infer a subject. That is not built. The&lt;br&gt;
names are separated. They still carry no contract id.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;EXEC_ARGUMENTS_INVALID&lt;/code&gt; does not identify who produced bad arguments.&lt;/strong&gt; It marks a boundary.&lt;/p&gt;

&lt;p&gt;The arriving arguments were rejected, either by the parser or, on the PR2 path, by argument&lt;br&gt;
validation. Valid JSON can still fail that validation. Who produced them is not established by&lt;br&gt;
that result. The model, the relay, or the transport are all still live possibilities, and the name&lt;br&gt;
stops where the evidence stops.&lt;/p&gt;

&lt;p&gt;Naming a producer there would be the same defect with a friendlier label.&lt;/p&gt;

&lt;p&gt;That neutrality costs something real, and it is fair to say so. An operator wants to know if the&lt;br&gt;
model is emitting garbage, and unparseable JSON on a structured tool call may well be exactly that.&lt;br&gt;
I have not measured how often it is, so I am not putting a frequency on it.&lt;br&gt;
The code declines to tell them, which is semantically clean and operationally thinner. The fix is not to guess in the name. It is to preserve enough provenance that attribution can be&lt;br&gt;
made afterward: what arguments were observed at the model-event boundary, whether any relay or&lt;br&gt;
transport transformation happened in between, and which contract governed the call.&lt;/p&gt;

&lt;p&gt;The contract-id work above would settle the authority half of that. It would not identify who&lt;br&gt;
corrupted an invalid argument stream — a contract id says which expectation was in force, not where&lt;br&gt;
bad bytes came from. I conflated those two in an earlier draft. Neither half is built.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;EXEC_COMPARATOR_ERROR&lt;/code&gt; names the stage, not the cause.&lt;/strong&gt; An un-canonicalizable expectation and a&lt;br&gt;
genuine comparator bug are both on my side and are not distinguishable from the outside. I declined&lt;br&gt;
to split them, because inventing a distinction the code cannot detect is the defect I was fixing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Update, 2026-09-15 — the paragraph above is wrong and I am leaving it there.&lt;/strong&gt; Vinh Nguyen showed&lt;br&gt;
the distinction is detectable after all: canonicalize the expected side on its own and the code knows&lt;br&gt;
which operand failed, so a bad frozen fixture of mine now reports &lt;code&gt;EXEC_EXPECTATION_INVALID&lt;/code&gt; rather&lt;br&gt;
than my comparator erroring, in both reducers, at&lt;br&gt;
&lt;a href="https://github.com/keniel13-ui/self-correcting-integration-maintainer/commit/872507f" rel="noopener noreferrer"&gt;&lt;code&gt;872507f&lt;/code&gt;&lt;/a&gt;.&lt;br&gt;
It still runs after the model call, so a bad expectation has already cost a relay by the time it is&lt;br&gt;
named, and validating the expected bytes at prepare time is the version that belongs and is not&lt;br&gt;
written.&lt;/p&gt;
&lt;h2&gt;
  
  
  The general shape, if you want to check your own
&lt;/h2&gt;

&lt;p&gt;Find every place your system writes a failure name, and ask what the name is a statement &lt;em&gt;about&lt;/em&gt;.&lt;br&gt;
If it names a thing rather than a party — arguments, response, payload, schema — check what else&lt;br&gt;
falls into the same branch. A name that describes the object can be read as a verdict on whoever produced it. Trace each error&lt;br&gt;
from the operation that raised it all the way to the receipt, and check whether the name it arrives&lt;br&gt;
under still tells the truth.&lt;/p&gt;

&lt;p&gt;The question that found this one: &lt;strong&gt;when this fires, whose fault does a reader assume it is, and is&lt;br&gt;
that always true?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If you want to check the claim rather than take it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/keniel13-ui/self-correcting-integration-maintainer
&lt;span class="nb"&gt;cd &lt;/span&gt;self-correcting-integration-maintainer
git checkout dd1a654
node &lt;span class="nt"&gt;--test&lt;/span&gt; &lt;span class="nb"&gt;test&lt;/span&gt;/exec-comparator-error.test.mjs        &lt;span class="c"&gt;# 6 pass&lt;/span&gt;

git checkout dd1a654~1 &lt;span class="nt"&gt;--&lt;/span&gt; scripts/judgment/live.mjs &lt;span class="se"&gt;\&lt;/span&gt;
    scripts/pr2/constants.mjs scripts/pr2/reducer.mjs
node &lt;span class="nt"&gt;--test&lt;/span&gt; &lt;span class="nb"&gt;test&lt;/span&gt;/exec-comparator-error.test.mjs        &lt;span class="c"&gt;# 5 fail, 1 passes&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last run is the one worth doing. The test that keeps passing is the control.&lt;/p&gt;

&lt;p&gt;If you have one of these in your own harness, I would genuinely like to see it. Different domains,&lt;br&gt;
same shape.&lt;/p&gt;

</description>
      <category>testing</category>
      <category>ai</category>
      <category>opensource</category>
      <category>debugging</category>
    </item>
    <item>
      <title>You Agreed to a use. Not to a Hallway That Didn't Exist Yet</title>
      <dc:creator>Self-Correcting Systems</dc:creator>
      <pubDate>Wed, 09 Sep 2026 15:27:22 +0000</pubDate>
      <link>https://dev.to/kenielzep97/you-agreed-to-a-use-not-to-a-hallway-that-didnt-exist-yet-650</link>
      <guid>https://dev.to/kenielzep97/you-agreed-to-a-use-not-to-a-hallway-that-didnt-exist-yet-650</guid>
      <description>&lt;p&gt;You posted a photograph in 2014 to show your friends a night out. You understood what you were&lt;br&gt;
sharing and who you were sharing it with. You did not agree to every future system that would read&lt;br&gt;
it, combine it with other records, and decide something about you.&lt;/p&gt;

&lt;p&gt;That is not a worry about surveillance. Surveillance is about being watched. This is about being&lt;br&gt;
&lt;strong&gt;judged from inputs you never knew were inputs&lt;/strong&gt;, by a rule you never saw, in a place you have&lt;br&gt;
never been.&lt;/p&gt;

&lt;p&gt;Here is what that actually looks like, in machines that exist right now.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four machines that get mistaken for one
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. An allegation that travels.&lt;/strong&gt; &lt;a href="https://patronscan.com/policies/" rel="noopener noreferrer"&gt;Patronscan&lt;/a&gt; sells ID scanning to bars and clubs. Venue flags stay&lt;br&gt;
with the venue. Network flags surface at other participating venues when your ID is scanned. The&lt;br&gt;
pitch is exactly what it sounds like: a patron banned at one venue cannot walk down the street and&lt;br&gt;
start fresh at the next.&lt;/p&gt;

&lt;p&gt;Say what a flag is. A venue recorded an allegation. It is a report, not a finding, and it travels&lt;br&gt;
with the authority of a finding.&lt;/p&gt;

&lt;p&gt;Their policies page publishes the network-versus-venue distinction, the conduct categories, the&lt;br&gt;
retention limits, who has access, the route for removal, and the response windows. It does not&lt;br&gt;
publish how often a dispute ends with a flag changed or removed, and I found no aggregate rate&lt;br&gt;
published anywhere else. The process is documented. Its aggregate outcomes are not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Affiliation, not conduct.&lt;/strong&gt; In January 2023 the New York Attorney General sought information&lt;br&gt;
about reports that MSG Entertainment used facial recognition to exclude attorneys whose firms were&lt;br&gt;
litigating against it. Nothing was alleged about anyone's behavior. The input was who employed them,&lt;br&gt;
and being on one side of a lawsuit is not wrongdoing.&lt;/p&gt;

&lt;p&gt;The litigation is worth two paragraphs, because it answers a question people assume they know.&lt;/p&gt;

&lt;p&gt;In &lt;a href="https://www.nycourts.gov/reporter/3dseries/2023/2023_01646.htm" rel="noopener noreferrer"&gt;&lt;code&gt;Hutcher v Madison Sq. Garden Entertainment Corp.&lt;/code&gt;&lt;/a&gt; (2023 NY Slip Op 01646), the Appellate&lt;br&gt;
Division agreed that Civil Rights Law § 40-b &lt;strong&gt;requires&lt;/strong&gt; admitting a valid ticket holder to a&lt;br&gt;
theatrical performance or concert. The plaintiffs were right about the statute. The court vacated&lt;br&gt;
the injunction anyway, on a rule it quoted from 1915: where a statute creates a right and&lt;br&gt;
prescribes a remedy, that remedy is exclusive. &lt;a href="https://www.nysenate.gov/legislation/laws/CVR/41" rel="noopener noreferrer"&gt;§ 41&lt;/a&gt; supplies it — a penalty of "not less than one&lt;br&gt;
hundred dollars nor more than five hundred dollars." The legislature set that number, not the&lt;br&gt;
venue. The March decision resolved the injunction, not the case; the § 40-b claim survived it.&lt;/p&gt;

&lt;p&gt;The statute also does not reach sporting events. It names theatres, music halls, opera houses,&lt;br&gt;
concert halls and circuses. At a basketball game the right in this story does not exist.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Inference from exhaust.&lt;/strong&gt; Tenant screening, credit-based insurance scoring, employment&lt;br&gt;
background checks. Frequently no incident anywhere — a score built from records you never saw,&lt;br&gt;
about something you have not done.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. The system acts.&lt;/strong&gt; Not a database that informs a person who decides. Software that observes,&lt;br&gt;
infers, decides, and operates the applications that make the decision real. In August, xAI launched&lt;br&gt;
a &lt;a href="https://x.ai/news/introducing-grok-bot" rel="noopener noreferrer"&gt;persistent agent&lt;/a&gt; with its own always-on cloud computer, browser and terminal, signed into your&lt;br&gt;
applications. In September, Anthropic shipped &lt;a href="https://support.claude.com/en/articles/14128542-let-claude-use-your-computer-in-cowork" rel="noopener noreferrer"&gt;background computer use&lt;/a&gt; — the model clicks and types&lt;br&gt;
while you keep working. These are shipped products.&lt;/p&gt;

&lt;p&gt;I have not found a public case of one of these deciding about a person at a door. That gap is the&lt;br&gt;
honest state of row four.&lt;/p&gt;

&lt;p&gt;Collapse those four and one counterexample takes the argument down. Keep them apart and the&lt;br&gt;
uncomfortable part shows up.&lt;/p&gt;

&lt;h2&gt;
  
  
  One of these happened to me
&lt;/h2&gt;

&lt;p&gt;I did not go looking for a case. It was not any product named above. It was a vendor identity gate:&lt;br&gt;
a photograph of my driver's license, a biometric scan of my face taken on my phone, an automatic "did not meet the&lt;br&gt;
requirements," and no author on the sentence.&lt;/p&gt;

&lt;p&gt;Nothing I was shown named a reviewer, a fact, or a rule. Run it through the three questions at the&lt;br&gt;
end of this piece and watch them all fail. &lt;strong&gt;Is the fact wrong?&lt;/strong&gt; I was never told which fact. &lt;strong&gt;Is&lt;br&gt;
the fact right and the rule wrong?&lt;/strong&gt; I was never shown the rule. &lt;strong&gt;If I win, what comes back?&lt;/strong&gt;&lt;br&gt;
I was not shown a path. That is not the same as no path&lt;br&gt;
existing, and I was never given the difference.&lt;/p&gt;

&lt;p&gt;That is not any of the four machines. It is the thing they have in common: &lt;strong&gt;a decision returned&lt;br&gt;
with no author on the sentence.&lt;/strong&gt; I submitted what was asked. I did not finish their process. I hit&lt;br&gt;
the end of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your footprint can do the job of an ID. Its protection depends on the transaction.
&lt;/h2&gt;

&lt;p&gt;A Social Security number is an obvious regulated identifier. Your footprint does some of the same&lt;br&gt;
linking work — the apartment, the job, the ride, the door — but what protects it is not fixed. It&lt;br&gt;
depends on who is using it, for what, and under which regime.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which protection attaches depends on the information, who is using it, for what, and under which&lt;br&gt;
regime. The photograph does not change. The transaction does.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Hold that 2014 photograph constant.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A screening firm compiles it and sells it to your employer.&lt;/strong&gt; In &lt;a href="https://www.ftc.gov/sites/default/files/documents/closing_letters/social-intelligence-corporation/110509socialintelligenceletter.pdf" rel="noopener noreferrer"&gt;2011 the FTC&lt;/a&gt; concluded a
company doing exactly that was a consumer reporting agency, because it "assembles or evaluates"
information furnished to third parties who use it for employment eligibility. Public posts, full
FCRA obligations: accuracy, notice, dispute.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your employer searches your name and finds it himself.&lt;/strong&gt; The consumer-report route may not
apply. That does not mean no law applies. &lt;a href="https://www.eeoc.gov/laws/guidance/background-checks-what-employers-need-know" rel="noopener noreferrer"&gt;EEOC and FTC guidance&lt;/a&gt; is explicit that background
information is subject to federal anti-discrimination law "regardless of how you got the
information." Remove the reporting company and the regime &lt;strong&gt;changes&lt;/strong&gt;. It does not disappear.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A venue network attaches it to a flag.&lt;/strong&gt; A different regime again, largely private.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An agent retains it, infers from it, and acts.&lt;/strong&gt; Automation does not delete the organization
that deployed it or the transaction underneath. Which duties attach depends on the workflow, and
you cannot read that off the architecture.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Four handlings of one photograph, four sets of rules — not because the photograph changed, and not&lt;br&gt;
because any of them is a lawless zone, but because &lt;strong&gt;which regime applies is a function of the&lt;br&gt;
transaction, and the transaction is invisible to you.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;You cannot know that when you post. You especially cannot know it in advance, because some of those&lt;br&gt;
hallways were built after you posted.&lt;/p&gt;

&lt;h2&gt;
  
  
  Correcting a fact and contesting an inference are different rights
&lt;/h2&gt;

&lt;p&gt;If a file has your birthday wrong, there is a procedure. If a system has your birthday, employer,&lt;br&gt;
travel and associates all &lt;em&gt;correct&lt;/em&gt;, and concludes from them that you are high risk, correcting the&lt;br&gt;
records accomplishes nothing. Every input was already right. The conclusion was the problem.&lt;/p&gt;

&lt;p&gt;California's Attorney General addressed this directly in &lt;a href="https://oag.ca.gov/system/files/opinions/pdfs/20-303.pdf" rel="noopener noreferrer"&gt;Opinion 20-303&lt;/a&gt;. It is an opinion, not a&lt;br&gt;
court holding, and that distinction matters. The reasoning does not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"when a business processes personal information to make an inference about the consumer's&lt;br&gt;
propensities, then the inference itself becomes part of the consumer's profile, and must be&lt;br&gt;
disclosed. A business might draw an inference about a consumer based in whole or in part on&lt;br&gt;
publicly available information... Under the CCPA, the inference must be disclosed to the&lt;br&gt;
consumer, &lt;strong&gt;even if the public information itself need not be disclosed&lt;/strong&gt;."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The conclusion drawn about you can carry more disclosure obligation than the public facts behind&lt;br&gt;
it. The inference is a separate object. A proprietary algorithm does not automatically make its&lt;br&gt;
individualized output a trade secret.&lt;/p&gt;

&lt;p&gt;Three limits. One state. An AG opinion. And California separately provides a right to request&lt;br&gt;
&lt;strong&gt;correction&lt;/strong&gt; of inaccurate information (&lt;a href="https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&amp;amp;sectionNum=1798.106." rel="noopener noreferrer"&gt;Civ. Code § 1798.106&lt;/a&gt;) — a third thing again. Access to an&lt;br&gt;
inference, correction of a wrong fact, and forcing a decision reversed are three different rights,&lt;br&gt;
and having the first does not get you the third.&lt;/p&gt;

&lt;p&gt;One date keeps this honest: California's automated-decision-making rules took effect January 1&lt;br&gt;
2026, and businesses using ADMT for significant decisions have until &lt;strong&gt;January 1 2027&lt;/strong&gt; to comply.&lt;br&gt;
This is a gap scheduled to partially close.&lt;/p&gt;

&lt;p&gt;"Human in the loop" is not a claim. A human can approve every action, approve only payments,&lt;br&gt;
approve one objective at the start, review afterward, or merely be able to intervene. All five ship&lt;br&gt;
under "human oversight." The question is &lt;em&gt;where exactly&lt;/em&gt; authority sits.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers, and where they are going
&lt;/h2&gt;

&lt;p&gt;This is not a forecast. These are dated facts about a system already running.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;California built a delete button and half a million people pressed it.&lt;/strong&gt; The Delete Act created&lt;br&gt;
DROP, a single platform where one request reaches every registered data broker. It opened January 1&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;By August 13 more than &lt;strong&gt;475,000&lt;/strong&gt; Californians had filed (&lt;a href="https://www.gov.ca.gov/2026/08/13/icymi-california-takes-historic-action-against-data-brokers/" rel="noopener noreferrer"&gt;Governor's office, Aug 13&lt;/a&gt;). By August 25 the agency reported
&lt;strong&gt;more than 500,000 registered&lt;/strong&gt; and &lt;strong&gt;654 data brokers&lt;/strong&gt; in the system (&lt;a href="https://privacy.ca.gov/2026/08/half-a-million-californians-have-signed-up-for-drop-to-delete-their-personal-information-from-data-brokers/" rel="noopener noreferrer"&gt;CalPrivacy, Aug 25&lt;/a&gt;).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Registration is a business and the state priced it.&lt;/strong&gt; A broker's annual registration fee is&lt;br&gt;
&lt;strong&gt;$6,000&lt;/strong&gt;. From August 1 2026 brokers must access DROP at least every &lt;strong&gt;45 days&lt;/strong&gt;. The &lt;a href="https://privacy.ca.gov/drop-for-data-brokers/" rel="noopener noreferrer"&gt;penalties&lt;/a&gt;&lt;br&gt;
are two separate meters: &lt;strong&gt;$200 for each day&lt;/strong&gt; a broker fails to register, and &lt;strong&gt;$200 for each&lt;br&gt;
deletion request for each day&lt;/strong&gt; it fails to delete.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Then read the outcome numbers carefully, because they are the whole lesson.&lt;/strong&gt; On August 25 the&lt;br&gt;
agency reported &lt;strong&gt;tens of millions of records deleted&lt;/strong&gt;, that &lt;strong&gt;99.9% of consumers had their&lt;br&gt;
profile deleted by at least one broker&lt;/strong&gt;, and that the typical user had been removed by &lt;strong&gt;over 40&lt;br&gt;
brokers&lt;/strong&gt;. It also reported that about &lt;strong&gt;25% of brokers had reported processing deletion requests.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Put those next to each other. &lt;strong&gt;99.9% having at least one deletion proves reach, not completion.&lt;/strong&gt;&lt;br&gt;
Being removed by more than 40 brokers is a real result. But 654 is the system census, not that&lt;br&gt;
person's denominator — DROP reports &lt;a href="https://privacy.ca.gov/drop/how-drop-works/" rel="noopener noreferrer"&gt;five separate statuses&lt;/a&gt;, including &lt;code&gt;Record not found&lt;/code&gt; for brokers&lt;br&gt;
that never held your data or could not match you from what you entered, &lt;code&gt;Exempted&lt;/code&gt; for records a broker may lawfully keep, and &lt;code&gt;Pending&lt;/code&gt;,&lt;br&gt;
and brokers have up to 90 days to report.&lt;/p&gt;

&lt;p&gt;So the honest reading is narrower and still hard: I have not found a published consumer-level&lt;br&gt;
measure of how many people reached a fully resolved state across every broker that actually held&lt;br&gt;
their data, or how long that took. Reach is published. Completion is not.&lt;/p&gt;

&lt;p&gt;And I made this mistake myself while drafting this section, which is the point: &lt;strong&gt;a number is only&lt;br&gt;
as good as the population it counted&lt;/strong&gt;, and the trap does not spare the person writing about it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The first enforcement action is the thesis in one sentence.&lt;/strong&gt; In August 2026 California brought&lt;br&gt;
its first action under both the CCPA and the Delete Act, against LocateSmarter LLC, for failing to&lt;br&gt;
register and for requiring Californians to provide unnecessary data — &lt;strong&gt;the last four digits of&lt;br&gt;
their Social Security number — before they could opt out.&lt;/strong&gt; Total: &lt;strong&gt;$116,490.&lt;/strong&gt;&lt;br&gt;
(&lt;a href="https://privacy.ca.gov/2026/08/calprivacy-brings-first-action-against-a-data-broker-under-both-the-ccpa-and-delete-act/" rel="noopener noreferrer"&gt;agency decision&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;To switch off the sale of a record the law already covered, you had to hand over part of a more&lt;br&gt;
sensitive identifier. California treated that extra demand as unlawful. &lt;strong&gt;The protection existed.&lt;br&gt;
The company put a disclosure in the way of using it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And read what the company was selling. Names, dates of birth, Social Security numbers, phones,&lt;br&gt;
emails, employment, driver's license, bankruptcy records, litigation history — and inferences about&lt;br&gt;
consumer characteristics, including whether a person is &lt;strong&gt;"litigious."&lt;/strong&gt; The agency's own line:&lt;br&gt;
&lt;em&gt;"Inferences are a protected form of personal information under California law."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The arena used who employed you. The broker sold whether you sue. Same family of input. Not the&lt;br&gt;
same machine.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And the next date is already set.&lt;/strong&gt; California's automated-decision-making &lt;a href="https://cppa.ca.gov/announcements/2025/20250923.html" rel="noopener noreferrer"&gt;rules&lt;/a&gt; take full effect&lt;br&gt;
for businesses using ADMT in significant decisions on &lt;strong&gt;January 1 2027&lt;/strong&gt;. Whatever the gap is today,&lt;br&gt;
part of it closes on a schedule, in one state, for covered businesses.&lt;/p&gt;

&lt;h2&gt;
  
  
  The claims that get to skip the evidence
&lt;/h2&gt;

&lt;p&gt;Everything above had to be sourced. A flag is an allegation until someone adjudicates it. An AG&lt;br&gt;
opinion is not a ruling. A staff letter is not a court order. That standard is not optional if you&lt;br&gt;
want to be believed.&lt;/p&gt;

&lt;p&gt;Now apply it to the loudest AI claim in circulation.&lt;/p&gt;

&lt;p&gt;On September 3, Senator Sanders and Rep. Casar announced &lt;a href="https://www.sanders.senate.gov/press-releases/news-sanders-casar-introduce-legislation-to-ban-artificial-superintelligence-and-temporarily-pause-advanced-ai-development/" rel="noopener noreferrer"&gt;legislation to ban artificial&lt;br&gt;
superintelligence&lt;/a&gt; and pause advanced AI development until a cabinet-level regulator exists, with a&lt;br&gt;
corporate death penalty and up to twenty years imprisonment for violations.&lt;/p&gt;

&lt;p&gt;Read the announcement for what it defines. It does define the class, broadly: systems that surpass&lt;br&gt;
human intelligence, systems capable of overthrowing governments, systems with dangerous abilities&lt;br&gt;
such as subverting shutdown commands. It also names who would decide: a proposed cabinet-level&lt;br&gt;
regulator with an advisory board. What I did not find in the announcement or the one-page summary&lt;br&gt;
is a &lt;strong&gt;benchmark or evaluation protocol&lt;/strong&gt; for deciding when a system has crossed the prohibited&lt;br&gt;
threshold. Up to twenty years in prison hangs on that threshold. The full bill text is still&lt;br&gt;
described as forthcoming.&lt;/p&gt;

&lt;p&gt;Read it for what happens next. The proposal names real instruments: "international agreements,&lt;br&gt;
allied coordination, and policies such as export controls." Those are not nothing. What the public&lt;br&gt;
material does not show is a &lt;strong&gt;mechanism that can compel a government or developer outside U.S.&lt;br&gt;
jurisdiction that refuses to cooperate.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And when someone says advanced AI will end the world, ask the next question: &lt;strong&gt;by what mechanism,&lt;br&gt;
and for what reason?&lt;/strong&gt; Not "it will be very smart." A stated pathway you can check.&lt;/p&gt;

&lt;p&gt;Notice what each side is required to produce. A person disputing a bar flag files into a documented&lt;br&gt;
process with a published response window. A prohibition on a technology class travels as news with a&lt;br&gt;
summary and no published threshold. I am not claiming anyone measured those two against each other.&lt;br&gt;
I am asking the same question of both: &lt;strong&gt;what evidence would turn this into an enforceable&lt;br&gt;
decision?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That asymmetry is the same defect this entire piece is about. &lt;strong&gt;How a claim gets treated depends on&lt;br&gt;
who is making it and where it lands, not on what is behind it.&lt;/strong&gt; A senator can announce a&lt;br&gt;
prohibition on a technology class with a broad definition, no published operational threshold, and no bill text and it is news. You will need&lt;br&gt;
receipts to get a listing corrected.&lt;/p&gt;

&lt;p&gt;None of that is an argument for building anything anyone wants. It is an argument that the danger&lt;br&gt;
worth legislating is the one you can describe: what is collected, what gets inferred from it, who&lt;br&gt;
gets to act on the inference, and what you can do about it. &lt;strong&gt;That is not a ban on the technology.&lt;br&gt;
It is a ban on a use.&lt;/strong&gt; The difference is the entire distance between a law that could work and a&lt;br&gt;
press release.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is actually happening while that argument runs
&lt;/h2&gt;

&lt;p&gt;Nobody has to coordinate any of this.&lt;/p&gt;

&lt;p&gt;Safety teams study models. Privacy teams study data. Utilities study load. Cities study zoning.&lt;br&gt;
Congress studies regulation. Each is competent inside its own boundary. The technology crosses all&lt;br&gt;
of them in a single product release.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Preparation is fragmented. Deployment is integrated.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Every participant has an independent reason to accelerate one layer. Each decision can be rational&lt;br&gt;
and the sum still moves faster than anyone chose.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three questions that are not the same question
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Is the fact wrong?&lt;/strong&gt; There is usually a process for that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is the fact right and the rule wrong?&lt;/strong&gt; A system can identify your employer perfectly. The
dispute is whether your employer should decide whether you get into a concert. Correcting the
record cannot touch that. Whether a route exists to challenge the rule itself is a different
question, and it is not the same in every hallway.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you win, what comes back, and when?&lt;/strong&gt; A published response deadline is measurable. Whether
the answer arrives before the apartment is rented, the job is filled, or the show ends is a
different measurement, and I have not found it published.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Before a record decides something about you, you should be able to ask why it belongs in that&lt;br&gt;
decision. Afterward, you need to know who can change the result.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you have one of these, I want it
&lt;/h2&gt;

&lt;p&gt;I am collecting real cases before building anything, not after. Mine is above.&lt;/p&gt;

&lt;p&gt;If a decision about you turned on a record — a rental, a job, a claim, a door, an account — I want&lt;br&gt;
four things, with every identifying detail removed:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;What information was used, and did you know it was an input?&lt;/li&gt;
&lt;li&gt;Was the problem a wrong fact, an unsupported inference, or the rule itself?&lt;/li&gt;
&lt;li&gt;Who could change the decision, and what happened when you challenged it?&lt;/li&gt;
&lt;li&gt;Did the correction reach everyone who got the error, and did it restore the opportunity?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Don't post anything that identifies you or anyone else. I am after the shape, not the file.&lt;/p&gt;

</description>
      <category>privacy</category>
      <category>security</category>
      <category>ai</category>
      <category>legal</category>
    </item>
    <item>
      <title>Compare Against the Schema They Shipped, Not the One You Expected</title>
      <dc:creator>Self-Correcting Systems</dc:creator>
      <pubDate>Sun, 06 Sep 2026 21:57:04 +0000</pubDate>
      <link>https://dev.to/kenielzep97/compare-against-the-schema-they-shipped-not-the-one-you-expected-3mb8</link>
      <guid>https://dev.to/kenielzep97/compare-against-the-schema-they-shipped-not-the-one-you-expected-3mb8</guid>
      <description>&lt;p&gt;My harness flagged the model for sending the wrong arguments. It compared what the model actually sent against what the run had committed to in advance, and they did not match.&lt;/p&gt;

&lt;p&gt;The mismatch was real. The conclusion I drew from it was wrong, and the comparator could not have told me so.&lt;/p&gt;

&lt;p&gt;Here is the check that was incomplete, the check that replaced it, why the fix is not "loosen the comparison," and the part of it that will rot.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;The harness prepares an &lt;code&gt;exec&lt;/code&gt; call before the model runs, freezes it, tells the model to send exactly that object, then compares the model's actual tool arguments against the frozen one. If they differ, the run fails closed. That comparison is the control.&lt;/p&gt;

&lt;p&gt;The expectation was built like this — &lt;a href="https://github.com/keniel13-ui/self-correcting-integration-maintainer/commit/5bf10acd7a6c0dd80e90d99217c2610df0d86d74" rel="noopener noreferrer"&gt;&lt;code&gt;5bf10ac&lt;/code&gt;&lt;/a&gt;, &lt;code&gt;scripts/judgment/candidate.mjs&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;expectedExecArguments&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;validateExecArguments&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;command&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`node &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nf"&gt;artifactPath&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;verifier&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;&lt;span class="s2"&gt; &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nf"&gt;artifactPath&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;payload&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;&lt;span class="s2"&gt; &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nf"&gt;artifactPath&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;manifest&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One key. From the run receipt, &lt;a href="https://github.com/keniel13-ui/self-correcting-integration-maintainer/blob/main/docs/freezes/RUN_004_RECEIPT.json" rel="noopener noreferrer"&gt;&lt;code&gt;RUN_004_RECEIPT.json&lt;/code&gt;&lt;/a&gt;, here is what the model actually sent, verbatim:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"intent"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Run candidate verification"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"node /opt/tf/uploads/candidate-verifier.cjs /opt/tf/uploads/candidate-payload.json /opt/tf/uploads/candidate-command-manifest.json"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Outcome in the same receipt: &lt;code&gt;EXEC_ARGUMENTS_MISMATCH&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two contracts, and they disagreed
&lt;/h2&gt;

&lt;p&gt;This is the part I got wrong on the first read, and it is more interesting than "the model deviated."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The provider's exec schema&lt;/strong&gt; — from the compiled artifact I ran, &lt;code&gt;@truefoundry/trueforge-core@0.1.4&lt;/code&gt;, &lt;code&gt;dist/core/sandbox/Sandbox.mjs&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;var&lt;/span&gt; &lt;span class="nx"&gt;sandboxExecSchema&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;object&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;intent&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;string&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;A brief, user-facing explanation of the purpose of this command. Avoid technical jargon and code details.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="na"&gt;command&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;string&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;The shell command to execute in the sandbox.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="na"&gt;cwd&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;string&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;optional&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Working directory for command execution.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;record&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;string&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;string&lt;/span&gt;&lt;span class="p"&gt;()).&lt;/span&gt;&lt;span class="nf"&gt;optional&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Additional environment variables to set.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Four fields. &lt;strong&gt;&lt;code&gt;intent&lt;/code&gt; and &lt;code&gt;command&lt;/code&gt; required. &lt;code&gt;cwd&lt;/code&gt; and &lt;code&gt;env&lt;/code&gt; optional.&lt;/strong&gt; This is the implementation schema wired into the sandbox &lt;code&gt;exec&lt;/code&gt; tool; I am quoting what the runtime accepts, not a tool descriptor captured from the session.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My harness's instruction&lt;/strong&gt;, &lt;code&gt;scripts/judgment/live.mjs&lt;/code&gt;, present before the fix and after it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Call truefoundry-system exec exactly once with this exact JSON argument object: &lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At that moment the object being interpolated had one key.&lt;/p&gt;

&lt;p&gt;So the model was handed two authorities that disagreed, and it followed the provider's. It satisfied the required-field schema and missed my exact-JSON instruction, &lt;strong&gt;because my instruction was asking for something the schema forbade.&lt;/strong&gt; My comparator saw one difference and reported it as deviation. It had no way to represent "both sides are internally consistent and one of them is wrong."&lt;/p&gt;

&lt;p&gt;That is the finding. Not that the model was right — that &lt;strong&gt;a mismatch establishes difference, not which operand is authoritative.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix, and the part that matters
&lt;/h2&gt;

&lt;p&gt;The tempting repair is to compare less — check only &lt;code&gt;command&lt;/code&gt;, ignore extra keys, move on. That makes the failure disappear and takes the control with it. An agent could then send any additional argument it liked and still pass.&lt;/p&gt;

&lt;p&gt;What landed instead, in &lt;a href="https://github.com/keniel13-ui/self-correcting-integration-maintainer/commit/0220a27" rel="noopener noreferrer"&gt;&lt;code&gt;0220a27&lt;/code&gt;&lt;/a&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;CANDIDATE_VERIFICATION_INTENT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Run candidate verification&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;expectedExecArguments&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;command&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;intent&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;CANDIDATE_VERIFICATION_INTENT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;intent&lt;/code&gt; value is &lt;strong&gt;harness-authored and constant, not copied from what the model sent.&lt;/strong&gt; Copying it would make the comparison check the model against itself.&lt;/p&gt;

&lt;p&gt;That commit is titled &lt;em&gt;"Implement adopted transport A and B controls"&lt;/em&gt; — the correction rode inside a larger transport change rather than shipping as a dedicated fix. Worth saying, since I am asking you to open it.&lt;/p&gt;

&lt;p&gt;The expected-object gate also got stricter:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;argumentKeys&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Object&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;keys&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;expectedArguments&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;sort&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;argumentKeys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;argumentKeys&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;command&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;argumentKeys&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;intent&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt;
    &lt;span class="nx"&gt;expectedArguments&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;intent&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="nx"&gt;CANDIDATE_VERIFICATION_INTENT&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt;
    &lt;span class="k"&gt;typeof&lt;/span&gt; &lt;span class="nx"&gt;command&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;string&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;command&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;failures&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;EXEC_ARGUMENTS_MISMATCH&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Three separate things, and only one of them changed:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The &lt;strong&gt;expectation&lt;/strong&gt; was corrected — one key became two.&lt;/li&gt;
&lt;li&gt;The &lt;strong&gt;expected-object gate&lt;/strong&gt; got stricter — that block is new.&lt;/li&gt;
&lt;li&gt;The &lt;strong&gt;actual-versus-expected comparison stayed exactly as it was.&lt;/strong&gt; It already read:
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;actual&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;parseStrictJson&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;call&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;function&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;arguments&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;...&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nf"&gt;canonicalJsonBytes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;actual&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;equals&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;canonicalJsonBytes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;prepared&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;expectedExecArguments&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That line is identical before and after. Keys are sorted during canonical serialization, so it is a canonical object compare rather than a raw-byte one, and JSON key order cannot cause a false mismatch. &lt;strong&gt;I did not repair a false failure by weakening the comparator.&lt;/strong&gt; That is the whole point.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I fixed is narrower than it looks
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The provider's contract:&lt;/strong&gt; &lt;code&gt;intent&lt;/code&gt; and &lt;code&gt;command&lt;/code&gt; required, &lt;code&gt;cwd&lt;/code&gt; and &lt;code&gt;env&lt;/code&gt; permitted.&lt;br&gt;
&lt;strong&gt;My frozen run contract:&lt;/strong&gt; exactly &lt;code&gt;command&lt;/code&gt; and &lt;code&gt;intent&lt;/code&gt;, nothing else, intent fixed to a constant.&lt;/p&gt;

&lt;p&gt;Mine is deliberately narrower. A permissive provider does not oblige a harness to accept every schema-valid variation — if the run precommitted to two specific arguments, rejecting a third key is a legitimate harness constraint. But it is &lt;em&gt;my&lt;/em&gt; policy, not TrueForge's requirement, and writing it as though the provider demanded it would be the same error in the other direction.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And here is what is not fixed.&lt;/strong&gt; &lt;code&gt;argumentKeys.length !== 2&lt;/code&gt; hardcodes two things at once: the provider's current required-field set, and my run's decision to prohibit the optional fields. If TrueForge adds a third &lt;em&gt;required&lt;/em&gt; field tomorrow, my harness will reject a model for complying with the new schema unless I edit the harness too. The durable version derives the provider-required fields from the actual tool schema, then applies the narrower harness policy explicitly on top. I have not built that. &lt;strong&gt;If you copy this pattern, copy the problem with it.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The check you can steal
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Where did the expected object come from?&lt;/strong&gt; If it came from your reading of the API, it encodes your assumptions. If it came from the provider's schema, it encodes theirs. Only one is authoritative about what a compliant call looks like.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Does a failing comparison tell you which side is wrong?&lt;/strong&gt; Mine did not. It printed a mismatch and I had to open the provider's source to find the expectation at fault. &lt;strong&gt;A mismatch establishes difference, not which operand is authoritative.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. When you fix a false failure, does the check get weaker?&lt;/strong&gt; This is the one that bites. The fastest way to clear a red comparison is to compare less, and every time you do it you buy a passing run by selling the control that made passing mean something.&lt;/p&gt;

&lt;p&gt;The general shape: &lt;strong&gt;a control that fires wrongly is not evidence the control is too strict.&lt;/strong&gt; It is evidence that something on one of its two sides is wrong, and you have to find out which before you touch it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it cost
&lt;/h2&gt;

&lt;p&gt;Two things, and I want both on the record. I drew the wrong first conclusion about which side deviated. And the run still did not verify — the same receipt carries &lt;code&gt;EXEC_RESPONSE_SHAPE_UNEXPECTED&lt;/code&gt; alongside the argument mismatch, and the sandbox turned out to have no JavaScript runtime at all. That half is written up separately.&lt;/p&gt;

&lt;p&gt;If you have a comparison sitting between you and a model, go read the schema you are comparing against and check that your expected object satisfies it. It takes a minute, and it prevents the afternoon you spend investigating a model that satisfied the provider's tool contract while my own instruction conflicted with it.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Before: &lt;a href="https://github.com/keniel13-ui/self-correcting-integration-maintainer/commit/5bf10acd7a6c0dd80e90d99217c2610df0d86d74" rel="noopener noreferrer"&gt;&lt;code&gt;5bf10ac&lt;/code&gt;&lt;/a&gt;. Correction contained in: &lt;a href="https://github.com/keniel13-ui/self-correcting-integration-maintainer/commit/0220a27" rel="noopener noreferrer"&gt;&lt;code&gt;0220a27&lt;/code&gt;&lt;/a&gt;. Model arguments and outcome from &lt;a href="https://github.com/keniel13-ui/self-correcting-integration-maintainer/blob/main/docs/freezes/RUN_004_RECEIPT.json" rel="noopener noreferrer"&gt;&lt;code&gt;RUN_004_RECEIPT.json&lt;/code&gt;&lt;/a&gt;. Provider schema read from the compiled &lt;code&gt;@truefoundry/trueforge-core@0.1.4&lt;/code&gt; artifact in my own &lt;code&gt;node_modules&lt;/code&gt;, not from upstream source. The runtime half of this same run is in &lt;a href="https://dev.to/kenielzep97/a-finding-is-not-a-discovery-gib"&gt;I Built an Agent That Marked Its Own Finding as Already Known&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>ai</category>
      <category>testing</category>
      <category>opensource</category>
    </item>
    <item>
      <title>The Detector Reported Zero Because It Only Had One Item.</title>
      <dc:creator>Self-Correcting Systems</dc:creator>
      <pubDate>Fri, 04 Sep 2026 05:06:35 +0000</pubDate>
      <link>https://dev.to/kenielzep97/the-detector-reported-zero-because-it-only-had-one-item-ni0</link>
      <guid>https://dev.to/kenielzep97/the-detector-reported-zero-because-it-only-had-one-item-ni0</guid>
      <description>&lt;p&gt;Two instructions went into an Auditor my agent collaborators and I built to surface conflicts in agent instruction files. Deployment authority is one of nine domains the tool explicitly knows how to judge.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Never deploy without human approval.
Auto-deploy the moment tests pass.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On &lt;code&gt;main&lt;/code&gt; at &lt;a href="https://github.com/keniel13-ui/memory-authority-auditor/commit/172d962f2d53c6870d551b5b69db533f30caacb1" rel="noopener noreferrer"&gt;&lt;code&gt;172d962&lt;/code&gt;&lt;/a&gt;, that returns:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;posture&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;low_observed_risk&lt;/span&gt;
&lt;span class="na"&gt;counts&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;items"&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;1&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;labels"&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;governs"&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;1&lt;/span&gt;&lt;span class="pi"&gt;},&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;risk_high"&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;0&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt;
         &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;conflicts"&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{},&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gates"&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;0&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;authority_categories"&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;0&lt;/span&gt;&lt;span class="pi"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;low_observed_risk&lt;/code&gt; is the product's own string, from &lt;code&gt;agents/report_writer.py:110&lt;/code&gt;. Not my summary of the output. The output.&lt;/p&gt;

&lt;p&gt;One item. The pairwise comparison step never received a pair, and my detector does not compare an item with itself.&lt;/p&gt;

&lt;p&gt;After the repair, same input, &lt;a href="https://memory-authority-auditor-web-qfppqeeedq-uc.a.run.app/" rel="noopener noreferrer"&gt;live service&lt;/a&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"severity"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"high"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"item_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"M001, M002"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"authority_collision"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"finding"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Conflicting governing instructions in deployment: require_human_approval vs allow_automatic."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"evidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Never deploy without human approval. | Auto-deploy the moment tests pass."&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;posture: needs_review&lt;/code&gt;. Two items, one high-severity collision, one verification gate.&lt;/p&gt;

&lt;p&gt;The difference between those two outputs is that two lines were touching.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the failure actually was
&lt;/h2&gt;

&lt;p&gt;Before anything compares instructions, something has to split the text into separate instructions. Mine joined unbulleted lines into one item whenever they sat on consecutive lines with no blank line between them.&lt;/p&gt;

&lt;p&gt;The join is in the tool's first commit, &lt;a href="https://github.com/keniel13-ui/memory-authority-auditor/commit/b71892d425db590d625020ee475bb70827fd5330" rel="noopener noreferrer"&gt;&lt;code&gt;b71892d&lt;/code&gt;&lt;/a&gt;, authored 2026-06-01 13:06:47 -0400, at &lt;code&gt;agents/memory_extractor.py&lt;/code&gt; lines 50–51:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;content&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;part&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;part&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;pending_paragraph&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;part&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;36&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both defects in this article are on those two lines, and they have been there since the first commit. Three months.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Zero findings was an answer about a collapsed population.&lt;/strong&gt; The pairwise loop behaved exactly as written. It never got a pair.&lt;/p&gt;

&lt;p&gt;I have been publishing about this class of defect for three months: a count of zero means nothing until you know what reached the counter. I wrote that, then shipped a tool that got it wrong, and did not find out for three months.&lt;/p&gt;

&lt;h2&gt;
  
  
  How wide the defect actually was
&lt;/h2&gt;

&lt;p&gt;The detector knows nine domains: deploy authority, secrets handling, database source of truth, access scope, customer response, log retention, billing records, refunds, escalation. All hand-written.&lt;/p&gt;

&lt;p&gt;(Seven live in a stance table you can read in one glance. Refunds and escalation are compared by threshold rather than opposing stance, so if you go looking for a list of nine you will find a list of seven and two functions.)&lt;/p&gt;

&lt;p&gt;The precise scope, because the wider version is wrong: &lt;strong&gt;any pair whose two sides were written as adjacent, unbulleted lines with no blank line between them could be collapsed before comparison, in any of the nine domains.&lt;/strong&gt; That is not "the nine domains were disabled." A bulleted pair, or a pair separated by a blank line, extracted fine and compared fine the whole time. The vulnerable thing was a writing shape, not a domain.&lt;/p&gt;

&lt;p&gt;My own commit message on the repair says it worse than this article does — that consecutive lines "silently disabled conflict detection" across the nine domains. That wording is too wide. I am correcting it here rather than rewriting the commit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three things about the repair worth more than the repair
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;One: the tests were themselves tested.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The repair added seven regression tests. Passing on repaired code would not show they distinguish old behaviour from new, so we made them face the defect: stash the repair, restore only the missing constant so imports resolve, rerun against the old logic.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Four of the seven failed.&lt;/strong&gt; Real behavioural failures, not import errors.&lt;/p&gt;

&lt;p&gt;But four red lines are not four proofs, and the reasons matter more than the count:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Test&lt;/th&gt;
&lt;th&gt;Fails on&lt;/th&gt;
&lt;th&gt;Is that the defect?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;adjacent_instructions_do_not_merge&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;assert 1 == 2&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Yes.&lt;/strong&gt; Two lines glued into one item.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;enumerated_domain_still_produces_a_real_collision&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;assert []&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Yes.&lt;/strong&gt; No pair survived, so no collision could fire.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;short_high_risk_instruction_is_not_silently_discarded&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;assert 0 == 1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Yes&lt;/strong&gt; — but only because the injected constant is &lt;code&gt;12&lt;/code&gt;. Inject &lt;code&gt;36&lt;/code&gt; and it dies on &lt;code&gt;assert 36 &amp;lt;= 16&lt;/code&gt;, failing on the constant before it ever reaches the discard.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;governing_instruction_..._reports_uncovered_domain&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;assert 'uncovered_domain' in set()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;No.&lt;/strong&gt; Old &lt;code&gt;main&lt;/code&gt; has no &lt;code&gt;uncovered_domain&lt;/code&gt; in the detector at all. That failure is missing new code, not collapsed extraction.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two of these four prove the extraction bug. One proves it only under the right constant. &lt;strong&gt;One does not prove it at all.&lt;/strong&gt; A negative control whose failures fail for the wrong reasons is the exact defect this article is about, so I would rather print the table than let four red lines carry more weight than they earned.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two: the first deploy succeeded on the wrong tier.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;gcloud&lt;/code&gt; reported the truth: the web service deployed and served 100% of traffic. Accurate. I read it as meaning the behaviour had changed. It had not.&lt;/p&gt;

&lt;p&gt;The web app is a router. Extraction runs in &lt;code&gt;memory-extractor-agent&lt;/code&gt;, a &lt;strong&gt;separate&lt;/strong&gt; Cloud Run service. The revision timestamps are the receipt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;memory-authority-auditor-web-00003-82f   2026-09-02T13:17:10.714842Z
memory-extractor-agent-00002-grk         2026-09-02T13:31:32.807334Z
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Fourteen minutes and twenty-two seconds&lt;/strong&gt; in which a correct success message sat on top of unchanged behaviour. We only caught it because we tested the endpoint instead of reading the deploy message. Same wrong-reason pattern this project studies, live in our own release process, minutes after fixing the tool. The receipt was not false. My reading of what it covered was.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Three: I added an absence instead of a domain.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The input that started this was a different pair — publish-versus-verify — and the obvious repair was to teach the tool about publishing. We did not.&lt;/p&gt;

&lt;p&gt;Tuning a ruleset to the case someone just handed you proves only that it catches the known case. So the detector now emits &lt;code&gt;uncovered_domain&lt;/code&gt; when a governing instruction matches no rule at all:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This instruction governs action but matched no contradiction rule, so it was NOT evaluated for conflicts. Absence of a conflict here is absence of a check, not evidence of agreement.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Before, &lt;em&gt;no conflicts&lt;/em&gt; and &lt;em&gt;never checked&lt;/em&gt; rendered as the same sentence. Now they do not.&lt;/p&gt;

&lt;h2&gt;
  
  
  The original pair still is not solved, and the live output says so
&lt;/h2&gt;

&lt;p&gt;Here is the input that started this, run against the repaired live service:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"items"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"M001"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Current policy: verify the live artifact before publishing."&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"M002"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Old note: publish immediately without checking."&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="err"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="nl"&gt;"classifications"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"M001"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"authority_label"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"governs"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="nl"&gt;"confidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.78&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"M002"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"authority_label"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"context_only"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"confidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.64&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="err"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="nl"&gt;"conflicts"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"severity"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"medium"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"item_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"M001"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"uncovered_domain"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
   &lt;/span&gt;&lt;span class="nl"&gt;"finding"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"This instruction governs action but matched no contradiction rule, so it was NOT evaluated for conflicts."&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;posture: usable_with_gates&lt;/code&gt;. Extraction is fixed — two items, correctly split. It is still &lt;strong&gt;not&lt;/strong&gt; an &lt;code&gt;authority_collision&lt;/code&gt;, and it never will be until publishing becomes a listed domain.&lt;/p&gt;

&lt;p&gt;And there is a second gap in that JSON I did not know about until I pasted it for this article. &lt;strong&gt;&lt;code&gt;M002&lt;/code&gt; is classified &lt;code&gt;context_only&lt;/code&gt; at confidence 0.64.&lt;/strong&gt; "Publish immediately without checking" is an imperative, and the classifier does not consider it strong enough to govern. So the &lt;code&gt;uncovered_domain&lt;/code&gt; warning fires on M001 only. The half of the contradiction that tells you to skip the check gets no warning at all, because &lt;code&gt;context_only&lt;/code&gt; items are not eligible for one.&lt;/p&gt;

&lt;p&gt;The extraction repair is real. It moved this input from one silent item to two visible items and one honest warning. It did not make the tool right about this pair.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is fixed and what is not
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Fixed:&lt;/strong&gt; finished, unbulleted instructions on adjacent lines no longer merge when the first line ends in terminal punctuation. The minimum item length dropped from 36 characters to 12, because the old floor silently discarded the unbulleted instruction &lt;code&gt;Delete all logs.&lt;/code&gt; — sixteen characters, high risk, dropped with no record. A bulleted line bypassed that floor entirely, so the same words survived as a list item and vanished as a sentence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Not fixed, with the receipts:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A standalone unbulleted fragment under 12 characters still disappears with no record. &lt;code&gt;Wipe logs.&lt;/code&gt; is ten characters and returns zero items. So does &lt;code&gt;See above.&lt;/code&gt; — the floor does not distinguish a command from a cross-reference, it just deletes both.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The split now over-fires on hard-wrapped prose.&lt;/strong&gt; I previously claimed wrapped text was safe because wrapped lines do not end in terminal punctuation. That is a bet on wrapping, not a proof, and here is the counterexample: &lt;code&gt;"Escalate to the on-call engineer within 15 min.\nThen page the team lead if unresolved."&lt;/code&gt; is one two-step escalation procedure and the repair returns it as two independent items. Over-splitting is safer than merging, because two items can still be compared. It is still wrong.&lt;/li&gt;
&lt;li&gt;Still only nine conflict domains. A governing instruction outside them produces &lt;code&gt;uncovered_domain&lt;/code&gt;, which names the gap but not the conflict. A &lt;code&gt;context_only&lt;/code&gt; item outside them produces nothing at all — see M002 above.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Reproduce it
&lt;/h2&gt;

&lt;p&gt;Both hostnames for the live app route to the same service; I checked with the same payload and the responses are identical, so there is no stale tier to trip over.&lt;/p&gt;

&lt;p&gt;The repair is public as a branch, not on &lt;code&gt;main&lt;/code&gt;: &lt;a href="https://github.com/keniel13-ui/memory-authority-auditor/commit/beae0bb753b7e2be5b46a8ee6a8cf6f392ff3f72" rel="noopener noreferrer"&gt;&lt;code&gt;beae0bb&lt;/code&gt;&lt;/a&gt; on &lt;a href="https://github.com/keniel13-ui/memory-authority-auditor/tree/fix/extraction-merge" rel="noopener noreferrer"&gt;&lt;code&gt;fix/extraction-merge&lt;/code&gt;&lt;/a&gt;. Three files, 146 additions and one deletion. &lt;strong&gt;&lt;code&gt;main&lt;/code&gt; is still &lt;code&gt;172d962&lt;/code&gt; and still carries &lt;code&gt;len(content) &amp;gt;= 36&lt;/code&gt;&lt;/strong&gt;, so a default clone gets the defect. I am saying that rather than letting a green link imply the whole repo moved.&lt;/p&gt;

&lt;p&gt;Suite on the branch: &lt;strong&gt;107 passed, 1 skipped, 1 xfailed&lt;/strong&gt; in a clean clone. My working copy reads 108 because one provenance test finds workspace files that do not exist in an isolated clone. The clone number is the honest one to print next to a clone command.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Run the negative control:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/keniel13-ui/memory-authority-auditor
&lt;span class="nb"&gt;cd &lt;/span&gt;memory-authority-auditor
git fetch origin fix/extraction-merge
git checkout FETCH_HEAD &lt;span class="nt"&gt;--&lt;/span&gt; tests/test_extraction_boundaries.py

&lt;span class="c"&gt;# Restore only the constant the new tests import, so collection succeeds.&lt;/span&gt;
&lt;span class="c"&gt;# The old extractor below still uses its own literal &amp;gt;= 36 logic.&lt;/span&gt;
python3 - &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="no"&gt;PY&lt;/span&gt;&lt;span class="sh"&gt;'
from pathlib import Path
p = Path("agents/memory_extractor.py")
t = p.read_text()
assert "def extract_memories(" in t
p.write_text(t.replace("def extract_memories(", "MIN_ITEM_CHARS = 12&lt;/span&gt;&lt;span class="se"&gt;\n\n\n&lt;/span&gt;&lt;span class="sh"&gt;def extract_memories(", 1))
&lt;/span&gt;&lt;span class="no"&gt;PY

&lt;/span&gt;python3 &lt;span class="nt"&gt;-m&lt;/span&gt; pytest &lt;span class="nt"&gt;-q&lt;/span&gt; tests/test_extraction_boundaries.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;4 failed, 3 passed&lt;/strong&gt; — with the caveats in the table above.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters past one tool
&lt;/h2&gt;

&lt;p&gt;Agent instruction files accumulate rules written at different times by different people. The contradictions get harder to hold in working memory as the files grow.&lt;/p&gt;

&lt;p&gt;The reason to build a tool like this is to help a human stay the operator — to make a growing instruction set easier to inspect, not to replace the inspection or certify that nothing was missed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which means a tool in that job needs the same scrutiny it applies.&lt;/strong&gt; Ours did not get it for three months, and the thing that finally found it was not the suite. It was Kairos, a separate live agent seat in this project, pasting two lines into the deployed service and reading the answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Go break it
&lt;/h2&gt;

&lt;p&gt;It is live and it takes text: &lt;strong&gt;&lt;a href="https://memory-authority-auditor-web-qfppqeeedq-uc.a.run.app" rel="noopener noreferrer"&gt;https://memory-authority-auditor-web-qfppqeeedq-uc.a.run.app&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Paste in an instruction file, a set of agent rules, a policy doc, anything with rules written at different times. No signup. The app does not persist what you paste — it processes in memory and returns the answer. I cannot promise Google logs nothing at the platform layer, so do not paste anything you would mind appearing in a cloud access log.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I want is the case it misses.&lt;/strong&gt; Two instructions that clearly contradict, where it returns &lt;code&gt;low_observed_risk&lt;/code&gt; or &lt;code&gt;uncovered_domain&lt;/code&gt; instead of a conflict. I already know four shapes that beat it, and every one of them is in this article: anything outside the nine domains, any unbulleted fragment under twelve characters, any imperative the classifier rates &lt;code&gt;context_only&lt;/code&gt;, and any procedure hard-wrapped after a period. &lt;strong&gt;There will be more.&lt;/strong&gt; The one that started this survived three months and a green suite.&lt;/p&gt;

&lt;p&gt;Post what you gave it and what it returned. A miss is worth more to me than a hit.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;The two-line input came from Kairos, a separate live agent seat testing the deployed service — not from an external user or customer. Implementation, testing, and deployment were collaborative agent work under my direction. None of the pre-existing tests exercised that input shape.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;One correction about my own commit, since I am asking you to read it. The repair commit message and the regression-test docstring both say the defects were *&lt;/em&gt;"found by an outside reader."** That was imprecise. I meant outside the tool's own test suite; a reader following the link would reasonably take it to mean an outside person. It was not. I am leaving the commit as written and correcting it here rather than force-pushing over it, because a rewritten history is a worse receipt than an inaccurate one with a published correction attached.*&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>I Followed the Appeal Path. There Was No Appeal.</title>
      <dc:creator>Self-Correcting Systems</dc:creator>
      <pubDate>Tue, 01 Sep 2026 02:45:49 +0000</pubDate>
      <link>https://dev.to/kenielzep97/i-followed-the-appeal-path-there-was-no-appeal-25e2</link>
      <guid>https://dev.to/kenielzep97/i-followed-the-appeal-path-there-was-no-appeal-25e2</guid>
      <description>&lt;p&gt;&lt;strong&gt;This is part four of the Defender Access series.&lt;/strong&gt; Each part is standalone, but here is the thread if you are landing here cold:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://dev.to/kenielzep97/openai-says-verified-defenders-get-more-access-im-going-to-test-that-1n82"&gt;Part one&lt;/a&gt;&lt;/strong&gt; — OpenAI says verified defenders get more access, so I preregistered a study to test it, and published the design and its outcome table before running anything.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://dev.to/kenielzep97/they-matched-the-slogan-the-decision-lived-in-the-undefined-word-36o0"&gt;Part two&lt;/a&gt;&lt;/strong&gt; — they matched the published slogan. The decision lived in the undefined white space between the rules.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://dev.to/kenielzep97/the-suite-was-green-four-binding-controls-hadnt-established-their-claims-5d01"&gt;Part three&lt;/a&gt;&lt;/strong&gt; — my own suite was green while four binding controls had never established their claims.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Part four, this one&lt;/strong&gt; — the study cannot run, and the reason is worth more than the numbers would have been.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;Part two set a condition for continuing, and part three quoted it rather than paraphrasing it. I am going to do the same thing:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Part four continues only if a live result earns it. Same condition part two set: nothing here gets extended on vibes, and a part that has nothing new to measure does not get written.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A live result arrived. It is not the result the series was designed to measure, and I want to be exact about that before anything else.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This part does not report an uplift measurement.&lt;/strong&gt; The study announced in part one compares what a verified defender account can do against a control account. That study did not run and, through this route, it cannot run. What I have instead is a measurement of something one layer up: whether the treatment arm is reachable through the gate in front of the study. Through this route, it is not. Two dated records expose the gap: the denial directs a mistaken applicant to Support, while Support says it cannot explain or reverse the result, and that the verification process offers no retry or appeal.&lt;/p&gt;

&lt;h2&gt;
  
  
  The timeline
&lt;/h2&gt;

&lt;p&gt;Every line here is a timestamped email in my inbox. I am redacting my organization ID and my address and nothing else.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;When (UTC)&lt;/th&gt;
&lt;th&gt;From&lt;/th&gt;
&lt;th&gt;What it said&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2026-08-11 01:20:19&lt;/td&gt;
&lt;td&gt;Persona&lt;/td&gt;
&lt;td&gt;
&lt;em&gt;"Verify your business to join OpenAI's **Daybreak&lt;/em&gt;* access program"*&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-08-11 01:20:20&lt;/td&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;
&lt;em&gt;"Thank you for your request"&lt;/em&gt; — subject line &lt;strong&gt;OpenAI Pilot: Trusted Access For Cyber&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-08-11 23:45:16&lt;/td&gt;
&lt;td&gt;Persona&lt;/td&gt;
&lt;td&gt;&lt;em&gt;"Persona was unable to verify your account for OpenAI's Trusted Access for Cyber program."&lt;/em&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-08-31 06:16:33&lt;/td&gt;
&lt;td&gt;me&lt;/td&gt;
&lt;td&gt;the appeal, to &lt;code&gt;support@openai.com&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-08-31 06:17:14&lt;/td&gt;
&lt;td&gt;OpenAI Support&lt;/td&gt;
&lt;td&gt;Case 14067641&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two things about that table before the substance.&lt;/p&gt;

&lt;p&gt;The first two rows are one second apart and give the program two different names. Persona calls it &lt;strong&gt;Daybreak&lt;/strong&gt;. OpenAI calls it &lt;strong&gt;Trusted Access for Cyber&lt;/strong&gt;. The support reply at the end uses both in one sentence. That is cosmetic and I mention it only because it is the first visible sign of different naming conventions across the handoff.&lt;/p&gt;

&lt;p&gt;The last two rows are &lt;strong&gt;forty-one seconds apart&lt;/strong&gt;. I will come back to that.&lt;/p&gt;

&lt;h2&gt;
  
  
  The two sentences this part exists for
&lt;/h2&gt;

&lt;p&gt;The denial, in full, is four lines long. Here is the operative one:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"If you believe this was a mistake, please contact OpenAI support at &lt;a href="mailto:support@openai.com"&gt;support@openai.com&lt;/a&gt;."&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That names a channel and the condition under which to use it. It does not say what the channel can do once you get there. I believed it was a mistake, so nineteen days later I used it. I asked two narrow questions: what documentation does the verification step accept, and should I reopen the existing request or file a new one.&lt;/p&gt;

&lt;p&gt;The answer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"For Trusted Access for Cyber / Daybreak, if Persona verification fails or is denied, Support can't manually override it or provide additional details on the specific result, and verification currently doesn't support retries or appeals."&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Read those two quotes next to each other. The denial routes you to a desk. &lt;strong&gt;The desk says it cannot override the result, cannot tell you why the result happened, and that the process has no retries and no appeals.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And there is a sentence in OpenAI's own organization-verification documentation that sharpens this into something better than a complaint:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"If retries or appeals are available for your verification flow, the notice or product experience will explain the next step."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That establishes the positive case, and only the positive case: &lt;strong&gt;if&lt;/strong&gt; retries or appeals are available, the notice or product experience is where the next step will appear. It does not obligate a notice to announce their absence, and I am not going to pretend it does.&lt;/p&gt;

&lt;p&gt;What it does let me say precisely is this. My notice named Support. The current Daybreak documentation separately states that this flow has no retries and no appeals. &lt;strong&gt;What the notice did not say is that the Support route it named could not become an appeal.&lt;/strong&gt; The information existed in one place and the referral existed in another, and the applicant is left to reconcile the two.&lt;/p&gt;

&lt;p&gt;The broad version of this complaint is wrong, so let me kill it first. Support is real and it does real work: technical failures, stuck verification status, access troubleshooting. It is implemented.&lt;/p&gt;

&lt;p&gt;What is not implemented is any recourse &lt;em&gt;after a denial&lt;/em&gt;. The three things you would go to that channel for once you have been denied — reversal, explanation, another attempt — are all ruled out by the reply, but for two different reasons and by two different owners. Support cannot explain or override the result. The verification process offers no retry and no appeal.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The support channel existed. Recourse did not.&lt;/strong&gt; The denial pointed at the first without disclosing the absence of the second.&lt;/p&gt;

&lt;p&gt;The cost of that is not all theirs, and I want it attributed properly. The referral gave me reason to believe contacting Support might produce recourse. It never told me to wait. &lt;strong&gt;My own experimental rule is what turned that belief into a nineteen-day delay&lt;/strong&gt; — I would not write until I thought the estimand was safe. Two costs, two owners.&lt;/p&gt;

&lt;p&gt;This series has spent three parts on one idea, and it is a rule I work under: &lt;strong&gt;a rule that depends on someone choosing to perform it is a request, not a control.&lt;/strong&gt; &lt;a href="https://dev.to/mansio"&gt;mansio&lt;/a&gt; pushed it somewhere I had not taken it, in a thread on a different post: relying on a prompt instruction delegates system integrity to the model being evaluated.&lt;/p&gt;

&lt;p&gt;What I have here is the mirror image of both. &lt;strong&gt;A referral that names a channel without naming its limits is a notice, not a remedy.&lt;/strong&gt; It has the grammar of recourse without saying that recourse is not among the things on offer.&lt;/p&gt;

&lt;h2&gt;
  
  
  About those forty-one seconds
&lt;/h2&gt;

&lt;p&gt;The reply arrived forty-one seconds after I sent the appeal, and it says so itself:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Hi Keniel — I'm AI-assisted support for OpenAI."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;and closes with:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"This response was generated with AI support which can make mistakes."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;I am not going to tell you a human refused me, and I am not going to tell you no human read it either.&lt;/strong&gt; I cannot establish that. What I can establish is that the reply arrived forty-one seconds after I sent mine, identified itself as AI-assisted, and disclaimed its own accuracy in the same breath. Whatever human involvement there was or was not, the answer I received is one that says it can be wrong.&lt;/p&gt;

&lt;p&gt;The disclosure limits what the forty-one-second reply can establish on its own, and it no longer has to carry the policy claim by itself. &lt;strong&gt;OpenAI's current Daybreak documentation states independently that Support cannot override a verification result or provide additional details about why a specific result occurred, and that verification does not support retries or appeals.&lt;/strong&gt; Its FAQ puts it in one line: &lt;em&gt;can support change a verification result? No.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;One caveat I have to state, because it is the same defect this series keeps finding. That is a current page, not an archived copy from August 11. It tells me what the policy is now. It is not evidence of what the policy was on August 11, when I was denied, and I have no archived copy of the page as it stood that day. So the division of labour is: the email establishes the speed and wording of what I was told, the documentation establishes the policy as it currently stands, and neither establishes the policy at the moment of the denial.&lt;/p&gt;

&lt;p&gt;Which leaves the thing this whole series keeps arriving at: &lt;strong&gt;there is no observer downstream of that decision that I can reach.&lt;/strong&gt; Nothing in the record available to me distinguishes an AI-only response from one a human touched. OpenAI may well have internal routing or audit data that settles it. I am not claiming their systems cannot know. I am saying the applicant cannot, from what the applicant is given.&lt;/p&gt;

&lt;h2&gt;
  
  
  What was actually checked
&lt;/h2&gt;

&lt;p&gt;Here is the part that changes what I am allowed to claim.&lt;/p&gt;

&lt;p&gt;I do not know which checks determined my denial.&lt;/p&gt;

&lt;p&gt;Persona is an identity and business verification vendor. The support reply describes what the step &lt;strong&gt;generally expects&lt;/strong&gt;:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"information that matches official records (legal business name, registered/operating address + contact info, and sometimes a tax ID/business registration number) plus clear, complete, current official business documents"&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Legal name. Registered address. Tax ID. Business registration number. &lt;strong&gt;Not one of those fields is about security research.&lt;/strong&gt; Nothing in that list asks what I defend, what I own, what I have published, or whether my intended work is authorized.&lt;/p&gt;

&lt;p&gt;That was Support's abbreviated description to me, and I should not let it stand in for the whole process. OpenAI's current organization-verification documentation is wider than that email: it also asks for &lt;em&gt;accurate information about your organization's activities or intended use of OpenAI&lt;/em&gt;, and says the review may include &lt;em&gt;document verification, and compliance or risk screening&lt;/em&gt;. So the claim I can defend is narrow and specific — &lt;strong&gt;the list I was given did not assess security research&lt;/strong&gt; — not the broader claim that nothing anywhere in the application ever does.&lt;/p&gt;

&lt;p&gt;And here is the part I have to state carefully, because this whole series is about people claiming causes their receipts do not carry.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I cannot establish why Persona denied me.&lt;/strong&gt; The process withholds that reason by policy — Support said so in the same message. What I can establish is that &lt;strong&gt;Self-Correcting Systems is not a registered company.&lt;/strong&gt; It is me. No LLC, no incorporation, no tax ID, no registered address, no filing of any kind. So I am missing exactly the records Support says the flow generally expects, which makes the absent entity my strongest available explanation and &lt;strong&gt;not a proven cause&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I own &lt;code&gt;primaworkflows.com&lt;/code&gt;. Controlling a domain is not a substitute for legal entity records, and I would not expect it to satisfy a check asking for a registration number.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A business-registration gate is a legitimate control.&lt;/strong&gt; If you are handing out elevated cyber capability, tying it to an accountable legal entity is a defensible design and I would probably build it that way too. My complaint is not that the gate exists.&lt;/p&gt;

&lt;p&gt;My complaint is what the applicant is left holding. The program is described in terms of trusted defenders. Support's reply foregrounded business records. Those may not be the criterion the program name implies, and &lt;strong&gt;I have no way to find out&lt;/strong&gt;, because the result arrives with no reason attached and the channel that could explain it is documented as unable to. Whatever the gate actually evaluated, it did not tell me, and an applicant who cannot see which criterion was applied cannot distinguish a considered rejection from a records lookup that came back empty.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this does to the experiment
&lt;/h2&gt;

&lt;p&gt;Part one published the design's outcome table in advance, including the failure branches. One of them reads:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Approval arrives before a valid T0 → primary estimand is lost; post-only observation.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The design anticipated approval being &lt;strong&gt;late&lt;/strong&gt;. It anticipated entitlement being &lt;strong&gt;unclear&lt;/strong&gt;. It did not enumerate a branch where &lt;strong&gt;the applicant fails the verification gate before the treatment arm can be provisioned&lt;/strong&gt;, and that is now the branch I am in. OpenAI describes verification as helping evaluate whether a request qualifies, so this was an evaluation. It is simply not the evaluation the study was designed around, and it terminates before the study can begin.&lt;/p&gt;

&lt;p&gt;So, plainly, for the record and for anyone who was following along expecting numbers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The treatment arm cannot be provisioned through this route.&lt;/strong&gt; No approval, no retries, no appeals.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;T1 does not exist and will not under this preregistered route.&lt;/strong&gt; There is no post-treatment measurement because there is no treatment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The uplift study announced in part one is terminated, not paused.&lt;/strong&gt; Reviving it would require a different route to eligibility, and I do not have one.&lt;/li&gt;
&lt;li&gt;Part one's ceiling was already an N-of-1 case study with entitlement confounded with account identity. That ceiling is now moot. There is no arm to confound.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The rule I held the appeal under was that sending it early could contaminate the before/after comparison. &lt;strong&gt;There was no after to protect.&lt;/strong&gt; I would rather say that plainly than let the timeline imply I was being disciplined.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would tell you to take from this
&lt;/h2&gt;

&lt;p&gt;If you are building an access gate, one line:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Name what the channel can actually do for the situation you are naming it in.&lt;/strong&gt; Support does real work, and none of it is recourse after a denial. One sentence in that notice would have told me so, and its absence is what nineteen days of mine were spent on. &lt;strong&gt;The delay is mine. The missing sentence is theirs.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If you are on the other side of a gate like this, two things I did not know before:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Find out who is actually checking you, and accept that you may not be able to.&lt;/strong&gt; The program name on the door and the vendor doing the verification can be evaluating different things. I assumed a cyber program was assessing my security work. The support response I received foregrounded business records, and it arrived only because I asked a question that was almost about something else. OpenAI's current documentation is broader than that answer: it also names organizational activities or intended use, and compliance or risk screening. &lt;strong&gt;What I still cannot recover is which of those checks actually determined my result.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A denial with no stated reason is not a judgment you can read.&lt;/strong&gt; I spent some of those nineteen days reading it as one. It might have been. It could also have been a records lookup that came back empty. The point is that the format gave me no way to tell, and I filled the silence in myself.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;The series, in order: *&lt;/em&gt;&lt;a href="https://dev.to/kenielzep97/openai-says-verified-defenders-get-more-access-im-going-to-test-that-1n82"&gt;part one&lt;/a&gt;** (the preregistered design, which &lt;a href="https://dev.to/devteam/top-7-featured-dev-posts-of-the-week-5c26"&gt;DEV featured&lt;/a&gt; on August 18 for announcing the instrument this part reports could not be run) · &lt;strong&gt;&lt;a href="https://dev.to/kenielzep97/they-matched-the-slogan-the-decision-lived-in-the-undefined-word-36o0"&gt;part two&lt;/a&gt;&lt;/strong&gt; (the decision lived in the undefined white space) · &lt;strong&gt;&lt;a href="https://dev.to/kenielzep97/the-suite-was-green-four-binding-controls-hadnt-established-their-claims-5d01"&gt;part three&lt;/a&gt;&lt;/strong&gt; (a green suite over four unestablished controls) · part four, this one.*&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The email quotations above come from five dated messages in my inbox: two from &lt;code&gt;no-reply@frompersona.com&lt;/code&gt;, one from &lt;code&gt;noreply@openai.com&lt;/code&gt;, one from &lt;code&gt;support@openai.com&lt;/code&gt;, and one I sent. Quotations from OpenAI documentation come from the two Help Center pages linked below. I am not publishing the raw messages — they carry my organization ID and address — so the quotes are on my word and the timestamps are what I would produce if anyone wants to check the intervals. The forty-one second reply identifies itself as AI-assisted and states it can make mistakes. Its central claim does not rest on that message alone: OpenAI's own published Daybreak troubleshooting documentation states that Support cannot override a verification result or disclose its specific reason, and that verification currently supports neither retries nor appeals. &lt;a href="https://help.openai.com/en/articles/10910291-api-organization-verification" rel="noopener noreferrer"&gt;API organization verification&lt;/a&gt; · &lt;a href="https://help.openai.com/en/articles/20001259" rel="noopener noreferrer"&gt;Daybreak troubleshooting&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>ai</category>
      <category>testing</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Remember CLAIM-24? Here's what happened.</title>
      <dc:creator>Self-Correcting Systems</dc:creator>
      <pubDate>Mon, 31 Aug 2026 03:27:16 +0000</pubDate>
      <link>https://dev.to/kenielzep97/my-tests-agreed-with-my-code-neither-of-them-checked-reality-124l</link>
      <guid>https://dev.to/kenielzep97/my-tests-agreed-with-my-code-neither-of-them-checked-reality-124l</guid>
      <description>&lt;p&gt;&lt;em&gt;This is the CLAIM-24 result. If you followed that series in June and then watched it go quiet, this is what happened in the eleven weeks since.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The timeline, with receipts
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;2026-06-04 → 06-14.&lt;/strong&gt; Nine posts, CLAIM-24 through CLAIM-31, published in eleven days. The last one was &lt;a href="https://dev.to/kenielzep97/the-rule-held-the-boundary-moved-up-ai-memory-judgment-claim-31-verified-carryover-11if"&gt;The Rule Held. The Boundary Moved Up.&lt;/a&gt; on June 14. Then nothing, for eleven weeks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2026-06-07.&lt;/strong&gt; I opened &lt;a href="https://github.com/tacoda/keystone/issues/4" rel="noopener noreferrer"&gt;&lt;code&gt;tacoda/keystone#4&lt;/code&gt;&lt;/a&gt; asking whether Cell 7 could be tested against a real signing API instead of a mock. That thread is public, it now has forty-one comments, and everything below came out of it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2026-08-05 12:12Z.&lt;/strong&gt; An independent breaker reproduced a cross-origin credential-forwarding path offline, using a sentinel value and a cross-origin &lt;code&gt;Location&lt;/code&gt;. The sentinel crossed. No live credential ever did. &lt;a href="https://github.com/keniel13-ui/ai-memory-judgment-demo/blob/4eb5a40cd39da54a07c19dc6f6d690280f842681/claim_24/MANDATE_CELL7_CODE_BREAKER_VERDICT_KAEL_2026-08-04.md" rel="noopener noreferrer"&gt;Breaker verdict&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2026-08-05 → 08-06.&lt;/strong&gt; Live baseline captured. Mandate narrowed through the real API. A defect surfaced: my parser required ISO 8601 timestamps, the API returns Unix seconds, and twenty-two passing tests never caught it because my fixtures carried the same assumption I did. &lt;a href="https://github.com/keniel13-ui/ai-memory-judgment-demo/commit/1959a5559f66eb6b2cb7852112012d8f953e33f4" rel="noopener noreferrer"&gt;Repair commit&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2026-08-13.&lt;/strong&gt; &lt;code&gt;POST /mandate&lt;/code&gt; came back with a Cloudflare 1010, &lt;code&gt;browser_signature_banned&lt;/code&gt;. The cycle stopped at issuance. No mandate, no CAPTURE, no PATCH, no EVALUATE.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2026-08-14.&lt;/strong&gt; The default &lt;code&gt;Python-urllib/3.13&lt;/code&gt; user agent emerged as the strongest lead, one of the most commonly flagged signatures in bot detection. A lead, not an established cause.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2026-08-16 17:35Z.&lt;/strong&gt; The provider confirmed in writing that nothing had changed on their side.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2026-08-19.&lt;/strong&gt; The controlled cycle ran once. &lt;code&gt;POST /mandate&lt;/code&gt; returned &lt;strong&gt;201&lt;/strong&gt;. CAPTURE 200, PATCH 200, EVALUATE 200. One attempt per stage, no retry, no redirect follow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2026-08-20 10:41Z.&lt;/strong&gt; Key revoked by the issuer, on the record.&lt;/p&gt;

&lt;h2&gt;
  
  
  What that does and does not establish
&lt;/h2&gt;

&lt;p&gt;It establishes that the route opened and the full cycle completed against a live external API.&lt;/p&gt;

&lt;p&gt;It does &lt;strong&gt;not&lt;/strong&gt; establish why. Three days sit between the provider's written statement and the run, and nobody controlled that window. The receipt that binds the 201 does not carry the User-Agent, so it cannot testify about its own independent variable. The result is labelled &lt;code&gt;CONTROLLED_EXTERNAL_API_FIXTURE&lt;/code&gt; because the project authored the PATCH, which makes it a fixture rather than independent lifecycle evidence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cell 7 is not closed.&lt;/strong&gt; This is the live result the series was waiting on, and it arrives with its ceiling attached rather than as a conclusion.&lt;/p&gt;

&lt;p&gt;The rest of this piece is the part I think is worth more than the 201: what the two defects in my own code actually were, and why my own test suite could not have found either one.&lt;/p&gt;




&lt;p&gt;I had twenty-two passing tests and two separate reviewers on a piece of code. None of it objected. Then I pointed it at a real API owned by somebody else and it broke on the first live read.&lt;/p&gt;

&lt;p&gt;The mismatch fit in one sentence: my parser required ISO 8601, the documented API returned Unix seconds. The &lt;a href="https://github.com/keniel13-ui/ai-memory-judgment-demo/commit/1959a5559f66eb6b2cb7852112012d8f953e33f4" rel="noopener noreferrer"&gt;repair&lt;/a&gt; was not one line. It touched five files, 74 lines of parser and 52 lines of tests. The assumption was small; making it safe was not.&lt;/p&gt;

&lt;p&gt;Here is why nothing caught it, and it is the part worth keeping:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;My tests used ISO because my code used ISO, so they agreed with each other and never checked reality.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The fixtures were written by the person who wrote the parser. They encoded the same assumption. The suite confirmed internal behaviour without ever challenging the ISO assumption, because both halves of it came from one head. Internally consistent is not the same claim as right, and nothing in that suite could tell the difference.&lt;/p&gt;

&lt;p&gt;Two separate reviewers missed it too. I cannot prove why, and I am not going to invent a reason. What I can show is that the parser and every fixture encoded the same ISO assumption, so none of the artifacts in front of anyone supplied the live contract that contradicted it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The second one was worse
&lt;/h2&gt;

&lt;p&gt;Working against a real system made redirect containment matter, so an independent breaker went at it. In Python 3.13 the default redirect handler rebuilds the redirected request from &lt;code&gt;req.headers&lt;/code&gt;, dropping only content length and type. My &lt;code&gt;X-API-Key&lt;/code&gt; sat in that header set, so the redirected request inherited it. Python has &lt;code&gt;Request.add_unredirected_header()&lt;/code&gt; for exactly this, which marks a header as one that will not be added to a redirected request. I was not using it. The &lt;a href="https://github.com/keniel13-ui/ai-memory-judgment-demo/blob/4eb5a40cd39da54a07c19dc6f6d690280f842681/claim_24/MANDATE_CELL7_CODE_BREAKER_VERDICT_KAEL_2026-08-04.md" rel="noopener noreferrer"&gt;breaker reproduced it offline&lt;/a&gt; with a sentinel value and a cross-origin &lt;code&gt;Location&lt;/code&gt;, and the sentinel crossed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No live FIPSign credential was ever shown to have crossed an origin.&lt;/strong&gt; The defect was real and unshipped. I did not find it by auditing my own code, and I did not find it myself — the live integration made the question worth asking, and a seat that was not mine answered it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it takes to actually get that
&lt;/h2&gt;

&lt;p&gt;On 2026-06-07 I opened an issue on &lt;a href="https://github.com/tacoda/keystone" rel="noopener noreferrer"&gt;tacoda/keystone&lt;/a&gt;, an agent charter framework, asking whether a narrow adapter test was possible against a real signing API. &lt;a href="https://fipsign.dev" rel="noopener noreferrer"&gt;FIPSign&lt;/a&gt; (&lt;code&gt;mobydickfinance&lt;/code&gt;) said yes and then did something I did not expect: he handed over disposable API keys, rotated them on request, and stayed in that thread for ten weeks.&lt;/p&gt;

&lt;p&gt;The collaboration produced a block I could not explain. &lt;code&gt;POST /mandate&lt;/code&gt; came back with a Cloudflare 1010, &lt;code&gt;browser_signature_banned&lt;/code&gt;. That stopped the cycle at issuance: no mandate was created, and CAPTURE, PATCH and EVALUATE never ran.&lt;/p&gt;

&lt;p&gt;I could have guessed. He didn't let either of us do that. His reply:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Checked both questions directly, not from assumption.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Then he listed what he actually did. Reviewed the CORS and origin handling in the Worker source and confirmed server-to-server calls authenticated by API key never enter that logic at all. Checked Bot Fight Mode specifically. Compared the public guide field by field against the real request validation in the backend — length limits, scope counts, budget types, expiry bounds, required headers — and confirmed nothing was undocumented. Then ran &lt;code&gt;POST /mandate&lt;/code&gt; himself from PowerShell on a residential Windows connection and got a clean signed mandate back.&lt;/p&gt;

&lt;p&gt;Both of my hypotheses were dead. It was not the CORS or origin policy I suspected, and it was not a missing requirement in the docs.&lt;/p&gt;

&lt;p&gt;What he asked me for next is the part I want other people to steal:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;the exact timestamp · the HTTP client/library + version · whether the request went through any proxy or VPN · what kind of network you are running it from — home connection, cloud VM, CI runner, corporate network&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;With the note that hosting-provider IPs get scored very differently from residential ones, independent of the client library.&lt;/p&gt;

&lt;p&gt;I gave him the client. &lt;code&gt;Python-urllib/3.13&lt;/code&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;that's the default, unmodified user agent for Python's &lt;code&gt;urllib&lt;/code&gt;, and it's one of the most commonly flagged signatures by automated bot detection precisely because so much low-effort scraping traffic uses it unchanged.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Then, before I spent another key:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;if you set a real user agent on your client, I'd bet that alone resolves it. Worth trying first — cheap to test.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;If you are hitting an endpoint from a script and getting bot-flagged for no visible reason, check what user agent your HTTP library is sending by default.&lt;/strong&gt; You probably never set one. That is a five-second check and it is the most portable thing in this entire post.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part where I refused the result
&lt;/h2&gt;

&lt;p&gt;One seat added the header, an independent seat confirmed it was present in the request &lt;code&gt;urllib&lt;/code&gt; assembled for transport, and I authorized a single controlled cycle that a separately assigned firing seat executed once. &lt;code&gt;POST /mandate&lt;/code&gt; returned &lt;strong&gt;201&lt;/strong&gt;. No 1010. CAPTURE 200, PATCH 200, EVALUATE 200. One attempt per stage, no retry, no redirect follow.&lt;/p&gt;

&lt;p&gt;That is the clean story: header missing, endpoint blocked, header added, endpoint opened.&lt;/p&gt;

&lt;p&gt;I am not going to tell you that.&lt;/p&gt;

&lt;p&gt;FIPSign's operator had confirmed nothing changed on his side — Bot Fight Mode off then and off now, never toggled, a Browser Integrity Check rule drafted during the investigation but never deployed. That statement is dated &lt;strong&gt;2026-08-16 17:35Z&lt;/strong&gt;. The firing seat ran the controlled cycle on &lt;strong&gt;2026-08-19&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Three days sit between a written statement about a provider's configuration and the run that depends on it. He was not being careless and neither was I. But nobody controlled that window, so a success is attributable to the header only as far as that statement reaches. One run, one header, three days later, is not causation. It is a result with an uncontrolled gap in front of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  And then the worse gap, which was mine
&lt;/h2&gt;

&lt;p&gt;That run's issuance receipt contains &lt;strong&gt;no User-Agent field at all.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The entire point of the run was the header. A narrative receipt written afterwards names it. The artifact that binds the 201 does not. The two share a run id, so software can associate them, but the User-Agent assertion sits outside the sealed receipt chain rather than inside it. A reader holding the artifact that binds the 201 still cannot recover which user agent produced it.&lt;/p&gt;

&lt;p&gt;We built the capture afterwards — it records the configured value, the mapping handed to the transport, and what &lt;code&gt;urllib&lt;/code&gt; assembled after &lt;code&gt;add_header&lt;/code&gt;, three witnesses at three different strengths, none of them TLS-level and none of them the edge's view. It exists because the previous run could not testify about its own independent variable, which is not a good reason for a thing to exist.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The route opened. The experiment did not establish why.&lt;/strong&gt; The receipt did not bind the one variable that was supposed to explain the change. That is a worse failure than the block was.&lt;/p&gt;

&lt;p&gt;The result carries &lt;code&gt;CONTROLLED_EXTERNAL_API_FIXTURE&lt;/code&gt; on its face. The project authored the PATCH, so it is not independent lifecycle evidence and I will not present it as any.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the ten weeks were actually worth
&lt;/h2&gt;

&lt;p&gt;Not the 201. The 201 is a fixture with a disclosed ceiling.&lt;/p&gt;

&lt;p&gt;What it was worth was two defects in my own code, found by two different paths, and collapsing them into one lesson would overclaim what the receipts show.&lt;/p&gt;

&lt;p&gt;The timestamp mismatch survived twenty-two green tests and two separate reviewers because the parser and every fixture carried the same wire-format assumption. The live response is what contradicted it.&lt;/p&gt;

&lt;p&gt;The redirect defect was different. It survived a green local suite and was found by an independent breaker attacking the adapter with a sentinel value, before any live credential could cross. &lt;strong&gt;One was reality contradicting my fixture. The other was an adversarial seat asking a question my tests had not asked.&lt;/strong&gt; One of them was a cross-origin credential-forwarding path.&lt;/p&gt;

&lt;p&gt;His closing note put it plainly, and he is right:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;it says something good about testing against a live system instead of only mocks.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For an external wire contract, a fixture you invented from your own reading of someone else's docs cannot independently validate that reading. My suite could not challenge the ISO assumption because every fixture already carried it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reality is a better adversary than your fixtures. It is not a substitute for a receipt.&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Thread is public: &lt;a href="https://github.com/tacoda/keystone/issues/4" rel="noopener noreferrer"&gt;&lt;code&gt;tacoda/keystone#4&lt;/code&gt;&lt;/a&gt;, opened 2026-06-07, forty-one comments when audited on 2026-08-30. Every quote above is from it. **The sanitized run receipt for 2026-08-19 is not currently public&lt;/em&gt;* — it is credential-adjacent and has not been through a publication scan, so the User-Agent gap described above is stated on my word rather than something you can open. I would rather say that than imply a receipt you cannot read. The keys used for the Mandate runs described here were disposable, delivered privately, reported spent, and revoked by the issuer.*&lt;/p&gt;

</description>
      <category>testing</category>
      <category>api</category>
      <category>opensource</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Tell Me About You</title>
      <dc:creator>Self-Correcting Systems</dc:creator>
      <pubDate>Sun, 30 Aug 2026 19:05:28 +0000</pubDate>
      <link>https://dev.to/kenielzep97/tell-me-about-you-1hi4</link>
      <guid>https://dev.to/kenielzep97/tell-me-about-you-1hi4</guid>
      <description>&lt;p&gt;After sixty-seven published posts here, the comments have become one of the best parts of the work. Some of you have run my code, found things I missed, challenged a claim, or brought me a project of your own. I also want to distinguish bots from real people.&lt;/p&gt;

&lt;p&gt;I know some of you by the exact problem we were discussing. I still do not know what many of you are building, learning, enjoying, or trying to become outside that thread.&lt;/p&gt;

&lt;p&gt;So this post is different. I am going to introduce myself first, and then I want to hear about you.&lt;/p&gt;

&lt;p&gt;If we have never spoken before, you are welcome here. If we have spoken ten times already, you are welcome to introduce yourself again. Tell me something I do not know yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who I am away from the terminal
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Basketball is a real part of my life.&lt;/strong&gt; On September 12 I am one of the official coaches at the Hugo Gonzalez Basketball Camp at Brandeis University in Waltham, Massachusetts. &lt;a href="https://www.nba.com/player/1642864/hugo-gonzalez" rel="noopener noreferrer"&gt;Hugo&lt;/a&gt; plays for the Boston Celtics. The brief the coaches were given is the part I liked: treat every player like they are your own kid, and no phones. That is genuinely one of the cooler things I get to say right now. Basketball was my first love i never played college or anything like that so i don’t wanna give the wrong impression either basketball is just a passion. &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I play PokeMMO, but only for PvP.&lt;/strong&gt; I think of it as modern-day chess. You are not just reacting to the board in front of you. You are playing several turns ahead of someone who is trying to do the same thing to you. You have to hold a model of what they think you are about to do, then decide whether to follow the obvious line or break it.&lt;/p&gt;

&lt;p&gt;That way of thinking is connected to everything else I work on. I expect the world to become much more automated, and I want to build now for the decisions I will need to understand later. I would rather learn how the machine reaches a conclusion than wait for a future where its decisions simply happen to me.&lt;/p&gt;

&lt;p&gt;I did not arrive here through scholarship or credentials. I am someone who genuinely enjoys learning, testing an idea, being wrong, and trying again with better evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I hope to build
&lt;/h2&gt;

&lt;p&gt;The long-term idea is &lt;strong&gt;Self-Correcting Systems&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The goal is not a machine that never fails. It is a machine that can notice when reality differs from what was expected, examine the anomaly, explain what it can and cannot establish, and help correct course without quietly hiding the failure from the person operating it.&lt;/p&gt;

&lt;p&gt;That asks machines to handle pieces of the critical judgment humans use every day: noticing that something feels wrong, comparing it with prior experience, changing an interpretation, and knowing when confidence is not enough. I know how far away that goal is. I am not claiming to have built it. I am telling you what all of this work is pointed toward.&lt;/p&gt;

&lt;p&gt;The next major system I want to develop is for cybersecurity. I expect the machinery underneath it to become intricate, but the controller has to remain understandable enough to operate almost like a video game. Complexity underneath cannot become confusion at the surface. If I cannot see what my own system observed, what it inferred, what it changed, and what authority it used, then I am no longer operating it. It is operating me.&lt;/p&gt;

&lt;p&gt;One day I want to build a company around this with people who think similarly: ambitious about what machines can become, skeptical of claims that outrun evidence, and serious about keeping humans able to understand and control what they create.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I actually work with AI
&lt;/h2&gt;

&lt;p&gt;I do not work with one assistant that agrees with everything I say. I use multiple AI agents in separate roles, and they are expected to challenge one another and challenge me.&lt;/p&gt;

&lt;p&gt;We try to work under a constraint system called the &lt;strong&gt;Drift Gauntlet&lt;/strong&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Close the current loop.&lt;/strong&gt; Do not escape into a new lane before the current one produces a real outcome.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Make something touch the world.&lt;/strong&gt; A reply, a test from another machine, a shipped artifact, a yes, or a no teaches more than another private plan.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Require a real edge.&lt;/strong&gt; “The AI will figure it out” is not a mechanism.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do the work before studying the idea of the work.&lt;/strong&gt; Research cannot become a hiding place from execution.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Record the receipt.&lt;/strong&gt; If the result is not preserved, nobody can learn from it or challenge it later.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Name the grandiosity.&lt;/strong&gt; A huge future goal does not excuse skipping the next checkable step.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The gauntlet is supposed to be blind to rank. My own named failure mode is impulsiveness: moving before I fully understand the practical use of what we are doing. One agent tends to skim and answer the last part of a problem. Another tends to over-produce and build more than the loop requires. We name those failure modes in advance because they rarely feel like mistakes while they are happening. They usually feel like working hard.&lt;/p&gt;

&lt;p&gt;This is not a perfect automated control. We still violate it. While this post was being built, one agent concluded that an unpublished draft was a forgotten article because its sentences barely overlapped with anything already live. Another agent opened the live article and found that the same thesis, commands, hashes, and public request had already been published two weeks earlier. The first agent withdrew the draft.&lt;/p&gt;

&lt;p&gt;That is the workflow at its best: disagreement is welcome, but the live object gets the final word.&lt;/p&gt;

&lt;h2&gt;
  
  
  Now I want to hear about you
&lt;/h2&gt;

&lt;p&gt;Who are you?&lt;/p&gt;

&lt;p&gt;Tell me whatever you are comfortable sharing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;what you are building, studying, or trying to learn&lt;/li&gt;
&lt;li&gt;a problem you cannot stop thinking about&lt;/li&gt;
&lt;li&gt;a hobby that matters to you&lt;/li&gt;
&lt;li&gt;something you are proud of&lt;/li&gt;
&lt;li&gt;a project link, if you have one&lt;/li&gt;
&lt;li&gt;or simply what brought you to this corner of the internet&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You do not need a project. You do not need a repository. You do not need to be technical. There is no required level of experience, and this invitation is not reserved for people I have never met.&lt;/p&gt;

&lt;p&gt;If you are a real person leaving an honest introduction, I will follow you back. The point is not to trade numbers. I want my feed to contain more people I have actually heard from and fewer accounts that are only numbers on a screen.&lt;/p&gt;

&lt;p&gt;So whether this is our first conversation or our twentieth: &lt;strong&gt;what is one thing you would like me to remember about you?&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>watercooler</category>
      <category>discuss</category>
      <category>community</category>
      <category>ai</category>
    </item>
    <item>
      <title>My Commit Said "Stop Trusting the Receipt." That Is the Commit Where the Trusting Started.</title>
      <dc:creator>Self-Correcting Systems</dc:creator>
      <pubDate>Sun, 30 Aug 2026 07:04:02 +0000</pubDate>
      <link>https://dev.to/kenielzep97/my-commit-said-stop-trusting-the-receipt-that-is-the-commit-where-the-trusting-started-306j</link>
      <guid>https://dev.to/kenielzep97/my-commit-said-stop-trusting-the-receipt-that-is-the-commit-where-the-trusting-started-306j</guid>
      <description>&lt;p&gt;On 2026-08-25 at 21:27 EDT I pushed this to &lt;a href="https://github.com/keniel13-ui/self-correcting-integration-maintainer" rel="noopener noreferrer"&gt;&lt;code&gt;self-correcting-integration-maintainer&lt;/code&gt;&lt;/a&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;fix: repair the four re-review findings; stop trusting the receipt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The message is accurate. It closed four real findings from an automated review. It is also the commit that added this line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;recomputed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;decide&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;receipt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;checks&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;receipt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;deciding_fields&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second argument is the receipt's own claim about which fields it should be judged on. So the validator recomputes its verdict over terms the subject supplied. A receipt carrying failing checks plus &lt;code&gt;deciding_fields: []&lt;/code&gt; recomputes over nothing, finds nothing failing, and validates clean.&lt;/p&gt;

&lt;p&gt;Absence reading as a pass, inside the fix for absence reading as a pass.&lt;/p&gt;

&lt;p&gt;The reviewer flagged it at 21:30:49 EDT. &lt;strong&gt;Three minutes and twenty-five seconds.&lt;/strong&gt; (The finding's &lt;code&gt;original_commit_id&lt;/code&gt; is &lt;code&gt;4a8e6c9&lt;/code&gt;; GitHub now displays it against the later head &lt;a href="https://github.com/keniel13-ui/self-correcting-integration-maintainer/commit/5d053c172320036de9d3b8e687f34322d956a717" rel="noopener noreferrer"&gt;&lt;code&gt;5d053c17&lt;/code&gt;&lt;/a&gt;, which is why the timestamps are worth stating rather than the UI position.)&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that is not a bug story
&lt;/h2&gt;

&lt;p&gt;The commit message is not a lie. It closed four findings. It just names the opposite of what the diff did on one line.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An audit by commit message passes this. An audit by diff summary passes this.&lt;/strong&gt; A reviewer reading "stop trusting the receipt" has been told the answer and will read the diff looking for confirmation of it. Only following the data catches it.&lt;/p&gt;

&lt;p&gt;That is the condition I did not have a name for before: a repair arrives carrying the credibility of a repair. It closed something real, it was reviewed, and it says so on the tin. That is precisely when nobody looks at it twice.&lt;/p&gt;

&lt;h2&gt;
  
  
  Readers named the class, not the commit
&lt;/h2&gt;

&lt;p&gt;On 2026-08-24 I published &lt;a href="https://dev.to/kenielzep97/the-tests-passed-the-contract-was-wrong-mp0"&gt;a piece&lt;/a&gt; about a contract that permitted the contradiction its tests were passing. Two commenters went past the instance.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/pm25coder/comment/3df2g"&gt;pm25coder&lt;/a&gt;, the same day:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Every repair moved authority to something 'better typed,' and the third contract's authority field is itself a derived value one level down."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;He was describing a different project — a grant-expiry contract in a Python file, not this repository. He named a &lt;strong&gt;class&lt;/strong&gt;: authority migrates one level down and the subject ends up supplying the terms it is judged by.&lt;/p&gt;

&lt;p&gt;The chronology is checkable end to end. His comment posted 2026-08-24 at 09:41 EDT. GitHub says this repository was created 2026-08-25 at 19:47:38 EDT. &lt;code&gt;4a8e6c9&lt;/code&gt; landed at 21:27 EDT that same night — one hour and forty minutes after the repository existed, in JavaScript rather than the Python he was reading, with a message claiming the opposite of what it did.&lt;/p&gt;

&lt;p&gt;I want to be exact about the credit, because getting it wrong would be the same defect one more time. &lt;strong&gt;He did not predict this commit.&lt;/strong&gt; He named a failure class, and the class recurred. That is more useful than prophecy and it is a weaker claim than prophecy, and the difference matters.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/anp2network/comment/3dfeg"&gt;anp2network&lt;/a&gt; went at the method rather than the instance, and I will come back to that.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same class on 2026-08-29, in prose
&lt;/h2&gt;

&lt;p&gt;On 2026-08-29 a submission document carried a stale count of review comments. The number had shipped wrong twice already. The correction read: &lt;strong&gt;"37 inline review comments across six merged pull requests."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Arithmetically right when written, and self-invalidating as a complete-set claim. Merging it creates a seventh merged pull request while the sentence still defines the universe as six. The total would have stayed 37 — PR #6 carries zero inline comments under the same endpoint — so the number never goes wrong. &lt;strong&gt;The set does.&lt;/strong&gt; The denominator named every pull request except the one doing the counting.&lt;/p&gt;

&lt;p&gt;Caught before merge by a second seat, not by anyone checking the arithmetic. The wording that shipped names the measured set instead:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Across the six pull requests merged before this correction (&lt;code&gt;#1&lt;/code&gt;, &lt;code&gt;#2&lt;/code&gt;, &lt;code&gt;#3&lt;/code&gt;, &lt;code&gt;#4&lt;/code&gt;, &lt;code&gt;#5&lt;/code&gt;, &lt;code&gt;#7&lt;/code&gt;), Qodo authored 37 inline review comments as of 2026-08-29.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Merged as &lt;a href="https://github.com/keniel13-ui/self-correcting-integration-maintainer/commit/3ee11d1d898c1d05be742e824c7b8d0c7370876e" rel="noopener noreferrer"&gt;&lt;code&gt;3ee11d1&lt;/code&gt;&lt;/a&gt;. Recomputed after the merge: still 37.&lt;/p&gt;

&lt;p&gt;Same shape as the commit above, moved out of code and into a sentence: &lt;strong&gt;an artifact supplied the terms of its own completeness.&lt;/strong&gt; Three attempts had fixed the arithmetic. The number was never the defect — the measurement boundary was.&lt;/p&gt;

&lt;h2&gt;
  
  
  Status of the repair, stated honestly
&lt;/h2&gt;

&lt;p&gt;The current line freezes the terms in the consumer and demotes the receipt's copy to evidence:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;CANONICAL_DECIDING_FIELDS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Object&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;freeze&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;node&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;trueforge&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;sdk&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
&lt;span class="p"&gt;...&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;recomputed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;decide&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;receipt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;checks&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;CANONICAL_DECIDING_FIELDS&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://github.com/keniel13-ui/self-correcting-integration-maintainer/blob/4a8e6c97ca3eca83710ac01ac403d15867fe0eb0/scripts/prerequisites.mjs#L103" rel="noopener noreferrer"&gt;Defect&lt;/a&gt; · &lt;a href="https://github.com/keniel13-ui/self-correcting-integration-maintainer/blob/2550a632c3dc118122664db9d953d5689014a999/scripts/prerequisites.mjs#L129" rel="noopener noreferrer"&gt;current&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The reviewer did come back to it. &lt;a href="https://github.com/keniel13-ui/self-correcting-integration-maintainer/commit/b5be3b7d9ad1940c0c2bdfdbc4b469b029624922" rel="noopener noreferrer"&gt;&lt;code&gt;b5be3b7&lt;/code&gt;&lt;/a&gt; is the repair — it adds the constant and swaps the argument — committed 2026-08-26 at 19:23:36 EDT, with Qodo's review updated to that exact commit at 19:26:04 EDT. &lt;strong&gt;A re-review at the repair head, and I am not going to omit it because it cuts against the shape of the story.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Qodo's review moved again ten minutes later to &lt;a href="https://github.com/keniel13-ui/self-correcting-integration-maintainer/commit/121a24ffcf5012208c77cc5a54143fea3f1113b0" rel="noopener noreferrer"&gt;&lt;code&gt;121a24f&lt;/code&gt;&lt;/a&gt;. That one is a different fix — counting providers instead of trusting that a response arrived — and it does not touch this file at all. It carries the repair only because it comes after it. Worth separating, because "the reviewer cleared it twice" would be a nicer sentence than the true one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I still do not call it fixed.&lt;/strong&gt; The patch is maker-authored, and no separately assigned breaker seat has adjudicated it. On this project a maker's own PASS does not count no matter who else looked, and the last two times I felt confident about a repair are the two stories above.&lt;/p&gt;

&lt;p&gt;One thing I did check, because a reviewer of this draft predicted a second hole in the same class: if a receipt simply &lt;em&gt;omits&lt;/em&gt; a canonical key from &lt;code&gt;checks&lt;/code&gt;, does absence read as a pass again? It does not. &lt;code&gt;decide()&lt;/code&gt; filters on &lt;code&gt;checks[field]?.observed !== true&lt;/code&gt;, so a missing key lands in &lt;code&gt;blocked_by&lt;/code&gt; and the receipt is rejected. Omitting &lt;code&gt;sdk&lt;/code&gt; yields &lt;code&gt;LOCAL_PREREQS_BLOCKED ["sdk"]&lt;/code&gt;; &lt;code&gt;checks: {}&lt;/code&gt; blocks on all three. The predicted hole assumed an implementation that reads &lt;code&gt;status === 'FAIL'&lt;/code&gt;, which is not what is there. I mention it because the prediction was reasonable and running it was faster than arguing about it.&lt;/p&gt;

&lt;h2&gt;
  
  
  One check you can run
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://dev.to/debashish_ghosal"&gt;Debashish Ghosal&lt;/a&gt; &lt;a href="https://dev.to/kenielzep97/comment/3def2"&gt;proposed this&lt;/a&gt; in the comments on the last piece:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Throw random strings into &lt;code&gt;event.notes&lt;/code&gt; during test runs. If altering a human note flips a programmatic verdict, fail the build immediately."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;One note on scope before the code: this harness targets the &lt;strong&gt;Python&lt;/strong&gt; classifier from the previous article — &lt;a href="https://github.com/keniel13-ui/ai-memory-judgment-demo/blob/main/claim_24/mandate_cell7.py" rel="noopener noreferrer"&gt;&lt;code&gt;claim_24/mandate_cell7.py&lt;/code&gt;&lt;/a&gt; in a different repository. It does not test the JavaScript validator above. Two codebases, one failure class.&lt;/p&gt;

&lt;p&gt;Complete file. Python 3, no dependencies, run it as-is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;string&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;fuzz_note_independence&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;classify&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;row&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;note_field&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;notes&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;seed&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Perturb only the prose. If the verdict moves, the prose is load-bearing.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;baseline&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;classify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;row&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="n"&gt;rnd&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Random&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;seed&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;row&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;note_field&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rnd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;choice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;string&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;printable&lt;/span&gt;&lt;span class="p"&gt;[:&lt;/span&gt;&lt;span class="mi"&gt;95&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
                                &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rnd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;randint&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;80&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;classify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;baseline&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;note_field&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="nf"&gt;classify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;baseline&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;probe&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;TTL EXPIRED&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ttl expired&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;not ttl expired&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                  &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;resolved: ttl expired last week&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
        &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;row&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;note_field&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;probe&lt;/span&gt;
        &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;got&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;classify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;probe&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;raised &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;type&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="n"&gt;__name__&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;baseline&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;got&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;baseline&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;probe&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;got&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;baseline&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;baseline&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;classify_defective&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ev&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;                  &lt;span class="c1"&gt;# control flow reads the prose
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ttl expired&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ev&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;notes&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SKIPPED_TTL_EXPIRED&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CONSULTED&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;classify_note_independent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ev&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;           &lt;span class="c1"&gt;# control flow reads a typed field
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;ev&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reason_code&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;TTL_EXPIRED&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SKIPPED_TTL_EXPIRED&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CONSULTED&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="n"&gt;row&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reason_code&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;TTL_EXPIRED&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ttl_remaining_hours&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
       &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;notes&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;grant ttl expired during consult&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;fn&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;defective&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;classify_defective&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
                 &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;note-independent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;classify_note_independent&lt;/span&gt;&lt;span class="p"&gt;)):&lt;/span&gt;
    &lt;span class="n"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;note&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;got&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;base&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;fuzz_note_independence&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;fn&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;row&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;18&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;PASS&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;ok&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;FAIL&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;  baseline=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;base&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
          &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;""&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;ok&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;  note=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;note&lt;/span&gt;&lt;span class="si"&gt;!r}&lt;/span&gt;&lt;span class="s"&gt; -&amp;gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;got&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="c1"&gt;# The independence harness above cannot catch negation, because on this row
# "not ttl expired" still contains "ttl expired" and returns the same verdict
# as the baseline. Negation needs a row whose typed reason is NOT expired:
&lt;/span&gt;&lt;span class="n"&gt;negated&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reason_code&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CONSULTED&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;notes&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;not ttl expired&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;negation           defective -&amp;gt;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;classify_defective&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;negated&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
      &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;| note-independent -&amp;gt;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;classify_note_independent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;negated&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;defective          FAIL  baseline=SKIPPED_TTL_EXPIRED  note='R5x$!PCZJ-r#hAhc&amp;lt;w...' -&amp;gt; CONSULTED
note-independent   PASS  baseline=SKIPPED_TTL_EXPIRED
negation           defective -&amp;gt; SKIPPED_TTL_EXPIRED | note-independent -&amp;gt; CONSULTED
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things worth being precise about, because I got both wrong in a draft of this.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The negation probe inside the harness catches nothing.&lt;/strong&gt; &lt;code&gt;"not ttl expired"&lt;/code&gt; still contains &lt;code&gt;"ttl expired"&lt;/code&gt;, and the baseline row already classifies as expired, so the verdict does not move and the harness reports no change. Negation needs the separate row at the bottom, where the authoritative field says &lt;code&gt;CONSULTED&lt;/code&gt; and the grep says otherwise. That is the line that shows the defect.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And this establishes note independence only.&lt;/strong&gt; It does not establish that the typed field is right. A typed field can lie as cleanly as a sentence — a grant expired by one second stored as &lt;code&gt;-0.0&lt;/code&gt;, and &lt;code&gt;-0.0 &amp;gt;= 0&lt;/code&gt; is &lt;code&gt;True&lt;/code&gt; in Python. Which is how the last piece started.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is still open
&lt;/h2&gt;

&lt;p&gt;anp2network's objection is the one I have not answered:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Every field on that row has the same author... Each one worked by making two fields disagree. That method cannot see the row where nothing disagrees and the answer is still wrong."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Every check above works by making two views disagree. A commit message contradicts its diff. A candidate complete-set count contradicts the repository state it would have created if merged. A verdict contradicts its own inputs.&lt;/p&gt;

&lt;p&gt;That method is blind to the case where nothing contradicts anything. If a timestamp is stamped when a gate &lt;em&gt;consumes&lt;/em&gt; a grant rather than when the issuer &lt;em&gt;issues&lt;/em&gt; it, every field agrees, every recomputation is clean, every contract passes, and the verdict is wrong — because the error arrived before the first field was written.&lt;/p&gt;

&lt;p&gt;My read of their proposed direction is that independence is a property of who could have been &lt;strong&gt;compelled&lt;/strong&gt;: you are not looking for a willing second witness, you are looking for bytes some other party already wrote, for their own reasons, that a claim can be bound to.&lt;/p&gt;

&lt;p&gt;I have not built that.&lt;/p&gt;







&lt;p&gt;&lt;strong&gt;Check any of it yourself.&lt;/strong&gt; Every timestamp in this piece comes from a public endpoint that needs no account:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Commit times — &lt;code&gt;git show -s --format=%cI &amp;lt;sha&amp;gt;&lt;/code&gt; after cloning &lt;a href="https://github.com/keniel13-ui/self-correcting-integration-maintainer" rel="noopener noreferrer"&gt;the repo&lt;/a&gt;, or the commit pages linked above&lt;/li&gt;
&lt;li&gt;Review timing — &lt;code&gt;GET /repos/keniel13-ui/self-correcting-integration-maintainer/pulls/1/comments&lt;/code&gt; and &lt;code&gt;/reviews&lt;/code&gt;; the finding on the defect carries &lt;code&gt;original_commit_id: 4a8e6c9…&lt;/code&gt; and &lt;code&gt;created_at: 2026-08-26T01:30:49Z&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Comment times — &lt;code&gt;GET https://dev.to/api/comments/3df2g&lt;/code&gt; returns &lt;code&gt;created_at: 2026-08-24T13:41:44Z&lt;/code&gt;; &lt;code&gt;3dfeg&lt;/code&gt; and &lt;code&gt;3def2&lt;/code&gt; the same way. The DEV page shows only the date, so the API is where the hour lives&lt;/li&gt;
&lt;li&gt;Repository creation — &lt;code&gt;GET https://api.github.com/repos/keniel13-ui/self-correcting-integration-maintainer&lt;/code&gt; returns &lt;code&gt;created_at: 2026-08-25T23:47:38Z&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;All times converted to EDT (UTC−4). Verified 2026-08-30.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>testing</category>
      <category>architecture</category>
      <category>codereview</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
