<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: John</title>
    <description>The latest articles on DEV Community by John (@hexisteme).</description>
    <link>https://dev.to/hexisteme</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3997679%2F1757a270-952f-4fbb-b529-231fade996c5.jpeg</url>
      <title>DEV Community: John</title>
      <link>https://dev.to/hexisteme</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/hexisteme"/>
    <language>en</language>
    <item>
      <title>pytest Failed. Every Collected Test Still Had a Passing Run.</title>
      <dc:creator>John</dc:creator>
      <pubDate>Wed, 30 Sep 2026 00:00:06 +0000</pubDate>
      <link>https://dev.to/hexisteme/pytest-failed-every-collected-test-still-had-a-passing-run-119p</link>
      <guid>https://dev.to/hexisteme/pytest-failed-every-collected-test-still-had-a-passing-run-119p</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://hexisteme.github.io/notes/pytest-full-run-failed-every-test-passed.html" rel="noopener noreferrer"&gt;hexisteme notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I had a red pytest command and a folder of green results. Choosing a color for the handoff would have thrown away information.&lt;/p&gt;

&lt;p&gt;The run that attempted everything returned exit code 1. After a carefully bounded retry, the records also showed a successful execution for each of the 1,269 named cases on the collection roster.&lt;/p&gt;

&lt;p&gt;Neither observation cancelled the other.&lt;/p&gt;

&lt;p&gt;This happened during a regression check dated 2026-09-27. The lesson is not a trick for rescuing red CI. It is about writing a completion statement whose nouns still mean what the evidence says.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the question behind the green label
&lt;/h2&gt;

&lt;p&gt;"Did this command finish successfully?" asks about a process. "Which cases have a successful result?" asks about a roster.&lt;/p&gt;

&lt;p&gt;Those questions often receive the same answer, so a dashboard can make them look interchangeable. Once retries enter the picture, they can diverge.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.pytest.org/en/stable/reference/exit-codes.html" rel="noopener noreferrer"&gt;pytest documents the meanings of its exit codes&lt;/a&gt;. A result assembled afterward cannot retroactively change the value returned by the shell.&lt;/p&gt;

&lt;p&gt;That gave me a useful constraint: keep the command failure visible even if a later attempt supplied the missing results. Otherwise, someone reading the handoff could mistake an accounting result for an observation of everything working together.&lt;/p&gt;

&lt;p&gt;The distinction is easier to keep when the report contains both answers, instead of forcing the reader to infer one from the other.&lt;/p&gt;

&lt;h2&gt;
  
  
  There was an environmental cause to investigate
&lt;/h2&gt;

&lt;p&gt;My initial attempt produced 1,199 passes alongside 31 skips, 3 failures, and 36 errors. The command returned 1. It had not been given the browser-runtime setting.&lt;/p&gt;

&lt;p&gt;An already installed, pinned Playwright runtime was available. I selected that cache and tried the complete roster again inside the restricted environment. This removed the skips, but it did not make the invocation successful: the results were 1,199 passes, 3 failures, and 67 errors, with exit code 1 again.&lt;/p&gt;

&lt;p&gt;At that point, calling the problem "environmental" still did not close it. I needed to find the operations being denied and establish what a retry would actually exercise.&lt;/p&gt;

&lt;p&gt;The browser component accounted for 69 cases. macOS denied its launch through MachPort. The other case belonged to an explorer fixture whose temporary web server was unable to bind localhost.&lt;/p&gt;

&lt;p&gt;Changing that boundary meant changing the conditions under which evidence would be collected. It deserved a precise scope, not an incidental mention buried under the final green number.&lt;/p&gt;

&lt;h2&gt;
  
  
  Retry the named group, not the permission model
&lt;/h2&gt;

&lt;p&gt;I selected the unsuccessful cases by their exact node IDs. Before permitting another attempt, I inspected the fixture behavior associated with that selection.&lt;/p&gt;

&lt;p&gt;The browser consumed fabricated pages held in memory. Fixture responses supplied the content; requests that would leave the test were stopped. The explorer served disposable data from the machine. This selection involved no real sign-in, no fetching through a production API, no use of actual credentials from Keychain, and no publishing.&lt;/p&gt;

&lt;p&gt;The restricted retry reproduced the problem. The authorized retry then allowed only the necessary browser launch and local bind operations.&lt;/p&gt;

&lt;p&gt;Here is the full sequence. A pass total is not the process result; the right-hand column preserves that difference.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Attempt&lt;/th&gt;
&lt;th&gt;Test outcomes&lt;/th&gt;
&lt;th&gt;Process exit code&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Before selecting the installed browser cache&lt;/td&gt;
&lt;td&gt;1199 passes; 31 skips; 3 failures; 36 errors&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Complete roster with the cache selected&lt;/td&gt;
&lt;td&gt;1199 passes; 0 skips; 3 failures; 67 errors&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Selected group inside the original restrictions&lt;/td&gt;
&lt;td&gt;0 passes; 0 skips; 3 failures; 67 errors&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Identical selection with bounded launch and bind permission&lt;/td&gt;
&lt;td&gt;70 passes; 0 skips; 0 failures; 0 errors&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;No assertions were softened. No test was deleted from the selection. No new browser installation was needed, and the broad command did not receive blanket permission.&lt;/p&gt;

&lt;p&gt;The last row answers a specific question: these 70 cases worked when those local operations were allowed. It does not answer whether they would work inside the restrictions that produced the previous row.&lt;/p&gt;

&lt;p&gt;That qualification belongs next to the outcome. A retry without its changed conditions is an incomplete observation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The sum was a clue, not the reconciliation
&lt;/h2&gt;

&lt;p&gt;Adding 1,199 and 70 gives 1,269. That arithmetic would still be true if both groups included the same case and another case never appeared.&lt;/p&gt;

&lt;p&gt;I could not infer membership from totals.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.pytest.org/en/stable/how-to/usage.html#specifying-which-tests-to-run" rel="noopener noreferrer"&gt;pytest accepts individual node IDs&lt;/a&gt;, including selections for parameterized cases. It also supports &lt;a href="https://docs.pytest.org/en/stable/how-to/output.html#creating-junitxml-format-files" rel="noopener noreferrer"&gt;JUnit XML output&lt;/a&gt;. Those mechanisms supply useful inputs, but they do not decide what a combined result means.&lt;/p&gt;

&lt;p&gt;The roster defined what needed a result. I associated each JUnit entry with a specific member, treating variations of a parameter as distinct cases rather than interchangeable names. Each outcome still pointed back to the attempt that produced it.&lt;/p&gt;

&lt;p&gt;The final comparison left 0 expected members without a pass, 0 deliberately excluded members, and 0 members covered only by a skip. In particular, the initial 31 skips earned no credit merely for having appeared in the output.&lt;/p&gt;

&lt;p&gt;A conceptual sketch is short:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;missing&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;collected_ids&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;passed_ids&lt;/span&gt;
&lt;span class="n"&gt;extra&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;passed_ids&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;collected_ids&lt;/span&gt;

&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;missing&lt;/span&gt;
&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;extra&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is explanatory pseudocode, not the parser I ran. Before any set operation, the mapping must reject ambiguous identifiers and duplicate entries within an attempt. Converting a questionable list to a set can conceal the very defect the check is supposed to catch.&lt;/p&gt;

&lt;p&gt;I wanted an answer to "where is this case's successful execution?", not just a plausible total at the bottom of a log.&lt;/p&gt;

&lt;h2&gt;
  
  
  A commit name would have described the wrong object
&lt;/h2&gt;

&lt;p&gt;The checkout was not pristine when this verification began. It already included two changes made concurrently with the surrounding work.&lt;/p&gt;

&lt;p&gt;If I had attached only a commit ID, a later reader could reasonably assume the measurements referred to that committed tree. They did not. The object under examination was the checkout as observed, including those modifications.&lt;/p&gt;

&lt;p&gt;The receipt included a fingerprint for each of 220 Git-tracked files, plus 7 markers used in operation. Comparing those fingerprints found no drift between attempts or when I checked the handoff.&lt;/p&gt;

&lt;p&gt;I did not alter application code or test code while resolving this verification. That statement is compatible with pre-existing modifications; it is not a claim that the worktree was clean.&lt;/p&gt;

&lt;p&gt;Nor do matching source bytes make the environments identical. Assigning the browser cache and permitting local launch and bind were real changes in execution conditions. A source fingerprint is useful precisely because it answers a narrower question, not because it proves every kind of sameness.&lt;/p&gt;

&lt;p&gt;This also limits reuse. A different checkout or changed dependency cannot inherit the result just because the old report is still nearby.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I could say, and what I could not
&lt;/h2&gt;

&lt;p&gt;The reconciliation gave every expected case a successful execution. It did not show the entire roster succeeding within a single process.&lt;/p&gt;

&lt;p&gt;A fixture may leave something behind for its neighbors. Scheduling or resource competition may change when cases share a process. Separate invocations can miss that behavior; this record does not measure it. Nor did I measure the likelihood of seeing the failures again.&lt;/p&gt;

&lt;p&gt;The denominator is a roster of named cases, not lines of executable Python. A statement about which cases succeeded cannot tell me which lines were exercised.&lt;/p&gt;

&lt;p&gt;And a synthetic browser fixture is not a live-account check. Permission to launch a browser for local content does not become permission to exercise credentials or publish something.&lt;/p&gt;

&lt;p&gt;I have written about a neighboring problem in &lt;a href="https://hexisteme.github.io/notes/green-tests-prove-behavior-not-reachability.html" rel="noopener noreferrer"&gt;Green Tests Prove Behavior, Not Reachability&lt;/a&gt;: a gate can behave correctly in isolation while being absent from the dispatch path. This case concerns a different object. I had an explicit roster, and needed an execution record for each member of it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://hexisteme.github.io/notes/a-check-that-never-ran-is-not-passing.html" rel="noopener noreferrer"&gt;A Check That Never Ran Is Not Passing&lt;/a&gt; addresses a missing evaluable input. Here, the successful observations already existed. The work was to reconcile them without crediting omissions, or claiming that a changed environment was the original one.&lt;/p&gt;

&lt;p&gt;Those distinctions help keep similar-sounding failure stories from turning into the same generic lesson.&lt;/p&gt;

&lt;h2&gt;
  
  
  The handoff needs more than a status word
&lt;/h2&gt;

&lt;p&gt;The custom result in my receipt was &lt;code&gt;FULL_COLLECTED_COVERAGE_BY_EXPLICIT_SHARDS&lt;/code&gt;. Alongside it, &lt;code&gt;whole_pytest_invocation_exit_zero&lt;/code&gt; stayed false.&lt;/p&gt;

&lt;p&gt;That is a description of this receipt, not a pytest feature or a reason to override the shell.&lt;/p&gt;

&lt;p&gt;For a future report, I would reserve room for the process outcome, the expected roster, the evidence attached to each member, the bytes examined, and the conditions changed by a retry. I would also state which decision the report supports.&lt;/p&gt;

&lt;p&gt;A verification statement and permission to merge or release are not interchangeable. Neither should disappear behind a green badge.&lt;/p&gt;

&lt;p&gt;This is a proposed reporting habit, not a claim that I have installed a general-purpose reconciliation framework across every project. Its useful property would be refusing completion when an identifier cannot be resolved or an expected result is absent.&lt;/p&gt;

&lt;p&gt;Preserving the earlier red attempts matters too. I retained the original logs and XML in a lossless archive and checked them against their hashes. A readable summary should sit beside those artifacts, not replace them with a tidied history.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep a way to revoke the conclusion
&lt;/h2&gt;

&lt;p&gt;The result stops being defensible if a required member has no successful execution, if an identifier points to more than one case, or if I discover that the observations came from different checkouts.&lt;/p&gt;

&lt;p&gt;The same applies if the supposedly local retry actually used an unauthorized production operation or live account. Its permitted scope is part of what makes the evidence usable.&lt;/p&gt;

&lt;p&gt;A later all-green invocation would be welcome additional evidence. It would not erase the red invocations already recorded, and these separate attempts do not authorize relabeling a failed CI job.&lt;/p&gt;

&lt;p&gt;The handoff I could defend was simple:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The complete command returned failure. The named cases all had successful executions under the recorded conditions. The record preserves both facts.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Prepared with AI assistance from my preserved regression-verification records; source and execution evidence rechecked on 2026-09-27.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Email list for these notes: &lt;a href="https://hexisteme.beehiiv.com/?modal=signup&amp;amp;utm_source=devto&amp;amp;utm_campaign=notes-engineering&amp;amp;utm_content=pytest-full-run-failed-every-test-passed" rel="noopener noreferrer"&gt;hexisteme.beehiiv.com&lt;/a&gt; — no issue has gone out yet, so you would be on it before the first one. No welcome sequence, no course, no upsell.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;More notes at &lt;a href="https://hexisteme.github.io/notes/" rel="noopener noreferrer"&gt;hexisteme.github.io/notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>testing</category>
      <category>python</category>
      <category>debugging</category>
    </item>
    <item>
      <title>I Rejected the Same dev.to Reliability Fix on One Path and Shipped It on Another</title>
      <dc:creator>John</dc:creator>
      <pubDate>Tue, 29 Sep 2026 00:00:06 +0000</pubDate>
      <link>https://dev.to/hexisteme/i-rejected-the-same-devto-reliability-fix-on-one-path-and-shipped-it-on-another-an7</link>
      <guid>https://dev.to/hexisteme/i-rejected-the-same-devto-reliability-fix-on-one-path-and-shipped-it-on-another-an7</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://hexisteme.github.io/notes/same-reliability-fix-opposite-decisions.html" rel="noopener noreferrer"&gt;hexisteme notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Readers made a reasonable suggestion about my publishing automation: verify delivery from outside the process that claims it succeeded. I had an article publisher and a comment replier. Both sent content to dev.to. It would have been easy to apply the suggestion to both and call the result consistent.&lt;/p&gt;

&lt;p&gt;My audit reached different decisions. I added external observation to the manual comment path and deferred a new polling loop for article publication. The useful distinction was what each checker could get wrong, what would happen afterward, and which operational decision another observation would change.&lt;/p&gt;

&lt;p&gt;This account is based on my delivery audit dated 2026-08-08. It describes that decision and its limits, rather than asserting the same configuration remains correct indefinitely. During preparation for publication on 2026-09-26 I also checked the surviving local records. That second pass narrowed a numerical claim in the original report.&lt;/p&gt;

&lt;h2&gt;
  
  
  Define what a failed check means
&lt;/h2&gt;

&lt;p&gt;The comment incident was a false negative: the reply had been posted, but the checker said it had not landed. Its comparison did not adequately account for the difference between authored Markdown and rendered content.&lt;/p&gt;

&lt;p&gt;That error had an operational consequence. A reply judged absent stayed pending. Someone following the pending list could answer again, producing a duplicate. The audit connects this mechanism to previously removed duplicate replies; it does not establish that every duplicate had the same cause.&lt;/p&gt;

&lt;p&gt;The surviving record preserves the useful evidence: a note saying that an apparent failure caused by Markdown versus rendering was corrected after a live API check. Its current verification flag is true. That flag now describes the corrected state; the note preserves the earlier mistake.&lt;/p&gt;

&lt;p&gt;The original report presented the comment result as 1/44 and attributed the false flag to this incident. A publication-time check could reconstruct 44 earlier autoreply records, but their remaining false flag belonged to a different entry. I therefore do not use 1/44 as a measured rate of this failure. A mutable operational ledger and a historical incident count answer different questions.&lt;/p&gt;

&lt;p&gt;The recorded false negative is enough to motivate fixing its comparison path. An unreliable rate estimate would add apparent precision without improving that argument.&lt;/p&gt;

&lt;h2&gt;
  
  
  The article audit asked a different question
&lt;/h2&gt;

&lt;p&gt;For articles, I compared success claims with external publication state. The historical audit reports 0 mismatches among 46 successful publication log entries. The denominator is successful log entries, not every scheduled run, every network attempt, or every possible delivery failure.&lt;/p&gt;

&lt;p&gt;The retained log still supports that success-entry count through the morning run on the audit date. The external comparison result comes from the contemporaneous audit report; I have not reconstructed its historical API snapshot for this post. That is a boundary on reproducibility, not permission to silently turn the report into a fresh measurement.&lt;/p&gt;

&lt;p&gt;The report also distinguishes known failures from false success. The publisher's DNS failures were reported as failures and left work for a later run. They were not examples of a success claim concealing an absent article.&lt;/p&gt;

&lt;p&gt;Zero observed mismatches does not establish zero risk. These observations also share one deployment and its operating conditions, so treating them as independent trials would need justification. I do not infer a general reliability guarantee from the sample.&lt;/p&gt;

&lt;p&gt;The severe counterexample remained possible: if the publisher falsely claimed success and moved the file out of the queue, an absent article could disappear from ordinary retry handling. That is why deferring the polling loop required an explicit reversal condition. Lack of an observed incident did not make silent loss harmless.&lt;/p&gt;

&lt;h2&gt;
  
  
  An extra observation should change an action
&lt;/h2&gt;

&lt;p&gt;The comment verifier had an immediate decision to make: had this reply appeared, or should the work remain pending? Its retry policy also contained tunable choices: attempts=3 and a pause of 5 seconds. Measuring when a reply became visible could inform that budget.&lt;/p&gt;

&lt;p&gt;The article path did not have a comparable short polling budget in the audited design. Reported failures waited for the next scheduled run. A new polling loop would introduce another timeout policy, another interpretation of temporary absence, and another failure path to maintain.&lt;/p&gt;

&lt;p&gt;I deferred that addition because the audit had found a concrete comment-checking defect and no demonstrated article success-report mismatch. This was a local prioritization decision. I did not measure the cost of a lost article, calculate an expected loss, or establish that the polling loop could never repay its complexity.&lt;/p&gt;

&lt;p&gt;That qualification matters. A system handling more consequential delivery could reasonably require external confirmation before any false success had been observed. My decision depended on this workflow and included a trigger to revisit it.&lt;/p&gt;

&lt;p&gt;The comparison I would reuse is small:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;Article publication in the audit&lt;/th&gt;
&lt;th&gt;Manual comment delivery in the audit&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;What was the observed defect?&lt;/td&gt;
&lt;td&gt;No mismatch reported in the checked success entries&lt;/td&gt;
&lt;td&gt;Posted content could be judged absent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What happened after the adverse result?&lt;/td&gt;
&lt;td&gt;A reported failure waited for a later run&lt;/td&gt;
&lt;td&gt;A reply stayed pending and could be duplicated&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What could new observations tune?&lt;/td&gt;
&lt;td&gt;No existing short polling budget&lt;/td&gt;
&lt;td&gt;Attempt count and pause duration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What did I choose?&lt;/td&gt;
&lt;td&gt;Defer the added polling loop&lt;/td&gt;
&lt;td&gt;Share and use the external reply verifier&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The table does not make the risks equivalent. It exposes why the same design suggestion led to different work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Share the comparison, then observe the retries
&lt;/h2&gt;

&lt;p&gt;The normalization fix already existed in the automatic reply path but had not reached the manual path. Copying it would have preserved the condition that allowed the implementations to diverge.&lt;/p&gt;

&lt;p&gt;I moved the comparison and landing verification into the shared comment module. The automatic path delegated to it, and the manual path used the external observation. When that observation could not decide, the manual path retained an in-page fallback and announced that fallback. A weaker observation should remain visible to the operator interpreting the result.&lt;/p&gt;

&lt;p&gt;The comparison also needed to handle links: the authored Markdown could contain a URL that was absent from the rendered visible text. Merely discarding punctuation did not make those strings equivalent. The fix addressed this mismatch in the shared implementation.&lt;/p&gt;

&lt;p&gt;I have left the original report's rendering-exposure count out of this post because I did not recover the corresponding historical sample. Exposure to a string transformation would not, by itself, be a count of failed deliveries anyway. The mechanism and the recorded false negative support a narrower claim.&lt;/p&gt;

&lt;p&gt;For retries, I kept the existing constants and started recording attempt_seen, attempts_budget, elapsed_s, and landed. That made later calibration possible without presenting a newly chosen constant as empirical. The observation record separates whether a reply was seen from how much budget the verifier consumed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make the deferral reversible
&lt;/h2&gt;

&lt;p&gt;The article decision has a sharp falsifier: observe one publication success report that disagrees with external publication state, and reopen the rejected verification work. Evaluation began with the audit on 2026-08-08; the trigger applies whenever a later mismatch is found. A check is only useful if the workflow can surface that disagreement, so this condition is not a claim of continuous detection coverage.&lt;/p&gt;

&lt;p&gt;The comment decision has a different review trigger: after at least 30 landing observations, if every observation succeeded on its first attempt, reconsider whether the three-attempt budget is excessive. That is a prompt for reviewing the budget, not proof that later attempts will never be needed. The observation mix and any failed checks still matter.&lt;/p&gt;

&lt;p&gt;These conditions keep the decisions attached to evidence. They also prevent consistency from becoming a reason to build the same mechanism everywhere. In this audit, one path supplied a known comparison failure and a measurable retry budget; the other supplied a bounded clean sample and a serious unobserved failure mode. Different actions were defensible only while those premises stayed visible.&lt;/p&gt;

&lt;p&gt;The reusable lesson is to specify the decision before adding a checker: which wrong claim are you trying to catch, what happens if you miss it, and what would a new observation make you do? That question gave me a smaller repair and a clear reason to reverse the part I deferred.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Prepared with AI assistance from my delivery audit and surviving operational records.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;License: CC BY 4.0.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Email list for these notes: &lt;a href="https://hexisteme.beehiiv.com/?modal=signup&amp;amp;utm_source=devto&amp;amp;utm_campaign=notes-engineering&amp;amp;utm_content=same-reliability-fix-opposite-decisions" rel="noopener noreferrer"&gt;hexisteme.beehiiv.com&lt;/a&gt; — no issue has gone out yet, so you would be on it before the first one. No welcome sequence, no course, no upsell.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;More notes at &lt;a href="https://hexisteme.github.io/notes/" rel="noopener noreferrer"&gt;hexisteme.github.io/notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>reliability</category>
      <category>debugging</category>
      <category>automation</category>
      <category>devops</category>
    </item>
    <item>
      <title>We Built Three NEXT ART AI Drafts Before Reading the Submission Rules</title>
      <dc:creator>John</dc:creator>
      <pubDate>Mon, 28 Sep 2026 00:00:06 +0000</pubDate>
      <link>https://dev.to/hexisteme/we-built-three-next-art-ai-drafts-before-reading-the-submission-rules-30n9</link>
      <guid>https://dev.to/hexisteme/we-built-three-next-art-ai-drafts-before-reading-the-submission-rules-30n9</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://hexisteme.github.io/notes/we-rendered-three-drafts-before-reading-the-submission-rules.html" rel="noopener noreferrer"&gt;hexisteme notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We had three AI artwork drafts when we checked whether our production method could satisfy the application.&lt;/p&gt;

&lt;p&gt;The work used attention from an encoder model. The submission materials asked us to identify the generative AI tool, its version, and the principal prompts. We had a method for making the image. We did not yet have an evidenced answer to the fields describing how the image had been made.&lt;/p&gt;

&lt;p&gt;That discovery changed the proposed production method. The expensive mistake was the order: we had started making the artifact before inspecting the evidence its destination would require.&lt;/p&gt;

&lt;p&gt;This account comes from my project specification dated September sixth. It records three drafts and a proposed change of method. It does not establish a finished submission, an organizer's rejection, or a measured amount of wasted rendering time.&lt;/p&gt;

&lt;h2&gt;
  
  
  The form exposed a design dependency
&lt;/h2&gt;

&lt;p&gt;The destination was NEXT ART AI. I rechecked its &lt;a href="https://www.next-art-ai.com/ko/apply/overview" rel="noopener noreferrer"&gt;official overview&lt;/a&gt; and &lt;a href="https://www.next-art-ai.com/ko/apply/new/terms" rel="noopener noreferrer"&gt;submission terms&lt;/a&gt; on 2026-09-26. The overview describes works made using generative AI platforms. The terms request the generative AI tools and versions, editing software, an account of AI use, and principal prompts.&lt;/p&gt;

&lt;p&gt;The overview also permits a broader set of ingredients, including external assets, game engines, AI frameworks, open source, and editing software. It describes AI-assisted creative processes broadly. That matters: I cannot turn the narrower wording into a definitive ruling that an organizer would reject every encoder-based work.&lt;/p&gt;

&lt;p&gt;Our actual problem was more concrete. The recorded method used &lt;code&gt;klue/roberta-small&lt;/code&gt; attention, and the specification did not record a generative stage or generation prompts for those drafts. We could describe the encoder inputs and visualization process honestly. We could not claim they were evidence of a generative stage we had not documented.&lt;/p&gt;

&lt;p&gt;Whether the organizer would accept that method remained a question for the organizer. Pretending a missing answer was ordinary final-day paperwork would not resolve it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Some application fields can only be answered upstream
&lt;/h2&gt;

&lt;p&gt;A title can usually be revised after rendering. A description can be edited. A field asking for the model and prompts that produced the work is different: its answer depends on what actually happened during production.&lt;/p&gt;

&lt;p&gt;The same is true of an asset license, a consent record, or a source-data snapshot. You can collect an existing record late. You cannot truthfully reconstruct a missing event by filling in a plausible string.&lt;/p&gt;

&lt;p&gt;This is why reading the rules early is too weak as an operational instruction. People read documents and still postpone their consequences. The useful step is to turn each destination field into a dependency before choosing the production method.&lt;/p&gt;

&lt;p&gt;For this project, a planning table would have looked like this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Destination asks for&lt;/th&gt;
&lt;th&gt;Evidence to retain&lt;/th&gt;
&lt;th&gt;Earliest useful check&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;AI tool and version&lt;/td&gt;
&lt;td&gt;The identity of the model actually used&lt;/td&gt;
&lt;td&gt;Before choosing the method&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Principal prompts&lt;/td&gt;
&lt;td&gt;Inputs from the actual generative stage&lt;/td&gt;
&lt;td&gt;During a minimal trial&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Account of AI use&lt;/td&gt;
&lt;td&gt;What the model did and what later editing changed&lt;/td&gt;
&lt;td&gt;During the trial&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Assets and editing tools&lt;/td&gt;
&lt;td&gt;An inventory linked to the work&lt;/td&gt;
&lt;td&gt;Before final assembly&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These are proposed checks, not a description of a gate that was already installed in our project.&lt;/p&gt;

&lt;p&gt;The important column is the last one. If a required answer depends on how the work is produced, discovering the requirement after production creates a dependency reversal. You are trying to write the provenance first and then change reality to match it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make unresolved eligibility visible before production
&lt;/h2&gt;

&lt;p&gt;A small manifest can expose that reversal. Here is an illustrative shape, not an official application schema:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;destination&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;next-art-ai&lt;/span&gt;
&lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;encoder-attention-visualization&lt;/span&gt;
&lt;span class="na"&gt;generation_evidence&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;absent&lt;/span&gt;
&lt;span class="na"&gt;prompt_record&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;absent&lt;/span&gt;
&lt;span class="na"&gt;eligibility_review&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;unresolved&lt;/span&gt;
&lt;span class="na"&gt;render_permission&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;blocked&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The blocked state is a local workflow decision. It does not claim the competition has rejected the work. It means we have neither evidence that our method satisfies the generative requirement nor an authoritative clarification accepting the method.&lt;/p&gt;

&lt;p&gt;An executable gate should preserve that distinction:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;allow_production&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;manifest&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;manifest&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eligibility_review&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;resolved&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Resolve the method&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s eligibility first&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;manifest&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;requires_generation_evidence&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;manifest&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;generation_evidence&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
            &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Missing record of the generative stage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;manifest&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prompt_record&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
            &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Missing prompts from the actual run&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This snippet only checks the presence of declared records. A production implementation would also have to load those records and bind them to the selected method. Otherwise, arbitrary nonempty text would pass. It cannot decide an organizer's interpretation or certify that an operator's account is truthful.&lt;/p&gt;

&lt;p&gt;The smallest useful experiment happens before a finished artwork: run a modest sample with the intended method, then attempt to complete the relevant application fields from its retained records. If doing so requires inventing a prompt history, the method is not ready for final production.&lt;/p&gt;

&lt;h2&gt;
  
  
  The repair affected the concept, not just the export
&lt;/h2&gt;

&lt;p&gt;Our specification proposed replacing the encoder-only approach with a generative instruct model. That was a design decision recorded in the plan, not proof that the replacement was completed.&lt;/p&gt;

&lt;p&gt;The change also offered a stronger version of the artwork. One concept represented a message degrading as it passed between speakers. In the draft method, the degradation was simulated. The proposed method would ask a model to retell the previous output and retain the resulting text.&lt;/p&gt;

&lt;p&gt;That would make the changing message an observed behavior of the model rather than a visual metaphor backed by a hand-written degradation rule. It would also produce actual inputs and outputs to describe in the application.&lt;/p&gt;

&lt;p&gt;But a plausible repair is still a hypothesis. The source record explicitly anticipates a failure: if the text collapses immediately, there may be no gradual change worth showing. That result would require another creative decision, not a declaration that meeting the application fields made the work successful.&lt;/p&gt;

&lt;p&gt;The change therefore carries a cost we can describe but have not measured: the method, its evidence trail, and the meaning of the artwork all need reconsideration. I have no supported figure for hours saved, render time wasted, or acceptance probability.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the preflight separate from the final artifact check
&lt;/h2&gt;

&lt;p&gt;I previously wrote about &lt;a href="https://hexisteme.github.io/notes/test-the-artifact-not-the-pipeline.html" rel="noopener noreferrer"&gt;a renderer dropping what its script requested&lt;/a&gt;. That incident required tests of the produced artifact. This one sits earlier: it asks whether the chosen production method can supply the destination's required evidence at all.&lt;/p&gt;

&lt;p&gt;Both checks are necessary. A preflight can approve a viable method while the renderer later produces a broken image. A beautiful final image can emerge from a method whose eligibility remains unresolved.&lt;/p&gt;

&lt;p&gt;For the next submission, I would put the field-to-evidence table beside the first concept sketch, attempt the smallest real run, and complete the consequential fields from that run. Only then would I invest in the final production pass.&lt;/p&gt;

&lt;h2&gt;
  
  
  What would change this conclusion?
&lt;/h2&gt;

&lt;p&gt;My evaluation date is 2026-09-26. The local eligibility block should be reconsidered if the organizer supplies a clear interpretation accepting our encoder-based method and its honest account of inputs. That is the trigger for revisiting the method decision, rather than treating our cautious interpretation as settled policy.&lt;/p&gt;

&lt;p&gt;The broader workflow recommendation also has a limit. If every consequential field can be completed truthfully from records the existing method already retains, this preflight may add little. It earns its place when a field exposes a missing event or a method choice that would be costly to revisit later.&lt;/p&gt;

&lt;p&gt;The lesson I can support is modest: our submission fields revealed a production dependency after three drafts existed. Bringing that dependency forward would have made the unresolved choice visible before we committed further work. It would not have guaranteed eligibility, artistic quality, or an award.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This article was prepared with AI assistance from my project specification and a recheck of the official submission pages.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;License: CC BY 4.0.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Email list for these notes: &lt;a href="https://hexisteme.beehiiv.com/?modal=signup&amp;amp;utm_source=devto&amp;amp;utm_campaign=notes-engineering&amp;amp;utm_content=we-rendered-three-drafts-before-reading-the-submission-rules" rel="noopener noreferrer"&gt;hexisteme.beehiiv.com&lt;/a&gt; — no issue has gone out yet, so you would be on it before the first one. No welcome sequence, no course, no upsell.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;More notes at &lt;a href="https://hexisteme.github.io/notes/" rel="noopener noreferrer"&gt;hexisteme.github.io/notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>testing</category>
    </item>
    <item>
      <title>Your SKILL.md Was Valid. Codex Ran Out of File Descriptors</title>
      <dc:creator>John</dc:creator>
      <pubDate>Sat, 26 Sep 2026 11:38:44 +0000</pubDate>
      <link>https://dev.to/hexisteme/your-skillmd-was-valid-codex-ran-out-of-file-descriptors-ln4</link>
      <guid>https://dev.to/hexisteme/your-skillmd-was-valid-codex-ran-out-of-file-descriptors-ln4</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://hexisteme.github.io/notes/your-skill-md-was-valid-codex-ran-out-of-file-descriptors.html" rel="noopener noreferrer"&gt;hexisteme notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The warning named my documents: Codex had skipped loading 13 skills because their &lt;code&gt;SKILL.md&lt;/code&gt; files were invalid. The detail beneath the heading named a different failure: &lt;code&gt;failed to read file: Too many open files (os error 24)&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That distinction determined the repair. A file that could not be opened had not necessarily failed a Markdown or schema check. Editing its contents would not answer why the process could not read it.&lt;/p&gt;

&lt;p&gt;This is a postmortem of my September incident, based on the remediation record dated 2026-09-21. The observations below belong to that run. They are not a claim about every Codex release or every warning carrying the word “invalid.”&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the nested error
&lt;/h2&gt;

&lt;p&gt;The useful signal was &lt;code&gt;EMFILE&lt;/code&gt;, the process-level “too many open files” error. “File” in this context includes descriptors consumed by communication channels. My setup had stdio MCP servers attached to a long-lived Desktop app-server, so ordinary skill documents were sharing a resource budget with pipes.&lt;/p&gt;

&lt;p&gt;The diagnostic record contained several different measurements:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Surface measured&lt;/th&gt;
&lt;th&gt;Observation before remediation&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Shell soft NOFILE limit&lt;/td&gt;
&lt;td&gt;256&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shell hard NOFILE limit&lt;/td&gt;
&lt;td&gt;unlimited&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kernel-reported maximum number of files per process&lt;/td&gt;
&lt;td&gt;92,160&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Desktop app-server numeric descriptors&lt;/td&gt;
&lt;td&gt;274&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Desktop app-server PIPE entries&lt;/td&gt;
&lt;td&gt;201&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Desktop app-server direct children&lt;/td&gt;
&lt;td&gt;67&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Globally enabled MCP servers&lt;/td&gt;
&lt;td&gt;23&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These are separate surfaces. In particular, the shell's soft limit was not a measurement of the already-running Desktop process's effective limit. Seeing 274 descriptors in that process does not establish that it was running under the shell's 256 limit. I would need the target process's own limit to make that claim.&lt;/p&gt;

&lt;p&gt;What the snapshot did show was a substantial pipe footprint alongside the read failures. The process inventory also showed MCP groups corresponding to several sessions. The resource hypothesis therefore had evidence beyond the top-level warning, even though the snapshot alone did not identify the exact descriptor allocation that failed.&lt;/p&gt;

&lt;p&gt;My inventory was also broad: 61 skills under the agent tree, 16 under the Codex skill directory, and 112 in plugin caches. Those counts describe the configured surface, not the number of simultaneously open files. I did not turn them into a concurrency measurement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two pressures met at the same boundary
&lt;/h2&gt;

&lt;p&gt;The incident record connected the symptom to parallel skill loading and the pipes associated with session-specific MCP children. It referenced the upstream discussions in &lt;a href="https://github.com/openai/codex/issues/36755" rel="noopener noreferrer"&gt;issue #36755&lt;/a&gt;, &lt;a href="https://github.com/openai/codex/issues/26984" rel="noopener noreferrer"&gt;issue #26984&lt;/a&gt;, and &lt;a href="https://github.com/openai/codex/issues/37971" rel="noopener noreferrer"&gt;issue #37971&lt;/a&gt;. Those are investigation pointers from the record; their current status is not evidence supplied by this post.&lt;/p&gt;

&lt;p&gt;My working explanation was that loading skills needed transient descriptors while the long-lived process already held a large communication footprint. A document happened to be the next resource consumer denied an open. Its filename made the failure look local to the document, while the nested error implicated the process.&lt;/p&gt;

&lt;p&gt;I treated that explanation as sufficient to test a resource intervention. I did not have a controlled experiment separating the contribution of loader concurrency from the contribution of retained MCP pipes. The successful intervention below supports the resource diagnosis, but it cannot assign each mechanism a share of the failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Change both the budget and the default demand
&lt;/h2&gt;

&lt;p&gt;I changed the CLI launch path to request a soft NOFILE limit of up to 65,536 before starting Codex. The wrapper respects a lower hard limit and preserves its existing behavior if the operating system refuses the change. The domain launcher received the same target.&lt;/p&gt;

&lt;p&gt;There is a limitation in my own evidence here. The verification note records a wrapper probe issuing &lt;code&gt;ulimit -Sn 4096&lt;/code&gt; from a parent whose soft limit was 256. That supports the narrower claim that the launch path could attempt a higher limit. It does not prove that the final 65,536 target was effective in every launched process. Configuration intent and an inherited runtime limit are different facts, and my record does not close that gap.&lt;/p&gt;

&lt;p&gt;The other change reduced routine demand. I cut the global MCP default from 23 servers to 8, disabled the Comfy and video-vision plugins by default, and retained domain profiles for work that needed a larger tool set. The full profile retained the previous broader configuration.&lt;/p&gt;

&lt;p&gt;That gave me an operational choice: a smaller baseline for ordinary work, with broader dependencies selected when required. A high ceiling alone would have left the original pipe demand in place. A smaller baseline alone would not have addressed the CLI's low inherited soft limit.&lt;/p&gt;

&lt;p&gt;The GUI path had its own boundary. My attempt to change the launchd maxfiles limit from the user session was rejected with &lt;code&gt;Operation not permitted&lt;/code&gt;. I did not record that as a successful Desktop limit increase. The changes actually applied were the reduced default demand and the CLI-side launch behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measure the process after the change
&lt;/h2&gt;

&lt;p&gt;After hot reload, the recorded Desktop app-server snapshot had 116 numeric descriptors, 78 PIPE entries, and 26 direct children. The before-and-after snapshots show that the retained footprint fell. They do not establish that an upstream lifetime-management defect was permanently fixed.&lt;/p&gt;

&lt;p&gt;The validation record was broader than the snapshot. Shell syntax checks passed for the wrapper and launcher. All 11 profile configurations passed their parsing and MCP-list checks. The strict doctor run reported 23 checks OK.&lt;/p&gt;

&lt;p&gt;I also ran fresh ephemeral executions with both the base profile and the full profile. Each returned exit 0, produced the requested response marker, and emitted no FD or skill warnings. The check after those runs found no new MCP orphans.&lt;/p&gt;

&lt;p&gt;Those fresh runs are the strongest recovery evidence in the record: the system performed the operation that had been failing. The full-profile run matters because reducing the default server set was not the only path exercised. Still, these were bounded smoke checks, not a long-duration test of session churn. “Recovered in the checked runs” is supported. “Cannot recur” is not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep a way to reject the diagnosis
&lt;/h2&gt;

&lt;p&gt;The remediation record includes a specific falsifier, evaluated on 2026-09-21: if a fresh process with the base profile and a verified NOFILE limit of 65,536 repeatedly fails on the same &lt;code&gt;SKILL.md&lt;/code&gt; files after &lt;code&gt;os error 24&lt;/code&gt; disappears, reclassify those files for a separate syntax or schema defect.&lt;/p&gt;

&lt;p&gt;The word “verified” matters because of the wrapper-probe gap. I would confirm the effective limit before treating that retry as a test of the falsifier. Otherwise, a failed attempt to apply the mitigation could be mistaken for evidence against the diagnosis.&lt;/p&gt;

&lt;p&gt;For recurrence, my runbook starts with the Desktop app-server's pipe count and direct-child count. Growth across sessions under the smaller baseline would reopen the lifetime-management question. Restarting the application is a recovery action in that situation, not proof that the cause has been removed.&lt;/p&gt;

&lt;p&gt;A fresh session also matters after a skipped-skill warning. The record's recovery procedure opens a new session rather than assuming the failed load has been repaired inside the old one. Saving active work and fully restarting Desktop is the stronger reset when the affected process persists.&lt;/p&gt;

&lt;h2&gt;
  
  
  The diagnostic process had a separate failure
&lt;/h2&gt;

&lt;p&gt;During this investigation, raw configuration diffs and adjacent configuration lines exposed credential values in tool output. The remediation record explicitly leaves credential rotation as an unresolved follow-up. I am not reproducing those values or treating the descriptor fix as a security repair.&lt;/p&gt;

&lt;p&gt;The operational lesson is concrete: configuration diagnosis should print server names, enabled flags, counts, and whether credentials are present. Narrow line ranges are not a sufficient filter when secrets can occupy neighboring lines. That mistake belongs in this postmortem because it happened while collecting the evidence, not because it explains &lt;code&gt;EMFILE&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The resource repair began when I stopped taking “invalid” as the diagnosis and followed the nested read error. The acceptance check was equally specific: fresh executions completed without the original warning, and the observed pipe footprint fell. Keeping those claims narrower than “all limits fixed” leaves the next investigation somewhere honest to start.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This post was prepared with AI assistance from my incident and remediation record.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;License: CC BY 4.0.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Email list for these notes: &lt;a href="https://hexisteme.beehiiv.com/?modal=signup&amp;amp;utm_source=devto&amp;amp;utm_campaign=notes-engineering&amp;amp;utm_content=your-skill-md-was-valid-codex-ran-out-of-file-descriptors" rel="noopener noreferrer"&gt;hexisteme.beehiiv.com&lt;/a&gt; — no issue has gone out yet, so you would be on it before the first one. No welcome sequence, no course, no upsell.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;More notes at &lt;a href="https://hexisteme.github.io/notes/" rel="noopener noreferrer"&gt;hexisteme.github.io/notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>debugging</category>
      <category>devtools</category>
      <category>ai</category>
      <category>programming</category>
    </item>
    <item>
      <title>Your LLM Telemetry Table Does Not Have One Denominator</title>
      <dc:creator>John</dc:creator>
      <pubDate>Mon, 21 Sep 2026 00:00:04 +0000</pubDate>
      <link>https://dev.to/hexisteme/your-llm-telemetry-table-does-not-have-one-denominator-e89</link>
      <guid>https://dev.to/hexisteme/your-llm-telemetry-table-does-not-have-one-denominator-e89</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://hexisteme.github.io/notes/stratified-comparison.html" rel="noopener noreferrer"&gt;hexisteme notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I had an LLM telemetry report with the usual reassuring furniture: medians, interquartile ranges, sample sizes, and bootstrap intervals. The rows were grouped by model. It looked ready for comparison.&lt;/p&gt;

&lt;p&gt;It was actually several different studies sharing a table.&lt;/p&gt;

&lt;p&gt;The core process metrics were attributed to model epochs inside threads. The completion proxy existed only at thread level. Main sessions and sidechains had different structures. Mixed-model threads could contribute to one table but fail the purity rule for another. Historical routing policy was often unknown, and task family was not observed at all.&lt;/p&gt;

&lt;p&gt;The dangerous mistake was no longer simply calling an association causal. It was treating every &lt;code&gt;n&lt;/code&gt; beside a model label as if it counted the same kind of thing.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;For readers auditing an agent harness, &lt;a href="https://github.com/hexisteme/hard-gate-hooks" rel="noopener noreferrer"&gt;hexisteme/hard-gate-hooks&lt;/a&gt; contains two MIT-licensed Stop-hook examples, their tests, and a read-only scanner. They are adjacent implementation examples, not the telemetry instrument described here.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A model column is not an analysis unit
&lt;/h2&gt;

&lt;p&gt;For the core metrics, attribution happened inside a thread. A multi-model thread could produce separate epoch rows because each assistant turn was assigned to the model epoch that produced it. Tool-error rate, re-edit rate, validation runs, recovery sequences, and output tokens therefore described epoch-attributed portions of work.&lt;/p&gt;

&lt;p&gt;The completion proxy had a different contract. It was computed once per thread, only for main threads with model purity at or above &lt;code&gt;0.9&lt;/code&gt;, and censored threads were excluded. The same model label could therefore have one sample size in the core table and a smaller one in the proxy table without either count being wrong.&lt;/p&gt;

&lt;p&gt;That distinction changes what a sentence is allowed to say:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;An epoch-attributed error rate describes the observed portion of work assigned to that epoch.&lt;/li&gt;
&lt;li&gt;A thread-level completion proxy describes eligible, sufficiently pure main threads.&lt;/li&gt;
&lt;li&gt;Neither can silently stand in for “tasks completed by this model.”&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Putting the values in neighboring sections does not make their denominators interchangeable. Before reading a difference, I now ask what one row represents: a turn-attributed epoch fragment, a whole thread, a delegation edge, or something else.&lt;/p&gt;

&lt;h2&gt;
  
  
  Role changes the meaning of the metric
&lt;/h2&gt;

&lt;p&gt;The report kept main sessions and sidechains separate. That was necessary because they do not end the same way and do not carry the same amount of interaction.&lt;/p&gt;

&lt;p&gt;A main thread can contain repeated reads, edits, recovery, and user turns. A sidechain is often a bounded delegated run. A re-edit rate in those two populations mixes model behavior with session structure. Calling the gap “rework” would add another unsupported interpretation: the metric sees repeated edits to a file, but it cannot tell productive iteration from repair.&lt;/p&gt;

&lt;p&gt;The completion proxy made the role mismatch even sharper. One sidechain harness commonly ended on a &lt;code&gt;tool_result&lt;/code&gt; line. In one recorded cell, that happened in &lt;code&gt;92&lt;/code&gt; of &lt;code&gt;99&lt;/code&gt; threads. A last-line heuristic could read those endings as incomplete even when the delegated work had finished. The report therefore excluded sidechains from completion comparison instead of repairing the headline with a caveat.&lt;/p&gt;

&lt;p&gt;That is the right direction for an invalid metric: narrow its valid population. A warning below a cross-role chart does not undo a denominator that never meant the same thing across roles.&lt;/p&gt;

&lt;h2&gt;
  
  
  Epoch boundaries belong in the key
&lt;/h2&gt;

&lt;p&gt;A model name can survive a relaunch, configuration change, or routing change. The treatment does not.&lt;/p&gt;

&lt;p&gt;The report split pre-relaunch and relaunch observations into separate model epochs and refused to pool them. This was more than a naming preference. Sequence metrics were calculated within each epoch fragment, so an edit performed by one model and validated after a model switch was not credited as a within-epoch validation sequence for either side.&lt;/p&gt;

&lt;p&gt;That limitation is useful because it is visible. Pooling both epochs under the same display name would create a larger sample by erasing the boundary that explains why the sample is heterogeneous.&lt;/p&gt;

&lt;p&gt;The practical key for a row is therefore at least:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;model_epoch × role × analysis_unit&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Add policy and task family only when those fields are actually observed. A friendly model name is presentation. It is not a stable experimental identity.&lt;/p&gt;

&lt;h2&gt;
  
  
  Missing strata do not become controls
&lt;/h2&gt;

&lt;p&gt;The report also grouped rows by a source-stamped dispatch-policy hash when that evidence existed. Many historical sessions had no known policy version. Those rows remained &lt;code&gt;UNKNOWN&lt;/code&gt;; they were not backfilled from the current router or guessed from which model happened to run.&lt;/p&gt;

&lt;p&gt;Task family was stricter still. Every populated row in that stratum was marked &lt;code&gt;NOT_OBSERVED&lt;/code&gt;. The correct output was an empty comparison, despite thousands of rows elsewhere in the report.&lt;/p&gt;

&lt;p&gt;This is the point where a telemetry system proves whether it is an instrument or a story generator. A missing assignment variable is not a neutral baseline. It is an unmeasured confounder. The table may still describe the sample under the routing policy that produced it, but it cannot claim to have held task family or policy constant.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bootstrap the metric you actually defined
&lt;/h2&gt;

&lt;p&gt;The interval calculation also had to follow the unit contract. Continuous process metrics used a difference in medians. Binary completion proxies used a difference in proportions. Both cells needed enough observations, and comparisons stayed within the same role.&lt;/p&gt;

&lt;p&gt;Several zero-heavy cells still produced a bootstrap interval of &lt;code&gt;[0, 0]&lt;/code&gt;. That did not mean the effect was known exactly. When most resamples contain the same tied value, the percentile bootstrap can collapse because the statistic does not move. The interval is describing the resampling behavior of a nearly constant cell, not granting the measurement infinite precision.&lt;/p&gt;

&lt;p&gt;An interval inherits every limitation of the metric beneath it. It cannot repair a proxy that changes meaning by role, a pooled epoch, or a policy field that was never observed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The comparison contract I keep beside every table
&lt;/h2&gt;

&lt;p&gt;Before I let a telemetry row influence routing, I record these fields with it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;population: which threads or fragments were eligible;&lt;/li&gt;
&lt;li&gt;unit: what one observation represents;&lt;/li&gt;
&lt;li&gt;role: main, sidechain, or another session structure;&lt;/li&gt;
&lt;li&gt;epoch: the version boundary used for attribution;&lt;/li&gt;
&lt;li&gt;assignment evidence: the source-stamped dispatch policy, or &lt;code&gt;UNKNOWN&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;metric semantics: process count, rate, additive total, or outcome proxy;&lt;/li&gt;
&lt;li&gt;exclusion rule: censored, mixed, unattributed, or below the comparison threshold;&lt;/li&gt;
&lt;li&gt;interpretation limit: the condition that would make the apparent difference disappear.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last field matters most. In routed production telemetry, an observed gap can reverse after task family, project, time, or policy is held fixed. The table is useful for monitoring drift and choosing where a controlled experiment would pay. It is not a model leaderboard.&lt;/p&gt;

&lt;p&gt;As of the &lt;code&gt;2026-09-14&lt;/code&gt; source snapshot, the table remains association-only. The testable prediction is that at least some apparent model gaps will shrink, disappear, or reverse after recorded policy, role, task family, and analysis unit are held fixed. That prediction is invalidated if adequately overlapping matched cohorts preserve the same gaps in direction and practical magnitude. The trigger is the first source-stamped task-family cohort large enough for that comparison. Persistence would justify a randomized dispatch experiment; only randomized assignment could support a causal attribution to the model.&lt;/p&gt;

&lt;p&gt;When you look at your own LLM telemetry, does every &lt;code&gt;n&lt;/code&gt; count the same kind of thing?&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Email list for these notes: &lt;a href="https://hexisteme.beehiiv.com/?modal=signup&amp;amp;utm_source=devto&amp;amp;utm_campaign=notes-engineering&amp;amp;utm_content=stratified-comparison" rel="noopener noreferrer"&gt;hexisteme.beehiiv.com&lt;/a&gt; — no issue has gone out yet, so you would be on it before the first one. No welcome sequence, no course, no upsell.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;More notes at &lt;a href="https://hexisteme.github.io/notes/" rel="noopener noreferrer"&gt;hexisteme.github.io/notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>analytics</category>
      <category>llm</category>
    </item>
    <item>
      <title>A Substring Is Not a Speech Act: My AI Agent Executed Questions and Quotes</title>
      <dc:creator>John</dc:creator>
      <pubDate>Sun, 20 Sep 2026 00:00:06 +0000</pubDate>
      <link>https://dev.to/hexisteme/a-substring-is-not-a-speech-act-my-ai-agent-executed-questions-and-quotes-27n5</link>
      <guid>https://dev.to/hexisteme/a-substring-is-not-a-speech-act-my-ai-agent-executed-questions-and-quotes-27n5</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://hexisteme.github.io/notes/a-substring-is-not-a-speech-act.html" rel="noopener noreferrer"&gt;hexisteme notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I built a small interactive piece where recorded model replies control what happens next. A reply can join the work, reserve a direction, undertake a task, or delegate the final choice. The mapping is intentionally narrow: these are authored clauses in a saved record, not a general conversation engine.&lt;/p&gt;

&lt;p&gt;Then a question joined the work.&lt;/p&gt;

&lt;p&gt;The first language binding looked for a few useful fragments. If a response contained &lt;code&gt;함께 결정&lt;/code&gt; or &lt;code&gt;같이 결정&lt;/code&gt;, the parser treated it as an offer to decide together. That seemed convenient because the recorded replies used those words. It also meant that these unrelated sentences became actions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;같이 결정해볼까?&lt;/code&gt; is a question, but it executed &lt;code&gt;join&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;같이 결정하지 않을래.&lt;/code&gt; is a negated proposal, but it executed &lt;code&gt;join&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;A quoted delegation followed by a negation still executed &lt;code&gt;delegate&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The bug was not that the model used an unusual phrase. The bug was that the executor confused a substring with a speech act. Seeing characters is not proof that a speaker made an affirmative offer, and an offer is not permission to mutate state.&lt;/p&gt;

&lt;h2&gt;
  
  
  The boundary I had failed to model
&lt;/h2&gt;

&lt;p&gt;There were two separate facts in every record entry:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;where the text came from, such as an authored line or a recorded model reply;&lt;/li&gt;
&lt;li&gt;what operation that complete, authored clause is allowed to perform in this work.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The old matcher used the first fact as if it had established the second. It also inspected fragments before it had decided whether the surrounding sentence was a question, a negation, or a quotation. Once a fragment had fired, later context could not take the action back.&lt;/p&gt;

&lt;p&gt;This is a familiar shape in application code. A feature flag checks whether a comment contains a word. A webhook accepts a payload because a nested string resembles a command. A moderation rule looks for a token and silently treats a quotation as the speaker's own statement. Each one has the same type error: character presence is being used as authorization.&lt;/p&gt;

&lt;h2&gt;
  
  
  Replace inference with an authored score
&lt;/h2&gt;

&lt;p&gt;The repair is deliberately finite. The saved record contains the exact clauses that the maker has approved, and a closed map gives each clause one operation. Text that is allowed to appear without an operation lives in a separate passive set.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;scored&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;
  &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;같이 결정해보자.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;join&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;내가 쓴 부분을 토대로 이야기의 방향을 정해볼 수 있어.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;reserve&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;좋아, 내가 초안을 해줄게.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;undertake&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;마지막 방향은 네가 정해도 돼.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;delegate&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;]);&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;sentences&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;some&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;sentence&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;scored&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;has&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;sentence&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;passive&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;has&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;sentence&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;operations&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[],&lt;/span&gt; &lt;span class="na"&gt;evidence&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important line is the whole-response check. The parser first proves that every sentence belongs to the closed score or the passive set. Only then does it collect operations. An unregistered clause anywhere in the response disables the response's operations, so a recognized fragment cannot launder a quotation, explanation, question, or negation into an action.&lt;/p&gt;

&lt;p&gt;The source kind is checked too. An unknown kind is an input error, not an invitation to guess. The implementation does not claim to understand Korean, infer a model's hidden intention, or classify arbitrary conversation. A new phrase becomes executable only after it is entered into the record and the authored score.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test the near misses, not just the happy path
&lt;/h2&gt;

&lt;p&gt;The saved browser check exercises the same function that drives the visible piece. At the first checkpoint, the recorded response has produced joint and reserve while labor and delegation remain zero. At the later checkpoint, the registered responses produce all four intended operations, and the browser reports no WebGL error. The check also keeps keyboard behavior outside the buttons and confirms that finished playback does not silently start a new game.&lt;/p&gt;

&lt;p&gt;The negative controls are the useful part. A question containing the right words must remain inert. A negated proposal must remain inert. A quotation with an affirmative sentence inside it must remain inert unless the complete quoted form is itself an authored clause. An unregistered paraphrase must remain inert even when a human reader thinks it means the same thing.&lt;/p&gt;

&lt;p&gt;This also gives the interface a clearer failure mode. The record can show that a response was received while the operation list stays empty, so an operator can distinguish “text arrived” from “the text had permission.” That distinction is useful in logs and review tools: retain the original clause, its source kind, and the rejection reason instead of replacing the text with a guessed intent. A future author can then extend the score deliberately, and a reviewer can see which negative control would have changed if the extension were unsafe.&lt;/p&gt;

&lt;p&gt;The same discipline closed a few neighboring holes. Marks are validated before time filtering so a non-finite value cannot disappear as if it were outside the sample. Coordinates and ranges are rejected at creation. Preview and result share one hinge function. An explicit new-game action is the only operation that clears a finished playback. None of these checks attempts to make the parser clever; they make its allowed surface smaller and observable.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means for text-to-action systems
&lt;/h2&gt;

&lt;p&gt;When text controls a side effect, a lexical hit is evidence about characters. It is not permission to act. Keep the source kind, preserve the complete authored clause, validate the whole response before executing anything, and make unknown text fail closed.&lt;/p&gt;

&lt;p&gt;This approach trades coverage for an honest contract. A finite grammar can tell you exactly which phrases are executable and exactly which near misses are rejected. A broad language classifier may accept more natural wording, but it also moves the decision into a probabilistic layer that is harder to audit and easier to confuse with intent.&lt;/p&gt;

&lt;p&gt;The falsifier is simple: if a newly recorded question, negation, quotation, or unregistered paraphrase produces a non-empty operation, the boundary has failed. The next useful test is therefore a growing corpus of negative controls, not a larger pile of positive examples.&lt;/p&gt;

&lt;p&gt;Where does an AI agent in your system still treat a substring as permission when it needs an authored clause?&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Email list for these notes: &lt;a href="https://hexisteme.beehiiv.com/?modal=signup&amp;amp;utm_source=devto&amp;amp;utm_campaign=notes-engineering&amp;amp;utm_content=a-substring-is-not-a-speech-act" rel="noopener noreferrer"&gt;hexisteme.beehiiv.com&lt;/a&gt; — no issue has gone out yet, so you would be on it before the first one. No welcome sequence, no course, no upsell.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;More notes at &lt;a href="https://hexisteme.github.io/notes/" rel="noopener noreferrer"&gt;hexisteme.github.io/notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>javascript</category>
      <category>testing</category>
      <category>debugging</category>
    </item>
    <item>
      <title>Entitlement Is Not Balance: Don't Authorize Google Flow Spend with Arithmetic</title>
      <dc:creator>John</dc:creator>
      <pubDate>Sat, 19 Sep 2026 10:13:26 +0000</pubDate>
      <link>https://dev.to/hexisteme/entitlement-is-not-balance-dont-authorize-google-flow-spend-with-arithmetic-3022</link>
      <guid>https://dev.to/hexisteme/entitlement-is-not-balance-dont-authorize-google-flow-spend-with-arithmetic-3022</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://hexisteme.github.io/notes/entitlement-is-not-balance.html" rel="noopener noreferrer"&gt;hexisteme notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A Google Flow account page said a plan included monthly credits and added daily credits. It was tempting to infer the remaining monthly balance by subtracting the daily amount from a displayed total. That arithmetic looked tidy, but it did not establish a balance, a billing cycle, or which credit pool would be spent next.&lt;/p&gt;

&lt;p&gt;The authorization code now treats those as separate facts. An entitlement says what a plan may include. A balance says what remains in a particular pool at a particular time. A priority rule says which pool a provider will debit. None follows from the others. Spend authorization requires evidence for the one it intends to consume.&lt;/p&gt;

&lt;h2&gt;
  
  
  The observation had a narrow scope
&lt;/h2&gt;

&lt;p&gt;On September 13, a browser UI capture for one account showed text describing 1,000 monthly included credits and 50 daily credits, with daily credits used first. It also showed 50 daily credits remaining. It did not show a current monthly balance or a billing-cycle boundary. The historical observation is useful as an input format and a test fixture; it is not a statement about today's provider policy, this account's current balance, or any other account.&lt;/p&gt;

&lt;p&gt;A content hash bound the local receipt to the captured bytes. That confirms the file later used by the code did not change. It does not authenticate the capture's origin or prove a forged capture could not have been created with a matching hash. Provenance and byte integrity are separate properties.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why subtraction is not authorization
&lt;/h2&gt;

&lt;p&gt;Suppose a UI names a monthly entitlement, shows a daily remainder, and displays an unrelated total. Subtracting one number from another silently assumes the total is a monthly balance, that no past debits were assigned differently, that the cycle is current, and that the provider will consume pools in the assumed order. Each assumption can be false while the arithmetic remains correct.&lt;/p&gt;

&lt;p&gt;The code consequently permits daily-only use from a fresh daily balance without inventing a reset instant or a monthly remainder. Monthly use requires a source-specific current remaining balance, a current cycle, and an explicit priority statement from the same capture. Unknown historical debits are reported as unclassified rather than reassigned to make the numbers add up.&lt;/p&gt;

&lt;p&gt;Freshness also belongs in the authorization contract. The captured balance was accepted for 15 minutes, then had to be rechecked. That rule does not predict a provider reset schedule. It limits how long the system is willing to act on an observation while external spending can change the account.&lt;/p&gt;

&lt;h2&gt;
  
  
  Count a reservation exactly once
&lt;/h2&gt;

&lt;p&gt;The other subtle failure is double subtraction. A pending request reserves credits before a provider response arrives. If the same request is then recognized as completed, its reservation must be removed exactly once. If it is uncertain or has a remote job identifier, ordinary retry is blocked and recovery is required.&lt;/p&gt;

&lt;p&gt;This is why authorization is more than &lt;code&gt;remaining &amp;gt;= planned&lt;/code&gt;. It is a state transition with evidence: fresh observation, durable intent, lock held across the pre-submit recheck, and a unique binding from reservation to spend. Duplicate identifiers, conflicting balances, malformed evidence, stale observations, and insufficient verified funds all stop rather than falling back to a guessed pool.&lt;/p&gt;

&lt;p&gt;The scoped test suite recorded 268 passing cases for parsing, freshness, malformed captures, reservation accounting, and legacy behavior. Tests establish that the implementation enforces these rules for the fixtures. They do not refresh the account, create a provider guarantee, or authorize real spending.&lt;/p&gt;

&lt;h2&gt;
  
  
  Design for a missing balance
&lt;/h2&gt;

&lt;p&gt;The safest useful outcome is often partial knowledge. A verified daily pool can authorize a daily-sized request even if monthly data is absent. A monthly entitlement alone cannot authorize a larger request. A UI total whose allocation is unknown should remain unknown, not be converted into a plausible answer for convenience.&lt;/p&gt;

&lt;p&gt;This pattern applies to prepaid APIs, cloud promotional credits, per-day quotas, and corporate allowances. Preserve the evidence source and capture time with every decision. Keep plan terms, available balance, pool order, and local reservations in different fields. When one of them is missing, reduce the authorization scope or stop.&lt;/p&gt;

&lt;p&gt;The falsifier is direct: reopen this design if a fresh, source-bound observation can pass while its declared pool balance or priority is wrong; if normal intent reobservation double-counts or drops a reservation; or if unknown historical rows are asserted as actual pool usage. The success condition is not a familiar-looking number. It is a traceable right to spend from a named balance.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Email list for these notes: &lt;a href="https://hexisteme.beehiiv.com/?modal=signup&amp;amp;utm_source=devto&amp;amp;utm_campaign=notes-engineering&amp;amp;utm_content=entitlement-is-not-balance" rel="noopener noreferrer"&gt;hexisteme.beehiiv.com&lt;/a&gt; — no issue has gone out yet, so you would be on it before the first one. No welcome sequence, no course, no upsell.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;More notes at &lt;a href="https://hexisteme.github.io/notes/" rel="noopener noreferrer"&gt;hexisteme.github.io/notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>billing</category>
      <category>security</category>
      <category>testing</category>
    </item>
    <item>
      <title>First Payoff Is Not the Final Answer: A Python Timing Gate</title>
      <dc:creator>John</dc:creator>
      <pubDate>Sat, 19 Sep 2026 10:11:57 +0000</pubDate>
      <link>https://dev.to/hexisteme/first-payoff-is-not-the-final-answer-a-python-timing-gate-ao2</link>
      <guid>https://dev.to/hexisteme/first-payoff-is-not-the-final-answer-a-python-timing-gate-ao2</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://hexisteme.github.io/notes/first-payoff-is-not-final-answer.html" rel="noopener noreferrer"&gt;hexisteme notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A timing check can be perfectly deterministic and still answer the wrong question. In a short-video pipeline, an existing field served two jobs at once: it identified the first visual payoff and was also treated as the timestamp for the final answer reveal. Those are different narrative events. The result was a semantic bug disguised as a duration rule.&lt;/p&gt;

&lt;p&gt;The repair was not to choose a more fashionable retention threshold. It was to give the final answer its own declared event and align that event to the narration. The early payoff remains useful for judging whether a video starts moving; it just cannot certify that the conclusion was withheld.&lt;/p&gt;

&lt;h2&gt;
  
  
  A timestamp has to name its event
&lt;/h2&gt;

&lt;p&gt;The original estimator could locate an early visual payoff. In a puzzle-style video, that may be the first diagram, contradiction, or explanatory movement. The final answer is the point where the viewer receives the solution. A single timestamp may coincide with both, but the contract must not assume that it does.&lt;/p&gt;

&lt;p&gt;The revised Python check reads the final answer from declared &lt;code&gt;reveal_beats&lt;/code&gt; and word alignment. If several beats exist, it takes the earliest answer-bearing word so a later declaration cannot hide an earlier leak. If the declaration or alignment cannot be resolved, the result is not a silent pass. It remains unmeasured and stops at the review boundary.&lt;/p&gt;

&lt;p&gt;This is a general testing pattern. “The first user-visible response” is not “the final state committed.” “The first partial result” is not “the answer returned.” Reusing the same field is attractive because it creates a neat test, but the neatness comes from collapsing two meanings that can diverge in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the new check actually showed
&lt;/h2&gt;

&lt;p&gt;The new local policy required the final reveal to be at least 40% of the runtime and 10 seconds in. Three earlier episodes were inspected without changing them. Their answer reveals were at 10.321 seconds of 40.623 seconds (25.4%), 13.156 of 43.743 (30.1%), and 9.441 of 32.284 (29.2%). Each would be classified as an early reveal by this new policy.&lt;/p&gt;

&lt;p&gt;Those ratios validate that the new check can distinguish the event it was built to inspect. They do not prove that early reveals caused weak retention, that viewers leave at a particular moment, or that 40% and 10 seconds are platform-wide optimums. The thresholds are a local policy for this format. The videos also predated the policy and were left grandfathered; a retrospective check should not rewrite historical evidence to make a new rule look successful.&lt;/p&gt;

&lt;h2&gt;
  
  
  Match preflight to the expensive check
&lt;/h2&gt;

&lt;p&gt;The implementation runs the same timing function before rendering and after rendering, using the same script and alignment inputs. That parity matters. If a check is only available after an expensive render, authors can spend the cost of an asset that was already disallowed by its declared structure. If preflight uses a weaker approximation, it merely moves the semantic bug earlier.&lt;/p&gt;

&lt;p&gt;Configuration also fails closed: the minimum ratio, delay, and enforcement fields must be finite and valid. A short runtime does not deactivate the absolute delay. The goal is not to pad videos with silence; it is to prevent a missing measurement from masquerading as approval.&lt;/p&gt;

&lt;p&gt;The focused implementation tests passed 30 cases, and the broader regression run passed 360 cases. That supports the code contract, not a claim about audience behavior. Engineering verification can show that the gate applies its declared rule consistently. It cannot convert an unmeasured causal story into an audience finding.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat the next episodes as the experiment
&lt;/h2&gt;

&lt;p&gt;The next three episodes must record first reveal, final answer reveal, 25/50/75% retention, average percentage viewed, narration pauses, audio RMS, screen changes, and upload interval. That list deliberately contains plausible alternatives: hooks, visual change, music, and topic repetition can all move a retention curve.&lt;/p&gt;

&lt;p&gt;The current intervention should be reconsidered if three episodes that place the final answer after the local 40%/10-second boundary do not improve median 50% retention. It should also be reconsidered if a normal script delivers an early visual payoff while withholding its final answer and the two gates interfere with each other. These falsifiers test the hypothesis, not whether the telemetry exists.&lt;/p&gt;

&lt;p&gt;The practical rule is simple: before sharing a timestamp, write the event name next to it. If the system cannot state which event it observed, it has not measured a substitute. It has changed the question.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Email list for these notes: &lt;a href="https://hexisteme.beehiiv.com/?modal=signup&amp;amp;utm_source=devto&amp;amp;utm_campaign=notes-engineering&amp;amp;utm_content=first-payoff-is-not-final-answer" rel="noopener noreferrer"&gt;hexisteme.beehiiv.com&lt;/a&gt; — no issue has gone out yet, so you would be on it before the first one. No welcome sequence, no course, no upsell.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;More notes at &lt;a href="https://hexisteme.github.io/notes/" rel="noopener noreferrer"&gt;hexisteme.github.io/notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>testing</category>
      <category>video</category>
      <category>ai</category>
    </item>
    <item>
      <title>My AI Agent Pipeline Cut Wall Time and Used More Tokens</title>
      <dc:creator>John</dc:creator>
      <pubDate>Sat, 19 Sep 2026 10:09:54 +0000</pubDate>
      <link>https://dev.to/hexisteme/my-ai-agent-pipeline-cut-wall-time-and-used-more-tokens-9</link>
      <guid>https://dev.to/hexisteme/my-ai-agent-pipeline-cut-wall-time-and-used-more-tokens-9</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://hexisteme.github.io/notes/multi-agent-wall-time-token-tradeoff.html" rel="noopener noreferrer"&gt;hexisteme notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I expected a multi-agent production pipeline to have one obvious efficiency story: divide work, finish sooner, spend less. One observed episode broke that story. It reached its first complete review candidate in 6:09:31, compared with 13:56 for an earlier preserved case. At the same time, total input including cached input rose from 197,682,422 to 246,251,458, and uncached input rose from 5,406,198 to 5,674,946.&lt;/p&gt;

&lt;p&gt;That is a 55.8% shorter elapsed path to the first candidate, alongside 24.6% more total input and 5.0% more uncached input. These are token counts, not currency or a bill estimate. The episodes also had different subject matter and development work, so this is an operational comparison, not a controlled causal experiment. But it was enough to retire the convenient claim that parallelism had made the work cheaper.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two clocks were telling different stories
&lt;/h2&gt;

&lt;p&gt;The first candidate arrived sooner. The logs do not isolate how much of that difference came from overlap among production, diagnosis, review, and implementation work. The successor candidate, after feedback, still arrived in 9:53:46, 29.0% faster than the older case.&lt;/p&gt;

&lt;p&gt;The second clock was less flattering. By the successor endpoint, total input including cached input was 387,745,989 and uncached input 10,635,589: 96.1% and 96.7% above the comparison case. Output also rose. The system reached a review candidate sooner without reducing the work measured across all paths.&lt;/p&gt;

&lt;p&gt;That distinction matters outside AI agents. A build farm can improve time to a green artifact while spending more compute. A larger incident rotation can reduce customer wait time while creating more handoffs. Neither result is automatically bad. They answer different questions. The mistake is naming one of them “efficiency” and allowing it to stand in for both.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recovery was a candidate source of overhead
&lt;/h2&gt;

&lt;p&gt;The later endpoint recorded 37 routes: 22 production routes, 15 development routes, and 17 blocked terminal events. Seven owners avoided simultaneous writes, but that protection did not prevent a single episode from being repeatedly divided, re-briefed, and resumed. Each blocked path created a fresh diagnosis, a new brief, and context transfer to another participant.&lt;/p&gt;

&lt;p&gt;It would be too strong to say that parallelism caused the increase. The observation contains confounders: pipeline defects were fixed during production, an external consultation timed out, and the comparison episode was not the same job. What it does support is a more useful hypothesis: route fragmentation, blocked-path recovery, and long-lived coordinator context are candidate costs worth measuring separately.&lt;/p&gt;

&lt;p&gt;The stage ledger added another warning. It recorded 1,327.702 seconds of active stage time, about 3.7% of the successor wall time. That is ledger coverage, not the process's true active percentage: model inference, qualification, external consultation, coordination, and waiting were outside that ledger. A partial ledger cannot settle where the rest of the time went; it can only show that the ledger is not yet an accounting system for the whole workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Put a budget at the handoff boundary
&lt;/h2&gt;

&lt;p&gt;The remediation did not try to outlaw agents or make every task serial. It made a small set of limits explicit: one logical episode job, at most two production stages, one development batch, two blocked terminations, two full renders, and no compaction for the target path. It also records session usage deltas at attach and finish instead of assuming a route's nominal model represents its cost.&lt;/p&gt;

&lt;p&gt;Those are not universal thresholds. They are an experiment design. A route that crosses its limit should trigger a scope decision or a batched development fix, instead of quietly becoming another re-briefing cycle. More importantly, the metric must join the cost of the work with the work that made it necessary. Render-only timing, or an account-wide utilization percentage, cannot do that.&lt;/p&gt;

&lt;p&gt;The next newly produced episode is the immediate qualification. The improvement claim fails if it misses a 90-minute first-review wall time, 1,621,859 uncached-input cap, 198,609 output cap, one job, at most two production stages, one development batch, two blocked terminations, two full renders, or zero compactions. Repeating that measurement on later episodes is necessary before treating one qualification as a stable improvement. Passing supports this particular control-plane design, not a general proof that AI agent systems are cheaper.&lt;/p&gt;

&lt;p&gt;The useful dashboard therefore has at least two lines: time to a reviewable result, and total work consumed to get there. When those lines move in opposite directions, do not average them into a success story. Decide which trade-off the product can actually afford, then make the handoffs observable enough to change it.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Email list for these notes: &lt;a href="https://hexisteme.beehiiv.com/?modal=signup&amp;amp;utm_source=devto&amp;amp;utm_campaign=notes-engineering&amp;amp;utm_content=multi-agent-wall-time-token-tradeoff" rel="noopener noreferrer"&gt;hexisteme.beehiiv.com&lt;/a&gt; — no issue has gone out yet, so you would be on it before the first one. No welcome sequence, no course, no upsell.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;More notes at &lt;a href="https://hexisteme.github.io/notes/" rel="noopener noreferrer"&gt;hexisteme.github.io/notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>devops</category>
    </item>
    <item>
      <title>One API Key Turned the Gateway's Cooldown Into a 60-Second Blackout, and I Blamed the Vendor for Months</title>
      <dc:creator>John</dc:creator>
      <pubDate>Thu, 17 Sep 2026 00:00:03 +0000</pubDate>
      <link>https://dev.to/hexisteme/one-api-key-turned-the-gateways-cooldown-into-a-60-second-blackout-and-i-blamed-the-vendor-for-3a9p</link>
      <guid>https://dev.to/hexisteme/one-api-key-turned-the-gateways-cooldown-into-a-60-second-blackout-and-i-blamed-the-vendor-for-3a9p</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://hexisteme.github.io/notes/one-key-cooldown-is-a-blackout-not-a-rotation.html" rel="noopener noreferrer"&gt;hexisteme notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;For about three months, I had a working theory about one leg of my model roster: the free tier was flaky. Every so often, a request routed through an NVIDIA NIM model would come back &lt;code&gt;503&lt;/code&gt;. My client would retry on a one-second backoff, then a three-second backoff — &lt;code&gt;retry_transient: [502, 503, 504]&lt;/code&gt; — and both would also fail. Then, roughly a minute later, with no intervention from me, the same request would go through fine. I filed it under "free tier, what do you expect," because a failure that heals itself on its own timetable is what a rate-limited endpoint is supposed to look like.&lt;/p&gt;

&lt;p&gt;The theory survived because the evidence fit it every time. What broke it wasn't a new outage — it was a probe result and one column in an access log I'd been reading past for months. The vendor wasn't the unreliable part. A gateway I run myself was — and, as it turned out, just one of four separate causes behind the unstable leg I was chasing that same day, the only one that actually lived inside the gateway.&lt;/p&gt;

&lt;h2&gt;
  
  
  The column I wasn't reading
&lt;/h2&gt;

&lt;p&gt;The gateway is CLIProxyAPI 7.1.58, a multi-credential proxy — the kind of thing grouped with LiteLLM-class routers. It runs locally; everything downstream of it, including an MCP tool server's custom provider and a council persona whose &lt;code&gt;OPENAI_BASE_URL&lt;/code&gt; points at it, reaches it on a local port. Its access log writes one line per request, with a latency column I'd been skimming past:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;[gin_logger.go:97] 503 |   6ms | 127.0.0.1 | POST /v1/chat/completions
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once I actually read that number across a run of failures, the pattern was obvious. The first &lt;code&gt;503&lt;/code&gt; in an episode carried a latency of 316 ms — a real round trip, long enough to actually reach NVIDIA and come back with a failure. Every &lt;code&gt;503&lt;/code&gt; that followed in the same episode came back in 6 ms, and in older episodes as fast as 1–2 ms — the gateway answering from memory before the request ever left the building. A direct probe to the vendor, at the same moment one of these episodes was running, came back &lt;code&gt;200&lt;/code&gt;. A status code alone doesn't say which layer failed — the same discipline applies in reverse, where an error envelope inside an HTTP 200 still reads as success if the only check is &lt;code&gt;response.ok&lt;/code&gt; (&lt;a href="https://hexisteme.github.io/notes/error-inside-http-200-poisoned-the-cache.html" rel="noopener noreferrer"&gt;An Error Inside HTTP 200 Poisoned My Cache&lt;/a&gt;). Here, a &lt;code&gt;503&lt;/code&gt; looked exactly like a vendor failure and wasn't one; the only way to tell them apart was to time it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a cooldown is for, and what it becomes at one key
&lt;/h2&gt;

&lt;p&gt;A gateway holding more than one credential for a provider uses a cooldown to protect itself: if key A errors, set it aside and route the next request through key B. In CLIProxyAPI's code, that logic lives in &lt;code&gt;sdk/cliproxy/auth/conductor.go&lt;/code&gt; — &lt;code&gt;case 408, 500, 502, 503, 504:&lt;/code&gt; sets &lt;code&gt;NextRetryAfter = now.Add(1 * time.Minute)&lt;/code&gt; unless &lt;code&gt;disable-cooling&lt;/code&gt; is set on that provider block (&lt;code&gt;429&lt;/code&gt; runs through a separate quota path). The per-provider switch lives at &lt;code&gt;internal/config/config.go:658&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That logic assumes a key B. My NVIDIA NIM provider block had exactly one API key configured — one &lt;code&gt;api-key-entries&lt;/code&gt; entry. So the same cooldown that's a sensible rotation with two keys instead put the entire provider to sleep for that minute: zero available auth, for the only credential it had. For the rest of that window the gateway answered every request itself, immediately, with &lt;code&gt;503 auth_unavailable&lt;/code&gt; / "no auth available," without ever contacting NVIDIA. That's the 6 ms.&lt;/p&gt;

&lt;p&gt;The arithmetic downstream was already decided before the cooldown even started. My retry policy backs off at 1 s and then 3 s, both well inside the 60-second window the cooldown holds open — not because the vendor failed again, but because a minute-long timer doesn't care how fast a client retries. The seed for a given episode was ordinary on its own — NVIDIA NIM's free tier returning "Service temporarily overloaded," or once a 302-second hang ending in a &lt;code&gt;504&lt;/code&gt; — but the blackout that followed had nothing more to do with how bad that seed event was.&lt;/p&gt;

&lt;h2&gt;
  
  
  Every fix I almost made
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Raise the consumer's retry count or timeout — the blackout is fixed at 60 seconds regardless of cause; three retries or five, all land inside the same window.&lt;/li&gt;
&lt;li&gt;Blame the free tier again — the direct probe answered &lt;code&gt;200&lt;/code&gt; at the same moment the gateway refused the request. The outage never left my machine.&lt;/li&gt;
&lt;li&gt;Restart the gateway to clear the cooldown — it works, because the cooldown is in-memory state, which is the problem: the restart clears the symptom and leaves the cause in place, and cuts every other consumer on the gateway for a fix only one leg needed.&lt;/li&gt;
&lt;li&gt;Turn &lt;code&gt;disable-cooling&lt;/code&gt; on globally — wrong for any provider that actually holds two or more keys, where the cooldown is doing its real job. The setting belongs per provider, sized to how many keys that provider has.&lt;/li&gt;
&lt;li&gt;Treat every non-2xx the same way. A &lt;code&gt;410&lt;/code&gt; means a model has genuinely reached end of life — permanent, needs its alias redirected, not retried. A &lt;code&gt;429&lt;/code&gt; runs through the gateway's own quota path. Folding either into "the leg is flaky" erases what tells you what to do next.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The fix
&lt;/h2&gt;

&lt;p&gt;For a provider with one key — the default case — the fix is &lt;code&gt;disable-cooling: true&lt;/code&gt; on that provider's block only:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;openai-compatibility&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nvidia-nim"&lt;/span&gt;
    &lt;span class="na"&gt;base-url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://integrate.api.nvidia.com/v1"&lt;/span&gt;
    &lt;span class="na"&gt;disable-cooling&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
    &lt;span class="na"&gt;api-key-entries&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;api-key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;&amp;lt;redacted&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;—&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;injected&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;from&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;env&amp;gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;CLIProxyAPI hot-reloads its config file via a watcher — the log records "config successfully reloaded" — no restart, no gap for other consumers. I applied this at 2026-09-05 20:41:58 and reran the probe: PASS 3/3 — nano in 3972 ms, super in 558 ms, and the 49b model correctly back with &lt;code&gt;410&lt;/code&gt;, its expected end-of-life answer, not a new failure.&lt;/p&gt;

&lt;p&gt;For a provider where I can hold two or more keys, the fix runs the other way: add another API key, let the cooldown become a real rotation, and leave it on. Custody matters here for the same reason it mattered when a metered API key, exported globally in a shell profile, turned out to be inherited by unrelated tools that had nothing to do with the one originally suspected (&lt;a href="https://hexisteme.github.io/notes/the-key-owner-was-not-the-spender.html" rel="noopener noreferrer"&gt;The Key's Owner Wasn't the Spender&lt;/a&gt;) — each key gets injected per consumer from the environment, never printed in the clear.&lt;/p&gt;

&lt;p&gt;Either way, retries get exactly one owner. Mine live in the consumer's roster — &lt;code&gt;retry_transient&lt;/code&gt; plus its backoff — and nowhere else. Stack a retry loop on top of a gateway-side cooldown and you have two layers with two different budgets deciding the same failure — and here the smaller budget was always going to lose.&lt;/p&gt;

&lt;p&gt;The diagnostic itself is just the latency column, grepped and tailed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s2"&gt;"gin_logger.go:[0-9]+&lt;/span&gt;&lt;span class="se"&gt;\]&lt;/span&gt;&lt;span class="s2"&gt; 503 "&lt;/span&gt; ~/path/to/gateway/launchd.out | &lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="nt"&gt;-F&lt;/span&gt;&lt;span class="s1"&gt;'|'&lt;/span&gt; &lt;span class="s1"&gt;'{print $1 "|" $2}'&lt;/span&gt; | &lt;span class="nb"&gt;tail&lt;/span&gt; &lt;span class="nt"&gt;-20&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run against the history, the same signature — one slow &lt;code&gt;503&lt;/code&gt;, then a run of 1–2 ms &lt;code&gt;503&lt;/code&gt;s — showed up on 2026-06-25, 2026-07-01, and 2026-08-26, and again on 2026-08-27 and 2026-09-02, the two episodes I'd gone looking for. The access log held 37 lines like that; most of them carried the signature. It had been there the whole time I was blaming the free tier.&lt;/p&gt;

&lt;h2&gt;
  
  
  A different layer, same misattribution
&lt;/h2&gt;

&lt;p&gt;This sits in the same family as a &lt;code&gt;tail -20&lt;/code&gt; that couldn't emit until EOF, with the pipe's write-end held open by an inherited descriptor — the tool was fine there too, and an observation pipeline manufactured the failure (&lt;a href="https://hexisteme.github.io/notes/the-pipe-was-the-bug-not-the-vendor.html" rel="noopener noreferrer"&gt;For Weeks I Logged a CLI as Flaky&lt;/a&gt;). The roles are flipped here: the vendor was fine, and my own gateway layer manufactured the outage. Same instinct to blame the far end of the wire first; the defect was one layer closer to home both times.&lt;/p&gt;

&lt;h2&gt;
  
  
  What generalizes
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;A protective mechanism built for N ≥ 2 of something — keys, replicas, nodes — can become the failure it exists to prevent once N = 1. The safety property was never in the code path; it was an assumption about the environment around it that nothing enforces.&lt;/li&gt;
&lt;li&gt;A status code from a proxy isn't automatically the vendor's status code. Time the request before deciding where it failed: milliseconds mean the network was never touched; hundreds of milliseconds to seconds mean somebody upstream answered.&lt;/li&gt;
&lt;li&gt;Retries and cooldowns are both timers with an opinion about failure. When a client's retry budget is smaller than a server-side cooldown window, every retry is a foregone conclusion — the fix is one layer owning retries, not a bigger timer.&lt;/li&gt;
&lt;li&gt;When several things break on the same day, check whether they even share a layer before treating them as one root cause. Symptoms carrying the same "unstable" label can still have separate homes.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where this stops being true
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;With two or more keys behind a provider, the cooldown is doing its actual job. Disabling it there removes a legitimate rotation, not a false one.&lt;/li&gt;
&lt;li&gt;If the &lt;code&gt;503&lt;/code&gt;'s latency in the log reads in the hundreds of milliseconds or seconds, that request went upstream and came back with a real failure — a vendor problem again, not a local blackout.&lt;/li&gt;
&lt;li&gt;If a gateway ever starts skipping cooldown by default when a provider has exactly one credential, or shrinks the cooldown window below a consumer's retry budget — a few seconds instead of a minute — this fix stops being necessary.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of the individual facts here were hidden. The code was right there to read, the log line carried a latency column the whole time, and the vendor answered honestly when I asked it directly. The signature sat in that log for about three months before I read it as a local blackout instead of vendor instability — not for lack of evidence, but because I hadn't yet thought to distrust the layer I'd installed to make the fleet steadier in the first place.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Email list for these notes: &lt;a href="https://hexisteme.beehiiv.com/?modal=signup&amp;amp;utm_source=devto&amp;amp;utm_campaign=notes-engineering" rel="noopener noreferrer"&gt;hexisteme.beehiiv.com&lt;/a&gt; — no issue has gone out yet, so you would be on it before the first one. No welcome sequence, no course, no upsell.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;More notes at &lt;a href="https://hexisteme.github.io/notes/" rel="noopener noreferrer"&gt;hexisteme.github.io/notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>devops</category>
      <category>debugging</category>
      <category>api</category>
      <category>ai</category>
    </item>
    <item>
      <title>The Observed Maximum Is Not the Structural Maximum: Set a Compliance Gate's Budget from Code, Not from Three Samples</title>
      <dc:creator>John</dc:creator>
      <pubDate>Wed, 16 Sep 2026 00:00:06 +0000</pubDate>
      <link>https://dev.to/hexisteme/the-observed-maximum-is-not-the-structural-maximum-set-a-compliance-gates-budget-from-code-not-5dda</link>
      <guid>https://dev.to/hexisteme/the-observed-maximum-is-not-the-structural-maximum-set-a-compliance-gates-budget-from-code-not-5dda</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://hexisteme.github.io/notes/observed-max-is-not-structural-max.html" rel="noopener noreferrer"&gt;hexisteme notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I run a compliance gate in front of every episode a short-form video pipeline produces, and one of its settings is a budget for checks the gate couldn't run at all: &lt;code&gt;compliance.max_unmeasured&lt;/code&gt;. Some episodes mix in stock clips, and without a clip-layer ledger recording where each clip came from, a handful of the gate's checks are impossible to evaluate — not approximately, but in principle; there's nothing to measure them against. The gate doesn't fail those checks, it records them as &lt;code&gt;unmeasured&lt;/code&gt;, and the budget exists to stop an episode where measurement itself has collapsed from still shipping green. It's a backstop against silent measurement failure, not a fine-grained discriminator between a good episode and a bad one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three episodes, one proposed number
&lt;/h2&gt;

&lt;p&gt;The setting needed a value, and I didn't have one yet. I measured the three hybrid episodes currently active in the pipeline and counted how many checks came back &lt;code&gt;unmeasured&lt;/code&gt; on each: 0, 1, and 4. The obvious move was to propose the observed maximum, so I proposed 4.&lt;/p&gt;

&lt;p&gt;I put that number up for review before committing it, and one reviewer left a single falsifying condition instead of an objection: if another configuration produced 5, then 4 wasn't a prescription, it was a coincidence. I didn't wait around to see whether some future episode would eventually produce a 5. I checked that condition against the code.&lt;/p&gt;

&lt;p&gt;That's a different kind of failure from a zero-width confidence interval mistaken for precision, a check pointed at the wrong part of the frame, a detector that simply never fires, a gate that measured everything correctly but still rubber-stamped the wrong artifact, or a control sample that turned out to be mislabeled. Here, the three measurements were accurate, correctly attributed, and the gate's own logic about them was sound. The problem was one level up: an accurate sample had simply never visited one branch of a pipeline that had at least two.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually produces the number
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;unmeasured&lt;/code&gt; isn't a property of an episode; it's a property of what the checker can reach. So instead of asking a fourth or fifth episode for a better answer, I went to find out, structurally, how many checks are even capable of coming back &lt;code&gt;unmeasured&lt;/code&gt; — by counting in the code, not by sampling more of the pipeline's output.&lt;/p&gt;

&lt;p&gt;Every check that depends on the clip-layer ledger consumes it through one function, &lt;code&gt;clip_ledger_coverage&lt;/code&gt;, so I counted every place &lt;code&gt;compliance_gate.py&lt;/code&gt; calls it. Five checks depend on that ledger: &lt;code&gt;role_collision_absent&lt;/code&gt;, &lt;code&gt;single_idea&lt;/code&gt;, &lt;code&gt;payoff_spoiler&lt;/code&gt;, &lt;code&gt;canvas_utilization&lt;/code&gt; (which averages over a window rather than a single frame), and &lt;code&gt;narration_scene_sync&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That fifth check is the one that mattered. &lt;code&gt;narration_scene_sync&lt;/code&gt; only ever fires for episodes in the &lt;code&gt;data_reveal&lt;/code&gt; lineage — on any other lineage it has nothing to check, ledger or no ledger. All three episodes in my sample were &lt;code&gt;sim&lt;/code&gt; lineage. Which means 4 was never the maximum number of checks a hybrid episode can fail to measure. It was the maximum for a &lt;code&gt;sim&lt;/code&gt;-lineage hybrid, specifically, and the first hybrid episode out of the &lt;code&gt;data_reveal&lt;/code&gt; lineage was always going to produce 5, independent of anything about that particular episode.&lt;/p&gt;

&lt;p&gt;The population here has at least two independent dimensions — lineage, and whether an episode is hybrid at all — and my sample of three had only ever visited one cell of that grid. When a sample covers one category out of a population that has more than one, the maximum it produces isn't the structural maximum. It's a lower bound on it. Setting a gate's threshold at that lower bound doesn't make the gate strict; it makes the gate a countdown timer. It stays green until the first episode from the category the sample never saw arrives, and then it turns red on what could be a perfectly ordinary day — the same shape as a test that fails specifically when something succeeds for the first time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the backstop didn't catch this
&lt;/h2&gt;

&lt;p&gt;This gate already had a positive control, and it still passed one. Set the ceiling at 5 against a real episode and the gate passes; set the same ceiling at 3 and it fails — proof the comparison is wired up and alive, not inert.&lt;/p&gt;

&lt;p&gt;But a positive control like that only exercises the categories the sample has already stepped into. Mine ran on a real episode, and the three episodes I had sampled were all &lt;code&gt;sim&lt;/code&gt;-lineage, so whatever the control proved, it proved inside that one category. It had nothing to say about &lt;code&gt;data_reveal&lt;/code&gt; hybrids, because none had passed through it. A control built from a given sample can't falsify a claim about a category that sample never visited — that takes something from outside the sample entirely. Here, the only thing that pointed outside the sample was the reviewer's falsifying condition, and closing it meant reading &lt;code&gt;compliance_gate.py&lt;/code&gt; once, not collecting a fourth or fifth episode and hoping it happened to land in the right lineage.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: write both numbers down
&lt;/h2&gt;

&lt;p&gt;I corrected the config value from 4 to 5. But the more important part of the fix wasn't the new number, it was what I wrote next to it. Two numbers now sit beside &lt;code&gt;compliance.max_unmeasured&lt;/code&gt;: the structural set counted from the code — 5 — and the observed maximum from the sample that produced the original proposal — 4, specifically for the &lt;code&gt;sim&lt;/code&gt; lineage. Neither number replaces the other. The structural count says what the gate can currently produce; the observed count says what's actually been seen, and the gap between the two is exactly what made the original proposal wrong.&lt;/p&gt;

&lt;p&gt;I also wrote the condition that would lower the value again as something a machine could evaluate, rather than a sentence for a future me to remember: compute the real maximum of unmeasured checks across active episodes, and if an external producer of clip geometry ever appears and that real measurement drops below 5, use that value instead. A threshold that can only be lowered by someone remembering a paragraph is a threshold that will drift.&lt;/p&gt;

&lt;p&gt;Then I checked that none of this had quietly broken anything. The gate's test suite — 315 tests — passed. I reconfirmed the positive control separately and directly, over the API, without touching &lt;code&gt;gate.yaml&lt;/code&gt;: the ceiling at 5 still passed, the ceiling at 3 still failed.&lt;/p&gt;

&lt;h2&gt;
  
  
  What generalizes
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Before a threshold candidate taken from observation becomes a config value, count the mechanism that produces it in the code, separately from the sample. Write down the size of that set, and which part of it your observation actually covered.&lt;/li&gt;
&lt;li&gt;Ask whether the sample walked every category dimension the population has — lineage, mode, path, whatever applies — before trusting its maximum. A handful of measurements from one category says nothing about a category they never touched.&lt;/li&gt;
&lt;li&gt;Write the structural count and the observed maximum down side by side, next to the value, and write the condition that would move the value as something computable, not something memorized.&lt;/li&gt;
&lt;li&gt;A positive control assembled from the same sample that produced a threshold cannot falsify that threshold for a category outside the sample. When someone hands you a falsifying condition shaped like "if another configuration produces one more than this," close it by reading the code, not by waiting for more data to show up — it's cheaper and it's certain.&lt;/li&gt;
&lt;li&gt;Keep backstops and discriminators separate in your head. A backstop that exists to catch measurement collapsing entirely belongs at the structural maximum, not fitted to whatever sample you happened to have. A discriminator needs an actual distribution behind it, and a few data points from one category isn't one.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where this stops being true
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;If another check that depends on the ledger gets added later, 5 goes stale, and the discipline of counting it in the code rots the moment whoever adds that check forgets to also update the number next to it. Without a test that cross-checks the counted set against the config value directly, this exact defect comes back.&lt;/li&gt;
&lt;li&gt;If the structural set turns out to be far larger than anything the gate could plausibly observe in practice, a backstop pinned to the structural maximum stops meaning much — at that point the prescription in this piece is the wrong one, and what's actually needed is a discriminator built from a real distribution, not a structural ceiling.&lt;/li&gt;
&lt;li&gt;If any of the five checks changes so that it no longer needs the ledger to run, the structural maximum drops below 5, and the value has to follow it down — the count is not a fact about the checks in general, it's a fact about what the code currently requires.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The number that was wrong here wasn't miscalculated. 4 really was the highest unmeasured count across every episode I had. The mistake was treating an accurate measurement of an incomplete sample as if it described the whole population the code could produce. The fix wasn't a bigger sample. It was reading the one place in the code that actually decides how big this number can get.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Email list for these notes: &lt;a href="https://hexisteme.beehiiv.com/?modal=signup&amp;amp;utm_source=devto&amp;amp;utm_campaign=notes-engineering" rel="noopener noreferrer"&gt;hexisteme.beehiiv.com&lt;/a&gt; — no issue has gone out yet, so you would be on it before the first one. No welcome sequence, no course, no upsell.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;More notes at &lt;a href="https://hexisteme.github.io/notes/" rel="noopener noreferrer"&gt;hexisteme.github.io/notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>softwareengineering</category>
      <category>statistics</category>
    </item>
    <item>
      <title>Exclusion Lists Managed by Name Always Rot. Mark the Population Instead.</title>
      <dc:creator>John</dc:creator>
      <pubDate>Tue, 15 Sep 2026 00:00:03 +0000</pubDate>
      <link>https://dev.to/hexisteme/exclusion-lists-managed-by-name-always-rot-mark-the-population-instead-1neg</link>
      <guid>https://dev.to/hexisteme/exclusion-lists-managed-by-name-always-rot-mark-the-population-instead-1neg</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://hexisteme.github.io/notes/exclusion-lists-rot-define-populations-structurally.html" rel="noopener noreferrer"&gt;hexisteme notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I've written before about checks that pass for reasons that have nothing to do with what they claim to verify: &lt;a href="https://hexisteme.github.io/notes/your-dead-code-detector-found-a-missing-rule.html" rel="noopener noreferrer"&gt;a dead-code detector that flagged six constants that turned out to be unenforced specifications, not dead code&lt;/a&gt;, &lt;a href="https://hexisteme.github.io/notes/detector-that-never-fires.html" rel="noopener noreferrer"&gt;a detector whose zero false positives were indistinguishable from a stub's until a positive control showed it could actually fire, which it did eight times&lt;/a&gt;, &lt;a href="https://hexisteme.github.io/notes/a-check-that-never-ran-is-not-passing.html" rel="noopener noreferrer"&gt;a rule that evaluated cleanly because the field it read had no producer anywhere in the codebase and was always null&lt;/a&gt;. This one is a different shape from all three. The check itself was correct on the day I wrote it — it told articles apart from non-articles without any ambiguity. What broke wasn't the logic. It was the definition of the population the logic ran over, and it broke the same day.&lt;/p&gt;

&lt;h2&gt;
  
  
  The list that broke the same day it was written
&lt;/h2&gt;

&lt;p&gt;I was building a check for a publishing pipeline: does every article on a hub site actually carry a conversion surface, or did one slip through empty? The check needed to walk every page and decide which ones counted as "articles" versus which ones were something else — an index page, a lead magnet, a policy page — that shouldn't be held to the same rule. My first pass did what most first passes do: a denylist of the slugs I already knew weren't articles.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;_EXCLUDED_PAGE_SLUGS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;frozenset&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;privacy&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;terms&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;disclosure&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="bp"&gt;...&lt;/span&gt;
&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;_EXCLUDED_PAGE_SLUGS&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;relative&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;parts&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;   &lt;span class="c1"&gt;# exact match
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three reasonable-looking strings. It was also already wrong, and the reason has nothing to do with which three strings they were — it would have been just as wrong with thirty.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same day, a fourth static page
&lt;/h2&gt;

&lt;p&gt;Later that same day I published one more static page: &lt;code&gt;privacy-policy&lt;/code&gt;. The set comparison is an exact match, and &lt;code&gt;privacy-policy&lt;/code&gt; is not &lt;code&gt;privacy&lt;/code&gt;. The page sailed straight through the denylist, got counted as an article, and the check reported &lt;code&gt;[PARTIAL] 13/14&lt;/code&gt; — a false positive claiming a real article was missing a conversion surface it never needed, because it wasn't an article in the first place.&lt;/p&gt;

&lt;p&gt;That's the more forgiving of the two directions a stale exclusion list can fail in. Under-exclusion produces false positives: a legitimate page gets permanently flagged, the check turns into background noise, and the next time it fires about something real, nobody's watching anymore because it already cried wolf. The other direction is worse. Over-exclusion produces false negatives — something that should have been checked quietly stops being checked, and nothing in the output says so. A list that's too narrow announces itself with noise. A list that's too broad doesn't announce anything at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same bug, independently, in a different language
&lt;/h2&gt;

&lt;p&gt;The part that made me stop treating this as a one-off typo: the identical defect already existed, independently, in a deploy script written in a different language against a different file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$slug&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="k"&gt;in &lt;/span&gt;lead-magnets|privacy&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;continue&lt;/span&gt; &lt;span class="p"&gt;;;&lt;/span&gt; &lt;span class="k"&gt;esac&lt;/span&gt;   &lt;span class="c"&gt;# counts privacy-policy as an article, deploy halts&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nobody copied this bug from the Python check into the shell script, or the other way around. Two separate implementations of "which pages don't count as articles" reached for the same shape of answer: write down the names you currently know about. Both were correct on the day they were written. Both were wrong within that same day, for the same reason — a name list has no way to know that a name it hasn't seen yet belongs to the same category as the names it does know.&lt;/p&gt;

&lt;h2&gt;
  
  
  Asking what to include, instead of what to exclude
&lt;/h2&gt;

&lt;p&gt;The fix wasn't a longer list, or a fuzzy match that would eventually hit the same wall with a different string. It was to stop defining the population by name at all. The site's renderer already leaves a mark on every real article and on nothing else: an &lt;code&gt;application/ld+json&lt;/code&gt; structured-data block. I checked this empirically before trusting it — every one of the 13 real articles carried the block, and every one of the 9 non-articles (the index page, the lead magnets, both policy pages) didn't. Thirteen for thirteen, nine for nine. That split is the population definition. It doesn't need to know that &lt;code&gt;privacy-policy&lt;/code&gt; exists as a string; it only needs to know whether a given page rendered the block, which is a fact about the page rather than a fact about whoever last updated the list.&lt;/p&gt;

&lt;p&gt;Once the check reads the marker instead of the name, adding a tenth non-article page the next morning requires touching neither the Python file nor the shell script. The page either renders the block or it doesn't, and the check answers correctly either way — for a page that didn't exist when the check was written.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this generalizes to
&lt;/h2&gt;

&lt;p&gt;The move is to flip the question. "What should this check skip?" invites a list, and every list has a day it stops matching the world. "What does a real member of this population look like, from the object's own output?" invites a marker, and a marker doesn't go stale just because the set of instances grows.&lt;/p&gt;

&lt;p&gt;Four things make that move stick instead of just sounding nicer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Define inclusion first, not exclusion. If there's any signal the object produces on its own — a class, a meta tag, a header line, a schema block — that a real member of the population always has and a non-member never has, that signal can replace the entire list.&lt;/li&gt;
&lt;li&gt;If no such signal exists yet, add one. A one-line marker the generator emits is cheaper to maintain forever than a name list is to keep updated even once.&lt;/li&gt;
&lt;li&gt;If a list is genuinely unavoidable, print how many items it excluded in every run's output. A silently stale list is invisible; a run that reports "excluded: 3" next to a population size that changed since last time is not. This applies to the marker approach too — its own failure mode is an item whose marker didn't render, dropping silently out of the population — so the excluded count needs to stay visible even after a marker has "solved" the problem.&lt;/li&gt;
&lt;li&gt;If the same rule is implemented in more than one place — a detector and a deploy check, in this case, in two different languages — use the identical marker in both, and leave a comment in each pointing at the other. Fixing one side without the other just lets the pair drift apart again, which is exactly what had already happened once.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where this doesn't apply
&lt;/h2&gt;

&lt;p&gt;This isn't a blanket argument against ever writing a list. It fails as advice in a domain where a marker can't attach reliably to every instance — where some real members of the population simply don't carry any distinguishing signal. There, switching to inclusion-by-marker just trades a visible false-positive problem for an invisible false-negative one, which is the worse trade. In that situation the right move is to keep the list and still hold onto the third rule above: expose the excluded count, so staleness is at least visible as a number instead of a silent gap.&lt;/p&gt;

&lt;p&gt;It's also not worth the trouble when the excluded population is small and provably fixed — five or fewer items that can't grow because of something structural about the domain, not just "I don't expect more right now." A short, closed list is more readable than a marker nobody else on the team knows to look for. The argument here is about lists that grow by nature — page types, file categories, route names — not about enumerations of a genuinely finite, closed set.&lt;/p&gt;

&lt;h2&gt;
  
  
  A second shape, same day
&lt;/h2&gt;

&lt;p&gt;The same day produced a second version of this failure in a completely different outfit. A regression test pinned only the sha256 hash of a build artifact — no copy of the artifact itself, just its checksum. The artifact it was supposed to compare against had quietly stopped existing, through three unrelated paths at once: the build directory was gitignored, the deployment process used an orphan commit plus a force push, and the live copy got overwritten on every redeploy. Each of those is individually a reasonable choice. Together they meant that the one time I actually needed to ask "is this difference the one I intended, or a regression," there was nothing left to diff against — only a hash with no file behind it.&lt;/p&gt;

&lt;p&gt;It isn't the same mechanism as the exclusion list — nothing here is a name versus a marker — but it's the same family of mistake: a check that quietly stopped verifying anything, because the thing it depended on could disappear without the check noticing. The fix was to keep the actual golden file in the repository, so a failure produces a real, readable diff, and to keep the hash only as a secondary anchor for noticing if the golden file itself gets silently edited.&lt;/p&gt;

&lt;h2&gt;
  
  
  How you'd know this is wrong
&lt;/h2&gt;

&lt;p&gt;The claim is specific enough to falsify. If a codebase's candidate markers turn out not to attach stably to every instance of the intended population — if genuine members sometimes render without the marker for legitimate reasons — then converting a name list into a marker check trades a visible failure mode for an invisible one, and that's a worse trade, not a better one. And if the excluded population genuinely can't grow past a handful of fixed names, the marker is solving a problem that was never going to occur. The 13/13 and 9/9 split is what made the marker trustworthy in this case; without a split that clean, the right fix is a list with its excluded count printed on every run, not a marker adopted on faith.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Email list for these notes: &lt;a href="https://hexisteme.beehiiv.com/?modal=signup&amp;amp;utm_source=devto&amp;amp;utm_campaign=notes-engineering" rel="noopener noreferrer"&gt;hexisteme.beehiiv.com&lt;/a&gt; — no issue has gone out yet, so you would be on it before the first one. No welcome sequence, no course, no upsell.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;More notes at &lt;a href="https://hexisteme.github.io/notes/" rel="noopener noreferrer"&gt;hexisteme.github.io/notes&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>testing</category>
      <category>codequality</category>
      <category>architecture</category>
      <category>softwareengineering</category>
    </item>
  </channel>
</rss>
