<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Aashish Bhandari</title>
    <description>The latest articles on DEV Community by Aashish Bhandari (@maxstravion).</description>
    <link>https://dev.to/maxstravion</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4127698%2F3da1070e-1903-4411-81da-cecde43b61a2.jpeg</url>
      <title>DEV Community: Aashish Bhandari</title>
      <link>https://dev.to/maxstravion</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/maxstravion"/>
    <language>en</language>
    <item>
      <title>The idle-parent trap on LLM Agents</title>
      <dc:creator>Aashish Bhandari</dc:creator>
      <pubDate>Thu, 17 Sep 2026 04:51:38 +0000</pubDate>
      <link>https://dev.to/maxstravion/the-idle-parent-trap-on-llm-agents-1ma5</link>
      <guid>https://dev.to/maxstravion/the-idle-parent-trap-on-llm-agents-1ma5</guid>
      <description>&lt;p&gt;Companion to &lt;a href="//2026-09-16_article_two-agents-one-night-158-million-tokens_Claude-Naruto.md"&gt;Two agents, one night, 158 million tokens&lt;/a&gt;. About four minutes to read.&lt;/p&gt;

&lt;h2&gt;
  
  
  The picture
&lt;/h2&gt;

&lt;p&gt;A parent model delegates a job to a worker and then has nothing else to do. It wants to know when the worker is finished. It has one tool for that: &lt;em&gt;wait up to N seconds for the worker, then tell me what happened.&lt;/em&gt; The default N is sixty.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;BUSY parent (rule holds)                IDLE parent (rule fails)

dispatch worker                         dispatch worker
work on own task ─┐                     wait 60 s ─── "not yet"   115k tokens
work on own task  │ worker finishes     wait 60 s ─── "not yet"   115k tokens
work on own task  │ notification lands  wait 60 s ─── "not yet"   115k tokens
read result ◄─────┘                     wait 60 s ─── "not yet"   115k tokens
                                          ... × 66 ...
1 wake-up, paid once                    wait 60 s ─── "done"      115k tokens
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each "not yet" is a full activation. The parent re-reads its entire conversation to learn that a minute has passed. The rule that says &lt;em&gt;do not poll&lt;/em&gt; is obeyed as long as the parent is busy, because a busy parent is woken by the worker's completion in the normal course of its next step. The moment the parent is idle, the only way it knows how to wait is the sixty-second wait, and the rule has nothing to grip.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the base case measured
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Measure&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Waits issued by the parent&lt;/td&gt;
&lt;td&gt;89&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;At sixty seconds&lt;/td&gt;
&lt;td&gt;88&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Timed out without a result&lt;/td&gt;
&lt;td&gt;66&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Returned a completion&lt;/td&gt;
&lt;td&gt;23&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tokens in activations that ended in a wait&lt;/td&gt;
&lt;td&gt;10,281,999&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Of which cached re-reads&lt;/td&gt;
&lt;td&gt;10,153,856&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mean tokens per waiting activation&lt;/td&gt;
&lt;td&gt;115,528&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Share of the whole run&lt;/td&gt;
&lt;td&gt;6.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Share of the run's list-price cost&lt;/td&gt;
&lt;td&gt;12.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The largest single turn issued 62 waits, 50 of which timed out. An earlier, smaller build showed the same pattern at five waits. This is a mechanism, not a one-off.&lt;/p&gt;

&lt;p&gt;Upper bound, labelled as a counterfactual: if every timed-out wait had been replaced by a silent notification, the 66 activations at the mean size would not have happened, about 7.6 million tokens. The 23 completions would still have cost a wake-up each. That is the most an event-driven design could remove here, not what it would save.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the two agents differ
&lt;/h2&gt;

&lt;p&gt;On &lt;strong&gt;Claude Code&lt;/strong&gt;, a background worker or a background shell command notifies the parent when it finishes. The harness wakes the parent once, on the event. There is nothing to poll, and the guardrail simply says: dispatch, then continue or end the turn.&lt;/p&gt;

&lt;p&gt;On &lt;strong&gt;Codex&lt;/strong&gt;, in the runtime we measured, a worker's completion message does not trigger a parent turn. A finished worker cannot wake an idle parent. The parent's only options are a bounded wait or a scheduled check, so a declarative "do not poll" rule cannot be followed by an idle parent. The fix has to be structural.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three ways to wait
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Model polls.&lt;/strong&gt; The parent waits N seconds, wakes, re-reads everything, decides to wait again. Cost: one activation per N seconds for as long as the worker runs. This is the trap.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One long wait.&lt;/strong&gt; The parent waits once, for as long as the worker could plausibly need. Cost: one activation if the estimate is right, two if it is not. Our interim rule sets N to five minutes, which cuts the polling cost by five with no other change.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zero-token watchdog.&lt;/strong&gt; A plain script, not a model, reads the worker's transcript counters every thirty to sixty seconds and stays silent while the worker is healthy. It speaks once: on completion, on a budget breach, or on a stuck signature such as the same failing call repeated. The parent is woken exactly once. Cost while waiting: zero model tokens. The design, with a tested reference sketch, is in the &lt;a href="//../../2026-09-15_efficiency_zero-token-watchdog-explainer_Claude-Naruto.md"&gt;zero-token watchdog explainer&lt;/a&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What the criteria will look for
&lt;/h2&gt;

&lt;p&gt;Two telemetry-only tests decide whether the fix worked, over three consecutive qualifying sessions per agent:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No repeated waiting.&lt;/strong&gt; Wait, timeout, wait again with no other work in between: zero occurrences.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No short checks.&lt;/strong&gt; Smallest wait timeout at least 300 seconds on Codex; status checks while a worker runs, zero on Claude.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Neither field exists in the telemetry yet. Adding them is fix number four in the article.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>chatgpt</category>
      <category>claude</category>
    </item>
    <item>
      <title>Two agents, one night, 158 million tokens: what it cost and what we are fixing</title>
      <dc:creator>Aashish Bhandari</dc:creator>
      <pubDate>Thu, 17 Sep 2026 03:08:01 +0000</pubDate>
      <link>https://dev.to/maxstravion/two-agents-one-night-158-million-tokens-what-it-cost-and-what-we-are-fixing-b3k</link>
      <guid>https://dev.to/maxstravion/two-agents-one-night-158-million-tokens-what-it-cost-and-what-we-are-fixing-b3k</guid>
      <description>&lt;p&gt;&lt;strong&gt;Draft for publication · 16 September 2026 · Owner: Aashish Bhandari (Max) · Author: Claude ("Naruto") · Measurements: Codex ("Goku") · Status: base case, before fixes.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Executive summary
&lt;/h2&gt;

&lt;p&gt;A design review of a web service had raised twelve findings: a missing transaction guard, retry and credential-rotation gaps, a fragile handoff route, thin diagnostics, portability problems and similar. On the evening of 15 September one AI coding agent was asked to fix all twelve, test them, integrate the work and produce a release package. It did. A second agent, from a different vendor, rebuilt the result from scratch the next morning and verified it independently.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What happened&lt;/th&gt;
&lt;th&gt;Measured&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Wall clock, first prompt to last&lt;/td&gt;
&lt;td&gt;13 h 4 min&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Of which active model time, approximately&lt;/td&gt;
&lt;td&gt;5.5 h&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Files changed, lines added&lt;/td&gt;
&lt;td&gt;118, 16,369&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tests reproduced by the second agent&lt;/td&gt;
&lt;td&gt;100 of 100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Browser checks reproduced&lt;/td&gt;
&lt;td&gt;73 of 73&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Release package&lt;/td&gt;
&lt;td&gt;byte-identical&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tokens processed by all models&lt;/td&gt;
&lt;td&gt;158,137,319&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Of which re-reads of already-cached context&lt;/td&gt;
&lt;td&gt;95.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Of which new text actually written by a model&lt;/td&gt;
&lt;td&gt;0.42%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The builder ran as one coordinating "parent" model with sixteen helper "workers" on cheaper models. Nobody typed code overnight. The verifier approved closure with no serious finding. Delivery at this pace is now routine for us. The cost is not yet acceptable, and this article is about why.&lt;/p&gt;

&lt;p&gt;One number to hold on to: the models read the equivalent of about two hundred copies of &lt;em&gt;War and Peace&lt;/em&gt; in order to write less than one. Almost all of that reading was the same conversation, sent again and again. That is the lever.&lt;/p&gt;

&lt;p&gt;A caution before the details. Reading the telemetry of a run like this is itself a model task, and our first attempt at it cost 60% of the run it was measuring. This time it cost 4.7%. Our target is under 2% of the run being assessed, measured the same way, or the analysis stops and asks for a budget.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the numbers come from
&lt;/h2&gt;

&lt;p&gt;Every figure here comes from a lifecycle hook that runs when a session pauses or ends. It reads the session transcript with a plain script, writes counters to a JSON file and regenerates a CSV and a summary. No model reads a transcript to do this. It costs zero model tokens and cannot be forgotten. The primer on this, and on why the cost of an agent is roughly &lt;em&gt;activations × context size&lt;/em&gt;, is in &lt;a href="//cost-model-in-one-picture.md"&gt;How an agent's cost adds up&lt;/a&gt; and &lt;a href="//zero-token-telemetry.md"&gt;Zero-token telemetry&lt;/a&gt;. The full fact sheet for this run is in &lt;a href="//base-case-fact-sheet.md"&gt;Base case fact sheet&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is already in place, and what each mechanism did
&lt;/h2&gt;

&lt;p&gt;Since early September both agents run under the same set of rules, installed in their global instruction files (&lt;code&gt;AGENTS.md&lt;/code&gt; for Codex, &lt;code&gt;CLAUDE.md&lt;/code&gt; for Claude Code), enforced where possible by hooks, and shared through a small set of skills. Here is each mechanism, and what this run says about it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Zero-token telemetry capture.&lt;/strong&gt; Installed as &lt;code&gt;Stop&lt;/code&gt; and &lt;code&gt;SessionEnd&lt;/code&gt; hooks for both agents. Before it, a model was asked to summarise its own usage at the end of a task, which cost one full-context activation and could not see its own final answer. Measured effect: capture of 55 components, 1,505 activations and 1,299 tool calls in this run for zero tokens. It is the reason the rest of this article exists.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Routing by work shape.&lt;/strong&gt; The parent is the strongest model. Workers default to the cheapest model that fits the job: mechanical and high-volume work to the smallest, broad reading to the mid-tier, demanding reasoning only to the top tier. Measured effect at frozen list prices: this run's named-model usage prices at $99.80 against $218.81 had every token run on the parent's model, a 54% reduction. The earlier, smaller build measured 43%. This is price substitution. The workers still processed 110 million tokens.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Delegation floor and one-shot worker packets.&lt;/strong&gt; A worker is dispatched only for a bounded workstream, with the objective, file paths, permissions, output shape and budget in one short packet, and told to return once. Measured effect: 70% of the tokens moved off the expensive parent. Partial failure: the parent reused four workers through 28 follow-up assignments, and those four long-lived workers carried 49% of the entire run, because a reused worker re-sends its accumulated history on every step.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Event-driven completion, no polling.&lt;/strong&gt; The rule says: dispatch, then do independent work or wait once; never ask "are you done?" on a timer. Measured effect: it held while the parent had work of its own, and failed the moment it went idle. The parent issued 89 waits, 88 of them at sixty seconds, and 66 timed out. Each of those checks re-sent the whole conversation: 10.3 million tokens, 6.5% of the run and 12.7% of its list-price cost, for the information "not yet". The mechanism is explained in &lt;a href="//the-idle-parent-trap.md"&gt;The idle-parent trap&lt;/a&gt;. On Claude Code the harness wakes the parent when a worker finishes, so the trap does not open there. On Codex today a finished worker cannot wake an idle parent, so a declarative rule alone is not enough.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Activation budgets and batching.&lt;/strong&gt; Claude counts per human prompt (target four, ledger at eight, deliver at twelve). Codex counts per task (ledger at eight, deliver at twelve). Both must batch independent reads and run one validation gate. Measured effect this run: 342 parent activations across twelve human turns, about 28 per turn. Not yet the target, and now measurable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Compaction.&lt;/strong&gt; The parent's context was compacted five times, three of them by hand. Each cut the next step's input by 36% to 84%. It is a relief valve, not a controller: the context regrew, and the longest post-compaction turn still spent 20 million tokens and issued 62 waits.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Analysis discipline.&lt;/strong&gt; Performance work lives in a separate workspace, reads the derived telemetry first, extracts missing fields with deterministic tools, and never assigns a model to parse raw logs. Measured effect: analysis overhead fell from 60% of the assessed run to 4.7%. Still above the 2% target, and the target is now written down.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Shared decision record and winning criteria.&lt;/strong&gt; Ten decisions and ten criteria, one file each, readable by both agents. A criterion is "won" only after three consecutive qualifying sessions per agent, judged from telemetry alone. This run is the &lt;em&gt;before&lt;/em&gt; row for every one of them.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mechanism&lt;/th&gt;
&lt;th&gt;Measured effect in this run&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Telemetry hooks&lt;/td&gt;
&lt;td&gt;0 tokens to count 158M&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Routing&lt;/td&gt;
&lt;td&gt;54% below single-model list price&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Delegation&lt;/td&gt;
&lt;td&gt;70% of tokens off the parent; 49% trapped in four reused workers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No-poll rule&lt;/td&gt;
&lt;td&gt;failed when idle: 66 timeouts, 6.5% of the run&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Compaction&lt;/td&gt;
&lt;td&gt;36–84% immediate context cut, no bound on total&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Analysis discipline&lt;/td&gt;
&lt;td&gt;60% → 4.7% overhead&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Work in progress
&lt;/h2&gt;

&lt;p&gt;Delivery was excellent. Orchestration was not. The two are separable, and separating them is the whole project. Six fixes are queued, each with a criterion that decides, from telemetry alone, whether it worked:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Five-minute waits&lt;/strong&gt; instead of sixty-second ones, with one progress line before the wait.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A zero-token watchdog&lt;/strong&gt;: a plain script that reads worker counters every minute and wakes the parent only on completion, budget breach or a stuck signature.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One worker, one stage&lt;/strong&gt;: a follow-up and activation gate per worker, with a fresh minimal-context worker for the next stage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Telemetry extensions&lt;/strong&gt; so waits, follow-ups, compaction tokens and terminal errors are counted by the hook, not reconstructed by hand.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mechanical work to the cheapest model, and parser failure as a stop condition&lt;/strong&gt;, with the analysis budget checked before an audit starts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Automatic approval reviewers&lt;/strong&gt;, 7.5% of this run at an unknown price, checked to see whether they are a setting rather than a design.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The same task shape will run again once these land, and the same hook will count it. If the numbers move, the criteria will say so. If they do not, this article stays the base case and we try the next six.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Further reading in this series: the primers linked above, the &lt;a href="//../../2026-09-15_efficiency_zero-token-watchdog-explainer_Claude-Naruto.md"&gt;zero-token watchdog explainer&lt;/a&gt;, and Goku's full &lt;a href="//../efficiency-management/reviews/2026-09-16-goku-phase-b-runtime-analysis.md"&gt;runtime analysis&lt;/a&gt; from which every measurement here is taken.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>development</category>
      <category>chatgpt</category>
      <category>claude</category>
    </item>
    <item>
      <title>[Codex feature request] Let a finished worker SubAgent wake an IDLE Parent Agent</title>
      <dc:creator>Aashish Bhandari</dc:creator>
      <pubDate>Thu, 17 Sep 2026 03:05:50 +0000</pubDate>
      <link>https://dev.to/maxstravion/codex-feature-request-let-a-finished-worker-subagent-wake-an-idle-parent-agent-3lf9</link>
      <guid>https://dev.to/maxstravion/codex-feature-request-let-a-finished-worker-subagent-wake-an-idle-parent-agent-3lf9</guid>
      <description>&lt;p&gt;&lt;strong&gt;Feature request for the Codex CLI · 17 September 2026 · Owner: Aashish Bhandari (Max) · Author: Claude ("Naruto") · Measurements and review: Codex ("Goku") · Environment measured: one GPT-6 Astra parent, sixteen GPT-5.6 workers; not yet reproduced on codex-cli 0.154.0.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Executive summary
&lt;/h2&gt;

&lt;p&gt;When a Codex parent delegates work to a sub-agent and then has nothing else to do, the only way it can learn that the worker has finished is to call &lt;code&gt;wait_agent&lt;/code&gt; with a timeout and ask again when it expires. Every ask is a full model activation that re-reads the parent's entire conversation. In one overnight build we measured that loop directly:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Measure&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Waits issued by the parent&lt;/td&gt;
&lt;td&gt;89&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Requesting a sixty-second timeout&lt;/td&gt;
&lt;td&gt;88&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Timed out with no result&lt;/td&gt;
&lt;td&gt;66&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tokens in activations that ended in a wait&lt;/td&gt;
&lt;td&gt;10,281,999&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Of which cached re-reads&lt;/td&gt;
&lt;td&gt;98.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Share of the whole run (158.1M tokens)&lt;/td&gt;
&lt;td&gt;6.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Share of known-price API equivalent, approval review excluded&lt;/td&gt;
&lt;td&gt;12.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Largest single turn&lt;/td&gt;
&lt;td&gt;62 waits, 50 timeouts, 7.4M tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The worker did nothing wrong. A longer timeout would have cut the count, and our interim rule now sets five minutes, but no timeout removes the loop, because the harness gives an idle parent no event-driven way to wait. We ask for one structural change: a worker's completion should be able to start a parent turn, so an idle parent can end its turn and be woken on the event. Claude Code's harness works this way, and the same rule set that fails on Codex holds there.&lt;/p&gt;

&lt;h2&gt;
  
  
  Background: what we were doing and why it matters
&lt;/h2&gt;

&lt;p&gt;We run two coding agents from different vendors under one shared set of efficiency rules, with zero-token telemetry hooks that count every activation. The full account is in &lt;a href="//2026-09-16_article_two-agents-one-night-158-million-tokens_Claude-Naruto.md"&gt;Two agents, one night, 158 million tokens&lt;/a&gt; and the mechanism in &lt;a href="//the-idle-parent-trap.md"&gt;The idle-parent trap&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The build fixed twelve design findings in a web service overnight, and the second agent reproduced 100 of 100 tests and a byte-identical release package the next morning. The cost was 158 million processed tokens, 95.4% of them re-reads of already-cached context. An agent's cost is roughly activations × context size, so the lever is how often the parent is re-sent, not how much it writes.&lt;/p&gt;

&lt;p&gt;The polling loop is the cleanest avoidable re-send we have found. It is a harness gap, not a model error, and it grows with the long autonomous runs we want more of. A parent holding 165k tokens of context that waits on a forty-minute worker at sixty-second intervals pays about 6.6 million tokens to hear "not yet" thirty-nine times. Five-minute waits divide that by five. They cannot remove it, because the model has no verb for "sleep until the event".&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it works on Claude Code
&lt;/h2&gt;

&lt;p&gt;A background sub-agent or background shell command is a tracked task. When it finishes, the harness delivers a notification that re-invokes the parent. The guardrail is therefore short and enforceable: dispatch, then do independent work or end the turn. An idle parent is not a running process; it costs nothing between dispatch and completion. In the Claude sessions we audited under the same rules we found no status checks during worker lifetimes. That is an observation on our own workloads, not a controlled comparison of equivalent runs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it does not work on Codex
&lt;/h2&gt;

&lt;p&gt;In the runtime we measured, a worker's final message joins the parent thread but does not start a parent turn. The runtime recorded 42 &lt;code&gt;SubagentStop&lt;/code&gt; events, and completions did reach the parent while it was busy, in the course of its next step. That is why the rule held while the parent was busy. Once idle, its only primitives were &lt;code&gt;wait_agent&lt;/code&gt; with a bounded &lt;code&gt;timeout_ms&lt;/code&gt;, or a shell &lt;code&gt;sleep&lt;/code&gt;. Both return to the model on expiry whether or not anything happened. The idle parent cannot end its turn, because nothing would resume it, so "do not poll" has nothing to grip.&lt;/p&gt;

&lt;p&gt;Two pressures made it worse. The parent chose sixty seconds on 88 of 89 waits, and the environment guidance discourages blocking beyond sixty seconds while asking for regular progress messages. That is a plausible contributor to the one-minute cadence; our analysis did not isolate its causal share.&lt;/p&gt;

&lt;p&gt;The gap is corroborated independently. &lt;a href="https://github.com/openai/codex/issues/15723" rel="noopener noreferrer"&gt;Issue 15723&lt;/a&gt;, open since March 2026, reports that background subprocesses and sub-agents do not wake the calling agent on completion. &lt;a href="https://github.com/openai/codex/issues/40932" rel="noopener noreferrer"&gt;Issue 40932&lt;/a&gt;, open since August on CLI 0.149.1, reports a parent turn ending before three running sub-agents returned, with their results surfacing only after the user's next message 32 minutes later. Both are firsthand user reports, not maintainer confirmation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Proposed solution
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Completion as a turn trigger.&lt;/strong&gt; Let a parent end its turn while sub-agents are pending. When a pending worker finishes, the harness starts a new parent turn with the worker's final message as input. This closes the trap with no new tool and no change to model behaviour beyond "end your turn".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;wait_agent&lt;/code&gt; without a hard ceiling.&lt;/strong&gt; Allow &lt;code&gt;timeout_ms&lt;/code&gt; to be omitted, returning only on completion or failure. Treat the timeout as a safety bound, not a cadence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wire completion events into the existing continuation primitives.&lt;/strong&gt; &lt;code&gt;codex exec resume --last "&amp;lt;prompt&amp;gt;"&lt;/code&gt; and the App Server &lt;code&gt;turn/start&lt;/code&gt; already start a turn from outside a session. What is missing is a reliable path from worker completion to that call, and a lifecycle hook a plain script can use. A plain script such as our zero-token watchdog could then wake the parent on a budget breach or stuck signature, as could CI or a deploy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Telemetry fields.&lt;/strong&gt; Expose wait count, outcome and timeout per session, so the effect is verifiable from hooks alone.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Acceptance, from telemetry over three consecutive qualifying sessions: zero occurrences of wait, timeout, wait again with no other work between, and zero parent activations whose only output is another wait. In the base case, the 66 timed-out activations sum to about 7.6 million tokens at the mean size. That is an upper bound on what an event-driven design could remove, not a measured saving: each completion still costs a wake-up, and the counterfactual has not been run.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Further reading: the &lt;a href="//base-case-fact-sheet.md"&gt;base case fact sheet&lt;/a&gt; and &lt;a href="//zero-token-telemetry.md"&gt;Zero-token telemetry&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>chatgpt</category>
      <category>eventdriven</category>
    </item>
    <item>
      <title>What a real product refactor revealed about AI coding agents</title>
      <dc:creator>Aashish Bhandari</dc:creator>
      <pubDate>Wed, 16 Sep 2026 09:21:18 +0000</pubDate>
      <link>https://dev.to/maxstravion/what-a-real-product-refactor-revealed-about-ai-coding-agents-2a9o</link>
      <guid>https://dev.to/maxstravion/what-a-real-product-refactor-revealed-about-ai-coding-agents-2a9o</guid>
      <description>&lt;p&gt;ReviewWithAI engineering case study · 15–16 September 2026**&lt;/p&gt;

&lt;p&gt;By Aashish Bhandari, with AI-assisted analysis by Goku (Codex, Astra) and independent review by Naruto (Claude, Fable 5.1).&lt;/p&gt;

&lt;h2&gt;
  
  
  Executive summary
&lt;/h2&gt;

&lt;p&gt;I worked with two AI coding agents to refactor ReviewWithAI, an alpha application for reviewing Markdown documents and handing changes to external agents. Goku acted as principal architect and implementing developer. Naruto independently reviewed the design and code. I set priorities, resolved material decisions and authorized progress through the review checkpoints.&lt;/p&gt;

&lt;p&gt;The work implemented eleven low-level designs across three milestones. It covered server and browser structure, authorization, persistence, testing, operational diagnostics and release preparation. The candidate passed independent checks, including tests, browser workflows and reproduction of the packaged application. Those results establish a reviewed engineering outcome; they do not establish production readiness or prove that the process was efficient.&lt;/p&gt;

&lt;p&gt;After implementation, we examined the runtime evidence to understand how the work had been orchestrated. The analysis used collected telemetry first, followed by a bounded deterministic inspection of fields missing from the collector. Its scope was Goku's implementation conversation, delegated workers and automatic approval reviewers. Naruto's separate review sessions and my time were outside that measurement.&lt;/p&gt;

&lt;p&gt;The clearest finding concerned waiting. When the parent agent had no independent work, it repeatedly resumed after short waits for its workers. Each resumption could carry conversation history back into a model invocation. Activations that issued waits accounted for more than a quarter of the parent's recorded tokens. Some waits returned completed work, so that share cannot simply be called waste or promised as recoverable savings.&lt;/p&gt;

&lt;p&gt;A second finding concerned worker reuse. Several delegated conversations accumulated assignments across multiple stages. Their recorded consumption became concentrated, but the evidence does not tell us whether starting fresh workers would have produced the same quality with less effort. Continuity may help correctness while increasing context costs. That tradeoff needs a controlled comparison.&lt;/p&gt;

&lt;p&gt;Compaction reduced the context entering the next invocation, but later context growth and recurring waits continued. Choosing cheaper worker models also lowered a calculation based on published rates, without demonstrating lower total work. Most recorded input was cached. Processed tokens therefore need to be distinguished from unique content, uncached computation and actual charges.&lt;/p&gt;

&lt;p&gt;The investigation exposed limitations in our own measurement process. The collector omitted compaction activity, and its interruption counter did not describe every terminal failure. Automatic approval reviewers also consumed resources separately from implementation workers. Even evaluating the session had a material cost: the analysis conversation included other work, preventing a clean estimate of analysis overhead. We report that uncertainty rather than presenting precision.&lt;/p&gt;

&lt;p&gt;The practical lesson is to evaluate the mechanisms around coding agents alongside their output. Instructions to avoid polling did not reliably prevent the observed behavior. Before testing optimizations, we need better deterministic counters, evaluation budgets and comparisons that include accepted quality, recovery and human effort. This report is a baseline from one project, with delivery and unresolved efficiency questions. The next step is to test changes against that baseline and publish the results, including any changes that fail to help. The detailed evidence below lets other developers inspect our reasoning and judge where these observations might apply to their workflows.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Boring stuff ahead. Proceed with caution.&lt;/strong&gt; The rest contains the methods, tables and caveats for readers who want to check the evidence.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The product and the development arrangement
&lt;/h2&gt;

&lt;p&gt;ReviewWithAI is a Markdown review application: a user selects text, attaches comments, hands work to an external coding agent, checks changed anchors, records repairs and accepts a specific source revision. The starting point was an alpha-grade, mid-sized product, as described by its developer. The work included server behavior, browser interactions, persistence, authorization, testing, documentation and release tooling. It was an existing-product refactor and remediation effort, rather than initial application generation.&lt;/p&gt;

&lt;p&gt;The eleven low-level designs (LLDs) addressed twelve review findings labelled Q1–Q12; Q3 and Q4 shared one design. Design preparation and agreement preceded the measured Phase B implementation task. The roles were:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Participant&lt;/th&gt;
&lt;th&gt;Responsibility&lt;/th&gt;
&lt;th&gt;Included in the measured task?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;One human developer, Aashish (Max)&lt;/td&gt;
&lt;td&gt;Product direction, scope, material decisions, milestone authorization and final candidate acceptance&lt;/td&gt;
&lt;td&gt;Human time is not measured&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Goku, Codex parent on GPT-6 Astra&lt;/td&gt;
&lt;td&gt;Principal architect, implementing developer, delegation, integration and verification&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sixteen Goku worker threads on Sol, Luna and Terra&lt;/td&gt;
&lt;td&gt;Bounded research, implementation, integration, tests and release work; some threads were later reused&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Naruto, Claude coding agent; final review identifies Fable 5.1&lt;/td&gt;
&lt;td&gt;Independent principal architect for design and code review, including reproduction of delivery evidence&lt;/td&gt;
&lt;td&gt;No; separate review sessions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Automatic Codex approval reviewers&lt;/td&gt;
&lt;td&gt;Review of eligible approval requests&lt;/td&gt;
&lt;td&gt;Yes; separate from implementation workers&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;“Two agents” describes the two principal collaborators. It does not mean two model processes: Goku's implementation task alone produced sixteen worker threads and 38 approval-review components.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    H["Human developer: scope, decisions and acceptance"] --&amp;gt; G["Goku: architecture and implementation"]
    H --&amp;gt; N["Naruto: independent design and code review"]
    subgraph measured["Measured Codex task"]
        G --&amp;gt; W["16 delegated worker threads"]
        G -. "approval requests across the task" .-&amp;gt; A["38 approval-review components"]
    end
    G --&amp;gt; E["Candidate code, package and verification evidence"]
    W --&amp;gt; E
    E --&amp;gt; N
    N --&amp;gt; R["Review findings and milestone verdicts"]
    R --&amp;gt; H
    R --&amp;gt; G&lt;/code&gt;&lt;/pre&gt;



&lt;h3&gt;
  
  
  Eleven designs, three review milestones
&lt;/h3&gt;

&lt;p&gt;The milestones were review checkpoints rather than individual commits. The &lt;a href="//../../SOURCE_REFERENCES.md#ref-10"&gt;LLD index and checkpoint records&lt;/a&gt; establish the sequence:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Milestone&lt;/th&gt;
&lt;th&gt;Implemented scope&lt;/th&gt;
&lt;th&gt;Review boundary&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1. Housekeeping and engineering controls&lt;/td&gt;
&lt;td&gt;Engineering-standard amendments, LLD navigation and citations, documentation checker and checker tests&lt;/td&gt;
&lt;td&gt;Human and Naruto approval before merge and runtime implementation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2. First implementation group&lt;/td&gt;
&lt;td&gt;Q5 handoff-route tests; Q10 transaction guard; combined Q3/Q4 typed operations, browser structure, retries and credential rotation&lt;/td&gt;
&lt;td&gt;Independent midpoint review; approved work integrated at &lt;code&gt;c9215c9&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3. Remaining implementation and release candidate&lt;/td&gt;
&lt;td&gt;Q2 diagnostics; Q6 portable tooling; Q8 strict agent inputs; Q9 health version; Q11 bounded discovery; Q12 handoff provenance; Q7 contributor docs; Q1 release curation&lt;/td&gt;
&lt;td&gt;Final independent closure review after integration, packaging and artifact checks&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Formatting and documentation-baseline amendments arrived after the initial Q1 freeze. They were applied and release curation was repeated at &lt;code&gt;64aa194&lt;/code&gt;; the final review explicitly accepted that sequence deviation. This additional work is included in the task's recorded consumption.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Design&lt;/th&gt;
&lt;th&gt;Finding(s)&lt;/th&gt;
&lt;th&gt;Engineering change&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Handoff endpoint verification&lt;/td&gt;
&lt;td&gt;Q5&lt;/td&gt;
&lt;td&gt;Seven route tests and expiry sensitivity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Non-nested transaction contract&lt;/td&gt;
&lt;td&gt;Q10&lt;/td&gt;
&lt;td&gt;Explicit transaction ownership, rollback and cleanup behavior&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Structural refactor&lt;/td&gt;
&lt;td&gt;Q3, Q4&lt;/td&gt;
&lt;td&gt;Typed handlers, browser decomposition, shared retry and credential-rotation ownership&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Operational diagnostics&lt;/td&gt;
&lt;td&gt;Q2&lt;/td&gt;
&lt;td&gt;Redacted, correlated and rate-bounded diagnostics&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Portable tooling&lt;/td&gt;
&lt;td&gt;Q6&lt;/td&gt;
&lt;td&gt;Browser checks and separate source/script typechecks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Strict agent inputs&lt;/td&gt;
&lt;td&gt;Q8&lt;/td&gt;
&lt;td&gt;Shared action policy and strict integer validation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Health version&lt;/td&gt;
&lt;td&gt;Q9&lt;/td&gt;
&lt;td&gt;Health response derived from package version&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bounded document discovery&lt;/td&gt;
&lt;td&gt;Q11&lt;/td&gt;
&lt;td&gt;Short-lived cache with capacity and recovery boundaries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Handoff provenance&lt;/td&gt;
&lt;td&gt;Q12&lt;/td&gt;
&lt;td&gt;Persisted provenance, authorization checks and schema migration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Public contributor documentation&lt;/td&gt;
&lt;td&gt;Q7&lt;/td&gt;
&lt;td&gt;Accurate contributor guides and explicit documentation gates&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Release curation&lt;/td&gt;
&lt;td&gt;Q1&lt;/td&gt;
&lt;td&gt;Curated source history, portable evidence and reproducible package&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Scope and method
&lt;/h2&gt;

&lt;p&gt;The target was the closed task &lt;code&gt;[Complete] [Phase B] Design fixes for Q1 to Q12&lt;/code&gt;, parent &lt;code&gt;01a0a59d-ca16-76b3-8a9a-5f95389d22e3&lt;/code&gt;. The analysis included its sixteen worker threads and associated approval reviewers. Prior design work, Naruto's independent review sessions, human effort and subsequent evaluation/editorial work fall outside that population.&lt;/p&gt;

&lt;p&gt;Routine totals come from the &lt;a href="//../../../telemetry/ai-review-v2/codex/main/sessions/01a0a59d-ca16-76b3-8a9a-5f95389d22e3.json"&gt;canonical derived record&lt;/a&gt;, captured by &lt;code&gt;codex-hook-v1&lt;/code&gt; and closed by &lt;code&gt;SessionEnd&lt;/code&gt; at 2026-09-16 04:13:19 UTC.&lt;/p&gt;

&lt;p&gt;The derived record does not retain worker task names, wait outcomes, compaction boundaries or terminal error reasons. A targeted accuracy audit therefore used deterministic &lt;code&gt;jq&lt;/code&gt;, &lt;code&gt;rg&lt;/code&gt;, &lt;code&gt;find&lt;/code&gt;, Git and arithmetic against only this parent's 15,355,576-byte (approximately 15.4 MB) runtime file and selected child metadata. No model worker parsed logs. Prompt bodies, source bodies, secrets and raw output were not copied into this report.&lt;/p&gt;

&lt;p&gt;Product outcome and authorship were checked against &lt;a href="//../../SOURCE_REFERENCES.md#ref-5"&gt;the contribution record&lt;/a&gt;, &lt;a href="//../../SOURCE_REFERENCES.md#ref-9"&gt;Goku's final response&lt;/a&gt; and &lt;a href="//../../SOURCE_REFERENCES.md#ref-8"&gt;Naruto's final review&lt;/a&gt;. Product files remained read-only.&lt;/p&gt;

&lt;p&gt;The measured wall window is 47,022 seconds, from 15 September 15:09:37 UTC to 16 September 04:13:19 UTC. It includes a roughly 7h33m gap between completed parent turns, external test waits and user pauses; it is not active model time.&lt;/p&gt;

&lt;h3&gt;
  
  
  Reading the counters
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Term&lt;/th&gt;
&lt;th&gt;Meaning in this report&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Activation&lt;/td&gt;
&lt;td&gt;A recorded model invocation, rather than a user prompt or a tool call&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Processed tokens&lt;/td&gt;
&lt;td&gt;Cached input + uncached input + output; cumulative across invocations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cached input&lt;/td&gt;
&lt;td&gt;Input tokens served from the model's prompt cache; still included in processed-token accounting&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output&lt;/td&gt;
&lt;td&gt;Recorded output tokens, including reasoning where the counter exposes it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Worker thread&lt;/td&gt;
&lt;td&gt;A delegated conversation that may contain multiple follow-up assignments&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Guardian&lt;/td&gt;
&lt;td&gt;An automatic approval-review component, distinct from a worker-health watchdog&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wait-generating activation&lt;/td&gt;
&lt;td&gt;A parent invocation that issued a wait; its entire usage is associated with that invocation, not isolated instruction-by-instruction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Frozen price equivalent&lt;/td&gt;
&lt;td&gt;Recorded token categories multiplied by a dated API rate schedule; an analytical comparison rather than a measured subscription charge&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Repeatedly presenting a large context contributes to the cumulative token count even when much of that input is cached. A 158M-token total therefore does not imply 158M tokens of unique code, prose or uncached processing.&lt;/p&gt;

&lt;p&gt;The analysis proceeded from derived aggregates to component/model reconciliation, then to a targeted deterministic audit of missing orchestration fields, and finally to comparison with independent product-review evidence. This editorial revision uses those existing findings; it does not reopen the raw transcripts or rerun the product tests.&lt;/p&gt;

&lt;h2&gt;
  
  
  Delivered outcome
&lt;/h2&gt;

&lt;p&gt;The compared Git range &lt;code&gt;69686a3..bb7af7d&lt;/code&gt; contains 27 commits and changes 118 files: 16,369 insertions and 3,229 deletions. The engineering-file subset covers 46 files, 11,903 insertions and 2,767 deletions. These figures include tests and tooling and the full range also includes reviews and records; neither is a pure backend-line count.&lt;/p&gt;

&lt;p&gt;Naruto's independent closure review found no P0 or P1 and reproduced:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;source and script typecheck plus build;&lt;/li&gt;
&lt;li&gt;100/100 TAP tests;&lt;/li&gt;
&lt;li&gt;73/73 browser checks;&lt;/li&gt;
&lt;li&gt;48 documentation files and 476 links with zero errors;&lt;/li&gt;
&lt;li&gt;clean Prettier and &lt;code&gt;git diff --check&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;a one-commit, 61-file curated repository with clean history scans; and&lt;/li&gt;
&lt;li&gt;a byte-identical 133,263-byte package with SHA-256 &lt;code&gt;485f87aceaa5f4bebf7c2c49e7d55d44b68d1cab00fec1de5fc29442db911ae3&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are candidate-level checks. They do not establish production reliability or absence of defects. Naruto recorded five non-blocking observations, including handoff-citation semantics, package reproduction instructions and retained module-size debt. Human candidate acceptance and publication remained separate decisions at the final-review snapshot.&lt;/p&gt;

&lt;h2&gt;
  
  
  Canonical accounting
&lt;/h2&gt;

&lt;p&gt;The collector records 1,299 tool calls across the measured task. Token and activation totals are separated below by component; the total excludes Naruto's separate review sessions and the later performance analysis.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component class&lt;/th&gt;
&lt;th&gt;Components&lt;/th&gt;
&lt;th&gt;Activations&lt;/th&gt;
&lt;th&gt;Processed&lt;/th&gt;
&lt;th&gt;Cached input&lt;/th&gt;
&lt;th&gt;Uncached input&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;th&gt;Share&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Parent Astra&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;342&lt;/td&gt;
&lt;td&gt;36,090,696&lt;/td&gt;
&lt;td&gt;34,588,032&lt;/td&gt;
&lt;td&gt;1,337,219&lt;/td&gt;
&lt;td&gt;165,445&lt;/td&gt;
&lt;td&gt;22.82%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Worker threads&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;1,011&lt;/td&gt;
&lt;td&gt;110,255,046&lt;/td&gt;
&lt;td&gt;106,565,888&lt;/td&gt;
&lt;td&gt;3,211,069&lt;/td&gt;
&lt;td&gt;478,089&lt;/td&gt;
&lt;td&gt;69.72%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Approval-review guardians&lt;/td&gt;
&lt;td&gt;38, of which 20 non-zero&lt;/td&gt;
&lt;td&gt;152&lt;/td&gt;
&lt;td&gt;11,791,577&lt;/td&gt;
&lt;td&gt;9,690,112&lt;/td&gt;
&lt;td&gt;2,081,891&lt;/td&gt;
&lt;td&gt;19,574&lt;/td&gt;
&lt;td&gt;7.46%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;55&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1,505&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;158,137,319&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;150,844,032&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;6,630,179&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;663,108&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Reasoning output is a 238,390-token subset of output, not an additional category. Cache writes are zero. Processed tokens are not bytes transmitted, inference FLOPs, energy, quota usage or a subscription bill.&lt;/p&gt;

&lt;h3&gt;
  
  
  Model mix and frozen price equivalent
&lt;/h3&gt;

&lt;p&gt;Rates are the dossier's 15 September 2026 frozen standard API rates. They may change. &lt;code&gt;codex-auto-review&lt;/code&gt; has no recorded public rate and is excluded from dollars.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Components&lt;/th&gt;
&lt;th&gt;Activations&lt;/th&gt;
&lt;th&gt;Processed&lt;/th&gt;
&lt;th&gt;Share&lt;/th&gt;
&lt;th&gt;Frozen standard equivalent&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Astra&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;342&lt;/td&gt;
&lt;td&gt;36,090,696&lt;/td&gt;
&lt;td&gt;22.82%&lt;/td&gt;
&lt;td&gt;$56.23&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Sol&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;592&lt;/td&gt;
&lt;td&gt;69,756,461&lt;/td&gt;
&lt;td&gt;44.11%&lt;/td&gt;
&lt;td&gt;$42.21&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Luna&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;414&lt;/td&gt;
&lt;td&gt;40,236,736&lt;/td&gt;
&lt;td&gt;25.44%&lt;/td&gt;
&lt;td&gt;$1.15&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Terra&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;261,849&lt;/td&gt;
&lt;td&gt;0.17%&lt;/td&gt;
&lt;td&gt;$0.21&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;codex-auto-review&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;20 non-zero&lt;/td&gt;
&lt;td&gt;152&lt;/td&gt;
&lt;td&gt;11,791,577&lt;/td&gt;
&lt;td&gt;7.46%&lt;/td&gt;
&lt;td&gt;Unknown&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Named-model total: &lt;strong&gt;146,345,742 tokens and $99.80&lt;/strong&gt;. Pricing the same named tokens entirely as Astra gives &lt;strong&gt;$218.81&lt;/strong&gt;, a &lt;strong&gt;54.39% price substitution reduction&lt;/strong&gt;. Worker-only usage is $43.57 at the actual model mix versus $162.58 at Astra rates, a &lt;strong&gt;73.20% price substitution reduction&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The frozen calculation assigns a lower price to the recorded model mix. It holds token counts constant and supplies no evidence about the amount of work an actual all-Astra run would perform; Sol and Luna together processed 109.99M tokens in the observed run.&lt;/p&gt;

&lt;h2&gt;
  
  
  All sixteen worker threads
&lt;/h2&gt;

&lt;p&gt;Static task names became incomplete descriptions because the parent issued &lt;strong&gt;28 &lt;code&gt;followup_task&lt;/code&gt; calls&lt;/strong&gt; and &lt;strong&gt;24 &lt;code&gt;send_message&lt;/code&gt; calls&lt;/strong&gt;. The “actual work” column follows the completed worker-turn records and product contribution evidence, not the original label alone.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Worker label&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Actual observed work and terminal state&lt;/th&gt;
&lt;th&gt;Processed&lt;/th&gt;
&lt;th&gt;Activations&lt;/th&gt;
&lt;th&gt;Cache read&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;design_evidence&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Terra&lt;/td&gt;
&lt;td&gt;Read-only Q1-Q12 evidence map; no files changed&lt;/td&gt;
&lt;td&gt;261,849&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;73.81%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;commit1_standard&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Sol&lt;/td&gt;
&lt;td&gt;B1-B4 engineering-standard document&lt;/td&gt;
&lt;td&gt;246,654&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;73.59%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;commit1_checker&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Sol&lt;/td&gt;
&lt;td&gt;Documentation checker, tests and bounded corrections&lt;/td&gt;
&lt;td&gt;1,500,331&lt;/td&gt;
&lt;td&gt;24&lt;/td&gt;
&lt;td&gt;94.83%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;commit1_navigation&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Luna&lt;/td&gt;
&lt;td&gt;Eleven-LLD index, plan link and citation corrections&lt;/td&gt;
&lt;td&gt;559,111&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;89.74%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;commit1_gate&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Luna&lt;/td&gt;
&lt;td&gt;Commit-1 integrated gate; 56/56 tests&lt;/td&gt;
&lt;td&gt;525,185&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;95.20%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;midpoint_merge&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Sol&lt;/td&gt;
&lt;td&gt;Reused for Q10 transaction guard and regressions&lt;/td&gt;
&lt;td&gt;1,871,363&lt;/td&gt;
&lt;td&gt;34&lt;/td&gt;
&lt;td&gt;96.66%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;q5_handoff&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Sol&lt;/td&gt;
&lt;td&gt;Seven Q5 handoff-route tests and expiry sensitivity&lt;/td&gt;
&lt;td&gt;2,775,134&lt;/td&gt;
&lt;td&gt;34&lt;/td&gt;
&lt;td&gt;96.53%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;q34_browser&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Sol&lt;/td&gt;
&lt;td&gt;Q3/Q4 browser split, mutation lifecycle and regressions&lt;/td&gt;
&lt;td&gt;9,609,041&lt;/td&gt;
&lt;td&gt;81&lt;/td&gt;
&lt;td&gt;98.07%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;q34_server&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Sol&lt;/td&gt;
&lt;td&gt;Q3/Q4 server types, handlers, retry and credential rotation&lt;/td&gt;
&lt;td&gt;9,473,162&lt;/td&gt;
&lt;td&gt;68&lt;/td&gt;
&lt;td&gt;96.72%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;midpoint_gates&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Luna&lt;/td&gt;
&lt;td&gt;Reused for integration, verification and midpoint records&lt;/td&gt;
&lt;td&gt;3,968,535&lt;/td&gt;
&lt;td&gt;68&lt;/td&gt;
&lt;td&gt;93.92%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;final_merge&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Luna&lt;/td&gt;
&lt;td&gt;Rebase/integration, Q2-Q12 gates, packaging, two long artifact runs and final freeze&lt;/td&gt;
&lt;td&gt;33,364,102&lt;/td&gt;
&lt;td&gt;287&lt;/td&gt;
&lt;td&gt;97.85%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;q6_portability&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Sol&lt;/td&gt;
&lt;td&gt;Reused for Q6, Q12 and B-F/B-D baseline work; final turn ended at model capacity after earlier deliverables&lt;/td&gt;
&lt;td&gt;20,249,883&lt;/td&gt;
&lt;td&gt;165&lt;/td&gt;
&lt;td&gt;96.05%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;q2_diagnostics&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Sol&lt;/td&gt;
&lt;td&gt;Reused for Q2, Q8, Q9 and integrated Q11 work&lt;/td&gt;
&lt;td&gt;10,394,239&lt;/td&gt;
&lt;td&gt;83&lt;/td&gt;
&lt;td&gt;96.43%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;q11_discovery&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Sol&lt;/td&gt;
&lt;td&gt;Reused for Q11, Q1 tooling and curation; final closing-record turn ended at model capacity after earlier deliverables&lt;/td&gt;
&lt;td&gt;13,636,654&lt;/td&gt;
&lt;td&gt;97&lt;/td&gt;
&lt;td&gt;96.51%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;closing_records&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Luna&lt;/td&gt;
&lt;td&gt;Six-document closing packet&lt;/td&gt;
&lt;td&gt;918,820&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;91.19%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;closure_merge&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Luna&lt;/td&gt;
&lt;td&gt;Final fast-forward while preserving five local edits&lt;/td&gt;
&lt;td&gt;900,983&lt;/td&gt;
&lt;td&gt;17&lt;/td&gt;
&lt;td&gt;94.19%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Four threads—&lt;code&gt;final_merge&lt;/code&gt;, &lt;code&gt;q6_portability&lt;/code&gt;, &lt;code&gt;q11_discovery&lt;/code&gt; and &lt;code&gt;q2_diagnostics&lt;/code&gt;—were used across multiple stages and together consumed 77,644,878 tokens: 70.42% of worker usage and 49.10% of the task. A follow-up can carry accumulated conversation context into later invocations. The aggregate concentration is observed; the fraction caused by reuse, and the benefit of starting fresh workers, remain unmeasured.&lt;/p&gt;

&lt;p&gt;Two Sol threads ended their last follow-up with &lt;code&gt;Selected model is at capacity&lt;/code&gt; after completing earlier stages. The derived record still reports zero interruptions. This is not a contradiction in the current schema—capacity failure is not counted as &lt;code&gt;interruptions&lt;/code&gt;—but it is a telemetry gap: component terminal outcomes should be explicit rather than inferred from raw child events.&lt;/p&gt;

&lt;p&gt;The approved telemetry-extension work, ED-6, is the relevant workstream. Its current enumerated fields do not explicitly include terminal outcomes or follow-up counts; those additions need an explicit specification rather than being described as already installed or already enumerated in the decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  Parent orchestration
&lt;/h2&gt;

&lt;p&gt;The parent recorded 330 tool calls:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Parent action&lt;/th&gt;
&lt;th&gt;Calls&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Shell/tool execution&lt;/td&gt;
&lt;td&gt;170&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Spawn worker&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Follow-up assignment&lt;/td&gt;
&lt;td&gt;28&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Worker message/steering&lt;/td&gt;
&lt;td&gt;24&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;wait_agent&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;89&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;sleep&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Orchestration and waiting were 160 of 330 tool calls, &lt;strong&gt;48.48%&lt;/strong&gt;. Tool-call count is not token cost, but it demonstrates how much of the parent's control flow became coordination rather than direct engineering.&lt;/p&gt;

&lt;p&gt;The canonical lifecycle record contains 12 parent &lt;code&gt;Stop&lt;/code&gt; events. Dividing 342 parent activations by those events gives &lt;strong&gt;28.5 activations per recorded stop&lt;/strong&gt;. This is a descriptive average, not a per-prompt compliance result: it hides the distribution and uses lifecycle events as its denominator.&lt;/p&gt;

&lt;h3&gt;
  
  
  The one-minute polling loop
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;wait_agent&lt;/code&gt;: 89 calls; 88 × 60 seconds, 1 × 10 seconds.&lt;/li&gt;
&lt;li&gt;Outcomes: 66 timeout-like, 23 completion-like.&lt;/li&gt;
&lt;li&gt;Wait-generating activations: 10,281,999 processed; 10,153,856 cached input; 96,038 uncached input; 32,105 output, of which 28,105 reasoning.&lt;/li&gt;
&lt;li&gt;Share: 28.49% of canonical parent tokens and 6.50% of the whole task.&lt;/li&gt;
&lt;li&gt;Mean per wait-generating activation: 115,528 processed tokens.&lt;/li&gt;
&lt;li&gt;Frozen Astra price equivalent: &lt;strong&gt;$12.72&lt;/strong&gt;, or 22.62% of parent Astra price equivalent and 12.74% of all known-price usage.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The largest parent work turn issued 62 waits, 50 of which timed out. Those wait-generating activations consumed 7,443,509 tokens. Another turn issued 23 waits, 14 timeouts, consuming 2,377,824 tokens.&lt;/p&gt;

&lt;p&gt;The runtime also recorded 42 &lt;code&gt;SubagentStop&lt;/code&gt; lifecycle events while only 23 waits returned completion-like output. These counts are not one-to-one because one worker thread can run multiple follow-ups and events can batch, but they establish that completion delivery existed independently of a fresh one-minute status check. The polling loop was not the only way completions reached the parent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Finding:&lt;/strong&gt; the global declarative no-poll guidance did not enforce event-driven behavior when the parent had no independent work. This task reproduces the previously observed active-versus-idle split at much larger scale. The causal contribution of platform instructions versus parent choice is not isolated here.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A["Parent activation"] --&amp;gt; B["Issue wait: usually 60 seconds"]
    B --&amp;gt; C{"Wait result"}
    C --&amp;gt;|"Timeout: 66 returns"| D["Parent resumes; may issue another wait"]
    D --&amp;gt; A
    C --&amp;gt;|"Completion-like: 23 returns"| E["Process worker result"]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The diagram shows the observed return categories; it does not count uninterrupted timeout-to-wait sequences. &lt;strong&gt;66 timeouts is not automatically 66 repeated waits with no intervening work.&lt;/strong&gt; The minimum requested timeout was &lt;strong&gt;10 seconds&lt;/strong&gt;, while the dominant cadence was 60 seconds.&lt;/p&gt;

&lt;p&gt;The 10.28M tokens and $12.72 cover all 89 wait-generating activations, including waits that returned completion-like output. Multiplying 66 by the all-wait mean of 115,528 gives approximately 7.6M tokens, but that estimate is neither an exact timeout-only total nor an upper bound on causal savings. An event-driven replacement still has to process completion, failures and recovery.&lt;/p&gt;

&lt;h2&gt;
  
  
  Compaction
&lt;/h2&gt;

&lt;p&gt;Five runtime compactions occurred. Turn boundaries indicate that numbers 1 and 3 happened inside long work turns and were likely automatic; numbers 2, 4 and 5 were separate compaction turns and were likely manual. That classification is an inference; the raw compaction record does not label its trigger.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Compaction&lt;/th&gt;
&lt;th&gt;Input before&lt;/th&gt;
&lt;th&gt;First input after&lt;/th&gt;
&lt;th&gt;Immediate drop&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;191,950&lt;/td&gt;
&lt;td&gt;32,181&lt;/td&gt;
&lt;td&gt;83.23%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;138,145&lt;/td&gt;
&lt;td&gt;33,523&lt;/td&gt;
&lt;td&gt;75.73%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;209,562&lt;/td&gt;
&lt;td&gt;34,271&lt;/td&gt;
&lt;td&gt;83.65%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;128,216&lt;/td&gt;
&lt;td&gt;32,921&lt;/td&gt;
&lt;td&gt;74.32%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;50,518&lt;/td&gt;
&lt;td&gt;32,097&lt;/td&gt;
&lt;td&gt;36.46%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Compaction clearly reduced the immediate context carried into the next activation. It did not reduce the need for architectural boundaries: the parent later rebuilt a large context, and the post-compaction long turn alone consumed 20.43M raw parent tokens while issuing 62 waits.&lt;/p&gt;

&lt;p&gt;The parent raw ledger contains 347 token-usage records totalling 36,832,258 tokens; the canonical parent component contains 342 activations and 36,090,696 tokens. The &lt;strong&gt;741,562-token difference&lt;/strong&gt; aligns with the five compaction activations omitted by the derived parent aggregate. Therefore &lt;code&gt;capture_completeness: complete&lt;/code&gt; means all expected worker components were found; it does not mean all model work is represented. Whole-task model work is at least &lt;strong&gt;158,878,881 tokens&lt;/strong&gt;, before checking compaction omissions inside workers and guardians. That broader raw reconstruction was deliberately stopped rather than assigning a model to parse every child log.&lt;/p&gt;

&lt;h2&gt;
  
  
  Approval-review guardians
&lt;/h2&gt;

&lt;p&gt;The collector associated 38 guardian components with this logical chat:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;20 non-zero &lt;code&gt;codex-auto-review&lt;/code&gt; components: 11,791,577 processed tokens and 152 activations;&lt;/li&gt;
&lt;li&gt;18 zero-token components, which appear as lifecycle stubs rather than model work;&lt;/li&gt;
&lt;li&gt;largest guardian: 6,340,410 tokens, 42 approval turns over 4,176 seconds, including two of its own compactions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are low-effort automatic approval reviewers. They have zero tool calls and one turn per approval decision. They are &lt;strong&gt;not&lt;/strong&gt; watchdogs supervising worker health, and their presence must not be used to claim watchdog coverage. Their price is unknown, but their 7.46% token share is operationally material.&lt;/p&gt;

&lt;p&gt;The review identified approval configuration as a further investigation: determine which approval requests caused these invocations and how the reviewer setting affected them. This report does not establish the task's configuration-to-cost relationship or evaluate the safety and human-review cost of changing that configuration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Comparison with the earlier build and evaluation
&lt;/h2&gt;

&lt;p&gt;The earlier isolated implementation audit counted 24,995,374 full-scope tokens after adding compaction and approval-review usage. This Phase B task's 158,137,319 canonical tokens are &lt;strong&gt;6.33×&lt;/strong&gt; that historical build total. The workloads are not matched: Phase B covered 12 findings, repeated integration, packaging, long artifact tests, curation and closing records.&lt;/p&gt;

&lt;p&gt;The earlier efficiency-analysis task captured 14,932,980 tokens. The owner remembers it as roughly half the original task; comparing that snapshot with the 24,995,374 full-scope build gives 59.74%, but the historical review warned that the populations were not strictly aligned. Within that analysis, the Sol worker consumed &lt;strong&gt;7,535,274 tokens, 50.46%&lt;/strong&gt; of the analysis total.&lt;/p&gt;

&lt;p&gt;The owner's clarification supplies the reported workflow cause: Sol was asked to do log parsing after the existing Python helpers failed to produce the needed answer. That causal account is owner-supplied; this report did not repeat the old raw-log investigation.&lt;/p&gt;

&lt;p&gt;The approved recovery decision, &lt;a href="//../../../DECISIONS.md#ed-2-instrumentation-gaps-and-audit-budget"&gt;ED-2&lt;/a&gt;, requires a budget check before starting, a stop and costed script-change proposal when instrumentation is insufficient, and no ad hoc raw reconstruction inside the requested analysis. If estimated work exceeds the budget, approval comes before starting. If a running task crosses the budget, it ends with completed work and a costed proposal for the remainder; the human decides in the next message.&lt;/p&gt;

&lt;p&gt;The original Phase B analysis used derived telemetry plus bounded deterministic extraction from one parent file and selected child metadata, with no analysis subagent. That describes its method; it does not establish compliance with the approved workflow or budget. The earlier sentence claiming that this report “followed that rule” has been removed.&lt;/p&gt;

&lt;h3&gt;
  
  
  The evaluator's own consumption
&lt;/h3&gt;

&lt;p&gt;Naruto's review identified a missing measurement: the original report did not disclose its enclosing analysis session's usage. That session also performed quick-reference preservation work, so derived aggregates cannot isolate the cost of writing the report.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Snapshot of Goku's evaluation session&lt;/th&gt;
&lt;th&gt;Processed&lt;/th&gt;
&lt;th&gt;Cached input&lt;/th&gt;
&lt;th&gt;Uncached input&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;th&gt;Share of assessed task&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Review snapshot, 16 September 08:07:40 UTC&lt;/td&gt;
&lt;td&gt;7,420,764&lt;/td&gt;
&lt;td&gt;6,378,112&lt;/td&gt;
&lt;td&gt;982,366&lt;/td&gt;
&lt;td&gt;60,286&lt;/td&gt;
&lt;td&gt;4.69%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Later snapshot, 16 September 08:21:09 UTC, after the response to Naruto's review&lt;/td&gt;
&lt;td&gt;8,746,798&lt;/td&gt;
&lt;td&gt;7,428,864&lt;/td&gt;
&lt;td&gt;1,253,044&lt;/td&gt;
&lt;td&gt;64,890&lt;/td&gt;
&lt;td&gt;5.53%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The first snapshot records Sol only, 62 activations, 55 tool calls, seven &lt;code&gt;Stop&lt;/code&gt; events and no workers. Reasoning accounts for 26,541 of its output tokens. The later snapshot records 74 activations, 66 tool calls, eight &lt;code&gt;Stop&lt;/code&gt; events and no workers; reasoning is 29,018 of its output tokens. The later record is available in the &lt;a href="//../../../telemetry/ai-review-v2-performance/codex/main/sessions/01a0a906-7e54-7770-bb20-9730aeb2e658.json"&gt;canonical evaluation session&lt;/a&gt;; the earlier counters were read and confirmed during the review response. This active record evolves, so it should not be expected to retain either snapshot as its latest total.&lt;/p&gt;

&lt;p&gt;Both ratios use the same 158,137,319-token assessed-task denominator. They are gross enclosing-session ratios, not isolated analysis overhead. The first is a conservative envelope for analysis work represented by that snapshot, subject to collector completeness. Neither includes the editorial work producing this revision or Naruto's cost of reviewing the analysis.&lt;/p&gt;

&lt;p&gt;The approved audit target, WC-7, is &lt;code&gt;min(2% of assessed tokens, 1,000,000)&lt;/code&gt;—&lt;strong&gt;1,000,000 tokens for this task&lt;/strong&gt;. The first enclosing-session snapshot is 7.42 times that target. Enforcement was not installed at the review snapshot; the mixed session prevents an exact analysis-only compliance calculation. Having no analysis worker is an observed orchestration fact, not proof of low evaluation cost.&lt;/p&gt;

&lt;p&gt;The historical approximately 60% ratio and this 4.69% snapshot have different populations and coverage. They are useful context, but do not establish a controlled reduction in evaluation overhead.&lt;/p&gt;

&lt;h2&gt;
  
  
  What remains to be tested
&lt;/h2&gt;

&lt;p&gt;The existing &lt;a href="//../../../DECISIONS.md"&gt;decision register&lt;/a&gt; separates approved direction from installed mechanisms. At this case-study snapshot, ED-1 approved an interim five-minute Codex wait; ED-4/ED-5 approved a deterministic watchdog; ED-2 approved the audit budget and instrumentation-gap stopping rule; and ED-6 approved telemetry extensions. This report implements none of them. A deterministic watchdog can inspect counters without invoking an LLM, but still consumes CPU, memory and I/O.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Follow-up&lt;/th&gt;
&lt;th&gt;Reason for examining it&lt;/th&gt;
&lt;th&gt;Evidence needed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Extend the collectors before another comparison&lt;/td&gt;
&lt;td&gt;Routine fields should not require another raw-log audit&lt;/td&gt;
&lt;td&gt;Derived wait counts/outcomes/repeats, compaction usage, context bands and dated price calculations; explicitly specify additional lineage, follow-up, terminal-error and external-wait fields&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Address recurring parent waits first&lt;/td&gt;
&lt;td&gt;All wait-generating activations account for 6.50% of recorded tokens and 12.74% of the known-price equivalent&lt;/td&gt;
&lt;td&gt;Matched active/idle-parent scenarios and reliable completion, failure and lost-notification handling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bound worker follow-ups and activations&lt;/td&gt;
&lt;td&gt;Four reused workers account for 49.10% of recorded tokens&lt;/td&gt;
&lt;td&gt;A comparison of retained versus fresh worker context, including rework and accepted quality&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Examine approval-review configuration&lt;/td&gt;
&lt;td&gt;Guardians account for 7.46% of recorded tokens at unknown price&lt;/td&gt;
&lt;td&gt;Trigger attribution, useful review outcomes, latency and security consequences of any alternative&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Budget the next evaluation before execution&lt;/td&gt;
&lt;td&gt;The evaluator's own session cost is material&lt;/td&gt;
&lt;td&gt;A declared budget, bounded measurement scope, stop behavior and a separately reported evaluation total&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The ordering reflects measurement confidence and implementability as well as price. Worker concentration and wait-generating activations measure different things; neither estimates removable waste. The large Luna &lt;code&gt;final_merge&lt;/code&gt; thread has a frozen price equivalent below one dollar, illustrating why token and price rankings can differ while neither captures latency or review effort.&lt;/p&gt;

&lt;p&gt;Validation would also need long external work, runaway behavior, parent restart and model-capacity failure cases. It should record total/cached/uncached/output tokens, elapsed time, guardian usage, rework and accepted quality. Fresh-worker benefits and net savings from event-driven orchestration remain hypotheses until that comparison exists.&lt;/p&gt;

&lt;h3&gt;
  
  
  Baseline measurements for that comparison
&lt;/h3&gt;

&lt;p&gt;These rows describe Phase B under the existing &lt;a href="//../../../WINNING-CRITERIA.md"&gt;winning-criteria definitions&lt;/a&gt;. They do not update or approve the shared criteria.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Criterion&lt;/th&gt;
&lt;th&gt;Defensible baseline from this analysis&lt;/th&gt;
&lt;th&gt;Limit&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;WC-1: repeated waits without other work&lt;/td&gt;
&lt;td&gt;89 waits and 66 timeouts; exact qualifying repeat count uncomputed&lt;/td&gt;
&lt;td&gt;Timeout count cannot substitute for sequence classification&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;WC-2: smallest requested agent-wait timeout&lt;/td&gt;
&lt;td&gt;10 seconds; 88 of 89 requests used 60 seconds&lt;/td&gt;
&lt;td&gt;Both are below the approved interim 300-second target&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;WC-5: parent step budget&lt;/td&gt;
&lt;td&gt;342 activations / 12 recorded stops = 28.5 average&lt;/td&gt;
&lt;td&gt;Does not establish the fraction of prompts within twelve steps&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;WC-7: audit budget&lt;/td&gt;
&lt;td&gt;7.42M / 4.69% at the first review snapshot; 1M target&lt;/td&gt;
&lt;td&gt;Enclosing session includes other work; later review/editorial cost must be reported separately&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;WC-9: complete counting&lt;/td&gt;
&lt;td&gt;At least 741,562 parent compaction tokens omitted&lt;/td&gt;
&lt;td&gt;Worker and guardian compaction omissions not fully audited&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Evidence status and limits
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Canonical totals and component counters: &lt;code&gt;LOCAL-CONFIRMED&lt;/code&gt; for the named closed session.&lt;/li&gt;
&lt;li&gt;Wait calls, outcomes, compaction boundaries, follow-up counts and two capacity errors: &lt;code&gt;LOCAL-CONFIRMED&lt;/code&gt; by targeted deterministic raw audit.&lt;/li&gt;
&lt;li&gt;Manual-versus-automatic compaction classification: inference from turn boundaries.&lt;/li&gt;
&lt;li&gt;Prior Sol parser assignment: owner-supplied current clarification; the old raw analysis was not repeated.&lt;/li&gt;
&lt;li&gt;Price figures: frozen 15 September 2026 standard API equivalents, not a bill; guardian price unknown.&lt;/li&gt;
&lt;li&gt;Causal savings, fresh-worker benefit and watchdog effectiveness: unproven until controlled validation.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is a single-task observational case study. There is no matched alternative orchestration run, and no claim of measured equal-quality savings. Product tests were not rerun for this performance report. No collector, raw log, product file, hook, global instruction or runtime policy was modified.&lt;/p&gt;

&lt;h2&gt;
  
  
  Evidence and revision notes
&lt;/h2&gt;

&lt;p&gt;Original runtime analysis: Goku, recorded as GPT-5.6 Sol. The engineering task being measured used Goku on GPT-6 Astra. Model names follow recorded identifiers. This revision incorporates Naruto's analysis review and the human developer's clarification of project context.&lt;/p&gt;

&lt;p&gt;The tables above retain the original component totals, all sixteen worker rows, the five compaction boundaries, the Git-range measurements and the independent delivery checks. Source links are local evidence references within the evaluation dossier, not claims that the underlying private records are publicly accessible.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Evidence&lt;/th&gt;
&lt;th&gt;What it supports&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="//../../../telemetry/ai-review-v2/codex/main/sessions/01a0a59d-ca16-76b3-8a9a-5f95389d22e3.json"&gt;Closed implementation telemetry&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Canonical component, token, activation and lifecycle totals&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="//../../SOURCE_REFERENCES.md#ref-5"&gt;Product contribution record&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Actual engineering contributions and model roles&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="//../../SOURCE_REFERENCES.md#ref-10"&gt;LLD and milestone source references&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Eleven designs, execution order and checkpoint boundaries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="//../../SOURCE_REFERENCES.md#ref-9"&gt;Implementation response&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Finding dispositions, implementation gates and candidate evidence&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="//../../SOURCE_REFERENCES.md#ref-8"&gt;Naruto's final product review&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Independent reproduction, closure verdict and remaining observations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="//2026-09-15-goku-mechanism-followup.md"&gt;Frozen pricing analysis&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Dated rate basis and price-substitution method&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;a href="//../../../DECISIONS.md"&gt;Decision register&lt;/a&gt; and &lt;a href="//../../../WINNING-CRITERIA.md"&gt;criteria&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;Approved direction and measurement definitions, separate from installation status&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Post-review corrections: the original analysis byline changed from Astra to recorded Sol; evaluator-session consumption was added with dated cutoffs; ED-2 wording was aligned; wait timeouts, repeats and savings were separated; the minimum wait was corrected to 10 seconds for WC-2; and the uninstalled collector extensions were distinguished from proposed additional fields. The human developer supplied the one-developer/two-principal-agent framing and the alpha/mid-sized product description. Naruto's independently authored reviews remain unchanged.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>chatgpt</category>
      <category>tutorial</category>
    </item>
  </channel>
</rss>
