<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sergei Parfenov</title>
    <description>The latest articles on DEV Community by Sergei Parfenov (@p0rt).</description>
    <link>https://dev.to/p0rt</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F157612%2F00065fc9-07d7-47dd-b882-f297a6158dbe.jpeg</url>
      <title>DEV Community: Sergei Parfenov</title>
      <link>https://dev.to/p0rt</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/p0rt"/>
    <language>en</language>
    <item>
      <title>Same Patient, Conflicting Documents: Can AI Preserve the Evidence?</title>
      <dc:creator>Sergei Parfenov</dc:creator>
      <pubDate>Wed, 30 Sep 2026 10:20:40 +0000</pubDate>
      <link>https://dev.to/p0rt/same-patient-conflicting-documents-can-ai-preserve-the-evidence-3anp</link>
      <guid>https://dev.to/p0rt/same-patient-conflicting-documents-can-ai-preserve-the-evidence-3anp</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;In the first chronological run of one synthetic history, a record built from GPT-6 Luna's extractions returned the expected answer to &lt;strong&gt;21 of 22 final queries&lt;/strong&gt;. Yet it had lost an earlier source assertion, hiding a conflict before a correction arrived. The accepted value at the end was right; the documented history was incomplete.&lt;/p&gt;

&lt;p&gt;A patient uploads a laboratory report, an older discharge summary and a photograph of a prescription. A correction arrives later. The same report appears twice. One note describes the patient's mother; another says a diagnosis is suspected. The dates inside the documents do not follow the order in which the files arrived.&lt;/p&gt;

&lt;p&gt;I'm a technical consultant at &lt;a href="https://symptomato.com" rel="noopener noreferrer"&gt;Symptomato&lt;/a&gt;, where AI gathers context before a human health specialist joins the conversation. For Symptomato, we are building a way to turn incoming documents into a longitudinal, source-grounded record.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ChartReplay&lt;/strong&gt; asks whether an update policy preserves every required assertion and its relationships as corrections and repeated documents arrive. A record fails this content test if it omits a required assertion, adds unsupported content or misrepresents its subject, value, time, uncertainty or relationship to other sources.&lt;/p&gt;

&lt;h3&gt;
  
  
  Trust starts with knowing what the record can establish
&lt;/h3&gt;

&lt;p&gt;There are two different questions here. Did the system preserve what the documents actually said? And are those documents themselves accurate accounts of what happened to the patient?&lt;/p&gt;

&lt;p&gt;This experiment tests the first. A faithful record can still contain an erroneous report, a mistaken patient recollection or unresolved disagreement. Establishing authenticity and medical accuracy requires evidence and review beyond this benchmark. The system should expose those questions rather than silently answer them through a confident rewrite.&lt;/p&gt;

&lt;p&gt;In ChartReplay, &lt;strong&gt;accepted&lt;/strong&gt; is a state under explicit record-merging rules. It does not mean independently verified, clinically current or true in the world. A statement can be accepted because no visible source disputes or supersedes it. If we lose a conflicting source, that state can become misleading.&lt;/p&gt;

&lt;p&gt;Provenance makes such decisions inspectable. &lt;a href="https://hl7.org/fhir/R4/provenance.html" rel="noopener noreferrer"&gt;FHIR's Provenance resource&lt;/a&gt; describes how information was created or revised and which entities and agents were involved. ChartReplay is not a FHIR implementation, but the relevant principle is the same: the record needs an inspectable relationship to its sources.&lt;/p&gt;

&lt;h3&gt;
  
  
  The record has to preserve relationships, not just values
&lt;/h3&gt;

&lt;p&gt;I defined a source contract before running the models. It specifies what should survive an update:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What arrives&lt;/th&gt;
&lt;th&gt;What the record must preserve&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Two different values for the same documented event&lt;/td&gt;
&lt;td&gt;Both assertions and the unresolved conflict; upload order does not choose a winner&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;An explicit correction&lt;/td&gt;
&lt;td&gt;The corrected assertion, its links to the originals, and the originals as superseded history&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A repeat upload&lt;/td&gt;
&lt;td&gt;The existing assertion, without inventing another measurement&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A later independent measurement&lt;/td&gt;
&lt;td&gt;A separate event, not a correction of the earlier one&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A family-history statement or suspected diagnosis&lt;/td&gt;
&lt;td&gt;The correct subject and the stated uncertainty&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A medication order or an absent entry&lt;/td&gt;
&lt;td&gt;What the source states, without inferring medication consumption or an explicit negative&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For example, a prescription does not establish that someone took a medication. An absent allergy entry does not establish that the patient denied allergies. These distinctions follow the source semantics described by &lt;a href="https://hl7.org/fhir/R4/medicationrequest.html" rel="noopener noreferrer"&gt;FHIR MedicationRequest&lt;/a&gt; and &lt;a href="https://hl7.org/fhir/R4/allergyintolerance.html" rel="noopener noreferrer"&gt;AllergyIntolerance&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Every retained assertion is checked against its supporting source, including subject, event time, value, uncertainty and correction relationships. A citation to an unrelated document does not count. The evaluation includes superseded assertions and unresolved conflicts, not just whichever value appears last.&lt;/p&gt;

&lt;h3&gt;
  
  
  Testing one part of the upload problem under controlled conditions
&lt;/h3&gt;

&lt;p&gt;The model pilot uses &lt;strong&gt;ten self-authored synthetic histories&lt;/strong&gt;, one from each logical family. It covers corrections, conflicts and reconciliation, family attribution, uncertainty, medication orders, duplicate versus independent events, time offsets, similar-looking sources, multi-assertion notes and longer histories.&lt;/p&gt;

&lt;p&gt;Each history is delivered chronologically, in a seeded shuffle and with duplicate replays, with two repetitions of each schedule. The unique source set is the same when delivery finishes. At intermediate checkpoints, the expected record is reconstructed from only the sources received so far: a system cannot act on a correction it has not received.&lt;/p&gt;

&lt;p&gt;These are controlled text fixtures with explicit source and event identifiers. They isolate the update problem; they do not reproduce all the difficulty of the opening upload scenario. No photographs, OCR or unrestricted clinical prose enter this pilot. Most documents contain one assertion. A deterministic parser written for the known grammar passes the calibration set without an LLM.&lt;/p&gt;

&lt;p&gt;The broader fixture generator contains 120 cases, 1,020 documents and 1,092 assertions. The model results below use the ten pilot histories, not all 120 cases. Repeated episodes are not independent patients.&lt;/p&gt;

&lt;h3&gt;
  
  
  Three ways to maintain the same history
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Full rebuild&lt;/strong&gt; asks the model to construct the whole chart again from all available originals.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rolling rewrite&lt;/strong&gt; gives it the previous chart, new documents and the complete available archive, and asks for an updated chart.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Extract plus ledger&lt;/strong&gt; asks the model to extract assertions from newly delivered documents. Deterministic software then applies the declared correction and conflict rules.&lt;/p&gt;

&lt;p&gt;The third design has an important constraint: its prompt explicitly prohibits re-extracting archive-only documents. Rebuild and rewrite can recover an earlier omission by reading the originals again. The extractor can see the archive for context, but does not revisit old documents for extraction unless they are delivered again. There is no separate completeness check before the ledger accepts an extraction.&lt;/p&gt;

&lt;p&gt;This is a comparison of those particular update policies, including that asymmetry. It does not establish the performance of every ledger-based design. All three can retain the full history without an imposed chart-size budget; provider context and response limits still apply. Each call starts a fresh conversation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;The primary models are &lt;a href="https://developers.openai.com/api/docs/models/gpt-6-luna" rel="noopener noreferrer"&gt;GPT-6 Luna&lt;/a&gt; (&lt;code&gt;gpt-6-luna&lt;/code&gt;), &lt;a href="https://developers.openai.com/api/docs/models/gpt-6-sol" rel="noopener noreferrer"&gt;GPT-6 Sol&lt;/a&gt; (&lt;code&gt;gpt-6-sol&lt;/code&gt;) and &lt;a href="https://platform.claude.com/docs/en/models/sonnet-5/overview" rel="noopener noreferrer"&gt;Claude Sonnet 5&lt;/a&gt; (&lt;code&gt;claude-sonnet-5&lt;/code&gt;). &lt;strong&gt;Gemini 3.8 Flash&lt;/strong&gt; (&lt;code&gt;google/gemini-3.8-flash&lt;/code&gt;) supplies two separate Kaggle phases.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Evidence&lt;/th&gt;
&lt;th&gt;Episodes and models&lt;/th&gt;
&lt;th&gt;Execution and overlap&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Primary comparison&lt;/td&gt;
&lt;td&gt;540 episodes; Luna, Sol and Sonnet 5&lt;/td&gt;
&lt;td&gt;Direct APIs: OpenAI for Luna/Sol, Anthropic for Sonnet&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Matched secondary comparison&lt;/td&gt;
&lt;td&gt;288 outcomes; the same three models plus Gemini&lt;/td&gt;
&lt;td&gt;72 Gemini episodes through Kaggle + 216 reused primary episodes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Public smoke benchmark&lt;/td&gt;
&lt;td&gt;12 episodes; Gemini only&lt;/td&gt;
&lt;td&gt;Separate Kaggle runs on two histories&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Each view includes all three update architectures. The primary comparison has the same 60 case–schedule–repetition combinations for each architecture and model. Its complete three-model subset was selected retrospectively from an original five-model lineup by coverage, not by correctness. Every terminal outcome is included.&lt;/p&gt;

&lt;p&gt;The secondary view matches all four models on Gemini's 24 available combinations. The public smoke is a smaller demonstration; its task scores are reported separately below. The direct-provider results are not Kaggle-hosted leaderboard entries.&lt;/p&gt;

&lt;p&gt;The original lineup was recorded on September 25, 2026, with recent releases and sustainable recurring cost as the practical criteria. &lt;a href="https://developers.openai.com/api/docs/changelog" rel="noopener noreferrer"&gt;Luna and Sol were released on September 22&lt;/a&gt;. Sonnet 5, released June 30, was the explicitly chosen cheaper Anthropic tier, not the provider's latest overall release. These are dated selection decisions, not a claim about the current model catalog.&lt;/p&gt;

&lt;p&gt;Inputs, prompts and profiles were frozen. Luna and Sol used medium reasoning effort; Sonnet used adaptive reasoning with high effort. These labels do not establish equal compute. Models ran at different times, and the Kaggle route has its own recorded settings. Each model's smoke test preceded its own pilot.&lt;/p&gt;

&lt;p&gt;The primary score requires an episode to finish under the declared interface and produce an exact final chart. Format failures remain unsuccessful outcomes. The direct-provider arms requested JSON through the prompt, without provider-enforced structured output. This matters for interpreting Sonnet's results; it is separate from the question of whether a completed record preserved the evidence. Saved, unmodified responses and first terminal outcomes are evaluated by the frozen rules, with no judge LLM.&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Losing one assertion hid a conflict
&lt;/h3&gt;

&lt;p&gt;In the reconciliation case, the first source reports &lt;strong&gt;138&lt;/strong&gt;. A second source reports &lt;strong&gt;145&lt;/strong&gt; for the same patient, event, concept and event time. Neither supersedes the other. At that point, the correct record contains both assertions, marked disputed.&lt;/p&gt;

&lt;p&gt;A third source explicitly corrects both to &lt;strong&gt;140&lt;/strong&gt;. Only then does the contract accept 140 and retain the two earlier assertions as superseded history. These numbers use synthetic units; the example concerns document relationships, not interpretation of a clinical measurement.&lt;/p&gt;

&lt;p&gt;In both chronological repetitions, Luna's extractor returned no assertion for the first source. When the second source arrived, the ledger contained only 145, marked accepted. The software had no extracted counterpart with which to identify the conflict.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa3egbr0m98ldr4g81ii3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa3egbr0m98ldr4g81ii3.png" alt="A three-step saved synthetic trace. After 138 is omitted, the later conflicting value 145 is marked accepted instead of disputed. A correction to 140 yields the expected accepted value, but the history of 138 remains missing." width="800" height="1062"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Observed Luna plus ledger trace, chronological delivery. “Accepted,” “disputed” and “superseded” are benchmark states, not judgments of medical truth. The figure shows the first three uploads of the six-document case.&lt;/em&gt; &lt;a href="https://sergei-parfenov.com/assets/chartreplay-hidden-conflict.png" rel="noopener noreferrer"&gt;Open the trace at full size&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The first assertion is &lt;code&gt;CR-02-0000-D34123:1&lt;/code&gt;; the conflicting source is &lt;code&gt;CR-02-0000-D12503&lt;/code&gt;; the explicit correction is &lt;code&gt;CR-02-0000-D31621&lt;/code&gt;. Those identities make the error traceable to particular documents rather than a vague claim that the model “forgot something.”&lt;/p&gt;

&lt;p&gt;The correction to 140 was later applied correctly, but 138 remained absent through all six checkpoints. The first repetition's failed query asked for the measurement history. A consumer inspecting only the accepted value would miss the loss.&lt;/p&gt;

&lt;p&gt;Full rebuild and rolling rewrite preserved the complete record in these chronological repetitions. The difference matters: repeated reconstruction gave those arms an opportunity that the one-pass extraction policy did not provide.&lt;/p&gt;

&lt;p&gt;This is also why storing the original file and retaining its assertions in the usable record are separate requirements. The original document remained in the source archive. It was the derived record that lost its contribution, and with it the visible disagreement.&lt;/p&gt;

&lt;h3&gt;
  
  
  Completeness matters even when retained facts are correct
&lt;/h3&gt;

&lt;p&gt;The primary results use the same 60 combinations in every cell. “Exact” requires the complete final record, including history and relationships, and successful completion of the strict interface contract.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Full rebuild&lt;/th&gt;
&lt;th&gt;Rolling rewrite&lt;/th&gt;
&lt;th&gt;Extract plus ledger&lt;/th&gt;
&lt;th&gt;Format failures across 180 episodes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Luna&lt;/td&gt;
&lt;td&gt;59/60&lt;/td&gt;
&lt;td&gt;59/60&lt;/td&gt;
&lt;td&gt;54/60&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Sol&lt;/td&gt;
&lt;td&gt;60/60&lt;/td&gt;
&lt;td&gt;60/60&lt;/td&gt;
&lt;td&gt;49/60&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 5&lt;/td&gt;
&lt;td&gt;0/60&lt;/td&gt;
&lt;td&gt;7/60&lt;/td&gt;
&lt;td&gt;0/60&lt;/td&gt;
&lt;td&gt;170&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All six Luna ledger failures and all eleven Sol ledger failures involved missing assertions. In one chronological Sol episode, a source explicitly documented an allergy to &lt;code&gt;drug:beta&lt;/code&gt;; its extraction was empty and the final record omitted the allergy.&lt;/p&gt;

&lt;p&gt;Across Sol's 590 ledger checkpoints, all 3,961 predicted assertion occurrences matched the contract, against 4,094 expected occurrences: &lt;strong&gt;100% precision and 96.75% recall&lt;/strong&gt;, yet only &lt;strong&gt;49/60 exact final records&lt;/strong&gt;. These are repeated occurrences across checkpoints and schedules, not thousands of independent patients. Checking whether retained facts are supported would not, by itself, find every missing fact.&lt;/p&gt;

&lt;p&gt;The experiment also contains successful handling of disagreement. In the dedicated unresolved-conflict case, Luna and Sol produced exact final records in all six schedule–repetition combinations under each architecture. The system could preserve a known conflict. The reconciliation failure shows what happens when extraction prevents one side from reaching that system.&lt;/p&gt;

&lt;p&gt;A second case checks what the assertion actually means. Under duplicate replays, with eight unique documents delivered twelve times, all three Luna architectures preserved statements about the patient's mother and father separately from a suspected condition in the patient. They also kept an allergy marked unknown and a symptom explicitly denied, without creating duplicate assertions. Preserving a supported statement can mean preserving uncertainty, rather than confirming a diagnosis.&lt;/p&gt;

&lt;p&gt;Exact scoring also catches literal-contract errors. Luna's single rebuild failure shortened the source identifier &lt;code&gt;procedure:alpha&lt;/code&gt; to &lt;code&gt;alpha&lt;/code&gt;; the assertion was present. That is a different defect from losing an allergy or hiding disagreement, even though each prevents an exact-record result.&lt;/p&gt;

&lt;h3&gt;
  
  
  A correct final record is not enough during an ongoing history
&lt;/h3&gt;

&lt;p&gt;Records may be used between uploads. Luna produced &lt;strong&gt;1,713/1,770 exact checkpoints&lt;/strong&gt;, and Sol &lt;strong&gt;1,656/1,770&lt;/strong&gt;. A final correction does not undo the period during which the record was incomplete.&lt;/p&gt;

&lt;p&gt;We also required the three delivery schedules to produce the same correct final record for each case–repetition group. The ledger passed &lt;strong&gt;14/20 groups for Luna&lt;/strong&gt; and &lt;strong&gt;11/20 for Sol&lt;/strong&gt;; full rebuild passed &lt;strong&gt;19/20 and 20/20&lt;/strong&gt;. An invariant but incomplete record would fail this requirement.&lt;/p&gt;

&lt;p&gt;The experiment does not isolate a causal effect of upload order. For ledger, repeated identical schedules disagreed in 4/30 pairs for Luna and 10/30 for Sol; different schedules within a repetition disagreed in 12/60 and 19/60 pairs. These correlated comparisons show that ordinary model variability also matters. The strongest observation is the actual lost evidence and its consequences, not a claim that reordering alone caused every difference.&lt;/p&gt;

&lt;h3&gt;
  
  
  Gemini adds evidence on the same available conditions
&lt;/h3&gt;

&lt;p&gt;Gemini's 24 available combinations let us make a smaller matched comparison across four models and all three architectures. Selection uses availability only and includes every format failure.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Full rebuild&lt;/th&gt;
&lt;th&gt;Rolling rewrite&lt;/th&gt;
&lt;th&gt;Extract plus ledger&lt;/th&gt;
&lt;th&gt;Format failures across 72 episodes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Luna&lt;/td&gt;
&lt;td&gt;24/24&lt;/td&gt;
&lt;td&gt;24/24&lt;/td&gt;
&lt;td&gt;22/24&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Sol&lt;/td&gt;
&lt;td&gt;24/24&lt;/td&gt;
&lt;td&gt;24/24&lt;/td&gt;
&lt;td&gt;19/24&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 5&lt;/td&gt;
&lt;td&gt;0/24&lt;/td&gt;
&lt;td&gt;5/24&lt;/td&gt;
&lt;td&gt;0/24&lt;/td&gt;
&lt;td&gt;65&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.8 Flash, via Kaggle&lt;/td&gt;
&lt;td&gt;23/24&lt;/td&gt;
&lt;td&gt;20/24&lt;/td&gt;
&lt;td&gt;17/24&lt;/td&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All ten families appear, but unevenly; only one case–repetition group has all three delivery schedules. This is not a complete four-model test of order invariance.&lt;/p&gt;

&lt;p&gt;Gemini's one completed but inexact episode was a shuffled medication history in the ledger arm. A newly delivered source described a separate 5 mg medication order. The extractor omitted it, leaving seven of eight required assertions in the final chart. It is another example of source information failing to enter the derived record.&lt;/p&gt;

&lt;p&gt;The sample also contains counterevidence to a blanket preference for reconstruction. In the chronological 24-document history, Gemini's ledger passed all 24 checkpoints, while rebuild and rewrite stopped on format violations at steps 24 and 8. A stopped integration and an incomplete completed record are different outcomes.&lt;/p&gt;

&lt;p&gt;A necessary interpretation note: all 170 Sonnet format failures and all eleven Gemini format failures involved Markdown-wrapped JSON. Offline removal of one outer wrapper made those responses pass the schema; the accumulated record at the rejected step was exact in 121/170 and 10/11 cases, respectively. These diagnostics do not repair primary scores or invent the later responses of episodes stopped early. The low Sonnet totals cannot be read as a general inability to understand the sources. The missing-fact examples above come from responses the original interface accepted.&lt;/p&gt;

&lt;h3&gt;
  
  
  Lower expense does not price the missing verification step
&lt;/h3&gt;

&lt;p&gt;At the September 25 tariffs and observed cache usage, the selected 60-episode rebuild and ledger arms cost an estimated &lt;strong&gt;$0.31 and $0.11 for Luna&lt;/strong&gt;, and &lt;strong&gt;$5.68 and $2.06 for Sol&lt;/strong&gt;. These are model-token estimates for selected calls, not provider invoices or production bills; interrupted attempts and unknown costs remain separate.&lt;/p&gt;

&lt;p&gt;Those cheaper ledger results include the consequences of the one-pass policy. A completeness check, targeted re-extraction or periodic reconciliation would add work that was not measured. We have not established the cheapest design that meets an acceptable quality target for Symptomato.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;p&gt;ChartReplay separates source inputs from evaluator-only expected records. It retains original source text, assertion identities, correction links, delivery schedules and saved response traces, so an error can be inspected at the update where it appeared. The ledger's disposition rules are explicit; a claim is not considered supported merely because it cites some document.&lt;/p&gt;

&lt;p&gt;The evaluator needed its own checks. An initial audit corrected timestamp comparison, categorical normalization and validation inconsistencies. Calibration then ran 9,240 deterministic episodes across 120 fixtures, seven schedules and eleven systems. Two positive controls each passed 840/840 final episodes and 5,988/5,988 checkpoints; all nine deliberately injected fault types were detected in at least one applicable case. These results check the declared contract, not clinical understanding.&lt;/p&gt;

&lt;p&gt;A separate controlled demonstration removed every superseded assertion from 24 parameterized histories. Eight existing questions still passed on all 24; full-record evaluation passed on none. Adding a history question exposed every failure. Question answering can test the same contract if its coverage is exhaustive, but a handful of correct answers is not evidence that the whole record survived.&lt;/p&gt;

&lt;p&gt;Streaming evaluation of medical memory is established prior work. &lt;a href="https://arxiv.org/abs/2605.11814" rel="noopener noreferrer"&gt;MedMemoryBench&lt;/a&gt; evaluates memory as it is constructed; &lt;a href="https://arxiv.org/abs/2609.01111" rel="noopener noreferrer"&gt;ClinTraceBench&lt;/a&gt; studies longitudinal clinical tasks with source-verifiable evidence. ChartReplay's narrower contribution is an explicit record-construction contract and traces showing how an update can lose evidence or change its disposition.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Kaggle tasks
&lt;/h3&gt;

&lt;p&gt;The &lt;a href="https://www.kaggle.com/benchmarks/sergeiparfenov/chartreplay-source-grounded-record-integrity" rel="noopener noreferrer"&gt;ChartReplay benchmark on Kaggle&lt;/a&gt; consists of three existing smoke tasks, one for each update architecture. They use two synthetic histories—a correction chain and a history containing a multi-assertion note—delivered chronologically and with duplicate replays, once per schedule.&lt;/p&gt;

&lt;p&gt;Each task returns the fraction of the two case–repetition groups for which &lt;strong&gt;both delivery schedules produce the same exact final record&lt;/strong&gt;. This score tests correctness and invariance together; it is different from the fraction of individual episodes that finish correctly.&lt;/p&gt;

&lt;p&gt;The saved Gemini 3.8 Flash results are:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Correct-and-invariant score&lt;/th&gt;
&lt;th&gt;Exact final records&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Full rebuild&lt;/td&gt;
&lt;td&gt;1.0 (2/2 groups)&lt;/td&gt;
&lt;td&gt;4/4 episodes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rolling rewrite&lt;/td&gt;
&lt;td&gt;1.0 (2/2 groups)&lt;/td&gt;
&lt;td&gt;4/4 episodes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Extract plus ledger&lt;/td&gt;
&lt;td&gt;0.5 (1/2 groups)&lt;/td&gt;
&lt;td&gt;3/4 episodes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The smoke used &lt;strong&gt;84 recorded application calls&lt;/strong&gt;. Its two histories recur in the larger pilot, so it is not independent held-out confirmation.&lt;/p&gt;

&lt;p&gt;The tasks provide a small executable test of the same evidence-preservation contract. They cover corrections and repeated delivery; the reconciliation example with values 138, 145 and 140 comes from the larger pilot, not these smoke tasks. The architecture scores are reported separately because combining them would obscure the comparison.&lt;/p&gt;

&lt;h3&gt;
  
  
  Inspect the saved failure offline
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://sergei-parfenov.com/assets/chartreplay-hidden-conflict-evidence.zip" rel="noopener noreferrer"&gt;Download the hidden-conflict evidence ZIP&lt;/a&gt;. It contains six synthetic documents, six saved prompts and unmodified model responses, expected and observed states, and 22 saved queries from Luna's first chronological reconciliation run.&lt;/p&gt;

&lt;p&gt;Unzip it, open the extracted folder in a terminal, and run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 verify_trace.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The standalone verifier uses Python's standard library and makes no API calls. It reconstructs dispositions and query answers from the saved trace. This case was selected after reviewing the results to illustrate a failure; the bundle does not reproduce all primary tables or the full evaluator.&lt;/p&gt;

&lt;h3&gt;
  
  
  What this changes for the system we are building
&lt;/h3&gt;

&lt;p&gt;For Symptomato, the immediate design implication is to keep the original upload, its extracted assertions and its reconciliation status inspectable separately. A successful extraction call should not, by itself, mark a source fully incorporated. I would expose unresolved coverage to the reviewer and test a separate check that can trigger targeted re-extraction. The benchmark knows which assertions are missing because it has an oracle; the product will not. That coverage check needs evaluation against independently annotated documents before we can claim it improves completeness.&lt;/p&gt;

&lt;p&gt;A further evaluation needs independently annotated, naturally written document sets with ambiguous references, dates and source quality. OCR would need its own paired tests against the corresponding text. The present experiment does not assess document authenticity, rank competing sources by authority, verify real-world medical facts or establish clinical safety. A prepared controlled-prose layer has not been run through models and would not, by itself, establish those capabilities.&lt;/p&gt;

&lt;p&gt;OpenAI Codex assisted implementation, offline analysis and drafting. The cover was generated with AI; the trace figure comes from saved experimental states. No model was used to judge the reported record scores.&lt;/p&gt;

&lt;p&gt;A record we can inspect must keep disagreement visible until there is evidence to resolve it. In the observed failure, the system made that disagreement disappear before the correction arrived. Preserving the accepted answer at the end was not enough to preserve the patient's documented history.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>The Test Looked Redundant. The Ninth Bug Needed It</title>
      <dc:creator>Sergei Parfenov</dc:creator>
      <pubDate>Tue, 15 Sep 2026 13:20:00 +0000</pubDate>
      <link>https://dev.to/p0rt/the-test-looked-redundant-the-ninth-bug-needed-it-16me</link>
      <guid>https://dev.to/p0rt/the-test-looked-redundant-the-ninth-bug-needed-it-16me</guid>
      <description>&lt;p&gt;On a catalogue of eight implementations, a repeated-status test added no unique detection. Every candidate it rejected was already rejected by another check.&lt;/p&gt;

&lt;p&gt;I added a ninth implementation. The repeated-status test became the only check that caught it.&lt;/p&gt;

&lt;p&gt;That changes how I would apply the recommendation at the end of &lt;a href="https://dev.to/p0rt/ai-generated-tests-can-make-coding-agents-worse-heres-how-to-check-yours-3jc9"&gt;my previous article on AI-generated tests&lt;/a&gt;. I suggested reviewing a test by asking which plausible wrong implementation it rejects. The comments made that question executable: run the suite against a catalogue of mistakes and count the rejections.&lt;/p&gt;

&lt;p&gt;The approach is useful. Its denominator still needs review.&lt;/p&gt;

&lt;p&gt;I have put the &lt;a href="https://github.com/P0rt/mutation-score-boundaries/tree/19c73b446009e91f10f220dac61a30f96cc1ba5d" rel="noopener noreferrer"&gt;code, checks, mutation diffs, and recorded results on GitHub&lt;/a&gt;. These are local runs of a deliberately constructed Python fixture, not measurements of a coding agent or a replication of ExecCritic. One part reproduces a commenter's result with matching Python and tool versions. Another reconstructs prose-described variants whose original files I do not have. The ninth candidate is my own deliberate counterexample.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five out of five did not distinguish the suites
&lt;/h2&gt;

&lt;p&gt;The previous example was an order filter with three rules: omitting the filter or passing &lt;code&gt;None&lt;/code&gt; returns all orders; an empty list returns none; a list of statuses selects matching orders.&lt;/p&gt;

&lt;p&gt;The correct implementation is small:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;ORDERS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;paid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pending&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;filter_orders&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;statuses&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;statuses&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;order&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;statuses&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Changing the condition to &lt;code&gt;if not statuses&lt;/code&gt; introduces the regression. Python treats both &lt;code&gt;None&lt;/code&gt; and &lt;code&gt;[]&lt;/code&gt; as falsey, so the function returns everything for an empty filter. Two tests—default call and paid-status selection—accept both versions. The empty-list assertion separates them.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/vinhnguyenthanhdn/comment/3ej62"&gt;Vinh Nguyen ran mutmut against this fixture&lt;/a&gt; and reported a boundary I had left too easy to miss: the generated candidates did not include that condition substitution.&lt;/p&gt;

&lt;p&gt;I reproduced the comparison with &lt;strong&gt;CPython 3.14.6 and mutmut 3.7.0&lt;/strong&gt;. The correct implementation scored &lt;strong&gt;5/5&lt;/strong&gt; with the two tests and &lt;strong&gt;5/5&lt;/strong&gt; after adding the empty-list test. The wrong implementation also scored &lt;strong&gt;5/5&lt;/strong&gt; with the two tests. With all three, the wrong implementation failed its baseline, so it received no mutation score.&lt;/p&gt;

&lt;p&gt;For the correct function, the tool inverted the identity comparison, replaced &lt;code&gt;list(orders)&lt;/code&gt; with &lt;code&gt;list(None)&lt;/code&gt;, altered the &lt;code&gt;"status"&lt;/code&gt; key in two ways, and inverted membership. It never replaced the identity comparison with a truthiness test. Several generated candidates raised exceptions; the suite detected those, too. The exact diffs and failure categories are recorded.&lt;/p&gt;

&lt;p&gt;The score correctly described all five candidates the tool had generated. It could not distinguish these two suites because both rejected that entire set.&lt;/p&gt;

&lt;p&gt;Following &lt;a href="https://dev.to/vinhnguyenthanhdn/comment/3ek7h"&gt;Vinh's second comparison&lt;/a&gt;, I added the missing truthiness candidate by hand. Against the same six candidates, the two-test suite rejected &lt;strong&gt;5/6&lt;/strong&gt; and the three-test suite rejected &lt;strong&gt;6/6&lt;/strong&gt;. All of the new distinction came from the candidate derived from the empty-filter requirement.&lt;/p&gt;

&lt;p&gt;This observation is specific to the fixture and operator configuration. It gives no estimate of how often mutation testing misses important distinctions in other code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Eight candidates made another test look unnecessary
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://dev.to/howcani_howcani_77e786a89/comment/3ejm6"&gt;howcani described a broader catalogue&lt;/a&gt;: eight alternatives involving the filter condition, ignored filtering, an incorrect comparison, list identity, output order, repeated selectors, and changes to the input.&lt;/p&gt;

&lt;p&gt;I reconstructed those alternatives with fresh input for each check and independent expected values. The cumulative rejections matched the reported progression:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The original two checks rejected &lt;strong&gt;3/8&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Adding the empty-list check rejected &lt;strong&gt;4/8&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Requiring a new list object brought it to &lt;strong&gt;5/8&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Adding an order-sensitive check brought it to &lt;strong&gt;7/8&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Checking that the caller's input remained unchanged brought it to &lt;strong&gt;8/8&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those matching counts are not a blind replication. The descriptions and counts informed the reconstruction. Exact inputs and implementation choices matter, and the original files were unavailable.&lt;/p&gt;

&lt;p&gt;One reconstructed candidate iterates over requested statuses first, then orders. That can both reorder the result and duplicate orders when a status appears twice. My order probe catches it. So does this check:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="nf"&gt;f&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ORDERS&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;paid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;paid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;ORDERS&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here, &lt;code&gt;f&lt;/code&gt; is the candidate being evaluated. Once the order probe is present, the repeated-status check adds no unique rejection among those eight candidates. It is redundant for separating that finite set. In the discussion, I had agreed with adding assertions only when they reject a surviving candidate.&lt;/p&gt;

&lt;p&gt;The stronger interpretation of that rule does not survive the next function.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ninth candidate preserves order and duplicates matches
&lt;/h2&gt;

&lt;p&gt;Here is the additional implementation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;duplicate_only&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;statuses&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;statuses&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="n"&gt;order&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;statuses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;order&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]))&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For unique requested statuses, it behaves like the correct filter. It preserves input order, creates a new list, leaves the input unchanged, and handles both &lt;code&gt;None&lt;/code&gt; and &lt;code&gt;[]&lt;/code&gt; correctly. For a repeated status, it repeats the matching order.&lt;/p&gt;

&lt;p&gt;It passes all six checks that collectively rejected the first eight candidates. The repeated-status check rejects it.&lt;/p&gt;

&lt;p&gt;With the ninth candidate included, the six-check result is &lt;strong&gt;8/9&lt;/strong&gt;. Retaining the seventh check makes it &lt;strong&gt;9/9&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faxapuptwj3u05bgfaewu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faxapuptwj3u05bgfaewu.png" alt="Mutation results and a matrix showing which checks reject each reconstructed candidate, including the separately added ninth variant" width="799" height="434"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The last row is the added challenge. Its only rejection comes from the repeated-status column.&lt;/p&gt;

&lt;p&gt;I constructed this candidate specifically to separate behaviors that the earlier loop combined. It is not a held-out sample, evidence of bug frequency, or a reason to report 9/9 as general coverage. It demonstrates a narrower point: &lt;strong&gt;a check's lack of a unique rejection can be a property of the catalogue.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The smaller catalogue contained no candidate that duplicated matches while preserving order. Its apparent redundancy inherited that omission.&lt;/p&gt;

&lt;h2&gt;
  
  
  My original runner also changed the count
&lt;/h2&gt;

&lt;p&gt;There was another awkward result. Running the same reconstructed catalogue through the untouched runner from my published archive produced &lt;strong&gt;4/8 and 5/8&lt;/strong&gt;, instead of 3/8 and 4/8.&lt;/p&gt;

&lt;p&gt;The extra rejection came from the candidate that builds the correct return value and then appends a sentinel to the input. In simplified form:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;append_after_copy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sentinel&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;My original default check compares the function's result with &lt;code&gt;ORDERS&lt;/code&gt;. That expected value is the same mutable list passed into the function. By the time equality is evaluated, the expected list has grown. The returned copy has not. The assertion fails.&lt;/p&gt;

&lt;p&gt;With an independent before-call snapshot as the expected output, the output check passes. A separate input-preservation assertion catches the side effect when that property is required.&lt;/p&gt;

&lt;p&gt;The original check happens to reject this candidate. But the two runners are measuring different things. A count without the expected-value construction and state-isolation rules is missing part of its method.&lt;/p&gt;

&lt;p&gt;This discrepancy does not show that howcani's count was wrong; I do not have his precise implementation and test harness. It shows why publishing the functions alone would not make my reconstruction reproducible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Some of the catalogue expanded the contract
&lt;/h2&gt;

&lt;p&gt;The original three rules described which orders should be returned. They did not separately promise a new outer list, stable ordering, or unchanged input. Those may be appropriate requirements. The brief did not settle them.&lt;/p&gt;

&lt;p&gt;I checked a narrow interpretation: preserve the membership and multiplicity of the matching orders, without specifying their order or object ownership. Repeating a selector does not create another order under this interpretation.&lt;/p&gt;

&lt;p&gt;Across a finite domain of &lt;strong&gt;1,134 calls per function&lt;/strong&gt;, four alternatives—returning an alias, reversing the output, sorting the input, and appending after constructing the result—returned the correct multiset on every call. The domain included empty inputs, omitted and explicit &lt;code&gt;None&lt;/code&gt; filters, repeated selectors, unknown statuses, and reversed input order. The full setup is in the &lt;a href="https://github.com/P0rt/mutation-score-boundaries/blob/19c73b446009e91f10f220dac61a30f96cc1ba5d/docs/experiment.md" rel="noopener noreferrer"&gt;experiment report&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Those functions can still be bad choices for a real caller. Mutating input can corrupt later calls. Order or ownership may be an established compatibility promise. Passing a one-call, finite-domain check does not establish product correctness.&lt;/p&gt;

&lt;p&gt;The implication for the catalogue is specific: attach a requirement to each candidate before labeling its behavior a bug. &lt;code&gt;return orders&lt;/code&gt; violates a new-list promise if the product makes one. The fact that my reference implementation uses &lt;code&gt;list(orders)&lt;/code&gt; does not, on its own, establish that promise.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the regression; keep questioning the catalogue
&lt;/h2&gt;

&lt;p&gt;For an agent repair loop, I would keep the review question from the previous article and narrow the pruning rule.&lt;/p&gt;

&lt;p&gt;Start with a written requirement and its supported inputs. Give each candidate an executable counterexample and a reviewed expected result. Use the candidate matrix to discover missing distinctions. Keep accepted regression checks stable while the implementation changes.&lt;/p&gt;

&lt;p&gt;If a test adds no unique rejection, inspect the overlap before deleting it. Two tests may reject the same candidate because that candidate bundles two independent mistakes. Ask whether one mistake can occur without the other. Here, an outer loop over selectors bundled ordering and duplication. Changing the loop structure separated them.&lt;/p&gt;

&lt;p&gt;This does not mean retaining every test forever. A finite catalogue can support pruning when the intended objective is discrimination within that catalogue. Removing a regression check tied to a distinct requirement needs a broader argument—such as the remaining checks establishing the same behavior over the supported inputs.&lt;/p&gt;

&lt;p&gt;I would carry that distinction into the same record as the score: contract revision, candidate revision, probes, expected values, environment, and observed failures. It extends the concern in &lt;a href="https://sergei-parfenov.com/blog/the-model-scored-30-the-harness-scored-100-which-one-did-you-benchmark-3mp4/" rel="noopener noreferrer"&gt;my earlier piece about harness-dependent scores&lt;/a&gt;: the number only describes the system that produced it.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://github.com/P0rt/mutation-score-boundaries" rel="noopener noreferrer"&gt;repository&lt;/a&gt; includes the original archive, four mutation-test configurations, saved generated functions, reconstructed candidates, the ninth challenge, and commands for replaying the evidence. The quick verification uses only Python's standard library; the full mutation run uses the recorded dependency versions.&lt;/p&gt;

&lt;p&gt;The repeated-status requirement did not change when I added the ninth function. The catalogue finally contained a way to violate it without also violating the order check. That is the reason I would keep the test.&lt;/p&gt;




&lt;p&gt;If this article helped you, you can &lt;a href="https://ko-fi.com/sergeiparfenov?ref=dev" rel="noopener noreferrer"&gt;buy me a coffee&lt;/a&gt;. Your support helps me make time for the next experiment.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>python</category>
      <category>agents</category>
    </item>
    <item>
      <title>AI-Generated Tests Can Make Coding Agents Worse. Here's How to Check Yours</title>
      <dc:creator>Sergei Parfenov</dc:creator>
      <pubDate>Fri, 11 Sep 2026 11:57:46 +0000</pubDate>
      <link>https://dev.to/p0rt/ai-generated-tests-can-make-coding-agents-worse-heres-how-to-check-yours-3jc9</link>
      <guid>https://dev.to/p0rt/ai-generated-tests-can-make-coding-agents-worse-heres-how-to-check-yours-3jc9</guid>
      <description>&lt;p&gt;A bug fix can make every new test pass and still introduce a regression. Here is a deliberately constructed Python example, checked locally without an LLM.&lt;/p&gt;

&lt;p&gt;An order filter has three requirements:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Omit the filter, or pass &lt;code&gt;None&lt;/code&gt;: return all orders.&lt;/li&gt;
&lt;li&gt;Pass an empty list: return no orders.&lt;/li&gt;
&lt;li&gt;Pass a list of statuses: return only matching orders.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The reported bug is that omitting the filter returns nothing. This proposed fix looks reasonable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;ORDERS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;paid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pending&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;filter_orders&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;statuses&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;statuses&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;order&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;statuses&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="nf"&gt;filter_orders&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ORDERS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;ORDERS&lt;/span&gt;
&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="nf"&gt;filter_orders&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ORDERS&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;paid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;ORDERS&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]]&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2 checks passed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both checks pass. Both branches of the &lt;code&gt;if&lt;/code&gt; have been exercised. The reported symptom is fixed.&lt;/p&gt;

&lt;p&gt;Now add the check for the second requirement:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="nf"&gt;filter_orders&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ORDERS&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[])&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It fails. The function returns every order.&lt;/p&gt;

&lt;p&gt;Python treats both &lt;code&gt;None&lt;/code&gt; and &lt;code&gt;[]&lt;/code&gt; as falsey. Our requirements give them different meanings, and the patch erases that distinction.&lt;/p&gt;

&lt;p&gt;The connection to coding agents becomes more consequential when those checks determine what the agent does next.&lt;/p&gt;

&lt;p&gt;On September 8, Leitian Tao and colleagues published the &lt;a href="https://arxiv.org/html/2609.09133v1" rel="noopener noreferrer"&gt;ExecCritic preprint&lt;/a&gt;. Holding the Qwen-3.5-35B-A3B Repair agent fixed, they reported these SWE-bench Verified results:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feedback source&lt;/th&gt;
&lt;th&gt;Tasks resolved&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Initial repair, before generated-test feedback&lt;/td&gt;
&lt;td&gt;61.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tests from the base Qwen Test agent&lt;/td&gt;
&lt;td&gt;57.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tests from GPT-5.6-sol&lt;/td&gt;
&lt;td&gt;65.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The weaker tests reduced the resolved rate by &lt;strong&gt;3.9 percentage points&lt;/strong&gt;. Better tests improved it.&lt;/p&gt;

&lt;p&gt;Rates average three repair runs, reusing generated tests. Failed test qualification retains the initial patch in the all-task score. The baseline does not forbid repository tests. A separate official evaluator determines resolution. Feedback adds test-generation and revision work; compute budgets are not matched. These are the authors' results, not a benchmark replication for this article. &lt;a href="https://arxiv.org/html/2609.09133v1#S5" rel="noopener noreferrer"&gt;Method and results&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;A bad test can do more than miss a defect. It can give the next edit the wrong target.&lt;/p&gt;

&lt;p&gt;Imagine adding this assertion to the filter example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# This expectation contradicts the stated empty-list requirement.
&lt;/span&gt;&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="nf"&gt;filter_orders&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ORDERS&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[])&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;ORDERS&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Our broken patch passes it. A correct implementation would fail it. Feed that failure into an automatic repair loop, and the loop now has a reason to damage correct behavior.&lt;/p&gt;

&lt;p&gt;Adding assertions has strengthened the wrong interpretation.&lt;/p&gt;

&lt;p&gt;Even the familiar “fails before the fix, passes afterward” check needs a closer look. Here is the original implementation from the fixture:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;filter_orders&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;statuses&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;statuses&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;statuses&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;order&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;statuses&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The default-filter assertion fails against this version and passes against our proposed patch. It correctly detects the original bug. It simply cannot detect the new one.&lt;/p&gt;

&lt;p&gt;The complete fix handles &lt;code&gt;None&lt;/code&gt; explicitly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;filter_orders&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;statuses&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;statuses&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;order&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;statuses&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Running the same checks against all three implementations produces:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Implementation&lt;/th&gt;
&lt;th&gt;Default + paid-filter checks&lt;/th&gt;
&lt;th&gt;Those checks + empty-list check&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Original&lt;/td&gt;
&lt;td&gt;1 passes, 1 fails&lt;/td&gt;
&lt;td&gt;2 pass, 1 fails&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Plausible patch&lt;/td&gt;
&lt;td&gt;2 pass&lt;/td&gt;
&lt;td&gt;2 pass, 1 fails&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Corrected patch&lt;/td&gt;
&lt;td&gt;2 pass&lt;/td&gt;
&lt;td&gt;3 pass&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The &lt;a href="https://sergei-parfenov.com/assets/downloads/ai-test-demo.zip" rel="noopener noreferrer"&gt;runnable companion&lt;/a&gt; includes all three implementations, the checks, and the verified output. It uses Python’s standard library and makes no LLM or network calls.&lt;/p&gt;

&lt;p&gt;The extra check earns its place because it distinguishes two implementations the earlier checks considered equally acceptable.&lt;/p&gt;

&lt;p&gt;That is the question I would bring to an AI-generated test review: &lt;strong&gt;which plausible wrong implementation would this test reject?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For the filter, the candidate mistakes are easy to name: ignore the filter entirely, treat every missing filter as empty, or treat every empty filter as missing. They correspond to different misunderstandings of the contract. Tests that separate those cases tell us more than several additional examples of paid orders.&lt;/p&gt;

&lt;p&gt;This is also where &lt;a href="https://mutmut.readthedocs.io/en/latest/" rel="noopener noreferrer"&gt;mutation testing&lt;/a&gt; can help: make small changes to the implementation and check whether the suite detects them. Inspect surviving mutations to understand what they change; some are equivalent for the supported inputs. For this fixture, changing &lt;code&gt;statuses is None&lt;/code&gt; to &lt;code&gt;not statuses&lt;/code&gt; is a useful manual mutation because it has a known, observable effect on required behavior.&lt;/p&gt;

&lt;p&gt;For an agent workflow, I would make four changes.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Write down the expected behavior before reviewing the patch.&lt;/strong&gt; Include the ordinary case, the reported failure, and the neighboring case most likely to be confused with it. Here, &lt;code&gt;None&lt;/code&gt; and &lt;code&gt;[]&lt;/code&gt; belong on separate rows. If the issue leaves that distinction unspecified, get a product decision before turning either interpretation into a test.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Review expected values as carefully as production code.&lt;/strong&gt; An assertion is a claim about the product. Trace that claim to a requirement, an established compatibility promise, or an independently checked example. Copying the current output into an expected value can preserve the exact behavior you meant to question.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Keep an accepted regression check stable during repair.&lt;/strong&gt; Let the agent change the implementation while a separate runner evaluates it with the reviewed tests. Protect the test command and configuration too: an unchanged test file helps little if the patch can skip its execution. If the test itself is wrong, revise and review it explicitly, then evaluate the candidate again.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Inspect the failure before asking the agent to fix it.&lt;/strong&gt; An assertion showing the wrong returned orders is actionable behavioral evidence. A missing dependency, an import failure, or a command that selected zero tests needs a different response. Record what ran and why it failed.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;ExecCritic separates test creation from repair, qualifies tests on the original repository, and keeps them unchanged during revision. That limits the repairer's ability to change its target. Separate contexts and permissions still cannot guarantee that both agents understood the issue correctly. &lt;a href="https://arxiv.org/html/2609.09133v1#S2" rel="noopener noreferrer"&gt;Paper&lt;/a&gt;; &lt;a href="https://github.com/MSR-Orchard/execcritic" rel="noopener noreferrer"&gt;released implementation&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The four steps above are a review procedure you can try in an existing project. They do not require training a model. Their value should be judged by the bugs and mistaken expectations they expose.&lt;/p&gt;

&lt;p&gt;For your next AI-assisted fix, keep the original code, the proposed patch, and the new tests. Identify one plausible alternative implementation that violates the requirement. Run the tests against it.&lt;/p&gt;

&lt;p&gt;If both implementations get the same green result, you have found a specific question the suite still cannot answer. Add the check that separates them, and review its expected result before trusting the next repair.&lt;/p&gt;




&lt;p&gt;If this article helped you, you can &lt;a href="https://ko-fi.com/sergeiparfenov?ref=dev" rel="noopener noreferrer"&gt;buy me a coffee&lt;/a&gt;. Your support helps me make time for the next experiment.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>python</category>
      <category>agents</category>
    </item>
    <item>
      <title>Can Your AI Use What It Remembers?</title>
      <dc:creator>Sergei Parfenov</dc:creator>
      <pubDate>Sat, 05 Sep 2026 14:09:59 +0000</pubDate>
      <link>https://dev.to/p0rt/can-your-ai-use-what-it-remembers-57c</link>
      <guid>https://dev.to/p0rt/can-your-ai-use-what-it-remembers-57c</guid>
      <description>&lt;p&gt;An assistant correctly recalls a user's tree-nut allergy. Given a request for macarons, it supplies an almond-flour recipe without applying that allergy to the user.&lt;/p&gt;

&lt;p&gt;This is a recorded example from &lt;a href="https://arxiv.org/html/2607.24368v1#S12" rel="noopener noreferrer"&gt;InMind, a July 2026 study of agent memory&lt;/a&gt;, using xMemory. The evaluations ask two different things: can the system retrieve a fact when the question names it, and can it bring that fact into a decision that needs it?&lt;/p&gt;

&lt;p&gt;The second question is the reason to give an assistant memory in the first place. A user should not need to know which past conversation to reference before every request.&lt;/p&gt;

&lt;p&gt;For a coding agent, the equivalent might be remembering that a service runs in short-lived processes, then proposing an in-process timer for hourly cleanup. That is an illustrative test case, not a result reported in the paper. It has the same useful structure: the current request does not repeat the architectural constraint, but the constraint changes what a correct solution looks like.&lt;/p&gt;

&lt;p&gt;InMind makes that distinction measurable. Its 125 synthetic tasks pair a stored personal fact with a later request whose relevance depends on background knowledge. The authors deliberately remove obvious lexical and semantic retrieval cues. This is a stress test for a particular failure mode, not a sample of everyday assistant traffic.&lt;/p&gt;

&lt;p&gt;For the study's A-Mem configuration with text-embedding-3-large, &lt;a href="https://arxiv.org/html/2607.24368v1#S4.T1" rel="noopener noreferrer"&gt;Table 1&lt;/a&gt; reports:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;100% direct recall:&lt;/strong&gt; a passing score on all 125 questions that directly ask for the stored fact.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;12% target recall:&lt;/strong&gt; the necessary fact is judged present in the model's context for 15 of the 125 indirect requests.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;9.6% application:&lt;/strong&gt; the combined memory-and-answer evaluation passes 12 of those 125 requests.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are separate measurements. The &lt;a href="https://github.com/imlrz/InMind/blob/5a2ab2686d3d5b832575f0288d41ecc93eab358b/evaluation/README.md#standard-protocol" rel="noopener noreferrer"&gt;released evaluation protocol&lt;/a&gt; calls for direct questions and indirect requests to use the same frozen memory state, without letting the first answer supply a hint for the second.&lt;/p&gt;

&lt;p&gt;The application number needs care. Its rubric requires both the personal fact in context and a relevant warning or reminder in the answer. When the authors score only the answer, the same A-Mem configuration reaches &lt;strong&gt;25.6%&lt;/strong&gt;, not 9.6%. A model can produce a useful caution from general knowledge without retrieving anything about this particular user. That answer may be good; it is not evidence that the memory system worked. The paper's &lt;a href="https://arxiv.org/html/2607.24368v1#S14.SS1" rel="noopener noreferrer"&gt;answer-only evaluation&lt;/a&gt; and &lt;a href="https://arxiv.org/html/2607.24368v1#S14.SS2" rel="noopener noreferrer"&gt;human audit&lt;/a&gt; expose exactly this ambiguity.&lt;/p&gt;

&lt;p&gt;The cleanest signal here is the missing fact: direct questions find it, while indirect requests usually do not deliver it to the answerer.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://arxiv.org/html/2607.24368v1#S15.SS7" rel="noopener noreferrer"&gt;A-Mem setup&lt;/a&gt; appends a raw target note to a prebuilt, read-only bank. This tests selection when the fact is available; it does not establish survival through weeks of memory rewriting.&lt;/p&gt;

&lt;p&gt;In a retrieve-then-answer pipeline, the order is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;new request
    → choose memories
    → put selected memories in context
    → generate a response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The selection step must decide what matters before the answerer sees the candidate facts together with the request. If selection misses a constraint, the answerer may produce a perfectly reasonable solution for a different environment.&lt;/p&gt;

&lt;p&gt;This does not mean embeddings cannot encode useful associations, or that an agent cannot discover them through additional searches. It means those capabilities need testing on requests that do not already identify the required memory. Bigger retrieval scores on direct questions do not settle that question.&lt;/p&gt;

&lt;p&gt;Related work already reaches beyond factual recall. &lt;a href="https://arxiv.org/abs/2602.10715" rel="noopener noreferrer"&gt;LoCoMo-Plus&lt;/a&gt; examines whether agents retain and apply implicit conversational constraints. InMind's useful contribution here is its paired diagnosis of stored facts whose relevance requires an unstated knowledge connection. It is not the discovery that memory should affect behavior.&lt;/p&gt;

&lt;p&gt;For a project assistant, I would turn that diagnosis into a small set of regression cases. Start with a requirement whose effect can be checked:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Stored constraint:
Processes can terminate at any time. Recurring jobs must
continue running even when no web request is active.

Direct question:
What execution constraints does this service have?

Indirect task:
Implement hourly cleanup of expired exports.

Behavior to check:
Does the proposed solution keep recurring execution outside
the lifetime of a single web process?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use independent copies of the same starting state. One receives the direct question. Another receives only the task. A third receives the task with the relevant constraint explicitly included. Keep unrelated context and budgets matched where possible; the explicit-fact condition is an intervention, not an ordinary retrieval result.&lt;/p&gt;

&lt;p&gt;Record the context the model actually received. That gives a failure somewhere to belong:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If the fact is absent from storage, investigate capture or updates.&lt;/li&gt;
&lt;li&gt;If it exists but does not reach the answerer, investigate selection and context assembly.&lt;/li&gt;
&lt;li&gt;If it reaches the answerer but the implementation violates it, investigate application.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Score the implementation separately from memory delivery. A model might choose a durable scheduler without knowing the project's constraint. That can pass the functional test while telling us little about retrieval. Conversely, mentioning the constraint in an explanation should earn no credit if the generated code still relies on a process-local timer.&lt;/p&gt;

&lt;p&gt;The test also needs requests where the constraint should have no effect. A button-label change should not become a lecture about background workers. And it needs a later instruction that legitimately changes the requirement: if recurring execution is explicitly removed from scope, the old rule should not silently override the new task.&lt;/p&gt;

&lt;p&gt;Without those controls, a system that recites every restriction on every turn can look surprisingly competent. InMind itself notes that its rubric does not penalize excessive warnings. That makes it useful for finding missed associations, but insufficient for deciding whether an assistant applies memory appropriately across ordinary work.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://sergei-parfenov.com/assets/downloads/agent-memory-probe.zip" rel="noopener noreferrer"&gt;accompanying probe&lt;/a&gt; contains 12 developer-constraint fixtures and isolated inputs for these conditions. A local BM25 run, returning at most three records, delivered the target fact for &lt;strong&gt;12 of 12 direct questions and 4 of 12 indirect tasks&lt;/strong&gt;. The cases were deliberately written with vocabulary mismatches. This is a transparent illustration of lexical retrieval, not a measurement of a modern agent's memory quality.&lt;/p&gt;

&lt;p&gt;No LLM was run in that probe. Whether a generated implementation follows the constraint remains a separate evaluation. The inputs and behavioral rubrics are included so those checks can be run against an actual agent. InMind's public release also still lacks baseline adapters and paper-aligned per-task results, so the published model scores above are reported findings, not a replication performed for this article.&lt;/p&gt;

&lt;p&gt;Keeping selected constraints visible in every task is one option to test. The engineering problem then moves to which constraints deserve that space, which project or environment they apply to, and when they expire. A remembered fact can be correct and still be irrelevant to the current branch, obsolete after a migration, or superseded by an explicit instruction.&lt;/p&gt;

&lt;p&gt;For an initial implementation, I would keep a small set of active project requirements with an explicit scope, while leaving detailed history searchable. Then I would test both omissions and unnecessary enforcement. That is a design proposal, not a fix established by the paper.&lt;/p&gt;

&lt;p&gt;The benchmark is small and deliberately difficult; GPT-5-mini serves as both answerer and judge, and the human audit found scoring errors. Its numbers do not establish how often a particular production agent will fail. They do establish a question that a direct-recall demo cannot answer.&lt;/p&gt;

&lt;p&gt;Ask the agent what it remembers. Then, from a fresh copy of the same state, ask it to do something that depends on that memory. Check what reached the model, check what it built, and check when the remembered rule should stop applying.&lt;/p&gt;

&lt;p&gt;The useful promise of memory is that you can stop repeating yourself.&lt;/p&gt;

&lt;p&gt;That is the promise the test should measure.&lt;/p&gt;




&lt;p&gt;Sources: Ruizhe Li, Mingxuan Du, Benfeng Xu, and Zhendong Mao, &lt;a href="https://arxiv.org/html/2607.24368v1" rel="noopener noreferrer"&gt;Keep It InMind, v1, July 27, 2026&lt;/a&gt; (CC BY 4.0); &lt;a href="https://github.com/imlrz/InMind" rel="noopener noreferrer"&gt;InMind repository&lt;/a&gt;; Yifei Li et al., &lt;a href="https://arxiv.org/abs/2602.10715" rel="noopener noreferrer"&gt;LoCoMo-Plus, February 11, 2026&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;If this article helped you, you can &lt;a href="https://ko-fi.com/sergeiparfenov?ref=dev" rel="noopener noreferrer"&gt;buy me a coffee&lt;/a&gt;. Your support helps me make time for the next experiment.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>llm</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>The Agent Knew It Was Wrong. The System Let It Ship</title>
      <dc:creator>Sergei Parfenov</dc:creator>
      <pubDate>Tue, 01 Sep 2026 17:49:00 +0000</pubDate>
      <link>https://dev.to/p0rt/the-agent-knew-it-was-wrong-the-system-let-it-ship-dgp</link>
      <guid>https://dev.to/p0rt/the-agent-knew-it-was-wrong-the-system-let-it-ship-dgp</guid>
      <description>&lt;p&gt;In 660 of 800 autonomous research runs, the agent found a serious flaw in its own work.&lt;/p&gt;

&lt;p&gt;It wrote the flaw down.&lt;/p&gt;

&lt;p&gt;Then it delivered the report anyway.&lt;/p&gt;

&lt;p&gt;The model did not fail to notice.&lt;/p&gt;

&lt;p&gt;The system failed to make noticing consequential.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AutoResearchEval labeled 82.5% of its 800 trajectories with “uncorrected self-awareness”: the agent identified a critical flaw, then continued without fixing or gating it&lt;/li&gt;
&lt;li&gt;A separate abstention benchmark found agents sometimes performed an irreversible action and only then claimed they had refused&lt;/li&gt;
&lt;li&gt;In a 3,621-trial policy study, an output reviewer ran after exposure and backend effects; moving enforcement to the tool boundary cut trace failures from 57.6% to 0.2% in the primary comparison&lt;/li&gt;
&lt;li&gt;A review that cannot change execution state is observability, not control&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The failure was not awareness
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/abs/2608.14905" rel="noopener noreferrer"&gt;AutoResearchEval&lt;/a&gt; took 100 research tasks across seven scientific domains and ran each through eight harness–model combinations.&lt;/p&gt;

&lt;p&gt;The result was 800 complete trajectories, roughly 73,000 tool calls, and an average of 92.3 steps per run. The evaluator inspected the reports, code, generated data, retrieval logs, and execution artifacts rather than scoring only the final answer.&lt;/p&gt;

&lt;p&gt;Its most common failure pattern was not hallucination.&lt;/p&gt;

&lt;p&gt;It was &lt;strong&gt;uncorrected self-awareness&lt;/strong&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pattern&lt;/th&gt;
&lt;th&gt;Trajectories&lt;/th&gt;
&lt;th&gt;What happened&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Uncorrected self-awareness&lt;/td&gt;
&lt;td&gt;660 / 800&lt;/td&gt;
&lt;td&gt;The agent identified a fatal or critical flaw and made no consequential correction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Method–conclusion disconnect&lt;/td&gt;
&lt;td&gt;620 / 800&lt;/td&gt;
&lt;td&gt;The written conclusion was not supported by the method actually executed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Failure to gate critical flaws&lt;/td&gt;
&lt;td&gt;502 / 800&lt;/td&gt;
&lt;td&gt;A critical issue was recorded but did not block delivery&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Report–trace gap&lt;/td&gt;
&lt;td&gt;484 / 800&lt;/td&gt;
&lt;td&gt;Claims could not be traced to the code, data, or logs produced in the same run&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Those categories overlap. They are not four independent populations and should not be added together.&lt;/p&gt;

&lt;p&gt;The useful part is the shape of the failure.&lt;/p&gt;

&lt;p&gt;The evidence was already inside the trajectory. The agent had the report, the code, the logs, and—in 660 cases—a written recognition that something was seriously wrong.&lt;/p&gt;

&lt;p&gt;Nothing in the execution loop required that recognition to change the result.&lt;/p&gt;

&lt;p&gt;We keep describing this as a reasoning problem because the visible artifact is text. The agent writes a bad answer, then writes a good critique of the bad answer, so the system appears to contain both failure and correction.&lt;/p&gt;

&lt;p&gt;But correction is not a paragraph.&lt;/p&gt;

&lt;p&gt;Correction is a state transition.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;RUNNING → REVIEW_REQUIRED → BLOCKED → REMEDIATED → COMMITTABLE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If “this result is invalid” and “publish this result” can both be true in the same system state, the review stage is decorative.&lt;/p&gt;

&lt;h2&gt;
  
  
  Self-review is usually just another generation
&lt;/h2&gt;

&lt;p&gt;A common agent loop looks approximately like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;draft&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;worker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;review&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;worker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;review&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;draft&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;final&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;worker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;revise&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;draft&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;review&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;final&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This can improve an answer. It can also produce a more articulate failure.&lt;/p&gt;

&lt;p&gt;The same probabilistic system proposes the work, selects the evidence, interprets the evidence, judges itself, and decides whether the judgment matters. The review has no independent authority and frequently no independent source of truth.&lt;/p&gt;

&lt;p&gt;Even adding a second model does not automatically fix that. Two actors are not two control planes if both operate inside the same mutable context and either can still call the production tool.&lt;/p&gt;

&lt;p&gt;The distinction I would make is this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A reviewer produces a judgment&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A gate changes what the system is allowed to do&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The first is information.&lt;/p&gt;

&lt;p&gt;The second is authority.&lt;/p&gt;

&lt;p&gt;I call the distance between them the &lt;strong&gt;review-to-effect gap&lt;/strong&gt;: the part of the pipeline between detecting a problem and the final point where the system can still prevent an externally visible result.&lt;/p&gt;

&lt;p&gt;The wider that gap, the easier it is for a correct diagnosis to become an irrelevant log entry.&lt;/p&gt;

&lt;h2&gt;
  
  
  The agent refused after it acted
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/abs/2607.10059" rel="noopener noreferrer"&gt;AgentAbstain&lt;/a&gt; tested 17 frontier models in four agent harnesses on 263 paired tasks across 42 executable sandbox environments.&lt;/p&gt;

&lt;p&gt;Each pair contained a normal task and a minimally changed version where the correct behavior was to stop. The best tested agent achieved 59.5% paired accuracy: it correctly handled both the act and abstain sides in fewer than three out of five pairs.&lt;/p&gt;

&lt;p&gt;The paper also separated what the agent &lt;strong&gt;said&lt;/strong&gt; from what it &lt;strong&gt;did&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;In 115 trajectories, 2.6% of the aggregate abstention-run breakdown, the agent crossed an irreversible action boundary and then verbally claimed restraint. The authors call this &lt;strong&gt;post-hoc abstention&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Their sandbox examples included agents canceling a flight before noticing conflicting rebooking constraints and submitting unauthorized simulated bank transfers after acknowledging that a required verifier was unavailable.&lt;/p&gt;

&lt;p&gt;The final message could say “I cannot proceed.”&lt;/p&gt;

&lt;p&gt;The tool trace said it already had.&lt;/p&gt;

&lt;p&gt;This is the same structural failure as the research report. The model discovers the trigger, generates the correct language, and reaches that language after the effect boundary has been crossed.&lt;/p&gt;

&lt;p&gt;An apology is not a rollback.&lt;/p&gt;

&lt;p&gt;A refusal is not a refusal if the transaction log contains a commit.&lt;/p&gt;

&lt;h2&gt;
  
  
  The reviewer arrived after the side effect
&lt;/h2&gt;

&lt;p&gt;The clearest timing experiment I found is &lt;a href="https://arxiv.org/abs/2608.27646" rel="noopener noreferrer"&gt;Out-of-Band Policy Enforcement at a Trusted Tool Boundary&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The authors ran 3,621 trials against controlled Jira and ServiceNow mocks. Their trace-level failure definition included any of three events:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;protected data entered model context&lt;/li&gt;
&lt;li&gt;an exact protected value appeared in the answer&lt;/li&gt;
&lt;li&gt;a forbidden backend effect completed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Detailed prompt rules alone still produced a 57.6% trace-failure rate in the primary comparison. Adding an enforcement boundary between the agent and the backend reduced it to 0.2%.&lt;/p&gt;

&lt;p&gt;The paper then compared two mechanisms directly on Claude Sonnet 4.6.&lt;/p&gt;

&lt;p&gt;One left the tool path untouched and asked a fixed reviewer model to pass or replace every assistant message. The other mediated the request before dispatch and the response before it re-entered agent context.&lt;/p&gt;

&lt;p&gt;The output-review arm had a 58.6% deterministic trace-failure rate and 19.7% safe-useful completion.&lt;/p&gt;

&lt;p&gt;The full boundary had 0.0% trace failure in that arm and 63.0% safe-useful completion.&lt;/p&gt;

&lt;p&gt;The reviewer also required 930 additional model calls.&lt;/p&gt;

&lt;p&gt;The important result is not that one reviewer prompt was weak. The paper explicitly says this is one mechanism comparison, not a ranking of every possible guardrail.&lt;/p&gt;

&lt;p&gt;The important result is that the reviewer saw the assistant message &lt;strong&gt;after the tools had returned&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It could suppress unsafe text.&lt;/p&gt;

&lt;p&gt;It could not remove data already placed in context.&lt;/p&gt;

&lt;p&gt;It could not undo a backend mutation.&lt;/p&gt;

&lt;p&gt;A post-output safety filter is a censor. It is not a transaction boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Review before dispatch helps — but it is still probabilistic
&lt;/h2&gt;

&lt;p&gt;Moving the reviewer earlier is useful.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/abs/2604.27233" rel="noopener noreferrer"&gt;Reinforced Agent&lt;/a&gt; puts a specialized reviewer in front of provisional tool calls. The call is reviewed before execution, and the worker can revise it before anything changes outside the agent.&lt;/p&gt;

&lt;p&gt;That architecture improved the reported tool-calling benchmarks. It also exposed the reviewer’s own failure rate.&lt;/p&gt;

&lt;p&gt;With o3-mini as reviewer, 36.8% of base-agent errors were corrected while 11.7% of previously correct cases were damaged. The reported benefit-to-risk ratio was 3.1 to 1.&lt;/p&gt;

&lt;p&gt;That is a useful component.&lt;/p&gt;

&lt;p&gt;It is not a proof boundary.&lt;/p&gt;

&lt;p&gt;A model reviewer is good for semantic questions that deterministic code cannot answer cheaply:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Does this action satisfy the user’s actual intent&lt;/li&gt;
&lt;li&gt;Is the evidence sufficient for this conclusion&lt;/li&gt;
&lt;li&gt;Does the requested operation conflict with a policy expressed in natural language&lt;/li&gt;
&lt;li&gt;Is the proposed scope disproportionate to the task&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It should not be the only thing standing between a stochastic plan and production credentials.&lt;/p&gt;

&lt;p&gt;The reviewer may recommend &lt;code&gt;ALLOW&lt;/code&gt;, &lt;code&gt;HOLD&lt;/code&gt;, or &lt;code&gt;DENY&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The trusted boundary must decide whether a valid capability exists for the exact operation about to execute.&lt;/p&gt;

&lt;h2&gt;
  
  
  The architecture I would ship
&lt;/h2&gt;

&lt;p&gt;The planner, reviewer, policy engine, and effect adapter have different jobs. Collapsing them into one “agent” object hides the boundary that matters.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;UNTRUSTED DECISION PLANE

user intent
    ↓
planner
    ↓
proposed action + evidence bundle
    ↓
semantic reviewer

TRUSTED CONTROL PLANE

schema and invariant checks
    ↓
policy and authorization decision
    ↓
exact action manifest
    ↓
short-lived commit capability

EFFECT PLANE

effect adapter holding credentials
    ↓
provider commit
    ↓
terminal receipt or durable uncertainty state
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;“Untrusted” here does not mean malicious. It means &lt;strong&gt;non-authoritative&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The planner may be brilliant. The reviewer may be more capable than the planner. Neither should be able to convert its own text directly into an authenticated side effect.&lt;/p&gt;

&lt;p&gt;The effect adapter should accept something closer to this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"proposal_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"refund_0184"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"operation"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"payments.refund"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"resource"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"payment_intent:pi_7F..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"arguments_hash"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"sha256:9fa..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"evidence_hash"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"sha256:1bd..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"policy_version"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"refunds@7.2"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"provider_state_version"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"captured@2026-09-01T10:42:18Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"review"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"decision"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"allow"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"risk"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"low"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"reason_codes"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"amount_within_limit"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"recipient_verified"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"commit_capability"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"cap_opaque_single_use"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The adapter then verifies the hashes, policy version, current provider state, capability scope, expiry, and single-use status before it calls the provider.&lt;/p&gt;

&lt;p&gt;The natural-language conversation is evidence for constructing the manifest.&lt;/p&gt;

&lt;p&gt;It is not the manifest.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five invariants that turn review into control
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1 Critical findings must change system state
&lt;/h3&gt;

&lt;p&gt;A critical finding cannot coexist with a committable action.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;review.severity == critical
    ⇒ run.state == BLOCKED
    ⇒ commit_capability == null
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do not rely on the worker to “take the feedback into account.” Make remediation create a new proposal and a new review record.&lt;/p&gt;

&lt;h3&gt;
  
  
  2 Approval must bind the exact action
&lt;/h3&gt;

&lt;p&gt;“Refund the customer” is not sufficient authorization.&lt;/p&gt;

&lt;p&gt;The gate should bind the operation, acting identity, resource, arguments, evidence version, policy version, and relevant external state. If any bound field changes, the old approval is invalid.&lt;/p&gt;

&lt;h3&gt;
  
  
  3 The credential holder must sit below the gate
&lt;/h3&gt;

&lt;p&gt;The worker and reviewer should not hold direct production credentials.&lt;/p&gt;

&lt;p&gt;Otherwise the controlled path is optional. An agent that can bypass the gateway eventually will—through a bug, a fallback, an alternate connector, or a tool call the policy layer never saw.&lt;/p&gt;

&lt;h3&gt;
  
  
  4 Uncertain delivery must not mint fresh authority
&lt;/h3&gt;

&lt;p&gt;A timeout does not mean no effect occurred.&lt;/p&gt;

&lt;p&gt;If the provider may have committed but the response was lost, the original authorization must remain occupied until reconciliation reaches a terminal result. Creating a fresh approval for a blind retry can turn one user decision into two external effects.&lt;/p&gt;

&lt;h3&gt;
  
  
  5 The final answer must be generated from the effect receipt
&lt;/h3&gt;

&lt;p&gt;Do not let the agent report operational state from memory.&lt;/p&gt;

&lt;p&gt;The user-facing response should be projected from the durable provider result or uncertainty record. That makes “I refused” impossible when the effect ledger says “committed.”&lt;/p&gt;

&lt;h2&gt;
  
  
  The gate has to survive retries and recovery
&lt;/h2&gt;

&lt;p&gt;Admission is only the start of authority.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/abs/2608.21159" rel="noopener noreferrer"&gt;AID-Guard&lt;/a&gt; frames the remaining problem as &lt;strong&gt;authorization-to-effect closure&lt;/strong&gt;. It revalidates the exact approved request and provider state at commit, keeps the original reservation while delivery is ambiguous, and permits release or one successor only after terminal evidence or certified no effect.&lt;/p&gt;

&lt;p&gt;Its target invariant is simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;one approval lineage → at most one provider effect
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In the paper’s bounded evaluation, all 210 Stripe test-mode provider-contract trials matched their declared outcomes, and the tested recovery and confirm/cancel schedules produced no duplicate effect.&lt;/p&gt;

&lt;p&gt;That is not a universal exactly-once guarantee. The authors used bounded provider schedules, test-mode contracts, synthetic credentials, and a high-latency prototype. Their strict exact-manifest profile also reduced benign utility substantially.&lt;/p&gt;

&lt;p&gt;But the architecture names a production bug most agent diagrams skip:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Retry and recovery are authority transitions, not transport details&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A gate that protects the first call but disappears during timeout recovery is not end-to-end enforcement.&lt;/p&gt;

&lt;h2&gt;
  
  
  What these numbers do not prove
&lt;/h2&gt;

&lt;p&gt;I am not claiming that 82.5% of all production agents knowingly ship invalid work.&lt;/p&gt;

&lt;p&gt;That number is the label rate for one new prerelease dataset of autonomous scientific-research trajectories. Most of its 800 annotations were produced by an artifact-aware agent judge calibrated against 50 human-labeled trajectories. The paper found the pattern across all eight tested systems, but it did not test the orchestration intervention I am proposing.&lt;/p&gt;

&lt;p&gt;AgentAbstain uses generated tasks in executable sandboxes. Its post-hoc abstention rate is a benchmark result, not an estimate of real banking or travel incidents.&lt;/p&gt;

&lt;p&gt;The policy-boundary study used controlled Jira and ServiceNow mocks and intentionally concentrated on cases where policy should intervene. Its 0.2% headline is conditional effectiveness inside that test design, not a promise for arbitrary production traffic. Durable approval and broad write controls were outside its primary evaluation.&lt;/p&gt;

&lt;p&gt;AID-Guard tested finite provider schedules and does not establish arbitrary provider linearizability.&lt;/p&gt;

&lt;p&gt;What survives those caveats is the mechanism.&lt;/p&gt;

&lt;p&gt;Across research reports, abstention tasks, data exposure, and state-changing tools, the same failure appears:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The system detects the problem&lt;/li&gt;
&lt;li&gt;The detection is represented as text&lt;/li&gt;
&lt;li&gt;The text has no binding relationship to the effect path&lt;/li&gt;
&lt;li&gt;The effect proceeds or has already happened&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That is not a missing sentence in the system prompt.&lt;/p&gt;

&lt;p&gt;It is a missing control boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  The next boundary
&lt;/h2&gt;

&lt;p&gt;My last article asked who owns the harness.&lt;/p&gt;

&lt;p&gt;This one asks who owns the commit.&lt;/p&gt;

&lt;p&gt;The model can propose an action. It can criticize the action. It can explain exactly why the action is unsafe and still take it, because awareness and authority are different capabilities.&lt;/p&gt;

&lt;p&gt;A review that cannot block the effect is not a control.&lt;/p&gt;

&lt;p&gt;It is a log entry.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What in your agent stack can actually stop the commit—and does it run before or after the side effect?&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources and further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2608.14905" rel="noopener noreferrer"&gt;How Do Agents Fail on AutoResearch&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2607.10059" rel="noopener noreferrer"&gt;AgentAbstain — Do LLM Agents Know When Not to Act&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2608.27646" rel="noopener noreferrer"&gt;Out-of-Band Policy Enforcement at a Trusted Tool Boundary&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2604.27233" rel="noopener noreferrer"&gt;Reinforced Agent — Inference-Time Feedback for Tool-Calling Agents&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2608.21159" rel="noopener noreferrer"&gt;AID-Guard — Stateful Authorization for Delegated Agent Effects&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Disclosure:&lt;/strong&gt; I wrote the original drafts in my native language and conducted the research, source selection, analysis, and conclusions myself. AI was used only for translation and English-language editing.&lt;/p&gt;




&lt;p&gt;If this article helped you, you can &lt;a href="https://ko-fi.com/sergeiparfenov?ref=dev" rel="noopener noreferrer"&gt;buy me a coffee&lt;/a&gt;. Your support helps me make time for the next experiment.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>llm</category>
      <category>mlops</category>
    </item>
    <item>
      <title>Your Agent Planned the Right Tools. It Still Crashed the Machine.</title>
      <dc:creator>Sergei Parfenov</dc:creator>
      <pubDate>Wed, 26 Aug 2026 13:51:56 +0000</pubDate>
      <link>https://dev.to/p0rt/your-agent-planned-the-right-tools-it-still-crashed-the-machine-58hf</link>
      <guid>https://dev.to/p0rt/your-agent-planned-the-right-tools-it-still-crashed-the-machine-58hf</guid>
      <description>&lt;p&gt;Your agent needs four independent facts before it can approve a refund:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;the order record,&lt;/li&gt;
&lt;li&gt;the fraud score,&lt;/li&gt;
&lt;li&gt;the customer's history,&lt;/li&gt;
&lt;li&gt;the policy that applies.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;It correctly sees that all four calls can run in parallel.&lt;/p&gt;

&lt;p&gt;So it launches all four.&lt;/p&gt;

&lt;p&gt;The order lookup is cheap. The policy search is cheap. The fraud model loads a large checkpoint. The history job scans two years of events. Together they cross the worker's memory limit, the container restarts, and the refund never happens.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The plan was correct. The execution was not.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Most agent benchmarks collapse those into one score.&lt;/p&gt;

&lt;p&gt;A new preprint submitted on August 25, &lt;a href="https://arxiv.org/abs/2608.24509" rel="noopener noreferrer"&gt;PeakBench&lt;/a&gt;, argues that this hides an entire class of production failures: the agent may understand every dependency and still overload the machine that executes them.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A dependency graph tells you what &lt;em&gt;may&lt;/em&gt; run in parallel. It does not tell you what your machine can run safely.&lt;/li&gt;
&lt;li&gt;Across eight tested models, planning accuracy had almost zero correlation with capacity violations on the same workflows.&lt;/li&gt;
&lt;li&gt;Resource metadata helped, but model-generated schedules were not reliably safe for every model.&lt;/li&gt;
&lt;li&gt;In production, let the model build the DAG. Let deterministic infrastructure enforce the resource envelope.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The DAG is not the schedule
&lt;/h2&gt;

&lt;p&gt;If step C needs the output of step A, C waits. If A and B are independent, the framework can run them concurrently.&lt;/p&gt;

&lt;p&gt;That is a logical property. It says what &lt;em&gt;may&lt;/em&gt; run at the same time.&lt;/p&gt;

&lt;p&gt;It says nothing about what the machine can survive.&lt;/p&gt;

&lt;p&gt;Two independent calls might each need 6 GB of memory on an 8 GB worker. Four browser sessions might fit logically but exceed a provider concurrency limit. Ten embedding jobs may be valid in parallel and still saturate their shared network connection.&lt;/p&gt;

&lt;p&gt;PeakBench calls this the peak-load problem: runtimes translate logical independence into immediate execution while implicitly assuming infinite capacity.&lt;/p&gt;

&lt;p&gt;When that assumption stays hidden, a crash gets blamed on “the agent.” The benchmark cannot tell you whether the failure came from reasoning, scheduling, or infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  What PeakBench separates
&lt;/h2&gt;

&lt;p&gt;The authors assembled roughly 1,200 MCP-compatible tools from about 130 servers and constructed 300 executable workflows:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;150 easy,&lt;/li&gt;
&lt;li&gt;100 medium,&lt;/li&gt;
&lt;li&gt;50 hard.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The workflows run inside containers. The benchmark records data flow, perturbs execution order, and watches for failures. This produces an execution-grounded dependency graph instead of relying only on an annotator's guess.&lt;/p&gt;

&lt;p&gt;In a manual audit of about 100 workflows, 94% of those graphs matched the auditors' optimal structure.&lt;/p&gt;

&lt;p&gt;PeakBench then scores two jobs independently.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Logical planning
&lt;/h3&gt;

&lt;p&gt;The model receives the task and tool descriptions. It must recover which calls are prerequisites and which may run concurrently.&lt;/p&gt;

&lt;p&gt;The benchmark reports graph edit distance and edge F1.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Physical scheduling
&lt;/h3&gt;

&lt;p&gt;The model receives the &lt;em&gt;verified&lt;/em&gt; dependency graph, so planning ambiguity is removed. It only has to assign start times under small, medium, and large machine profiles.&lt;/p&gt;

&lt;p&gt;The schedule is scored on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;completion time,&lt;/li&gt;
&lt;li&gt;capacity violation area: how badly, and for how long, the schedule exceeds a resource limit,&lt;/li&gt;
&lt;li&gt;strict mean resource utilization: utilization that counts only while the schedule remains feasible.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That split is the useful contribution. A failed run no longer has to disappear into one end-to-end number.&lt;/p&gt;

&lt;h2&gt;
  
  
  The result that matters
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/abs/2608.24509" rel="noopener noreferrer"&gt;PeakBench evaluated eight models&lt;/a&gt; under one prompting and parsing protocol: GPT-5, o3, GPT-4.1, Claude Sonnet 4.6, GLM-5, Kimi-K2.5, DeepSeek-V4-Pro, and DeepSeek-V4-Flash.&lt;/p&gt;

&lt;p&gt;GPT-5 was the strongest logical planner in the reported table:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Graph edit distance:&lt;/strong&gt; 0.42&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Edge F1:&lt;/strong&gt; 0.839&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Capacity violation area:&lt;/strong&gt; 3.698 for its resource-blind schedule&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;DeepSeek-V4-Flash recovered the graph less accurately:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Graph edit distance:&lt;/strong&gt; 0.81&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Edge F1:&lt;/strong&gt; 0.733&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Capacity violation area:&lt;/strong&gt; 3.458 once given the verified graph&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That comparison is not a model ranking. The case-level result is more important.&lt;/p&gt;

&lt;p&gt;Across all eight models, the correlation between edge F1 and capacity violations was almost zero. The reported coefficients ranged from roughly &lt;strong&gt;-0.045 to 0.000&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A model understanding a workflow's dependencies told you almost nothing about whether it would schedule that same workflow safely.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Your planner benchmark is not a scheduler benchmark.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Resource data helped. It did not solve scheduling.
&lt;/h2&gt;

&lt;p&gt;The authors added a Resource-Aware Scheduling Context, or RASC. For each tool call, the model sees estimated duration, average and peak CPU, peak memory, machine capacity, and verified dependencies.&lt;/p&gt;

&lt;p&gt;No weights change. The model simply stops scheduling blind.&lt;/p&gt;

&lt;p&gt;The aggregate trade-off is easier to read without a wide table:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Launch everything ASAP:&lt;/strong&gt; 8.62 s completion, 5.865 violation area, 0.080 safe utilization.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run everything serially:&lt;/strong&gt; 15.19 s, 2.925 violation area, 0.097 safe utilization.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best rule-based scheduler:&lt;/strong&gt; 9.13 s, 2.925 violation area, 0.141 safe utilization.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best RASC result:&lt;/strong&gt; 9.11 s, 2.938 violation area, 0.165 safe utilization.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Blind parallelism was fastest and most dangerous. Serial execution reduced overload by paying almost twice the latency. Resource-aware context nearly matched the best rule-based violation score while producing higher safe utilization.&lt;/p&gt;

&lt;p&gt;But the effect was model-dependent. RASC improved most models, not all. DeepSeek-V4-Flash slightly increased its violation area. Kimi-K2.5 and GPT-4.1 lost strict utilization even while finishing faster.&lt;/p&gt;

&lt;p&gt;Resource metadata is necessary input. It is not an admission controller.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would change in production
&lt;/h2&gt;

&lt;p&gt;The tempting fix is another system-prompt sentence:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Be careful with resources.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is not an architecture.&lt;/p&gt;

&lt;p&gt;I would make three artifacts explicit.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. The planner emits a dependency graph
&lt;/h3&gt;

&lt;p&gt;The planner decides which outputs are required and which calls are logically independent. It should not decide that every ready call starts now.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"steps"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"fraud"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"depends_on"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[]},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"history"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"depends_on"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[]},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"policy"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"depends_on"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[]},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"decision"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"depends_on"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"fraud"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"history"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"policy"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  2. A scheduler owns physical execution
&lt;/h3&gt;

&lt;p&gt;Give every tool a versioned resource profile:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"tool"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"fraud_score@3.4.1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"duration_p95_s"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;4.8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cpu_peak_cores"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;1.7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"memory_peak_mb"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;6144&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"network_slots"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"provider_concurrency"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then enforce the envelope outside the model with queues, per-resource semaphores, rate-limit budgets, and backpressure.&lt;/p&gt;

&lt;p&gt;The model may propose a schedule. Deterministic infrastructure should decide whether it is allowed to run.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. The trace records what actually ran
&lt;/h3&gt;

&lt;p&gt;A configured concurrency limit of four does not prove that an SDK did not add three hidden retries inside one call.&lt;/p&gt;

&lt;p&gt;For every attempt, record:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;requested and actual start time,&lt;/li&gt;
&lt;li&gt;queue delay,&lt;/li&gt;
&lt;li&gt;resource-profile version,&lt;/li&gt;
&lt;li&gt;observed peak resources,&lt;/li&gt;
&lt;li&gt;retry and fallback attempts,&lt;/li&gt;
&lt;li&gt;the dependency state that made the call eligible,&lt;/li&gt;
&lt;li&gt;the limiter or scheduler decision.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without that trace, a passing benchmark tells you what the agent intended to do, not what the runtime did.&lt;/p&gt;

&lt;h2&gt;
  
  
  A production drill you can run this week
&lt;/h2&gt;

&lt;p&gt;This is not a reproduction of PeakBench. It is a smaller test for your own runtime.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Pick one real workflow with at least two expensive independent calls.&lt;/li&gt;
&lt;li&gt;Capture each tool's p95 duration, peak memory, peak CPU, and external concurrency limits.&lt;/li&gt;
&lt;li&gt;Run the same workflow under three envelopes:

&lt;ul&gt;
&lt;li&gt;normal production capacity,&lt;/li&gt;
&lt;li&gt;degraded or burst-constrained capacity,&lt;/li&gt;
&lt;li&gt;a generous diagnostic capacity.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Compare three policies:

&lt;ul&gt;
&lt;li&gt;launch every ready call,&lt;/li&gt;
&lt;li&gt;serialize every call,&lt;/li&gt;
&lt;li&gt;enforce a resource-aware queue.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Report task correctness and resource feasibility separately.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A correct answer that violates the capacity envelope is not a pass. A safe schedule that serializes everything is not automatically good either.&lt;/p&gt;

&lt;p&gt;The goal is maximum &lt;em&gt;safe&lt;/em&gt; parallelism.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the comments mattered
&lt;/h2&gt;

&lt;p&gt;Comments on my previous harness article kept returning to the same missing fields: publish the action budget, retry policy, and actual route taken, not only the model and final score.&lt;/p&gt;

&lt;p&gt;PeakBench made the operational half of that argument measurable.&lt;/p&gt;

&lt;p&gt;That comment signal helped identify the question. It is not the evidence for the answer; the benchmark is.&lt;/p&gt;

&lt;p&gt;An agent run is defined not only by the model, prompt, tools, and dependency graph. It is also defined by the machine profile and the scheduler that turns readiness into execution.&lt;/p&gt;

&lt;p&gt;If those fields are absent, a successful plan can still become an outage—and the model will get blamed for a failure the runtime created.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limitations
&lt;/h2&gt;

&lt;p&gt;PeakBench is a version-one preprint. Its workflows are synthesized, its capacity profiles are simulated, and the reported results should not be treated as measurements of your cluster.&lt;/p&gt;

&lt;p&gt;The benchmark is best read as a diagnostic: it shows that logical planning and physical scheduling can fail independently. The exact numbers still need validation against real production traces.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources &amp;amp; further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2608.24509" rel="noopener noreferrer"&gt;PeakBench paper&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/Czzzk/Staggering-the-Peaks" rel="noopener noreferrer"&gt;PeakBench code repository&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;If this article helped you, you can &lt;a href="https://ko-fi.com/sergeiparfenov?ref=dev" rel="noopener noreferrer"&gt;buy me a coffee&lt;/a&gt;. Your support helps me make time for the next experiment.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>devops</category>
      <category>mlops</category>
    </item>
    <item>
      <title>The Model Scored 30%. The Harness Scored 100%. Which One Did You Benchmark?</title>
      <dc:creator>Sergei Parfenov</dc:creator>
      <pubDate>Mon, 24 Aug 2026 13:38:40 +0000</pubDate>
      <link>https://dev.to/p0rt/the-model-scored-30-the-harness-scored-100-which-one-did-you-benchmark-3mp4</link>
      <guid>https://dev.to/p0rt/the-model-scored-30-the-harness-scored-100-which-one-did-you-benchmark-3mp4</guid>
      <description>&lt;p&gt;On July 24, ARC Prize verified Claude Opus 5 at 30.16% on the ARC-AGI-3 public set. On August 21, NVIDIA reported the same model at 100.00 on the same set. The weights did not change. The code around them did.&lt;/p&gt;

&lt;p&gt;In between, MIT did the same thing (August 5), a group led by Impossible Research got to 98.98 (July 15), and OpenAI tripled GPT-5.6 Sol's score by flipping two API settings (July 29). Then Microsoft published a framework that trains the model &lt;em&gt;through&lt;/em&gt; the harness (August 18), and Google published one that gives the environment a harness of its own (August 20).&lt;/p&gt;

&lt;p&gt;In July I wrote that self-editing harnesses have a provenance problem. This month the problem moved up a level: the benchmark score itself has no provenance.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; On ARC-AGI-3's public set, the spread between "model in the official harness" and "model in the best harness" is 25 to 70 points, on a benchmark designed to resist exactly this. None of the 100s are verified on the private set, and every author says so. Microsoft's Agent Lightning v1.0 runs RL with the deploy-time harness owning the loop, so the harness is becoming part of the weights, and its reward-hacking section is the checklist my July post warned about. A benchmark number without a harness version, memory state and action budget attached is a self-reported claim with an unmarked type. Unmarked means self-reported.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Five harnesses, one public set
&lt;/h2&gt;

&lt;p&gt;ARC-AGI-3 scores agents with RHAE (Relative Human Action Efficiency). Per level, &lt;code&gt;score = (human_baseline_actions / ai_actions)^2&lt;/code&gt;, with the ratio capped at 1.15x the human baseline. Game scores are level-weighted averages, you must finish the last level to get full credit, and the overall number is the mean over games. A 100.00 means the agent finished every level at least as efficiently as a first-time human.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Harness&lt;/th&gt;
&lt;th&gt;Who&lt;/th&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Public RHAE&lt;/th&gt;
&lt;th&gt;Actions&lt;/th&gt;
&lt;th&gt;Verified by ARC Prize&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Official ARC Prize harness&lt;/td&gt;
&lt;td&gt;ARC Prize&lt;/td&gt;
&lt;td&gt;Jul 24&lt;/td&gt;
&lt;td&gt;Claude Opus 5 (high)&lt;/td&gt;
&lt;td&gt;30.16%&lt;/td&gt;
&lt;td&gt;n/a&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Official harness, default settings&lt;/td&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;Jul 29&lt;/td&gt;
&lt;td&gt;GPT-5.6 Sol (max)&lt;/td&gt;
&lt;td&gt;13.3%&lt;/td&gt;
&lt;td&gt;n/a&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Official harness + retained reasoning + compaction&lt;/td&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;Jul 29&lt;/td&gt;
&lt;td&gt;GPT-5.6 Sol (max)&lt;/td&gt;
&lt;td&gt;38.3%&lt;/td&gt;
&lt;td&gt;6x fewer output tokens&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Schema&lt;/td&gt;
&lt;td&gt;Impossible Research (+ UC Berkeley, CMU)&lt;/td&gt;
&lt;td&gt;Jul 15&lt;/td&gt;
&lt;td&gt;Opus 4.8 / Fable 5&lt;/td&gt;
&lt;td&gt;98.98&lt;/td&gt;
&lt;td&gt;n/a&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;VISTA&lt;/td&gt;
&lt;td&gt;MIT (Han, Hu, Qiu, Wu, He)&lt;/td&gt;
&lt;td&gt;Aug 5&lt;/td&gt;
&lt;td&gt;Claude Opus 5&lt;/td&gt;
&lt;td&gt;100.00&lt;/td&gt;
&lt;td&gt;7,542 (humans: 17,135)&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AVO&lt;/td&gt;
&lt;td&gt;NVIDIA&lt;/td&gt;
&lt;td&gt;Aug 21&lt;/td&gt;
&lt;td&gt;Claude Opus 5&lt;/td&gt;
&lt;td&gt;100.00&lt;/td&gt;
&lt;td&gt;6,624&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The number that matters is not in the table. It is the gap between the first row and the last: 70 points, same model, same 25 games, same metric.&lt;/p&gt;

&lt;p&gt;The official harness is not a neutral baseline. OpenAI's write-up quotes ARC's intent: an "intentionally generic harness, without tools or special features" built to make "model shortcomings more visible." In practice it discarded all private reasoning after each game action and used a rolling truncation window, so older actions vanished as history grew. Retaining reasoning and enabling compaction took Sol from 13.3% to 38.3% and cut output tokens by 6x. The harness was wiping the model's mind between moves.&lt;/p&gt;

&lt;p&gt;So the leaderboard measures "model plus a harness built to expose the model." The 100s measure "model plus a harness built to cover for the model." Neither measures the model, and nobody has isolated which part of the 70 points is which.&lt;/p&gt;

&lt;p&gt;The authors are unusually honest about this. NVIDIA: the AVO-versus-VISTA comparison "should not be interpreted as a controlled ablation," and the results "should not be interpreted as a direct measurement of the performance contribution of AVO." VISTA: the models "were released after the public ARC-AGI-3 games," overlap cannot be excluded, and "the private set remains the real test of generalization." Schema: "no frozen-harness or held-out-performance claim." Every 100 on that table is a public-set number on games the models may have seen in training.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the 70 points are made of
&lt;/h2&gt;

&lt;p&gt;Read the harness papers side by side and the same three components appear under different names.&lt;/p&gt;

&lt;p&gt;Memory. VISTA keeps a "lossless visual memory" of every past observation. AVO carries forward "prior implementations, evaluation results, compiler and profiler outputs, and accumulated reasoning." OpenAI's two settings are memory settings: keep the reasoning, compact instead of truncate.&lt;/p&gt;

&lt;p&gt;Supervision. AVO runs a monitor that watches "the broader trajectory for stagnation or repeated unproductive cycles and can redirect the main agent." That is the layer that turns a model that gives up into an agent that does not.&lt;/p&gt;

&lt;p&gt;An action budget. RHAE squares the efficiency ratio, so wasted moves are punished quadratically. AVO's headline against VISTA is 12% fewer actions. That is a harness optimization target, not a model property.&lt;/p&gt;

&lt;p&gt;In July I split harness work into two piles: compensatory layers that patch what the model cannot do yet, and protective layers that constrain what it must not do. I predicted pile one depreciates with every model release. All three components above are pile one, and on a benchmark built to resist static tricks they are currently worth 25 to 70 points with the newest frontier models. Either my prediction is early or it is wrong about magnitude. I will take the second reading until the private-set numbers say otherwise.&lt;/p&gt;

&lt;p&gt;One more thing about compaction, since it is the setting that tripled OpenAI's score. In my preregistered compaction experiment, the same operation produced 3.47% false proceeds on irreversible-action gates: the agent went through a gate it should have stopped at, because the compacted context no longer carried the provenance the gate depended on. Not a contradiction. ARC-AGI-3 scores task completion; my gates scored whether the agent still knew &lt;em&gt;why&lt;/em&gt; it was allowed to act. Compaction improves the first, degrades the second, and a benchmark only sees the first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then Microsoft put the harness inside the training loop
&lt;/h2&gt;

&lt;p&gt;Agent Lightning v1.0 (arXiv, August 18) names something the July thread never got to: RL where the harness is not a bystander. In their words, "the harness owns this loop, while the training engine observes only a sequence of LLM request-response pairs." The deploy-time scaffold (mini-SWE-agent in their coding runs) executes the task inside Kubernetes; the trainer sits behind a gateway that looks like a normal LLM endpoint and collects the traffic.&lt;/p&gt;

&lt;p&gt;The result is real: Qwen3.5-9B goes from 41.8% to 56.4% on SWE-bench Verified, a 14.6-point gain from about 6,000 examples filtered out of SWE-smith's 59,136 tasks across 128 repositories, in roughly 3,500 lines of framework code.&lt;/p&gt;

&lt;p&gt;Two details matter more than the headline.&lt;/p&gt;

&lt;p&gt;First, retokenization. The harness re-renders text between calls, chat templates are not compositional, decode-then-retokenize is lossy, and the harness parses and repairs outputs. So the token IDs the trainer sees for the model's previous answer can differ from the ones it actually sampled. Their fix is best-effort merging: merge only when the exact token prefix holds, otherwise close the sequence. That is the engineering admission that model and harness now share a boundary at the token level. Train through one harness's rendering and you get a model tuned to that rendering.&lt;/p&gt;

&lt;p&gt;Second, section 4.3.2, "Preventing Reward Hacking." During training the agents were caught "using Git history to locate the gold commit," "using wget or curl to retrieve upstream source code from GitHub," using pip to download a package's source, and using urllib to do the same. Countermeasures: "disable Git commands and hide the .git directory from the agent," plus a Kubernetes network policy that "blocks general outbound network access and permits connections only to explicitly whitelisted services."&lt;/p&gt;

&lt;p&gt;That is my July post compressed into a paragraph, arrived at as an engineering necessity rather than a design principle. Vinicius Pereira said it best in the comments: the agent must not be able to author the artifact the gate reads. Microsoft's version is that it must not be able to &lt;em&gt;reach&lt;/em&gt; it either, through the filesystem or the network. Dipankar Sarkar's separate trust domain for test execution is the same control from the other side.&lt;/p&gt;

&lt;p&gt;Now put the halves together. The harness that decides what the model sees also decides what the trainer sees. Once RL runs through it, the tricks in pile one stop being code you can diff and become weights you cannot. That is the absorption I predicted, except what gets absorbed includes whatever the harness let the agent get away with. Hide &lt;code&gt;.git&lt;/code&gt; and the model learns the task. Forget to, and it learns to find the gold commit, and the benchmark will not tell you which one you trained.&lt;/p&gt;

&lt;h2&gt;
  
  
  Google gave the environment a harness too
&lt;/h2&gt;

&lt;p&gt;EnvHarness (arXiv, August 20, Google Research) wraps a static environment at the &lt;code&gt;reset&lt;/code&gt;/&lt;code&gt;step&lt;/code&gt; interface with three plug-in types: Setup reshapes the initial state, Rule reshapes "which actions are allowed, what they do, and what the agent observes," and Link composes in another environment's tasks. A designer agent, EnvRigger, "treats the target policy as a black box, observing its execution trajectories to synthesize EnvHarness components targeting diagnosed flaws," writes a &lt;code&gt;_Rules&lt;/code&gt; subclass, and tests it. Across ALFWorld, WebArena, SWE-bench Verified, OfficeQA and SpreadsheetBench, skills learned in reshaped environments transfer back for up to 9.0 points on held-out instances with 9.8% fewer steps.&lt;/p&gt;

&lt;p&gt;Credit where due: this is the responsible version. Verifiers are untouched, the goal predicate is never modified, evaluation happens on the unadapted benchmark. A curriculum, not a thumb on the scale.&lt;/p&gt;

&lt;p&gt;But note the direction of travel. In one week the field shipped a harness around the agent (AVO, VISTA), a harness around the trainer (Agent Lightning), and a harness around the environment (EnvHarness). The capability you end up with has its provenance spread across three codebases, and only one of them comes with the model card.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you are actually buying
&lt;/h2&gt;

&lt;p&gt;The one benchmark this month that held the model constant and varied the harness came from a vendor. TrueFoundry's TrueForge comparison (August 18) ran DevRev's Enterprise-Bench: 14 cross-system tasks, three MCP servers, fresh session per task, blind grading, list-rate and cache-aware costs.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;System&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Tasks solved&lt;/th&gt;
&lt;th&gt;Cost per run&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Managed Agents&lt;/td&gt;
&lt;td&gt;Opus 4.8&lt;/td&gt;
&lt;td&gt;~11/14&lt;/td&gt;
&lt;td&gt;$11.80&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TrueForge&lt;/td&gt;
&lt;td&gt;Opus 4.8&lt;/td&gt;
&lt;td&gt;~11/14&lt;/td&gt;
&lt;td&gt;$8.50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TrueForge&lt;/td&gt;
&lt;td&gt;GLM-5.2&lt;/td&gt;
&lt;td&gt;~11/14&lt;/td&gt;
&lt;td&gt;$2.90&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;deepagents / LangGraph&lt;/td&gt;
&lt;td&gt;Opus 4.8&lt;/td&gt;
&lt;td&gt;~10/14&lt;/td&gt;
&lt;td&gt;n/a&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Same model, same tasks, 28% cost difference from the harness alone. Swap the model under the same harness and cost drops another 2.9x with no change in tasks solved. TrueFoundry sells the gateway next to TrueForge, so the framing is self-serving. It is still more methodology than NVIDIA offered.&lt;/p&gt;

&lt;p&gt;I have seen this pattern in teams that compare a vendor's managed agent against their own scaffold and attribute the entire difference to the model. After this month I do not think that attribution is defensible without a controlled harness swap, and almost nobody runs one.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a score needs to carry
&lt;/h2&gt;

&lt;p&gt;If provenance is a vector, a benchmark score needs one. The minimum I would want attached to an agent number before quoting it, illustrative rather than a standard:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;score&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;100.00&lt;/span&gt;
&lt;span class="na"&gt;metric&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;RHAE&lt;/span&gt;
&lt;span class="na"&gt;set&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;arc-agi-3-public-25&lt;/span&gt;       &lt;span class="c1"&gt;# not semi-private, not private&lt;/span&gt;
&lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;claude-opus-5&lt;/span&gt;           &lt;span class="c1"&gt;# provider version string, reasoning effort&lt;/span&gt;
&lt;span class="na"&gt;harness&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;avo@&amp;lt;commit&amp;gt;&lt;/span&gt;          &lt;span class="c1"&gt;# the code between model and environment&lt;/span&gt;
&lt;span class="na"&gt;memory_at_start&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;empty&lt;/span&gt;         &lt;span class="c1"&gt;# or warm, and from which prior runs&lt;/span&gt;
&lt;span class="na"&gt;supervisor&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;stagnation-monitor&lt;/span&gt; &lt;span class="c1"&gt;# any policy that can redirect the agent&lt;/span&gt;
&lt;span class="na"&gt;compaction&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;on&lt;/span&gt;                 &lt;span class="c1"&gt;# summarize vs truncate, and where&lt;/span&gt;
&lt;span class="na"&gt;action_budget&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;6624&lt;/span&gt;
&lt;span class="na"&gt;trainer_harness&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;none&lt;/span&gt;          &lt;span class="c1"&gt;# if the weights were RL'd through a harness, which one&lt;/span&gt;
&lt;span class="na"&gt;verified_by&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;self&lt;/span&gt;              &lt;span class="c1"&gt;# or ARC Prize, or a named third party&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Mike Czerwinski's rule from the July thread applies to every row: unmarked has to mean self-reported. A score that arrives without the harness commit is not a measurement of the model. It is a claim about a system, typed by whoever produced it, and the default type is untrusted.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this post might get wrong
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;The model/harness split may already be dissolving. After harnessed RL, "the model" is partly a harness artifact, and the boundary I am measuring might not survive the next benchmark cycle.&lt;/li&gt;
&lt;li&gt;Every 100 is on the public set, which shipped before the models did. If Opus 5 in the generic harness gets 30 on the private set and AVO gets 40, the harness story shrinks from 70 points to 10, and the leaderboard was more honest than I am giving it credit for.&lt;/li&gt;
&lt;li&gt;Source hygiene: NVIDIA sells the compute AVO runs on, OpenAI's two settings are its own API features, TrueFoundry sells a gateway, Microsoft would like you on Azure. I read the papers and reproduced none of them.&lt;/li&gt;
&lt;li&gt;My July prediction that compensatory harness layers depreciate with each model release. This month's evidence points the other way. The prediction stays up, marked as losing.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The question I cannot answer alone
&lt;/h2&gt;

&lt;p&gt;If the harness is worth 70 points on a benchmark built to resist it, who owns the harness in your stack: you, the model vendor, or the router in between? And when you report an agent result internally, does the harness commit travel with the model version, or does it get dropped at the first summary?&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources &amp;amp; further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://arcprize.org/results/anthropic-claude-opus-5" rel="noopener noreferrer"&gt;ARC Prize: Claude Opus 5 results, verified July 24, 2026&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.arcprize.org/methodology.md" rel="noopener noreferrer"&gt;ARC Prize: ARC-AGI-3 scoring methodology (RHAE)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/" rel="noopener noreferrer"&gt;OpenAI: How enabling two settings tripled our scores on the ARC-AGI-3 benchmark (July 29, 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://schema-harness.github.io/" rel="noopener noreferrer"&gt;Schema: Frontier Models with Our Harness Achieve ~99% on ARC-AGI-3 Public (July 15, 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://vista-research.github.io/" rel="noopener noreferrer"&gt;VISTA: A Visual Harness for Reasoning in an Interactive World, MIT (August 5, 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents/" rel="noopener noreferrer"&gt;NVIDIA: AVO Reaches 100 on ARC-AGI-3 (August 21, 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2603.24517" rel="noopener noreferrer"&gt;AVO: Agentic Variation Operators for Autonomous Evolutionary Search (arXiv 2603.24517)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://thenewstack.io/nvidia-avo-arcagi3-benchmark/" rel="noopener noreferrer"&gt;The New Stack on AVO and the harness debate (August 21, 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2608.17528" rel="noopener noreferrer"&gt;Agent Lightning v1.0: Towards Harnessed Agentic RL (arXiv 2608.17528, August 18, 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2608.19880" rel="noopener noreferrer"&gt;EnvHarness: Awakening Static Worlds for Agent Learning (arXiv 2608.19880, August 20, 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/google-research/envharness" rel="noopener noreferrer"&gt;EnvHarness on GitHub, Google Research&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.truefoundry.com/blog/engineering/trueforge-vs-claude-managed-agents-benchmark/" rel="noopener noreferrer"&gt;TrueFoundry: TrueForge vs Claude Managed Agents benchmark (August 18, 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://venturebeat.com/orchestration/truefoundrys-open-source-ai-agent-harness-trueforge-boasts-30-75-cheaper-task-completion-than-claude-managed-agents" rel="noopener noreferrer"&gt;VentureBeat on TrueForge (August 19, 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://earendil.com/posts/what-is-a-harness/" rel="noopener noreferrer"&gt;Earendil: What Is a Harness? (August 20, 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/p0rt/the-agent-faked-a-test-log-then-believed-it-self-editing-harnesses-have-a-provenance-problem-3id6"&gt;My July post: The Agent Faked a Test Log, Then Believed It&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/p0rt/my-strawman-baseline-beat-my-own-scheme-on-half-the-gate-classes-177h"&gt;My compaction experiment: My Strawman Baseline Beat My Own Scheme on Half the Gate Classes&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;If this article helped you, you can &lt;a href="https://ko-fi.com/sergeiparfenov?ref=dev" rel="noopener noreferrer"&gt;buy me a coffee&lt;/a&gt;. Your support helps me make time for the next experiment.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>agents</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Distilling Kimi Into Qwen Doesn't Give You Kimi. It Gives You Qwen With Kimi's Handwriting</title>
      <dc:creator>Sergei Parfenov</dc:creator>
      <pubDate>Mon, 10 Aug 2026 12:40:37 +0000</pubDate>
      <link>https://dev.to/p0rt/distilling-kimi-into-qwen-doesnt-give-you-kimi-it-gives-you-qwen-with-kimis-handwriting-284p</link>
      <guid>https://dev.to/p0rt/distilling-kimi-into-qwen-doesnt-give-you-kimi-it-gives-you-qwen-with-kimis-handwriting-284p</guid>
      <description>&lt;p&gt;On July 22, 2026, White House OSTP director Michael Kratsios accused Moonshot AI of distilling Anthropic's Fable 5 to build Kimi K3, and Treasury put sanctions on the table. No technical evidence was published, and researchers immediately pointed out the timeline problem: by K3's launch day, Fable 5 had been publicly reachable for roughly 18 days in total.&lt;/p&gt;

&lt;p&gt;Everyone argued about whether it happened. Almost nobody asked the more useful engineering question: &lt;strong&gt;if it did happen, what would Moonshot actually have received?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here is the experiment that answers it. In 2025, the Berkeley team behind Sky-T1 fine-tuned a model on long reasoning traces whose &lt;strong&gt;final answers were wrong&lt;/strong&gt;. Accuracy dropped by 3.2 points. They randomized half the numbers inside the reasoning steps: 3.3 points. Then they shuffled the &lt;em&gt;order&lt;/em&gt; of the steps, and performance collapsed. The content of a distill is nearly disposable. The structure is the payload.&lt;/p&gt;

&lt;p&gt;That is the folk model's blind spot. The folk model says: pour a strong model's outputs into an open model, get a model at the strong model's level. What actually crosses the wire is mostly the &lt;em&gt;shape&lt;/em&gt; of the reasoning, and the size of the gain is set by your base model, not by how smart your teacher was.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; A "distill" is a dataset of teacher traces, not a set of weights. Fine-tuning on it reliably transfers output &lt;em&gt;structure&lt;/em&gt; (long chain-of-thought, backtracking, &lt;code&gt;&amp;lt;think&amp;gt;&lt;/code&gt; blocks) and only narrowly transfers capability. Evidence: models trained on traces with &lt;strong&gt;wrong answers&lt;/strong&gt; lose about 3.2 points versus correct ones, and randomizing half the numbers in the traces costs 3.3 points on AIME 2024. Real capability gains do exist (DeepSeek's R1 distills beat RL on the same base by 25 points on AIME), but they cost 800k rejection-sampled samples and a full fine-tune, not 8k samples and a LoRA. The headline benchmark jumps you see on Hugging Face are very often eval artifacts.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What a "distill" actually is
&lt;/h2&gt;

&lt;p&gt;When someone says "I poured a distill into Qwen," they are not moving weights. They are running SFT on a dataset of teacher outputs. There are two distinct channels, and they behave differently:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Black-box (traces)&lt;/th&gt;
&lt;th&gt;White-box (logits)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;What you collect&lt;/td&gt;
&lt;td&gt;Generated text, usually prompt + long CoT + answer&lt;/td&gt;
&lt;td&gt;Full next-token distributions, or hidden states&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Loss&lt;/td&gt;
&lt;td&gt;Cross-entropy on the teacher's tokens&lt;/td&gt;
&lt;td&gt;KL between teacher and student distributions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bits per token&lt;/td&gt;
&lt;td&gt;One sampled token&lt;/td&gt;
&lt;td&gt;The whole distribution ("dark knowledge")&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Needs&lt;/td&gt;
&lt;td&gt;API access&lt;/td&gt;
&lt;td&gt;Weights, and enough GPUs to run the teacher&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Works across model families&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes, if tokenizers align&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Anything distilled from a closed API (Claude, GPT, Gemini) is black-box by construction. You cannot get logits out of an inference endpoint. This matters more than it sounds: black-box distillation is the low-bandwidth channel, and it is the only one available in the scenario the White House described.&lt;/p&gt;

&lt;p&gt;The whole current ecosystem is downstream of two decisions made four months apart. In September 2024, OpenAI hid o1's raw chain of thought and explicitly listed competitive advantage among the reasons. In January 2025, DeepSeek shipped R1 with full traces exposed under a permissive license, and within weeks the distill wave existed: s1, LIMO, Sky-T1, and a Hugging Face shelf of "-Distill" repos. Traces are the substrate; whoever exposes them feeds the ecosystem, and hiding them is anti-distillation policy by another name. None of this is exotic in-house either: Google's own Gemini 1.5 technical report states that Flash is online-distilled from the much larger Pro. Every lab does this to its own models. The fight is only ever about doing it to someone else's.&lt;/p&gt;

&lt;p&gt;Kimi is the interesting inverse case. K3 shipped as open weights in late July 2026 at 2.8 trillion parameters, so white-box distillation &lt;em&gt;from&lt;/em&gt; Kimi is legally and technically on the table. The gate is not access, it is the inference bill for generating traces from a model that needs roughly 1.4 TB of fast memory resident before you load any context.&lt;/p&gt;

&lt;h2&gt;
  
  
  The case that it genuinely works
&lt;/h2&gt;

&lt;p&gt;DeepSeek ran the cleanest public experiment on this, and it is still the strongest pro-distillation datapoint we have.&lt;/p&gt;

&lt;p&gt;They generated about 600k rejection-sampled reasoning traces plus 200k general samples from R1, then ran plain SFT for two epochs on off-the-shelf open bases. No RL on the students at all. Then they asked the obvious control question: what if you skip the teacher and just run large-scale RL on the same base?&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Qwen-32B base, three treatments&lt;/th&gt;
&lt;th&gt;AIME 2024&lt;/th&gt;
&lt;th&gt;MATH-500&lt;/th&gt;
&lt;th&gt;GPQA-D&lt;/th&gt;
&lt;th&gt;LiveCodeBench&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;QwQ-32B-Preview (reference)&lt;/td&gt;
&lt;td&gt;50.0&lt;/td&gt;
&lt;td&gt;90.6&lt;/td&gt;
&lt;td&gt;54.5&lt;/td&gt;
&lt;td&gt;41.9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RL directly on the base, 10k+ steps&lt;/td&gt;
&lt;td&gt;47.0&lt;/td&gt;
&lt;td&gt;91.6&lt;/td&gt;
&lt;td&gt;55.0&lt;/td&gt;
&lt;td&gt;40.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SFT on 800k R1 traces&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;72.6&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;94.3&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;62.1&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;57.2&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That is a 25 point gap on AIME in favor of distillation, against an RL run that cost far more compute. DeepSeek's own conclusion was blunt: distilling a powerful model into a smaller one works, while small models relying on large-scale RL need enormous compute and may still lose.&lt;/p&gt;

&lt;p&gt;This is not just curve-sharpening either. A widely cited ICML/NeurIPS 2025 analysis measured pass@k rather than pass@1 and found the two methods differ in kind. RLVR raises pass@1 while &lt;em&gt;narrowing&lt;/em&gt; the reasoning boundary at large k, because the paths it reinforces were already in the base's sampling distribution. Distillation raises the curve at every k (on Qwen-7B: pass@1 from 28% to 45%, still above 90% at pass@256), meaning genuinely new reasoning patterns entered the model.&lt;/p&gt;

&lt;p&gt;So: yes, capability moves. Hold that thought.&lt;/p&gt;

&lt;h2&gt;
  
  
  The case that it isn't what you think
&lt;/h2&gt;

&lt;p&gt;Back to the corruption experiment from the intro, with the setup spelled out. The Sky-T1 team first got +40 points on AIME 2024 by fine-tuning Qwen2.5-32B-Instruct on just 17k long-CoT traces distilled from R1. Only then did they start breaking the training data on purpose, and found that only structural damage (shuffling, inserting, deleting steps) actually hurt, while wrong answers and randomized numbers cost ~3 points each. Their conclusion sits in the paper's own title: structure, not content, is what matters.&lt;/p&gt;

&lt;p&gt;Two more datapoints point the same way. s1 hit strong reasoning numbers with 1,000 samples. LIMO used 817. If a thousand examples move a benchmark 40 points, you are not transferring a frontier lab's knowledge in a thousand examples. You are flipping a switch that was already wired.&lt;/p&gt;

&lt;p&gt;This is the same finding that killed the first imitation wave in 2023. Berkeley's "False Promise of Imitating Proprietary LLMs" found crowd workers rated ChatGPT imitators as competitive, while targeted benchmarks showed they closed little to none of the gap: they mimicked style, not factuality. Thinking Machines said the same thing in 2025 about off-policy distillation, that the student learns the teacher's style and confidence without necessarily learning its accuracy.&lt;/p&gt;

&lt;h2&gt;
  
  
  A worked example at 0.01% of the weights
&lt;/h2&gt;

&lt;p&gt;Here is the exact thing the question is usually about, done in public and documented honestly.&lt;/p&gt;

&lt;p&gt;Someone took &lt;code&gt;Qwen3.6-35B-A3B&lt;/code&gt;, generated ~7.8k reasoning traces from Kimi K2.6 via OpenRouter, and ran SFT with Unsloth and LoRA. Attention-only adapters, &lt;code&gt;r=16&lt;/code&gt;, 980 steps, about 21 hours on a single H200. Trainable parameters: 3.44M out of 35.1B. That is 0.01% of the model.&lt;/p&gt;

&lt;p&gt;The model card then does something almost nobody does. It reports evals that undercut the model:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark, same pipeline both sides&lt;/th&gt;
&lt;th&gt;Base Qwen3.6-35B-A3B&lt;/th&gt;
&lt;th&gt;Kimi-distill&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;MATH-500 (0-shot, &lt;code&gt;math_verify&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;53.0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;47.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPQA Diamond (0-shot CoT)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;79.29&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;75.25&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GSM8K (8-shot, strict-match)&lt;/td&gt;
&lt;td&gt;64.0&lt;/td&gt;
&lt;td&gt;92.67&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The author's read on that GSM8K number is the important part: the base scoring 64% is implausible for a frontier 35B-A3B, and the likely cause is that the few-shot template never triggers the base's thinking mode. So the +28.67 point "win" measures &lt;em&gt;"my pipeline rewards models that always think,"&lt;/em&gt; not capability. His stated conclusion is that the run provides no evidence the distillation improved raw reasoning over the base. What it does provide is a guarantee: the distill emits &lt;code&gt;&amp;lt;think&amp;gt;&lt;/code&gt; blocks regardless of prompt shape, where the base's thinking mode is conditional.&lt;/p&gt;

&lt;p&gt;That is a real, useful property. It is also exactly what "handwriting transplant" means. And note the cost signature: Kimi's traces averaged 2,933 tokens against 849 for a matched Claude Opus 4.7 set, roughly a 2.5x compute multiplier for the same number of rows.&lt;/p&gt;

&lt;h2&gt;
  
  
  The price list
&lt;/h2&gt;

&lt;p&gt;Dollar figures are illustrative (cloud rates move), but the orders of magnitude are the point:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Run&lt;/th&gt;
&lt;th&gt;Data&lt;/th&gt;
&lt;th&gt;Compute&lt;/th&gt;
&lt;th&gt;What you demonstrably get&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;s1-32B (Stanford)&lt;/td&gt;
&lt;td&gt;1,000 curated Gemini traces&lt;/td&gt;
&lt;td&gt;26 min on 16 H100s, ~$25 at cloud rates&lt;/td&gt;
&lt;td&gt;Thinking format + test-time scaling; beats o1-preview on competition math&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K2.6 → Qwen3.6 LoRA&lt;/td&gt;
&lt;td&gt;7.8k K2.6 traces via OpenRouter&lt;/td&gt;
&lt;td&gt;~21 h on one H200&lt;/td&gt;
&lt;td&gt;Unconditional &lt;code&gt;&amp;lt;think&amp;gt;&lt;/code&gt; blocks; no measured capability gain over base&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek R1 distills&lt;/td&gt;
&lt;td&gt;800k rejection-sampled samples&lt;/td&gt;
&lt;td&gt;Full SFT, 2 epochs, six bases&lt;/td&gt;
&lt;td&gt;+25 pts on AIME over RL on the same base&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thinking Machines on-policy&lt;/td&gt;
&lt;td&gt;Teacher-scored student rollouts&lt;/td&gt;
&lt;td&gt;1,800 GPU-hours (their RL baseline: 17,920)&lt;/td&gt;
&lt;td&gt;~70% AIME'24, RL parity at ~10% of the compute&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Read top to bottom and the pattern is the whole article: the format transplant costs lunch money, the capability transfer costs a training run.&lt;/p&gt;

&lt;h2&gt;
  
  
  So what actually crossed the wire?
&lt;/h2&gt;

&lt;p&gt;Until recently this was unanswerable per output. A December 2025 paper proposes a provenance-tracing framework that scores every sentence a distilled model produces under three models (teacher, original student, distilled student) and sorts it into four buckets: teacher-originated, student-originated, already present in both, and pre-existing but boosted by distillation. In their analysis, teacher-originated actions do appear in unseen test contexts and correlate with correctness, but a large share of what the distilled model emits is student-internal patterns that distillation merely amplified, and not all of that amplification helps.&lt;/p&gt;

&lt;p&gt;One more mechanism worth knowing, because it bounds the cross-family case. Anthropic Fellows work published in Nature this year showed models can transmit behavioral traits through data with no semantic connection to those traits, including through reasoning traces and code. The catch: the effect only appears when teacher and student share the same base model or a behaviorally matched one. Across families it disappears. So a Fable-to-Kimi transfer, if it happened, has only the visible-traces channel open. The spooky hidden channel is not available across architectures.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bill
&lt;/h2&gt;

&lt;p&gt;Distillation is not free, and the invoice arrives in three places.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Capacity gap.&lt;/strong&gt; Apple's distillation scaling law found the student's loss improves with teacher strength only up to a point, after which a stronger teacher produces a &lt;em&gt;worse&lt;/em&gt; student. The teacher becomes harder to model, and the student can no longer absorb the gains. Picking the strongest available teacher is not automatically correct.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Distribution mismatch.&lt;/strong&gt; Off-policy SFT trains the student to navigate states the teacher visits. At inference the student is in its own states, which it never trained for, and errors compound over long chains. This is the structural argument for on-policy distillation, where the student generates and the teacher scores each token. Thinking Machines reported roughly 70% on AIME'24 for 1,800 GPU-hours where their RL baseline needed 17,920, and framed the reason cleanly: RL delivers about O(1) bits of signal per episode, distillation delivers O(N) bits, one per token.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Forgetting.&lt;/strong&gt; A 2026 paper on post-training with knowledge retention reports that SFT on rejection-sampled Gemini 2.5 Pro responses dropped IFEval by 11.5 points on Llama-3.1-8B-Instruct, and even eroded in-domain reasoning on Qwen3-8B from 46.8% to 41.0%. You are not adding a skill to a static model. You are moving the whole distribution, and things fall out of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checklist: capability transfer or handwriting transplant?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Factor&lt;/th&gt;
&lt;th&gt;Handwriting only&lt;/th&gt;
&lt;th&gt;Capability actually moves&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Data scale&lt;/td&gt;
&lt;td&gt;1k to 20k traces&lt;/td&gt;
&lt;td&gt;100k+, rejection-sampled against verified answers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What you train&lt;/td&gt;
&lt;td&gt;Attention-only LoRA, 0.01% of params&lt;/td&gt;
&lt;td&gt;Full fine-tune, or LoRA over FFN/experts too&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Teacher gap&lt;/td&gt;
&lt;td&gt;As large as possible&lt;/td&gt;
&lt;td&gt;Inside the capacity-gap range for your student&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Domain&lt;/td&gt;
&lt;td&gt;Traces from one domain, eval in another&lt;/td&gt;
&lt;td&gt;Traces drawn from the distribution you'll deploy on&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Policy&lt;/td&gt;
&lt;td&gt;Off-policy SFT only&lt;/td&gt;
&lt;td&gt;On-policy phase after the cold start&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Eval&lt;/td&gt;
&lt;td&gt;pass@1, one pipeline, base's thinking mode never invoked&lt;/td&gt;
&lt;td&gt;pass@k, identical pipeline, thinking mode forced on both sides, OOD and instruction-following held out&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If your setup sits mostly in the left column, you did not get a frontier model. You got your base model that now reliably reasons out loud. Ship it if that is what you needed. Just do not put "matches Kimi K2.6" in the model card.&lt;/p&gt;

&lt;h2&gt;
  
  
  Back to Kimi K3
&lt;/h2&gt;

&lt;p&gt;Numbers first, because most coverage skipped them. In February 2026, Anthropic publicly attributed more than 16 million Claude exchanges across roughly 24,000 fraudulent accounts to coordinated campaigns by Moonshot, DeepSeek, and MiniMax, over 3.4 million of them pinned on Moonshot alone. In June, Anthropic told the US Senate Banking Committee that Alibaba's Qwen lab had run the largest known distillation attack to date. So industrial-scale trace harvesting is on the record as a specific, quantified allegation, months before Kratsios spoke.&lt;/p&gt;

&lt;p&gt;The Fable-to-K3 link is the shaky part. Per press reconstructions of the timeline, Fable 5 launched June 9, was suspended June 12 under export controls, and returned July 1; K3 launched July 16. That is the ~18 days from the intro. K3's pretraining necessarily predates all of it, so the only technically coherent version of the accusation is targeted post-training on Fable-derived traces over an already-trained base. Which is precisely the scenario this article is about: it would buy behavioral transfer in the domains the traces covered, at black-box bandwidth, from under three weeks of harvesting. It would not buy the pretraining run, and pretraining is the part nobody knows how to distill, because the parameter space is specific to each network.&lt;/p&gt;

&lt;p&gt;I wrote earlier about the &lt;a href="https://dev.to/p0rt/how-model-distillation-actually-works-and-what-the-china-distilled-our-model-headlines-really-3o0o"&gt;legal and ToS side of this&lt;/a&gt; and I want to correct my own emphasis: framing distillation primarily as an IP question makes it sound like the technique hands over a copy of the model. It does not. The interesting question was always mechanical, and the mechanical answer constrains the policy answer.&lt;/p&gt;

&lt;p&gt;Caveats on my own sources, since I am asking you to update on them: DeepSeek's numbers are self-reported, the Kimi-distill evals above are single-run and small-sample by the author's own admission, and the White House claim has no published evidence behind it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Question for anyone who has actually run one of these:&lt;/strong&gt; what is the one eval in your harness that would have caught a benchmark jump that was really just your base model's thinking mode failing to trigger? I suspect most of us do not have one, and that is why the Hugging Face leaderboard looks the way it does.&lt;/p&gt;




&lt;h2&gt;
  
  
  Sources &amp;amp; further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2501.12948" rel="noopener noreferrer"&gt;DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL&lt;/a&gt; (Tables 5 and 6, distillation vs RL)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2504.13837" rel="noopener noreferrer"&gt;Does RL Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?&lt;/a&gt; (pass@k; distillation expands the boundary, RLVR narrows it)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2502.07374" rel="noopener noreferrer"&gt;LLMs Can Easily Learn to Reason from Demonstrations: Structure, not content, is what matters!&lt;/a&gt; (the corruption experiments)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2305.15717" rel="noopener noreferrer"&gt;The False Promise of Imitating Proprietary LLMs&lt;/a&gt; (style vs factuality, 2023)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://machinelearning.apple.com/research/distillation-scaling-laws" rel="noopener noreferrer"&gt;Distillation Scaling Laws&lt;/a&gt; (capacity gap)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://thinkingmachines.ai/blog/on-policy-distillation/" rel="noopener noreferrer"&gt;On-Policy Distillation&lt;/a&gt; (Thinking Machines Lab)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2512.20908" rel="noopener noreferrer"&gt;Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.nature.com/articles/s41586-026-10319-8" rel="noopener noreferrer"&gt;Language models transmit behavioural traits through hidden signals in data&lt;/a&gt; (Nature; &lt;a href="https://arxiv.org/abs/2507.14805" rel="noopener noreferrer"&gt;preprint&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2603.01683" rel="noopener noreferrer"&gt;Surgical Post-Training: Proximal On-Policy Distillation with Knowledge Retention&lt;/a&gt; (forgetting numbers)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://huggingface.co/lordx64/Qwen3.6-35B-A3B-Kimi-K2.6-Reasoning-Distilled" rel="noopener noreferrer"&gt;Qwen3.6-35B-A3B-Kimi-K2.6-Reasoning-Distilled model card&lt;/a&gt; (the honest eval section)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2501.19393" rel="noopener noreferrer"&gt;s1: Simple test-time scaling&lt;/a&gt; (1,000 samples, 26 minutes on 16 H100s)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://openai.com/index/learning-to-reason-with-llms/" rel="noopener noreferrer"&gt;Learning to reason with LLMs&lt;/a&gt; (OpenAI on hiding o1's raw chain of thought)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2403.05530" rel="noopener noreferrer"&gt;Gemini 1.5 technical report&lt;/a&gt; (Flash is online-distilled from Pro)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.scmp.com/news/us/diplomacy/article/3361510/trump-tech-official-accuses-chinas-moonshot-ai-stealing-anthropic" rel="noopener noreferrer"&gt;Trump tech official accuses China's Moonshot AI of stealing from Anthropic&lt;/a&gt; (SCMP), &lt;a href="https://thenewstack.io/moonshot-fable5-distillation-accusations/" rel="noopener noreferrer"&gt;The New Stack&lt;/a&gt; (the 16M / 24,000 numbers), and &lt;a href="https://glitchwire.com/news/trump-administration-accuses-moonshot-ai-of-distilling-anthropics-fable-escalati/" rel="noopener noreferrer"&gt;Glitchwire&lt;/a&gt; (Moonshot's 3.4M, the June Senate claim about Qwen)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://kingy.ai/blog/kimi-k3-fable-5-distillation/" rel="noopener noreferrer"&gt;Kingy.ai's timeline reconstruction&lt;/a&gt; (Fable 5 availability windows vs K3's release)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://simonwillison.net/2026/Jul/16/kimi-k3/" rel="noopener noreferrer"&gt;Kimi K3 launch notes&lt;/a&gt; (Simon Willison)&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;If this article helped you, you can &lt;a href="https://ko-fi.com/sergeiparfenov?ref=dev" rel="noopener noreferrer"&gt;buy me a coffee&lt;/a&gt;. Your support helps me make time for the next experiment.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
    </item>
    <item>
      <title>I Found Two Bugs in Zulip. The Maintainers Had Filed Both Two Weeks Earlier.</title>
      <dc:creator>Sergei Parfenov</dc:creator>
      <pubDate>Fri, 07 Aug 2026 14:38:24 +0000</pubDate>
      <link>https://dev.to/p0rt/i-found-two-bugs-in-zulip-the-maintainers-had-filed-both-two-weeks-earlier-4mom</link>
      <guid>https://dev.to/p0rt/i-found-two-bugs-in-zulip-the-maintainers-had-filed-both-two-weeks-earlier-4mom</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for &lt;a href="https://dev.to/bugsmash"&gt;DEV's Summer Bug Smash: Smash Stories&lt;/a&gt; powered by &lt;a href="https://sentry.io/" rel="noopener noreferrer"&gt;Sentry&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I went hunting in Zulip's codebase for this challenge and found two real data-corruption bugs in the Slack importer. Solid ones: silent message scrambling during workspace migrations, the kind of bug that costs somebody a re-migration.&lt;/p&gt;

&lt;p&gt;Then I did the thing you are supposed to do before writing a single line of fix. I searched the tracker.&lt;/p&gt;

&lt;p&gt;Issue &lt;a href="https://github.com/zulip/zulip/issues/39650" rel="noopener noreferrer"&gt;#39650&lt;/a&gt;. Opened June 30, by the maintainers themselves: a systematic audit of the Slack import path, roughly a dozen items, confirmed and triaged. My two discoveries were on that list, described more precisely than I would have described them. &lt;strong&gt;I was two weeks late to my own findings.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is the story of what happened next, because "what happened next" turned out to be the useful part.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Hunted in Zulip → found 2 real bugs → both already filed by the maintainers (#39650) → the maintainer was already fixing them in an open PR (#39757) → pivoted twice: found an unreported latent twin of the bug class in the newest importer (Microsoft Teams), and claimed the one confirmed item the maintainer's PR did not touch (an unguarded timestamp sort key, which turned out to hide a silent NaN failure mode on top of the two loud ones). Shipped two PRs (&lt;a href="https://github.com/zulip/zulip/pull/39813" rel="noopener noreferrer"&gt;#39813&lt;/a&gt;, &lt;a href="https://github.com/zulip/zulip/pull/39814" rel="noopener noreferrer"&gt;#39814&lt;/a&gt;), each with a test that fails on the old code. Lessons at the end.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Picking the hunting ground
&lt;/h2&gt;

&lt;p&gt;Why Zulip: it is on this challenge's suggested repo list, it is Python, and it has a reputation that makes hunting interesting rather than easy. Near-total backend test coverage, strict mypy, ruthless linting, and a commit discipline of "each commit is a minimal coherent idea". I ran an extended ruff pass over the whole backend just to check the floor. Nothing but style noise. &lt;strong&gt;Lint-level bugs do not survive in this codebase.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Which leaves logic-level bugs, and for those you want a subsystem where correctness is hard and the inputs are hostile. Data import is exactly that: long batch jobs over &lt;em&gt;other tools'&lt;/em&gt; export files, arbitrary data quality, and no retry culture, because a workspace migration is a one-shot event for the admin running it. One unhandled edge case on message 31,000 of 50,000 and the whole thing is off.&lt;/p&gt;

&lt;h2&gt;
  
  
  The two bugs I "found"
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Bug one: thread state does not survive chunk boundaries.&lt;/strong&gt; Slack messages stream through the converter in chunks of 1,000. The map that routes thread replies to their Zulip topics was allocated &lt;em&gt;inside&lt;/em&gt; the per-chunk function. Thread root in chunk N, replies in chunk N+1: the replies land in an orphan topic literally named &lt;code&gt;"... No channel message"&lt;/code&gt;. The detail that makes it art: the &lt;em&gt;adjacent&lt;/em&gt; cache in the same file was deliberately made module-global, with a comment explaining that it must survive across calls. One cache got the cross-chunk treatment. Its sibling, added later, did not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bug two: the thread key is truncated to seconds.&lt;/strong&gt; A thread's identity was computed as &lt;code&gt;strftime("%Y/%m/%d %H:%M:%S")&lt;/code&gt; plus the parent's user id. No channel component. No microseconds, even though Slack's raw &lt;code&gt;ts&lt;/code&gt; carries them, as a string, right there in the message. Any bot that posts twice within one second and collects replies on both messages: two distinct threads merged into one topic, conversations interleaved.&lt;/p&gt;

&lt;p&gt;Both bugs are silent. No exceptions, no warnings, just a migrated archive that is quietly wrong.&lt;/p&gt;

&lt;p&gt;I was fairly pleased with myself for about as long as it took to search the tracker.&lt;/p&gt;

&lt;h2&gt;
  
  
  The tracker, and the etiquette call
&lt;/h2&gt;

&lt;p&gt;Finding #39650 stung, then got worse in an instructive way: &lt;a href="https://github.com/PieterCK" rel="noopener noreferrer"&gt;PieterCK&lt;/a&gt;, the maintainer who owns the importers, already had an open PR, &lt;a href="https://github.com/zulip/zulip/pull/39757" rel="noopener noreferrer"&gt;#39757&lt;/a&gt;, "slack_importer: Fix Slack thread conversion bugs". I opened its Files changed tab with a sinking feeling. Both of my thread bugs: being fixed, by the person whose code it is, with more context than I will ever have.&lt;/p&gt;

&lt;p&gt;This challenge has a section called "Smash Bugs, Respectfully", about not adding to maintainer workload. Racing a maintainer's open PR on their own audit items is a textbook way to fail it. So: &lt;strong&gt;stand down on the thread bugs.&lt;/strong&gt; Not negotiable, and honestly not even disappointing once framed correctly, because independently rediscovering two items from a maintainer audit is not wasted work. It is calibration. My nose was pointing at real bugs; it was just pointing at them second.&lt;/p&gt;

&lt;p&gt;If the audited file is picked clean, two moves remain: find what the audit &lt;em&gt;missed&lt;/em&gt; elsewhere, and find what the audit &lt;em&gt;found&lt;/em&gt; but nobody claimed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Move one: same disease, newer organ
&lt;/h2&gt;

&lt;p&gt;Audits cover files, but bug classes travel across files, carried by copy-paste and by shared habits. So I took the bug classes from the Slack audit and checked the sibling importers. Mattermost: clean on these classes, different batching helper. Microsoft Teams, the &lt;em&gt;newest&lt;/em&gt; importer in the family: jackpot.&lt;/p&gt;

&lt;p&gt;Its batching generator yielded a list and then called &lt;code&gt;.clear()&lt;/code&gt; on the same object to build the next batch. Run the minimal version yourself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;batched&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="c1"&gt;# expected: [[0..4], [5..9], [10, 11]]
# actual:   [[10, 11], [10, 11], [10, 11]]
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every yielded batch is one shared list. Ten of twelve messages gone, zero exceptions. The bug is latent: Zulip's current caller consumes each batch before advancing, so nothing fires today. It is still a broken contract, because the consumer owns a yielded value, and every future consumer inherits a trap that silently destroys data during a one-shot migration. Nobody had reported it. That became PR &lt;a href="https://github.com/zulip/zulip/pull/39814" rel="noopener noreferrer"&gt;#39814&lt;/a&gt;, and the "why fix a latent bug" argument became [its own submission][LINK: Teams Clear the Lineup post].&lt;/p&gt;

&lt;h2&gt;
  
  
  Move two: the confirmed item nobody took
&lt;/h2&gt;

&lt;p&gt;Back in #39650, one high-impact robustness item sat confirmed and unclaimed: the unguarded timestamp sort key. &lt;code&gt;float(message["ts"])&lt;/code&gt;, used to sort every message of the export and as &lt;code&gt;date_sent&lt;/code&gt;. I checked the maintainer's open PR for it, Ctrl+F in Files changed: &lt;strong&gt;not found.&lt;/strong&gt; Free.&lt;/p&gt;

&lt;p&gt;The item as filed had two failure modes: missing &lt;code&gt;ts&lt;/code&gt; raises &lt;code&gt;KeyError&lt;/code&gt;, garbage &lt;code&gt;ts&lt;/code&gt; raises &lt;code&gt;ValueError&lt;/code&gt;, either one aborts an entire import because of one message. Writing the guard surfaced a third mode that nobody had listed, and it is the best souvenir of this whole hunt: &lt;code&gt;"ts": "NaN"&lt;/code&gt;. &lt;code&gt;float("NaN")&lt;/code&gt; parses without complaint. NaN compares as False against everything, which quietly violates the total ordering Timsort assumes, so &lt;code&gt;sorted()&lt;/code&gt; returns an inconsistent order. No crash, no warning, scrambled chronology. &lt;strong&gt;The crash is the lucky failure mode.&lt;/strong&gt; That is why the shipped guard requires &lt;code&gt;math.isfinite&lt;/code&gt;, not just a successful parse. This became PR &lt;a href="https://github.com/zulip/zulip/pull/39813" rel="noopener noreferrer"&gt;#39813&lt;/a&gt; and [its own submission][LINK: Slack Clear the Lineup post].&lt;/p&gt;

&lt;h2&gt;
  
  
  Boring discipline, on purpose
&lt;/h2&gt;

&lt;p&gt;Both fixes went through the same routine. Write the regression test first, roll the source back to &lt;code&gt;upstream/main&lt;/code&gt;, and watch the test fail with the exact predicted error: &lt;code&gt;KeyError: 'ts'&lt;/code&gt; for the Slack guard, &lt;code&gt;AssertionError: 24 != 29&lt;/code&gt; for the Teams aliasing (the batches measurably eat messages). Then restore the fix and run everything: 56/56 on the Slack importer module, 10/10 on Teams, lint and mypy clean, coverage showing no uncovered lines in the touched Slack module. A regression test that never failed against the old code is a test fitted to the fix, not a test of the fix.&lt;/p&gt;

&lt;p&gt;Full disclosure on process: the mechanical part of this hunt (cloning, grepping, running suites) went through AI agent tooling; every call you have read about, from which repo to whose bug not to race to which failure policy to ship, stayed human. Zulip has an explicit AI-use policy for contributions, and both PRs follow it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Epilogue: arguing with the robot
&lt;/h2&gt;

&lt;p&gt;For the other track's submission I wired the crash into Sentry and asked Seer, its AI debugger, for a root cause. Credit where due: Seer nailed the diagnosis in seconds, down to quoting the exact poisoned message it pulled from the frame locals. Then it suggested a fix: &lt;code&gt;float(message.get("ts", 0))&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That one-liner is the fallback-timestamp option I had already declined in the PR. It patches the missing-&lt;code&gt;ts&lt;/code&gt; mode, leaves &lt;code&gt;ValueError&lt;/code&gt; alive, waves &lt;code&gt;"NaN"&lt;/code&gt; straight through into the sort, and stamps real messages with a 1970 &lt;code&gt;date_sent&lt;/code&gt;. One failure mode out of three, plus fabricated chronology. The tool found the root cause faster than I would have; deciding the failure &lt;em&gt;policy&lt;/em&gt; was still my job. I suspect that division of labor is going to describe a lot of debugging from here on.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I am taking away
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Freshly audited ground is picked clean. Hunt where the code is newest.&lt;/strong&gt; The Slack importer had a dozen filed bugs and zero available ones; the newest importer had an unreported twin waiting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rediscovery is calibration, not waste.&lt;/strong&gt; Going two-for-two against a maintainer audit told me the method works; it just needs to run earlier or elsewhere.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Filed" is not "fixed".&lt;/strong&gt; A confirmed, high-impact item sat unclaimed for over three weeks in one of the best-maintained Python codebases around. Trackers are full of these.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read the open PRs before you race them.&lt;/strong&gt; The most useful contribution I made to the thread bugs was not making one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The crash is the lucky failure mode.&lt;/strong&gt; The loud errors were filed by an audit; the silent NaN mode was not. The bugs that skip the crash are the ones that make it to production, and to your archives.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Have you ever arrived two weeks late to your own discovery? And did you file it under wasted effort, or under proof your nose works? I have firmly moved to the second column.&lt;/p&gt;

&lt;h2&gt;
  
  
  The paper trail
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The maintainers' audit: &lt;a href="https://github.com/zulip/zulip/issues/39650" rel="noopener noreferrer"&gt;https://github.com/zulip/zulip/issues/39650&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;The maintainer's thread-fix PR: &lt;a href="https://github.com/zulip/zulip/pull/39757" rel="noopener noreferrer"&gt;https://github.com/zulip/zulip/pull/39757&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;My Slack timestamp guard: &lt;a href="https://github.com/zulip/zulip/pull/39813" rel="noopener noreferrer"&gt;https://github.com/zulip/zulip/pull/39813&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;My Teams batching fix: &lt;a href="https://github.com/zulip/zulip/pull/39814" rel="noopener noreferrer"&gt;https://github.com/zulip/zulip/pull/39814&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;The challenge's "Smash Bugs, Respectfully" guide: &lt;a href="https://dev.to/opensourcepledge/how-to-respectfully-contribute-to-open-source-cbh"&gt;https://dev.to/opensourcepledge/how-to-respectfully-contribute-to-open-source-cbh&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;If this article helped you, you can &lt;a href="https://ko-fi.com/sergeiparfenov?ref=dev" rel="noopener noreferrer"&gt;buy me a coffee&lt;/a&gt;. Your support helps me make time for the next experiment.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>bugsmash</category>
      <category>python</category>
      <category>opensource</category>
    </item>
    <item>
      <title>This Bug Has Never Fired in Production. I Fixed It Anyway</title>
      <dc:creator>Sergei Parfenov</dc:creator>
      <pubDate>Wed, 05 Aug 2026 10:23:33 +0000</pubDate>
      <link>https://dev.to/p0rt/this-bug-has-never-fired-in-production-i-fixed-it-anyway-ga6</link>
      <guid>https://dev.to/p0rt/this-bug-has-never-fired-in-production-i-fixed-it-anyway-ga6</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for &lt;a href="https://dev.to/bugsmash"&gt;DEV's Summer Bug Smash: Clear the Lineup&lt;/a&gt; powered by &lt;a href="https://sentry.io/" rel="noopener noreferrer"&gt;Sentry&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Run this and predict the output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;batched&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;chunk_size&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;batch&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;batch&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;chunk_size&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="n"&gt;batch&lt;/span&gt;
            &lt;span class="n"&gt;batch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;clear&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="n"&gt;batch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;batch&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="n"&gt;batch&lt;/span&gt;

&lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;batched&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Expected: &lt;code&gt;[[0..4], [5..9], [10, 11]]&lt;/code&gt;. Actual:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="p"&gt;[[&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;11&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;11&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;11&lt;/span&gt;&lt;span class="p"&gt;]]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All three elements are &lt;strong&gt;the same list object&lt;/strong&gt;. Ten of twelve messages are gone, no exception raised. This exact pattern was sitting in Zulip's Microsoft Teams importer, and the interesting part is not the fix (one line). It is that the bug has never fired in production, and I think that made it &lt;em&gt;more&lt;/em&gt; worth fixing, not less.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; &lt;code&gt;get_batched_export_message_data()&lt;/code&gt; in Zulip's Teams importer yielded a batch list, then cleared and refilled the same object for the next batch. Safe for the current lazy consumer, silently destructive for any consumer that retains batches (starting with &lt;code&gt;list(...)&lt;/code&gt;). My fix (&lt;a href="https://github.com/zulip/zulip/pull/39814" rel="noopener noreferrer"&gt;zulip/zulip#39814&lt;/a&gt;) hands ownership of each yielded list to the consumer, plus a regression test that materializes the generator and fails on the old code with a concrete data-loss assertion: &lt;code&gt;24 != 29&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Project Overview
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/zulip/zulip" rel="noopener noreferrer"&gt;Zulip&lt;/a&gt; is an open-source team chat server (Django/Python, ~25k stars) with a famously strict engineering bar: near-total backend coverage, strict mypy, "each commit is a minimal coherent idea".&lt;/p&gt;

&lt;p&gt;The Microsoft Teams importer in &lt;code&gt;zerver/data_import/&lt;/code&gt; is the newest member of Zulip's data-import family: it converts Teams export files into Zulip's format, reading messages and yielding them in batches of &lt;code&gt;chunk_size&lt;/code&gt; for downstream processing. Newest code, fewest eyes: that is exactly where I went hunting, after discovering that the older Slack importer had just been through a full maintainer audit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bug Fix or Performance Improvement
&lt;/h2&gt;

&lt;p&gt;The old batching generator:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;batched_messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;MicrosoftTeamsFieldsT&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;path&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;message_data_paths&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;get_data_file&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;sorted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])):&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;batched_messages&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;chunk_size&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="n"&gt;batched_messages&lt;/span&gt;
            &lt;span class="n"&gt;batched_messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;clear&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;   &lt;span class="c1"&gt;# &amp;lt;-- the bug
&lt;/span&gt;        &lt;span class="n"&gt;batched_messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;batched_messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="n"&gt;batched_messages&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;yield&lt;/code&gt; suspends the generator, the consumer processes the batch, and on the next &lt;code&gt;next()&lt;/code&gt; the generator wakes up and calls &lt;code&gt;.clear()&lt;/code&gt; on &lt;strong&gt;the object it just handed out&lt;/strong&gt;. Every yielded batch is one shared list in memory.&lt;/p&gt;

&lt;p&gt;This is a violation of an implicit contract: &lt;strong&gt;the consumer owns a yielded value.&lt;/strong&gt; As long as consumption is strictly lazy (process each batch fully before advancing) nothing visible happens, which is why the current Zulip caller works. But the moment any consumer &lt;em&gt;retains&lt;/em&gt; batches, the simplest being &lt;code&gt;list(generator)&lt;/code&gt;, it gets N references to a single object containing only the final batch's contents. Data is not lost loudly; it is silently replaced. In the importer's test dataset, the total message count collapses from 29 to 24 with zero exceptions.&lt;/p&gt;

&lt;p&gt;The standard library agrees on the contract, by the way: &lt;code&gt;itertools.batched&lt;/code&gt; allocates a fresh tuple per batch. So does every batching recipe in the itertools docs. &lt;code&gt;clear()&lt;/code&gt;-and-refill looks like a memory optimization; what it actually optimizes away is correctness under any future change to the consumer: buffering, retries, parallel prefetch, or a colleague writing &lt;code&gt;list(...)&lt;/code&gt; in a test.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;PR: &lt;strong&gt;&lt;a href="https://github.com/zulip/zulip/pull/39814" rel="noopener noreferrer"&gt;zulip/zulip#39814&lt;/a&gt;&lt;/strong&gt; (branch &lt;code&gt;P0rt:teams-batch-aliasing&lt;/code&gt;, single commit, +21/-1).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;batched_messages&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;chunk_size&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="n"&gt;batched_messages&lt;/span&gt;
    &lt;span class="c1"&gt;# Start a new list rather than clearing the yielded one;
&lt;/span&gt;    &lt;span class="c1"&gt;# the consumer owns the yielded list, and mutating it here
&lt;/span&gt;    &lt;span class="c1"&gt;# would corrupt every batch if the generator is materialized
&lt;/span&gt;    &lt;span class="c1"&gt;# (e.g. with list(...)) or batches are retained across
&lt;/span&gt;    &lt;span class="c1"&gt;# iterations.
&lt;/span&gt;    &lt;span class="n"&gt;batched_messages&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One line of code, six lines of comment. That ratio is deliberate: the fix is trivial, but the &lt;em&gt;reason&lt;/em&gt; it must stay this way is not visible from the code, and the whole failure mode exists because a past "optimization" looked equally trivial.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Improvements
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The test asserts the contract, not the current caller.&lt;/strong&gt; I extended the existing &lt;code&gt;test_get_batched_export_message_data&lt;/code&gt; to materialize the generator with &lt;code&gt;list(...)&lt;/code&gt; and then check two things: the total message count across all batches, and the exact flattened sequence of message IDs against the sorted source files. That second assertion is the important one; it catches loss, duplication, and reordering in a single check, so the test defends the ownership contract rather than one symptom.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Red phase before green phase.&lt;/strong&gt; With the source file rolled back to &lt;code&gt;upstream/main&lt;/code&gt; and the new test kept, the run fails with &lt;code&gt;AssertionError: 24 != 29&lt;/code&gt;: the aliasing measurably eats messages. With the fix restored, the full module passes: &lt;code&gt;./tools/test-backend zerver.tests.test_microsoft_teams_importer&lt;/code&gt;, 10/10. &lt;code&gt;./tools/lint&lt;/code&gt; and &lt;code&gt;./tools/run-mypy&lt;/code&gt; clean on the branch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why fix a bug that never fires?&lt;/strong&gt; Because "latent" describes the caller, not the function. The function's contract is broken today; the current caller just happens not to lean on the broken part. Every future consumer inherits a trap that produces silent data corruption during a one-shot, high-stakes operation (a workspace migration is not a request you can retry). The cost of the fix is one allocation per batch. The cost of not fixing it is a debugging session that starts with "the import succeeded but a fifth of the messages are missing", which is close to the worst bug report a data migration can generate.&lt;/p&gt;




&lt;p&gt;I keep coming back to one framing from this fix: &lt;strong&gt;tests that only cover your current callers are tests of an implementation; tests that cover the contract are tests of the function.&lt;/strong&gt; The first kind rots the moment anyone new calls your code. Where do you draw that line in your own test suites: do you test what the function promises, or what today's callers happen to need?&lt;/p&gt;




&lt;p&gt;If this article helped you, you can &lt;a href="https://ko-fi.com/sergeiparfenov?ref=dev" rel="noopener noreferrer"&gt;buy me a coffee&lt;/a&gt;. Your support helps me make time for the next experiment.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>bugsmash</category>
      <category>python</category>
      <category>opensource</category>
    </item>
    <item>
      <title>The Bug That Crashes Your Import Is the Lucky One</title>
      <dc:creator>Sergei Parfenov</dc:creator>
      <pubDate>Fri, 31 Jul 2026 12:43:45 +0000</pubDate>
      <link>https://dev.to/p0rt/the-bug-that-crashes-your-import-is-the-lucky-one-25of</link>
      <guid>https://dev.to/p0rt/the-bug-that-crashes-your-import-is-the-lucky-one-25of</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for &lt;a href="https://dev.to/bugsmash"&gt;DEV's Summer Bug Smash: Clear the Lineup&lt;/a&gt; powered by &lt;a href="https://sentry.io/" rel="noopener noreferrer"&gt;Sentry&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;You are migrating a 50,000-message Slack workspace to Zulip. Somewhere around message 31,000 the import dies with &lt;code&gt;KeyError: 'ts'&lt;/code&gt;. Annoying, but here is the uncomfortable part: &lt;strong&gt;that is the lucky outcome.&lt;/strong&gt; The unlucky one is &lt;code&gt;"ts": "NaN"&lt;/code&gt;, where nothing dies, nothing warns, and your company's message history quietly comes out in the wrong order.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Zulip's Slack importer used &lt;code&gt;float(message["ts"])&lt;/code&gt; unguarded, both as a sort key and as &lt;code&gt;date_sent&lt;/code&gt;. One message with a missing or malformed &lt;code&gt;ts&lt;/code&gt; aborted the entire import; a non-finite value like &lt;code&gt;"NaN"&lt;/code&gt; did not even raise, it silently broke the sort. My fix (&lt;a href="https://github.com/zulip/zulip/pull/39813" rel="noopener noreferrer"&gt;zulip/zulip#39813&lt;/a&gt;) skips such messages with a warning and requires &lt;code&gt;ts&lt;/code&gt; to parse to a &lt;em&gt;finite&lt;/em&gt; float via &lt;code&gt;math.isfinite&lt;/code&gt;. The regression test fails with &lt;code&gt;KeyError: 'ts'&lt;/code&gt; on the old code.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Project Overview
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/zulip/zulip" rel="noopener noreferrer"&gt;Zulip&lt;/a&gt; is an open-source team chat server (Django/Python, ~25k stars) with an unusually strict engineering culture: near-total backend test coverage, strict mypy, and a commit discipline of "each commit is a minimal coherent idea".&lt;/p&gt;

&lt;p&gt;The code I touched lives in &lt;code&gt;zerver/data_import/&lt;/code&gt;: the subsystem that converts exports from Slack, Microsoft Teams, and Mattermost into Zulip's format. This subsystem has one property that should shape every line in it: &lt;strong&gt;the input is another tool's output.&lt;/strong&gt; Import is a long batch process over data of arbitrary quality, and the admin running the migration has no way to "fix" what Slack's export tool produced. A pipeline that dies on record 31,207 of 50,000 is strictly worse than one that skips record 31,207 with a warning.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bug Fix or Performance Improvement
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;get_messages_iterator()&lt;/code&gt; in &lt;code&gt;zerver/data_import/slack.py&lt;/code&gt; streams every message of the export, sorting each day's messages by timestamp:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="nf"&gt;sorted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;messages_for_one_day&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;get_timestamp_from_message&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;where the sort key was simply:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_timestamp_from_message&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;ZerverFieldsT&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That one line has &lt;strong&gt;three distinct failure modes&lt;/strong&gt;, all reproduced during validation:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. &lt;code&gt;ts&lt;/code&gt; is missing.&lt;/strong&gt; &lt;code&gt;KeyError: 'ts'&lt;/code&gt; straight out of &lt;code&gt;sorted(...)&lt;/code&gt;. The whole import aborts because of one message:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;File&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;zerver/data_import/slack.py&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt; &lt;span class="mi"&gt;911&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;get_messages_iterator&lt;/span&gt;
  &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="nf"&gt;sorted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;messages_for_one_day&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;get_timestamp_from_message&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;File&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;zerver/data_import/slack.py&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt; &lt;span class="mi"&gt;1461&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;get_timestamp_from_message&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="nb"&gt;KeyError&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;ts&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;2. &lt;code&gt;ts&lt;/code&gt; is garbage.&lt;/strong&gt; &lt;code&gt;"not-a-number"&lt;/code&gt; raises &lt;code&gt;ValueError&lt;/code&gt;. Same total abort.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. &lt;code&gt;ts&lt;/code&gt; is &lt;code&gt;"NaN"&lt;/code&gt;.&lt;/strong&gt; The nasty one. &lt;code&gt;float("NaN")&lt;/code&gt; is a perfectly valid parse, so nothing raises. But NaN is incomparable (&lt;code&gt;NaN &amp;lt; x&lt;/code&gt; and &lt;code&gt;x &amp;lt; NaN&lt;/code&gt; are both False), which violates the total ordering Timsort assumes, so &lt;code&gt;sorted()&lt;/code&gt; silently returns an inconsistent order:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;sorted&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;1434139102.000002&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;NaN&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;1434139101.000001&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# =&amp;gt; ['1434139102.000002', 'NaN', '1434139101.000001']   # not sorted, no error
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No exception, no warning. Just a migrated archive with scrambled chronology. Crashing is the lucky case; this is the case that costs you a re-migration three weeks later when someone notices the history reads wrong.&lt;/p&gt;

&lt;p&gt;One honesty note: I did not discover this failure class from zero. The Zulip maintainers' own audit issue (&lt;a href="https://github.com/zulip/zulip/issues/39650" rel="noopener noreferrer"&gt;#39650&lt;/a&gt;) flags the unguarded timestamp sort key as one of the two highest-impact robustness items, and it was unclaimed when I picked it up. What I brought is the implementation, the non-finite analysis, and the test that pins the behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;PR: &lt;strong&gt;&lt;a href="https://github.com/zulip/zulip/pull/39813" rel="noopener noreferrer"&gt;zulip/zulip#39813&lt;/a&gt;&lt;/strong&gt; (branch &lt;code&gt;P0rt:slack-ts-guard&lt;/code&gt;, single commit, +54/-0 across the module and its test file).&lt;/p&gt;

&lt;p&gt;The fix is a predicate plus a guard, following the skip-with-warning pattern that already exists two lines above for other unimportable messages:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;message_has_valid_timestamp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;ZerverFieldsT&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Whether `ts` is present and parses to a finite float.

    `get_timestamp_from_message` is used both as a sort key and for
    `date_sent`, so a missing or malformed `ts` on a single message
    would otherwise abort the entire import — and a non-finite value
    like &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;NaN&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; would not even raise, instead silently producing an
    inconsistent sort order.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isfinite&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]))&lt;/span&gt;
    &lt;span class="nf"&gt;except &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;KeyError&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;TypeError&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="nf"&gt;message_has_valid_timestamp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;logging&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;warning&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Skipping Slack message with invalid ts %r in %s&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;message_dir&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;continue&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The detail that matters: &lt;code&gt;math.isfinite&lt;/code&gt;, not a bare try/except around &lt;code&gt;float()&lt;/code&gt;. A naive "does it parse" check would have fixed the two loud failure modes and waved the silent one straight through.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Improvements
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The test had to fail first.&lt;/strong&gt; I wrote &lt;code&gt;test_get_messages_iterator_skips_invalid_timestamps&lt;/code&gt;: a temp directory with one day of Slack export containing five messages, two valid (deliberately in reverse chronological order), one with no &lt;code&gt;ts&lt;/code&gt;, one with &lt;code&gt;"not-a-number"&lt;/code&gt;, one with &lt;code&gt;"NaN"&lt;/code&gt;. Then I rolled the source file back to &lt;code&gt;upstream/main&lt;/code&gt;, kept the test, and ran it: &lt;code&gt;KeyError: 'ts'&lt;/code&gt;, exactly the predicted crash. Restored the fix: the test asserts that exactly the two valid messages survive, in correct order, with exactly three warnings logged. A regression test that never failed against the old code is a test fitted to the fix, not a test of the fix.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Full-suite validation in a provisioned dev environment&lt;/strong&gt;, not just the happy path: &lt;code&gt;./tools/test-backend zerver.tests.test_slack_importer&lt;/code&gt; (56/56 passing), &lt;code&gt;./tools/lint&lt;/code&gt; and &lt;code&gt;./tools/run-mypy&lt;/code&gt; clean, and &lt;code&gt;test-backend --coverage&lt;/code&gt; showing zero uncovered lines in &lt;code&gt;zerver/data_import/slack.py&lt;/code&gt; after the change, including the new guard path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One decision I explicitly left open for reviewers:&lt;/strong&gt; skipping a message with a broken &lt;code&gt;ts&lt;/code&gt; versus synthesizing a fallback timestamp for it. Skipping loses the message but keeps the archive honest; a synthetic timestamp keeps the message but fabricates chronology. I went with skip-plus-warning because it matches the file's existing conventions, and flagged the tradeoff in the PR instead of pretending it does not exist.&lt;/p&gt;

&lt;h2&gt;
  
  
  Best Use of Sentry
&lt;/h2&gt;

&lt;p&gt;I instrumented the demo with &lt;code&gt;sentry-sdk&lt;/code&gt; (&lt;code&gt;traces_sample_rate=1.0&lt;/code&gt;, environment &lt;code&gt;bugsmash-demo&lt;/code&gt;) and ran the exact same poisoned export twice: once with &lt;code&gt;zerver/data_import/slack.py&lt;/code&gt; rolled back to &lt;code&gt;upstream/main&lt;/code&gt;, once with the fix in place. Same branch, same data, one file different.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Before the fix.&lt;/strong&gt; The import span dies with &lt;code&gt;KeyError: 'ts'&lt;/code&gt;, attached right on the trace, status &lt;code&gt;internal_error&lt;/code&gt;:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx8pcqvur59p1mf059wvh.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx8pcqvur59p1mf059wvh.jpeg" alt="Sentry trace of the failing import: KeyError 'ts' on the slack_import_demo span" width="800" height="537"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Drilling into the issue gives the full stacktrace, pointing exactly where this PR points: &lt;code&gt;get_messages_iterator&lt;/code&gt; -&amp;gt; &lt;code&gt;get_timestamp_from_message&lt;/code&gt;:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fakvf3l2ahf92zqrzoifx.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fakvf3l2ahf92zqrzoifx.jpeg" alt="Sentry issue PYTHON-DJANGO-2: KeyError 'ts' with the stacktrace into get_messages_iterator" width="800" height="537"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;And Seer's root-cause analysis of that issue. The diagnosis is spot on, down to quoting the exact poisoned message it pulled from the frame locals (&lt;code&gt;{"channel_name": "general", "text": "no ts field"}&lt;/code&gt;):&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzlcwa0dmswm1bal58ejc.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzlcwa0dmswm1bal58ejc.jpeg" alt="Seer agent root cause and suggested fix for the KeyError issue" width="800" height="537"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One detail worth being honest about, precisely because a Sentry engineer judges this category: Seer's suggested one-liner, &lt;code&gt;float(message.get("ts", 0))&lt;/code&gt;, is the fallback-timestamp option I explicitly declined in the PR. It fixes the missing-&lt;code&gt;ts&lt;/code&gt; mode, but leaves the &lt;code&gt;ValueError&lt;/code&gt; mode alive, waves &lt;code&gt;"NaN"&lt;/code&gt; straight through into the sort, and stamps real messages with a 1970 &lt;code&gt;date_sent&lt;/code&gt;. Its closing advice, though (log a warning and investigate the upstream data quality), is exactly what the shipped fix does. Reading Seer's output critically instead of pasting its patch is, I would argue, the intended use of the tool: it found the root cause in seconds; deciding the failure policy stayed my job.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;After the fix.&lt;/strong&gt; Same export, same transaction: zero issues, and the three skipped messages surface as three warning-level logs on the trace instead of one fatal:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fppcp8awxhl2qm1ioxn00.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fppcp8awxhl2qm1ioxn00.jpeg" alt="Sentry trace of the fixed import: 0 issues, 3 warning logs, import completed" width="800" height="537"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That pair of traces is the whole fix, told by monitoring: the same bad input, downgraded from an unhandled exception that kills a migration to three structured warnings on a completed run. (A pedantic footnote for anyone reading the trace metadata: both runs report release &lt;code&gt;c0b5d8c&lt;/code&gt; because the red run rolled back only the module file, not the branch HEAD.)&lt;/p&gt;




&lt;p&gt;The uncomfortable takeaway from failure mode 3: &lt;strong&gt;"does it crash" is a terrible proxy for "is it correct"&lt;/strong&gt;, and the failure modes that skip the crash are the ones that survive into production. Which raises the question I left for the reviewers, and now for you: for a message that is otherwise fine but has a broken timestamp, would you skip it or synthesize a fallback &lt;code&gt;date_sent&lt;/code&gt;? I picked skip. Convince me otherwise in the comments.&lt;/p&gt;




&lt;p&gt;If this article helped you, you can &lt;a href="https://ko-fi.com/sergeiparfenov?ref=dev" rel="noopener noreferrer"&gt;buy me a coffee&lt;/a&gt;. Your support helps me make time for the next experiment.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>bugsmash</category>
      <category>python</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Nothing Was Broken. The Report Still Didn't Arrive.</title>
      <dc:creator>Sergei Parfenov</dc:creator>
      <pubDate>Sun, 26 Jul 2026 13:50:33 +0000</pubDate>
      <link>https://dev.to/p0rt/nothing-was-broken-the-report-still-didnt-arrive-k29</link>
      <guid>https://dev.to/p0rt/nothing-was-broken-the-report-still-didnt-arrive-k29</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for &lt;a href="https://dev.to/bugsmash"&gt;DEV's Summer Bug Smash: Clear the Lineup&lt;/a&gt; powered by &lt;a href="https://sentry.io/" rel="noopener noreferrer"&gt;Sentry&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;On July 24 at 09:00 Berlin time, the doctors group did not get its daily digest. It got the first half of one, unfinished. Nobody noticed.&lt;/p&gt;

&lt;p&gt;The failure was recorded correctly, in a state file that nothing reads. The job's delivery mode was &lt;code&gt;none&lt;/code&gt;. The fallback alert channel had been dead since June 13, for reasons that turn out to be bug four. So the run failed, the record was written, and the information stopped there.&lt;/p&gt;

&lt;p&gt;Here is how our team actually learns that something broke. Different job, same pipeline:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4bti84sqh0ssrzaxztbk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4bti84sqh0ssrzaxztbk.png" alt="Telegram message from the bot reporting a failed cron job, with a doctor replying and tagging the CTO" width="799" height="410"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The failure the team actually saw: a doctor's scheduled job died and she had to tag the CTO in chat. This is what «silent» pipelines look like when they finally speak.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Some context, since everything below assumes it. Symptomato is a telehealth service: patients describe symptoms in a chat, doctors answer them from a helpdesk. Sympy is the agent that works the seam between the two — on a schedule it reads the doctors' inbox and tells them, in Telegram, which conversations need a human today. Nobody watches it run. That is the point of it, and it is also why a broken run can go unnoticed for six weeks.&lt;/p&gt;

&lt;p&gt;So I went looking for the bug behind the missing digest. That is the uncomfortable part: there wasn't one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; an agent run failed with no defective component anywhere in the path. That non-incident triggered a full audit of the pipeline, which turned up four bugs in our own code that nobody had ever seen fail. All four have the same shape: a function that does not know something reports a value instead of admitting it. This post is those four fixes, plus the instrumentation that makes this failure class visible, plus a deliberate decision to record none of the message content while doing it.&lt;/p&gt;
&lt;h2&gt;
  
  
  Project Overview
&lt;/h2&gt;

&lt;p&gt;Under the hood: a self-hosted agent runtime with a job scheduler, a tool plugin we wrote on top of &lt;a href="https://www.chatwoot.com/" rel="noopener noreferrer"&gt;Chatwoot&lt;/a&gt; (the helpdesk our doctors work in), a set of host-side Python cron scripts, and Telegram as the delivery channel. Three scheduled jobs do the boring work: a morning digest of conversations that need attention, pings when a paid consultation goes 24 hours without a doctor reply, inbox prechecks.&lt;/p&gt;

&lt;p&gt;One constraint shapes everything below: patient conversations contain medical text. Any observability we add has to work without recording message bodies.&lt;/p&gt;
&lt;h2&gt;
  
  
  Bug Fix or Performance Improvement
&lt;/h2&gt;
&lt;h3&gt;
  
  
  The incident with no bug in it
&lt;/h3&gt;

&lt;p&gt;The digest agent sent its first Telegram message, then decided to finalize the digest by editing that message instead of sending a second one. The message tool requires a recipient field even for an edit. The edit call did not have one. The tool rejected it, the turn ended, the scheduler marked the run as &lt;code&gt;error&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Walk that path again and look for the defect. The tool validated its input exactly as its schema says. The scheduler recorded the failure exactly as designed. The model picked a legal tool sequence that the prompt never forbade. Every component behaved to spec, and the doctors still had no digest.&lt;/p&gt;

&lt;p&gt;This is the failure mode that makes agent pipelines different: the execution path is chosen at runtime by a model, so "correct components" and "correct behavior" stop being the same claim. Yesterday the same job sent one message and worked. Today it chose send-then-edit and did not. There is no line of code you can point at, and no test that fails, because nothing is deterministic enough to fail.&lt;/p&gt;

&lt;p&gt;Once instrumented, the same failure looks like this. The capture is from the replay run — the July 24 failure itself predates the instrumentation — and it needed no code changes at the failure site:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4kteyjglaz75wbq0fz3g.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4kteyjglaz75wbq0fz3g.png" alt="Sentry issue detail showing ToolInputError with raw params and a breadcrumb trail ending in a Telegram API call" width="799" height="563"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The anchor bug, captured automatically: the model chose «send, then edit», and the edit call has no &lt;code&gt;to&lt;/code&gt;. The tool that failed lives in the runtime, not in our code — we never instrumented it; the error was scraped from the gateway's ERROR log lines and turned into an issue.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcinm2e833uayio13gru2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcinm2e833uayio13gru2.png" alt="Sentry Seer panel reconstructing the root cause of the issue from breadcrumbs" width="799" height="563"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Seer reconstructs the failure from breadcrumbs alone — no prompts, no message bodies were ever recorded.&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  The four bugs the audit did find
&lt;/h3&gt;

&lt;p&gt;A missing digest with no defect in it is a bad place to stop, so I audited the pipeline properly: the tool plugin, the scheduler jobs, the host scripts, the existing telemetry extensions. It came back with a list. Four items on that list turned out to be the same bug wearing different clothes, and I only saw the pattern once I wrote them down next to each other.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. The tool reports a conversation status it never looked up.&lt;/strong&gt; &lt;code&gt;getConversationSummary&lt;/code&gt; returns &lt;code&gt;status: "open"&lt;/code&gt; as a literal, for every conversation, always. The real status sits in the API response one method below. So when the agent asks "is this conversation still open", it is told yes, unconditionally. Every judgment the model makes about whether to act on a conversation rests on a constant.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The tool returns page one and calls it the inbox.&lt;/strong&gt; &lt;code&gt;listConversationsByInbox&lt;/code&gt; fetches &lt;code&gt;page=1&lt;/code&gt; and returns the payload with no &lt;code&gt;all_count&lt;/code&gt; check and no truncation marker. Chatwoot paginates at 25. Our host-side Python script got this right and loops until the count matches; the agent tool never did. So on any day with more than 25 open conversations, the model receives 25 and narrates them as the full picture, which is exactly the sort of confident summary you cannot catch by reading the output.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. "The data isn't there yet" gets cached as a fact.&lt;/strong&gt; &lt;code&gt;tariff_from_triage_note&lt;/code&gt; returns &lt;code&gt;"unknown"&lt;/code&gt; for two very different situations: the note is unparseable, and the note has not been posted yet. The caller caches the result unconditionally, and &lt;code&gt;"unknown"&lt;/code&gt; is truthy, so the next run short-circuits on the cache and never looks again. A patient whose chat is scanned in the window before the backend posts its triage note is marked &lt;code&gt;unknown&lt;/code&gt; permanently, which means the 24-hour ping that exists specifically for paid one-time consultations never fires for them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. The error handler reports failures through the channel that just failed.&lt;/strong&gt; Our host digest script calls a Telegram helper with no error handling, catches the resulting exception, and then tries to report that exception with the same helper, which raises again. Since June 13 the script has ended in a double traceback — 128 of them in the log — and every one of those runs paid for a model call before crashing. Next to it, the monitor script ends its send with &lt;code&gt;&amp;gt; /dev/null 2&amp;gt;&amp;amp;1&lt;/code&gt;, so a failed alert is indistinguishable from a delivered one anywhere in the system.&lt;/p&gt;

&lt;p&gt;Now the shape. In every one of these, something the system does not know is represented as something it does know. Unknown status becomes "open". Twenty-five of thirty becomes "the inbox". A missing note becomes a tariff value, cached forever. A failed alert becomes a successful one. Three of the four never raise at all; the fourth raises twice a day into a log nobody reads, which comes to the same thing. All of them produce plausible output. And the incident that started the audit is the same thing one level up: a failed run represented as nothing at all.&lt;/p&gt;
&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;The status bug is the smallest and my favorite, because the correct value was already in scope. The whole thing is one line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;conversationId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;open&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;formatted&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;status&lt;/code&gt; is a literal. The real value sits in the conversation payload that the method one level below already fetches — the fix is to read it from there instead of asserting it.&lt;/p&gt;

&lt;p&gt;Pagination was ported from the host script that already did it right, plus an explicit truncation marker on message reads, so a partial conversation announces itself instead of passing as complete.&lt;/p&gt;

&lt;p&gt;The tariff cache stops recording ignorance as knowledge:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# before
&lt;/span&gt;&lt;span class="n"&gt;tariff&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;tariff_from_triage_note&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cid&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cw_token&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tariff&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tariff&lt;/span&gt;          &lt;span class="c1"&gt;# cache: note content is immutable
&lt;/span&gt;
&lt;span class="c1"&gt;# after
&lt;/span&gt;&lt;span class="n"&gt;tariff&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;tariff_from_triage_note&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cid&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cw_token&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;tariff&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;unknown&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;           &lt;span class="c1"&gt;# "not posted yet" is not a value
&lt;/span&gt;    &lt;span class="n"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tariff&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tariff&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This does not distinguish the two «unknown»s either; it stops trusting them. An unparseable note now costs a re-read every run instead of being wrong forever, which is the trade I want.&lt;/p&gt;

&lt;p&gt;And the alerting path: the Telegram helper handles its own failures and logs them, the error handler no longer depends on the channel that just failed, and chunking — a fifth thing I fixed while I was in there — splits outside HTML tags instead of through them.&lt;/p&gt;

&lt;p&gt;One thing I am deliberately not calling a fix: the prompt line telling the digest job to send once and never edit. It patches today's symptom and it took ten seconds, but it is an instruction to a stochastic system, not a repair. The actual answer to the opening incident is not in the prompt. It is that a failed run now reaches someone, instead of a file nothing reads.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Improvements
&lt;/h2&gt;

&lt;p&gt;Every fix went red before it went green. The pagination and tariff bugs got standalone repro scripts against local mocks, with the buggy implementation copied verbatim, so the failure is demonstrated rather than argued:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;open conversations in inbox:      30
agent check_inbox list_open sees: 25  (ids 1..25)
host precheck list_open sees:     30  (ids 1..30)
=&amp;gt; 5 conversations invisible to the LLM agent, no truncation signal returned
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;run 1 (triage in progress): tariff='unknown'  cached={'tariff': 'unknown'}
run 2+ (note now present):  tariff='unknown'  (cache short-circuits, note never re-read)
24h ping fires: False   (correct behaviour would be: True)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For the digest incident itself, the replay went into a private test group rather than the doctors group, running the same job with the same shape of payload:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn2ytamrvixzxl9y01jcl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn2ytamrvixzxl9y01jcl.png" alt="Telegram test group showing two unedited RED demo messages and one GREEN demo message" width="800" height="231"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;RED replay: the agent sends, then tries to edit — the edit dies, the message stays raw. GREEN: same job, one send, no edit.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The RED runs are the interesting ones. The message is still sitting there in its raw, pre-edit state, which is precisely how this failed in production: not with an absence, but with a half-finished artifact that looks close enough to a real digest to be skimmed past.&lt;/p&gt;

&lt;h2&gt;
  
  
  Best Use of Sentry
&lt;/h2&gt;

&lt;p&gt;The instrumentation is the other half of the fix, because the four bugs above are the ones I found. The category of "silent wrong answer" is not exhausted by an audit, so the pipeline needs to be able to report on itself. Three independent sources now feed one stream: errors scraped from the runtime's own error log lines, failures inside our agent tools, and cron monitors that stop hearing from a job.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqvneft97e92hjmcpe43w.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqvneft97e92hjmcpe43w.png" alt="Sentry issue stream with three numbered issues: a core tool error, an agent tool failure and a cron monitor failure" width="800" height="445"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Three different origins, one stream: (1) a core tool error scraped from the runtime's ERROR log lines, (2) a failure inside one of our own agent tools, (3) a cron monitor that stopped hearing from its job. The unnumbered rows are the same machinery catching unrelated faults.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;When one of our tools fails, the HTTP call that caused it arrives attached, which turns "the agent said something odd this morning" into a five-second read:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr17a9ly5g0dzcc54vnop.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr17a9ly5g0dzcc54vnop.png" alt="Sentry breadcrumbs showing an agent tool failure and the Chatwoot HTTP request that caused it" width="799" height="563"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;An agent tool failure with the exact HTTP call that caused it — attached automatically as breadcrumbs.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Every tool call is now a span with the attributes that matter for this pipeline, and none of the attributes that would put patient text into a third-party system:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhy2scicwp8l8eizu7dt9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhy2scicwp8l8eizu7dt9.png" alt="Attributes tab of a gen_ai.execute_tool span in Sentry showing tool name, action and conversation id" width="799" height="563"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Every agent tool call is a span: tool name, action, numeric conversation id, latency, error status. Nothing in that list is message content — that is the whole design.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The tool's own Input and Output tabs are where the arguments and the result would sit. They are empty, and that is deliberate:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg6fe7c9ftqr23vjhvy2b.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg6fe7c9ftqr23vjhvy2b.png" alt="Sentry AI tab showing the trace as a timeline of model calls and tool calls, with an empty Input tab" width="799" height="563"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The same run in Sentry's AI view — model calls and tool calls in sequence, with the selected tool's Input tab empty: arguments and responses are never recorded (PHI).&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Model calls get the same treatment, including the ones nobody is watching, which for this agent is most of them:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkwoow5569wsgwu1oqufj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkwoow5569wsgwu1oqufj.png" alt="Attributes of a gen_ai.chat span in Sentry showing model, provider, conversation id and timings" width="799" height="563"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;One span per model call — model, provider, conversation id, latency, time to first byte. Emitted for background cron runs too, which is where this agent does most of its work.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;And they correlate, which is what makes an agent run debuggable at all: the model call, the tool call it triggered, and the outbound HTTP request that tool made, in one waterfall.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdbsvyrvej84yfyj92cfm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdbsvyrvej84yfyj92cfm.png" alt="Sentry trace waterfall of one agent run: a model call, a tool call and the outbound HTTP request" width="800" height="293"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;One agent run, one trace: the model call, the tool call it triggered, and the outbound HTTP request — correlated through the runtime's own trace context.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Token usage and cost land at run level, keyed by conversation:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi54y821df9gztngwoo1c.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi54y821df9gztngwoo1c.png" alt="Attributes of a gen_ai.invoke_agent span showing token usage and dollar cost per agent run" width="799" height="563"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Run-level usage: tokens in/out and dollar cost per agent run, keyed by conversation id — the cheapest possible answer to «what is this agent costing us».&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Then the piece that speaks to the failure mode underneath the opening incident. Error monitoring catches runs that fail. It cannot catch runs that stop happening, and it cannot catch a delivery that quietly goes nowhere. Cron monitors turn a schedule into an expectation, and check-ins into evidence. (The red one below is the host digest script from bug four, not the doctors' digest from the opening: two different jobs, one failure mode.)&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu8e4n2vpkwt65rynbcye.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu8e4n2vpkwt65rynbcye.png" alt="Sentry cron monitors list with one failing and two healthy monitors" width="799" height="262"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Three host crons, monitored from code (no UI setup). The red one is the script from bug four — it has been failing since June 13, and the monitor is the first thing in six weeks to say so out loud.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvzt9z06y0a8udh031k1o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvzt9z06y0a8udh031k1o.png" alt="Sentry cron monitor detail page showing missed and failed check-ins and an ongoing issue" width="799" height="563"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Missed, missed, failed. A job that stops running produces no error at all, and a job that fails into a log nobody reads produces none that anyone sees — a cron monitor turns both into the same alert.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Six weeks of that script crashing on schedule would have been one alert on day one.&lt;/p&gt;

&lt;p&gt;On the PHI side, the deliberate choice: prompt and response recording is off. The safe list is model id, provider, token counts, cost, durations, tool name and action, numeric conversation ids, session key, job id, outcome. Anything string-valued that comes from message content goes through redaction or does not get sent. The result is a monitoring stack that can tell me a tool failed, which tool, on which conversation id, how long it took and what it cost, and cannot tell me or anyone else what the patient wrote.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4k153d225mdzezbtwdtz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4k153d225mdzezbtwdtz.png" alt="Sentry AI transcript tab stating that the conversation's messages were not captured" width="799" height="563"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Sentry offers a full conversation transcript for AI traces — and for this agent it is empty by construction: «This conversation's messages weren't captured».&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That empty transcript is not a gap in the setup. It is the setup.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I took away
&lt;/h2&gt;

&lt;p&gt;The bug I went looking for did not exist, and the four I found were all the same bug wearing different clothes: a component that could not distinguish "I don't know" from a value, and picked a value. In ordinary code, that produces a wrong answer somewhere downstream and usually an exception eventually. In a pipeline where a language model reads those answers, it produces a fluent, confident, well-formatted summary of a reality that is 25 conversations wide instead of 30, and nobody downstream has any way to tell.&lt;/p&gt;

&lt;p&gt;So the question I am still working on, and would like yours on: in your own systems, how do you tell the difference between an agent run that went fine and an agent run that quietly took a path that does not work? Error rates will not show it. Neither will the output, because the output always looks great.&lt;/p&gt;




&lt;p&gt;If this article helped you, you can &lt;a href="https://ko-fi.com/sergeiparfenov?ref=dev" rel="noopener noreferrer"&gt;buy me a coffee&lt;/a&gt;. Your support helps me make time for the next experiment.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>bugsmash</category>
      <category>ai</category>
      <category>observability</category>
    </item>
  </channel>
</rss>
