<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Chen Yuan</title>
    <description>The latest articles on DEV Community by Chen Yuan (@chenyuan20509).</description>
    <link>https://dev.to/chenyuan20509</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3935918%2F55c92f67-ea0a-42da-a9f2-f44b2d4c60b2.png</url>
      <title>DEV Community: Chen Yuan</title>
      <link>https://dev.to/chenyuan20509</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/chenyuan20509"/>
    <language>en</language>
    <item>
      <title>How to Build a Post-Launch Eval Canary That Tells a Real LLM Regression From Sampling Noise</title>
      <dc:creator>Chen Yuan</dc:creator>
      <pubDate>Wed, 30 Sep 2026 13:56:53 +0000</pubDate>
      <link>https://dev.to/chenyuan20509/how-to-build-a-post-launch-eval-canary-that-tells-a-real-llm-regression-from-sampling-noise-bp9</link>
      <guid>https://dev.to/chenyuan20509/how-to-build-a-post-launch-eval-canary-that-tells-a-real-llm-regression-from-sampling-noise-bp9</guid>
      <description>&lt;p&gt;Is the model actually getting worse, or did I just get unlucky on a handful of prompts?&lt;/p&gt;

&lt;p&gt;That question is why threads like "is it just me or is it dumber today" keep recurring, and it is the question a post-launch eval canary has to answer with a number instead of a feeling. The reference implementation here is &lt;a href="https://github.com/ninjahawk/livenerf" rel="noopener noreferrer"&gt;livenerf&lt;/a&gt;, a long-running, deterministic-as-possible benchmark tracking whether Claude Opus 5.5 (released 2026-09-22) gets quietly worse after launch. This article teaches the reusable method behind it: how to freeze prompts, pin the harness, calibrate a panel of "sometimes right" questions, compute paired per-item statistics with clustered standard errors, run a control arm, watch output-token counts as an early signal, and pre-register the rule that decides when you are allowed to say "regression."&lt;/p&gt;

&lt;p&gt;You do not need a frontier model or a cluster. You need a few hundred labeled questions, one machine, and the discipline to fix everything except the model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Accept that the harness is part of the treatment
&lt;/h2&gt;

&lt;p&gt;A canary compares a model to itself across time. Every other moving part becomes a confound. The livenerf v0 runner executes a Claude Max subscription through headless Claude Code (&lt;code&gt;claude -p&lt;/code&gt;), pins the CLI version, disables the auto-updater, and refuses to run when the version no longer matches — because a changed harness looks exactly like a changed model. The same logic applies to your system prompt, your sampling parameters, your tool list, and your working directory.&lt;/p&gt;

&lt;p&gt;The cheapest enforcement is a guard that fails loudly. The code below is illustrative and was not executed for this article; treat the version string and command shape as placeholders for your own stack.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# illustrative, not executed
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;sys&lt;/span&gt;

&lt;span class="n"&gt;PINNED_CLI&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-code 2.4.1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;assert_pinned_harness&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--version&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;capture_output&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;check&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;
    &lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="n"&gt;stdout&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;PINNED_CLI&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;harness drift: expected &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;PINNED_CLI&lt;/span&gt;&lt;span class="si"&gt;!r}&lt;/span&gt;&lt;span class="s"&gt;, got &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="si"&gt;!r}&lt;/span&gt;&lt;span class="s"&gt;. &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Refusing to run; a harness change is indistinguishable from a model change.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Version pinning only works if the run itself is hermetic. livenerf uses a short frozen system prompt, no tools, no MCP, no &lt;code&gt;CLAUDE.md&lt;/code&gt; or memory, one turn, and a fixed empty working directory. Code answers are graded by hidden tests in a sandbox outside the model. That last decision matters more than it sounds: if an LLM judge scores your outputs, the judge can drift too, and you will be measuring two moving objects at once.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Freeze a prompt file, not a prompt idea
&lt;/h2&gt;

&lt;p&gt;A canary needs an artifact you can hash. Keep the system prompt, the item text, and the scoring function in version control, and make the runner read them from disk rather than from a database that someone can quietly edit.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# illustrative, not executed
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;

&lt;span class="n"&gt;PANEL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;panel/frozen_panel_v1.jsonl&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;SYSTEM&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;panel/system_v1.txt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;load_frozen&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;system&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;SYSTEM&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;items&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;PANEL&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;splitlines&lt;/span&gt;&lt;span class="p"&gt;()]&lt;/span&gt;
    &lt;span class="n"&gt;digest&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sha256&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;SYSTEM&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_bytes&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;PANEL&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_bytes&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;hexdigest&lt;/span&gt;&lt;span class="p"&gt;()[:&lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;system&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;items&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;digest&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two files, one digest, no edits without a new panel version. The digest is what you log next to every run so a later comparison can prove it used the same questions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Calibrate toward "sometimes right" questions
&lt;/h2&gt;

&lt;p&gt;Here is the part most homegrown evals skip, and it is the part that buys you statistical power. Under a logit-shift model, a question with pass rate &lt;code&gt;p&lt;/code&gt; carries information &lt;code&gt;p(1-p)&lt;/code&gt; per sample. Questions the model always answers correctly contribute nearly nothing, and questions it always fails contribute nearly nothing. Only questions near the middle of the difficulty curve move when capability moves.&lt;/p&gt;

&lt;p&gt;livenerf screened 2,336 GPQA Diamond, MMLU-Pro, competition-math and AIME 2025–26 questions with 4 samples each. Opus 5.5 got roughly 93% right on the first try, and 97% of questions turned out to be always right or always wrong. The surviving 78 "sometimes right" questions became the frozen panel. A second effect showed up: questions selected for being sometimes-right regress toward the mean. On fresh samples their pass rate rose from 54.7% to 62.0%, so the power calculation has to use the fresh rates rather than the selection rates.&lt;/p&gt;

&lt;p&gt;If you calibrate your own panel, do the same: sample each candidate question several times, keep the ones with intermediate pass rates, then re-estimate their rates on data you did not use for selection. Budget the panel around that middle band, because that is where the information is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: Compare items, not just averages
&lt;/h2&gt;

&lt;p&gt;A raw accuracy comparison between two time windows carries question difficulty inside it. If day 30 happens to sample harder items, the score drops for reasons that have nothing to do with the model. The fix is to compare each item to itself: for every question, take the score in the current window minus the score in the launch-week baseline, then average those per-item deltas.&lt;/p&gt;

&lt;p&gt;Because each item is measured repeatedly, the per-item deltas are not independent, and naive standard errors will be too small. That is what clustered standard errors are for, and it is the approach livenerf takes, following Evan Miller's &lt;a href="https://arxiv.org/abs/2411.00640" rel="noopener noreferrer"&gt;"Adding Error Bars to Evals"&lt;/a&gt;. Cluster by question; with one observation per item per window, the per-item delta is the cluster.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# illustrative, not executed
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;numpy&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;paired_delta_with_clustered_se&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;baseline&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ndarray&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;current&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ndarray&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;baseline/current are item-level scores in [0,1], same item order.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;current&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;baseline&lt;/span&gt;
    &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;mean&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;mean&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="c1"&gt;# cluster-robust variance with one cluster per item:
&lt;/span&gt;    &lt;span class="c1"&gt;# var(mean) = sum_i (d_i - mean)^2 / (n * (n - 1))
&lt;/span&gt;    &lt;span class="n"&gt;var&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;mean&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="n"&gt;se&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sqrt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;var&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;n_items&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;delta&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;mean&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;se&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;se&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ci99&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;mean&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mf"&gt;2.576&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;se&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;mean&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mf"&gt;2.576&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;se&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The choice of 99% rather than 95% is deliberate: a canary fires repeatedly, so it needs a higher bar per look to keep false alarms rare. With one run per day of the whole panel, livenerf's own power calculation puts the detectable accuracy change at roughly 7.5 points per 10-day window. That is the instrument's resolution, and knowing it is what stops you from over-reading a 2-point wobble.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5: Add a control arm and watch tokens first
&lt;/h2&gt;

&lt;p&gt;The control arm is the cheapest way to avoid blaming the model for something else. livenerf runs &lt;code&gt;claude-opus-5&lt;/code&gt; on the same harness alongside the tracked model. If the control arm moves by the same amount, the change is probably infrastructure, a shared dependency, or the harness — not a silent downgrade of the tracked model.&lt;/p&gt;

&lt;p&gt;Output-token counts are the early signal. In livenerf's validation, lowering effort to low cut output tokens by 62% and accuracy by 8.3 ± 4.5 points; effort medium cut tokens by 26% and accuracy by 4.2 ± 3.9 points. Tokens moved before accuracy did, which makes them a leading indicator worth logging on every run even when accuracy looks flat.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 6: Pre-register the decision rule
&lt;/h2&gt;

&lt;p&gt;A canary that decides what counts as a regression after seeing the data is a vibes machine with extra steps. The livenerf decision rule was committed to public git before any series data existed, so the timestamp is meaningful. It calls a change only if three conditions hold together: the 99% interval excludes zero in two consecutive 10-day windows, the effect is at least 3 points, and the control arm does not show the same move. Null results and improvements get published just as loudly as regressions.&lt;/p&gt;

&lt;p&gt;The two-window requirement is what protects against a single unlucky stretch. Here is a compact check you can adapt; it assumes you already computed intervals per window for the tracked model and the control.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# illustrative, not executed
&lt;/span&gt;&lt;span class="n"&gt;MIN_EFFECT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.03&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;call_regression&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;windows&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;control&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;windows: per-10-day dicts with &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;ci99&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt; and &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;delta&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt; for the tracked model.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;windows&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;insufficient data&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;last_two&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;windows&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;:]&lt;/span&gt;
    &lt;span class="n"&gt;excludes_zero&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;all&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;w&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ci99&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;w&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ci99&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;w&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;last_two&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;big_enough&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;all&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;abs&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;w&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;delta&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;MIN_EFFECT&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;w&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;last_two&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;control_moved&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;any&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ci99&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ci99&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;control&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;:]&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;excludes_zero&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;big_enough&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;control_moved&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;regression called&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;no call&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Note what the rule does not do. It does not treat the launch-week baseline as ground truth. livenerf's README is explicit that launch week is a reference point, not a ceiling: launch week could be the worst week because of capacity strain, and past quality incidents turned out to be infrastructure bugs rather than deliberate downgrades. The baseline is the thing you compare against, not the thing you trust.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 7: Publish the limits alongside the numbers
&lt;/h2&gt;

&lt;p&gt;An instrument that hides its blind spots will be misused. livenerf documents a concrete one: a same-family swap (Opus 5 in place of Opus 5.5) was not distinguishable in a validation's worth of samples, at -3.8 ± 6.3 points and -23% tokens. That is a known limit of the instrument, written down in advance so a null result is not later spun as a clean bill of health.&lt;/p&gt;

&lt;p&gt;Every canary has a floor. The honest move is to state yours: how many points you can detect, over what window, with what false-alarm rate, and which kinds of changes you would miss entirely. Then keep the raw logs append-only, as livenerf does with Inspect &lt;code&gt;.eval&lt;/code&gt; files, so anyone can re-analyze without re-running.&lt;/p&gt;

&lt;h2&gt;
  
  
  The method in one paragraph
&lt;/h2&gt;

&lt;p&gt;Freeze the prompts and the harness, calibrate toward questions that are sometimes right, compare each item against its own baseline, cluster your standard errors by question, run a control arm, log tokens as a leading indicator, and pre-register the rule that authorizes the word "regression." None of that requires a lab. It requires deciding, before you look, what would count as evidence — and then publishing the null results with the same energy as the alarming ones. The alternative is the loop we already know: vibes versus vibes, one launch week at a time.&lt;/p&gt;




&lt;p&gt;Originally published on &lt;a href="https://dispatch-blog.hashnode.dev/how-to-build-a-post-launch-eval-canary-that-tells-a-real-llm-regression-from-sampling-noise" rel="noopener noreferrer"&gt;Dispatch&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>python</category>
      <category>programming</category>
      <category>tutorial</category>
      <category>ai</category>
    </item>
    <item>
      <title>Your Go Module Path Should Outlive GitHub: Designing Stable Import Paths</title>
      <dc:creator>Chen Yuan</dc:creator>
      <pubDate>Tue, 29 Sep 2026 12:30:10 +0000</pubDate>
      <link>https://dev.to/chenyuan20509/your-go-module-path-should-outlive-github-designing-stable-import-paths-2e0n</link>
      <guid>https://dev.to/chenyuan20509/your-go-module-path-should-outlive-github-designing-stable-import-paths-2e0n</guid>
      <description>&lt;p&gt;Iain Cambridge recently described a company that ended up using GitLab, GitHub, and Azure DevOps at the same time because moving code between hosts had become too expensive. The code itself was not trapped by a proprietary build system. The trap was simpler: the hosting provider's domain had become part of the Go module path.&lt;/p&gt;

&lt;p&gt;That is an architectural dependency hiding inside a naming convention.&lt;/p&gt;

&lt;p&gt;Go makes repository-shaped module paths convenient. A module such as &lt;code&gt;github.com/acme/widgets&lt;/code&gt; tells the toolchain where the code probably lives, gives readers a familiar place to inspect it, and works with almost no setup. The cost appears later, when the repository owner wants to change hosts, split infrastructure, mirror code, or move to an internal forge.&lt;/p&gt;

&lt;p&gt;A stable module path changes that tradeoff. Instead of naming the current hosting company, it names a domain you control and lets that domain tell the Go toolchain where the repository lives today.&lt;/p&gt;

&lt;p&gt;This is not about avoiding GitHub. It is about making GitHub an implementation detail rather than part of your public package identity.&lt;/p&gt;

&lt;h2&gt;
  
  
  The coupling is in the module path, not the Git remote
&lt;/h2&gt;

&lt;p&gt;A Go module is identified by the path declared in &lt;code&gt;go.mod&lt;/code&gt;. That path also becomes the prefix for package imports inside the module. If you publish a public library with this declaration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="n"&gt;module&lt;/span&gt; &lt;span class="n"&gt;github&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;com&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;acme&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;widgets&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;consumer code naturally imports packages from that namespace:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="s"&gt;"github.com/acme/widgets/parser"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Changing the Git remote on your own workstation does not change that identity. You can point &lt;code&gt;origin&lt;/code&gt; at another server and keep developing, but consumers still ask the Go toolchain for &lt;code&gt;github.com/acme/widgets&lt;/code&gt;. The public name and the original host remain attached.&lt;/p&gt;

&lt;p&gt;That distinction matters because repository location and module identity solve different problems. A Git remote answers where maintainers push code. A module path answers what downstream builds request.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a custom domain changes the boundary
&lt;/h2&gt;

&lt;p&gt;The Go modules reference allows a module path to begin with a domain you control. When the path does not directly identify a supported repository host, the &lt;code&gt;go&lt;/code&gt; command performs an HTTP lookup using the module path and looks for a &lt;code&gt;go-import&lt;/code&gt; meta tag.&lt;/p&gt;

&lt;p&gt;That makes a path such as this possible:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="n"&gt;module&lt;/span&gt; &lt;span class="k"&gt;go&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;example&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;com&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;widgets&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The public package identity is now &lt;code&gt;go.example.com/widgets&lt;/code&gt;, while the repository can still live on GitHub. The domain answers the lookup request and points the Go toolchain at the current source repository.&lt;/p&gt;

&lt;p&gt;A minimal response can contain a tag like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;meta&lt;/span&gt; &lt;span class="na"&gt;name=&lt;/span&gt;&lt;span class="s"&gt;"go-import"&lt;/span&gt;
      &lt;span class="na"&gt;content=&lt;/span&gt;&lt;span class="s"&gt;"go.example.com/widgets git https://github.com/acme/widgets"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the repository later moves to another Git server, the module path does not need to change. The operator updates the repository URL returned by the domain. New consumers continue to request &lt;code&gt;go.example.com/widgets&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;This is the same reason organizations use stable DNS names in front of replaceable infrastructure. A name controlled by the application owner can stay fixed while the service behind it changes.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Go toolchain actually asks for
&lt;/h2&gt;

&lt;p&gt;The custom domain is not magic. It participates in a defined resolution protocol.&lt;/p&gt;

&lt;p&gt;When the &lt;code&gt;go&lt;/code&gt; command needs a module directly from version control, it can request a URL based on the module path with the query parameter &lt;code&gt;go-get=1&lt;/code&gt;. The response must place the &lt;code&gt;go-import&lt;/code&gt; meta tag in the document head, early enough for the restricted parser to find it.&lt;/p&gt;

&lt;p&gt;For a module named &lt;code&gt;go.example.com/widgets&lt;/code&gt;, the lookup is conceptually equivalent to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://go.example.com/widgets?go-get=1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The meta tag contains three important pieces: the module root path, the version-control type, and the repository URL. The root path must match the module being requested or be a valid prefix that can be verified by another lookup.&lt;/p&gt;

&lt;h2&gt;
  
  
  A small redirect service is enough
&lt;/h2&gt;

&lt;p&gt;You do not need a package registry to own the namespace. A tiny HTTP handler can answer module discovery requests while normal browser visits redirect to project documentation.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="n"&gt;moduleHandler&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;w&lt;/span&gt; &lt;span class="n"&gt;http&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ResponseWriter&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;http&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;URL&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Query&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"go-get"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="s"&gt;"1"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;w&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Header&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Content-Type"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"text/html; charset=utf-8"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;fmt&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Fprint&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;w&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;`&amp;lt;html&amp;gt;&amp;lt;head&amp;gt;&amp;lt;meta name="go-import" content="go.example.com/widgets git https://github.com/acme/widgets"&amp;gt;&amp;lt;/head&amp;gt;&amp;lt;/html&amp;gt;`&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="n"&gt;http&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Redirect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;w&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"https://github.com/acme/widgets"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;http&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;StatusFound&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important property is ownership. Your DNS, TLS certificate, and HTTP response define the stable name. GitHub, GitLab, a self-hosted forge, or another supported VCS endpoint can sit behind it.&lt;/p&gt;

&lt;p&gt;For a larger organization, this endpoint can be generated from a table of module roots and repository URLs. The routing layer should remain boring. It is naming infrastructure, not application logic.&lt;/p&gt;

&lt;h2&gt;
  
  
  Migration gets easier when the stable name exists first
&lt;/h2&gt;

&lt;p&gt;The best time to choose a durable module path is before the first public release. Once downstream modules import a path, changing it becomes an ecosystem migration rather than a repository migration.&lt;/p&gt;

&lt;p&gt;Suppose version one is published as &lt;code&gt;github.com/acme/widgets&lt;/code&gt;. Moving the repository later does not automatically rename existing imports. A &lt;code&gt;replace&lt;/code&gt; directive can help a maintainer test an alternate source locally, but it is not a global redirect for every consumer. Each downstream module controls its own dependency graph.&lt;/p&gt;

&lt;p&gt;By contrast, if version one starts as &lt;code&gt;go.example.com/widgets&lt;/code&gt;, the host can move without asking every consumer to edit source code. The module identity survives because consumers never imported the hosting provider's namespace.&lt;/p&gt;

&lt;p&gt;There is still operational work. The new repository must contain the expected tags, module contents, and &lt;code&gt;go.mod&lt;/code&gt;. Access controls must work. The vanity domain must continue serving the correct metadata. The point is not zero migration work. The point is that the migration stays behind an interface you own.&lt;/p&gt;

&lt;h2&gt;
  
  
  Public modules and private modules have different failure modes
&lt;/h2&gt;

&lt;p&gt;For public modules, the main concerns are durable naming, tag continuity, and making the discovery endpoint available. Module proxies may cache released versions, which helps consumers keep building older releases even if the repository later moves. Future versions still need a resolvable module path and a reachable source.&lt;/p&gt;

&lt;p&gt;Private modules add authentication and proxy policy. The Go reference documents &lt;code&gt;GOPRIVATE&lt;/code&gt; for module prefixes that should not use the public proxy or checksum database. A private organization might publish modules under a namespace such as &lt;code&gt;go.corp.example.com/team/service&lt;/code&gt; while routing discovery to an internal Git server.&lt;/p&gt;

&lt;p&gt;The naming idea remains the same: expose a stable application-owned prefix, then keep repository credentials and transport details behind it.&lt;/p&gt;

&lt;p&gt;One useful rule is to separate three questions during design review:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;What name will downstream source code import?&lt;/li&gt;
&lt;li&gt;Which service currently stores the Git repository?&lt;/li&gt;
&lt;li&gt;Which credentials and proxy rules are required to fetch it?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If the same string is answering all three questions, the design is probably more coupled than it needs to be.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stable import paths do not remove semantic versioning rules
&lt;/h2&gt;

&lt;p&gt;Owning the domain does not let a module ignore Go's version rules. Major versions after version one still require the expected path suffix. A version two module should use a path ending in &lt;code&gt;/v2&lt;/code&gt;, regardless of whether the prefix is a GitHub domain or your own.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="n"&gt;module&lt;/span&gt; &lt;span class="k"&gt;go&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;example&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;com&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;widgets&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;v2&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The custom domain protects the repository boundary. It does not change how module versions identify incompatible API lines.&lt;/p&gt;

&lt;p&gt;The same warning applies to repository subdirectories. A module's declared path, repository root, subdirectory, and tags still have to agree with the module rules. A vanity domain is a routing layer, not permission to make version layout ambiguous.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test the namespace like production infrastructure
&lt;/h2&gt;

&lt;p&gt;A stable import domain can become more important than the Git host because every clean environment depends on it during resolution. Treat it as infrastructure.&lt;/p&gt;

&lt;p&gt;At minimum, test the discovery response, TLS renewal, repository target, and a clean module download. Do the test from an environment that does not already have the module in its cache. A warm developer machine can hide a broken discovery endpoint.&lt;/p&gt;

&lt;p&gt;A simple CI check can request the discovery URL, assert that the expected &lt;code&gt;go-import&lt;/code&gt; tag is present, then create a temporary module and resolve a released version. The check should fail before a DNS or routing change reaches users.&lt;/p&gt;

&lt;h2&gt;
  
  
  The common shortcuts fail for predictable reasons
&lt;/h2&gt;

&lt;p&gt;One shortcut is to keep the GitHub module path and assume a repository transfer will solve future moves. That works only while the desired destination remains compatible with the old public name. A transfer inside GitHub may preserve useful redirects for web traffic, but the module identity is still a GitHub-owned namespace. Moving to a different host is a different problem.&lt;/p&gt;

&lt;p&gt;Another shortcut is to plan a mass search-and-replace later. That changes your own repository, not every consumer repository. Public libraries may have forks, tutorials, generated code, internal mirrors, and old services importing the earlier path. A source rewrite is therefore a compatibility event.&lt;/p&gt;

&lt;p&gt;A third shortcut is to use &lt;code&gt;replace&lt;/code&gt; directives as a migration mechanism. They are excellent for local development and controlled builds, but each main module owns its replacements. A library cannot publish a &lt;code&gt;replace&lt;/code&gt; directive that globally rewires all of its consumers.&lt;/p&gt;

&lt;p&gt;A fourth shortcut is to make the vanity endpoint depend on a large application stack. That turns a simple naming dependency into another fragile service. The discovery response should be small enough to serve from a static host, edge rule, or minimal handler.&lt;/p&gt;

&lt;h2&gt;
  
  
  When a GitHub-shaped path is still reasonable
&lt;/h2&gt;

&lt;p&gt;Not every Go repository needs a custom domain.&lt;/p&gt;

&lt;p&gt;A short-lived internal tool may never become a dependency. A prototype may be intentionally tied to one organization and one host. A personal project may value zero infrastructure more than future portability. In those cases, &lt;code&gt;github.com/owner/repo&lt;/code&gt; is direct and easy to understand.&lt;/p&gt;

&lt;p&gt;The calculation changes when the module becomes a durable public API. If other teams will import it for years, if multiple repositories share an organizational namespace, or if host migration is a realistic possibility, the cost of owning a small stable domain is easier to justify.&lt;/p&gt;

&lt;p&gt;The decision is similar to choosing a public API hostname. You can expose the current server name directly, but once clients depend on it, renaming becomes coordination work.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical rollout checklist
&lt;/h2&gt;

&lt;p&gt;For a new public Go module, the rollout can stay simple.&lt;/p&gt;

&lt;p&gt;First, choose a module prefix under a domain the organization intends to keep. Second, configure the &lt;code&gt;go-import&lt;/code&gt; response before publishing the first tagged release. Third, verify resolution from a clean environment. Fourth, keep the repository target configurable rather than hard-coded across many pages. Fifth, document who owns the DNS and discovery endpoint so the module does not become orphaned during an infrastructure handoff.&lt;/p&gt;

&lt;p&gt;For an existing module already published under GitHub, do not pretend a rename is free. Decide whether the compatibility cost is worth paying. If it is, publish a migration plan, keep old documentation available, and avoid moving the repository and renaming the module in the same opaque step. Consumers need to understand whether they are changing a source location, a module identity, or both.&lt;/p&gt;

&lt;p&gt;The deeper principle is small but useful: names that appear in other people's source code should be controlled by the party promising their stability.&lt;/p&gt;

&lt;p&gt;GitHub can remain the repository host. The module path does not have to advertise that fact forever.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Iain Cambridge, "Don't couple your Go code to GitHub": &lt;a href="https://iain.rocks/blog/dont-couple-your-go-code-to-github" rel="noopener noreferrer"&gt;https://iain.rocks/blog/dont-couple-your-go-code-to-github&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Go Modules Reference, module paths and repository discovery: &lt;a href="https://go.dev/ref/mod" rel="noopener noreferrer"&gt;https://go.dev/ref/mod&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The Hacker News discussion surfaced the topic on September 28, 2026. The implementation details above follow the Go module reference rather than assuming a hosting provider can transparently rename a public module namespace.&lt;/p&gt;




&lt;p&gt;Originally published on &lt;a href="https://dispatch-blog.hashnode.dev/your-go-module-path-should-outlive-github-designing-stable-import-paths" rel="noopener noreferrer"&gt;Dispatch&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>go</category>
      <category>github</category>
      <category>devops</category>
      <category>opensource</category>
    </item>
    <item>
      <title>When Code Is Cheap, Understanding Becomes the Bottleneck</title>
      <dc:creator>Chen Yuan</dc:creator>
      <pubDate>Fri, 25 Sep 2026 12:53:02 +0000</pubDate>
      <link>https://dev.to/chenyuan20509/when-code-is-cheap-understanding-becomes-the-bottleneck-29hf</link>
      <guid>https://dev.to/chenyuan20509/when-code-is-cheap-understanding-becomes-the-bottleneck-29hf</guid>
      <description>&lt;p&gt;A coding agent can produce a large branch faster than a human can build a reliable mental model of it. That changes the review problem. The limiting factor is no longer typing speed or even raw code generation. It is whether a reviewer can reconstruct intent, architecture, tradeoffs, and risk before approving a change.&lt;/p&gt;

&lt;p&gt;That is the interesting idea behind Whiteboard, an open-source desktop app from dev.fast that appeared on Hacker News this week. It connects coding agents such as Claude Code and Codex to a shared visual workspace. Agents can draw architecture, sequence diagrams, traces, and explanations next to the code they are changing, while the reviewer can jump from those artifacts back to source.&lt;/p&gt;

&lt;p&gt;The project is still early, but the design points at a broader engineering shift: review tooling has to become better at compression. A thousand changed lines may contain only three important decisions. If the interface cannot surface those decisions, a faster agent simply creates a larger verification queue.&lt;/p&gt;

&lt;h2&gt;
  
  
  The review bottleneck moved from syntax to intent
&lt;/h2&gt;

&lt;p&gt;Traditional review assumes that the code diff is the primary artifact. A human reads changed files, infers the goal, reconstructs the control flow, and checks whether the implementation matches the intended behavior. That works reasonably well when a developer wrote the branch over hours or days and can explain it during review.&lt;/p&gt;

&lt;p&gt;Agent-written branches break that assumption in two ways. First, the amount of generated code can grow much faster than reviewer attention. Second, the agent may make dozens of local choices that were never stated explicitly in the task. A diff shows the result of those choices but not the reasoning path that produced them.&lt;/p&gt;

&lt;p&gt;The useful review object therefore becomes larger than a patch. It includes the request, the architectural choices, the agent trace, the changed symbols, and the tests that claim to validate the result. Whiteboard is interesting because it treats these as connected objects rather than separate tabs.&lt;/p&gt;

&lt;p&gt;A minimal machine-readable review record could look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"goal"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"add idempotent retry handling"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"decisions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"store request keys before side effects"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"reuse existing transaction boundary"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"reject duplicate payload mismatch"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"changedSymbols"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"createJob"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"JobRepository.insert"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"evidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"retry_test"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"duplicate_payload_test"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The value of such a record is not that JSON is better than prose. The value is that every review surface can point back to the same small set of claims.&lt;/p&gt;

&lt;h2&gt;
  
  
  A semantic diff is more useful than a shorter diff
&lt;/h2&gt;

&lt;p&gt;Whiteboard includes an AST-aware semantic diff viewer written in Rust. The project describes it as a way to hide noise, summarize large added functions as pseudocode, and collapse categories such as tests or documentation when they are not the current focus.&lt;/p&gt;

&lt;p&gt;That distinction matters. A normal diff compressor usually removes lines. A semantic diff tries to preserve meaning while reducing visual volume.&lt;/p&gt;

&lt;p&gt;Consider a refactor that renames a helper, moves a function, and changes one branch condition. A line diff may show three files with dozens of changed lines. A semantic view should answer a different question: what behavior actually changed?&lt;/p&gt;

&lt;p&gt;One possible intermediate representation is a list of symbol-level edits:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;SemanticChange&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;moved&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;symbol&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;from&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;to&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;renamed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;before&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;after&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;behavior&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;symbol&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;summary&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;test&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;covers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once a tool can classify changes at that level, the UI can prioritize behavioral edits and de-emphasize movement or formatting. That is a much better match for how senior reviewers think.&lt;/p&gt;

&lt;h2&gt;
  
  
  Diagrams become useful when they are linked to code
&lt;/h2&gt;

&lt;p&gt;Architecture diagrams often decay because they are separate documentation. The boxes survive while the implementation moves on. Whiteboard takes a different approach: visualizations such as sequence diagrams and entity relationships can link back to underlying code.&lt;/p&gt;

&lt;p&gt;That connection is important because a diagram should not merely explain a system. It should help a reviewer test claims about the system.&lt;/p&gt;

&lt;p&gt;Suppose an agent claims that a new webhook path is idempotent. A useful diagram can show request entry, key lookup, transaction start, side effect, and response. The reviewer should then be able to jump from each node to the exact implementation that supports it.&lt;/p&gt;

&lt;p&gt;A simple Mermaid sequence could capture the claim:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;sequenceDiagram
  participant C as Client
  participant A as API
  participant R as Repository
  participant W as Worker
  C-&amp;gt;&amp;gt;A: POST job with idempotency key
  A-&amp;gt;&amp;gt;R: reserve key
  R--&amp;gt;&amp;gt;A: existing or new
  A-&amp;gt;&amp;gt;W: enqueue only if new
  A--&amp;gt;&amp;gt;C: stable result&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The important part is not the drawing. It is the traceability. If the "reserve key" node links to code that performs the lookup outside the transaction, the reviewer can immediately challenge the diagram instead of trusting it.&lt;/p&gt;

&lt;p&gt;This turns a visual artifact into an executable review index: every box is a claim, and every claim should have code or test evidence behind it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agent traces should expose decisions, not just activity
&lt;/h2&gt;

&lt;p&gt;Whiteboard also focuses on decision logs. That addresses another common problem in agent workflows: traces are usually too verbose to review directly.&lt;/p&gt;

&lt;p&gt;A raw coding-agent trace may contain searches, file reads, failed attempts, tool calls, and intermediate plans. Keeping the trace is useful for auditability, but asking a reviewer to read the whole thing defeats the purpose.&lt;/p&gt;

&lt;p&gt;The better abstraction is a decision ledger. A decision is worth surfacing when it changes behavior, risk, or maintainability. Examples include choosing a new dependency, changing a transaction boundary, adding a fallback path, skipping an existing abstraction, or accepting a compatibility tradeoff.&lt;/p&gt;

&lt;p&gt;A compact decision schema might be:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;decision-17&lt;/span&gt;
&lt;span class="na"&gt;topic&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;retry ownership&lt;/span&gt;
&lt;span class="na"&gt;choice&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;worker owns retry scheduling&lt;/span&gt;
&lt;span class="na"&gt;alternatives&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;API schedules retry&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;queue policy schedules retry&lt;/span&gt;
&lt;span class="na"&gt;reason&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;existing worker already records attempt state&lt;/span&gt;
&lt;span class="na"&gt;evidence&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;worker/retry.ts&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;worker/retry.test.ts&lt;/span&gt;
&lt;span class="na"&gt;risk&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;duplicate scheduling if API fallback remains enabled&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is reviewable because it separates an engineering choice from the mechanical work used to implement it. It also gives future maintainers a reason for the shape of the code instead of only a commit hash.&lt;/p&gt;

&lt;h2&gt;
  
  
  The best review surface is reversible
&lt;/h2&gt;

&lt;p&gt;A review tool should make it easy to move from explanation back to evidence without changing the branch. That sounds obvious, but many agent interfaces blur review and execution. A reviewer asks a question, the agent edits the code, and the evidence changes while it is being inspected.&lt;/p&gt;

&lt;p&gt;A safer pattern is to separate review mode from implementation mode. In review mode, tools may read the repository, inspect traces, render diagrams, and compare branches. They should not silently mutate files.&lt;/p&gt;

&lt;p&gt;That boundary can be represented explicitly in an agent tool contract:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;ReviewContext&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;mode&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;read-only&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;baseRef&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;headRef&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;allowFileWrite&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;allowGitMutation&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is not merely a permissions detail. Reversibility improves reasoning. A reviewer can explore alternative explanations without worrying that a question has already changed the object being reviewed.&lt;/p&gt;

&lt;p&gt;Whiteboard currently states that files cannot be edited inside the app. That limitation may look inconvenient, but it also creates a useful separation: the canvas is for understanding and review, while code mutation remains in the connected coding agent or editor.&lt;/p&gt;

&lt;h2&gt;
  
  
  Local-first matters more once traces contain repository context
&lt;/h2&gt;

&lt;p&gt;The project is MIT licensed and works against local checkouts. Its README also says anonymous telemetry excludes code, diffs, Whiteboard text, prompts, and model output, and that telemetry can be disabled.&lt;/p&gt;

&lt;p&gt;That model fits a practical constraint of agent-assisted development: review artifacts can contain more sensitive information than a normal diff. A trace may reveal rejected designs, internal paths, debugging output, prompts, or architecture notes that were never intended to leave the workstation.&lt;/p&gt;

&lt;p&gt;A local review surface reduces the number of systems that need access to that context. It does not remove the need to evaluate the connected model provider, because Claude Code, Codex, or another agent may still send data according to its own configuration. But it keeps the visualization layer from automatically becoming another hosted copy of the repository conversation.&lt;/p&gt;

&lt;p&gt;For teams evaluating similar tools, the useful privacy checklist is concrete:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Where is the repository read?&lt;/li&gt;
&lt;li&gt;Where are traces stored?&lt;/li&gt;
&lt;li&gt;Can telemetry be disabled?&lt;/li&gt;
&lt;li&gt;Does sharing create a new server-side copy?&lt;/li&gt;
&lt;li&gt;Which component sends prompts to model providers?&lt;/li&gt;
&lt;li&gt;Can a review be reproduced without network access?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those questions are more useful than a generic "local-first" label because they expose the actual data boundaries.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a production review loop should verify
&lt;/h2&gt;

&lt;p&gt;A visual review tool is only useful if it leads to a deterministic approval decision. The interface may be a canvas, but the final questions are still engineering questions.&lt;/p&gt;

&lt;p&gt;For an agent-generated branch, a production review loop should verify at least four layers.&lt;/p&gt;

&lt;p&gt;First, intent: does the branch solve the requested problem, and are the major autonomous decisions visible?&lt;/p&gt;

&lt;p&gt;Second, behavior: which code paths changed, and what tests or runtime evidence cover those paths?&lt;/p&gt;

&lt;p&gt;Third, blast radius: which callers, schemas, permissions, migrations, queues, or external contracts could be affected?&lt;/p&gt;

&lt;p&gt;Fourth, reversibility: if the change is wrong, can it be disabled, rolled back, or isolated without another emergency rewrite?&lt;/p&gt;

&lt;p&gt;A small review manifest can force those questions before approval:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;intent_verified&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="na"&gt;behavior_tests&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;retry_test&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;duplicate_payload_test&lt;/span&gt;
&lt;span class="na"&gt;blast_radius&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;jobs table&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;worker queue&lt;/span&gt;
&lt;span class="na"&gt;rollback&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;feature flag&lt;/span&gt;
  &lt;span class="na"&gt;owner&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;platform team&lt;/span&gt;
&lt;span class="na"&gt;open_questions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A tool can visualize this manifest, but it should not invent the answers. The reviewer still owns the decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real product is a shorter path from claim to evidence
&lt;/h2&gt;

&lt;p&gt;Whiteboard still has clear limitations. The project says it does not currently support editing files in the app, multi-repository review is not well supported, and shared reviews do not automatically receive later updates. Those constraints matter for teams with large service graphs or stacked changes.&lt;/p&gt;

&lt;p&gt;But the architectural direction is useful even if a team never adopts this particular application. The review system should minimize the distance between a claim and the evidence that can falsify it.&lt;/p&gt;

&lt;p&gt;If an agent says a change is backward compatible, the reviewer should be one action away from the schema diff and compatibility test. If it says a retry is idempotent, the reviewer should be one action away from the transaction boundary and duplicate-request test. If it says an architectural choice was required, the alternatives and tradeoff should be recorded instead of reconstructed from chat history.&lt;/p&gt;

&lt;p&gt;That suggests a practical rule for agent tooling: do not optimize only for faster code generation. Optimize for faster human verification of generated decisions.&lt;/p&gt;

&lt;p&gt;The winning interface may look less like an editor and more like an evidence map. Code, diagrams, traces, tests, and decisions remain separate artifacts, but the reviewer can move among them without rebuilding context from scratch.&lt;/p&gt;

&lt;p&gt;As coding agents become capable of producing broader changes, that compression layer becomes part of software quality. More generated code is only useful when a human can still understand what changed, why it changed, and where to look when the explanation is wrong.&lt;/p&gt;




&lt;p&gt;Originally published on &lt;a href="https://dispatch-blog.hashnode.dev/when-code-is-cheap-understanding-becomes-the-bottleneck" rel="noopener noreferrer"&gt;Dispatch&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>python</category>
      <category>programming</category>
      <category>tutorial</category>
      <category>ai</category>
    </item>
    <item>
      <title>CUDA Rust: What the Two GPU Kernel Tracks Actually Guarantee</title>
      <dc:creator>Chen Yuan</dc:creator>
      <pubDate>Thu, 17 Sep 2026 15:55:01 +0000</pubDate>
      <link>https://dev.to/chenyuan20509/cuda-rust-what-the-two-gpu-kernel-tracks-actually-guarantee-22f0</link>
      <guid>https://dev.to/chenyuan20509/cuda-rust-what-the-two-gpu-kernel-tracks-actually-guarantee-22f0</guid>
      <description>&lt;p&gt;On September 8, 2026, NVIDIA published a technical blog post that introduced CUDA Rust, a pair of projects that let developers write GPU kernels in Rust and compile them natively to PTX. The first, cuda-oxide, is a custom rustc codegen backend that routes kernel functions through Rust MIR, the Pliron IR framework and LLVM before handing everything else to the standard backend. The second, cutile-rs, is published on crates.io, runs on stable Rust 1.89 or newer, and JIT-compiles tile-based kernels through CUDA Tile IR. Neither project is production-ready, and NVIDIA describes both as work that will keep maturing into 2027 and beyond.&lt;/p&gt;

&lt;p&gt;The interesting part is not the syntax. It is where each track draws the line between what the compiler proves and what the programmer still has to prove by hand.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed in September 2026
&lt;/h2&gt;

&lt;p&gt;The systems layer of AI has been moving toward Rust for a while. The Nova Linux driver is written in Rust, NVIDIA Dynamo is built on a Rust core, and NVTX has Rust bindings. The GPU kernel stayed the exception. Kernels could be launched from Rust, but the kernel body itself usually had to be written in another language. CUDA Rust closes that gap by making the kernel a Rust compilation target rather than a wrapper around code produced somewhere else.&lt;/p&gt;

&lt;p&gt;Two tracks match the two programming models CUDA already exposes. SIMT is the model used in CUDA C++ and numba-cuda: the programmer states what one thread does and launches thousands of them. Tile is the newer model, also available in C++ and Python, where the programmer states what one tile of data does and the compiler decides how that maps onto the hardware. cuda-oxide is the SIMT track. cutile-rs is the Tile track.&lt;/p&gt;

&lt;p&gt;NVIDIA's advice on picking between the two models is explicit. Reach for Tile first, because the compiler decides how tiles map onto each architecture and the source does not encode architecture-specific choices. Drop to SIMT when that control is needed, or when the kernel wants to manage memory and threads directly.&lt;/p&gt;

&lt;p&gt;Both projects are early. cuda-oxide is early alpha. cutile-rs is further along, published on crates.io and already used outside NVIDIA in the Grout inference engine at Hugging Face and in mistral.rs. Coverage is incomplete in both, and APIs will move. NVIDIA also says it plans to support inter-language interop between CUDA Rust, CUDA C++ and CUDA Python, so picking one frontend does not cut a team off from the others.&lt;/p&gt;

&lt;h2&gt;
  
  
  The SIMT track: one thread at a time
&lt;/h2&gt;

&lt;p&gt;cuda-oxide is a custom rustc codegen backend. It intercepts compilation, routes functions marked as kernels through Rust MIR, Pliron and LLVM IR down to PTX, and hands everything else to the standard backend. The GPU dialects layered on top of Pliron are NVIDIA's own. Host and device code live in one file, build with one command, and need no separate kernel crate.&lt;/p&gt;

&lt;p&gt;The safety argument lives in the kernel signature. A mutable slice is the wrong shape for an output buffer, because every thread would need the same mutable borrow and the borrow checker refuses that. cuda-oxide uses a type called &lt;code&gt;DisjointSlice&amp;lt;f32&amp;gt;&lt;/code&gt; instead, which splits one mutable borrow into per-thread pieces, giving each thread exclusive access to its own element and nothing else. The index returned by &lt;code&gt;thread::index_1d()&lt;/code&gt; is not a bare integer, and &lt;code&gt;get_mut&lt;/code&gt; only accepts that index type. It hands back an &lt;code&gt;Option&lt;/code&gt;, so the out-of-bounds case becomes a branch the kernel handles rather than a memory error discovered later.&lt;/p&gt;

&lt;p&gt;Here is the vector addition kernel from the announcement, attributes included:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;use&lt;/span&gt; &lt;span class="nn"&gt;cuda_device&lt;/span&gt;&lt;span class="p"&gt;::{&lt;/span&gt;&lt;span class="n"&gt;kernel&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;launch_bounds&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;launch_contract&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;thread&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;DisjointSlice&lt;/span&gt;&lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="nd"&gt;#[kernel]&lt;/span&gt;
&lt;span class="nd"&gt;#[launch_bounds(&lt;/span&gt;&lt;span class="mi"&gt;256&lt;/span&gt;&lt;span class="nd"&gt;)]&lt;/span&gt;
&lt;span class="nd"&gt;#[launch_contract(domain&lt;/span&gt; &lt;span class="nd"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="nd"&gt;,&lt;/span&gt; &lt;span class="nd"&gt;block&lt;/span&gt; &lt;span class="nd"&gt;=&lt;/span&gt; &lt;span class="nd"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;256&lt;/span&gt;&lt;span class="nd"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="nd"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="nd"&gt;))]&lt;/span&gt;
&lt;span class="k"&gt;pub&lt;/span&gt; &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;vecadd&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;f32&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;f32&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="k"&gt;mut&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;DisjointSlice&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nb"&gt;f32&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;idx&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nn"&gt;thread&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;index_1d&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;idx_raw&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;idx&lt;/span&gt;&lt;span class="nf"&gt;.get&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nf"&gt;Some&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;c_elem&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="nf"&gt;.get_mut&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;idx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;c_elem&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;idx_raw&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;idx_raw&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;#[launch_bounds(256)]&lt;/code&gt; tells the compiler how many threads per block to budget registers for. &lt;code&gt;#[launch_contract]&lt;/code&gt; declares that this kernel indexes in one dimension with 256-thread blocks. The inputs &lt;code&gt;a&lt;/code&gt; and &lt;code&gt;b&lt;/code&gt; are ordinary shared slices that every thread can read. The output &lt;code&gt;c&lt;/code&gt; is the exclusive piece.&lt;/p&gt;

&lt;p&gt;The host side has to satisfy that declaration rather than assert it. The launcher validates the requested geometry against the contract and against the live device limits, and the safe launch method only accepts the token that validation returns:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;prepared&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;module&lt;/span&gt;&lt;span class="nf"&gt;.prepare_vecadd&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nn"&gt;LaunchConfig1D&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;new&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="n"&gt;N&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nb"&gt;u32&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="nf"&gt;.div_ceil&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;256&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;256&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;&lt;span class="o"&gt;?&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="n"&gt;module&lt;/span&gt;&lt;span class="nf"&gt;.vecadd&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;prepared&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;a_dev&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;b_dev&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="k"&gt;mut&lt;/span&gt; &lt;span class="n"&gt;c_dev&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;?&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A kernel without a contract exposes only raw unsafe launch methods, because a bare launch configuration says nothing about the kernel it is launching. The declaration and the check are what turn a launch from a convention into something the toolchain can reject.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Tile track: one logical thread per tile
&lt;/h2&gt;

&lt;p&gt;cutile-rs works one level higher. Each tile block runs the kernel body once, as a single logical thread over one sub-tensor of data, and the compiler decides how many real GPU threads back that logical thread. The &lt;code&gt;#[cutile::module]&lt;/code&gt; macro embeds the kernel AST in the host binary and JIT-compiles it through CUDA Tile IR the first time the kernel is actually launched.&lt;/p&gt;

&lt;p&gt;The exclusivity mechanism is different. There is no &lt;code&gt;DisjointSlice&lt;/code&gt; in this track. Mutable tensors are partitioned on the host before launch, and each tile block receives one writable sub-tensor that no other tile block can overlap. That is what a mutable reference already guarantees in Rust, extended across the GPU launch boundary.&lt;/p&gt;

&lt;p&gt;The same elementwise addition, written for tiles, looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;use&lt;/span&gt; &lt;span class="nn"&gt;cutile&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nn"&gt;prelude&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="nd"&gt;#[cutile::module]&lt;/span&gt;
&lt;span class="k"&gt;mod&lt;/span&gt; &lt;span class="n"&gt;kernel&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;use&lt;/span&gt; &lt;span class="nn"&gt;cutile&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nn"&gt;core&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

    &lt;span class="nd"&gt;#[cutile::entry()]&lt;/span&gt;
    &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="n"&gt;add&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="n"&gt;B&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;i32&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;z&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="k"&gt;mut&lt;/span&gt; &lt;span class="n"&gt;Tensor&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nb"&gt;f32&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;B&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;Tensor&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nb"&gt;f32&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;Tensor&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nb"&gt;f32&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;tx&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;load_tile_like&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;z&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;ty&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;load_tile_like&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;z&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="n"&gt;z&lt;/span&gt;&lt;span class="nf"&gt;.store&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tx&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;ty&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The body loads the tile of each input that lines up with the output sub-tensor, adds them, and stores the result across the whole tile. The minus one in the input shapes is a sentinel rather than a size: that dimension is read off the tensor at launch, so the shape can vary without recompiling. The kernel body runs once per sub-tensor, and there are no thread indices to compute.&lt;/p&gt;

&lt;p&gt;The host side is where the geometry is fixed, and one call does three jobs at once:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;z&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nn"&gt;api&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nn"&gt;zeros&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nb"&gt;f32&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;&lt;span class="nf"&gt;.partition&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="mi"&gt;128&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;Vec&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nb"&gt;f32&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nn"&gt;kernel&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;z&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;.first&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="nf"&gt;.unpartition&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="nf"&gt;.to_host_vec&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="nf"&gt;.sync_on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;?&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;.partition([128])&lt;/code&gt; gives each tile exclusive ownership of a 128-element chunk, fixes the grid at 1024 divided by 128, which is 8 tiles, and supplies the const generic tile width that never appears at the call site. Nothing runs until &lt;code&gt;.sync_on(&amp;amp;stream)&lt;/code&gt;. Everything before it, including the allocations, the kernel call and the copy back to the host, is a lazy description recorded rather than submitted, which is why the whole program is one chain with a single synchronization point.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the two tracks actually differ
&lt;/h2&gt;

&lt;p&gt;The difference is not only syntax. It is the level at which the safety contract is expressed, and how much of the execution model the programmer keeps.&lt;/p&gt;

&lt;p&gt;In cuda-oxide, the programmer thinks in threads. The kernel body describes one thread, and the launch contract describes how many threads exist and how they are organized. Shared memory, barriers, atomics and warp-level operations stay available as explicit primitives. That control is what makes the SIMT track usable for kernels with unusual access patterns or hardware-specific tricks, and it is also what keeps the shared memory path on unsafe code today. Shared memory is the bedrock of fast SIMT kernels, and making that path safe is described as active work.&lt;/p&gt;

&lt;p&gt;In cutile-rs, the programmer thinks in tiles. The kernel body describes operations on blocks of data, and the compiler decides how those operations map onto threads, shared memory and hardware resources. There is no thread index and no explicit shared memory allocation in the source. A tile block is a single logical thread, so there are no threads for the programmer to race. The trade is less control over scheduling, and in exchange, less opportunity to get the scheduling wrong.&lt;/p&gt;

&lt;p&gt;Both tracks catch the aliasing mistake that is hardest to debug: passing an output buffer as one of its own inputs. They draw the line in different places. cuda-oxide checks each launch call. cutile-rs follows ownership of the tensors across the launch boundary, which is the stronger of the two claims.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the compiler catches, and what it cannot
&lt;/h2&gt;

&lt;p&gt;The guarantees are real but bounded. Both toolchains reject aliasing and ownership violations before the kernel runs, and cuda-oxide turns an out-of-bounds index into an &lt;code&gt;Option&lt;/code&gt; the kernel must handle. Neither toolchain catches a kernel that reads the wrong input, performs an incorrect reduction, or walks memory in an order that destroys performance. Memory safety and data-race freedom are not computational correctness.&lt;/p&gt;

&lt;p&gt;The aliasing rejections look like this in practice. Passing the output buffer as one of its own inputs does not compile, whether or not that kernel would actually race:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;error[E0502]: cannot borrow `c_dev` as mutable because it is also borrowed as immutable
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same mistake on the Tile side does not compile either, because the tensor has already moved:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;error[E0382]: use of moved value: `z`
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Neither error depends on a runtime check or a race detector. The first is the borrow checker looking at one call site. The second is ownership that survives the launch boundary, which is why the Tile track can make the claim without a purpose-built type.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reading the two examples side by side
&lt;/h2&gt;

&lt;p&gt;The two kernels above compute the same thing for the same input, and both print the same line when the host side is attached. The SIMT version processes one element per thread, reads the index from a helper that returns an index type, and writes through a &lt;code&gt;DisjointSlice&lt;/code&gt;. The Tile version processes one tile per block, takes the tile width as a const generic parameter, and writes through a mutable tensor reference.&lt;/p&gt;

&lt;p&gt;The signatures carry the difference. The SIMT kernel takes two shared slices and one exclusive slice-like type. The Tile kernel takes two shared tensors with a dynamic dimension and one mutable tensor whose static width is the partition size. The SIMT kernel needs a launch contract so the host side can validate the geometry. The Tile kernel needs a host-side partition so exclusivity is established and the grid follows from it.&lt;/p&gt;

&lt;p&gt;Where cuda-oxide makes the programmer state the launch geometry and then checks it, cutile-rs derives the geometry from the data layout and removes the geometry from the source entirely. That is the same trade as before, restated at the level of the signature, and it is the clearest way to see what each track is optimizing for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trying it today: requirements and rough edges
&lt;/h2&gt;

&lt;p&gt;Both tracks need Linux and a GPU with compute capability 8.0 or later. The similarity ends there. cuda-oxide also needs a CUDA toolkit of version 12.x or newer, clang with its libclang headers, and the pinned nightly toolchain. Installing the driver subcommand and scaffolding a project looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cargo +nightly-2026-04-03 install --git https://github.com/NVlabs/cuda-oxide.git cargo-oxide
cargo oxide new vecadd_demo
cd vecadd_demo
cargo oxide doctor
cargo oxide run
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first run builds the codegen backend, so it takes a while, and later runs reuse the cache. &lt;code&gt;cargo oxide doctor&lt;/code&gt; checks the whole environment, including an optional system LLVM.&lt;/p&gt;

&lt;p&gt;cutile-rs asks for less. It needs stable Rust 1.89 or newer, CUDA 13.3, and a GPU of the same compute capability, with no nightly toolchain and no LLVM of the programmer's own. It is on crates.io, so there is nothing to clone:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cargo new vecadd_demo
cd vecadd_demo
cargo add cutile
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Adoption follows the same split. cutile-rs is already used outside NVIDIA, while cuda-oxide remains early alpha. The pinned nightly in the SIMT track is the kind of requirement NVIDIA says it would like to stop asking for, and until then the toolchain moves on NVIDIA's schedule rather than the team's.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to choose a track for a real workload
&lt;/h2&gt;

&lt;p&gt;The choice is about where complexity should live. Tile programming is easier to write and safe by construction for elementwise work, reductions and matrix operations, because the compiler owns thread mapping and memory layout, and the partition removes aliasing without asking the programmer to reason about thread indices. SIMT programming keeps control for kernels that need specific thread coordination, custom shared memory, or access patterns that the tile compiler cannot express.&lt;/p&gt;

&lt;p&gt;The heuristic NVIDIA offers is short: find the Tile track first, use it while the operations map naturally to tiles, and reach for SIMT when control over memory and threads is needed. That ordering also matches maturity. The Tile track is the one with a published crate and stable Rust support, so it is the reasonable place to prototype, and the SIMT track is where a team lands when the prototype shows that the compiler's scheduling decisions are the bottleneck.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means for existing CUDA C++ and Python stacks
&lt;/h2&gt;

&lt;p&gt;CUDA Rust does not ask for a rewrite. The tracks are additive, and the stated plan for inter-language interop means a Rust kernel can sit next to C++ and Python kernels rather than replacing them. A team with a working CUDA C++ codebase can move one hot kernel at a time and keep the rest of the pipeline where it is.&lt;/p&gt;

&lt;p&gt;The nearer problem is churn. cuda-oxide pins a nightly toolchain, which means the toolchain updates on NVIDIA's schedule, and both projects are early enough that APIs will move between releases. Coverage is incomplete, so a kernel that depends on a specific hardware feature may not have a safe path yet. The honest reading of the announcement is that the destination is clear and the road is not finished.&lt;/p&gt;

&lt;p&gt;What the announcement does settle is where the ownership guarantees stop. For years, Rust on the CPU could prove that two writers never touched the same memory at the same time, and the same program had to hand its GPU work to another language and lose that proof at the boundary. The SIMT track restores it with a purpose-built type that splits a mutable borrow across threads. The Tile track restores it by partitioning the output before the launch, so exclusivity is established on the host and carried into the kernel. Either way, the aliasing mistakes that used to surface as flaky production failures become compile errors or rejected launches, and what remains for the programmer is the part that was always a human responsibility, which is getting the computation right.&lt;/p&gt;




&lt;p&gt;Originally published on &lt;a href="https://dispatch-blog.hashnode.dev/cuda-rust-what-the-two-gpu-kernel-tracks-actually-guarantee" rel="noopener noreferrer"&gt;Dispatch&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>python</category>
      <category>programming</category>
      <category>tutorial</category>
      <category>ai</category>
    </item>
    <item>
      <title>Testing Race Conditions: Making Nondeterminism Reproducible</title>
      <dc:creator>Chen Yuan</dc:creator>
      <pubDate>Sat, 12 Sep 2026 12:39:15 +0000</pubDate>
      <link>https://dev.to/chenyuan20509/testing-race-conditions-making-nondeterminism-reproducible-3bkd</link>
      <guid>https://dev.to/chenyuan20509/testing-race-conditions-making-nondeterminism-reproducible-3bkd</guid>
      <description>&lt;p&gt;A passing concurrency test usually proves that one schedule was safe, not that the program was safe. If the failure depends on two operations landing in a narrow order, adding more loop iterations often repeats the same harmless order. The useful change is to control the schedule, record it, or explore it systematically.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a Passing Stress Test Proves Almost Nothing
&lt;/h2&gt;

&lt;p&gt;The usual stress test starts several workers, runs the operation many times, and checks the final result. That catches bugs when the operating system happens to choose the needed interleaving. It does not ask the scheduler to choose that interleaving, and it does not preserve the choice when a run fails.&lt;/p&gt;

&lt;p&gt;A green stress test therefore has a narrow meaning: the observed executions did not violate the assertion. The Go &lt;a href="https://go.dev/doc/articles/race_detector" rel="noopener noreferrer"&gt;race detector documentation&lt;/a&gt; states the same limit for its dynamic detector: it can find races that happen at runtime, but cannot find a race in code that the test never executes. Code coverage and schedule coverage are different measurements.&lt;/p&gt;

&lt;p&gt;A failing stress test is useful when it preserves its seed and inputs. A green stress test that consumes minutes has not shown that the schedule space was examined.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Failure Mode: Interleavings, Not Code Paths
&lt;/h2&gt;

&lt;p&gt;The first diagnostic step is to name the failure precisely. A data race is an unsynchronized concurrent access to the same memory location where at least one access writes. A race condition is broader: the program's result is wrong because events occur in an order the design did not permit. A program can avoid a data race with a mutex and still have an ordering bug inside the locked operations.&lt;/p&gt;

&lt;p&gt;Four categories show up often in test reports:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A data race has conflicting memory accesses without a valid synchronization edge.&lt;/li&gt;
&lt;li&gt;An atomicity violation splits a check and an update that should behave as one operation.&lt;/li&gt;
&lt;li&gt;An ordering assumption expects event A before event B, but no barrier, channel, or condition establishes it.&lt;/li&gt;
&lt;li&gt;A lost wakeup lets a notification occur before a waiter records that it is waiting.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The detector must match the category. A sanitizer is suited to unsynchronized memory accesses. A barrier can force both sides of a check-then-act window. Events and conditions can make a required order explicit. A small model checker can explore alternative transitions. No single green result covers all four categories.&lt;/p&gt;

&lt;p&gt;This distinction also prevents a common bad fix: adding a sleep. A sleep changes timing without specifying the order. It may make a failure disappear on one machine and return on another. A synchronization primitive states what the test requires and gives the test a point at which to assert it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Instrument the Scheduler Instead of the Clock
&lt;/h2&gt;

&lt;p&gt;Replace guessed delays with a protocol. In the example below, request A writes one tenant, request B waits until that write happened and overwrites the shared client, and A then reads the tenant. The events force the bad order without depending on how long either thread sleeps.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;threading&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;concurrent.futures&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;ThreadPoolExecutor&lt;/span&gt;


&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Client&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_tenant&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;set_tenant&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tenant&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_tenant&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tenant&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_tenant&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_tenant&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_request_state_is_not_shared&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Client&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;first_write&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;threading&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Event&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;second_write&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;threading&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Event&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;request_a&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_tenant&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tenant-a&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;first_write&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="n"&gt;second_write&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;wait&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_tenant&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;request_b&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="n"&gt;first_write&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;wait&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_tenant&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tenant-b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;second_write&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nc"&gt;ThreadPoolExecutor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;max_workers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;observed_a&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pool&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;submit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request_a&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;pool&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;submit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request_b&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;result&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;observed_a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;result&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tenant-a&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The assertion fails on every run because the test has specified the interleaving. The production repair would normally make tenant state local to a request or protect a complete operation, not add another delay. The test is useful because it names the invariant: request A must not observe state written for request B.&lt;/p&gt;

&lt;p&gt;Use a barrier for phases, an event for milestones, and a condition for state predicates. Each primitive makes the scheduler part of the test fixture.&lt;/p&gt;

&lt;p&gt;For a notification path, start the waiter first and acknowledge that it has reached the waiting point before sending the signal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;threading&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_waiter_observes_a_notification&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;started&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;threading&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Event&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;notified&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;threading&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Event&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;finished&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;threading&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Event&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;values&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;waiter&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="n"&gt;started&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;notified&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;wait&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;AssertionError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;notification was lost&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;values&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;finished&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="n"&gt;thread&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;threading&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;target&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;waiter&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;thread&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;start&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;started&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;wait&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;notified&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;finished&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;wait&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;thread&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;values&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The test does not claim that every notification implementation is correct. It checks one contract with an explicit order and a bounded failure. A real condition-variable test should apply the same idea around the predicate and the wait, rather than relying on a sleep that merely makes the waiter likely to run first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deterministic Replay: Turn a Rare Failure into a Test Fixture
&lt;/h2&gt;

&lt;p&gt;A replayable failure needs more than a random seed. Record the seed, the generated input, the schedule decisions, the build identity, and any injected faults that can change the path. A seed is sufficient only when every decision comes from the same deterministic scheduler and input generator.&lt;/p&gt;

&lt;p&gt;Wall-clock reads, operating-system randomness, network responses, and uncontrolled background tasks can break replay. If they matter to the bug, replace them with recorded inputs or a virtual interface. Otherwise, a test that says it is replaying a failure is only rerunning a similar experiment.&lt;/p&gt;

&lt;p&gt;The manifest should be written before the test process exits, even when the assertion fails. A small helper can make the contract visible:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;random&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;


&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Schedule&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;seed&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;seed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;seed&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;rng&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Random&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;seed&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;decisions&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;choose&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;runnable&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;index&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;rng&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;randrange&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;runnable&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
        &lt;span class="n"&gt;choice&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;runnable&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;decisions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;choice&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;choice&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;save_replay&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;schedule&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;input_data&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;build_id&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;manifest&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;schema_version&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;seed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;schedule&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;seed&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;decisions&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;schedule&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;decisions&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;input&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;input_data&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;build_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;build_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;write_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;manifest&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;indent&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important property is that replay consumes the record rather than silently generating new choices. Keep the manifest as a regression fixture when the defect is fixed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Data Race Detection With Sanitizers in CI
&lt;/h2&gt;

&lt;p&gt;ThreadSanitizer, or TSan, instruments the program at compile time and uses a runtime library to detect data races in executions that actually occur. The &lt;a href="https://clang.llvm.org/docs/ThreadSanitizer.html" rel="noopener noreferrer"&gt;Clang documentation&lt;/a&gt; describes typical TSan slowdown as about 5x-15x and typical memory overhead as about 5x-10x. Those figures are tool guidance, not a promise for a particular service.&lt;/p&gt;

&lt;p&gt;TSan does not enumerate schedules. It can miss a race in an unexecuted path, and Clang warns that code generally needs to be compiled with the sanitizer flag. Precompiled or uninstrumented libraries can cause missed reports or false positives because the detector cannot see all synchronization. A clean run is evidence about that instrumented execution, not a proof about every possible execution.&lt;/p&gt;

&lt;p&gt;Go has a related detector behind &lt;code&gt;-race&lt;/code&gt;. Its documentation requires cgo and, on non-Darwin systems, a C compiler, and it lists the supported platform combinations. The same page documents the &lt;code&gt;GORACE&lt;/code&gt; option &lt;code&gt;halt_on_error&lt;/code&gt;, which can make the first report terminate the process instead of allowing a CI job to continue.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;env &lt;/span&gt;&lt;span class="nv"&gt;GORACE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;halt_on_error&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 go &lt;span class="nb"&gt;test&lt;/span&gt; &lt;span class="nt"&gt;-race&lt;/span&gt; ./...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run this stage on code paths that matter, not only on a tiny smoke test. Keep sanitizer jobs separate when needed, but make their failure visible and actionable. Do not link the sanitizer runtime into a production executable as a substitute for testing; Clang explicitly warns that its runtime is not intended for production security constraints.&lt;/p&gt;

&lt;h2&gt;
  
  
  Model Checking and Bounded Exhaustive Search
&lt;/h2&gt;

&lt;p&gt;Dynamic tools observe schedules. Model-checking tools generate schedules. For a small state machine, systematic exploration can be cheaper than trying to make a large integration test fail by chance.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/tokio-rs/loom" rel="noopener noreferrer"&gt;Rust Loom&lt;/a&gt; runs concurrent Rust tests while permuting executions under its supported portion of the C11 memory model, and it uses state reduction to limit combinatorial growth. Its own documentation also lists unsupported behaviour and warns that a passing result is not a sound proof for every C11 execution. Tests must use Loom's instrumented synchronization types so the checker can see the choices.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/awslabs/shuttle" rel="noopener noreferrer"&gt;AWS Shuttle&lt;/a&gt; takes a different trade-off. It controls and randomizes scheduling, and its documentation describes deterministic reproduction of failing tests, but it is not an exhaustive checker. That makes it useful for larger test cases where full exploration is too expensive, while leaving a clear limit on what a green run means.&lt;/p&gt;

&lt;p&gt;The boundary is the model. A semaphore accounting algorithm, connection-pool admission rule, or circuit-breaker transition may fit in a bounded model. A database server, network, external queue, or multi-service deployment does not become exhaustively checked merely because the client has a model-checking test. Keep the model small enough that its state and assumptions can be reviewed.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Practical Workflow for a Concurrency Bug Report
&lt;/h2&gt;

&lt;p&gt;Production reports usually describe symptoms, not interleavings: a response is truncated, a request receives another tenant's data, or a counter is occasionally wrong. Capture the request identifiers, input, build version, relevant logs, synchronization events, and the first invalid observation. Then reduce the report to a local invariant.&lt;/p&gt;

&lt;p&gt;Cloudflare's &lt;a href="https://blog.cloudflare.com/hyper-bug/" rel="noopener noreferrer"&gt;account of a hyper HTTP bug&lt;/a&gt; is a useful example. Its Images service saw intermittent truncation for larger images while responses still returned HTTP 200 and no application error. Cloudflare reports spending six weeks tracing the issue; in one case, roughly 200 KB arrived when the response was expected to be 3.3 MB. The investigation eventually isolated an incomplete flush before connection shutdown and added a deterministic test; the article says the fix took four lines of code.&lt;/p&gt;

&lt;p&gt;A practical sequence is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Identify the shared state and the operation that should protect it.&lt;/li&gt;
&lt;li&gt;Classify the defect as a data race, atomicity violation, ordering bug, or lost wakeup.&lt;/li&gt;
&lt;li&gt;Replace timing guesses with a barrier, event, condition, or scheduler hook.&lt;/li&gt;
&lt;li&gt;Record the input and schedule before rerunning the test.&lt;/li&gt;
&lt;li&gt;Verify that the unfixed version fails and the repaired version passes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A lost-update report can become a deterministic regression test with two barriers, one after both workers read and one after both workers write:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;threading&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_counter_update_is_atomic&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;counter&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;read_barrier&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;threading&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Barrier&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;write_barrier&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;threading&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Barrier&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;iterations&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;increment&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;iterations&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;counter&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
            &lt;span class="n"&gt;read_barrier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;wait&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
            &lt;span class="n"&gt;counter&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
            &lt;span class="n"&gt;write_barrier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;wait&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="n"&gt;threads&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;threading&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Thread&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;target&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;increment&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;thread&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;threads&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;thread&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;start&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;thread&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;threads&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;thread&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;counter&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;iterations&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The unsynchronized version fails because both workers read the same value in each round. A fixed implementation can protect the read-modify-write operation with a lock, after which this test should be adjusted so the synchronization in the fixture does not create an artificial deadlock. The key is that the regression test describes the lost update directly.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to Keep, What to Delete, and What to Watch
&lt;/h2&gt;

&lt;p&gt;An ordinary application repository should keep deterministic tests for critical shared state, replay manifests for injected schedules, and a sanitizer job when the language supports one. Delete long stress loops that only repeat a schedule; retain short stress runs when they complement, rather than replace, a forced-interleaving test.&lt;/p&gt;

&lt;p&gt;A library repository should add model-checking tests for small concurrent algorithms. A bug fixed in a reusable queue, pool, or state machine can affect every caller, so the extra constraints and instrumented primitives are worth maintaining. Keep a deterministic fixture for every concurrency defect that reached the issue tracker.&lt;/p&gt;

&lt;p&gt;For authorization, payment, replication, and other components where a race can corrupt trust or data, use all three layers where practical: forced schedules for known invariants, sanitizer coverage for observed memory races, and bounded exploration for small core algorithms. The watch item is always the same: a fix is merged after a green stress run, but the failing schedule was never captured. A race becomes a regression test only when its order is part of the test's data.&lt;/p&gt;

&lt;p&gt;The goal is not to remove nondeterminism from the whole system. It is to put nondeterminism behind an interface that a test can control, record, and challenge. Once the schedule is visible, concurrency debugging becomes an engineering task instead of a request to get lucky.&lt;/p&gt;




&lt;p&gt;Originally published on &lt;a href="https://dispatch-blog.hashnode.dev/testing-race-conditions-making-nondeterminism-reproducible" rel="noopener noreferrer"&gt;Dispatch&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>python</category>
      <category>programming</category>
      <category>tutorial</category>
      <category>ai</category>
    </item>
    <item>
      <title>Object Storage Can Replace a Database Under Four Contracts</title>
      <dc:creator>Chen Yuan</dc:creator>
      <pubDate>Thu, 10 Sep 2026 15:00:24 +0000</pubDate>
      <link>https://dev.to/chenyuan20509/object-storage-can-replace-a-database-under-four-contracts-5dbe</link>
      <guid>https://dev.to/chenyuan20509/object-storage-can-replace-a-database-under-four-contracts-5dbe</guid>
      <description>&lt;p&gt;Object storage is not a database. It stores bytes under keys. Yet it can serve as a database-like control-plane store when the workload fits a narrow set of contracts. The useful test is not a yes-or-no replacement claim. It is the set of guarantees the application needs, the ones the storage service supplies, and the ones the application must rebuild. The discussion around Ampbase and Tigris makes the tradeoffs concrete, because the system in question did not start from a desire to be clever. It started from a desire to avoid running a database it did not need, and the work was in rebuilding four guarantees that a database would have provided by default.&lt;/p&gt;

&lt;h2&gt;
  
  
  The data shape decides the storage engine
&lt;/h2&gt;

&lt;p&gt;The first thing to establish is that object storage is not a general replacement for a database. It is a replacement for a particular shape of data and access. In the Ampbase case, the data partitioned cleanly per organization, the writes were low volume and mostly uncontended, the reads were point lookups on keys the application controlled, and the interesting history was append-only by nature. Those properties are not incidental. They are the reason the design works at all.&lt;/p&gt;

&lt;p&gt;Change one of those assumptions and the design changes. Contention on one key turns CAS retries into a possible livelock. A transaction across objects cannot be built from per-object preconditions. Ad-hoc queries turn each new question into a hand-written backfill and index, with application code responsible for correctness.&lt;/p&gt;

&lt;p&gt;Storage selection follows the data shape and the questions the application asks. Small, append-only, pre-partitioned, already-aggregated records can fit object storage. High-volume, frequently updated, cross-referenced state is a database workload, and object storage will make you rebuild one badly.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four guarantees behind the headline
&lt;/h2&gt;

&lt;p&gt;When people reach for a database engine, they often need four behaviors: uniqueness, transactions, indexes, and history. Object storage supplies none of those as database features. Some services do supply two useful primitives: strong read-after-write consistency and conditional writes. The application can build the other behaviors only within the limits of those primitives.&lt;/p&gt;

&lt;p&gt;Strong read-after-write consistency means that a successful write is immediately visible to a later read. Amazon S3 documents this behavior for new objects, overwrites, deletes, and listings, and dates its change to December 2020. Without a strong consistency model, a stale read could make an application misjudge uniqueness or compare-and-swap state.&lt;/p&gt;

&lt;p&gt;Conditional writes turn compare-and-swap into HTTP preconditions. &lt;code&gt;If-None-Match: *&lt;/code&gt; creates only when a key is absent. &lt;code&gt;If-Match: {etag}&lt;/code&gt; replaces only the version whose ETag was read. These checks run against the object's latest state within the bucket's consistency model. They provide a concurrency primitive, not a complete transaction system.&lt;/p&gt;

&lt;p&gt;These primitives map to the four behaviors in narrow ways. Uniqueness is conditional creation on a content-derived key. Transaction-like mutation is a read, pure computation, and conditional replacement of one object. An index is a key designed around a frequent lookup. History is an append-only key space whose order matches the question being asked.&lt;/p&gt;

&lt;p&gt;Each contract has a boundary. Every writer must use conditional creation for uniqueness. A mutation retry must be pure. An index answers only its planned question, and an append-only log records only the events that were successfully appended.&lt;/p&gt;

&lt;h2&gt;
  
  
  A control plane built from immutable objects
&lt;/h2&gt;

&lt;p&gt;An object-first control plane can use two bucket layers. A global directory bucket records organizations; each organization gets a scoped bucket. Credentials for one customer cannot address another customer's objects, so tenant isolation is an infrastructure boundary rather than an application predicate. There is no &lt;code&gt;WHERE org_id = ?&lt;/code&gt; for a developer to forget.&lt;/p&gt;

&lt;p&gt;Within a customer bucket, a few key families provide the database-like behavior: a membership key for point lookup, a mutable pointer to the active configuration, immutable version objects, and append-only event objects. The application controls the paths, so these records do not require a general query engine.&lt;/p&gt;

&lt;p&gt;Version objects and events are never rewritten, so their audit property comes from the key space rather than from an &lt;code&gt;UPDATE&lt;/code&gt; path. Pointers are different: deploying overwrites the active pointer, and rollback moves it to an older version. The history remains because version objects are retained.&lt;/p&gt;

&lt;p&gt;The first block shows the key pattern: derive paths from the lookup, then create immutable records conditionally.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;ulid&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;membership_key&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;digest&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sha256&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;hexdigest&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;members/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;digest&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;.json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;version_key&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;channel_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;versions/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;channel_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;created_at&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;.pb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;event_key&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;org_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;event_time&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;event_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;events/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;org_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;event_time&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;event_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;.pb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The membership key is an index expressed as a path. Hashing the email lets the application compute one &lt;code&gt;GetObject&lt;/code&gt; without a listing or secondary index. Version keys use ULIDs; because object listings are lexicographic, a prefix or time range can return history in order. That works only for access patterns designed into the key names. It is not an &lt;code&gt;ORDER BY&lt;/code&gt; substitute for arbitrary queries.&lt;/p&gt;

&lt;h2&gt;
  
  
  What conditional writes can and cannot protect
&lt;/h2&gt;

&lt;p&gt;Conditional writes protect one key from concurrent mutation. They do not coordinate multiple keys or create cross-key atomicity. That boundary is easy to forget when the primitive feels strong.&lt;/p&gt;

&lt;p&gt;The second block shows the two forms: create-if-absent and replace-if-unchanged.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;precondition&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;etag&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;tuple&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;etag&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;*&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;etag&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;put_with_precondition&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;bucket&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                          &lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bytes&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;etag&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;if_match&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;if_none_match&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;precondition&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;etag&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;put_object&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;Bucket&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;bucket&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Body&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;IfMatch&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;if_match&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;IfNoneMatch&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;if_none_match&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Concurrent writers to one key receive one success and one precondition error. The loser must re-read: the existing record may be its own retried write or a real conflict. That idempotency check belongs in the application contract.&lt;/p&gt;

&lt;p&gt;A mutation function runs again after a conflict, so it must be pure. Side effects inside it, such as an email, counter, or payment call, may happen once per attempt. The compiler will not enforce this; the interface and review must.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why indexes and history change the design
&lt;/h2&gt;

&lt;p&gt;An index precomputes lookup paths; a unique index also rejects duplicate values. Object-first systems encode both in key names and conditional creation. The tradeoff is fixed access patterns: the application can ask only questions it designed for.&lt;/p&gt;

&lt;p&gt;Every access pattern becomes a precomputed key. A new question requires a new key and backfill, which is a hand-written index migration. That is the largest ongoing tax, and it returns with every feature.&lt;/p&gt;

&lt;p&gt;An append-only history can resemble event sourcing, but this design stores current state directly and uses events to explain changes. The log is an audit record, not a replay engine.&lt;/p&gt;

&lt;p&gt;The third block shows an index record whose path is the lookup and whose conditional creation enforces uniqueness.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;write_index_record&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;bucket&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                       &lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;payload&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sort_keys&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;put_object&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;Bucket&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;bucket&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Body&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;IfNoneMatch&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;*&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without multi-object transactions, state and event writes can split. If state is written first and the process dies before the event, current state is correct but history has a gap. Every log entry is real, but completeness requires reconciliation.&lt;/p&gt;

&lt;p&gt;Deletion is also different. Removing a key erases the distinction between never written and deleted unless a tombstone remains. Keep version objects and move pointers when history matters; storage then grows until a lifecycle policy removes old data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Multi-region writes need an explicit conflict policy
&lt;/h2&gt;

&lt;p&gt;Multi-region writes are the hardest boundary. A conditional write is evaluated against the state visible in its region. Two regions can therefore accept the same compare-and-swap before replication converges, after which only one version remains current.&lt;/p&gt;

&lt;p&gt;A read-after-write check cannot repair this because the lost update happens at the write. The safer rule is to adjudicate each CAS in one primary region and reject dependent writes elsewhere with a client-side guard. Replicas can serve reads, but the write policy must be explicit.&lt;/p&gt;

&lt;p&gt;The fourth block makes the retry policy visible: the caller supplies a pure mutation and the loop has a hard attempt limit.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;mutate_with_retry&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;bucket&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                      &lt;span class="n"&gt;mutate&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_attempts&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;delay&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.01&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;attempt&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;max_attempts&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;current&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;etag&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;read_object&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;bucket&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;new_body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;mutate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;current&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="nf"&gt;put_with_precondition&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;bucket&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;new_body&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;etag&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt;
        &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="n"&gt;PreconditionFailed&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;attempt&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;max_attempts&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;raise&lt;/span&gt;
            &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;delay&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;delay&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;delay&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The example uses a short exponential backoff and five attempts. If contention persists, the caller should surface a conflict rather than loop forever.&lt;/p&gt;

&lt;p&gt;Hot keys expose the workload assumption: continuous contention can turn retries into a livelock. Read amplification is the other cost. A directory read per request is repeated by every instance, and listing members before fetching each object creates many round trips. A read-through cache can reduce latency, but it cannot become the source of truth; cold-cache operation must remain correct.&lt;/p&gt;

&lt;h2&gt;
  
  
  A decision checklist for object-first systems
&lt;/h2&gt;

&lt;p&gt;Before deciding that object storage can replace a database, check the contracts. The data should partition cleanly per tenant, so that isolation is a credential boundary rather than a predicate. Writes should be low volume and mostly uncontended, so that compare-and-swap with retry resolves on the first attempt. Reads should be point lookups on keys you control, so that the access pattern is a key you chose rather than a query you wrote. The interesting history should be append-only, so that immutability is a property of the key space rather than a feature you implemented.&lt;/p&gt;

&lt;p&gt;Then check what you are giving up. You are giving up multi-key atomicity, so any invariant that spans objects has to be enforced by the application or abandoned. You are giving up ad-hoc queries, so every new question is a new key and a backfill job. You are giving up the guarantee that the audit log is complete, so if completeness matters you need a reconciliation process. You are giving up simple deletion, so retention has to be an explicit lifecycle policy. And you are giving up single-region simplicity, so multi-region writes need an explicit conflict policy that names the primary region and treats writes elsewhere as errors.&lt;/p&gt;

&lt;p&gt;If the workload fits and the contracts are acceptable, object storage can replace a database for that workload. If the workload does not fit, the honest move is to use a database and stop trying to build one from preconditions. The practical rule is this: let the workflow pick the storage engine, and only notice when the data shape has changed enough that the choice should change with it. The original discussion is at &lt;a href="https://news.ycombinator.com/item?id=49618450" rel="noopener noreferrer"&gt;https://news.ycombinator.com/item?id=49618450&lt;/a&gt; and the Tigris writeup is at &lt;a href="https://www.tigrisdata.com/blog/object-storage-all-need/" rel="noopener noreferrer"&gt;https://www.tigrisdata.com/blog/object-storage-all-need/&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;Originally published on &lt;a href="https://dispatch-blog.hashnode.dev/object-storage-can-replace-a-database-under-four-contracts" rel="noopener noreferrer"&gt;Dispatch&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>python</category>
      <category>programming</category>
      <category>tutorial</category>
      <category>ai</category>
    </item>
    <item>
      <title>When Safe Browsing Says Clean but Google Ads Says Malware</title>
      <dc:creator>Chen Yuan</dc:creator>
      <pubDate>Wed, 09 Sep 2026 15:10:00 +0000</pubDate>
      <link>https://dev.to/chenyuan20509/when-safe-browsing-says-clean-but-google-ads-says-malware-a7n</link>
      <guid>https://dev.to/chenyuan20509/when-safe-browsing-says-clean-but-google-ads-says-malware-a7n</guid>
      <description>&lt;p&gt;How can a website and a download look clean in several public checks while an advertising account is still rejected for malware? The contradiction is easier to understand when the ad campaign is treated as a delivery system rather than as a single web page.&lt;/p&gt;

&lt;p&gt;A current Hacker News discussion points to a useful case. The author of RACE, a native macOS terminal multiplexer written in Rust, reported that a Google Ads account was suspended after spending USD 500 on a campaign. The notice named "Malicious software" and "Compromised Site." The author then checked the website, the download infrastructure, the application, and several public security services, but the appeals still did not reveal which observation caused the decision.&lt;/p&gt;

&lt;p&gt;That report does not prove that Google made a mistake, and it does not reveal how the classifier worked. It does show why a publisher needs a record of the whole path from ad impression to first launch. A clean result from one layer is evidence about that layer. It is not a universal verdict about every layer that a visitor may cross.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the RACE case actually shows
&lt;/h2&gt;

&lt;p&gt;RACE is described as a terminal multiplexer that keeps shell sessions alive across application restarts. Its website is a static Bridgetown site, with downloads hosted separately. The author says it had no custom server-side application at submission time.&lt;/p&gt;

&lt;p&gt;The advertising attempt produced a suspension rather than a useful diagnostic. The notice named malicious software and a compromised site. Repeated appeals were rejected, and one path was blocked for a week. The review covered Safe Browsing, Search Console, VirusTotal, signatures, notarization, JavaScript bundles, logs, and different user agents.&lt;/p&gt;

&lt;p&gt;One detail matters more than the number of checks. RACE starts and manages background shell processes because process persistence is part of a terminal multiplexer. That behavior can look unusual to a detector even when documented. The report says the application did not inject code into other applications, change browser behavior, or conceal its process activity.&lt;/p&gt;

&lt;p&gt;Those are observations from the publisher's investigation, not the platform's internal reason. "The public checks were clean" is defensible; "the campaign was proven safe" is not supported by the same evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why clean scanners do not clear an ad campaign
&lt;/h2&gt;

&lt;p&gt;The checks in this case answer different questions. Google Safe Browsing examines URLs and warns users in Search and browsers when it detects unsafe sites. Its public status tool can tell a publisher what it reports about a URL at the time of the check. It does not promise that the URL will be treated the same way by an advertising review.&lt;/p&gt;

&lt;p&gt;Search Console's Security issues report has a different purpose. Google says the report presents findings when an evaluation determines that a site was hacked or shows behavior that could harm a visitor or their computer. The report can include sample affected URLs, but Google also says the sample may not be complete and that some issues have no example URL. A green result in that report is therefore useful evidence about the report, not a signed clearance for every ad destination and binary.&lt;/p&gt;

&lt;p&gt;Apple notarization answers a third question for macOS software. Apple describes its notary service as an automated system that scans for malicious components, checks code-signing issues, and returns a ticket that Gatekeeper can find. Apple also says notarization is not App Review. A valid Developer ID signature and ticket help a user evaluate an application, but they do not determine whether an ad platform accepts a campaign.&lt;/p&gt;

&lt;p&gt;Google Ads documentation says reviews may use multiple sources, including the ad, website, accounts, and third-party sources. It does not identify the RACE signal, but it explains why independent reports can disagree. A policy system can inspect more than a URL scanner or notarization service.&lt;/p&gt;

&lt;p&gt;The outputs are partial observations. An ad review can consider the relationship between the account, creative, destination, redirects, download, and prior activity. Treating the outputs as interchangeable creates false confidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Model the download path as a changing state machine
&lt;/h2&gt;

&lt;p&gt;A publisher should record the path as a sequence of states, not as a screenshot of a landing page. A practical sequence is search query, ad, landing URL, HTML response, script and redirect chain, download host, archive, signed application, first launch, and process or network behavior.&lt;/p&gt;

&lt;p&gt;Each state needs a timestamp, URL or artifact identity, observation method, and result. If the download changes while the ad remains unchanged, the reviewed object may no longer be the delivered object. A redirect can also vary by referrer, region, user agent, or cookie. These are engineering possibilities, not claims about the RACE classifier.&lt;/p&gt;

&lt;p&gt;A small record type makes the distinction concrete:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;dataclasses&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;dataclass&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timezone&lt;/span&gt;

&lt;span class="nd"&gt;@dataclass&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;frozen&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Observation&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;stage&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="n"&gt;artifact_sha256&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="n"&gt;observed_at&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;timestamp&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;timezone&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;utc&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;isoformat&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;


&lt;span class="n"&gt;record&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Observation&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;stage&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;landing-page&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://downloads.example.invalid/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;artifact_sha256&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;observed_at&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;timestamp&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="n"&gt;method&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;isolated-http-client&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The record does not make a verdict. It preserves context for comparing a later fetch with the earlier one, which matters when an appeal asks which release or response was examined.&lt;/p&gt;

&lt;p&gt;A state-machine view prevents a category error. A signed file can be safe at rest while a landing page is compromised, or a clean page can redirect to a different download. A legitimate application can also launch a helper that was not in the original hash.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verify redirects and artifacts without trusting one layer
&lt;/h2&gt;

&lt;p&gt;The first pass should collect metadata without launching the program. Google’s security guidance warns against opening infected pages directly in a browser and recommends safer methods for examining responses. Use an isolated HTTP client and disposable environment, not a daily workstation.&lt;/p&gt;

&lt;p&gt;The redirect chain should be preserved, not reduced to the final URL. Record status codes, location headers, response headers, the request method, and the user agent. Run the same capture from a clean environment when the result matters, then compare the chains rather than choosing the most reassuring one.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;urllib.request&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;build_opener&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;HTTPRedirectHandler&lt;/span&gt;


&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;TraceRedirects&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;HTTPRedirectHandler&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;redirect_request&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;file&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;code&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;new_url&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;events&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;from&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;full_url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;code&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;location&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;new_url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;})&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;super&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;redirect_request&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;file&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;code&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;new_url&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;fetch_with_trace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;tuple&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;]]:&lt;/span&gt;
    &lt;span class="n"&gt;tracer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;TraceRedirects&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;opener&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;build_opener&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tracer&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;request&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;User-Agent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;release-audit/1.0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;opener&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;4096&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;geturl&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;tracer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;events&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a capture tool, not a bypass for an advertising review. Do not rotate identities to evade a policy control. Document the ordinary path a reviewer can reproduce and note whether the destination changes.&lt;/p&gt;

&lt;p&gt;The file itself needs an identity. A filename is not enough because a publisher can replace the bytes while keeping the same name. Hash the exact archive that was uploaded, hash the extracted executable when appropriate, and keep the command output with the release record.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;sha256&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;sha256_file&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;digest&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sha256&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;block&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;iter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="sa"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;digest&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;update&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;block&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;digest&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;hexdigest&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;


&lt;span class="n"&gt;archive_hash&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sha256_file&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;RACE.dmg&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;file&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;RACE.dmg&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sha256&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;archive_hash&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a macOS release, keep the signing identity, entitlements, notarization result, and hash of the notarized deliverable. If a ticket is stapled after hashing, the reviewed object may differ from the object offered to users.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build an evidence bundle for a manual review
&lt;/h2&gt;

&lt;p&gt;A useful appeal package is a reproducible bundle, not a pile of screenshots. Identify the ad text, destination, time window, redirect chain, response headers, HTML and script hashes, download hash, signature details, notarization status, and observation environment.&lt;/p&gt;

&lt;p&gt;Separate facts from interpretation. An &lt;code&gt;observed&lt;/code&gt; field can contain a status code or hash; a &lt;code&gt;hypothesis&lt;/code&gt; field can say that a background process may have looked unusual. This keeps an explanation from being mistaken for a platform finding.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;dataclasses&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;asdict&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;dataclass&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;


&lt;span class="nd"&gt;@dataclass&lt;/span&gt;
&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ReviewBundle&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;ad_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;landing_url&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;final_url&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;artifact_sha256&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;observed&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;hypotheses&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;write_bundle&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bundle&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;ReviewBundle&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;write_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;asdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bundle&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;indent&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sort_keys&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;


&lt;span class="n"&gt;bundle&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;ReviewBundle&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;ad_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;campaign-record-2026-09-09&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;landing_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://race-term.example.invalid/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;final_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://downloads.example.invalid/RACE.dmg&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;artifact_sha256&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;record-after-download&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;observed&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;source&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;isolated-http-client&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
    &lt;span class="n"&gt;hypotheses&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;documented process persistence may need an explicit explanation&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;write_bundle&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bundle&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;review-bundle.json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The example values are placeholders and must not be submitted as evidence. In a real bundle, do not redact the artifact identity or the time of observation, but do remove cookies, access tokens, and private account data. Keep the original response bytes and logs in a protected archive so the summary can be regenerated.&lt;/p&gt;

&lt;p&gt;This format also makes version changes visible. If the binary, redirect chain, or JavaScript bundle changes after an appeal, create a new bundle rather than editing the old one. A reviewer should be able to tell whether the publisher is defending the same object or a later release.&lt;/p&gt;

&lt;h2&gt;
  
  
  What ad platforms should expose to developers
&lt;/h2&gt;

&lt;p&gt;The RACE report points to a product problem as well as a publisher problem: a generic label leaves no path to test the alleged fault. A review system should expose the category of object that triggered the decision, even if it cannot reveal a sensitive classifier: ad, destination, redirect, download, executable, account context, or third-party report.&lt;/p&gt;

&lt;p&gt;It should provide the observation time, sampled URL, response status, artifact hash when available, and environment class. It should distinguish a page warning from an account suspension. If a third-party signal was involved, the publisher needs enough information to identify the asset without seeing private detection rules.&lt;/p&gt;

&lt;p&gt;Search Console acknowledges that sample URLs may be incomplete or absent. An appeal interface should show that limit rather than treating a missing sample as proof that nothing was checked, or a clean public report as proof that the ad review had no other input.&lt;/p&gt;

&lt;p&gt;That would help both sides: a real compromise would be easier to reproduce and fix, while a false positive would be easier to isolate. The platform could protect its detection logic while giving the publisher a stable object to inspect.&lt;/p&gt;

&lt;h2&gt;
  
  
  A safer publishing checklist
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Freeze the release. Record the exact ad text, destination URL, redirect chain, archive hash, executable hash, signature identity, and notarization result before submitting a campaign.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Capture the normal path with an isolated HTTP client. Preserve redirects, response headers, and final URLs. Do not open an unknown page or program on the workstation used for credentials and daily work.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Check the layers separately. Run Safe Browsing, review Search Console security findings, inspect the download, and verify the platform-specific signing process. Store each result with its timestamp and scope.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Explain unusual behavior plainly. A terminal multiplexer that keeps shell processes alive should document that behavior, its configuration options, and its cleanup path. Documentation does not prove safety, but silence makes a legitimate behavior harder to review.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Submit a factual appeal. State what was observed, what was not observed, and what remains unknown. Do not call a clean scan a certificate, and do not claim to know the classifier's reason without a platform response.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Keep every appeal tied to an immutable bundle. If the binary or destination changes, start a new record, and ask which object and observation must change before a new review can be meaningful.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The RACE case does not establish that Google Ads misclassified a legitimate application. It establishes a narrower engineering lesson: security evidence is scoped to the system that produced it. Safe Browsing, Search Console, Apple notarization, file scanners, and ad review can each observe a different part of the same delivery path. Publishers need reproducible artifacts and a clear separation between observation and theory; platforms need review signals that let a legitimate publisher find the disputed state.&lt;/p&gt;

&lt;p&gt;Further reading: &lt;a href="https://xlii.space/eng/malicious-software-on-google-ads/" rel="noopener noreferrer"&gt;the RACE case report&lt;/a&gt;, &lt;a href="https://support.google.com/adspolicy/answer/6015406?hl=en" rel="noopener noreferrer"&gt;Google Ads policy documentation&lt;/a&gt;, &lt;a href="https://support.google.com/webmasters/answer/9044101?hl=en" rel="noopener noreferrer"&gt;Search Console security issues&lt;/a&gt;, &lt;a href="https://developer.apple.com/documentation/security/notarizing_macos_software_before_distribution" rel="noopener noreferrer"&gt;Apple notarization documentation&lt;/a&gt;, and &lt;a href="https://transparencyreport.google.com/safe-browsing/overview" rel="noopener noreferrer"&gt;Google Safe Browsing status&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;Originally published on &lt;a href="https://dispatch-blog.hashnode.dev/when-safe-browsing-says-clean-but-google-ads-says-malware" rel="noopener noreferrer"&gt;Dispatch&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>python</category>
      <category>programming</category>
      <category>tutorial</category>
      <category>ai</category>
    </item>
    <item>
      <title>Coding Agents Do Not Test Like Experts: What Agentic Verification Actually Needs</title>
      <dc:creator>Chen Yuan</dc:creator>
      <pubDate>Tue, 08 Sep 2026 13:13:03 +0000</pubDate>
      <link>https://dev.to/chenyuan20509/coding-agents-do-not-test-like-experts-what-agentic-verification-actually-needs-laj</link>
      <guid>https://dev.to/chenyuan20509/coding-agents-do-not-test-like-experts-what-agentic-verification-actually-needs-laj</guid>
      <description>&lt;p&gt;A large evaluation of coding agents produced an awkward result. Naming a testing technique did not reliably make the agents use that technique well. In the Zstd implementation study described by &lt;a href="https://danluu.com/agentic-testing/" rel="noopener noreferrer"&gt;Dan Luu&lt;/a&gt;, agents were given instructions for TDD, QuickCheck, property-based testing, fuzzing, differential testing, mutation testing, formal methods, and other tools. The default condition, with no extra testing instruction, performed above average. Several named techniques performed worse.&lt;/p&gt;

&lt;p&gt;The point is not that testing is useless. The point is that a test library name is not a testing method. Agents can produce tests that compile, run, and miss the bug that matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  The uncomfortable result is not a missing test command
&lt;/h2&gt;

&lt;p&gt;Coding agents already know that a finished implementation should have tests. They can create a test module, run the project command, repair a failing assertion, and report a green result. That workflow looks like verification from a distance. It often checks only whether the agent's implementation agrees with the agent's own assumptions.&lt;/p&gt;

&lt;p&gt;Luu's evaluation makes that gap visible. Agents did not simply ignore every instruction. They often used the requested library or wrote something that resembled the requested technique. The problem was the missing judgment underneath it. A property-based test with a weak generator is still a weak test. A formal proof about an irrelevant invariant does not protect the code path where the defect lives. A differential test that implements the same mistake twice creates agreement, not an oracle.&lt;/p&gt;

&lt;p&gt;That distinction matters more as agents take on larger changes. A human reviewer can ask whether a test attacks the risky assumption. An agent tends to optimize for a local signal: a command finished successfully, the assertion passed, or the requested tool appeared somewhere in the diff. The harness must make the stronger question visible.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the evaluation actually measured
&lt;/h2&gt;

&lt;p&gt;The primary evaluation reused a Rust implementation task for Zstd. It compared 26 prompt conditions, including Default, TDD, QuickCheck, property-based testing, fuzzing, differential testing, mutation testing, Lean 4, Verus, Alloy, and other formal or testing tools. Four additional skills were tested. Correctness came from hidden tests rather than from the tests the agent wrote for itself.&lt;/p&gt;

&lt;p&gt;The main graph averaged 80 runs for each condition and effort level and compared medium with higher-effort runs. The article also discusses an IMAP RFC evaluation. The exact ranking should not be treated as a universal league table. The task, model, harness, prompt wording, and available dependencies all affect the result. The author repeatedly warns that the observed differences are noisy and that agents often failed to apply a technique in a meaningful way.&lt;/p&gt;

&lt;p&gt;The broad pattern is more useful than the ordering. Default performed above average. At higher effort, fuzzing and property-based conditions did somewhat better than formal methods on average, while the medium-effort picture was mixed. The testing skills recommended by the model underperformed, while a small custom skill did better because it pushed agents toward examining likely mistakes instead of presenting a long tutorial.&lt;/p&gt;

&lt;p&gt;This is a study of minimal instructions, not a study of what an expert can achieve with Verus, QuickCheck, or TDD. It asks what happens when an agent receives a technique label and limited guidance. That is close to how many teams deploy skills today, which is why the limitation is also the practical lesson.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why naming a technique is not using it
&lt;/h2&gt;

&lt;p&gt;QuickCheck was a clear example. The agents generally wrote simple smoke tests, used random inputs that often followed rejection paths, and checked few meaningful properties. In 63 of 160 QuickCheck runs discussed in the study, only one property was checked. The library was present, but the reasoning that makes property-based testing useful was absent.&lt;/p&gt;

&lt;p&gt;The formal tools showed a similar split. Verus agents often proved abstract arithmetic facts instead of properties tied to the implementation. Some proofs were effectively of the form &lt;code&gt;A =&amp;gt; A&lt;/code&gt;: valid, but not useful for finding a Zstd defect. Lean, Alloy, and related conditions also relied heavily on ordinary Rust tests while adding proofs or models that did not reach the risky behavior.&lt;/p&gt;

&lt;p&gt;TDD changed the workflow more visibly. Agents wrote about twice as many tests and created failing tests earlier, but the condition still underperformed. More test-first activity did not guarantee better cases. In the four-stream jump-table feature, agents often made all four streams identical. That is a legal fixture for a simple path, but it cannot reveal a bug that appears only when the streams differ.&lt;/p&gt;

&lt;p&gt;None of this is an indictment of the tools. It is an indictment of the assumption that a tool name contains its own operating knowledge. An expert knows how to choose generators, construct an oracle, shrink a failure, inspect a counterexample, and decide whether an invariant touches the risk. A prompt that says “use QuickCheck” supplies none of that.&lt;/p&gt;

&lt;h2&gt;
  
  
  The difference between test presence and test power
&lt;/h2&gt;

&lt;p&gt;Test presence is easy to count. Test power is harder to observe. A suite can contain hundreds of assertions and still leave the important behavior unconstrained.&lt;/p&gt;

&lt;p&gt;The Zstd examples show why. Agents noticed that bitstream reversal could be risky, then used palindromic inputs for a relevant test. Reversing a palindrome produces the same sequence, so an encode/decode implementation with the wrong order can pass. The test is related to the feature but powerless against the defect.&lt;/p&gt;

&lt;p&gt;The same problem appeared in tests for four Huffman streams. Making every stream identical exercises the shape of the input without exercising the relationship between distinct streams. A fixture that looks realistic can still erase the dimension where the bug occurs.&lt;/p&gt;

&lt;p&gt;Differential testing adds another warning. It works when independent implementations receive the same inputs and disagree. In the study, most agents that attempted it did not build two genuinely independent implementations. They wrote the same idea twice, allowing the same wrong assumption to appear in both versions. Agreement then became a false signal of correctness.&lt;/p&gt;

&lt;p&gt;A passing test proves only that one execution satisfied one assertion. It does not prove that the assertion represents the specification, that the fixture reaches the risky state, or that the oracle is independent from the code under test. Those are separate questions, and an agent needs a separate artifact for each one.&lt;/p&gt;

&lt;h2&gt;
  
  
  A verification loop an agent can follow
&lt;/h2&gt;

&lt;p&gt;A useful verification loop starts with a claim instead of a command. The agent states what must remain true, identifies the failure mode that would violate it, chooses an oracle, and records the evidence. A small artifact can make that sequence explicit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;verification&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;claim&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;decode(encode(data)) equals data for valid inputs&lt;/span&gt;
  &lt;span class="na"&gt;risk&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;bitstream order can be reversed in one direction&lt;/span&gt;
  &lt;span class="na"&gt;targeted_cases&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;empty input&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;one-symbol input&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;non-palindromic multi-block input&lt;/span&gt;
  &lt;span class="na"&gt;adversarial_case&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;distinct values in every stream&lt;/span&gt;
  &lt;span class="na"&gt;independent_oracle&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;reference decoder or specification&lt;/span&gt;
  &lt;span class="na"&gt;required_evidence&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;test result&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;mutation result&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;failure log when a case is rejected&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The claim gives the agent a target that can be reviewed. The targeted cases connect the claim to boundaries and ordinary behavior. The adversarial case tries to falsify the agent's interpretation. The independent oracle prevents the implementation from grading its own homework. The evidence record makes it possible to inspect what the agent actually checked.&lt;/p&gt;

&lt;p&gt;The adversarial case should be shaped around the suspected failure, not chosen because it is convenient to serialize. For a stream-order bug, a non-palindromic input is more informative than another randomly generated input that happens to collapse to the same path. For a parser, malformed delimiters and ambiguous boundaries may matter more than another valid example.&lt;/p&gt;

&lt;p&gt;A small Rust test can express the difference without pretending to be a complete proof:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="nd"&gt;#[test]&lt;/span&gt;
&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;distinct_streams_exercise_ordering&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;input&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nn"&gt;StreamSet&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;new&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;
        &lt;span class="nd"&gt;vec!&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="nd"&gt;vec!&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="nd"&gt;vec!&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;13&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;21&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;34&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="nd"&gt;vec!&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;55&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;]);&lt;/span&gt;

    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;encoded&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;input&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;decoded&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;reference_decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;encoded&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="nd"&gt;assert_eq!&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;decoded&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;input&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This test is stronger than a palindrome because each stream carries different structure. It still does not prove that the encoder is correct for every valid input. It only supplies one independent, reviewable check aimed at a particular failure mode.&lt;/p&gt;

&lt;p&gt;The loop should also have a stop condition. If a mutation survives, the agent must either add a case that kills it or explain why the mutation is outside the specification. A green command with surviving mutations is not a green verification result.&lt;/p&gt;

&lt;h2&gt;
  
  
  Turn verification into an observable interface
&lt;/h2&gt;

&lt;p&gt;A prompt or skill can nudge an agent, but the environment decides what the agent can learn from failure. The evaluation found that agents could name risky areas such as FSE, Huffman coding, and bit readers without testing them effectively. That is a feedback problem as much as a reasoning problem.&lt;/p&gt;

&lt;p&gt;A harness should separate implementation context from verification context where practical. It should provide bounded tools, expose mutation results, preserve failing inputs, and make the independent oracle easy to call. It should also stop runaway work and attach each artifact to a change identifier. The agent then receives information that a longer instruction cannot provide.&lt;/p&gt;

&lt;p&gt;Here is a compact interface sketch:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;VerificationHarness&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;oracle&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;mutation_tool&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;fuzzer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;threshold&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;oracle&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;oracle&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;mutation_tool&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;mutation_tool&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;fuzzer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;fuzzer&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;threshold&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;threshold&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;verify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;implementation&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;specification&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;claim&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;propose_claim&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;specification&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;tests&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;generate_targeted_tests&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;implementation&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;claim&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="nf"&gt;oracle_checks&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;oracle&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tests&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;specification&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;Reject&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tests do not check the specification&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="n"&gt;mutation_score&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;mutation_tool&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;implementation&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tests&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;mutation_score&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;threshold&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;Reject&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;important mutations survived&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="n"&gt;fuzz_result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;fuzzer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;implementation&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;fuzz_result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;found_failure&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;Reject&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;fuzzer found: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;fuzz_result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;case&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;Accept&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;claim&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;claim&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tests&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;tests&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;mutation_score&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;mutation_score&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The mutation score is not a correctness certificate. It is a pressure signal: tests that merely mirror the implementation should fail to kill useful mutations. Fuzzing is not a certificate either. It adds cases and failure modes that the agent did not choose by hand. The oracle check asks a different question again.&lt;/p&gt;

&lt;p&gt;Technique selection should follow the risk and the available oracle. Differential testing fits a protocol with a trusted reference implementation. Property-based testing fits a domain with a clear invariant and a useful generator. Fuzzing fits parsers and state machines with robust crash or semantic oracles. Mutation testing evaluates the tests themselves. Formal methods fit a specification that is precise enough to prove. TDD can structure feedback, but it does not replace hard-case design.&lt;/p&gt;

&lt;h2&gt;
  
  
  What teams should change first
&lt;/h2&gt;

&lt;p&gt;Start with one high-risk path, not an attempt to impose every testing technique on every task. Write down the independent oracle before asking an agent to implement the path. If the oracle cannot be described, the team is not ready to treat a green agent-generated suite as strong evidence.&lt;/p&gt;

&lt;p&gt;Require an adversarial case tied to the likely defect. Do not accept “edge cases included” as a summary. Name the input dimension, state transition, permission boundary, or protocol ambiguity that the case exercises. A reviewer should be able to see why the case could distinguish a correct implementation from a plausible wrong one.&lt;/p&gt;

&lt;p&gt;Record verification artifacts with the change. Keep the claim, test rationale, adversarial input, oracle source, mutation result, and preserved failures. This turns review from “the agent says tests pass” into an inspection of the evidence. It also gives a later agent something concrete to revisit when a test fails in production.&lt;/p&gt;

&lt;p&gt;Measure escaped defects and holdout failures instead of generated test counts. The study's TDD condition wrote more tests without improving correctness. Counting files or assertions rewards activity, not constraint. Run holdout cases that the agent did not see, mutate the implementation, and compare results against a reference where one exists.&lt;/p&gt;

&lt;p&gt;The practical fix is not a longer prompt. It is a verification system that makes weak tests fail visibly, gives the agent an oracle it cannot rewrite, and leaves a reviewer enough evidence to judge the gap between tests that exist and tests that matter.&lt;/p&gt;




&lt;p&gt;Originally published on &lt;a href="https://dispatch-blog.hashnode.dev/coding-agents-do-not-test-like-experts-what-agentic-verification-actually-needs" rel="noopener noreferrer"&gt;Dispatch&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>python</category>
      <category>programming</category>
      <category>tutorial</category>
      <category>ai</category>
    </item>
    <item>
      <title>Rust Trait Objects Are Just a Pointer and a Vtable</title>
      <dc:creator>Chen Yuan</dc:creator>
      <pubDate>Sun, 06 Sep 2026 06:51:27 +0000</pubDate>
      <link>https://dev.to/chenyuan20509/rust-trait-objects-are-just-a-pointer-and-a-vtable-400l</link>
      <guid>https://dev.to/chenyuan20509/rust-trait-objects-are-just-a-pointer-and-a-vtable-400l</guid>
      <description>&lt;p&gt;A trait object in Rust is a fat pointer. That statement is precise, not metaphorical. Every &lt;code&gt;&amp;amp;dyn Trait&lt;/code&gt;, &lt;code&gt;Box&amp;lt;dyn Trait&amp;gt;&lt;/code&gt;, and &lt;code&gt;*const dyn Trait&lt;/code&gt; occupies two machine words: one word holds the address of the concrete value, and the other holds the address of a vtable. The vtable is a static structure generated by the compiler for each concrete type that implements the trait. This layout is the entire mechanism behind dynamic dispatch in Rust.&lt;/p&gt;

&lt;p&gt;The Hacker News discussion around the article "Visualizing Rust's Vtables: How dyn Trait Works In Memory" surfaced a common point of confusion: developers coming from C++ expect the vtable pointer to live inside the object itself. In Rust, it does not. The vtable pointer travels with the reference, not with the data. This distinction changes how you reason about memory layout, object size, and API design.&lt;/p&gt;

&lt;h2&gt;
  
  
  The two words that explain a trait object
&lt;/h2&gt;

&lt;p&gt;A trait object is not a type in the same sense as a struct or an enum. It is a dynamically sized type (DST). The compiler does not know its size at compile time, so you cannot store a &lt;code&gt;dyn Trait&lt;/code&gt; directly on the stack or in a struct field without indirection. You must always use a pointer: &lt;code&gt;&amp;amp;dyn Trait&lt;/code&gt;, &lt;code&gt;Box&amp;lt;dyn Trait&amp;gt;&lt;/code&gt;, &lt;code&gt;Rc&amp;lt;dyn Trait&amp;gt;&lt;/code&gt;, or &lt;code&gt;*const dyn Trait&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The pointer itself is the key. A regular pointer in Rust, such as &lt;code&gt;&amp;amp;Circle&lt;/code&gt;, is a single machine word containing an address. A trait object pointer, such as &lt;code&gt;&amp;amp;dyn Draw&lt;/code&gt;, is two machine words. The first word points to the concrete data. The second word points to the vtable for that concrete type's implementation of the trait.&lt;/p&gt;

&lt;p&gt;This two-word representation is why &lt;code&gt;std::mem::size_of::&amp;lt;&amp;amp;dyn Draw&amp;gt;()&lt;/code&gt; returns 16 on a 64-bit platform while &lt;code&gt;std::mem::size_of::&amp;lt;&amp;amp;Circle&amp;gt;()&lt;/code&gt; returns 8. The data pointer and the vtable pointer together form what the Rust community calls a fat pointer.&lt;/p&gt;

&lt;p&gt;The fat pointer does not contain a length, unlike a slice fat pointer which stores a pointer and a length. For trait objects, the metadata is a vtable pointer, not a count. The &lt;code&gt;Pointee&lt;/code&gt; trait in the standard library formalises this: for &lt;code&gt;dyn Trait&lt;/code&gt;, the &lt;code&gt;Metadata&lt;/code&gt; associated type is &lt;code&gt;DynMetadata&amp;lt;dyn Trait&amp;gt;&lt;/code&gt;, which is a pointer to the vtable.&lt;/p&gt;

&lt;h2&gt;
  
  
  What lives in the fat pointer
&lt;/h2&gt;

&lt;p&gt;The first component of a fat pointer is straightforward: it points to the concrete value. That value can be anywhere—stack, heap, or static memory. The pointer is raw address information, nothing more.&lt;/p&gt;

&lt;p&gt;The second component points to a vtable. The vtable is a static structure, typically placed in read-only data segments, that contains the function pointers for all the trait's methods, plus some additional metadata required by the language.&lt;/p&gt;

&lt;p&gt;The vtable layout is an implementation detail of the Rust compiler, not a stable language guarantee. However, the general structure is well understood from compiler source and documentation. A vtable begins with three entries:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A pointer to the &lt;code&gt;drop_in_place&lt;/code&gt; implementation for the concrete type. This is how &lt;code&gt;Box&amp;lt;dyn Trait&amp;gt;&lt;/code&gt; knows how to drop the value when it goes out of scope.&lt;/li&gt;
&lt;li&gt;The size of the concrete type, in bytes.&lt;/li&gt;
&lt;li&gt;The alignment of the concrete type, in bytes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;After these three header entries, the vtable contains one function pointer for each method defined in the trait, in declaration order. For traits with supertraits, the vtable may also contain pointers to the supertrait vtables to support trait upcasting coercion.&lt;/p&gt;

&lt;p&gt;Each concrete type gets its own vtable for each trait it implements. If a type implements both &lt;code&gt;Draw&lt;/code&gt; and &lt;code&gt;Clone&lt;/code&gt;, the compiler generates a &lt;code&gt;Draw&lt;/code&gt; vtable and a separate &lt;code&gt;Clone&lt;/code&gt; vtable. When you coerce a &lt;code&gt;&amp;amp;Circle&lt;/code&gt; to &lt;code&gt;&amp;amp;dyn Draw&lt;/code&gt;, the fat pointer uses the &lt;code&gt;Draw&lt;/code&gt; vtable. When you coerce the same &lt;code&gt;&amp;amp;Circle&lt;/code&gt; to &lt;code&gt;&amp;amp;dyn Clone&lt;/code&gt;, it uses the &lt;code&gt;Clone&lt;/code&gt; vtable.&lt;/p&gt;

&lt;p&gt;All instances of the same concrete type share the same vtable. The vtable is per-(type, trait) pair, not per-instance.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reading a vtable without guessing
&lt;/h2&gt;

&lt;p&gt;The standard library documentation exposes &lt;code&gt;std::ptr::DynMetadata&lt;/code&gt; as the conceptual type for trait-object metadata, but the pointer-metadata APIs are still nightly-only in the current stable documentation. That distinction matters: the type is useful vocabulary for understanding what the compiler carries, but it is not a portable stable inspection interface. On stable Rust, &lt;code&gt;std::mem::size_of::&amp;lt;&amp;amp;dyn Draw&amp;gt;()&lt;/code&gt; and ordinary coercions let you reason about the representation without extracting metadata.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;use&lt;/span&gt; &lt;span class="nn"&gt;std&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;mem&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;trait&lt;/span&gt; &lt;span class="n"&gt;Speak&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;speak&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="k"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;struct&lt;/span&gt; &lt;span class="n"&gt;Dog&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;impl&lt;/span&gt; &lt;span class="n"&gt;Speak&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;Dog&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;speak&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="k"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="s"&gt;"woof"&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;dog&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Dog&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;trait_obj&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="k"&gt;dyn&lt;/span&gt; &lt;span class="n"&gt;Speak&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;dog&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

    &lt;span class="nd"&gt;println!&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"trait object reference: {} bytes"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nn"&gt;mem&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;size_of_val&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;trait_obj&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
    &lt;span class="nd"&gt;println!&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"concrete reference: {} bytes"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nn"&gt;mem&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nn"&gt;size_of&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;Dog&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On a 64-bit target, the first line reports 16 and the second reports 8. The example measures the reference representation, not the size of the &lt;code&gt;Dog&lt;/code&gt; value. It is stable and safe, but it does not expose the vtable pointer or its function slots.&lt;/p&gt;

&lt;p&gt;Reading the actual function pointers from a vtable is not a stable operation. The vtable layout is not guaranteed across Rust versions or target platforms. Any attempt to read vtable slots by offset is a debugging experiment, not production code. Even comparing vtable pointer values is unreliable: the compiler may duplicate or merge vtables during code generation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Static dispatch and dynamic dispatch side by side
&lt;/h2&gt;

&lt;p&gt;Static dispatch, using generics, produces a separate monomorphised copy of the function for each concrete type. The compiler knows exactly which &lt;code&gt;draw&lt;/code&gt; implementation to call at each call site. There is no vtable, no fat pointer, and no runtime overhead for dispatch.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;trait&lt;/span&gt; &lt;span class="n"&gt;Draw&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;draw&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="k"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;struct&lt;/span&gt; &lt;span class="n"&gt;Circle&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;struct&lt;/span&gt; &lt;span class="n"&gt;Square&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;impl&lt;/span&gt; &lt;span class="n"&gt;Draw&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;Circle&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;draw&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="k"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="s"&gt;"Circle"&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;impl&lt;/span&gt; &lt;span class="n"&gt;Draw&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;Square&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;draw&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="k"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="s"&gt;"Square"&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// Static dispatch: monomorphised per type.&lt;/span&gt;
&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="n"&gt;draw_statically&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;T&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Draw&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;T&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;shape&lt;/span&gt;&lt;span class="nf"&gt;.draw&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Circle&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Square&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="nd"&gt;println!&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"{}"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;draw_statically&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
    &lt;span class="nd"&gt;println!&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"{}"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;draw_statically&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The compiler generates &lt;code&gt;draw_statically::&amp;lt;Circle&amp;gt;&lt;/code&gt; and &lt;code&gt;draw_statically::&amp;lt;Square&amp;gt;&lt;/code&gt; as separate functions. Each function contains a direct call to the corresponding &lt;code&gt;draw&lt;/code&gt; implementation.&lt;/p&gt;

&lt;p&gt;Dynamic dispatch, using &lt;code&gt;dyn Trait&lt;/code&gt;, uses a single function that accepts a fat pointer and dispatches through the vtable.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;draw_dynamically&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="k"&gt;dyn&lt;/span&gt; &lt;span class="n"&gt;Draw&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;shape&lt;/span&gt;&lt;span class="nf"&gt;.draw&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Circle&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Square&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;shapes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="k"&gt;dyn&lt;/span&gt; &lt;span class="n"&gt;Draw&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;shape&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;shapes&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nd"&gt;println!&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"{}"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;draw_dynamically&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;draw_dynamically&lt;/code&gt; function is not monomorphised. There is one copy of the function. Each call to &lt;code&gt;shape.draw()&lt;/code&gt; loads the vtable pointer from the fat pointer, finds the method entry for &lt;code&gt;Draw::draw&lt;/code&gt;, and calls that function with the data pointer as the &lt;code&gt;self&lt;/code&gt; argument.&lt;/p&gt;

&lt;p&gt;Dynamic dispatch adds an indirect call and can limit inlining compared with a statically known concrete type. The actual cost depends on the surrounding code, cache behaviour, and whether the compiler can remove or specialise the abstraction. It should be measured in the workload that matters, not reduced to a universal number of loads.&lt;/p&gt;

&lt;p&gt;The choice between static and dynamic dispatch is a trade-off between optimisation opportunities and flexibility. Static dispatch enables inlining and devirtualisation when the compiler can see the concrete type, but requires compile-time knowledge of the participating types. Dynamic dispatch allows heterogeneous collections and runtime polymorphism, but gives up some of those opportunities.&lt;/p&gt;

&lt;h2&gt;
  
  
  Object safety is an API boundary
&lt;/h2&gt;

&lt;p&gt;Not every trait can be used as a trait object. The compiler enforces a set of rules called object safety, also referred to in recent compiler versions as "dyn compatibility". These rules determine whether a trait can be converted to a &lt;code&gt;dyn Trait&lt;/code&gt; type.&lt;/p&gt;

&lt;p&gt;A trait is dyn-compatible if its methods can be called through a trait object. The most important rules are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Methods cannot have generic type parameters unless the method is restricted with &lt;code&gt;where Self: Sized&lt;/code&gt;. A vtable entry must represent one callable signature, while a generic method would require monomorphisation for each type argument.&lt;/li&gt;
&lt;li&gt;Methods cannot return &lt;code&gt;Self&lt;/code&gt; by value unless they are similarly restricted. A trait object does not know the erased concrete size needed for such a return.&lt;/li&gt;
&lt;li&gt;A method cannot use &lt;code&gt;Self&lt;/code&gt; in a position that requires the erased type's size, such as a by-value argument, unless the method is restricted with &lt;code&gt;where Self: Sized&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Associated constants are not dyn-compatible. Associated types are allowed when the trait object specifies the associated type, as in &lt;code&gt;dyn Iterator&amp;lt;Item = u8&amp;gt;&lt;/code&gt;.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Dyn-compatible.&lt;/span&gt;
&lt;span class="k"&gt;trait&lt;/span&gt; &lt;span class="n"&gt;Safe&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;do_something&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="k"&gt;self&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;do_another&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="k"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;i32&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// Not dyn-compatible: generic method.&lt;/span&gt;
&lt;span class="k"&gt;trait&lt;/span&gt; &lt;span class="n"&gt;UnsafeGeneric&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="n"&gt;generic&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;T&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="k"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;T&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// Not dyn-compatible: associated constant.&lt;/span&gt;
&lt;span class="k"&gt;trait&lt;/span&gt; &lt;span class="n"&gt;UnsafeConst&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="n"&gt;KIND&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="k"&gt;'static&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// Dyn-compatible when the object supplies Output.&lt;/span&gt;
&lt;span class="k"&gt;trait&lt;/span&gt; &lt;span class="n"&gt;HasOutput&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;type&lt;/span&gt; &lt;span class="n"&gt;Output&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="k"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="k"&gt;Self&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;Output&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="c1"&gt;// A usable object type can be written as dyn HasOutput&amp;lt;Output = i32&amp;gt;.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The compiler will refuse to coerce a value to &lt;code&gt;&amp;amp;dyn UnsafeGeneric&lt;/code&gt; or &lt;code&gt;Box&amp;lt;dyn UnsafeConst&amp;gt;&lt;/code&gt;. A trait with an associated type is not automatically rejected; the object type may need to bind that type before it can be used.&lt;/p&gt;

&lt;p&gt;The object safety rules are an API boundary, not a performance optimisation. They exist because the compiler cannot generate one callable vtable entry for operations that still require compile-time type information. A vtable does not choose generic method instantiations or provide storage for a size-varying return. An associated type can work when the trait object fixes its value in the object type.&lt;/p&gt;

&lt;p&gt;The Rust reference explicitly defines object safety in terms of what operations can be performed on a trait object. A trait object permits "late binding" of methods, dispatched using vtables. If a method cannot be dispatched through a vtable, the trait is not object safe.&lt;/p&gt;

&lt;h2&gt;
  
  
  A small memory probe in Rust
&lt;/h2&gt;

&lt;p&gt;You can observe the fat pointer layout directly using &lt;code&gt;std::mem::transmute&lt;/code&gt; or by casting to a raw pointer and reading the two words. This is an unsafe debugging technique, not a stable API. The following example demonstrates the layout on a 64-bit platform.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="k"&gt;use&lt;/span&gt; &lt;span class="nn"&gt;std&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;mem&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;trait&lt;/span&gt; &lt;span class="n"&gt;Identify&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;id&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="k"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;struct&lt;/span&gt; &lt;span class="nf"&gt;Item&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;u64&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;impl&lt;/span&gt; &lt;span class="n"&gt;Identify&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;Item&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;id&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="k"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="s"&gt;"Item"&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;Item&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;trait_ptr&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="k"&gt;dyn&lt;/span&gt; &lt;span class="n"&gt;Identify&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

    &lt;span class="c1"&gt;// A fat pointer is two words.&lt;/span&gt;
    &lt;span class="nd"&gt;println!&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"size of &amp;amp;dyn Identify: {} bytes"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nn"&gt;mem&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nn"&gt;size_of&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&amp;amp;&lt;/span&gt;&lt;span class="k"&gt;dyn&lt;/span&gt; &lt;span class="n"&gt;Identify&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;
    &lt;span class="c1"&gt;// On 64-bit: 16 bytes.&lt;/span&gt;

    &lt;span class="c1"&gt;// Unsafe inspection: read the two words of the fat pointer.&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data_ptr&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;vtable_ptr&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="p"&gt;())&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;unsafe&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nn"&gt;mem&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;transmute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;trait_ptr&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;};&lt;/span&gt;

    &lt;span class="nd"&gt;println!&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"data pointer: {:p}"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;data_ptr&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="nd"&gt;println!&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"vtable pointer: {:p}"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;vtable_ptr&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="c1"&gt;// The data pointer points to the `item` variable.&lt;/span&gt;
    &lt;span class="c1"&gt;// The vtable pointer points to a static vtable for Item + Identify.&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;transmute&lt;/code&gt; call is unsafe because it bypasses Rust's type system. The language does not promise a stable vtable ABI or a stable ordering for private vtable slots. This code is useful when investigating a particular compiler build, but it is not a general FFI contract. The old &lt;code&gt;core::raw::TraitObject&lt;/code&gt; pattern should not be treated as a supported public API.&lt;/p&gt;

&lt;p&gt;The pointer-metadata APIs describe the same split more explicitly, but they are still nightly-only in the current standard-library documentation. A nightly experiment can reconstruct a raw trait-object pointer from matching parts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="nd"&gt;#![feature(ptr_metadata)]&lt;/span&gt;

&lt;span class="k"&gt;use&lt;/span&gt; &lt;span class="nn"&gt;std&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;ptr&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;inspect_again&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;trait_ptr&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="k"&gt;dyn&lt;/span&gt; &lt;span class="n"&gt;Identify&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;metadata&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nn"&gt;ptr&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;trait_ptr&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;trait_ptr&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;reconstructed&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="k"&gt;dyn&lt;/span&gt; &lt;span class="n"&gt;Identify&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nn"&gt;ptr&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;from_raw_parts&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="c1"&gt;// The metadata must belong to the same erased trait-object type.&lt;/span&gt;
    &lt;span class="nd"&gt;assert_eq!&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;trait_ptr&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;reconstructed&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="p"&gt;());&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The functions are safe in the narrow sense that constructing the raw pointer does not dereference it, but a caller must still ensure that the data pointer and metadata describe a valid &lt;code&gt;Identify&lt;/code&gt; object before dereferencing the result. That is a compiler-experiment boundary, not a stable application interface.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rules for using dyn Trait deliberately
&lt;/h2&gt;

&lt;p&gt;Use &lt;code&gt;dyn Trait&lt;/code&gt; when you need runtime polymorphism in a collection or when the set of types is not known at compile time. A &lt;code&gt;Vec&amp;lt;Box&amp;lt;dyn Draw&amp;gt;&amp;gt;&lt;/code&gt; can hold circles, squares, and triangles together. A &lt;code&gt;Vec&amp;lt;Circle&amp;gt;&lt;/code&gt; cannot.&lt;/p&gt;

&lt;p&gt;Avoid &lt;code&gt;dyn Trait&lt;/code&gt; when the concrete type is known at compile time and you do not need heterogeneous storage. Static dispatch through generics produces faster code and enables inlining.&lt;/p&gt;

&lt;p&gt;When designing a trait that you intend to use as a trait object, ensure object safety from the start. Generic methods and &lt;code&gt;Self&lt;/code&gt;-returning methods are the most common blockers. If you need a trait to be object safe and also support generic methods, consider splitting the trait into an object-safe core and a separate generic extension trait.&lt;/p&gt;

&lt;p&gt;The vtable pointer in a fat pointer is not a stable ABI. Different Rust versions or different compiler flags may change the vtable layout. Code that relies on a specific vtable offset is fragile and should be confined to unsafe debugging or specialised crates that track compiler internals.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;DynMetadata&lt;/code&gt; API provides stable access to size and alignment. Use it when you need to know the size of the concrete type behind a trait object. Do not attempt to read the function pointer slots directly in production code.&lt;/p&gt;

&lt;p&gt;The cost of dynamic dispatch is two additional memory indirections per method call. In performance-critical code paths, consider using an enum over a fixed set of types instead of &lt;code&gt;dyn Trait&lt;/code&gt;. An enum allows static dispatch through pattern matching while still supporting heterogeneous values.&lt;/p&gt;

&lt;p&gt;Object safety is a property of the trait definition, not the implementation. You cannot make an object-unsafe trait object-safe by changing the implementation. The trait's method signatures must satisfy the object safety rules for any &lt;code&gt;dyn Trait&lt;/code&gt; to exist. If you control the trait, design it to be object safe. If you do not control the trait, use generics or wrapper types.&lt;/p&gt;

&lt;p&gt;The fat pointer representation of trait objects is one of Rust's most elegant design decisions. It separates the mechanism of polymorphism from the data itself, keeps objects small, and enables zero-cost abstractions where static dispatch is sufficient. A trait object is just a pointer and a vtable. Everything else follows from that fact.&lt;/p&gt;




&lt;p&gt;Originally published on &lt;a href="https://dispatch-blog.hashnode.dev/rust-trait-objects-are-just-a-pointer-and-a-vtable" rel="noopener noreferrer"&gt;Dispatch&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>python</category>
      <category>programming</category>
      <category>tutorial</category>
      <category>ai</category>
    </item>
    <item>
      <title>Inside the Sandbox Is Still Code Execution: Chromium CVE-2026-85046 for Agent Operators</title>
      <dc:creator>Chen Yuan</dc:creator>
      <pubDate>Sat, 05 Sep 2026 14:17:14 +0000</pubDate>
      <link>https://dev.to/chenyuan20509/inside-the-sandbox-is-still-code-execution-chromium-cve-2026-85046-for-agent-operators-enm</link>
      <guid>https://dev.to/chenyuan20509/inside-the-sandbox-is-still-code-execution-chromium-cve-2026-85046-for-agent-operators-enm</guid>
      <description>&lt;p&gt;Is a renderer compromise only a containment failure, or is it already a breach of the agent's trust boundary? CVE-2026-85046 makes that contrast concrete. The bug is a type confusion flaw in Chromium's V8 engine. Google patched it in Chrome 152.0.7977.82 after confirming it was exploited in the wild. The NVD description is narrower than the headlines that followed: a crafted HTML page can run arbitrary code inside the sandbox. That sentence is easy to read as a relief. For an AI agent that browses the open web with a logged-in session, it is the opposite.&lt;/p&gt;

&lt;p&gt;A sandbox is a containment layer. It limits what a renderer process may do to the rest of the machine. A trust boundary is the set of actions the agent is allowed to take as the user. Those two lines are not the same. In-sandbox code execution crosses the second line without needing to cross the first.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the advisory actually says
&lt;/h2&gt;

&lt;p&gt;The NVD entry for CVE-2026-85046 names the component, the class of bug, the fixed build, and the impact. The component is V8. The class is type confusion, catalogued as CWE-843. The fixed build is Chrome 152.0.7977.82. The impact is remote code execution inside the sandbox from a crafted HTML page. Chromium rated the issue High. Public writeups credit Salvatore Gulizia, also known as Serotav, with the report.&lt;/p&gt;

&lt;p&gt;Type confusion is not a logic error in a web API. V8 represents JavaScript values with hidden classes and typed slots. A successful exploit creates a value under one type and later forces the engine to read those slots as another type. Once that happens, attacker-controlled JavaScript can treat a small integer as a pointer, or a pointer as a length. From there the usual memory-corruption path is available inside the renderer: arbitrary read, arbitrary write, then control of the instruction pointer in that process.&lt;/p&gt;

&lt;p&gt;The important qualifier is the process, not the machine. Chromium splits work across a browser process and one or more renderers. Site isolation puts different sites in different renderers when it can. V8 lives in the renderer, so the bug starts in the process that already has the page's DOM and the cookies that page is allowed to touch.&lt;/p&gt;

&lt;p&gt;"Inside the sandbox" describes the operating-system jail around that renderer. It does not mean the page, the session, or the agent remains honest.&lt;/p&gt;

&lt;h2&gt;
  
  
  In-sandbox code execution is not a sandbox escape
&lt;/h2&gt;

&lt;p&gt;The sandbox is an operating-system policy around the renderer: restricted tokens and job objects on Windows, seatbelt on macOS, namespaces and seccomp on Linux. Those controls are meant to stop a hostile renderer from opening arbitrary files or talking to the rest of the desktop as an equal.&lt;/p&gt;

&lt;p&gt;CVE-2026-85046 does not remove that jail. Code that runs after the type confusion still runs as the renderer. It still sees the page and still speaks to the browser process through the same IPC channels a legitimate renderer uses.&lt;/p&gt;

&lt;p&gt;That is a different bug class from a sandbox escape. The same Chrome release also fixed CVE-2026-85050, an out-of-bounds write in WebGL on Android that NVD describes as code execution outside the sandbox. A full browser exploit chain often wants both stages: first a renderer primitive, then a second bug that breaks the jail. CVE-2026-85046 is the first stage. CVE-2026-85050, on the platforms where it applies, is the second.&lt;/p&gt;

&lt;p&gt;Treating "no sandbox escape" as "no incident" collapses the two stages into one. For Chrome's security team the distinction is real. For an agent whose job is to click and read as the user, the first stage already steals the product. The attacker does not need a new host process if the existing renderer can drive the same session.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why agent browsers inherit the same bug
&lt;/h2&gt;

&lt;p&gt;Playwright, Puppeteer, Chrome DevTools Protocol clients, and extension-based automation do not ship a private JavaScript engine. They start or attach to Chromium. The V8 in that binary is the V8 in that Chrome build. If the build is older than 152.0.7977.82, the advisory applies.&lt;/p&gt;

&lt;p&gt;There are at least three common ways an agent gets a browser, and they do not update together.&lt;/p&gt;

&lt;p&gt;The first is the user's daily Chrome. Auto-update may patch it on a schedule the operator does not control. That is the binary people mean when they say "I already updated Chrome."&lt;/p&gt;

&lt;p&gt;The second is Chrome for Testing, or the Chromium revision Playwright downloads into a cache directory. Teams pin that binary so CI is reproducible. Pinning last month's revision is a stability choice. It is also a decision to keep last month's V8.&lt;/p&gt;

&lt;p&gt;The third is attaching to an already running profile. Hermes-style "use my logged-in browser" automation, Playwright's channel attach, and some CDP scripts skip the pinned testing binary and speak to the browser that already has cookies. That path inherits whatever build the user actually launched. It also inherits the session.&lt;/p&gt;

&lt;p&gt;A patched desktop Chrome therefore does not repair a pinned testing binary. A patched testing binary does not repair an old Electron shell. An updated library package does not replace a browser downloaded once and left in a cache. The bug lives in the executable that runs V8, not in the Python that called &lt;code&gt;browser.new_page()&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;

&lt;span class="n"&gt;MIN_CHROME&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;152&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;7977&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;82&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;parse_version&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;match&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;(\d+)\.(\d+)\.(\d+)\.(\d+)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;match&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;tuple&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;part&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;part&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;match&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;groups&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;chromium_version&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;browser_path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;browser_path&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--version&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;capture_output&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;check&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;returncode&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;parse_version&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stdout&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stderr&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;assert_patched&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;browser_path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;version&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;chromium_version&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;browser_path&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;version&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;SystemExit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;refusing to start: browser version is unknown&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;version&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;MIN_CHROME&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;SystemExit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;refusing to start: Chromium %s is older than 152.0.7977.82&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;part&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;part&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;version&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;version&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  A pinned browser is a security decision
&lt;/h2&gt;

&lt;p&gt;Pinning a browser revision is not a missing update. It is a choice to prefer a known binary over a moving one. That choice is reasonable for a test suite that only opens fixtures on localhost. It is a poor default for an agent that follows links, opens email HTML, or visits sites chosen by a model.&lt;/p&gt;

&lt;p&gt;Playwright's install step downloads the Chromium revision that matches the package. Upgrading the Python package without re-running the install leaves the old binary in the cache. Docker tags have the same shape: a tag that was current in August is not a patch for a September V8 advisory.&lt;/p&gt;

&lt;p&gt;The other pinning trap is flags. &lt;code&gt;--no-sandbox&lt;/code&gt; is still common in containers because it makes Chrome start under root. That flag does not create CVE-2026-85046, but it removes the containment layer the advisory still relies on. &lt;code&gt;--disable-web-security&lt;/code&gt; and &lt;code&gt;--disable-site-isolation-trials&lt;/code&gt; widen what a hostile renderer can reach after it already runs code. A version gate that ignores flags is a version gate that lies.&lt;/p&gt;

&lt;p&gt;A production agent should treat the binary path, the four-part version, the launch flags, and the profile as one reviewable object. If any field is unknown, the process should not open a URL. Unknown is not "probably fine." Unknown is "the check did not run."&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;sys&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;

&lt;span class="c1"&gt;# Reuses assert_patched() from the previous listing.
&lt;/span&gt;&lt;span class="n"&gt;FORBIDDEN_FLAGS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--no-sandbox&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--disable-web-security&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--disable-site-isolation-trials&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;assert_launch_safe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;browser_path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;launch_flags&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;version&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;assert_patched&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;browser_path&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;forbidden&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;FORBIDDEN_FLAGS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;intersection&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;launch_flags&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;forbidden&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;SystemExit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;refusing to start: forbidden flags %s&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="nf"&gt;sorted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;forbidden&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chromium %s accepted&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;part&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;part&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;version&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nb"&gt;file&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stderr&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;open_untrusted_url&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;page&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;browser_path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;launch_flags&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;assert_launch_safe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;browser_path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;launch_flags&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;startswith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;startswith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;http://&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;SystemExit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;refusing to open a non-http URL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;goto&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Containment is not the same as trust
&lt;/h2&gt;

&lt;p&gt;After a renderer compromise, the attacker has the same view the page had, plus the ability to ignore the page's own JavaScript and talk to the renderer internals. That is enough to read cookies the renderer is allowed to send, fill forms, click buttons, inspect the DOM, and intercept traffic that already terminates in that process.&lt;/p&gt;

&lt;p&gt;For a personal browsing agent the blast radius is the user's session. The agent was going to use that session anyway. The attacker now uses it first. Anything the renderer could already send, including cookies and filled forms, is in scope.&lt;/p&gt;

&lt;p&gt;For an attached profile the blast radius is larger still. The automation did not create a throwaway login. It borrowed the user's real Chrome. In-sandbox code execution then sits next to every other tab in that profile's renderer topology. Site isolation helps, but it is not a promise that a hostile origin cannot be opened, and it is not a promise that the agent's privileged tab is in a different process from the attacker's page.&lt;/p&gt;

&lt;p&gt;The sandbox still does useful work. It is why a renderer bug is not automatically a kernel bug, and why "steal the session" and "install a host implant" should stay different tickets. The mistake is to file the first one as "not an incident because the sandbox held."&lt;/p&gt;

&lt;p&gt;An agent is a program whose feature is to act. The renderer is where that acting happens. Control of the renderer is control of the feature. "Inside the sandbox" answers a different question: whether that control also reaches the rest of the disk as the host user.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a version check should prove
&lt;/h2&gt;

&lt;p&gt;A useful check answers four questions with evidence, not with a changelog.&lt;/p&gt;

&lt;p&gt;Which executable started? The path must be the path the agent will launch or attach to, not a &lt;code&gt;google-chrome --version&lt;/code&gt; from some other install on the same machine.&lt;/p&gt;

&lt;p&gt;Which four-part version did that executable print? Compare it as a tuple against 152.0.7977.82. String equality against a single build is too brittle, and a major-only check is too loose.&lt;/p&gt;

&lt;p&gt;Which flags were actually passed? Parse the launch argument list and fail closed on &lt;code&gt;--no-sandbox&lt;/code&gt;, &lt;code&gt;--disable-web-security&lt;/code&gt;, and &lt;code&gt;--disable-site-isolation-trials&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Which profile is attached? A throwaway user-data-dir is a different decision from the person's daily profile. If the policy forbids attaching to the real profile, a running Chrome with that profile is a failed start.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;browser_policy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;min_version&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;152.0.7977.82"&lt;/span&gt;
  &lt;span class="na"&gt;attach_user_profile&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
  &lt;span class="na"&gt;allowed_launch_flags&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--headless=new"&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--disable-dev-shm-usage"&lt;/span&gt;
  &lt;span class="na"&gt;disallowed_launch_flags&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--no-sandbox"&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--disable-web-security"&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--disable-site-isolation-trials"&lt;/span&gt;
  &lt;span class="na"&gt;on_unknown_version&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;block&lt;/span&gt;
  &lt;span class="na"&gt;on_forbidden_flag&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;block&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The policy is small on purpose. It does not detect exploits. It only refuses to start a browser whose version or flags the operator has not accepted, including when CI reuses an old cache.&lt;/p&gt;

&lt;h2&gt;
  
  
  The boundary you can actually review
&lt;/h2&gt;

&lt;p&gt;You cannot review whether a random page contains a V8 exploit. You can review which Chromium binary the agent starts, whether that binary is at least 152.0.7977.82, whether the sandbox flags are intact, and whether the session is a disposable profile or the user's own Chrome.&lt;/p&gt;

&lt;p&gt;Those four facts belong in the same place as the rest of the agent's contract: a file, a startup assertion, and an exit status. They do not belong in a wiki reminder to "keep Chrome updated." The desktop auto-updater and the agent cache are different programs. Treating them as one program is how a patched laptop still runs a vulnerable renderer for the model.&lt;/p&gt;

&lt;p&gt;The sandbox remains worth keeping. It is the reason CVE-2026-85046 is not described as code execution on the host. Keep the flag that enables it. Keep site isolation.&lt;/p&gt;

&lt;p&gt;Then keep the other line in view. An agent that browses untrusted HTML hands a renderer a session and asks it to act. Code execution inside that renderer is already enough to steal cookies, drive the page, and act as the logged-in user. Chrome's security team will still care whether the jail held. The operator should care whether the agent still belonged to them. For this class of product, "inside the sandbox" is already a total loss of the session the agent was built to use.&lt;/p&gt;




&lt;p&gt;Originally published on &lt;a href="https://dispatch-blog.hashnode.dev/inside-the-sandbox-is-still-code-execution-chromium-cve-2026-85046-for-agent-operators" rel="noopener noreferrer"&gt;Dispatch&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>python</category>
      <category>programming</category>
      <category>tutorial</category>
      <category>ai</category>
    </item>
    <item>
      <title>The Skill Is the Product: Making AI Coding Agents Obey Real Repositories</title>
      <dc:creator>Chen Yuan</dc:creator>
      <pubDate>Thu, 03 Sep 2026 11:33:07 +0000</pubDate>
      <link>https://dev.to/chenyuan20509/the-skill-is-the-product-making-ai-coding-agents-obey-real-repositories-3ckn</link>
      <guid>https://dev.to/chenyuan20509/the-skill-is-the-product-making-ai-coding-agents-obey-real-repositories-3ckn</guid>
      <description>&lt;p&gt;A Hacker News post on September 2, 2026 linked Matt Pocock's public repository, &lt;a href="https://github.com/mattpocock/skills" rel="noopener noreferrer"&gt;Skills For Real Engineers&lt;/a&gt;. The repository is not a new agent runtime. It is a collection of small Markdown skills taken from the author's &lt;code&gt;.agents&lt;/code&gt; directory, with an installer for putting selected files into a project. The discussion around the post made the useful question sharper: which instructions belong in a general repository guide, and which deserve their own reusable workflow?&lt;/p&gt;

&lt;p&gt;That distinction matters once an agent edits a real codebase. A short prompt can suggest an approach. A repository guide can describe local rules. A skill can package a repeatable operation with its trigger, supporting files, and expected checkpoints. The last item is valuable only when the operation has a boundary that a human can inspect. Otherwise, a skill is just a longer prompt with a more impressive filename.&lt;/p&gt;

&lt;h2&gt;
  
  
  The repository is part of the agent
&lt;/h2&gt;

&lt;p&gt;An agent does not arrive at a repository with the same assumptions as its author. It must learn the project’s names, tests, issue tracker, architecture, and definition of done. The README for Pocock’s collection describes this problem directly: agents are often asked to discover a project’s vocabulary as they work, then spend many words restating concepts that the project could have named once.&lt;/p&gt;

&lt;p&gt;The collection’s &lt;code&gt;CONTEXT.md&lt;/code&gt; example shows the practical fix. A project-specific phrase can replace a long explanation of a domain event. That is not decoration. It gives the agent a stable search term for functions, files, tests, and design notes. The repository becomes part of the skill’s input rather than a blank directory that the agent may interpret differently on every run.&lt;/p&gt;

&lt;p&gt;This does not mean that every local rule belongs in a skill. A global rule such as “run the formatter before committing” fits an &lt;code&gt;AGENTS.md&lt;/code&gt; or &lt;code&gt;CLAUDE.md&lt;/code&gt; file. A workflow such as “turn an ambiguous request into a spec, then create tickets” has a trigger, a sequence, and a visible result. That workflow is a better skill candidate. The test is simple: could a teammate invoke it by name and know what artifact should exist when it finishes?&lt;/p&gt;

&lt;h2&gt;
  
  
  What a skill actually contains
&lt;/h2&gt;

&lt;p&gt;Pocock’s repository uses ordinary Markdown files, grouped into directories such as &lt;code&gt;skills/engineering&lt;/code&gt; and &lt;code&gt;skills/productivity&lt;/code&gt;. Its README separates user-invoked skills from model-invoked skills. User-invoked entries include flows such as &lt;code&gt;/grill-me&lt;/code&gt;, &lt;code&gt;/to-spec&lt;/code&gt;, and &lt;code&gt;/implement&lt;/code&gt;. Model-invoked entries include reusable practices such as test-driven development, bug diagnosis, code review, and codebase design. The distinction is about who may start the flow, not about whether one kind of file is magically more intelligent.&lt;/p&gt;

&lt;p&gt;The repository also shows why a skill is different from a README. A README explains a collection and links to its members. A skill file describes the work to perform, the questions to ask, the files to consult, and the checks to complete. The surrounding repository supplies scripts, templates, reference documents, and installation metadata where needed. Together, these parts form an operational package.&lt;/p&gt;

&lt;p&gt;It is useful to keep the contract small. A skill needs a trigger, inputs, ordered actions, and an observable verification rule. Those fields do not claim to be the schema of Pocock’s repository. They are a local convention for teams that want to make a file-based workflow easier to review and test.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;add-health-route&lt;/span&gt;
&lt;span class="na"&gt;trigger&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;add a health route&lt;/span&gt;
&lt;span class="na"&gt;inputs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;handler_file&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;path&lt;/span&gt;
    &lt;span class="na"&gt;required&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
  &lt;span class="na"&gt;route_path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;string&lt;/span&gt;
    &lt;span class="na"&gt;required&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
    &lt;span class="na"&gt;allowed_prefix&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/"&lt;/span&gt;
&lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;action&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;append_if_missing&lt;/span&gt;
    &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;{{&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;handler_file&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;}}"&lt;/span&gt;
    &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
      &lt;span class="s"&gt;@app.get("{{ route_path }}")&lt;/span&gt;
      &lt;span class="s"&gt;def health():&lt;/span&gt;
          &lt;span class="s"&gt;return {"status": "ok"}&lt;/span&gt;
&lt;span class="na"&gt;verify&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;action&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;contains&lt;/span&gt;
    &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;{{&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;handler_file&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;}}"&lt;/span&gt;
    &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;@app.get(&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s"&gt;{{&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;route_path&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;}}&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s"&gt;)"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The example is deliberately boring. It names one file, accepts two values, performs one mutation, and checks one result. It does not promise to understand a whole web framework. That modesty makes review possible. A maintainer can ask whether the route belongs in that file, whether the decorator is correct, and whether the verification is strong enough before the agent runs it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with a narrow contract
&lt;/h2&gt;

&lt;p&gt;A broad instruction such as “improve this service” gives an agent too many legal interpretations. A narrow contract removes choices that do not belong to the task. It can state the permitted directories, the files that must already exist, the command that may run, and the result that proves completion. When an assumption is false, the skill should stop before mutation and report the missing condition.&lt;/p&gt;

&lt;p&gt;Narrow does not mean trivial. Pocock’s collection includes workflows for grilling through a plan, writing a test-driven change, diagnosing a difficult bug, and reviewing a diff. Each workflow is substantial, but each has a recognizable seam. &lt;code&gt;tdd&lt;/code&gt; is about a red-green-refactor loop. &lt;code&gt;diagnosing-bugs&lt;/code&gt; is about building a feedback loop around one hard bug. &lt;code&gt;code-review&lt;/code&gt; examines a bounded diff against standards and the originating specification.&lt;/p&gt;

&lt;p&gt;The contract should also say what the skill will not do. A route skill should not silently redesign authentication. A ticket skill should not publish a specification to an issue tracker without making the destination explicit. A test skill should not claim success because a file changed. Negative boundaries protect the repository from enthusiastic automation.&lt;/p&gt;

&lt;p&gt;The input contract is where safety begins. Reject absolute paths if the operation is meant to stay inside the repository. Reject a missing required value. Restrict enumerated options. Resolve defaults before the first write. These checks are cheap, and they turn an ambiguous request into a visible failure instead of a guess.&lt;/p&gt;

&lt;h2&gt;
  
  
  Turn the contract into executable checks
&lt;/h2&gt;

&lt;p&gt;Written instructions explain intent, but an exit status or an exact file check gives the agent something firmer to observe. The verification need not be elaborate. It can assert that a generated file exists, a required string appears once, a test command exits successfully, or a diff stays inside an allowed directory.&lt;/p&gt;

&lt;p&gt;A useful check tests the outcome rather than the action. “The append step ran” is weak evidence. “The route decorator appears in the intended file, the test passes, and no file outside the allowlist changed” is stronger. The more consequential the operation, the more the verification should inspect the external state that matters.&lt;/p&gt;

&lt;p&gt;The same rule applies to skills copied from a public collection. Installation is not verification. The README for &lt;code&gt;mattpocock/skills&lt;/code&gt; offers two paths: a Claude Code plugin or the &lt;code&gt;skills.sh&lt;/code&gt; installer, which can copy selected editable files into a repository. After installation, a team still needs to confirm that the chosen files are in the expected directory, that their triggers are reachable in the chosen agent, and that project-specific paths and terminology have been adapted. A copied workflow is a starting point, not proof that it fits.&lt;/p&gt;

&lt;p&gt;A small runner can make the local convention testable without introducing a framework. The important details are path confinement, input validation, bounded commands, and structured results. The runner below treats the YAML contract as data. It never passes a rendered command through a shell.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;yaml&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;SkillRunner&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;contract_path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;repo_path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;contract_path&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;contract_path&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;repo_path&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;repo_path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;resolve&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;contract_path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;contract&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;yaml&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;safe_load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_inputs&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;provided&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;values&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;spec&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;contract&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;inputs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{}).&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;provided&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;provided&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
            &lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;default&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;spec&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;spec&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;default&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
            &lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;spec&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;required&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;missing required input: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;continue&lt;/span&gt;
            &lt;span class="n"&gt;allowed_prefix&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;spec&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;allowed_prefix&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;allowed_prefix&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;startswith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;allowed_prefix&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;invalid value for input: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;values&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;values&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_render&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;values&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;values&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
            &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;{{ &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; }}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;candidate&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;repo_path&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;resolve&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;candidate&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;repo_path&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;repo_path&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;candidate&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;parents&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;path leaves repository&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;candidate&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;provided&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;values&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_inputs&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;provided&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;step&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;contract&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;steps&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[]):&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;action&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;append_if_missing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;path&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_render&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;path&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;values&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
                &lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;parent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;mkdir&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;parents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;exist_ok&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="n"&gt;content&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_render&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;values&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="n"&gt;current&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exists&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;
                &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;content&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;current&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;current&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;action&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;command&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;argv&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_render&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;values&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;argv&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]]&lt;/span&gt;
                &lt;span class="n"&gt;completed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                    &lt;span class="n"&gt;argv&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cwd&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;repo_path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;capture_output&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                    &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;timeout&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
                    &lt;span class="n"&gt;check&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;completed&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;returncode&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ok&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;phase&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;step&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;returncode&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;completed&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;returncode&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;stderr&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;completed&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stderr&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
            &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;unsupported action: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;action&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;check&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;contract&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;verify&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[]):&lt;/span&gt;
            &lt;span class="n"&gt;path&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_render&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;check&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;path&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;values&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
            &lt;span class="n"&gt;expected&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_render&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;check&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;values&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;check&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;action&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;contains&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;expected&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ok&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;phase&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;verify&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;check&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;check&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;unnamed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ok&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;phase&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;verify&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is not a general agent framework. It supports two file-level actions and one command action, and it returns early on a failed command or check. That is enough to demonstrate the property a skill needs: the procedure is explicit, and the result is not inferred from a confident paragraph.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep context local and explicit
&lt;/h2&gt;

&lt;p&gt;Context should be close to the workflow that consumes it. A repository-wide guide can define language, package manager, test command, and protected directories. A skill can point to the small set of files needed for its operation. A reference document can explain a domain term without forcing every task to reread the whole project.&lt;/p&gt;

&lt;p&gt;Pocock’s README makes this separation concrete through &lt;code&gt;CONTEXT.md&lt;/code&gt;, ADRs, and skills that update or use them. The point is not to put every fact into one enormous instruction file. Large instruction files become difficult to audit, and agents may spend attention on rules unrelated to the current change. A pointer to a focused document is often better than a duplicate explanation.&lt;/p&gt;

&lt;p&gt;Local context also improves portability. The public collection can be installed as editable files, but the README warns that a team should choose one installation philosophy rather than installing the same skills twice through different mechanisms. After installation, replace generic assumptions with the project’s actual issue tracker, documentation location, labels, and vocabulary. A skill copied unchanged is not automatically a team process.&lt;/p&gt;

&lt;p&gt;The context boundary should be visible in the contract. List the files the skill may read. List the paths it may write. Name commands and give them time limits. If a workflow needs a human decision, produce a proposal and stop at that seam. Do not hide a product decision inside a filesystem helper.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test failure paths, not demos
&lt;/h2&gt;

&lt;p&gt;The happy path proves that the example was arranged correctly. Failure tests prove that the boundary exists. For a file-writing skill, test a missing required input, a path outside the repository, an unsupported option, a command timeout, and a verification mismatch. The assertion should cover both the returned status and the absence of an unsafe mutation.&lt;/p&gt;

&lt;p&gt;A failure result should be useful to the next action. “Failed” is not enough. Include the phase, the named check, the return code, or the rejected input. Avoid copying an entire environment into the result; a short diagnostic is easier for a human and an agent to inspect.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;yaml&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;runner&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;SkillRunner&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;make_contract&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;expected&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;@app.get(&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s"&gt;/health&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s"&gt;)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;contract&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;inputs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;handler_file&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;path&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;required&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;}},&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;steps&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;action&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;append_if_missing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;path&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;{{ handler_file }}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;@app.get(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/health&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;)
&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;}],&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;verify&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;route&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;action&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;contains&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;path&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;{{ handler_file }}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;expected&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;}],&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;yaml&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;safe_dump&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;contract&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_successful_run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tmp_path&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;contract&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tmp_path&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;skill.yaml&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="nf"&gt;make_contract&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;contract&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;SkillRunner&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;contract&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tmp_path&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;handler_file&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;app/routes.py&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ok&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;phase&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;verify&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;@app.get(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/health&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;)&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tmp_path&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;app/routes.py&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;read_text&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_failed_verification_is_reported&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tmp_path&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;contract&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tmp_path&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;skill.yaml&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="nf"&gt;make_contract&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;contract&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;expected&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;marker-that-is-not-written&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;SkillRunner&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;contract&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tmp_path&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;handler_file&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;app/routes.py&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ok&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;phase&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;verify&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;check&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;route&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second test is intentionally unglamorous. The write happens, but the contract still reports failure because the stated outcome is absent. That distinction prevents a runner from confusing “the step executed” with “the task succeeded.” In a real repository, the verification would usually call the project’s existing tests rather than inventing a second test system.&lt;/p&gt;

&lt;h2&gt;
  
  
  A small Python skill runner
&lt;/h2&gt;

&lt;p&gt;The runner’s most important design choice is what it refuses to do. It does not interpret free-form prose as a command. It does not accept a path that escapes the repository. It does not turn a missing input into a guessed default unless the contract declares that default. It does not report success until every verification entry passes.&lt;/p&gt;

&lt;p&gt;The command action is still a sharp tool. A fixed argument list is safer than a string assembled for a shell, but it can still delete data or contact an external service. Give the skill an allowlist of commands, use a short timeout, capture output, and run it in a temporary checkout when the operation is experimental. For deployment, billing, credentials, or irreversible migrations, make the human approval an explicit boundary instead of pretending that a local exit code proves the remote state.&lt;/p&gt;

&lt;p&gt;This is also where existing engineering skills help. Pocock’s collection treats test-driven development, diagnosis, architecture, and code review as separate disciplines that can be combined around a change. A local runner should not reproduce all of those practices. It should invoke the project’s established tools and verify the seam that belongs to the skill.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this approach stops helping
&lt;/h2&gt;

&lt;p&gt;A skill contract is a good fit when the task has a stable trigger, a small input surface, and a deterministic check. It is a poor fit for work whose correctness depends on an unresolved product decision, a broad architectural tradeoff, or a remote system that cannot be inspected from the repository. In those cases, the skill should gather evidence, produce a proposal, and stop before the decision point.&lt;/p&gt;

&lt;p&gt;The Hacker News comments on Pocock’s repository raise a fair boundary question: some practices could fit in a general &lt;code&gt;AGENTS.md&lt;/code&gt; or &lt;code&gt;CLAUDE.md&lt;/code&gt;, while a skill is more useful for a named workflow, a project-specific script, unusual domain knowledge, or a sequence that benefits from explicit gates. That is a better rule than treating every useful paragraph as a new command.&lt;/p&gt;

&lt;p&gt;The practical recommendation is to start with one narrow skill that touches a known seam. Define its inputs and refusal conditions. Reuse the repository’s own tests. Add failure cases before sharing the file. If the skill grows a long list of exceptions, stop adding branches and move the complex reasoning into a normal tool or a human review step. The product is the bounded, inspectable workflow. The agent is the caller that applies it to the repository in front of it.&lt;/p&gt;




&lt;p&gt;Originally published on &lt;a href="https://dispatch-blog.hashnode.dev/the-skill-is-the-product-making-ai-coding-agents-obey-real-repositories" rel="noopener noreferrer"&gt;Dispatch&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>python</category>
      <category>programming</category>
      <category>tutorial</category>
      <category>ai</category>
    </item>
    <item>
      <title>Why uv Now Deduplicates Every Wheel in Its Cache (and How It Does It)</title>
      <dc:creator>Chen Yuan</dc:creator>
      <pubDate>Mon, 31 Aug 2026 15:51:22 +0000</pubDate>
      <link>https://dev.to/chenyuan20509/why-uv-now-deduplicates-every-wheel-in-its-cache-and-how-it-does-it-2gh</link>
      <guid>https://dev.to/chenyuan20509/why-uv-now-deduplicates-every-wheel-in-its-cache-and-how-it-does-it-2gh</guid>
      <description>&lt;h2&gt;
  
  
  The Problem: Duplicate Wheels in the Cache
&lt;/h2&gt;

&lt;p&gt;A developer working on a monorepo with two Python projects—a FastAPI web service and a data pipeline—notices something odd. Each project has a &lt;code&gt;requirements.txt&lt;/code&gt; file listing around 200 dependencies. When they run &lt;code&gt;uv pip install&lt;/code&gt; for the first project, uv downloads and extracts 200 wheels into its global cache. Then they run &lt;code&gt;uv pip install&lt;/code&gt; for the second project, and uv does the same thing again.&lt;/p&gt;

&lt;p&gt;But here's the kicker: these two projects share roughly 150 of the same dependencies. &lt;code&gt;pydantic==2.10.4&lt;/code&gt;, &lt;code&gt;httpx==0.27.2&lt;/code&gt;, &lt;code&gt;typing-extensions==4.12.2&lt;/code&gt;—the list goes on. Yet when they check their disk usage, they find that uv has stored 400 copies of extracted wheels on disk. The same &lt;code&gt;typing-extensions&lt;/code&gt; wheel, byte-for-byte identical, exists twice in the cache. They're effectively paying double the storage cost for a dependency they use in both projects.&lt;/p&gt;

&lt;p&gt;This isn't an edge case. In any Python environment with multiple projects, virtual environments, or CI pipelines, the same packages get pulled down and extracted repeatedly. The waste compounds quickly. A monorepo with 10 microservices might store 10 copies of &lt;code&gt;numpy&lt;/code&gt;, 10 copies of &lt;code&gt;pandas&lt;/code&gt;, and 10 copies of &lt;code&gt;scikit-learn&lt;/code&gt;—each taking up hundreds of megabytes.&lt;/p&gt;

&lt;p&gt;Why does this happen? Prior to the content-addressed cache feature, uv's cache stored extracted wheels in &lt;code&gt;archive-v0&lt;/code&gt; under randomly generated IDs. If the same wheel was reached through different cache entries—for example, fetched from two different indexes—uv stored duplicate copies because it identified extracts by their source rather than their contents.&lt;/p&gt;

&lt;p&gt;This identity-based approach is simple and fast—you can look up a package by name and version instantly. But it leads to significant redundancy when many projects share the same dependencies.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before: How uv's Cache Organized Data
&lt;/h2&gt;

&lt;p&gt;Before the content-addressed cache features landed, uv's cache had a two-tier structure. Downloaded wheels lived in the &lt;code&gt;wheels-v0&lt;/code&gt; or &lt;code&gt;wheels-v1&lt;/code&gt; bucket, keyed by a combination of the package name, version, and a hash of the wheel file itself. Extracted wheels lived in &lt;code&gt;archive-v0&lt;/code&gt;, and each extract got its own randomly generated directory ID.&lt;/p&gt;

&lt;p&gt;You can inspect your current uv cache location with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uv cache &lt;span class="nb"&gt;dir&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On Linux and macOS, this typically outputs something like &lt;code&gt;$HOME/.cache/uv&lt;/code&gt;. On Windows, it's &lt;code&gt;%LOCALAPPDATA%\uv\cache&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;If you peek inside the &lt;code&gt;archive-v0&lt;/code&gt; directory, you'll see a bunch of seemingly random directory names:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;ls&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;uv cache &lt;span class="nb"&gt;dir&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;/archive-v0/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On a real system, this might show entries like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;0VSKRv7KPsU6OVxKhRKyA/
3vlR8U66f3qKdzSrxDtyR/
gekN10oQ_Vp-g2TxEYqsD/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each of these opaque directories contains the full extracted contents of a single wheel. The IDs are randomly generated—they carry no relationship to the files inside. This means that if you install the same wheel from PyPI and then from a local directory, uv treats them as different cache entries and stores two copies of the extracted files.&lt;/p&gt;

&lt;p&gt;To see just how much duplication exists, you can use tools like &lt;code&gt;du&lt;/code&gt; to measure the total size of the cache and compare it against the unique file count:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;du&lt;/span&gt; &lt;span class="nt"&gt;-sh&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;uv cache &lt;span class="nb"&gt;dir&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="nb"&gt;ls&lt;/span&gt; &lt;span class="nt"&gt;-R&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;uv cache &lt;span class="nb"&gt;dir&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;/archive-v0/ | &lt;span class="nb"&gt;wc&lt;/span&gt; &lt;span class="nt"&gt;-l&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a project with many dependencies, the &lt;code&gt;archive-v0&lt;/code&gt; directory can contain tens of thousands of files. A significant portion of these are duplicates—the same &lt;code&gt;.py&lt;/code&gt; files, the same compiled extensions, the same native libraries, stored over and over again.&lt;/p&gt;

&lt;h2&gt;
  
  
  Content-Addressable Storage: A Brief Primer
&lt;/h2&gt;

&lt;p&gt;Content-addressable storage (CAS) is a pattern where data is stored and retrieved by its content hash rather than by a name or location. Instead of asking for "file &lt;code&gt;typing-extensions-4.12.2.dist-info&lt;/code&gt;," you ask for "the file whose BLAKE3 hash is &lt;code&gt;b7a3...&lt;/code&gt;." If two files have identical contents, they produce identical hashes, and the storage system returns the same object for both requests.&lt;/p&gt;

&lt;p&gt;This concept is not new. Git uses content-addressable storage for its object database. Every blob, tree, and commit is identified by its SHA-1 hash, which is why Git can detect duplicate files across your repository and only store them once. Docker layer caching works on a similar principle—if two images share the same base layer, they can reuse it without storing a second copy. npm deduplicates tarballs in its cache using content hashes, ensuring that if you install the same package in different projects, you only download it once.&lt;/p&gt;

&lt;p&gt;The key insight is that content-addressable storage turns the problem of "how do I avoid storing the same data twice" into a simple lookup: compute the hash of the content, check if an object with that hash already exists, and if so, reuse it. The hard part is managing the links—making sure that when you ask for the contents of a package, the system knows which hashed objects to assemble, and when a package is removed, the system can safely garbage-collect unreferenced objects.&lt;/p&gt;

&lt;h2&gt;
  
  
  How uv's Content-Addressed Cache Works
&lt;/h2&gt;

&lt;p&gt;uv's content-addressed cache evolution happened in two stages. The foundational work landed in PR #19693, which introduced content-based directory hashes for entire extracted wheels. The idea was simple: instead of storing an extracted wheel under a randomly generated ID, compute a hash of the directory's entire contents, and store it under that hash. If two different wheels extract to the exact same directory tree, they'll produce the same hash, and uv can deduplicate them.&lt;/p&gt;

&lt;p&gt;The more significant optimization came in PR #21327, which moved deduplication from the wheel level to the file level. In this newer implementation, every file extracted from a wheel is stored individually in a &lt;code&gt;files-v0&lt;/code&gt; bucket, keyed by its BLAKE3 content hash. The original directory structure is preserved through hardlinks—the &lt;code&gt;archive-v0&lt;/code&gt; directory contains hardlinks pointing to the actual file objects in &lt;code&gt;files-v0&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Here's how the flow works:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;When uv needs to extract a wheel, it iterates through every file in the archive.&lt;/li&gt;
&lt;li&gt;For each file, it computes a BLAKE3 hash of the file's contents.&lt;/li&gt;
&lt;li&gt;It checks if an object with that hash already exists in the &lt;code&gt;files-v0&lt;/code&gt; bucket.&lt;/li&gt;
&lt;li&gt;If the object exists, uv creates a hardlink from the extracted wheel's location to the existing object.&lt;/li&gt;
&lt;li&gt;If the object doesn't exist, uv writes the file to the &lt;code&gt;files-v0&lt;/code&gt; bucket under its hash, then creates a hardlink to it.&lt;/li&gt;
&lt;li&gt;The extracted wheel directory in &lt;code&gt;archive-v0&lt;/code&gt; now contains hardlinks to shared file objects, with each object stored exactly once on disk.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This approach preserves the existing installation pipeline—the wheel installation step still reads files from their expected paths in &lt;code&gt;archive-v0&lt;/code&gt;—but the underlying storage is now fully deduplicated.&lt;/p&gt;

&lt;p&gt;The PR benchmarks showed impressive savings. On a local machine with a typical Python cache, the file-level deduplication saved 545.2 MiB of disk space, or about 10% of the total cache. The optimization covered all payload files in the cache—not just executables and native libraries, but every &lt;code&gt;.py&lt;/code&gt;, &lt;code&gt;.so&lt;/code&gt;, &lt;code&gt;.dll&lt;/code&gt;, and data file.&lt;/p&gt;

&lt;p&gt;The performance impact was also carefully measured. Cold installs—where the cache is empty and files must be written for the first time—showed a slowdown of less than 4% on median. Warm installs, where the cache is already populated, showed no statistically significant difference. The team found this tradeoff worthwhile—a small performance hit on cold installs for a noticeable reduction in disk usage.&lt;/p&gt;

&lt;p&gt;The feature uses BLAKE3 as its hashing algorithm, which is known for its speed and parallelizability, making it well-suited for hashing many files during wheel extraction.&lt;/p&gt;

&lt;h2&gt;
  
  
  Enabling the Preview Feature
&lt;/h2&gt;

&lt;p&gt;The content-addressed cache features are currently behind preview flags. To use them, you need to explicitly opt in. uv provides several ways to enable preview features.&lt;/p&gt;

&lt;p&gt;First, you can enable the feature via the &lt;code&gt;--preview-features&lt;/code&gt; flag on individual commands:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uv pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;--preview-features&lt;/span&gt; content-addressed-cache &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To enable it for all commands in a session, you can use the &lt;code&gt;UV_PREVIEW_FEATURES&lt;/code&gt; environment variable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;UV_PREVIEW_FEATURES&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;content-addressed-cache
uv pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You can also enable preview features in your &lt;code&gt;uv.toml&lt;/code&gt; file or under &lt;code&gt;[tool.uv]&lt;/code&gt; in your &lt;code&gt;pyproject.toml&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight toml"&gt;&lt;code&gt;&lt;span class="nn"&gt;[tool.uv]&lt;/span&gt;
&lt;span class="py"&gt;preview-features&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"content-addressed-cache"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For the physical space accounting feature—which accounts for hardlinks when reporting cache size—you need the &lt;code&gt;cache-physical-space&lt;/code&gt; preview feature as well. You can enable multiple features simultaneously:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uv cache size &lt;span class="nt"&gt;--preview-features&lt;/span&gt; content-addressed-cache,cache-physical-space
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or set the environment variable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;UV_PREVIEW_FEATURES&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;content-addressed-cache,cache-physical-space
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once enabled, uv will start using content-addressed storage for new wheel extractions. Existing cache entries will remain in their current format—uv doesn't automatically migrate old cache data to the new structure. You can either let new installs populate the new format gradually, or you can clear your cache and start fresh:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uv cache clean
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Note that preview features are subject to change. uv explicitly warns that they may be modified or removed without notice, and they should not be relied upon in production environments until stabilized.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measuring the Savings
&lt;/h2&gt;

&lt;p&gt;To measure how much space your cache is actually using on disk—accounting for hardlinks—you need the &lt;code&gt;cache-physical-space&lt;/code&gt; preview feature. Without it, &lt;code&gt;uv cache size&lt;/code&gt; reports the apparent size of the cache, which counts the same file multiple times if it's hardlinked in different locations. With the physical space accounting, the command reports the actual disk usage.&lt;/p&gt;

&lt;p&gt;Run the following command to see your cache's physical space usage:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uv cache size &lt;span class="nt"&gt;--preview-features&lt;/span&gt; content-addressed-cache,cache-physical-space &lt;span class="nt"&gt;--human&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;--human&lt;/code&gt; flag displays the size in a human-readable format (e.g., &lt;code&gt;1.2 GiB&lt;/code&gt; instead of raw bytes).&lt;/p&gt;

&lt;p&gt;The actual savings you'll see depend on your workload. The benchmark in PR #21327 showed a 545.2 MiB reduction on a local machine, which was about 10% of the total cache. However, that benchmark was on a single machine with a specific set of packages. In practice, the savings could be higher or lower depending on how much overlap exists between your dependencies.&lt;/p&gt;

&lt;p&gt;For workloads with many large packages that share common files—like multiple machine learning frameworks that bundle the same native libraries—the savings could be significantly more. The per-file selection table from the PR shows the pattern: selecting only executables and native libraries saves 275.7 MiB across 3,336 files, while selecting all payload files saves 545.2 MiB across 134,222 files.&lt;/p&gt;

&lt;p&gt;To see the breakdown, you can also use system-level tools to compare the apparent size of the extracted tree against the actual disk usage of the object store:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;du&lt;/span&gt; &lt;span class="nt"&gt;-sb&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;uv cache &lt;span class="nb"&gt;dir&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;/archive-v0   &lt;span class="c"&gt;# Apparent size: every hardlink counts full&lt;/span&gt;
&lt;span class="nb"&gt;du&lt;/span&gt; &lt;span class="nt"&gt;-s&lt;/span&gt;  &lt;span class="si"&gt;$(&lt;/span&gt;uv cache &lt;span class="nb"&gt;dir&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;/files-v0     &lt;span class="c"&gt;# Disk usage: each unique object once&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The difference between these numbers represents the space saved by file-level deduplication.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Means for CI and Local Development
&lt;/h2&gt;

&lt;p&gt;For local development, the primary benefit is reduced disk usage. If you work on many Python projects simultaneously, a smaller cache means you're less likely to run out of disk space on your development machine. It also means less SSD wear—hardlinking does not write new data to disk, so the cache's write amplification is lower.&lt;/p&gt;

&lt;p&gt;For CI pipelines, the impact is more nuanced. CI systems often cache the &lt;code&gt;~/.cache/uv&lt;/code&gt; directory between runs to avoid re-downloading packages. But the cache's &lt;code&gt;archive-v0&lt;/code&gt; directory contains many small files that take time to compress and decompress. With file-level deduplication, the cache contains fewer distinct file objects in &lt;code&gt;files-v0&lt;/code&gt;, which may improve compression ratios and reduce cache transfer times.&lt;/p&gt;

&lt;p&gt;However, the best practice for CI is still to cache only the downloaded wheels (&lt;code&gt;wheels-v1&lt;/code&gt;) and not the extracted archives (&lt;code&gt;archive-v0&lt;/code&gt;), since the archive can be regenerated from the wheels. The content-addressed cache doesn't change this recommendation—if anything, it makes it more important to be selective about what gets cached.&lt;/p&gt;

&lt;p&gt;For monorepo setups, where many projects share dependencies, the benefits compound. A single &lt;code&gt;typing-extensions&lt;/code&gt; wheel extracted once into &lt;code&gt;files-v0&lt;/code&gt; and hardlinked across 10 project caches would save 9 copies of the extracted files. Over hundreds of shared dependencies across dozens of projects, the savings can add up quickly.&lt;/p&gt;

&lt;p&gt;It's worth noting that this feature is still in preview. Both PR #19693 (wheel-level dedup, merged 2026-08-25) and PR #21327 (file-level dedup, merged 2026-08-31) are in the development branch and have been released under the &lt;code&gt;content-addressed-cache&lt;/code&gt; preview flag in uv 0.12.7. Early adopters should expect some rough edges, especially around cross-platform compatibility. The initial implementation was validated on Linux with the ext4 filesystem. Supporting macOS, Windows, and other filesystems may require additional work. The team has considered edge cases like reflink support on copy-on-write filesystems (e.g., Btrfs, ZFS, APFS), but these features are still under development.&lt;/p&gt;




&lt;p&gt;Originally published on &lt;a href="https://dispatch-blog.hashnode.dev/why-uv-now-deduplicates-every-wheel-in-its-cache-and-how-it-does-it" rel="noopener noreferrer"&gt;Dispatch&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>python</category>
      <category>programming</category>
      <category>tutorial</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
