<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Youfu Hsu</title>
    <description>The latest articles on DEV Community by Youfu Hsu (@youfuhsu).</description>
    <link>https://dev.to/youfuhsu</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4111810%2F08dd26e1-4ba0-4ac7-8fd8-535e566479f1.png</url>
      <title>DEV Community: Youfu Hsu</title>
      <link>https://dev.to/youfuhsu</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/youfuhsu"/>
    <language>en</language>
    <item>
      <title>Giving an AI agent memory is easy. Keeping it true is the hard part.</title>
      <dc:creator>Youfu Hsu</dc:creator>
      <pubDate>Wed, 09 Sep 2026 04:38:16 +0000</pubDate>
      <link>https://dev.to/youfuhsu/giving-an-ai-agent-memory-is-easy-keeping-it-true-is-the-hard-part-3f02</link>
      <guid>https://dev.to/youfuhsu/giving-an-ai-agent-memory-is-easy-keeping-it-true-is-the-hard-part-3f02</guid>
      <description>&lt;p&gt;Giving a coding agent memory that survives the session is not a hard engineering problem. A directory of markdown files outside any repo, one fact per file, an index that gets loaded at the start of every session. That is the whole mechanism, and it takes an afternoon.&lt;/p&gt;

&lt;p&gt;I have been running one for about four months across roughly forty projects. The mechanism has never been the problem. &lt;strong&gt;In a single working day last week, four separate memory entries turned out to be confidently wrong&lt;/strong&gt;, and each one had been steering decisions for weeks.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;"Schema fix committed locally, not yet pushed."&lt;/strong&gt; Read as: there is unpushed work sitting on this machine. Actually: that repo has no remote at all and never did. The deploy path was a script hitting the platform's API directly, so nothing was ever supposed to be pushed. The entry had turned a fact about &lt;em&gt;one&lt;/em&gt; repo's setup into a phantom to-do that survived several audits.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Template fatigue in the post generator: fixed and deployed."&lt;/strong&gt; Green checkmark, verification notes, an A/B comparison in the file. All true. The fix was deployed to a queue worker whose last log line was from six weeks earlier and which had no scheduled task pointing at it. The pipeline that actually posts every morning is a different script that never imports the fixed module. The fix was real, verified, deployed, and had &lt;strong&gt;zero effect on anything anyone could see&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"600 platform credits banked, use before November."&lt;/strong&gt; Actually the credits had already been converted into a different resource with a &lt;strong&gt;30-day&lt;/strong&gt; expiry, and the deadline was five weeks earlier than the memory claimed. The entry was written on the day of purchase and described the plan, not the outcome.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Two recordings missed last night."&lt;/strong&gt; Written by an automated job that judged success from a process exit code. The recordings were on disk, complete, and had already been transcribed. The exit code was a red herring from a downloader that returns non-zero while cleaning up after a perfectly good capture.&lt;/p&gt;

&lt;p&gt;Different projects, different weeks, one shape.&lt;/p&gt;

&lt;h2&gt;
  
  
  What they have in common
&lt;/h2&gt;

&lt;p&gt;None of these were sloppy notes. They were all written carefully, at the moment of the work, by someone (or something) with full context.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Every one of them recorded a judgment and formatted it as a fact.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;"Not yet pushed" is an inference from "I did not push." "Fixed and deployed" is an inference from "the deploy script exited zero." "Missed" is an inference from "exit code was 1." The inference was reasonable when it was written. The problem is that once it is a line in a file, the reasoning is gone and only the conclusion is left — and conclusions do not carry their own expiry date.&lt;/p&gt;

&lt;p&gt;This gets worse specifically because the memory works. An agent that reads its memory and acts on it is doing the right thing. A wrong entry does not cause an error; it causes confident, efficient work in the wrong direction. My agent spent weeks not investigating the post generator, because the file said it was fixed.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually helps
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Write the re-check, not just the claim
&lt;/h3&gt;

&lt;p&gt;Any entry that asserts a state — done, fixed, deployed, live, expired — should carry the command that re-establishes it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;✅ Subtitle pipeline deployed and verified 08-18.
Re-check: &lt;span class="sb"&gt;`grep "card1:" &amp;lt;log&amp;gt; | tail -3`&lt;/span&gt; — openers should not
all share one skeleton. If they do, the fix is not on the live path.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the claim is falsifiable by anyone who reads it, including a future session with no context. Half of my four would have been caught in seconds by their own re-check line.&lt;/p&gt;

&lt;h3&gt;
  
  
  Distinguish observed from inferred
&lt;/h3&gt;

&lt;p&gt;I now try to write what was seen, and separately what it was taken to mean:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;Observed: deploy script exit 0, target file contains the new function.
Inferred: the fix is live.   ← this is the part that can rot
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The inference is usually still right. But when something later contradicts the memory, you know which half to attack first.&lt;/p&gt;

&lt;h3&gt;
  
  
  Let files outrank memory
&lt;/h3&gt;

&lt;p&gt;When memory and the filesystem disagree, the filesystem wins. This is worth writing as code rather than as a habit. Correcting that "two recordings missed" entry was six lines:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;slug&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;rec&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="c1"&gt;# source of truth is the artifact, not what the job concluded at the time
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exists&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;WORK&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;slug&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/summary.md&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;rec&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;done&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;slug&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;rec&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;done&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                       &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;note&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;corrected: exit code misjudged; output is complete&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Note what it does &lt;em&gt;not&lt;/em&gt; do: regenerate the state file from scratch. It patches only the rows contradicted by evidence and leaves everything else alone. I have destroyed hand-maintained data by "regenerating" it from rules before, and that is a genuinely bad afternoon.&lt;/p&gt;

&lt;h3&gt;
  
  
  Verify from outside the system that made the claim
&lt;/h3&gt;

&lt;p&gt;The strongest habit of the four. A deploy script reporting success is the deploy script's opinion. Ask something that has no stake in the answer.&lt;/p&gt;

&lt;p&gt;This cuts both ways, and the false negative is the one that wastes your time. On the same day I fixed the four entries above, I updated 23 listings on a platform, got 23 success responses back, then checked the public search index and saw &lt;strong&gt;zero&lt;/strong&gt; of the changes. Looked exactly like "claimed success, not actually live." It was a caching layer on the search endpoint; the per-item public endpoint had all 23 updates with the correct timestamp.&lt;/p&gt;

&lt;p&gt;So: verify externally, but when the external check disagrees, suspect the checker before you rip out the work. The rule is not "the external check is right." It is &lt;strong&gt;"one source is never enough."&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The part I did not expect
&lt;/h2&gt;

&lt;p&gt;I assumed the risk of agent memory was the model hallucinating something into a file. Four months in, that has not happened once. Everything in my memory directory was written from real work.&lt;/p&gt;

&lt;p&gt;The risk is &lt;strong&gt;decay&lt;/strong&gt;. Every entry is a snapshot of a moment, and the world keeps going. A memory system does not fail by filling with lies. It fails by filling with things that were carefully, honestly true in July.&lt;/p&gt;

&lt;p&gt;Which means the maintenance job is not "keep the notes tidy." It is: for anything currently steering a decision, go look again.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I write about running unattended automation and AI agents in production, including the parts where the agent — or I — confidently get it wrong. The full system I use for this is the &lt;a href="https://alphatech4.gumroad.com/l/claude-code-automation-playbook" rel="noopener noreferrer"&gt;Claude Code Automation Playbook&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>automation</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Your prompt is feeding the model its own cliché</title>
      <dc:creator>Youfu Hsu</dc:creator>
      <pubDate>Tue, 08 Sep 2026 01:08:32 +0000</pubDate>
      <link>https://dev.to/youfuhsu/your-prompt-is-feeding-the-model-its-own-cliche-4pj4</link>
      <guid>https://dev.to/youfuhsu/your-prompt-is-feeding-the-model-its-own-cliche-4pj4</guid>
      <description>&lt;p&gt;I have a small pipeline that turns a daily market report into four subtitle cards for a short video. It runs unattended every morning. I checked its output after two weeks and the wording never repeated once.&lt;/p&gt;

&lt;p&gt;It still looked machine-generated. Here is what the opening card said on ten consecutive days:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;the coin that pumped hardest, but volume is cold?
old coin FIL back on the gainers board?
small caps exploding, majors flat?
up 38% and nobody's buying?
the top gainer is the fakest one?
the coin that rose least is the popularity king
falling price but volume spiking, who's rotating out?
everything red, but this one is up
meme DeFi AI all pumping the same day?
DEX names all agreed to pump together?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every word is different. &lt;strong&gt;Eleven out of fourteen are the same skeleton&lt;/strong&gt;: a contrast, then a question mark. Read one, it's fine. Read a week, and you can see the machine.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it came from
&lt;/h2&gt;

&lt;p&gt;The relevant line of my system prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;4. Card 1 is the hook (a cliffhanger or counterintuitive point),
   cards 2-3 carry the concrete numbers, card 4 is a conclusion or a question.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;"Cliffhanger." "Counterintuitive." I wrote those words to describe the &lt;em&gt;goal&lt;/em&gt;, and the model read them as a description of the &lt;em&gt;form&lt;/em&gt;. A counterintuitive point, rendered as a sentence, is "X but Y?" — so that is what came out, every single day.&lt;/p&gt;

&lt;p&gt;The model was not being lazy. &lt;strong&gt;I put the cliché in the prompt and it did what I asked.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There is a second cause underneath: the model cannot see what it wrote yesterday. Each run is a fresh context. Even if the wording rule were perfect, nothing stops it from reaching for the same structure every morning, because from its point of view every morning is the first one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fix 1: stop describing the shape
&lt;/h2&gt;

&lt;p&gt;Replace the adjective with a menu of options, and name the failure mode explicitly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;4. Card 1 is the opener, cards 2-3 carry the concrete numbers, card 4 concludes.
5. Card 1 must open differently every day. Angles to rotate through
   (do not keep picking the same one):
   - name a specific ticker and its number
   - lead with a magnitude (volume, market cap, share of total)
   - lead with a change over time (three days running, week to date)
   - describe the state of a whole group or sector
   - state a fact that is happening, with no evaluation
   Do not write "A is up but B is..." or "...and nobody's buying?" every day.
   A reader who scrolls past three of these should not be able to tell
   they came from the same template.
   Card 1 should not end in a question mark unless the content really is
   an open question.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things matter here. Naming the bad pattern in the prompt works better than only describing the good one — the model needs something to steer away from. And listing angles rather than one instruction gives it somewhere to go.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fix 2: show it what it already said
&lt;/h2&gt;

&lt;p&gt;The structural half needs memory. Read back your own recent output and attach it to the request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;recent_hooks&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Last n days of card 1, as an avoid-list. Returns &lt;/span&gt;&lt;span class="sh"&gt;''&lt;/span&gt;&lt;span class="s"&gt; on any failure.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;seen&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;                      &lt;span class="c1"&gt;# date -&amp;gt; hook, so a same-day rerun overwrites
&lt;/span&gt;        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;LOG&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;match&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;^\[(\d{4}-\d{2}-\d{2})T[^\]]*\] card1: (.+)$&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;seen&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;group&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;group&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="n"&gt;hooks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;seen&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;values&lt;/span&gt;&lt;span class="p"&gt;())[&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;:]&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;hooks&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;
        &lt;span class="nf"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="s"&gt;These are the openers already used in the last &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;hooks&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; days. Today&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s must be clearly different — &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;not just different words, a different sentence skeleton:&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;- &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;h&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;h&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;hooks&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then &lt;code&gt;messages=[{"role": "user", "content": draft + recent_hooks()}]&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Three details that are easy to get wrong:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Read from a log you are already writing.&lt;/strong&gt; My first instinct was a new &lt;code&gt;history.json&lt;/code&gt; that the pipeline would append to. That version would have been correct and useless — it starts empty, so the avoid-list does nothing for the first week. The daily log already had &lt;code&gt;card1: ...&lt;/code&gt; in it going back a month, so the feature worked on the first run. I had made this exact mistake before on a different pipeline and shipped a de-duplication function that read from an empty directory and therefore returned an empty string forever.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Return empty on failure, never raise.&lt;/strong&gt; A cosmetic feature must not be able to take down a job that runs while you sleep.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key by date, not by line.&lt;/strong&gt; A same-day rerun should replace that day's entry, not add a second one and push a real day off the end of the window.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to know it worked
&lt;/h2&gt;

&lt;p&gt;Run the old and new prompt against &lt;strong&gt;the same input&lt;/strong&gt;. Same day's report, two runs:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Card 1&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Old&lt;/td&gt;
&lt;td&gt;&lt;code&gt;DEX names all agreed to pump together?&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;New&lt;/td&gt;
&lt;td&gt;&lt;code&gt;RAY up 41% in a day&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I added a &lt;code&gt;--dry-cards&lt;/code&gt; flag to the production script that stops after the model call — no video render, no upload, no notification — so this comparison runs the code that actually ships rather than a copy of it I pasted into a scratch file. A copy would have proved something about the copy.&lt;/p&gt;

&lt;p&gt;That flag introduced one more problem worth mentioning, because it is the kind of thing that bites three weeks later. The dry run writes its cards to the same log that &lt;code&gt;recent_hooks()&lt;/code&gt; reads. Test output would have become "history" and been fed back to the model as something to avoid. So the dry mode tags its lines:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;[dry] &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;dry_cards&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;''&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;card&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and the regex in &lt;code&gt;recent_hooks()&lt;/code&gt; requires &lt;code&gt;card1:&lt;/code&gt; to come straight after the timestamp, so &lt;code&gt;[dry] card1:&lt;/code&gt; never matches. I checked both directions: the test line is in the file (grep finds it), and the parser returns zero rows for it. &lt;strong&gt;A guard you only tested in one direction is half-tested&lt;/strong&gt; — I have shipped a filter that correctly blocked the bad case and silently ate a good one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The general version
&lt;/h2&gt;

&lt;p&gt;If your prompt contains an adjective describing the output, expect that adjective to show up in the output. "Punchy," "surprising," "counterintuitive," "engaging" — these read as instructions about tone, but they land as instructions about form.&lt;/p&gt;

&lt;p&gt;And a stateless model cannot avoid repeating itself. Variety is not something you can ask for; it is something you have to give it the inputs for.&lt;/p&gt;

&lt;p&gt;The tell is worth internalising: &lt;strong&gt;no repeated words, same repeated skeleton&lt;/strong&gt;. Any per-token novelty check will pass that. You have to read a week of output in one sitting to see it, which is exactly what nobody does with a pipeline that has been quietly working for a month.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I write about running unattended automation and AI agents in production, including the parts where the agent — or I — confidently get it wrong. The full system I use for this is the &lt;a href="https://alphatech4.gumroad.com/l/claude-code-automation-playbook" rel="noopener noreferrer"&gt;Claude Code Automation Playbook&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>promptengineering</category>
      <category>python</category>
      <category>automation</category>
    </item>
    <item>
      <title>Your auto-captioner is measuring the wrong thing: a ratio that fired on a perfectly good clip</title>
      <dc:creator>Youfu Hsu</dc:creator>
      <pubDate>Mon, 07 Sep 2026 04:54:19 +0000</pubDate>
      <link>https://dev.to/youfuhsu/your-auto-captioner-is-measuring-the-wrong-thing-a-ratio-that-fired-on-a-perfectly-good-clip-1dh2</link>
      <guid>https://dev.to/youfuhsu/your-auto-captioner-is-measuring-the-wrong-thing-a-ratio-that-fired-on-a-perfectly-good-clip-1dh2</guid>
      <description>&lt;p&gt;I burn subtitles onto short AI-generated ad videos. The pipeline is small: generate the clip, transcribe it, write an &lt;code&gt;.ass&lt;/code&gt; file, let ffmpeg burn it in.&lt;/p&gt;

&lt;p&gt;It broke three times, and each break taught me something narrower and more useful than the last. The third one is the interesting one, because my fix was wrong in a way that looked completely right.&lt;/p&gt;




&lt;h2&gt;
  
  
  Break 1: hand-written timings never match generated speech
&lt;/h2&gt;

&lt;p&gt;The first version used the timings from my own script. I wrote the lines, I knew roughly how long each one took to say, so I typed them in:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"A1_partner_snore"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;3.2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"He snored through our entire honeymoon."&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mf"&gt;3.2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;7.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"I stopped hearing it on night one."&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;]]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every video was out of sync. Not slightly — by a second or more, drifting worse toward the end.&lt;/p&gt;

&lt;p&gt;The reason is obvious in hindsight: &lt;strong&gt;I was timing my own reading of the line, and a generative video model has its own pacing.&lt;/strong&gt; It pauses where it wants. It rushes short clauses. Nothing about my estimate was connected to the audio that actually exists in the file.&lt;/p&gt;

&lt;p&gt;The only reliable source of timing is the audio itself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;whisper&lt;/span&gt;
&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;whisper&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;small&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;transcribe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;wav&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;language&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;en&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;fp16&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;segs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[(&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;start&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;end&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;res&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;segments&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Break 2: the transcript is the wrong text
&lt;/h2&gt;

&lt;p&gt;Now the timing was right and the words were wrong.&lt;/p&gt;

&lt;p&gt;Whisper mishears. In one clip "earplugs" came out as "Urplugs." Separately, the video model itself sometimes drops a word from the line it was handed.&lt;/p&gt;

&lt;p&gt;So there are three versions of every sentence in play:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;th&gt;Timing&lt;/th&gt;
&lt;th&gt;Wording&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;My script&lt;/td&gt;
&lt;td&gt;wrong&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Whisper transcript&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;td&gt;wrong, twice over&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What I want on screen&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Neither source is usable alone, and the fix is to stop treating them as competing answers.&lt;br&gt;
&lt;strong&gt;Take timing from the audio, take text from the script.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;align&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;heard&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;script_lines&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;heard&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;script_lines&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[(&lt;/span&gt;&lt;span class="n"&gt;heard&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;heard&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;script_lines&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
                &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;script_lines&lt;/span&gt;&lt;span class="p"&gt;))]&lt;/span&gt;
    &lt;span class="c1"&gt;# segment counts disagree: spread the script across the spoken span,
&lt;/span&gt;    &lt;span class="c1"&gt;# weighted by word count
&lt;/span&gt;    &lt;span class="n"&gt;t0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;t1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;heard&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;heard&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;weights&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;()))&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;script_lines&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;total&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;weights&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="p"&gt;[],&lt;/span&gt; &lt;span class="n"&gt;t0&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;wt&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;zip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;script_lines&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;weights&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;d&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;t1&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;t0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;wt&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;total&lt;/span&gt;
        &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
        &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;d&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This generalizes past captions. When two sources each have a known-good field and a known-bad field, joining on the good fields beats picking a winner.&lt;/p&gt;




&lt;h2&gt;
  
  
  Break 3: the frozen caption
&lt;/h2&gt;

&lt;p&gt;Two clips came out with a single subtitle sitting on screen for about eight seconds, unmoving, while nothing was being said.&lt;/p&gt;

&lt;p&gt;Those clips barely had dialogue. When speech is sparse, whisper happily extends a segment's &lt;code&gt;end&lt;/code&gt; through the trailing silence. The segment is still correct &lt;em&gt;as a transcript&lt;/em&gt;. It just isn't a caption cue.&lt;/p&gt;

&lt;p&gt;I needed a guard that says &lt;em&gt;this timing is untrustworthy, fall back to hand-written cues&lt;/em&gt;. Here was my first attempt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# if any single segment covers more than 40% of the clip, distrust the timing
&lt;/span&gt;&lt;span class="n"&gt;longest&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;segs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;longest&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;duration&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mf"&gt;0.40&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It caught both broken clips. It also rejected a clip that was completely fine.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why the ratio was the wrong quantity
&lt;/h2&gt;

&lt;p&gt;The clip it wrongly rejected: 6.4 seconds inside a 15-second video, so 43% — over my threshold. But that segment was one long line of dialogue, delivered at &lt;strong&gt;1.72 words per second&lt;/strong&gt;, which is just... a person talking.&lt;/p&gt;

&lt;p&gt;I spent a while wondering whether 40% should have been 50%, or 60%. That was the wrong question, and it's the part worth generalizing:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A long sentence taking a long time is not a defect.&lt;/strong&gt; My ratio couldn't tell "someone spoke&lt;br&gt;
for six seconds" apart from "someone spoke for one second and then there was five seconds of&lt;br&gt;
nothing," because duration was in the numerator and nothing in the formula represented how&lt;br&gt;
much was actually &lt;em&gt;said&lt;/em&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;No threshold fixes that. Under that measurement the two cases are genuinely identical, so any cutoff that catches one catches the other. Moving the number only chooses which failure you get.&lt;/p&gt;

&lt;p&gt;What separates them is &lt;strong&gt;speech rate&lt;/strong&gt;. Normal delivery runs about 2–3 words per second. Padding a segment with silence adds seconds without adding words, so the rate collapses:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;MIN_WORDS_PER_SEC&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;1.5&lt;/span&gt;

&lt;span class="n"&gt;slowest&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
              &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;segs&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;slowest&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;MIN_WORDS_PER_SEC&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;slowest segment is &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;slowest&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; words/sec &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
          &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;(&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;slowest&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;slowest&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;s of audio) &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
          &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;- silence merged into speech, falling back to manual cues&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same shape of check, still one threshold, completely different behaviour — because the denominator now normalizes by the thing that actually causes the variance.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Clip&lt;/th&gt;
&lt;th&gt;Longest segment&lt;/th&gt;
&lt;th&gt;% of clip&lt;/th&gt;
&lt;th&gt;Words/sec&lt;/th&gt;
&lt;th&gt;Ratio guard&lt;/th&gt;
&lt;th&gt;Rate guard&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Long single line (good)&lt;/td&gt;
&lt;td&gt;6.4s&lt;/td&gt;
&lt;td&gt;43%&lt;/td&gt;
&lt;td&gt;1.72&lt;/td&gt;
&lt;td&gt;❌ rejected&lt;/td&gt;
&lt;td&gt;✅ accepted&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sparse dialogue (broken)&lt;/td&gt;
&lt;td&gt;8.1s&lt;/td&gt;
&lt;td&gt;54%&lt;/td&gt;
&lt;td&gt;0.49&lt;/td&gt;
&lt;td&gt;✅ rejected&lt;/td&gt;
&lt;td&gt;✅ rejected&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The heuristic that survives is the one whose units mean something. "Fraction of clip" is a number about the clip. "Words per second" is a number about speech — and speech was what I was trying to judge all along.&lt;/p&gt;




&lt;h2&gt;
  
  
  The habit that caught my bad fix
&lt;/h2&gt;

&lt;p&gt;I only found the false positive because of a rule I now apply to every guard I write:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Run the regression in &lt;strong&gt;both&lt;/strong&gt; directions. The bad input must be rejected &lt;strong&gt;and&lt;/strong&gt; a&lt;br&gt;
known-good input must still be accepted.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Testing only the bug you just saw is how you ship a fix that breaks something adjacent. My ratio guard passed the test I wrote for it — it rejected the broken clip perfectly. The clip it silently damaged was one I wasn't looking at, because it had never been broken.&lt;/p&gt;

&lt;p&gt;That's the same failure mode as writing your own test fixtures and being pleased when they all pass. If the input and the expected output come out of the same head at the same moment, agreement between them proves nothing. Run the real corpus. Look at what changed for the items you didn't touch.&lt;/p&gt;




&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Never time generated speech by hand.&lt;/strong&gt; Transcribe it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Timing and wording can come from different sources.&lt;/strong&gt; Join them; don't pick a winner.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;When a heuristic uses a proportion, interrogate the denominator.&lt;/strong&gt; If it doesn't represent
the mechanism you're trying to detect, no threshold will save you.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Regression-test in both directions,&lt;/strong&gt; on real inputs, before and after.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Number 3 is the one I keep re-learning. A badly tuned threshold is loud and easy to fix. A well-tuned threshold on the wrong quantity looks like a working system, right up until it quietly throws away something good.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I write about running unattended automation and AI pipelines in production, including the&lt;br&gt;
parts where the fix turns out to be the bug. The full system is the &lt;a href="https://alphatech4.gumroad.com/l/claude-code-automation-playbook" rel="noopener noreferrer"&gt;Claude Code Automation Playbook&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>ai</category>
      <category>machinelearning</category>
      <category>programming</category>
    </item>
    <item>
      <title>Cost doesn't set price: how I got a product decision wrong 4 times in one day</title>
      <dc:creator>Youfu Hsu</dc:creator>
      <pubDate>Mon, 07 Sep 2026 04:52:19 +0000</pubDate>
      <link>https://dev.to/youfuhsu/cost-doesnt-set-price-how-i-got-a-product-decision-wrong-4-times-in-one-day-2nhe</link>
      <guid>https://dev.to/youfuhsu/cost-doesnt-set-price-how-i-got-a-product-decision-wrong-4-times-in-one-day-2nhe</guid>
      <description>&lt;p&gt;I spent a day letting an AI agent research a product decision for me. It changed its recommendation four times. Every reversal was caused by the same mistake, and I didn't see it until the fourth one.&lt;/p&gt;

&lt;p&gt;The mistake is embarrassingly basic: &lt;strong&gt;I was estimating the selling price from the cost.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here's how that plays out when you're moving fast, and the three data sources that fixed it — all free, all public.&lt;/p&gt;




&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;I was screening physical products to advertise. Two questions matter:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Is anyone making money in this category? (demand)&lt;/li&gt;
&lt;li&gt;Could &lt;em&gt;we&lt;/em&gt; make money in it? (unit economics)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I built a gate for question 1 first: scrape the ad library, count how many distinct sellers have been running ads for 30+ days, filter for keyword relevance. Ads that run for a year are a good proxy for "this is profitable for someone."&lt;/p&gt;

&lt;p&gt;That gate worked. Then I needed selling prices for question 2, and I didn't have them. So I did the natural thing: took the supplier cost and multiplied by a typical retail markup.&lt;/p&gt;

&lt;p&gt;That single shortcut caused every reversal below.&lt;/p&gt;




&lt;h2&gt;
  
  
  Reversal 1 → 2: the tax code moved
&lt;/h2&gt;

&lt;p&gt;First recommendation: run the test in the US and UK.&lt;/p&gt;

&lt;p&gt;Then I checked import rules properly. The US &lt;code&gt;de minimis&lt;/code&gt; exemption — which let sub-$800 parcels enter duty-free — ended for all countries in August 2025. Small parcels from China now pay duty &lt;em&gt;plus&lt;/em&gt; a flat customs entry fee of roughly $17 per parcel.&lt;/p&gt;

&lt;p&gt;That flat fee is the killer, and it took me a while to see why:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Product&lt;/th&gt;
&lt;th&gt;Goods cost&lt;/th&gt;
&lt;th&gt;Duty (37.5%)&lt;/th&gt;
&lt;th&gt;Flat entry fee&lt;/th&gt;
&lt;th&gt;Fee as % of goods&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A&lt;/td&gt;
&lt;td&gt;$4.95&lt;/td&gt;
&lt;td&gt;$1.86&lt;/td&gt;
&lt;td&gt;$17.00&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;344%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B&lt;/td&gt;
&lt;td&gt;$14.09&lt;/td&gt;
&lt;td&gt;$5.28&lt;/td&gt;
&lt;td&gt;$17.00&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;121%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C&lt;/td&gt;
&lt;td&gt;$24.80&lt;/td&gt;
&lt;td&gt;$9.30&lt;/td&gt;
&lt;td&gt;$17.00&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;69%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The percentage duty is small — it's charged on the cheap goods value. The &lt;strong&gt;flat fee doesn't care what the item costs&lt;/strong&gt;, so it lands hardest on exactly the low-price items that direct-to-consumer import was built on.&lt;/p&gt;

&lt;p&gt;The UK still has its £135 relief (until October 2028). So: UK, not US.&lt;/p&gt;

&lt;p&gt;Fine. Real finding. But notice it's about &lt;em&gt;landed cost&lt;/em&gt; — still the cost side.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reversal 2 → 3: I trusted stale data
&lt;/h2&gt;

&lt;p&gt;Second recommendation: three specific products, ranked.&lt;/p&gt;

&lt;p&gt;Except two of them had been scanned &lt;em&gt;before&lt;/em&gt; I'd deployed a relevance filter on the ad scraper. Without it, the scraper counted every ad the keyword search returned — including romance novels and migraine coaches that happened to contain the word "ice" or "face."&lt;/p&gt;

&lt;p&gt;I rescanned with the filter on. Two of my top picks turned out to be red oceans (3,000 and 6,500 active ads). The "sellers" in the old data included what were obviously personal accounts and coupon-spam pages, not brands.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Lesson, cheap version:&lt;/strong&gt; when you fix a data collection bug, the old rows don't fix themselves. Write the backfill. I now have a &lt;code&gt;rescan --stale&lt;/code&gt; mode that finds every row collected before the fix and re-runs it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reversal 3 → 4: the estimate was doing all the work
&lt;/h2&gt;

&lt;p&gt;Third recommendation, with real conviction this time: one product, clean on both gates, priced at $102, break-even ROAS under 2.&lt;/p&gt;

&lt;p&gt;Then I went to look at what competitors &lt;em&gt;actually charge&lt;/em&gt; — and the leading brand in that category sells the equivalent product for &lt;strong&gt;$32.99&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I had computed $102 from &lt;code&gt;supplier_cost × 3&lt;/code&gt;. There is no world in which that product sells for $102. At $33, after VAT and payment fees, the ad budget per order is about four dollars. It's not a marginal business, it's an impossible one.&lt;/p&gt;

&lt;p&gt;That was reversal four, and it finally made the pattern visible.&lt;/p&gt;




&lt;h2&gt;
  
  
  The actual lesson
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Cost does not set price. The market sets price. Cost only tells you what's left.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Look at what my markup assumption implied versus reality:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Product&lt;/th&gt;
&lt;th&gt;Supplier cost&lt;/th&gt;
&lt;th&gt;My estimate (×3)&lt;/th&gt;
&lt;th&gt;What the leader actually charges&lt;/th&gt;
&lt;th&gt;Real multiple&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Sleep mask&lt;/td&gt;
&lt;td&gt;$6.95&lt;/td&gt;
&lt;td&gt;$21&lt;/td&gt;
&lt;td&gt;$127 (catalog median)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;18×&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bidet attachment&lt;/td&gt;
&lt;td&gt;$34&lt;/td&gt;
&lt;td&gt;$102&lt;/td&gt;
&lt;td&gt;$33&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1×&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One product supports an 18× markup. Another supports 1×. Same formula, opposite errors — and no amount of care about the cost side would have caught either.&lt;/p&gt;

&lt;p&gt;The fix isn't a better multiplier. It's to &lt;strong&gt;stop estimating a number you can go and measure.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Three free ways to measure the real price
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Shopify's &lt;code&gt;/products.json&lt;/code&gt; is public
&lt;/h3&gt;

&lt;p&gt;Most DTC brands run Shopify, and Shopify exposes the full catalog as JSON with no auth:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://&amp;lt;domain&amp;gt;/products.json?limit=250
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You get every product, every variant, every price:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;urllib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;statistics&lt;/span&gt;

&lt;span class="n"&gt;req&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;urllib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;domain&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/products.json?limit=250&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                             &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;User-Agent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Mozilla/5.0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="n"&gt;products&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;urllib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;urlopen&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;))[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;products&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="n"&gt;prices&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;v&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;price&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
          &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;products&lt;/span&gt;
          &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;variants&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[])&lt;/span&gt;
          &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;price&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;v&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;price&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;products&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;products, median&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;statistics&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;median&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prices&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That five-line query is what ended reversal 4. It's also what showed me something I'd have missed entirely: the category leader has &lt;strong&gt;95 SKUs&lt;/strong&gt;, and most of them are consumables — dryer sheets, cleaning solution, refills. They're not selling one gadget. They're selling a razor-and-blades business, which is why they can afford customer acquisition I can't.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Where do you get the domains?
&lt;/h3&gt;

&lt;p&gt;From the ad copy. Advertisers put their own URL in the ad text, so a regex over scraped ad bodies gives you the competitor list for free:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;DOMAIN&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;\b((?:[a-z0-9][a-z0-9-]*\.)+(?:com|co\.uk|shop|store))\b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;I&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Filter out the obvious noise (&lt;code&gt;facebook.com&lt;/code&gt;, &lt;code&gt;amazon.com&lt;/code&gt;, your ad platform) and rank by frequency — an advertiser's own domain repeats across their ads.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Traffic tells you if the category is growing or dying
&lt;/h3&gt;

&lt;p&gt;Free-tier traffic estimators (3 lookups/day is plenty) give month-over-month change and channel mix. Two competitors, same day:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Leader A&lt;/th&gt;
&lt;th&gt;Leader B&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Monthly visits&lt;/td&gt;
&lt;td&gt;~232K&lt;/td&gt;
&lt;td&gt;~52K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MoM change&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+38.6%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;−10.9%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Top channel&lt;/td&gt;
&lt;td&gt;Paid social 32%&lt;/td&gt;
&lt;td&gt;Paid &lt;strong&gt;search&lt;/strong&gt; 37%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bounce / pages / duration&lt;/td&gt;
&lt;td&gt;42.9% / 3.18 / 2m07s&lt;/td&gt;
&lt;td&gt;74.8% / 1.82 / 31s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two things fell out of that table that nothing else told me:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The category I'd ranked &lt;em&gt;first&lt;/em&gt; on search-trend data had a &lt;strong&gt;shrinking&lt;/strong&gt; leader. Search volume and actual purchase traffic disagreed, and purchase traffic is closer to the money.&lt;/li&gt;
&lt;li&gt;Leader B's biggest channel is paid &lt;em&gt;search&lt;/em&gt;, not paid social. I'd been studying the wrong battlefield for that product.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What I changed in the pipeline
&lt;/h2&gt;

&lt;p&gt;The screening gate now refuses to be confident on estimated prices:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;retail_price&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;score&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;PASS_THRESHOLD&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;verdict&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;WATCH&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;   &lt;span class="c1"&gt;# never PASS on an estimate
&lt;/span&gt;    &lt;span class="n"&gt;reasons&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;price is estimated, not measured — go fetch the real one&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That one rule would have prevented three of the four reversals. Not because the estimate was badly calibrated, but because &lt;strong&gt;an estimate and a measurement should never be allowed to look the same downstream.&lt;/strong&gt; Once &lt;code&gt;$102 (estimated)&lt;/code&gt; becomes just &lt;code&gt;102&lt;/code&gt; in a table, every decision after it inherits a confidence the number never had.&lt;/p&gt;

&lt;p&gt;If you're building anything that ranks options, tag every input with how you got it. The ones you guessed are the ones that will move.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I write about running unattended automation and AI agents in production, including the parts where the agent — or I — confidently get it wrong. The full system I use for this is the &lt;a href="https://alphatech4.gumroad.com/l/claude-code-automation-playbook" rel="noopener noreferrer"&gt;Claude Code Automation Playbook&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>python</category>
      <category>productivity</category>
      <category>devops</category>
    </item>
    <item>
      <title>Exit code 0 is a lie: 7 ways my unattended automation silently did nothing</title>
      <dc:creator>Youfu Hsu</dc:creator>
      <pubDate>Sun, 06 Sep 2026 03:44:16 +0000</pubDate>
      <link>https://dev.to/youfuhsu/exit-code-0-is-a-lie-7-ways-my-unattended-automation-silently-did-nothing-501j</link>
      <guid>https://dev.to/youfuhsu/exit-code-0-is-a-lie-7-ways-my-unattended-automation-silently-did-nothing-501j</guid>
      <description>&lt;p&gt;I run about thirty scheduled jobs on a single Windows box. Some are scrapers, some generate content, some are trading bots, some just check that the other jobs are alive. Most of them were written and are maintained by an AI coding agent that I let run unattended.&lt;/p&gt;

&lt;p&gt;Over three months, every one of the failures below reported &lt;strong&gt;success&lt;/strong&gt;. The scheduler said &lt;code&gt;LastTaskResult = 0&lt;/code&gt;. The logs looked fine or didn't exist. And nothing had happened.&lt;/p&gt;

&lt;p&gt;If you only take one thing from this post: &lt;strong&gt;stop checking exit codes, start checking artifacts.&lt;/strong&gt; I'll get to why at the end. First, the seven ways I got lied to.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. The wrapper that always returns 0
&lt;/h2&gt;

&lt;p&gt;To stop console windows flashing on my desktop every few minutes, I wrapped each scheduled task in a tiny VBScript launcher:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight vb"&gt;&lt;code&gt;&lt;span class="k"&gt;Set&lt;/span&gt; &lt;span class="n"&gt;WshShell&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;CreateObject&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"WScript.Shell"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;WshShell&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Run&lt;/span&gt; &lt;span class="s"&gt;"cmd /c &lt;/span&gt;&lt;span class="se"&gt;""&lt;/span&gt;&lt;span class="s"&gt;python job.py &amp;gt;&amp;gt; job.log 2&amp;gt;&amp;amp;1&lt;/span&gt;&lt;span class="se"&gt;""&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;True&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;0&lt;/code&gt; hides the window. &lt;code&gt;True&lt;/code&gt; waits for completion. I assumed &lt;code&gt;True&lt;/code&gt; also meant the exit code came back. &lt;strong&gt;It does not.&lt;/strong&gt; &lt;code&gt;WshShell.Run&lt;/code&gt; used as a &lt;em&gt;statement&lt;/em&gt; discards the return value, so &lt;code&gt;wscript.exe&lt;/code&gt; exits 0 no matter what the child did.&lt;/p&gt;

&lt;p&gt;I found this because a content pipeline had been dead for five days while the scheduler reported green every single day.&lt;/p&gt;

&lt;p&gt;The fix is to call &lt;code&gt;Run&lt;/code&gt; as a &lt;em&gt;function&lt;/em&gt; and pass the value out:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight vb"&gt;&lt;code&gt;&lt;span class="k"&gt;Set&lt;/span&gt; &lt;span class="n"&gt;WshShell&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;CreateObject&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"WScript.Shell"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;exitCode&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;WshShell&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"cmd /c &lt;/span&gt;&lt;span class="se"&gt;""&lt;/span&gt;&lt;span class="s"&gt;python job.py &amp;gt;&amp;gt; job.log 2&amp;gt;&amp;amp;1&lt;/span&gt;&lt;span class="se"&gt;""&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;WScript&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Quit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;exitCode&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Note the parentheses — required when you're taking a return value. After fixing this across&lt;br&gt;
17 launchers, one task showed a non-zero result &lt;strong&gt;for the first time in its life&lt;/strong&gt;. It had&lt;br&gt;
been failing for weeks.&lt;/p&gt;


&lt;h2&gt;
  
  
  2. The last line of your batch file overwrites the exit code
&lt;/h2&gt;

&lt;p&gt;Fixed the launcher, still got false greens. The next layer down was a &lt;code&gt;.cmd&lt;/code&gt; shim:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight batchfile"&gt;&lt;code&gt;&lt;span class="kd"&gt;node&lt;/span&gt; &lt;span class="kd"&gt;pipeline&lt;/span&gt;.js &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;run&lt;/span&gt;.log &lt;span class="m"&gt;2&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&amp;amp;&lt;/span&gt;&lt;span class="m"&gt;1&lt;/span&gt;
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="kd"&gt;done&lt;/span&gt;&lt;span class="o"&gt;]&lt;/span&gt; &lt;span class="k"&gt;exit&lt;/span&gt; &lt;span class="kd"&gt;code&lt;/span&gt; &lt;span class="nv"&gt;%errorlevel%&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;run&lt;/span&gt;.log
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That &lt;code&gt;echo&lt;/code&gt; is the last command, &lt;code&gt;echo&lt;/code&gt; always succeeds, so the batch file returns &lt;strong&gt;its&lt;/strong&gt; exit code — zero — regardless of what &lt;code&gt;node&lt;/code&gt; did. The log even contained the correct non-zero errorlevel. It just never made it out.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight batchfile"&gt;&lt;code&gt;&lt;span class="kd"&gt;node&lt;/span&gt; &lt;span class="kd"&gt;pipeline&lt;/span&gt;.js &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;run&lt;/span&gt;.log &lt;span class="m"&gt;2&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&amp;amp;&lt;/span&gt;&lt;span class="m"&gt;1&lt;/span&gt;
&lt;span class="kd"&gt;set&lt;/span&gt; &lt;span class="kd"&gt;NODE_EXIT&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;%errorlevel%&lt;/span&gt;
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="kd"&gt;done&lt;/span&gt;&lt;span class="o"&gt;]&lt;/span&gt; &lt;span class="k"&gt;exit&lt;/span&gt; &lt;span class="kd"&gt;code&lt;/span&gt; &lt;span class="nv"&gt;%NODE_EXIT%&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;run&lt;/span&gt;.log
&lt;span class="k"&gt;exit&lt;/span&gt; &lt;span class="na"&gt;/b &lt;/span&gt;&lt;span class="nv"&gt;%NODE_EXIT%&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Debugging heuristic:&lt;/strong&gt; when you fix one layer and still get false greens, assume there's&lt;br&gt;
another layer. Mine was three deep: scheduler → vbs → cmd → python.&lt;/p&gt;


&lt;h2&gt;
  
  
  3. No console, no stdout, no error, no service
&lt;/h2&gt;

&lt;p&gt;A Flask service was set to start at login through a VBS launcher using &lt;code&gt;pythonw.exe&lt;/code&gt; (the GUI-subsystem Python with no console). The script began with a routine encoding guard:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stdout&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;encoding&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stdout&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;reconfigure&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Under &lt;code&gt;pythonw&lt;/code&gt;, &lt;code&gt;sys.stdout&lt;/code&gt; isn't a usable stream. That line raised, uncaught, before the server ever bound its port. No console existed to print the traceback to, so there was no error anywhere. The visible symptom was "the port isn't open after reboot," which sends you straight to firewall and networking — the wrong place entirely.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fix:&lt;/strong&gt; don't use &lt;code&gt;pythonw&lt;/code&gt;. Use regular &lt;code&gt;python.exe&lt;/code&gt; with output redirected to a file, so&lt;br&gt;
stdout is a real stream:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight vb"&gt;&lt;code&gt;&lt;span class="n"&gt;exitCode&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;WshShell&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"cmd /c &lt;/span&gt;&lt;span class="se"&gt;""&lt;/span&gt;&lt;span class="s"&gt;python.exe app.py &amp;gt;&amp;gt; app.log 2&amp;gt;&amp;amp;1&lt;/span&gt;&lt;span class="se"&gt;""&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Heuristic:&lt;/strong&gt; if a background service won't come up, check &lt;code&gt;tasklist&lt;/code&gt; for the process&lt;br&gt;
&lt;em&gt;first&lt;/em&gt;. If it isn't there, the program died on startup — it's not a network problem.&lt;/p&gt;


&lt;h2&gt;
  
  
  4. Two scheduler defaults that quietly kill laptop jobs
&lt;/h2&gt;

&lt;p&gt;Windows Task Scheduler ships with defaults that are reasonable for a desktop and lethal on a laptop:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Setting&lt;/th&gt;
&lt;th&gt;Default&lt;/th&gt;
&lt;th&gt;What it does to you&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;DisallowStartIfOnBatteries&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;true&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Job doesn't run at all when unplugged. No error.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;StopIfGoingOnBatteries&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;true&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Job dies mid-run if you unplug.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;StopOnIdleEnd&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;true&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Job is &lt;strong&gt;killed when you touch the machine&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;StartWhenAvailable&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;false&lt;/td&gt;
&lt;td&gt;A missed run is skipped, not retried&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That third row cost me the most. A job would start on schedule, I'd sit down and move the mouse, idle state would end, and the scheduler would terminate the task mid-flight. The result code was "terminated," which is easy to misread as a crash in your own code.&lt;/p&gt;

&lt;p&gt;Every new task I create now gets all four flipped:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$s&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;New-ScheduledTaskSettingsSet&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-AllowStartIfOnBatteries&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="se"&gt;`
&lt;/span&gt;&lt;span class="w"&gt;     &lt;/span&gt;&lt;span class="nt"&gt;-DontStopIfGoingOnBatteries&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-StartWhenAvailable&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-MultipleInstances&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;IgnoreNew&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;Set-ScheduledTask&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-TaskName&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;$name&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Settings&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;$s&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="nv"&gt;$t&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;Get-ScheduledTask&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-TaskName&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;$name&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="nv"&gt;$t&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Settings&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;IdleSettings&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;StopOnIdleEnd&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="bp"&gt;$false&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;Set-ScheduledTask&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-TaskName&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;$name&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Settings&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;$t&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;Settings&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;StopOnIdleEnd&lt;/code&gt; isn't exposed on &lt;code&gt;New-ScheduledTaskSettingsSet&lt;/code&gt;, which is why it needs the second, uglier step — and why it's the one people miss.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. A comment warning about a trap became the trap
&lt;/h2&gt;

&lt;p&gt;This one is my favourite, because it's so stupid.&lt;/p&gt;

&lt;p&gt;A batch file had a comment reminding future-me about a previous bug:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight batchfile"&gt;&lt;code&gt;&lt;span class="c"&gt;rem 坑備忘：exit /b %errorlevel% 必留&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(It's in Chinese — "trap memo: &lt;code&gt;exit /b %errorlevel%&lt;/code&gt; must stay" — which matters, because that's what makes it multi-byte.)&lt;/p&gt;

&lt;p&gt;The file is UTF-8. &lt;code&gt;cmd.exe&lt;/code&gt; decodes batch files using the system codepage, which on this machine is CP950, not UTF-8. Multi-byte characters get re-paired at the wrong boundaries, the comment gets truncated partway through, and &lt;strong&gt;the remainder of the line is executed as a command.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every morning at 08:30 a console popped up: &lt;code&gt;'errorlevel' is not recognized as an internal or external command&lt;/code&gt;. The pipeline itself succeeded. The exit code was 0. The only symptom was a comment's corpse trying to run.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rule I follow now:&lt;/strong&gt; comments in &lt;code&gt;.cmd&lt;/code&gt;/&lt;code&gt;.bat&lt;/code&gt; files that run unattended are ASCII-only.&lt;br&gt;
If you truly need non-ASCII, &lt;code&gt;chcp 65001 &amp;gt;nul&lt;/code&gt; must be the &lt;em&gt;first&lt;/em&gt; line — it can't rescue anything above itself.&lt;/p&gt;

&lt;p&gt;The same class of bug bites &lt;code&gt;.vbs&lt;/code&gt; (WSH reads it as ANSI — save as UTF-16 LE if it contains non-ASCII) and &lt;code&gt;.ps1&lt;/code&gt; (non-ASCII comments can desync the parser so badly that source lines end up inside string variables).&lt;/p&gt;


&lt;h2&gt;
  
  
  6. Your OS silently blocks the CLI tool you pip-installed
&lt;/h2&gt;

&lt;p&gt;A publishing job died on:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OSError: [WinError 4551] This file is blocked by application control policy
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It was calling &lt;code&gt;yt-dlp&lt;/code&gt; via &lt;code&gt;subprocess&lt;/code&gt;. Windows Smart App Control had decided that this particular unsigned, freshly-updated standalone &lt;code&gt;.exe&lt;/code&gt; had no reputation yet and blocked it. The queue silently backed up for a day.&lt;/p&gt;

&lt;p&gt;The fix is not to weaken your security policy. It's to stop invoking the standalone executable at all:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;executable&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;-m&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;yt_dlp&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;...])&lt;/span&gt;   &lt;span class="c1"&gt;# goes through trusted python.exe
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Most pip-installed CLI tools support &lt;code&gt;-m&lt;/code&gt; invocation. Prefer it in anything unattended.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. The audit script you don't have
&lt;/h2&gt;

&lt;p&gt;Every failure above shares a property: &lt;strong&gt;the status channel said one thing and reality said another.&lt;/strong&gt; So I stopped trusting the status channel.&lt;/p&gt;

&lt;p&gt;I now run two scripts, and they answer deliberately different questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Did it run, and what happened?&lt;/strong&gt; Enumerate every scheduled task, resolve where its log
actually goes (mine hide behind &lt;code&gt;set LOG=&lt;/code&gt; variables, nested vbs→cmd calls, and relative redirects), and read the tail.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Did anything come out?&lt;/strong&gt; For each job, define the artifact it's supposed to produce and
the maximum age that's acceptable. Anything staler than that is flagged, regardless of what the exit code says.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The second one is the important one. It's how I found out a video pipeline had been running "successfully" every day for weeks while producing zero files.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;CHECKS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="c1"&gt;# (name, artifact glob, max age in hours)
&lt;/span&gt;    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;daily-report&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;C:\...\memory\daily_*.md&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;26&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;product-radar&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;C:\...\data\radar_*.csv&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;26&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;video-pipeline&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;C:\...\videos\*.mp4&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;26&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Fifteen lines of config, and it catches an entire class of failure that no amount of exit-code checking will.&lt;/p&gt;




&lt;h2&gt;
  
  
  The pattern
&lt;/h2&gt;

&lt;p&gt;Six of these seven are the same bug wearing different clothes: &lt;strong&gt;a layer between you and the work reported on itself instead of on the work.&lt;/strong&gt; The wrapper reported on the wrapper. The batch file reported on its last &lt;code&gt;echo&lt;/code&gt;. The scheduler reported that it launched something, not that the something did anything.&lt;/p&gt;

&lt;p&gt;So the habit that actually protects you isn't better error handling. It's:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Define, in advance, the observable thing that proves the job did its work.&lt;br&gt;
Then check for &lt;em&gt;that&lt;/em&gt;, on a schedule, and alert when it goes stale.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Exit codes tell you a process ended. Artifacts tell you a job happened. Only one of those is what you actually care about at 3am when you're asleep and the machine is working.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I package the full version of this — 22 documented traps, the two audit scripts, wrapper&lt;br&gt;
templates, and the memory system that stops an AI agent re-learning all of it every session — as the &lt;a href="https://alphatech4.gumroad.com/l/claude-code-automation-playbook" rel="noopener noreferrer"&gt;Claude Code Automation Playbook&lt;/a&gt;. But everything above is standalone; take it and go fix your own false greens.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>claude</category>
      <category>sre</category>
      <category>devops</category>
      <category>python</category>
    </item>
  </channel>
</rss>
