<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Marvin Okafor</title>
    <description>The latest articles on DEV Community by Marvin Okafor (@marvinoka4).</description>
    <link>https://dev.to/marvinoka4</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3374130%2F2d22fef9-b8b3-4b12-b24c-55f422c5e7d4.jpg</url>
      <title>DEV Community: Marvin Okafor</title>
      <link>https://dev.to/marvinoka4</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/marvinoka4"/>
    <language>en</language>
    <item>
      <title>I Use AI Every Day. I'm Also Googling Crop Rotation.</title>
      <dc:creator>Marvin Okafor</dc:creator>
      <pubDate>Sat, 26 Sep 2026 09:48:36 +0000</pubDate>
      <link>https://dev.to/marvinoka4/i-use-ai-every-day-im-also-googling-crop-rotation-2cip</link>
      <guid>https://dev.to/marvinoka4/i-use-ai-every-day-im-also-googling-crop-rotation-2cip</guid>
      <description>&lt;p&gt;I have been a software engineer for a little over eight years.&lt;/p&gt;

&lt;p&gt;For most of those years I did not really use AI. I wrote the code. I read the docs. I sat in the bugs until they made sense. That was the job.&lt;/p&gt;

&lt;p&gt;Then the last year and a half happened.&lt;/p&gt;

&lt;p&gt;I started using AI the way some people start drinking coffee again after years of "I'm fine." Quietly, then constantly. Drafts. Refactors. Tests. That awkward first version of an idea I didn't want to stare at for forty minutes. I use it every day now. I like it. I'm not pretending otherwise.&lt;/p&gt;

&lt;p&gt;It made me faster. It made some boring work less boring. It made me better at the part of the job that was never just typing: deciding what is worth building, what is sloppy, what will break at 2am.&lt;/p&gt;

&lt;p&gt;I am still young. I am still a senior engineer. I am still very much in this AI season of the industry. I am not writing this from a cabin after I "quit tech."&lt;/p&gt;

&lt;p&gt;I am writing this because I caught myself planning a farm.&lt;/p&gt;

&lt;p&gt;Not as a joke. Well, first as a joke. I actually laughed at myself. "Building an agricultural farm." Who talks like that. I meant land. I meant growing things. Crop rotation on one side, and I kept calling it rotation farming until I looked it up. Vegetables and fruit on another. A fishery in a corner. Cocoa somewhere in the mix, at least enough to test whether I can export something that does not live in a Git repo.&lt;/p&gt;

&lt;p&gt;People have to eat.&lt;/p&gt;

&lt;p&gt;That sentence is doing a lot of work in my head lately.&lt;/p&gt;

&lt;p&gt;AI can write a frightening amount of the code I used to treat as my identity. It does not 100% replace a software engineer. Anyone who has shipped a real system knows the last 10% is still full of judgment, context, ugly constraints, and people. But it almost does the part we spent years getting paid for. "Almost" is a strange place to stand when you are still good at the job.&lt;/p&gt;

&lt;p&gt;So I am not shutting the laptop.&lt;/p&gt;

&lt;p&gt;I am also not going to pretend the only rational next move is to become the world's best prompt engineer and hope the market stays kind.&lt;/p&gt;

&lt;p&gt;The thing I don't want to lose is the entrepreneurial part. The part that likes creating something and actually delivering it. Value you can point at. Not a slide. Not a sprint demo that evaporates after the next rewrite.&lt;/p&gt;

&lt;p&gt;Code is wonderful. Code is also weightless. A model can swallow a lot of weightless work.&lt;/p&gt;

&lt;p&gt;A pond full of fish is not weightless. Cocoa that has to be harvested, dried, and sold is not weightless. Food is an old industry with new tools, but the product still has to exist in the world. Manufacturing is like that. Farming is like that. A lot of "boring" sectors are like that. They can use AI. They are not made of AI.&lt;/p&gt;

&lt;p&gt;That is the thought, not a manifesto.&lt;/p&gt;

&lt;p&gt;Stay in software while it still makes sense. Use the tools. Get better with them. Keep the craft.&lt;/p&gt;

&lt;p&gt;And in the same season, put some of that restlessness into something AI cannot fully eat. Land. Water. Crops that follow a calendar instead of a changelog. A small business that fails or works in daylight.&lt;/p&gt;

&lt;p&gt;I don't have the farm yet. I have tabs. I have questions about soil and seasons I cannot autocomplete my way through. I have the slightly embarrassing feeling of a senior engineer discovering that "retirement plan" for me does not mean a beach. It means I might want to grow food and sell cocoa and still write software on the side because I like both.&lt;/p&gt;

&lt;p&gt;If you have been in this industry long enough to remember working without Copilot, and you now work with it every day, you probably know the feeling. Gratitude and unease in the same week.&lt;/p&gt;

&lt;p&gt;I got sick recently and almost let a whole week pass with nothing published. That is a small version of the same fear. Time moving, output thinning, the sense that if you don't plant something on purpose the season just goes.&lt;/p&gt;

&lt;p&gt;So this is the plant for today.&lt;/p&gt;

&lt;p&gt;I am not announcing that engineers should all become farmers. Please do not @ me with a tractor.&lt;/p&gt;

&lt;p&gt;I am saying that if AI is doing more of the work you used to guard with your chest, it is not crazy to ask a second question next to "how do I stay employable?"&lt;/p&gt;

&lt;p&gt;The second question is: what can I build that still needs a human when the code is cheap?&lt;/p&gt;

&lt;p&gt;For me, right now, that question smells like soil. And fish. And cocoa. And a software career I have not finished.&lt;/p&gt;

&lt;p&gt;The engines are still on. I am just pointing one of them at a field.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>career</category>
      <category>discuss</category>
      <category>watercooler</category>
    </item>
    <item>
      <title>69 Tests. All Passing. Zero Bugs Caught.</title>
      <dc:creator>Marvin Okafor</dc:creator>
      <pubDate>Thu, 17 Sep 2026 22:19:31 +0000</pubDate>
      <link>https://dev.to/marvinoka4/69-tests-all-passing-zero-bugs-caught-27k5</link>
      <guid>https://dev.to/marvinoka4/69-tests-all-passing-zero-bugs-caught-27k5</guid>
      <description>&lt;p&gt;An AI model wrote 69 tests for a Python module. Every one passed. Together they caught zero of the eleven bugs I had deliberately planted in that module.&lt;/p&gt;

&lt;p&gt;A second setup, pointed at the specific bugs rather than at the module, used 17 attempts and caught ten.&lt;/p&gt;

&lt;p&gt;That contrast is the whole project.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I built
&lt;/h2&gt;

&lt;p&gt;A mutation-testing harness. It changes source code in small ways, flips a comparison, alters a constant, deletes a &lt;code&gt;raise&lt;/code&gt;, then runs the existing test suite and records which changes the suite fails to notice. A change nothing catches is a fault your tests cannot detect.&lt;/p&gt;

&lt;p&gt;Coverage tells you a line ran. This tells you whether anything would have failed if that line had been wrong. Those are very different numbers. A toy module with one happy-path test showed 47% line coverage and caught 2 of 21 mutations.&lt;/p&gt;

&lt;p&gt;Then I pointed an agent at the mutations the tests missed. For each one it writes a single test, and that test is kept only if it both passes on clean code and fails on that specific mutation. Ground truth is a subprocess exit code. No model judges any outcome.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I found
&lt;/h2&gt;

&lt;p&gt;Three approaches, same model, same token ceiling, twelve widely used Python libraries including cachetools, toolz, tenacity and boltons. 455 mutations generated, 53 of which survive the existing tests and sit on lines those tests actually execute.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;approach&lt;/th&gt;
&lt;th&gt;caught&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;one prompt, "write more tests" (556 tests)&lt;/td&gt;
&lt;td&gt;9 of 53&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;one test per call, no targeting (53 tests)&lt;/td&gt;
&lt;td&gt;2 of 53&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;targeted at the specific fault, with a pass/fail gate&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;44 of 53&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I hand-audited the nine it missed. Seven are provably unkillable, six of those being type annotations inside &lt;code&gt;TYPE_CHECKING&lt;/code&gt; blocks that never execute at runtime. So 44 of 46 that could be caught at all.&lt;/p&gt;

&lt;p&gt;Three findings I think matter more than that number.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Most undetected faults are unreached code, not weak assertions.&lt;/strong&gt; Only 53 of 133 surviving mutations sit on a line the tests execute at all. I suspected my test commands were scoped too narrowly, so I widened every one of them by 6 to 40 times. The count went from 54 to 53. Down. In mature, human-written code, tests do not mostly execute code without checking it. They mostly do not execute it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The gate never once rejected a broken test.&lt;/strong&gt; Across 74 attempts, every rejected draft was a valid, passing test that simply failed to detect the fault. Not one was broken. The gate turns out to have exactly one job in practice, and it is not the job I designed it for.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The tests it writes catch the fault they were shown and almost nothing else.&lt;/strong&gt; Zero cross-function transfer across 44 kept tests. 36 of the 44 catch exactly one mutation. That is the uncomfortable result, and it belongs next to the 44 rather than underneath it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it actually mattered
&lt;/h2&gt;

&lt;p&gt;Building the agent took a weekend. The rest of the time went into discovering that my measuring instrument kept lying to me.&lt;/p&gt;

&lt;p&gt;Eleven bugs in the harness itself. Editable installs that made mutations invisible. Parallel execution corrupting one target. A classifier running on the wrong unit. A bytecode cache serving stale results. An outcome bucket that had never once filled, because it matched on a string my version of pytest does not emit. I had published that empty bucket as a finding.&lt;/p&gt;

&lt;p&gt;Two things they had in common.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Every one made the results look better, or made an absence look like evidence.&lt;/strong&gt; That is selection rather than conspiracy. Debugging is triggered by surprise, and a pleasing result is not surprising. So the filter that removes measurement bugs gets applied unevenly, hard against results you dislike and softly against results you like. The ones that flatter you survive to publication.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Not one was found by reading the code.&lt;/strong&gt; Every single one was caught by running a check whose outcome I had predicted in advance and getting the wrong answer.&lt;/p&gt;

&lt;p&gt;Then I published, and three readers found three more. All of them by pointing at my checks, never at my numbers. Nobody has disputed a single result. Every correction landed on the instrument.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would tell you
&lt;/h2&gt;

&lt;p&gt;If you build evaluations for your own work, the harness is the part worth publishing. A result is a claim people can take or leave. An instrument is something they can attack, and the attacks are what tell you whether it works.&lt;/p&gt;

&lt;p&gt;I got more out of three comment threads than out of forty hours of my own review.&lt;/p&gt;

&lt;p&gt;Repo, with all eleven bugs documented and every result reproducible from a clean clone: &lt;a href="https://github.com/marvinoka4/killcheck" rel="noopener noreferrer"&gt;killcheck repo&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The longer write-ups, if you want them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dev.to/marvinoka4/nine-bugs-in-my-own-evaluation-harness-every-one-made-my-results-look-better-4fkj"&gt;The eleven bugs, and why they all pointed one way&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/marvinoka4/six-checks-to-run-before-you-trust-a-number-your-own-harness-produced-2fbk"&gt;Six checks to run before trusting your own harness&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/marvinoka4/two-readers-broke-my-checks-not-one-of-my-numbers-moved-nmd"&gt;What happened when readers audited the checks&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Thanks to Vinh Nguyen (&lt;a href="https://dev.to/vinhnguyenthanhdn"&gt;@vinhnguyenthanhdn&lt;/a&gt;), Ahmet Özel (&lt;a href="https://dev.to/ahmetozel"&gt;@ahmetozel&lt;/a&gt;) and Zain Dana Harper (&lt;a href="https://dev.to/zaindanaharper"&gt;@zaindanaharper&lt;/a&gt;). Three findings, three checks improved, zero numbers moved.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>python</category>
      <category>programming</category>
    </item>
    <item>
      <title>Six Checks to Run Before You Trust a Number Your Own Harness Produced.</title>
      <dc:creator>Marvin Okafor</dc:creator>
      <pubDate>Sun, 13 Sep 2026 21:27:38 +0000</pubDate>
      <link>https://dev.to/marvinoka4/six-checks-to-run-before-you-trust-a-number-your-own-harness-produced-2fbk</link>
      <guid>https://dev.to/marvinoka4/six-checks-to-run-before-you-trust-a-number-your-own-harness-produced-2fbk</guid>
      <description>&lt;p&gt;I built a measurement harness. I found ten bugs in it. Every one made my results look better than they were, and not one was found by reading the code.&lt;/p&gt;

&lt;p&gt;Here is the checklist that came out of that. It is meant to be used rather than admired, so each check comes with what it proves, what it does not prove, and roughly what it costs to write.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the bugs point one way
&lt;/h2&gt;

&lt;p&gt;Debugging is triggered by surprise. A pleasing result is not surprising.&lt;/p&gt;

&lt;p&gt;So the filter that removes measurement bugs from your work is applied unevenly. Hard against results you dislike, softly against results you like. Every pass removes more unflattering bugs than flattering ones, and after a few days your instrument has drifted in one direction while nothing in your process is designed to notice.&lt;/p&gt;

&lt;p&gt;That is the whole mechanism. It does not require anyone to be dishonest, which is why being careful does not fix it.&lt;/p&gt;

&lt;p&gt;I wrote that argument up in full, with all ten bugs and the direction each one pushed, &lt;a href="https://dev.to/marvinoka4/nine-bugs-in-my-own-evaluation-harness-every-one-made-my-results-look-better-4fkj"&gt;here&lt;/a&gt;. The rest of this post is the checklist rather than the case for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check 1: the canary
&lt;/h2&gt;

&lt;p&gt;Replace the thing under test with something unparseable and assert that your harness notices.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_canary_reaches_interpreter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;target&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;swapped_source&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;target&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;this is not valid python&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;run_suite&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;target&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;failed&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;harness did not notice a broken module&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Proves:&lt;/strong&gt; your change actually reaches the interpreter.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does not prove:&lt;/strong&gt; that the change is isolated, or that your scoring is correct, or that a &lt;em&gt;valid&lt;/em&gt; change reaches the interpreter. Mine passed happily while a same-length mutation was being silently ignored, because unparseable garbage has a different byte length and therefore invalidates Python's bytecode cache. Vinh Nguyen (&lt;a href="https://dev.to/vinhnguyenthanhdn"&gt;@vinhnguyenthanhdn&lt;/a&gt;) found that one. A byte-size-preserving variant of the canary closes it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Caught in my project:&lt;/strong&gt; editable installs. &lt;code&gt;pip install -e&lt;/code&gt; on src-layout packages resolved imports back to the original checkout, so modifications written to a temp copy never executed. Three targets scored 0.000, which reads as "these test suites are terrible" rather than "my harness is broken."&lt;/p&gt;

&lt;p&gt;Write this before you write any measurement at all. It is the cheapest check on the list, and it catches the most embarrassing class of failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check 2: the determinism gate
&lt;/h2&gt;

&lt;p&gt;Run the same scoring three times, serially, and require byte-identical output.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;runs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;target&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt;
&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;runs&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;runs&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;runs&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nondeterministic: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;runs&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Proves:&lt;/strong&gt; execution is isolated between runs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does not prove:&lt;/strong&gt; that you are executing the right thing. Three identical runs of the wrong file are still perfectly deterministic.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Caught in my project:&lt;/strong&gt; parallel execution. Running mutants concurrently produced three different results across four runs on the one target doing real async I/O. I had already recorded an improvement off that data. It was noise.&lt;/p&gt;

&lt;p&gt;The important thing is that checks 1 and 2 do not substitute for each other. The canary passes happily while concurrency corrupts your results. The determinism gate passes happily while your imports resolve to the wrong file. You need both, and you need to write down what each one actually proves so you do not talk yourself into believing one covers the other.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check 3: the negative control
&lt;/h2&gt;

&lt;p&gt;Build one case where the score &lt;em&gt;should&lt;/em&gt; be near zero.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# a suite that imports, calls, and asserts nothing meaningful
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_negative_control&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;score&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;run_scoring&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;vacuous_suite&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;target&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;score&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;THRESHOLD&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;vacuous suite scored &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Proves:&lt;/strong&gt; your scoring is not generous in a way you would otherwise never look for.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does not prove:&lt;/strong&gt; anything about the top of the range.&lt;/p&gt;

&lt;p&gt;This is the one I did not have, and it came from Ahmet Özel (&lt;a href="https://dev.to/ahmetozel"&gt;@ahmetozel&lt;/a&gt;). The logic is worth stating carefully. Everywhere else in a harness, high is good and low prompts investigation. On a negative control, high is the alarm. That gives you exactly one place where a flattering failure is the surprising one, which is the only condition under which debugging fires reliably.&lt;/p&gt;

&lt;p&gt;There is a critical design constraint. &lt;strong&gt;The control must fail loudly if it passes for a boring reason.&lt;/strong&gt; A near-zero score is also what you get when nothing loaded, which is precisely what my first bug produced. So the control has to assert separately that the suite passes on clean source, that the module actually imported and executed, and that your other checks still fire.&lt;/p&gt;

&lt;p&gt;Mine did not come back at zero. It came back at 7 of 51, and the explanation was a genuine calibration fact: &lt;code&gt;assert x is not None&lt;/code&gt; is a real detector, just an extremely narrow one. It catches exactly one class of fault. All seven kills were that one operator and zero were anything else.&lt;/p&gt;

&lt;p&gt;So a suite of vacuous tests has a non-zero floor rather than a zero one. A control that came back clean would have taught me nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check 4: assert the property, not the proxy
&lt;/h2&gt;

&lt;p&gt;When you write a regression test, check the thing you care about rather than a signal correlated with it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# proxy: correlated with correctness, until it isn't
&lt;/span&gt;&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="nf"&gt;count_cache_files&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;workdir&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;

&lt;span class="c1"&gt;# property: the thing you actually need to be true
&lt;/span&gt;&lt;span class="n"&gt;before&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;observe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;target&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;mutated_source&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;target&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;after&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;observe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;target&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;after&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;before&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;source changed, but behaviour did not&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Proves:&lt;/strong&gt; the behaviour you depend on actually holds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does not prove:&lt;/strong&gt; that you picked the right property. But it removes a failure mode the proxy has for free.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Caught in my project:&lt;/strong&gt; my fix for the bytecode problem asserted a cache-file count of zero in the temp copy. Set &lt;code&gt;PYTHONPYCACHEPREFIX&lt;/code&gt; and bytecode goes to a central tree keyed on the copy's absolute path. Zero cache files arrive in the copy, and the stale read happens anyway. My test would have passed while the harness lied. Vinh found that too, and ran three branches on his own machine to isolate which variable mattered.&lt;/p&gt;

&lt;p&gt;The proxy held under my configuration and failed under a supported environment variable. The property costs the same to assert and does not have that failure mode. That is the entire argument.&lt;/p&gt;

&lt;p&gt;Add a meta-test that reproduces the broken case, so you know the regression test has teeth. A regression test nobody has ever seen fail belongs in the same category as a canary that cannot detect what it claims to.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check 5: predict the outcome before you run it
&lt;/h2&gt;

&lt;p&gt;This is not a check you write. It is a habit, and it caught four of my ten.&lt;/p&gt;

&lt;p&gt;Before running any diagnostic, write down what you expect to see. One line is enough. Then run it, and treat any contradiction as stop-and-investigate rather than as something to rationalise on the spot.&lt;/p&gt;

&lt;p&gt;There is no code for this one, which is precisely why it is easy to skip. The prediction has to exist before the output does, and nothing in your tooling will remind you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proves:&lt;/strong&gt; nothing on its own.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does what nothing else does:&lt;/strong&gt; it manufactures surprise where none would otherwise exist. A check you run without a prediction just produces another number, and you will interpret that number the same way you interpret all your other numbers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Caught in my project:&lt;/strong&gt; I knew, independently, that one specific generated test was broken. So I predicted that removing it would make the suite go green. It did not. That contradiction is the only reason I found a bug where a reconstruction step was dropping shared imports and manufacturing failures that were not real. The scores it produced looked entirely plausible, and plausible in a direction I liked.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check 6: run it where you don't develop it
&lt;/h2&gt;

&lt;p&gt;Install into a fresh environment, outside your own tree, and point it at a repository you did not build around.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;mktemp&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
pip &lt;span class="nb"&gt;install&lt;/span&gt; &amp;lt;your-tool&amp;gt;
&amp;lt;your-tool&amp;gt; run &lt;span class="nt"&gt;--target&lt;/span&gt; some-repo-you-never-tested-on
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Proves:&lt;/strong&gt; the thing works for someone who is not you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does not prove:&lt;/strong&gt; correctness. It proves that your correctness is reachable by someone else, which is a separate problem. Your checks can all be sound and still never run for a user, because the path they take through your tool is not the path you take.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Caught in my project:&lt;/strong&gt; my tenth bug, and the one I find most instructive. I had fixed a problem across twelve curated repositories. Then I shipped a CLI where the check that catches that problem was not wired into the command my own quickstart tells people to run first. A stranger pointing it at a normal project would get a confident 0.0000 with no warning at all.&lt;/p&gt;

&lt;p&gt;Everything looked correct from inside the repository where I develop it. That is the point. Your development environment is the one configuration you have accidentally optimised for, and it is the one your users are least likely to reproduce.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this checklist does not cover
&lt;/h2&gt;

&lt;p&gt;Every check above validates the execution path. That your change reached the interpreter, that runs are isolated, that the behaviour actually differs. None of them validates the scorer: the code that reads a result and decides what it means, then aggregates that across everything else.&lt;/p&gt;

&lt;p&gt;That gap is not theoretical. Three of my ten bugs were scorer-level. A classifier run on the wrong unit, a reconstruction step dropping shared imports, an assertion style the classifier couldn't see. Not one was caught by a check on this list. All three came from check 5, predicting an outcome and hitting a contradiction, which is a habit rather than a control.&lt;/p&gt;

&lt;p&gt;Zain Dana Harper (&lt;a href="https://dev.to/zaindanaharper"&gt;@zaindanaharper&lt;/a&gt;) put the general version of this better than I can: an intact artifact tells you nothing about whether the thing interpreting it is correct. A report can be perfectly well-formed, every byte verified, and carry a wrong number.&lt;/p&gt;

&lt;p&gt;I'm building three checks for this now. Known-outcome fixtures that assert the pipeline reports an answer you know by construction. Conservation invariants so counts have to reconcile at every stage. Explicit unit metadata on anything carrying a count, so a consumer can assert what it's being handed rather than assume.&lt;/p&gt;

&lt;p&gt;The first one found a bug on its first run. My outcome classifier decided a result was a collection error only if the output contained a specific phrase, one that the pytest version I'm on never actually emits. So that branch had never fired, for any target, ever. Worse than the bug: I had published the empty bucket as a finding. "Zero error outcomes across all twelve targets" read as reassurance about data quality when it was the signature of dead code. A uniform zero across twelve independent targets should have been suspicious on its own. Real data is rarely that clean.&lt;/p&gt;

&lt;p&gt;That one doesn't fit the pattern in the rest of this post, and it's worth naming why. It didn't inflate a metric, and the primary number is unaffected. What it did was make a null result look like evidence. Until the scorer checks are finished, treat this list as covering execution and leave interpretation to check 5.&lt;/p&gt;

&lt;h2&gt;
  
  
  The order to do these in
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Canary&lt;/strong&gt; first, before any measurement exists at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Determinism gate&lt;/strong&gt; before you trust any number that comes out.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Negative control&lt;/strong&gt; before you interpret a good result.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fresh-environment test&lt;/strong&gt; before anyone else runs your tool.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Property-not-proxy&lt;/strong&gt; every single time you write a regression test.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Predict-then-run&lt;/strong&gt; on every diagnostic, forever.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The ordering matters because each one gets harder to run honestly once you have results you are attached to.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three traps that no check catches
&lt;/h2&gt;

&lt;p&gt;These are scoring decisions rather than checks. Nothing on the list above will catch them, so you have to decide them deliberately and then say what you decided.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Batch versus per-test scoring.&lt;/strong&gt; If one broken item invalidates a whole batch, your number is a floor rather than a measurement. Both are defensible. Publishing one while implying the other is not. Say which you have.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Budget matching.&lt;/strong&gt; When you compare approaches that spend resources differently, there is no neutral unit. Matching calls starves one arm of output. Matching output tokens effectively rebuilds a different arm. Every choice advantages someone. Report the full resource vector, and say plainly that the unit was chosen before you saw the results, assuming it was.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Controls die at small n.&lt;/strong&gt; I built two versions of a held-out control and abandoned both. One ended up with a denominator of 1. The right move is to state that you have no control, not to dress up something weaker and call it one. A weak control that nobody flags is worse than an admitted absence, because it transfers confidence you have not earned.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this costs and what it bought
&lt;/h2&gt;

&lt;p&gt;Ten bugs, all flattering, none found by inspection.&lt;/p&gt;

&lt;p&gt;Then two readers found two more by pointing at my checks rather than at my numbers. Neither time did a result move. Nobody disputed a single finding. All of the scrutiny landed on the instrument, and both times the instrument was wrong in a way that could not be seen by reading it.&lt;/p&gt;

&lt;p&gt;That is the argument for publishing the instrument rather than the finding. A result is a claim people can take or leave. An instrument is something they can attack, and the attacks are what tell you whether it works.&lt;/p&gt;

&lt;p&gt;I am considering putting the full version of this together, with the checks as drop-in code, the pre-registration and commit-ordering templates, and the eval set as a worked example.&lt;/p&gt;

&lt;p&gt;If you build evaluations, what would you want in it that isn't here, and what would you not bother with? The second question is the more useful one to me.&lt;/p&gt;

</description>
      <category>testing</category>
      <category>python</category>
      <category>ai</category>
      <category>programming</category>
    </item>
    <item>
      <title>Two Readers Broke My Checks. Not One of My Numbers Moved.</title>
      <dc:creator>Marvin Okafor</dc:creator>
      <pubDate>Fri, 11 Sep 2026 23:20:40 +0000</pubDate>
      <link>https://dev.to/marvinoka4/two-readers-broke-my-checks-not-one-of-my-numbers-moved-nmd</link>
      <guid>https://dev.to/marvinoka4/two-readers-broke-my-checks-not-one-of-my-numbers-moved-nmd</guid>
      <description>&lt;p&gt;Last week I published &lt;a href="https://dev.to/marvinoka4/nine-bugs-in-my-own-evaluation-harness-every-one-made-my-results-look-better-4fkj"&gt;a post about nine measurement bugs I found in my own evaluation harness&lt;/a&gt;, all of which happened to favour my results. Two readers replied with things I hadn't found.&lt;/p&gt;

&lt;p&gt;Both were about my checks rather than my numbers. I verified both and fixed both. My results did not change at all.&lt;/p&gt;

&lt;p&gt;That turned out to be the interesting part.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reader one: a canary that could not detect what it claimed to
&lt;/h2&gt;

&lt;p&gt;Some setup. The harness mutates a source file in a temp copy and runs the test suite against it. To prove that mutations actually reach the interpreter, I had a canary: overwrite the module with unparseable garbage and assert the suite fails.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/vinhnguyenthanhdn"&gt;Vinh Nguyen&lt;/a&gt; pointed out that CPython's default bytecode invalidation keys on source mtime in whole seconds plus source size. A mutation flipping &lt;code&gt;&amp;lt;&lt;/code&gt; to &lt;code&gt;&amp;gt;&lt;/code&gt; is byte-size identical. If a stale &lt;code&gt;.pyc&lt;/code&gt; is sitting there with a matching header, Python skips recompilation and imports the original bytecode.&lt;/p&gt;

&lt;p&gt;The mutation is on disk. The original behaviour executes. The suite passes. The mutant gets recorded as surviving.&lt;/p&gt;

&lt;p&gt;Here is why my canary could not catch that. Unparseable garbage has a different byte length. So the pyc genuinely does invalidate, so the canary fires, so it reports success. It was answering a question I wasn't asking.&lt;/p&gt;

&lt;p&gt;I checked the mechanism first.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# compile a pyc from this
&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;

&lt;span class="c1"&gt;# overwrite with a byte-identical mutation, force mtime back
&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;

&lt;span class="c1"&gt;# import: stale bytecode executes, original comparison runs
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Reproduced on the first try.&lt;/p&gt;

&lt;p&gt;Then I checked whether it reaches anything in my harness. It doesn't. Every temp-copy call site passes &lt;code&gt;ignore_patterns("__pycache__", "*.pyc", ...)&lt;/code&gt;, so no pyc ever arrives in the copy. I verified it by copying a target with 19 real pyc files and confirming zero arrive.&lt;/p&gt;

&lt;p&gt;Here is the part I keep thinking about. I wrote that line to avoid copying junk. It turned out to be load-bearing for correctness, and I had no idea until someone pointed at the mechanism.&lt;/p&gt;

&lt;p&gt;The reason an instrument is correct is often not the reason you wrote the code that makes it correct. A line added for tidiness that silently does safety work is a line someone removes during a cleanup, with no way to know what it was holding up.&lt;/p&gt;

&lt;p&gt;I added a byte-size-preserving canary alongside the original. Eight of twelve targets are eligible and all pass. The other four report N/A rather than being silently skipped.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reader two: the negative control I didn't have
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://dev.to/ahmetozel"&gt;Ahmet Özel&lt;/a&gt; came at it from a different direction. His reading of the original post was that "debugging is triggered by surprise" explains a one-directional bias without anyone needing to be dishonest, which is exactly the argument I was making.&lt;/p&gt;

&lt;p&gt;His counter-measure was one I didn't have. Keep a deliberate negative control. A case where the score should be near zero.&lt;/p&gt;

&lt;p&gt;Everywhere else in a harness, high is good and low prompts investigation. On a negative control, high is the alarm. That gives you one place where a flattering failure is the surprising one, which is the only condition under which debugging reliably fires.&lt;/p&gt;

&lt;p&gt;He also made a sharper point about pre-registration, which I &lt;em&gt;was&lt;/em&gt; doing. Writing down your expected number helps when you're wrong. It does nothing when the instrument is broken in a way that produces the number you predicted.&lt;/p&gt;

&lt;p&gt;One of my original nine was exactly that. A classifier misread &lt;code&gt;self.assertEqual&lt;/code&gt; as "no assertion" while I was testing a hypothesis that models write assertion-free tests. It would have handed me my own prediction with a confirmation attached.&lt;/p&gt;

&lt;p&gt;So I built the control. A suite of deliberately vacuous tests. Import the module, call things, assert nothing meaningful. I expected near zero.&lt;/p&gt;

&lt;p&gt;It killed 7 of 51 mutants. 13.7%.&lt;/p&gt;

&lt;p&gt;That is not a bug. &lt;code&gt;assert x is not None&lt;/code&gt; is a real detector, just a very narrow one. It catches exactly one class of mutation: a function that started returning &lt;code&gt;None&lt;/code&gt;. All seven kills were that operator. Zero were constant mutations and zero were comparison mutations.&lt;/p&gt;

&lt;p&gt;So a suite of vacuous tests has a non-zero floor rather than a zero one. That is a calibration fact about my scoring that I would never have learned from a control that came back clean.&lt;/p&gt;

&lt;p&gt;The design constraint is the practical part, and it took me a second attempt to get right. A negative control has to fail loudly if it passes for a boring reason. A near-zero score is also what you get when the module never loaded, which is precisely what bug 1 in my last post produced. Three targets at 0.000 that read as a finding rather than as a fault.&lt;/p&gt;

&lt;p&gt;So the control asserts separately that the suite passes on clean source, that the module actually imported and executed (coverage-confirmed, 37 lines), and that both canaries still fire.&lt;/p&gt;

&lt;h2&gt;
  
  
  Round two: my fix asserted a proxy
&lt;/h2&gt;

&lt;p&gt;Vinh came back after I described the fix and found a hole in it.&lt;/p&gt;

&lt;p&gt;My regression test asserted &lt;code&gt;pyc_count == 0&lt;/code&gt; in the temp copy. But set &lt;code&gt;PYTHONPYCACHEPREFIX&lt;/code&gt; and bytecode goes to a central tree keyed on the copy's absolute path. Zero pyc files arrive in the copy and the stale read happens anyway.&lt;/p&gt;

&lt;p&gt;My test would have passed while the harness was lying.&lt;/p&gt;

&lt;p&gt;He ran three branches on his own machine to isolate the real dependency:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;no prefix, fixed work path: fresh&lt;/li&gt;
&lt;li&gt;prefix, same fixed work path: stale, zero pyc in the copy&lt;/li&gt;
&lt;li&gt;prefix, unique directory per mutant: fresh&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I reproduced all three. What actually protects me is a second accidental thing. Every copy site uses &lt;code&gt;tempfile.TemporaryDirectory()&lt;/code&gt;, which gives a unique path per mutant, which is branch three. I verified it across every call site and ran a full twelve-target sweep with the prefix set. It diffs byte-identical against the committed results.&lt;/p&gt;

&lt;p&gt;Two accidental protections. Neither written for the reason it works. &lt;code&gt;ignore_patterns&lt;/code&gt; to avoid junk. &lt;code&gt;TemporaryDirectory&lt;/code&gt; because it cleans itself up. Both are now commented as load-bearing and name the environment variable explicitly, because anyone adding a &lt;code&gt;--work-dir&lt;/code&gt; flag would remove one of them and have no way to know.&lt;/p&gt;

&lt;p&gt;The rule underneath this is the most useful thing I got out of the whole exchange.&lt;/p&gt;

&lt;p&gt;My test asserted a proxy. Pyc count is a proxy. "The observed behaviour actually changed" is the property. The proxy held under my configuration and failed under a supported environment variable. The property costs the same to assert and does not have that failure mode.&lt;/p&gt;

&lt;p&gt;So I rewrote the test to push a byte-size-preserving mutant through the real code path and assert the observed outcome isn't "survived." I added a meta-test that reproduces the stale branch to prove the check has teeth. A regression test nobody has ever seen fail sits in the same category as the canary that couldn't detect this in the first place.&lt;/p&gt;

&lt;h2&gt;
  
  
  One result that came along the way
&lt;/h2&gt;

&lt;p&gt;While I was in there I closed something I had deferred: hand-labelling the nine mutants the agent couldn't kill. Seven are provably equivalent, six of them type-hint mutations inside &lt;code&gt;TYPE_CHECKING&lt;/code&gt; blocks that are never evaluated at runtime. Two are genuine misses, traced to the agent testing the wrong call pattern and the wrong function. So the claim moves from 44 of 53 with an unbounded equivalence caveat to 44 of 46 killable, 95.7%, with the per-mutant arguments published.&lt;/p&gt;

&lt;h2&gt;
  
  
  Update: the tenth bug, three weeks later.
&lt;/h2&gt;

&lt;p&gt;After publishing this, I packaged the harness as a CLI so other people could run it, then field-tested it against three repos I hadn't built it around: pytest, attrs, and jsonschema.&lt;br&gt;
It returned &lt;code&gt;0.0000&lt;/code&gt; on pytest. Zero of 35 mutations detected, reported confidently, with no warning that anything was wrong.&lt;br&gt;
It's the src-layout bug from item one of this list. The editable install resolves imports back to the original checkout, so the mutated copy is never executed. Exactly the bug I said I'd found and fixed, which I had - for the twelve repositories I curated by hand.&lt;br&gt;
What I hadn't done was wire the check into the command people actually run. The canary exists. It works. It catches this every time. But it only ran under &lt;code&gt;verify&lt;/code&gt;, and &lt;code&gt;score&lt;/code&gt; is the command my own quickstart puts first. So a stranger installing the tool, pointing it at a normal src-layout project, would get a confidently wrong number reading "your test suite detects nothing" when the truth is "my harness never executed your code."&lt;br&gt;
The shape is worth sitting with. Both reader findings in this post were about checks that couldn't detect what they claimed. This one is a check that could, and simply wasn't called. Same outcome for the user. And I found it the same way as all the others - not by reading the code, but by running it somewhere I hadn't run it before and getting an answer I couldn't explain.&lt;br&gt;
&lt;code&gt;score&lt;/code&gt; and &lt;code&gt;harden&lt;/code&gt; now abort on canary failure rather than returning a number. There's a &lt;code&gt;--skip-canary&lt;/code&gt; for people who know what they're doing, and it says loudly that the results are unverified. The error message names src-layout as the likely cause and suggests the fix, because "mutations are not reaching this target's test process" tells a user nothing.&lt;br&gt;
One more thing that came out of it. pytest is a hard case - it imports itself before test collection runs, so the usual &lt;code&gt;-o pythonpath=src&lt;/code&gt; fix doesn't work. Setting &lt;code&gt;PYTHONPATH&lt;/code&gt; before the interpreter starts does. With that, pytest scores 0.657 instead of aborting. So the tool needed a way to set environment variables, not just command-line flags, which I'd never have discovered without a target that broke in an unusual way.&lt;br&gt;
Nine bugs became ten. The tenth is the one I'd have been least likely to find on my own, because everything looked correct from inside the repository where I develop it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I actually took from this
&lt;/h2&gt;

&lt;p&gt;Two rounds. Two readers. Four checks improved. Zero numbers moved.&lt;/p&gt;

&lt;p&gt;Nobody disputed a finding. All of the scrutiny landed on the instrument, and both times the instrument was wrong in a way I could not have seen by reading it.&lt;/p&gt;

&lt;p&gt;That is the argument for publishing the harness rather than the headline. A result is a claim people can take or leave. An instrument is something they can attack, and the attacks are what tell you whether it works. I got more out of two comment threads than out of forty hours of my own review.&lt;/p&gt;

&lt;p&gt;Two things worth stealing:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Assert the property, not the proxy.&lt;/strong&gt; Same cost. One fewer failure mode.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Build one check where a good number is the alarm.&lt;/strong&gt; Everywhere else, surprise is what triggers debugging, and a flattering bug never surprises you.&lt;/p&gt;

</description>
      <category>testing</category>
      <category>python</category>
      <category>ai</category>
      <category>programming</category>
    </item>
    <item>
      <title>Nine Bugs in My Own Evaluation Harness. Every One Made My Results Look Better.</title>
      <dc:creator>Marvin Okafor</dc:creator>
      <pubDate>Thu, 10 Sep 2026 21:24:21 +0000</pubDate>
      <link>https://dev.to/marvinoka4/nine-bugs-in-my-own-evaluation-harness-every-one-made-my-results-look-better-4fkj</link>
      <guid>https://dev.to/marvinoka4/nine-bugs-in-my-own-evaluation-harness-every-one-made-my-results-look-better-4fkj</guid>
      <description>&lt;p&gt;I built an evaluation harness. I found nine bugs in it. Every single one would have made my results look better than they were.&lt;/p&gt;

&lt;p&gt;Nine out of nine, all pointing the same way. That is not a coincidence, and I do not think it is specific to me or to my project. I think it is a structural problem with any evaluation you build for yourself.&lt;/p&gt;

&lt;p&gt;Here is the mechanism, before any of the evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Debugging is triggered by surprise
&lt;/h2&gt;

&lt;p&gt;You do not audit numbers. You audit numbers that bother you.&lt;/p&gt;

&lt;p&gt;When a result comes back disappointing, you go looking for the reason. You check the setup, you re-run it, you add logging, you find the bug. The bug gets fixed and the number moves.&lt;/p&gt;

&lt;p&gt;When a result comes back good, none of that fires. Nothing feels wrong. There is no surprise to investigate. You write it up.&lt;/p&gt;

&lt;p&gt;So the filter that removes measurement bugs from your work is applied unevenly. Hard against results you dislike. Softly against results you like. Every pass through that filter removes more unflattering bugs than flattering ones.&lt;/p&gt;

&lt;p&gt;Run that loop for a few days and your instrument has drifted in one direction, and nothing in your process is designed to notice. The bugs that survive to publication are disproportionately the ones that helped you. Not because anyone was dishonest. Because they never triggered the thing that catches bugs.&lt;/p&gt;

&lt;p&gt;I knew this argument in the abstract before I started. It did not stop me from writing nine of them.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the harness did
&lt;/h2&gt;

&lt;p&gt;Only enough context for the bugs to make sense.&lt;/p&gt;

&lt;p&gt;Mutation testing changes your source code in small ways. Flip a comparison. Change a constant. Delete a &lt;code&gt;raise&lt;/code&gt;. Then it runs your test suite and checks whether anything failed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;  def withdraw(balance, amount):
&lt;span class="gd"&gt;-     if amount &amp;lt;= 0:
&lt;/span&gt;&lt;span class="gi"&gt;+     if amount &amp;lt; 0:
&lt;/span&gt;          raise ValueError("amount must be positive")
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the suite stays green, that is a fault your tests cannot detect.&lt;/p&gt;

&lt;p&gt;Coverage tells you a line ran. This tells you whether anything would have complained if the line were wrong. Those are very different questions. A toy module with one happy-path test sits at 47% line coverage and a 9.5% mutation kill score.&lt;/p&gt;

&lt;p&gt;The harness generated mutations, ran suites against them, and scored how many got caught. Roughly 40 hours of work.&lt;/p&gt;

&lt;h2&gt;
  
  
  The nine bugs, and which way each one pushed
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Editable installs made mutations invisible.&lt;/strong&gt; &lt;code&gt;pip install -e&lt;/code&gt; on src-layout packages resolved imports back to the original checkout, so mutations written to a temp copy never executed. Three targets scored 0.000. &lt;em&gt;Direction: reads as "these test suites are terrible" rather than "my harness is broken." It made the problem I was solving look bigger.&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Parallel execution corrupted one target.&lt;/strong&gt; Running mutants concurrently produced three different results across four runs on the one target doing real async I/O. A recorded improvement of 0.27 to 0.77 was noise. Spurious failures get counted as the mutation being detected, and detection was the number I was maximising. &lt;em&gt;Direction: inflated the result.&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;A file picker chose an unrelated test file&lt;/strong&gt; for the hardest target, feeding the model irrelevant context exactly where context mattered most. &lt;em&gt;Direction: made a baseline look worse than it was.&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;A classifier categorised batches instead of individual tests.&lt;/strong&gt; One good test in a batch of 69 would have marked all 69 as good. &lt;em&gt;Direction: inflated quality.&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;A reconstruction step dropped shared imports&lt;/strong&gt; and manufactured test failures that were not real. &lt;em&gt;Direction: understated a baseline's capability.&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;An extractor only scanned top-level functions&lt;/strong&gt;, so a valid &lt;code&gt;unittest.TestCase&lt;/code&gt; response was discarded as "no test found." The retry loop then received a harness error instead of real pytest output, which disabled the exact mechanism I was measuring. &lt;em&gt;Direction: understated the agent.&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;&lt;code&gt;self.assertEqual(...)&lt;/code&gt; was classified as "no assertion."&lt;/strong&gt; I was testing a hypothesis about models writing assertion-free tests at the time. &lt;em&gt;Direction: would have manufactured my own hypothesis and handed it back to me.&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;A pre-registered metric was not computable&lt;/strong&gt; on dunder-dispatched code. It read as a real near-zero rate instead of as undefined. &lt;em&gt;Direction: false signal.&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;A documentation figure that had already survived two audits.&lt;/strong&gt; One arm was recorded as having zero clean-pass failures when it had three. &lt;em&gt;Direction: flattered a comparison.&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Note that they do not all inflate the headline number. Three of them understate a baseline or an arm. That still counts as favourable, because a worse baseline makes the thing I built look better by comparison. "Favourable" means favourable to the story, not favourable to one metric.&lt;/p&gt;

&lt;p&gt;Bug 9 is the one I find hardest to be relaxed about. It had been looked at twice. Two audits, both of which read past it, because the figure was consistent with what we expected to see and nothing about it invited a third look.&lt;/p&gt;

&lt;h2&gt;
  
  
  None of them was found by reading code
&lt;/h2&gt;

&lt;p&gt;This is the part I would most want someone to take away.&lt;/p&gt;

&lt;p&gt;Not one of those nine was caught by re-reading the function. I had already read the functions. Reading code that you wrote, looking for a bug you do not yet believe exists, is close to useless.&lt;/p&gt;

&lt;p&gt;Every single one was caught the same way: by running a check whose outcome I had predicted in advance, and getting a different answer.&lt;/p&gt;

&lt;p&gt;The clearest case was bug 5. I knew, independently and from earlier output, that one specific generated test was broken. So the prediction was simple. Remove that one test, and the suite goes green.&lt;/p&gt;

&lt;p&gt;I removed it. The suite did not go green.&lt;/p&gt;

&lt;p&gt;That contradiction is the only reason I found the dropped-imports bug before its numbers went into anything. There was no other signal. The scores it produced looked entirely plausible. They were plausible in a direction I liked, which is why nothing else would have prompted me to look.&lt;/p&gt;

&lt;p&gt;The prediction is what does the work. A check you run without a prediction just produces another number, and you will interpret that number the same way you interpret all the others.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two checks worth stealing
&lt;/h2&gt;

&lt;p&gt;Both are cheap. Both caught things.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The canary.&lt;/strong&gt; Overwrite the file under test with unparseable garbage and assert that the suite fails. This proves your mutation actually reaches the interpreter. If the suite passes while the file is syntactically invalid, you are not testing what you think you are testing. This is what would have caught bug 1 on day one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The determinism gate.&lt;/strong&gt; Run the same scoring three times, serially, and require byte-identical output. This proves execution is isolated. This is what catches bug 2.&lt;/p&gt;

&lt;p&gt;The important part is that these prove different properties, and neither substitutes for the other.&lt;/p&gt;

&lt;p&gt;The canary passes happily while concurrency silently corrupts your results. Your mutation reached the interpreter, so the canary is satisfied, and the numbers are still garbage.&lt;/p&gt;

&lt;p&gt;The determinism gate passes happily while imports resolve to the wrong file. Three identical runs of the wrong thing are still perfectly deterministic. The gate is satisfied and the score is meaningless.&lt;/p&gt;

&lt;p&gt;You need both, and you need to write down what each one actually proves, so you do not talk yourself into believing one covers the other.&lt;/p&gt;

&lt;h2&gt;
  
  
  The check that caught one of my own targets
&lt;/h2&gt;

&lt;p&gt;Three hours before my deadline I did a clean-clone reproduction run to verify the reproducibility claim.&lt;/p&gt;

&lt;p&gt;The determinism gate quarantined one of the project's own targets. Eleven of twelve reproduced exactly. The twelfth varied.&lt;/p&gt;

&lt;p&gt;I reported it in the README instead of fixing it.&lt;/p&gt;

&lt;p&gt;A check that has never caught anything is indistinguishable from a check that cannot catch anything. Mine had just caught something, and the something was mine. Removing that from the record would have made the project look better and the instrument look worse, which is exactly the trade this entire post is about.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to actually do
&lt;/h2&gt;

&lt;p&gt;Before you measure anything you built:&lt;/p&gt;

&lt;p&gt;Write down what your instrument would look like if it were lying to you. Then build the check that catches specifically that, and run it before you have any results you are attached to.&lt;/p&gt;

&lt;p&gt;Then write down which direction each possible lie would push your result.&lt;/p&gt;

&lt;p&gt;That second list is the one that matters, because it is a list of the checks you will be least motivated to run. Every item on it is a place where a bug will feel like a finding. You will not notice those on your own. Nobody does. That is the whole reason the nine came out nine for nine.&lt;/p&gt;

</description>
      <category>testing</category>
      <category>ai</category>
      <category>programming</category>
      <category>python</category>
    </item>
    <item>
      <title>I Grade AI Agent Code for a Living. Here's the 12-Point Checklist I Run Before Trusting Any of It.</title>
      <dc:creator>Marvin Okafor</dc:creator>
      <pubDate>Mon, 07 Sep 2026 23:17:07 +0000</pubDate>
      <link>https://dev.to/marvinoka4/i-grade-ai-agent-code-for-a-living-heres-the-12-point-checklist-i-run-before-trusting-any-of-it-5adp</link>
      <guid>https://dev.to/marvinoka4/i-grade-ai-agent-code-for-a-living-heres-the-12-point-checklist-i-run-before-trusting-any-of-it-5adp</guid>
      <description>&lt;p&gt;Your agent produced 400 lines in nine seconds. It compiles. The tests pass. The demo works.&lt;/p&gt;

&lt;p&gt;None of that tells you whether it's correct.&lt;/p&gt;

&lt;p&gt;I evaluate agentic AI coding output against structured rubrics professionally - correctness, instruction adherence, edge-case handling, the whole grid. Separately, I've spent seven years shipping production systems: high-traffic e-commerce, multi-tenant healthcare data, a national eLearning platform handling tens of thousands of concurrent applications. The checklist below is what happens when those two things collide. It's the pass I run before I'll put my name on anything an agent wrote.&lt;/p&gt;

&lt;p&gt;None of it is exotic. All of it is stuff agents get wrong constantly, and reviewers skip because the code &lt;em&gt;looks&lt;/em&gt; fine.&lt;/p&gt;




&lt;h3&gt;
  
  
  1. Does deleting the error handling break any test?
&lt;/h3&gt;

&lt;p&gt;If not, your error paths are untested decoration. This is the single fastest way to find out whether a test suite is real. Comment out a &lt;code&gt;catch&lt;/code&gt; block and run the suite. Green? You have no coverage of the thing most likely to hurt you.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. What happens on the second call?
&lt;/h3&gt;

&lt;p&gt;Agents write beautiful single-execution logic. Ask: if this runs twice - retry, duplicate message, user double-click - does it produce one side effect or two? Idempotency is seldom in the generated code unless you asked for it explicitly, and it's almost always required in production.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Is every dependency pinned, and does the lockfile exist?
&lt;/h3&gt;

&lt;p&gt;"Latest" is not a version. I've root-caused a bug where a &lt;em&gt;minor&lt;/em&gt; version bump silently collapsed TypeScript types to &lt;code&gt;never&lt;/code&gt; across a monorepo - semver protects runtime behaviour; it promises nothing about type inference or subtle behavioural edges. Agents love unpinned ranges. Pin them.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Is it mocking the thing it's supposed to be testing?
&lt;/h3&gt;

&lt;p&gt;The most common fake-coverage pattern I see: mock the database, assert the mock was called correctly, declare the data layer tested. If the correctness property is enforced by the database - constraints, row-level security policies, transaction isolation - a mock verifies your &lt;em&gt;assumption&lt;/em&gt; about the policy, not the policy. Run against a real, disposable instance. I maintain a security suite that validates RLS policies against actual Postgres for exactly this reason.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Does it check what should be &lt;em&gt;invisible&lt;/em&gt;, not just what's visible?
&lt;/h3&gt;

&lt;p&gt;Access-control tests written by agents almost always assert "the authorised user can see their data." The security-relevant half is "the unauthorised user &lt;em&gt;cannot&lt;/em&gt;." Test negative cases explicitly - and remember that some systems (RLS being the classic) filter silently rather than erroring, so "no exception raised" is not the same as "correctly denied." Assert the actual result set.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. How many database round-trips does the happy path make?
&lt;/h3&gt;

&lt;p&gt;ORMs plus agents produce N+1 queries at an impressive rate, because each line looks perfectly reasonable. Log the queries for one request. Count them. The number is usually higher than anyone guessed, and it's where a large share of real latency wins hide.&lt;/p&gt;

&lt;h3&gt;
  
  
  7. What are the timeout and retry policies - explicitly?
&lt;/h3&gt;

&lt;p&gt;Not "does it retry," but: how many times, with what backoff, and what happens when retries are exhausted? Unbounded retries against a struggling downstream service is how a partial outage becomes a full one. Agents default to either no retries or naive infinite ones.&lt;/p&gt;

&lt;h3&gt;
  
  
  8. Does it handle partial failure, or only total failure?
&lt;/h3&gt;

&lt;p&gt;Total failure is easy - the call throws, you catch it. Partial failure is the hard case: three of five writes succeeded, the response timed out but the operation actually completed, the queue delivered twice. Generated code is overwhelmingly written as though operations either fully succeed or fully fail.&lt;/p&gt;

&lt;h3&gt;
  
  
  9. Are the IAM permissions scoped, or is it wildcards?
&lt;/h3&gt;

&lt;p&gt;Agents reach for permissive policies because permissive policies make the demo work. Any wildcard in a generated permission set is a finding, not a default. Same for over-broad database roles - a correct policy can still leak data if the connecting role has privileges that sidestep it.&lt;/p&gt;

&lt;h3&gt;
  
  
  10. Is there anything in here that only works because of a race the tests never trigger?
&lt;/h3&gt;

&lt;p&gt;Concurrency is where "semantically wrong about failure" gets most expensive. Check-then-act patterns, non-atomic read-modify-write, missing transaction boundaries. Tests run sequentially; production doesn't.&lt;/p&gt;

&lt;h3&gt;
  
  
  11. Can someone else reproduce this from a clean clone?
&lt;/h3&gt;

&lt;p&gt;One command, fresh machine, same result. If setup requires tribal knowledge or an undocumented sequence of steps, your correctness is unverifiable by anyone but you - which, for anything that will outlive your attention span, means it's unverified.&lt;/p&gt;

&lt;h3&gt;
  
  
  12. Is the reasoning written down anywhere?
&lt;/h3&gt;

&lt;p&gt;Which ambiguity did you resolve, and how? What did you deliberately not handle? Agents produce code without provenance - no record of what was considered and rejected. That gap is a real maintenance liability, because the next person can't distinguish a deliberate decision from an accident. Write down the &lt;em&gt;why&lt;/em&gt;, especially for the non-obvious calls.&lt;/p&gt;




&lt;h2&gt;
  
  
  The pattern underneath all twelve
&lt;/h2&gt;

&lt;p&gt;Every item is a variant of the same thing: &lt;strong&gt;agents are excellent at code and unreliable about consequences.&lt;/strong&gt; They handle the path you described and quietly assume the paths you didn't. The failure mode people complain loudest about - hallucinated APIs - is the easy one, because it's loud and any test catches it. The dangerous one is syntactically perfect code that's confidently wrong about what happens when something breaks.&lt;/p&gt;

&lt;p&gt;Which means the review skill that matters now isn't "can you spot bad code." It's "can you enumerate the failure modes nobody wrote down." That's not a new skill. It's the thing senior engineers have always done. It just got a lot more load-bearing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'm building from this
&lt;/h2&gt;

&lt;p&gt;I'm turning this checklist into something executable - a fault-injection harness that stress-tests agent-generated infrastructure code against realistic failure conditions (retries, partial outages, IAM misconfigurations, concurrent access) with deterministic pass/fail checks instead of eyeballing. Chaos engineering, pointed at AI output. It'll go up on my GitHub and portfolio as I build it in the open, along with what breaks and why.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What's on your list that isn't on mine?&lt;/strong&gt; I'm collecting failure patterns for the harness's scenario set, and the ones that come from people who've been burned in production are worth more than anything I can invent. Drop them in the comments.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Part of an ongoing series on production debugging, performance engineering, and evaluation infrastructure for AI systems.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>codereview</category>
      <category>devops</category>
      <category>webdev</category>
    </item>
    <item>
      <title>I Built an Agent to Fix Bad Tests. I Found Eight Bugs in My Own Ruler.</title>
      <dc:creator>Marvin Okafor</dc:creator>
      <pubDate>Tue, 01 Sep 2026 18:01:34 +0000</pubDate>
      <link>https://dev.to/marvinoka4/i-built-an-agent-to-fix-bad-tests-i-found-eight-bugs-in-my-own-ruler-1eap</link>
      <guid>https://dev.to/marvinoka4/i-built-an-agent-to-fix-bad-tests-i-found-eight-bugs-in-my-own-ruler-1eap</guid>
      <description>&lt;p&gt;Here is a Python function and a test for it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;withdraw&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;balance&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;amount&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;amount must be positive&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;amount&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;balance&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;insufficient funds&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;balance&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;amount&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_withdraw&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="nf"&gt;withdraw&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;70&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That test gives you 47% line coverage. It gives you a 9.5% mutation kill score.&lt;/p&gt;

&lt;p&gt;The harness generates 21 small breakages of that module. The suite notices 2 of them.&lt;/p&gt;

&lt;p&gt;That gap is the whole reason I started this project. Line coverage is the default test-quality signal in most of the industry, and it measures whether a line ran. It does not measure whether anything would have complained if the line were wrong. AI-generated tests are unusually good at producing that shape: high coverage, low detection.&lt;/p&gt;

&lt;p&gt;Mutation testing measures the real thing. You break the code in small ways and check whether the tests notice.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;  def withdraw(balance, amount):
&lt;span class="gd"&gt;-     if amount &amp;lt;= 0:
&lt;/span&gt;&lt;span class="gi"&gt;+     if amount &amp;lt; 0:
&lt;/span&gt;          raise ValueError("amount must be positive")
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If no test fails, that is a bug your suite cannot detect. It has been sitting there the whole time.&lt;/p&gt;

&lt;p&gt;Mutation testing never went mainstream, and I think the reason is simple. It hands you a wall of surviving mutants and no path to fixing any of them. It tells you that you have a problem and then leaves.&lt;/p&gt;

&lt;p&gt;So the idea was an agent that closes the loop. Find the survivors. Write tests that kill them. Gate each generated test on a hard criterion: keep it only if it passes on clean code and fails on the mutant. Ground truth is a subprocess exit code. No model judges any outcome.&lt;/p&gt;

&lt;p&gt;I built it over about 30 hours for the micro1 Frontier Engineering Challenge, which had around 7,800 registrants.&lt;/p&gt;

&lt;p&gt;The tool works, sort of. It is incomplete and I will get to the numbers. But that is not the interesting part of the weekend.&lt;/p&gt;

&lt;p&gt;The interesting part is that my measuring instrument kept lying to me, and it lied in a consistent direction.&lt;/p&gt;

&lt;h2&gt;
  
  
  The finding that broke my own premise
&lt;/h2&gt;

&lt;p&gt;Before I ran a single agent call, I ran the harness across 12 widely-used, well-maintained Python libraries: cachetools, validators, natsort, dictdiffer, toolz, voluptuous, python-slugify, python-dotenv, shortuuid, boltons, aiofiles, tenacity.&lt;/p&gt;

&lt;p&gt;455 mutants generated. 133 survived the existing test suites.&lt;/p&gt;

&lt;p&gt;Then I checked something I had assumed I would not need to check. Of those 133 survivors, how many sit on a line the tests actually execute?&lt;/p&gt;

&lt;p&gt;53.&lt;/p&gt;

&lt;p&gt;The rest were never run at all. Not weakly tested. Not vacuously tested. Untested.&lt;/p&gt;

&lt;p&gt;I suspected my test commands were scoped too narrowly, so I widened them per target, between 6 and 40 times more test code. If the "executes but does not assert" category was real and I was just missing it, that number should climb.&lt;/p&gt;

&lt;p&gt;It went from 54 to 53. Down.&lt;/p&gt;

&lt;p&gt;And where widening changed anything, it converted unreachable mutations directly into kills. It did not move them into the middle category. It skipped it.&lt;/p&gt;

&lt;p&gt;So in mature, human-written Python, the vacuous test failure mode is rare. Where these suites fail, they fail by not running the code at all.&lt;/p&gt;

&lt;p&gt;The story I had absorbed about tests that execute everything and assert nothing is a story about AI-generated tests. It is not a story about human ones. I had to reframe the project before I had built the main part of it.&lt;/p&gt;

&lt;p&gt;That was the first sign that what I was actually building was a measuring instrument, and that I had not been treating it like one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Eight bugs in the instrument
&lt;/h2&gt;

&lt;p&gt;Every one of these would have produced a confident, wrong, publishable number.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Editable installs made mutations invisible.&lt;/strong&gt; &lt;code&gt;pip install -e&lt;/code&gt; on src-layout packages resolves imports back to the original checkout. My mutations were written to a temp copy, so they never executed. Three targets silently scored 0.000. That would have read as "the agent fails on src-layout packages," which is a finding. It is just not a true one.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Concurrency corrupted one target.&lt;/strong&gt; Running mutants in parallel gave me three different survivor sets across four runs, on the one target doing real async I/O. I had already drafted a result of "0.27 to 0.77" off that. It was noise. Note which way it pointed: spurious failures get counted as kills, and kills are the number every arm is trying to increase.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;A file picker chose the wrong test file.&lt;/strong&gt; On the hardest target. Which means the model would have been shown irrelevant context in exactly the place where context mattered most.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;A classifier was about to run on the wrong unit.&lt;/strong&gt; It classified batches, not individual tests. One strong test in a batch of 69 would have marked all 69 as strong.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;A reconstruction step dropped shared imports.&lt;/strong&gt; This manufactured test failures that were not real failures.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;An extractor only scanned top-level functions.&lt;/strong&gt; So a perfectly valid &lt;code&gt;unittest.TestCase&lt;/code&gt; response got discarded as "no test found." Worse, the agent's retry loop then received a harness error instead of real pytest output. That quietly disabled the exact mechanism I was trying to measure.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;&lt;code&gt;self.assertEqual(...)&lt;/code&gt; was classified as "no assertion."&lt;/strong&gt; This one would have manufactured precisely the finding I was hypothesising. It would have handed me my own conclusion.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;A pre-registered metric was not computable&lt;/strong&gt; on dunder-dispatched code like &lt;code&gt;__call__&lt;/code&gt; and &lt;code&gt;__or__&lt;/code&gt;. It read as a real near-zero rate rather than as undefined.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The two things they had in common
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;They all pointed the same way.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every single one of those eight would have made my result look better, cleaner, or more publishable. Not one of them would have made the agent look worse than it was.&lt;/p&gt;

&lt;p&gt;I do not think that is a coincidence, and I do not think it is a conspiracy either. It is attention. When a number disappoints you, you go looking for the reason. When a number pleases you, you write it up. So the measurement bugs that survive all the way to publication are disproportionately the ones that helped you.&lt;/p&gt;

&lt;p&gt;That is a selection effect operating on your own debugging, and you cannot fix it by being careful. Careful people are exactly as motivated to stop investigating when the number looks good.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;None of them was found by reading code.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every one was caught by running a check whose outcome I had predicted in advance, and getting the wrong answer.&lt;/p&gt;

&lt;p&gt;The clearest case was bug 5. My prediction was: remove this one known-bad test and the suite goes green. It did not go green. That contradiction is the only reason I found the dropped-imports bug before I trusted the numbers it was feeding me.&lt;/p&gt;

&lt;p&gt;I would not have found it by rereading the function. I had already read the function.&lt;/p&gt;

&lt;h2&gt;
  
  
  The results, with the caveats attached
&lt;/h2&gt;

&lt;p&gt;Three arms, same model, same token ceiling.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Arm&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;th&gt;Mutants killed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;A&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;One prompt: "write as many tests as warranted." The brief's specified baseline.&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;B&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;One test per call, same call count as the agent. No mutation hint, no gate, no retry.&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;C&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The agent: mutation diff in context, execution gate, one retry with real pytest output fed back.&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On the 15 mutants the agent covered, B killed 1 and C killed 9. Keep rate was 60%, so the gate was genuinely filtering rather than rubber-stamping. Nine retries fired and three succeeded.&lt;/p&gt;

&lt;p&gt;Now the parts that matter just as much.&lt;/p&gt;

&lt;p&gt;The agent ran on 2 of 10 targets. The API budget ran out mid-run. I did not swap in a substitute model to finish the sweep, because then the comparison would not be a comparison.&lt;/p&gt;

&lt;p&gt;Those two targets are the ones where the baseline performed worst. That is not a random sample, and I have no way to argue it is representative.&lt;/p&gt;

&lt;p&gt;All 9 kills were on the two cheapest mutation types. There was zero cross-function transfer. Seven of the nine kept tests kill exactly the one mutation they were written for, and nothing else. That is a real limitation, not a rounding error.&lt;/p&gt;

&lt;h2&gt;
  
  
  One result that inverted my hypothesis
&lt;/h2&gt;

&lt;p&gt;I had pre-registered a prediction that gate-passing tests would mostly be vacuous. Bare existence checks. Tests with no assertion that "kill" a mutant by crashing rather than by detecting anything.&lt;/p&gt;

&lt;p&gt;That is not what happened. The &lt;code&gt;none&lt;/code&gt; category was empty. Eight of the nine kills were real assertion failures.&lt;/p&gt;

&lt;p&gt;And of the six discarded drafts, zero failed on clean code. All six were valid, passing tests that simply did not detect the bug.&lt;/p&gt;

&lt;p&gt;So the gate was not catching broken tests. It was catching working tests that miss. That is a more interesting failure mode than the one I predicted, and I only know it because the prediction was written down first and was wrong in a specific way.&lt;/p&gt;

&lt;h2&gt;
  
  
  The determinism check that quarantined my own target
&lt;/h2&gt;

&lt;p&gt;Three hours before the deadline I did a clean-clone reproduction run to verify the reproducibility claim.&lt;/p&gt;

&lt;p&gt;The determinism check runs each target three times serially and requires the survivor sets to be byte-identical. Eleven of twelve reproduced exactly. The twelfth varied.&lt;/p&gt;

&lt;p&gt;I reported it in the README instead of fixing it. A check that has never caught anything is indistinguishable from a check that cannot catch anything, and the first thing mine ever caught was one of my own targets. Removing that from the record would have made the project look better and the instrument look worse.&lt;/p&gt;

&lt;p&gt;I missed the submission by 11 minutes.&lt;/p&gt;

&lt;p&gt;That is annoying in a way I do not want to dress up. But the reproduction run is what found the twelfth target, and running it is the reason I trust the other eleven.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would tell you to take from this
&lt;/h2&gt;

&lt;p&gt;Before you measure an agent, write down what your instrument would look like if it were lying to you. Then build the check that catches exactly that.&lt;/p&gt;

&lt;p&gt;And write down which direction each possible lie would push your result. That second list is the important one, because it tells you which checks you will be least motivated to run.&lt;/p&gt;

&lt;p&gt;The repo is private while I finish the write-up. If you build evaluations for agents and you have hit this, I would like to compare notes, particularly on catching measurement bias before it reaches a number you have already started believing.&lt;/p&gt;

</description>
      <category>testing</category>
      <category>ai</category>
      <category>python</category>
      <category>showdev</category>
    </item>
    <item>
      <title>5,900 Engineers Just Registered for a Hackathon Where Using AI Is the Point. Here's How It Will Actually Be Judged.</title>
      <dc:creator>Marvin Okafor</dc:creator>
      <pubDate>Fri, 28 Aug 2026 11:16:11 +0000</pubDate>
      <link>https://dev.to/marvinoka4/5900-engineers-just-registered-for-a-hackathon-where-using-ai-is-the-point-heres-how-it-will-1bdd</link>
      <guid>https://dev.to/marvinoka4/5900-engineers-just-registered-for-a-hackathon-where-using-ai-is-the-point-heres-how-it-will-1bdd</guid>
      <description>&lt;p&gt;Most hackathons treat coding agents as something between a crutch and cheating. micro1's Frontier Engineering Challenge, which kicks off today, inverts that entirely: you're &lt;em&gt;expected&lt;/em&gt; to use coding agents. The competition isn't about whether you can generate code. It's about whether you can generate code that survives scrutiny.&lt;/p&gt;

&lt;p&gt;That's a fundamentally different game, and judging by how most engineers approach AI-assisted work, a lot of the field is about to optimise for the wrong thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The details
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;What:&lt;/strong&gt; micro1 Frontier Engineering Challenge 2026 - &lt;cite&gt;a free, global, online competition, run as a three-day sprint where you use coding agents to tackle a real-world software engineering problem&lt;/cite&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;When:&lt;/strong&gt; &lt;cite&gt;August 28–31, 2026. The full challenge statement is released at kickoff - August 28 at 15:00 UTC&lt;/cite&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Format:&lt;/strong&gt; Online, individual (team size 1), free&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Registrations so far:&lt;/strong&gt; ~5,900&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Why it matters beyond the prize:&lt;/strong&gt; micro1 has indicated that top-performing participants get considered for paid opportunities with them&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The problem statement is deliberately withheld until kickoff, so &lt;cite&gt;nobody gets to pre-build&lt;/cite&gt;. Everyone starts cold.&lt;/p&gt;

&lt;h2&gt;
  
  
  The line in the announcement that tells you everything
&lt;/h2&gt;

&lt;p&gt;Buried in the challenge description is the sentence that should reframe your entire strategy: &lt;cite&gt;AI can produce convincing code in seconds - real engineering begins when convincing is not enough: incomplete requirements, hidden dependencies, difficult edge cases, failure modes, and decisions that require technical judgment.&lt;/cite&gt;&lt;/p&gt;

&lt;p&gt;And the deliverable: &lt;cite&gt;a solution that is correct, reproducible, testable and clearly explained.&lt;/cite&gt;&lt;/p&gt;

&lt;p&gt;Read those four words again - &lt;strong&gt;correct, reproducible, testable, explained&lt;/strong&gt;. Not "impressive." Not "feature-complete." Not "shipped fastest." If you've ever built or graded an evaluation rubric, you recognise immediately what that list is: it's a rubric where &lt;em&gt;three of the four criteria have nothing to do with whether your code runs&lt;/em&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I think most of the field will lose points
&lt;/h2&gt;

&lt;p&gt;I evaluate agentic AI coding outputs against structured rubrics professionally - rubric construction, adversarial prompt design, deciding which checks can be programmatic and which need human judgment. So I'm reading this challenge less as "what should I build" and more as "where does this rubric bite." My honest read:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Reproducibility is the silent killer.&lt;/strong&gt; "It works on my machine after four hours of undocumented agent conversation" is not reproducible. If a judge can't clone your repo and get the same result, correctness is unverifiable - and unverifiable correctness scores as zero, not as partial credit. Pin your dependencies. Commit your lockfile. Containerize. Make the setup a single command.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. "Testable" doesn't mean "has tests."&lt;/strong&gt; It means the tests actually prove the thing that matters. A test suite that only covers the happy path proves your agent could write a happy path. The edge cases and failure modes the challenge explicitly calls out - those need tests that would &lt;em&gt;fail&lt;/em&gt; if the behaviour were wrong. If deleting your error-handling doesn't break any test, you don't have error-handling coverage; you have decoration.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Agents are confidently wrong about failure, not about syntax.&lt;/strong&gt; This is the pattern I keep seeing in evaluation work: the generated code is syntactically perfect and semantically wrong about what happens when things break. Retries, timeouts, partial writes, duplicate messages, concurrent access. Given the challenge explicitly names failure modes as part of the frontier, this is almost certainly where the separation happens.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. "Clearly explained" is a scored deliverable, not a README afterthought.&lt;/strong&gt; When ambiguity is baked into the problem statement - and the announcement promises incomplete requirements - the judges cannot read your mind about which interpretation you chose. Your write-up needs to name the ambiguity, state the interpretation you picked, and justify it. An engineer who documents "the spec didn't define whether delivery is at-least-once or exactly-once; I assumed at-least-once and made the consumer idempotent, here's why" is demonstrating exactly the technical judgment being tested. An engineer who silently picks one and says nothing looks identical to an engineer who never noticed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Time will go where you don't expect.&lt;/strong&gt; With agents, generating a working first draft is the fast part. Verification, reproducibility, and documentation are where three days actually go. Budget accordingly - a working solution with no test suite and no write-up will lose to a slightly narrower solution that's fully verified and clearly reasoned.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I'd structure the three days
&lt;/h2&gt;

&lt;p&gt;Rough plan, adjust to the actual problem:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Day 1 - Specify before you generate.&lt;/strong&gt; Read the problem twice. Write down every ambiguous term &lt;em&gt;before&lt;/em&gt; you touch an agent, because every ambiguity is a decision you'll otherwise make accidentally. Decide what "correct" means, concretely, in a form you could test. Set up the reproducible environment (container, lockfile, one-command setup) first, not last.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Day 2 - Generate, then attack your own output.&lt;/strong&gt; Use agents aggressively for the implementation. Then switch hats: try to break it. Inject the failure modes. Write tests that would fail if the behaviour were wrong. Every defect you find yourself is one a judge doesn't find for you.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Day 3 - Write the explanation as if the reader is sceptical.&lt;/strong&gt; Architecture, edge cases considered and rejected, ambiguities and your resolutions, known limitations. Stating a limitation honestly reads as judgment. Hiding one reads as an oversight when someone finds it - and someone will.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The thing I'd most want to internalise: &lt;strong&gt;stating what you deliberately did not do, and why, is a strength.&lt;/strong&gt; Rubrics reward demonstrated judgment. Scope honesty is judgment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this format is a signal about where engineering is going
&lt;/h2&gt;

&lt;p&gt;There's a real argument that this is what technical hiring looks like in a couple of years. Not "can you invert a binary tree without autocomplete," but "given agents that generate plausible code instantly, can you specify, verify, and defend a solution?" That's a senior-engineering skill set, and it's notably &lt;em&gt;not&lt;/em&gt; the skill set that LeetCode grinding builds.&lt;/p&gt;

&lt;p&gt;Registration and the full brief are on HackerEarth - the problem statement drops at 15:00 UTC today. I'm going in, and I'll write up what I learn regardless of how I place, including anything I get wrong in the read above.&lt;/p&gt;

&lt;p&gt;If you're competing too, say so in the comments - I'd like to compare approaches afterwards, especially on how people handled the ambiguity-documentation piece.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I write about production debugging, performance work, and building evaluation environments for AI systems. Previous posts in this series cover deterministic RL environments for cloud infrastructure and what months of grading agentic code taught me about where models actually fail.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>hackathon</category>
      <category>ai</category>
      <category>devops</category>
      <category>career</category>
    </item>
    <item>
      <title>I've Spent Months Grading AI Agents' Code for a Living. Here's the Pattern Nobody's Talking About</title>
      <dc:creator>Marvin Okafor</dc:creator>
      <pubDate>Thu, 06 Aug 2026 18:20:52 +0000</pubDate>
      <link>https://dev.to/marvinoka4/ive-spent-months-grading-ai-agents-code-for-a-living-heres-the-pattern-nobodys-talking-about-34bf</link>
      <guid>https://dev.to/marvinoka4/ive-spent-months-grading-ai-agents-code-for-a-living-heres-the-pattern-nobodys-talking-about-34bf</guid>
      <description>&lt;p&gt;Everyone's talking about agentic AI shipping production code. Nobody's talking about what happens when you actually sit down and grade thousands of lines of it against a rubric, line by line, for months.&lt;/p&gt;

&lt;p&gt;I have. And the failure pattern that shows up over and over isn't the one Twitter/X is arguing about.&lt;/p&gt;

&lt;h2&gt;
  
  
  The job title that didn't exist two years ago
&lt;/h2&gt;

&lt;p&gt;"AI evaluator." "AI trainer." "Expert contributor to frontier model training data." None of these existed as job titles when I started my career. Now they're where a chunk of the most interesting engineering signal in the industry is actually happening — quietly, behind NDAs, far from the demo videos.&lt;/p&gt;

&lt;p&gt;Here's what the job actually is: agentic coding outputs land on your desk, and you grade them against a structured rubric — correctness, instruction adherence, quality, edge-case handling. You design adversarial prompts to find where the model's reasoning breaks. You decide which checks can be programmatic and deterministic, and which genuinely need a human who's shipped production systems to make the call. This is &lt;strong&gt;RL environment design&lt;/strong&gt; and &lt;strong&gt;LLMOps&lt;/strong&gt; in its rawest form, and it's a completely different skill from "prompt engineer" or "ML researcher." It's closer to being a QA lead for a junior engineer who never sleeps, never gets embarrassed, and will confidently ship the wrong answer with perfect syntax.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern: agents are great at code, bad at consequences
&lt;/h2&gt;

&lt;p&gt;Here's the uncomfortable part. The failure mode people are loudest about — hallucinated APIs, made-up library functions — is the &lt;em&gt;easy&lt;/em&gt; failure mode. It's loud, it's obvious, and any decent test suite catches it in seconds.&lt;/p&gt;

&lt;p&gt;The failure mode that actually matters, the one that slips past a surface read and even past a naive test suite, looks like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The code is syntactically perfect and semantically wrong about failure.&lt;/strong&gt; It handles the happy path beautifully and quietly assumes the retry, the timeout, the partial write, the duplicate message never happens.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It optimises for the metric, not the intent&lt;/strong&gt; — the agentic-AI version of Goodhart's Law. Give a model a rubric that checks "does the deploy succeed," and you'll occasionally get a solution that technically satisfies the check while doing something no engineer would sign off on. Evaluators call this &lt;strong&gt;reward hacking&lt;/strong&gt;, and it's a far more common failure than outright hallucination once you're grading real-world infra tasks instead of leetcode.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It's confidently wrong about IAM, concurrency, and distributed state&lt;/strong&gt; — exactly the areas where production engineering experience matters most and where a rubric written by someone who's never operated a real system will miss the defect entirely.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this is a knock on the models. It's a knock on how we evaluate them. You cannot catch consequence-blindness with a rubric written by someone who has never had a database silently corrupt state under a race condition at 2am. This is the actual bottleneck in scaling agentic AI into production-grade infrastructure work: not model capability, &lt;strong&gt;evaluation quality&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why "vibe coding" breaks down at the infra layer
&lt;/h2&gt;

&lt;p&gt;"Vibe coding" — accepting AI-generated code because it &lt;em&gt;looks&lt;/em&gt; right and the demo works — is fine for a prototype. It is not fine for anything touching IAM policies, message queues, durable storage, or disaster recovery. The gap between "looks right" and "is right" is exactly the gap that &lt;strong&gt;golden reference solutions&lt;/strong&gt; and &lt;strong&gt;deterministic validation tests&lt;/strong&gt; exist to close — the same discipline I wrote about in my last post on building RL environments for cloud infrastructure evaluation.&lt;/p&gt;

&lt;p&gt;The uncomfortable truth for the "AI writes all our code now" crowd: the more production-grade the system, the more the bottleneck shifts from &lt;em&gt;generating&lt;/em&gt; the code to &lt;em&gt;specifying and verifying&lt;/em&gt; it. That's a systems-engineering problem, not a model-scaling problem. It's also, not coincidentally, exactly what senior backend engineers have spent their careers getting good at — writing test suites against real databases instead of mocks, root-causing defects that hide three layers deep, documenting edge cases precisely enough that someone else can reproduce the reasoning. That skill set didn't get less valuable when agents showed up. It became the thing standing between "the demo worked" and "it survived contact with production."&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'm building next
&lt;/h2&gt;

&lt;p&gt;I'm turning this into a real project instead of just a hot take: a lightweight harness for stress-testing AI-agent-generated infrastructure code against realistic, injected failure conditions — retries, partial outages, IAM misconfigurations, the exact defect classes I've been describing above — with &lt;strong&gt;deterministic pass/fail checks&lt;/strong&gt; instead of vibes. Think chaos engineering meets AI evals: inject one fault at a time, assert invariants instead of traces, and see whether an agent's "working" solution is actually working or just golden-path lucky.&lt;/p&gt;

&lt;p&gt;It'll live on my portfolio and GitHub as I build it in the open — seed scenarios, the fault-injection harness, and a write-up of what breaks and why. If you're working on anything adjacent (RL environments, AI evals, chaos engineering, or you've just been burned by AI-generated infra code in production), I want to hear about it — drop it in the comments or find me on GitHub.&lt;/p&gt;

&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;Agentic AI isn't going to be stopped by a model that can't write a for-loop. It's going to be shaped by whether the industry gets serious about &lt;strong&gt;evaluation infrastructure&lt;/strong&gt; — golden solutions, deterministic tests, adversarial failure scenarios — as fast as it's getting excited about generation. That's the unglamorous, unsexy, extremely fundable problem hiding behind every "AI wrote our whole backend" headline.&lt;/p&gt;

&lt;p&gt;If you've seen this pattern too — agents that ace the demo and fail the disaster-recovery drill — I'd genuinely like to compare notes.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This is the fifth in a series on production engineering, debugging, and building evaluation environments for AI systems. Follow for the harness build-in-public over the coming weeks.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llmops</category>
      <category>agenticai</category>
      <category>careerdev</category>
    </item>
    <item>
      <title>Building Deterministic RL Environments for Cloud Infrastructure Evaluation: What Actually Transfers From Production Engineering</title>
      <dc:creator>Marvin Okafor</dc:creator>
      <pubDate>Thu, 06 Aug 2026 18:16:20 +0000</pubDate>
      <link>https://dev.to/marvinoka4/building-deterministic-rl-environments-for-cloud-infrastructure-evaluation-what-actually-transfers-52jm</link>
      <guid>https://dev.to/marvinoka4/building-deterministic-rl-environments-for-cloud-infrastructure-evaluation-what-actually-transfers-52jm</guid>
      <description>&lt;p&gt;AI models are increasingly being asked to design, deploy, secure, and recover production-grade infrastructure — not just write functions that pass a unit test. That shift changes what a training environment has to be. It's no longer enough to check that output "looks right." You need a &lt;strong&gt;golden reference solution&lt;/strong&gt;, a &lt;strong&gt;deterministic validation suite&lt;/strong&gt;, and a set of &lt;strong&gt;intentionally broken variants&lt;/strong&gt; that probe exactly where a model's reasoning breaks down under failure.&lt;/p&gt;

&lt;p&gt;I've spent the last several months on the other side of this problem — evaluating agentic coding outputs against structured rubrics, designing programmatic verification checks, and writing up edge cases that rubric authors hadn't considered. Before that, I spent seven years building and debugging the kind of systems these environments are meant to simulate: multi-tenant SaaS platforms with row-level security, AWS infrastructure serving high-traffic e-commerce, and data pipelines processing tens of thousands of concurrent requests. This post is about where those two things meet — what it actually takes to build an infrastructure RL environment that's reproducible, fair to evaluate, and hard to game.&lt;/p&gt;

&lt;h2&gt;
  
  
  A golden solution is only as good as its ambiguity budget
&lt;/h2&gt;

&lt;p&gt;The first mistake in building any evaluation environment is under-specifying the scenario and over-specifying the solution. If the task says "deploy a fault-tolerant queue consumer" without pinning down delivery semantics, retry policy, and what "fault-tolerant" means operationally, you'll end up with a golden solution that's just &lt;em&gt;one&lt;/em&gt; valid interpretation among several — and you'll penalize a model for a decision the spec never actually made.&lt;/p&gt;

&lt;p&gt;This is the same discipline as writing rubrics for coding evaluations: every ambiguous term is a future dispute. In practice that means, before writing a single line of infrastructure code, defining:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The exact failure modes in scope (node loss, network partition, message duplication, clock skew)&lt;/li&gt;
&lt;li&gt;The invariants that must hold regardless of implementation (at-least-once vs exactly-once delivery, idempotency guarantees, RTO/RPO targets)&lt;/li&gt;
&lt;li&gt;What counts as "recovered" — not just "the service is up," but that state is consistent&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Deterministic validation means designing against non-determinism
&lt;/h2&gt;

&lt;p&gt;Distributed systems are inherently non-deterministic — that's exactly what makes them hard to evaluate. A validation suite that works by re-running the same sequence of API calls and diffing output will produce flaky, unfair results the moment retries, timeouts, or eventual consistency are involved.&lt;/p&gt;

&lt;p&gt;The pattern that's worked for me, most recently writing a security test suite that validates row-level-security policies against a real Postgres instance rather than mocks, is to validate &lt;strong&gt;invariants&lt;/strong&gt;, not &lt;strong&gt;traces&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Anti-pattern: asserting on the exact sequence of events
&lt;/span&gt;&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;consumer_started&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;message_received&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;message_processed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ack_sent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="c1"&gt;# Better: assert the invariant the system must uphold,
# regardless of retries, ordering, or timing
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_no_duplicate_side_effects&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;env&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;inject_fault&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;redeliver_message&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;count&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run_until_settled&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;side_effect_count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;charge_customer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;final_state_is_consistent&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Testing against a real, disposable instance of the actual dependency — a real queue, a real Postgres, a real IAM policy engine — rather than a mock catches the class of bug that mocks are structurally blind to: the ones where the &lt;em&gt;contract&lt;/em&gt; you assumed doesn't match the &lt;em&gt;behaviour&lt;/em&gt; you get. I've root-caused production bugs (a minor dependency version bump silently collapsing TypeScript types to &lt;code&gt;never&lt;/code&gt;, an OAuth token-refresh edge case that only reproduced under real API rate limits) that a mocked test suite would have sailed straight past. The same principle applies at the infrastructure layer, just with higher stakes per failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Defective variants need to fail for the &lt;em&gt;right&lt;/em&gt; reason
&lt;/h2&gt;

&lt;p&gt;The point of an intentionally broken variant isn't just "does the model notice something is wrong" — it's "does the model correctly diagnose &lt;em&gt;why&lt;/em&gt;." A variant with a misconfigured IAM policy that happens to also have a network misconfiguration will teach a model to pattern-match on the wrong signal.&lt;/p&gt;

&lt;p&gt;Building these well means treating each defect as a single, isolated hypothesis:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;One fault per variant.&lt;/strong&gt; Resist the urge to combine a broken retry policy and an under-provisioned autoscaling group into one scenario "for efficiency." You'll never know which one the model actually reasoned about.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fail loud enough to be observable, quiet enough to require diagnosis.&lt;/strong&gt; A defect that immediately crashes the deploy is a smoke test, not an evaluation. A defect that silently corrupts data under specific timing conditions is where reasoning gets tested — but it has to be &lt;em&gt;reliably&lt;/em&gt; reproducible, or you're evaluating luck.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Document the intended diagnosis path.&lt;/strong&gt; If you can't write down, in advance, the sequence of observability signals (logs, metrics, traces) that should lead a correct reasoner to the root cause, the variant isn't ready yet.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where this comes from, concretely
&lt;/h2&gt;

&lt;p&gt;None of this is theoretical for me. A few data points from production work that map directly onto this kind of environment design:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Wrote and maintain a 14-test security suite validating row-level-security policies against a real Postgres database — the reproducible-environment-over-mocks principle, applied.&lt;/li&gt;
&lt;li&gt;Root-caused a silent dependency-resolution failure where a minor version bump collapsed TypeScript types to &lt;code&gt;never&lt;/code&gt; across a monorepo — the kind of defect class that's genuinely worth encoding as a training scenario, because it's realistic and painful to diagnose.&lt;/li&gt;
&lt;li&gt;Built and operated AWS infrastructure (EC2, S3, Lambda) for a high-traffic e-commerce platform, including the profiling and query-optimisation work behind a 30% latency reduction — direct experience with the observability and performance-diagnosis loop these environments are meant to simulate.&lt;/li&gt;
&lt;li&gt;Currently evaluate agentic coding outputs against structured rubrics professionally, including rubric construction, adversarial prompt design, and selecting which checks are actually programmatically verifiable versus which need human judgment — the exact skill set "golden reference solution + deterministic test" environment design draws on.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The uncomfortable part: most of the work is writing, not coding
&lt;/h2&gt;

&lt;p&gt;The infrastructure code for a good RL environment is often the easy part. The hard part is the documentation — writing down, precisely enough that someone else could reproduce your reasoning, what each scenario is testing, what "correct" means, what edge cases were considered and rejected, and why. That's not a side task. For environments meant to train reasoning about production systems, the documentation &lt;em&gt;is&lt;/em&gt; the specification the golden solution is graded against. Treating it as an afterthought is how you end up with an environment that's internally inconsistent and nobody notices until a model exploits the gap.&lt;/p&gt;

&lt;p&gt;If you're building or evaluating environments in this space, I'd be glad to compare notes — particularly on validating distributed invariants without flaking, and on keeping golden solutions honest about which parts of a spec are actually unambiguous.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>backend</category>
      <category>devops</category>
      <category>testing</category>
    </item>
    <item>
      <title>Postmortem: How a Minor Version Bump Silently Collapsed Our Types to `never`</title>
      <dc:creator>Marvin Okafor</dc:creator>
      <pubDate>Thu, 06 Aug 2026 18:15:45 +0000</pubDate>
      <link>https://dev.to/marvinoka4/postmortem-how-a-minor-version-bump-silently-collapsed-our-types-to-never-485e</link>
      <guid>https://dev.to/marvinoka4/postmortem-how-a-minor-version-bump-silently-collapsed-our-types-to-never-485e</guid>
      <description>&lt;p&gt;Some bugs announce themselves with a stack trace. Others just make your code quietly stop meaning anything. This is about the second kind — a dependency resolution issue in a TypeScript monorepo that didn't throw, didn't fail CI in an obvious way, and took real digging to root-cause, because the symptom looked nothing like the cause.&lt;/p&gt;

&lt;h2&gt;
  
  
  The symptom
&lt;/h2&gt;

&lt;p&gt;We're a Next.js + TypeScript monorepo using npm workspaces, on top of Supabase/Postgres with row-level-security policies. One afternoon, a teammate opened a PR, and the type checker flagged what looked like an unrelated function as having a parameter of type &lt;code&gt;never&lt;/code&gt;. Not &lt;code&gt;unknown&lt;/code&gt;, not &lt;code&gt;any&lt;/code&gt; — &lt;code&gt;never&lt;/code&gt;. That's the type TypeScript uses to say "this code path is unreachable" or "no value can satisfy this type."&lt;/p&gt;

&lt;p&gt;The function in question was very much reachable. It was called constantly, in production, successfully. TypeScript was simply wrong about it — or rather, TypeScript was being correctly told the wrong thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why &lt;code&gt;never&lt;/code&gt; is the worst possible symptom
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;never&lt;/code&gt; is a black hole in the type system. Once a type collapses to &lt;code&gt;never&lt;/code&gt;, everything downstream that touches it also tends to become uninferrable or nonsensical, because there's no value that can inhabit &lt;code&gt;never&lt;/code&gt;. This means the &lt;em&gt;first&lt;/em&gt; place you see an error is rarely anywhere near the actual cause — the type has already been wrong for several hops by the time the checker has no choice but to complain.&lt;/p&gt;

&lt;p&gt;That property makes this class of bug genuinely dangerous to debug by intuition. Staring at the flagged function tells you almost nothing, because the function is innocent. You have to work backwards through the type's provenance instead of forwards from the symptom.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tracing it back
&lt;/h2&gt;

&lt;p&gt;The investigation went roughly like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Confirm it's not a logic bug.&lt;/strong&gt; Runtime behaviour was correct. This ruled out anything in our own function bodies and pointed at the type layer specifically — either our type definitions or something upstream of them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bisect the type, not the code.&lt;/strong&gt; I traced the offending type backwards through each intermediate type alias and generic constraint until I found the layer where it stopped being sane. It turned out to originate from a shared type re-exported from one of our internal workspace packages.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check what changed.&lt;/strong&gt; &lt;code&gt;git log&lt;/code&gt; on the lockfile, not the source, is the move here — the source hadn't changed. A transitive dependency had bumped a &lt;em&gt;minor&lt;/em&gt; version, which under semver should have been safe.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reproduce in isolation.&lt;/strong&gt; I pinned every workspace package to the exact versions from before the bump, confirmed the type error disappeared, then bumped dependencies one at a time until it reappeared. This isolated the exact package and version.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Understand the actual mechanism.&lt;/strong&gt; The new minor version had changed an internal conditional type in a way that was backwards-compatible at the &lt;em&gt;value&lt;/em&gt; level but not at the &lt;em&gt;type inference&lt;/em&gt; level for one specific generic usage pattern we relied on. Semver protects you from breaking runtime behaviour; it says nothing about breaking type inference for edge-case generic usage. That's the real lesson here.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The fix, and the more important fix
&lt;/h2&gt;

&lt;p&gt;The immediate fix was a version pin with a comment explaining exactly why, linking to the isolated repro. That's necessary but not sufficient — a pin with no explanation just becomes a mystery for the next person, and eventually someone "cleans it up" and reintroduces the bug.&lt;/p&gt;

&lt;p&gt;The more important fix was writing it up: what the symptom looked like, why it was misleading, the bisection method that found it, and the underlying mechanism (type-level semver violations in transitive dependencies). That doc is what turns a three-hour debugging session into a five-minute fix the next time a teammate hits something that smells similar.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd generalise from this
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;When a type error implicates innocent code, suspect provenance, not logic.&lt;/strong&gt; &lt;code&gt;never&lt;/code&gt; and its cousins are usually downstream symptoms of an upstream problem.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bisect the dependency graph, not just your own commits.&lt;/strong&gt; Runtime-correct, type-broken changes are invisible to normal "did my code change" instincts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Semver is a promise about behaviour, not about type inference.&lt;/strong&gt; Especially in libraries that lean on conditional or mapped types, a "minor" bump can break inference for edge-case usages that the maintainers never tested against.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The write-up is part of the fix.&lt;/strong&gt; A pinned version with no documented reasoning is technical debt with a delay timer on it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Debugging this kind of thing is oddly satisfying once you see the shape of it — the bug isn't in the obvious place; it's in the second-order effects of a decision made three layers away. If you've hit something similar in a large TypeScript monorepo, I'd be curious to compare notes on tooling for catching type-level regressions from dependency bumps before they reach a PR.&lt;/p&gt;

</description>
      <category>typescript</category>
      <category>webdev</category>
      <category>debugging</category>
    </item>
    <item>
      <title>Cutting Latency by 30% on a High-Traffic E-Commerce Platform: A Profiling-First Approach</title>
      <dc:creator>Marvin Okafor</dc:creator>
      <pubDate>Thu, 06 Aug 2026 18:15:07 +0000</pubDate>
      <link>https://dev.to/marvinoka4/cutting-latency-by-30-on-a-high-traffic-e-commerce-platform-a-profiling-first-approach-e2b</link>
      <guid>https://dev.to/marvinoka4/cutting-latency-by-30-on-a-high-traffic-e-commerce-platform-a-profiling-first-approach-e2b</guid>
      <description>&lt;p&gt;"Make it faster" is one of the least actionable instructions in engineering, right up until you have data telling you exactly where the time is going. Over about two and a half years on a high-traffic e-commerce platform, we reduced platform latency by roughly 30% while maintaining 99.9% uptime during peak trading periods — including the highest-traffic days of the year, when the margin for error is smallest, and the cost of getting it wrong is highest. None of it came from a single dramatic rewrite. It came from a repeatable process, applied consistently.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start by refusing to guess
&lt;/h2&gt;

&lt;p&gt;The instinct when something feels slow is to optimise the thing you &lt;em&gt;assume&lt;/em&gt; is slow — usually whatever you last touched, or whatever looks inefficient on a quick read. That instinct is wrong often enough to be actively dangerous, because "optimising" code that wasn't the bottleneck adds complexity and risk for zero latency benefit, and it burns the time you should have spent on the actual bottleneck.&lt;/p&gt;

&lt;p&gt;The fix is boring and non-negotiable: profile first, in an environment that resembles production traffic patterns, before writing a single optimisation. Real bottlenecks are frequently unglamorous — an N+1 query hiding behind an ORM abstraction, a synchronous call to a downstream service that could have been parallelised, a cache that's technically present but missing on the hot path that actually matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the 30% actually came from
&lt;/h2&gt;

&lt;p&gt;Roughly, the wins broke down into three categories:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Query optimisation.&lt;/strong&gt; The single highest-leverage category. Several code paths were making sequential database round-trips with no data dependencies between them — classic N+1 patterns introduced incrementally as features were added over time, none individually alarming, collectively expensive. Batching these, and in a few cases denormalising specific hot-path reads, produced the largest single latency improvements we measured.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Targeted refactoring of performance-critical paths.&lt;/strong&gt; Not a rewrite — a refactor scoped specifically to the paths the profiler flagged. This mattered for maintainability as much as speed: refactoring only what the data justified meant the team could reason about &lt;em&gt;why&lt;/em&gt; each change existed, instead of inheriting a large diff whose performance rationale had to be taken on faith.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Infrastructure-level tuning.&lt;/strong&gt; Some latency wasn't in application code at all. Right-sizing EC2 instances and Lambda configurations for actual observed load, rather than defaults inherited from an earlier stage of the business, closed gaps that no amount of query optimisation would have touched.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that's easy to skip: measuring in production, safely
&lt;/h2&gt;

&lt;p&gt;A profiler on a laptop tells you about your laptop. Production traffic has a shape — concurrency patterns, cache warmth, data skew — that's genuinely hard to fake in staging. The way to close that gap without risking the platform is incremental rollout with real metrics at every step:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1. Ship the change behind a flag or to a small traffic percentage
2. Compare p50/p95/p99 latency against the control group, not just the mean
3. Watch error rate and downstream load, not just the metric you're optimising
4. Roll forward only when the data says so — not the calendar
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The p95/p99 discipline matters specifically because averages hide the experience of your worst-off users, and in e-commerce, the worst-off users during peak trading are disproportionately the ones checking out at the exact moment load is highest — which is to say, the moment you can least afford to be wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Automation as a latency and reliability strategy
&lt;/h2&gt;

&lt;p&gt;Alongside the latency work, we built internal automation that cut manual data processing by 80%. That's not directly a latency number, but it's related: manual processes are a hidden source of both delay and error, and every manual step you remove is one less place where a human under time pressure introduces a mistake during exactly the high-traffic periods when mistakes are costliest. Reliability and performance work end up reinforcing each other more than people expect.&lt;/p&gt;

&lt;h2&gt;
  
  
  What held up across the whole effort
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Profile before you touch anything.&lt;/strong&gt; Every optimisation that didn't start here either did nothing measurable or made something else worse.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Optimize what the data flags, not what looks inefficient.&lt;/strong&gt; These are not the same list, and the gap between them is where wasted effort lives.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure tail latency, not averages.&lt;/strong&gt; Peak-trading experience is a p99 problem, not a p50 problem.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ship incrementally, with rollback as a real option, not a theoretical one.&lt;/strong&gt; 99.9% uptime through peak periods isn't compatible with big-bang deploys.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The unglamorous truth is that most of the 30% came from patient, boring, well-measured work — not a single clever trick. That's usually how sustainable performance work goes, and it's the part that's easy to leave out of the story.&lt;/p&gt;

</description>
      <category>performance</category>
      <category>postgres</category>
      <category>aws</category>
    </item>
  </channel>
</rss>
