<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: SmartWord World</title>
    <description>The latest articles on DEV Community by SmartWord World (@smartword_world_0b5b38003).</description>
    <link>https://dev.to/smartword_world_0b5b38003</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4115445%2F0e6fc6c7-a52e-498b-a1ec-9e1f01771045.png</url>
      <title>DEV Community: SmartWord World</title>
      <link>https://dev.to/smartword_world_0b5b38003</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/smartword_world_0b5b38003"/>
    <language>en</language>
    <item>
      <title>My CI check passed every My CI check passed every build for four days. The JSON said {}</title>
      <dc:creator>SmartWord World</dc:creator>
      <pubDate>Thu, 24 Sep 2026 17:04:38 +0000</pubDate>
      <link>https://dev.to/smartword_world_0b5b38003/my-ci-check-passed-every-my-ci-check-passed-every-build-for-four-days-the-json-said--bgm</link>
      <guid>https://dev.to/smartword_world_0b5b38003/my-ci-check-passed-every-my-ci-check-passed-every-build-for-four-days-the-json-said--bgm</guid>
      <description>&lt;p&gt;This week I found out the GitHub Action for my open-source tool never failed a single build.&lt;/p&gt;

&lt;p&gt;The tool is DocsWatcher. It scans code for API calls that already have a shutdown date (OpenAI models, Stripe API versions, that kind of thing) and fails CI when it finds one. So a check that never fails is the one bug it cannot have.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the Action decides
&lt;/h2&gt;

&lt;p&gt;The CLI is Java, compiled to a native binary with GraalVM. The Action runs it, reads the JSON report, counts the findings whose severity is &lt;code&gt;breaking&lt;/code&gt;, and fails the step if that count isn't zero.&lt;/p&gt;

&lt;p&gt;That design was deliberate. The CLI exits 1 when a breaking finding exists, and the Action treats exit 1 as "a result, not a failure". It wants the count from the report, not a bare exit code.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the binary actually printed
&lt;/h2&gt;

&lt;p&gt;For every finding, this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Four findings, no fields. The count of breaking findings was always zero, so the step always passed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why
&lt;/h2&gt;

&lt;p&gt;GraalVM builds a closed world. Anything reached only by reflection is dropped unless it is registered. Jackson writes a Java record by reading its accessor methods reflectively, and my &lt;code&gt;Finding&lt;/code&gt; record was never registered. So the native binary serialised every finding as an empty object. No exception, no warning.&lt;/p&gt;

&lt;p&gt;The fix is a few lines of reachability metadata:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"dev.docswatcher.engine.Finding"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"methods"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"severity"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"parameterTypes"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"change"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"parameterTypes"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(The real entry lists every accessor.)&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the tests missed it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The unit tests run on the JVM, where reflection just works. They were green the whole time.&lt;/li&gt;
&lt;li&gt;The release smoke test did run the native binary. It checked that a repository with a breaking finding exits 1. It did. The exit code was right all along. Only the JSON was empty.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What changed
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The release smoke test now checks the JSON itself, on each platform's own runner:
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'%s\n'&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$json&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-q&lt;/span&gt; &lt;span class="s1"&gt;'"severity": "breaking"'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"::error::the findings carry no severity"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;1&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;The Action refuses a report whose findings have no severity, instead of counting it as clean. A broken binary now fails loudly.&lt;/li&gt;
&lt;li&gt;The Action's &lt;code&gt;v0&lt;/code&gt; tag points at the fixed release, so existing users get the fix without changing anything.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Looking for the same class of bug found a second one: the native binary crashed on any repository where a finding had more than one location, because an array type (&lt;code&gt;Evidence[]&lt;/code&gt;) wasn't registered either.&lt;/p&gt;

&lt;h2&gt;
  
  
  The lesson
&lt;/h2&gt;

&lt;p&gt;Test the artifact you ship, not the build you test. My tests proved the JVM build was right. Nobody had checked the bytes users actually downloaded.&lt;/p&gt;

&lt;p&gt;A question for anyone shipping GraalVM native images or other AOT builds: how do you catch missing reflection metadata before your users do? The tracing agent helps, but it only sees the paths your tests exercise.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;DocsWatcher is open source: &lt;a href="https://github.com/jameskomo/docswatcher" rel="noopener noreferrer"&gt;github.com/jameskomo/docswatcher&lt;/a&gt;. The fix is in v0.3.0 and later.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>automation</category>
      <category>cicd</category>
      <category>github</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Anthropic shipped an official eval runner for Claude Code. I made my tool run on it the same day. Field notes.</title>
      <dc:creator>SmartWord World</dc:creator>
      <pubDate>Sat, 12 Sep 2026 05:38:20 +0000</pubDate>
      <link>https://dev.to/smartword_world_0b5b38003/anthropic-shipped-an-official-eval-runner-for-claude-code-i-made-my-tool-run-on-it-the-same-day-425k</link>
      <guid>https://dev.to/smartword_world_0b5b38003/anthropic-shipped-an-official-eval-runner-for-claude-code-i-made-my-tool-run-on-it-the-same-day-425k</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk1lsghp2mj8l16l50gat.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk1lsghp2mj8l16l50gat.png" alt=" " width="800" height="749"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This week Claude Code 2.1.269 made &lt;code&gt;claude plugin eval&lt;/code&gt; public: an official runner for plugin eval suites, with generated starter cases, a no-plugin ablation arm by default, an HTML report and JSON output. I maintain config-drift-checker, a tool that had shipped its own compatible runner while the official one was gated. So release day was migration day, and these are the notes.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gotchas
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. The frontmatter whitelist rejects, it doesn't warn.&lt;/strong&gt; My cases carried a custom &lt;code&gt;covers:&lt;/code&gt; key for coverage tracking. Under my own lenient parser, fine. Under the official runner, every case carrying it failed to load with "unknown frontmatter key". If you attach custom metadata to cases, put it in a sidecar file next to prompt.md; the runner ignores files it doesn't know.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The YAML parser is strict.&lt;/strong&gt; &lt;code&gt;description: Proves the skill shapes generated Java: envelope, injection&lt;/code&gt; dies on the second colon. Two of my six cases failed this way. Quote any description containing ": ".&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. &lt;code&gt;--case&lt;/code&gt; matches the name, which defaults to the directory.&lt;/strong&gt; My case lived in &lt;code&gt;evals/spring-work-triggers-skill/&lt;/code&gt; with a pretty &lt;code&gt;name: Spring work triggers the conventions skill&lt;/code&gt;, and filtering by the directory matched nothing. The real fix (credit to a commenter below) is better than any workaround: set no &lt;code&gt;name:&lt;/code&gt; at all. It defaults to the folder name, so the filter and the report keys align with your directories. My suite dropped its explicit names the same day.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. The transcript is ephemeral, but grader verdicts survive it.&lt;/strong&gt; Tool calls and the full response go to a temp trace file, deleted when the command exits. The JSON you keep has scores, turns, cost, and (another correction from the comments) the per-run verdicts of the cheap graders with their explanations: &lt;code&gt;regex&lt;/code&gt;, &lt;code&gt;tool_used&lt;/code&gt; and &lt;code&gt;file_exists&lt;/code&gt; are computed from the transcript before it vanishes, and lines like "Skill called 1x" persist forever. What you lose is the raw material for anything you didn't write a grader for. If you need that too, run with &lt;code&gt;--keep-temp&lt;/code&gt; and copy the trace out before you exit.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I deleted, and what I didn't
&lt;/h2&gt;

&lt;p&gt;The migration was also a strategy decision. The official runner made my runner a fallback for old versions, and that's fine: the runner was never the product. What the official command deliberately doesn't do (the docs tell you to pin your model so results stay comparable) is remember anything. No stored baseline. No history. No answer to "did this get worse since last month", "is this case just flaky", or "which Claude Code release broke it".&lt;/p&gt;

&lt;p&gt;So that's the product: same cases, official format, official runner underneath, and on top a pinned baseline, history across every release, per-case noise bands learned from each case's own past (one of mine naturally swings 0.75, a fixed threshold would either alarm daily or never), a canary track that runs when a Claude Code release actually ships, and a bump PR once a new model proves itself green twice.&lt;/p&gt;

&lt;p&gt;Their command answers "does my plugin work right now on my machine". The layer on top answers "did anything stop working since the baseline, across every release, without me watching".&lt;/p&gt;

&lt;p&gt;v0.5.0 with all of the above shipped the same day their release did: &lt;a href="https://github.com/jameskomo/config-drift-checker/releases/tag/v0.5.0" rel="noopener noreferrer"&gt;https://github.com/jameskomo/config-drift-checker/releases/tag/v0.5.0&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The repo, with a full comparison table of built-in vs added: &lt;a href="https://github.com/jameskomo/config-drift-checker" rel="noopener noreferrer"&gt;https://github.com/jameskomo/config-drift-checker&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you're building your own layer on their JSON instead, I'd genuinely like to compare notes; the noise-band problem in particular is deeper than it looks.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>llm</category>
      <category>testing</category>
    </item>
    <item>
      <title>Your CLAUDE.md has zero tests. Here's what happened when I broke mine on purpose.</title>
      <dc:creator>SmartWord World</dc:creator>
      <pubDate>Tue, 08 Sep 2026 10:03:18 +0000</pubDate>
      <link>https://dev.to/smartword_world_0b5b38003/your-claudemd-has-zero-tests-heres-what-happened-when-i-broke-mine-on-purpose-58mp</link>
      <guid>https://dev.to/smartword_world_0b5b38003/your-claudemd-has-zero-tests-heres-what-happened-when-i-broke-mine-on-purpose-58mp</guid>
      <description>&lt;p&gt;Claude Code shipped roughly 25 releases last month. My setup, 400 lines of &lt;code&gt;CLAUDE.md&lt;/code&gt;, four skills and a guard hook, had exactly zero tests against any of them.&lt;/p&gt;

&lt;p&gt;That bothered me, because this configuration has behavior, and behavior breaks. A new model version ships and suddenly a skill stops triggering. Nothing errors. Nothing goes red. The agent just quietly stops doing the thing you taught it, and you notice three weeks later when the code review comes back weird.&lt;/p&gt;

&lt;p&gt;So I built CI for it: &lt;a href="https://github.com/jameskomo/config-drift-checker" rel="noopener noreferrer"&gt;config-drift-checker&lt;/a&gt; turns your &lt;code&gt;CLAUDE.md&lt;/code&gt;, skills and hooks into eval cases and re-runs them on every Claude Code release and every PR that touches the setup.&lt;/p&gt;

&lt;p&gt;But a tester you've never seen fail is worthless. So I sabotaged my own setup to test the tester.&lt;/p&gt;

&lt;h2&gt;
  
  
  Attempt one: the sabotage that did nothing
&lt;/h2&gt;

&lt;p&gt;I deleted the &lt;code&gt;skills&lt;/code&gt; entry from my plugin manifest, expecting everything to break.&lt;/p&gt;

&lt;p&gt;Nothing broke.&lt;/p&gt;

&lt;p&gt;Claude Code auto-discovers the skills directory, so the manifest key does nothing. My first sabotage was a no-op, which is exactly the kind of thing you only learn by trying to break your own system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Attempt two: the realistic break
&lt;/h2&gt;

&lt;p&gt;This time I rewrote one skill's trigger description the way a careless PR would. The skill still existed, still had all its content, but its description now talked about Terraform instead of Spring services.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After sabotage&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Suite score&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.36&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tripwire case (did the skill fire?)&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.00&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Content cases&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;0.33 to 0.60&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The content cases dropped because the agent no longer followed conventions it used to follow. And one case went to exactly zero: the tripwire.&lt;/p&gt;

&lt;h2&gt;
  
  
  The tripwire
&lt;/h2&gt;

&lt;p&gt;A tripwire case has a single grader:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Was the skill actually invoked?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Not "did the output look right". Just "did the trigger fire". When a trigger breaks, that case reads 0 out of 3 runs, every run, every day, and no amount of model randomness produces that pattern.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hard part is noise, not scoring
&lt;/h2&gt;

&lt;p&gt;One failed run doesn't mean your rule broke. Models are stochastic; sometimes they just ignore an instruction once. Telling "rule stopped firing" apart from "model ignored it run" turned out to be the real engineering problem.&lt;br&gt;
                                                                                                                                                                               Three things made the signal trustworthy:&lt;br&gt;
                                                                                                                                                                               1. &lt;strong&gt;Three runs per case.&lt;/strong&gt; A genuine flake usually recovers within the same three runs. A broken trigger never does.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Noise bands learned from history.&lt;/strong&gt; Each case's expected spread comes from its own past runs. One of my cases naturally swings by 0.75, so a fixed threshold would either alarm daily or catch nothing. A drop inside the band gets an amber "noisy" flag, never a red check.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Guards so real breaks can't hide in the band.&lt;/strong&gt; If no current run reaches the baseline, or the case has been below the bar for consecutive runs, it escalates to red anyway. Five failures in a row is not noise.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Likely refusals (zero tool calls, one turn, short reply) get labeled separately, so they don't masquerade as broken rules.&lt;/p&gt;
&lt;h2&gt;
  
  
  See the failure yourself
&lt;/h2&gt;

&lt;p&gt;The whole broken run is published, unedited:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://jameskomo.github.io/config-drift-checker/example-break/report.html" rel="noopener noreferrer"&gt;The red report from the sabotage run&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It runs as a GitHub Action on your runner with your key:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;jameskomo/config-drift-checker/action@v0&lt;/span&gt;
  &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;plugin-dir&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Costs $0 on a Claude Pro or Max plan via a subscription token.&lt;/p&gt;

&lt;p&gt;Repo: &lt;a href="https://github.com/jameskomo/config-drift-checker" rel="noopener noreferrer"&gt;https://github.com/jameskomo/config-drift-checker&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you've built your own version of this in-house (I keep meeting people who have), I'd genuinely like to compare notes, especially on how you separate "rule stopped firing" from "model ignored it this run".&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>programming</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
