<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Nadir Ali</title>
    <description>The latest articles on DEV Community by Nadir Ali (@menadirali).</description>
    <link>https://dev.to/menadirali</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4157218%2Fdce30261-e949-4afc-9b1c-893c34144bd0.jpg</url>
      <title>DEV Community: Nadir Ali</title>
      <link>https://dev.to/menadirali</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/menadirali"/>
    <language>en</language>
    <item>
      <title>I A/B-tested 9 popular AI agent skills. 4 of them did nothing.</title>
      <dc:creator>Nadir Ali</dc:creator>
      <pubDate>Sun, 04 Oct 2026 12:29:23 +0000</pubDate>
      <link>https://dev.to/menadirali/i-ab-tested-9-popular-ai-agent-skills-4-of-them-did-nothing-5ba5</link>
      <guid>https://dev.to/menadirali/i-ab-tested-9-popular-ai-agent-skills-4-of-them-did-nothing-5ba5</guid>
      <description>&lt;p&gt;Agent skills are everywhere this year. A skill is a &lt;code&gt;SKILL.md&lt;/code&gt; file that teaches a coding agent&lt;br&gt;
(Claude Code, Codex, Cursor, Gemini CLI…) how to do something: verify before saying "done",&lt;br&gt;
keep diffs small, review code for real bugs. Some skill repos have hundreds of thousands of stars.&lt;/p&gt;

&lt;p&gt;I noticed nobody measures them. A skill is a prompt, and whether a prompt helps depends on the&lt;br&gt;
model reading it. A skill written for last year's model might do nothing on this year's, or&lt;br&gt;
make it worse. So I tested them.&lt;/p&gt;
&lt;h2&gt;
  
  
  What I did
&lt;/h2&gt;

&lt;p&gt;I took 9 of the most popular skill ideas and rewrote them for current models: short, calm, no&lt;br&gt;
walls of &lt;code&gt;MUST&lt;/code&gt; and &lt;code&gt;NEVER&lt;/code&gt;. Then I gave each one an eval suite using Anthropic's&lt;br&gt;
&lt;a href="https://code.claude.com/docs/en/plugin-evals" rel="noopener noreferrer"&gt;&lt;code&gt;claude plugin eval&lt;/code&gt;&lt;/a&gt;, which runs every task&lt;br&gt;
twice:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;with&lt;/strong&gt; the skill installed, and&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;without&lt;/strong&gt; it, as a baseline.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The difference between the two scores (Δ) is what the skill actually adds. That came to 19 test&lt;br&gt;
cases, 3 runs each per arm, on &lt;strong&gt;Sonnet 5.5&lt;/strong&gt; and &lt;strong&gt;Haiku 4.5&lt;/strong&gt;. The prompts read like what a&lt;br&gt;
real user would type, and they never name the skill. The graders check outcomes ("were the&lt;br&gt;
unrelated lines left untouched?", "did it admit the fix was untested?"), not whether the reply&lt;br&gt;
followed the skill's own formatting.&lt;/p&gt;
&lt;h2&gt;
  
  
  The results
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Skill&lt;/th&gt;
&lt;th&gt;Sonnet 5.5 Δ&lt;/th&gt;
&lt;th&gt;Haiku 4.5 Δ&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;grill&lt;/code&gt;: interview me before coding, one question at a time, each with a recommended answer&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+75&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+58&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;✅ keep&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;bug-hunt-review&lt;/code&gt;: report only real bugs, each with a concrete failing input&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+10&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+17&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;✅ keep&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;handoff&lt;/code&gt;: write a note a fresh session can resume from&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+13&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;✅ keep (small models)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;prove-it&lt;/code&gt;: don't say "fixed" without the command that shows it&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+11&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;✅ keep (small models)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;root-cause&lt;/code&gt;: fix the bug where it starts, not where it was reported&lt;/td&gt;
&lt;td&gt;+13&lt;/td&gt;
&lt;td&gt;−11&lt;/td&gt;
&lt;td&gt;⚠️ on probation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;surgical&lt;/code&gt;: smallest possible diff&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;✂️ cut&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;stdlib-first&lt;/code&gt;: built-ins before new packages&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;✂️ cut&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;answer-first&lt;/code&gt;: first sentence is the answer&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;+3&lt;/td&gt;
&lt;td&gt;✂️ cut&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;secure-defaults&lt;/code&gt;: parameterized SQL, no shell strings&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;−8&lt;/td&gt;
&lt;td&gt;✂️ cut&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Four of nine were cut. They're still in the repo under &lt;code&gt;retired/&lt;/code&gt;, with their evals, so anyone&lt;br&gt;
can re-test them on a future model.&lt;/p&gt;
&lt;h2&gt;
  
  
  What I learned
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Skills that add a workflow help. Skills that restate good habits don't.&lt;/strong&gt;&lt;br&gt;
&lt;code&gt;grill&lt;/code&gt; makes the model do something it wouldn't choose on its own: ask one question at a time&lt;br&gt;
and recommend an answer for each. Without it, Sonnet asked five or more questions at once and didn't recommend&lt;br&gt;
an answer for any of them. That's a +75 point difference. But "keep your diff small" and "parameterize&lt;br&gt;
your SQL"? Sonnet 5.5 already does that. The skill adds nothing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Most "be careful" skills never even loaded.&lt;/strong&gt;&lt;br&gt;
On natural prompts, &lt;code&gt;surgical&lt;/code&gt;, &lt;code&gt;stdlib-first&lt;/code&gt;, &lt;code&gt;answer-first&lt;/code&gt; and &lt;code&gt;secure-defaults&lt;/code&gt; were loaded&lt;br&gt;
in &lt;strong&gt;0 of 6&lt;/strong&gt; runs. The model decided they weren't relevant, and it was right: it already behaved&lt;br&gt;
that way. A skill that never fires still costs context on every turn, because its description&lt;br&gt;
is always loaded.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Smaller models benefit more.&lt;/strong&gt;&lt;br&gt;
&lt;code&gt;handoff&lt;/code&gt; and &lt;code&gt;prove-it&lt;/code&gt; did nothing for Sonnet, which already writes accurate handoffs and&lt;br&gt;
admits when it couldn't run the tests. Haiku gained 11–13 points from them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. A skill's description alone can change behavior.&lt;/strong&gt;&lt;br&gt;
&lt;code&gt;root-cause&lt;/code&gt; never loaded on either model, yet scored +13 on Sonnet and −11 on Haiku. The only&lt;br&gt;
part of it the model saw was its one-line description in the skill list. That's a weak, noisy&lt;br&gt;
effect, so it stays on probation instead of claiming a win.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Check your graders before you trust your numbers.&lt;/strong&gt;&lt;br&gt;
My first run showed &lt;code&gt;handoff&lt;/code&gt; &lt;em&gt;hurting&lt;/em&gt; Sonnet by 13 points. The cause was my grader: it&lt;br&gt;
required the first "next step" to name a function to change, and it failed the correct answer,&lt;br&gt;
"re-run the tests first". After I fixed the grader, the effect was 0. The fix is noted in the&lt;br&gt;
changelog, and the README table only uses the corrected run.&lt;/p&gt;
&lt;h2&gt;
  
  
  Caveats
&lt;/h2&gt;

&lt;p&gt;Three runs per arm is noisy, and some of my cases are probably too easy: when the baseline&lt;br&gt;
already scores 100%, a skill can't show a benefit. Harder eval cases are the most useful thing&lt;br&gt;
anyone could contribute.&lt;/p&gt;
&lt;h2&gt;
  
  
  Bonus: skills are code you run with your agent's permissions
&lt;/h2&gt;

&lt;p&gt;While doing this I built &lt;strong&gt;&lt;code&gt;skill-vet&lt;/code&gt;&lt;/strong&gt;, a zero-dependency scanner you can point at any skill&lt;br&gt;
repo &lt;em&gt;before&lt;/em&gt; installing it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @menadirali/skill-vet vet owner/repo
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It checks for download-and-execute (&lt;code&gt;curl … | sh&lt;/code&gt;), hidden Unicode, prompt-injection phrasing,&lt;br&gt;
credential access, the skill spec, and how many tokens a skill pack adds to every session. I ran&lt;br&gt;
it on 115 skills from 9 of the most popular skill repos. It found no download-and-execute,&lt;br&gt;
prompt-injection, or hidden-Unicode problems, a couple of spec errors, and lots of all-caps&lt;br&gt;
"shouting" that current models don't need.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it, or help
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Install the five surviving skills: &lt;code&gt;npx skills add nadirali1350/vetted&lt;/code&gt;, or in Claude Code,
&lt;code&gt;/plugin marketplace add nadirali1350/vetted&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Repo, full results, and every eval case: &lt;strong&gt;&lt;a href="https://github.com/nadirali1350/vetted" rel="noopener noreferrer"&gt;https://github.com/nadirali1350/vetted&lt;/a&gt;&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;There are good-first issues open for Hacktoberfest: new scanner rules, eval cases,
translations. Five people have contributed so far, and PRs get reviewed within a day.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you've seen your coding agent repeatedly get something wrong on a current model, I'd love to&lt;br&gt;
hear it. That's how the next skill gets written, with its eval first.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>claude</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
