DEV Community

Cover image for I A/B-tested 9 popular AI agent skills. 4 of them did nothing.
Nadir Ali
Nadir Ali

Posted on

I A/B-tested 9 popular AI agent skills. 4 of them did nothing.

Agent skills are everywhere this year. A skill is a SKILL.md file that teaches a coding agent
(Claude Code, Codex, Cursor, Gemini CLI…) how to do something: verify before saying "done",
keep diffs small, review code for real bugs. Some skill repos have hundreds of thousands of stars.

I noticed nobody measures them. A skill is a prompt, and whether a prompt helps depends on the
model reading it. A skill written for last year's model might do nothing on this year's, or
make it worse. So I tested them.

What I did

I took 9 of the most popular skill ideas and rewrote them for current models: short, calm, no
walls of MUST and NEVER. Then I gave each one an eval suite using Anthropic's
claude plugin eval, which runs every task
twice:

  • with the skill installed, and
  • without it, as a baseline.

The difference between the two scores (Δ) is what the skill actually adds. That came to 19 test
cases, 3 runs each per arm, on Sonnet 5.5 and Haiku 4.5. The prompts read like what a
real user would type, and they never name the skill. The graders check outcomes ("were the
unrelated lines left untouched?", "did it admit the fix was untested?"), not whether the reply
followed the skill's own formatting.

The results

Skill Sonnet 5.5 Δ Haiku 4.5 Δ Verdict
grill: interview me before coding, one question at a time, each with a recommended answer +75 +58 ✅ keep
bug-hunt-review: report only real bugs, each with a concrete failing input +10 +17 ✅ keep
handoff: write a note a fresh session can resume from 0 +13 ✅ keep (small models)
prove-it: don't say "fixed" without the command that shows it 0 +11 ✅ keep (small models)
root-cause: fix the bug where it starts, not where it was reported +13 −11 ⚠️ on probation
surgical: smallest possible diff 0 0 ✂️ cut
stdlib-first: built-ins before new packages 0 0 ✂️ cut
answer-first: first sentence is the answer 0 +3 ✂️ cut
secure-defaults: parameterized SQL, no shell strings 0 −8 ✂️ cut

Four of nine were cut. They're still in the repo under retired/, with their evals, so anyone
can re-test them on a future model.

What I learned

1. Skills that add a workflow help. Skills that restate good habits don't.
grill makes the model do something it wouldn't choose on its own: ask one question at a time
and recommend an answer for each. Without it, Sonnet asked five or more questions at once and didn't recommend
an answer for any of them. That's a +75 point difference. But "keep your diff small" and "parameterize
your SQL"? Sonnet 5.5 already does that. The skill adds nothing.

2. Most "be careful" skills never even loaded.
On natural prompts, surgical, stdlib-first, answer-first and secure-defaults were loaded
in 0 of 6 runs. The model decided they weren't relevant, and it was right: it already behaved
that way. A skill that never fires still costs context on every turn, because its description
is always loaded.

3. Smaller models benefit more.
handoff and prove-it did nothing for Sonnet, which already writes accurate handoffs and
admits when it couldn't run the tests. Haiku gained 11–13 points from them.

4. A skill's description alone can change behavior.
root-cause never loaded on either model, yet scored +13 on Sonnet and −11 on Haiku. The only
part of it the model saw was its one-line description in the skill list. That's a weak, noisy
effect, so it stays on probation instead of claiming a win.

5. Check your graders before you trust your numbers.
My first run showed handoff hurting Sonnet by 13 points. The cause was my grader: it
required the first "next step" to name a function to change, and it failed the correct answer,
"re-run the tests first". After I fixed the grader, the effect was 0. The fix is noted in the
changelog, and the README table only uses the corrected run.

Caveats

Three runs per arm is noisy, and some of my cases are probably too easy: when the baseline
already scores 100%, a skill can't show a benefit. Harder eval cases are the most useful thing
anyone could contribute.

Bonus: skills are code you run with your agent's permissions

While doing this I built skill-vet, a zero-dependency scanner you can point at any skill
repo before installing it:

npx @menadirali/skill-vet vet owner/repo
Enter fullscreen mode Exit fullscreen mode

It checks for download-and-execute (curl … | sh), hidden Unicode, prompt-injection phrasing,
credential access, the skill spec, and how many tokens a skill pack adds to every session. I ran
it on 115 skills from 9 of the most popular skill repos. It found no download-and-execute,
prompt-injection, or hidden-Unicode problems, a couple of spec errors, and lots of all-caps
"shouting" that current models don't need.

Try it, or help

  • Install the five surviving skills: npx skills add nadirali1350/vetted, or in Claude Code, /plugin marketplace add nadirali1350/vetted
  • Repo, full results, and every eval case: https://github.com/nadirali1350/vetted
  • There are good-first issues open for Hacktoberfest: new scanner rules, eval cases, translations. Five people have contributed so far, and PRs get reviewed within a day.

If you've seen your coding agent repeatedly get something wrong on a current model, I'd love to
hear it. That's how the next skill gets written, with its eval first.

Top comments (1)

Collapse
 
reidmarlow profile image
Reid Marlow •

The split between workflow skills and habit reminders matches what happens to prompt files over a few months of maintenance.

Rules like surgical diffs or preferring stdlib don't shift frontier model behavior because modern base training already covers that distribution. The instruction takes up context tokens without altering the tool trajectory. The reason grill or bug-hunt-review move the needle is that they change the execution graph. Forcing an extra question turn or demanding a concrete failing input before writing a patch creates an external artifact the agent has to reconcile against.

The drop on Haiku with secure-defaults and root-cause is also telling. Smaller models get derailed when negative constraints compete with the primary task instructions.