Claude Haiku 5.5 came out on 7 October. Within a day I ran three popular Claude Code skills on it, and on Haiku 4.5 next to it. Three runs each, same task, same grader, same Claude Code version.
Here's the result that made me laugh.
The git workflow skill on Haiku 5.5, scored 0 to 1:
| Run | Without the skill | With the skill | Gap |
|---|---|---|---|
| 1 | 0.300 | 0.849 | +0.549 |
| 2 | 0.849 | 0.839 | -0.010 |
| 3 | 0.613 | 0.848 | +0.235 |
Run 1 is the screenshot that goes viral. "This one skill makes Haiku 5.5 55 points better."
Run 2 says the skill does nothing.
Same setup both times. And look at where the swing actually is: the skill arm sits at about 0.84 every time. It's Haiku 5.5 without the skill that jumped from 0.300 to 0.849 between runs. So the "+55 points" was mostly one bad baseline run.
Run it once and you would have published a fake win, with a real number attached.
The rest of it
- No skill clearly helped Haiku 5.5 in two or more of its three runs.
- The documentation skill clearly helped Haiku 4.5 in two of three runs (+0.14, +0.19, then +0.13 that was too noisy to call). On Haiku 5.5, all three runs had too few answers to tell either way. So does it still help on the new model? I can't tell you. Neither can anyone who ran it once.
- The code review skill on Haiku 4.5: run 1 said +0.34, clearly helped. Runs 2 and 3 said +0.09 and +0.02, too few answers to tell.
- Four of the six model and skill pairs changed their verdict between repeats.
That last line is the whole story. A lot of "I tried skill X and Claude got way better" posts are one run. Some popular skill-eval tools run a single attempt by default. Any one-run test would have reported run 1 above as a big win.
What this does not show
One task per skill, and only 3 to 10 answers per side in each run, so "too few answers to tell" means these runs couldn't see a 0.05 gap. It does not mean the skill is useless. The Haiku 5.5 calls also ran with Claude Code's per-turn effort setting and the Haiku 4.5 calls didn't, so the two models differ in more than the model. Every answer was scored three times by Claude Opus 5. This is not a model ranking.
The skills are from Addy Osmani's agent-skills pack, and I'm not knocking them. The point is that one run can't tell you whether any skill works, good or bad.
Full report, every number and every receipt: driftproofhq.com/reports/014
So, honest question: how many times do you run a skill before you decide it works?
*I maintain Driftproof, the open-source Claude Code plugin that ran this.
Top comments (1)
The swing being entirely in the baseline arm is the pattern that breaks most agent evals. The skill held the model steady at roughly 0.85 across all three attempts, providing an execution rail that kept runs from spiraling on dirty git worktree states.
When a small model fails an unassisted agent loop, the failure is rarely an incremental score dip. An early bad command sends it down a catastrophic branch where it burns its remaining turns trying to recover. A skill bounds that floor. Evaluating against a single baseline run measures the unassisted model's tail risk instead of any capability increase.