DEV Community

Driftproofhq
Driftproofhq

Posted on AI-assisted

The same git skill added almost nothing to Sonnet 5.5 and took Opus 5.5 from about 0.5 to 0.86

Claude Sonnet 5.5 came out on 28 September. I started testing it at 03:53 UTC the next day: the same three agent skills and tasks I used on Opus 5.5 last week, with fresh Sonnet 5 and Opus 5.5 runs next to it, and the whole set repeated three times.

Two things came out of it. The same skill can matter a lot on one model and barely at all on another. And one run on its own would often have told a different story.

What I ran

Three public skills from addyosmani/agent-skills: code review (code-review-and-quality), git workflow (git-workflow-and-versioning) and writing docs (documentation-and-adrs). One task per skill, the same tasks as my earlier reports.

Each model answered each task several times with the skill's SKILL.md in the prompt, and several times without it. Claude Opus 5 graded every answer on a 0 to 1 scale, three times per answer. All of it ran on Claude Code 2.1.284, and I repeated the full set three times back to back with the same script and settings. That gives 27 result files (Driftproof calls them receipts), all published with the report.

What the skill added

Score with the skill minus score without it, lowest to highest across the three runs:

Skill Sonnet 5 Sonnet 5.5 Opus 5.5
Code review +0.085 to +0.302 +0.009 to +0.051 +0.038 to +0.190
Git workflow -0.004 to +0.020 +0.000 to +0.013 +0.276 to +0.371
Docs -0.085 to -0.006 +0.208 to +0.272 +0.116 to +0.322

The git row is the one that surprised me. Without the skill, Sonnet 5.5 already scored 0.85 to 0.86, and Opus 5.5 scored 0.49 to 0.58. With it, both landed at about 0.86. So the skill did almost nothing for one model and a lot for the other, and they ended up in the same place.

Docs goes the other way round: Sonnet 5.5 gained 0.21 to 0.27 in every run, while Sonnet 5 came out slightly lower with the skill than without it in every run.

Sonnet 5.5 against Opus 5.5

With the skill loaded, I couldn't find a clear difference between Sonnet 5.5 and Opus 5.5 on any of the three skills, in any of the three runs.

That isn't the same as "they're equal". Driftproof only calls a difference when the spreads of the two models' scores don't overlap and their averages are at least 0.05 apart. In 5 of those 9 comparisons there weren't enough answers to spot a 0.05 gap even if one was there.

A reader also made a fair point about the git task: Sonnet 5.5 (with or without the skill) and Opus 5.5 with it all land at 0.85 to 0.86. That looks like it could be the most this task's grading gives out. If so, a tie there means both models hit the top, not that they're equally good.

Against Sonnet 5, Sonnet 5.5 with the skill scored clearly higher on docs in 2 of 3 runs, and showed no clear difference on the other two skills.

Why three runs

Because the runs didn't agree with each other:

  • The code-review skill added +0.076 to Opus 5.5 in run 1, +0.038 in run 2 and +0.190 in run 3.
  • Opus 5.5 without the skill scored 0.831, 0.857 and then 0.713 on code review.
  • Sonnet 5 with the code-review skill scored 0.880, 0.881, and then 0.657 in run 3.
  • Sonnet 5.5 pulled clearly ahead of Sonnet 5 on docs in runs 1 and 3, but not in run 2.

Counting it up, 5 of the 9 skill-and-comparison pairs didn't get the same verdict in all three runs (here "no clear difference" and "too few answers to tell" count as different verdicts). Any one of these runs on its own would have made a confident-looking blog post.

Effort settings

I didn't set effort or thinking for any model, same as my earlier reports, so each ran on its Claude Code default. Claude Code applied a per-turn effort on every Sonnet 5.5 and Opus 5.5 call and on none of the Sonnet 5 calls. That's what you get out of the box, but it means this isn't a pinned-effort head-to-head.

What it cost

The runs went through a Claude subscription, so nothing was metered. At API list prices, the answers themselves came to an estimated $2.30 for Sonnet 5.5 across all three runs, $3.30 for Sonnet 5 and $5.74 for Opus 5.5. Grading was most of the bill: about $54 of the $65.48 total.

What this doesn't show

One task per skill, and 3 to 9 answers per model per run. It says nothing about how these models code in general, or how these skills will do on your own work. The grader is Claude Opus 5, as in my earlier reports, so these numbers compare with those and nothing else.

Full report, all 27 receipts and the raw evidence: https://driftproofhq.com/reports/013/

Disclosure: I build and maintain Driftproof (open source, Apache 2.0).

Top comments (0)