DEV Community

Cover image for We scored the skills inside the most-starred agent repos. Star count is not a score.
SKILL123.me
SKILL123.me

Posted on Originally published at skill123.me

We scored the skills inside the most-starred agent repos. Star count is not a score.

Star counts are the only signal most people use when picking an agent skill. We have per-skill evaluations for 74 skills spread across the seven most-starred agent repos. The two numbers barely talk to each other.

Here's the whole finding in one table. Stars are live from the GitHub API on 2026-10-10; scores are ours, from a six-dimension rubric, published per skill with the written rationale.

Repo Stars Skills we've scored Score range Highest
obra/superpowers 297,128 14 7.9 – 9.1 verification-before-completion (9.1)
anthropics/skills 180,294 19 6.7 – 9.8 xlsx (9.8)
DietrichGebert/ponytail 160,220 6 8.0 – 9.0 ponytail-review (9.0)
addyosmani/agent-skills 104,373 24 7.9 – 9.5 constraint-driven-development (9.5)
Egonex-AI/Understand-Anything 85,826 9 7.0 – 9.0 understand-explain (9.0)
ayghri/i-have-adhd 56,254 1 9.0 i-have-adhd (9.0)
OthmanAdi/planning-with-files 27,374 1 8.3 planning-with-files (8.3)

A repo is not a skill

Look at the third column. anthropics/skills has 180k stars and a 6.7 sitting at the bottom of its own range. addyosmani/agent-skills spans 7.9 to 9.5 across 24 skills. Understand-Anything spans 7.0 to 9.0.

That's what a star count can't express: a repo is a bundle, and the skills inside it were written at different times, by different people, with different amounts of care. Starring the repo tells you the bundle got attention. It tells you nothing about whether the skill you're about to install is the 9.8 or the 6.7.

We see this everywhere, not just here. The weak entries inside popular repos tend to be the ones added last — after the repo was already popular, when contributions arrived faster than review.

The individual skills that scored highest

  • xlsx (9.8) and claude-api (9.8) — the official spreadsheet skill and the API reference. Both win on the same dimension: they check their own output rather than declaring success.
  • constraint-driven-development (9.5) — binds your agent to a written quality contract for the whole session. The most interesting idea in the set.
  • verification-before-completion (9.1) — the discipline half of the superpowers suite, and the skill most worth copying if you write your own.
  • i-have-adhd (9.0) and ponytail / ponytail-review (9.0) — two very different fixes for the same failure mode: agents that over-explain or over-build.

What this table is not

It is not a ranking of the biggest repos. Three of the largest — mattpocock/skills (283k), affaan-m/ECC (276k), JuliusBrussee/caveman (110k) — aren't here because we haven't evaluated their skills yet. Same for claude-mem, impeccable, and scientific-agent-skills. A missing repo means unscored, not badly scored.

It is not a complete count of the repos that are here. We've scored 24 of addyosmani's skills and 14 from the superpowers suite; both have more. Stars also move daily — treat this as a snapshot with a date on it.

And a score is an opinion, not a fact. Ours is one rubric applied consistently — trigger quality, structure, workflow design, content, engineering, security — with the rationale written out on every skill page, so you can argue with a specific number instead of the whole system.

The practical version

Pick skills, not repos. Open the repo, find the one skill that does what you need, and check whether anyone has actually read its scripts. If you use scorecards as the filter, use them at that granularity — per skill, per version — because that's the only level at which the number means anything.

Cross-posted from skill123.me, where every score above has its written rationale, and the adversarial injection-test results for 47 of these skills are open data.

Top comments (0)