What 600 Skill Evaluations Taught Us About Writing Good Agent Skills
After scoring 600+ agent skills on the same six-dimension rubric — trigger quality, structure, workflow design, content, engineering, and security — patterns emerge. Some confirm what skill authors already believe. Some don't.
This is the data-backed version of "how to write a good skill," with the numbers behind every claim.
Pattern 1: The top 1% verify their own output
Skills scoring 9+ almost universally include a verification step: a check the agent runs after doing the work, before claiming done. The official document skills (xlsx, docx, pptx — all 9.7+) recalculate and validate after every write. The best skill we've evaluated in the finance category simulates trades before executing them.
The correlation is strong enough to be a heuristic: no verification step, no top score. It's worth 10-15 points on our rubric, and in practice it's the difference between an agent that says "done" and one that says "done, and here's the evidence."
Pattern 2: Trigger descriptions are the weakest link
The single most common failure across 600 skills: trigger descriptions that don't say when not to fire. Over 60% of evaluated skills lack any When-Not boundary.
Why it matters: a skill without boundaries loads itself into every half-relevant conversation, diluting the agent's attention. Your skill becomes noise.
The fix is one sentence in the description: "Do not use this when X." The top-rated skills in our directory all have it. It's the cheapest quality improvement in the entire ecosystem.
Pattern 3: Code-drawn beats model-generated
In creative categories — video, diagrams, images — skills that draw output in code (SVG, HTML, Canvas, Remotion, manim) consistently outscore skills that call image-generation APIs. Not by a little: by roughly a full tier.
The reason is reproducibility. Code-drawn output renders the same way every time, can be version-controlled, and can be branded. API-generated output varies between calls and can't be verified. For business contexts, that's the whole ballgame.
Pattern 4: Structure quality is bimodal
Skills either have progressive disclosure or they don't. There's no middle. The distribution has two peaks: skills with a lean SKILL.md plus well-organized references, and skills with one bloated file. The gap between the peaks maps almost perfectly onto our structure scores.
The skill authors who get it right treat SKILL.md like a README and references/ like the docs. The ones who get it wrong treat SKILL.md like the docs.
Pattern 5: Popularity and safety are uncorrelated
The credential harvester we caught had 60,000 GitHub stars. Benign skills sit at 12 stars. Across the full corpus, we found no correlation between star count and security score.
This is the pattern with the most consequences. The skill ecosystem inherited npm's trust model — stars as a proxy for safety — without npm's mitigations, while giving skills more access than npm packages (your shell, your files, your agent's context). If you take one thing from 600 evaluations: read the scripts, or use a directory that does.
The one-sentence summary
If we had to compress 600 evaluations into one sentence for skill authors:
A great skill tells the agent exactly what to do, when to fire, when not to fire, and how to check that it worked — and everything it touches is declared.
Everything in our rubric is downstream of that sentence.
The rubric is public at our methodology page, every scorecard has written rationale, and the full directory is here. If you write skills, we'd love to evaluate yours.
Top comments (0)