We ran eight SKILL.md files on a strong and a weak model. Both mostly passed our checks; the weak one failed in ways checks cannot see, and used more tokens.
Mostly not where you would look. On our eight skills the weaker model passed the mechanical checks about as often as the stronger one. What it got wrong was quieter: claims nobody gave it, conditions rewritten into something else, a formal address dropped, a Bulgarian misspelling left in place every time, severity inflated. It also spent several times more output tokens on the same task, and most of those were reasoning tokens.
Read the full report on AISkills402: https://aiskills402.com/blog/weak-model-eight-skills
Top comments (0)