DEV Community

Cover image for skillcheck Update: Scorer Fixes, Cleaner Failures, Honest Token Numbers
Brad Kinnard
Brad Kinnard Subscriber

Posted on

skillcheck Update: Scorer Fixes, Cleaner Failures, Honest Token Numbers

skillcheck is a static analyzer for SKILL.md files, the format agents like Claude Code, Copilot, Codex, and Cursor use to load reusable skills. It validates frontmatter, scores description discoverability, checks file references, enforces token budgets, and flags cross-agent compatibility issues. No network calls, no LLM calls, no file mutations. Runs as a CLI, a GitHub Action, or a pre-commit hook.

pip install skillcheck
skillcheck skills/
Enter fullscreen mode Exit fullscreen mode

Latest pass was hardening and accuracy, not features. Here's what changed and why.

Description scores went up. Skills that were scoring low because the scorer was broken will now see a jump in scoring. Median across the reference corpus went from 75 to 90. --explain-score also now tells you which pattern hits or misses instead of just a number. The score exists to predict whether an agent will actually find and trigger your skill, so a scorer that under-credits good descriptions defeats the point. The fix was validated against real-world skills, and the separation held: filler still scores 28-65, well-written descriptions 85-100.

Corrupt files now fail cleanly instead of crashing. Before, a bad history ledger or non-UTF-8 skillcheck.toml above the skill dumped a Python traceback. It's now a clear error naming the file and byte offset (exit code 2). Config discovery walks up the directory tree, so one bad file could break every scan under it. Now every untrusted read (ingest, history, config) goes through the same guard before parsing, so they all reject the same way.

README has been corrected in regards to token estimates. Without tiktoken, expect roughly 20-30% over-estimation, so install the extra if you're near a budget limit. The offline heuristic feeds the budget checks and its accuracy had never actually been measured, just assumed. It's benchmarked against tiktoken across the full corpus now, and the documented numbers are the measured ones.

pip install "skillcheck[tiktoken]"
Enter fullscreen mode Exit fullscreen mode

The rest of the pass is invisible on purpose: flag-conflict logic consolidated to one source of truth, golden-file tests pinning exact diagnostic output, coverage floor raised from 75% to 80% (actual sits at 90%). Diagnostic output across the corpus verified byte-for-byte identical before and after. Nothing changed except what's above.

GitHub logo moonrunnerkc / tracemantle

Validate SKILL.md files, track agent-skill bundle changes, and check release evidence against trusted policies.

TraceMantle

Validate agent skills and check the evidence behind a release.

CLI reference · Report a bug or request a feature

CI status PyPI version Python versions MIT license
Contents
  1. About the project
  2. Getting started
  3. Usage
  4. Integrations
  5. Contributing
  6. License and contact

About the project

TraceMantle is a Python CLI and library for checking agent skill bundles before they are committed or released. It provides:

  • Skill validation: checks SKILL.md frontmatter, file references, size limits and compatibility advice against the Agent Skills specification.
  • Bundle tracking: hashes skill files and packaged resources so helper changes are detected too.
  • Evidence checks: imports evaluator results and compares bundles against trusted policies. Missing, stale or incompatible evidence blocks the gate.

TraceMantle analyzes files; it does not execute skills or evaluators. Static checks and imported model judgments do not prove that a skill works in a live agent.

Getting started

Requires Python 3.10 or later. Install in a virtual environment:

python -m venv .venv
source
…
Enter fullscreen mode Exit fullscreen mode

If skillcheck flags something in your skills that looks wrong, open an issue. The reference corpus grows from real-world cases and the scorer improves with them.

Top comments (0)