Back in January 2026, security researchers audited a major skill marketplace and found 341 malicious skills out of 2,857 — 11.9%. One coordinated campaign. By February the count of confirmed malicious skills crossed a thousand.
But here's the part that stuck with me: the audits all happened at publish time.
The gap nobody closes
Every skill you install — a SKILL.md, some scripts, maybe a helper module — lands as plain files on your disk. The marketplace scan checks them once, at the moment of publication. After that, the files are yours, and nothing watches them.
Meanwhile skills get updated. Upstream repos push changes. You pull. A helper script that was clean when you installed it in March doesn't have to be clean in June — and no re-scan fires, because from the registry's point of view nothing happened. The rug-pull class of attack isn't theoretical; it's documented, with hashing-based detection named as the mitigation and then left as an exercise for the reader.
I looked for a tool that closes this gap: something that says "the skills on your disk are not what you installed." It didn't exist as a standalone, cross-platform thing — detection research is all publish-time or LLM-based triage. So I built the boring version.
skill-integrity
One Python file, zero dependencies, no network calls:
python3 skill_integrity.py snapshot # right after installing skills
python3 skill_integrity.py check # any time later
The check output:
skills scanned: 23 (snapshot taken 2026-10-10T21:44:32+0800)
[~] pdf-extractor — MODIFIED
changed: SKILL.md
added: fetch_helper.sh
1 skill(s) changed since snapshot. Review diffs before trusting them.
That's the whole product. SHA-256 every file, keep a local manifest, diff against it. Modified files, added files, deleted skills, brand-new skills — all four show up. --json for machines, --fail-on-change for CI:
0 */12 * * * python3 ~/skill-integrity/skill_integrity.py check --fail-on-change --json >> ~/.skill-integrity/checks.log
What it deliberately doesn't do
A hash catches any byte change. It cannot judge intent — a legitimate upstream update and a malicious rewrite look identical. That's not a limitation I'm hiding; it's the shape of the problem. The tool's job is to be the tripwire: something changed, here's the file list, go read the diff. Reading the diff before trusting the skill is your half of the contract.
This is the same honesty boundary I keep hitting in this suite: signature systems are a step behind novel attacks, and pretending otherwise is how security tools lose trust. agent-ledger says "novel evasions remain an open problem" in its README; skill-integrity says "hashes don't read minds."
Try it
git clone https://github.com/ZidoCode/skill-integrity
python3 skill-integrity/skill_integrity.py snapshot
If you run Claude Code or Hermes with a pile of community skills installed, run check right now — I'd genuinely like to know how many people find their skills have drifted.
Building in public: agent security and reliability tooling. Also in the suite: agent-ledger (audit + policy enforcement) and agent-janitor (process hygiene). github.com/ZidoCode
Top comments (0)