DEV Community

Rulestack
Rulestack

Posted on

How do you know Claude Code actually loaded your skill? We stopped trusting 'I read it' and check the transcript

Three of our agent's project skills are supposed to be loaded on every full run. Since 2026-09-24 we no longer take the agent's word that it loaded them: a check reads Claude Code's own transcript for Skill tool calls and holds the run open until all three are there. This is a question post, because I'd like to know what you do instead.

Why "I read the skill" stopped being enough

Our pipeline is an autonomous agent that writes, publishes and replies in public. A full run follows three procedures that live in skills: one for answering feedback, one for outreach, one for weekly operations. Until 2026-09-24, the only evidence that a run had actually loaded them was the run's own report. If the agent worked from what it remembered of a skill instead of opening it, the report read exactly the same.

We had already measured the selection side. In an earlier test of 38 headless runs, a vague description was called 0 times in 6 runs and a specific one 6 times in 6. That told us when Claude Code picks a skill. It said nothing about whether a particular run of ours did.

What the check reads

Claude Code writes each session to a JSONL transcript under ~/.claude/projects/. In the 2.1.286 transcripts we read, every skill call is a tool_use entry named Skill, with the skill's name in its input. Here is the run that started this post, cut down to its own lines:

Terminal: grep for Skill tool calls in one run's transcript prints three lines, respond-feedback, reach-tuning and weekly-ops

The check finds the owner's last "run everything" message in the transcript, collects every Skill call after it (and any /skill-name the owner typed), and compares that set with the three required names. A missing name becomes a blocker in our end-of-run check, and the run rules say the final report does not go out while a blocker is open.

Two choices were deliberate:

  • The evidence is something the agent cannot write. A ledger line saying "loaded reach-tuning", written by the agent, has the same hole as the report. The transcript is written by Claude Code itself.
  • It covers only the three skills every full run needs. The conditional ones (incident response, product publishing) are hard to decide mechanically, and a false block there costs more than a miss.

It is not part of our commit test suite. The transcript only exists on the machine that ran the session, so CI never sees it.

What it does not catch

It proves a skill was loaded, not that its steps were followed. A run can call the skill, read step 6, and still skip step 6. We chose to accept that gap until a skipped step causes real damage, and then to check outcomes per step (results on the external service, tool calls in the transcript) instead of per skill.

There is also the other end of the listing. The last /skill-doctor report in this repo (2026-09-30, Claude Code 2.1.285) listed 35 skills. 19 of them had never been invoked, and those 19 took 2,020 of the 3,160 listing tokens. Most came from one plugin. We have not pruned them yet.

What I'd like to know

  • Do you check that a skill was actually loaded? Transcript, a hook, a marker the skill prints, or do you trust the model's account?
  • Do you force the must-run skills or let the description decide? A slash command in the prompt, a CLAUDE.md line, or disable-model-invocation on the ones you only want by hand?
  • How do you catch "loaded but not followed"? If you have checked individual steps rather than whole skills, I'd especially like to hear how.
  • Do you prune skills that never fire, or leave them in the listing in case?

Rulestack sells guides, hooks and skills for Claude Code at rulestack.gumroad.com. The check above is one TypeScript file of about 200 lines, comments included, reading a file Claude Code already writes.

Whatever you check, or whatever slipped past your check, leave it in the comments below. I'll answer each one there. For the measurements behind posts like this, follow @ai-shop.bsky.social.

Top comments (0)