DEV Community

Rulestack
Rulestack

Posted on

With 160 skills installed, Claude Code kept 20 descriptions, and the model found 0 of 3 skills that lost theirs

With Opus 5.5's 1M-token window, Claude Code 2.1.289 held the skill listing to 30,000 characters. Of 160 project skills with 1,000-character descriptions, 20 kept their descriptions (22 of 40, 21 of 80) and every name stayed, but the model invoked 0 of 3 target skills that were listed by name only, against 7 of 7 that still had descriptions.

Every skill you install adds a line to the model's context on every turn: the skill's name and its description. When the model decides on its own whether a request calls for a skill, the description is all it has to go on. The skills page of the Claude Code docs says this listing has a character budget that "scales at 1% of the model's context window", and that once the listing overflows, Claude Code drops descriptions "starting with the skills you invoke least". We wanted three things the page doesn't give: the budget as an actual number for a default Opus session, the order in which descriptions really go, and what happens to a request that should have triggered a skill whose description was dropped. So we built a lab with 10, 40, 80 and 160 dummy skills, ran Claude Code headless 38 times, and read the listing each session recorded in its transcript.

Everything below ran on 2026-10-05 (UTC) with Claude Code 2.1.289 (claude --version) and claude-opus-5-5, selected with --model opus. The documentation quotes come from https://code.claude.com/docs/en/skills.md, settings-reference.md, env-vars.md and model-config.md, all fetched that day between 16:20 and 16:21 UTC.

What the docs say about the listing

The skills page covers the listing in its troubleshooting section, under "Skill descriptions are cut short". These are the sentences we tested, quoted as fetched:

  • "The listing always contains every skill name, but if you have many skills, Claude Code drops some descriptions to fit the listing's character budget, which removes the keywords Claude needs to match your request. The budget scales at 1% of the model's context window. When the listing overflows, Claude Code drops descriptions starting with the skills you invoke least, so the skills you use most keep their full text."
  • "When the listing exceeds its budget, Claude Code also writes a warning to the debug log, visible with --debug."
  • "To raise the budget, set the skillListingBudgetFraction setting (for example, 0.02 = 2%) or the SLASH_COMMAND_TOOL_CHAR_BUDGET environment variable to a fixed character count. To free budget for other skills, set low-priority entries to "name-only" in skillOverrides so they list without a description. You can also trim the description and when_to_use text at the source: put the key use case first, since each entry's combined text is capped at 1,536 characters regardless of budget. The cap is configurable with skillListingMaxDescChars."

The settings reference adds what a dropped description costs. Under skillListingBudgetFraction: "When the listing is over the cap, Claude Code keeps every skill's name but drops the descriptions of the least-used skills, so Claude can still invoke those skills but is less likely to choose one on its own." Its default is given as "0.01, which reserves 1% of the context window", and skillListingMaxDescChars defaults to 1536. The environment variable reference describes SLASH_COMMAND_TOOL_CHAR_BUDGET this way: "Override the character budget for skill metadata shown to the Skill tool. The budget scales dynamically at 1% of the context window, with a fallback of 8,000 characters. Legacy name kept for backwards compatibility".

The changelog has some history behind those numbers. Cut at its own heading, the 2.1.32 section says "Skill character budget now scales with context window (2% of context), so users with larger context windows can see more skill descriptions without truncation", and the 2.1.105 section says "raised the listing cap from 250 to 1,536 characters". The docs and the default setting now say 1%; we found no changelog line announcing a move from 2% to 1%.

OpenAI's Codex documents a similar budget. Its skills page at https://developers.openai.com/codex/skills (fetched 2026-10-05 16:21 UTC) says: "In Codex, the initial list also includes each skill's file path. To avoid crowding out the rest of the prompt, this list uses at most 2% of the model's context window, or 8,000 characters when the context window is unknown. If many skills are installed, Codex shortens skill descriptions first. For large skill sets, Codex may omit some skills from the initial list and show a warning." we measured only Claude Code, so everything we say about Codex in this article comes from that page.

What neither product's page tells you is how many characters "1% of the context window" comes to, given that the window is counted in tokens. That was the first thing the lab had to answer.

The lab

The lab is a set of fresh directories, each holding nothing but .claude/skills/. A small Python script writes N skills named lab-sk-001 to lab-sk-N. Each SKILL.md has a name, a 1,000-character description and a one-line body:

---
name: lab-sk-001
description: "CW001A Recalibrates the amber ledger for the build. Use when the user asks to recalibrate the amber ledger, mentions a drifted amber ledger, or wants the amber ledger checked before a release. It reads the files the user points at, checks them against the team's written conventions, [filler repeated up to 1,000 characters] CW001Z"
---

When this skill runs, reply with the single token BODY001 and nothing else.
Enter fullscreen mode Exit fullscreen mode

Every description starts and ends with a codeword of its own (CW001A and CW001Z), so we could tell from the model's answer whether a description arrived whole, arrived cut short, or didn't arrive at all. Every skill also gets its own trigger phrase ("amber ledger", "basalt ledger" and so on), built from twenty adjectives and eight nouns. One skill, lab-sk-002, gets a 2,000-character description to test the 1,536-character cap. We built the set at 10, 40, 80 and 160 skills.

Each listing run sent the same prompt:

Do not use any tools. Look at everything in your context for this turn, including the list of
available skills. List every token that begins with the two letters CW followed by three digits
and one letter, in the order they appear, separated by single spaces. If there are none, reply
NONE. Reply with that single line only.
Enter fullscreen mode Exit fullscreen mode

And each run was one invocation:

claude -p "$PROMPT" --setting-sources project,local --strict-mcp-config \
  --tools Skill --max-turns 1 --model opus --output-format stream-json --verbose \
  --session-id "$SID" --debug-file "runs/$NAME/debug.txt" < /dev/null
Enter fullscreen mode Exit fullscreen mode

--setting-sources project,local keeps our user settings, and the plugins they enable, out of the runs. --strict-mcp-config with no config file means no MCP servers. We also unset the environment variables that the Claude Code session we were working from exports (CLAUDECODE, CLAUDE_CODE_SESSION_ID, CLAUDE_EFFORT, CLAUDE_CODE_AUTO_COMPACT_WINDOW and about ten more), so the child runs did not inherit them.

--tools Skill matters. Our first pilot used --tools "" to keep the model from reading files, and the model answered NONE. That transcript had no skill listing at all, and the first request was only 2,748 tokens. The listing is attached only when the Skill tool is available, so every later run allows that one tool and nothing else.

The evidence for each run comes from three places. The first is the transcript at ~/.claude/projects/<project>/<session-id>.jsonl, where the listing is an attachment record of type skill_listing. Its content field holds the listing text, skillCount and names give the count and order of entries, and its rendered field holds what was sent to the model: the same text inside a system reminder that begins "The following skills are available for use with the Skill tool:". The second is the debug log, which states the budget whenever the listing is over it. The third is the model's reply, which we checked against the codewords in the listing, plus the size of the first request (input_tokens + cache_creation_input_tokens + cache_read_input_tokens from the first assistant message's usage).

On the context window: the result event's modelUsage reported "contextWindow": 1000000 for claude-opus-5-5 in every run that used the default settings, and the model configuration page says that "On the Anthropic API, Fable 5.1, Fable 5, Sonnet 5 and later, and Opus 4.7 and later run with the 1M window on every plan". With CLAUDE_CODE_DISABLE_1M_CONTEXT=1, which the environment variable reference says holds models with a native 1M window to 200K, the same field read 200000.

One mistake to own. A quoting bug in our shell loop sent nine runs, each meant for a different condition, into the lab's root directory, which has no skills. They became nine runs of the zero-skill baseline, which are in the table below, and we reran the intended conditions.

The listing runs in one table

Condition Runs Entries Project skills with a description Listing sent, characters Budget named in the debug log First request, tokens
no Skill tool (--tools "") 1 — — no listing — 2,748
0 project skills 9 14 — 6,629 no warning 6,151
10 skills 2 24 10 of 10 17,313 no warning 9,153
40 skills 2 54 22 of 40 29,723 30,000 12,741
80 skills 2 94 21 of 80 29,242 30,000 12,825
160 skills 2 174 20 of 160 29,280 30,000 13,272
160, after invoking three once 2 174 20 of 160 29,280 30,000 13,277
160, skillListingBudgetFraction: 0.02 2 174 50 of 160 59,326 60,000 21,554
160, SLASH_COMMAND_TOOL_CHAR_BUDGET=12000 2 174 3 of 160 11,714 12,000 8,434
160, CLAUDE_CODE_DISABLE_1M_CONTEXT=1 2 174 0 of 160 8,575 6,000 7,562

The 14 extra entries in every row are Claude Code's own: on this install, dataviz, update-config, keybindings-help, code-review, simplify, fewer-permission-prompts, loop, schedule, claude-api, workflow-authoring, run, plugin-authoring, init and security-review. They always came after the project skills in the listing. The two runs of each condition produced byte-identical listings, all nine baseline runs produced the same one, and in all 25 runs the codewords in the model's reply matched the codewords in the recorded listing exactly, in the same order.

The budget is 30,000 characters, not 40,000

The debug log names the budget whenever the listing is over it. At 160 skills it read:

[WARN] Skill listing over budget: 174 skills, 169503 chars > 30000 budget — descriptions will be truncated. Run /skills to disable some, or raise skillListingBudgetFraction in settings.
Enter fullscreen mode Exit fullscreen mode

169,503 is the size the listing would have reached if every description had gone in; 29,280 is what Claude Code actually sent. All four budgets we saw fit one formula: window × fraction × 3 characters. That gave 30,000 at 1M with the default 0.01, 60,000 at 1M with 0.02, and 6,000 at 200K. SLASH_COMMAND_TOOL_CHAR_BUDGET=12000 replaced the computed value with exactly 12,000. The 8,000-character fallback in the environment variable reference did not appear in any of our runs. The page calls it a fallback, presumably for a window Claude Code can't determine, and every one of our runs had a known window.

In tokens, the 1% is about right. Going from the 14-entry baseline to 160 skills added 22,651 characters to the listing and 7,121 tokens to the first request, about 3.2 characters per token. That makes a 30,000-character listing roughly 9,400 tokens, a little under 1% of a 1M window. But if you turn 1% of a million tokens into characters with the usual four characters per token, you get 40,000 and overestimate your room by a third.

The debug log line for 160 skills: Skill listing over budget, 174 skills, 169503 chars against a 30000 budget, descriptions will be truncated; the listing actually sent was 29,280 characters with 20 of 160 skills described

The listing itself is capped, not just measured. In every over-budget run on the 1M window, the sent listing landed within a few hundred characters under the budget (29,723, 29,242, 29,280, 59,326 and 11,714). The exception is the 200K case, which comes up below.

What gets dropped: whole descriptions, never names

The page says Claude Code "drops some descriptions", and that is the mechanism we saw. Descriptions weren't shortened. In every over-budget run, each project skill appeared in one of two forms: with its full description, carrying both codewords, or as a bare 12-character line such as - lab-sk-021. No description was cut partway to make it fit. Every name was present in every run: in all 37 runs that had a listing, the entries in content, the names array and skillCount agreed.

With no usage history, the skills that kept descriptions were the first ones in name order: lab-sk-001 through lab-sk-022 at 40 skills, through 021 at 80, and through 020 at 160. The count falls as the set grows because a bare name still costs about 13 characters with its line break: 160 of them take about 2,000 characters, which is two descriptions' worth.

Claude Code's own entries were treated differently. Twelve of the fourteen kept their descriptions in all 37 runs that had a listing, including four that had no usage record on this machine at all (keybindings-help, fewer-permission-prompts, workflow-authoring and plugin-authoring). The 200K runs show how far that goes. Those twelve descriptions take 6,297 characters, more than the whole 6,000-character budget. Claude Code kept them anyway, kept every name, dropped every project description, and sent a listing of 8,575 characters, 43% over the budget it had just logged. The other two built-ins, init and security-review, behaved like project skills. They kept their descriptions where there was room left over and lost them in the 200K runs and in our second trigger lab, described below, where 22 project descriptions used up the space. By that arithmetic, a session with a 200K window on this install has no room for any project skill description, however few skills you have; we measured it with 160. The model configuration page mentions "Opus 4.8 and later when they run with a 200K context window, such as on Amazon Bedrock, Google Cloud's Agent Platform, and Microsoft Foundry", and CLAUDE_CODE_DISABLE_1M_CONTEXT=1 holds any session to 200K. We measured only the second.

The 1,536-character cap matched the docs too. lab-sk-002's 2,000-character description arrived as its first 1,535 characters followed by …, 1,536 in all. Its closing codeword sat at offset 1,994 and never arrived. In the 12 listing runs where lab-sk-002 had a description, the model listed CW002A and not CW002Z all 12 times.

One invocation was enough to move a skill up

To test "the skills you invoke least", we ran one session in the 160-skill project that told the model to invoke lab-sk-150, lab-sk-155 and lab-sk-160 with the Skill tool. All three were listed by name only at that point. All three calls went through, each tool result reading Launching skill: lab-sk-150 and so on, so a skill without a description can still be invoked by name. Afterwards ~/.claude.json held a skillUsage entry for each of the three, usageCount: 1.

In the next two listing runs, all three had their descriptions, and lab-sk-018, 019 and 020, the last three never-invoked skills in name order, had lost theirs. It was a one-for-one swap, 2 of 2 runs. The history is also per machine, not per project. skillUsage is keyed by skill name in the global config file, and our skillListingBudgetFraction: 0.02 runs, in a separate directory with a copy of the same skills, still put 150, 155 and 160 among their 50 described skills. The trigger tests below show the same thing from the other direction: a single invocation of lab-sk-018 gave it back its description in the very next run and pushed out lab-sk-017.

This ranking has a consequence that is easy to miss. The skill most likely to lose its description is one nobody has invoked yet and whose name sorts late, and that describes a skill you just wrote and want the model to discover on its own.

What a missing description does to triggering

The listing runs show what the model can see. To see what it does with it, we sent natural requests that matched exactly one skill's trigger phrase, again with only the Skill tool and --max-turns 4: "The garnet ledger drifted again. Please recalibrate the garnet ledger before the release."

lab-sk-007 (garnet ledger) had its description, and the model invoked it on the first call in 3 of 3 runs, replying BODY007.

The name-only test went wrong in an instructive way. Our target was lab-sk-018 (russet ledger), which had just lost its description. The model invoked it anyway and wrote, before the call: "The skill descriptions follow an alphabetical pattern (001 amber … 017 quartz), so the russet ledger is most likely lab-sk-018. Its description is cut off in the listing, so I'll load it and check that it's the right one before I follow it." Our adjectives ran in alphabetical order, the descriptions of 001 to 017 gave that away, and the model extrapolated. That invocation also gave lab-sk-018 a usage record, so it had its description back in the next two runs, where the model invoked it 2 of 2 times. Those two runs count as described, not name-only.

So we built a second 160-skill set, named tq-001 to tq-160, with the 160 trigger phrases shuffled by a fixed random seed so that no skill's phrase could be inferred from its neighbors. Every name was fresh, with no usage history. In that set, 22 skills kept their descriptions (tq-001 to tq-022). We picked three name-only targets (tq-061 rustic anvil, tq-097 ivory ferry, tq-133 golden hopper) and two described ones (tq-004 misty beacon, tq-011 quiet cistern), one run each.

The described targets were invoked on the first call, 2 of 2. The name-only targets were invoked 0 of 3 times. In all three the model made no tool call at all, said it could not find the skill, and asked us for the skill's name. Each time it picked out the described skills with similar phrases and said that none of them was the target:

The model's reply when its target skill was listed by name only: skills tq-023 through tq-160 show only their names, with no descriptions; one of them may be the golden hopper skill, but it would have to guess

The ivory ferry reply ran along the same lines: "The closest are tq-013 (lunar ferry) and tq-009 (ivory beacon), and both are for different parts." It added that "Running them at random to find it would mean following instructions written for other parts, so I'd rather not before a release." So a skill that has lost its description isn't triggered by a near match or by trial and error. It just isn't triggered, and the model says so. That is the better of the possible failures, but in an unattended run nobody is there to answer the question.

Taken together: with descriptions, 7 of 7 requests went to the right skill on the first call (3 for lab-sk-007, 2 for lab-sk-018 once it was described again, 2 in the shuffled set). With only an opaque name, 0 of 3. With only a name plus a guessable pattern, 1 of 1. Our names were deliberately meaningless, so this is the worst case. A skill named after what it does still gives the model something to match after its description is gone, but we didn't test that.

What it costs per turn

The budget does its job on cost. The first request grew from 6,151 tokens with no project skills to 9,153 with 10, then flattened: 12,741 at 40, 12,825 at 80 and 13,272 at 160, for a set whose full listing would have run 169,503 characters. Each extra name-only skill cost a handful of tokens. Raising skillListingBudgetFraction to 0.02 put 30 more descriptions in front of the model and added 8,282 tokens to the first request. That is the trade the settings page describes ("at the cost of more context per turn"), and the extra text sits in the context of every request in the session. In dollars, runs cost between $0.0016 and $0.17; the first run of each configuration paid for writing the prompt cache, and identical repeats cost under two cents.

What we would do with these numbers

Work out your room before you add skills. On this install, at the default budget, Claude Code's own entries took 6,629 characters, which left 23,371 for project skills: about 13 characters per skill for the name line, plus the description and 2 more for each skill that keeps one. If every skill had a description of length D and a name of about ten characters, roughly 23,371 / (D + 15) would fit with descriptions. That is about 15 at the 1,536-character cap, 23 at 1,000, 45 at 500 and 108 at 200. That is arithmetic from the measured budget, not a measurement. The 22, 21 and 20 that we measured at 1,000 characters fit it once the name-only lines and the one 1,536-character description are counted. Plugin skills and user-level skills would join the same listing; our runs had neither.

In our runs the order was twelve of the built-ins first, then skills that had been invoked, then the rest in name order. If a skill must be found by description, make sure it has been used, or keep the set small enough that it never needs to be. Put the trigger words in the name too, because the name is the only part that always stays. For skills you only ever start yourself with /name, the docs offer "name-only" in skillOverrides or disable-model-invocation: true, which hands their space to the skills the model has to find on its own. We didn't run either.

If your sessions use a 200K window, check the listing before anything else. On this install the built-in descriptions alone fill the 6,000-character budget, so the model sees none of yours until you raise skillListingBudgetFraction or set SLASH_COMMAND_TOOL_CHAR_BUDGET. At 12,000 characters our 160-skill set got 3 descriptions, the three we had invoked.

Keep what matters in the first 1,535 characters of description plus when_to_use. And run with --debug or --debug-file once after adding skills. The Skill listing over budget line names both numbers, and in a headless run it is the only warning you get. Nothing in the stream-json output or the result said a listing had been cut.

Codex, by its documentation

Going only by the page quoted above, Codex differs in three ways. Its budget is 2% of the context window instead of 1%. It "shortens skill descriptions first" where Claude Code, in our runs, dropped whole descriptions and never shortened one to fit the budget (the 1,536-character cap is a separate, per-skill limit). And it "may omit some skills from the initial list", where Claude Code kept every name in all 37 of our runs that had a listing. Both products mention an 8,000-character figure: Codex for an unknown window, Claude Code as its "fallback". We ran no Codex sessions, so we can't say how its shortening looks in practice.

What we did not measure

Every run was headless. We didn't open /context, /doctor or /skill-doctor, which the docs say report the listing's cost. We didn't test skillListingMaxDescChars, "name-only" in skillOverrides, disable-model-invocation, or descriptions split between description and when_to_use. We didn't load plugin skills or user-level skills, and we didn't check whether CLAUDE_CODE_AUTO_COMPACT_WINDOW changes the budget. We never produced the "unknown window" case that would use the 8,000-character fallback. We gave each skill at most a few invocations, so we can't say whether the ranking weighs count or recency. We didn't test descriptive skill names after their descriptions were dropped, other models, a session where skills change mid-way, or Windows. The trigger tests are small and used one phrasing per skill: 7 described requests, 3 name-only requests in the clean set, and 1 in the guessable set.

Reproduce it

The generator, trimmed:

def description(i, phrase, length=1000):
    head, tail = f"CW{i:03d}A", f"CW{i:03d}Z"
    body = (f"{head} Recalibrates the {phrase} for the build. Use when the user asks to "
            f"recalibrate the {phrase}, mentions a drifted {phrase}, or wants the {phrase} "
            f"checked before a release. ")
    while len(body) + len(tail) + 1 < length:
        body += FILLER
    return body[: length - len(tail) - 1].rstrip() + " " + tail

for i in range(1, n + 1):
    d = f"{dest}/.claude/skills/lab-sk-{i:03d}"
    os.makedirs(d)
    open(f"{d}/SKILL.md", "w").write(
        f"---\nname: lab-sk-{i:03d}\ndescription: \"{description(i, phrase(i))}\"\n---\n\n"
        f"When this skill runs, reply with the single token BODY{i:03d} and nothing else.\n")
Enter fullscreen mode Exit fullscreen mode

Run the listing prompt with the command line shown earlier from inside the project directory, then read the listing, the name-only count and the first-request size from the transcript:

T=~/.claude/projects/<project>/<session-id>.jsonl
jq -r 'select(.attachment.type=="skill_listing") | .attachment.content' "$T" | wc -c
jq -r 'select(.attachment.type=="skill_listing") | .attachment.content' "$T" | grep -c '^- [^:]*$'
jq 'select(.type=="assistant") | .message.usage
    | .input_tokens + .cache_creation_input_tokens + .cache_read_input_tokens' "$T" | head -1
grep 'over budget' runs/<name>/debug.txt
Enter fullscreen mode Exit fullscreen mode

For the other conditions, put {"skillListingBudgetFraction": 0.02} in the project's .claude/settings.json, or prefix the run with SLASH_COMMAND_TOOL_CHAR_BUDGET=12000 or CLAUDE_CODE_DISABLE_1M_CONTEXT=1. Skill usage lives in ~/.claude.json under skillUsage and is shared by every project on the machine, so give each experiment fresh skill names. Otherwise one run's invocations change the next run's listing, as they did in ours.

Numbers, for the record

There were 39 claude -p invocations on 2026-10-05 between 16:23 and 16:32 UTC, with Claude Code 2.1.289 on macOS and claude-opus-5-5. Thirty-eight reached the model; the first exited before starting because our script passed it an empty prompt. The 38 were 25 listing runs (the table above, nine of them the accidental baseline), one pilot without the Skill tool, one usage run and 11 trigger runs. The reported cost for all 38 was $1.75. All 25 listing replies matched their recorded listings codeword for codeword.


This lab is one generator script, a shell loop and a few jq lines, so it is quick to rerun whenever your skill count or your model's context window changes.

If your install names a different budget in its Skill listing over budget line, or keeps a different number of descriptions, post the numbers and your Claude Code version in the comments below.

Top comments (0)