DEV Community

Sungwoo Lee
Sungwoo Lee

Posted on

Claude Code Skills: 96% of My Command's Output Was Thinking

I have a custom Claude Code command called /his. At the end of a work session it writes a short entry into the project's HISTORY.md: what was decided, why, what was tried and abandoned, and what the next session should pick up. It's the one piece of memory that survives /clear and context compaction, so I run it a lot.

It also felt slow and expensive for what it produced. So I measured it.

The measurement

Claude Code keeps a JSONL transcript of every session, and each assistant turn carries a usage block. I wrote a small script that finds every /his invocation in a project's transcripts and adds up what happened between the command and the next user turn.

One trap before the numbers: the transcript records the same API request two or three times. If you sum usage naively you double- or triple-count. Deduplicate by requestId first.

Across six runs in one project:

Tokens
Total output 138,701
Thinking 132,721
Share that was thinking 96%
Actual document written per run ~1,000

The command wrote about a thousand tokens of history per run. Everything else was the model reasoning about how to write it.

Why a command file makes the model think

Across six runs the command produced 138,701 output tokens, of which 132,721 (96%) were thinking, while each run wrote only about 1,000 tokens of history

The instruction file for /his had grown to 16.6 KB. None of it was wrong. Each rule had been added after a real mistake: check the file size, move old entries into an index, don't duplicate the git log, verify links, handle the case where the session was just cleared, and so on.

But every branch in an instruction file is a decision the model makes again, from scratch, on every run. "If the history file is over 8 KB, compress the third-oldest entry into a one-line index" is trivial for a script. For a model it means reading the file, estimating its size, finding the entries, deciding which one is third-oldest, deciding what one line should say, and then second-guessing all of that in its reasoning. Multiply by forty such rules.

A skill or command body is loaded into context when you invoke it. After that, its size isn't the main cost. The main cost is how many judgment calls it asks for.

What I changed

I split the work by a single question: does this step need judgment, or just execution?

Everything mechanical moved into two Python scripts:

  • his_prep.py <slug> collects the facts (changed files, commits, git state) into a separate file and creates a skeleton entry with four empty sections. It also refuses to run in two situations and exits with code 2. The first is when the context is close to the compaction threshold, because a compaction mid-write erodes exactly the reasoning I'm trying to save. The second is when the session is effectively empty right after /clear, since there's no "why" left in context to record.
  • his_finish.py <slug> refuses to finish if any of the four sections is blank. Then it does the bookkeeping: it inserts the new entry, demotes the previous one, compresses the one before that into a one-line index, and checks links, file size and uncommitted changes.

The command file itself went from 16.6 KB to 5.3 KB, and most of what's left explains what not to do. The model's job is now exactly one thing: fill in four sections.

  1. Key decisions and why
  2. Alternatives rejected and why
  3. Approaches that failed (or "none")
  4. What the next session should do first

It no longer copies file lists or commit hashes into the entry, because the facts file already has them and the skeleton links to it. Copying those was a large part of the old 96%.

Two mistakes I made along the way

The /his instruction file shrank from 16.6 KB to 5.3 KB after mechanical steps moved into two Python scripts

Automating the "why" killed it. An earlier version pre-filled the skeleton with plausible text generated from the diff. It looked complete, so the real reasoning never got written. I rolled that back. The rule now: the script can fill in facts, but only the model that just did the work writes the reasons, even if that's one honest line.

Copies drift. I had copied the scripts into three projects. When I fixed the measurement in one, the other two kept using the broken version, and nobody noticed. The tools now live in one global folder and the command calls them by absolute path. If the script is missing, the command falls back to a legacy manual procedure kept in a separate file, which is loaded only when it's needed.

What I haven't measured

I haven't re-run the same six-run measurement after the change with the same rigor, so I won't quote an "after" percentage. What I can say is structural: the model used to make dozens of formatting and bookkeeping decisions per run, and now it makes four writing decisions, and the scripts reject the output if any of them are skipped.

The rule I took away

If you're writing a Claude Code skill or command, go through it line by line and ask whether each line needs judgment. If it doesn't, turn it into a script and let the skill call that script. Keep the skill for the parts only the model can do. A long skill isn't mainly a context cost. It's the reasoning cost of making the same decisions again on every run.

I write about running Claude Code day to day, mostly in Korean, at my-blog.org/claude-code.


Get new posts by email — subscribe on my-blog.org.


Comparing coding assistants? Copilot, Cursor, Claude Code & ChatGPT compared.

Top comments (1)

Collapse
 
promptalo profile image
PromptAlo •

Automating the "why" from the diff made the skeleton look complete, so the real reasoning never got written. Keeping the four sections empty until the model that just did the work fills them is what makes HISTORY.md usable next session. Are empty sections rejected by the scripts, or only noticed when the next session opens the file?