DEV Community

Cover image for Caveman vs Ponytail vs Chisle: I benchmarked the Claude Code token-saving plugins on 20 tasks
Jay Pokale
Jay Pokale

Posted on

Caveman vs Ponytail vs Chisle: I benchmarked the Claude Code token-saving plugins on 20 tasks

If you use Claude Code, you've probably seen the two popular plugins that promise to cut your token bill: caveman, which makes Claude talk like a caveman, and ponytail, which pushes it to write less code. I built a third one, Chisle, and benchmarked all three against Claude with no plugin at all: 20 live tasks, 59+ model runs on Haiku and Sonnet, with the arms differing only in the injected ruleset.

TL;DR: caveman cut the bill to 80%, ponytail to 68%, and Chisle to 52%, nearly half. Chisle also had the smallest worst day (173% vs 424% and 227%) and backfired once in 20 tasks, where caveman backfired 6 times and ponytail 8.

Total billed output across 20 tasks as percent of the no-plugin baseline: caveman 80% (worst day 424%), ponytail 68% (worst day 227%), Chisle 52% (worst day 173%)

20 live tasks, billed output vs no plugin total bill average task worst case backfires
no plugin (baseline) 100% 100%
caveman 80% 98% 424% 6 / 20
ponytail 68% 91% 227% 8 / 20
Chisle 52% 69% 173% 1 / 20

In the July re-verification run, every answer from every arm was graded correct. None of these tools buys its savings with wrong answers. Every number comes from committed raw transcripts in the Chisle repo.

Where the savings come from: coding vs explanation

tasks caveman ponytail Chisle
coding (wants working code) 12 74% 59% 44%
explanation (wants prose) 8 103% 104% 87%

On coding, Chisle bills 44% of a bare model, a third less than ponytail, the closest thing to a dedicated "lazy code" tool. On explanation prompts both specialists go above 100%: tools built to write less made Claude write more than using nothing. Chisle is the only one that stays under.

Split by answer length, the gap widens: on long answers Chisle bills 45%, caveman 79%, ponytail 59%.

Caveman alternative: great prose compressor, no engineering judgment

caveman is genuinely good at compressing prose, and on some short prose prompts it's a hair leaner than Chisle. But it has no judgment about what to build. Asked to "add caching", it produced three implementations (330 tokens). Chisle gave one @cache decorator and a one-line upgrade path (151 tokens). Its worst day cost 4.2× a bare model.

Ponytail alternative: right instinct on code, padded prose

ponytail has the right instinct: smallest thing that works. But it pads prose so much that it backfires: on a "retry logic" prompt it ran 227% of baseline. A tool whose whole job is writing less wrote more than twice as much.

Installing both to cover both axes doesn't fix it either: they fight over prose style and double per-session overhead. On one task the pair did worse (605 tokens) than Chisle alone (595).

Chisle compresses three things; caveman and ponytail compress one each

prose code judgment input / context publishes failures
caveman ✅ ❌ ❌ ❌
ponytail ❌ ✅ ❌ ❌
Chisle ✅ ✅ ✅ ✅
  1. Output prose: no filler, no hedging, no manufactured structure.
  2. Output code: a YAGNI "efficiency ladder" (use what exists, ask instead of guessing, skip speculative abstractions). Safety, error handling, validation and accessibility are never cut.
  3. Input context, the axis neither rival touches. Across 171 real Claude Code sessions, tool output was 67.5% of the context window, re-billed on every later request. Chisle's PostToolUse hook trims oversized tool output by ~46% before it re-enters context (Read/Edit/Write are never touched).

Same prompt, same model

"Add debounce to a search input that currently fires an API call on every keystroke." Verbatim committed output:

  • No plugin: 142 lines, 1,506 tokens. A generic useDebounce<T> hook in its own file, then Option 2, Option 3, a comparison table and caveats.
  • Chisle: 35 lines, 602 tokens. setTimeout in the effect you already have, two lines on why, and "use lodash.debounce if it's already installed."
useEffect(() => {
  const timer = setTimeout(async () => {
    if (query.trim()) { /* fetch */ }
  }, 300);
  return () => clearTimeout(timer);
}, [query]);
Enter fullscreen mode Exit fullscreen mode

Not golfed, just boring: one less file, one less abstraction, same behaviour.

The fine print (because the repo publishes it)

  • The 20-task total includes one cache task where the bare model wrote an unusually long answer (4,910 tokens, against 375–813 in later runs), and a few cells that an older ruleset example may have primed. Without those cells, Chisle's 20-task total is 70%.
  • A later, smaller rerun (13 prompts × 2 seeds on the newest plugin versions) again put Chisle lowest: 83%, vs caveman 102% and ponytail 105%. It was the only one of the three below a bare model.
  • On short answers, every tool breaks even or worse; there's little to cut in a three-line reply.

Chisle is the only tool in this class that publishes the runs where it lost: benchmarks/results.

Install Chisle in Claude Code

npx chisle             # installs for Claude Code and any other agents it finds
npx chisle --dry-run   # preview first
Enter fullscreen mode Exit fullscreen mode

Zero dependencies, MIT licensed. The same ruleset ships to Cursor, Codex, Gemini CLI, GitHub Copilot, Windsurf, Cline, OpenCode, Kiro, Antigravity, Hermes and Pi. Claude Code and Pi also get the input-side compressor and live modes (lite, full, ultra). /chisle-audit flags over-engineered code and bloated prose/docs in one ranked report; ponytail's audit is code-only and caveman has none.

FAQ

Is Chisle better than caveman?
Across 20 tasks: 52% vs 80% of the bare-model bill, worst case 173% vs 424%, 1 vs 6 backfires. On coding prompts 44% vs 74%. caveman is a little leaner on some short prose prompts.

Is Chisle better than ponytail?
Across 20 tasks: 52% vs 68%, worst case 173% vs 227%, 1 vs 8 backfires. On explanation prompts ponytail goes above 100% (104%); Chisle stays at 87%.

Can I install caveman and ponytail together instead?
You can, but they fight over prose style and double the overhead. On one task the pair did worse (605 tokens) than Chisle alone (595).

Does it make answers wrong?
In the July re-verification every answer from every arm graded correct. The rule is necessary, not fewest characters, and safety is never cut.

Where's the raw data?
All transcripts are committed: benchmarks/results. Full tables: docs/benchmarks.md.

GitHub logo JayPokale / Chisle

Cut your AI coding agent's token bill on three axes: terse prose, YAGNI-first code, and tool-output compression. Claude Code, Pi, Cursor, Codex, Gemini + 4 more. Zero deps, published benchmarks including the runs it loses.

Chisle

Chisle

Your AI talks less, builds less, reads less, and says more. Like a senior dev who bills by the syllable.

The only tool in this class that publishes the runs where it lost. Here's why.

npm version Works with 12 agents CI Zero deps MIT Star Chisle on GitHub

Built for Claude Code: coding answers come back 33% shorter and 24% cheaper, while caveman and ponytail make them longer · 12 agents · zero dependencies · one command

Chisle is a Claude Code plugin that makes Claude cheaper to run without making it dumber. It cuts what Claude writes: no filler, no hedging, no speculative abstractions, just the smallest code that works. It also cuts what Claude reads: a PostToolUse hook trims oversized tool output by ~46% before it re-enters the context window, where it would be re-billed on every later request. In agent-loop tests Claude with Chisle passes the same tasks as Claude without it. One npx chisle installs it, and the…




Site: chisle.jaypokale.me. If you run your own comparison against caveman or ponytail, I'd like to see it, especially where Chisle loses.

Top comments (0)