If you use an AI coding agent every day, you have probably noticed how much of its output is padding. Phrases like "The reason this is happening is likely because..." or "I'd recommend that you consider..." appear before the useful part of the answer. You pay for those tokens, and you spend time reading them.
Caveman is an open source project that fixes this. Its tagline is "why use many token when few do trick", and it does what it says: your agent answers in short, blunt sentences without the filler.
What it does
Here is the example from the README. The question is why a React component keeps re-rendering.
Normal agent (69 tokens):
The reason your React component is re-rendering is likely because you're creating a new object reference on each render cycle. When you pass an inline object as a prop, React's shallow comparison sees it as a different object every time, which triggers a re-render. I'd recommend using useMemo to memoize the object.
Caveman agent (19 tokens):
New object ref each render. Inline object prop = new ref = re-render. Wrap in
useMemo.
The diagnosis and the fix are the same. Only the throat-clearing is gone.
Code, commands, file paths, and exact error messages are never rewritten. Only the prose around them gets shortened, so the answers stay accurate and copy-pasteable.
Install in one command
The simplest version is a skill, which is a rule file your agent loads:
npx skills add JuliusBrussee/caveman
It works with 30+ agents, including Claude Code, Codex, Gemini CLI, Cursor, Windsurf, Cline, and Copilot. If your agent does not pick it up automatically, type /caveman in the chat.
Intensity levels and commands
You control how terse the agent gets:
/caveman lite
/caveman full
/caveman ultra
/caveman off
The skill also adds a few useful commands:
-
/caveman-commitwrites short Conventional Commit messages. -
/caveman-reviewgives one-line, actionable code review comments. -
/caveman-compress <file>shrinks Markdown memory files and keeps a backup of the original. -
/caveman-statsshows local token usage and estimated savings in Claude Code.
The numbers
The author ran ten ordinary coding prompts through the Claude API, with and without the skill. Output tokens dropped by 65% on average. The range is wide, though:
| Task | Normal | Caveman | Saved |
|---|---|---|---|
| Implement React error boundary | 3454 | 456 | 87% |
| Set up PostgreSQL connection pool | 2347 | 380 | 84% |
| Explain git rebase vs merge | 702 | 292 | 58% |
| Refactor callback to async/await | 387 | 301 | 22% |
The pattern is simple. When the agent would have written an essay, you save a lot. When the answer was mostly code, you save very little.
Read this before quoting the 65%
The project is unusually honest about its limits, and you should be too:
- The skill only shortens output. Input and reasoning tokens do not change.
- The skill's own rules add roughly 1 to 1.5k input tokens on every turn.
- Whole-session savings will be lower than the benchmark table, and on work that was already terse you can end up paying more.
The README says speed and readability are the real product, and the discount is a bonus. Measure your own setup before promising anyone a 65% cut on the bill.
The optional proxy
There is also a second, bigger piece: a local proxy that sits between your agent and the AI provider. It shrinks what the agent reads, such as logs, diffs, and test output, before each call, and it keeps a backup of every original on your machine.
npm install -g @caveman-ai/cli && caveman setup --install
caveman claude
In a 54-run Claude Code benchmark, total input tokens dropped by about 33%. One case, an HTML dashboard, actually got 9.9% worse, and the README leaves that row in the table. It is a bigger commitment than the skill, so start with the skill.
Things to know before you install
-
Telemetry: the
cavemanCLI sends anonymous usage stats by default (commands run and token counts, never prompts or code). Turn it off withcaveman telemetry offorDO_NOT_TRACK=1. The skill itself runs entirely on your machine. - Licensing: the skill and CLI are MIT. The engine and proxy are BSL-1.1, which is source-available but not OSI open source. Self-hosting for your own use is fine.
Is it worth trying?
Yes, with realistic expectations. The skill is a one-line install, easy to turn off, and does not touch your code. If your agent is wordy, you get faster, easier-to-scan answers right away. Treat any token savings as a bonus and check them with /caveman-stats.
Top comments (0)