If you use Cursor, Claude Desktop, or Cline daily,
you have probably watched your rate limits evaporate
because of a single runtime crash.
A Next.js build fails or a Python script panics, and
your terminal vomits 500 lines of stack traces. The
LLM eagerly ingests all 45,000 tokens of internal
node_modules machinery, webpack bundles, and event
loop frames.
You just paid $0.15 for an AI model to read code it
can't edit, your prompt cache is wiped, and your agent
is now hallucinating because its context window is
full of garbage.
Worse, your terminal stderr probably just leaked
DATABASE_URL=postgres://admin:password@... straight
to an external API.
I got tired of paying for framework noise, so I
built an open-source Rust Model Context Protocol (MCP)
server called Tokenectomy Razor to fix it locally
before logs ever touch the LLM.
## What is eating your context window?
A typical Next.js error trace looks like this:
text
TypeError: Cannot read properties of undefined
(reading 'digest')
at Object.<anon> (/node_modules/next/bundle5.
js:142:31)
at __webpack_require__
(/node_modules/next/bundle5.js:198:12)
at Object.execute (/node_modules/next/dev-
server.js:412:19)
at processTicksAndRejections (task_queues:95:5)
Database connection failed:
postgresql://admin:super_secret_password@db.prod.
internal:5432/primary
API key leaked: sk-ant-api03-
abcdef1234567890abcdef1234567890
[... 480 internal dependency frames flooding
context ...]
Notice three things:
1. 98% of those frames are inside node_modules. Your
AI agent is not going to edit webpack's internal
bundle logic. It only cares about the one line in
src/components/Header.tsx:42 where you missed a
parenthesis.
2. The database password and API key are sitting
unmasked in plain text.
3. The raw token count for this single error dump was
45,820 tokens.
## What happens after Tokenectomy runs
When the AI agent invokes get_error_context via MCP,
Tokenectomy intercepts the log, strips framework
internals, redacts all secrets locally using a
deterministic DFA regex, and grabs bounded source code
lines around the actual crash:
[:TOKENECTOMY:M2M_CONTROL_PLANE:v1.3.0]
[STATE=FRAMEWORK_NOISE_PURGED]
[STRATEGY_APPLIED=AGGRESSIVE]
[ORIGINAL_BYTES=45820 | CLEAN_BYTES=118 |
REDUCTION=99%]
[PRIMARY_CRASH_COORDINATES=src/components/Header.
tsx:42]
[COGNITIVE_DIRECTIVE=INSPECT_CALLER_AT_src/components/
Header.tsx:42]
[:END_CONTROL_PLANE]
src/components/Header.tsx:42:15 - SyntaxError
42 | const user = useSession( ;
| ^ Expected ')'
🛡️ [CONNECTION_STRING_REDACTED]
🛡️ [REDACTED]
ANTHROPIC_API_KEY=[REDACTED_SECRET_KEY]
Final token count: 118 tokens.
Reduction: 99.7%.
Zero credentials leaked to the cloud.
## Why Rust and why local-first?
I did not want another slow node script or cloud proxy
adding 300ms of network latency to an agent loop.
1. Sub-millisecond latency: The core log surgery and
secret redaction pipeline runs in <0.2 milliseconds on
an Intel i5 CPU.
2. Zero cloud leaks: It runs 100% locally over stdio.
Your error logs, environment variables, and
proprietary code never touch an external server.
3. AST verification & rollback: The apply_code_patch
tool parses modified code with Tree-sitter before
saving to disk. If the agent generates invalid syntax,
it immediately rolls back with zero dirty git diff.
4. Lightweight footprint: Baseline process memory is
3.45 MB VmRSS.
Glama.ai audited the server definition under their
Tool Definition Quality Score (TDQS) and awarded it
Grade A (4.7 / 5.0) across all tools.
## How to set it up (Takes 30 seconds)
You don't need to install Rust or compile anything. We
distribute pre-built native binaries via npm for
Linux, macOS (Apple Silicon & Intel), and Windows.
Add this to your claude_desktop_config.json or Cursor
MCP settings:
{
"mcpServers": {
"tokenectomy": {
"command": "npx",
"args": ["-y", "tokenectomy-razor", "--mcp"]
}
}
}
Or if you prefer native cargo:
cargo install tokenectomy
Then configure:
{
"mcpServers": {
"tokenectomy": {
"command": "razor",
"args": ["--mcp"]
}
}
}
## Open Source & Repositories
Everything is open source under the MIT license:
1. GitHub: https://github.com/Tokenectomy-
Labs/Tokenectomy
2. Glama: https://glama.ai/mcp/servers/Tokenectomy-
Labs/Tokenectomy
3. npm: https://www.npmjs.com/package/tokenectomy-
razor
4. crates.io: https://crates.io/crates/tokenectomy
Give it a run next time your agent is about to eat a
40,000-token crash log. Your token bill will thank
you.
Top comments (5)
we had our version of this problem in production, except ours was quieter
and way more expensive. our RAG pipeline was feeding full stack traces
into the LLM context when document parsing failed — thinking "more context
helps the model debug." turned out those traces averaged 8-12k tokens each,
and with ~50 parse failures per hour, we were burning $40-60/day just on
error context that the model couldn't even use.
the fix was dumb simple: a preprocessor that strips everything outside
src/ and redacts anything matching secret patterns (api keys, connection
strings, bearer tokens). went from "full trace dump" to "the three lines
around the failure + the error type." token usage dropped 95%, and
suddenly our logs were actually readable by humans too, not just models.
your point about local-first is where most teams miss the boat. we tried
a cloud proxy for this at first, and the latency killed our agent loop —
each error added 200-300ms, and when you're doing 10 tool calls per task,
that's seconds of pure overhead. moving it to a local rust binary was
exactly your play: <1ms latency, zero network, zero leaks.
the tree-sitter rollback is clever. we never got that fancy — our apply
step just writes to a temp file, runs the linter, and only overwrites if
it passes. less elegant, but it caught maybe 80% of the garbage the model
tried to commit. what's your false positive rate on the redaction regex?
we kept hitting false positives on base64-encoded values that looked like
secrets but weren't.
Really appreciate the deep dive, that $40-60/day number is exactly the kind of silent cost most teams never audit until someone actually measures it.
Honest answer on false positives: right now the redaction is a deterministic regex/DFA pass, not entropy-based, so yeah — high-entropy base64 blobs that aren't secrets can get flagged. It's a known tradeoff (favoring over-redaction over leaking a real key), but I'm looking at adding an entropy threshold + length heuristic to cut down the noise without loosening the false-negative side. If you've got example patterns that tripped it, I'd genuinely love to see them.
Your temp-file + linter approach for the apply step is a solid pragmatic call btw, might actually steal that as a fallback layer.
glad the temp-file trick earned its keep — steal it freely, it's survived
two years of model-generated garbage so far.
on false positives, the three patterns that hit us most:
what eventually cut the noise: shape allowlists (trace ids have a fixed
dash format, checksums are always 64 hex and always sit next to the word
"checksum" in our logs), plus routing flagged hits into a quarantine list
instead of hard-masking, so a human skims it once a week. annoying, but that
queue caught two real keys in the first month, so nobody argues about it
anymore.
entropy threshold + length heuristic sounds right for your case — maybe
weight it by context: the same string next to "authorization:" vs next to
"traceparent:" is a different question. happy to send anonymized samples
from our quarantine list if useful, just need to scrub them first.
Do you resolve source maps before pruning frames? With generated JSX or SVG components, the useful frame may point to generated code rather than the original source
Good catch — currently it doesn't resolve source maps before pruning, it's filtering based on path heuristics (excluding node_modules/build output). So yeah, for generated JSX/SVG or anything from a bundler transform, the "useful" frame could absolutely point at generated code instead of the real source. That's a gap, adding it to the roadmap — appreciate you flagging it.