DEV Community

Vidyax
Vidyax

Posted on

How I Cut 45,000 Next.js Error Tokens to 118 in <0.2ms Using a Rust MCP Server

If you use Cursor, Claude Desktop, or Cline daily,
you have probably watched your rate limits evaporate
because of a single runtime crash.

A Next.js build fails or a Python script panics, and
Enter fullscreen mode Exit fullscreen mode

your terminal vomits 500 lines of stack traces. The
LLM eagerly ingests all 45,000 tokens of internal
node_modules machinery, webpack bundles, and event
loop frames.

You just paid $0.15 for an AI model to read code it
Enter fullscreen mode Exit fullscreen mode

can't edit, your prompt cache is wiped, and your agent
is now hallucinating because its context window is
full of garbage.

Worse, your terminal stderr probably just leaked
Enter fullscreen mode Exit fullscreen mode

DATABASE_URL=postgres://admin:password@... straight
to an external API.

I got tired of paying for framework noise, so I
Enter fullscreen mode Exit fullscreen mode

built an open-source Rust Model Context Protocol (MCP)
server called Tokenectomy Razor to fix it locally
before logs ever touch the LLM.

## What is eating your context window?

A typical Next.js error trace looks like this:
Enter fullscreen mode Exit fullscreen mode

text
    TypeError: Cannot read properties of undefined
  (reading 'digest')
        at Object.<anon> (/node_modules/next/bundle5.
  js:142:31)
        at __webpack_require__
  (/node_modules/next/bundle5.js:198:12)
        at Object.execute (/node_modules/next/dev-
  server.js:412:19)
        at processTicksAndRejections (task_queues:95:5)
        Database connection failed:
  postgresql://admin:super_secret_password@db.prod.
  internal:5432/primary
        API key leaked: sk-ant-api03-
  abcdef1234567890abcdef1234567890
        [... 480 internal dependency frames flooding
  context ...]

  Notice three things:

  1. 98% of those frames are inside node_modules. Your
  AI agent is not going to edit webpack's internal
  bundle logic. It only cares about the one line in
  src/components/Header.tsx:42 where you missed a
  parenthesis.
  2. The database password and API key are sitting
  unmasked in plain text.
  3. The raw token count for this single error dump was
  45,820 tokens.

  ## What happens after Tokenectomy runs

  When the AI agent invokes get_error_context via MCP,
  Tokenectomy intercepts the log, strips framework
  internals, redacts all secrets locally using a
  deterministic DFA regex, and grabs bounded source code
  lines around the actual crash:

    [:TOKENECTOMY:M2M_CONTROL_PLANE:v1.3.0]
    [STATE=FRAMEWORK_NOISE_PURGED]
    [STRATEGY_APPLIED=AGGRESSIVE]
    [ORIGINAL_BYTES=45820 | CLEAN_BYTES=118 |
  REDUCTION=99%]
    [PRIMARY_CRASH_COORDINATES=src/components/Header.
  tsx:42]

  [COGNITIVE_DIRECTIVE=INSPECT_CALLER_AT_src/components/
  Header.tsx:42]
    [:END_CONTROL_PLANE]

    src/components/Header.tsx:42:15 - SyntaxError
      42 |   const user = useSession( ;
         |                           ^ Expected ')'
    🛡️ [CONNECTION_STRING_REDACTED]
    🛡️ [REDACTED]
  ANTHROPIC_API_KEY=[REDACTED_SECRET_KEY]

  Final token count: 118 tokens.
  Reduction: 99.7%.
  Zero credentials leaked to the cloud.

  ## Why Rust and why local-first?

  I did not want another slow node script or cloud proxy
  adding 300ms of network latency to an agent loop.

  1. Sub-millisecond latency: The core log surgery and
  secret redaction pipeline runs in <0.2 milliseconds on
  an Intel i5 CPU.
  2. Zero cloud leaks: It runs 100% locally over stdio.
  Your error logs, environment variables, and
  proprietary code never touch an external server.
  3. AST verification & rollback: The apply_code_patch
  tool parses modified code with Tree-sitter before
  saving to disk. If the agent generates invalid syntax,
  it immediately rolls back with zero dirty git diff.
  4. Lightweight footprint: Baseline process memory is
  3.45 MB VmRSS.

  Glama.ai audited the server definition under their
  Tool Definition Quality Score (TDQS) and awarded it
  Grade A (4.7 / 5.0) across all tools.

  ## How to set it up (Takes 30 seconds)

  You don't need to install Rust or compile anything. We
  distribute pre-built native binaries via npm for
  Linux, macOS (Apple Silicon & Intel), and Windows.

  Add this to your claude_desktop_config.json or Cursor
  MCP settings:

    {
      "mcpServers": {
        "tokenectomy": {
          "command": "npx",
          "args": ["-y", "tokenectomy-razor", "--mcp"]
        }
      }
    }

  Or if you prefer native cargo:

    cargo install tokenectomy

  Then configure:

    {
      "mcpServers": {
        "tokenectomy": {
          "command": "razor",
          "args": ["--mcp"]
        }
      }
    }

  ## Open Source & Repositories

  Everything is open source under the MIT license:

  1. GitHub: https://github.com/Tokenectomy-
  Labs/Tokenectomy
  2. Glama: https://glama.ai/mcp/servers/Tokenectomy-
  Labs/Tokenectomy
  3. npm: https://www.npmjs.com/package/tokenectomy-
  razor
  4. crates.io: https://crates.io/crates/tokenectomy

  Give it a run next time your agent is about to eat a
  40,000-token crash log. Your token bill will thank
  you.
Enter fullscreen mode Exit fullscreen mode

Top comments (5)

Collapse
 
kaziava profile image
Hardcore Engineer •

we had our version of this problem in production, except ours was quieter
and way more expensive. our RAG pipeline was feeding full stack traces
into the LLM context when document parsing failed — thinking "more context
helps the model debug." turned out those traces averaged 8-12k tokens each,
and with ~50 parse failures per hour, we were burning $40-60/day just on
error context that the model couldn't even use.

the fix was dumb simple: a preprocessor that strips everything outside
src/ and redacts anything matching secret patterns (api keys, connection
strings, bearer tokens). went from "full trace dump" to "the three lines
around the failure + the error type." token usage dropped 95%, and
suddenly our logs were actually readable by humans too, not just models.

your point about local-first is where most teams miss the boat. we tried
a cloud proxy for this at first, and the latency killed our agent loop —
each error added 200-300ms, and when you're doing 10 tool calls per task,
that's seconds of pure overhead. moving it to a local rust binary was
exactly your play: <1ms latency, zero network, zero leaks.

the tree-sitter rollback is clever. we never got that fancy — our apply
step just writes to a temp file, runs the linter, and only overwrites if
it passes. less elegant, but it caught maybe 80% of the garbage the model
tried to commit. what's your false positive rate on the redaction regex?
we kept hitting false positives on base64-encoded values that looked like
secrets but weren't.

Collapse
 
daffa2555 profile image
Vidyax •

Really appreciate the deep dive, that $40-60/day number is exactly the kind of silent cost most teams never audit until someone actually measures it.
Honest answer on false positives: right now the redaction is a deterministic regex/DFA pass, not entropy-based, so yeah — high-entropy base64 blobs that aren't secrets can get flagged. It's a known tradeoff (favoring over-redaction over leaking a real key), but I'm looking at adding an entropy threshold + length heuristic to cut down the noise without loosening the false-negative side. If you've got example patterns that tripped it, I'd genuinely love to see them.
Your temp-file + linter approach for the apply step is a solid pragmatic call btw, might actually steal that as a fallback layer.

Collapse
 
kaziava profile image
Hardcore Engineer •

glad the temp-file trick earned its keep — steal it freely, it's survived
two years of model-generated garbage so far.

on false positives, the three patterns that hit us most:

  1. w3c traceparent ids — every span log line carries a 32-hex trace id, and our "looks like an api key" regex ate them all. thousands of flags per day, all harmless.
  2. base64 data urls in error payloads — frontend logs included screenshot thumbnails as data:image/png;base64,... and the redactor masked the whole blob, which also destroyed the one useful part of the log.
  3. sha256 checksums of uploaded files — we log them for integrity checks, and a 64-hex string is indistinguishable from a secret to a dumb regex.

what eventually cut the noise: shape allowlists (trace ids have a fixed
dash format, checksums are always 64 hex and always sit next to the word
"checksum" in our logs), plus routing flagged hits into a quarantine list
instead of hard-masking, so a human skims it once a week. annoying, but that
queue caught two real keys in the first month, so nobody argues about it
anymore.

entropy threshold + length heuristic sounds right for your case — maybe
weight it by context: the same string next to "authorization:" vs next to
"traceparent:" is a different question. happy to send anonymized samples
from our quarantine list if useful, just need to scrub them first.

Collapse
 
svgicons profile image
Svg/icons •

Do you resolve source maps before pruning frames? With generated JSX or SVG components, the useful frame may point to generated code rather than the original source

Collapse
 
daffa2555 profile image
Vidyax •

Good catch — currently it doesn't resolve source maps before pruning, it's filtering based on path heuristics (excluding node_modules/build output). So yeah, for generated JSX/SVG or anything from a bundler transform, the "useful" frame could absolutely point at generated code instead of the real source. That's a gap, adding it to the roadmap — appreciate you flagging it.