<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Vidyax</title>
    <description>The latest articles on DEV Community by Vidyax (@daffa2555).</description>
    <link>https://dev.to/daffa2555</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4125363%2F2f80f83a-db73-4345-a6ea-b98779317d43.png</url>
      <title>DEV Community: Vidyax</title>
      <link>https://dev.to/daffa2555</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/daffa2555"/>
    <language>en</language>
    <item>
      <title>Why Does Your AI Coding Agent Start Forgetting What It Was Doing?</title>
      <dc:creator>Vidyax</dc:creator>
      <pubDate>Tue, 22 Sep 2026 23:19:03 +0000</pubDate>
      <link>https://dev.to/daffa2555/why-does-your-ai-coding-agent-start-forgetting-what-it-was-doing-fpp</link>
      <guid>https://dev.to/daffa2555/why-does-your-ai-coding-agent-start-forgetting-what-it-was-doing-fpp</guid>
      <description>&lt;p&gt;Why Does Your AI Coding Agent Start Forgetting What It Was Doing?&lt;/p&gt;

&lt;p&gt;I've been working with AI coding agents for quite a while, and there's one behavior that keeps bothering me.&lt;/p&gt;

&lt;p&gt;At first, everything looks fine.&lt;/p&gt;

&lt;p&gt;The agent understands the task, reads the error, finds the relevant file, makes a change, runs the test, and moves forward.&lt;/p&gt;

&lt;p&gt;Then, after several iterations, something strange can happen.&lt;/p&gt;

&lt;p&gt;The agent starts going back to things that have already been fixed.&lt;/p&gt;

&lt;p&gt;It reads old errors again.&lt;/p&gt;

&lt;p&gt;It investigates files that are no longer relevant.&lt;/p&gt;

&lt;p&gt;Sometimes, it even starts working on something that was already completed instead of focusing on the part that is still broken.&lt;/p&gt;

&lt;p&gt;And the first thing we usually think is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The model is getting stupid."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But I'm not convinced that's always the real problem.&lt;/p&gt;

&lt;p&gt;The context gets messy&lt;/p&gt;

&lt;p&gt;When an AI coding agent works on a real software project, it doesn't only see the code.&lt;/p&gt;

&lt;p&gt;It also sees a lot of other information:&lt;/p&gt;

&lt;p&gt;framework logs&lt;/p&gt;

&lt;p&gt;stack traces&lt;/p&gt;

&lt;p&gt;dependency output&lt;/p&gt;

&lt;p&gt;tool responses&lt;/p&gt;

&lt;p&gt;terminal output&lt;/p&gt;

&lt;p&gt;previous errors&lt;/p&gt;

&lt;p&gt;repeated information&lt;/p&gt;

&lt;p&gt;files from previous investigations&lt;/p&gt;

&lt;p&gt;debugging attempts that are already finished&lt;/p&gt;

&lt;p&gt;And this happens over and over again.&lt;/p&gt;

&lt;p&gt;A long debugging session can look something like this:&lt;/p&gt;

&lt;p&gt;Task&lt;br&gt;
  ↓&lt;br&gt;
Error&lt;br&gt;
  ↓&lt;br&gt;
Tool call&lt;br&gt;
  ↓&lt;br&gt;
Framework logs&lt;br&gt;
  ↓&lt;br&gt;
Stack trace&lt;br&gt;
  ↓&lt;br&gt;
Relevant code&lt;br&gt;
  ↓&lt;br&gt;
Patch&lt;br&gt;
  ↓&lt;br&gt;
Test&lt;br&gt;
  ↓&lt;br&gt;
New error&lt;br&gt;
  ↓&lt;br&gt;
More logs&lt;br&gt;
  ↓&lt;br&gt;
More tool output&lt;br&gt;
  ↓&lt;br&gt;
Another patch&lt;br&gt;
  ↓&lt;br&gt;
Another test&lt;br&gt;
  ↓&lt;br&gt;
...&lt;/p&gt;

&lt;p&gt;The agent keeps accumulating information.&lt;/p&gt;

&lt;p&gt;Eventually, the problem may not be that the context window is too small.&lt;/p&gt;

&lt;p&gt;It may simply be that the context has become too noisy.&lt;/p&gt;

&lt;p&gt;A bigger context window doesn't automatically solve this&lt;/p&gt;

&lt;p&gt;This is something I've been thinking about for a while.&lt;/p&gt;

&lt;p&gt;Even if an LLM has a very large context window, an agent can still fill that context with information that is no longer useful.&lt;/p&gt;

&lt;p&gt;So the question isn't only:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"How much information can the model handle?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is also:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"How much of that information is actually useful right now?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Imagine an agent is debugging a large application.&lt;/p&gt;

&lt;p&gt;Early in the process, it discovers a problem in the authentication middleware.&lt;/p&gt;

&lt;p&gt;It fixes the issue.&lt;/p&gt;

&lt;p&gt;The tests pass.&lt;/p&gt;

&lt;p&gt;The agent moves on.&lt;/p&gt;

&lt;p&gt;Ten iterations later, that entire investigation is still sitting somewhere in the conversation alongside old logs, tool outputs, stack traces, and previous debugging attempts.&lt;/p&gt;

&lt;p&gt;The agent can still see it.&lt;/p&gt;

&lt;p&gt;But the useful state is much simpler:&lt;/p&gt;

&lt;p&gt;Authentication middleware&lt;br&gt;
→ DONE&lt;/p&gt;

&lt;p&gt;Database transaction&lt;br&gt;
→ STILL BROKEN&lt;/p&gt;

&lt;p&gt;That distinction matters.&lt;/p&gt;

&lt;p&gt;If everything remains in the context without a clear representation of what has already been resolved, the agent has to continuously work through information that may no longer be relevant to the current problem.&lt;/p&gt;

&lt;p&gt;And that's where things get interesting.&lt;/p&gt;

&lt;p&gt;This is why I started building Tokenectomy&lt;/p&gt;

&lt;p&gt;Tokenectomy started from a simple idea:&lt;/p&gt;

&lt;p&gt;What if we process the information before giving it to the agent?&lt;/p&gt;

&lt;p&gt;Instead of blindly passing everything produced by the environment back into the model, we can try to:&lt;/p&gt;

&lt;p&gt;remove unnecessary noise&lt;/p&gt;

&lt;p&gt;keep the relevant information&lt;/p&gt;

&lt;p&gt;extract useful context&lt;/p&gt;

&lt;p&gt;preserve important state&lt;/p&gt;

&lt;p&gt;redact sensitive information&lt;/p&gt;

&lt;p&gt;reduce irrelevant output&lt;/p&gt;

&lt;p&gt;The idea looks roughly like this:&lt;/p&gt;

&lt;p&gt;Raw Environment&lt;br&gt;
               │&lt;br&gt;
               ▼&lt;br&gt;
   ┌─────────────────────────┐&lt;br&gt;
   │ logs / stack traces      │&lt;br&gt;
   │ tool output / code       │&lt;br&gt;
   │ terminal output          │&lt;br&gt;
   └────────────┬────────────┘&lt;br&gt;
                │&lt;br&gt;
                ▼&lt;br&gt;
           Tokenectomy&lt;br&gt;
                │&lt;br&gt;
                ▼&lt;br&gt;
      Relevant Context + State&lt;br&gt;
                │&lt;br&gt;
                ▼&lt;br&gt;
           AI Agent&lt;br&gt;
                │&lt;br&gt;
                ▼&lt;br&gt;
          Action / Patch&lt;br&gt;
                │&lt;br&gt;
                ▼&lt;br&gt;
            Test / Run&lt;br&gt;
                │&lt;br&gt;
                └──────────► Feedback&lt;/p&gt;

&lt;p&gt;The goal isn't to make the underlying model magically smarter.&lt;/p&gt;

&lt;p&gt;The goal is to make the information surrounding the model more useful.&lt;/p&gt;

&lt;p&gt;Tokenectomy isn't just about saving tokens&lt;/p&gt;

&lt;p&gt;This is an important distinction.&lt;/p&gt;

&lt;p&gt;At first glance, something that removes unnecessary context sounds like a token optimization tool.&lt;/p&gt;

&lt;p&gt;But I'm more interested in what happens after the cleanup.&lt;/p&gt;

&lt;p&gt;If an agent receives less irrelevant information, does it make fewer unnecessary tool calls?&lt;/p&gt;

&lt;p&gt;Does it repeat fewer investigations?&lt;/p&gt;

&lt;p&gt;Does it stay focused on the remaining problem for longer?&lt;/p&gt;

&lt;p&gt;Does its performance degrade less as the debugging session becomes longer?&lt;/p&gt;

&lt;p&gt;Those are much more interesting questions to me than simply:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"How many tokens did we save?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Tokenectomy is evolving into an infrastructure/tooling layer for AI agents, with things like MCP tools, context processing, code analysis, patching, and an AI Gateway.&lt;/p&gt;

&lt;p&gt;But the technology itself isn't really the interesting part.&lt;/p&gt;

&lt;p&gt;The interesting part is the hypothesis behind it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can better context management make an AI agent more reliable during long-running tasks?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I'm not saying context noise is always the problem&lt;/p&gt;

&lt;p&gt;This is important.&lt;/p&gt;

&lt;p&gt;I'm not claiming that every AI agent failure is caused by messy context.&lt;/p&gt;

&lt;p&gt;Agents can fail for many different reasons.&lt;/p&gt;

&lt;p&gt;The model can misunderstand the task.&lt;/p&gt;

&lt;p&gt;A tool can return bad information.&lt;/p&gt;

&lt;p&gt;The generated patch can be incorrect.&lt;/p&gt;

&lt;p&gt;The test environment can be broken.&lt;/p&gt;

&lt;p&gt;The agent can simply make a bad reasoning decision.&lt;/p&gt;

&lt;p&gt;Context noise is only one possible factor.&lt;/p&gt;

&lt;p&gt;But I've repeatedly seen situations where an agent starts revisiting old work during long debugging sessions.&lt;/p&gt;

&lt;p&gt;So instead of saying:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"AI agents forget because their context gets messy."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I'd rather ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What happens if we deliberately remove irrelevant context during a long debugging trajectory?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's something we can actually test.&lt;/p&gt;

&lt;p&gt;Let's measure it&lt;/p&gt;

&lt;p&gt;If this idea is real, it should show up in the data.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;How many tokens are sent to the model?&lt;/p&gt;

&lt;p&gt;How many iterations are required?&lt;/p&gt;

&lt;p&gt;How often does the agent repeat completed work?&lt;/p&gt;

&lt;p&gt;How often does it call irrelevant tools?&lt;/p&gt;

&lt;p&gt;How often does it revisit previously resolved errors?&lt;/p&gt;

&lt;p&gt;Does the final patch pass the tests?&lt;/p&gt;

&lt;p&gt;How does performance change as the debugging session gets longer?&lt;/p&gt;

&lt;p&gt;Does context cleaning reduce that degradation?&lt;/p&gt;

&lt;p&gt;The goal isn't to make a nice-looking demo.&lt;/p&gt;

&lt;p&gt;The goal is to find out whether this actually changes the behavior of an agent.&lt;/p&gt;

&lt;p&gt;That's where Kronumos comes in&lt;/p&gt;

&lt;p&gt;I'm also building Kronumos, a specialized software-repair agent, to experiment with this idea in a more focused environment.&lt;/p&gt;

&lt;p&gt;Kronumos isn't meant to be another general-purpose coding assistant.&lt;/p&gt;

&lt;p&gt;The idea is much narrower:&lt;/p&gt;

&lt;p&gt;Bug&lt;br&gt;
 ↓&lt;br&gt;
Understand the failure&lt;br&gt;
 ↓&lt;br&gt;
Get relevant context&lt;br&gt;
 ↓&lt;br&gt;
Analyze the code&lt;br&gt;
 ↓&lt;br&gt;
Generate a patch&lt;br&gt;
 ↓&lt;br&gt;
Run tests&lt;br&gt;
 ↓&lt;br&gt;
Inspect the result&lt;br&gt;
 ↓&lt;br&gt;
Fix again if necessary&lt;/p&gt;

&lt;p&gt;This gives me a controlled environment where I can experiment with different combinations of:&lt;/p&gt;

&lt;p&gt;model + context + tools + execution feedback.&lt;/p&gt;

&lt;p&gt;And that's the part I'm really interested in.&lt;/p&gt;

&lt;p&gt;I'm not trying to claim that a small model is inherently smarter than a much larger model.&lt;/p&gt;

&lt;p&gt;That's not the question.&lt;/p&gt;

&lt;p&gt;The question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How much can system design, specialized training, better context, and a proper feedback loop affect the performance of an AI agent?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Because sometimes the model might not be the only problem.&lt;/p&gt;

&lt;p&gt;Sometimes, maybe, we're just giving it too much garbage.&lt;/p&gt;

&lt;p&gt;And that's the problem I'm trying to solve with Tokenectomy.&lt;/p&gt;




&lt;p&gt;I'd love to hear from other developers&lt;/p&gt;

&lt;p&gt;Have you ever had an AI coding agent suddenly start working on something it had already finished?&lt;/p&gt;

&lt;p&gt;Or had an agent repeatedly investigate an error that you thought was already resolved?&lt;/p&gt;

&lt;p&gt;I'm curious whether other people are seeing the same behavior in long-running coding sessions.&lt;/p&gt;

&lt;p&gt;GitHub: [&lt;a href="https://github.com/Tokenectomy-Labs/Tokenectomy" rel="noopener noreferrer"&gt;https://github.com/Tokenectomy-Labs/Tokenectomy&lt;/a&gt;]&lt;/p&gt;

</description>
      <category>ai</category>
    </item>
    <item>
      <title>How I Cut 45,000 Next.js Error Tokens to 118 in &lt;0.2ms Using a Rust MCP Server</title>
      <dc:creator>Vidyax</dc:creator>
      <pubDate>Tue, 15 Sep 2026 03:19:09 +0000</pubDate>
      <link>https://dev.to/daffa2555/how-i-cut-45000-nextjs-error-tokens-to-118-in-02msusing-a-rust-mcp-server-1c7j</link>
      <guid>https://dev.to/daffa2555/how-i-cut-45000-nextjs-error-tokens-to-118-in-02msusing-a-rust-mcp-server-1c7j</guid>
      <description>&lt;p&gt;If you use Cursor, Claude Desktop, or Cline daily,&lt;br&gt;
  you have probably watched your rate limits evaporate&lt;br&gt;
  because of a single runtime crash.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A Next.js build fails or a Python script panics, and
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;your terminal vomits 500 lines of stack traces. The&lt;br&gt;
  LLM eagerly ingests all 45,000 tokens of internal&lt;br&gt;
  &lt;code&gt;node_modules&lt;/code&gt; machinery, webpack bundles, and event&lt;br&gt;
  loop frames.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You just paid $0.15 for an AI model to read code it
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;can't edit, your prompt cache is wiped, and your agent&lt;br&gt;
  is now hallucinating because its context window is&lt;br&gt;
  full of garbage.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Worse, your terminal stderr probably just leaked
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;code&gt;DATABASE_URL=postgres://admin:password@...&lt;/code&gt; straight&lt;br&gt;
  to an external API.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;I got tired of paying for framework noise, so I
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;built an open-source Rust Model Context Protocol (MCP)&lt;br&gt;
  server called &lt;strong&gt;Tokenectomy Razor&lt;/strong&gt; to fix it locally&lt;br&gt;
  before logs ever touch the LLM.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;## What is eating your context window?

A typical Next.js error trace looks like this:
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
text
    TypeError: Cannot read properties of undefined
  (reading 'digest')
        at Object.&amp;lt;anon&amp;gt; (/node_modules/next/bundle5.
  js:142:31)
        at __webpack_require__
  (/node_modules/next/bundle5.js:198:12)
        at Object.execute (/node_modules/next/dev-
  server.js:412:19)
        at processTicksAndRejections (task_queues:95:5)
        Database connection failed:
  postgresql://admin:super_secret_password@db.prod.
  internal:5432/primary
        API key leaked: sk-ant-api03-
  abcdef1234567890abcdef1234567890
        [... 480 internal dependency frames flooding
  context ...]

  Notice three things:

  1. 98% of those frames are inside node_modules. Your
  AI agent is not going to edit webpack's internal
  bundle logic. It only cares about the one line in
  src/components/Header.tsx:42 where you missed a
  parenthesis.
  2. The database password and API key are sitting
  unmasked in plain text.
  3. The raw token count for this single error dump was
  45,820 tokens.

  ## What happens after Tokenectomy runs

  When the AI agent invokes get_error_context via MCP,
  Tokenectomy intercepts the log, strips framework
  internals, redacts all secrets locally using a
  deterministic DFA regex, and grabs bounded source code
  lines around the actual crash:

    [:TOKENECTOMY:M2M_CONTROL_PLANE:v1.3.0]
    [STATE=FRAMEWORK_NOISE_PURGED]
    [STRATEGY_APPLIED=AGGRESSIVE]
    [ORIGINAL_BYTES=45820 | CLEAN_BYTES=118 |
  REDUCTION=99%]
    [PRIMARY_CRASH_COORDINATES=src/components/Header.
  tsx:42]

  [COGNITIVE_DIRECTIVE=INSPECT_CALLER_AT_src/components/
  Header.tsx:42]
    [:END_CONTROL_PLANE]

    src/components/Header.tsx:42:15 - SyntaxError
      42 |   const user = useSession( ;
         |                           ^ Expected ')'
    🛡️ [CONNECTION_STRING_REDACTED]
    🛡️ [REDACTED]
  ANTHROPIC_API_KEY=[REDACTED_SECRET_KEY]

  Final token count: 118 tokens.
  Reduction: 99.7%.
  Zero credentials leaked to the cloud.

  ## Why Rust and why local-first?

  I did not want another slow node script or cloud proxy
  adding 300ms of network latency to an agent loop.

  1. Sub-millisecond latency: The core log surgery and
  secret redaction pipeline runs in &amp;lt;0.2 milliseconds on
  an Intel i5 CPU.
  2. Zero cloud leaks: It runs 100% locally over stdio.
  Your error logs, environment variables, and
  proprietary code never touch an external server.
  3. AST verification &amp;amp; rollback: The apply_code_patch
  tool parses modified code with Tree-sitter before
  saving to disk. If the agent generates invalid syntax,
  it immediately rolls back with zero dirty git diff.
  4. Lightweight footprint: Baseline process memory is
  3.45 MB VmRSS.

  Glama.ai audited the server definition under their
  Tool Definition Quality Score (TDQS) and awarded it
  Grade A (4.7 / 5.0) across all tools.

  ## How to set it up (Takes 30 seconds)

  You don't need to install Rust or compile anything. We
  distribute pre-built native binaries via npm for
  Linux, macOS (Apple Silicon &amp;amp; Intel), and Windows.

  Add this to your claude_desktop_config.json or Cursor
  MCP settings:

    {
      "mcpServers": {
        "tokenectomy": {
          "command": "npx",
          "args": ["-y", "tokenectomy-razor", "--mcp"]
        }
      }
    }

  Or if you prefer native cargo:

    cargo install tokenectomy

  Then configure:

    {
      "mcpServers": {
        "tokenectomy": {
          "command": "razor",
          "args": ["--mcp"]
        }
      }
    }

  ## Open Source &amp;amp; Repositories

  Everything is open source under the MIT license:

  1. GitHub: https://github.com/Tokenectomy-
  Labs/Tokenectomy
  2. Glama: https://glama.ai/mcp/servers/Tokenectomy-
  Labs/Tokenectomy
  3. npm: https://www.npmjs.com/package/tokenectomy-
  razor
  4. crates.io: https://crates.io/crates/tokenectomy

  Give it a run next time your agent is about to eat a
  40,000-token crash log. Your token bill will thank
  you.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>rust</category>
      <category>nextjs</category>
    </item>
  </channel>
</rss>
