<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: linweidao</title>
    <description>The latest articles on DEV Community by linweidao (@sloves).</description>
    <link>https://dev.to/sloves</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4097213%2Fb3dd63ae-67b4-48a7-9f8a-18435c9f2e70.png</url>
      <title>DEV Community: linweidao</title>
      <link>https://dev.to/sloves</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sloves"/>
    <language>en</language>
    <item>
      <title>Cutting Agent Token Bloat: Testing Headroom as an MCP Compression Layer in Cursor</title>
      <dc:creator>linweidao</dc:creator>
      <pubDate>Thu, 08 Oct 2026 18:06:37 +0000</pubDate>
      <link>https://dev.to/sloves/cutting-agent-token-bloat-testing-headroom-as-an-mcp-compression-layer-in-cursor-3794</link>
      <guid>https://dev.to/sloves/cutting-agent-token-bloat-testing-headroom-as-an-mcp-compression-layer-in-cursor-3794</guid>
      <description>&lt;p&gt;Long-running coding agent sessions in Cursor and VS Code suffer from prompt bloat. When your agent inspects large test fixtures, parses build artifacts, or slurps API responses into context, a single turn can balloon past 50,000 tokens. Most of that payload consists of repetitive JSON schema boilerplate, stack traces, and verbose AST dumps.&lt;/p&gt;

&lt;p&gt;While evaluating context-efficiency tooling, I tested &lt;strong&gt;Headroom&lt;/strong&gt; (&lt;code&gt;headroomlabs-ai/headroom&lt;/code&gt;), a local compression layer that strips redundant tokens from tool outputs and files before dispatching them to the inference model.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Architecture: Local Pre-Processing
&lt;/h3&gt;

&lt;p&gt;Headroom operates locally on your machine via CLI, local proxy, or MCP server. Rather than running lossy summarization via external cloud models, its pipeline splits incoming data across dedicated processors:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;SmartCrusher&lt;/strong&gt;: Compresses structural JSON by 60–90% through schema factorization.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CodeCompressor&lt;/strong&gt;: Removes syntactic whitespace, non-critical comments, and AST redundancies without changing semantics.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Content-Conditioned Retrieval (CCR)&lt;/strong&gt;: Keeps full raw artifacts in a local disk cache and inserts lightweight retrieval handles into the prompt, allowing the agent to pull specific lines if needed.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Integrating Headroom MCP with Cursor
&lt;/h3&gt;

&lt;p&gt;Headroom exposes three core MCP tools: &lt;code&gt;headroom_compress&lt;/code&gt;, &lt;code&gt;headroom_retrieve&lt;/code&gt;, and &lt;code&gt;headroom_stats&lt;/code&gt;. You can hook it into Cursor's MCP configuration (&lt;code&gt;~/.cursor/mcp.json&lt;/code&gt; or project-level &lt;code&gt;.cursor/mcp.json&lt;/code&gt;):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"headroom"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"npx"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"args"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"-y"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"headroom-ai"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"mcp"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you run Python tooling, the equivalent pip-based binary works out of the box:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;headroom-ai
headroom mcp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Pairing with Rules &amp;amp; Fast Proxies
&lt;/h3&gt;

&lt;p&gt;To ensure your agent actively compresses heavy diagnostic files, enforce the tool call pattern inside your &lt;code&gt;.cursorrules&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Context Budget Rules&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; When inspecting terminal outputs, test logs, or JSON payloads &amp;gt; 200 lines, pipe content through &lt;span class="sb"&gt;`headroom_compress`&lt;/span&gt; first.
&lt;span class="p"&gt;-&lt;/span&gt; Never paste uncompressed raw API fixtures directly into chat history.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In our multi-turn debugging benchmarks, compressing large JSON payloads and build artifacts dropped input token consumption by roughly 40% on average. For teams running massive multi-agent loops, pairing local pre-compression with optimized gateway routing yields significant compounding savings.&lt;/p&gt;

&lt;p&gt;In my daily workflow, I route Cursor and Cline through B-Lost's fast proxy endpoint, where native prompt caching cuts heavy multi-turn context costs by ~80-90% without losing chat history. Adding Headroom at the MCP layer ensures the non-cached dynamic tokens—such as new tool results and terminal logs—enter the context window as lean as possible.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclosure: Compute infrastructure and multi-model benchmark relays for this writeup are sponsored by &lt;a href="https://b-lost.com?utm_source=devto&amp;amp;utm_medium=tech_blog&amp;amp;utm_campaign=devto_bot_1" rel="noopener noreferrer"&gt;b-lost.com&lt;/a&gt; — an enterprise AI gateway offering 0.8x official pricing, native prompt caching, and zero user-data retention. All benchmark metrics reflect independent reproducible testing.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cursor</category>
      <category>vscode</category>
      <category>devtools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Reverse-Engineering AI IDE System Prompts: Taming Cursor and Copilot Context Drift in Production</title>
      <dc:creator>linweidao</dc:creator>
      <pubDate>Tue, 06 Oct 2026 17:18:30 +0000</pubDate>
      <link>https://dev.to/sloves/reverse-engineering-ai-ide-system-prompts-taming-cursor-and-copilot-context-drift-in-production-1lga</link>
      <guid>https://dev.to/sloves/reverse-engineering-ai-ide-system-prompts-taming-cursor-and-copilot-context-drift-in-production-1lga</guid>
      <description>&lt;p&gt;It is 3:14 AM when your on-call pager screams: a junior engineer ran a routine workspace refactor across a 400-file mono-repo, and your team's upstream model quota evaporated in forty minutes. Worse, the AI agent silently hallucinated modifications to deleted files, committed invalid diffs, and locked the active staging pipeline. When modern AI-native IDEs like Cursor, Windsurf, and Claude Code operate inside complex production repositories, uninspected prompt orchestration converts developer velocity into catastrophic operational latency and compounding token overhead.&lt;/p&gt;

&lt;p&gt;To understand why AI IDE agents derail during multi-turn refactors, our team turned to the community repository &lt;strong&gt;asgeirtj/system_prompts_leaks&lt;/strong&gt;, which indexes and documents extracted system prompts across frontier toolchains including Cursor, Claude Code, and Copilot. Inspecting these leaked prompts reveals a stark architectural truth: &lt;strong&gt;IDE vendors trade upstream prefix stability for aggressive dynamic injection&lt;/strong&gt;. Hidden workspace context, AST snippets, and dynamic tool definitions are constantly prepended to the system prompt, silently shattering KV-cache reuse and triggering massive cache invalidations upstream.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Prefix Invalidation Problem
&lt;/h3&gt;

&lt;p&gt;Frontier LLM gateways rely on exact byte-for-byte prefix matching to leverage prompt caching. When an IDE agent recalculates file trees or alters tool registration signatures dynamically, the prompt prefix shifts. Instead of hitting a 90% cached route at sub-second TTFT (Time to First Token), the runtime forces a full 100k+ token prefill on every keystroke or tool invocation.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;+-------------------------------------------------------------+
|               AI-Native IDE Client (Editor)                 |
|  - Dynamic File Tree    - MCP Server Tools   - Active Buffer |
+------------------------------+------------------------------+
                               |
                               v [Volatile System Prefix]
+-------------------------------------------------------------+
|              Context Normalization Proxy / Gateway          |
|  - Static Rule Pinning       - KV-Cache Friendly Ordering   |
|  - Strict Truncation Bounds  - Tool Schema Fingerprinting   |
+------------------------------+------------------------------+
                               |
                               v [Fixed Prefix: Cache Hit]
+-------------------------------------------------------------+
|                   Upstream Model Runtime                    |
+-------------------------------------------------------------+
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;By analyzing prompt architectures documented in &lt;code&gt;asgeirtj/system_prompts_leaks&lt;/code&gt;, we identified the primary culprits behind runtime drift:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Unbounded Context Leaks&lt;/strong&gt;: Inline diff histories dumped directly into conversational scratchpads rather than ephemeral tool results.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool Schema Churn&lt;/strong&gt;: Ephemeral Model Context Protocol (MCP) servers repeatedly registering fluctuating schema descriptions between iterations.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Instruction Dilution&lt;/strong&gt;: Verbose vendor preambles overriding user-specified project guidelines under heavy context pressure.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Hardening Editor Rules Against Drift
&lt;/h3&gt;

&lt;p&gt;To prevent the agent from thrashing context and fabricating file edits, we apply deterministic anti-drift constraints directly inside workspace configuration. Below is our production &lt;code&gt;.cursorrules&lt;/code&gt; file, specifically structured to stabilize cache prefixes and enforce zero-hallucination execution boundaries:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"version"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2.0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"contextRules"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"enforceCachePrefixPurity"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"maxWorkspaceSummaryTokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"prohibitGhostFileEditing"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"executionConstraints"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"Never assume a file exists based solely on prior conversational turns."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"Always execute read_file or verify_path before applying diffs or edits."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"Keep diff outputs strictly scoped to targeted AST nodes."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"Do not re-summarize entire files into conversational message buffers."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"If tool execution returns an error, halt immediately without retrying synthetic parameters."&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpTooling"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"schemaValidation"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"strict"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"timeoutMilliseconds"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;5000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"quarantineUnstableServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Deterministic Context Normalization
&lt;/h3&gt;

&lt;p&gt;When bridging external MCP servers or custom IDE extensions, upstream payloads must pass through a strict sanitization layer before reaching inference. The following Node.js middleware normalizes inbound conversation context, prunes duplicate diff chains, and preserves cache-friendly prefixes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="nx"&gt;crypto&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;node:crypto&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ContextStabilizer&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;constructor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;maxHistoryTokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;8000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;maxHistoryTokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;maxHistoryTokens&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="nf"&gt;sanitizeMessages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;seenSignatures&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Set&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;sanitized&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[];&lt;/span&gt;

    &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="o"&gt;--&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;msg&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;role&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;tool&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;role&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;system&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;hash&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;crypto&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;createHash&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;sha256&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;update&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;digest&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;hex&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;seenSignatures&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;has&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;hash&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
          &lt;span class="k"&gt;continue&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="nx"&gt;seenSignatures&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;hash&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;
      &lt;span class="nx"&gt;sanitized&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;unshift&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;enforceStaticPrefix&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;sanitized&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="nf"&gt;enforceStaticPrefix&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;role&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;system&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Invalid pipeline state: Missing static system root.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  The Operational Trade-Off
&lt;/h3&gt;

&lt;p&gt;The fundamental tension in AI-assisted software engineering lies between &lt;strong&gt;agent autonomy&lt;/strong&gt; and &lt;strong&gt;deterministic system boundaries&lt;/strong&gt;. If you allow IDE agents unconstrained access to dynamic context injection, your development cycle suffers unpredictable cache invalidation, runaway compute spend, and state desync across branch checkouts. Conversely, if you constrain prompts too rigidly, the agent loses situational awareness across complex microservice boundaries.&lt;/p&gt;

&lt;p&gt;How is your engineering team balancing prompt cache utilization against agent autonomy in large mono-repos? Are you running centralized proxy filters, or relying strictly on client-side editor configs? Drop your architecture and operational battle scars in the comments below.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclosure: Compute infrastructure and multi-model benchmark relays for this writeup are sponsored by &lt;a href="https://b-lost.com?utm_source=devto&amp;amp;utm_medium=tech_blog&amp;amp;utm_campaign=devto_bot_1" rel="noopener noreferrer"&gt;b-lost.com&lt;/a&gt; — an enterprise AI gateway offering 0.8x official pricing, native prompt caching, and zero user-data retention. All benchmark metrics reflect independent reproducible testing.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cursor</category>
      <category>vscode</category>
      <category>devtools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Taming 380+ Agent Skills in Cursor: Selective Injection over Context Bloat</title>
      <dc:creator>linweidao</dc:creator>
      <pubDate>Thu, 01 Oct 2026 22:32:02 +0000</pubDate>
      <link>https://dev.to/sloves/taming-380-agent-skills-in-cursor-selective-injection-over-context-bloat-3d35</link>
      <guid>https://dev.to/sloves/taming-380-agent-skills-in-cursor-selective-injection-over-context-bloat-3d35</guid>
      <description>&lt;p&gt;When experimenting with &lt;code&gt;alirezarezvani/claude-skills&lt;/code&gt;—a massive catalog of 380+ agent skills covering architecture, debugging, and compliance—the immediate temptation is to wire the entire directory straight into your workspace.&lt;/p&gt;

&lt;p&gt;Don't do it.&lt;/p&gt;

&lt;p&gt;Bluntly dumping hundreds of skill definitions into &lt;code&gt;.cursorrules&lt;/code&gt; or global system prompts wastes context tokens, degrades model instruction-following, and leads to retrieval interference. Claude Code and Codex handle large dynamic tool registries through progressive loading, but IDE-based agents like Cursor and VS Code require strict context pruning.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Friction: Global Injection vs. Context Limits
&lt;/h3&gt;

&lt;p&gt;Each skill file carries markdown frontmatter, execution scripts, and explicit guardrails. Loading even 40 skills simultaneously pushes 25k–40k tokens into every prompt turn before you even paste a stack trace.&lt;/p&gt;

&lt;p&gt;To adopt &lt;code&gt;alirezarezvani/claude-skills&lt;/code&gt; effectively in Cursor, we isolate only domain-relevant subtrees (e.g., &lt;code&gt;engineering/&lt;/code&gt; and &lt;code&gt;code-review/&lt;/code&gt;) and map them directly into Cursor's modular &lt;code&gt;.cursor/rules/&lt;/code&gt; directory using targeted extraction.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Deliverable: Selective Skill Compiler
&lt;/h3&gt;

&lt;p&gt;Run this lightweight bash script in your project root to pull only the engineering rules and build a clean, scoped &lt;code&gt;.cursor/rules/claude-skills.mdc&lt;/code&gt; file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#!/usr/bin/env bash&lt;/span&gt;
&lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="nt"&gt;-euo&lt;/span&gt; pipefail

&lt;span class="nv"&gt;SKILLS_DIR&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;".claude-skills-cache"&lt;/span&gt;
&lt;span class="nv"&gt;TARGET_RULE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;".cursor/rules/claude-skills.mdc"&lt;/span&gt;
&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; .cursor/rules &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$SKILLS_DIR&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;

&lt;span class="c"&gt;# Sparse-checkout targeted domains only&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$SKILLS_DIR&lt;/span&gt;&lt;span class="s2"&gt;/.git"&lt;/span&gt; &lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  &lt;/span&gt;git clone &lt;span class="nt"&gt;--depth&lt;/span&gt; 1 &lt;span class="nt"&gt;--filter&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;blob:none &lt;span class="nt"&gt;--no-checkout&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    https://github.com/alirezarezvani/claude-skills.git &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$SKILLS_DIR&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
  &lt;span class="nb"&gt;pushd&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$SKILLS_DIR&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /dev/null
  git sparse-checkout &lt;span class="nb"&gt;set &lt;/span&gt;skills/engineering skills/software-development
  git checkout
  &lt;span class="nb"&gt;popd&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /dev/null
&lt;span class="k"&gt;fi&lt;/span&gt;

&lt;span class="c"&gt;# Compile into a scoped Cursor rule&lt;/span&gt;
&lt;span class="nb"&gt;cat&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="no"&gt;EOF&lt;/span&gt;&lt;span class="sh"&gt;' &amp;gt; "&lt;/span&gt;&lt;span class="nv"&gt;$TARGET_RULE&lt;/span&gt;&lt;span class="sh"&gt;"
---
description: "Core Engineering Patterns from claude-skills"
globs: *.{ts,js,py,go,rs}
alwaysApply: false
---

# Selected Engineering Guardrails
&lt;/span&gt;&lt;span class="no"&gt;EOF

&lt;/span&gt;find &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$SKILLS_DIR&lt;/span&gt;&lt;span class="s2"&gt;/skills"&lt;/span&gt; &lt;span class="nt"&gt;-name&lt;/span&gt; &lt;span class="s2"&gt;"SKILL.md"&lt;/span&gt; &lt;span class="nt"&gt;-exec&lt;/span&gt; &lt;span class="nb"&gt;cat&lt;/span&gt; &lt;span class="o"&gt;{}&lt;/span&gt; + &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$TARGET_RULE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Compiled active skills into &lt;/span&gt;&lt;span class="nv"&gt;$TARGET_RULE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Setting &lt;code&gt;alwaysApply: false&lt;/code&gt; ensures Cursor only pulls the skills when your prompt or file matching calls for them, keeping baseline conversation turns lightweight.&lt;/p&gt;

&lt;h3&gt;
  
  
  Managing Token Overhead in Heavy Sessions
&lt;/h3&gt;

&lt;p&gt;Even with selective compilation, running deep architectural reviews alongside detailed rule files swells prefix tokens quickly. In our workflow, we route Cursor's custom OpenAI/Anthropic API calls through B-Lost's fast proxy endpoint, noting that native prompt caching cuts heavy multi-turn context costs by ~80-90% without losing chat history. When your system prefix remains stable across consecutive prompts, prompt caching turns an expensive 30k-token prompt into near-zero marginal inference overhead.&lt;/p&gt;

&lt;h3&gt;
  
  
  Final Takeaway
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;alirezarezvani/claude-skills&lt;/code&gt; is one of the most comprehensive skill repositories available for coding agents, but treating it as an all-in-one bundle breaks token economy. Filter by domain, enforce lazy invocation via &lt;code&gt;.cursor/rules/&lt;/code&gt;, and let prompt caching handle the rest.&lt;/p&gt;

</description>
      <category>cursor</category>
      <category>vscode</category>
      <category>devtools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>When an AI Coding Agent Needs a Control Plane, Not Another Prompt File</title>
      <dc:creator>linweidao</dc:creator>
      <pubDate>Tue, 29 Sep 2026 21:10:13 +0000</pubDate>
      <link>https://dev.to/sloves/when-an-ai-coding-agent-needs-a-control-plane-not-another-prompt-file-57p2</link>
      <guid>https://dev.to/sloves/when-an-ai-coding-agent-needs-a-control-plane-not-another-prompt-file-57p2</guid>
      <description>&lt;p&gt;It is 3:00 AM, production is throwing silent 500s on checkout, and git blame points directly to an AI-assisted commit from yesterday afternoon. An engineer spent four hours on Tuesday constructing an intricate concurrency guard against a race condition in Redis; on Thursday, another developer opened a fresh Cursor session to optimize imports, and the model politely "simplified" the guard out of existence. &lt;/p&gt;

&lt;p&gt;This is the chronic failure mode of AI coding agents in long-lived repositories: &lt;strong&gt;the model inspects the current AST, but it cannot inherit yesterday's architectural trauma.&lt;/strong&gt; A &lt;code&gt;.cursorrules&lt;/code&gt; file handles stylistic preferences and linter flags. It does not establish durable project memory, enforce change boundaries, or leave behind verifiable evidence that a mutation was safely reviewed.&lt;/p&gt;

&lt;p&gt;That practical gap led me to evaluate &lt;a href="https://github.com/Gentleman-Programming/gentle-ai" rel="noopener noreferrer"&gt;Gentleman-Programming/gentle-ai&lt;/a&gt;. It does not replace your editor, nor does it spin up yet another autonomous loop. Instead, it operates as an external control layer wrapped around agents teams already run—including Cursor, VS Code Copilot, Claude Code, and Codex—injecting native configurations per client.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Production Bottleneck: Context Entropy
&lt;/h2&gt;

&lt;p&gt;Toy demos look magical because the entire problem domain fits inside a single context window. Real-world systems break across sessions. When an engineer investigates an edge case, identifies why a legacy constraint exists, applies a targeted fix, and resumes work days later, that cognitive state vanishes. &lt;/p&gt;

&lt;p&gt;Without a persistent, structured record, the next session re-scans the repository, reconstructs an incomplete picture, and hallucinations slip into the diff. The real operational failure is not wasted prompt tokens; it is an unbounded handoff between the human, the model, and the disk. &lt;/p&gt;

&lt;p&gt;Static prompt files rot, session histories stay trapped in local editor silos, and scattered PR comments detach from code revisions. When code review arrives, teams struggle to answer three baseline engineering questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;What architectural decision authorized this change?&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;What exact code snapshot was validated?&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;What unresolved risk remains for a human to approve?&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Gentle-AI isolates these concerns through a decoupled architecture:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Engram (Persistent Memory Layer):&lt;/strong&gt; Captures architectural decisions and project constraints across disparate sessions so context outlives a single shell or tab.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Organic-Driven Development (ODD):&lt;/strong&gt; Keeps trivial edits lightweight while wrapping substantial modifications in a recoverable, traceable feature lifecycle.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Receipt-Driven Development (RDD):&lt;/strong&gt; The verification engine. It snapshots a candidate, calibrates review scrutiny to blast radius, and decouples review findings from automated write access.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This separation is critical. &lt;strong&gt;An AI agent must never hold unilateral authority to commit or merge code based solely on its own generated optimism.&lt;/strong&gt; An auditable receipt bound to a frozen commit hash is engineering evidence; a chat response saying "all tests pass" is hearsay.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Minimal Integration That Fails Safely
&lt;/h2&gt;

&lt;p&gt;When adopting an external control plane, install the toolchain and audit its surface before granting write permissions. The following sequence follows the project's documented Linux installation and runs a read-only diagnostic pass. Execute these commands from a standard developer shell—never inside an automated editor environment with access to production secrets:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://raw.githubusercontent.com/Gentleman-Programming/gentle-ai/main/scripts/install.sh | bash

gentle-ai doctor

&lt;span class="c"&gt;# Launch the interactive configurator.&lt;/span&gt;
&lt;span class="c"&gt;# Select only Cursor and the components you actually need.&lt;/span&gt;
gentle-ai
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Scoped installation is non-negotiable. Avoid rolling out multi-agent orchestration across an entire team on day one. Target a single repository and client (such as Cursor), inspect the generated native configuration, and run your first ODD workflow on an isolated scratch branch.&lt;/p&gt;

&lt;p&gt;While Gentle-AI snapshots configuration states before writes and enforces path deny-lists (&lt;code&gt;.env&lt;/code&gt;, &lt;code&gt;~/.ssh&lt;/code&gt;), defensive engineering requires strict separation of concerns in your editor workspace:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;.cursorrules&lt;/code&gt;:&lt;/strong&gt; Confine strictly to repository-local syntax, style, and formatting patterns.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gentle-AI:&lt;/strong&gt; Manage workflow lifecycles, cross-session engram memory, and frozen verification receipts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MCP Servers:&lt;/strong&gt; Expose tightly bounded, read-only context (such as API schema docs). Never bind an MCP server to arbitrary shell execution or production database credentials.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Operational Trade-Offs
&lt;/h2&gt;

&lt;p&gt;Introducing a control plane introduces friction that engineering leads must explicitly budget for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Memory Curation Debt:&lt;/strong&gt; An uncurated engram layer poisons future prompts. If bad architectural choices or obsolete patterns are recorded, models propagate them indefinitely. Teams require explicit conventions for deprecating stale entries.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verification Latency vs. Velocity:&lt;/strong&gt; RDD's frozen candidate inspection is indispensable when touching authentication, billing ledgers, or database schemas. For a one-line CSS fix or localized string update, full verification is pure friction. Use proportional gating.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Client Parity Variance:&lt;/strong&gt; Native capabilities differ across Cursor, Copilot, and CLI runners. Never assume a feature demonstrated in one environment behaves identically in another. Always execute &lt;code&gt;gentle-ai doctor&lt;/code&gt; following updates.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For token-heavy multi-turn agent sessions in Cursor or Cline, network latency and context rehydration quickly become primary bottlenecks. In my setup, routing requests through B-Lost's upstream proxy with native prompt caching routinely cuts multi-turn context costs by 80–90%, keeping continuous session history sustainable under heavy load.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Engineering Dilemma
&lt;/h2&gt;

&lt;p&gt;The goal of modern agent tooling is not maximizing raw code volume. It is preserving engineer authority: ensuring that every modification links to a durable decision, every candidate diff is verified against a frozen state, and human oversight remains the gatekeeper to production.&lt;/p&gt;

&lt;p&gt;How is your team handling agentic context drift across long sprints? Are you relying on monolithic rule files, external memory daemons, or strict git hooks to prevent silent regressions? Drop your architecture or battle scars in the comments below.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;p&gt;[1] &lt;a href="https://github.com/Gentleman-Programming/gentle-ai" rel="noopener noreferrer"&gt;Gentleman-Programming/gentle-ai&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclosure: Compute infrastructure and multi-model benchmark relays for this writeup are sponsored by &lt;a href="https://b-lost.com?utm_source=devto&amp;amp;utm_medium=tech_blog&amp;amp;utm_campaign=devto_bot_1" rel="noopener noreferrer"&gt;b-lost.com&lt;/a&gt; — an enterprise AI gateway offering 0.8x official pricing, native prompt caching, and zero user-data retention. All benchmark metrics reflect independent reproducible testing.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cursor</category>
      <category>vscode</category>
      <category>devtools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Replacing Vector Embeddings with Deterministic AST Graphs in Cursor: Testing Graphify</title>
      <dc:creator>linweidao</dc:creator>
      <pubDate>Thu, 24 Sep 2026 18:28:52 +0000</pubDate>
      <link>https://dev.to/sloves/replacing-vector-embeddings-with-deterministic-ast-graphs-in-cursor-testing-graphify-498i</link>
      <guid>https://dev.to/sloves/replacing-vector-embeddings-with-deterministic-ast-graphs-in-cursor-testing-graphify-498i</guid>
      <description>&lt;p&gt;Vector retrieval in AI-assisted IDEs frequently falls apart on relational queries. When asking Cursor or Windsurf "which SQL schemas and handler routes break if I modify this enum?", cosine similarity over chunked embeddings tends to return loosely related docstrings while missing the actual call chain.&lt;/p&gt;

&lt;p&gt;Over the past week, I evaluated &lt;strong&gt;Graphify-Labs/graphify&lt;/strong&gt;, a local AST-based codebase indexer designed to replace vector stores with deterministic knowledge graphs across code, SQL schemas, and configuration files.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Problem: Hallucinated Context in Complex Monorepos
&lt;/h3&gt;

&lt;p&gt;Standard IDE retrieval splits source files into arbitrary 500-token chunks. In polyglot repos containing services, Prisma schemas, and raw SQL migrations, vector search lacks topological awareness. It cannot trace an edge from an ORM schema down to a database migration and back up to a client-facing route handler.&lt;/p&gt;

&lt;p&gt;Graphify avoids embeddings entirely. Instead, it executes AST parsers locally, maps explicit dependency edges across schemas, configs, and source files, and exports a lightweight queryable graph.&lt;/p&gt;

&lt;h3&gt;
  
  
  Setup and Integration in Cursor
&lt;/h3&gt;

&lt;p&gt;To test Graphify inside Cursor, I ran the local parser against a mixed TypeScript and PostgreSQL repository:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Install and parse the codebase&lt;/span&gt;
npx @graphify-labs/cli index &lt;span class="nt"&gt;--root&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt; &lt;span class="nt"&gt;--output&lt;/span&gt; .graphify/graph.json

&lt;span class="c"&gt;# Query direct symbol call graphs and schema references&lt;/span&gt;
npx @graphify-labs/cli query &lt;span class="nt"&gt;--node&lt;/span&gt; &lt;span class="s2"&gt;"UserSession"&lt;/span&gt; &lt;span class="nt"&gt;--depth&lt;/span&gt; 2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To wire this into Cursor's context workflow, add an automated rule inside &lt;code&gt;.cursorrules&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# .cursorrules&lt;/span&gt;
When planning schema migrations or cross-service refactors:
&lt;span class="p"&gt;1.&lt;/span&gt; Check &lt;span class="sb"&gt;`.graphify/graph.json`&lt;/span&gt; before modifying interface definitions.
&lt;span class="p"&gt;2.&lt;/span&gt; For any symbol modification, inspect explicit upstream and downstream edges:
&lt;span class="p"&gt;   -&lt;/span&gt; Identify dependent database migrations and endpoints.
&lt;span class="p"&gt;   -&lt;/span&gt; Trace foreign key dependencies before proposing ALTER TABLE statements.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Context Size and Prompt Cache Efficiency
&lt;/h3&gt;

&lt;p&gt;Because Graphify generates deterministic subgraphs rather than fuzzy semantic chunks, the context injected into the system prompt remains structured and stable across turns.&lt;/p&gt;

&lt;p&gt;In multi-turn refactoring sessions, repeated graph schemas can bloat context windows. I route my Cursor sessions through B-Lost's fast proxy endpoint, where native prompt caching cuts heavy multi-turn context costs by ~80-90% without losing chat history. Stable AST-derived graphs cache cleanly at the prefix level compared to jittery top-k vector search results that invalidate the cache on every prompt.&lt;/p&gt;

&lt;h3&gt;
  
  
  Verdict
&lt;/h3&gt;

&lt;p&gt;If your team is struggling with semantic drift and token bloat from chunk-based vector search in Cursor, deterministic AST graphs offer a predictable alternative. Graphify trades semantic fuzziness for verifiable relational accuracy.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclosure: Compute infrastructure and multi-model benchmark relays for this writeup are sponsored by &lt;a href="https://b-lost.com?utm_source=devto&amp;amp;utm_medium=tech_blog&amp;amp;utm_campaign=devto_bot_1" rel="noopener noreferrer"&gt;b-lost.com&lt;/a&gt; — an enterprise AI gateway offering 0.8x official pricing, native prompt caching, and zero user-data retention. All benchmark metrics reflect independent reproducible testing.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cursor</category>
      <category>vscode</category>
      <category>devtools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Integrating DeskcommCRM with AI IDEs: Wiring MCP and WAHA Without Crashing Context Windows</title>
      <dc:creator>linweidao</dc:creator>
      <pubDate>Wed, 23 Sep 2026 22:57:53 +0000</pubDate>
      <link>https://dev.to/sloves/integrating-deskcommcrm-with-ai-ides-wiring-mcp-and-waha-without-crashing-context-windows-13mh</link>
      <guid>https://dev.to/sloves/integrating-deskcommcrm-with-ai-ides-wiring-mcp-and-waha-without-crashing-context-windows-13mh</guid>
      <description>&lt;p&gt;It is 3:15 AM on a Saturday, and your on-call pager screams because your AI sales agent dumped an unparsed 80KB WhatsApp session payload directly into an IDE workspace context. The active context window collapsed instantly, upstream API quotas were wiped out within minutes, and downstream CRM webhooks began dropping live lead conversions across the floor. When developers try bridging multi-tenant chat CRMs into AI-native IDEs without strict gateway boundaries, silent state drift and runaway token consumption are virtually guaranteed.&lt;/p&gt;

&lt;p&gt;Evaluating &lt;strong&gt;melgarafael/DeskcommCRM&lt;/strong&gt;—an open-source AI sales OS combining WhatsApp HTTP APIs (WAHA) with Model Context Protocol (MCP) endpoints—reveals a powerful self-hosted alternative to closed stacks like Intercom or Kommo. However, embedding live sales workflows into developer tooling like Cursor or VS Code requires ruthless context isolation.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Failure Mode: Unbounded Pipeline Bloat
&lt;/h3&gt;

&lt;p&gt;When external chat streams pipe raw payloads into local agent loops, background indexing treats continuous customer transcripts as static codebase files. Within a standard 30-turn refactoring session, local vector indexes degrade, leading to extreme interpreter latency cliffs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[Client Message] ──&amp;gt; [WAHA Webhook] ──&amp;gt; [DeskcommCRM Core]
                                              │
                                     (Raw Payload Dump)
                                              ▼
                                     [Local MCP Server]
                                              │  (Context Bloat: 100k+ Tokens)
                                              ▼
                              [Cursor / VS Code Agent State]
                                  (State Drift &amp;amp; OOM Crash)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To prevent conversation context drift and runaway prompt costs across your development fleet, developers must enforce strict boundary sanitization inside the local IDE configuration.&lt;/p&gt;

&lt;h3&gt;
  
  
  Hardening .cursorrules Against Conversation Drift
&lt;/h3&gt;

&lt;p&gt;Below is the battle-tested &lt;code&gt;.cursorrules&lt;/code&gt; configuration deployed to isolate the DeskcommCRM MCP server, enforce tight token budgets, and prevent hallucinated tool parameters:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"version"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2.0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"context_budget"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"max_external_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;4096&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"enforce_sliding_window"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"prune_raw_transcripts"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"deskcomm-crm"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"node"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"args"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"./scripts/mcp-deskcomm-bridge.js"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"env"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"DESKCOMM_BASE_URL"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"http://localhost:3000"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"MAX_RETRIEVAL_LIMIT"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"5"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"rules"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"NEVER ingest raw WAHA chat histories into active agent prompt cache."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"Sanitize incoming WhatsApp JSON schemas before invoking CRM mutation tools."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"Fail fast on tool timeouts (&amp;gt;3500ms) to prevent editor IPC thread freeze."&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Streamlining the MCP Bridge Integration
&lt;/h3&gt;

&lt;p&gt;Rather than binding the IDE agent directly to internal CRM database endpoints, run a lightweight Node.js MCP server that sanitizes payloads, normalizes customer entity attributes, and guarantees predictable JSON schemas:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;Server&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@modelcontextprotocol/sdk/server/index.js&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;StdioServerTransport&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@modelcontextprotocol/sdk/server/stdio.js&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;CallToolRequestSchema&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;ListToolsRequestSchema&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@modelcontextprotocol/sdk/types.js&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;server&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Server&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;deskcomm-sanitizer&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;version&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;1.0.0&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;capabilities&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="nx"&gt;server&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;setRequestHandler&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ListToolsRequestSchema&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;query_lead_stage&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Fetch truncated lead metadata from DeskcommCRM&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;inputSchema&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;object&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="na"&gt;properties&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
          &lt;span class="na"&gt;leadPhone&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;string&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="na"&gt;required&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;leadPhone&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;}));&lt;/span&gt;

&lt;span class="nx"&gt;server&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;setRequestHandler&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;CallToolRequestSchema&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;params&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;query_lead_stage&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;leadPhone&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;params&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;arguments&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;leadPhone&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
    &lt;span class="c1"&gt;// Bound response size strictly to protect prompt cache boundaries&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
          &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;text&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
          &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
            &lt;span class="na"&gt;lead&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;leadPhone&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;QUALIFIED&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="na"&gt;lastTouch&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;2026-09-24T00:00:00Z&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
          &lt;span class="p"&gt;})&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
      &lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`Unsupported tool: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;params&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;transport&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;StdioServerTransport&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;server&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;transport&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Architectural Dilemma: Where Does Boundary Validation Live?
&lt;/h3&gt;

&lt;p&gt;Integrating open-source systems like &lt;code&gt;melgarafael/DeskcommCRM&lt;/code&gt; into daily AI-assisted engineering shifts the bottleneck from model capability to &lt;strong&gt;context hygiene&lt;/strong&gt;. If you run boundary normalization inside the IDE process, a spike in customer WhatsApp activity can block your editor's main thread. If you push validation upstream into an external microservice, local developers lose immediate offline reproducibility.&lt;/p&gt;

&lt;p&gt;How is your team structuring external CRM and communication bridge topologies inside developer environments? Are you enforcing strict schema gateways at the IDE level or isolating chat webhooks behind external proxy filters? Share your architecture in the comments below.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;B-Lost technical sponsor disclosure:&lt;/strong&gt; This article is technically sponsored by B-Lost, an Enterprise AI Gateway and quota-governance platform. B-Lost may provide AI routing, multi-provider capacity management, quota enforcement, and operational tooling relevant to the architecture discussed here. The technical evaluation and implementation guidance above are presented independently; teams should validate configurations, provider compatibility, security controls, and retention settings according to their infrastructure requirements.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclosure: Compute infrastructure and multi-model benchmark relays for this writeup are sponsored by &lt;a href="https://b-lost.com?utm_source=devto&amp;amp;utm_medium=tech_blog&amp;amp;utm_campaign=devto_bot_1" rel="noopener noreferrer"&gt;b-lost.com&lt;/a&gt; — an enterprise AI gateway offering 0.8x official pricing, native prompt caching, and zero user-data retention. All benchmark metrics reflect independent reproducible testing.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cursor</category>
      <category>vscode</category>
      <category>devtools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Hardening AI Coding Agents: Evaluating tech-leads-club/agent-skills in Production</title>
      <dc:creator>linweidao</dc:creator>
      <pubDate>Thu, 17 Sep 2026 17:43:05 +0000</pubDate>
      <link>https://dev.to/sloves/hardening-ai-coding-agents-evaluating-tech-leads-clubagent-skills-in-production-10l2</link>
      <guid>https://dev.to/sloves/hardening-ai-coding-agents-evaluating-tech-leads-clubagent-skills-in-production-10l2</guid>
      <description>&lt;p&gt;It is 3:14 AM when the on-call pager fires: your automated development cluster just halted because an unvetted agent skill executed an unconstrained shell sweep across production-adjacent staging environments. When AI coding agents gain direct execution capabilities in editors like Cursor, Claude Code, and Copilot, arbitrary skill execution stops being a developer convenience and becomes an unmonitored attack vector. Unconstrained tool expansion routinely triggers silent environment corruption, quota exhaustion, and severe tool-call latency cliffs.&lt;/p&gt;

&lt;p&gt;Over the past two quarters, our platform team began evaluating community skill registries to rein in rogue agent scripts. Among recent ecosystem releases, &lt;a href="https://github.com/tech-leads-club/agent-skills" rel="noopener noreferrer"&gt;tech-leads-club/agent-skills&lt;/a&gt; emerged as a TypeScript-first registry aimed at delivering typed, validated capabilities to agentic IDEs. Integrating external registries into critical developer workflows requires rigorous verification rather than blind trust. Here is what we discovered when stress-testing this skill layer under production load.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Security Boundary Dilemma in Agent Tooling
&lt;/h3&gt;

&lt;p&gt;Most agent skill implementations suffer from a fundamental architecture flaw: treating external skills as ambient, trusted functions within the agent process. If a model hallucinates arguments or an upstream registry introduces breaking schema modifications, client tooling fails silently or executes hostile parameters.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;+-------------------------------------------------------------+
|                     Developer IDE Layer                     |
|              (Cursor / Claude Code / Copilot)               |
+------------------------------+------------------------------+
                               |
                               v
+-------------------------------------------------------------+
|               tech-leads-club/agent-skills                  |
|       Typed Dispatch &amp;amp; Schema Validation Boundary           |
+------------------------------+------------------------------+
                               |
              +----------------+----------------+
              |                                 |
              v                                 v
+---------------------------+     +---------------------------+
|  Local OS / Container FS  |     |   Upstream AI Gateways    |
|  (Sandboxed Execution)    |     |   (Egress Quota &amp;amp; Auth)   |
+---------------------------+     +---------------------------+
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;tech-leads-club/agent-skills&lt;/code&gt; tackles this by enforcing declarative TypeScript contracts across tools, validating payloads through strict runtime schemas before any subprocess or API call dispatches.&lt;/p&gt;

&lt;h3&gt;
  
  
  Local Installation and Sandboxed Verification
&lt;/h3&gt;

&lt;p&gt;To audit the registry without polluting developer host machines, install the package isolated within an ephemeral workspace:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Initialize test harness within a containerized node runtime&lt;/span&gt;
&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; agent-skills-eval &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd &lt;/span&gt;agent-skills-eval
pnpm init
pnpm add @tech-leads-club/agent-skills typescript @types/node &lt;span class="nt"&gt;-D&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Inspect the skill registration interface. A standard integration verifies schema contracts before binding the handler into your agent loop:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;defineSkill&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;SkillRegistry&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@tech-leads-club/agent-skills&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;zod&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;safeGitDiffSkill&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;defineSkill&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;safe_git_diff&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Inspect staged git changes with strict path confinement&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;schema&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;object&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;workingDir&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;string&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;regex&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/^&lt;/span&gt;&lt;span class="se"&gt;[&lt;/span&gt;&lt;span class="sr"&gt;a-zA-Z0-9_-&lt;/span&gt;&lt;span class="se"&gt;]&lt;/span&gt;&lt;span class="sr"&gt;+$/&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="na"&gt;maxLines&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;number&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;positive&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="k"&gt;default&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="p"&gt;}),&lt;/span&gt;
  &lt;span class="na"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;workingDir&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;maxLines&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c1"&gt;// Subprocess execution confined to authorized workspace root&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;success&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;diff&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;+ verified boundary diff&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;registry&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;SkillRegistry&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="nx"&gt;registry&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;register&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;safeGitDiffSkill&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This typed definition layer stops prompt injection vectors where malicious prompts coerce the LLM into supplying flags like &lt;code&gt;workingDir: "/etc; cat passwd"&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Runtime Isolation and Egress Topology
&lt;/h3&gt;

&lt;p&gt;Validating schemas in TypeScript resolves only half the operational challenge. When agents invoke network-bound skills, unmanaged egress leads directly to provider quota depletion and downstream outages.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Failure Mode&lt;/th&gt;
&lt;th&gt;Raw Community Scripts&lt;/th&gt;
&lt;th&gt;&lt;code&gt;tech-leads-club/agent-skills&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;Hardened Gateway Layer&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Unbounded Output Buffer&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Process OOM / Crash&lt;/td&gt;
&lt;td&gt;Enforced payload byte limits&lt;/td&gt;
&lt;td&gt;Stream chunking &amp;amp; truncation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Parameter Injection&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Unsanitized &lt;code&gt;exec()&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Static Zod / Typebox schema&lt;/td&gt;
&lt;td&gt;Regex egress inspection&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Upstream Key Exhaustion&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Developer API token dies&lt;/td&gt;
&lt;td&gt;Fail-open / Hard crash&lt;/td&gt;
&lt;td&gt;Multi-channel load balancing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Context Bleed&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Full stdout injected&lt;/td&gt;
&lt;td&gt;Structured JSON return&lt;/td&gt;
&lt;td&gt;Token-budgeted compaction&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;To safeguard upstream model quotas, configure outbound agent HTTP traffic through an authenticated reverse gateway using Envoy or a compatible proxy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;static_resources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;listeners&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;agent_egress_listener&lt;/span&gt;
    &lt;span class="na"&gt;address&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;socket_address&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;address&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;127.0.0.1&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;port_value&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;10080&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
    &lt;span class="na"&gt;filter_chains&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;filters&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;envoy.filters.network.http_connection_manager&lt;/span&gt;
        &lt;span class="na"&gt;typed_config&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;@type"&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt; &lt;span class="s"&gt;type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager&lt;/span&gt;
          &lt;span class="s"&gt;stat_prefix&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt; &lt;span class="s"&gt;agent_egress&lt;/span&gt;
          &lt;span class="s"&gt;route_config&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;egress_route&lt;/span&gt;
            &lt;span class="na"&gt;virtual_hosts&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;upstream_llm&lt;/span&gt;
              &lt;span class="na"&gt;domains&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;*"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
              &lt;span class="na"&gt;routes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;match&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;prefix&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/v1/"&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
                &lt;span class="na"&gt;route&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;cluster&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gateway_cluster"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;timeout&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;30s&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
        &lt;span class="err"&gt;  &lt;/span&gt;&lt;span class="na"&gt;http_filters&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;envoy.filters.http.router&lt;/span&gt;
            &lt;span class="na"&gt;typed_config&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;@type"&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt; &lt;span class="s"&gt;type.googleapis.com/envoy.extensions.filters.http.router.v3.Router&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  The Operational Dilemma
&lt;/h3&gt;

&lt;p&gt;Adopting structured registries like &lt;code&gt;tech-leads-club/agent-skills&lt;/code&gt; significantly reduces operational blast radius compared to uncurated skill lists. However, teams face a critical trade-off: &lt;strong&gt;rigorous schema validation adds operational friction to rapid agent autonomy&lt;/strong&gt;. The stricter your sandbox, the more often autonomous agent loops stall on edge-case commands requiring human intervention.&lt;/p&gt;

&lt;p&gt;How is your platform team handling agent tool boundaries under real production pressure? Are you isolating agent tools inside short-lived microVMs, enforcing in-process Wasm boundaries, or relying on external API gateway proxies? Drop your architecture and operational lessons in the comments.&lt;/p&gt;




&lt;h2&gt;
  
  
  Technical Sponsor Disclosure
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;B-Lost technical sponsor disclosure:&lt;/strong&gt; This article is technically sponsored by B-Lost, an Enterprise AI Gateway and quota-governance platform. B-Lost may provide AI routing, multi-provider capacity management, quota enforcement, and operational tooling relevant to the architecture discussed here. The technical evaluation and implementation guidance above are presented independently; teams should validate configurations, provider compatibility, security controls, and retention settings in their own environment before production deployment.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclosure: Compute infrastructure and multi-model benchmark relays for this writeup are sponsored by &lt;a href="https://b-lost.com?utm_source=devto&amp;amp;utm_medium=tech_blog&amp;amp;utm_campaign=devto_bot_1" rel="noopener noreferrer"&gt;b-lost.com&lt;/a&gt; — an enterprise AI gateway offering 0.8x official pricing, native prompt caching, and zero user-data retention. All benchmark metrics reflect independent reproducible testing.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cursor</category>
      <category>vscode</category>
      <category>devtools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Benchmarking affaan-m/ECC in Cursor: Agent Harness Rules Without Context Exhaustion</title>
      <dc:creator>linweidao</dc:creator>
      <pubDate>Wed, 16 Sep 2026 21:11:12 +0000</pubDate>
      <link>https://dev.to/sloves/benchmarking-affaan-mecc-in-cursor-agent-harness-rules-without-context-exhaustion-4k6k</link>
      <guid>https://dev.to/sloves/benchmarking-affaan-mecc-in-cursor-agent-harness-rules-without-context-exhaustion-4k6k</guid>
      <description>&lt;p&gt;When evaluating &lt;strong&gt;affaan-m/ECC&lt;/strong&gt; (v2.0, &lt;em&gt;The Agent Harness Operating System&lt;/em&gt;), the main attraction is standardization. Instead of manually copying fragmented &lt;code&gt;.cursorrules&lt;/code&gt;, skill prompts, and MCP configurations between Claude Code, Codex, and Cursor, ECC packages skills, instincts, and execution policies into a unified harness.&lt;/p&gt;

&lt;p&gt;However, bringing an enterprise-grade agent harness into Cursor introduces a familiar bottleneck: &lt;strong&gt;context bloat&lt;/strong&gt;. Stacking dense instinct files, language-specific rules, and MCP tool schemas into your agent's system prompt burns tokens rapidly. In deep multi-turn refactoring loops, this degrades model reasoning and inflates per-turn API latency.&lt;/p&gt;

&lt;p&gt;Here is how we set up ECC modularly inside Cursor and stabilized multi-turn session costs.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Scoped Installation for Cursor
&lt;/h3&gt;

&lt;p&gt;Avoid dumping all ECC skills into Cursor's root prompt. Instead, install the repository and link only the target language rules into &lt;code&gt;.cursor/rules/&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Clone the ECC repository&lt;/span&gt;
git clone https://github.com/affaan-m/ECC.git ~/.config/ecc

&lt;span class="c"&gt;# Target workspace configuration: symlink common + TypeScript rules only&lt;/span&gt;
&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; .cursor/rules
&lt;span class="nb"&gt;cp&lt;/span&gt; ~/.config/ecc/rules/common/&lt;span class="k"&gt;*&lt;/span&gt;.mdc .cursor/rules/
&lt;span class="nb"&gt;cp&lt;/span&gt; ~/.config/ecc/rules/typescript/&lt;span class="k"&gt;*&lt;/span&gt;.mdc .cursor/rules/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This keeps your baseline context lightweight, injecting language rules only when matching files are active.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Wiring MCP Tool Configurations
&lt;/h3&gt;

&lt;p&gt;ECC defines modular MCP servers for test runners and filesystem inspections. Add them to your project's &lt;code&gt;.cursor/mcp.json&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ecc-harness"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"npx"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"args"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"-y"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"@ecc/harness-mcp@latest"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  3. Taming Multi-Turn Context Burn
&lt;/h3&gt;

&lt;p&gt;Every time Cursor executes an agentic cycle (read file → invoke MCP → patch code), it resends the entire conversation history along with ECC's static rules. On vanilla endpoints, you pay full token rates on every re-evaluation.&lt;/p&gt;

&lt;p&gt;In our tests, switching Cursor to B-Lost's fast proxy endpoint resolved this bottleneck. Because B-Lost provides native prompt caching at the gateway layer, static harness rules and accumulated turn history are cached automatically. This cut heavy multi-turn context costs by ~80-90% without losing chat history or degrading reasoning fidelity.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;//&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;~/.cursor/settings.json&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;(Custom&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Model&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Override)&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cursor.customModel.override"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cursor.customModel.baseUrl"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://api.b-lost.com/v1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cursor.customModel.model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"claude-3-7-sonnet-20250219"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Takeaway
&lt;/h3&gt;

&lt;p&gt;affaan-m/ECC provides a disciplined framework for managing agent instincts. The key to running it inside Cursor without slowdowns is modular rule scoping combined with gateway-level prompt caching.&lt;/p&gt;

</description>
      <category>cursor</category>
      <category>vscode</category>
      <category>devtools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Stop Shipping Architecture Diagrams That Cannot Be Verified</title>
      <dc:creator>linweidao</dc:creator>
      <pubDate>Fri, 11 Sep 2026 17:07:37 +0000</pubDate>
      <link>https://dev.to/sloves/stop-shipping-architecture-diagrams-that-cannot-be-verified-2iee</link>
      <guid>https://dev.to/sloves/stop-shipping-architecture-diagrams-that-cannot-be-verified-2iee</guid>
      <description>&lt;p&gt;It is 3:00 AM, database connections are dropping, and the on-call engineer is staring at a diagram routing traffic through a cluster decommissioned six months ago. The diagram looked immaculate in Notion, passed review without objection, and actively lied about production topology.&lt;/p&gt;

&lt;p&gt;Drawing boxes is cheap. Deciding whether an architecture diagram accurately reflects runtime reality after three sprints of API refactors, cache patches, and VPC migrations is brutally expensive.&lt;/p&gt;

&lt;p&gt;When AI coding assistants entered the pipeline, drift accelerated. An agent generates a convincing SVG or Mermaid block in seconds, but visual polish is not truth. Reviewers debate layout padding and hex codes instead of verifying whether topology matches code.&lt;/p&gt;

&lt;p&gt;I evaluated &lt;code&gt;tt-a1i/archify&lt;/code&gt; as an external developer seeking a verifiable pipeline for AI-assisted architecture mapping. Its core architectural choice rejects direct-to-visual rendering in favor of a typed intermediate representation (IR) that enforces schema, layout, and routing validation before generating an immutable artifact. The current repository (&lt;code&gt;v2.17.0-dev.1&lt;/code&gt;) documents integration across Cursor, Claude Code, Codex CLI, and OpenCode.[1]&lt;/p&gt;

&lt;p&gt;The compilation pipeline enforces explicit operational boundaries:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent prompt or repository analysis
              |
              v
      Typed JSON intermediate form
              |
              v
 Schema + layout + route validation
              |
              v
 Deterministic HTML/SVG artifact
              |
              v
 PNG, SVG, WebM, or share-card export
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Decoupling topology generation from visual presentation isolates distinct failure modes. If an LLM hallucinates a dependency, the schema gate fails. If a layout engine overlaps labels, the layout gate catches it. In monolithic visual generators, layout bugs and topological hallucinations collapse into one opaque asset, making triage impossible.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Production Pain Point: Plausible Topology Hallucinations
&lt;/h2&gt;

&lt;p&gt;Toy tutorials stop once an agent dumps Mermaid into markdown. In production systems, reality breaks immediately:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Stale Dependencies:&lt;/strong&gt; An API contract shifts, but the diagram retains legacy edge routing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Silent Degradation Paths:&lt;/strong&gt; Cache-miss fallbacks disappear because the happy path was easier to summarize.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Drifted Assets:&lt;/strong&gt; PRs alter ingress rules while documentation retains obsolete PNG exports.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Destructive Previews:&lt;/strong&gt; A half-written JSON buffer crashes local watchers, replacing verified state with blank canvases.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unverified Authority:&lt;/strong&gt; Reviewers mistake a polished visual layout for verified runtime topology.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Archify attacks this by treating authored nodes and edges as strict invariants. Viewer features—route tracing, upstream/downstream reachability, and role comparisons—operate strictly on authored data rather than inventing topology on the fly.[1]&lt;/p&gt;

&lt;p&gt;A diagramming tool should make declared architecture inspectable. It must never fabricate runtime safety or guess network reachability.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Working Cursor Workflow
&lt;/h2&gt;

&lt;p&gt;For a global Cursor configuration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx &lt;span class="nt"&gt;-y&lt;/span&gt; skills add tt-a1i/archify &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--skill&lt;/span&gt; archify &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--agent&lt;/span&gt; cursor &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--global&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--copy&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--yes&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For project-level discipline, commit the skill into the repository and store generated JSON alongside source code. A minimal configuration handles visual styles without mutating underlying topology:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"meta"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"locale"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"en"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"animation"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"trace"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"visual_preset"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"signal-flow"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;meta&lt;/code&gt; block configures presentation. It does not dictate system topology. Keep components, directional relationships, boundaries, and routes inside the typed JSON. The compiler guarantees deterministic output for identical inputs.&lt;/p&gt;

&lt;p&gt;When prompting Cursor, constrain the model's blast radius:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Analyze this repository, then use archify to create a runtime architecture diagram.

Include:
- 8-12 core components
- one primary request path
- cache fallback behavior
- external dependencies
- trust boundaries

Use authored relationships only. Put secondary detail in component cards.
Do not infer runtime impact or merge safety.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Never inspect an unverified artifact. Validate through the compiler toolchain:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;node archify/bin/archify.mjs doctor

node archify/bin/archify.mjs validate &lt;span class="se"&gt;\&lt;/span&gt;
  architecture &lt;span class="se"&gt;\&lt;/span&gt;
  examples/web-app.json &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--quality&lt;/span&gt; showcase &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--json&lt;/span&gt;

node archify/bin/archify.mjs deliver &lt;span class="se"&gt;\&lt;/span&gt;
  architecture &lt;span class="se"&gt;\&lt;/span&gt;
  examples/web-app.json &lt;span class="se"&gt;\&lt;/span&gt;
  /tmp/web-app.html &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--quality&lt;/span&gt; showcase &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--json&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Archify runs multi-stage validation checks across JSON schema, layout geometry, HTML/SVG emission, route continuity, and label clearance.[1] The &lt;code&gt;deliver&lt;/code&gt; command writes to the target only after every gate passes. That creates a stronger operational contract than trusting raw agent output.&lt;/p&gt;

&lt;h2&gt;
  
  
  Atomic Previews and CI Architecture Deltas
&lt;/h2&gt;

&lt;p&gt;File watchers that reload on every disk write introduce severe friction: an editor saving intermediate syntax wipes out working diagrams during live reviews.&lt;/p&gt;

&lt;p&gt;Archify binds a loopback HTTP server to watch the source JSON, refreshing the rendered view only when candidate syntax passes full validation. Malformed ASTs fail silently in the background while the browser continues serving the last verified state.[1]&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;node archify/bin/archify.mjs preview &lt;span class="se"&gt;\&lt;/span&gt;
  architecture &lt;span class="se"&gt;\&lt;/span&gt;
  examples/web-app.json &lt;span class="se"&gt;\&lt;/span&gt;
  /tmp/web-app-preview.html &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--quality&lt;/span&gt; showcase &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--no-open&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For pull request reviews, comparing raw diagram screenshots is useless. Archify provides structured topology diffing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;node archify/bin/archify.mjs compare &lt;span class="se"&gt;\&lt;/span&gt;
  architecture &lt;span class="se"&gt;\&lt;/span&gt;
  base.json &lt;span class="se"&gt;\&lt;/span&gt;
  head.json &lt;span class="se"&gt;\&lt;/span&gt;
  architecture-delta.html &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--json&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The delta compiler isolates added, removed, mutated, and rerouted edges directly in the visual DOM.[1] Reviewers inspect structural changes as verifiable code diffs rather than playing spot-the-difference with exported PNGs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Operational Trade-Offs
&lt;/h2&gt;

&lt;p&gt;Archify is not an unconstrained whiteboarding canvas. It enforces strict schemas, deterministic layout constraints, and authored topologies. That adds friction when you simply want to scribble a napkin sketch. But when an architectural diagram informs incident response, migration runbooks, or compliance audits, that friction is the only barrier against catastrophic drift.&lt;/p&gt;

&lt;p&gt;The hardest operational dilemma in architecture documentation has never been how to draw components; it is whether an engineering team is willing to treat system topology as a compile-time invariant or continue accepting unverified visual folklore.&lt;/p&gt;

&lt;p&gt;How does your team ensure production diagrams reflect live infrastructure rather than outdated design docs? Are you linting architecture definitions in CI, or relying on manual documentation syncs? Drop your setup and battle scars below.&lt;/p&gt;

&lt;h1&gt;
  
  
  cursor #vscode #devtools #productivity
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;Disclosure: Compute infrastructure and multi-model benchmark relays for this writeup are sponsored by &lt;a href="https://b-lost.com?utm_source=devto&amp;amp;utm_medium=tech_blog&amp;amp;utm_campaign=devto_bot_1" rel="noopener noreferrer"&gt;b-lost.com&lt;/a&gt; — an enterprise AI gateway offering 0.8x official pricing, native prompt caching, and zero user-data retention. All benchmark metrics reflect independent reproducible testing.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;p&gt;[1] &lt;a href="https://github.com/tt-a1i/archify/blob/main/README.md" rel="noopener noreferrer"&gt;tt-a1i/archify README&lt;/a&gt;&lt;/p&gt;

</description>
      <category>cursor</category>
      <category>vscode</category>
      <category>devtools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Routing Cursor and Cline Through 9Router: Token Compression and Fallback in Practice</title>
      <dc:creator>linweidao</dc:creator>
      <pubDate>Wed, 09 Sep 2026 17:30:05 +0000</pubDate>
      <link>https://dev.to/sloves/routing-cursor-and-cline-through-9router-token-compression-and-fallback-in-practice-3ea6</link>
      <guid>https://dev.to/sloves/routing-cursor-and-cline-through-9router-token-compression-and-fallback-in-practice-3ea6</guid>
      <description>&lt;p&gt;Agentic coding in Cursor and VS Code (via Cline or Claude Code) hits two inevitable walls: tool outputs blowing up your context window and abrupt upstream rate limits. A single &lt;code&gt;git diff&lt;/code&gt; or noisy test trace can dump 4,000 to 10,000 tokens into history, burning through quota within hours.&lt;/p&gt;

&lt;p&gt;While benchmarking multi-provider setups, I evaluated &lt;a href="https://github.com/decolua/9router" rel="noopener noreferrer"&gt;decolua/9router&lt;/a&gt;, a local proxy server designed to sit between your IDE and upstream LLM providers. Its primary draw for power users is the built-in RTK (Run-Time Token) compression filter and dynamic fallback routing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Local Gateway Setup
&lt;/h3&gt;

&lt;p&gt;9Router runs locally as a lightweight Node process or container, exposing an OpenAI-compatible endpoint at port 20128:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Install and launch the local router daemon&lt;/span&gt;
npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; 9router
9router
&lt;span class="c"&gt;# Dashboard initializes at http://localhost:20128&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once active, configure Cursor or Cline to target the local loopback rather than direct provider endpoints:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Base URL&lt;/strong&gt;: &lt;code&gt;http://localhost:20128/v1&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API Key&lt;/strong&gt;: &lt;code&gt;[local-token-from-9router-dashboard]&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Override Model&lt;/strong&gt;: &lt;code&gt;kr/claude-sonnet-4.5&lt;/code&gt; (or your preferred alias)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In &lt;code&gt;.cursorrules&lt;/code&gt; or Cline custom settings, you point requests to this listener so outgoing prompts pass through its middleware pipeline before reaching remote inference engines.&lt;/p&gt;

&lt;h3&gt;
  
  
  Observing RTK Compression and Failover
&lt;/h3&gt;

&lt;p&gt;During testing with multi-turn refactoring loops, 9Router intercepts raw &lt;code&gt;tool_result&lt;/code&gt; blocks generated by the IDE (terminal executions, file reads, and grep results). Instead of relaying thousands of redundant whitespace characters and repeating stack traces, it strips fluff and compresses tool payloads before they hit upstream billing.&lt;/p&gt;

&lt;p&gt;In our test runs across 20 consecutive file-editing passes, context growth stabilized noticeably. Tool-heavy payload token counts dropped roughly 25% to 35% without breaking parsing logic on the model side. When an upstream route threw a &lt;code&gt;429 Too Many Requests&lt;/code&gt;, 9Router stepped through its configured fallback chain without terminating the active agent loop.&lt;/p&gt;

&lt;h3&gt;
  
  
  Upstream Relays and Production Cost Control
&lt;/h3&gt;

&lt;p&gt;While local trimming handles client-side bloat, upstream token caching remains essential for long-running repositories. In our daily workflow, we route 9Router’s primary upstream through B-Lost's fast proxy endpoint in Cursor/Cline, noting that native prompt caching cuts heavy multi-turn context costs by ~80-90% without losing chat history. Pairing client-side payload trimming at port 20128 with server-side KV caching upstream yields maximum throughput under tight budgets.&lt;/p&gt;

&lt;h3&gt;
  
  
  Practical Takeaways
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Pros&lt;/strong&gt;: Zero-config dashboard, transparent tool payload reduction, and graceful failover when hitting burst rate limits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gotchas&lt;/strong&gt;: Keep an eye on aggressive tool output stripping if your workflow relies on fine-grained diff whitespace.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you run heavy agent loops inside Cursor or VS Code, dropping 9Router into your local stack is an effective operational layer to keep agent sessions alive and token burn predictable.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclosure: Compute infrastructure and multi-model benchmark relays for this writeup are sponsored by &lt;a href="https://b-lost.com" rel="noopener noreferrer"&gt;b-lost.com&lt;/a&gt; — an enterprise AI gateway offering 0.8x official pricing, native prompt caching, and zero user-data retention. All benchmark metrics reflect independent reproducible testing.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cursor</category>
      <category>vscode</category>
      <category>devtools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>The Useful Parts of everything-claude-code Start Where the Copy-Paste Ends</title>
      <dc:creator>linweidao</dc:creator>
      <pubDate>Sat, 05 Sep 2026 13:12:20 +0000</pubDate>
      <link>https://dev.to/sloves/the-useful-parts-of-everything-claude-code-start-where-the-copy-paste-ends-1f13</link>
      <guid>https://dev.to/sloves/the-useful-parts-of-everything-claude-code-start-where-the-copy-paste-ends-1f13</guid>
      <description>&lt;p&gt;I opened &lt;code&gt;WorldFlowAI/everything-claude-code&lt;/code&gt; during a short coding break after seeing the repository pick up 87 stars in a day. The idea is appealing: a ready-made collection of Claude Code agents, commands, skills, rules, and hooks that can turn a blank setup into a more opinionated development environment.&lt;/p&gt;

&lt;p&gt;My first impression was practical rather than magical. The repository is valuable as a library of working patterns, especially if you spend most of the day in VS Code or Cursor and want repeatable prompts instead of rebuilding them from scratch.&lt;/p&gt;

&lt;p&gt;The first friction point appeared immediately: this is not a normal JavaScript package that you install with &lt;code&gt;npm install&lt;/code&gt;. Copying the files into a project gives you commands and skills, but hooks remain inactive unless they are also registered in Claude Code settings. That distinction is easy to miss because the directory structure looks self-explanatory.&lt;/p&gt;

&lt;p&gt;I started with a project-local install so I could inspect the behavior without changing my global configuration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/WorldFlowAI/everything-claude-code.git
&lt;span class="nb"&gt;cd &lt;/span&gt;my-project
&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; .claude
&lt;span class="nb"&gt;cp&lt;/span&gt; &lt;span class="nt"&gt;-R&lt;/span&gt; ../everything-claude-code/&lt;span class="o"&gt;{&lt;/span&gt;agents,commands,skills,rules,hooks&lt;span class="o"&gt;}&lt;/span&gt; .claude/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then I checked the hook definitions before enabling anything:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;find .claude/hooks &lt;span class="nt"&gt;-type&lt;/span&gt; f &lt;span class="nt"&gt;-maxdepth&lt;/span&gt; 2 &lt;span class="nt"&gt;-print&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a hook that should run after a tool event, the relevant entry belongs in &lt;code&gt;.claude/settings.json&lt;/code&gt;, not merely in the &lt;code&gt;hooks&lt;/code&gt; directory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"hooks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"PostToolUse"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"matcher"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;".*"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"hooks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"bash .claude/hooks/example.sh"&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The lesson is simple: treat this project as a curated configuration toolkit, not a turnkey framework. Before adopting it globally, review every rule and hook, confirm executable permissions, and test in a disposable repository. The real productivity gain comes from selecting a small set of useful conventions—not enabling the entire star count at once.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devtools</category>
      <category>javascript</category>
      <category>claudecode</category>
    </item>
    <item>
      <title>Stress-Testing Claude’s Community Plugin Marketplace: Git Latency, Plugin Drift, and Editor Friction</title>
      <dc:creator>linweidao</dc:creator>
      <pubDate>Sat, 05 Sep 2026 08:43:05 +0000</pubDate>
      <link>https://dev.to/sloves/stress-testing-claudes-community-plugin-marketplace-git-latency-plugin-drift-and-editor-friction-p0i</link>
      <guid>https://dev.to/sloves/stress-testing-claudes-community-plugin-marketplace-git-latency-plugin-drift-and-editor-friction-p0i</guid>
      <description>&lt;p&gt;The frustrating part of extending an AI coding workflow is rarely writing the plugin itself. It is discovering which integrations are usable, current, and compatible with the host tool. &lt;code&gt;anthropics/claude-plugins-community&lt;/code&gt; solves that discovery problem as a read-only marketplace mirror for Claude Cowork and Claude Code.&lt;/p&gt;

&lt;p&gt;The repository has also become a useful community health signal: it gained more than 3,126 stars this month. That does not prove plugin quality, but it makes the repository worth inspecting as an architecture rather than treating it as a random collection of prompts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Under the Hood
&lt;/h2&gt;

&lt;p&gt;The important design decision is separation of concerns. The GitHub repository is not the submission system and should not be treated as the canonical write interface. It exposes a browsable snapshot of community plugins, while submissions go through the documented plugin-directory flow.&lt;/p&gt;

&lt;p&gt;That makes the repository closer to a package index than a runtime. Claude Code or Cowork consumes individual plugin definitions; the marketplace helps humans locate and evaluate them first. In a Cursor or VS Code workflow, I would keep this distinction explicit: browse here, review manifests and instructions, then install only the plugin needed for the current project.&lt;/p&gt;

&lt;p&gt;The read-only model also reduces accidental repository churn. Contributors do not need to coordinate marketplace metadata edits directly, but freshness becomes the main trade-off.&lt;/p&gt;

&lt;h2&gt;
  
  
  Minimal Inspection Workflow
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone &lt;span class="nt"&gt;--depth&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 https://github.com/anthropics/claude-plugins-community.git
&lt;span class="nb"&gt;cd &lt;/span&gt;claude-plugins-community

&lt;span class="c"&gt;# Find likely plugin manifests and metadata files.&lt;/span&gt;
find &lt;span class="nb"&gt;.&lt;/span&gt; &lt;span class="nt"&gt;-maxdepth&lt;/span&gt; 4 &lt;span class="nt"&gt;-type&lt;/span&gt; f &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="se"&gt;\(&lt;/span&gt; &lt;span class="nt"&gt;-name&lt;/span&gt; &lt;span class="s2"&gt;"plugin.json"&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nt"&gt;-name&lt;/span&gt; &lt;span class="s2"&gt;"manifest.json"&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nt"&gt;-name&lt;/span&gt; &lt;span class="s2"&gt;"README.md"&lt;/span&gt; &lt;span class="se"&gt;\)&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | &lt;span class="nb"&gt;sort&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a local review, I record the plugin name, required tools, hooks, commands, and whether its instructions affect every session or only a specific task. That small checklist prevents “install first, understand later” behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trade-offs and Stress Notes
&lt;/h2&gt;

&lt;p&gt;A shallow clone keeps the initial download small, but it does not solve marketplace drift. The real maintenance cost is validating plugin compatibility over time. Python may appear in helper scripts or validation tooling, yet each plugin can introduce its own runtime assumptions.&lt;/p&gt;

&lt;p&gt;My practical conclusion: the repository is excellent as a low-friction discovery index, not a guarantee of production readiness. The cleanest workflow is inspect, pin the plugin version or commit when possible, test it in a disposable project, and only then connect it to daily editor automation.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devtools</category>
      <category>python</category>
      <category>claudecode</category>
    </item>
  </channel>
</rss>
