<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: hao li</title>
    <description>The latest articles on DEV Community by hao li (@haoli).</description>
    <link>https://dev.to/haoli</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4151730%2Fbab6238b-564c-40ce-9e39-15ee06e05f63.png</url>
      <title>DEV Community: hao li</title>
      <link>https://dev.to/haoli</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/haoli"/>
    <language>en</language>
    <item>
      <title>My agent changed `toBe(10)` to `toBe(9)` and called it done. I built hooks to stop that</title>
      <dc:creator>hao li</dc:creator>
      <pubDate>Wed, 30 Sep 2026 18:04:21 +0000</pubDate>
      <link>https://dev.to/haoli/my-agent-changed-tobe10-to-tobe9-and-called-it-done-i-built-hooks-to-stop-that-8c4</link>
      <guid>https://dev.to/haoli/my-agent-changed-tobe10-to-tobe9-and-called-it-done-i-built-hooks-to-stop-that-8c4</guid>
      <description>&lt;p&gt;A few days ago someone on r/ClaudeCode posted the exact failure mode I'd been dreading. Their agent was asked to fix an indexing bug. Instead of fixing the source file, it quietly edited the test — &lt;code&gt;expect(page.items.length).toBe(10)&lt;/code&gt; became &lt;code&gt;toBe(9)&lt;/code&gt; — re-ran the suite, saw green, and reported the refactor complete. "All 14 tests passing" nearly sailed through review.&lt;/p&gt;

&lt;p&gt;Same week, a 136-point thread asked whether Claude Code's unit tests are useless at all. The top-voted answer: &lt;em&gt;"every generated test should be forced to prove it can fail."&lt;/em&gt; And in a separate thread about stopping agents from doing things they shouldn't in production, the line that stuck with me: &lt;em&gt;"the real risk is not a bad answer but a bad action"&lt;/em&gt; — with the follow-up that the boundary has to sit &lt;em&gt;below the prompt layer&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Prompt-level instructions don't survive contact with a determined agent. So I built &lt;code&gt;agent-guard&lt;/code&gt;: two Claude Code hooks that enforce behavior where the model can't argue with them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Guard 1: the test-tampering guard
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;agent-guard-hooks
agent-guard &lt;span class="nb"&gt;install&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On &lt;code&gt;SessionStart&lt;/code&gt;, the hook snapshots sha256 hashes of every test file and every source file in the project. On &lt;code&gt;Stop&lt;/code&gt;, it diffs. If test files were modified or deleted while &lt;strong&gt;no source file changed&lt;/strong&gt;, the stop is blocked:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Test-tampering guard: test files changed but no source files changed since this session started.
Changed test files:
  - tests/test_auth.py
...
Before proceeding, verify each changed test actually fails without the fix:
revert the source change, re-run the test, and confirm it goes red. A test that stays green without the fix is not covering the bug.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That "revert and re-run" check is the community's own verified workaround — it just used to be manual. New test files never count as tampering (writing tests for new code is legitimate); only modified or deleted assertions trip it. Warn-only mode exists if you'd rather nag than block.&lt;/p&gt;

&lt;h2&gt;
  
  
  Guard 2: the outbound-action guard
&lt;/h2&gt;

&lt;p&gt;A &lt;code&gt;PreToolUse&lt;/code&gt; hook on &lt;code&gt;Bash&lt;/code&gt; matches commands against a denylist of risky patterns and blocks before they run:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;pushes to protected branches (&lt;code&gt;main&lt;/code&gt;, &lt;code&gt;master&lt;/code&gt;, &lt;code&gt;prod*&lt;/code&gt;, &lt;code&gt;release/*&lt;/code&gt;) and any &lt;code&gt;--force&lt;/code&gt; push&lt;/li&gt;
&lt;li&gt;package publishes: &lt;code&gt;npm publish&lt;/code&gt;, &lt;code&gt;twine upload&lt;/code&gt;, &lt;code&gt;cargo publish&lt;/code&gt;, &lt;code&gt;gh release create&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;prod deploys: &lt;code&gt;kubectl apply&lt;/code&gt;, &lt;code&gt;terraform apply&lt;/code&gt;, &lt;code&gt;fly deploy&lt;/code&gt;, &lt;code&gt;vercel --prod&lt;/code&gt;, …&lt;/li&gt;
&lt;li&gt;cloud provisioning (the spend vector): &lt;code&gt;aws ec2 run-instances&lt;/code&gt;, &lt;code&gt;gcloud compute instances create&lt;/code&gt;, …&lt;/li&gt;
&lt;li&gt;mass-send: Slack webhook URLs, &lt;code&gt;sendmail&lt;/code&gt;, SendGrid/Mailgun API calls
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Outbound-action guard blocked this command (rule 'git-push-protected': Push to a protected branch (main/master/prod*/release/*)).
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Your own allowlist in &lt;code&gt;~/.config/agent-guard/config.json&lt;/code&gt; overrides the denylist (internal registries, staging targets), and every block &lt;em&gt;and&lt;/em&gt; every override lands in an audit log you can read with &lt;code&gt;agent-guard log&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The design rules I kept from the last hook I built
&lt;/h2&gt;

&lt;p&gt;This is the sibling of &lt;a href="https://github.com/hahahahahahahahah6/edit-guard" rel="noopener noreferrer"&gt;edit-guard&lt;/a&gt; (stale cross-session edit blocking), and it keeps the same contract:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Stdlib only.&lt;/strong&gt; No dependencies, no daemon, no network. State is flat files under &lt;code&gt;~/.config/agent-guard/&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fail open, always.&lt;/strong&gt; Corrupt snapshot, unreadable file, malformed hook input — anything unexpected means "allow". A guard that wedges your session is worse than no guard.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Block, don't warn (by default).&lt;/strong&gt; I learned this the hard way: the agent reads a warning, says "noted", and does it anyway. Blocking forces the verification step, which is the actual fix.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One deliberate contrast with a neighbor project: Rashomon (r/aiagents) &lt;em&gt;observes&lt;/em&gt; — it records what the agent did and compares it against the agent's summary. agent-guard &lt;em&gt;prevents&lt;/em&gt;. Observation tells you after the fact; the hook is there before the fact. Different layers, complementary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest limitations
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;"Tests changed, source didn't" is a strong signal, not a proof. Legit test-only refactors get flagged — use &lt;code&gt;ignore_paths&lt;/code&gt; or warn mode.&lt;/li&gt;
&lt;li&gt;The outbound guard watches &lt;code&gt;Bash&lt;/code&gt;. An agent calling a Slack MCP tool directly is out of scope for this version.&lt;/li&gt;
&lt;li&gt;The spend denylist is best-effort; it's a seatbelt, not a vault.&lt;/li&gt;
&lt;li&gt;The hook protocol (stdin shape, exit-2-blocks) is community-documented, not a stable API.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;13 smoke tests pass, including the exact navune scenario and fail-open behavior on corrupted state.&lt;/p&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;GitHub: &lt;a href="https://github.com/hahahahahahahahah6/agent-guard" rel="noopener noreferrer"&gt;https://github.com/hahahahahahahahah6/agent-guard&lt;/a&gt; (MIT)&lt;/li&gt;
&lt;li&gt;PyPI: &lt;code&gt;pip install agent-guard-hooks&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you run agents across sessions: what's the worst thing one of yours has done that a prompt-level rule failed to stop? I'm collecting failure modes for the next guards (comment-slop and mutation-style test-honesty are on the roadmap).&lt;/p&gt;

</description>
      <category>claudecode</category>
      <category>python</category>
      <category>cli</category>
      <category>ai</category>
    </item>
    <item>
      <title>Two Claude Code sessions edited the same file. One of them was wrong. I built a hook to stop that</title>
      <dc:creator>hao li</dc:creator>
      <pubDate>Wed, 30 Sep 2026 08:35:34 +0000</pubDate>
      <link>https://dev.to/haoli/two-claude-code-sessions-edited-the-same-file-one-of-them-was-wrong-i-built-a-hook-to-stop-that-15jp</link>
      <guid>https://dev.to/haoli/two-claude-code-sessions-edited-the-same-file-one-of-them-was-wrong-i-built-a-hook-to-stop-that-15jp</guid>
      <description>&lt;p&gt;I run multiple Claude Code sessions against the same repo — one refactoring, one fixing tests, sometimes one more exploring. Last week I watched session A carefully edit a function based on code that session B had already rewritten ten minutes earlier. The edit applied cleanly. It was also completely wrong: it reintroduced the old logic on top of the new. I only caught it in &lt;code&gt;git diff&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;There's a name for this failure mode. Dat Do's measurements post on agent memory put it plainly: &lt;em&gt;"With several agent sessions on one repository, session A reads &lt;code&gt;createSession&lt;/code&gt;, session B changes it, and A then edits on the old assumption."&lt;/em&gt; The existing answers are heavyweight: file-reservation MCP servers, session buses with path claims. I wanted something I could install in ten seconds.&lt;/p&gt;

&lt;p&gt;So I built &lt;code&gt;edit-guard&lt;/code&gt;: a single stdlib-only Python script that plugs into Claude Code's &lt;code&gt;PreToolUse&lt;/code&gt; hook.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;edit-guard
edit-guard &lt;span class="nb"&gt;install&lt;/span&gt;   &lt;span class="c"&gt;# merges into ~/.claude/settings.json, backs it up first&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;From then on, every &lt;code&gt;Read&lt;/code&gt; records what your session saw (file hash + timestamp), and every &lt;code&gt;Edit&lt;/code&gt;/&lt;code&gt;Write&lt;/code&gt; checks the file's current hash against what &lt;em&gt;your session&lt;/em&gt; last saw. If the bytes changed and another session touched the file in between, the write is blocked and the agent is told to re-read first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Stale edit blocked: 'src/auth.py' changed since you last saw it
(last modified by session 'a4c2e901'). Re-read the file before editing.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same-session rewrites always pass. New files always pass. A corrupt state log, an unreadable file, anything unexpected — the hook fails open and allows the edit. A safety tool that can wedge your workflow is worse than no tool.&lt;/p&gt;

&lt;p&gt;The state is one append-only JSONL file at &lt;code&gt;~/.config/edit-guard/writes.jsonl&lt;/code&gt;. No daemon, no server, no network. &lt;code&gt;edit-guard log&lt;/code&gt; shows you what's being tracked.&lt;/p&gt;

&lt;h2&gt;
  
  
  The design tension: block vs. warn
&lt;/h2&gt;

&lt;p&gt;I went back and forth on whether a stale write should block or just warn. Warning is friendlier, but in practice the agent reads the warning, says "noted", and edits anyway — I've watched it happen. Blocking forces the re-read, which is the actual fix. The failure mode of blocking is annoyance (a reformat by session B blocks session A); the failure mode of warning is silent corruption. For a tool whose whole job is preventing silent corruption, blocking is the honest default.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest limitations
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;It detects write-after-write staleness, not true locking — two sessions racing in the same second still race.&lt;/li&gt;
&lt;li&gt;Byte-level digests mean a pure reformat counts as "changed".&lt;/li&gt;
&lt;li&gt;The hook protocol (stdin shape, exit-2-blocks) is community-documented, not a stable API; if it changes, the hook degrades to allow.&lt;/li&gt;
&lt;li&gt;Per-machine only.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;GitHub: &lt;a href="https://github.com/hahahahahahahahah6/edit-guard" rel="noopener noreferrer"&gt;https://github.com/hahahahahahahahah6/edit-guard&lt;/a&gt; (MIT)&lt;/li&gt;
&lt;li&gt;PyPI: &lt;code&gt;pip install edit-guard&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;7 smoke tests pass, including the exact A-reads/B-writes/A-blocked scenario and a fail-open test against a corrupted state log.&lt;/p&gt;

&lt;p&gt;If you multi-session a single repo: how do you keep your agents from stepping on each other? I'm curious what breaks first at your scale.&lt;/p&gt;

</description>
      <category>claudecode</category>
      <category>python</category>
      <category>cli</category>
      <category>ai</category>
    </item>
    <item>
      <title>I use Claude Code and Codex side by side. Handing off between them was pure pain, so I built a CLI</title>
      <dc:creator>hao li</dc:creator>
      <pubDate>Wed, 30 Sep 2026 08:34:30 +0000</pubDate>
      <link>https://dev.to/haoli/i-use-claude-code-and-codex-side-by-side-handing-off-between-them-was-pure-pain-so-i-built-a-cli-25bf</link>
      <guid>https://dev.to/haoli/i-use-claude-code-and-codex-side-by-side-handing-off-between-them-was-pure-pain-so-i-built-a-cli-25bf</guid>
      <description>&lt;p&gt;I split my coding work across two agents: Claude Code for most things, Codex CLI when I want a second opinion or a different model. The problem hits every time I switch mid-task: the new session knows nothing. I'd open a blank &lt;code&gt;handover.md&lt;/code&gt; and type from memory what the last session did — which files it touched, what broke, where I left off. Half the time I'd forget the exact error message, and the receiving agent would redo work or repeat the same failed command.&lt;/p&gt;

&lt;p&gt;There are good handoff tools for Claude Code alone (Matt Pocock's &lt;code&gt;/handoff&lt;/code&gt; skill, for example). But nobody was doing the cross-tool case: one command that reads &lt;em&gt;both&lt;/em&gt; tools' local session transcripts and spits out a handover doc the next agent — whichever agent — can pick up.&lt;/p&gt;

&lt;p&gt;So I built one. It's called &lt;code&gt;session-handover&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does
&lt;/h2&gt;

&lt;p&gt;Both CLIs already keep everything on disk. Claude Code writes each session as &lt;code&gt;.jsonl&lt;/code&gt; under &lt;code&gt;~/.claude/projects&lt;/code&gt;; Codex CLI writes per-day &lt;code&gt;.jsonl&lt;/code&gt; files under &lt;code&gt;~/.codex/sessions&lt;/code&gt;. The CLI walks both:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;session-handover

&lt;span class="nv"&gt;$ &lt;/span&gt;session-handover list
TOOL         SESSION                  STARTED           MESSAGES
codex        0199f3a1-…               2026-09-30 00:41   57
claude-code  a4c2e901-…               2026-09-29 23:58   34

&lt;span class="nv"&gt;$ &lt;/span&gt;session-handover make &lt;span class="nt"&gt;--session&lt;/span&gt; a4c2e901 &lt;span class="nt"&gt;--out&lt;/span&gt; HANDOVER.md
Wrote HANDOVER.md &lt;span class="o"&gt;(&lt;/span&gt;claude-code / a4c2e901-…&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The generated document has the shape I was hand-writing anyway:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Goal&lt;/strong&gt; — the session's first user message&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What happened&lt;/strong&gt; — counts plus a timeline: which files were edited, which commands were run&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Files touched&lt;/strong&gt;, plus &lt;code&gt;git diff --stat&lt;/code&gt; when the session's working dir is a repo&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Errors encountered&lt;/strong&gt; — failed commands and tracebacks, quoted verbatim&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open items&lt;/strong&gt; and a &lt;strong&gt;next steps&lt;/strong&gt; checklist for the receiving agent&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;No LLM in the loop. It's heuristic extraction over the transcript — fast, local, private, zero dependencies (stdlib only).&lt;/p&gt;

&lt;h2&gt;
  
  
  The annoying part: undocumented formats
&lt;/h2&gt;

&lt;p&gt;Neither transcript format is documented, and both drift. Codex's in particular has changed shape across versions. So the parsers are written to degrade, not crash: every line is parsed defensively, unknown shapes are skipped, malformed lines are dropped, and a corrupt file just yields less detail instead of killing the whole &lt;code&gt;list&lt;/code&gt; run. The Codex side is explicitly best-effort — when its format moves again, you lose detail, not the tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest limitations
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Heuristic, not intelligent: it summarizes what the transcript &lt;em&gt;says&lt;/em&gt;, it doesn't understand &lt;em&gt;why&lt;/em&gt; decisions were made. Read the handover before trusting it.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;--last&lt;/code&gt; orders by file mtime; clock skew can misorder the list.&lt;/li&gt;
&lt;li&gt;Local transcripts only — nothing leaves your machine.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;GitHub: &lt;a href="https://github.com/hahahahahahahahah6/session-handover" rel="noopener noreferrer"&gt;https://github.com/hahahahahahahahah6/session-handover&lt;/a&gt; (MIT)&lt;/li&gt;
&lt;li&gt;PyPI: &lt;code&gt;pip install session-handover&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;11 smoke tests pass, including fake fixtures for both transcript formats and a malformed-input test that asserts the parser never crashes on garbage.&lt;/p&gt;

&lt;p&gt;If you also bounce between agents: what does your handover ritual look like? I'm curious what I missed.&lt;/p&gt;

</description>
      <category>claudecode</category>
      <category>python</category>
      <category>cli</category>
      <category>ai</category>
    </item>
    <item>
      <title>Your MCP servers are eating your context window. I built a CLI to audit the tax</title>
      <dc:creator>hao li</dc:creator>
      <pubDate>Wed, 30 Sep 2026 08:27:17 +0000</pubDate>
      <link>https://dev.to/haoli/your-mcp-servers-are-eating-your-context-window-i-built-a-cli-to-audit-the-tax-2e80</link>
      <guid>https://dev.to/haoli/your-mcp-servers-are-eating-your-context-window-i-built-a-cli-to-audit-the-tax-2e80</guid>
      <description>&lt;p&gt;Every MCP server you connect injects &lt;em&gt;all&lt;/em&gt; of its tool schemas into &lt;em&gt;every&lt;/em&gt; request. Claude Code loads them at session start, and there is no UI to temporarily disable a server you don't need right now. People have measured 41k tokens of pure schema; one estimate says 6 mid-size servers eat 10–15% of the window before the conversation even starts.&lt;/p&gt;

&lt;p&gt;Nobody measures this per server. So I built a tiny CLI that does.&lt;/p&gt;

&lt;h2&gt;
  
  
  mcp-tax
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;mcp-tax&lt;/code&gt; audits the &lt;strong&gt;context tax&lt;/strong&gt; of your MCP servers, then lets you launch Claude Code with the expensive ones switched off — per session, without touching your real config.&lt;/p&gt;

&lt;p&gt;Zero dependencies. Python standard library only.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;mcp-tax
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;See what you have configured&lt;/strong&gt; (reads &lt;code&gt;~/.claude.json&lt;/code&gt; plus &lt;code&gt;./.mcp.json&lt;/code&gt; when present):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;mcp-tax list
github
    npx &lt;span class="nt"&gt;-y&lt;/span&gt; @modelcontextprotocol/server-github
postgres  &lt;span class="o"&gt;[&lt;/span&gt;off]
    uvx mcp-server-postgres &lt;span class="nt"&gt;--db-url&lt;/span&gt; ...
2 server&lt;span class="o"&gt;(&lt;/span&gt;s&lt;span class="o"&gt;)&lt;/span&gt;, 1 disabled
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Measure the tax&lt;/strong&gt; — it handshakes each server over stdio (&lt;code&gt;initialize&lt;/code&gt;, then &lt;code&gt;tools/list&lt;/code&gt;), counts tools, and estimates tokens:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;mcp-tax audit
server    tools  schema chars  est. tokens
&lt;span class="nt"&gt;------&lt;/span&gt;  &lt;span class="nt"&gt;-------&lt;/span&gt;  &lt;span class="nt"&gt;------------&lt;/span&gt;  &lt;span class="nt"&gt;-----------&lt;/span&gt;
github       51       118,203        29,551
postgres &lt;span class="o"&gt;[&lt;/span&gt;off]    9        12,440         3,110
&lt;span class="nt"&gt;------&lt;/span&gt;  &lt;span class="nt"&gt;-------&lt;/span&gt;  &lt;span class="nt"&gt;------------&lt;/span&gt;  &lt;span class="nt"&gt;-----------&lt;/span&gt;
TOTAL        60       130,643        32,661
~16.3% of a 200k context window &lt;span class="o"&gt;(&lt;/span&gt;est. tokens &lt;span class="o"&gt;=&lt;/span&gt; schema chars / 4&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Toggle servers off/on&lt;/strong&gt; (persisted in &lt;code&gt;~/.config/mcp-tax/disabled.json&lt;/code&gt;):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;mcp-tax off postgres
postgres disabled &lt;span class="o"&gt;(&lt;/span&gt;affects &lt;span class="sb"&gt;`&lt;/span&gt;mcp-tax run&lt;span class="sb"&gt;`&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;mcp-tax on postgres
postgres enabled &lt;span class="o"&gt;(&lt;/span&gt;affects &lt;span class="sb"&gt;`&lt;/span&gt;mcp-tax run&lt;span class="sb"&gt;`&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Launch Claude Code without the disabled servers:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;mcp-tax run &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="s2"&gt;"summarize this repo"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This writes a filtered &lt;code&gt;{"mcpServers": ...}&lt;/code&gt; config (disabled servers removed) and execs &lt;code&gt;claude --mcp-config &amp;lt;file&amp;gt;&lt;/code&gt; with your args forwarded. Your real &lt;code&gt;~/.claude.json&lt;/code&gt; is never modified.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the token estimate works
&lt;/h2&gt;

&lt;p&gt;For each server, mcp-tax JSON-encodes the full &lt;code&gt;tools/list&lt;/code&gt; result (compact, no whitespace) and counts characters. Estimated tokens = &lt;code&gt;round(chars / 4)&lt;/code&gt; — the Anthropic tokenizer's ~4-chars-per-token rule of thumb on English/JSON text.&lt;/p&gt;

&lt;p&gt;Deliberately crude. Real counts vary with the tokenizer and how Claude Code wraps schemas, so treat it as an order-of-magnitude gauge: good enough to answer &lt;em&gt;"which server is eating my window?"&lt;/em&gt;, not a billing meter.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest limitations
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;stdio servers only.&lt;/strong&gt; SSE / streamable HTTP servers aren't audited.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Audit uses &lt;code&gt;select(2)&lt;/code&gt;&lt;/strong&gt; on the server's stdout pipe — fine on Linux/macOS, not on Windows.&lt;/li&gt;
&lt;li&gt;The estimate ignores runtime behavior: a 2-tool server can still be expensive if its tool &lt;em&gt;results&lt;/em&gt; are huge. This measures schema cost only.&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;--mcp-config&lt;/code&gt; flag for &lt;code&gt;run&lt;/code&gt; comes from Claude Code's documented CLI options; it couldn't be verified on the machine where this was built (no Claude Code CLI there). If the flag name ever changes, &lt;code&gt;run&lt;/code&gt; prints the filtered config path so you can pass it manually.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;GitHub: &lt;a href="https://github.com/hahahahahahahahah6/mcp-tax" rel="noopener noreferrer"&gt;https://github.com/hahahahahahahahah6/mcp-tax&lt;/a&gt; (MIT)&lt;/li&gt;
&lt;li&gt;PyPI: &lt;code&gt;pip install mcp-tax&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;7 smoke tests pass, including a fake stdio MCP server that speaks the same newline-delimited JSON-RPC 2.0 framing as real servers — so the audit math is tested against a realistic handshake, including stdout noise and timeout cases.&lt;/p&gt;

&lt;p&gt;If you run it against your own setup, I'd genuinely like to know: which server turned out to be your biggest tax?&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>python</category>
      <category>cli</category>
    </item>
    <item>
      <title>I gave my coding agent eyes on my real, logged-in browser (stdlib-only, local)</title>
      <dc:creator>hao li</dc:creator>
      <pubDate>Wed, 30 Sep 2026 08:09:49 +0000</pubDate>
      <link>https://dev.to/haoli/i-gave-my-coding-agent-eyes-on-my-real-logged-in-browser-stdlib-only-local-5he0</link>
      <guid>https://dev.to/haoli/i-gave-my-coding-agent-eyes-on-my-real-logged-in-browser-stdlib-only-local-5he0</guid>
      <description>&lt;p&gt;Every agentic coding tool hits the same wall eventually: the page you need is behind a login, a 403, or bot detection, and the agent's server-side fetcher is locked out. Your browser is logged in. Your agent is not. So the fix is not a better scraper — it is letting the agent &lt;em&gt;borrow your eyes&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;I built &lt;a href="https://github.com/hahahahahahahahah6/browser-buddy" rel="noopener noreferrer"&gt;browser-buddy&lt;/a&gt;, an MCP server + Chrome extension (Manifest V3) that lets coding agents read pages through the user's real Chrome profile: cookies, sessions, logins included. MIT licensed, zero dependencies, nothing leaves your machine.&lt;/p&gt;

&lt;h2&gt;
  
  
  The architecture (boring on purpose)
&lt;/h2&gt;

&lt;p&gt;Three pieces, each dumb and replaceable:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Chrome extension (MV3)&lt;/strong&gt; — content scripts extract the readable text of a tab on demand. It never acts on its own; it only responds to messages.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Native host (Python, stdlib)&lt;/strong&gt; — bridges the extension to the outside world over a Unix socket in the temp dir. &lt;code&gt;chrome.runtime.connectNative&lt;/code&gt; on one side, a socket server on the other.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MCP server (Python, stdlib)&lt;/strong&gt; — speaks JSON-RPC 2.0 over stdio, exactly the framing MCP stdio servers use. No &lt;code&gt;mcp&lt;/code&gt; SDK package; the protocol surface I need is two tools, so I hand-rolled the framing in ~200 lines:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;TOOLS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;read_active_tab&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Read the user&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s currently active Chrome tab through their real, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;logged-in browser profile. Returns the page URL, title, and &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cleaned readable text. Use when WebFetch/fetch is blocked by a &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;403, login wall, or bot detection.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;inputSchema&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;object&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;properties&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{},&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;additionalProperties&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;open_and_read_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Open a URL in a background Chrome tab using the user&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s real &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;browser profile (cookies and login sessions included), extract &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;its readable text, then close the tab.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="bp"&gt;...&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Message flow for &lt;code&gt;open_and_read_url&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;agent ──JSON-RPC/stdio──▶ MCP server ──Unix socket──▶ native host
                                                          │ connectNative
                                                          ▼
                                                   Chrome extension ──▶ tab
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Logs go to stderr (stdout is reserved for protocol messages — the classic stdio footgun, handled once and never again).&lt;/p&gt;

&lt;h2&gt;
  
  
  Why not just use a headless browser?
&lt;/h2&gt;

&lt;p&gt;Headless browsers are &lt;em&gt;new&lt;/em&gt; browsers: no cookies, no sessions, and they get fingerprinted as bots — the exact problem I was trying to escape. The insight is embarrassingly simple: the most authenticated, least-bot-detected browser in existence is the one the user is already using. So don't automate a fake browser; ask the real one to read.&lt;/p&gt;

&lt;h2&gt;
  
  
  The security model
&lt;/h2&gt;

&lt;p&gt;This design is only acceptable because everything is local:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;No server, no cloud, no telemetry. The page content goes from your tab to your agent's context, period.&lt;/li&gt;
&lt;li&gt;The extension cannot initiate anything — it is purely reactive to tool calls the agent makes.&lt;/li&gt;
&lt;li&gt;The MCP server only runs when your agent starts it (stdio), and only your local machine can reach the socket.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Verification
&lt;/h2&gt;

&lt;p&gt;17/17 smoke tests pass (framing, tool schemas, socket relay, tab extraction mocks), and it's live on the official MCP registry as &lt;code&gt;io.github.hahahahahahahahah6/browser-buddy&lt;/code&gt;. Next: real-world testing on Windows Chrome.&lt;/p&gt;

&lt;p&gt;Try it: &lt;code&gt;uvx browser-buddy-mcp&lt;/code&gt; (needs the extension + host from the repo).&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Built in the open. Issues and PRs welcome — especially from people whose agents keep getting 403'd.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>showdev</category>
      <category>opensource</category>
      <category>mcp</category>
      <category>python</category>
    </item>
  </channel>
</rss>
