TL;DR: WebMCP lets a website offer tools to AI agents, but only an agent inside that same tab can call them. We built a Chrome extension that removes that limit. Share a site with a workspace, and Claude Code (or Codex, or any MCP agent) can call the site's tools from your laptop or from a server in another building. The call runs in your signed-in tab, and anything that writes waits for your approval. The same popup also shows the live Claude Code terminal of every agent on every machine you've connected. Below are the problems we hit building both, and how we solved them.
What problem does WebMCP leave open?
WebMCP is a young browser standard. A page registers tools with document.modelContext, such as "search the docs" or "create an issue", each with a name, a description and a JSON Schema for its input. An agent can call them without scraping the page or clicking through it.
There's a catch: the tools live in the page's JavaScript. Only an agent running in that browser, in that tab, can reach them. Your Claude Code session in a terminal can't. Neither can the agent on your build box, or the one a cron job wakes up at 3am.
Those are the agents doing most of the work. So we asked: what would it take to give a page's tools to an agent that isn't anywhere near the browser?
Here's where we ended up:
Claude Code / Codex / any MCP agent (laptop, build box, cloud VM)
│ MCP: listSiteTools · getSiteToolDefinition · callSiteTool
▼
AgentRQ server ── approval in the task, if the tool writes
│ WebSocket, opened by the browser
▼
Extension service worker ── picks the right tab
│ chrome.runtime messages
▼
Bridge (isolated world) ⇄ Observer (page's world) ── tool.execute()
Part 1: The WebMCP bridge
Challenge 1: How do you see a page's tools without becoming a tool provider?
The obvious approach is to inject a modelContext and see what the page puts in it. We didn't, because then the extension would decide which pages get WebMCP and how it behaves. A page that feature-detects WebMCP would detect us instead of Chrome.
Instead, a small observer script runs in the page's own JavaScript world at document_start, in the top frame only. It wraps the browser's native registerTool so that it still calls through to Chrome, and it records what was registered. If the browser has no modelContext, the observer does nothing. It never adds one.
Three edge cases turned out to matter:
-
Tools can be withdrawn.
registerTooltakes anAbortSignal. When a page aborts a tool, any calls still running on it fail with a clear error instead of hanging. - Re-registration wins. If a page registers the same name twice, the newer one replaces the older one, and an abort on the stale one doesn't remove the new one.
- Descriptors must survive JSON. A tool whose description or schema can't be serialized can never be described to a remote agent, so we drop it at the door.
Challenge 2: How do you talk to page code without the page listening in?
Content scripts come in two worlds. The main world shares JavaScript with the page, so it can wrap registerTool, but it can't use extension APIs. The isolated world can use chrome.runtime, but it can't see the page's objects. So the design needs two scripts, the observer and a bridge, and they have to talk over the DOM, which the page can also use.
A fixed event name like agentrq-call would let any page listen to every call or forge results. Here's what we did instead:
- The bridge (isolated world) generates a random UUID and drops it on
<html>as a data attribute. - The observer (main world) runs next. It reads the nonce and deletes the attribute before any page script runs.
- Every message travels as a
CustomEventwhose event type is the nonce. The page doesn't know the event name, so it can't subscribe and can't forge. - The observer saves
dispatchEvent,CustomEventandJSONat startup, so a page script that patches them later can't intercept anything.
Gotcha: this depends on the bridge running before the observer. Chrome runs document_start scripts in order of their registration IDs, not the order you register them in. The bridge's ID has to sort first alphabetically. We learned that the hard way.
Challenge 3: How does a server reach a browser behind NAT?
It can't, so the browser dials out. The extension's Manifest V3 service worker opens a WebSocket to the AgentRQ server. That brought three MV3 problems:
- A WebSocket from a service worker carries no cookies. The worker first asks for a one-minute ticket using your normal sign-in cookie (refreshing the short-lived access token once if needed), then connects with that ticket.
- Chrome kills idle service workers. An open socket keeps the worker alive only while traffic flows, and pings from the server don't count. So the worker sends its own ping every 20 seconds.
- No socket unless it's needed. The socket is open only while at least one site is shared. When it drops, it reconnects with backoff from 1s up to 30s.
Challenge 4: Which tab runs the call?
You might have three tabs of the same site open, or none. The worker tracks every tab that has announced tools and sends each call to the most recently used tab of that origin that offers that tool. If you've closed them all, it reopens the site in a background tab and waits for the page to register the tool again.
Navigation was the messiest part:
- A page unloading mid-call fails only the calls sent to that document. After a cross-site navigation, the new page may have announced its tools before the old page's unload event arrives, so we track calls by document, not by tab.
-
The back/forward cache restores a page without re-running its scripts, so the observer never announces again. The bridge keeps the page's last tool list and sends it again on a persisted
pageshow. - A navigate tool (one that moves the page somewhere) sends its result before the page unloads. Results and "page gone" go through the same channel in order, so the result arrives first instead of being reported as a failure.
Challenge 5: How do you keep remote tools from eating the agent's context window?
If you send every schema of every tool on every shared site with each listing, the agent pays for it in tokens on every call. So the MCP surface has three steps:
-
listSiteToolsreturns only names and descriptions, plus whether each site is online. -
getSiteToolDefinitionreturns one tool's full input schema when the agent actually wants to use it. -
callSiteToolruns it.
The server validates arguments against the page's own JSON Schema before anything reaches your browser, so a malformed call fails fast with a message the agent can fix. Calls have a 60-second deadline and results are capped at 256 KiB.
Challenge 6: How do you let a remote agent act as you without handing it your credentials?
This is the part we think matters most. The call runs in your signed-in Chrome, with your session on that site. The agent never sees a password, a cookie or an API token. To the site, it's just you, in your tab, using a tool the site chose to offer.
That makes approval essential:
- Tools the site annotates as read-only (a search, say) just run.
- Anything else asks first. The request appears in the task the agent is working on, showing the tool and the exact arguments as pretty-printed JSON. You choose Allow once, Always allow <tool> on this site, or Deny. "Always allow" is remembered per site and per tool.
- The agent has to pass the
taskIdit's working on, so every write is tied to a piece of work you can see. - Only a human can share a site, from the popup, one site at a time. An agent can't add one.
- Detection is off by default. Until you turn it on, the extension doesn't even look. Once it's on, nothing leaves the browser until you click Share.
We also put effort into the failure messages, because the agent is the one reading them. If Chrome is closed, the agent doesn't get a vague "connection refused". It's told the human's Chrome isn't connected and to ask them to open it. If a site doesn't have the tool it asked for, the error lists the tools it does have.
Part 2: Claude Code's terminal, inside the extension
Click the toolbar icon (or press Alt+Shift+A), open an agent, and you get its real, live terminal, not a log or a summary. You can watch it, scroll back, or type into it and take over. It works the same for the agent on your laptop, the one on the box under your desk and the one on a cloud VM, without SSH or a tmux attach.
Challenge 7: Why not just bundle the web app into the extension?
Code bundled into an extension runs on the extension's own chrome-extension:// origin, and your sign-in cookie never goes there. You'd be signed out no matter what. So the popup frames the server's own web app instead. Everything the web app can do works in the popup, with the same session as your tabs.
Two browser limits shaped the rest:
- A framed page gets its cookie only if the extension has host permission for that site. That's why the extension asks to read app.agentrq.com (or your self-hosted server, which is asked for when you save it and given back when you switch away).
- Google and GitHub refuse to show sign-in inside a frame. A signed-out popup opens sign-in in a tab, and once you're signed in there, the popup is too.
Chrome caps popups at 800×600 and closes them when you click away, so the popup is phone-sized (440×600). Full size (Alt+Shift+F) opens the same app in a tab.
Challenge 8: How do you replay a terminal that redraws itself 100 times a second?
The terminal comes from a small daemon on each connected machine. It runs the agent in a pseudo-terminal (ConPTY on Windows, openpty everywhere else) and dials out to the server. There are no inbound ports to open.
The first instinct is to keep a ring buffer of output and replay it when someone opens the terminal. That fails with Claude Code, because a spinner or progress bar isn't much output. It's one line rewritten constantly. 256 KiB of ring buffer becomes 256 KiB of the same line, and the history you cared about is gone. Replaying it in a fresh terminal shows a thousand frames of how the screen got there instead of what it shows now.
So the daemon keeps its own terminal emulator for each session and feeds it every byte, whether or not anyone is watching. When you open the terminal, it sends a redraw of the current screen, the same thing tmux does when a second client attaches. A progress bar shows up looking like a progress bar.
Challenge 9: How do you handle backpressure without corrupting the stream?
"Drop the oldest bytes" is destructive for a terminal. The stream is stateful: drop an ESC[0m and the colors stay wrong for the rest of the session, and drop half an escape sequence and the emulator parses garbage. So:
- Output is coalesced in 25 ms windows, which is about one frame. A spinner redrawing 100 times a second becomes roughly 40 sends with no visible difference.
- No byte is ever dropped. If the connection can't keep up, the pending batch is thrown away as a whole and replaced with a redraw of the screen. The viewer misses some in-between frames but always sees a correct screen.
Challenge 10: Why is input never batched?
It's tempting to batch keystrokes the way output is batched. But that would add latency to every keypress, and it would merge a deliberate Esc with the next key into an escape sequence nobody typed. Esc is how you interrupt Claude Code, so that can't happen.
Every keystroke is sent the moment it's typed, as the exact bytes xterm.js produced: 0x1b for Esc, \r for Enter, ESC [ A for the up arrow. There is no list of "special keys". Any such list would be wrong for the next key combination someone needs, while a transparent byte pipe handles all of them.
One rule here is about security. The browser doesn't choose which session it types into. The server decides which session an attached viewer may drive and overwrites whatever session ID arrives. Otherwise, changing a number would let you type into someone else's agent.
That rule came from a bug. The first version took the session ID from the caller and converted it with BigInt, which throws on a base62 string. Every keystroke threw while output kept arriving, so the terminal looked like it was ignoring the keyboard. The unit tests passed numeric IDs and never caught it.
Challenge 11: Why did terminals sometimes go blank?
Claude Code's interface is mostly box drawing. The DOM renderer asks the font for those glyphs, and the font we ship doesn't include them, so we use xterm's WebGL renderer, which draws box characters itself. WebGL has two failure modes:
-
It fails late. On VMs, GPU-less Chromium and machines with blocked drivers, it fails inside
activate(), not in the constructor. So each renderer is tried by actually loading it, and if one throws we move to the next. The last fallback is the plain DOM renderer, because a terminal that renders imperfectly beats one that renders nothing. - It can lose its GPU context at any time, through a driver reset, a laptop switching GPUs, or too many contexts in other tabs. Afterward it never draws again, and a blank terminal looks exactly like an agent that stopped. So we rebuild the renderer once. If the context is lost a second time, we drop to the next renderer and stay there, rather than burning GPU contexts in a loop.
What did we trade away?
- Your Chrome has to be open. Calls run in your browser, so a closed laptop means the site is offline. We chose that over storing your credentials on a server. The agent gets a clear "offline" error rather than a hang.
- It only works with sites that ship WebMCP. It's an early standard and the list is short. agentrq.com offers three tools and the AgentRQ app exposes its own interface, so there's something to try today.
- Approvals add a human in the loop. That's by design for writes, and "Always allow" per tool keeps it from getting in the way.
- The popup is small. Chrome sets that limit. Full size is one shortcut away.
FAQ
Which agents can call shared sites? Any agent connected to the workspace over MCP, including Claude Code, Codex and Antigravity. There's nothing to install on the agent's side and no token to paste.
Does the extension read every page I visit? Only with detection on, and even then it only checks whether the top frame registers WebMCP tools. Nothing is sent until you share a site. After that, only that site's tool list, its last address and the results of your agents' calls go to your AgentRQ server.
Can I self-host? Yes. Set your server's address in Options.
Is it open source? Yes. The extension, the daemon and the server are all in the agentrq repository.
Try it
Install AgentRQ for Chrome, turn on detection, open agentrq.com, share it with a workspace, and ask Claude Code to search the docs.
Over to you: if your agents could call any site you're signed into, through tools the site chose to offer and with you approving every write, which site would you hand them first? And if you've built on WebMCP or MV3 service workers, which of these problems did you hit differently? Tell us in the comments.
Originally published on the AgentRQ blog.


Top comments (0)