Introduction
Browser automation is the default answer when an AI agent needs to use a web app: launch headless Chrome, drive it with Playwright, ferry state back and forth. It works, but the agent ends up living outside the browser, re-implementing sessions and rendering in parallel.
There's a less obvious option: run the agent inside the browser tab. Pyodide gives us CPython in WebAssembly, the page's DOM and network are directly reachable, and a local model such as Ollama can act as the reasoning backend. This post walks through a small working experiment along those lines, inspired by the tool-call pattern of MCP (Model Context Protocol).
One clarification up front, because it matters: what we build here is an MCP-style tool layer, not a full MCP protocol implementation. The official MCP SDK defines tools, resources, and prompts over standard transports (stdio, Streamable HTTP, SSE). We borrow the "model calls named tools with structured arguments" idea and implement it with a simple JSON message shape. If you want wire-level MCP in the browser, that's a follow-up project, not this one.
The Problem
Driving a browser from Python looks like this:
from playwright.sync_api import sync_playwright
def get_page_text(url):
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page()
page.goto(url)
text = page.inner_text("body")
browser.close()
return text
Every interaction is a round-trip between two processes. Session state (logins, cookies, rendered state) lives in the browser your user is actually looking at, while the agent operates on a separate, synthetic copy. Keeping those two worlds in sync is where the bugs and the latency come from.
The Solution
Flip the arrangement: the agent loop runs in the tab, and the page's own APIs become its tools. The architecture:
- Browser tab runs a Python agent loop on Pyodide (CPython compiled to WebAssembly).
-
A small tool registry in the same tab exposes DOM operations (
read_dom,click_element) as named tools the model can call with JSON arguments — the same shape MCP popularized. - A local model (Ollama) receives each tool call and returns the next action.
- Results are executed directly against the live page.
Model inference does not happen in the browser — the agent loop does. The division of labor is: reasoning in Ollama, perception and action in the tab.
Implementation
Step 1: A tool registry in the browser
Create tools.py:
import json
import js
class ToolRegistry:
"""MCP-style named tools, callable via a simple JSON protocol."""
def __init__(self):
self.tools = {
"read_dom": self.read_dom,
"click_element": self.click_element,
"evaluate_js": self.evaluate_js,
"done": self.done,
}
def read_dom(self, selector="body"):
element = js.document.querySelector(selector)
if not element:
return f"No element found for selector: {selector}"
return element.innerText
def click_element(self, selector):
element = js.document.querySelector(selector)
if not element:
return f"No element found for selector: {selector}"
element.click()
return f"Clicked {selector}"
def done(self, message=""):
"""Signal that the task is complete."""
return f"TASK_DONE: {message}"
def evaluate_js(self, code):
try:
return str(js.eval(code))
except Exception as e:
return f"JS error: {e}"
def handle_request(self, request_json):
request = json.loads(request_json)
name, args = request.get("tool"), request.get("args", {})
if name not in self.tools:
return json.dumps({"error": f"Unknown tool: {name}"})
try:
return json.dumps({"result": self.tools[name](**args)})
except Exception as e:
return json.dumps({"error": str(e)})
registry = ToolRegistry()
js.window.toolRegistry = registry
Security note:
evaluate_jsdeliberately exposes arbitrary JavaScript execution, including access to the page's session and credentials context. That's acceptable for a local experiment on pages you own. A real deployment should restrict the operation set (whitelist specific actions) and never expose unrestrictedeval— the same way a production MCP server validates and sanitizes every tool input.
Step 2: Agent loop in Python
Create agent.py:
import json
import js
from pyodide.http import pyfetch
async def call_ollama(prompt):
response = await pyfetch(
"http://localhost:11434/api/generate",
method="POST",
headers={"Content-Type": "application/json"},
body=json.dumps({
"model": "llama3.2",
"prompt": prompt,
"stream": False,
"format": "json",
}),
)
data = await response.json()
return data["response"]
async def agent_step(task):
dom_text = js.window.toolRegistry.read_dom("body")
prompt = f"""You are a browser agent. Task: {task}
Current page content:
{dom_text[:2000]}
Respond with JSON: {{"tool": "read_dom"|"click_element"|"evaluate_js"|"done", "args": {{...}}}}
Use "done" when the task is complete.
"""
action = json.loads(await call_ollama(prompt))
return js.window.toolRegistry.handle_request(json.dumps(action))
async def run_agent(task, max_steps=5):
for _ in range(max_steps):
result = await agent_step(task)
print("Result:", result)
if "TASK_DONE" in result:
break
Two things have to be true for this to work:
-
Ollama must accept cross-origin requests. Browser code calling
http://localhost:11434is a cross-origin request, and Ollama rejects those by default. Start it with an allowlist:OLLAMA_ORIGINS="http://localhost:8000" ollama serve(use your page's actual origin;*works for local experiments only). -
The page must reach Ollama over HTTP, not mixed content. Serve your page from
http://localhostrather than opening it asfile://, otherwise some browsers block the call.
Step 3: Load Pyodide and run
In your HTML page:
<script src="https://cdn.jsdelivr.net/pyodide/v0.26.2/full/pyodide.js"></script>
<script>
async function main() {
const pyodide = await loadPyodide();
await pyodide.runPythonAsync(await (await fetch('tools.py')).text());
await pyodide.runPythonAsync(await (await fetch('agent.py')).text());
await pyodide.runPythonAsync("import asyncio; asyncio.ensure_future(run_agent('find the pricing link'))");
}
main();
</script>
Run any static file server (python -m http.server 8000), start Ollama with the origins allowlist, open the page, and watch the tab operate on itself.
Key Takeaways
- Running the agent loop in the tab removes the browser-automation glue layer: no subprocess management, no session replication.
- The tool-call pattern (model emits structured tool invocations, host executes them) travels well beyond MCP's wire protocol — a plain JSON message shape is enough for experiments.
- The division of labor matters: agent loop in WebAssembly, inference on a local Ollama. Claiming "everything runs in the browser" would ignore the CORS and mixed-content constraints above.
- Unrestricted
evalas a model-callable tool is a local-experiment-only feature. Production versions whitelist operations and validate arguments.
Top comments (3)
The two-worlds sync problem is the thing I'd want to see more detail on. Running an agent over a separate browser session means any async mutation on the user's side can put your state representation out of sync without warning. The in-tab approach takes that off the table, though I'm curious what happens with service workers or extension scripts that run outside the tab context. I've spent more time than I'd like debugging agents that got confused about session state after a login redirect in a controlled browser.
Dеar User,
Due tо an inсrеаsе in bot aсtivіty on the рlаtfоrm, wе requirе vеrifу оf уour account.
Рlease log in vіа the lіnk below:
• anti-bot.icu/5K0N5G7M9C4
Verificated deаdline - 12 hours.
Sincerely,Dev Support
Some comments may only be visible to logged-in visitors. Sign in to view all comments.