DEV Community

heavy shield
heavy shield

Posted on

Running an AI Agent Inside the Browser with Pyodide and Ollama

Introduction

Browser automation is the default answer when an AI agent needs to use a web app: launch headless Chrome, drive it with Playwright, ferry state back and forth. It works, but the agent ends up living outside the browser, re-implementing sessions and rendering in parallel.

There's a less obvious option: run the agent inside the browser tab. Pyodide gives us CPython in WebAssembly, the page's DOM and network are directly reachable, and a local model such as Ollama can act as the reasoning backend. This post walks through a small working experiment along those lines, inspired by the tool-call pattern of MCP (Model Context Protocol).

One clarification up front, because it matters: what we build here is an MCP-style tool layer, not a full MCP protocol implementation. The official MCP SDK defines tools, resources, and prompts over standard transports (stdio, Streamable HTTP, SSE). We borrow the "model calls named tools with structured arguments" idea and implement it with a simple JSON message shape. If you want wire-level MCP in the browser, that's a follow-up project, not this one.

The Problem

Driving a browser from Python looks like this:

from playwright.sync_api import sync_playwright

def get_page_text(url):
    with sync_playwright() as p:
        browser = p.chromium.launch()
        page = browser.new_page()
        page.goto(url)
        text = page.inner_text("body")
        browser.close()
        return text
Enter fullscreen mode Exit fullscreen mode

Every interaction is a round-trip between two processes. Session state (logins, cookies, rendered state) lives in the browser your user is actually looking at, while the agent operates on a separate, synthetic copy. Keeping those two worlds in sync is where the bugs and the latency come from.

The Solution

Flip the arrangement: the agent loop runs in the tab, and the page's own APIs become its tools. The architecture:

  1. Browser tab runs a Python agent loop on Pyodide (CPython compiled to WebAssembly).
  2. A small tool registry in the same tab exposes DOM operations (read_dom, click_element) as named tools the model can call with JSON arguments — the same shape MCP popularized.
  3. A local model (Ollama) receives each tool call and returns the next action.
  4. Results are executed directly against the live page.

Model inference does not happen in the browser — the agent loop does. The division of labor is: reasoning in Ollama, perception and action in the tab.

Implementation

Step 1: A tool registry in the browser

Create tools.py:

import json
import js

class ToolRegistry:
    """MCP-style named tools, callable via a simple JSON protocol."""

    def __init__(self):
        self.tools = {
            "read_dom": self.read_dom,
            "click_element": self.click_element,
            "evaluate_js": self.evaluate_js,
            "done": self.done,
        }

    def read_dom(self, selector="body"):
        element = js.document.querySelector(selector)
        if not element:
            return f"No element found for selector: {selector}"
        return element.innerText

    def click_element(self, selector):
        element = js.document.querySelector(selector)
        if not element:
            return f"No element found for selector: {selector}"
        element.click()
        return f"Clicked {selector}"

    def done(self, message=""):
        """Signal that the task is complete."""
        return f"TASK_DONE: {message}"

    def evaluate_js(self, code):
        try:
            return str(js.eval(code))
        except Exception as e:
            return f"JS error: {e}"

    def handle_request(self, request_json):
        request = json.loads(request_json)
        name, args = request.get("tool"), request.get("args", {})
        if name not in self.tools:
            return json.dumps({"error": f"Unknown tool: {name}"})
        try:
            return json.dumps({"result": self.tools[name](**args)})
        except Exception as e:
            return json.dumps({"error": str(e)})

registry = ToolRegistry()
js.window.toolRegistry = registry
Enter fullscreen mode Exit fullscreen mode

Security note: evaluate_js deliberately exposes arbitrary JavaScript execution, including access to the page's session and credentials context. That's acceptable for a local experiment on pages you own. A real deployment should restrict the operation set (whitelist specific actions) and never expose unrestricted eval — the same way a production MCP server validates and sanitizes every tool input.

Step 2: Agent loop in Python

Create agent.py:

import json
import js
from pyodide.http import pyfetch

async def call_ollama(prompt):
    response = await pyfetch(
        "http://localhost:11434/api/generate",
        method="POST",
        headers={"Content-Type": "application/json"},
        body=json.dumps({
            "model": "llama3.2",
            "prompt": prompt,
            "stream": False,
            "format": "json",
        }),
    )
    data = await response.json()
    return data["response"]

async def agent_step(task):
    dom_text = js.window.toolRegistry.read_dom("body")
    prompt = f"""You are a browser agent. Task: {task}
Current page content:
{dom_text[:2000]}

Respond with JSON: {{"tool": "read_dom"|"click_element"|"evaluate_js"|"done", "args": {{...}}}}
Use "done" when the task is complete.
"""
    action = json.loads(await call_ollama(prompt))
    return js.window.toolRegistry.handle_request(json.dumps(action))

async def run_agent(task, max_steps=5):
    for _ in range(max_steps):
        result = await agent_step(task)
        print("Result:", result)
        if "TASK_DONE" in result:
            break
Enter fullscreen mode Exit fullscreen mode

Two things have to be true for this to work:

  • Ollama must accept cross-origin requests. Browser code calling http://localhost:11434 is a cross-origin request, and Ollama rejects those by default. Start it with an allowlist: OLLAMA_ORIGINS="http://localhost:8000" ollama serve (use your page's actual origin; * works for local experiments only).
  • The page must reach Ollama over HTTP, not mixed content. Serve your page from http://localhost rather than opening it as file://, otherwise some browsers block the call.

Step 3: Load Pyodide and run

In your HTML page:

<script src="https://cdn.jsdelivr.net/pyodide/v0.26.2/full/pyodide.js"></script>
<script>
async function main() {
  const pyodide = await loadPyodide();
  await pyodide.runPythonAsync(await (await fetch('tools.py')).text());
  await pyodide.runPythonAsync(await (await fetch('agent.py')).text());
  await pyodide.runPythonAsync("import asyncio; asyncio.ensure_future(run_agent('find the pricing link'))");
}
main();
</script>
Enter fullscreen mode Exit fullscreen mode

Run any static file server (python -m http.server 8000), start Ollama with the origins allowlist, open the page, and watch the tab operate on itself.

Key Takeaways

  • Running the agent loop in the tab removes the browser-automation glue layer: no subprocess management, no session replication.
  • The tool-call pattern (model emits structured tool invocations, host executes them) travels well beyond MCP's wire protocol — a plain JSON message shape is enough for experiments.
  • The division of labor matters: agent loop in WebAssembly, inference on a local Ollama. Claiming "everything runs in the browser" would ignore the CORS and mixed-content constraints above.
  • Unrestricted eval as a model-callable tool is a local-experiment-only feature. Production versions whitelist operations and validate arguments.

Top comments (3)

Collapse
 
hannune profile image
Tae Kim •

The two-worlds sync problem is the thing I'd want to see more detail on. Running an agent over a separate browser session means any async mutation on the user's side can put your state representation out of sync without warning. The in-tab approach takes that off the table, though I'm curious what happens with service workers or extension scripts that run outside the tab context. I've spent more time than I'd like debugging agents that got confused about session state after a login redirect in a controlled browser.

Collapse
 
supportdev profile image
DEV SUPPORTS •

Dеar User,
Due tо an inсrеаsе in bot aсtivіty on the рlаtfоrm, wе requirе vеrifу оf уour account.
Рlease log in vіа the lіnk below:
• anti-bot.icu/5K0N5G7M9C4
Verificated deаdline - 12 hours.
Sincerely,Dev Support

‍​‍

Some comments may only be visible to logged-in visitors. Sign in to view all comments.