DEV Community

Emre Sokullu
Emre Sokullu

Posted on

I built an open-source browser agent that runs on your local LLM (and now plugs into Claude Code via MCP)

Most "AI in the browser" extensions I tried had the same three problems: closed source, locked to one vendor's model, and a monthly subscription. So I built WebBrain — a free, GPL-3.0 browser extension that puts an agent in your sidebar and lets you pick the model, including a fully local one.

What it does

It has three explicit modes, because I don't want an agent clicking things I didn't ask for:

Mode Can do
Ask (default) Read-only: read the page, answer questions, extract tables/lists, read PDFs
Act Click, type, navigate, fill forms, upload/download — and it asks before consequential actions
Dev Page source, styles, console, network, reversible page edits

Typical things I use it for: "summarize this thread", "pull every product name and price on this page into a table", "fill this form from what you remember about me, but stop before submit".

Bring your own model (or none at all)

Any OpenAI-compatible endpoint works. For local:

llama-server -m your-model.gguf --port 8080   # llama.cpp
ollama serve                                   # Ollama  -> :11434/v1
vllm serve your-model --port 8000              # vLLM
Enter fullscreen mode Exit fullscreen mode

LM Studio, Jan, LocalAI, GPT4All, etc. work the same way. Cloud options (OpenAI, Claude, OpenRouter, Gemini, DeepSeek, Mistral, Grok) are there if you want them, with your own key.

A few things I learned making it usable on small local models:

  • Context is the bottleneck, not intelligence. Give the model at least a 16k context window (8k works in Compact mode). WebBrain trims history and caps tool output automatically as it gets close to the limit.
  • Keep the tool surface small. Local models get confused by 40 tools. Narrow, stable tool schemas matter more than clever prompts.
  • Screenshots are expensive. It leans on the accessibility tree first and only sends screenshots when needed.

With a local model, no page data leaves your machine. No telemetry, no account.

New: use it as an MCP server

This is the part I'm most excited about. Coding agents like Claude Code, Codex, Cursor or OpenCode can now drive your real, already-signed-in browser through WebBrain — no cookie export, no headless login, no credential replay.

Claude Code --stdio--> webbrain-mcp --ws://127.0.0.1:17374--> WebBrain extension --> your tabs
Enter fullscreen mode Exit fullscreen mode
claude mcp add --transport stdio webbrain -- npx -y @webbrain/mcp-server
Enter fullscreen mode Exit fullscreen mode

Then in WebBrain → Settings → General → Advanced → MCP, point it at ws://127.0.0.1:17374/extension. The permission gate stays in the browser, so the coding agent still can't do consequential things without your approval. (Chromium-only for now; Firefox MV3 has no offscreen document to host the bridge.)

What I'd love feedback on

  • Which local models are you running for agentic/tool-use tasks? I'm collecting real-world results.
  • Sites where it fails — issues with a URL and the task are gold.
  • Anything about the safety model (Ask/Act/Dev + confirmations) that feels too loose or too annoying.

Issues and PRs welcome: https://github.com/webbrain-one/webbrain/issues

Top comments (0)