Most "AI in the browser" extensions I tried had the same three problems: closed source, locked to one vendor's model, and a monthly subscription. So I built WebBrain — a free, GPL-3.0 browser extension that puts an agent in your sidebar and lets you pick the model, including a fully local one.
- Site: https://webbrain.one
- Code: https://github.com/webbrain-one/webbrain
- Works in Chrome, Edge, Brave, Firefox
What it does
It has three explicit modes, because I don't want an agent clicking things I didn't ask for:
| Mode | Can do |
|---|---|
| Ask (default) | Read-only: read the page, answer questions, extract tables/lists, read PDFs |
| Act | Click, type, navigate, fill forms, upload/download — and it asks before consequential actions |
| Dev | Page source, styles, console, network, reversible page edits |
Typical things I use it for: "summarize this thread", "pull every product name and price on this page into a table", "fill this form from what you remember about me, but stop before submit".
Bring your own model (or none at all)
Any OpenAI-compatible endpoint works. For local:
llama-server -m your-model.gguf --port 8080 # llama.cpp
ollama serve # Ollama -> :11434/v1
vllm serve your-model --port 8000 # vLLM
LM Studio, Jan, LocalAI, GPT4All, etc. work the same way. Cloud options (OpenAI, Claude, OpenRouter, Gemini, DeepSeek, Mistral, Grok) are there if you want them, with your own key.
A few things I learned making it usable on small local models:
- Context is the bottleneck, not intelligence. Give the model at least a 16k context window (8k works in Compact mode). WebBrain trims history and caps tool output automatically as it gets close to the limit.
- Keep the tool surface small. Local models get confused by 40 tools. Narrow, stable tool schemas matter more than clever prompts.
- Screenshots are expensive. It leans on the accessibility tree first and only sends screenshots when needed.
With a local model, no page data leaves your machine. No telemetry, no account.
New: use it as an MCP server
This is the part I'm most excited about. Coding agents like Claude Code, Codex, Cursor or OpenCode can now drive your real, already-signed-in browser through WebBrain — no cookie export, no headless login, no credential replay.
Claude Code --stdio--> webbrain-mcp --ws://127.0.0.1:17374--> WebBrain extension --> your tabs
claude mcp add --transport stdio webbrain -- npx -y @webbrain/mcp-server
Then in WebBrain → Settings → General → Advanced → MCP, point it at ws://127.0.0.1:17374/extension. The permission gate stays in the browser, so the coding agent still can't do consequential things without your approval. (Chromium-only for now; Firefox MV3 has no offscreen document to host the bridge.)
What I'd love feedback on
- Which local models are you running for agentic/tool-use tasks? I'm collecting real-world results.
- Sites where it fails — issues with a URL and the task are gold.
- Anything about the safety model (Ask/Act/Dev + confirmations) that feels too loose or too annoying.
Issues and PRs welcome: https://github.com/webbrain-one/webbrain/issues
Top comments (0)