Most browser automation tools built for AI agents start from an inconvenient premise: they assume your task wants a clean, empty browser window.
Tools like Playwright or Puppeteer boot disposable profiles in headless mode. That's fine for running test suites in continuous integration. It falls flat the moment you want an agent to handle daily engineering work, like approving a pull request behind corporate SAML or pulling a report from an internal intranet portal.
Standard headless browsers don't have your session cookies. They trigger Cloudflare turnstiles, get blocked by bot detectors, and can't pass hardware key 2FA prompts.
I built edge-agent-bridge to sidestep this entirely. Instead of fighting bot walls in a sandbox, it drives the browser tabs you already have open: Edge and Chrome both work.
Why Drive the Live Browser?
When an agent controls your open Edge or Chrome tabs:
- Authenticated sessions stay valid. You log in once via SSO; your agent uses that existing state.
- Actions dispatch through Chrome DevTools Protocol as real, trusted input events (
isTrusted: true), not injected synthetic JavaScript clicks. - Corporate VPN tunnels and client certificates work without extra configuration.
Everything runs locally. A small background daemon listens on 127.0.0.1:18999 and coordinates requests between your agent tools and a companion browser extension.
Quick Setup
1. Install the PyPI Package
The client requires Python 3.10+ and has zero third-party dependencies:
pip install edge-agent-bridge
2. Add the Companion Extension
The extension ships on both stores:
- Edge: Agent Browser Bridge on Microsoft Edge Add-ons
- Chrome: Agent Browser Bridge on Chrome Web Store
If you prefer building from source, run:
edge-bridge extension open
That command displays the extension directory path. From there, open edge://extensions, toggle Developer mode on, and pick "Load unpacked".
Verify the background service anytime:
edge-bridge daemon status
Three Ways to Use It
1. Element-Ref CLI
The CLI uses an accessibility snapshot model. It maps visible interactive nodes to short ref tokens (e1, e2):
# Capture the active tab
edge-bridge snapshot
# Output format:
# - textbox "Search" [e1] value=""
# - button "Submit" [e2]
# Click or type by ref
edge-bridge fill e1 "agentic automation"
edge-bridge click e2
Refs prevent accidental clicks. If the page reloads or DOM nodes shift before your command finishes, the bridge flags the stale ref and stops rather than clicking the wrong element.
2. Python Scripting
For standalone automation scripts, import the Edge client:
from edge_agent_bridge import Edge
with Edge() as browser:
snap = browser.snapshot()
browser.fill("e1", "edge-agent-bridge")
browser.click("e2")
browser.wait(text="Done")
The context manager locks onto the active tab from your first query. If you click into another desktop window while the script runs, execution won't wander to the wrong tab.
3. MCP Server for Coding Assistants
You can wire Edge control straight into Claude Code, Cursor, Windsurf, or Gemini CLI:
edge-bridge setup
This registers 29 tools (prefixed with edge_) using the Model Context Protocol stdio transport. Your coding assistant can inspect the DOM tree, fill forms, press key shortcuts, and capture targeted element screenshots.
Security and Privacy
Because the extension interacts with your live tabs, the safety boundaries are intentionally conservative:
- The daemon binds strictly to loopback (
127.0.0.1). It inspects theHostheader and rejects external origins or DNS rebinding with HTTP 421. - Mutating actions require local authorization tokens stored in your user directory.
- No network requests leave your machine. There are no telemetry pings, no remote trackers, no crash reporters, and no hosted cloud relays.
Check out the project:
- PyPI: edge-agent-bridge
- Edge Add-ons: Agent Browser Bridge
- Chrome Web Store: Agent Browser Bridge
Top comments (0)