I run an AI agent in a browser all day. Every "look at this page" is a context-window decision, and most agent stacks make the expensive choice by default: they screenshot the page and let a vision model figure it out.
There is a cheaper way that also clicks better. The browser already maintains a description of the page for screen readers: the accessibility tree. Hand that to the agent instead. This post is the measured math, the failure modes, and when you genuinely still need pixels.
What the agent actually sees
An accessibility-tree snapshot of a page reads like a text outline of the interface. Hacker News looks like this:
table "Hacker News new | past | comments | ask | show | jobs | submit" @e1
link "Hacker News" @e5
link "new" @e6
link "submit" @e12
link "login" @e13
Every line is an element the agent can act on. The @e1, @e5 refs are the important part: the agent does not compute "the submit button is at x=830, y=210" from pixels. It reads link "submit" @e12 and clicks @e12. Coordinates die when the page re-renders or the window resizes. A ref to the named element survives.
The measured math
Same page, two formats, measured during a live session:
- accessibility tree of the full page: 2.4 KB of text
- 1000px-wide JPEG of the same page: 16 KB
A correction, forced by a reader who did the math properly: 6-10x was a bytes-to-bytes comparison, and bytes are not the unit on the image side. Image token cost scales with pixels, not file size: this 1000x713 frame is roughly 950 tokens whether the JPEG weighs 16 KB or 185 KB. Tokenize the 2.4 KB tree and it lands near 840 tokens. On a light page, token count is close to parity. The savings are real, but they live somewhere else. Text runs on any model tier, a screenshot needs a vision tier: that is the price-per-token win. Refs survive re-renders, pixel coordinates do not, and an agent that clicks coordinates sometimes clicks confidently wrong: that is the correctness win. And the tree is queryable: a scoped snap of one story row on this page measured 147 bytes, on the order of 50 tokens at the same bytes-per-token ratio. Cost per decision is a slice of the tree; a screenshot is all or nothing.
A browsing task is not one look. It is snap, decide, act, re-snap, twenty times. Twenty full-page screenshots bill a vision tier at ~950 tokens a look. Twenty scoped tree reads bill a text tier at a slice each. The gap that matters shows up on the bill, not in the token count.
What the tree does not see
Honesty time, because this is where the pitch usually gets oversold.
Div soup, gone. The accessibility tree contains semantic elements. I measured a LinkedIn feed page: 2911 DOM elements, 784 of them divs. The tree had zero divs and 121 buttons, every single one with an accessible name. The tree is not "the DOM but smaller", it is the page minus everything that was never information in the first place.
Canvas is invisible. Excalidraw, Figma, anything that draws: the tree sees a rectangle. If your target app is a canvas, you need pixels, full stop.
Some state only lives in pixels. In my Excalidraw test the agent double-clicked (the page requires a trusted event, more on that below) and created a text element. The new element appeared in the tree. The text it contained only appeared after commit, and reading it back required a screenshot. Tree for structure, pixels for verification: that is the actual division of labor.
Synthetic events get ignored. Some apps check isTrusted and quietly drop synthetic clicks. The honest answer is a mode that drives the same events DevTools sends, isTrusted=true, and the equally honest caveat: those events and the DevTools protocol are detectable by the page. There is no stealth mode in what I built, on purpose.
Verdicts beat silence
The underrated part of tree-first agents is failure reporting. Because the tree is cheap, you can snapshot before and after every action and answer the only question that matters: did that click actually do anything?
The scheme I use returns a verdict per action: succeeded, needs_human (a login wall appeared, call the human), blocked (rate limit, back off), or uncertain (the event fired, nothing observably changed, check before retrying). In the Excalidraw case the single click came back uncertain, a screenshot confirmed nothing had happened, and a trusted double-click got through. Compare that to a screenshot agent that clicks, screenshots again, and hopes.
Try it
Two paths, pick either:
Playwright MCP has an accessibility-snapshot mode; if you are already in that ecosystem, turn it on and stop screenshotting by default.
chrome-bridge, my tool, is built tree-first. A tiny Chrome extension plus a zero-dependency Node CLI, and the agent drives the Chrome you are already logged into (the sessions and 2FA you already have, instead of a fresh profile that hits a login wall on step one). Setup is one paste: install the extension, click its toolbar button, copy the setup prompt, hand it to your agent, done. It works with any agent that can run a shell command: Claude Code, Codex, Cursor, GLM, Kimi, local models. No MCP server, no account, everything on 127.0.0.1.
node cli.mjs snap example.com # the tree, ~2.4KB
node cli.mjs click example.com @e14 # click by ref
node cli.mjs shot example.com out.png # when you truly need pixels
GitHub: https://github.com/siropkin/chrome-bridge
Chrome Web Store: https://chromewebstore.google.com/detail/chrome-bridge/kmhjlnokjigmnimgjjmiahlinjbcebkg
The rule of thumb
Snap first, shot last. The tree answers "what is on this page and what can I click" for a tenth of the price, and it cannot misclick a coordinate it never guessed. Pixels are for canvas, for visual verification, and for the moments the tree says uncertain. Treat them that way and the token bill takes care of itself.
Top comments (7)
@siropkin, “tree for structure, pixels for verification” is an excellent operating rule. I especially like the per-action verdicts (
succeeded,needs_human,blocked,uncertain) because they keep an agent from treating lack of visible change as permission to retry blindly. I’m exploring a similar before/after evidence boundary in agent-inspect; do you retain the accessibility snapshots as a compact trajectory so failures can be replayed or regression-tested later?Thanks, and nice framing on the evidence boundary. Current state: verdicts are computed live per action (succeeded / needs_human / blocked / uncertain) and land in the command history with ok/fail.
history --batch outexports that log as a replayable batch script, failed lines commented out, secrets redacted. The trees themselves aren't persisted by the bridge: the agent owns them, and the command log is the stored trajectory. Honest caveat: replay isn't idempotent, pages move. A stored snapshot per step that replays as a diff is the natural next version of that, and two people asking in one week is a strong signal. It's on the list.The 6 to 10x is a bytes-to-bytes comparison, and bytes are not the unit on the image side: image cost is derived from pixel dimensions, so the JPEG size never enters it. I took the same page you used, a 1000x713 frame of Hacker News in Chrome 152, and re-encoded the identical frame at different qualities: 55.3 KB at quality 30, 184.7 KB at quality 95. That is a 3.3x swing in bytes and zero change in context cost either way, 951 tokens under the width-times-height-over-750 rule, or 4 tiles and 765 tokens under the 85-plus-170-per-tile scheme. Then run your own 2.4 KB tree through a tokenizer: about 840 tokens on o200k_base, which lands it level with the screenshot rather than an order of magnitude under it. Boundary on that last number - my snapshot tool emits a much fatter tree than yours on the same page, 13.2 KB and 4631 tokens, so I used your figure rather than mine, and the tile formula is one vendor's. None of this touches the parts that carry the post: a ref survives a re-render and a coordinate does not, and text goes to any model tier while pixels need a vision tier. The second one looks like the real saving, and it is an argument about price per token rather than about token count.
Fair catch, and you did the work I should have: I compared bytes to bytes and called it tokens. On your numbers it's ~951 vs ~840 on a light page, so token count is close to parity and the headline unit was wrong. I've corrected the framing in the post. What survives is what you already named: text runs on any model tier (price per token, not token count), and a ref survives a re-render while a coordinate doesn't. One thing your math doesn't cover: the tree is queryable. I just measured a scoped snap of one story row on HN: 147 bytes, on the order of 50 tokens at your bytes-per-token ratio, vs ~951 for the screenshot. Cost per decision is a slice of the tree; a screenshot is all or nothing. Thanks for grading your own tool in the same breath.
The token math is what made this click for me: a screenshot is an opaque blob that only a vision model can read, while an accessibility snapshot is already text, so the cheap path and the correct path are the same path. Running browser automation against a headless session, I hit the same wall — screenshots cost tokens and still hallucinate fine print — and swapping to a text representation of the live DOM was the single biggest context saving I made.
The per-action verdicts (succeeded / needs_human / blocked / uncertain) are the part I want to steal: most agents treat "nothing visibly changed" as permission to retry blind. Do you snapshot the accessibility tree per step and replay it as a trajectory when a run fails, or is the verdict computed live and then discarded? Replaying failed runs on a stored snapshot would turn regressions into diffs.
Glad the math landed. Answer on replay, same as above but shorter: verdicts are live, the command log is what's stored,
history --batchturns a failed run into a replayable script. Snapshots aren't kept by the bridge today, so regression-diffing trees is on the list rather than in it. If agent-inspect lands the stored trajectory first, write it up, I'll be reading.Some comments may only be visible to logged-in visitors. Sign in to view all comments.