DEV Community

Cover image for How our browser agent stopped spending 58% of its tokens
Qweezyy
Qweezyy

Posted on Fully Autonomous

How our browser agent stopped spending 58% of its tokens

TL;DR. A browser agent sees a page as a text accessibility tree. After every click it used to get the whole tree again, even when two lines of the page had changed. We now send the model only the changes (a diff) and strip what the model never uses. On a real 98-step browsing chat, page states went from 355K to 151K tokens (−58%). A reproducible benchmark on four sites gave −63%, and a live A/B run of the agent on the same task gave −54% on browser tool results. Here is how it works, where it nearly broke, and a script to check the numbers yourself.

About this post. The project is Altair, an open-source (Apache-2.0) AI agent for your PC and phone. I'm the author and build it largely with Claude Code. This post, the measurements and the screenshots were prepared by that same AI assistant; I reviewed it and stand behind it. It will also answer the comments while I watch and approve every reply, so treat it as a live test: if something reads as dishonest or spammy, say so and I'll fix it. Every number below is reproducible.

How the agent "sees" a page

A screenshot is expensive and imprecise: to press a button the model has to guess coordinates. So many browser agents, ours included, read the page as its accessibility tree — what a screen reader sees. Playwright returns it in a YAML-like form, and every element gets a short ref:

- banner [ref=e9]:
  - heading "Navigation Menu" [level=2] [ref=e10]
- button "Code" [ref=e180]
- textbox "Search Hacker News" [ref=e684]
Enter fullscreen mode Exit fullscreen mode

The model calls browser_click(ref="e180") and the agent clicks exactly that element.

The question is when the model gets this tree. Every action changes the page, and the model needs to know what it looks like now. The simplest answer is to return the whole tree again. That is what we did.

What it costs

A GitHub repo page is about 20K tokens as a tree, a Wikipedia article 136K. We cut a snapshot at 14,000 characters (≈3.5–4K tokens), and the agent finds the rest with an in-page search. Even so, every click cost 3.5–4K tokens.

Then the easy-to-miss part kicks in: the history is resent on every step. After 30 actions on a page, the history holds 30 near-identical copies of it, and every following request sends them all. Prompt caching lowers the price of resending, but not the room they take in the context window.

In our logs, browser tools made up more than half of all tool tokens. One real 98-step chat had 355K tokens of page states.

The clearest example comes from the benchmark below. The agent clicks the search field on Hacker News. It used to get the whole front page back — 3,690 tokens. Exactly one word changed:

             - text: "Search:"
-            - textbox [ref=e684]
+            - textbox [active] [ref=e684]
Enter fullscreen mode Exit fullscreen mode

Now that answer costs 80 tokens.

Trick 1: drop what the model never uses

  • Unnamed wrappers. - generic [ref=e42]: with no name is a layout div. The model never clicks it. The line goes; its children stay.
  • Redundant [cursor=pointer]. A link, button or tab is clickable by its role. The mark stays only where it says something: on a clickable generic.
  • Long link targets. Tracking links run to 1,000+ characters. Past 150 characters we drop the query string.

Refs are never touched, or clicks would break. On untruncated snapshots this saves 7–27% (Hacker News −15%, GitHub −7%, MDN −27%, Wikipedia −9%). It's modest, and that's the honest result: the main win is elsewhere. It also fits more real content under the 14K-character cut.

Trick 2: after an action, send only the changes

The agent remembers the page the model saw last: tab, URL and tree. After an action, if the tab and URL are the same, the model gets a unified diff with one line of context:

def page_changes(old: str, new: str) -> str | None:
    """"" = nothing changed; None = too much changed, send the whole page."""
    lines = []
    for line in difflib.unified_diff(old.splitlines(), new.splitlines(), lineterm="", n=1):
        if line.startswith(("---", "+++")):
            continue
        lines.append("…" if line.startswith("@@") else line)
    text = "\n".join(lines)
    if not text:
        return ""
    return None if len(text) > 0.5 * max(len(new), 1) else text
Enter fullscreen mode Exit fullscreen mode

The rules around it:

  • The whole page comes on navigation, a new tab, or an explicit browser_read. A diff between two different pages is meaningless.
  • If more than half of the page changed, send it whole. A diff of 80% of a page is longer than the page and harder to read.
  • "Nothing visible changed" is an answer too. It's a useful signal: the click probably didn't work.
  • The header explains the format: - lines were removed, + added, the rest of the page and its refs are as before. No extra system-prompt instructions were needed.

Playwright refs are stable within a page, so refs from the old snapshot keep working after a diff; we tested that on a real browser.

A live run. The task: open the GitHub repo, open the green Code menu and give me the HTTPS clone URL, then open the branch picker and list the branches.

The second click (branch picker) without diffs — the model gets the whole page again, from the dialog to the site header and the file list. 3,386 tokens:

A click result without diffs: the whole page

The first click (Code menu) with diffs — only the menu that appeared, with the HTTPS URL right there. 568 tokens:

A click result with diffs: only the changes

Where it nearly broke: the history

A diff only makes sense while the page it was taken against is still in the history. We already had two mechanisms that clean the history:

  1. Superseding stale snapshots: old page snapshots became a one-line note, the newest two stayed. With diffs, "the newest two" could be two diffs whose base was gone — changes to a page the model can't see.
  2. Budget clearing: near the context limit, old large tool outputs become a note. The largest output in a browsing chat is exactly the whole page.

The fix: superseding keeps the newest two whole pages with every diff after them, and budget clearing never takes the latest whole page. Clearing is batched, because every history edit invalidates the prompt cache from that point.

Measurements

1. A real 98-step chat (measured during development): page states 355K → 270K tokens with leaner snapshots → 151K with diffs. That's the −58%. I can't re-run that exact chat — its stored outputs were already compacted — so here are two independent checks.

2. A reproducible benchmark. Real Chromium, four sites (Wikipedia, GitHub, Hacker News, MDN), a fixed script: open, click search, type, Escape, open a menu. Each step is rendered with the same PageView.render the agent uses, same truncation. Tokens via tiktoken o200k_base (tokenizers differ, the ratios hold).

Benchmark: 17 steps on 4 real sites

17 steps: 60,117 → 22,255 tokens, −63%. Actions that stay on the same page only: 37,449 → 6,865, −82%. Opening a page still costs full price; the win is everything after.

3. A live A/B. Built Altair 0.1.3, model glm-5.3-flash, the same GitHub task twice — diffs on, then off (BROWSER_SNAPSHOT_DIFF=false; lean snapshots in both). Both runs took the same path (navigate + two clicks) and gave the same correct answer.

Live A/B: browser tool results

Browser results: 10,314 → 4,726 tokens, −54%. By the provider's own count, the last request was 20,923 tokens without diffs and 15,113 with them — and that includes the system prompt and tool schemas, which diffs don't touch. On a 30-step task the gap grows with every step, since each saved page is not resent in all the requests after it.

What this doesn't solve

  • The first look at a page costs full price. An agent that mostly follows links gains less.
  • The model has to read diffs. In our runs it did, without instructions — unified diffs are everywhere in training data. But we don't yet have a systematic "task success with vs. without diffs" measurement, only manual checks and ref-stability tests. If your model gets confused, BROWSER_SNAPSHOT_DIFF=false brings back the old behavior.
  • Small samples. Four sites, one live A/B, one historical chat. Shown as they are, no cherry-picked averages.
  • tiktoken is a proxy for GLM, Claude or Gemini tokenizers.

Check it yourself

from playwright.sync_api import sync_playwright
from core.browser_session import PageView, compact_tree, page_changes, SNAPSHOT_LIMIT
import tiktoken

enc = tiktoken.get_encoding("o200k_base")
tok = lambda s: len(enc.encode(s))
render = lambda url, title, tree, ch=None: PageView(
    url=url, title=title, tree=tree[:SNAPSHOT_LIMIT],
    truncated=len(tree) > SNAPSHOT_LIMIT, changes=ch).render()

with sync_playwright() as pw:
    page = pw.chromium.launch().new_page()
    seen = None
    for action in [lambda: page.goto("https://news.ycombinator.com/"),
                   lambda: page.locator("input[name=q]").click(),
                   lambda: page.keyboard.type("agent", delay=40)]:
        action(); page.wait_for_timeout(1200)
        raw = page.aria_snapshot(mode="ai")
        lean = compact_tree(raw)
        ch = page_changes(seen[1], lean) if seen and seen[0] == page.url else None
        seen = (page.url, lean)
        print(tok(render(page.url, page.title(), raw)), "→",
              tok(render(page.url, page.title(), lean, ch)))
Enter fullscreen mode Exit fullscreen mode

Run it from the repo's pc/ folder: it imports the functions straight from the agent, so it measures exactly what ships. Code: https://github.com/Qweezyy/AltairAgent (core/browser_session.py). Diffs are on by default since 0.1.2.

If you've built a browser agent: do you send the model the whole tree, a diff, screenshots, or something else? Have you seen models get confused by diffs?

Top comments (0)