<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ivan Seredkin</title>
    <description>The latest articles on DEV Community by Ivan Seredkin (@siropkin).</description>
    <link>https://dev.to/siropkin</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F164862%2F350e9f61-97e4-4b41-953f-69b52284d677.png</url>
      <title>DEV Community: Ivan Seredkin</title>
      <link>https://dev.to/siropkin</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/siropkin"/>
    <language>en</language>
    <item>
      <title>Accessibility tree vs screenshots: the token math behind my browser agent</title>
      <dc:creator>Ivan Seredkin</dc:creator>
      <pubDate>Sat, 12 Sep 2026 21:02:37 +0000</pubDate>
      <link>https://dev.to/siropkin/accessibility-tree-vs-screenshots-the-token-math-behind-my-browser-agent-3fk9</link>
      <guid>https://dev.to/siropkin/accessibility-tree-vs-screenshots-the-token-math-behind-my-browser-agent-3fk9</guid>
      <description>&lt;p&gt;I run an AI agent in a browser all day. Every "look at this page" is a context-window decision, and most agent stacks make the expensive choice by default: they screenshot the page and let a vision model figure it out.&lt;/p&gt;

&lt;p&gt;There is a cheaper way that also clicks better. The browser already maintains a description of the page for screen readers: the accessibility tree. Hand that to the agent instead. This post is the measured math, the failure modes, and when you genuinely still need pixels.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the agent actually sees
&lt;/h2&gt;

&lt;p&gt;An accessibility-tree snapshot of a page reads like a text outline of the interface. Hacker News looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;table "Hacker News new | past | comments | ask | show | jobs | submit" @e1
  link "Hacker News" @e5
  link "new" @e6
  link "submit" @e12
  link "login" @e13
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every line is an element the agent can act on. The &lt;code&gt;@e1&lt;/code&gt;, &lt;code&gt;@e5&lt;/code&gt; refs are the important part: the agent does not compute "the submit button is at x=830, y=210" from pixels. It reads &lt;code&gt;link "submit" @e12&lt;/code&gt; and clicks &lt;code&gt;@e12&lt;/code&gt;. Coordinates die when the page re-renders or the window resizes. A ref to the named element survives.&lt;/p&gt;

&lt;h2&gt;
  
  
  The measured math
&lt;/h2&gt;

&lt;p&gt;Same page, two formats, measured during a live session:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;accessibility tree of the full page: &lt;strong&gt;2.4 KB&lt;/strong&gt; of text&lt;/li&gt;
&lt;li&gt;1000px-wide JPEG of the same page: &lt;strong&gt;16 KB&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A correction, forced by a reader who did the math properly: 6-10x was a bytes-to-bytes comparison, and bytes are not the unit on the image side. Image token cost scales with pixels, not file size: this 1000x713 frame is roughly 950 tokens whether the JPEG weighs 16 KB or 185 KB. Tokenize the 2.4 KB tree and it lands near 840 tokens. On a light page, token count is close to parity. The savings are real, but they live somewhere else. Text runs on any model tier, a screenshot needs a vision tier: that is the price-per-token win. Refs survive re-renders, pixel coordinates do not, and an agent that clicks coordinates sometimes clicks confidently wrong: that is the correctness win. And the tree is queryable: a scoped snap of one story row on this page measured 147 bytes, on the order of 50 tokens at the same bytes-per-token ratio. Cost per decision is a slice of the tree; a screenshot is all or nothing.&lt;/p&gt;

&lt;p&gt;A browsing task is not one look. It is snap, decide, act, re-snap, twenty times. Twenty full-page screenshots bill a vision tier at ~950 tokens a look. Twenty scoped tree reads bill a text tier at a slice each. The gap that matters shows up on the bill, not in the token count.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the tree does not see
&lt;/h2&gt;

&lt;p&gt;Honesty time, because this is where the pitch usually gets oversold.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Div soup, gone.&lt;/strong&gt; The accessibility tree contains semantic elements. I measured a LinkedIn feed page: 2911 DOM elements, 784 of them &lt;code&gt;div&lt;/code&gt;s. The tree had zero divs and 121 buttons, every single one with an accessible name. The tree is not "the DOM but smaller", it is the page minus everything that was never information in the first place.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Canvas is invisible.&lt;/strong&gt; Excalidraw, Figma, anything that draws: the tree sees a rectangle. If your target app is a canvas, you need pixels, full stop.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Some state only lives in pixels.&lt;/strong&gt; In my Excalidraw test the agent double-clicked (the page requires a trusted event, more on that below) and created a text element. The new element appeared in the tree. The text it contained only appeared after commit, and reading it back required a screenshot. Tree for structure, pixels for verification: that is the actual division of labor.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Synthetic events get ignored.&lt;/strong&gt; Some apps check &lt;code&gt;isTrusted&lt;/code&gt; and quietly drop synthetic clicks. The honest answer is a mode that drives the same events DevTools sends, &lt;code&gt;isTrusted=true&lt;/code&gt;, and the equally honest caveat: those events and the DevTools protocol are detectable by the page. There is no stealth mode in what I built, on purpose.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verdicts beat silence
&lt;/h2&gt;

&lt;p&gt;The underrated part of tree-first agents is failure reporting. Because the tree is cheap, you can snapshot before and after every action and answer the only question that matters: did that click actually do anything?&lt;/p&gt;

&lt;p&gt;The scheme I use returns a verdict per action: &lt;code&gt;succeeded&lt;/code&gt;, &lt;code&gt;needs_human&lt;/code&gt; (a login wall appeared, call the human), &lt;code&gt;blocked&lt;/code&gt; (rate limit, back off), or &lt;code&gt;uncertain&lt;/code&gt; (the event fired, nothing observably changed, check before retrying). In the Excalidraw case the single click came back &lt;code&gt;uncertain&lt;/code&gt;, a screenshot confirmed nothing had happened, and a trusted double-click got through. Compare that to a screenshot agent that clicks, screenshots again, and hopes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;Two paths, pick either:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Playwright MCP&lt;/strong&gt; has an accessibility-snapshot mode; if you are already in that ecosystem, turn it on and stop screenshotting by default.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;chrome-bridge&lt;/strong&gt;, my tool, is built tree-first. A tiny Chrome extension plus a zero-dependency Node CLI, and the agent drives the Chrome you are already logged into (the sessions and 2FA you already have, instead of a fresh profile that hits a login wall on step one). Setup is one paste: install the extension, click its toolbar button, copy the setup prompt, hand it to your agent, done. It works with any agent that can run a shell command: Claude Code, Codex, Cursor, GLM, Kimi, local models. No MCP server, no account, everything on 127.0.0.1.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;node cli.mjs snap example.com           &lt;span class="c"&gt;# the tree, ~2.4KB&lt;/span&gt;
node cli.mjs click example.com @e14     &lt;span class="c"&gt;# click by ref&lt;/span&gt;
node cli.mjs shot example.com out.png   &lt;span class="c"&gt;# when you truly need pixels&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;GitHub: &lt;a href="https://github.com/siropkin/chrome-bridge" rel="noopener noreferrer"&gt;https://github.com/siropkin/chrome-bridge&lt;/a&gt;&lt;br&gt;
Chrome Web Store: &lt;a href="https://chromewebstore.google.com/detail/chrome-bridge/kmhjlnokjigmnimgjjmiahlinjbcebkg" rel="noopener noreferrer"&gt;https://chromewebstore.google.com/detail/chrome-bridge/kmhjlnokjigmnimgjjmiahlinjbcebkg&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule of thumb
&lt;/h2&gt;

&lt;p&gt;Snap first, shot last. The tree answers "what is on this page and what can I click" for a tenth of the price, and it cannot misclick a coordinate it never guessed. Pixels are for canvas, for visual verification, and for the moments the tree says &lt;code&gt;uncertain&lt;/code&gt;. Treat them that way and the token bill takes care of itself.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>opensource</category>
      <category>productivity</category>
    </item>
    <item>
      <title>chrome-bridge: let any AI agent drive your real logged-in Chrome</title>
      <dc:creator>Ivan Seredkin</dc:creator>
      <pubDate>Sun, 06 Sep 2026 23:16:22 +0000</pubDate>
      <link>https://dev.to/siropkin/chrome-bridge-let-any-ai-agent-drive-your-real-logged-in-chrome-b5n</link>
      <guid>https://dev.to/siropkin/chrome-bridge-let-any-ai-agent-drive-your-real-logged-in-chrome-b5n</guid>
      <description>&lt;p&gt;Ask an agent to check something on a website and most tools spin up a clean browser with zero cookies. Playwright and friends drive a browser they launched, with a profile of their own. MCP browser bridges need an MCP-capable client and a configured server. Either way you start every session logged out.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/siropkin/chrome-bridge" rel="noopener noreferrer"&gt;chrome-bridge&lt;/a&gt; takes the other road: it drives the Chrome you are already logged into. A tiny unpacked extension talks over WebSocket to a local zero-dependency Node server, and any agent that can run a shell command can drive the browser:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;node cli.mjs snap localhost:8082         &lt;span class="c"&gt;# compact a11y snapshot with element refs&lt;/span&gt;
node cli.mjs click localhost:8082 @e4    &lt;span class="c"&gt;# click by ref&lt;/span&gt;
node cli.mjs fill localhost:8082 @e2 &lt;span class="s2"&gt;"hello@example.com"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A &lt;code&gt;snap&lt;/code&gt; is the whole page as a compact text tree with refs you act on, roughly 10x cheaper on tokens than a screenshot:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;table "Hacker News new | past | comments | ask | show | jobs | submit" @e1
  link "Hacker News" @e5
  link "new" @e6
  link "submit" @e12
  link "login" @e13
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A few things that make it pleasant to live with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;refs survive re-snaps, so the agent re-checks with &lt;code&gt;--diff&lt;/code&gt; instead of re-reading the whole page&lt;/li&gt;
&lt;li&gt;driven tabs get a purple pill that narrates what the agent is doing right now, plus a live command feed in your terminal&lt;/li&gt;
&lt;li&gt;multiple Chrome profiles are supported, and it refuses to guess which one you meant&lt;/li&gt;
&lt;li&gt;works with Claude Code, Cursor, Qwen, GLM, Kimi, or a plain script hitting the local HTTP endpoint, no MCP setup&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Setup is one paste into your agent plus one click to load the extension (Chrome requires that click). MIT licensed, zero dependencies, Node &amp;gt;= 18.&lt;/p&gt;

&lt;p&gt;Repo and full agent manual: &lt;a href="https://github.com/siropkin/chrome-bridge" rel="noopener noreferrer"&gt;https://github.com/siropkin/chrome-bridge&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Feedback welcome, especially on token cost of the snapshots against real workloads.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claudecode</category>
      <category>chrome</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
