DEV Community

Gayatri Kakumanu
Gayatri Kakumanu

Posted on Originally published at gayatrikakumanu.hashnode.dev AI-assisted

Your <div onClick> Is Invisible to AI Agents

I Found Out What an AI Agent Actually Sees When It Looks at My Buttons

Imagine describing your app to someone over the phone. You can't send a screenshot. You just read out what's on the page: "There's a heading that says Checkout. A text box for email. A button that says Place order."

That's roughly what an AI agent gets when it uses your app.

I wanted to know exactly what that description looks like, so I built a page with nine different buttons and printed what the agent receives for each one. Some of it surprised me.

Same button, two very different stories

First, a quick split

Browser agents come in two flavors.

Some look at pictures. They take a screenshot and click where the button appears. Claude computer use and OpenAI's computer-using agent work this way.

Some read a description. They never see pixels at all. They get a text list of what's on the page — each thing's type, its label, and its current state. Then they act on items from that list. browser-use, Stagehand, Microsoft's Playwright MCP and Vercel's Agent Browser all work this way.

The second group is growing fast, for a boring reason: a text description is way smaller than a screenshot. Cheaper per click, and more precise about what got clicked. Over a fifty-step task, that adds up.

This article is about what that description contains. I used Playwright 1.56 and Chromium 141 to generate it.

Finding 1: Your div isn't invisible. It's a guess.

I expected a <div onClick> to be completely missing from the list. It isn't. Here's what came back:

generic [ref=e2] [cursor=pointer]: Place order
Enter fullscreen mode Exit fullscreen mode

The agent can see it, and can click it. It even knows the mouse cursor turns into a pointer, so the CSS leaks through a little.

What it doesn't know is what the thing is. "generic" means "some box." So the agent sees a box with the words "Place order" in it, and a pointer cursor, and has to work out that this is probably a button.

Compare the real button:

button "Place order" [ref=e3] [cursor=pointer]
Enter fullscreen mode Exit fullscreen mode

That one says it outright. It's a button, and it's called "Place order." No working out required.

That's the whole difference: one is a fact, the other is a guess. Guesses usually land. That's exactly what makes them miserable to debug when they don't.

Good news if you can't use a real <button> for styling reasons: adding role="button" and tabindex="0" to your div produces an identical line. The escape hatch works. You just have to actually take it.

Finding 2: An icon button can be a button that does nothing knowable

This one made me go back and re-run the test because I didn't believe it.

button [ref=e5]:
  img
Enter fullscreen mode Exit fullscreen mode

The agent knows there's a button. It has no idea what the button does. Not "a rough idea" — the name is empty.

This is worse than the div. At least the div had its text. A trash-can icon is obvious to you and is literally nothing in the description.

One attribute fixes it:

<button aria-label="Delete item" onClick={remove}>
  <TrashIcon />
</button>
Enter fullscreen mode Exit fullscreen mode

Now it reads button "Delete item".

Small trap I hit: if your icon is an emoji instead of an SVG, the emoji becomes the name. Your test looks fine and your real users are still stuck.

Finding 3: ARIA on a plain div gets thrown away

I put aria-disabled="true" on a div, expecting the agent to see a disabled thing. Here's what it got:

generic [ref=e13] [cursor=pointer]: Submit
Enter fullscreen mode Exit fullscreen mode

No mention of disabled anywhere. The real one, for comparison:

button "Submit" [disabled] [ref=e14]
Enter fullscreen mode Exit fullscreen mode

Here's why. ARIA states need something to attach to. A plain div has no role, so there's nothing to hold "disabled," and it gets dropped.

Which means: a button that looks disabled, is styled disabled, and has aria-disabled on it is, to an agent, just a normal clickable box. It will click it.

The fix isn't to add more ARIA. It's to use an element that has a role in the first place.

All nine, side by side

Nine buttons, nine snapshots

Why this matters now

Accessibility usually gets argued two ways: it's the right thing to do, and in a lot of places it's legally required. Both true. Neither one has ever beaten a deadline. "A11y pass" goes on the board and then gets cut.

Here's a third reason, and it's the one that survives sprint planning: semantic HTML is turning into the interface your product gives to automation.

Automated tests. Agents doing things for your users. Any tool that drives your UI instead of calling your API. More and more of that runs through this exact description. If your components are styled divs, those flows don't crash — they guess, they're occasionally wrong, and nobody can reproduce the bug.

Try it on your own app (2 minutes)

I packaged this as one file so you can see what your own components look like.

Step 1 — make a folder and install Playwright

mkdir agent-snapshot && cd agent-snapshot
npm init -y
npm i playwright
npx playwright install chromium
Enter fullscreen mode Exit fullscreen mode

That last line downloads a browser, about 150 MB. It only happens once. (Node 18+.)

Step 2 — grab the script

curl -O https://gist.githubusercontent.com/kakumanu-gayatri/eefa6cc43e83460b612ae678bcaeacd1/raw/agent-snapshot.js
Enter fullscreen mode Exit fullscreen mode

Windows PowerShell:

curl.exe -O https://gist.githubusercontent.com/kakumanu-gayatri/eefa6cc43e83460b612ae678bcaeacd1/raw/agent-snapshot.js
Enter fullscreen mode Exit fullscreen mode

Step 3 — run the demo

node agent-snapshot.js
Enter fullscreen mode Exit fullscreen mode

This runs the same nine buttons from this article, so you can see the output for yourself before pointing it anywhere real.

Step 4 — point it at your app

Start your dev server, then:

node agent-snapshot.js http://localhost:3000
Enter fullscreen mode Exit fullscreen mode

Reading your results

Scan the output for three things:

generic where you meant "button." Your element has no role. The agent is guessing.

button [ref=e5] with nothing in quotes. A button with no name. The agent knows it can click it and has no idea what happens.

Missing [expanded] or [disabled]. Your state lives in a CSS class where nothing can read it.

If you get ERR_CONNECTION_REFUSED, your dev server isn't running yet.

One honest note: _snapshotForAI() is an internal Playwright API, so it can change between versions. The script falls back to the public ariaSnapshot() if it's missing. I measured everything here on Playwright 1.56.0 and Chromium 141.

One honest caveat from my testing: the public ariaSnapshot() and the internal format Playwright MCP actually sends agents don't match exactly. The public one showed my div as plain text; the agent-facing one showed generic with a cursor. If you're debugging a specific agent, check the format it really uses — and note that the internal API can change between versions.

The short version

A <button> tells an agent what it is. A styled div makes it guess.

Most of the time the guess is right. The rest of the time you get a bug report saying "the agent couldn't find the button" — about a button sitting right there in plain sight.


The test page, the script, and the raw output are here: Agent Snapshot Gist. Run it on your own app and see what comes back.

Top comments (1)

Collapse
 
jonrandy profile image
Jon Randy 🎖️ •

Small thing - it should really be onclick not onClick. Ultimately it doesn't matter to the browser, but HTML element attributes should really be lowercase.