Ask a frontier vision model for the pixel coordinates of a "Submit" button in a 1920x1080 screenshot and it will often miss by dozens of pixels. That miss is the difference between an agent that completes a checkout flow and one that clicks empty whitespace and loops until it hits a timeout. The fix is not a bigger model. It is a better interface between the model and the screen.
That interface has two parts. Screen painting draws numbered marks onto the screenshot before the model sees it. Visual grounding maps the model's answer back to a real, clickable element. Together they turn a coordinate-regression problem into a multiple-choice question, and models handle multiple-choice questions far more reliably.
Why Raw Coordinates Fail
Vision-language models are trained mostly to describe images, not to regress precise coordinates. Anthropic's Computer Use and OpenAI's Operator both shipped with explicit caveats about accuracy on dense interfaces. The failure modes are predictable:
- Resolution scaling. Many APIs downscale images before inference. Claude's documentation, for example, recommends keeping screenshots at or below roughly XGA (1024x768) and scaling coordinates yourself. If you skip that step, every click lands off-target by the scale factor.
- Dense targets. Toolbars, table rows, and dropdown items sit 20–30 pixels apart. A small error picks the wrong row.
- Visually identical elements. Five "Edit" buttons in a list are indistinguishable without context.
Takeaway: Before you blame the model, log every predicted coordinate alongside the screenshot and the intended target. Most "model errors" turn out to be scaling bugs.
Screen Painting with Set-of-Mark Prompting
Set-of-Mark (SoM) prompting, introduced by Microsoft Research in 2023, overlays numbered labels on segmented regions of an image. Instead of asking "where is the login button?", you ask "which number is the login button?" The model only has to read a label. It no longer has to estimate a position.
For web UIs, you don't need a segmentation model. The DOM already knows where everything is. Here is a minimal Playwright implementation that paints interactive elements:
from playwright.sync_api import sync_playwright
PAINT_JS = """
() => {
const sel = 'a, button, input, select, textarea, [role=button], [onclick]';
const els = [...document.querySelectorAll(sel)].filter(e => {
const r = e.getBoundingClientRect();
return r.width > 0 && r.height > 0 && r.top < innerHeight;
});
return els.map((e, i) => {
const r = e.getBoundingClientRect();
const tag = document.createElement('div');
tag.textContent = i;
tag.style.cssText = `position:fixed;left:${r.left}px;top:${r.top}px;
background:#e11;color:#fff;font:bold 12px monospace;
padding:1px 3px;z-index:2147483647;pointer-events:none`;
tag.className = '__som';
document.body.appendChild(tag);
return {id: i, x: r.left + r.width/2, y: r.top + r.height/2,
text: (e.innerText || e.value || e.ariaLabel || '').slice(0, 40)};
});
}
"""
with sync_playwright() as p:
page = p.chromium.launch().new_page(viewport={"width": 1280, "height": 800})
page.goto("https://news.ycombinator.com")
marks = page.evaluate(PAINT_JS)
page.screenshot(path="painted.png")
page.evaluate("document.querySelectorAll('.__som').forEach(e => e.remove())")
You now have a painted screenshot and a lookup table from mark ID to center coordinates. Send both the image and a compact text list ([12] "login") to the model. Then require it to answer with an ID.
Takeaway: Always remove the overlay before executing the action. Set pointer-events:none so the labels can never intercept clicks.
Grounding Beyond the Browser
Desktop apps, Citrix sessions, and canvas-heavy UIs like Figma have no DOM to query. You have three options for finding elements there:
-
Accessibility trees. On Windows, use UI Automation (via
pywinautowithbackend="uia"). On macOS, use the AX API. On Linux, use AT-SPI. These expose bounding boxes for native controls. Electron apps often expose the most data once accessibility is enabled. - Detection models. Microsoft's OmniParser is open source on GitHub and Hugging Face. It combines a fine-tuned YOLO detector for interactable regions with a captioning model, and it outputs SoM-ready boxes from pixels alone.
- OCR fallback. Tesseract or PaddleOCR can locate text labels when nothing else works. This handles a surprising share of enterprise forms.
In production, layer these sources. Query the accessibility tree first because it is fast and exact. Fill the gaps with detection. Use OCR to disambiguate.
Takeaway: Run pip install pywinauto and dump app.window().print_control_identifiers() on your target app. If most controls appear, you can skip vision-based detection entirely for that app.
Making It Production-Grade
Grounding solves targeting. Reliability comes from what you build around it:
- Verify every action. After each click, take a new screenshot and check for an expected change: a URL change, a DOM mutation, or a pixel diff above a threshold. If nothing changed, retry with the next-best candidate rather than repeating the same click.
- Cap mark density. Painting 300 labels makes the screenshot unreadable. Crop to the active region, or let the model zoom: first pick a quadrant, then pick a mark inside it.
- Cache by layout hash. Hash the set of element roles and texts. If the layout matches a page you have seen before, reuse the earlier grounding decision and skip inference. This cuts both latency and cost on repetitive workflows.
- Evaluate on public benchmarks. ScreenSpot measures pure grounding accuracy across web, mobile, and desktop. OSWorld and WebArena measure end-to-end task success. Report both, because good grounding does not guarantee task completion.
Takeaway: Add a post-action verification step before adding any other feature. It converts silent failures into recoverable retries.
Here is something you can do today. Take the Playwright snippet above and run it against one internal web app your team automates. Then send painted.png plus the mark list to your current model. Ask it to complete five real tasks by returning mark IDs only, and record how many succeed. Run the same five tasks with raw coordinate prompting. That side-by-side number will tell you whether to invest in screen painting for your agent, and it will take less than an hour to get.
Top comments (3)
One addition from working on the accessibility-tree layer on macOS: a few traps make AX look emptier than it really is, and people give up and reach for vision too early.
On the layout-hash cache: hash roles plus stable labels only, and drop anything that looks like a count, timestamp or price. Otherwise an inbox with an unread badge never hits the cache twice.
Great write-up — the shift from coordinate regression to multiple-choice really clicks. Curious about a concrete case: on a dense data table with 30+ rows of near-identical action links, did the numbered marks still hold up, or did you have to scope the painted region first?
The 'verify every action' line deserves the most weight here — in our experience, reliability is won or lost at the verification step far more than in grounding accuracy. One caveat: verification screenshots can mislead on pages with region-local animation, so comparing against a masked region matters. That also pairs nicely with capping mark density by zooming, since a smaller region is cheaper to re-check.