DEV Community

lucanu
lucanu

Posted on

Visual Grounding for AI Agents: Set-of-Mark Prompting and Reliable UI Clicks

Ask a frontier vision model for the pixel coordinates of a "Submit" button in a 1920x1080 screenshot and it will often miss by dozens of pixels. That miss is the difference between an agent that completes a checkout flow and one that clicks empty whitespace and loops until it hits a timeout. The fix is not a bigger model. It is a better interface between the model and the screen.

That interface has two parts. Screen painting draws numbered marks onto the screenshot before the model sees it. Visual grounding maps the model's answer back to a real, clickable element. Together they turn a coordinate-regression problem into a multiple-choice question, and models handle multiple-choice questions far more reliably.

Why Raw Coordinates Fail

Vision-language models are trained mostly to describe images, not to regress precise coordinates. Anthropic's Computer Use and OpenAI's Operator both shipped with explicit caveats about accuracy on dense interfaces. The failure modes are predictable:

  • Resolution scaling. Many APIs downscale images before inference. Claude's documentation, for example, recommends keeping screenshots at or below roughly XGA (1024x768) and scaling coordinates yourself. If you skip that step, every click lands off-target by the scale factor.
  • Dense targets. Toolbars, table rows, and dropdown items sit 20–30 pixels apart. A small error picks the wrong row.
  • Visually identical elements. Five "Edit" buttons in a list are indistinguishable without context.

Takeaway: Before you blame the model, log every predicted coordinate alongside the screenshot and the intended target. Most "model errors" turn out to be scaling bugs.

Screen Painting with Set-of-Mark Prompting

Set-of-Mark (SoM) prompting, introduced by Microsoft Research in 2023, overlays numbered labels on segmented regions of an image. Instead of asking "where is the login button?", you ask "which number is the login button?" The model only has to read a label. It no longer has to estimate a position.

For web UIs, you don't need a segmentation model. The DOM already knows where everything is. Here is a minimal Playwright implementation that paints interactive elements:

from playwright.sync_api import sync_playwright

PAINT_JS = """
() => {
  const sel = 'a, button, input, select, textarea, [role=button], [onclick]';
  const els = [...document.querySelectorAll(sel)].filter(e => {
    const r = e.getBoundingClientRect();
    return r.width > 0 && r.height > 0 && r.top < innerHeight;
  });
  return els.map((e, i) => {
    const r = e.getBoundingClientRect();
    const tag = document.createElement('div');
    tag.textContent = i;
    tag.style.cssText = `position:fixed;left:${r.left}px;top:${r.top}px;
      background:#e11;color:#fff;font:bold 12px monospace;
      padding:1px 3px;z-index:2147483647;pointer-events:none`;
    tag.className = '__som';
    document.body.appendChild(tag);
    return {id: i, x: r.left + r.width/2, y: r.top + r.height/2,
            text: (e.innerText || e.value || e.ariaLabel || '').slice(0, 40)};
  });
}
"""

with sync_playwright() as p:
    page = p.chromium.launch().new_page(viewport={"width": 1280, "height": 800})
    page.goto("https://news.ycombinator.com")
    marks = page.evaluate(PAINT_JS)
    page.screenshot(path="painted.png")
    page.evaluate("document.querySelectorAll('.__som').forEach(e => e.remove())")
Enter fullscreen mode Exit fullscreen mode

You now have a painted screenshot and a lookup table from mark ID to center coordinates. Send both the image and a compact text list ([12] "login") to the model. Then require it to answer with an ID.

Takeaway: Always remove the overlay before executing the action. Set pointer-events:none so the labels can never intercept clicks.

Grounding Beyond the Browser

Desktop apps, Citrix sessions, and canvas-heavy UIs like Figma have no DOM to query. You have three options for finding elements there:

  1. Accessibility trees. On Windows, use UI Automation (via pywinauto with backend="uia"). On macOS, use the AX API. On Linux, use AT-SPI. These expose bounding boxes for native controls. Electron apps often expose the most data once accessibility is enabled.
  2. Detection models. Microsoft's OmniParser is open source on GitHub and Hugging Face. It combines a fine-tuned YOLO detector for interactable regions with a captioning model, and it outputs SoM-ready boxes from pixels alone.
  3. OCR fallback. Tesseract or PaddleOCR can locate text labels when nothing else works. This handles a surprising share of enterprise forms.

In production, layer these sources. Query the accessibility tree first because it is fast and exact. Fill the gaps with detection. Use OCR to disambiguate.

Takeaway: Run pip install pywinauto and dump app.window().print_control_identifiers() on your target app. If most controls appear, you can skip vision-based detection entirely for that app.

Making It Production-Grade

Grounding solves targeting. Reliability comes from what you build around it:

  • Verify every action. After each click, take a new screenshot and check for an expected change: a URL change, a DOM mutation, or a pixel diff above a threshold. If nothing changed, retry with the next-best candidate rather than repeating the same click.
  • Cap mark density. Painting 300 labels makes the screenshot unreadable. Crop to the active region, or let the model zoom: first pick a quadrant, then pick a mark inside it.
  • Cache by layout hash. Hash the set of element roles and texts. If the layout matches a page you have seen before, reuse the earlier grounding decision and skip inference. This cuts both latency and cost on repetitive workflows.
  • Evaluate on public benchmarks. ScreenSpot measures pure grounding accuracy across web, mobile, and desktop. OSWorld and WebArena measure end-to-end task success. Report both, because good grounding does not guarantee task completion.

Takeaway: Add a post-action verification step before adding any other feature. It converts silent failures into recoverable retries.

Here is something you can do today. Take the Playwright snippet above and run it against one internal web app your team automates. Then send painted.png plus the mark list to your current model. Ask it to complete five real tasks by returning mark IDs only, and record how many succeed. Run the same five tasks with raw coordinate prompting. That side-by-side number will tell you whether to invest in screen painting for your agent, and it will take less than an hour to get.

Top comments (3)

Collapse
 
autenai profile image
Auten •

One addition from working on the accessibility-tree layer on macOS: a few traps make AX look emptier than it really is, and people give up and reach for vision too early.

  • Chrome and Electron apps only build their full AX tree when an assistive client asks for it. Set AXEnhancedUserInterface (Chrome) or AXManualAccessibility (Electron) on the app element first, otherwise you get a handful of nodes and conclude the app is opaque.
  • AX frames are in global screen points (origin at the top-left of the main display), screenshots are in pixels. On Retina that is a 2x factor, and a second monitor can have negative coordinates. It's the same scaling bug you describe, one layer down.
  • For the five identical "Edit" buttons, label each mark with the nearest ancestor's text (the row title) instead of the button's own text. They become distinct choices without any vision call.
  • Where the control supports it, perform AXPress or set AXValue instead of clicking the center point. It often still works when the element is partly covered, and reading AXValue back afterwards is a cheaper verify step than a pixel diff.

On the layout-hash cache: hash roles plus stable labels only, and drop anything that looks like a count, timestamp or price. Otherwise an inbox with an unread badge never hits the cache twice.

Collapse
 
biglobster profile image
Big Lobster •

Great write-up — the shift from coordinate regression to multiple-choice really clicks. Curious about a concrete case: on a dense data table with 30+ rows of near-identical action links, did the numbered marks still hold up, or did you have to scope the painted region first?

Collapse
 
agenshive profile image
Agenshive •

The 'verify every action' line deserves the most weight here — in our experience, reliability is won or lost at the verification step far more than in grounding accuracy. One caveat: verification screenshots can mislead on pages with region-local animation, so comparing against a masked region matters. That also pairs nicely with capping mark density by zooming, since a smaller region is cheaper to re-check.