DEV Community

Cover image for Building a Selector-Free Autonomous Desktop RPA Engine with Python & Vision LLMs
Ruesch Manny
Ruesch Manny

Posted on Originally published at mannyverse767.gumroad.com

Building a Selector-Free Autonomous Desktop RPA Engine with Python & Vision LLMs

Building a Selector-Free Autonomous Desktop RPA Engine with Python & Vision LLMs

Legacy enterprise software—on-premise SAP GUI installations, AS/400 terminal emulators, desktop-only accounting suites, and specialized medical record software—rarely provides accessible REST APIs. For the past two decades, the standard remedy has been traditional RPA (Robotic Process Automation) platforms like UiPath, Automation 360, or Power Automate Desktop.

In practice, traditional RPA is notoriously brittle. It relies on accessibility element trees, Win32 handle introspection, or static DOM selectors. When an update alters an internal UI hierarchy, a display scale shifts from 100% to 125%, or dynamic IDs change, automation jobs break.

This guide demonstrates how to architect a selector-free, OS-agnostic RPA engine in Python that leverages Multimodal Vision LLMs (such as Anthropic Claude 3.5 Sonnet / Computer Use) to perceive screen buffers and drive low-level OS input events natively across Windows, macOS, and Linux.


1. The Bottleneck: Why Traditional RPA Breaks

Traditional RPA frameworks rely on accessibility trees (Microsoft UI Automation, Apple Accessibility API, AT-SPI). This architecture incurs three major vulnerabilities:

  1. Fragile Tree Traversal: Virtualized grids (like SAP or legacy WPF apps) render contents via hardware-accelerated canvases without exposing child accessibility nodes.
  2. Resolution & DPI Vulnerability: Pixel-offset OCR fallbacks shatter when remote desktop windows resize, rendering absolute coordinates invalid.
  3. Licensing Costs: Commercial enterprise runners charge $5,000 to $15,000 per unattended bot license annually, locking infrastructure into closed ecosystems.

The Vision-Grounding Alternative

Instead of querying internal OS handle structures, a Vision-Driven RPA engine operates identically to a human engineer:

  1. Capture the raw desktop framebuffer.
  2. Pass the downscaled/encoded buffer into a multimodal reasoning model alongside an action objective.
  3. Extract normalized visual coordinate predictions (x_pct, y_pct) alongside an action token (click, double_click, type, key_press, scroll).
  4. Map those coordinates to native hardware events.
  5. Verify the state delta across consecutive frames with a self-healing loop.
+--------------------+      +-------------------------+      +-------------------------+
| Desktop Framebuffer| ---> | Vision Coordinate Model | ---> | Action Dispatcher       |
| (X11 / Quartz / OS)|      | (Claude / Qwen-VL)      |      | (pyautogui / pynput)    |
+--------------------+      +-------------------------+      +-------------------------+
          ^                                                               |
          |                 +-------------------------+                   |
          +---------------- | Visual Verification Loop|<------------------+
                            | (Delta / SSIM Check)    |
                            +-------------------------+
Enter fullscreen mode Exit fullscreen mode

2. System Architecture

The engine consists of four core decoupling boundaries:

  • Display Capture Pipeline: Intercepts screens via fast C-bindings (mss) and handles multi-monitor resolution normalization.
  • Coordinate Normalizer: Converts screen coordinates to a standardized 1000x1000 grid to avoid prompt overhead, then re-scales back to client coordinates.
  • Action Execution Engine: Dispatches hardware inputs using pyautogui or virtual input device drivers (/dev/uinput on Linux).
  • Audit & Self-Healing Loop: Records state deltas using Structural Similarity Index (SSIM) and saves structured JSONL execution logs.

3. Implementation: Code & Logic

Step 1: Screen Capture and Coordinate Rescaling

We capture the screen buffer with mss and rescale coordinates between raw monitor resolution and normalized model space (0–1000).

import mss
from PIL import Image
import io

class ScreenCapture:
    def __init__(self, monitor_index: int = 1):
        self.sct = mss.mss()
        self.monitor = self.sct.monitors[monitor_index]
        self.width = self.monitor["width"]
        self.height = self.monitor["height"]

    def capture_frame(self, target_width: int = 1280) -> tuple[bytes, float]:
        sct_img = self.sct.grab(self.monitor)
        raw_image = Image.frombytes("RGB", sct_img.size, sct_img.bgra, "raw", "BGRX")

        # Maintain aspect ratio for LLM visual processing
        aspect_ratio = self.height / self.width
        target_height = int(target_width * aspect_ratio)
        resized_image = raw_image.resize((target_width, target_height), Image.Resampling.LANCZOS)

        buffer = io.BytesIO()
        resized_image.save(buffer, format="PNG", optimize=True)

        # Scale factor relative to native resolution
        scale_factor = target_width / self.width
        return buffer.getvalue(), scale_factor

    def denormalize_coordinates(self, norm_x: int, norm_y: int, grid_size: int = 1000) -> tuple[int, int]:
        """Translates normalized grid coordinates (0-1000) to actual monitor pixels."""
        real_x = self.monitor["left"] + int((norm_x / grid_size) * self.width)
        real_y = self.monitor["top"] + int((norm_y / grid_size) * self.height)
        return real_x, real_y
Enter fullscreen mode Exit fullscreen mode

Step 2: LLM Prompt Structuring & Tool Schema

We enforce structured JSON tool responses so the LLM outputs unambiguous actions and coordinates.

TOOL_DEFINITION = {
    "name": "desktop_action",
    "description": "Execute an input action on the desktop target UI.",
    "input_schema": {
        "type": "object",
        "properties": {
            "action": {
                "type": "string",
                "enum": ["click", "double_click", "right_click", "type", "press_key", "hotkey", "wait", "terminate"]
            },
            "coordinates": {
                "type": "array",
                "items": {"type": "integer"},
                "description": "[x, y] coordinates mapped to a 0-1000 grid. Required for click events."
            },
            "text": {
                "type": "string",
                "description": "Text string to type into input fields."
            },
            "key": {
                "type": "string",
                "description": "Key or key combination to trigger (e.g., 'enter', 'tab', 'ctrl+s')."
            },
            "reasoning": {
                "type": "string",
                "description": "Step-by-step rationale for why this UI element was selected."
            }
        },
        "required": ["action", "reasoning"]
    }
}
Enter fullscreen mode Exit fullscreen mode

Step 3: Hardware Action Execution & State Verification

The dispatcher executes hardware events and captures state confirmation immediately afterward.

import pyautogui
import time
import json
from datetime import datetime

pyautogui.FAILSAFE = True
pyautogui.PAUSE = 0.05

class ActionDispatcher:
    def __init__(self, capture: ScreenCapture, log_path: str = "audit_log.jsonl"):
        self.capture = capture
        self.log_path = log_path

    def dispatch(self, action_payload: dict) -> bool:
        action = action_payload.get("action")
        coords = action_payload.get("coordinates")
        text = action_payload.get("text")
        key = action_payload.get("key")

        real_x, real_y = None, None
        if coords and len(coords) == 2:
            real_x, real_y = self.capture.denormalize_coordinates(coords[0], coords[1])

        # Execute action
        if action in ["click", "double_click", "right_click"] and real_x and real_y:
            pyautogui.moveTo(real_x, real_y, duration=0.2)
            if action == "click":
                pyautogui.click()
            elif action == "double_click":
                pyautogui.doubleClick()
            elif action == "right_click":
                pyautogui.rightClick()

        elif action == "type" and text:
            pyautogui.write(text, interval=0.02)

        elif action == "press_key" and key:
            pyautogui.press(key)

        elif action == "hotkey" and key:
            keys = [k.strip() for k in key.split("+")]
            pyautogui.hotkey(*keys)

        elif action == "wait":
            time.sleep(2.0)

        elif action == "terminate":
            return False

        self._log_audit(action_payload, (real_x, real_y))
        return True

    def _log_audit(self, payload: dict, absolute_coords: tuple):
        log_entry = {
            "timestamp": datetime.utcnow().isoformat(),
            "payload": payload,
            "absolute_coords": absolute_coords
        }
        with open(self.log_path, "a") as f:
            f.write(json.dumps(log_entry) + "\n")
Enter fullscreen mode Exit fullscreen mode

Step 4: The Closed-Loop Execution Loop

This loop orchestrates the screenshot capture, calls the Anthropic API, executes the action, and checks whether the visual state changed.

import base64
from anthropic import Anthropic

def run_autonomous_task(task_instruction: str, max_steps: int = 25):
    client = Anthropic()
    capture = ScreenCapture(monitor_index=1)
    dispatcher = ActionDispatcher(capture=capture)

    messages = []

    for step in range(max_steps):
        print(f"[*] Executing step {step + 1}/{max_steps}")
        img_bytes, _ = capture.capture_frame(target_width=1280)
        base64_img = base64.b64encode(img_bytes).decode("utf-8")

        step_prompt = (
            f"Current task: {task_instruction}\n"
            f"Analyze the provided desktop screenshot. Identify the target element "
            f"and invoke the desktop_action tool. If the task is finished, invoke terminate."
        )

        messages.append({
            "role": "user",
            "content": [
                {"type": "text", "text": step_prompt},
                {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": base64_img}}
            ]
        })

        response = client.messages.create(
            model="claude-3-5-sonnet-20241022",
            max_tokens=1024,
            tools=[TOOL_DEFINITION],
            tool_choice={"type": "tool", "name": "desktop_action"},
            messages=messages
        )

        tool_call = next(c for c in response.content if c.type == "tool_use")
        action_data = tool_call.input

        # Append assistant turn to prevent context drift
        messages.append({"role": "assistant", "content": response.content})

        continue_running = dispatcher.dispatch(action_data)
        if not continue_running or action_data.get("action") == "terminate":
            print("[+] Task reported complete.")
            break

        time.sleep(1.0) # Allow application state to settle
Enter fullscreen mode Exit fullscreen mode

4. Deployment, Rate Limits, and Error Handling

1. Frame Buffer Throttling

Sending entire screen frames per token step can exhaust LLM rate limits (especially TPM limits on images). To handle this:

  • SSIM Gating: Compute the Structural Similarity Index between consecutive frames. If visual delta is < 1%, do not re-send an image; tell the LLM that the UI hasn't visually updated yet.
  • Sub-region Cropping: If an active window handle is known, crop the capture strictly to the target bounding box rather than the whole 4K desktop.

2. Guardrails & Failsafes

Always ensure native emergency interrupts remain functional:

  • Keep pyautogui.FAILSAFE = True enabled. Slamming the mouse into any screen corner immediately raises pyautogui.FailSafeException and aborts execution.
  • Run tasks within isolated sandboxes, dedicated virtual displays (via Xvfb on Linux), or isolated VMs when automating irreversible accounting or ERP entries.

5. Conclusion & Production Blueprint

Selector-free RPA turns brittle, element-bound automations into reliable computer-use pipelines. By treating UI like human operators do—through visual feedback—you can automate legacy software without maintenance overhead or costly commercial runner licenses.

You can implement this architecture using the snippets above, or deploy our complete, production-hardened template:

The production repository includes:

  • Full cross-platform window management (X11, macOS Quartz, Win32).
  • Local SSIM frame-caching to cut LLM token costs by over 60%.
  • Pre-built error recovery routines with step-level visual debugging.
  • Complete JSONL schema validators and visual heatmap audit tools.

Top comments (0)