DEV Community

Paul Crinigan
Paul Crinigan

Posted on

How AI Browser Agents Actually Drive a Browser, and When a Script Still Wins

For about twenty years, browser automation meant telling the browser exactly what to do. Selenium, then Puppeteer, then Playwright each made that easier, but every one of them still needed a selector for every click and a wait for every load, and every site redesign meant fixing broken scripts. AI browser agents flip that around. You describe the outcome, and a language model works out the clicks. This article walks through how they work under the hood, where they are strong, and where a plain script is still the better tool.

The Loop Behind Every Browser Agent

Every agent runs the same four step cycle: perception, reasoning, action and evaluation.

Perception is how the agent sees the page. DOM based agents read the HTML, strip out what does not matter, and label each interactive element with a number so the model can refer to it. Browser Use works this way. Vision based agents take a screenshot and send it to a vision model instead, which still works when the DOM is obfuscated or the interface is drawn on a canvas. Skyvern leads with this approach. The strongest setups mix both, reading the DOM when it is clean and falling back to a screenshot when it is not.

Reasoning is the model call. The agent sends the page state, the original goal and the history of steps so far, and the model picks the next move. This is where the flexibility comes from. If a Submit button changes from a button tag to a link, a script fails, while the model just sees a submit control and clicks it. It can also deal with cookie banners and login prompts that were never part of the plan.

Action turns that decision into a real command, almost always through Playwright or Puppeteer underneath. Evaluation captures the new page state and checks whether the step worked, so the agent can retry or try another path when it did not.

What the Numbers Say

The best open source agents now complete around 85 to 89 percent of tasks on the WebVoyager benchmark, with Browser Use at 89.1 percent and Skyvern at 85.85 percent. That is impressive, and it also means roughly one task in eight fails. Real production workflows tend to score lower than benchmarks.

Speed and cost follow from the loop. Every step needs at least one model call, so an action a script finishes in 2 seconds can take an agent 30 seconds or more. A complex task can need 20 to 50 calls at a few cents each, while a script costs close to nothing per run once it is written. Vision agents also use more tokens than DOM agents, because screenshots are expensive to send.

Choosing Between an Agent and a Script

The decision usually comes down to who controls the website. If it is your own application, or a stable site you hit at high volume, a deterministic script is faster, cheaper and repeatable. If the sites belong to someone else, change often, or number in the dozens, an agent saves you from writing and maintaining a scraper for each one.

Most teams end up in the middle. Stagehand is built around this idea: write normal Playwright style code for the steps you can predict, and call the model only for the parts of the page that are messy. LaVague takes another route and has the model generate Selenium or Playwright code, which leaves you with a script you can read and audit.

A few limits are worth planning for either way. Bot detection looks at fingerprints, mouse movement and typing rhythm. An agent that logs in for you is handling your cookies and passwords, which matters more when page content goes to an external model. Very dense pages can overflow the model's context window, so the agent has to read them in pieces. The complete guide to AI browser agents goes deeper on each framework, real world use cases from price monitoring to compliance checks, and how to match a tool to your requirements.

The Takeaway

Treat agents and scripts as two tools, not a replacement for one another. Use an agent where the page is out of your control and adapting is the whole job, and keep scripts for the stable, high volume work where speed and repeatability matter most.

Top comments (0)