I'm a computer engineering student in Manaus, Brazil, and for my research project I've been building ANCHOR, a web automation agent: you describe a task in plain language ("Register Maria Silva, IT department, contractor, accept the terms and save"), and it runs the task in the browser. Everything runs locally, with small models (Qwen 3.5, 4B and 9B, through Ollama) on a GTX 1050 Ti with 4 GB.
This post is about the one design decision that made the biggest difference, and a few things I learned by measuring everything.
The problem
Classic RPA depends on fixed selectors written by a developer:
page.locator('[data-test="add-to-cart"]').click()
It works until the page changes, and every automation needs someone who codes, far from the person who actually knows the task.
The trendy answer is to put an LLM in control: it looks at the page and decides where to click. With small models, that has a serious problem, which I measured (below): the model says it's done when it isn't.
The decision: the model only plans
In ANCHOR, the model never clicks anything. It gets the request and a summary of the page (the elements, numbered), and returns a plan: goals and steps, using the elements' names ("fill Full name", "check Contractor", "click Save registration").
A deterministic engine, with no AI, carries it out:
- A heuristic finds each element from the step's description, combining text, label, role and context (synonyms, nearby text), and refuses when it isn't sure.
- When it refuses, the user decides: the candidates are numbered on the page, and you pick one (or click the right element).
- The choice is remembered, so the next run doesn't ask.
- Each action's effect is checked (did the page change? did the field keep the value? did an error appear?), and the end of the run is checked against the request.
In the GIF, the site changed after the automation was saved ("Add to cart" became "Add to bag"). Instead of guessing, ANCHOR asks which button to use, remembers the answer, and the next run goes on its own.
Learn once, replay without the model
A task that worked can be saved. The first run plans with the model (about 70 seconds on my GPU); the next ones replay the approved plan without calling the model, in about 3 seconds, even once per row of a spreadsheet. When the site changes and a saved step breaks, the automation recovers and records what it changed, and that record can be reviewed and undone.
What the measurements showed
I built a resilience benchmark: 30 tasks on pages altered in 5 levels (renamed IDs, synonyms, restructured layouts, cookie banners and hidden sections, lookalike decoys with injected instructions), 150 runs per executor and model, with no user to help:
| Executor | Tasks done | False successes |
|---|---|---|
| Fixed-selector script (classic RPA) | 53% | 0 |
| LLM in control (4B / 9B) | 68% / 72% | 33 / 41 |
| ANCHOR (4B / 9B) | 82% / 84% | 2 / 1 |
The column that surprised me most was false successes: the LLM-in-control agent claimed it had done the task, without doing it, in about one run in four. For whoever relies on the automation, that's worse than failing, because nobody finds out.
ANCHOR's weak spot is pages with part of the content hidden behind a "Show more" button (47%, vs 57–60% for the LLM in control). And an honest caveat: these are development numbers, where I tuned the system looking at the tasks. For 1.0, there's a closed set of 80 requests written by colleagues who never saw the project, which will run only once, at the end.
Three lessons
1. Re-measure everything after every change. The first full benchmark run revealed four bugs no test had caught (for example, the barrier against destructive actions let a "Delete all" through when the request was "delete Bruno Lima"). And two of my first fixes created new problems, which only showed up because I re-ran the other evaluations too.
2. A prompt change affects tasks that have nothing to do with it. When I added table extraction, I put two lines in the prompt explaining the new actions. The 4B model started stopping a search after filling the box, and the 9B one left the contract type out of a registration. The fix was to show those lines only when the request is about reading data; for every other request, the prompt is identical, byte for byte, to the previous version.
3. Never put your password in a request to an AI. Testing on LinkedIn, I wrote my e-mail and password in the request, and it typed them. The password never left my computer, but it went through the model, which shouldn't happen. That became version 0.3.1: ANCHOR never types a password, refuses a request with a password before it reaches the model, and when a task needs a login, it stops and waits for you to log in by hand (the session is kept for the next runs).
Try it
The project is on GitHub, with the setup steps in the README:
https://github.com/LucasDantas2701/ANCHOR
Testers and opinions are very welcome, especially failure cases on real sites: that's the most useful feedback. The next version (0.4) adds file downloads and uploads, new tabs and iframes.
Disclosure: the code was written with help from an AI assistant (Claude), which also helped me write this post; the project, the decisions and the evaluation are mine. The code is source-available (PolyForm Strict): you can read and use it, but not modify or redistribute it.


Top comments (2)
tr.ee/dev-to
Happy to answer questions about the design or the benchmark. I’m especially curious: would you trust an automation that recovers on its own if every fix is recorded and can be undone?
Some comments have been hidden by the post's author - find out more