Qwen3-0.6B (400 MB) on a 2017 Note 8: 10/10 with SiFR, 0/3 without it
TL;DR. A 400 MB model on a 2017 Samsung Note 8 drives a live desktop browser and passes its tasks 10 out of 10. The same model on the same phone, without our layer, never gets through a Wikipedia page. It's the same model every time. The only thing that changes is what it sees.
First, who is writing this: we build this layer (e2llm, the SiFR format). So there is nothing below you can't check yourself. Scripts, raw logs, model versions and an offline replay are in an open repo.
Start with the failure
The task: start on an unrelated site, go to the Wikipedia page for the Samsung Galaxy Note series, pick exactly "Note 8" among similar links (the traps next to it are "Note 8.0", "Galaxy Note 8.0" and "Note FE"), and read the release date from the infobox. The expected answer is 15 September 2017.
Without the layer. The model gets raw HTML. The page is 466,744 characters. A 16k context fits 40,000 of them, about 9%. Result: 0 out of 3. Every time, the model returned the address of the page it was already on. Cost per run: about 12,300 tokens and 35 to 44 minutes.
With the layer. Same model, same phone, same task. 10 out of 10, median 49 seconds per run. All ten times it picked "Note 8", even though node ids changed between runs. So it chooses by text, not by a memorized path.
One model, one phone. The difference is about 40 minutes against 49 seconds, and zero against ten.
Where raw HTML copes
To be fair, here is a second task, the books.toscrape.com sandbox: find a book and return its price, rating and availability. The page is small, and raw HTML passes 3 out of 4. Cost per run: about 12,200 tokens and about 1,360 seconds.
With the layer, the same model on the same Note 8 takes 25 seconds per run and passes 10 out of 10.
That's about 25 times fewer tokens, about 50 times less time, and a stable result instead of a lottery. We don't claim this is impossible without the layer. We claim that with the layer it's cheaper, faster and repeatable.
What the model sees
The model never sees the whole page. It gets a short list of candidates:
div036: message box, "Enter a prompt for Gemini"
and returns one decision:
{"target": "div036"}
Everything else is plain Python: capture, click, paste, recapture, check. The model's only job is to choose. Facts come off the page through the layer, not through the model. For example, a book's rating on books.toscrape lives in a CSS class. It isn't in the page text at all.
Three phones, one kit
| Device | Year | Chip | RAM | Sec per run* |
|---|---|---|---|---|
| Galaxy Note 8 | 2017 | Exynos 8895 | 5.2 GB | 25 |
| Galaxy S21 | 2021 | Exynos 2100 | 7 GB | 29 |
| Galaxy A04e | 2022 | Helio P35 | 2.7 GB | 62 |
* books.toscrape task, Qwen3-0.6B. Full tables for both tasks are in HARDWARE.md.
Every device runs the same llama.cpp commit, the same model files and the same prompt. Qwen3-0.6B is the only model with 10/10 on all three devices, including a phone that costs about $100 and has 2.7 GB of RAM.
Four years of flagship hardware give the small model nothing: 19 seconds of model time on both the Note 8 and the S21. The gap shows up with size. Ministral 3 3B takes 193 seconds on the Note 8 and 47 on the S21, four times as long.
The model is swappable
We ran the same task on the Note 8 with 14 models. Six models from five vendors got 10 out of 10: Qwen3-0.6B, Qwen2.5-1.5B, GLM-Edge-1.5B, Gemma-2-2B, Llama-3.2-3B and Ministral 3 3B. The full table is in HARDWARE.md.
Size doesn't set the floor. Qwen2.5-0.5B gets 6 out of 10, while Qwen3-0.6B, almost the same size, gets 10 out of 10. The models that scored zero (Llama-3.2-1B, Gemma-3-1B, LFM2.5-1.2B) answered with a placeholder "ID" or clicked at random. That's the limit of their training on the format, not a shortage of parameters.
Relay: a conversation across two browsers
The latest series goes past fixed navigation. Ministral 3 3B on the S21 carries a live conversation between Gemini in Chrome and Z.ai in Firefox. Every step works the same way: capture the page, pick an element, act, wait for the new reply, and check that it isn't an echo of what was sent.
Here is what one hop looks like. The layer gives the model candidates from the current Gemini page: the input box, the send button, the "new chat" button. The model returns {"target": "div036"}, the input box. The harness pastes the message there, recaptures the page, picks the send button from the fresh candidates, clicks, and waits for the reply to appear and stop changing. Then it checks that the reply isn't an echo of what was sent and carries it to Firefox, to Z.ai. The same thing happens there, in the other direction.
Result: 10 out of 10 runs, 20 out of 20 rounds, 40 out of 40 hops between browsers. All 40 replies were not echoes, and all 20 chat resets were clean. The services know nothing about each other. The phone just works with their ordinary web pages. The model only makes UI decisions. Reply text is carried over as is, and the model doesn't replace the cloud models.
Run time: minimum 7:59, median 9:29, maximum 11:55. The median of the second half of the series is 11.6% higher than the first half (9:14 to 10:18). Battery over the series went from 79% to 63%, temperature from 32.5 to 35.4 °C. Run 7 had one paste retry: the first paste didn't commit, and there was no click until the retry. You can see it in the logs.
What we don't claim
- This isn't a benchmark. These are measurements on specific tasks, and all the logs are open.
- Results are reported per device. The same model on different hardware can give a different result.
- We don't replace the cloud model. We show how much work is left for a small model once someone has already read the page for it.
Check it yourself
- Repo: github.com/e2llm/edge-browser-agent
- Note 8 series video, one take: youtu.be/-7OC7sge4bA
- Full relay series video: youtu.be/T9fJtp1z3-A
Don't take our word for it. Replay the logs yourself: replay.py reproduces the series from the open JSONL, with no phone, no extension and no browser.
A small model doesn't have to be smart. It has to get the world already translated into its language.
Top comments (0)