DEV Community

Cover image for Agent Harness: Turning Browser Clicking Into Multiple Choice With jev-ultrafast v0.1.0
HIROKI II
HIROKI II

Posted on

Agent Harness: Turning Browser Clicking Into Multiple Choice With jev-ultrafast v0.1.0

You ask an AI to search flights. Every step it takes a screenshot, reads the image, thinks, and clicks. One misread button means starting over. By the time it finishes, you would have found the flight yourself. The method is the problem, and Browser Use's new project attacks exactly that.

jev-ultrafast v0.1.0, open-sourced in September 2026 by the Browser Use team, does not show the model screenshots by default. It turns visible controls into a numbered list and lets a judgment model pick “which operation” and “on which element” in one shot. The official recording completes a Google Flights search (Zurich to London, one-way, September 20) in 7.073 seconds. Facts here are checked against the official repository, the performance page, and TypeSafe documentation (checked 2026-10-02); I have not run the project myself.

Try the core idea in two chat sessions

No installation needed. Grab a screenshot of a page without login state (a search or settings page works), open two chats, and ask both the same question: “Which control should be operated next?”

Chat 1: screenshot only. Ask the model to point at the next control. Then verify: does that control exist on the page? Did it invent a button that isn't there?

Chat 2: numbered list. Describe the same page as a numbered table and ask again:

[1] button    Trip type · Round trip
[2] combobox  From · empty
[3] textbox   Departure date · empty
[4] button    Search
Question: change the trip to one-way. Answer with the number and the operation only.
Enter fullscreen mode Exit fullscreen mode

Side by side, the list version is easier to check, and the model rarely invents options that are not in the list. The lesson: a browser agent lists the page state as checkable options first, then lets a model pick the operation and the target. This compares two ways of representing a page; it says nothing about real agent speed or success rates.

What jev-ultrafast actually does

The README calls it “a browser agent with a dynamic, indexed action space.” Four steps:

  1. Read the page into a table. On every page change, visible common controls become a numbered element table. There are eight operations — CLICK, TYPE_TEXT, SELECT, SCROLL_UP, SCROLL_DOWN, WAIT, DONE, BLOCKED — and each operation only gets compatible targets.
  2. One request, two answers. The table goes to TypeSafe's Jev, a System One judgment model. One network round trip returns the operation and the element, with per-candidate probabilities and confidence.
  3. A small LLM writes text only when needed. Jev makes judgments; it does not write text. When “Zurich” must be typed, a small language model generates it (the demo uses inception/mercury-2.5 over OpenRouter with reasoning disabled). The recording generated “Zurich” in 581 ms.
  4. Validate before every click. The executor re-checks page state, the target element, and occlusion before acting. Model output never becomes selectors, coordinates, shell commands, or executable JavaScript. Offscreen article bodies and footers never enter the model context.

Where the Agent Harness concept fits

Recent official documentation gives this design a name-shaped box: an agent harness — the software layer that connects a model to context and tools, coordinates the agent loop, and maintains session state (VS Code's definition; OpenAI's April 2026 Agents SDK announcement discusses harness, agent loop, and sandboxes together). Read that way, the four steps are a browser harness at work: observe the page, list options, ask the model, verify before acting. One distinction matters: the project's connection tool is literally named Browser Harness — that is a proper noun. Agent Harness here is the general concept, not the project's self-description.

What the numbers prove, and what they don't

  • The 7.073-second recording starts after the initial page observation and includes model calls, generated text, browser work, stale decisions, and loading waits. Browser setup, initial navigation, and post-run verification are outside the clock. It contains 17 Jev requests with a 178 ms median latency.
  • Six alternating runs on the same task: median 9.450 s → 7.092 s (25% lower), median TypeSafe requests 22 → 17, median browser protocol calls 1,092 → 101. The repo's own caveat: three pairs are too few for a strong statistical claim; this is not a general benchmark.
  • Two independent smoke checks: a Wikipedia article in 2.798 s, a local hotel search in 1.896 s. They are not part of the speed comparison.
  • Cost: TypeSafe lists Jev 1.13 at $0.042 per million input tokens, output free. The recording's 90,558 input tokens estimate to ≈$0.0038; add the actual $0.00006272 billed for the two text calls and the two models together are ≈$0.0039. That is an estimate from public pricing — browser costs excluded, not a bill.

What to skip for now

  • Don't rush to install anything or configure two API keys (TypeSafe + OpenRouter). Do the two-chat exercise first.
  • frames, canvas, file uploads, new tabs, and arbitrary keyboard widgets are unsupported in this version — don't build on them yet.
  • Don't read DONE as success. The official example re-verifies results after the run; the agent shares your Chrome profile and login state, so keep logins, orders, and payments human.
  • Don't promise 7 seconds for everything — it is one recording of one flight scenario.

Sources: jev-ultrafast, performance.md, TypeSafe models, VS Code: Understand agent harnesses.

Top comments (0)