DEV Community

Cover image for Let the AI click. Don't let it grade.
Ian
Ian

Posted on

Let the AI click. Don't let it grade.

I broke the Add button on a tiny todo page and ran my agent test against it, with a scripted stand-in where the model would go.

The agent step came back green. The test still failed, one line later, on a plain locator that expected "1 open" and read "0 open".

That one-line gap is the whole design question for AI end-to-end tests: which steps get a model, and which stay code.

Goals in English, checks in code

e2e, from TesterArmy, was #2 on GitHub trending on Oct 5, with 1,398 stars in a day. It is an Apache-2.0 TypeScript test runner for the web (Playwright underneath) and for iOS and Android. A test mixes three kinds of step. agent.act hands a goal in plain English to a model that drives the app. agent.assert asks a model to judge the screen. screen and expect are ordinary locators that never call a model.

Here is the test I wrote against my page:

test('agent step: add a todo', async ({ app, agent, screen }) => {
  await app.open('/');
  await agent.act('add a todo named {title}', { params: { title: 'Buy milk' } });
  await expect(screen.getByRole('status')).toHaveText('1 open');
});
Enter fullscreen mode Exit fullscreen mode

Nothing new so far. Plenty of AI testing tools can do the first run. The interesting part is the second.

The second run is the product

When the company launched its hosted product on HN in June, the sharpest question came from poisonborz: hand-written E2E tests are "deterministic AND cheap to run", so how does an agent compete? The open-source framework answers with a replay cache.

An agent.act that a later check verified gets recorded: the actions, the controls by role and name, and what the screen should look like afterward. On the next run, the runner replays that recording with no model call. If a control is gone or the end state does not match, the model picks up from wherever the app is.

I wanted to watch this without a model. I had no API key for this post, so I used the custom executor hook from the docs and wrote a 17-line scripted stand-in. It finds the textbox and the button in the redacted screen text, types and taps through the runner's checked actions, and logs every call to a file.

  • Run 1: Cache 1 missed, executor called once.
  • Runs 2 and 3: Cache 1 replayed, executor never called.

The recording is a 1.4 KB JSON file. It holds two actions (type "Buy milk" into textbox "New todo", then tap button "Add"), the anchors that appeared (status "1 open", listitem "Buy milk"), and the one that went away (status "0 open"). With a real model, that act costs calls once, then nothing until the UI moves.

What the cache cannot replay

Here is the catch, and it decides where the AI belongs. agent.assert, agent.waitFor and agent.extract always run live. A judgment is never cached.

Some rough math, using the sample cost table in the project's debugging docs: one act at $0.0198 over three model calls, one assert at $0.0114. Those are their example figures, not my measurements. A suite with 40 acts and 40 asserts comes to about $1.25 on a cold run. Replay every act perfectly and it still costs $0.46, on every run, forever. All of that is grading. Swap the asserts for locator checks and a warm run makes no model calls at all.

So my claim is this. Let the model do the clicking, because clicking is the part that gets recorded, and the part that breaks with every redesign. Keep the grading in code wherever an exact value exists. Save agent.assert for outcomes no locator can state, like whether a chatbot's reply answers the question, and pay for it on purpose.

My broken button shows the second reason. After the replay hit an end-mismatch, my stand-in took over, tapped the dead button again, and reported success, the way any agent grades its own work. The expect caught it. In issue #884, a user hit the real version with a decision-model executor: it claimed a tab step was done without clicking in 4 of 5 runs, and "only our toHaveURL assertion caught it."

Three ways a run goes red

Determinism is not one thing, and e2e splits failures by exit code. I hit three of them.

  • The product broke. My dead button: exit 1, ASSERTION_FAILED, expected "1 open", observed "0 open", with a text dump of the screen.
  • The model was unreachable. My first run used the default config with no key. The locator test passed. The agent test failed with MODEL_PROVIDER_FAILED, filed under category infrastructure, exit 3. A dead key does not pose as a bug in your app, which is exactly what CI needs.
  • The recording went stale. I renamed "Add" to "Create". Under --strict-cache, the run failed with REPLAY_STALE and exit 2, and never called the executor.

Without the flag, that same rename passed. The replay typed the text, could not find the old button, and handed off (Cache 1 handed off). The step took 14.6 seconds instead of about 0.2; the docs give a missing control 15 seconds to show up. The docs warn about this default themselves: a broken recording slides quietly into a live model run on every build until someone re-records it. In CI, I would turn strict on from day one.

The defaults that cost money

Two settings matter more than which model you pick.

First, e2e init adds .e2e/cache/ to .gitignore, and the cache defaults to read-only in CI. Put together, CI starts with no recordings and cannot save any, so every act runs live on every pull request until you commit the cache folder. The docs explain how. It is opt-in.

Second, CI retries a failed test once by default, and retries never replay. That is the honest call, since a replayed retry would just repeat what failed. It also means a flaky test pays for its second attempt in model calls, not just minutes.

Then there is issue #850, still open. After a hand-off, the wait a recording stores can grow on each run until it hits 120 seconds. The reporter found 8 of 26 entries pinned there, and a two-line patch took their cached suite from 6 min 51 s to 1 min 35 s.

What I ran, and what I didn't

I ran it on Node 24.15 in a throwaway folder: npx e2e init --yes, then npm install (40 packages, 72 MB). I served a static page with Python, and Chromium came from my existing Playwright cache. The locator test passed 20 times in a row with --repeat-each 20, in 6 seconds. The agent test ran once with no key, to see the failure, then with my scripted executor, to watch the cache record, replay, go stale and catch a regression. I also set E2E_TELEMETRY_DISABLED=1; the CLI sends anonymous usage data unless you opt out.

I did not run a real model. So I have no numbers of my own on how often the built-in agent finishes a goal, how flaky a live act is, or what a step really costs. The dollar figures above are the docs' sample. The project is also pre-1.0, and 0.17.0 shipped two breaking changes on Oct 4.

Where I'd put the AI

If I adopt this, the rule fits in four lines:

  1. agent.act to get somewhere: sign in, fill the wizard, reach the screen under test.
  2. An expect right after every act. It is what makes the recording exist, and what catches an agent grading its own homework.
  3. agent.assert only where the right answer is a meaning, not a string.
  4. --strict-cache in CI, with the cache folder committed.

Which check in your suite could only a model make, and which one have you been handing to a model out of habit?

Top comments (0)