DEV Community

Vijay Javvadi
Vijay Javvadi

Posted on

Auditable AI in QA: what Playwright MCP taught me about scaffolding tests that don't drift

A Playwright test fails in staging. The button locator misses. You start the bisect, half an hour in you realise you've never seen the file before — an LLM generated it, someone shipped it, and nobody can tell you which model, which prompt, or which run produced any of the four broken lines.

A test suite is a contract with the future. When you can't say where a line came from, you can't say what it means.

I spent most of this year building @vijaypjavvadi/bdd2pw (currently v4.3.2) as a way to think through that failure mode. It's an MIT-licensed Gherkin .feature → Playwright TypeScript scaffolder built on the Microsoft Playwright MCP server. A paper describing it was published in Elsevier SoftwareX in October 2026 (DOI 10.1016/j.softx.2026.102933).

This post isn't the paper. It's the three design intuitions behind it that I had to learn the hard way, and would have paid real money to be told at the start.

1. Hash the accessibility tree, not the DOM

The Playwright MCP server exposes a browser's accessibility tree to a caller — role, name, id, label, landmark structure. Ordinary Playwright codegen reads the DOM once, picks a locator heuristically, and hands you a spec. MCP lets you keep the tree around and interrogate it.

I used to think the interesting part of MCP was locator picking. It isn't. The interesting part is that the accessibility tree gives you a stable, canonical representation of the page — an artifact you can hash. Two trees that describe the same UI hash to the same sha256, regardless of what CSS or whitespace churned underneath.

That property is what I would build change-detection on top of in any future UI-testing tool. The DOM churns for reasons unrelated to what the page actually is: class names get regenerated, whitespace shifts, hidden comment nodes appear and disappear. Hash the DOM and every drift alarm has a 90% chance of being noise. Hash the accessibility tree and every drift alarm is a real drift alarm, because the accessibility tree is what your tests actually use.

If you take one thing from this post: when you want to answer "has this page changed in a way that matters," ignore the markup and hash the accessibility tree.

2. LLMs will emit code that lies — rewrite before you reject

If you point a model at "generate a Playwright binding for this Gherkin step against this Page Object", it will produce something. That something will look right. It will typecheck about half the time. On the apps I've tested informally, it runs against the actual page roughly 60% of the time.

The failures aren't random. There are repeatable patterns — the model invents a fill(locator, value) helper on the Page Object that doesn't exist, drops a required await, imports from the wrong path, confuses expect(page) with expect(page.locator(...)). Each of these looks like code and compiles against ambient types; at runtime it crashes.

Early on I treated these as flat failures and dropped them on the floor. The step landed as a // TODO and a human had to come back. That was correct — but it was too eager.

Most of the common emission mistakes are mechanically recoverable before you ever ask a human. The model wants loginPage.usernameInput.fill("standard_user"); it emits loginPage.fill("usernameInput", "standard_user"). The intent is unambiguous; the surface form is wrong. You can lift the intent out and rewrite the AST to the correct form, then run the rewritten version through the same gate that would have rejected it originally. If the rewrite survives, you ship it. If it doesn't, you drop it.

Rewrite before you reject. Fail-closed is a floor, not a ceiling. If you're going to spend a model call producing a binding, spend a millisecond trying to repair the output before you throw it out.

3. Attribution turns a scaffold into something you can audit

The hardest bug I had wasn't a wrong locator. It was a locator that was right six months ago, generated by a model I'd since swapped out, and I had no record of which version of the pipeline had produced it. Every "why" question — why did this fail now, which pack owns this rule, does the LLM actually earn its keep on this suite — was unanswerable because the emitted code carried no provenance.

The minimum attribution surface I'd now insist on in any AI-assisted codegen tool, from day one:

  • Every emitted artifact says who generated it. Not "the system" — a specific rule id (rule-14, pack banking) or a specific LLM call (provider anthropic, model claude-sonnet-4-6, cached or not, governance-sanitised or not, UTC timestamp).
  • Every scenario that touched the LLM carries a visible tag. In Playwright-land, @llm-generated next to the existing tags is enough — then your trusted-only lane is --grep-invert @llm-generated with no new framework.
  • The attribution lives in a sidecar file, not a comment that will drift. One <spec>.meta.json per spec, machine-readable, grep-able.

Attribution has almost no complexity cost at write time (one optional field in the emission record). The migration cost of adding it later is enormous — I know because I did it later and it hurt. Build it in on day one of any AI-involved tool.

And one about self-healing

A tangent that saved me a lot of pain: the first attempt at self-healing locators was a 450-line inline

Top comments (0)