DEV Community

Abhishek Jain
Abhishek Jain

Posted on

3 Ways I Use AI to Auto-Generate Playwright Locators

If you've maintained a Playwright suite for more than a few months, you know the real cost isn't writing tests — it's the constant locator rot. A designer tweaks a class name, a component gets refactored, and suddenly a dozen tests are red for reasons that have nothing to do with actual bugs.

After 18 years in test automation — including building enterprise-scale frameworks — I've found that AI-assisted locator generation is one of the highest-leverage places to bring LLMs into a QA workflow. Here are three approaches I actually use, not just demo-ware.

1. AI-Assisted Locator Discovery from the DOM

Instead of hand-picking a data-testid or fighting brittle CSS selectors, I feed a snapshot of the relevant DOM section to an LLM and ask it to propose the most resilient locator strategy — prioritizing accessible roles and text over implementation-specific attributes.

// Instead of guessing at a selector manually:
const button = page.locator('.btn.btn-primary.mt-2.submit-btn-v2');

// Ask an AI-assisted step to suggest something resilient:
const button = page.getByRole('button', { name: 'Submit Order' });
Enter fullscreen mode Exit fullscreen mode

The win isn't that AI "knows" your app — it's that it consistently nudges you toward Playwright's built-in resilient locator patterns (getByRole, getByLabel, getByText) instead of the CSS-selector habits many teams fall back into under deadline pressure.

2. Self-Healing Locators via Fallback Chains

When a primary locator fails during a run, rather than failing the test outright, I use an AI step to analyze the current DOM and suggest a fallback locator that matches the original intent of the interaction — then logs the drift so a human confirms it later.

This isn't "magic self-healing" that silently rewrites your test suite (which I'd actually caution against — silent healing can mask real regressions). It's a flagged suggestion: "Your original locator broke; here's what looks like the same element now; confirm before I update the source."

3. Locator Generation from Natural-Language Test Steps

For teams writing BDD-style scenarios ("User clicks the 'Add to Cart' button"), I use AI to translate that natural language directly into a first-draft Playwright locator + action, which a human then reviews and commits.

When the user clicks "Add to Cart"
Enter fullscreen mode Exit fullscreen mode

becomes a suggested:

await page.getByRole('button', { name: 'Add to Cart' }).click();
Enter fullscreen mode Exit fullscreen mode

This dramatically speeds up onboarding less-experienced testers into a Playwright framework, since they're writing intent, not fighting Playwright's API surface on day one.


The common thread across all three: AI isn't replacing test design judgment — it's removing the tedious, error-prone parts of locator selection so testers can focus on what to test, not how to write a selector.

I go deeper into all of this — plus patterns for AI-assisted test generation, flaky test triage, and building a production-grade framework end to end — in my book, Playwright Test Automation with AI.

Top comments (6)

Collapse
 
jo-do profile image
Jo Do •

Locator rot is the correct thing to optimize because it is maintenance cost, not authoring cost - a dozen tests red for reasons unrelated to actual bugs is the exact moment teams stop trusting the suite, and a suite nobody trusts is worse than none. The leverage of AI here is that the regeneration loop is mechanical: the intent of the test did not change, only the selector did. The part I would watch is the failure mode where the model "fixes" a red test by picking a locator that targets a different element - the suite goes green and the coverage silently changed. Reviewing regenerated locators as a diff against intent, not just as green CI, is what keeps this honest.

Collapse
 
abhishek_jain_cd78ebe598d profile image
Abhishek Jain •

Exactly — "diff against intent, not green CI" is the better framing. The closest I've gotten in practice: surface the old selector, the new one, and a screenshot of what it resolves to, side by side, before merge. Doesn't automate the judgment call, but turns silent coverage drift into something a reviewer actually catches in 10 seconds instead of 6 months later.

Are you doing this at the PR review stage, or have you found a way to flag it automatically when the selector changes but the assertion doesn't?

Collapse
 
eternaclarity profile image
Jesse Gamble •

The accessible-role-first part is what makes this practical. AI can help propose selectors, but keeping the fallback chain biased toward user-visible semantics should reduce the chance of a test "healing" into a selector that still passes while pointing at the wrong element.

Collapse
 
abhishek_jain_cd78ebe598d profile image
Abhishek Jain •

Really good point — and honestly, even role-based locators aren't fully safe from this. Two buttons both named "Submit" (common in modal-over-form UIs) can still let getByRole('button', { name: 'Submit' }) resolve to the wrong one after a change — just with a more "trustworthy-looking" failure than a CSS selector would.

One mitigation I use: log the resolved element's bounding box or nearby text alongside pass/fail, so a silently-wrong match still leaves a visible trail in review, even if the test itself stays green.

Curious if you've hit the multi-match version of this in your own suites.

Collapse
 
eternaclarity profile image
Jesse Gamble •

More than once, usually duplicate labels in modals. I'd rather a test fail loudly than pass quietly.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.