Self-healing test automation repairs broken element locators while a test is running. When a selector stops matching, the tool looks for the element that resembles the original the most and replaces the old reference so the test can continue. That makes brittle tests cheaper to maintain, but it doesn’t make them any less brittle. And when it comes to critical user flows, cheaper isn’t always safer.
This downside used to be something vendors could sell you as a separate tool. As of Playwright 1.56, a healer ships inside the runner most of this audience already uses alongside a planner and a generator, at no additional cost. So the question is no longer whether to buy a self-healing tool. It’s whether you’re comfortable enough letting the agent already sitting in your repo rewrite your assertions.
Here are five things people get wrong about that bargain, built on QA.tech's very own breakdown of the misconceptions.
How Self-Healing Tests Work
Self-healing is a feature in script-based testing tools that automatically fixes broken locators. The classic approach works by recording a fingerprint of each element at authoring time. That fingerprint can include its id, visible text, ARIA role, position in the DOM, nearby siblings.
When the primary selector fails at runtime, an algorithm compares the current elements against that fingerprint, picks the best match above a confidence threshold, and carries on. That’s the mechanism inside commercial platforms and some open-source solutions like Healenium for Selenium and Appium (Healenium catches the NoSuchElement exception and runs the longest-common-subsequence match against the last known good path).
The newer variant swaps that scorer for a model. Playwright's healer replays the failing steps, inspects the current UI, and suggests a patch (a locator update, a wait adjustment, a data fix) before re-running the test until it passes or a guardrail stops the loop.
It’s the same job, only the mechanism is smarter.
1. "Healing" Is a Re-Scored Guess
When a selector changes, the value your test was pointing at comes back empty. Instead of failing, the self-healing tool finds an element with the most similar text, role, and neighbors, and rewrites the reference.
But the test is still a recorded sequence of implementation details: this selector, this path, this position. Implementation changes regularly, sometimes every sprint, sometimes every day. So the script keeps breaking and the tool keeps patching it.
Now think about how you'd build an agent to do the same job. You wouldn't hardcode a button's coordinates. You'd hand it a goal (like "reset a password from the login screen") and let it read the page and figure out what to do on each run.
2. Self-Healing Moves the Maintenance Out of Sight
Self-healing promises to eliminate your maintenance burden, but what it actually does is relocate that burden somewhere else.
Locator scoring is probabilistic. Sometimes it grabs the right element. Sometimes it picks the wrong one, and your test ends up exercising a different flow without you knowing. On occasion, it may even "stabilize" a flaky test by papering over a real timing bug. So, teams develop a new ritual: reviewing the heals, reading the diffs, approving locator changes in batches.
And there you have it. You’ve simply traded selector maintenance for heal maintenance. The difference is that the new kind fails silently, and those kinds of failures can be much more expensive to catch.
3. Healing a Test That Should Have Failed Defeats the Purpose
A test has one job: to go red when something is wrong. A feature designed to keep tests green after the application changes works against that.
If you've ever watched an AI coding tool "fix" a failing unit test by rewriting the test instead of the code, you've seen this failure mode already. The check adapts to the bug rather than catching it. Self-healing can do the same thing to your end-to-end suite. The app breaks a user journey, the tool heals the test to match the new (broken) behavior, and the check that was supposed to catch the regression reports success.
This isn’t just some theoretical risk. You can see it in the healer that now ships in Playwright. Its loop re-runs a failing test until it passes or a guardrail stops it. Here are a few ways it can get to green:
It can adjust a wait to get past a race condition, masking the underlying issue.
It can mark a test as skipped when it decides the underlying feature is broken, quietly pulling that test out of your signal.
Because it’s trying to get the test to pass, it can change what the test actually checks. Sometimes, the easiest way to make a failing assertion pass is to assert less; for instance, trading a text check for a visibility check. The message copy may have changed, or the healer may have stopped matching it. Either way, the test goes green and you have no way of knowing whether it still verifies what you need it to.
There is a deeper problem underneath. Self-healing focuses on tests that break, but the more costly failure is the tests that are missing. For example, a change ships, and no test is there to catch a problem. A healed test that goes green tells you nothing about whether the change was intended, and if no test covered that part of the application, then there’s nothing to heal.
The answer isn’t to heal everything or ban healing altogether. Some journeys can tolerate changes. For example, if a user can still finish signup, the signup test should pass whatever the layout looks like this week. Others, like a checkout total, a consent screen, or a permissions boundary, must never drift, though.
A serious testing system lets you make that distinction for each test. In QA.tech, that’s the difference between a plain-English goal and pinned, explicit steps within the same test. A healer that optimizes for green makes that decision for you, and green is always the goal.
4. Self-Healing and Agentic Testing Are Different Architectures
Three very different things fall under the "AI testing" label right now:
| Category | How it handles a UI change | What it's actually made of | Best for |
|---|---|---|---|
| Scripted automation (Selenium, Cypress, Playwright) | Breaks, waits for a human | Recorded steps, hardcoded locators | Deterministic flows you own end to end |
| Scripted + repair layer (most "self-healing" tools) | Re-scores the locator and swaps it | The same scripts + a scorer or an LLM | Aging suites you can't rewrite yet; cosmetic churn |
| Agentic testing | Re-reads the screen, re-plans toward the goal | A perception-action-verify loop with memory | Fast-moving products where journeys matter more than markup |
That middle row deserves a closer look. Now that a healer ships in the Playwright runner, "scripted + repair layer" is no longer something you have to buy separately. It is the default state of any Playwright suite on a current version. Most of you who are reading this are already sitting in that row, without making a conscious choice.
On a change, the split looks like this:
Something worth noting: most tools sold as "agentic" are actually just that middle row with agentic branding on top. A simple way to tell the difference is by asking what happens when the test reaches a state that wasn’t part of the original script.
A repair layer reaches for another locator, while an agent reasons about the journey again and figures out what to do next. If every answer circles back to finding a replacement selector, you're looking at a patch. (For a more detailed taxonomy, QA.tech's testing pyramid for the agentic age and its Selenium-to-agentic migration guide are both useful digests.)
5. If a Tool Leads with Self-Healing, Ask Why the Tests Break
Why do the tests break in the first place? For script-based tools, the honest answer is that they test clicks rather than intent. Nobody opens your product hoping to fire button#checkout-submit. They want to buy the thing and move on.
Bring sharper questions to the demo instead. QA.tech has a full buyer's guide for agentic testing tools, but you should start with these four:
- Does it verify the outcome, or just follow the click path?
- Can I pin the journeys that must never change and let the rest flex?
- When something fails, do I see expected-versus-actual, or just a missing selector?
- Does it get smarter about my product across runs, or does every run start cold?
The Only Honest Job for Self-Healing
Self-healing has one legitimate use. For instance, if you're sitting on a large Selenium suite that works and you can't afford to rewrite it this quarter, a repair layer makes perfect sense. It’s faster than fixing selectors by hand and doesn’t pretend to be more than it is.
| Self-healing is fine for... | Falling apart when... |
|---|---|
| A suite you have already decided to retire | It becomes the reason you never retire it |
| Cosmetic churn (renamed classes, moved nodes) | The UI change is the bug, and the healer agreed with it |
| Buying time during a migration | You point it at regulated or high-stakes flows and let it choose green |
The problem was never the healing itself. It's mistaking a patch for an architecture.
What Replaces the Repair Layer
If the pain points in this article sound familiar, that’s the exact problem QA.tech has set out to address. Tests are written as goals in plain language (say, reset a password and confirm the user can sign in again), and an agent runs them through the interface the way a person would, by looking, deciding, acting, and checking the result.
There are no locators to heal, because the test is a goal rather than a path. But that doesn’t mean no more maintenance, as agents still need direction and correction. What changes is the kind of maintenance involved.
You also get dynamic verification focused on the parts of the application a change can actually affect. Tests run on the pull request before a human reviewer reads it, with evidence attached to each result.
If you want to try it on your own app, the demo takes about half an hour.
FAQ
What’s the difference between self-healing tests and agentic testing?
Self-healing repairs broken locators inside a scripted test, while an agentic tester works from a goal and finds a fresh path each run, so it has no locators to heal.
Do self-healing tests fix flaky tests?
Not really. They often mask flakiness by finding a way to pass, which hides the timing bug or race condition underneath.
Can self-healing tests hide real bugs?
Yes. If the tool heals a test to match a broken journey, the suite stays green even though the product is broken.
Which tools offer self-healing?
Locator repair is now a standard feature across the established script-based platforms. It also ships in open-source Healenium and Playwright's Healer agent, which is part of its test-agents toolkit.

Top comments (0)