<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Luthfi Ferdian</title>
    <description>The latest articles on DEV Community by Luthfi Ferdian (@ferdianluthfi).</description>
    <link>https://dev.to/ferdianluthfi</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3850144%2Fe84b7f52-9a75-4a34-9e7f-d0f2f2049f13.png</url>
      <title>DEV Community: Luthfi Ferdian</title>
      <link>https://dev.to/ferdianluthfi</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ferdianluthfi"/>
    <language>en</language>
    <item>
      <title>Stop Healing Your Tests: Why Throwaway Automation Fits the AI Era</title>
      <dc:creator>Luthfi Ferdian</dc:creator>
      <pubDate>Wed, 30 Sep 2026 02:11:48 +0000</pubDate>
      <link>https://dev.to/ferdianluthfi/stop-healing-your-tests-why-throwaway-automation-fits-the-ai-era-45g5</link>
      <guid>https://dev.to/ferdianluthfi/stop-healing-your-tests-why-throwaway-automation-fits-the-ai-era-45g5</guid>
      <description>&lt;p&gt;Self-healing tests are solving the wrong problem.&lt;/p&gt;

&lt;p&gt;When a "healer" swaps a broken locator for a "close enough" element, the run stays green, but the test might now be asserting something nobody intended. You haven't fixed the test. You've traded a loud failure for a quiet false positive.&lt;/p&gt;

&lt;p&gt;Here's the alternative I keep coming back to: &lt;strong&gt;don't maintain a script at all.&lt;/strong&gt; If a model can read a test case and execute it in a browser by itself, the script stops being the thing you need. The test case is. You can't break a script that doesn't exist.&lt;/p&gt;

&lt;h2&gt;
  
  
  Self-healing fixes the symptom
&lt;/h2&gt;

&lt;p&gt;A self-healing locator answers "the button moved, how do I still click it?" It never asks "should this script still exist?"&lt;/p&gt;

&lt;p&gt;Worse, a healer can swap a broken selector for a nearby element that looks plausible. The test passes, but it now checks something the author never meant. You've traded a loud failure for a quiet false positive, and those are the most expensive kind of test debt. A locator that keeps breaking is also telling you something (unstable hooks, a churning screen), and auto-repair mutes that message.&lt;/p&gt;

&lt;h2&gt;
  
  
  The throwaway loop: test cases in, execution out
&lt;/h2&gt;

&lt;p&gt;Here's the alternative I keep coming back to. Your source of truth is the &lt;strong&gt;test case in plain language&lt;/strong&gt;, the same one you'd hand to a manual tester. Instead of turning it into a Playwright script that someone has to own, you let an AI agent read it and run it.&lt;/p&gt;

&lt;p&gt;With a browser-control tool such as the Playwright MCP server, the agent can open a headless browser, follow the steps, look at the page, and report what happened. Because it works from the page as it is right now, there is nothing to heal. If the button moved, it finds the button, the way a human would.&lt;/p&gt;

&lt;p&gt;A test case in this world is just this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## TC-CART-014: Promo banner shows for a new promo SKU&lt;/span&gt;
Precondition: logged in as a standard test user, cart is empty
&lt;span class="p"&gt;1.&lt;/span&gt; Add the item with SKU PROMO-001 to the cart
&lt;span class="p"&gt;2.&lt;/span&gt; Open the cart page
Expected: a promo banner is visible above the item list and mentions the discount
Also check: no error toast, page has no layout overlap on mobile width
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the instruction to the agent is equally short:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Run TC-CART-014 against the staging URL in a headless browser.
For each step, note what you did and what you saw.
Take a screenshot at the end. Report PASS/FAIL per expected result, with evidence.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The output is a short report with screenshots. &lt;strong&gt;You verify visually.&lt;/strong&gt; A human skims the evidence in seconds, which is often faster than debugging a red CI job. No script exists, so no script can rot.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this beats healing for the long tail
&lt;/h2&gt;

&lt;p&gt;Most suites have a long tail of checks that matter for a moment: a release, a migration, a risky refactor, a single bug report you want to reproduce. Writing and maintaining scripts for those is a bad trade.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No maintenance.&lt;/strong&gt; The "automation" is regenerated on every run from the test case. UI redesigns don't break anything.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Runs in the background.&lt;/strong&gt; Kick it off in headless mode and come back to a report. It doesn't need your attention while it runs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Visual verification.&lt;/strong&gt; Screenshots catch things scripted assertions never asked about: overlapping elements, a broken layout, a wrong image.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test cases stay the durable asset.&lt;/strong&gt; Everyone can read and review them, including product and manual testers.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Promote only what earns it
&lt;/h2&gt;

&lt;p&gt;Here's the part I like most: this isn't a one-way door. Those background runs double as a &lt;strong&gt;discovery pipeline for your maintained suite&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;When a scenario keeps showing up, catches real bugs, or guards something critical, that's a signal it deserves a proper, deterministic script. At that point you promote it: take what the agent did (the steps, the selectors it found, the state it needed) and turn it into a reviewed test with real assertions.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;expect&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@playwright/test&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;promo banner shows for a new promo SKU&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;request&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/api/cart&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;data&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;sku&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;PROMO-001&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;qty&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;()).&lt;/span&gt;&lt;span class="nf"&gt;toBeTruthy&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;cartId&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;goto&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`/cart/&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;cartId&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByTestId&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;promo-banner&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;toBeVisible&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;My rule: &lt;strong&gt;if you wouldn't urgently fix a script when it breaks, don't promote it.&lt;/strong&gt; Everything else stays as a test case the agent runs on demand. Your maintained suite stays small and trusted, because every script in it was chosen, not accumulated.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this goes wrong
&lt;/h2&gt;

&lt;p&gt;I don't want to oversell it. Some honest limits:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;They aren't a merge gate.&lt;/strong&gt; Agent runs are non-deterministic, slower than a compiled script, and cost tokens. Two runs may take different paths or judge "visible" differently. Fine for exploration and release checks, not for the thousand-test regression run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agents can be too forgiving.&lt;/strong&gt; An agent that "works around" a broken flow may report a pass where a real user would be stuck. Write expected results that are strict, and read the evidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Core flows need code.&lt;/strong&gt; Audit and compliance trails need deterministic, durable scripts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test data and environment knowledge&lt;/strong&gt; still come from you. AI doesn't know your seeded accounts or your staging quirks unless you tell it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Some failures are real bugs.&lt;/strong&gt; Never re-run until it goes green.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Bottom Line
&lt;/h2&gt;

&lt;p&gt;Self-healing optimizes for keeping every script alive. The throwaway approach optimizes for keeping the &lt;em&gt;right&lt;/em&gt; scripts alive: let AI read your test cases and run them headless, verify the evidence visually, and only turn a scenario into maintained code when it has proven its worth.&lt;/p&gt;

&lt;p&gt;Your value as an SDET isn't how many scripts you maintain. It's the judgment about which checks deserve to become code at all.&lt;/p&gt;

&lt;p&gt;So here's my question: which part of your current suite could be replaced by a well-written test case and an agent, and what would you need to see before you trusted it?&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://luthfiferdian.com/blog/throwaway-automation-instead-of-self-healing" rel="noopener noreferrer"&gt;luthfiferdian.com&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>sdet</category>
      <category>ai</category>
      <category>automation</category>
      <category>testing</category>
    </item>
    <item>
      <title>I Let AI Write My Entire Test Suite — Here's What It Missed</title>
      <dc:creator>Luthfi Ferdian</dc:creator>
      <pubDate>Wed, 15 Apr 2026 14:07:00 +0000</pubDate>
      <link>https://dev.to/ferdianluthfi/i-let-ai-write-my-entire-test-suite-heres-what-it-missed-4jff</link>
      <guid>https://dev.to/ferdianluthfi/i-let-ai-write-my-entire-test-suite-heres-what-it-missed-4jff</guid>
      <description>&lt;h3&gt;
  
  
  Introduction
&lt;/h3&gt;

&lt;p&gt;As an SDET, writing test cases is one of my core responsibilities. What we test and how we test it directly shapes the quality of what we ship. So when AI tools started promising to automate test case generation, I had to try it. I believe writing the test case 100% manually will be outdated and counterproductive, so I gave an AI my PRD and TCD, and asked it to generate the full suite. What came back in seconds would have taken me two to three days. Some of it was genuinely impressive — but some of it would have let real bugs slip into production.&lt;/p&gt;

&lt;h3&gt;
  
  
  What AI Got Right
&lt;/h3&gt;

&lt;p&gt;The result stuns me. I wouldn't say we can trust it with 100% of the creation. Yes AI can also be wrong, yes AI can put its incorrect assumptions on the table, and yes absolutely it can create a very generic test case that could slow the test case creation rather than speed it up. As far as I know, AI is mostly about context. Less context means your result will be generic as per AI's assumptions. Too much context on the other hand may also reduce AI's accuracy and results as too many things it should consider. Not to mention not all the context is actually needed to perform the action. With proper context that we give to the AI, I would say at least 80% of the test case creation effort could be done by the AI.&lt;/p&gt;

&lt;h3&gt;
  
  
  What AI Missed
&lt;/h3&gt;

&lt;p&gt;Though it helps a lot, I found that human intervention is still required (at least at this time 😂). Test cases generated by the AI still lack in some areas. First is the &lt;strong&gt;integration&lt;/strong&gt; of the current PRD &lt;strong&gt;with the whole system&lt;/strong&gt;. Many PRDs are essentially about enhancing current capabilities of our product. It means enhancing some part of the large system it is in. AI still faces difficulty in understanding how the current PRD could benefit the whole system and how the whole flow could be affected by it. Second is &lt;strong&gt;generating&lt;/strong&gt; the proper &lt;strong&gt;test steps and test data&lt;/strong&gt;. Since it is difficult for the AI to understand the whole system flow, sometimes it means very generic test steps. Steps that fail to capture the personalized flow of the system. Test data on the other hand is really important to the testing activities. Without it, it is simply impossible to do the testing. Some test cases need specific test data to be set. AI still lacks knowledge on how to prepare the test data.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Bottom Line
&lt;/h3&gt;

&lt;p&gt;So should you use AI to write your test cases? Absolutely — but treat it as a first draft, not a final product. Here's my rule of thumb: let AI handle the happy-path and standard validation cases (it's great at those), but always review what it generated. AI got me about 80% of what I need. That last 20% — the part that requires system knowledge, domain context, and real test data — is exactly where SDETs prove their value. The testers who learn to work with AI on the 80% and focus their energy on the 20% will be the ones who thrive.&lt;/p&gt;

&lt;h3&gt;
  
  
  Call to Action
&lt;/h3&gt;

&lt;p&gt;Not leveraging AI in your daily work as an SDET would be a missed opportunity. Even though there's room for improvement, AI already saves significant time on the repetitive parts of test creation. Learning to work with it — trying, iterating, and figuring out where it fits in your workflow — is a process that I find joy in. How's your experience using AI in your daily tasks?&lt;/p&gt;

</description>
      <category>qa</category>
      <category>testing</category>
      <category>ai</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
