<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Pranta Kundu</title>
    <description>The latest articles on DEV Community by Pranta Kundu (@prantakunduqa).</description>
    <link>https://dev.to/prantakunduqa</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4118484%2F2b56fd78-c883-4be7-9c5c-9681e07becd9.jpg</url>
      <title>DEV Community: Pranta Kundu</title>
      <link>https://dev.to/prantakunduqa</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/prantakunduqa"/>
    <language>en</language>
    <item>
      <title>Race Conditions Don't Happen Once. They Return With Every New Feature.</title>
      <dc:creator>Pranta Kundu</dc:creator>
      <pubDate>Tue, 29 Sep 2026 06:24:00 +0000</pubDate>
      <link>https://dev.to/prantakunduqa/race-conditions-dont-happen-once-they-return-with-every-new-feature-40fl</link>
      <guid>https://dev.to/prantakunduqa/race-conditions-dont-happen-once-they-return-with-every-new-feature-40fl</guid>
      <description>&lt;p&gt;&lt;em&gt;Your feature passed every test. Then two users clicked at the same moment. Why concurrency bugs keep coming back, and what developers and QA should each do about it.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Feature Worked. The System Failed.
&lt;/h2&gt;

&lt;p&gt;Here is a story I have seen in many forms. (This is a hypothetical scenario built from common real-world patterns.)&lt;/p&gt;

&lt;p&gt;A team ships a "last item in stock" feature. It passes functional testing, API testing, automation, regression, and staging. QA signs off.&lt;/p&gt;

&lt;p&gt;On the first busy evening, two customers press "Buy" within a few milliseconds of each other. Both get a success message. The warehouse has one item and two orders.&lt;/p&gt;

&lt;p&gt;The feature worked. The system failed under concurrency.&lt;/p&gt;

&lt;p&gt;Next sprint, a coupon limit ships with the same bug: one coupon, used twice. Then a wallet balance. Then a booking slot. Different features, same root problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A race condition is not a bug you fix once. It returns whenever a new feature touches shared state.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Let's take it in order: when concurrent work happens, how it becomes a race condition, what database and design decisions create the risk, and what developers and QA should each do.&lt;/p&gt;




&lt;h2&gt;
  
  
  Part 1: When Does Concurrent Work Happen?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Concurrent work means more than one thing is happening at the same time on the same system.&lt;/strong&gt; This is normal. It is not a bug. Every real product is concurrent all day.&lt;/p&gt;

&lt;p&gt;A simple picture: two people stand at two different ATMs and withdraw money from the same joint account at the same moment. The bank must handle both correctly. That is concurrency, and the shared account is the "shared state."&lt;/p&gt;

&lt;p&gt;Real situations in software:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Two users, one resource.&lt;/strong&gt; Two customers buy the last item. Two agents book the same seat.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One user, two requests.&lt;/strong&gt; A user double-clicks "Pay," or a mobile app retries after a slow network. The server receives the same action twice.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Two admins, one record.&lt;/strong&gt; Two people open the same order and both change its status.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;User and background job.&lt;/strong&gt; A customer cancels an order while a scheduled job is shipping it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Two workers, one event.&lt;/strong&gt; A queue delivers the same message and two workers process it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Many servers, one database.&lt;/strong&gt; Production runs several copies of your app, all writing to the same rows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A slow job that overlaps itself.&lt;/strong&gt; A nightly job runs longer than expected and the next run starts before it ends.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these needs a huge traffic spike. Two requests are enough. Concurrency is not only a performance topic. It is a correctness topic.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Concurrent work is normal. It becomes a problem only when it touches shared state without protection.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Part 2: How Concurrent Work Becomes a Race Condition
&lt;/h2&gt;

&lt;p&gt;A race condition happens when &lt;strong&gt;the result depends on which request finishes first, or on how the steps of two requests interleave.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Three ingredients must be present:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Shared state:&lt;/strong&gt; something both requests read or change (stock, balance, status).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Overlapping timing:&lt;/strong&gt; both requests are in progress at the same time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A gap between checking and changing:&lt;/strong&gt; the code reads a value, decides, and writes later, with no protection in between.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Remove any one ingredient and the race disappears.&lt;/p&gt;

&lt;h3&gt;
  
  
  Example 1: Overselling the last item
&lt;/h3&gt;

&lt;p&gt;Stock is 1. The code does this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Step 1 → Read the stock (returns 1)
Step 2 → Check: is stock greater than 0? Yes, continue
Step 3 → Update stock to 0
Step 4 → Create the order
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Request A and Request B both finish Step 1 before either reaches Step 3. Both see stock = 1. Both pass the check. Both create an order.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;sequenceDiagram
    participant A as Request A
    participant DB as Database (stock = 1)
    participant B as Request B

    A-&amp;gt;&amp;gt;DB: Read stock
    DB--&amp;gt;&amp;gt;A: stock = 1
    B-&amp;gt;&amp;gt;DB: Read stock
    DB--&amp;gt;&amp;gt;B: stock = 1
    Note over A,B: Both pass the "stock &amp;gt; 0" check
    A-&amp;gt;&amp;gt;DB: Update stock = 0
    A-&amp;gt;&amp;gt;DB: Create order A
    B-&amp;gt;&amp;gt;DB: Update stock = 0
    B-&amp;gt;&amp;gt;DB: Create order B
    Note over DB: Result: stock = 0, but 2 orders exist for 1 item&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The time between the read (Step 1) and the write (Step 3) is called the &lt;strong&gt;race window&lt;/strong&gt;. It may last only a few milliseconds, but production traffic will find it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Example 2: The lost update
&lt;/h3&gt;

&lt;p&gt;A wallet has a balance of 100. Two deposits of 50 arrive together.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request A reads 100 → Request B reads 100 → A writes 150 → B writes 150
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The customer should have 200. They have 150. No error was shown. &lt;strong&gt;A lost update is dangerous because it fails silently.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Example 3: Invalid state
&lt;/h3&gt;

&lt;p&gt;An order is PAID. A customer cancels while the warehouse marks it SHIPPED. Both updates succeed. The order ends as CANCELLED, but the package is already on a truck. Each update was valid alone. The combination is not.&lt;/p&gt;

&lt;h3&gt;
  
  
  Example 4: Duplicate processing
&lt;/h3&gt;

&lt;p&gt;A payment request times out. The client retries. But the first request was still running and finished. The customer is charged twice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why is it so hard to catch?&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It depends on exact timing, so it is intermittent.&lt;/li&gt;
&lt;li&gt;Sequential tests never create the overlap.&lt;/li&gt;
&lt;li&gt;It often passes in staging, where traffic is low.&lt;/li&gt;
&lt;li&gt;It leaves quiet damage (wrong numbers, duplicates, bad state), not crashes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why does it keep returning?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Each new feature has its own code path and its own read-check-write. The fix for the stock feature protects only the stock feature. The coupon feature starts with no protection, because nobody asked the concurrency question for it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Every feature that changes shared state should trigger one question: what happens if two actors do this at the same time?&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Part 3: How It Arises: Database and System Design Concerns
&lt;/h2&gt;

&lt;p&gt;A race condition is created in two places: the database layer and the system design.&lt;/p&gt;

&lt;h3&gt;
  
  
  Database concerns
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;1. Read-modify-write outside a safe boundary.&lt;/strong&gt; The most common cause. The read and the write are separate steps with a gap between them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Transaction boundary in the wrong place.&lt;/strong&gt; If the read is outside the transaction and the write is inside, the gap is still open. A transaction that is too wide also hurts: it holds locks longer and causes waiting and deadlocks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Isolation level.&lt;/strong&gt; Isolation controls what one running transaction can see of another. Different levels allow different anomalies. A test that passes on one setting can fail on another. Check that your test database uses the same isolation setting as production.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Missing constraints.&lt;/strong&gt; If a rule matters ("one coupon per user," "one active booking per slot"), the database should enforce it with a unique constraint. Application code can have bugs or forgotten code paths. A constraint cannot be bypassed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Locking choices.&lt;/strong&gt; Row-level locks block other writers on the same row. They protect data but make requests wait. Two transactions can also lock each other, causing a deadlock, where the database kills one of them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. Optimistic vs pessimistic locking.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Pessimistic:&lt;/strong&gt; lock first, work after. Safe, but slower under contention.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Optimistic:&lt;/strong&gt; no lock, but a version number. If the version changed, the update is rejected. Fast, but the app must handle the conflict error.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  System design concerns
&lt;/h3&gt;

&lt;p&gt;A request travels through several layers, and each one carries its own concurrency risk:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    U[User] --&amp;gt; F[Frontend]
    F --&amp;gt; G[API Gateway]
    G --&amp;gt; AI["App instances&amp;lt;br/&amp;gt;(multiple copies)"]
    AI --&amp;gt; BL[Business logic]
    BL --&amp;gt; DB[(Database)]
    BL --&amp;gt; Q["Queue &amp;amp; workers"]
    BL --&amp;gt; EXT[External service]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;&lt;strong&gt;1. Multiple app instances.&lt;/strong&gt; A lock in one server's memory does nothing for the other servers. Protection must live in shared places, like the database or a shared lock service.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Queues that deliver at least once.&lt;/strong&gt; Many message systems can deliver the same message more than once. Consumers must be safe to run twice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Retries and timeouts.&lt;/strong&gt; When a call times out, the caller does not know if the work happened. A blind retry can duplicate it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Missing idempotency.&lt;/strong&gt; Idempotent means doing something twice has the same effect as doing it once. Without an idempotency key or a natural unique rule, the system cannot recognize a repeated request.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Cache and database out of sync.&lt;/strong&gt; If the cache says 5 in stock and the database says 0, a decision from the cache is wrong. That is a race between two copies of the truth.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. Async flows and external services.&lt;/strong&gt; A payment provider or warehouse system adds delay, and delay widens race windows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;7. Unclear source of truth.&lt;/strong&gt; If two services can both change the same data, races between services are very hard to prevent.&lt;/p&gt;




&lt;h2&gt;
  
  
  Part 4: Developer Concerns
&lt;/h2&gt;

&lt;p&gt;Developers own the implementation. Here is what I would expect, and what I would gladly review.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Write down the rule first.&lt;/strong&gt; "Stock can never go below zero." "One coupon per user." If the rule is not clear, nobody can protect it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Make the check and the change one step.&lt;/strong&gt; For the stock example, do it in one atomic database statement:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;UPDATE&lt;/span&gt; &lt;span class="n"&gt;products&lt;/span&gt;
&lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;stock&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;stock&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;101&lt;/span&gt; &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;stock&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then check how many rows were updated. One row means success. Zero rows means someone else got it, so return a clear "out of stock" response.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Choose the right control for the problem.&lt;/strong&gt; There is no universal answer. Common options:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Atomic update:&lt;/strong&gt; best for counters and stock.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Row lock (pessimistic):&lt;/strong&gt; useful when several steps must happen together.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;  &lt;span class="k"&gt;BEGIN&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;products&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;101&lt;/span&gt; &lt;span class="k"&gt;FOR&lt;/span&gt; &lt;span class="k"&gt;UPDATE&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="c1"&gt;-- check stock, then update&lt;/span&gt;
  &lt;span class="k"&gt;UPDATE&lt;/span&gt; &lt;span class="n"&gt;products&lt;/span&gt; &lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;stock&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;stock&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;101&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;COMMIT&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Optimistic version check:&lt;/strong&gt; useful for records people edit.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;  &lt;span class="k"&gt;UPDATE&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;
  &lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'SHIPPED'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;version&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;version&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
  &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;55&lt;/span&gt; &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="k"&gt;version&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If zero rows changed, someone else edited first.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Unique constraint plus idempotency key:&lt;/strong&gt; best for "do this only once." A repeat request fails on the unique index, and you return the first result.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Serialize through a queue:&lt;/strong&gt; all changes for one item go through one ordered path.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The right choice depends on your database, transaction boundary, business rule, isolation level, traffic, and how strict consistency must be. One SQL pattern does not solve every case.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Do not rely on in-memory locks&lt;/strong&gt; in a system with multiple instances.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Handle conflict errors kindly.&lt;/strong&gt; A losing request should get a clear response, not a 500 error. Decide when it is safe to retry.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. Keep transactions short.&lt;/strong&gt; Long transactions hold locks and cause waiting and deadlocks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;7. Make consumers and retries safe.&lt;/strong&gt; Any handler that a queue or client may call twice must be idempotent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;8. Add request IDs to logs.&lt;/strong&gt; When something breaks at 2 a.m., correlation IDs are how you reconstruct what happened.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;9. Tell QA where the risk is.&lt;/strong&gt; A short note in the pull request, like "this updates stock inside a transaction using X," helps QA target tests.&lt;/p&gt;




&lt;h2&gt;
  
  
  Part 5: QA Concerns
&lt;/h2&gt;

&lt;p&gt;This is where I spend most of my time. The SDET does not have to write the locking code. The SDET has to notice the risk, design the tests, and check the truth in the data.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1: Spot the risk before testing
&lt;/h3&gt;

&lt;p&gt;The most valuable sentence an SDET can say in grooming:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"This feature changes shared state, so concurrency is part of its risk profile."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Use this checklist for every feature.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Shared state&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Do multiple users touch the same resource?&lt;/li&gt;
&lt;li&gt;Is stock, balance, or status involved?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Parallel requests&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Can two requests arrive at nearly the same time?&lt;/li&gt;
&lt;li&gt;Can a user double-click, or a client retry?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Async processing&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Are queues or workers involved?&lt;/li&gt;
&lt;li&gt;Can a job run twice?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Database&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Is there read → modify → write logic?&lt;/li&gt;
&lt;li&gt;Which constraint protects the rule?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;External dependencies&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What if a payment call times out? Can the client retry?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Distributed system&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Are there multiple instances? Is a cache involved?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Several "yes" or "not sure" answers mean the feature needs concurrency tests.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Test in four levels
&lt;/h3&gt;

&lt;p&gt;Most teams stop at level 2.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Level 1, Functional:&lt;/strong&gt; Request A → Success&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Level 2, Sequential:&lt;/strong&gt; Request A completes → then Request B completes&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Level 3, Concurrent:&lt;/strong&gt; Request A and Request B hit the API at the same moment&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Level 4, High concurrency:&lt;/strong&gt; 10 users → 50 requests → 100 requests → 500 requests&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Measure more than the status code: correctness, duplicate records, lost updates, wrong state, transaction failures, deadlocks, timeouts, retry behavior, error rate, response time, and throughput.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A concurrency test is not passed because the API returned HTTP 200.&lt;/strong&gt; Two 200 responses for one item is a failure. The final database state and business outcome are the real result.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: A practical test design
&lt;/h3&gt;

&lt;p&gt;Scenario: only one user can buy the final item.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Initial state:&lt;/strong&gt; Stock = 1&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Concurrent requests:&lt;/strong&gt; User A and User B both send &lt;code&gt;POST /orders&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expected:&lt;/strong&gt; one request succeeds, one fails gracefully (for example, "out of stock")&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Database:&lt;/strong&gt; stock = 0, exactly 1 successful order&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There must be no negative stock, no duplicate reservation, no inconsistent order status, and no double payment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to run it:&lt;/strong&gt; write a small script (JavaScript or Python) that fires 20 requests at the same instant using a shared start time, and prints every status code. You do not need a full load-testing platform for a first reproduction.&lt;/p&gt;

&lt;p&gt;Then verify the truth in the database:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Query the stock for the product. It must be 0, never negative.&lt;/li&gt;
&lt;li&gt;Count the successful orders for the product. It must be exactly 1.&lt;/li&gt;
&lt;li&gt;Group orders by user and look for any user with more than one.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Check five things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;API response:&lt;/strong&gt; one success, the rest handled cleanly, no 500 errors&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Database state:&lt;/strong&gt; stock is 0, never negative&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Queue or event state:&lt;/strong&gt; one event published, not two&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Logs:&lt;/strong&gt; no deadlocks or unexpected exceptions&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Business outcome:&lt;/strong&gt; one order, one charge, one shipment&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Step 4: Reproduce it reliably
&lt;/h3&gt;

&lt;p&gt;Race conditions are intermittent, so make them appear on purpose:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Start all requests at the same instant&lt;/li&gt;
&lt;li&gt;Use more parallel workers&lt;/li&gt;
&lt;li&gt;Repeat the test many times&lt;/li&gt;
&lt;li&gt;Add a small artificial delay between the read and the write in a test build, to widen the race window&lt;/li&gt;
&lt;li&gt;Hit the same row from many requests&lt;/li&gt;
&lt;li&gt;Simulate retries and network delay&lt;/li&gt;
&lt;li&gt;Reset test data to a known state before each run&lt;/li&gt;
&lt;li&gt;Tag every request with a unique ID for log tracing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Flaky test or race condition?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Run once → Pass. Run 100 times → Failure. This does not automatically mean the test is flaky.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A flaky test&lt;/strong&gt; is unreliable because of the test itself: bad waits, shared test data, an unstable environment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A race condition&lt;/strong&gt; is real product behavior that depends on timing. The test is reliable and is telling you the truth.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;To tell them apart, fix the test setup (isolated data, deterministic waits) and rerun. If it still fails under concurrency, it is the product. Never label a concurrency failure "flaky, ignore." That is how these bugs reach production.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 5: Use tools by purpose
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;To generate concurrent requests&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A small custom script is often best for the first reproduction&lt;/li&gt;
&lt;li&gt;k6 or JMeter when you need many virtual users and metrics&lt;/li&gt;
&lt;li&gt;Postman/Newman or Playwright API tests when your suite already exists&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;To investigate the database&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Lock and transaction views for waiting queries&lt;/li&gt;
&lt;li&gt;Query, slow query, and deadlock logs&lt;/li&gt;
&lt;li&gt;Direct queries comparing final state to expected state&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;To investigate the application&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Request and correlation IDs&lt;/li&gt;
&lt;li&gt;Timestamps in logs&lt;/li&gt;
&lt;li&gt;Distributed tracing, if available&lt;/li&gt;
&lt;li&gt;Queue dashboards and worker logs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;In CI/CD&lt;/strong&gt;&lt;br&gt;
A small concurrency check can run before release. Even 20 concurrent requests that verify final database state can catch a broken feature on every deployment.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 6: Follow a feature-level workflow
&lt;/h3&gt;



&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[New feature] --&amp;gt; B[Identify shared state]
    B --&amp;gt; C[Identify concurrent actors]
    C --&amp;gt; D[Identify race window]
    D --&amp;gt; E[Review system design]
    E --&amp;gt; F[Review database &amp;amp; transaction strategy]
    F --&amp;gt; G[Create concurrent test]
    G --&amp;gt; H[Run controlled reproduction]
    H --&amp;gt; I[Increase concurrency]
    I --&amp;gt; J[Verify API, database, and business state]
    J --&amp;gt; K[Analyze logs and tracing]
    K --&amp;gt; L[Regression]
    L --&amp;gt; M[Release decision]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The first four steps happen before any test is written. Concurrency thinking is a design-time activity. If you start at "create concurrent test," you are already late.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 7: Ask these in feature grooming
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;What shared state does this feature modify?&lt;/li&gt;
&lt;li&gt;Can two actors modify it at the same time?&lt;/li&gt;
&lt;li&gt;What prevents duplicate processing?&lt;/li&gt;
&lt;li&gt;Is this operation idempotent?&lt;/li&gt;
&lt;li&gt;Where does the transaction start and end?&lt;/li&gt;
&lt;li&gt;What happens if the request times out after the server has processed it?&lt;/li&gt;
&lt;li&gt;Can the queue deliver the same message twice?&lt;/li&gt;
&lt;li&gt;Which database constraint protects this business rule?&lt;/li&gt;
&lt;li&gt;What happens when two workers pick up the same entity?&lt;/li&gt;
&lt;li&gt;How can QA reproduce the expected concurrency behavior?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If nobody can answer question 10, you have found a testability problem. Better to find it now than after release.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 8: Release readiness
&lt;/h3&gt;

&lt;p&gt;Before release, a feature that changes shared state should have answers to these:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What happens when two requests arrive at once?&lt;/li&gt;
&lt;li&gt;What happens when the same request is retried?&lt;/li&gt;
&lt;li&gt;What happens when a worker processes an event twice?&lt;/li&gt;
&lt;li&gt;What happens when an update conflicts?&lt;/li&gt;
&lt;li&gt;What happens when an external service times out and the client retries?&lt;/li&gt;
&lt;li&gt;Is the operation idempotent where it needs to be?&lt;/li&gt;
&lt;li&gt;Can invalid state or duplicate records exist?&lt;/li&gt;
&lt;li&gt;Is the final database state correct?&lt;/li&gt;
&lt;li&gt;Are failures visible in logs or alerts?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Release readiness is not only "does the happy path work?" It is "does the system keep its business rules true when concurrency goes wrong?"&lt;/p&gt;

&lt;p&gt;Be honest about where this effort is not worth it. A single-user internal tool, data written by only one process, or a read-only feature does not need the same depth. Match the effort to the risk.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who Owns What
&lt;/h2&gt;

&lt;p&gt;Nobody owns this alone. It is built across the chain:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Product:&lt;/strong&gt; states the business rule clearly&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Architect:&lt;/strong&gt; decides the source of truth and how requests are ordered&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Developer:&lt;/strong&gt; implements the control and documents the transaction boundary&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Database engineer:&lt;/strong&gt; reviews constraints, indexes, isolation, and locking&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;QA / SDET:&lt;/strong&gt; spots the risk, designs the concurrent tests, verifies final data&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DevOps / SRE:&lt;/strong&gt; makes failures visible with logs, metrics, alerts, and correct retry and timeout settings&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;Every new feature is a new opportunity for a concurrency bug.&lt;/p&gt;

&lt;p&gt;That is why QA should ask two questions, not one. The first is the one we always ask: "Does this feature work?" The second is the one that catches these bugs: "What happens when multiple actors try to change the same thing at the same time?"&lt;/p&gt;

&lt;p&gt;Good SDET work is not only proving that the system works. It is discovering the conditions under which the system stops being correct, and making those conditions visible before production does.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;What was the last concurrency bug your team shipped, and which check would have caught it earlier?&lt;/strong&gt; I'd like to compare notes in the comments.&lt;/p&gt;

</description>
      <category>softwaretesting</category>
      <category>concurrencytesting</category>
      <category>sre</category>
      <category>qa</category>
    </item>
    <item>
      <title>From Automation Frameworks to Agentic QA: What Changes When AI Can Take Action?</title>
      <dc:creator>Pranta Kundu</dc:creator>
      <pubDate>Mon, 14 Sep 2026 18:12:24 +0000</pubDate>
      <link>https://dev.to/prantakunduqa/from-automation-frameworks-to-agentic-qa-what-changes-when-ai-can-take-action-377a</link>
      <guid>https://dev.to/prantakunduqa/from-automation-frameworks-to-agentic-qa-what-changes-when-ai-can-take-action-377a</guid>
      <description>&lt;p&gt;&lt;strong&gt;&lt;em&gt;A Staff SDET's field guide to the architectural shift — for juniors who want to actually understand it, not just hear the buzzword.&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Your test fails at 2 AM. The pipeline goes red. A human wakes up, opens the trace viewer, stares at a screenshot, greps the logs, guesses it's a selector change, fixes it, reruns, goes back to sleep.&lt;/p&gt;

&lt;p&gt;Now imagine the pipeline doesn't wait for that human. It looks at the failure itself, forms a hypothesis, checks the DOM, checks the API response, decides whether the app broke or the test broke, patches the test if it's safe to, reruns it, and only wakes a human up if it's &lt;em&gt;not&lt;/em&gt; sure.&lt;/p&gt;

&lt;p&gt;That's not "AI will replace testers." That's a new layer sitting on top of the framework you already built. This article is about exactly what that layer is made of, what it can and can't be trusted with, and what you — as a QA engineer — need to learn next.&lt;/p&gt;

&lt;p&gt;No hype. No "70% faster" stats pulled from nowhere. Just architecture, trade-offs, and one worked example you can map onto your own codebase.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. The Shift in One Picture
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Traditional automation loop:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A[Write] --&amp;gt; B[Run]
    B --&amp;gt; C[Fail]
    C --&amp;gt; D["Human investigates"]
    D --&amp;gt; E[Fix]
    E --&amp;gt; F[Rerun]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Every arrow after "Fail" requires a person. The framework can tell you &lt;em&gt;that&lt;/em&gt; something broke. It cannot tell you &lt;em&gt;why&lt;/em&gt;, and it definitely can't decide &lt;em&gt;what to do about it&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agentic QA loop:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A[Goal] --&amp;gt; B[Plan]
    B --&amp;gt; C[Act]
    C --&amp;gt; D[Observe]
    D --&amp;gt; E[Diagnose]
    E --&amp;gt; F[Modify]
    F --&amp;gt; G[Retest]
    G --&amp;gt; H[Report]
    G -.loop until confident or escalated.-&amp;gt; C&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The difference isn't "more AI." It's that the system now has a &lt;strong&gt;goal&lt;/strong&gt; instead of a &lt;strong&gt;script&lt;/strong&gt;, and a &lt;strong&gt;loop&lt;/strong&gt; instead of a &lt;strong&gt;dead end&lt;/strong&gt;. A traditional test is a fixed sequence of steps someone wrote in advance. An agent is given an objective — "verify checkout completes for a logged-in user" — and decides, step by step, how to pursue it, using the actual state of the application in front of it.&lt;/p&gt;

&lt;p&gt;Keep that distinction in your head for the rest of this article: &lt;strong&gt;script = predetermined steps. agent = goal + reasoning + tools + feedback loop.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  2. What a Traditional Automation Framework Actually Does
&lt;/h2&gt;

&lt;p&gt;Before touching agents, it's worth being precise about what you already have, because the agent doesn't replace any of this — it sits on top of it.&lt;/p&gt;

&lt;p&gt;A mature framework is a stack of responsibilities:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Responsibility&lt;/th&gt;
&lt;th&gt;Example&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Test code&lt;/td&gt;
&lt;td&gt;Encodes expected behavior as assertions&lt;/td&gt;
&lt;td&gt;&lt;code&gt;expect(page.locator('#total')).toHaveText('$49.99')&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Page Objects / Components&lt;/td&gt;
&lt;td&gt;Abstract UI structure from test logic&lt;/td&gt;
&lt;td&gt;&lt;code&gt;CheckoutPage.submitOrder()&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Utilities&lt;/td&gt;
&lt;td&gt;Shared helpers (auth, data setup, waits)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;loginAs(user)&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Test data&lt;/td&gt;
&lt;td&gt;Inputs and fixtures&lt;/td&gt;
&lt;td&gt;JSON/DB seed data, factories&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Assertions&lt;/td&gt;
&lt;td&gt;Pass/fail decision logic&lt;/td&gt;
&lt;td&gt;Playwright's web-first &lt;code&gt;expect&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Configuration&lt;/td&gt;
&lt;td&gt;Environments, browsers, retries&lt;/td&gt;
&lt;td&gt;&lt;code&gt;playwright.config.ts&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reporting&lt;/td&gt;
&lt;td&gt;Human-readable results&lt;/td&gt;
&lt;td&gt;HTML report, Allure, trace viewer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CI/CD&lt;/td&gt;
&lt;td&gt;Trigger, execute, gate&lt;/td&gt;
&lt;td&gt;GitHub Actions, Jenkins&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This stack is deterministic by design. That determinism is a &lt;em&gt;feature&lt;/em&gt; — it's why automation is trustworthy. An agent doesn't remove any of these layers. It adds a reasoning layer that can &lt;strong&gt;operate&lt;/strong&gt; these layers instead of a human doing it manually when things go wrong.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. What Changes When an Agent Enters the Architecture
&lt;/h2&gt;

&lt;p&gt;Here's the honest list of new capabilities — and new problems — that show up the moment you add an agent.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Reasoning&lt;/strong&gt; — the system forms a hypothesis ("this looks like a timing issue, not a locator change") instead of just reporting a stack trace.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool use&lt;/strong&gt; — the agent can call real tools: run a test, read a log file, query an API, inspect the DOM, open a git diff — not just generate text about them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context&lt;/strong&gt; — it needs to know the application's structure, the test's intent, recent commits, and prior failures, not just the current error.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Observation&lt;/strong&gt; — it reads back the &lt;em&gt;result&lt;/em&gt; of its own actions (a screenshot, an accessibility snapshot, an HTTP status) and updates its plan.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;State&lt;/strong&gt; — it tracks what it has already tried across a multi-step investigation, instead of starting fresh each time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decision-making&lt;/strong&gt; — it chooses between several next actions (retry / inspect further / modify code / escalate) based on confidence, not a fixed script.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Controlled action&lt;/strong&gt; — and this is the one that actually matters — it can &lt;em&gt;change things&lt;/em&gt;: rerun a suite, edit a locator, open a PR, post to Slack. This is what separates "AI-assisted" from "agentic," and it's exactly where guardrails become non-negotiable.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last point is the whole ballgame. An LLM that &lt;em&gt;suggests&lt;/em&gt; a fix in a chat window is low risk. An agent that &lt;em&gt;applies&lt;/em&gt; the fix, commits it, and reruns the pipeline is a different risk category entirely — even if the underlying model is identical.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Traditional vs. AI-Assisted vs. Agentic vs. Self-Healing — The Comparison Table
&lt;/h2&gt;

&lt;p&gt;This is the table people usually get wrong, because "self-healing" and "agentic" get used interchangeably. They are not the same thing.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Traditional Automation&lt;/th&gt;
&lt;th&gt;AI-Assisted Automation&lt;/th&gt;
&lt;th&gt;Self-Healing Automation&lt;/th&gt;
&lt;th&gt;Agentic QA&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Autonomy&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;None — runs exactly what's written&lt;/td&gt;
&lt;td&gt;None — suggests, human applies&lt;/td&gt;
&lt;td&gt;Narrow — auto-fixes locators only&lt;/td&gt;
&lt;td&gt;Broad — plans and executes multi-step investigation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Decision-making&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Fixed pass/fail assertion&lt;/td&gt;
&lt;td&gt;Human decides after AI suggestion&lt;/td&gt;
&lt;td&gt;Rule-based similarity match (e.g., nearest DOM match)&lt;/td&gt;
&lt;td&gt;Model reasons over observed state to choose next action&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Tool usage&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Test runner only&lt;/td&gt;
&lt;td&gt;None (chat/IDE suggestion)&lt;/td&gt;
&lt;td&gt;Internal locator-matching heuristic&lt;/td&gt;
&lt;td&gt;Multiple tools: browser, runner, logs, API, source code, git&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Debugging&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Human reads logs/trace&lt;/td&gt;
&lt;td&gt;AI explains logs, human interprets&lt;/td&gt;
&lt;td&gt;N/A — doesn't debug, just adapts selectors&lt;/td&gt;
&lt;td&gt;Agent inspects logs, DOM, network, diffs to form root cause&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Retesting&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Manual rerun&lt;/td&gt;
&lt;td&gt;Manual rerun&lt;/td&gt;
&lt;td&gt;Automatic (same test, healed locator)&lt;/td&gt;
&lt;td&gt;Agent decides if/how to rerun, possibly with modified test&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Risk profile&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Low — deterministic, predictable&lt;/td&gt;
&lt;td&gt;Low — advisory only&lt;/td&gt;
&lt;td&gt;Medium — can mask a real UI regression as a "fix"&lt;/td&gt;
&lt;td&gt;Medium–High — can take unreviewed action if unguarded&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Human involvement&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Full ownership of triage &amp;amp; fix&lt;/td&gt;
&lt;td&gt;Full ownership, AI as assistant&lt;/td&gt;
&lt;td&gt;Configuration + periodic audit&lt;/td&gt;
&lt;td&gt;Approval gates, escalation review, strategy decisions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Best for&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Stable, well-understood flows&lt;/td&gt;
&lt;td&gt;Speeding up authoring/debugging&lt;/td&gt;
&lt;td&gt;Cosmetic UI drift (renamed classes, reordered DOM)&lt;/td&gt;
&lt;td&gt;Ambiguous failures needing multi-step investigation&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Trade-off to internalize:&lt;/strong&gt; self-healing automation solves a narrow, well-bounded problem (locator drift) with a narrow, well-bounded mechanism (similarity matching). Agentic QA solves a much wider problem (arbitrary failure investigation) with a much wider mechanism (an LLM planning over tools) — which means wider power &lt;em&gt;and&lt;/em&gt; wider blast radius. Don't reach for an agent to solve what self-healing already solves cheaply and safely.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. One Failure, Three Ways to Handle It — Login Workflow
&lt;/h2&gt;

&lt;p&gt;Scenario: your CI run just failed on &lt;code&gt;should log in with valid credentials&lt;/code&gt;. Here's how each generation of automation handles it, with actual JS/Playwright code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The test:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// login.spec.js&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;expect&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@playwright/test&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;should log in with valid credentials&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;goto&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/login&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByLabel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Email&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;fill&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;user@example.com&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByLabel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Password&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;fill&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Secret123!&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByRole&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;button&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Sign in&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;click&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByTestId&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;dashboard-welcome&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;toBeVisible&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It fails with: &lt;code&gt;TimeoutError: element not found — [data-testid="dashboard-welcome"]&lt;/code&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  A) Traditional Automation
&lt;/h3&gt;

&lt;p&gt;The pipeline goes red. That's it. The framework's job ends at "assertion failed." A human opens the HTML report, downloads the trace, scrubs through screenshots, and manually decides: is this a real regression, a flaky wait, or a renamed test-id? Time to resolution depends entirely on who's on call and how fast they can context-switch.&lt;/p&gt;

&lt;h3&gt;
  
  
  B) AI-Assisted Automation
&lt;/h3&gt;

&lt;p&gt;Same failure, but now an engineer pastes the stack trace and screenshot into a chat interface. The model reads the trace and says something like: &lt;em&gt;"The welcome banner's &lt;code&gt;data-testid&lt;/code&gt; may have changed — check the latest frontend diff for the dashboard component."&lt;/em&gt; Useful. Faster than manual grepping. But the human still opens the repo, checks the diff, decides, and edits the test by hand. The AI never touched the pipeline.&lt;/p&gt;

&lt;h3&gt;
  
  
  C) Agentic QA Workflow
&lt;/h3&gt;

&lt;p&gt;The agent is wired into the pipeline with tool access. On failure it:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Observes&lt;/strong&gt; — pulls the Playwright trace, screenshot, and accessibility snapshot at the point of failure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Diagnoses&lt;/strong&gt; — queries the app's current DOM via a Playwright MCP-style tool call and compares it against the expected &lt;code&gt;data-testid&lt;/code&gt;. Finds &lt;code&gt;dashboard-welcome-v2&lt;/code&gt; now exists where &lt;code&gt;dashboard-welcome&lt;/code&gt; used to.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Checks intent&lt;/strong&gt; — pulls the last few commits touching the dashboard component to see if this was an intentional rename (it was — part of a rebrand PR merged two hours ago).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decides&lt;/strong&gt; — this is a cosmetic, intentional rename, not a functional regression. Confidence is high enough to act.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Acts&lt;/strong&gt; — updates the locator in the test file, opens a draft PR (does &lt;strong&gt;not&lt;/strong&gt; merge it), and reruns the test against the fix.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reports&lt;/strong&gt; — posts to the team channel: &lt;em&gt;"login.spec.js failed due to a test-id rename in PR #482 (dashboard rebrand). Draft fix opened in PR #611, test now passing locally. Awaiting review."&lt;/em&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Notice what the agent did &lt;strong&gt;not&lt;/strong&gt; do: it did not merge its own fix, it did not silently loosen the assertion, and it did not decide on its own that the rebrand was "fine" from a product standpoint — it only judged that the &lt;em&gt;locator&lt;/em&gt; mismatch was cosmetic. That distinction — technical judgment vs. product judgment — is exactly the line covered in Section 8.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Practical Architecture: How the Pieces Actually Fit Together
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj5y5tu5csgg09436sg3u.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj5y5tu5csgg09436sg3u.png" alt=" " width="800" height="451"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why each box matters:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Tools, not vision:&lt;/strong&gt; In 2026, the dominant pattern for browser-based agents is reading a structured &lt;strong&gt;accessibility snapshot&lt;/strong&gt; rather than a screenshot — Microsoft's open-source Playwright MCP server exposes dozens of browser-control tools this way, which makes agent actions deterministic and auditable instead of vision-model guesswork.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Observe is not optional:&lt;/strong&gt; an agent that acts without reading back the result of its action is just automation with extra steps and extra risk. The observe step is what turns "action" into a controlled feedback loop instead of a blind write.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evaluate is the guardrail chokepoint:&lt;/strong&gt; this is where confidence thresholds, allow-lists of permitted actions, and human-approval gates live. Everything downstream of "Evaluate" should be scoped by policy, not by the model's own judgment alone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Escalate is a feature, not a failure:&lt;/strong&gt; a good agent architecture treats "I'm not sure, ask a human" as a successful outcome, not a fallback of last resort. Teams that punish escalation (by treating it as "the AI didn't work") end up training the system, implicitly, to guess more confidently — which is the opposite of what you want.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  7. What "Taking Action" Actually Means
&lt;/h2&gt;

&lt;p&gt;"Agentic" gets thrown around loosely. Concretely, action means the agent can do some subset of the following, through real tool calls:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Running tests&lt;/strong&gt; — triggering a specific spec or suite, not just reading results.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Inspecting failures&lt;/strong&gt; — pulling traces, console logs, network waterfalls, accessibility snapshots.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Navigating the application&lt;/strong&gt; — clicking through the live app to check whether a flow still works, independent of the failing test's script.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Collecting artifacts&lt;/strong&gt; — screenshots, HAR files, DB state snapshots, timestamps — building an evidence trail, not just a verdict.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Modifying test code&lt;/strong&gt; — updating a locator, adjusting a wait strategy, adding a missing &lt;code&gt;await&lt;/code&gt; — scoped to test code, ideally as a draft PR, never a direct merge.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rerunning tests&lt;/strong&gt; — after a change, to confirm the fix actually resolves the failure rather than just silencing it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Escalating uncertain failures&lt;/strong&gt; — routing to a human with a structured summary when confidence is low or the change touches assertions/business logic.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A simplified version of a modify-and-retest tool call, using a Playwright MCP–style setup, looks like this conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// agent-tools.js — simplified illustrative tool wrapper&lt;/span&gt;
&lt;span class="c1"&gt;// (structure mirrors how an MCP-based agent would call Playwright tools)&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;inspectFailure&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;testId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;snapshot&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;accessibility&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;snapshot&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;trace&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;context&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nx"&gt;tracing&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stop&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`traces/&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;testId&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;.zip`&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;snapshot&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;trace&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;proposeLocatorFix&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;oldLocator&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;domSnapshot&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// agent reasoning step happens here (LLM call, not shown),&lt;/span&gt;
  &lt;span class="c1"&gt;// returns a candidate replacement based on structural similarity&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;findClosestMatch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;oldLocator&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;domSnapshot&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;applyFixAsDraftPR&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;filePath&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;oldLocator&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;newLocator&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// scoped to test files only — never application source code&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;git&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;checkoutBranch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;agent/fix-login-locator&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;editFile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;filePath&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;oldLocator&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;newLocator&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;git&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;commit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;fix(test): update stale locator after rebrand (agent-proposed)&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;git&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;openDraftPR&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;reviewers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;qa-team&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;rerunAndConfirm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;specFile&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;runPlaywright&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;specFile&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;passed&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice the shape: &lt;strong&gt;every action is narrow, logged, and reversible.&lt;/strong&gt; That's not incidental — it's the actual design requirement for any agent you let near your pipeline.&lt;/p&gt;




&lt;h2&gt;
  
  
  8. What Agents Still Cannot Reliably Determine
&lt;/h2&gt;

&lt;p&gt;This is the section junior engineers should bookmark. An agent can be extremely good at &lt;em&gt;technical&lt;/em&gt; diagnosis and still be the wrong entity to make these calls:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Business intent&lt;/strong&gt; — whether a behavior change is "expected" often depends on a roadmap conversation, a support ticket, or a product decision the agent has no visibility into.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Acceptable risk&lt;/strong&gt; — a flaky test on a marketing page and a flaky test on a payment confirmation step are not the same risk, even if the failure signature looks identical.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Whether a test should exist at all&lt;/strong&gt; — an agent can tell you a test is failing; it can't reliably tell you the test is testing the wrong thing, or that a whole flow is no longer worth covering.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Whether changing an assertion masks a real defect&lt;/strong&gt; — this is the single most dangerous action an agent can take unsupervised. Loosening &lt;code&gt;toHaveText('$49.99')&lt;/code&gt; to &lt;code&gt;toBeVisible()&lt;/code&gt; makes the test pass. It does not mean checkout still charges the right amount.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Whether a passing test provides meaningful coverage&lt;/strong&gt; — a green suite full of tests that got quietly "healed" into shallow assertions is a worse state than a red suite, because it's a false signal that's actively hiding risk.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Staff-level framing:&lt;/strong&gt; the job isn't shrinking, it's moving. The scarce skill is no longer &lt;em&gt;writing automation steps&lt;/em&gt; — it's &lt;strong&gt;engineering reliable systems that can reason and act around automation, and knowing exactly where to put the fence around them.&lt;/strong&gt; That's a systems-design skill, not a scripting skill, and it's a more senior skill than what most of us started our careers doing.&lt;/p&gt;




&lt;h2&gt;
  
  
  9. What SDETs Need to Learn Next
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Area&lt;/th&gt;
&lt;th&gt;Why it matters&lt;/th&gt;
&lt;th&gt;Where to start&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Agent fundamentals&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Understand planner/reasoning loops, not just prompts&lt;/td&gt;
&lt;td&gt;Study ReAct-style agent loops, ADK/agent-framework docs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Tool calling&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Agents are only as good as the tools they can invoke&lt;/td&gt;
&lt;td&gt;Practice writing well-scoped, single-purpose tool functions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;MCP (Model Context Protocol)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The emerging standard for exposing tools/context to LLMs&lt;/td&gt;
&lt;td&gt;Read the MCP spec; try Microsoft's Playwright MCP server locally&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Browser automation internals&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;You need to know what the agent is actually calling&lt;/td&gt;
&lt;td&gt;Deepen Playwright fundamentals — accessibility tree, tracing, network interception&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Context &amp;amp; state management&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Agents fail silently when context is stale or too large&lt;/td&gt;
&lt;td&gt;Learn context-window budgeting, memory/session design&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Agent evaluation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;"It worked once" isn't evidence&lt;/td&gt;
&lt;td&gt;Learn eval harnesses — golden datasets of known failures, pass/fail scoring for agent decisions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Observability &amp;amp; tracing&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;You'll debug the agent, not just the app&lt;/td&gt;
&lt;td&gt;OpenTelemetry basics, structured logging for multi-step agent runs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Guardrails &amp;amp; policy&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;This is the actual engineering discipline here&lt;/td&gt;
&lt;td&gt;Allow-lists, approval gates, sandboxed branches, rollback design&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;CI/CD integration&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Agents live in pipelines, not just chat windows&lt;/td&gt;
&lt;td&gt;GitHub Actions/Jenkins integration patterns, draft-PR workflows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Test strategy&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The highest-leverage skill, and the one AI can't do for you&lt;/td&gt;
&lt;td&gt;Risk-based test design, coverage strategy, what &lt;em&gt;not&lt;/em&gt; to automate&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  10. What Not to Automate With an Agent
&lt;/h2&gt;

&lt;p&gt;Concrete no-go list, not a vague warning:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Direct writes to production data or production environments&lt;/strong&gt; — an agent's "checkout works" investigation should never run against real customer accounts or real payment rails.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Merging its own pull requests&lt;/strong&gt; — draft only, human-reviewed, always.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deleting or disabling failing tests to make CI green&lt;/strong&gt; — this is the agentic equivalent of &lt;code&gt;.skip()&lt;/code&gt;-ing everything, and it's worse when it's automated because it happens silently and at scale.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Loosening assertions on regulated or financial flows&lt;/strong&gt; (pricing, tax calculation, payment confirmation) without a named human sign-off.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security- and auth-sensitive test paths&lt;/strong&gt; — session handling, permission boundaries, PII fields — where a "plausible fix" can quietly create a real vulnerability.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Irreversible destructive actions&lt;/strong&gt; — dropping test databases, rotating shared credentials, force-pushing to shared branches.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anything where "looks the same" isn't good enough&lt;/strong&gt; — visual-only comparisons on financial documents, legal text, or compliance-critical copy, where subtle wording changes matter more than layout.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Rule of thumb: &lt;strong&gt;if a wrong decision here is expensive, irreversible, or invisible until a customer hits it — keep a human explicitly in the loop, not just "available."&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  11. Five-Minute Checklist: Is My Framework Ready for Agentic QA?
&lt;/h2&gt;

&lt;p&gt;Score yourself honestly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Do my tests produce &lt;strong&gt;structured, machine-readable artifacts&lt;/strong&gt; (traces, accessibility snapshots, JSON reports) — not just pass/fail in a console log?&lt;/li&gt;
&lt;li&gt;Is my test code organized well enough (Page Objects, clear naming, isolated specs) that an automated diff to one file wouldn't ripple unpredictably elsewhere?&lt;/li&gt;
&lt;li&gt;Do I have &lt;strong&gt;CI/CD hooks&lt;/strong&gt; where an external process could safely trigger a rerun or open a draft PR without direct merge rights?&lt;/li&gt;
&lt;li&gt;Do I have a clear, written &lt;strong&gt;list of what should never be auto-modified&lt;/strong&gt; (assertions on financial/security flows, production config)?&lt;/li&gt;
&lt;li&gt;Is there a &lt;strong&gt;human-review gate&lt;/strong&gt; (required PR review, branch protection) that a bot account cannot bypass?&lt;/li&gt;
&lt;li&gt;Do I have &lt;strong&gt;observability&lt;/strong&gt; into &lt;em&gt;why&lt;/em&gt; a test failed beyond the assertion message (network logs, DOM state, recent commits)?&lt;/li&gt;
&lt;li&gt;Could I explain, in one sentence, the difference between what my &lt;strong&gt;self-healing tooling&lt;/strong&gt; already handles and what would require full agentic reasoning?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Five or more checked: you're structurally ready to pilot agentic workflows on a low-risk test suite. Fewer than five: fix the framework fundamentals first — an agent layered on a fragile framework just automates the fragility faster.&lt;/p&gt;




&lt;h2&gt;
  
  
  12. The Maturity Path
&lt;/h2&gt;



&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A["Traditional SDET"] --&amp;gt; B["AI-Assisted SDET&amp;lt;br/&amp;gt;adds AI as a debugging/authoring assistant"]
    B --&amp;gt; C["AI-Augmented SDET&amp;lt;br/&amp;gt;integrates AI suggestions into daily workflow, still manual application"]
    C --&amp;gt; D["Agentic Test Engineer&amp;lt;br/&amp;gt;builds/configures tool-using agents for bounded tasks, e.g. locator healing"]
    D --&amp;gt; E["AI QA Agent Engineer&amp;lt;br/&amp;gt;designs multi-step agent workflows, guardrails, escalation policy"]
    E --&amp;gt; F["Owns the architecture: planners, tool contracts, evals, observability, policy"]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;You don't skip stages by reading one article — but you can see, clearly, that every stage is still a &lt;em&gt;quality engineering&lt;/em&gt; skill. The tools changed. The judgment requirement didn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Actual Point
&lt;/h2&gt;

&lt;p&gt;Agentic QA doesn't remove the need for a human who understands the product, the risk, and the business — it removes the tedium of &lt;em&gt;manually operating&lt;/em&gt; the investigation loop a skilled QA engineer already runs in their head. The framework gets more capable. The judgment stays yours. That's the honest version of this shift, and it's a more interesting career than "write more scripts, faster."&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Discussion:&lt;/strong&gt; If an AI agent could run, debug, modify, and retest your automation today, which action would you trust it to take &lt;em&gt;without&lt;/em&gt; human approval?&lt;/p&gt;

</description>
      <category>playwright</category>
      <category>ai</category>
      <category>testing</category>
      <category>qa</category>
    </item>
    <item>
      <title>Agentic Testing Is Changing QA: When AI Agents Can Run, Debug, and Retest Your Tests</title>
      <dc:creator>Pranta Kundu</dc:creator>
      <pubDate>Sat, 12 Sep 2026 14:25:34 +0000</pubDate>
      <link>https://dev.to/prantakunduqa/agentic-testing-is-changing-qa-when-ai-agents-can-run-debug-and-retest-your-tests-2ik4</link>
      <guid>https://dev.to/prantakunduqa/agentic-testing-is-changing-qa-when-ai-agents-can-run-debug-and-retest-your-tests-2ik4</guid>
      <description>&lt;p&gt;&lt;em&gt;&lt;strong&gt;From CI failure to verified fix: how AI agents can run, diagnose, repair, and retest automated tests.&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Your test fails at 2 AM in the CI pipeline. You wake up, open the logs, scroll past 400 lines of stack trace, squint at a screenshot, and mutter "is this a locator issue or did the button actually move?" Twenty minutes later you find it — a &lt;code&gt;data-testid&lt;/code&gt; changed during a refactor. You fix it, push, and go back to your actual work.&lt;/p&gt;

&lt;p&gt;If that paragraph felt uncomfortably familiar, you already understand the problem agentic testing is trying to solve.&lt;/p&gt;

&lt;p&gt;This isn't another "AI will replace QA" post. It's a practical breakdown — from one SDET to another — of what agentic testing actually is, what it can realistically do today, and where it still needs you.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Agentic Testing Is Different Now (Not Just Rebranded RPA)
&lt;/h2&gt;

&lt;p&gt;If you've been in QA for more than a few years, you've seen "smart" testing tools come and go — rule-based self-healing, record-and-playback tools that claimed to "understand" your app, RPA scripts marketed as AI. It's fair to be skeptical that this is just another rebrand.&lt;/p&gt;

&lt;p&gt;Here's what's genuinely different this time, in concrete terms:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tool use, not just chat.&lt;/strong&gt; Earlier "AI in testing" meant a chatbot you pasted logs into. Current LLMs can call functions directly — run a command, read a file, query a DOM, hit an API — and decide which tool to call based on what they find. That's the mechanical difference between "an assistant you operate" and "an agent that operates tools."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Multi-step reasoning that holds up across steps.&lt;/strong&gt; Rule-based self-healing followed one fixed heuristic ("find nearest matching element"). Today's models can chain several reasoning steps — read a log, form a hypothesis, check a second source to confirm it, revise the hypothesis — which is closer to how you actually debug, not a single pattern match.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Context windows large enough to hold real evidence.&lt;/strong&gt; Diagnosing a failure well means holding the test code, the DOM snapshot, the network log, and recent git history in view at once. That was previously impractical; it's now routine.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cheaper, faster inference.&lt;/strong&gt; Running this reasoning loop on every CI failure would have been cost-prohibitive a few years ago. It's now realistic to run on a meaningful slice of your failures, not just a curated demo.&lt;/p&gt;

&lt;p&gt;To be clear about what this is not: it's not evidence that these systems reason the way humans do, and it's not a claim that failure rates have dropped by some measurable percentage — no credible industry-wide numbers exist yet, and be skeptical of anyone quoting one. What's changed is narrower and more mechanical: agents now have the tool access and reasoning chain-length to make the loop from earlier in this article actually work end-to-end, where before it had to be done manually or with brittle heuristics.&lt;/p&gt;




&lt;h2&gt;
  
  
  What "Agentic Testing" Actually Means
&lt;/h2&gt;

&lt;p&gt;Strip away the buzzword, and agentic testing is this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;An AI agent that can reason about a testing task, use tools to act on it, observe the results, and decide what to do next — instead of just executing a fixed script.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A traditional test does exactly what you wrote, every time, in the same order. It has no awareness that anything went wrong beyond a pass/fail signal.&lt;/p&gt;

&lt;p&gt;An agentic test system behaves more like a junior engineer sitting next to you: it runs the test, notices the failure, opens the logs itself, forms a hypothesis ("this looks like a timing issue, not a real bug"), checks the DOM to confirm, tries an adjusted action, and reports back with reasoning — not just a red X.&lt;/p&gt;

&lt;p&gt;The key word is &lt;strong&gt;loop&lt;/strong&gt;. Agentic testing isn't one clever trick (like auto-healing a broken locator). It's a repeatable reasoning loop applied to the whole investigation process.&lt;/p&gt;




&lt;h2&gt;
  
  
  Five Terms People Keep Mixing Up
&lt;/h2&gt;

&lt;p&gt;This confusion is the single biggest source of AI-testing hype. Let's separate them cleanly.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftxpsgn8afhtvx9osn981.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftxpsgn8afhtvx9osn981.png" alt=" " width="646" height="409"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Junior engineers often think self-healing is agentic testing. It isn't. Self-healing swaps a broken selector; it doesn't ask "did the product actually break, or did my test just get stale?" Agentic testing asks that question — and that's the whole point.&lt;/p&gt;




&lt;h2&gt;
  
  
  LLM/AI vs AI-Assisted Testing vs Agentic Testing — What Came First, and What Each One Actually Does
&lt;/h2&gt;

&lt;p&gt;This is where most articles blur three very different things into one "AI testing" blob. They're not interchangeable — they're layers, and each one entered your testing life in a different order, doing a different job.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The order they showed up in your actual workflow:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8lh3qqwkxhmhywyvgndb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8lh3qqwkxhmhywyvgndb.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;LLM/AI came first&lt;/strong&gt; — it's just a reasoning engine. It has no idea your test suite exists unless you copy-paste context into it manually. This is where most SDETs already live today: pasting a stack trace into an AI chat and asking "what's likely going on here?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI-assisted testing came next&lt;/strong&gt; — someone wired that same reasoning engine into your actual tooling (IDE plugin, test generator, log summarizer), but a human is still the one clicking "run," reading the suggestion, and deciding what to do with it. The AI never acts on its own.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agentic testing is the third layer&lt;/strong&gt; — the same reasoning engine, now given tools it can call itself (run the test, read the DOM, check the network tab, open a diff) and permission to move through the investigate → fix → validate loop without you manually feeding it each piece of context at every step. You still set the boundaries; it works inside them.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv8i69yj9b8wkv5qs05a0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv8i69yj9b8wkv5qs05a0.png" alt=" " width="619" height="362"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The progression is really about &lt;em&gt;who holds the tools and who presses "go" at each step&lt;/em&gt; — not about the AI getting "smarter" in some vague sense. Same underlying reasoning capability, increasing levels of tool access and autonomy.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Agentic Testing Loop
&lt;/h2&gt;

&lt;p&gt;Here's the loop in plain terms — the same one you already run manually when debugging a failure, except an agent runs it as a first pass. Each stage below is tagged with who's actually doing it, so it's clear where the "agentic" part really is:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8q0gu99g6gwbczsmxmkq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8q0gu99g6gwbczsmxmkq.png" alt=" " width="800" height="731"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Where the human sits in this loop: &lt;strong&gt;goal-setting&lt;/strong&gt; at the start (you decide what "correct" means), and &lt;strong&gt;approval/audit&lt;/strong&gt; at the two decision points — Change/Retry and Validate — especially for anything business-critical. Everything in between (Plan → Act → Observe → Diagnose) is where the agent is doing genuinely new work compared to a traditional script: it's reasoning across evidence, not just executing steps.&lt;/p&gt;

&lt;p&gt;Notice: &lt;strong&gt;Report is not optional&lt;/strong&gt;, and &lt;strong&gt;Validate always exists&lt;/strong&gt; — a responsible agentic system never silently commits a "fix" without confirming it actually holds up, and never hides its reasoning from the human reviewing it.&lt;/p&gt;




&lt;h2&gt;
  
  
  A Realistic Example: Login Test Failure
&lt;/h2&gt;

&lt;p&gt;Let's make this concrete instead of theoretical. Say you have a straightforward login test.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;user can log in with valid credentials&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;goto&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/login&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fill&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;#email&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;userk@example.com&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fill&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;#password&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Test1234!&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;click&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;#login-submit-btn&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;locator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;.dashboard-header&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;toBeVisible&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It fails in CI. Here's what a traditional pipeline gives you:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;FAIL: user can log in with valid credentials
Error: Timeout 30000ms exceeded waiting for locator '.dashboard-header'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's it. You now have to go do the investigation yourself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What an agentic system does instead&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 1 — Observe.&lt;/strong&gt; The agent captures the DOM snapshot at failure time, the network log, and a screenshot instead of just the timeout message.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 2 — Diagnose.&lt;/strong&gt; It compares the captured DOM against the expected state and finds this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Expected element: .dashboard-header
DOM at failure time contains: .dashboard__header  (double underscore, renamed in last deploy)
Network log: POST /api/login → 200 OK, redirect to /dashboard confirmed
Screenshot: Dashboard page IS visible, correctly rendered
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 3 — Form a hypothesis.&lt;/strong&gt; Login actually succeeded — the assertion is broken, not the product. The class name changed in a recent CSS refactor (a real PR diff would confirm this, if the agent has repo access).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 4 — Propose a change (not silently apply it).&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent suggestion:
  Confidence: High (87%)
  Cause: Locator drift (class renamed, unrelated to auth logic)
  Suggested fix: update locator to '.dashboard__header'
  Action taken: none yet - awaiting human approval (or auto-applied
  only if this test is flagged as "low-risk, auto-fixable" in config)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 5 — Validate.&lt;/strong&gt; If approved, it reruns the test against the new locator, confirms a clean pass, and logs the change with a diff — visible in the PR, not buried in a black box.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 6 — Report.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SUMMARY
Test: user can log in with valid credentials
Result: Locator drift, not a functional defect
Evidence: network 200 OK, dashboard rendered, class renamed in commit abc123
Fix applied: locator updated, rerun passed 3/3
Human review: recommended (auth-related test, high business risk)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the entire value proposition in one example: faster, evidence-backed triage — not unsupervised decision-making.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Compare that to what a self-healing tool alone would have done:&lt;/strong&gt; it would have just swapped &lt;code&gt;.dashboard-header&lt;/code&gt; for the nearest matching element and moved on — no diagnosis, no report explaining why, no distinction made between "cosmetic class rename" and "the dashboard didn't actually load." That distinction is exactly why agentic reasoning matters more than pattern-matching self-healing.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Full CI Workflow: From Failure to Verified Fix, With a Real Git Diff
&lt;/h2&gt;

&lt;p&gt;The login example above shows the reasoning. This section shows how that reasoning actually plugs into a pipeline you'd run in production — because "the agent diagnoses it" means nothing until you see where it sits relative to your CI, your git history, and your merge gate.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feeoktbt90h3hw0lxhid3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feeoktbt90h3hw0lxhid3.png" alt=" " width="800" height="1200"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The part worth underlining: the agent never pushes directly to your main branch. Its output is a draft PR with an attached evidence bundle — the same artifact a human contributor would produce, reviewable the same way. This is what makes "agent modifies the test" fundamentally different from "agent silently mutates CI state." If your implementation skips the draft-PR step and lets an agent commit straight to a protected branch, you've removed the one checkpoint that makes this whole workflow trustworthy.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Agents Can Realistically Do Today
&lt;/h2&gt;

&lt;p&gt;Framed as illustrative capability, not a benchmarked claim:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Inspect DOM/page state, accessibility tree, and console errors at the moment of failure&lt;/li&gt;
&lt;li&gt;Parse logs, stack traces, and network requests to narrow down a likely cause&lt;/li&gt;
&lt;li&gt;Correlate a failure with a recent code or config change, if given repo/CI access&lt;/li&gt;
&lt;li&gt;Distinguish common failure classes: locator drift, timing/race conditions, stale test data, environment issues&lt;/li&gt;
&lt;li&gt;Propose or apply a scoped fix (locator update, wait strategy, data refresh)&lt;/li&gt;
&lt;li&gt;Rerun the test and compare before/after outcomes&lt;/li&gt;
&lt;li&gt;Produce a structured, evidence-based failure summary instead of a raw stack trace&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What Agents Silently Cannot Guarantee
&lt;/h2&gt;

&lt;p&gt;This is the section most AI-testing content skips, and it's the most important one for your credibility as an SDET:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Business correctness.&lt;/strong&gt; An agent doesn't know that "checkout total should include a loyalty discount for tier-3 customers" — unless that rule is explicitly encoded somewhere it can see.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Detecting every functional defect.&lt;/strong&gt; It's diagnosing why a test failed, not independently discovering bugs the test wasn't written to catch.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Knowing intended product behavior.&lt;/strong&gt; If the requirement was never documented or the ticket is vague, the agent has no ground truth to compare against.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Safely changing tests without masking real defects.&lt;/strong&gt; An agent that "fixes" a failing assertion by loosening it can accidentally hide a genuine regression. This risk is real and needs guardrails (below).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Replacing human risk judgment.&lt;/strong&gt; Deciding whether a payment flow test failure is worth blocking a release is a judgment call about business risk — not a pattern-matching task.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;An agent that can fix a test is not the same as an agent that can prove the product is correct. Fixing a broken locator tells you the test is executable again. It tells you nothing about whether the feature behind that locator does what the business needs it to do. Treat every auto-applied fix as "test restored to a runnable state," not "feature verified correct."&lt;/p&gt;




&lt;h2&gt;
  
  
  Comparison Table: Autonomy, Reasoning, and Risk
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F50c9vu8rqghq40z88pku.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F50c9vu8rqghq40z88pku.png" alt=" " width="631" height="570"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Architecture of a Practical Agentic Testing System
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsj62j7w80nz29k2ej8h4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsj62j7w80nz29k2ej8h4.png" alt=" " width="800" height="960"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The guardrails box is not decoration — it's the component that decides which failures an agent may auto-fix versus which must always stop and wait for a human. Without it, you don't have agentic testing; you have an unsupervised script with extra confidence.&lt;/p&gt;




&lt;h2&gt;
  
  
  What "Tools" Actually Mean for an Agent (No Magic Here)
&lt;/h2&gt;

&lt;p&gt;"The agent inspects the DOM" sounds like it understands your app the way you do. It doesn't. An agent only has a defined, finite list of functions it's allowed to call — nothing more. If you didn't wire up a tool for it, it cannot do that thing, full stop. This matters because it demystifies the whole system: there's no hidden capability, just a list of functions with a name, inputs, and outputs.&lt;/p&gt;

&lt;p&gt;A realistic tool list for a testing agent looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Tool: run_test(test_name)
  → executes one test, returns pass/fail + exit code

Tool: get_dom_snapshot(page)
  → returns the current HTML/accessibility tree at time of call

Tool: read_logs(run_id)
  → returns console, network, and framework logs for a run

Tool: get_screenshot(run_id)
  → returns a screenshot captured at failure time

Tool: git_diff(since_commit)
  → returns code changes since a given commit/tag

Tool: read_file(path)
  → returns file contents (read-only, scoped to repo paths)

Tool: propose_fix(file, change)
  → does NOT apply the change - stages it for human approval

Tool: apply_fix(file, change)   [gated - only callable if
                                 autonomy config allows it
                                 for this specific test]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each tool is just a function with a schema — the agent picks which one to call based on its reasoning, the same way you'd decide "let me check the network tab" versus "let me check git blame" while debugging. The agent isn't smarter than the tools you gave it access to. If &lt;code&gt;git_diff&lt;/code&gt; isn't wired up, it can't correlate a failure with a recent code change — it'll guess instead, and guessing is exactly the failure mode you're trying to avoid. The quality of your agentic testing setup is mostly a function of how good and how scoped your tool list is — not how good the underlying model is.&lt;/p&gt;

&lt;p&gt;Notice &lt;code&gt;apply_fix&lt;/code&gt; is deliberately gated separately from &lt;code&gt;propose_fix&lt;/code&gt;. That split — "the agent can always suggest, but can only act where explicitly permitted" — is the single most important design decision in the whole system, and it's what your instruction file (next section) actually configures.&lt;/p&gt;




&lt;h2&gt;
  
  
  Hands-On: How to Actually Implement This in Your Framework
&lt;/h2&gt;

&lt;p&gt;Everything above explains the concept. This section is for the engineer who closes this article and asks "okay, but how do I actually set this up on Monday?" Here's a concrete, three-part breakdown.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Pick an AI Coding Agent: Claude Code / Codex / Local LLM
&lt;/h3&gt;

&lt;p&gt;You need a "coding agent" — something that can read your repo, run commands, and call tools — not just a chat window. The three realistic options today:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0enx0t6rvepdo4d5ed3m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0enx0t6rvepdo4d5ed3m.png" alt=" " width="632" height="410"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to actually decide:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Does your test data or codebase contain sensitive/regulated information? → Local LLM, or a vendor with a data-retention agreement you've verified.&lt;/li&gt;
&lt;li&gt;Do you need strong reasoning over messy logs and DOM diffs? → Hosted agent (Claude Code or Codex) generally outperforms smaller local models on this kind of multi-step reasoning today.&lt;/li&gt;
&lt;li&gt;Is budget the constraint, not data sensitivity? → Start with whichever hosted agent your team already pays for elsewhere — don't add a second subscription just for this.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Whichever you pick, verify the current subscription tiers and rate limits directly on the provider's docs before committing a budget line — pricing and included usage change frequently.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. The Instruction/Agent File — and Its Lifecycle
&lt;/h3&gt;

&lt;p&gt;This is the single most important artifact in the whole setup. It's a file (commonly named &lt;code&gt;AGENTS.md&lt;/code&gt;, &lt;code&gt;CLAUDE.md&lt;/code&gt;, or a system-prompt config) that tells the agent: what its job is, what tools it may use, what it must never do without approval, and what format to report in.&lt;/p&gt;

&lt;p&gt;A minimal version looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Test Agent Instructions&lt;/span&gt;

&lt;span class="gu"&gt;## Role&lt;/span&gt;
You are a test-failure triage agent for the checkout and login test suites.

&lt;span class="gu"&gt;## You MAY:&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Read test code, logs, DOM snapshots, and network traces
&lt;span class="p"&gt;-&lt;/span&gt; Rerun a failing test up to 2 times
&lt;span class="p"&gt;-&lt;/span&gt; Propose locator/wait/data fixes with a confidence score

&lt;span class="gu"&gt;## You MUST NOT:&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Modify test files without an approval flag set to true
&lt;span class="p"&gt;-&lt;/span&gt; Touch tests tagged &lt;span class="sb"&gt;`payment-critical`&lt;/span&gt; or &lt;span class="sb"&gt;`compliance`&lt;/span&gt; without human sign-off
&lt;span class="p"&gt;-&lt;/span&gt; Loosen an assertion without explicitly flagging it as a "risk: masking possible defect"

&lt;span class="gu"&gt;## Report format&lt;/span&gt;
Always output: cause, evidence, confidence %, suggested fix, risk flag.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You don't write this once and forget it — it has a lifecycle, the same way test code does:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyuow9d74kjf5hfv92nyy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyuow9d74kjf5hfv92nyy.png" alt=" " width="800" height="436"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Treat step 6 as non-negotiable. Every time an agent gets something wrong — masks a real bug, misdiagnoses a cause — that's a signal to update the instruction file, not just to override the output once and move on. Version-control this file the same way you version-control test code, with PR review on changes.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Multi-Agent Prompt Setup: Splitting the Work Across Specialized Agents
&lt;/h3&gt;

&lt;p&gt;A single "do everything" agent gets unreliable fast. In practice, this works better as a small team of narrowly-scoped agents that hand off to each other — similar to how you'd split responsibilities across QA roles on a real team.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feezjxm5b4v1vwiqlcged.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feezjxm5b4v1vwiqlcged.png" alt=" " width="800" height="436"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Illustrative role prompts for each agent (adapt scope/tools to your stack):
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fed3ik4xeaw34lpv4wpgu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fed3ik4xeaw34lpv4wpgu.png" alt=" " width="640" height="645"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A few more worth adding as your setup matures:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Flaky Test Detector&lt;/strong&gt; — tracks pass/fail history over time and flags tests that fail intermittently regardless of code changes, so they get fixed at the root instead of being repeatedly "healed."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test Data Manager&lt;/strong&gt; — checks whether a failure traces back to stale/expired test fixtures or seed data, a very common false-positive source that's easy to under-diagnose.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Regression Risk Scorer&lt;/strong&gt; — cross-references a failing test against recent deploys/PRs to estimate whether this failure correlates with a specific change, speeding up root-cause attribution.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Keep each agent's prompt scoped to one job. The temptation is to write one giant prompt that does analysis, scripting, healing, and escalation together — resist it. Narrow scope is what makes each agent's output reviewable, and reviewability is what makes the whole system trustworthy enough to actually put in CI.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. A First Experiment You Can Actually Run This Week
&lt;/h3&gt;

&lt;p&gt;Don't start by wiring all five agents into CI. Start with one test and one narrow question. Here's a scoped experiment using a Playwright suite you already have:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Pick one flaky or recently-broken test — not a payment or auth test. A low-risk UI check is ideal for a first run.&lt;/li&gt;
&lt;li&gt;Manually gather the evidence bundle the way an agent would: the failure's stack trace, a DOM snapshot at failure time (&lt;code&gt;page.content()&lt;/code&gt; in Playwright), the relevant network log, and git diff since the last known-green commit.&lt;/li&gt;
&lt;li&gt;Paste all of it into a single prompt to whichever AI coding agent you chose in step 1 (Claude Code, Codex, or a local model), something close to:&lt;/li&gt;
&lt;/ol&gt;

&lt;blockquote&gt;
&lt;p&gt;Here is a failing Playwright test, its stack trace, a DOM snapshot at failure time,&lt;br&gt;
   and the git diff since the last passing run. Classify the failure as: locator drift,&lt;br&gt;
   timing issue, data issue, or likely real defect. State your confidence and cite the&lt;br&gt;
   specific evidence you used. Do not modify any files - diagnosis only.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;ol&gt;
&lt;li&gt;Judge the output yourself — not on whether it "sounds smart," but on whether its cited evidence actually supports its conclusion. This is the reviewer skill from the "What This Means for SDETs" section, and this experiment is where you practice it.&lt;/li&gt;
&lt;li&gt;Only if the diagnosis holds up, take the second step: ask it to propose (not apply) a minimal fix, and manually verify that fix against a rerun yourself.&lt;/li&gt;
&lt;li&gt;Log what happened — was the diagnosis right, wrong, or partially right? That log is the seed of your instruction file's guardrails from the lifecycle diagram earlier in this section.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This costs you one afternoon, no CI changes, and no new tooling investment — and it tells you, concretely, whether your test suite and your chosen agent are actually a good match before you build any automation around it. If the diagnosis is consistently solid on your first five attempts, you have a real candidate for the CI-integrated workflow above. If it's shaky, that's useful information too — it tells you your logs/evidence capture needs work before an agent (or a human) can diagnose reliably from them.&lt;/p&gt;




&lt;h2&gt;
  
  
  What This Means for SDETs
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fruzejowrqsvqa1e2ifst.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fruzejowrqsvqa1e2ifst.png" alt=" " width="631" height="344"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For a junior engineer, this is genuinely good news day-to-day: instead of spending 40 minutes tracing why a test failed, you review a structured summary and decide in 5 minutes whether the fix is sound. But that 5 minutes requires you to actually understand why the agent reached its conclusion — you're becoming a reviewer of reasoning, not just a writer of scripts. That's a skill you build the same way you built debugging skill: by doing it, deliberately, on real failures.&lt;/p&gt;




&lt;h2&gt;
  
  
  5-Minute Checklist: Is This Test a Good Candidate for Agentic Testing?
&lt;/h2&gt;

&lt;p&gt;Use this before handing any test suite to an agentic workflow:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Is the failure mode usually mechanical (locators, timing, stale data) rather than deep business logic?&lt;/li&gt;
&lt;li&gt;Does the test have clear, deterministic pass/fail criteria (not a fuzzy or flaky assertion)?&lt;/li&gt;
&lt;li&gt;Is there low-to-medium business risk if a wrong auto-fix temporarily masks a regression?&lt;/li&gt;
&lt;li&gt;Do you have log/DOM/network access rich enough for the agent to actually diagnose, not guess?&lt;/li&gt;
&lt;li&gt;Is there an audit trail requirement you can satisfy (every change logged, reviewable)?&lt;/li&gt;
&lt;li&gt;Can you define a human approval gate for anything touching payments, auth, or compliance-sensitive flows?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Score 5–6 yes:&lt;/strong&gt; good candidate for agent-assisted auto-fix with light review.&lt;br&gt;
&lt;strong&gt;Score 3–4 yes:&lt;/strong&gt; good candidate for agent-assisted diagnosis only, human applies the fix.&lt;br&gt;
&lt;strong&gt;Score 0–2 yes:&lt;/strong&gt; keep this test fully manual — the risk of a silent wrong fix outweighs the time saved.&lt;/p&gt;




&lt;h2&gt;
  
  
  Decision Matrix: When NOT to Use Agentic Testing
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbtyhwpdw52c2dwu567uj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbtyhwpdw52c2dwu567uj.png" alt=" " width="630" height="376"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  A Realistic Adoption Path
&lt;/h2&gt;

&lt;p&gt;Nobody should go from "manual debugging" to "fully autonomous agent commits fixes to main" in one step. A sane rollout looks like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Start AI-assisted.&lt;/strong&gt; Use AI to summarize failures and suggest causes; a human still fixes everything.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Automate the repetitive investigation.&lt;/strong&gt; Let the agent gather DOM/logs/network evidence automatically, even if a human still decides the fix.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Introduce controlled agents on low-risk suites.&lt;/strong&gt; Smoke tests, non-critical UI checks — auto-fix with logged rationale, human spot-checks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Add evaluation and guardrails.&lt;/strong&gt; Define autonomy scopes, blocked test lists, and confidence thresholds below which the agent must always escalate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Integrate with CI/CD.&lt;/strong&gt; Agent runs as part of the pipeline, gated by the guardrails from step 4, not bypassing them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure outcomes.&lt;/strong&gt; Track false-fix rate, time-to-triage, and — critically — how often a "fix" later correlated with a real regression that got masked. Adjust scope based on that data, not on vibes.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  The Takeaway
&lt;/h2&gt;

&lt;p&gt;Agentic testing is a real shift in how the debugging half of QA work gets done — not because AI is infallible, but because a structured reasoning loop over logs, DOM state, and code history can triage mechanical failures faster than a human doing it manually every time.&lt;/p&gt;

&lt;p&gt;It does not replace the parts of QA that were never about typing test steps in the first place: understanding what "correct" means for the business, deciding what risk is acceptable to ship, and knowing when a green checkmark is lying to you.&lt;/p&gt;

&lt;p&gt;Treat agentic testing as augmentation of your engineering judgment, not a substitute for it. The agents get faster at the loop. The judgment about when to trust the loop — that's still yours.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Would you trust an AI agent to debug, modify, and retest your automation without human approval? Where would you draw the line?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Drop your take below — especially if you've drawn that line somewhere different from where you expected to.&lt;/p&gt;

</description>
      <category>playwright</category>
      <category>ai</category>
      <category>testing</category>
      <category>qa</category>
    </item>
    <item>
      <title>When Should AI Write Your Tests? A Staff SDET's Decision Framework + Prompts</title>
      <dc:creator>Pranta Kundu</dc:creator>
      <pubDate>Thu, 10 Sep 2026 17:12:25 +0000</pubDate>
      <link>https://dev.to/prantakunduqa/when-should-ai-write-your-tests-a-staff-sdets-decision-framework-prompts-48a2</link>
      <guid>https://dev.to/prantakunduqa/when-should-ai-write-your-tests-a-staff-sdets-decision-framework-prompts-48a2</guid>
      <description>&lt;p&gt;&lt;strong&gt;The most expensive mistake in AI-assisted QA isn't a bad test. It's not knowing why you wrote it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every week I watch engineers open ChatGPT, paste a user story, and type "write test cases for this." Ten seconds later they have 40 lines of confident-looking test code. Nobody asked what actually needed to be tested. Nobody asked what could actually break. The AI didn't fail — the &lt;em&gt;question&lt;/em&gt; did. A generated test that verifies the wrong thing is worse than no test at all, because it gives you a green checkmark and a false sense of safety.&lt;/p&gt;

&lt;p&gt;This isn't an anti-AI article. I use AI to write tests every single day. But after 8+ years designing automation frameworks and watching teams adopt AI tooling, I've learned that the skill that matters now isn't "can you prompt an LLM" — it's "can you tell, before you prompt it, whether this is even a job for AI." That's the decision most junior engineers skip. Let's fix that.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Core Question Nobody Asks
&lt;/h2&gt;

&lt;p&gt;Not "can AI write this test?" — it almost always &lt;em&gt;can&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;The real question is: &lt;strong&gt;"Should AI own this decision, or just the typing?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Writing a test involves two separate jobs that get collapsed into one prompt:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Test strategy&lt;/strong&gt; — deciding &lt;em&gt;what&lt;/em&gt; needs to be verified, under what conditions, and why it matters.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test implementation&lt;/strong&gt; — turning that decision into code.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;AI is frequently excellent at #2. It is inconsistent — and sometimes dangerous — at #1. The framework below exists to help you separate these two jobs before you ever open a prompt window.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Staff SDET Decision Framework: 8 Dimensions to Evaluate First
&lt;/h2&gt;

&lt;p&gt;Before you decide whether AI writes, assists, or stays out of a test, run it through these eight lenses. This takes less time to do than to read.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Question to Ask&lt;/th&gt;
&lt;th&gt;Why It Matters&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Business risk&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;If this fails in production, what breaks — a typo or a transaction?&lt;/td&gt;
&lt;td&gt;High-risk flows (payments, auth, data integrity) need human-owned assertions, not AI guesses.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Test complexity&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Is this a single assertion, or a multi-step state machine?&lt;/td&gt;
&lt;td&gt;AI handles linear flows well; it loses the thread in deep conditional logic.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Requirements clarity&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Is the expected behavior documented, or does it live in someone's head?&lt;/td&gt;
&lt;td&gt;AI cannot infer undocumented business rules — it will invent plausible-sounding ones instead.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Repetition/volume&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Am I writing 1 unique test or 30 variations of the same pattern?&lt;/td&gt;
&lt;td&gt;High-volume, low-variance work is where AI has the best cost/benefit ratio.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Framework maturity&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Do we have established patterns, page objects, fixtures AI can follow?&lt;/td&gt;
&lt;td&gt;AI writing into a mature framework is safe. AI &lt;em&gt;designing&lt;/em&gt; the framework from scratch is not.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Test-data complexity&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Does this need realistic, interdependent, or regulated data?&lt;/td&gt;
&lt;td&gt;AI often invents data that "looks right" but violates real constraints (e.g. invalid state combinations).&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Debugging/maintenance cost&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;If this test is flaky in 3 months, who has to understand it?&lt;/td&gt;
&lt;td&gt;Code no one on the team understands is a liability, no matter who — or what — wrote it.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Need for human judgment&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Does "correct" depend on product intuition, UX nuance, or edge-case tradeoffs?&lt;/td&gt;
&lt;td&gt;Judgment calls are exactly where AI produces confident, wrong answers.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Run through these in under a minute. You're not filling out a form — you're building the instinct a Staff SDET already has.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Decision Tree: "Should AI Write This Test?"
&lt;/h2&gt;

&lt;p&gt;Here's the shape that instinct takes once it's explicit.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqgqg9ntrbizkucy184mh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqgqg9ntrbizkucy184mh.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Four honest outcomes — not "AI vs. human," but a spectrum:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;AI Writes&lt;/strong&gt; — generates the test independently; you review before merge.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI Assists&lt;/strong&gt; — you design the test, AI helps with syntax, boilerplate, or edge-case suggestions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Human-Led&lt;/strong&gt; — you write and own it; AI might draft a skeleton, nothing more.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do Not Automate Yet&lt;/strong&gt; — the real problem is unclear requirements, not test authorship.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last branch is the one juniors skip most often — and it's usually where the actual bug in the &lt;em&gt;process&lt;/em&gt; lives.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where AI Is Genuinely Excellent
&lt;/h2&gt;

&lt;p&gt;No hype needed here — these are the cases where AI reliably saves real hours:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Repetitive scenario expansion&lt;/strong&gt; — same flow, different inputs (valid email formats, currency values, locale variations).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API test variations&lt;/strong&gt; — status codes, header combinations, pagination edge cases.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Boundary and negative cases&lt;/strong&gt; — empty strings, max-length fields, null values, off-by-one boundaries.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Boilerplate framework code&lt;/strong&gt; — Playwright/Selenium page objects, fixture setup, wait handling, config scaffolding.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data-driven test tables&lt;/strong&gt; — generating CSV/JSON test matrices once the &lt;em&gt;logic&lt;/em&gt; is defined by a human.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Regression suite expansion&lt;/strong&gt; — converting an existing manual test case into automated coverage using an established pattern.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Common thread: &lt;strong&gt;the decision about what to test was already made by a human.&lt;/strong&gt; AI is filling in the matrix, not drawing it.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where AI Should NOT Be Trusted to Decide Independently
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Unclear or conflicting requirements&lt;/strong&gt; — AI will resolve ambiguity by guessing, and the guess will look confident.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Critical business workflows&lt;/strong&gt; — checkout, billing, auth, data migrations. The cost of a wrong assumption is too high.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Complex state transitions&lt;/strong&gt; — multi-step workflows with conditional branches (e.g., order states: draft → pending → paid → refunded → disputed).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security-sensitive behavior&lt;/strong&gt; — auth bypass attempts, permission boundaries, injection handling. Getting the &lt;em&gt;threat model&lt;/em&gt; wrong is worse than not testing at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anything where a wrong assertion creates false confidence&lt;/strong&gt; — a test that passes but checks the wrong thing is more dangerous than a missing test, because it hides the gap.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The risk isn't that AI writes broken code. It's that AI writes code that &lt;strong&gt;runs, passes, and proves nothing.&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvbkdrgiavtx7defcwumu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvbkdrgiavtx7defcwumu.png" alt=" " width="800" height="489"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  7 Practical Prompts You Can Copy-Paste (Browser ChatGPT)
&lt;/h2&gt;

&lt;p&gt;Each prompt includes &lt;strong&gt;when to use it&lt;/strong&gt;, &lt;strong&gt;what to feed it&lt;/strong&gt;, &lt;strong&gt;what to expect back&lt;/strong&gt;, and &lt;strong&gt;what you must verify yourself&lt;/strong&gt; — because the output is a draft, not a decision.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Generating test cases from requirements
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Prompt:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Act as a QA analyst. Here are the requirements for [feature]:
[paste requirements / user story / acceptance criteria]

List test cases grouped by: positive, negative, boundary, and edge cases.
For each test case, state the precondition, the action, and the expected result.
Do not write automation code yet — this is test design only.
Flag any requirement that is ambiguous or missing information instead of guessing.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;When to use:&lt;/strong&gt; Early in test design, right after requirements are written or refined.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Input needed:&lt;/strong&gt; The actual requirement or user story — not a summary of it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expected output:&lt;/strong&gt; A structured list of test case ideas with expected results.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You must verify:&lt;/strong&gt; That every "ambiguous" flag AI raises is actually resolved by a human before automation starts. Check for missing business rules AI wouldn't know (permissions, regional logic, edge policies).&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. Converting test cases into Playwright tests
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Prompt:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;Convert&lt;/span&gt; &lt;span class="nx"&gt;the&lt;/span&gt; &lt;span class="nx"&gt;following&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt; &lt;span class="nx"&gt;cases&lt;/span&gt; &lt;span class="nx"&gt;into&lt;/span&gt; &lt;span class="nx"&gt;Playwright&lt;/span&gt; &lt;span class="nx"&gt;tests&lt;/span&gt; &lt;span class="nx"&gt;using&lt;/span&gt; &lt;span class="nx"&gt;TypeScript&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;
&lt;span class="nx"&gt;Follow&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt; &lt;span class="nx"&gt;existing&lt;/span&gt; &lt;span class="nx"&gt;pattern&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="nx"&gt;our&lt;/span&gt; &lt;span class="nx"&gt;framework&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;paste&lt;/span&gt; &lt;span class="nx"&gt;an&lt;/span&gt; &lt;span class="nx"&gt;existing&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt; &lt;span class="nx"&gt;file&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nx"&gt;a&lt;/span&gt; &lt;span class="nx"&gt;style&lt;/span&gt; &lt;span class="nx"&gt;reference&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="nx"&gt;Use&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;testid&lt;/span&gt; &lt;span class="nx"&gt;selectors&lt;/span&gt; &lt;span class="nx"&gt;where&lt;/span&gt; &lt;span class="nx"&gt;possible&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt; &lt;span class="nx"&gt;Do&lt;/span&gt; &lt;span class="nx"&gt;not&lt;/span&gt; &lt;span class="nx"&gt;invent&lt;/span&gt; &lt;span class="nx"&gt;selectors&lt;/span&gt; &lt;span class="err"&gt;—&lt;/span&gt; &lt;span class="nx"&gt;mark&lt;/span&gt; &lt;span class="kr"&gt;any&lt;/span&gt; &lt;span class="nx"&gt;selector&lt;/span&gt;
&lt;span class="nx"&gt;you&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;re unsure about with a TODO comment instead of guessing.

Test cases:
[paste test cases]
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;When to use:&lt;/strong&gt; Once test cases are approved and you're ready to implement.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Input needed:&lt;/strong&gt; Approved test cases + a real example from your codebase for pattern consistency.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expected output:&lt;/strong&gt; Playwright test code matching your team's conventions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You must verify:&lt;/strong&gt; Every selector against the real DOM — AI will hallucinate plausible-looking selectors that don't exist. Confirm wait strategies aren't just &lt;code&gt;page.waitForTimeout()&lt;/code&gt; hacks.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3. Finding missing edge cases
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Prompt:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Here is a feature description and the test cases we already have:
[paste feature description]
[paste existing test list]

Identify edge cases or scenarios that are NOT covered.
Focus on: concurrency, empty/null states, permission boundaries, and unusual but
realistic user sequences. Do not repeat cases we already have.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;When to use:&lt;/strong&gt; As a pre-review step before finalizing test coverage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Input needed:&lt;/strong&gt; Your existing test list — the value here is in the &lt;em&gt;gap&lt;/em&gt;, not a fresh list.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expected output:&lt;/strong&gt; A short, targeted list of genuinely new scenarios.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You must verify:&lt;/strong&gt; Relevance — AI sometimes proposes edge cases that don't apply to your actual architecture (e.g., suggesting race-condition tests for a single-threaded batch job).&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  4. Generating negative API test scenarios
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Prompt:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Here is an API contract: [paste OpenAPI spec / endpoint docs / example request-response]

Generate negative test scenarios covering: invalid payloads, missing required fields,
wrong data types, unauthorized access, and unexpected status codes.
For each scenario, state the request, the expected status code, and expected error behavior
per the documented contract only — do not assume behavior that isn't documented.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;When to use:&lt;/strong&gt; When you have a documented API contract (Swagger/OpenAPI/internal spec).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Input needed:&lt;/strong&gt; The actual contract, not a paraphrase of it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expected output:&lt;/strong&gt; A scenario list mapped to documented expected behavior.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You must verify:&lt;/strong&gt; That expected status codes match the &lt;em&gt;real&lt;/em&gt; API, not just what's "typical" — AI defaults to REST conventions even if your API doesn't follow them.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  5. Reviewing AI-generated (or any) tests
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Prompt:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Review this test code as a senior QA engineer would:
[paste test code]

Check for: weak or missing assertions, hardcoded waits, brittle selectors,
tests that would pass even if the feature were broken, and duplicated coverage
with tests I already have (if I paste them below).
Be specific about what's wrong and why — don't just say "looks good."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;When to use:&lt;/strong&gt; Before merging any AI- or human-written test.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Input needed:&lt;/strong&gt; The test code, plus related existing tests if checking for duplication.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expected output:&lt;/strong&gt; A critique — ideally uncomfortable, not a rubber stamp.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You must verify:&lt;/strong&gt; Run the critique against the actual application behavior. AI can flag a weak assertion in theory but can't confirm whether the &lt;em&gt;strengthened&lt;/em&gt; assertion is actually correct for your app.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  6. Improving weak assertions
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Prompt:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Here is a test with an assertion I think is too weak:
[paste test code]

Suggest a stronger assertion that would actually fail if [describe the real bug
this test is supposed to catch]. Explain what the current assertion misses.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;When to use:&lt;/strong&gt; When a test "passes" but you suspect it wouldn't catch a real regression.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Input needed:&lt;/strong&gt; The test &lt;em&gt;and&lt;/em&gt; a clear description of the failure mode you're worried about.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expected output:&lt;/strong&gt; A tightened assertion plus reasoning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You must verify:&lt;/strong&gt; That the new assertion doesn't just become brittle instead of weak — overly strict assertions fail on harmless changes and erode trust in the suite.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  7. Identifying redundant or flaky tests
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Prompt:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Here is a list of test names and short descriptions from our suite:
[paste list]

Identify: (1) tests that likely overlap in coverage, (2) tests that sound like they
depend on timing/network state and are flaky-prone, (3) tests that test implementation
details rather than behavior. Explain your reasoning for each flag.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;When to use:&lt;/strong&gt; Periodic suite health reviews, or before a big regression run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Input needed:&lt;/strong&gt; Test names/descriptions — full code helps but isn't required for a first pass.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expected output:&lt;/strong&gt; Flags with reasoning, not automatic deletions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You must verify:&lt;/strong&gt; Every flag manually. "Sounds redundant" isn't the same as "is redundant" — two tests can look similar but cover different failure paths.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Before / After: A Real Prompt Upgrade
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Junior engineer's prompt:&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Write Playwright tests for the login page."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;What comes back:&lt;/strong&gt; Generic tests for a login form that may not match your actual UI, using invented selectors, testing only the happy path plus one wrong-password case, with no negative testing, no consideration of lockout policy, and no connection to your existing framework patterns.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Staff SDET-quality prompt:&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Here's our login page requirements doc [pasted], our existing Playwright page object for the signup flow as a style reference [pasted], and our account lockout policy (5 failed attempts locks for 15 minutes) [pasted]. Generate Playwright tests in TypeScript covering: successful login, invalid password, non-existent email, account lockout after 5 failed attempts, and session expiry redirect. Use &lt;code&gt;data-testid&lt;/code&gt; selectors matching our existing pattern. Mark any selector you're not certain about with a TODO instead of guessing. Do not test password strength rules — that's covered in the signup suite."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Why the difference matters:&lt;/strong&gt; The second prompt gives AI real constraints (existing patterns, actual business rules, explicit scope boundaries) instead of asking it to invent them. The output quality gap isn't about "better AI" — it's about the human doing the strategy work &lt;em&gt;before&lt;/em&gt; the prompt, not hoping the prompt does it for them.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Principle to Internalize
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;AI can generate test code. It cannot automatically determine whether the test proves the right thing.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Correctness of &lt;em&gt;syntax&lt;/em&gt; and correctness of &lt;em&gt;intent&lt;/em&gt; are two different problems. AI is increasingly reliable at the first. The second still requires someone who understands the business, the risk, and the failure modes — because "the test passed" and "the feature works" are not the same sentence.&lt;/p&gt;

&lt;p&gt;The risks worth naming plainly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hallucinated selectors&lt;/strong&gt; — code that looks right, targets nothing real.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Weak assertions&lt;/strong&gt; — tests that pass regardless of whether the feature is broken.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Duplicated coverage&lt;/strong&gt; — AI doesn't know what already exists unless you tell it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Incorrect assumptions&lt;/strong&gt; — AI fills gaps in unclear requirements with plausible guesses, silently.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;False confidence&lt;/strong&gt; — the most expensive risk of all: a green suite that isn't actually protecting you.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  The 5-Minute AI Test Generation Checklist (Save This)
&lt;/h2&gt;

&lt;p&gt;Before you accept any AI-generated test:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Do I know &lt;em&gt;why&lt;/em&gt; this test exists — what specific failure it should catch?&lt;/li&gt;
&lt;li&gt;Are the requirements behind this test documented, not assumed?&lt;/li&gt;
&lt;li&gt;Have I checked every selector against the real DOM/UI?&lt;/li&gt;
&lt;li&gt;Would this test actually fail if the real bug happened? (Mentally break the feature and check.)&lt;/li&gt;
&lt;li&gt;Does this duplicate existing coverage?&lt;/li&gt;
&lt;li&gt;Is the test data realistic and valid for our actual system constraints?&lt;/li&gt;
&lt;li&gt;If this test is flaky in 3 months, can someone on my team debug it without me?&lt;/li&gt;
&lt;li&gt;Did I make the risk/judgment call, or did the AI make it for me by default?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you can't check every box, the test isn't ready — regardless of who wrote it.&lt;/p&gt;




&lt;h2&gt;
  
  
  How This Skill Evolves
&lt;/h2&gt;

&lt;p&gt;AI-assisted testing isn't a single skill — it's a maturity curve most teams move through:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Manual test design&lt;/strong&gt; → &lt;strong&gt;AI-assisted test design&lt;/strong&gt; → &lt;strong&gt;AI-generated automation&lt;/strong&gt; → &lt;strong&gt;AI-reviewed automation&lt;/strong&gt; → &lt;strong&gt;Agentic testing&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Manual test design&lt;/strong&gt; — humans decide what and how, no AI involved.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI-assisted test design&lt;/strong&gt; — AI helps brainstorm cases and gaps; humans still decide.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI-generated automation&lt;/strong&gt; — AI writes the code for well-scoped, well-understood cases.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AI-reviewed automation&lt;/strong&gt; — AI flags weak assertions, duplication, and flakiness risk across the suite.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agentic testing&lt;/strong&gt; — AI systems that can plan, execute, and adapt test runs with human-defined guardrails.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Most teams today are somewhere between stage 2 and 3. Stage 5 gets a lot of hype, but it doesn't remove the need for stage 1's judgment — it just moves that judgment further upstream, into how you define the guardrails. The dimensions in this framework don't disappear as the tooling gets more capable. They become more important, because the cost of an unexamined mistake scales with how much autonomy you hand over.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Staff-Level Takeaway
&lt;/h2&gt;

&lt;p&gt;AI didn't change what makes a good test. It changed how fast a bad decision can get typed into code. The engineers who get real leverage from AI in QA aren't the ones prompting the most — they're the ones who can look at a testing problem for thirty seconds and know exactly which parts belong to a human and which parts belong to a machine. That judgment call is now the actual skill. The code was never the hard part.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;So here's the discussion I want to have:&lt;/strong&gt; What would you never let AI decide when writing your tests? Drop it in the comments — I'll respond to the interesting ones.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>automation</category>
      <category>qa</category>
    </item>
    <item>
      <title>The AI-in-QA Decision Framework, With Real Prompts</title>
      <dc:creator>Pranta Kundu</dc:creator>
      <pubDate>Thu, 10 Sep 2026 08:56:19 +0000</pubDate>
      <link>https://dev.to/prantakunduqa/the-ai-in-qa-decision-framework-with-real-prompts-16kh</link>
      <guid>https://dev.to/prantakunduqa/the-ai-in-qa-decision-framework-with-real-prompts-16kh</guid>
      <description>&lt;p&gt;&lt;strong&gt;&lt;em&gt;A Staff SDET's field guide to knowing when to use AI in testing — and the exact prompts to use when you do.&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Most "AI for QA" content shows you a shiny demo: paste a requirement, get 20 test cases, applause.&lt;/p&gt;

&lt;p&gt;Nobody shows you the moment &lt;em&gt;after&lt;/em&gt; that — when the AI hallucinates a login flow that doesn't exist, or generates a Playwright test that passes locally and fails in CI because it never accounted for a loading spinner.&lt;/p&gt;

&lt;p&gt;I wrote an earlier article introducing a 7-question framework for deciding &lt;strong&gt;whether&lt;/strong&gt; AI should touch a QA problem in the first place. This is the follow-up I promised: every question turned into a real, copy-paste-ready prompt, with the engineering thinking behind it — the part most "AI + testing" content skips entirely.&lt;/p&gt;

&lt;p&gt;If you're a junior QA engineer wondering how to actually use AI on Monday morning without breaking your pipeline, this is for you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why This Framework Exists
&lt;/h2&gt;

&lt;p&gt;AI didn't remove the hardest part of QA. It just moved it.&lt;/p&gt;

&lt;p&gt;The hard part used to be: &lt;em&gt;"How do I write this test?"&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The hard part now is: &lt;em&gt;"Should AI write this test, and how do I know if it's right?"&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A junior engineer asks: &lt;strong&gt;"Can AI do this?"&lt;/strong&gt; A Staff SDET asks: &lt;strong&gt;"Should AI do this, what risk does it introduce, how will I validate it, and what happens when the AI is wrong?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That gap — between &lt;em&gt;can&lt;/em&gt; and &lt;em&gt;should&lt;/em&gt; — is the entire framework.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 7-Question Decision Framework (Quick View)
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;1. What does the AI technology/tool actually do?&lt;/code&gt;&lt;br&gt;
&lt;code&gt;2. What QA problem does it solve?&lt;/code&gt;&lt;br&gt;
&lt;code&gt;3. Is AI actually necessary, or would deterministic automation be better?&lt;/code&gt;&lt;br&gt;
&lt;code&gt;4. What AI capability is being used?&lt;/code&gt;&lt;br&gt;
&lt;code&gt;5. What tools/integration are required?&lt;/code&gt;&lt;br&gt;
&lt;code&gt;6. How do I validate the AI's output?&lt;/code&gt;&lt;br&gt;
&lt;code&gt;7. What are the risks, limitations, and human responsibilities?&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Below, each question gets a real scenario, a real prompt, and a real validation step. No generic "generate test cases for login" prompts here.&lt;/p&gt;
&lt;h2&gt;
  
  
  Question 1 — What does the AI technology/tool actually do?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why an SDET should ask it:&lt;/strong&gt; You cannot evaluate output you don't understand the mechanics of. An LLM predicts plausible text; it does not "know" your application. If you don't internalize that, you will trust output you shouldn't.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scenario:&lt;/strong&gt; A junior engineer wants to use an LLM to generate test cases for a new "Apply Coupon" feature from a Jira ticket.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI tool/use case:&lt;/strong&gt; General-purpose LLM (Claude, GPT-class model) used for natural-language requirement analysis.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are a Senior SDET reviewing a requirement before test design.

Requirement:
"Users can apply one coupon code per order at checkout. Invalid or
expired codes should show an error. Coupon discount should apply to
subtotal before tax."

Do the following, in order:
1. List all assumptions you are making about this requirement.
2. List any ambiguities or missing information a developer/QA should
   clarify before test design begins.
3. Do NOT generate test cases yet. Only output assumptions and gaps.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Example input/context:&lt;/strong&gt; One Jira ticket, no linked designs, no acceptance criteria attached.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Expected AI output type:&lt;/strong&gt; A list like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Assumption: "one coupon per order" means additional codes are silently rejected, not queued.&lt;/li&gt;
&lt;li&gt;Gap: What happens if the coupon is valid but the cart becomes empty before submit?&lt;/li&gt;
&lt;li&gt;Gap: Is there a limit on coupon value relative to order value (can discount exceed subtotal)?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;How to verify:&lt;/strong&gt; Cross-check every assumption against the actual ticket, designs, and product owner — not against your memory of "how coupons usually work." Treat the list as a starting point for a 10-minute clarification conversation, not a finished artifact.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When NOT to use AI here:&lt;/strong&gt; If the requirement is trivial and unambiguous (e.g., "field must be required"), skip this — you're adding process overhead for zero risk reduction.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key lesson:&lt;/strong&gt; AI is strongest at surfacing ambiguity, not at resolving it. Use it to ask better questions, not to skip asking them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Question 2 — What QA problem does it solve?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why an SDET should ask it:&lt;/strong&gt; "We should use AI" is not a QA problem. "Our regression suite takes 40 minutes to triage after each nightly run" is. Tie every AI use case to a named, measurable pain point.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scenario:&lt;/strong&gt; Your team's CI pipeline produces 200+ automated test results nightly, and someone manually scans logs every morning to find real failures vs. flaky noise.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI tool/use case:&lt;/strong&gt; LLM-based log summarization and clustering, either via a script feeding failure logs to an LLM API, or a CI-integrated AI log tool.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are a CI triage assistant for a test automation pipeline.

Below are failure logs from tonight's regression run. For each failure:
1. Extract the failing test name.
2. Classify the failure type: [Assertion Failure, Timeout,
   Element Not Found, Network/API Error, Environment Issue, Unknown].
3. Group failures that share the same likely root cause.
4. Flag any failure pattern that resembles known flaky behavior
   (e.g., timing-related, intermittent selector issues).
5. Output a ranked list: which failures need human investigation
   FIRST, and why.

Logs:
&amp;lt;paste raw CI failure logs here&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Example input/context:&lt;/strong&gt; Raw stack traces and console logs from a Jenkins/GitHub Actions run, 15–30 failed tests.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Expected AI output type:&lt;/strong&gt; A clustered table — e.g., "12 failures share a TimeoutError on #checkout-button, likely a shared root cause (staging environment slow to load) — investigate first."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to verify:&lt;/strong&gt; Spot-check 2–3 clusters manually against the raw logs. Confirm the "likely root cause" is a hypothesis, not a diagnosis — the AI is pattern-matching text, not executing your app.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When NOT to use AI here:&lt;/strong&gt; If your suite only has 5–10 tests, manual triage is faster and more reliable than setting up an AI pipeline for it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key lesson:&lt;/strong&gt; AI earns its place on volume and repetition, not novelty. Triage at scale is a textbook fit; triage of five tests is not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Question 3 — Is AI actually necessary, or would deterministic automation be better?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why an SDET should ask it:&lt;/strong&gt; This is the question most engineers skip, and it's the one that separates a Staff-level thinker from someone chasing a trend. Deterministic code is cheaper, faster, and 100% reproducible for problems that don't require language understanding or judgment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scenario:&lt;/strong&gt; You need to verify that an API response always returns &lt;code&gt;status: "active"&lt;/code&gt; for a newly created user.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why AI is the WRONG tool here:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;❌ Don't do this:&lt;br&gt;
"Ask an LLM to check if the API response looks correct."&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;An LLM introduces non-determinism, latency, and cost into a check that a one-line assertion solves perfectly:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;// ✅ Deterministic automation - correct choice&lt;br&gt;
expect(response.body.status).toBe("active");&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When AI legitimately enters this scenario:&lt;/strong&gt; Generating the initial draft of the test scaffolding, or generating edge-case ideas for the response schema (e.g., "what other states might a user API expose that we haven't tested?") — then a human converts those into deterministic assertions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key lesson (say this out loud in code review):&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"If the correct answer is a fixed rule, write a rule. If the correct answer requires judgment or language understanding, that's where AI earns its seat."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is the single most important discipline in the entire framework. AI-assisted testing should never replace an &lt;code&gt;expect()&lt;/code&gt; statement that can be written in five seconds.&lt;/p&gt;
&lt;h2&gt;
  
  
  Question 4 — What AI capability is being used?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why an SDET should ask it:&lt;/strong&gt; "AI" is not one thing. Code generation, natural language reasoning, visual/image analysis, and autonomous multi-step agents have different failure modes. Naming the capability tells you what kind of validation you need.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scenario:&lt;/strong&gt; You want AI to generate a Playwright test for a multi-step checkout flow from a manual test case.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI capability being used:&lt;/strong&gt; Code generation grounded in natural-language instructions (not visual testing, not agentic browsing).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;You&lt;/span&gt; &lt;span class="nx"&gt;are&lt;/span&gt; &lt;span class="nx"&gt;a&lt;/span&gt; &lt;span class="nx"&gt;Senior&lt;/span&gt; &lt;span class="nx"&gt;SDET&lt;/span&gt; &lt;span class="nx"&gt;writing&lt;/span&gt; &lt;span class="nx"&gt;a&lt;/span&gt; &lt;span class="nx"&gt;Playwright&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="nx"&gt;TypeScript&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;

&lt;span class="nx"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;Framework&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Playwright&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;TypeScript&lt;/span&gt;
&lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;Page&lt;/span&gt; &lt;span class="nb"&gt;Object&lt;/span&gt; &lt;span class="nx"&gt;Model&lt;/span&gt; &lt;span class="k"&gt;is&lt;/span&gt; &lt;span class="nx"&gt;used&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt; &lt;span class="nx"&gt;objects&lt;/span&gt; &lt;span class="nx"&gt;live&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="sr"&gt;/page&lt;/span&gt;&lt;span class="err"&gt;s
&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;Existing&lt;/span&gt; &lt;span class="nx"&gt;selector&lt;/span&gt; &lt;span class="nx"&gt;convention&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;testid&lt;/span&gt; &lt;span class="nx"&gt;attributes&lt;/span&gt; &lt;span class="nx"&gt;only&lt;/span&gt;
&lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;Test&lt;/span&gt; &lt;span class="nx"&gt;must&lt;/span&gt; &lt;span class="nx"&gt;run&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="nx"&gt;CI&lt;/span&gt; &lt;span class="nx"&gt;headless&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;must&lt;/span&gt; &lt;span class="nx"&gt;not&lt;/span&gt; &lt;span class="nx"&gt;use&lt;/span&gt; &lt;span class="nx"&gt;hardcoded&lt;/span&gt; &lt;span class="nf"&gt;waits &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;no&lt;/span&gt;
  &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;waitForTimeout&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nx"&gt;Manual&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt; &lt;span class="k"&gt;case&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;1. Go to checkout page with 1 item in cart.
 2. Apply coupon code 'SAVE10'.
 3. Verify discounted total reflects 10% off subtotal (not total
    with tax).
 4. Click 'Place Order'.
 5. Verify order confirmation page shows correct order ID format.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;

&lt;span class="nx"&gt;Write&lt;/span&gt; &lt;span class="nx"&gt;the&lt;/span&gt; &lt;span class="nx"&gt;Playwright&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt; &lt;span class="nx"&gt;Requirements&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;Use&lt;/span&gt; &lt;span class="nx"&gt;Page&lt;/span&gt; &lt;span class="nb"&gt;Object&lt;/span&gt; &lt;span class="nx"&gt;Model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;assume&lt;/span&gt; &lt;span class="nx"&gt;a&lt;/span&gt; &lt;span class="nx"&gt;CheckoutPage&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;exists&lt;/span&gt; &lt;span class="kd"&gt;with&lt;/span&gt;
  &lt;span class="nx"&gt;methods&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;applyCoupon&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;code&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nf"&gt;getDiscountedTotal&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="nf"&gt;placeOrder&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;Use&lt;/span&gt; &lt;span class="nx"&gt;explicit&lt;/span&gt; &lt;span class="nx"&gt;waits&lt;/span&gt; &lt;span class="nx"&gt;tied&lt;/span&gt; &lt;span class="nx"&gt;to&lt;/span&gt; &lt;span class="nx"&gt;network&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="nx"&gt;DOM&lt;/span&gt; &lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;never&lt;/span&gt; &lt;span class="nx"&gt;fixed&lt;/span&gt; &lt;span class="nx"&gt;timeouts&lt;/span&gt;
&lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;Add&lt;/span&gt; &lt;span class="nx"&gt;one&lt;/span&gt; &lt;span class="nx"&gt;negative&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;path&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;invalid&lt;/span&gt; &lt;span class="nx"&gt;coupon&lt;/span&gt; &lt;span class="nx"&gt;code&lt;/span&gt;
&lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;Call&lt;/span&gt; &lt;span class="nx"&gt;out&lt;/span&gt; &lt;span class="kr"&gt;any&lt;/span&gt; &lt;span class="nx"&gt;missing&lt;/span&gt; &lt;span class="nx"&gt;selectors&lt;/span&gt; &lt;span class="nx"&gt;or&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt; &lt;span class="nx"&gt;object&lt;/span&gt; &lt;span class="nx"&gt;methods&lt;/span&gt; &lt;span class="nx"&gt;you&lt;/span&gt; &lt;span class="nx"&gt;assumed&lt;/span&gt;
  &lt;span class="nx"&gt;exist&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;at&lt;/span&gt; &lt;span class="nx"&gt;the&lt;/span&gt; &lt;span class="nx"&gt;end&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nx"&gt;a&lt;/span&gt; &lt;span class="nx"&gt;comment&lt;/span&gt; &lt;span class="nx"&gt;block&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Example input/context:&lt;/strong&gt; One manual test case + a short description of existing project conventions (this context matters enormously — a prompt without your project's real conventions produces unusable code).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Expected AI output type:&lt;/strong&gt; A &lt;code&gt;.spec.ts&lt;/code&gt; file using your POM structure, plus a clearly labeled comment block listing assumed methods/selectors that don't exist yet — this is the AI being honest about its own guesses, which you explicitly asked it to do.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to verify:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Run it locally before it ever touches CI.&lt;/li&gt;
&lt;li&gt;Check every assumed selector/method against the real codebase.&lt;/li&gt;
&lt;li&gt;Confirm the "no hardcoded waits" instruction was actually followed — AI models frequently insert &lt;code&gt;waitForTimeout&lt;/code&gt; even when told not to.&lt;/li&gt;
&lt;li&gt;Review assertions for false positives (e.g., checking element exists instead of checking element value).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;When NOT to use AI here:&lt;/strong&gt; For extremely simple tests (a single field validation), writing it yourself is faster than prompting, reviewing, and correcting AI output.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key lesson:&lt;/strong&gt; The quality of AI-generated test code is a direct function of how much real project context you give it. A prompt with no conventions produces demo-quality code, not production-quality code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Question 5 — What tools/integration are required?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why an SDET should ask it:&lt;/strong&gt; A good prompt is useless if it doesn't fit your actual pipeline. Before adopting an AI use case, map exactly where it plugs into your stack — IDE, CI runner, test management tool, or a standalone script.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scenario:&lt;/strong&gt; You want AI-generated API test scenarios to plug directly into your existing Postman/Newman or REST-assured suite, not live in a separate doc nobody maintains.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI tool/use case:&lt;/strong&gt; LLM prompt run through a script (or IDE plugin) that outputs structured JSON, consumed directly by your test generator.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;You&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;are&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;generating&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;API&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;test&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;scenarios&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;for&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;an&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;SDET's&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;automated&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;suite.&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;Endpoint:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;POST&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;/api/v&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="err"&gt;/orders&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;Request&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;schema:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"userId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"string (UUID)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"items"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"sku"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"string"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"quantity"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"integer &amp;gt; 0"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"couponCode"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"string, optional"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;Generate&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;test&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;scenarios&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;covering:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;-&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Valid&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;requests&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;(happy&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;path)&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;-&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Boundary&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;values&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;(quantity&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="err"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;quantity&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;max&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;int,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;empty&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;items&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;array)&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;-&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Invalid&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;types&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;(quantity&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;as&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;string,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;missing&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;userId)&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;-&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Security-relevant&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;edge&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;cases&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;(SQL-injection-style&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;strings&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;in&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;sku,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="err"&gt;oversized&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;payloads)&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;-&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Business-rule&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;edge&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;cases&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;(couponCode&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;present&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;but&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;user&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;has&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;no&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;cart)&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;Output&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;as&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;a&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;JSON&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;array,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;each&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;object&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;with&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;fields:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"scenario_name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"request_body"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{},&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"expected_status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"expected_behavior"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"risk_level"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"low|medium|high"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;Only&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;output&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;valid&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;JSON.&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;No&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;explanation&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;text.&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Integration point:&lt;/strong&gt; This structured JSON output is designed to feed directly into a script that converts each object into a REST-assured or Newman test case — this is what makes it "integration-ready" instead of a wall of text you retype by hand.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to verify:&lt;/strong&gt; Validate the JSON schema programmatically before trusting it (a malformed field breaks your generator silently). Manually review the &lt;code&gt;risk_level&lt;/code&gt; field — AI risk-ranking is a starting opinion, not a security audit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When NOT to use AI here:&lt;/strong&gt; If your API surface is small and stable (under ~10 endpoints, rarely changing), manually maintained scenarios are easier to keep accurate than an AI-regeneration pipeline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key lesson:&lt;/strong&gt; The real ROI of AI in QA comes from integration, not novelty. A brilliant prompt that produces output you copy-paste by hand doesn't scale — structure your prompts to feed your existing tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  Question 6 — How do I validate the AI's output?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why an SDET should ask it:&lt;/strong&gt; This is the question that determines whether you're doing "AI-assisted testing" or "AI-assisted incident creation." Validation is not optional, and it is not the same as "it compiled."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scenario:&lt;/strong&gt; A test that was passing for three months starts failing intermittently — 1 in 8 runs — and you want AI to help investigate before you spend an afternoon on it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI tool/use case:&lt;/strong&gt; LLM-based root-cause reasoning over logs, diffs, and historical run data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are helping an SDET investigate a flaky test.

Test: "should display updated cart total after quantity change"
Framework: Cypress
Failure rate: ~1 in 8 CI runs, always in headless CI, never
reproduces locally.

Here is the failing run's log/error output:
&amp;lt;paste error + stack trace&amp;gt;

Here is the relevant test code:
&amp;lt;paste test code&amp;gt;

Here is the relevant application code being tested (cart update
handler):
&amp;lt;paste app code&amp;gt;

Analyze and respond in this structure:
1. Most likely root cause category (race condition, animation/timing,
   test isolation issue, environment resource contention, incorrect
   assertion, backend non-determinism) - with confidence level
   (low/medium/high).
2. Specific evidence from the logs/code that supports this.
3. Two alternative hypotheses, ranked by likelihood.
4. A minimal experiment I could run to confirm or rule out the top
   hypothesis (not a rewrite - an experiment).
5. What you are NOT confident about, explicitly.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Example input/context:&lt;/strong&gt; Real logs, real test code, real handler code — vague or partial context produces vague or partial (and confidently wrong) diagnoses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Expected AI output type:&lt;/strong&gt; A ranked hypothesis (e.g., "high confidence: race condition — the assertion checks total before the debounced API call resolves") plus a small, falsifiable experiment ("add a network-idle wait tied to the specific PATCH call and rerun 20x in CI").&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to verify — this is the core discipline:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Run the suggested experiment. Do not accept the hypothesis until it's been tested against the real, flaky test at least 10–20 times.&lt;/li&gt;
&lt;li&gt;If the "fix" makes the test pass once, that is not validation — flaky tests pass "by luck" too. Require a meaningfully higher pass rate over repeated runs.&lt;/li&gt;
&lt;li&gt;Never merge an AI-suggested fix into a shared test without your own understanding of why it works. If you can't explain it in code review, don't ship it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;When NOT to use AI here:&lt;/strong&gt; If the flake is already well-understood by your team (e.g., "we know this suite has animation timing issues on Fridays"), skip the AI investigation and apply the known fix.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key lesson:&lt;/strong&gt; AI-generated root cause analysis is a hypothesis engine, not a verdict. Validation is the SDET's job, every time, with no exceptions — this is the line between engineering and guessing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Question 7 — What are the risks, limitations, and human responsibilities?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why an SDET should ask it:&lt;/strong&gt; Every AI use case has a failure mode. Naming it in advance is how you prevent it from becoming an incident.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scenario:&lt;/strong&gt; Using AI to generate synthetic test data for a staging environment load test.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI tool/use case:&lt;/strong&gt; LLM-based synthetic data generation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight csvs"&gt;&lt;code&gt;&lt;span class="k"&gt;Generate&lt;/span&gt; &lt;span class="mf"&gt;50&lt;/span&gt; &lt;span class="k"&gt;synthetic&lt;/span&gt; &lt;span class="k"&gt;user&lt;/span&gt; &lt;span class="k"&gt;records&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="k"&gt;load&lt;/span&gt;&lt;span class="err"&gt;-&lt;/span&gt;&lt;span class="k"&gt;testing&lt;/span&gt; &lt;span class="k"&gt;a&lt;/span&gt; &lt;span class="k"&gt;signup&lt;/span&gt; &lt;span class="k"&gt;API&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;

&lt;span class="nv"&gt;Constraints:&lt;/span&gt;
&lt;span class="err"&gt;-&lt;/span&gt; &lt;span class="nv"&gt;Fields:&lt;/span&gt; &lt;span class="k"&gt;fullName&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;email&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;dateOfBirth&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;must&lt;/span&gt; &lt;span class="k"&gt;include&lt;/span&gt; &lt;span class="k"&gt;users&lt;/span&gt; &lt;span class="k"&gt;aged&lt;/span&gt;
  &lt;span class="mf"&gt;17&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;18&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;and&lt;/span&gt; &lt;span class="mf"&gt;100&lt;/span&gt;&lt;span class="err"&gt;+&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="k"&gt;boundary&lt;/span&gt; &lt;span class="k"&gt;testing&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="k"&gt;country&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;mix&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="k"&gt;at&lt;/span&gt; &lt;span class="k"&gt;least&lt;/span&gt;
  &lt;span class="mf"&gt;8&lt;/span&gt; &lt;span class="k"&gt;countries&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;including&lt;/span&gt; &lt;span class="k"&gt;ones&lt;/span&gt; &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="k"&gt;non&lt;/span&gt;&lt;span class="err"&gt;-&lt;/span&gt;&lt;span class="k"&gt;Latin&lt;/span&gt; &lt;span class="k"&gt;characters&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="k"&gt;password&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;must&lt;/span&gt; &lt;span class="k"&gt;include&lt;/span&gt; &lt;span class="k"&gt;realistic&lt;/span&gt; &lt;span class="k"&gt;weak&lt;/span&gt; &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="k"&gt;strong&lt;/span&gt; &lt;span class="k"&gt;examples&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="err"&gt;-&lt;/span&gt; &lt;span class="k"&gt;Do&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;use&lt;/span&gt; &lt;span class="k"&gt;real&lt;/span&gt; &lt;span class="k"&gt;people&lt;/span&gt;&lt;span class="err"&gt;'&lt;/span&gt;&lt;span class="k"&gt;s&lt;/span&gt; &lt;span class="k"&gt;names&lt;/span&gt; &lt;span class="k"&gt;or&lt;/span&gt; &lt;span class="k"&gt;real&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;resolvable&lt;/span&gt; &lt;span class="k"&gt;email&lt;/span&gt; &lt;span class="k"&gt;domains&lt;/span&gt; &lt;span class="err"&gt;-&lt;/span&gt;
  &lt;span class="k"&gt;use&lt;/span&gt; &lt;span class="k"&gt;example&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="k"&gt;com&lt;/span&gt; &lt;span class="k"&gt;or&lt;/span&gt; &lt;span class="k"&gt;example&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="k"&gt;org&lt;/span&gt; &lt;span class="k"&gt;only&lt;/span&gt;
&lt;span class="err"&gt;-&lt;/span&gt; &lt;span class="k"&gt;Include&lt;/span&gt; &lt;span class="mf"&gt;5&lt;/span&gt; &lt;span class="k"&gt;intentionally&lt;/span&gt; &lt;span class="k"&gt;malformed&lt;/span&gt; &lt;span class="k"&gt;records&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;invalid&lt;/span&gt; &lt;span class="k"&gt;email&lt;/span&gt; &lt;span class="k"&gt;format&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;future&lt;/span&gt; &lt;span class="k"&gt;date&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="k"&gt;birth&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;empty&lt;/span&gt; &lt;span class="k"&gt;required&lt;/span&gt; &lt;span class="k"&gt;field&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;explicitly&lt;/span&gt; &lt;span class="k"&gt;labeled&lt;/span&gt;
  &lt;span class="s2"&gt;"INVALID - for negative testing"&lt;/span&gt;
&lt;span class="err"&gt;-&lt;/span&gt; &lt;span class="k"&gt;Output&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="k"&gt;CSV&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Named risks for this use case:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fksta5gvaxjn5nq5iramm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fksta5gvaxjn5nq5iramm.png" alt=" " width="640" height="271"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to verify:&lt;/strong&gt; Scan the CSV for real-looking PII patterns, confirm invalid records are actually invalid in the ways intended, and confirm the file never gets referenced outside the test/staging scope.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When NOT to use AI here:&lt;/strong&gt; Regulated domains (healthcare, finance) where synthetic data must meet specific compliance patterns — use a purpose-built synthetic data tool with compliance guarantees, not a general LLM prompt.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key lesson:&lt;/strong&gt; The risk isn't that AI gets something wrong. The risk is a team building a habit of not checking because AI usually gets it right. &lt;em&gt;Usually&lt;/em&gt; is where incidents live.&lt;/p&gt;

&lt;h2&gt;
  
  
  AI Prompt ≠ AI Decision
&lt;/h2&gt;

&lt;p&gt;This is the sentence I'd want every junior SDET to tattoo on their monitor:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;A good SDET does not blindly execute whatever AI recommends.&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The real workflow — the one that separates "using AI" from "engineering with AI" — looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[Requirement] --&amp;gt; B["Problem: What are we actually trying to verify or solve?"]
    B --&amp;gt; C["AI Suitability Check — Would deterministic automation be better?"]
    C --&amp;gt; D["Prompt: structured, project-aware, explicit about constraints"]
    D --&amp;gt; E["AI Output (draft — never treated as final)"]
    E --&amp;gt; F["Validation: run it, test it, cross-check assumptions, check risk"]
    F --&amp;gt; G["Human Decision: accept / modify / reject — with a reason"]
    G --&amp;gt; H["Automation / CI Integration — only validated, understood work ships"]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Every step in this chain is a checkpoint. Skipping &lt;strong&gt;AI Suitability Check&lt;/strong&gt; gets you AI-generated tests for problems a one-line assertion would've solved. Skipping &lt;strong&gt;Validation&lt;/strong&gt; gets you flaky CI. Skipping &lt;strong&gt;Human Decision&lt;/strong&gt; gets you tests nobody on the team can explain in six months.&lt;/p&gt;

&lt;p&gt;One-liner for your team's Slack channel:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;"AI can write the test. It can't own the test."&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fom69h473ad5qx27yv7rg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fom69h473ad5qx27yv7rg.png" alt=" " width="640" height="320"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Staff-level read of this table:&lt;/strong&gt; notice reliability goes down and required human involvement goes up as autonomy increases. That's not a flaw in AI tooling — that's the correct engineering relationship between autonomy and oversight. Anyone selling you "fully autonomous QA, zero human review" is selling you an incident.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 5-Minute AI Adoption Checklist
&lt;/h2&gt;

&lt;p&gt;Before you introduce any AI tool into a QA workflow, run through this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Named problem&lt;/strong&gt; — Can I state the QA problem in one sentence without mentioning AI?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deterministic check&lt;/strong&gt; — Have I confirmed a fixed rule/script can't solve this more reliably?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Capability match&lt;/strong&gt; — Do I know exactly which AI capability I'm using (text generation, code generation, reasoning, agentic action)?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context supplied&lt;/strong&gt; — Does my prompt include real project conventions, not generic instructions?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Validation plan&lt;/strong&gt; — Do I know before I run the prompt how I will verify the output?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Blast radius&lt;/strong&gt; — If this output is wrong and I miss it, what's the worst that happens (flaky CI vs. a shipped bug vs. leaked data)?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Owner&lt;/strong&gt; — Is there a named human who reviews and approves before this reaches CI/production?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reversibility&lt;/strong&gt; — Can I easily disable/roll back this AI-generated artifact if it causes problems?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you can't check every box in five minutes, you're not ready to automate this with AI yet — and that's a useful, not embarrassing, conclusion.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Would Explain This to a Junior SDET
&lt;/h2&gt;

&lt;p&gt;Imagine you hire a very fast, very well-read intern. They've read almost every testing blog, every framework doc, every Stack Overflow answer — but they've never actually run your app, never seen your production incidents, and they will confidently guess when they don't know something instead of saying "I don't know."&lt;/p&gt;

&lt;p&gt;Would you let that intern:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Draft your test cases? &lt;strong&gt;Yes&lt;/strong&gt; — review before you use them.&lt;/li&gt;
&lt;li&gt;Write your first-pass Playwright script? &lt;strong&gt;Yes&lt;/strong&gt; — review every selector.&lt;/li&gt;
&lt;li&gt;Decide alone that a flaky test is now fixed? &lt;strong&gt;No&lt;/strong&gt; — make them prove it, ten times, in CI.&lt;/li&gt;
&lt;li&gt;Push straight to your production test suite unsupervised? &lt;strong&gt;Never.&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That intern is exactly what AI is in your QA workflow. Brilliant draft-writer. Terrible unsupervised decision-maker. Treat it exactly that way and you'll get real value without real risk.&lt;/p&gt;

&lt;h2&gt;
  
  
  How a Staff SDET Thinks Differently
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F13tyruswx7aizr85hyif.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F13tyruswx7aizr85hyif.png" alt=" " width="640" height="310"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That last row is the whole game. AI doesn't get blamed for a bad test in a postmortem. The engineer who merged it does. Staff-level thinking means you never let a prompt make a decision you haven't personally signed off on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where This Is Heading (Without the Hype)
&lt;/h2&gt;

&lt;p&gt;None of this means testers get replaced. It means the skill set keeps layering:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A["AI-assisted SDET — uses AI as a tool inside existing manual/automated workflows"] --&amp;gt; B["AI-augmented SDET — AI embedded across the workflow: test design, debugging, data, triage - with strong validation habits"]
    B --&amp;gt; C["Agentic Test Engineer — designs and supervises AI agents that execute multi-step testing tasks with defined guardrails"]
    C --&amp;gt; D["AI QA Agent Engineer — builds, evaluates, and governs the AI systems/agents QA teams rely on"]
    D --&amp;gt; E["AI Quality Engineering — quality engineering discipline expands to include AI system quality itself: evaluating models, prompts, and agents as first-class testable systems"]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;To be clear: these are evolving skill directions, not standardized job titles you'll see on every job board tomorrow. Nobody should treat this as a fixed career ladder — treat it as a map of where the skills are heading, and start building the ones nearest to where you already are.&lt;/p&gt;

&lt;p&gt;The testers who do well here aren't the ones who trust AI the most. They're the ones who validate it the best.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try This Today
&lt;/h2&gt;

&lt;p&gt;Pick one task from this article — requirement ambiguity analysis, log triage, or flaky test investigation — and run the real prompt on a real problem you have this week. Don't use it on a toy example. Use it on something in your actual backlog, then apply the validation step before you trust a single line of it.&lt;/p&gt;

&lt;p&gt;That's the whole framework in practice: not "AI or no AI," but &lt;strong&gt;problem → suitability → prompt → validation → decision&lt;/strong&gt;, every single time.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Which QA task would you trust AI with first — and which one would you never hand over to AI?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Drop it in the comments. I'll respond with how I'd validate (or reject) it.&lt;/p&gt;

</description>
      <category>qualityengineering</category>
      <category>testautomation</category>
      <category>testing</category>
      <category>softwaretesting</category>
    </item>
  </channel>
</rss>
