DEV Community

Cover image for Use case: Creating an agentic workflow from frontend to test automation
Mellina Yonashiro
Mellina Yonashiro

Posted on

Use case: Creating an agentic workflow from frontend to test automation

Introduction

Every time our Frontend (FE) changed a UI component, a little nightmare started: tests broke, the Test Automation (TA) team was called and what seemed to be endless back-and-forth happened.

I saw an opportunity to reduce this friction by introducing a headless agent into our workflow - one that picks up the broken tests and tries to fix them.

Developing automations like this is a hard process though - designing a meaningful flow, dealing with cross-dependencies and accesses that might turn into dead-ends, and other unknown issues along the way. In this article, I'll walk you through the problem, the design decisions and trade-offs, and how it was implemented end-to-end.

The Problem

When working in frontend, we have some CI steps set up. One of them is to run test automation. More often than not, these tests fail because frontend changed the HTML structure or user flow altogether. When this happens, we have to go through these steps shown in the diagram.

Flow diagram showing a cross-team loop between Frontend and UI Test Automation when UI tests fail. A Frontend pull request triggers tests A, B, and C, which fail. UI Test Automation responds by skipping tests A–C and later creating and merging a PR to address them. Meanwhile, Frontend’s CI passes because the tests are skipped, so the Frontend PR is merged to main. Afterward, UI Test Automation runs again and creates a fix PR; during this period, there’s a warning state where automation “passes” but tests remain skipped on main, until the fixes are merged and tests are stabilized on main.

Summarizing:

  1. FE: makes a breaking change. Tests break.
  2. TA: skip related tests.
  3. FE: CI pipeline now passes and merges the feature branch.
  4. TA: unskip the test, fix them, open a pull request and merge.

This workflow triggers for every feature branch opened by FE. The steps have to be done in sequence, and the flow involves at least four engineers (the ones writing the code and the ones reviewing it), context switching, and cross-team dependency. Also, there is a period in between that we lose test coverage: the tests pass, but only because they're skipped.

For our experiment, we decided to narrow our focus to something specific and easy to fix: when FE changes a UI component, needing a selector change; and only smoke tests would be used as the gate. I’ll give you more context below.

What are selectors in UI testing?

In frontend applications, we implement components that users interact with. These components are accessible - either by their text, input placeholder, role or by non-visible screen-reader aria-* attributes, falling back to data-testid. These identifiers, selectors (or locators, as some libraries call them), are used by testing libraries, such as Testing Library, Playwright or Cypress to conduct automated tests, which simulate user flows.

Whenever a text, label, or whatever is used to identify the testing component is changed, the tests break. You may think this is a rare problem, but it can happen in some situations:

  • If the frontend is using a design system and it’s moving to another one, this will likely break selectors. This will happen a lot during migrations.
  • If the text is changed and sometimes its position, this will likely break as well.
  • Any HTML change can break the selector.

What are smoke tests?

Smoke tests are a small set of end-to-end tests that cover only the most critical flows of an application. The idea is not to test everything, but to quickly answer one question: is the app still working? Because they are fewer and faster than the full suite, they are cheaper to run on every feature branch - and for this experiment, that also meant fewer (and less flaky) failures for the agent to deal with. If you want to go deeper, this article is a good starting point.

Constraints

This use case is not a one-size-fits-all solution, but it can give you some ideas. We had this specific setup:

  • Two separate codebases: some projects keep the frontend and test automation suite in the same codebase. This is not the case for our project - there is one repository for the frontend and another for test automation.
  • FE can’t be built locally in TA CI: this means that the TA pipeline can’t test on the same conditions as FE does. FE runs on a local build and triggers smoke tests there. On the other hand, TA can’t - it can only run against some deployed QA environments. Bottom line: because of this, the same tests that fail in the FE pipeline can pass in TA (and vice-versa). This is why one of the agent's outcomes is Not reproduced (more on that below).
  • Different CI/CD tools per repository: FE runs its CI pipeline in Drone and its CD (deployments) in GitHub Actions (GA). TA only has a CI pipeline in Drone - it doesn't deploy anything, just some reports, which are managed by Drone as well.

Proposed Solution

Flowchart showing cross-team loop - on the Frontend side, a pull request is created, tests A/B/C fail, and Drone sends a  raw `repository_dispatch` endraw  event to notify UI Test Automation. On the UI Test Automation side, the notification triggers a process where an automated agent (shown inside a dashed box labeled “automated”) updates broken tests, checks a decision point “Tests fixed?”, and if fixed opens a PR for review, after which a human reviews and merges the PR, feeding the result back so the overall pipeline can proceed. After it, back to Frontend, CI passes leading to the PR being merged.

The overall idea is that whenever there is a change in FE, the CI pipeline runs smoke tests, and if they fail, it would trigger a workflow agent in TA by sending a repository_dispatch event to the TA repo (through the GitHub REST API), which starts a GitHub Actions workflow there. This agent would then try to fix the non-passing tests. If it is successful, it would create a PR for a human to check. If this PR was merged, then the FE CI pipeline would pass.

Some notes to consider here:

  • When TA PR is merged, FE CI is not triggered automatically. A user has to re-trigger the job in Drone. This could be tightened in the future, by adding another task to do it when the branch is merged to main in TA.
  • Also when TA PR is merged to main, it will cause FE main and other branches to fail the tests (because they would not up-to-date yet with the feature branch). This is a known trade-off, which is acceptable, considering we are speeding up the process, at the same time controlling the flow (because a human will need to push the button to merge TA PR into main).

After coming to this high-level design (after a lot of back and forth), I started to think about the details.

Designing the automation flow

To implement this, some questions had to be answered:

  1. Timing: FE can trigger the agent every time tests fail, but should it?
  2. Platform: Where should the agentic workflow run - Drone or GitHub Actions?
  3. Cross-repo/CI communication: How would the connection between FE and TA repositories happen?
  4. Outcomes: What happens if the agent can’t fix the test? What if it can? What if it can’t reproduce (tests pass)?

Timing: when should the agent run?

Considering that every feature branch in FE triggers a CI pipeline and because of cost-efficiency and having more control, especially during the testing of this automation, I chose to trigger the agent under certain conditions. The feature branch changes would need to be deployed into a specific environment (we have more than 4 environments set up, so we can test multiple feature branches at the same time). Let’s call this environment the QA Env.

Having this environment is also needed because of the constraint that TA can only run smoke tests against deployed environments.

Platform: where should the agent run?

There are multiple steps of the automation design, and one of them is to decide where to trigger the agentic flow. As mentioned in the Constraints, we were using both Drone and GA.

Our Drone was becoming crowded already - it was running processes on all branches, including main. And I realised that with GA we would have more control on when and how to trigger - their interface was more friendly for testing and manual triggering. Also, I could create a separated flow only for the agent, not mixing with other processes.

For those reasons, I decided to go with GA for the agentic workflow, therefore the point of connection should be GA for triggering the agent.

Connecting the dots: how do the repositories communicate?

The initiator of this automation is FE when two conditions are met: smoke tests fail AND it is deployed to QA Env. Since FE CI runs in Drone and FE CD runs in GA, these two conditions are checked in different places. I had two options:

  1. Run smoke tests in GA after QA Env deployment, and then send the dispatch event to TA.
  2. In Drone, after smoke tests step, check if QA Env is deployed and then send the dispatch event to TA.

The first option would be more costly - we would need to bring the smoke test step into the GA, and it would duplicate the step in both CI and CD pipelines.

The second option, on the other hand, would be a simple check. That’s because we use GitHub labels to track which environment is being deployed. We already had a flow set up where adding a label to a PR triggers a deployment to that environment.

Given that, we could check the GitHub label right after the Smoke tests step. If both smoke tests failed and the deployment label was set as QA Env, then Drone would send the dispatch event to TA.

Given that CI and CD happen in different environments, one trade-off here is that it is possible that Drone will receive and identify the label, but the GA deployment process can fail, or Drone check can happen before the branch is actually deployed. I accepted these trade-offs, because the worst that could happen is that the tests run against an environment without FE change, triggering a false negative.

For the race condition, as soon as the label is added, usually the CD happens in 5 minutes. The entire CI pipeline usually takes 15 minutes, making it a comfortable pass, since the check would be as the last step of the CI pipeline.

Outcomes: what should the agent report?

The FE created a new feature that broke smoke tests, it was deployed to QA Env, and the dispatch event was sent. Now what? What should the agent do?

The entire goal is for the broken smoke tests to be fixed, by updating the selectors. For that, I took some decisions:

  • We had more than 40 written smoke tests in total. For the purposes of this automation, we don’t need the entire suite to run. Therefore, there should be a payload carrying the information of which smoke tests have failed in the FE side.
  • When attempting to fix the test, I see some outcomes, overall:
    • The agent is able to fix the test.
    • The agent is not able to fix the test.
    • The tests pass.

So, with three main outcomes, plus two edge cases we would have:

  • If tests are fixed, then a PR is created in TA side, and a comment is sent to FE PR with the message of “Already repaired”, along with the link to the TA PR.
  • If tests were already fixed by a prior run, then a comment is sent to FE PR with the message of “Already repaired”, without a link to a TA PR.
  • If tests are not fixed, then no PR is created, but a message is sent back to FE PR with the message of “Possible regression”.
  • If tests pass, then a message is sent back to FE PR with the message of “Not reproduced”.
  • If there is an environment failure (e.g., QA Env is unreachable), the agent aborts and no message is sent back to FE PR - the error is only logged in the GA run.

Although it is possible to see all the logs in GA workflows page, I wanted to make provide as much up front transparency as possible, by ending the flow in FE PR, with a message on what was the outcome of the entire process.

Then, with most of the definitions set up, we can finally look at how the agent does this, and the harnesses around it, in the next sections.

Setting up the prerequisites

As you will see, there are multiple steps to implement the flow end-to-end.

Creating the agent harness

A harness is everything around the model that shapes how the agent works - instruction files, skills, commands, scripts and guardrails - so it behaves predictably without someone guiding it step by step.

The initial state of the codebase was not ready for AI at all. To avoid increasing the blast radius of the project, I decided to just create enough harnesses to implement this POC. Two assets were created:

  • AGENTS.md: with CLAUDE.md pointing to it, which enables uses for both Claude and OpenAI models to read. Normal instructions file that is injected to every agent session (not going too deep into this as there's plenty of documentation on the web).
  • Selector conventions skill: defines the guidelines on how to choose the selector. There is a hierarchy the agent has to follow. For example, the first attempt should be with getByRole, then by getByLabel and so on.

With those, I tested how it performed using my own agent (local Claude Code CLI).

Other commands were created too, but they belong to the automation per se, so they will be specified in the following sections.

Making tests environment-agnostic

Another problem was that the tests were not able to run against QA Env. Data was coupled into a specific environment, as well as some selectors. This isn’t related to AI nor to automation, but it would impact the ability to develop it. Therefore, I had an additional prep work: creating functions that would setup/create and teardown/delete data for every test run and making selectors environment agnostic.

Implementing the flow

A visual diagram of the entire flow is shown below:

This flowchart illustrates an automated test selector repair system with two interconnected workflows. On the frontend side, after a PR is opened, and smoke test fails, the system then checks whether the code is deployed to the QA environment. If so, the system dispatches a

Frontend: dispatching the event

As the flow starts in FE, naturally I had to start there, but the work was slim: I’ve added a new step in drone.yml, after the Smoke test step:

- name: Smoke test
  commands:
# Smoke tests instructions

- name: Dispatch selector repair
  environment:
    GITHUB_APP_TOKEN:
      from_secret: GITHUB_APP_TOKEN
    REPAIR_ENVIRONMENTS: QA_Env
  commands:
    - sh ci/drone/dispatch-selector-repair.sh
  depends_on:
    - Smoke test
  when:
    status:
      - failure
Enter fullscreen mode Exit fullscreen mode

(Note: this is an example piece of code and should not be used as it will not work).

As variables, a GitHub App token was needed to enable cross-repository communication and the REPAIR_ENVIRONMENTS is a mirror list of which environments are enabled. For now, only QA Env is enabled to trigger this automation.

I could add all the instructions in the YAML file, but I had extracted them into ci/drone/dispatch-selector-repair.sh to separate concerns and improve readability. In this script, a few things were done:

  • Check if deployed environment is valid
  • Check if GitHub token is valid
  • Collect the failed tests. Exit if no JSON report of failed tests is found.
  • Fetch commit SHA, PR id number
  • Prepare payload and dispatch frontend-smoke-failed event

Then, TA would receive this event and payload n its end.

Test Automation: running the agent

To receive the event, a GA workflow was created. selector-repair.yml is triggered when the event is received:

repository_dispatch:
  types: [frontend-smoke-failed]
Enter fullscreen mode Exit fullscreen mode

It can also be trigger manually, although this was mostly enabled for testing purposes, when evaluating the performance of the agent.

This workflow:

  1. Runs deterministic checks before calling the agent: validates the payload from the dispatcher (environment, FE PR, SHA and failed specs) with regex and against injection, checks API keys, re-checks the deployed environment and confirms portal access. Failing early here avoids unnecessary agent calls and adds a security layer.
  2. If everything’s ok, then installs OpenAI Codex CLI, for headless agent. The chosen model was GPT-5.4. The goal is for the harness be good enough that we don’t need more capable/high reasoning (and more expensive) models, such as Opus and Astra.
  3. Triggers the agent to run repair, by using codex exec command, which runs the Repair selector command (described below, officially called repair-selectors-ci)
  4. Outputs repair summary and uploads evidence (HTML files and images)
  5. Comments back in FE PR

As you can see, the LLM is only involved in in one specific step. Before and after, deterministic scripts were written, both to reduce AI costs and improve quality. The agent is only triggered if it can actually run the tests in an enabled environment; and after the agent finishes, a script puts together the message body for the PR comment back to FE.

The Repair selector command

On top of the shared harness from the setup section, the CI run adds two commands. Namely:

  • Repair selector command: the command that is triggered by automation itself.
  • Open repair PR command: a separate command that is triggered inside the Repair selector command.

The Repair selector command is where the core agent logic lives. In summary, it does:

1. Reads inputs from environment variables: which environment to test, which FE PR triggered the run, and a list of failed specs.

2. Manages a persistent branch: one branch per frontend PR, reused and appended to across multiple runs rather than recreated each time, as each commit in the same branch triggers a new run. It never rebases or force-pushes, since the commit history works as an audit log. Also, if there is a new commit from FE side, GA cancels the current run via concurrency: cancel-in-progress, so it doesn’t keep working on outdated code. If the commit was already pushed, it stays there.

3. Determines what to test: either the specific failing specs or runs the whole smoke suite itself, in case of a manual run. If a previous repair run already committed a fix for it on the persistent branch, it's categorized as “Already repaired”.

4. Triages each failure into one of three buckets:

  • Rot - the element is still there, just addressed differently (changed role/name/DOM position) → safe to auto-fix.
  • Regression — the element is missing or the underlying data is wrong → escalate to a human, reported as Possible regression.
  • Environment failure — the portal itself was unreachable/broken, so failures aren't meaningful at all → abort triage entirely rather than misreporting every spec as broken. Nothing is sent back to FE; the error is only logged in the GA run.

5. Repairs rot-classified selectors by probing the live page for accessible signals (role, label, text, placeholder, etc.), picking the most precise selector that resolves to exactly one element, and verifying the fix actually turns the test green before keeping it.

6. Runs quality gates — lint/prettier checks, cleanup of any scratch/probe files — before allowing anything to be committed.

7. Opens or updates a draft PR (never merges) - this is the Open repair PR command.

8. Writes a structured JSON summary (repair-summary.json) as its only communication channel back to the calling workflow — listing what was fixed, what was escalated as a regression, what couldn't be reproduced, what a prior run already fixed, and any environment error — with strict rules about which fields must stay empty in which scenarios, so the human-facing report is never misleading.

After that, the payload is sent back to selector-repair.yml to be picked up by the last step of "Comment back on FE PR". This is a deterministic code that executes the feedback to the FE repository as we already discussed in "Outcomes: what should the agent report?" section.

Conclusion

Before this experiment, a single breaking UI change meant skipping tests, merging with gaps in coverage, and coordinating at least four engineers across two teams. With this flow, FE gets feedback directly in its PR, TA reviews a draft fix instead of writing it from scratch, and, when the agent is successful, tests no longer need to sit skipped on main.

The biggest lesson for me was that the agent was the smallest part of the work. Most of the effort went into the pieces around it: deciding when it should run, connecting two repositories and two CI tools, making the tests run against a shared environment, and writing deterministic checks so the LLM is only called when it can actually help.

There's still room to improve and I believe this can serve as a case study for autonomous systems that go beyond using AI in local development. The future is agents working in the cloud without a human driving each step - and for that, we need solid harnesses and guardrails, and to keep observing their output to improve quality, reduce costs and speed things up.

Top comments (0)