DEV Community

Katja
Katja

Posted on

My testing agent remembers what it learned. Mostly.

TL;DR: An open-source agent that reads a Jira ticket, writes test steps, runs them in a real Chrome window, judges the result from screenshots, and reports back to Jira. The new version has a memory: it reuses browser automation code that worked and drops code that broke. Setup is at the end.


A little while ago I wrote about an AI QA agent that tests your Jira tickets in a real browser. You give it a ticket ID and a URL. It reads the bug, writes test steps, drives Chrome through them, judges the result from screenshots, and reports back to Jira. If everything passes, it moves the ticket to Done.

I've updated it since. The big change: it has a memory now.

The old version started from zero on every ticket. Every step meant an LLM writing fresh browser code, even for a click it had already done a hundred times. That works, but it's slow, it costs tokens, and it's a bit like a colleague who forgets everything over the weekend.

The new version remembers automation code that worked and reuses it when a similar step comes up again, even on a different ticket and phrased differently. It also forgets code that broke, so it doesn't keep trusting something broken.

Here's how the updated version works and why I built it this way. If you want to run it yourself, skip to How to try it.

Two agents, one phone call between them

There are two agents.

The main one, the orchestrator, turns a Jira ticket into a list of test steps and coordinates running them. The second one, step_execution, does exactly one thing: it executes a single test step against a live browser and tells you how it went.

The orchestrator calls the sub-agent as a tool, the same way it would call any function. Give it a step, get back a verdict. It never sees a selector, a screenshot comparison or a line of browser code. That separation turned out to matter more than I expected: I can rework the sub-agent without touching the planner.

Here's the rough shape:

Jira ticket ID ──▶ [Main agent]
                     ├─ Start browser session (BaaS)
                     ├─ Log in to the app under test
                     ├─ Fetch Jira issue + comments
                     ├─ LLM: generate ordered test steps
                     └─ Orchestrator LLM
                          │  one tool call per step
                          ▼
                   [Sub-agent: step_execution]
                     ├─ Screenshot + HTML snapshot
                     ├─ LLM: which page is this? check memory
                     ├─ Reuse a stored program OR write a new one
                     ├─ Run the program in the browser
                     ├─ LLM: compare screenshots with expected result
                     └─ Memory: keep good programs, drop bad ones
                          │  verdict per step
                          ▼
                   back to main agent
                     ├─ Broken step? Retry it
                     ├─ Test finished -> Post results as a Jira comment
                     ├─ All passed? Move ticket to Done
                     └─ Stop browser session
Enter fullscreen mode Exit fullscreen mode

The main agent: from ticket to test plan

The input is small: a URL, a Jira ID and a Jira user email.

It starts a real browser. The agent calls our Browser-as-a-Service (BaaS) API to get an isolated browser session, then logs in. BaaS accepts a program string in a tiny automation language, with commands like navigate(...), clickN(...), sendKeysToElement(...) and takeScreenshot(...), sent as JSON to the BaaS runner. What comes back is the result: a sessionID, which travels through the whole run, and a screenshot of the freshly logged-in app.

It reads the whole ticket, comments included. Two calls to the Jira REST API, one for the issue and one for its comments. Small JavaScript nodes clean both up. The comments matter. That's often where the real reproduction steps are hiding.

It plans like a QA engineer. The LLM gets the bug summary, the description, the comments, the current screenshot, and a library of reusable procedures for the app we test, which is a platform for running EOS/L10 meetings. For example, "start an L10 meeting" has two variants, depending on whether a Resume button is already on screen.

It has to return strict JSON, an ordered list of steps:

{ "stepDescription": "...", "stepExpectedResult": "...", "sessionID": "..." }
Enter fullscreen mode Exit fullscreen mode

The sessionID is filled in by a template, and the prompt forbids the model from inventing or changing it.

An LLM runs the loop. The orchestrator gets the whole list and one tool: the sub-agent. Its instructions are short:

Execute each step by calling the tool. Wait for the result before proceeding. If status is test_broken, retry once. If still broken, or if failed, stop execution.

It reports back. The orchestrator's transcript of tool results goes to Jira as a comment. A JavaScript node checks the status of the last step. If it passed, the agent finds the ticket's "Done" transition and fires it. If not, the ticket stays open and waits for me. At the end the browser session is shut down.

The sub-agent: run one step, remember it

This is the part I'm happy about. step_execution gets a session ID, a step description and an expected result. It does five things.

1. Look before touching anything. One BaaS call returns a screenshot and the page's HTML. A second call asks BaaS which automation commands it currently supports. That list goes straight into the code-generation prompt, so the language isn't frozen into my prompt text. If we add a command to BaaS tomorrow, the agent can use it tomorrow.

2. Work out where it is, then check its memory. An LLM classifies the screenshot into one of a fixed set of pages, the old Page Object idea: Active Meeting, Issues, Scorecard and so on. Then it has to search a memory collection called test_steps, filtered by that page. It's looking for a step that means the same thing, the same action with the same expected outcome, even if it's worded differently. It returns the step plus one flag: from_memory, true or false.

3. Reuse or write. A branch splits on that flag.

  • Found: pull the stored program from memory. No code generation at all.
  • Not found: a second LLM writes a new program from scratch. It can only use the commands discovered in step 1, and it has house rules. Use clickN for repeated selectors. Fall back to llmClickElement when there's no reliable selector. Output one line, no template literals.

That's the bet. The first time the agent sees something like "click Raise Issue on a to-do," an LLM writes the code. The next time a similar step shows up, on a different ticket, phrased differently, it should come straight from memory.

4. Run it and look again. The program, stored or new, goes to the browser, followed by another screenshot. A third LLM compares before and after against the expected result. It returns one of four verdicts: success, failed, broken or test_broken.

5. Remember or forget. A final branch decides what happens to memory.

  • success or failed: store the program that was used, tagged by page.
  • broken or test_broken: delete the stored program that just produced a bad result.

So code that works gets reused, and code that misbehaves gets thrown out. The agent shouldn't keep trusting automation that just let it down. That's the theory, anyway.

How to try it

The whole thing is open source and runs on Docker. It's built from two pieces. Agent Factory is the studio where you design, edit and chat with agents. BaaS (Browser as a Service) is what actually drives Chrome: it opens pages, clicks, waits and takes screenshots for the agent.

GitHub logo Ursa-Minor-Beta / agent-factory-docker-api-ui

Docker deployment for Agent Factory (agent-factory + agent-factory-ui).

Agent Factory Deploy

Docker deployment for Agent Factory (agent-factory + agent-factory-ui).

Quick Start

  1. Copy environment file:
cp .env.example .env
Enter fullscreen mode Exit fullscreen mode
  1. Edit .env with your configuration (at minimum set ADMIN_EMAIL and ADMIN_PASSWORD)

  2. Run with local MongoDB:

docker-compose --profile with-db up --build
Enter fullscreen mode Exit fullscreen mode

Or with external MongoDB (set MONGODB_URI in .env):

docker-compose up --build
Enter fullscreen mode Exit fullscreen mode
  1. Access

Deploy Specific Version

API_VERSION=<branch/tag> UI_VERSION=<branch/tag> docker-compose up --build
Enter fullscreen mode Exit fullscreen mode

Update After Branch Changes

docker-compose --profile with-db build --no-cache && docker-compose --profile with-db up
Enter fullscreen mode Exit fullscreen mode

Configuration


































































Variable Description Default
API_VERSION API branch/tag to deploy master
UI_VERSION UI branch/tag to deploy master
API_PORT API exposed port 3000
UI_PORT UI exposed port 8080
MONGO_PORT MongoDB exposed port 27017
MONGO_USER MongoDB root username -
MONGO_PASSWORD MongoDB root password -
JWT_SECRET JWT signing secret -
ENCRYPTION_KEY Encryption key (32 bytes) -
MONGODB_URI MongoDB connection string mongodb://mongo:27017/agent-factory
VITE_API_URL API URL for UI http://localhost:3000

…




GitHub logo Ursa-Minor-Beta / baas

Browser as a Service — browser automation and document extraction over HTTP

BaaS — Browser as a Service

BaaS drives a real Chrome instance behind an HTTP API. You POST a small JavaScript-like program describing what to do in the browser; BaaS runs it and returns the page, a screenshot, cookies, a downloaded file, or Markdown extracted from whatever the page served you.

It also parses documents without a browser: PDF, DOCX, XLSX, PPTX, CSV and plain text all convert to Markdown through the same API.

  • Browser automation — navigate, click, type, wait, upload, download screenshot, run arbitrary JS, all through ChromeDP.
  • LLM-assisted actions — llmClick('the login button') llmLogin(user, pass), llmText('one-time code'), for pages whose selectors you do not know in advance. Backed by OpenAI, Azure OpenAI or Ollama.
  • Sync and async sessions — one-shot programs, or a long-lived session you send commands to over time.
  • Readability extraction — article text, tables and metadata from a URL.
  • Document conversion…

The short version:

  1. Start Agent Factory. Clone the repo, set an admin email and password in .env, and run:

    docker-compose --profile with-db up --build
    
  2. Run BaaS. Clone the repo and set an API key, BROWSER_HEADFUL=true and your LLM credentials in its .env. If you run BaaS locally, also set your Chrome path. Locally it operates your own browser, and you get to watch the agent click around in a real window. Map host.docker.internal to 127.0.0.1 so the two can find each other, then start BaaS. It will complain about missing Pandoc and friends. Ignore it, that's what success looks like.

  3. Connect and smoke test. In the studio at localhost:8080, add BAAS_API_KEY under Secrets and your LLM key under Provider. Then run the Browser Screenshot agent on google.com. If you get a screenshot back, you're wired up.

  4. Hook up Jira. Create an Atlassian API token, base64-encode email:token, and save it as the Jira_auth secret:

    echo -n "you@example.com:your-api-token" | base64
    

    In the Test Orchestrator agent, set your test environment URL and Jira subdomain, and hardcode your login in the http-1 node. Login rarely changes, so there's no point paying tokens for it every run.

  5. Run it and tune it. Chat with Test Orchestrator, give it a ticket ID like SCRUM-165 and your URL, and watch. Then teach it your app: describe it in the llm-generate-steps prompt (your user documentation works well here), and adjust the three Test Step Executor prompts. llm-0 adapts steps to the screen, llm-1 writes the BaaS commands, and llm-2 decides pass, fail or broken.

Every command and config value is in my previous post: Building an AI QA Agent That Tests Your Jira Tickets in a Real Browser.

This is a side project, and I'd honestly like to know whether to keep building it. If it's useful, star the repo, send me a DM, or leave a comment. If it breaks on your app, tell me that too. That's often the more useful message.

Keep humans on the loop, and don't ship on Fridays.

Top comments (0)