The short version
Tessvia is a service where you write the steps a user takes on screen once, as a single scenario (the one source everything comes from), and then mechanically derive four different deliverables from it. The four are: tutorial playback, automatic manual generation, E2E (end-to-end — reproducing a user's actions exactly as they happen) tests, and liveness monitoring of the screen. You no longer write the same operation four times.
We built this tool while handing our own development over to AI, one piece at a time. Tests, manuals, and monitoring all come down to the same thing: writing out "the user operates the screen like this." If that's true, then write those steps once and reuse them. Tessvia starts from that plain idea. This series describes how we built Tessvia and what we learned along the way, in order. This article is the entry point: it explains, with the actual screens, what Tessvia does.
Why combine four things into one
What we want to do is one thing. You write the steps — "the user operates this screen like this" (the scenario) — just once. You manage that one scenario as the single source. After that, you mechanically derive four different deliverables from the same scenario.
- Tutorial playback: Overlay callouts and highlights on the actual screen to guide the user through the operation.
- Automatic manual generation: AI actually traces the same operation, captures it, and assembles it into manual pages.
- E2E tests: Run the same operation as a Playwright (a tool for automating browser operations) test to check that the screen isn't broken.
- Liveness monitoring (at the design stage): Run the same operation on a schedule to watch whether the production screen keeps working. This fourth use is still at the design stage (implementation came after the other three).
Normally, these four are written separately by separate people. And so, bit by bit, they diverge. Take a concrete example. Suppose a "Submit" button on some screen changes to "Save and submit." The engineer fixes the button name in the E2E test. But the image pasted into the manual isn't re-captured, so it keeps the old button name. The tutorial callout still has the old wording too. The monitoring script, if unlucky, can't find the button and reports a false alarm. One change ripples to four places, but only the one place someone noticed gets fixed. That is the real cost of writing them separately.
If everything comes out of one scenario, fixing the button name in that scenario changes all four at once. You don't write the same operation four times. Drift caused by forgetting to fix something can't happen, in principle. That is also the motivation for building it. We just wanted an easier way to avoid writing the same work over and over. That's all.
What a scenario looks like
The source "scenario" doesn't require learning any special notation. You just line up steps like "press this button" or "type this into this field" one at a time, in plain language (everyday words). For each step, you write a title (the name of that step) and a body (what it does).
When you run it, you don't have to rush straight to the end — you can advance one step at a time, stopping as you go. Because you can confirm which step changes the screen's state, you can also check step by step, with your own eyes, whether the steps the AI proposed are correct.
Steps can branch, too. For example: "if already logged in, proceed; otherwise log in first." Real operations don't always run in the same order. Because you can write the points where the path forks by state into one scenario, both the test and the manual can express the flows that can actually occur, as they are.
Why it's hard to break: the AI finds the elements
This is the most important part of the tool. How do steps written in plain language turn into actual operations?
With traditional E2E tests — tools like Playwright or MagicPod — to make a machine perform the operation "press the login button," a person had to specify exactly where that button is on the screen, as a DOM selector (a way of pointing to an element within the HTML structure). It meant tracing the element in the developer tools and teaching the machine "the login button is this element." The problem is that the moment the screen's construction changes a little — the button moves, the HTML structure gets rearranged — that specification stops working and the test breaks. Nothing about the feature itself is broken, yet the test goes red just because the selector slipped. This is the single biggest reason E2E tests have been called "hard to write and hard to fix."
Tessvia hands that part to the AI. All a person writes is the step in everyday words: "press the submit button." When it runs, the AI compares that sentence against the screen currently open, finds which element is the "submit button" on its own, and operates it. There is no step where a person writes and fixes a DOM selector. So even if the button's position or the HTML structure changes somewhat, as long as the intent "submit button" is the same, the AI can find the same button again. The situation where you remodel the screen a little and the tests all go red, one after another, becomes much less likely.
This form — "a person writes the intent, the AI finds the element" — produces two effects at once. One is that the scenario is easy for people to read. It isn't a list of selectors but the manual itself — "log in," "press the submit button" — so even someone who didn't write it can understand its contents. The other is that the test is hard to break, because the fixed selector that was the source of fragility has been replaced by the AI's search at run time. Readability and being hard to break usually don't go together, but here both come out of the same mechanism. And these same steps flow on not only to the test but to the tutorial, the manual, and the monitoring — which is the point made at the start of this article.
Seeing it on screen (1): running a test from the extension's console
Words alone are hard to convey, so let me show you the actual screens. Tessvia has a Chrome extension (an add-on) that you install in your browser. With the screen you want to test still open, you open a tall panel over it, and from there you run scenarios, create manuals, and report bugs.
The screen below is that extension's panel. At the top left it says "Test & Record Console," and "Connected" at the top right indicates that it's connected to our own (EarthLink Network) workspace. Tessvia organizes work by organization (a company or team) and project (the target app). Here, a project called "promptflow CI catalog" is selected. The panel is divided into three tabs — "Test Execution," "Manual Authoring," and "Bugs" — and the Test Execution tab is currently open. (The tab names appear in English on screen.)
The "Test Execution" tab lists the registered tests. In this example, a suite (a grouping of tests) called "CI full suite (promptflow)" contains 361 cases, divided into per-feature categories like "chat (23)," "auth (62)," and "billing (59)." Open a category and you can see each case name one by one, along with how many steps of operation each is made of (for example, "5 steps"). Every one of them is simply a scenario a person wrote in plain language, laid out as is.
Seeing it on screen (2): open a case and let the AI run it
Open a single case and the scenario — what it does — is laid out step by step. The screen below is a case for "can send a message and receive an AI response," where five steps read straight through: "type a prompt into the chat input field," "press the submit button," "the response is displayed as streaming," and so on. There's no special notation. The steps a person wrote in everyday words become the units of execution as they are.
When you run it, as described earlier, the AI traces these steps on the screen from top to bottom. Where it says "press the submit button," the AI finds the submit button on the current screen and presses it. It works without a person specifying selectors, and it follows along even when the screen changes somewhat. Whether each step passed is something the AI checks automatically. It's fully automatic by default.
On top of that, from "JUDGE THIS CASE" at the bottom of the screen, a person can also manually set "Pass," "Fail," or "Skip." When the automatic result doesn't match reality — for instance, when it nominally ran through to the end but, to a human eye, should be a fail — the position is that a person can override it at the end. The AI does the hands-on work of operating the screen; a person can hold final responsibility for the pass/fail result. In other words, the hands-on work and the responsibility for pass/fail are separated.
Seeing it on screen (3): report a bug right where you found it
If you spot something wrong with the screen in the middle of testing or authoring a manual, you can report it on the spot, without switching to another tool. The screen below is that "bug report" modal (a small window). It captures the screen you're looking at as is, and you can overlay annotations — Arrow, Circle, Square, Text — in a color of your choice. Instead of writing out in words what's wrong and where, you can mark it directly on the image. After that, you just write a title and a description and press "Register report."
Reported bugs accumulate in the list on the same "Bugs" tab. In the screen below, "RECENTLY REPORTED BUGS" lists items that were actually reported, like "model switch reverts to default on reload" and "individual cost of some models shows NaN in comparison mode." Each is linked to the scenario it was reported from, so you don't have to recall later "which screen, in what state, was this."
How to present test results (passed or failed, which step it stopped at) and how to actually compose test scenarios are covered in detail in the E2E installment of the series. In this article, just take away the overall picture — "four things come out of the same scenario" — and the fact that these operations are completed right on top of the target screen.
Why a Chrome extension
Testing and manual-authoring tools usually run in a screen separate from the target app. That means you open the manual and the actual app screen in separate windows. You end up going back and forth, retracing memory — "this operation was, I think, on that screen…" — and as that memory goes stale, the manual's contents and the real screen drift apart.
The extension form is a choice to eliminate that back and forth. With the page under test still open, you layer a panel over that same screen. You can point to the button you're looking at, right there, and record it, right there. The person writing the steps and the screen being operated are in the same place. So the steps and reality are less likely to drift apart. Another benefit is that the screenshots on each step of the manual are captured from this very "screen you're looking at." There's no need to stand up a separate environment for captures and shoot them again.
An AI agent traces the actual screen
When you say "automatically generate a manual," people tend to think it's just lining up images captured in advance. Tessvia is different. An AI agent (a mechanism that operates the screen automatically) actually traces the steps written in plain language and captures the real screen at each step. So even if the app's appearance changes, re-running the scenario re-builds the manual with the latest screen. The entire step of manually taking screenshots and swapping them in is unnecessary.
We provide two ways to actually operate the screen, depending on the use. One is "generate here," where Tessvia's server itself drives the browser and captures. It's self-contained at hand and requires no extra setup. The other is "via a runner," where a capture agent running on a separate machine operates on your behalf. You use this when you've placed the server somewhere and want to capture an internal app from the machine at hand. You can choose where to run it based on where the screen you want to capture is.
Apps that require login can be captured too. If you prepare and pass in session information for an already-logged-in state, the operation can start from that state. Authentication information is passed only as environment variables (configuration values passed into a program from outside), never written down in files or logs. We've kept to a design, from the start, that puts as little of the test target's contents outside as possible.
Liveness monitoring, the fourth use (at the design stage)
As a fourth use, we are designing liveness monitoring. The concept is to run an existing scenario automatically on a schedule and watch whether the production screen is still working. Tests run on every change, but monitoring runs separately, repeating on a set schedule. If one morning the login steps have started stopping partway, you can notice it before the users do — that's the shape monitoring aims for.
What matters here is, again, the "one scenario." Instead of writing new steps for monitoring, you simply run an existing scenario on a schedule as is. The flow checked by the test and the flow watched by monitoring can be said to be the same thing. You can keep all four — tutorial, manual, test, and monitoring — coming out of the same single scenario, right to the end. Note that this liveness monitoring was implemented later than the other three and is still at the stage of preparing the integration design. What's actually running is the tutorial playback, automatic manual generation, and E2E tests shown on screen so far.
The flow of using it
The overall flow is simple.
- Write the scenario once. Write the steps — "operate this screen like this" — in plain language (everyday words). Just line up a title and a body for each step.
- Have the AI actually trace it. Following the steps you wrote, the AI agent operates the actual screen and captures a screenshot of each step. In a few minutes you have a manual with screenshots, which you can also download as is.
- Expand it to four uses. Play the same scenario as a tutorial, assemble it into a manual, run it as an E2E test, or run it on a schedule as monitoring (monitoring is at the design stage) — dispatch it according to the use.
What a human does here is two things: deciding the steps, and confirming that what comes out matches the intent. The work of writing out the same operation four times is shifted to the machine's side. More than the speed of your hands, it's the speed at which you decide what to make and in what order — and the quality of that judgment — that determines the result.
Who the tool is for
Tessvia is for people who want to automate their own development and operations as much as possible. We have in mind, first of all, developers running multiple services with few people, and people doing business on their own. In this position, the person who writes the tests, the person who makes the manuals, and the person who watches the monitoring are usually the same one person. When one person wears four hats, the burden of writing the same operation four times hits directly. So the value of a mechanism where you write once and expand to four shows up largest here.
We ourselves needed this tool from exactly that position. Developing and operating our own several services with few people, we had no room to prepare tests, manuals, and monitoring separately. If hands are short, the only way is to build things so you don't do the same work over and over. Tessvia is a tool born of that need. So this series is not a clean product introduction but a record of the process of building something out of necessity.
When you automate one piece of work, the next piece you want to automate comes into view. You want to create tests automatically — then you want to create manuals automatically — then you want to watch automatically whether it's running. That chain, gathered into one scenario, is Tessvia today.
How to read this series
Tessvia today is actually two projects that started separately and later merged into one. One side was "making scenarios," the other "delivering (distributing) scenarios." In hindsight, these two were two faces of a single service, but at the time they started out in separate contexts, so for a while they grew as separate products. Why we were slow to notice this duplicated development, and how we folded them into one, is something this series will also cover.
In this series, we trace that history in order. The next article describes the first three days when the name Tessvia was born (a record of its birth) — the story of how, when we handed over the spec all at once, a monorepo (a mechanism for handling multiple packages together in one repository) MVP (minimum viable product) came up the same day. After that, we continue one by one: how we composed the browser extension's E2E, how we stopped the AI's mistaken generation, how we folded a duplicated implementation into one, how we realized automatic manual generation, and so on.
This is not a clean, finished design written up after the fact. We'll write the process as it was — making mistakes, fixing them, and sometimes rebuilding the whole thing — as we went. We hope it conveys what actually happens to those who, like us, are trying to automate their own work with AI.
A lesson worth carrying over
- Write the same operation only once. Tutorial, manual, test, and monitoring look like different things, but their contents are the same "steps for operating the screen." Gather them into one scenario and drift from forgetting to fix something can't happen, in principle.
- Hold the scenario in a machine-readable form. An operation where a person fixes four places by hand will, someday, drift somewhere for certain. Build it so fixing one place changes everything, and there's no room left to drift.
- Do the work right on top of the target screen. The more you switch to another tool, the more you handle the manual and the actual screen separately, and divergence is born. When you can run it right on top of the screen you're looking at, the steps and reality are less likely to drift apart.
- Humans concentrate on steps and judgment. Shift the hands-on work of writing it out four times to the machine, and let humans decide what to make and confirm that what comes out matches the intent. The further automation goes, the more this separation pays off.





Top comments (0)