<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Bartek Szafranow</title>
    <description>The latest articles on DEV Community by Bartek Szafranow (@b_szafranow).</description>
    <link>https://dev.to/b_szafranow</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4166585%2Fe1bd4ae7-43e2-4df0-a3ae-bd6f706492f2.jpg</url>
      <title>DEV Community: Bartek Szafranow</title>
      <link>https://dev.to/b_szafranow</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/b_szafranow"/>
    <language>en</language>
    <item>
      <title>Introducing e2e: open source agentic testing for web, iOS, and Android</title>
      <dc:creator>Bartek Szafranow</dc:creator>
      <pubDate>Thu, 08 Oct 2026 18:12:39 +0000</pubDate>
      <link>https://dev.to/testerarmy/introducing-e2e-open-source-agentic-testing-for-web-ios-and-android-4flb</link>
      <guid>https://dev.to/testerarmy/introducing-e2e-open-source-agentic-testing-for-web-ios-and-android-4flb</guid>
      <description>&lt;p&gt;&lt;em&gt;Written by Oskar Kwasniewski, CTO at TesterArmy. Originally published on the &lt;a href="https://tester.army/blog/introducing-e2e" rel="noopener noreferrer"&gt;TesterArmy blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We're open sourcing e2e, a TypeScript testing framework that lets agent steps and exact assertions live in the same test. It runs on web, iOS, and Android with one API, it works with the AI model or subscription you already pay for, and the setup wizard takes most projects from install to a first passing test in about a minute:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx e2e init
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;e2e comes out of the work we do every day at TesterArmy, where our testing agent tests other teams' apps. We built it based on the knowledge and feedback we got from our customers, and it will be the open foundation that powers TesterArmy in the future.&lt;/p&gt;

&lt;p&gt;Here's what you can expect from this first release:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Agent steps next to exact checks.&lt;/strong&gt; You write role-based locators and auto-retrying assertions where precision matters, and use &lt;code&gt;agent.act()&lt;/code&gt;, &lt;code&gt;agent.assert()&lt;/code&gt;, and &lt;code&gt;agent.extract()&lt;/code&gt; where describing a goal is easier than scripting it. Because both live in one test, the report always says exactly what was verified.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The model runs once per step.&lt;/strong&gt; When an agent step passes with a recorded check, e2e caches its actions and replays them on later runs without calling the model, so a cached step runs at the speed of a scripted test and costs no tokens.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One suite for web and mobile.&lt;/strong&gt; Browsers run through &lt;a href="https://playwright.dev" rel="noopener noreferrer"&gt;Playwright&lt;/a&gt;, and iOS simulators and Android emulators run through &lt;a href="https://github.com/callstack/agent-device" rel="noopener noreferrer"&gt;agent-device&lt;/a&gt;, with the same fixtures, locators, and assertions on every platform.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The model and subscription you already have.&lt;/strong&gt; Agent steps run on any AI SDK model, including local ones, or on a ChatGPT, GitHub Copilot, or SuperGrok subscription, and we add no markup on tokens.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Support for coding agents.&lt;/strong&gt; e2e ships an agent skill and an MCP server, so Claude Code, Cursor, or Codex can explore your app and write tests with real locators.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;iframe class="tweet-embed" id="tweet-2105675143464763540-763" src="https://platform.twitter.com/embed/Tweet.html?id=2105675143464763540"&gt;
&lt;/iframe&gt;

  // Detect dark theme
  var iframe = document.getElementById('tweet-2105675143464763540-763');
  if (document.body.className.includes('dark-theme')) {
    iframe.src = "https://platform.twitter.com/embed/Tweet.html?id=2105675143464763540&amp;amp;theme=dark"
  }



&lt;/p&gt;

&lt;h2&gt;
  
  
  Why we built it
&lt;/h2&gt;

&lt;p&gt;Testing so many different apps keeps surfacing the same trade-off. An agent can complete a goal like "buy the cheapest item on this list" without a single selector, but when the test passes, it's hard to say exactly what was checked. A scripted end-to-end test tells you precisely what it verified, and you pay for that precision by rewriting a selector every time someone renames a button or the flow in your app changes.&lt;/p&gt;

&lt;p&gt;Teams usually pick one approach for the whole suite and live with its costs. e2e lets you make that choice per step instead of per suite.&lt;/p&gt;

&lt;p&gt;My favorite thing about e2e is that you can gradually adopt agentic APIs where it makes sense. Migration from frameworks like Playwright to e2e is super simple: you port your tests using the same familiar APIs, then add agent steps where they help.&lt;/p&gt;

&lt;p&gt;On top of that, e2e can run exploration bug bashes via &lt;code&gt;e2e explore&lt;/code&gt; before you open a pull request, which makes a great verification step in your software factory. More on that is coming soon.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agent steps and exact checks in one test
&lt;/h2&gt;

&lt;p&gt;If you've written Playwright tests, most of e2e will feel familiar: role-based locators, auto-retrying assertions, and the same &lt;code&gt;test()&lt;/code&gt; API. The three agent steps cover the rest. &lt;code&gt;agent.act()&lt;/code&gt; carries out a goal you describe, &lt;code&gt;agent.assert()&lt;/code&gt; checks a condition you describe in plain language, and &lt;code&gt;agent.extract()&lt;/code&gt; reads data off the screen so you can use it later in the test.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;expect&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;e2e&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;a member upgrades to Pro&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;screen&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/settings/billing&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;act&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;upgrade the workspace to the Pro plan&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;screen&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByRole&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;status&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;toContainText&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Pro&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;assert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;the invoice preview shows the Pro price&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In this test, the agent works out the upgrade flow on its own, so a reworked billing page is far less likely to break it. The &lt;code&gt;expect&lt;/code&gt; line then pins down the result with an exact locator, which means a passing run tells you the status really reads "Pro", whatever path the agent took to get there.&lt;/p&gt;

&lt;h2&gt;
  
  
  What agent steps cost
&lt;/h2&gt;

&lt;p&gt;Agent steps are slower and more expensive than scripted ones on their first run, because each action needs a model call. We designed e2e so that you pay this cost once per step rather than on every run.&lt;/p&gt;

&lt;p&gt;When an agent step passes with a recorded check, e2e saves the actions it took to a trace cache. On the next run, it replays those actions directly, without calling the model, so the step runs at the speed of a scripted test and uses no tokens. If the UI changes enough that the replay fails, the runner hands the step back to the live agent to find a new path, and that run costs model calls again. In practice, this means your token spend follows how often your UI changes rather than how often your tests run.&lt;/p&gt;

&lt;h2&gt;
  
  
  One API for web and mobile
&lt;/h2&gt;

&lt;p&gt;e2e drives browsers through &lt;a href="https://playwright.dev" rel="noopener noreferrer"&gt;Playwright&lt;/a&gt;, and iOS simulators and Android emulators through &lt;a href="https://github.com/callstack/agent-device" rel="noopener noreferrer"&gt;agent-device&lt;/a&gt;. Your tests use the same fixtures, locators, and assertions on every platform, so your web app and mobile app can share one suite and one config:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;E2EConfig&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;e2e&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;web&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@e2e-dev/web&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;mobile&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@e2e-dev/mobile&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;gateway&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ai&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;targets&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;web&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;engine&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;web&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
      &lt;span class="na"&gt;app&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;url&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;http://127.0.0.1:3000&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ios&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;engine&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;mobile&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;platform&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ios&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}),&lt;/span&gt;
      &lt;span class="na"&gt;app&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;bundleId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;com.example.app&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="na"&gt;agents&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;default&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;gateway&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;openai/gpt-6-luna-fast&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="nx"&gt;satisfies&lt;/span&gt; &lt;span class="nx"&gt;E2EConfig&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you'd rather your CI run on infrastructure someone else maintains, e2e has first-class support for hosted browsers from &lt;a href="https://kernel.sh" rel="noopener noreferrer"&gt;Kernel&lt;/a&gt; and mobile simulators from &lt;a href="https://expo.dev" rel="noopener noreferrer"&gt;Expo&lt;/a&gt;. Your tests stay the same either way, and only the config changes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models, keys, and secrets
&lt;/h2&gt;

&lt;p&gt;Agent steps run on any &lt;a href="https://ai-sdk.dev" rel="noopener noreferrer"&gt;AI SDK&lt;/a&gt; model. You can bring your own key through &lt;a href="https://vercel.com/ai-gateway" rel="noopener noreferrer"&gt;Vercel AI Gateway&lt;/a&gt; or &lt;a href="https://openrouter.ai" rel="noopener noreferrer"&gt;OpenRouter&lt;/a&gt;, point e2e at a local model server, or sign in with the ChatGPT, GitHub Copilot, or SuperGrok subscription you already have. You pay your provider's price for tokens, with no added markup.&lt;/p&gt;

&lt;p&gt;Test credentials stay in environment variables, outside the model's context. The agent can type a password into a login form without the password ever appearing in a prompt, which matters when the model runs on a third-party API.&lt;/p&gt;

&lt;h2&gt;
  
  
  Working with coding agents
&lt;/h2&gt;

&lt;p&gt;e2e ships an agent skill and an MCP server (Model Context Protocol, the standard most coding agents use to call external tools). With them, Claude Code, Cursor, or Codex can open your app, explore it, write tests with locators taken from the real UI, and read the failure report when something breaks:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add tester-army/e2e
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is how tests get written in our own repo: the coding agent that built a feature also explores it and adds the regression test in the same change.&lt;/p&gt;

&lt;h2&gt;
  
  
  See it in action
&lt;/h2&gt;

&lt;p&gt;The two-minute explainer below shows how agent goals and deterministic checks fit together in one test, then sets up e2e from scratch on a Next.js app. It's a quick way to see the whole wizard before you run it on your own project.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/NEawZb6Hssw" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  Availability
&lt;/h2&gt;

&lt;p&gt;e2e is available today on npm as &lt;code&gt;e2e&lt;/code&gt;, under the Apache 2.0 license. It's still pre-1.0: the core API is the one we use every day, and we expect parts of it to change as more teams run it on their own apps.&lt;/p&gt;

&lt;p&gt;Web, iOS, and Android are supported now. Desktop apps and other platforms aren't covered yet. Engines are pluggable, so you can write your own and keep the test API unchanged, and we're working on more platforms ourselves.&lt;/p&gt;

&lt;p&gt;To get started, run this in your project:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx e2e init
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;a href="https://e2e.tester.army/docs" rel="noopener noreferrer"&gt;docs&lt;/a&gt; cover writing tests, choosing a model, and running in CI, and &lt;a href="https://tester.army/e2e" rel="noopener noreferrer"&gt;tester.army/e2e&lt;/a&gt; has the full overview. The code lives at &lt;a href="https://github.com/tester-army/e2e" rel="noopener noreferrer"&gt;github.com/tester-army/e2e&lt;/a&gt;. &lt;/p&gt;

&lt;p&gt;If e2e saves you from rewriting a selector or two, a star helps more people find it, and issues and PRs are very welcome.&lt;/p&gt;

</description>
      <category>testing</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
