<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Jakob Norlin</title>
    <description>The latest articles on DEV Community by Jakob Norlin (@jakobnorlin).</description>
    <link>https://dev.to/jakobnorlin</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4125858%2F62ae09e2-7bf3-4b91-9b69-638dcbeba5e8.jpeg</url>
      <title>DEV Community: Jakob Norlin</title>
      <link>https://dev.to/jakobnorlin</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/jakobnorlin"/>
    <language>en</language>
    <item>
      <title>How To Write Playwright tests in minutes with Playwright MCP and Claude Code</title>
      <dc:creator>Jakob Norlin</dc:creator>
      <pubDate>Mon, 05 Oct 2026 09:44:22 +0000</pubDate>
      <link>https://dev.to/jakobnorlin/how-to-write-playwright-tests-in-minutes-with-playwright-mcp-and-claude-code-1o0d</link>
      <guid>https://dev.to/jakobnorlin/how-to-write-playwright-tests-in-minutes-with-playwright-mcp-and-claude-code-1o0d</guid>
      <description>&lt;p&gt;&lt;em&gt;This article was originally published on the&lt;/em&gt; &lt;a href="https://endform.dev/blog/playwright-mcp-claude-code" rel="noopener noreferrer"&gt;&lt;em&gt;Endform blog&lt;/em&gt;&lt;/a&gt;&lt;em&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3egdamuwg279l903vgqt.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3egdamuwg279l903vgqt.webp" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;By the end of this walkthrough you'll have a single Playwright test that Claude Code wrote against your live app, that you've reviewed in two passes, run repeatedly to check for flakiness, and committed next to the plain-English scenario it came from. The trick that makes it trustworthy: the agent reads locators off the running page through the Playwright MCP server instead of guessing them from your source code.&lt;/p&gt;

&lt;p&gt;You need four things installed before step 1, so start there.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before you start
&lt;/h2&gt;

&lt;p&gt;Confirm all four of these first. Skipping one is the most common reason step 1 fails.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Node.js 20 or newer.&lt;/strong&gt; Check with &lt;code&gt;node --version&lt;/code&gt;. The MCP server launches through &lt;code&gt;npx&lt;/code&gt;, so Node is required no matter how you installed Claude Code.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Claude Code, installed and signed in.&lt;/strong&gt; Check with &lt;code&gt;claude --version&lt;/code&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;A Playwright project.&lt;/strong&gt; You need &lt;code&gt;playwright.config.ts&lt;/code&gt; at the repo root. No project yet? Run &lt;code&gt;npm init playwright@latest&lt;/code&gt; to scaffold one.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;An app to test.&lt;/strong&gt; Local dev server or staging URL, either works. Playwright can start it for you when the config says so.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Got all four? Connect Claude Code to a browser.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem this setup solves
&lt;/h2&gt;

&lt;p&gt;Here's a failure pattern you've probably lived through: Claude Code writes an end-to-end test, it passes locally, you merge, and CI goes red a day later on what looks like flakiness. The test was wrong from the start, it just had no way to show it.&lt;/p&gt;

&lt;p&gt;The reason is where the locators came from. Reading your components without ever opening the app, an agent spots a button with a class like &lt;code&gt;.btn-primary&lt;/code&gt; and writes a selector against it. Then a component library or some runtime logic rewrites that class on its way to the DOM, and by the time the browser renders, the selector points at a hashed string, a restructured node, or nothing. Nobody caught it because nobody ran the test against a real page. CI is the first thing that does.&lt;/p&gt;

&lt;p&gt;Compare the two ways the same button gets targeted:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Hallucinated: guessed from training data, does not exist on this page&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;locator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;#submit-btn&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;click&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="c1"&gt;// Grounded: read from the accessibility tree Claude Code can see&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByRole&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;button&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Place order&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;click&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Anthropic's Model Context Protocol is what closes the gap. The Playwright MCP server hands Claude Code a live browser: it can open the app, walk the accessibility tree, and build locators out of the roles and names the page actually exposes. It's working from a structured snapshot rather than a screenshot, which keeps the locators stable and the token cost low, because the model reasons over text it can already read. Same model, better inputs. Nothing here is a capability upgrade, only an access one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Register the Playwright MCP server
&lt;/h2&gt;

&lt;p&gt;A single command wires the server into Claude Code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp add playwright &lt;span class="nt"&gt;--&lt;/span&gt; npx &lt;span class="nt"&gt;-y&lt;/span&gt; @playwright/mcp@latest
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That entry gets saved to your local config, and the server process spins up the next time you open a session. &lt;code&gt;-y&lt;/code&gt; skips the install confirmation &lt;code&gt;npx&lt;/code&gt; would otherwise wait on, and &lt;code&gt;@latest&lt;/code&gt; grabs the newest published build so you're not stuck on a stale cache.&lt;/p&gt;

&lt;p&gt;By default this is a personal, single-project entry. Working on a team? Add &lt;code&gt;--scope project&lt;/code&gt;, which writes the same config to a &lt;code&gt;.mcp.json&lt;/code&gt; at the repo root so everyone shares one server without redoing setup:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"playwright"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"stdio"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"npx"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"args"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"-y"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"@playwright/mcp@latest"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Don't trust it until you've checked it. Run &lt;code&gt;claude mcp list&lt;/code&gt; and look for &lt;code&gt;✔ Connected&lt;/code&gt;. Immediately after adding, you may catch a &lt;code&gt;✘ Failed to connect&lt;/code&gt; while &lt;code&gt;npx&lt;/code&gt; is still pulling the package in the background; run the command again a few seconds later and it usually clears.&lt;/p&gt;

&lt;p&gt;If retrying doesn't clear it, the problem is elsewhere. Claude Code runs the server as a subprocess in its own environment, and that subprocess can't always resolve &lt;code&gt;npx&lt;/code&gt; the way your interactive shell does. Homebrew-installed Node on macOS is the usual culprit. Give it the absolute binary path instead of trusting PATH:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp remove playwright
claude mcp add playwright &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;which npx&lt;span class="si"&gt;)&lt;/span&gt; &lt;span class="nt"&gt;-y&lt;/span&gt; @playwright/mcp@latest
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run &lt;code&gt;claude mcp list&lt;/code&gt; once more and wait for &lt;code&gt;✔ Connected&lt;/code&gt; before continuing.&lt;/p&gt;

&lt;p&gt;Connected only means the process is alive. It doesn't prove Claude Code can actually drive a browser yet. Open a session and call the tool by name, otherwise Claude Code might reach for a Bash command instead of the MCP server:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using Playwright MCP, open [your app's local URL] and tell me the page title and the first heading you see.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If it answers with something concrete off the page, not a hedge and not a generic description, the tool works. That answer is coming from a live snapshot rather than memory, which is the entire reason you connected it.&lt;/p&gt;

&lt;p&gt;These commands are current as of August 2026. MCP tooling changes quickly, so if a command here stops matching what you see, check the official Playwright MCP repo. Server verified and driving a browser, the next hurdle is login.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Get past the login wall with storage state
&lt;/h2&gt;

&lt;p&gt;Nearly every test worth writing sits behind authentication. A checkout, a settings screen, an admin panel, all of them dead ends if the &lt;a href="https://endform.dev/blog/agentic-ai-testing" rel="noopener noreferrer"&gt;agent&lt;/a&gt; can't clear the sign-in page.&lt;/p&gt;

&lt;p&gt;The MCP server has no idea about your suite's existing auth. It either carries its own browser profile across sessions or, in isolated mode, boots logged out every time. Either way it's disconnected from however your tests currently authenticate.&lt;/p&gt;

&lt;p&gt;Fix it by launching the server with a session already loaded:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @playwright/mcp@latest &lt;span class="nt"&gt;--isolated&lt;/span&gt; &lt;span class="nt"&gt;--storage-state&lt;/span&gt; .auth/user.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;--isolated&lt;/code&gt; holds the profile in memory rather than writing it to disk, so every session starts fresh. &lt;code&gt;--storage-state&lt;/code&gt; reads a saved authenticated session from a file, and anything that changes mid-session gets thrown away at the end. But that file has to exist first.&lt;/p&gt;

&lt;p&gt;Playwright's auth docs suggest a dedicated setup test that signs in and saves the session:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nx"&gt;setup&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@playwright/test&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;authFile&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;.auth/user.json&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;password&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;TEST_USER_PASSWORD&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;password&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;TEST_USER_PASSWORD is not set&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nf"&gt;setup&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;authenticate&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;goto&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://your-app.example.com/login&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByLabel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Email&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;fill&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;test-user@example.com&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByLabel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Password&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;fill&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;password&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByRole&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;button&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Log in&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;click&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;waitForURL&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;**/dashboard&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;context&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;storageState&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;authFile&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Wire it to run automatically by adding a setup project in &lt;code&gt;playwright.config.ts&lt;/code&gt; that the browser projects depend on:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;projects&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;setup&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;testMatch&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sr"&gt;/auth&lt;/span&gt;&lt;span class="se"&gt;\.&lt;/span&gt;&lt;span class="sr"&gt;setup&lt;/span&gt;&lt;span class="se"&gt;\.&lt;/span&gt;&lt;span class="sr"&gt;ts/&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;chromium&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;use&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;storageState&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;.auth/user.json&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="na"&gt;dependencies&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;setup&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;],&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run it once:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx playwright &lt;span class="nb"&gt;test&lt;/span&gt; &lt;span class="nt"&gt;--project&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;setup
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When it finishes, &lt;code&gt;.auth/user.json&lt;/code&gt; holds a live authenticated session, and &lt;code&gt;auth.setup.ts&lt;/code&gt; is now part of the suite, so login logic lives in exactly one place.&lt;/p&gt;

&lt;p&gt;Worth doing if your app allows it: create a throwaway test user through an API or seed script before login and tear it down after. Claude Code pokes at the app while it explores, and a disposable account keeps it from mutating a shared login someone else is on.&lt;/p&gt;

&lt;p&gt;With the file in place, re-register the server so it loads that session:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp remove playwright
claude mcp add playwright &lt;span class="nt"&gt;--&lt;/span&gt; npx &lt;span class="nt"&gt;-y&lt;/span&gt; @playwright/mcp@latest &lt;span class="nt"&gt;--isolated&lt;/span&gt; &lt;span class="nt"&gt;--storage-state&lt;/span&gt; .auth/user.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For teams, mirror it in &lt;code&gt;.mcp.json&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"playwright"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"stdio"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"npx"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"args"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"-y"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"@playwright/mcp@latest"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"--isolated"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"--storage-state"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;".auth/user.json"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For messier auth flows, &lt;a href="https://endform.dev/blog/playwright-mcp" rel="noopener noreferrer"&gt;Endform's Playwright MCP guide&lt;/a&gt; goes deeper on this.&lt;/p&gt;

&lt;p&gt;Now that Claude Code browses as a logged-in user, it's time to write the prompt that turns a described flow into a committable test.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Build the prompt in three parts
&lt;/h2&gt;

&lt;p&gt;This step is where a shippable test and a throwaway one diverge. Rather than one rambling paragraph that asks for a test and hopes, split the prompt into three parts that each do one job.&lt;/p&gt;

&lt;p&gt;Part one is environment setup: the app-specific facts the agent can't deduce on its own.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The website under test is at https://staging.yourapp.com.
You're already logged in as a test user through the storage state configured earlier.
The test user's cart currently holds one item, a placeholder t-shirt priced at $24.99.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Part two is the scenario, phrased the way you'd brief a teammate, not as pseudocode:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# Guest checkout with a saved card
1. Open the cart page.
2. Proceed to checkout.
3. Confirm the shipping address shown is the default one.
4. Select the saved Visa card ending in 4242.
5. Place the order.
6. Confirm the order confirmation page shows an order number and the correct total.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep this as a markdown file in the repo next to the tests it drives, not buried in a chat log. Playwright Test Agents work the same way: a planner writes the scenario as markdown, and a later step compiles it into code. Splitting the two into separate, reviewable artifacts is worth carrying over here.&lt;/p&gt;

&lt;p&gt;This part is also the one nobody on the team can write better than you. The environment facts are just facts, and the system prompt below is boilerplate you reuse everywhere. The scenario is the only part carrying judgment: does the address get confirmed before payment, does the saved card matter more than a fresh one, is the confirmation total worth asserting or is a loaded page enough? An agent can't rank those. Someone who knows the product has to.&lt;/p&gt;

&lt;p&gt;Part three is the system prompt, the standing rules that ride along on every test and describe what "good" looks like, not just what to do:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are a Playwright test generator.
Explore the app using the Playwright MCP tools before writing any code, don't generate steps from assumption alone.
Prefer getByRole, getByLabel, and getByTestId locators over CSS selectors or XPath.
Don't add manual waitForTimeout calls, rely on Playwright's built-in auto-waiting and retrying assertions instead.
Group related steps with test.step for readability in the trace viewer.
Assert on outcomes a user would actually see, not on incidental implementation details.
Save the finished test to the tests directory, run it, and keep iterating until it passes.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You don't have to write this cold. Debbie O'Brien, a longtime Playwright advocate and ex-member of Microsoft's Playwright team, keeps a public set of prompt files for exactly this, a better base than reinventing it.&lt;/p&gt;

&lt;p&gt;Stacked together, this is the full message to Claude Code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are a Playwright test generator.
Explore the app using the Playwright MCP tools before writing any code, don't generate steps from assumption alone.
Prefer getByRole, getByLabel, and getByTestId locators over CSS selectors or XPath.
Don't add manual waitForTimeout calls, rely on Playwright's built-in auto-waiting and retrying assertions instead.
Group related steps with test.step for readability in the trace viewer.
Assert on outcomes a user would actually see, not on incidental implementation details.
Save the finished test to the tests directory, run it, and keep iterating until it passes.

The website under test is https://staging.yourapp.com.
You're already logged in as a test user through the storage state configured earlier.
The test user's cart currently holds one item, a placeholder t-shirt priced at $24.99.

# Guest checkout with a saved card
1. Open the cart page.
2. Proceed to checkout.
3. Confirm the shipping address shown is the default one.
4. Select the saved Visa card ending in 4242.
5. Place the order.
6. Confirm the order confirmation page shows an order number and the correct total.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The proof this prompt works is watching the agent call the MCP tools to explore before it writes any test code. A "explore first" instruction only counts if the agent obeys it. Three parts, three jobs, one message. Next, the generation itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: Let it generate the test
&lt;/h2&gt;

&lt;p&gt;One look at the page isn't where Claude Code stops. It keeps navigating and reading through the MCP server until it has actually seen every element the scenario names. That's what separates this from feeding a model a plain-English description and hoping.&lt;/p&gt;

&lt;p&gt;Watch a run and the sequence is plain: open the page, click through the scenario's steps, read back what each returned, and only then start writing assertions. It writes the spec last, runs it, and keeps tweaking until it passes instead of handing you untested code.&lt;/p&gt;

&lt;p&gt;A real example: on a logout scenario, two headings both matched "Secure Area," so Playwright threw a strict mode violation, because a locator meant to act on one element matched several. Claude Code read the error, diagnosed it, and appended &lt;code&gt;exact: true&lt;/code&gt; so the locator demanded the full heading text rather than a substring. That correction landed in the same pass, no nudge from the developer.&lt;/p&gt;

&lt;p&gt;A checkout with a cart, saved card, and confirmation page runs the same loop, step by step, until green. Here's a second flow to show it's not a one-off.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://endform.dev/blog/playwright-mcp" rel="noopener noreferrer"&gt;Endform's guide to shipping quality end-to-end tests with Playwright MCP&lt;/a&gt; has a scenario worth reusing: check that a newly created team shows up in an activity log. Sign in, confirm a signup event is already logged, create a team, confirm the new event lands. Endform walks through the scenario and the review, not the code, so here's a plausible test around that flow, in the shape Claude Code would generate it against a live app:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;expect&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@playwright/test&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;new team activity appears in the activity log&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;step&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;open the activity log&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;goto&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://staging.yourapp.com/dashboard&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByRole&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;link&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Activity&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;click&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;step&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;confirm the signup event is already logged&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByText&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;you signed up&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;toBeVisible&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;step&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;create a new team&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByRole&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;button&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Create a new team&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;click&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByLabel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Team name&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;fill&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;QA Playground&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByRole&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;button&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Create team&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;click&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;step&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;confirm the new team event appears in the log&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByText&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;you created a new team&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;toBeVisible&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice how deliberate the file is. Every link and button name is text Claude Code read off the page. The &lt;code&gt;test.step&lt;/code&gt; blocks track the scenario one to one, so a failure drops you straight on the broken step. No manual waits anywhere, because the retrying assertions cover timing.&lt;/p&gt;

&lt;p&gt;You end up with a runnable file that clears its first run more often than not, precisely because each locator was verified against the page before a line got written. That still isn't the same as a good test, which is what step 5 is for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5: Review it in two passes
&lt;/h2&gt;

&lt;p&gt;Passing once doesn't earn a test a place in the suite. Run two review passes, because they catch different things: one on the code, one on the meaning.&lt;/p&gt;

&lt;p&gt;Pass one is the code. Read it and ask if you'd have written it this way. Generated tests tend toward the baroque, redundant checks, extra steps, logic that folds down to less, and trimming that is usually the first win. Hold the locators to the priority you set (&lt;code&gt;getByRole&lt;/code&gt; and &lt;code&gt;getByTestId&lt;/code&gt; over CSS or XPath), swap out anything brittle, and make sure each assertion proves something a user would notice rather than just confirming an action didn't throw. The &lt;code&gt;test.step&lt;/code&gt; groupings should read the way a person would narrate the flow.&lt;/p&gt;

&lt;p&gt;Pass two is the one no linter will ever do for you, and it's the one that matters more. Put the scenario next to the finished test and ask whether the thing being verified still matches what the scenario meant, not just whether it's green. Tests drift here quietly. On that logout test, Claude Code noticed the login banner showed once and got eaten by the stored session, so it retargeted the assertion to the permanent heading before wrapping up. Useful that it caught it, but it happened during generation, not review, and that's the point: next time it might not, and this pass is your backstop.&lt;/p&gt;

&lt;p&gt;The first draft gets much faster; the judgment moves into refinement. And that judgment, what the product does, what risk you're covering, what the scenario is really asserting, is yours to bring. Both passes done, one last check remains.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 6: Verify, then commit
&lt;/h2&gt;

&lt;p&gt;The commands below use this tutorial's logout test as the example. A test that passed while being generated still hasn't earned your trust. Run it again, clean, outside the generation loop:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx playwright &lt;span class="nb"&gt;test &lt;/span&gt;tests/logout.spec.ts
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A single green run tells you almost nothing about reliability. Fire it several times in a row:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx playwright &lt;span class="nb"&gt;test &lt;/span&gt;tests/logout.spec.ts &lt;span class="nt"&gt;--repeat-each&lt;/span&gt; 5
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;--repeat-each&lt;/code&gt; reruns the same test N times in one invocation. It's among the fastest ways to smoke out a test that's green most runs and red occasionally, the exact flakiness that tends to surface only once it hits CI.&lt;/p&gt;

&lt;p&gt;Repeats clean? Commit the test together with the markdown scenario behind it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git add tests/logout.spec.ts tests/scenarios/logout.md
git commit &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="s2"&gt;"Add logout test with scenario spec"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Committing them together keeps intent attached to implementation. Whoever reads the diff later sees what the test was meant to prove, not just the assertions.&lt;/p&gt;

&lt;p&gt;One thing to check before that first commit: your &lt;code&gt;.gitignore&lt;/code&gt;. A freshly scaffolded Playwright project ignores &lt;code&gt;playwright/.auth/&lt;/code&gt;, its default spot for storage state. But this setup writes to &lt;code&gt;.auth/&lt;/code&gt; at the repo root, a different path the default rule doesn't cover, which means a live authenticated session can slip into history unnoticed.&lt;/p&gt;

&lt;p&gt;Add the right path first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;".auth/"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; .gitignore
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That one line keeps the session out of your repo history by intent rather than luck.&lt;/p&gt;

&lt;p&gt;Verified, stable over repeats, and committed alongside its scenario, the test is a dependable addition. What gets harder from here: longer flows, data that shifts between runs, and auth beyond a plain login form.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this workflow falls down
&lt;/h2&gt;

&lt;p&gt;Trust comes from being straight about where Claude Code struggles, not just where it shines. Four things break it fairly predictably.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Long multi-step flows.&lt;/strong&gt; The longer the flow, the more the agent loses the thread, since each step starts from a fresh snapshot rather than the whole sequence before it, and long MCP sessions dropping browser context is a known issue. Splitting the flow into smaller scenarios and stitching them later holds up better.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Data that changes between runs.&lt;/strong&gt; The agent tends to assert against whatever it saw at generation time, a timestamp, an order ID, a count, and those move on the next run. Assertions last longer when they target something stable: a confirmation state, a completed action, the presence of a result, not the exact value from one run.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Deep conditionals.&lt;/strong&gt; While generating, the agent only travels one branch, one role, one account state, one flag. Branches it never walks stay invisible to it. Prompt each branch on its own, or write the conditional logic yourself, rather than expecting one pass to find every path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OAuth and third-party auth.&lt;/strong&gt; Usually the first wall you hit. Redirects out to Google, Microsoft, or any external identity provider leave your app, and the agent's control loop isn't built to follow through the handoff. One developer automating &lt;a href="https://endform.dev/blog/playwright-github-actions" rel="noopener noreferrer"&gt;GitHub's&lt;/a&gt; OAuth earned a temporary IP ban for it, and other providers can react to automated logins the same way, which makes this genuinely fragile. The move is to authenticate in a separate setup step, save the session with &lt;code&gt;storageState&lt;/code&gt;, and generate against an app that's already logged in.&lt;/p&gt;

&lt;p&gt;None of these shrink the workflow's value. They just mark where a scenario needs prep before you hand it over.&lt;/p&gt;

&lt;h2&gt;
  
  
  What comes next: keeping the suite fast
&lt;/h2&gt;

&lt;p&gt;We opened with a test that looked done and broke in CI regardless. Claude Code explored the real app through Playwright MCP, checked what was on the page, and produced a test you reviewed, verified, and committed.&lt;/p&gt;

&lt;p&gt;That changes the economics of test creation. Once a trustworthy test costs minutes instead of hours, teams write more of them, and a bigger suite brings its own problem. Those tests still run in CI, and as the count climbs, the bottleneck moves from writing tests to running them fast without flakiness.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://endform.dev" rel="noopener noreferrer"&gt;Endform&lt;/a&gt; runs every Playwright test on its own isolated machine in parallel, which keeps suite duration predictable no matter how many you add.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>mcp</category>
      <category>testing</category>
    </item>
    <item>
      <title>The Playwright + GitHub Actions setup that actually runs fast</title>
      <dc:creator>Jakob Norlin</dc:creator>
      <pubDate>Tue, 15 Sep 2026 14:53:37 +0000</pubDate>
      <link>https://dev.to/jakobnorlin/the-playwright-github-actions-setup-that-actually-runs-fast-317i</link>
      <guid>https://dev.to/jakobnorlin/the-playwright-github-actions-setup-that-actually-runs-fast-317i</guid>
      <description>&lt;p&gt;&lt;em&gt;This article was originally published on the&lt;/em&gt; &lt;a href="https://endform.dev/blog/playwright-github-actions" rel="noopener noreferrer"&gt;&lt;em&gt;Endform blog&lt;/em&gt;&lt;/a&gt;&lt;em&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;By the end of this you'll have a Playwright pipeline that gives useful feedback on every PR and stays under five minutes on a single runner, no sharding required. You need a Playwright suite, a GitHub repo, and the default Actions workflow that &lt;code&gt;npm init playwright@latest&lt;/code&gt; generated. Two changes carry most of the weight: caching the browser binaries and fixing parallelism. Both are config edits you can land in an afternoon.&lt;/p&gt;

&lt;p&gt;Here's the complete setup first, then the reasoning behind each piece.&lt;/p&gt;

&lt;h2&gt;
  
  
  The complete workflow
&lt;/h2&gt;

&lt;p&gt;Two files. Worker count, &lt;code&gt;fullyParallel&lt;/code&gt;, and the reporters go in &lt;code&gt;playwright.config.ts&lt;/code&gt;. Caching, the Chromium-on-PR split, and the artifact upload go in the workflow YAML.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;playwright.config.ts&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;typescript&lt;/span&gt;

&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;defineConfig&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;devices&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@playwright/test&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="nf"&gt;defineConfig&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;testDir&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;./tests&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;fullyParallel&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// parallelize tests within a file, not just across files&lt;/span&gt;
  &lt;span class="na"&gt;forbidOnly&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;!!&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;CI&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// fail the build if a stray test.only is committed&lt;/span&gt;
  &lt;span class="na"&gt;retries&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;CI&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// retry flaky tests in CI only&lt;/span&gt;
  &lt;span class="na"&gt;workers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;CI&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;50%&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;undefined&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// scale workers to the runner; measure and raise from here&lt;/span&gt;
  &lt;span class="na"&gt;reporter&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;html&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="c1"&gt;// full report, uploaded as an artifact&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;github&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="c1"&gt;// inline annotations on the run summary&lt;/span&gt;
  &lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="na"&gt;use&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;trace&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;on-first-retry&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// capture a trace when a test retries&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;projects&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;chromium&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;use&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt;&lt;span class="nx"&gt;devices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Desktop Chrome&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;firefox&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;use&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt;&lt;span class="nx"&gt;devices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Desktop Firefox&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;webkit&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;use&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt;&lt;span class="nx"&gt;devices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Desktop Safari&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;.github/workflows/playwright.yml&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="s"&gt;yaml&lt;/span&gt;

&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Playwright Tests&lt;/span&gt;
&lt;span class="na"&gt;on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;pull_request&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;branches&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;main&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
  &lt;span class="na"&gt;push&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;branches&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;main&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;

&lt;span class="na"&gt;jobs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;test&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;timeout-minutes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;60&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-latest&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/checkout@v7&lt;/span&gt;

      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/setup-node@v7&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;node-version&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;lts/*&lt;/span&gt;
          &lt;span class="na"&gt;cache&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;npm"&lt;/span&gt; &lt;span class="c1"&gt;# caches ~/.npm, the safe npm cache, not node_modules&lt;/span&gt;

      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Install dependencies&lt;/span&gt;
        &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npm ci&lt;/span&gt;

      &lt;span class="c1"&gt;# Cache the browser binaries. The lockfile is a proxy for the&lt;/span&gt;
      &lt;span class="c1"&gt;# Playwright version, so it usually turns over when you update Playwright.&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Cache Playwright browsers&lt;/span&gt;
        &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;playwright-cache&lt;/span&gt;
        &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/cache@v6&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;~/.cache/ms-playwright&lt;/span&gt;
          &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ runner.os }}-playwright-${{ hashFiles('package-lock.json') }}&lt;/span&gt;

      &lt;span class="c1"&gt;# Cache miss: download browsers and system libraries together&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Install browsers and system dependencies&lt;/span&gt;
        &lt;span class="na"&gt;if&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;steps.playwright-cache.outputs.cache-hit != 'true'&lt;/span&gt;
        &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npx playwright install --with-deps&lt;/span&gt;

      &lt;span class="c1"&gt;# Cache hit: binaries are restored, but system libraries still need installing&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Install OS dependencies for browsers&lt;/span&gt;
        &lt;span class="na"&gt;if&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;steps.playwright-cache.outputs.cache-hit == 'true'&lt;/span&gt;
        &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npx playwright install-deps&lt;/span&gt;

      &lt;span class="c1"&gt;# Chromium only on PRs for fast feedback; the full browser set on merges to main&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Run Playwright tests&lt;/span&gt;
        &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
          &lt;span class="s"&gt;if [ "${{ github.event_name }}" = "pull_request" ]; then&lt;/span&gt;
            &lt;span class="s"&gt;npx playwright test --project=chromium&lt;/span&gt;
          &lt;span class="s"&gt;else&lt;/span&gt;
            &lt;span class="s"&gt;npx playwright test&lt;/span&gt;
          &lt;span class="s"&gt;fi&lt;/span&gt;

      &lt;span class="c1"&gt;# Upload the report whether tests pass or fail, so failures are never lost&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Upload Playwright report&lt;/span&gt;
        &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/upload-artifact@v7&lt;/span&gt;
        &lt;span class="na"&gt;if&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ !cancelled() }}&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;playwright-report&lt;/span&gt;
          &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;playwright-report/&lt;/span&gt;
          &lt;span class="na"&gt;retention-days&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;14&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Drop both in, push, and the first run populates the cache while every run after it skips the download. The rest of this post explains why each block earns its place.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the time actually goes
&lt;/h2&gt;

&lt;p&gt;The default workflow from &lt;code&gt;npm init playwright@latest&lt;/code&gt; works, it just spends most of its time on setup and a serial test run. On a public repo with about 40 lightweight tests, the baseline came in at 3m 18s. Treat that as a fixed number for comparing fixes, not a prediction of your own run times.&lt;/p&gt;

&lt;p&gt;Only two steps mattered. &lt;code&gt;Install Playwright Browsers&lt;/code&gt; took 42 seconds because it downloads Chromium, Firefox, and WebKit plus their system libraries on every single run, whether or not those browsers changed. &lt;code&gt;Run Playwright tests&lt;/code&gt; took 2m 25s, and that number is about parallelism. The generated config pins CI to one worker:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;typescript&lt;/span&gt;

&lt;span class="nx"&gt;workers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;CI&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;undefined&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So the suite runs close to one test at a time even with idle cores. Checkout, Node setup, and &lt;code&gt;npm ci&lt;/code&gt; came in at a second or two each and aren't worth touching.&lt;/p&gt;

&lt;p&gt;Two independent bottlenecks: repeated download overhead, and serial execution. Here's each fix.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fix 1: cache the browser binaries
&lt;/h2&gt;

&lt;p&gt;Playwright installs its browsers to &lt;code&gt;~/.cache/ms-playwright&lt;/code&gt;, which on the Linux runner resolves to &lt;code&gt;/home/runner/.cache/ms-playwright&lt;/code&gt;. That directory is what you cache. The binaries only change when your Playwright version changes, so keying the cache to a hash of &lt;code&gt;package-lock.json&lt;/code&gt; is a convenient proxy. Updating Playwright turns the cache over, though any unrelated dependency edit does too, which forces the occasional slow run you didn't strictly need.&lt;/p&gt;

&lt;p&gt;One thing most guides skip: the cache stores the browser binaries but not the OS libraries they need, because those get installed system-wide with apt and never land in &lt;code&gt;~/.cache/ms-playwright&lt;/code&gt;. On a cache hit the binaries are present but the libraries are missing, and tests fail with &lt;code&gt;Host system is missing dependencies to run browsers&lt;/code&gt;. So you install the libraries on a cache hit too, just without the download:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="s"&gt;yaml&lt;/span&gt;

&lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Cache Playwright browsers&lt;/span&gt;
    &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;playwright-cache&lt;/span&gt;
    &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/cache@v6&lt;/span&gt;
    &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;~/.cache/ms-playwright&lt;/span&gt;
      &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ runner.os }}-playwright-${{ hashFiles('package-lock.json') }}&lt;/span&gt;

  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Install Playwright browsers and system dependencies&lt;/span&gt;
    &lt;span class="na"&gt;if&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;steps.playwright-cache.outputs.cache-hit != 'true'&lt;/span&gt;
    &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npx playwright install --with-deps&lt;/span&gt;

  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Install OS dependencies for browsers&lt;/span&gt;
    &lt;span class="na"&gt;if&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;steps.playwright-cache.outputs.cache-hit == 'true'&lt;/span&gt;
    &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npx playwright install-deps&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;id&lt;/code&gt; on the cache step is what lets the two install steps read &lt;code&gt;cache-hit&lt;/code&gt; and pick a branch. On a cache miss, &lt;code&gt;playwright install --with-deps&lt;/code&gt; downloads browsers and libraries together. On a cache hit, &lt;code&gt;playwright install-deps&lt;/code&gt; runs instead, every time, installing only the libraries. It's fast because there's no download, just a dependency check apt has already satisfied. Each job starts from a clean machine, so this isn't a one-off step you skip later.&lt;/p&gt;

&lt;p&gt;One alternative skips caching: run the job inside the official &lt;code&gt;mcr.microsoft.com/playwright&lt;/code&gt; image, which ships with browsers and OS packages preinstalled. You're trading the browser download for an image pull rather than removing it. It's a common choice for visual regression work where consistent rendering matters, but for most suites the cache is lighter.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fix 2: run tests in parallel with workers
&lt;/h2&gt;

&lt;p&gt;The 2m 25s test run is Playwright doing exactly what the config told it to. Two changes fix it.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;fullyParallel: true&lt;/code&gt; lets tests within a single file run across workers. By default Playwright parallelizes separate files but runs tests inside one file in order, so a suite concentrated in a few large files leaves most workers idle. The second change replaces the hard one-worker cap with a percentage:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;typescript&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="nf"&gt;defineConfig&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;fullyParallel&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// top-level, applies to every project&lt;/span&gt;
  &lt;span class="na"&gt;workers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;CI&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;50%&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;undefined&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// top-level, not inside a project&lt;/span&gt;
  &lt;span class="na"&gt;use&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="cm"&gt;/* ... */&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;projects&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="cm"&gt;/* chromium, firefox, webkit */&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both keys sit at the top level of &lt;code&gt;defineConfig&lt;/code&gt;, alongside &lt;code&gt;use&lt;/code&gt; and &lt;code&gt;projects&lt;/code&gt;, not nested inside any one project.&lt;/p&gt;

&lt;p&gt;On the worker count, one assumption gets repeated: that GitHub's free runners are 2-core. That's true for private repos, but public repos now get 4-core Linux runners for free. On a 2-core runner &lt;code&gt;undefined&lt;/code&gt; resolves to one worker, so the default is effectively serial. On a 4-core runner the same &lt;code&gt;undefined&lt;/code&gt; gives you two workers out of the box.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Runner&lt;/th&gt;
&lt;th&gt;Cores&lt;/th&gt;
&lt;th&gt;50% gives you&lt;/th&gt;
&lt;th&gt;Worth trying&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Private repo, ubuntu-latest&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;1 worker&lt;/td&gt;
&lt;td&gt;Set to 2 explicitly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Public repo, ubuntu-latest&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;2 workers&lt;/td&gt;
&lt;td&gt;75% or 100%, then measure&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Larger runner&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;4 workers&lt;/td&gt;
&lt;td&gt;75% and up, more headroom&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Start at 50%, look at wall-clock time on your suite, and raise it in steps until the run stops getting faster or the &lt;a href="https://endform.dev/blog/playwright-flaky-tests" rel="noopener noreferrer"&gt;tests start going flaky&lt;/a&gt;. There's no single correct number, since it depends on core count, how heavy your tests are, and how much memory each browser instance eats. The thing not to do is push workers well above your core count. Once workers contend for the same cores and RAM, you trade a faster run for memory pressure and intermittent failures that didn't exist serially.&lt;/p&gt;

&lt;p&gt;With &lt;code&gt;fullyParallel&lt;/code&gt; on and the cap lifted, the test run on a 4-core runner drops to roughly half, and the install step is already cached. Those two fixes together pull the whole job under five minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fix 3: run only the browsers you need
&lt;/h2&gt;

&lt;p&gt;The baseline was hiding something. The default config runs your tests against Chromium, Firefox, and WebKit, so those 40 tests were actually 120 test runs on every push. A good share of that 2m 25s was the same assertions executing twice more against browsers that rarely surface a bug the first one missed.&lt;/p&gt;

&lt;p&gt;If your product doesn't need every browser verified on every change, tier it by event. Run Chromium on every PR, where the feedback loop needs to be fast and most rendering and logic bugs show up. Save the full three-browser sweep for merges to main or a nightly schedule. It's a single flag switched on the triggering event:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="s"&gt;yaml&lt;/span&gt;

&lt;span class="na"&gt;on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;pull_request&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;push&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;branches&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;main&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;

&lt;span class="na"&gt;jobs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;test&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-latest&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="c1"&gt;# checkout, setup-node, cache, and install from the earlier sections&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Run Playwright tests&lt;/span&gt;
        &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
          &lt;span class="s"&gt;if [ "${{ github.event_name }}" = "pull_request" ]; then&lt;/span&gt;
            &lt;span class="s"&gt;npx playwright test --project=chromium&lt;/span&gt;
          &lt;span class="s"&gt;else&lt;/span&gt;
            &lt;span class="s"&gt;npx playwright test&lt;/span&gt;
          &lt;span class="s"&gt;fi&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On a PR this runs Chromium alone. On a push to main it falls through to the full set. This keeps everything in one job. Local runs are unaffected, since the conditionals key off the CI event.&lt;/p&gt;

&lt;p&gt;If you'd rather run the three browsers as separate parallel jobs on main, a conditional matrix does that:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="s"&gt;yaml&lt;/span&gt;

&lt;span class="na"&gt;jobs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;test&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-latest&lt;/span&gt;
    &lt;span class="na"&gt;strategy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;fail-fast&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
      &lt;span class="na"&gt;matrix&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;project&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ github.event_name == 'push'&lt;/span&gt;
          &lt;span class="s"&gt;&amp;amp;&amp;amp; fromJSON('["chromium", "firefox", "webkit"]')&lt;/span&gt;
          &lt;span class="s"&gt;|| fromJSON('["chromium"]') }}&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="c1"&gt;# same setup steps&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npx playwright test --project=${{ matrix.project }}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A push to main fans out into three jobs, one per browser; a PR collapses to a single Chromium job. It stays bounded at three jobs, so it doesn't drift into open-ended job-splitting.&lt;/p&gt;

&lt;p&gt;One caveat on WebKit: the build Playwright runs on Linux tracks WebKit's main development branch, often ahead of what Apple ships in Safari, and it lacks Apple's platform integrations. A green Linux WebKit run is a good signal for layout and JavaScript bugs, but it's not proof the page looks right in Safari. When that distinction matters you need macOS runners, which is a separate decision from this workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Get the failure report off the runner
&lt;/h2&gt;

&lt;p&gt;Playwright writes an HTML report on every run, but a CI runner is torn down when the job ends, so by default that report vanishes and you never see it. Teams usually discover this on their first failing run, when they go looking for the detail behind a red X and find nothing left to open.&lt;/p&gt;

&lt;p&gt;Upload it as an artifact before the job exits:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="s"&gt;yaml&lt;/span&gt;

&lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Upload Playwright report&lt;/span&gt;
    &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/upload-artifact@v7&lt;/span&gt;
    &lt;span class="na"&gt;if&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ !cancelled() }}&lt;/span&gt;
    &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;playwright-report&lt;/span&gt;
      &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;playwright-report/&lt;/span&gt;
      &lt;span class="na"&gt;retention-days&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;14&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;if: ${{ !cancelled() }}&lt;/code&gt; is doing real work. Without it, the step only runs when everything before it passed, which is backwards, since the run you most want the report for is the one that failed. This uploads whether tests passed or failed and skips only on a manual cancel. Use &lt;code&gt;retention-days: 14&lt;/code&gt; for active development and &lt;code&gt;30&lt;/code&gt; on release branches.&lt;/p&gt;

&lt;p&gt;The artifact gives you the full report but costs a few clicks. You can't open the HTML file directly, since it loads its data over HTTP, so you download the zip, unzip it, and run &lt;code&gt;npx playwright show-report path/to/playwright-report&lt;/code&gt;. Fine when you need a trace, more friction than most failures deserve.&lt;/p&gt;

&lt;p&gt;That's where a second reporter comes in. The &lt;code&gt;github&lt;/code&gt; reporter writes failures as inline annotations in the Actions UI, so a failed assertion shows up on the run summary and against the offending line in the diff, no download in the loop. Add it alongside the HTML reporter rather than replacing it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;typescript&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="nf"&gt;defineConfig&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;reporter&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;html&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;github&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]],&lt;/span&gt;
  &lt;span class="c1"&gt;// ... rest of your config&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now a failed run annotates the summary with what broke and where, and the full HTML report is still attached for the times you need the trace, screenshots, and step timeline.&lt;/p&gt;

&lt;h2&gt;
  
  
  Passing config and secrets
&lt;/h2&gt;

&lt;p&gt;If your tests need a base URL or an API token, pass it through the test step with an &lt;code&gt;env&lt;/code&gt; block. Non-sensitive values can go straight in the workflow; anything secret goes in a repository secret and gets referenced:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="s"&gt;yaml&lt;/span&gt;

&lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Run Playwright tests&lt;/span&gt;
    &lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;BASE_URL&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://staging.example.com&lt;/span&gt;
      &lt;span class="na"&gt;API_TOKEN&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.API_TOKEN }}&lt;/span&gt;
    &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="c1"&gt;# ... as above&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  When to add sharding (and when not to)
&lt;/h2&gt;

&lt;p&gt;Sharding splits your suite across several CI jobs running in parallel, each handling a slice, with a final step merging the reports. It's the standard answer to a slow Playwright suite, and it works, but it solves a problem you don't have until a single runner genuinely can't keep up.&lt;/p&gt;

&lt;p&gt;The decision rule: start with workers and &lt;code&gt;fullyParallel&lt;/code&gt;, tune the worker count, and measure your best single-runner time. If you've maxed that and still need more speed, try a bigger runner before sharding. An 8 or 16-core GitHub-hosted runner gives you more workers without the merge-reports step or shard-count maintenance. Only reach for sharding when one runner, even a larger one, can't finish inside your PR review window.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Suite size&lt;/th&gt;
&lt;th&gt;Single-runner time&lt;/th&gt;
&lt;th&gt;What to do&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Under 50 tests&lt;/td&gt;
&lt;td&gt;Well under 5 min&lt;/td&gt;
&lt;td&gt;Workers and fullyParallel are plenty&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;50–200 tests&lt;/td&gt;
&lt;td&gt;Around 3 to 5 min&lt;/td&gt;
&lt;td&gt;Move to a bigger runner before you shard&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Over 200 tests, or 10+ min&lt;/td&gt;
&lt;td&gt;Past your PR window&lt;/td&gt;
&lt;td&gt;Shard, or move the run off your own CI&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Run time is what matters here, not test count. A suite of heavy end-to-end flows hits the ceiling sooner than the same number of light page checks. Sharding buys wall-clock time; it costs more CI minutes, a merge-reports step, and a shard count you revisit as the suite grows. That's maintenance you take on deliberately, not a default the moment a run feels slow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common setup mistakes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Installing browsers without &lt;code&gt;--with-deps&lt;/code&gt;.&lt;/strong&gt; &lt;code&gt;npx playwright install&lt;/code&gt; alone fetches binaries but not the system libraries, and runners don't ship those. Tests fail with a "Host system is missing dependencies" error that says nothing about the cause. Use &lt;code&gt;--with-deps&lt;/code&gt; on a cache miss and &lt;code&gt;install-deps&lt;/code&gt; on a cache hit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Caching &lt;code&gt;node_modules&lt;/code&gt; instead of the npm cache.&lt;/strong&gt; A restored &lt;code&gt;node_modules&lt;/code&gt; can carry platform-specific or partially built dependencies that don't match the runner, surfacing as subtle failures. Cache &lt;code&gt;~/.npm&lt;/code&gt; by setting &lt;code&gt;cache: 'npm'&lt;/code&gt; on the setup-node step and let &lt;code&gt;npm ci&lt;/code&gt; rebuild cleanly each run.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Setting workers too high for the runner.&lt;/strong&gt; More workers stop helping once they outnumber the cores, and past that point the browser instances contend for CPU and memory, trading speed for flakiness. Scale to the runner and raise only as far as your measurements stay stable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Forgetting the &lt;code&gt;github&lt;/code&gt; reporter.&lt;/strong&gt; Without it, a failed run tells you something broke but not what. Adding &lt;code&gt;['github']&lt;/code&gt; alongside &lt;code&gt;['html']&lt;/code&gt; puts the failure inline on the summary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrapping up
&lt;/h2&gt;

&lt;p&gt;A &lt;a href="https://endform.dev/blog/the-fastest-playwright-runner" rel="noopener noreferrer"&gt;fast Playwright pipeline&lt;/a&gt; on GitHub Actions doesn't require sharding. Caching the browser binaries and tuning parallelism for your runner are both afternoon-sized config changes, and they pay off on every PR rather than once. The single-runner setup holds for a good while as a suite grows, and outgrowing it is a scaling decision to make on purpose when your numbers cross the line, not evidence something was done wrong.&lt;/p&gt;

&lt;p&gt;If your suite has crossed that line and you don't want to own shard counts, matrix jobs, and report merging, &lt;a href="https://endform.dev" rel="noopener noreferrer"&gt;Endform&lt;/a&gt; gives you the same parallelism without the CI plumbing.&lt;/p&gt;

</description>
      <category>playwright</category>
      <category>github</category>
      <category>testing</category>
    </item>
  </channel>
</rss>
