DEV Community

AutomationDataCamp
AutomationDataCamp

Posted on Originally published at automationdatacamp.com Fully Autonomous

Playwright MCP: let an AI assistant explore your app, then write the test (Claude Code, Copilot, Cursor)

Playwright MCP is Microsoft's official Model Context Protocol server for Playwright. It gives an AI assistant such as Claude Code, GitHub Copilot or Cursor the tools to drive a real browser: navigate, click, fill forms and read the page. The interesting part for testers: the assistant reads the page through structured accessibility snapshots, not screenshots, so no vision model is needed and it naturally "sees" the roles and accessible names that good Playwright locators are built on.

This post walks through installation, a first explore-then-generate scenario, how MCP compares with Playwright CLI and the Playwright Test Agents, and the rules we teach to keep AI-generated tests trustworthy.

Install it in one command

Node.js is the only prerequisite (the server runs with npx). The browser is headed by default, which is handy to watch what the assistant does.

Claude Code

claude mcp add playwright npx @playwright/mcp@latest
Enter fullscreen mode Exit fullscreen mode

VS Code (GitHub Copilot agent mode)

code --add-mcp '{"name":"playwright","command":"npx","args":["@playwright/mcp@latest"]}'
Enter fullscreen mode Exit fullscreen mode

Cursor and other MCP clients accept the standard JSON configuration:

{
  "mcpServers": {
    "playwright": {
      "command": "npx",
      "args": ["@playwright/mcp@latest"]
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

Options worth knowing (append them after @playwright/mcp@latest):

  • --headless: no visible browser (CI machines, containers)
  • --isolated: keep the profile in memory, start from a clean session every time
  • --storage-state <path>: start already logged in
  • --secrets <path>: a dotenv file, so credentials never appear in your prompt
  • --test-id-attribute: defaults to data-testid
  • --caps: optional capabilities such as testing (assertion and locator-generation tools), network (route mocking), storage, pdf, vision, devtools

A first scenario: explore, then generate

Step 1, exploration. Be precise in the prompt:

Using Playwright MCP, open http://localhost:3000/login.
Describe the fields, buttons and messages on the page.
Then try to log in with a wrong password and report the exact error message.
Do not submit any other form.
Enter fullscreen mode Exit fullscreen mode

The assistant calls browser_navigate, browser_snapshot, browser_fill_form / browser_type and browser_click, then reads the page again. You get a report of the elements (with roles and accessible labels) and of the observed behaviour: a small, documented exploratory session.

Step 2, generation.

From what you observed, write tests/login.spec.ts with Playwright Test in TypeScript:
one successful login and one wrong-password case. Use getByRole and getByLabel only,
no CSS selectors, read credentials from environment variables.
Then run the tests and fix them if they fail.
Enter fullscreen mode Exit fullscreen mode

A typical result (labels come from the demo app; yours will come from the real exploration):

import { test, expect } from '@playwright/test';

test.describe('Login', () => {
  test.beforeEach(async ({ page }) => {
    await page.goto('/login');
  });

  test('a valid user reaches the dashboard', async ({ page }) => {
    await page.getByLabel('Email').fill(process.env.TEST_USER_EMAIL!);
    await page.getByLabel('Password').fill(process.env.TEST_USER_PASSWORD!);
    await page.getByRole('button', { name: 'Sign in' }).click();

    await expect(page).toHaveURL(/\/dashboard/);
    await expect(page.getByRole('heading', { name: 'Dashboard' })).toBeVisible();
  });

  test('a wrong password shows an error', async ({ page }) => {
    await page.getByLabel('Email').fill(process.env.TEST_USER_EMAIL!);
    await page.getByLabel('Password').fill('wrong-password');
    await page.getByRole('button', { name: 'Sign in' }).click();

    await expect(page.getByRole('alert')).toContainText('Invalid credentials');
    await expect(page).toHaveURL(/\/login/);
  });
});
Enter fullscreen mode Exit fullscreen mode

Step 3, human review. Read it like a colleague's pull request. Do the assertions check a business outcome or just that a click happened? Does the error case also check the user is not logged in? Is an edge case missing (empty field, locked account)? Run it yourself and make it fail once on purpose to prove it catches a regression.

MCP, CLI or Test Agents?

Approach How it works Best for
Playwright MCP The AI calls browser_* tools and reads accessibility snapshots Exploration, self-healing tests, long autonomous loops with a persistent browser context
Playwright CLI + Skills The coding agent runs short commands described by Skills Coding agents juggling a large codebase: more token-efficient
Playwright Test Agents (1.56+) Three agents: planner, generator, healer Building or maintaining a Playwright Test suite with an assistant

The Playwright MCP README is explicit: for coding agents, CLI + Skills is often the better choice because it avoids loading large tool schemas and verbose accessibility trees into the context. MCP stays relevant when the agent needs to reason over page structure with persistent state.

Test Agents are set up in an existing project:

npx playwright init-agents --loop=claude   # or --loop=vscode, --loop=opencode
Enter fullscreen mode Exit fullscreen mode

For Claude Code this writes agent definitions in .claude/agents/ and a .mcp.json that starts npx playwright run-test-mcp-server. The planner explores the app and writes a Markdown test plan, the generator turns it into test files, the healer runs the suite and repairs failing tests (as a last resort it marks a test test.fixme() with a comment). Re-run the command after each Playwright upgrade.

Rules we teach for AI-generated tests

  1. Review every generated line. A passing test is not necessarily a good test. The AI checks what it sees; the tester defines what the business expects.
  2. Demand accessible locators (getByRole, getByLabel, getByTestId) and say so in the prompt.
  3. Own your test data. Dedicated accounts and datasets, recreated before each run.
  4. Keep secrets out of prompts and code: --secrets for exploration, environment variables for tests.
  5. Never point the agent at production. An assistant that clicks on its own can submit, delete or pay. Use a test environment and --isolated.
  6. Watch token cost. Every snapshot and tool description fills the context. Scope the exploration to one flow.
  7. Keep CI deterministic. AI writes and repairs tests; plain Playwright Test runs them in the pipeline.

What it changes for testers

Less time writing selectors and boilerplate, more time on test design (equivalence partitions, boundary values, decision tables), test data and code review. You still need to read and fix TypeScript: you can't review what you don't understand.


This article is adapted from our French guide Playwright MCP : tester avec l'IA. It was written with AI assistance; commands and options were checked against the official Playwright MCP and Playwright documentation.

AutomationDataCamp is an online software testing academy (Playwright, API testing, CI, AI-assisted testing, ISTQB prep). A ready-to-run Playwright + TypeScript starter with CI is on GitHub; more on our site.

Sources: microsoft/playwright-mcp · microsoft/playwright-cli · Playwright Test Agents · Playwright locators

Top comments (0)