DEV Community

Tayguara Reis
Tayguara Reis

Posted on AI-assisted

BDD with Playwright: Gherkin scenarios that business people can actually read

I have written and maintained end-to-end suites with Cucumber on top of Playwright, using the cucumber-js runner. It works, but two risks show up in every Gherkin suite. Either the feature files turn into click-scripts with a Given/When/Then prefix ("When I click the button with id continue"), so nobody outside engineering reads them, or the step layer grows into a second framework with its own world object, its own retries and its own reporting, rebuilt next to a test runner that already does all of that.

Gherkin is worth it only when a product owner can read a scenario and say "yes, that is the rule." Everything below comes from a small public repository I maintain, playwright-qa-showcase, which runs 10 Gherkin scenarios (16 tests once the outlines expand) against the SauceDemo store with Playwright 1.63 and playwright-bdd 9.2.1.

Why playwright-bdd instead of cucumber-js

playwright-bdd does not run your scenarios in a separate runner. bddgen compiles the .feature files into ordinary Playwright spec files, and the Playwright test runner executes them. That one decision gives you everything the runner already does well: fixtures, full parallelism, traces, the HTML report, retries, sharding and blob-report merging in CI.

The whole BDD setup is a few lines in playwright.config.ts:

// BDD is used for the UI project only. `bddgen` turns the .feature files into Playwright specs.
const uiTestDir = defineBddConfig({
  features: 'features/**/*.feature',
  steps: ['features/steps/**/*.ts', 'src/fixtures/ui.fixtures.ts'],
  outputDir: '.features-gen/ui',
});
Enter fullscreen mode Exit fullscreen mode

The ui project then points testDir at uiTestDir, and the same config keeps trace: 'on-first-retry' in CI and 'retain-on-failure' locally. Gherkin tags become native Playwright tags, so npx playwright test --grep @smoke selects the two smoke scenarios with no extra tooling.

Where BDD pays off, and where it does not

The repository has four Playwright projects, and only one of them uses Gherkin:

  • UI business flows (login, inventory, checkout): Gherkin. A readable scenario is useful to non-engineers, and the steps are reused across features.
  • API tests: plain Playwright. Their value is in assertions on status codes, headers and payload contracts. "Then the response status is 200" adds a translation layer without adding clarity.
  • Accessibility checks: plain Playwright. The interesting output is a list of axe rule ids per page, not a sentence.
  • Unit tests for test helpers: plain Playwright, no browser.

This is the first judgment call I make on any engagement. BDD is a communication tool. Where there is no one to communicate with outside engineering, it is overhead.

Writing steps in business language

Here is the core checkout scenario:

@smoke
Scenario: A shopper completes a purchase end to end
  Given the cart contains the following products:
    | product               |
    | Sauce Labs Backpack   |
    | Sauce Labs Bike Light |
  When I open the cart
  And I start the checkout
  And I submit valid shipping information
  Then the order overview lists the same products
  And the item total equals the sum of the item prices
  And the order total equals the item total plus tax
  When I finish the order
  Then I see the order confirmation "Thank you for your order!"
  And the cart badge is not shown
Enter fullscreen mode Exit fullscreen mode

There are no selectors, no field names and no test data that does not matter to the rule. "I submit valid shipping information" is declarative: the step knows what valid means (a synthetic Ada / Lovelace / 12345 record in src/data/checkoutData.ts). The imperative version, three "When I fill ..." lines, would tell a reader nothing new and break the scenario every time the form changes.

When the variation is the rule, I spell it out, and a Scenario Outline keeps it to one scenario:

Scenario Outline: Required shipping fields are validated: <case>
  ...
  When I submit the shipping form with first name "<first_name>", last name "<last_name>" and postal code "<postal_code>"
  Then I see the checkout error "<error>"
  And I am still on the shipping information step

  Examples:
    | case                | first_name | last_name | postal_code | error                          |
    | missing first name  |            | Lovelace  | 12345       | Error: First Name is required  |
    | missing last name   | Ada        |           | 12345       | Error: Last Name is required   |
    | missing postal code | Ada        | Lovelace  |             | Error: Postal Code is required |
Enter fullscreen mode Exit fullscreen mode

Two smaller choices matter as much as the wording. In the sorting outline, "Name (A to Z)" is deliberately left out of the examples, with a comment explaining why: it is the default order, so it would pass even if the sort control did nothing. And the login outline says "a password" (valid, wrong, empty) instead of putting the password in the feature file. The step maps the kind to a value, and the real password stays in config/env.ts.

Fixtures and page objects, injected into steps

Steps never construct page objects. They receive them as Playwright fixtures, which createBdd binds to Given/When/Then:

export const test = base.extend<UiFixtures>({
  inventoryPage: async ({ page }, use) => {
    await use(new InventoryPage(page));
  },
  checkoutOverviewPage: async ({ page }, use) => {
    await use(new CheckoutOverviewPage(page));
  },
  // ...one fixture per page object, plus the header component
  scenario: async ({}, use) => {
    await use({ cartProducts: [] });
  },
});

export const { Given, When, Then } = createBdd(test);
Enter fullscreen mode Exit fullscreen mode

The scenario fixture replaces the Cucumber "world". It holds state shared between the steps of one scenario, here the products added to the cart, and because it is test-scoped it cannot leak into another test. A step asks only for what it needs:

When('I add {string} to the cart', async ({ inventoryPage, scenario }, productName: string) => {
  await inventoryPage.addToCart(productName);
  scenario.cartProducts.push(productName);
});

Then('the order overview lists the same products', async ({ checkoutOverviewPage, scenario }) => {
  await expect(checkoutOverviewPage.productNames).toHaveText(scenario.cartProducts);
});
Enter fullscreen mode Exit fullscreen mode

No globals, no shared mutable module state, and every test runs in its own browser context, so fullyParallel: true is safe. Page objects own locators (role and data-test attributes, never CSS paths), and assertions are web-first, so there are no hard waits.

Free-form strings from Gherkin are also narrowed at the boundary. toSauceUser(username) turns "standard_user" into a typed union member and throws with the list of known users if a scenario has a typo, instead of failing later on a confusing login error.

Login: through the UI only where login is under test

Logging in through the form in every scenario is slow and makes the login page a failure point for tests that have nothing to do with it. SauceDemo keeps its session in a plain session-username cookie, so outside login.feature the suite sets the cookie and opens the inventory page:

export async function loginViaSession(page: Page, username: SauceUser): Promise<void> {
  await page
    .context()
    .addCookies([{ name: 'session-username', value: username, url: env.sauce.baseUrl }]);
  await page.goto('/inventory.html');
}
Enter fullscreen mode Exit fullscreen mode

I validated this against the live site before relying on it, and the Given I am logged in as "..." step still asserts that the Products page is shown. If the site ever stops honoring the cookie, only this function changes. login.feature keeps a separate step, Given I am logged in as "standard_user" using the login form, for the logout scenario, where the real form is part of what is being tested. In a real application the equivalent would be an API login or a stored storageState. The principle is the same.

Asserting business rules, not just screens

"The order total equals the item total plus tax" is a business rule, so the step must check the arithmetic, not only that a total is visible. Money is never compared as floats (0.1 + 0.2 !== 0.3). Everything goes through integer cents:

export function toCents(amount: number): number {
  return Math.round(amount * 100);
}

export function sumCents(amounts: readonly number[]): number {
  return amounts.reduce((total, amount) => total + toCents(amount), 0);
}
Enter fullscreen mode Exit fullscreen mode
Then('the order total equals the item total plus tax', async ({ checkoutOverviewPage }) => {
  const itemTotal = await checkoutOverviewPage.itemTotal();
  const tax = await checkoutOverviewPage.tax();

  expect(toCents(await checkoutOverviewPage.total())).toBe(sumCents([itemTotal, tax]));
});
Enter fullscreen mode Exit fullscreen mode

The item total step does the same against the sum of the line prices. These helpers are pure functions, so they have their own unit tests in the plain Playwright unit project.

Known defects: tracked with @fail, not skipped

SauceDemo's problem_user is broken on purpose: every product shows the same placeholder image. Instead of skipping a test, I wrote the scenario that asserts the correct behavior and tagged it @fail:

@fail
Scenario: problem_user sees a distinct image for each product
  Given I am logged in as "problem_user"
  Then every product shows its own image
Enter fullscreen mode Exit fullscreen mode

playwright-bdd turns @fail into test.fail() in the generated spec. Today the scenario is reported as an expected failure. If the defect is fixed, Playwright reports it as "unexpectedly passed", which is the signal to remove the tag. The defect stays visible in every run.

The honest limitation: @fail accepts any failure. If the site is down or a selector drifts, this scenario still "passes" as an expected failure. It relies on the other UI tests to prove the site is up. A stricter version would assert the specific failure, but for one tracked defect in a public sandbox I chose the simpler mechanism and documented the trade-off in docs/KNOWN_ISSUES.md, along with the problem_user defects I verified by hand but did not automate.

Keeping the suite honest in CI

A Gherkin suite rots quietly in two ways: a step loses its definition, or someone disables a scenario "for now". The CI pipeline has a dedicated Gherkin job, which needs no browser and runs in parallel with lint and typecheck:

# Fails fast if a Gherkin step has no definition.
- run: npx bddgen
# Lint covers test.skip in TypeScript. This covers the same escape hatch in Gherkin.
- name: Reject disabled scenarios
  run: |
    if grep -rnE '(^|[[:space:]])@(skip|fixme|only)([[:space:]]|$)' features --include='*.feature'; then
      echo '::error::A scenario is disabled with @skip, @fixme or @only. Fix it or track it with @fail.'
      exit 1
    fi
Enter fullscreen mode Exit fullscreen mode

The error message states the policy: fix it, or track it with @fail. Only after this and the other fast checks pass do the browser tests run, sharded in two, with blob reports merged into a single HTML report.

Takeaways

  • Use Gherkin where someone outside engineering will read it. Keep API, accessibility and unit tests in plain Playwright.
  • Run Gherkin on the Playwright runner. You keep fixtures, parallelism, traces, reports and sharding instead of rebuilding them.
  • Write declarative steps. Spell out values only when the variation is the business rule, and use Scenario Outlines for that.
  • Inject page objects and scenario state as fixtures. No globals, no shared world object.
  • Never skip silently: track known defects with @fail, know its limits, and let CI reject @skip, @fixme and @only.

The full project, including the API and accessibility suites, is on GitHub: github.com/tayguara/playwright-qa-showcase.


Tayguara Dias Reis is a Lead QA / SDET with 14+ years of experience and ISTQB CTFL certification. The code in this article is from playwright-qa-showcase.

Top comments (0)