DEV Community

Cover image for Running Playwright in GitHub Actions: named checks, sharding, one merged report, and a flakiness policy
Tayguara Reis
Tayguara Reis

Posted on AI-assisted

Running Playwright in GitHub Actions: named checks, sharding, one merged report, and a flakiness policy

A test pipeline has one job: give the team a fast signal it can trust. When a pull request goes red, the author should know what broke without opening a log. When it goes green, nobody should wonder whether a retry quietly hid a problem.

This article walks through the pipeline of playwright-qa-showcase, a small public repository I use to show how I set up test automation when I lead QA on a project. It has 56 tests: BDD UI flows against SauceDemo, typed API tests against Restful Booker, axe-core accessibility scans, and unit tests for the test-support code. Every snippet below is taken from the real files, shortened for reading. Every number comes from a real run.

One check per validation

My first version had a single Lint, typecheck, format job. It worked, but a red check told you only that "something in quality" failed. So I split it into five parallel jobs, each with its own name:, and each shows up as its own check on the pull request:

jobs:
  lint:
    name: Lint
    runs-on: ubuntu-latest
    timeout-minutes: 10
    steps:
      - uses: actions/checkout@v7
        with:
          persist-credentials: false
      - uses: actions/setup-node@v7
        with:
          node-version-file: .nvmrc
          cache: npm
      - run: npm ci
      - run: npm run lint
  # typecheck (Typecheck), format (Format), unit (Unit tests), gherkin (Gherkin): same shape
Enter fullscreen mode Exit fullscreen mode

None of the five needs a browser. Unit tests runs the pure-Node Playwright project that covers the helpers (money math, the accessibility baseline diff, a test data builder). Gherkin runs bddgen, which fails when a step has no definition, and then a guard I will come back to. In the real run of the PR that introduced this split, the five jobs finished in 10 to 18 seconds each.

The browser tests wait for all of them:

  test:
    name: Tests (shard ${{ matrix.shardIndex }}/${{ matrix.shardTotal }})
    needs: [lint, typecheck, format, unit, gherkin]
Enter fullscreen mode Exit fullscreen mode

Duplicating checkout, setup-node and npm ci five times looks wasteful, and on paper it is. In practice the npm cache makes each copy cheap, and the payoff is a PR page that reads like a checklist. A composite action could remove the repetition. At five short jobs I chose readability over DRY.

Sharding, and an honest note about it

The browser suites (ui, a11y, api) run as a two-shard matrix:

    strategy:
      fail-fast: false
      matrix:
        shardIndex: [1, 2]
        shardTotal: [2]
    steps:
      # checkout, setup-node, npm ci
      - run: npx playwright install --with-deps chromium
      - run: npx bddgen
      - name: Run Playwright tests
        run: npx playwright test --project=ui --project=a11y --project=api --shard=${{ matrix.shardIndex }}/${{ matrix.shardTotal }}
      - name: Upload blob report
        if: ${{ !cancelled() }}
        uses: actions/upload-artifact@v7
        with:
          name: blob-report-${{ matrix.shardIndex }}
          path: blob-report
          retention-days: 1
          if-no-files-found: error
Enter fullscreen mode Exit fullscreen mode

Three details matter. fail-fast: false keeps one failing shard from cancelling the other, so you see every failure in one run instead of discovering them one at a time. The unit project is left out of the shards because it needs no browser and already ran in its own job. And if-no-files-found: error makes a shard that produced no report fail loudly instead of silently shrinking the merged report.

Now the honest part. At this size, sharding demonstrates the pattern; it does not save time. Each shard ran 14 tests. In that same run, the Playwright step took 7.3 seconds of test time per shard, while playwright install --with-deps chromium took 58 seconds on one shard and 26 on the other. A single job would finish sooner than two jobs that each install a browser. I kept the matrix because the pattern is the point of the repository, and because scaling it is a one-line change: shardIndex: [1, 2, 3, 4] with shardTotal: [4]. On a real project I would add shards when test time, not setup time, dominates the job.

One merged report, plus a summary nobody has to download

Sharding creates a reporting problem: two partial reports are worse than one. In CI the config uses the blob reporter (reporter: isCI ? [['blob'], ['github'], ['list']] : ...), and a final job merges the blobs:

  merge-reports:
    name: Merge reports
    if: ${{ !cancelled() && needs.test.result != 'skipped' }}
    needs: test
    steps:
      # checkout, setup-node, npm ci
      - uses: actions/download-artifact@v8
        with:
          path: all-blob-reports
          pattern: blob-report-*
          merge-multiple: true
      - name: Merge into one HTML report
        env:
          PLAYWRIGHT_HTML_OPEN: never
          PLAYWRIGHT_JSON_OUTPUT_NAME: ${{ runner.temp }}/merged-results.json
        run: npx playwright merge-reports --reporter html,json ./all-blob-reports
Enter fullscreen mode Exit fullscreen mode

The if: condition is the part people get wrong. The default success() would skip this job exactly when you need the report most, after a shard failed. always() goes too far: when Lint fails, the shards are skipped and there is nothing to merge. !cancelled() && needs.test.result != 'skipped' covers both cases.

The HTML report is uploaded as an artifact kept for 14 days. Downloading an artifact is friction, though, so a short inline Node script reads the merged JSON and writes a table to the job summary:

const { stats } = JSON.parse(fs.readFileSync(process.env.RESULTS_JSON, "utf8"));
const rows = [
  ["Passed", stats.expected],
  ["Failed", stats.unexpected],
  ["Flaky (passed on retry)", stats.flaky],
  ["Skipped", stats.skipped],
];
Enter fullscreen mode Exit fullscreen mode

The "Flaky" row is there on purpose. It is the bridge to the flakiness policy.

A flakiness policy, enforced by config and lint

"We don't tolerate flaky tests" is a slogan. A policy is something the tooling enforces. From playwright.config.ts:

forbidOnly: isCI,
retries: isCI ? 2 : 0,
use: {
  // Retries are off locally, so keep the trace of a failure instead of waiting for a retry.
  trace: isCI ? 'on-first-retry' : 'retain-on-failure',
  screenshot: 'only-on-failure',
},
Enter fullscreen mode Exit fullscreen mode

Retries exist only in CI, so a flaky test is visible while you develop instead of being absorbed. on-first-retry would never fire locally with zero retries, so locally the trace is kept on failure. forbidOnly stops a stray test.only from turning a full run into a one-test run.

The trade-off I state openly: a test that passes on retry leaves the run green. That is why the summary shows the flaky count, and why the policy says it gets investigated. Someone has to own that number.

Lint closes the usual escape hatches, with eslint --max-warnings=0:

rules: {
  'playwright/no-wait-for-timeout': 'error',
  'playwright/no-skipped-test': 'error',
  'playwright/no-force-option': 'error',
},
Enter fullscreen mode Exit fullscreen mode

Hard waits, forced clicks and skipped tests are the three most common ways a flaky test gets "fixed" without being fixed. Gherkin has the same escape hatch in tags, so the Gherkin job greps for it:

      - name: Reject disabled scenarios
        run: |
          if grep -rnE '(^|[[:space:]])@(skip|fixme|only)([[:space:]]|$)' features --include='*.feature'; then
            echo '::error::A scenario is disabled with @skip, @fixme or @only. Fix it or track it with @fail.'
            exit 1
          fi
Enter fullscreen mode Exit fullscreen mode

A known product defect is tracked with @fail, an expected failure that asserts the correct behavior, rather than hidden with @skip. If the site gets fixed, Playwright reports it as unexpectedly passing.

Supply chain and security, sized for a test repo

Test repositories run code on every pull request, so they deserve the same hygiene as product code:

  • permissions: contents: read at the workflow level. Only the CodeQL job gets security-events: write.
  • persist-credentials: false on every checkout, so the token is not left in .git/config for later steps.
  • No pull_request_target.
  • Pinned action versions. dependency-review-action is pinned to v5.0.0 because the project publishes no floating v5 tag.
  • A concurrency group with cancel-in-progress: true, so a new push cancels the superseded run.

Three extra workflows add checks. codeql.yml analyzes both javascript-typescript and actions (the workflow files themselves), on PRs, on pushes to main and weekly. dependency-review.yml fails a PR that adds a dependency with a known vulnerability of moderate severity or worse. actionlint.yml lints the workflows.

Dependency review taught me a real lesson. Its first run failed with "Dependency review is not supported on this repository. Please ensure that Dependency graph is enabled". The YAML was correct; a repository setting was off. After enabling the Dependency graph, the re-run passed. A security check can depend on configuration that lives outside the repository, so check the settings page, not only the diff.

actionlint is path-filtered to .github/workflows/**, and that is exactly why it is not a required check. A workflow skipped by a path filter never reports, and a required check that never reports leaves every unrelated PR waiting on a pending status.

Branch protection, nightly runs and an honest badge

main requires ten checks: Lint, Typecheck, Format, Unit tests, Gherkin, Tests (shard 1/2), Tests (shard 2/2), CodeQL (javascript-typescript), CodeQL (actions) and Dependency review. Dependency review can be required because it runs on every pull request; actionlint cannot, for the reason above. Named jobs make this list readable, and it is also why stable job names matter: rename a job and branch protection waits for a check that no longer exists. The shard names embed the shard count, so raising the matrix means updating this list too. Merge reports stays optional, since the shards themselves already block the merge.

The Playwright workflow also runs nightly (cron: '0 6 * * *') to catch drift in the public sandboxes. Those sandboxes go down from time to time, and I do not want that to make the project look broken. The README badge is filtered to push events:

[![Playwright](https://github.com/tayguara/playwright-qa-showcase/actions/workflows/playwright.yml/badge.svg?branch=main&event=push)](...)
Enter fullscreen mode Exit fullscreen mode

The nightly run still reports failures in the Actions tab, where the team looks, while the badge reflects the state of the code on main. GitHub also disables scheduled workflows after 60 days without repository activity, which is worth knowing before trusting a nightly job on a quiet repo.

Dependabot, with two deliberate exceptions

Dependabot updates npm packages and GitHub Actions weekly, grouped into one PR per ecosystem. Two major updates are ignored:

    ignore:
      # typescript-eslint does not support TypeScript 7 yet; revisit when it does.
      - dependency-name: typescript
        update-types: ['version-update:semver-major']
      # The project runs on Node 22 (.nvmrc); @types/node majors must follow the Node version.
      - dependency-name: '@types/node'
        update-types: ['version-update:semver-major']
Enter fullscreen mode Exit fullscreen mode

The @types/node rule came from a real PR: Dependabot proposed 22.20.4 to 26.6.3, and CI passed. Green CI did not make it correct. Types for Node 26 on a Node 22 runtime let you compile calls to APIs that do not exist at runtime. I closed the PR, and closing a grouped PR does not ignore future versions, so the rule went into config. Automated updates need a human to decide which versions the project actually targets.

Takeaways

  1. Give every validation its own named job. A PR check list that reads like a checklist saves more time than the duplicated setup costs.
  2. Shard when test time dominates, not before. Measure the browser install against the test step; here it was 26 to 58 seconds of setup against about 7 seconds of tests.
  3. Merge sharded reports with !cancelled() && needs.test.result != 'skipped', and put a results table, including flaky count, in the job summary.
  4. Enforce the flakiness policy in tooling: retries only in CI, forbidOnly, lint bans on hard waits, forced clicks and skips, and a guard for Gherkin tags.
  5. Treat the pipeline as product code: least-privilege tokens, pinned actions, required checks with stable names, and settings you verify, not assume.

Tayguara Dias Reis is a Tech Lead with 14+ years in software quality and ISTQB CTFL certification. The full pipeline is at github.com/tayguara/playwright-qa-showcase. For the Jenkins side of the same ideas, see From scripted to declarative: what 10+ years of Jenkins pipelines taught me about shared libraries and quality gates.

Top comments (0)