DEV Community

Silviu Technology
Silviu Technology

Posted on

Playwright Retry Budgets Need a Failure Receipt

Retries are useful when a browser test meets a temporary network problem. They become harmful when every failure gets three more attempts with no record of what happened. A retry can turn one actionable defect into a green build that nobody trusts.

For Playwright suites, I find a small retry budget and a consistent failure receipt much more useful than unlimited patience. The goal is not to eliminate retries. It is to make every retry explain itself.

The problem with unlimited retries

Imagine a signup test that creates a user, waits for an email, and completes verification. The first attempt fails because the message is delayed. The second attempt passes. CI reports green, but the suite has hidden a latency problem and may have left a user or mailbox behind.

This gets harder when the fixture uses a temp mail so address or a free throwaway email. The test may be checking three systems at once: the browser, the application API, and the email provider. A generic retry: 2 setting does not tell us which boundary was unstable.

Retries also multiply side effects. A test that sends an invitation, creates a payment intent, or consumes a one-time token is not automatically safe to repeat. Before increasing retries, ask whether the operation is idempotent and whether cleanup runs after a failed attempt. Otherwise the retry is a second experiment with different data, not a repeat of the first one.

Define a retry budget per failure class

Start with a simple classification:

  • Browser timing: a locator appeared late or an animation was still running. One retry can be reasonable.
  • External dependency: an email, webhook, or sandbox API was slow. Retry only after recording the wait and response details.
  • Product assertion: the UI showed the wrong state. Retry zero times until the failure is understood.
  • Test isolation: another run reused data or a mailbox. Fix the fixture and retry zero times.

This makes the budget a diagnostic decision. A flaky locator may deserve one controlled retry, while a failed business assertion should produce an immediate failure receipt. It also keeps the test report honest: a passing retry is evidence of instability, not evidence that the scenario is healthy.

Capture one useful failure receipt

For each attempt, save enough information to answer five questions:

  1. Which test, project, and CI run was executing?
  2. Which attempt number failed?
  3. What was the last meaningful application event?
  4. Which data and email fixture belonged to this run?
  5. What artifacts should a developer open first?

Playwright already provides a strong base with traces, screenshots, videos, and test attachments. The missing piece is often the small JSON receipt around those files. Keep secrets out of it: never store an inbox password, a verification token, or a full authorization header. A redaction boundary for OAuth callback logs is useful for test artifacts too.

A receipt can be as small as:

{
  "test": "signup verifies email",
  "attempt": 1,
  "fixture_id": "signup-1842",
  "last_event": "verification message not visible after polling",
  "artifacts": ["trace.zip", "page.png", "events.json"],
  "retryable": true
}
Enter fullscreen mode Exit fullscreen mode

The fixture_id should identify the run without exposing the mailbox contents. When teams search for tempail mail or tamp mail com during fixture setup, that search phrase should never become a secret or a test credential in the receipt.

A Playwright implementation

Keep retry policy close to the test runner, but keep classification in the test fixture or helper. Here is a compact pattern for attaching a receipt after a failure:

import { test as base } from '@playwright/test';

export const test = base.extend({
  failureReceipt: async ({}, use, testInfo) => {
    const receipt = {
      test: testInfo.title,
      attempt: testInfo.retry,
      project: testInfo.project.name,
      retryable: testInfo.status !== 'passed',
    };

    await use(receipt);
    await testInfo.attach('failure-receipt', {
      body: JSON.stringify(receipt, null, 2),
      contentType: 'application/json',
    });
  },
});
Enter fullscreen mode Exit fullscreen mode

In a real suite, add the fixture ID and last event from your application test harness. Do not catch every error and mark it retryable. That shortcut makes the report look better, but the debugging gets slower.

Set a modest global retry count, then override it only for tests with a documented reason:

// playwright.config.ts
export default defineConfig({
  retries: process.env.CI ? 1 : 0,
  reporter: [['html'], ['list']],
});
Enter fullscreen mode Exit fullscreen mode

The first run should collect the detailed trace. If the retry passes, preserve the first attempt artifacts and label the test as recovered. A recovered test is a maintenance signal for the team.

How to classify the next failure

Read the receipt before reading the final assertion. If the last event says the verification message was never observed, inspect polling intervals, provider latency, and fixture cleanup. If the trace shows the button was clicked twice, inspect the action's idempotency. If the request returned a stable 4xx response, remove the retry and fix the contract.

This is also where operational habits help. Tests that exercise deployment environments need clear signal routing, much like signal routing for EKS rollbacks. The failure should reach the person who can fix the boundary, not disappear into a generic flaky-test bucket.

CI checklist

  • Give each run a unique fixture ID and isolated email data.
  • Keep the default retry budget at zero locally and one in CI only when useful.
  • Attach a trace, screenshot, and small redacted JSON receipt.
  • Preserve artifacts from the first failed attempt when a retry passes.
  • Classify assertion failures separately from dependency timeouts.
  • Make cleanup idempotent and safe to run after partial setup.
  • Review recovered tests weekly; they are not fully healthy tests.

A retry budget turns flaky behavior into visible evidence. Once every attempt has a small, privacy-aware receipt, Playwright failures stop being mysterious red screens and become a repeatable QA conversation.

Top comments (0)