A Minimal Playwright Failure Review Setup: Traces, Screenshots, and Logs
A red Playwright test is rarely enough to explain what went wrong.
The terminal might show a timeout. A screenshot might show an unexpected page. A trace might reveal a missing network response. But without a consistent failure setup, every debugging session starts with the same question:
Do we have enough evidence to decide what to check next?
This article shows a small Playwright configuration that keeps successful runs lightweight while collecting useful evidence whenever a test fails.
Start with evidence, not guesses
A failed end-to-end test can have several causes:
- A real product defect
- An outdated test expectation
- An unstable locator
- Missing or incorrect test data
- A problem in CI, the network, or an external dependency
The first debugging goal should not be to guess the root cause immediately. It should be to collect enough context to classify the failure and assign a clear next step.
For most teams, three artifacts are enough to start:
- A screenshot of the failed state
- A trace that shows the steps before the failure
- Logs or error details that explain what the test was waiting for
A practical Playwright configuration
The following configuration records screenshots, videos, and traces only when a test fails.
import { defineConfig, devices } from '@playwright/test';
export default defineConfig({
testDir: './tests',
timeout: 30_000,
expect: {
timeout: 5_000,
},
reporter: [
['html', { open: 'never' }],
['list'],
],
use: {
baseURL: process.env.BASE_URL ?? 'http://localhost:3000',
// Keep useful artifacts only for failed tests.
screenshot: 'only-on-failure',
video: 'retain-on-failure',
trace: 'retain-on-failure',
},
projects: [
{
name: 'chromium',
use: { ...devices['Desktop Chrome'] },
},
],
});
This setup avoids storing artifacts for every successful test. At the same time, a failed run produces the evidence needed for an initial review.
Why traces matter
A screenshot tells you what the page looked like at one moment. A trace provides the surrounding sequence:
- Which actions the test performed
- Which locators were used
- What the page looked like before and after each step
- Console messages and network activity
- The exact error produced by Playwright
This is especially helpful when a test fails in CI but cannot be reproduced immediately on a local machine.
To inspect a trace locally:
npx playwright show-trace path/to/trace.zip
The HTML reporter also links to available traces, screenshots, and videos.
Add meaningful test steps
A trace is only as useful as the test flow it records. Named steps make a long test easier to review.
import { test, expect } from '@playwright/test';
test('a signed-in user can update a profile', async ({ page }) => {
await test.step('Sign in as a standard user', async () => {
await page.goto('/login');
await page.getByLabel('Email').fill('user@example.test');
await page.getByLabel('Password').fill('password');
await page.getByRole('button', { name: 'Sign in' }).click();
});
await test.step('Update the display name', async () => {
await page.goto('/profile');
await page.getByLabel('Display name').fill('Alex Taylor');
await page.getByRole('button', { name: 'Save changes' }).click();
});
await test.step('Verify the confirmation message', async () => {
await expect(
page.getByText('Changes saved')
).toBeVisible();
});
});
When this test fails, the trace makes it easier to see whether the problem happened during login, data entry, saving, or the final assertion.
Use stable locators
Many false alarms originate from brittle selectors.
Prefer locators that reflect how users interact with the application:
page.getByRole('button', { name: 'Save changes' });
page.getByLabel('Email');
page.getByText('Changes saved');
For components without an accessible role or label, stable test IDs are often a better option than CSS classes or DOM structure.
page.getByTestId('profile-save-button');
A failed locator should be reviewed as a possible automation problem before it is treated as a product defect.
A simple failure review template
A test report becomes much more useful when the team records the decision that follows it.
Use a short template in an issue, pull request, or team chat:
Test:
Environment:
Observed behavior:
Expected behavior:
Available evidence:
- Screenshot:
- Trace:
- Relevant log or network response:
Likely category:
- Product
- Test
- Locator
- Test data
- Environment
Confidence:
- High / Medium / Low
Owner:
Next action:
This does not need to become paperwork. Its purpose is to prevent vague handovers such as “the login test is broken” and replace them with a concrete next step.
Keep the review lightweight
Not every failed test needs a full investigation.
A useful rule is:
- If the cause is obvious, fix it and move on.
- If the cause is unclear, classify the failure before assigning work.
- If the same class of failure repeats, improve the test setup rather than solving each occurrence in isolation.
For example, repeated failures caused by stale test data point to a data-management problem. Repeated locator failures may indicate that the application needs stable test IDs. Repeated CI-only failures may require better environment observability.
Final thought
The value of Playwright artifacts is not in collecting more files. It is in helping a team answer one practical question:
What is the next decision, and who owns it?
Screenshots, traces, and logs provide the evidence. A small, consistent review workflow turns that evidence into action.
About the authors
Melanie Kinzel is co-founder of XYVA and works on traceable QA workflows for Playwright teams.
Dominic Kinzel is co-founder of XYVA and leads its technical development.
Top comments (0)