DEV Community

Silviu Technology
Silviu Technology

Posted on

Playwright Mail Tests Need a Consumption Ledger

An email verification test can pass for the wrong reason. A shared inbox may already contain a valid message, a retry may consume the message from the first attempt, or two parallel workers may both read the same link. The final assertion is green, but the test did not prove that this run delivered the message it requested.

In QA automation, I treat every received email as a small piece of evidence with a lifecycle. A consumption ledger records which test discovered a message, which test used it, and whether the verification link was already consumed. This is a simple addition to a Playwright fixture, but it makes stale inbox failures much easier to diagnose.

Why a passing email test can still be wrong

Consider a signup test that creates an account, waits for a verification email, and opens its link. If the mailbox is shared, the test may find yesterday's message. If the subject is the only filter, another worker can provide a plausible match. If a retry starts with the same address, the second attempt can use the first attempt's evidence.

Those are false positives, not merely flaky tests. The test passed while skipping the boundary it was meant to check: that this signup produced this message for this run.

The problem becomes more visible when a team uses a temp mail so service to create isolated test data. A throwaway email generator can help with address creation, but it does not replace message identity. A mailbox can still hold multiple messages, and a polling helper can still return the wrong one.

Define the message consumption ledger

The ledger does not need to be a database. For most end-to-end suites, a run-scoped object or JSON attachment is enough. Give each expected message a correlation ID and track a few transitions:

type MailReceipt = {
  runId: string;
  mailboxId: string;
  messageId?: string;
  correlationId: string;
  state: 'requested' | 'observed' | 'claimed' | 'opened' | 'expired';
  observedAt?: string;
  claimedBy?: string;
};
Enter fullscreen mode Exit fullscreen mode

The useful distinction is between observed and claimed. An inbox helper can observe that a message exists, but only one test attempt should claim it for verification. If a retry or another worker sees the same message, the receipt can explain that it was already claimed instead of silently returning it again.

A ledger also gives the failure report a better vocabulary. “Message not found” is vague. “A message was observed, but it was already claimed by attempt 1” points to retry isolation. “No message matched correlation ID” points to delivery or application behavior.

Implement cursor-based polling in Playwright

Start polling after recording the mailbox cursor. The cursor can be a provider message ID, a created-at timestamp, or a local set of IDs already seen. A provider ID is best when available because timestamps can have coarse resolution.

async function waitForNewMessage(
  mailbox: Mailbox,
  correlationId: string,
  seenIds: Set<string>,
  timeoutMs = 30_000,
): Promise<Message> {
  const deadline = Date.now() + timeoutMs;

  while (Date.now() < deadline) {
    const messages = await mailbox.list({ limit: 20 });
    const match = messages.find((message) =>
      !seenIds.has(message.id) &&
      message.text.includes(correlationId),
    );

    if (match) {
      seenIds.add(match.id);
      return match;
    }

    await new Promise((resolve) => setTimeout(resolve, 750));
  }

  throw new Error(`No new message for ${correlationId}`);
}
Enter fullscreen mode Exit fullscreen mode

The helper does three important things. It rejects messages already seen by this fixture, it matches a run-specific marker, and it uses a bounded wait. A fixed sleep can be useful for a quick experiment, but it doesnt tell you whether the message arrived early or late. The cursor gives the test a repeatable observation rule.

After the message is found, claim it before opening the link:

const message = await waitForNewMessage(mailbox, runId, seenIds);
ledger.claim(message.id, testInfo.testId);

const verificationUrl = extractVerificationUrl(message.text);
await page.goto(verificationUrl);
ledger.markOpened(message.id);
Enter fullscreen mode Exit fullscreen mode

In a real fixture, make claim atomic when multiple workers can access the same ledger. If atomic storage is not available, unique mailboxes per worker are the safer boundary. The ledger should never become a new shared race condition.

Make parallel workers own their messages

Each Playwright worker should have an address, correlation ID, and ledger scope that it can explain. A worker ID alone is not enough because a retry may reuse the worker while changing the test attempt. Include both values in the marker, for example signup-${workerIndex}-${testInfo.retry}-${randomId}.

Keep the raw message out of normal CI logs. Store the provider message ID, subject hash, timestamp, and claim owner in the receipt. This is enough for failure analysis without leaking verification tokens. A pattern for less flaky email verification also starts with separating delivery evidence from browser assertions.

Be explicit about malformed-input tests too. A value such as tepm mail com or dummy e mail can be useful for client-side validation, but it should never be allowed into the happy-path fixture. Test data that looks almost valid are easy to misread in a report.

Cleanup should close the ledger even when the browser assertion fails:

try {
  await runSignupFlow(page, mailbox.address, runId);
  await verifyMessage(mailbox, ledger, runId);
} finally {
  await ledger.close();
  await mailbox.cleanup();
}
Enter fullscreen mode Exit fullscreen mode

The cleanup contract is seperate from message consumption. A test can claim the right message and still leave an address behind, or clean the address while losing the receipt needed to understand a failure.

Failure-analysis checklist

When a test fails, read the receipt in this order:

  1. Was a unique correlation ID generated for this attempt?
  2. What was the mailbox cursor before the signup request?
  3. Did the application record a send request for that ID?
  4. Was a new message observed, or was an old message returned?
  5. Who claimed the message, and was it claimed twice?
  6. Did link extraction succeed before the browser navigation?
  7. Did cleanup run after the failure?

For CI runs, attach the receipt beside the trace. A low-noise CI email check can help when the application reports success but the mailbox has no matching message. Retain the boundary evidence, not just the last failed assertion.

Final takeaways

A consumption ledger turns email testing from “find the newest message” into a small, reviewable protocol. Record the cursor, match a unique marker, claim exactly one message, and preserve the receipt when the browser flow fails.

This pattern does not make delivery instantaneous, and it wont fix a broken verification endpoint. It does make the next failure honest: stale message, duplicate claim, missing delivery, or browser interaction. That is the kind of signal a QA team can act on.

Top comments (0)