Signup tests often fail in a way that is hard to reproduce. The form works on a laptop, but the CI job sometimes reads an old verification message, waits for a code that belongs to another run, or leaves a test inbox behind. The usual response is to increase the timeout. That makes the symptom quieter, not the test more reliable.
The better fix is a small contract between the test, the application, and the email fixture. Every run gets an identity, every message is matched to that identity, and cleanup has an owner even when the assertion fails. This pattern is useful whether the inbox comes from a fake emails generator, a local mail catcher, or a service that can create temporary mail.
The flaky test is usually a missing contract
Consider a Playwright test that submits a signup form and then searches an inbox for the latest message. “Latest” is not a stable identifier. Parallel CI jobs can send messages close together, retries can create duplicates, and a previous run may still be visible.
There is another subtle failure: the test may assert that a verification page loaded without proving that its code belonged to the user just created. The browser test goes green while the product behavior is only partially checked.
I treat the email fixture like any other test resource. It needs a run ID, a known lifecycle, and an observable result. The same thinking behind low-risk email fixtures for CI applies inside a React test suite: isolation is a correctness feature, not just a privacy preference.
Define the run-scoped email contract
Before writing the test, decide what the fixture promises:
- Ownership: one test or worker owns the inbox for one run.
- Correlation: the message includes a recipient or run identifier that the test can match.
- Freshness: the test records a message cursor or creation time before sending.
- Bounded waiting: polling stops with a useful error, not an infinite retry.
-
Cleanup: the inbox is deleted or expired in a
finallypath.
The run identifier can be short and opaque. For example, the test can create signup-${workerIndex}-${randomSuffix} and use that value in the email address when the provider supports aliases. If it cannot, store the generated address and associate it with the test in a small fixture object.
Do not make the component know about this machinery. The signup UI should receive an email address and expose its normal states. The test harness owns provisioning, polling, and teardown.
Model the React test as a small state machine
The test becomes easier to debug when it records what it expects from the UI at each stage. A TypeScript discriminated union is enough:
type SignupStep =
| { kind: "form" }
| { kind: "submitted"; address: string; runId: string }
| { kind: "message-found"; messageId: string; code: string }
| { kind: "verified" };
The important detail is not the type itself. It is that the transition from submitted to message-found carries evidence from the fixture. A test failure can now say whether the form submission failed, the message never arrived, the wrong message was selected, or the code was rejected.
For React Testing Library, keep the UI assertions close to the user action:
await user.type(screen.getByLabelText(/email/i), fixture.address);
await user.click(screen.getByRole("button", { name: /create account/i }));
expect(await screen.findByText(/check your email/i)).toBeVisible();
Then use the fixture client to find the message. Mixing inbox polling into a component helper usually make the failure harder to understand, because a browser assertion and an external timeout become one opaque promise.
Keep the fixture outside the component
A simple fixture API can expose the lifecycle without leaking provider details into every test:
type EmailFixture = {
address: string;
since: string;
waitForVerification(): Promise<{ id: string; code: string }>;
dispose(): Promise<void>;
};
async function createEmailFixture(runId: string): Promise<EmailFixture> {
// Provision an isolated address and capture the pre-send cursor.
return provisionFixture({ runId });
}
The since cursor matters. Filtering only by recipient is not enough when a retry sends two messages. Match the recipient, message type, and cursor, then assert that the code is for the current signup attempt. A temp mailid can be a fine value in a deliberately negative test, but it should not accidentally become the shared fixture name. Likewise, a typo such as tempail mail in a search term should be surfaced as test data, not silently normalized into a passing result.
Make polling and cleanup explicit
Polling should report the last useful observation. Instead of “timeout after 30 seconds”, return something like: “saw 2 messages for the fixture, 0 matching verification events, last provider status: queued.” That message points the next investigation toward the sender, worker, or matcher.
Use a deadline and a small interval with backoff. Keep the total budget below the surrounding test timeout so the test can still collect logs and dispose of resources. If the app has an asynchronous worker, the deployment context for email checks is a useful operational reminder: a test can only diagnose delivery when it knows which environment and deployment produced the message.
Always clean up in a finally block:
const fixture = await createEmailFixture(runId);
try {
await completeSignup(fixture.address);
const message = await fixture.waitForVerification();
await verifyCode(message.code);
} finally {
await fixture.dispose();
}
This is a bit more clear than relying on a global after-all hook. A global hook may not know which worker owns a leaked resource, and it can hide the original failure when cleanup also errors. If disposal fails, record it as a secondary failure with the run ID.
A practical checklist for CI
Before calling the signup test stable, check these points:
- Does every parallel worker receive a distinct fixture?
- Is there a cursor or timestamp captured before the message is sent?
- Does the matcher reject messages from another run?
- Can the timeout explain the last observed provider state?
- Does cleanup run after assertion failures and browser crashes?
- Are email addresses, codes, and message bodies redacted from ordinary logs?
The payoff is bigger than fewer flaky tests. A run-scoped email contract makes the failure boundary visible. Product code owns the signup state; the test harness owns the inbox lifecycle; CI owns the final evidence. Once those responsibilities are separate, increasing a timeout is no longer the main strategy. The suite can move faster, fail more honestly, and tell you what to fix next.
Top comments (0)