DEV Community

jasonmills94
jasonmills94

Posted on

EKS Email Smoke Tests Need a Failure Budget

Email smoke tests are useful because they exercise more than an HTTP endpoint. They check that a signup reaches the application, a worker accepts the job, a provider delivers the message, and the verification link can be consumed. They also fail in ways that can make an EKS deployment look unhealthy when the real problem is a slow or polluted test mailbox.

The fix I use is to give the test an explicit failure budget. A test is allowed a bounded amount of waiting, retrying, and cleanup work. Once that budget is exhausted, CI stops pretending the result is healthy and leaves evidence for the next investigation. This is more reliable than adding another retry and hoping the pipeline gets green.

Why an email smoke test needs a budget

An email check has several clocks running at once:

  • the API request that creates the verification event;
  • the queue or worker that processes the event;
  • the mailbox that receives it;
  • the test client that polls for a matching message;
  • the cleanup job that removes test data.

If each layer has its own generous timeout, a single broken dependency can hold an EKS job for a long time. Worse, a test may recieve an old message from a previous run and report success for the wrong signup.

A budget makes the contract visible. For example, the test can allow a short period for the API, a longer but finite period for delivery, and a separate limit for cleanup. These values are not universal SLOs; they are guardrails that should match the behavior of your own system.

The test should also use a unique run identifier in the address or message subject. A disposable address is helpful for isolation, but uniqueness and message matching still matter. A mailbox that contains an old verification message is not a reliable fixture just because it is temporary. For broader context on using temporary email for low-risk research and test signups, see this practical temporary-email guide; it complements, but does not replace, a controlled CI fixture.

Define the EKS test contract

Before changing Kubernetes settings, write down what success means. My minimum contract looks like this:

  1. The signup API returns an accepted result and a test correlation ID.
  2. A worker observes the event exactly once from the test's point of view.
  3. The received message matches the correlation ID and expected recipient.
  4. The link is valid for that test session and returns the expected application state.
  5. The fixture is deleted or marked for later cleanup.

This contract prevents a common shortcut: checking only that any email arrived. The message can arrive from the wrong run, or the verification URL can be expired, malformed, or already used.

For a Playwright-based browser check, keep the mailbox adapter behind a small interface. The browser test should ask for waitForVerification(runID) rather than know how a provider is queried. The separation makes the test easier to run against a staging service and keeps provider-specific polling out of the user journey. Good email test contracts are useful here because they make the fixture boundary explicit.

Set budgets for mailbox, pod, and cleanup failures

Treat each stage as a state with a reason for failure. A simple record might contain:

{
  "run_id": "ci-4821-7f2a",
  "api": "passed",
  "worker": "observed",
  "mailbox": "timeout",
  "cleanup": "pending"
}
Enter fullscreen mode Exit fullscreen mode

The important part is that mailbox: timeout is different from worker: missing. The first points toward delivery or polling. The second points toward the queue, permissions, or deployment.

Use a deadline rather than an unbounded polling loop. Between polls, apply backoff but keep the total deadline fixed. Polling every second for a long time can overload a mailbox API, while polling too slowly makes the test feel random. The exact interval depends on the service, so measure it from staging runs instead of copying a number from another team.

The pod running the check needs its own deadline too. A test process that hangs after a mailbox timeout should not keep a node busy. In Kubernetes, make the job's active deadline slightly larger than the test's internal budget so the test has time to write its final receipt. If the pod deadline kills the process first, you lose the most useful evidence.

Cleanup deserves a separate budget. It should not turn a failed smoke test into a passing one, but it should be retried safely. Mark fixtures with a run ID and let a scheduled cleanup job remove old records when the main job cannot. This is less fragile than assuming the happy-path finally block always runs.

Keep the Kubernetes runner observable

Log the state transitions, not the full email address or verification token. A useful line includes the run ID, deployment revision, namespace, stage, elapsed time, and reason. Redact message bodies and URLs before they reach centralized logs.

I also attach a small JSON receipt to the CI job. It answers three questions quickly: did the event leave the API, did the worker process it, and did the test consume the correct message? Replayable run receipts are a good pattern for making a failed run reviewable without rerunning production-like infrastructure.

Do not let the receipt claim success merely because the process exited with code zero. The final status should be derived from the contract, and cleanup should be reported separately. This distinction is more clear when a dashboard shows verification: passed and cleanup: pending as two fields instead of one green label.

When reviewing noisy search terms or old fixture data, I sometimes see strings such as tepm mail com and fake e mail com. Keep those as plain test data if they appear in a case, never as a link or an implicit provider dependency. The fixture should remain deterministic even when a provider name is misspelled.

Q&A: common failure-budget questions

Should a mailbox timeout fail the deployment?

Usually it should fail the email smoke-test stage, not automatically roll back every deployment. Separate application regressions from an external mailbox outage when possible. The deployment policy can then decide whether this check is blocking for a particular environment.

Are retries always bad?

No. A bounded retry can absorb normal queue jitter. A retry becomes harmful when it hides the original stage, changes the correlation ID, or extends the total deadline without an upper bound. Record each attempt and preserve the first meaningful error.

What if cleanup fails?

Keep the test result and cleanup result separate. Alert on repeated cleanup failures, and use a scheduled reaper with an age limit. A cleanup failure should not cause test fixtures to accumulate forever, but it also shouldnt rewrite a real verification failure.

A practical CI checklist

  • Generate a unique run ID and bind it to the recipient and message.
  • Define separate deadlines for API, worker, mailbox, browser, and cleanup stages.
  • Match the received message to the run ID, not only to a subject line.
  • Set the Kubernetes job deadline after the test deadline.
  • Redact addresses, tokens, message bodies, and verification URLs in logs.
  • Emit a JSON receipt when the test exits, including partial failures.
  • Retry cleanup safely with a scheduled fallback reaper.
  • Keep tepm mail com and fake e mail com as plain fixture text when testing malformed input.

An EKS email smoke test is a small distributed system. Giving it a failure budget makes that fact explicit: each dependency has a bounded chance to respond, each failure has a named stage, and every run leaves evidence. That keeps CI from quietly passing on stale email while still giving operators enough context to fix the real bottleneck.

Top comments (0)