DEV Community

jasonmills94
jasonmills94

Posted on

Kubernetes Email Fixtures Need a Namespace Budget

Email verification tests are deceptively expensive to operate. A test creates an account, requests a message, waits for a link, and then moves on. In a Kubernetes-based CI system, that flow also creates pods, secrets, inbox state, logs, and sometimes a forgotten disposable email account.

The failure mode I have seen most often is not that the email provider is down. It is that test resources outlive the test that created them. A slow suite then shares a namespace with an older run, cleanup races with a retry, and the final failure gives no clue which inbox or pod was involved.

The useful fix is to give every email fixture run a namespace budget: a clear owner, a maximum lifetime, and a small set of resources that can be inspected before deletion. This keeps Kubernetes test infrastructure predictable without pretending that email delivery is instant.

Why an email fixture needs a namespace budget

A fixture should have the same lifecycle as the test run that owns it. If a CI job starts at 10:00, its namespace should not still contain an inbox secret at 14:00. The budget is not only about cost. It limits stale credentials, prevents cross-run message matches, and makes cluster capacity easier to reason about.

For a typical email smoke test, the budget can be expressed as:

owner: ci-run-1842
maximum lifetime: 30 minutes
allowed resources: deployment, job, secret, configmap
cleanup: delete on success and expiry
Enter fullscreen mode Exit fullscreen mode

The inbox address is another resource, not merely a string in an environment variable. I prefer generating a unique address per run and passing it through a Kubernetes Secret. A service such as temporary email can be useful for isolated test work, but the address still needs an owner and an expiry in the test system.

This is also where application policy matters. The policy objects for signup checks are a good reminder that a check should have explicit inputs and outcomes. The cluster fixture deserves the same treatment.

Define the namespace contract

Create one namespace per workflow attempt, rather than one shared email-tests namespace for the whole repository. The name should contain a short commit or run identifier, not an email address. Email addresses can leak into pod names, events, and monitoring labels if they are used carelessly.

Here is a small manifest pattern:

apiVersion: v1
kind: Namespace
metadata:
  generateName: email-fixture-
  labels:
    owner: ci
    cleanup: required
  annotations:
    fixture.example.com/ttl: "30m"
    fixture.example.com/run-id: "1842"
Enter fullscreen mode Exit fullscreen mode

The cleanup controller or CI finally step can use these labels. Do not rely on the finally step alone: cancelled jobs and lost runners are normal events. A scheduled controller should find namespaces whose TTL has passed and remove them.

Add a resource quota as a second line of defense. An email test normally does not need dozens of pods or gigabytes of ephemeral storage. A small quota turns a runaway retry into one failed test, instead of a noisy neighbor problem for every build in the cluster.

apiVersion: v1
kind: ResourceQuota
metadata:
  name: fixture-limit
spec:
  hard:
    pods: "6"
    requests.cpu: "1"
    requests.memory: 1Gi
    secrets: "8"
Enter fullscreen mode Exit fullscreen mode

The exact numbers depend on the test harness. Start with observed usage and leave enough room for a diagnostic job. A quota that blocks the log collector is not a useful safety limit.

Run isolated fixtures in Kubernetes

The test job should create the namespace, apply the fixture resources, wait for readiness, and record the run ID before it calls the application. Readiness is important: an HTTP 200 from a pod does not prove that the mailbox consumer is ready to see the message.

The sequence I use is:

  1. Create the namespace and apply labels.
  2. Create the unique inbox Secret with the smallest required permissions.
  3. Start the mailbox adapter and wait for its readiness probe.
  4. Run the signup or verification test with the run ID attached.
  5. Save the message ID, pod logs, and Kubernetes events.
  6. Delete the namespace, or leave it briefly when the test failed.

The run ID must be present in application logs and in the subject or metadata used to find the message. Otherwise a retry can read the first matching email, which makes a green test less trustworthy. Email verification is a signal, not an identity, so the assertion should prove the intended account and run, not just that a link exists.

Avoid putting full message bodies into ordinary pod logs. Verification links may contain tokens. Redact them before upload, and keep the raw message only in a short-lived, access-controlled artifact when debugging requires it.

Make cleanup and failure evidence explicit

Cleanup should be idempotent. Deleting a namespace that is already gone is success for cleanup purposes. The fixture are easier to trust when a retry can run the same teardown without special cases.

On success, delete immediately. On failure, preserve only a bounded evidence bundle: test output, sanitized adapter logs, events, pod status, and the namespace name. Do not preserve the mailbox forever just because a test failed. Thirty minutes is often enough for a human to inspect the evidence; after that, the controller should finish the job.

A simple CI shell shape looks like this:

set -euo pipefail

namespace="email-fixture-${CI_RUN_ID}"
kubectl create namespace "$namespace"
kubectl label namespace "$namespace" owner=ci cleanup=required

cleanup() {
  kubectl -n "$namespace" get events,pods -o wide > artifacts/k8s-state.txt || true
  kubectl delete namespace "$namespace" --ignore-not-found=true
}
trap cleanup EXIT

kubectl -n "$namespace" apply -f fixture-resources/
kubectl -n "$namespace" wait --for=condition=Ready pod -l app=mail-adapter --timeout=120s
./run-email-smoke-test --run-id "$CI_RUN_ID"
Enter fullscreen mode Exit fullscreen mode

In a real pipeline, create the artifact directory first and make the failure collector run before deletion. The example is intentionally small; the important contract is ownership, bounded lifetime, and evidence before cleanup. A dummy e mail address is not a strategy for isolation if every parallel job still reads the same inbox.

A deployment checklist

Before adding another email smoke test, check these items:

  • Does every run get a unique namespace and inbox?
  • Is the namespace labeled with an owner and expiry?
  • Is there a quota for pods, memory, and secrets?
  • Can the test match a message to its run ID?
  • Are tokens and message bodies redacted in logs?
  • Does cleanup work after cancellation and retry?
  • Is the failure bundle useful before the namespace disappears?

This pattern costs a few more Kubernetes objects, but it buys something more important: a test result that can be explained. Namespaces keeps the blast radius small, quotas stop accidental growth, and expiry handles the runs that CI never gets to clean up. That is a much better foundation for reliable cloud email testing than a shared inbox and a best-effort delete.

Top comments (0)