Email fixtures are easy to add to a CI pipeline and surprisingly hard to retire. A test creates a disposable inbox, starts a Kubernetes Job, waits for a verification message, and then the pipeline moves on. If the runner is canceled, the namespace is deleted half way, or the worker loses its lease, the fixture can stay around longer than the test that created it.
That is not just wasted cluster capacity. Old inboxes and test payloads can expose data to later runs, while orphaned Jobs make a busy EKS cluster harder to operate. The pattern that worked for me is simple: every run gets one bounded identity, one cleanup owner, and one receipt that proves what happened.
The cleanup problem is an ownership problem
Teams often treat cleanup as a final shell step:
kubectl delete job email-check-$RUN_ID -n $NAMESPACE
This works until the shell step never runs. A canceled GitHub Actions runner does not care that a trap was configured. A node failure does not wait for the final command. Cleanup which depends on the same process that created the resource will eventually miss a case.
I still use an explicit delete in the happy path, but it is only the first layer. The cluster must have enough metadata to clean up a forgotten run without human archaeology. It also need a clear answer to the question: who owns this resource after the CI process disappears?
The answer should be a run-scoped controller boundary, not a developer's memory. A useful identity includes the repository, commit, environment, and a short run ID. Do not put the full branch name in a Kubernetes name; it create awkward truncation and can produce collisions.
Give every test run a bounded identity
Use labels and annotations consistently on the namespace, Job, Docker workload, and mailbox record:
metadata:
labels:
app.kubernetes.io/name: email-fixture
ci.example.com/run-id: "1842-7f3a"
ci.example.com/owner: "checkout-api"
annotations:
ci.example.com/expires-at: "2026-10-04T04:00:00Z"
The run-id is the join key for logs and cleanup. The expiry annotation is useful when an operator needs to find stale objects, but it is not a substitute for a controller. Clock formats must be UTC, and the producer should reject an expiry that is already in the past.
For short-lived integration tests, I prefer one namespace per run when the cluster budget allows it. It makes network policy and resource accounting easier to reason about. If namespace-per-run is too expensive, use a shared namespace with a strict label selector and never issue an unscoped delete.
A run using a tempmailso-style disposable inbox should store only the mailbox identifier needed by the test. Avoid placing message bodies, tokens, or full addresses in labels because labels are widely visible and have size limits. Values with misspellings such as tempail or temp gamil com should still be treated as untrusted test input, not normalized into a shared fixture.
Use Kubernetes TTL and explicit finalizers
Kubernetes Jobs support a time-to-live after completion. Set it in the manifest so a successful or failed Job has a second cleanup path:
apiVersion: batch/v1
kind: Job
metadata:
name: email-fixture-1842-7f3a
spec:
ttlSecondsAfterFinished: 900
backoffLimit: 1
template:
spec:
restartPolicy: Never
containers:
- name: verifier
image: ghcr.io/example/email-verifier:2026.10.04
The TTL is deliberately short, but not zero. Fifteen minutes gives an operator time to inspect a failed pod and still prevents a forgotten Job from becoming permanent. In one cluster, the cleanup controller was delayed during an API-server incident, so we also used a daily sweep that selected expired ci.example.com/run-id labels. Defense in depth matters here.
Use a finalizer only when a controller really needs to revoke an external resource, such as a mailbox lease. A finalizer that is added without a reliable controller is a stuck deletion with a fancy name. The controller should remove it after it has recorded the external cleanup result, and it should retry with backoff.
Keep Docker and mailbox cleanup in one receipt
The Kubernetes Job is not the entire fixture. A Docker test container may have a local volume, and the mailbox service may have a remote lease. Cleanup is complete only when all three states are known:
- The Job reached a terminal state.
- The Docker container and temporary volume were removed.
- The mailbox was released or its expiry was recorded.
Write a small JSON receipt for each run. It can live as a CI artifact and as a structured log event:
{
"run_id": "1842-7f3a",
"job": "email-fixture-1842-7f3a",
"mailbox": "fixture-1842-7f3a",
"job_status": "succeeded",
"mailbox_status": "released",
"cleanup_at": "2026-10-04T03:32:00Z"
}
This is the same operational idea as pod budget context during an EKS drain: an event is much easier to manage when its important state is explicit. For deployments, I also keep an AWS receipt contract for Docker releases so the build result is not confused with the runtime result.
What to alert on
Alert on cleanup age, not just Job failure. A failed test is normal enough; a fixture that remains for two hours is usually an operational defect. Useful signals are:
- expired runs with a live Kubernetes object;
- mailbox leases past their expiry time;
- Jobs stuck in
Terminatingbecause of a finalizer; - cleanup attempts that retry more than a small threshold;
- orphaned Docker volumes on self-hosted runners.
Keep the alert dimensions small. run_id, repository, and environment are usually enough. Do not put email addresses in metric labels. That decision save both cardinality and privacy headaches.
Questions that come up in production
Should cleanup happen before the test result is published? Usually no. Publish the test result first, then cleanup, and attach the cleanup state to the receipt. Otherwise a cleanup outage can hide a useful test failure.
What if the mailbox provider is unavailable? Record release_pending, let the lease expire on its own, and keep retrying from a durable queue. A deleted Kubernetes namespace does not cancel an external lease.
Do I need a separate operator? Not always. Kubernetes TTL plus a scheduled sweep is enough for small teams. Add a controller when mailbox leases, multiple clusters, or compliance retention make the state transitions more complex.
The main lesson is boring but important: disposable test data still need lifecycle controls. Give each run a bounded identity, make cleanup recoverable after process failure, and leave a receipt that tells the next operator what was actually released. That keeps Kubernetes, Docker, and AWS CI work predictable when the happy path stop being happy.
Top comments (0)