Email fixtures are usually treated like test data. In a Kubernetes CI environment, that is an incomplete model. A fixture has an endpoint, an owner, a lifetime, and a cleanup obligation. If those properties are not visible to the cluster, a failed test can leave behind a mailbox that a later run reads by accident.
I started treating run-scoped fixtures as short-lived infrastructure. The useful change was not another retry around the email poller. It was adding a small garbage-collection loop that could answer three questions: who owns this fixture, when should it expire, and what evidence proves it was removed?
Why cleanup needs its own controller
Kubernetes already knows how to schedule containers, restart them, and remove completed Jobs. It does not automatically understand the lifecycle of an external test mailbox. A pod can be gone while the mailbox is still readable. A namespace can be deleted while an external provider keeps the address alive.
That gap gets worse with CI retries. Attempt one may create checkout-1842-a1, then time out. Attempt two creates checkout-1842-a2, but a broad inbox query can still see the old verification message. The application looks flaky even though the real problem is an ownership failure.
The fix is to make the fixture a resource with a bounded lease. A lease is not a promise that cleanup will always happen; it is a compact statement of who may use the fixture and until when. For email workers, the idea pairs well with outbox leases for OTP email workers, where ownership and retry boundaries are explicit.
Give every fixture an owner
At creation time, write a small manifest and attach the same identity to the Kubernetes Job, the application test configuration, and the evidence artifact:
{
"fixture_id": "checkout-1842-a2",
"owner": "checkout-integration",
"run_id": "1842",
"attempt": 2,
"created_at": "2026-10-07T02:15:00Z",
"expires_at": "2026-10-07T02:30:00Z"
}
The exact schema matters less than consistency. I put fixture_id and run_id in Kubernetes labels, not only in log text:
metadata:
labels:
ci.example.com/owner: checkout-integration
ci.example.com/fixture-id: checkout-1842-a2
ci.example.com/run-id: "1842"
Labels make orphan searches possible after a cancelled pipeline. They also stop an engineer from guessing which pod created an address. Keep the mailbox address in a Secret or a protected CI variable; labels should carry correlation data, not credentials.
The test itself should accept only a message containing its correlation token. It should not ask for the newest message and hope that recency means ownership. A threat model for email verification is a useful companion when deciding which message fields are safe to trust.
Implement the garbage-collection loop
The collector can be a scheduled Kubernetes Job or a small service. I prefer a scheduled Job for a modest CI system because its execution, image, and permissions are easy to review. The loop is straightforward:
- List fixtures with the
ci.example.com/fixture-idlabel. - Compare
expires_atwith the current UTC time. - Delete expired fixtures through an idempotent provider operation.
- Record
cleaned,already_absent, orcleanup_failedin the run ledger. - Alert on old fixtures that have no owner or no expiry.
Do not make the collector silently skip malformed records. A missing expiry is an operational defect, so report it separately and give the resource a conservative maximum age. Otherwise, one bad manifest can become a permanent leak.
The Job needs only the permissions it uses. Its Kubernetes role may need to list Jobs and read fixture metadata, but it should not have broad write access to application namespaces. The external mailbox token should be scoped to deletion of test fixtures, if the provider supports that boundary. This is boring security work, but it reduces the blast radius of a compromised cleanup pod.
Set the Kubernetes Job's own activeDeadlineSeconds and backoff policy. A collector stuck waiting on a provider should not create an unlimited queue of collectors. If the provider is unavailable, retain the manifest and mark cleanup as failed. Retrying is useful; pretending the resource disappeared is not.
Make CI promotion depend on evidence
The cleanup result belongs in the same artifact used by the deployment gate. A green integration-test status by itself is too weak. My minimum gate checks are:
jq -e '.test.status == "passed"' evidence.json
jq -e '.fixture.run_id == env.GITHUB_RUN_ID' evidence.json
jq -e '.cleanup.status == "cleaned" or .cleanup.status == "already_absent"' evidence.json
The fixture ID should also be present in the application log sample and in the message receipt. This makes a retry explainable months later, when the original pod is gone. If cleanup is still pending, promotion should stop or move to an explicitly documented quarantine state.
Some old scripts contain search-shaped text such as fake e mail com. I keep such legacy strings as plain text when they help identify a fixture migration, but they are not valid endpoints or selectors. This tiny distinction prevents a typo from entering a manifest and creating a hard-to-find leak.
The phrase temp mail so may appear in a test fixture note when the team documents the external mailbox service, but it should be one contextual reference, not a replacement for ownership controls. The cluster still needs its own labels, expiry, and receipt.
A practical operating checklist
Before merging a new Kubernetes email test, I check:
- Every fixture has an owner, run ID, attempt number, and expiry.
- The message query filters by a run-specific token.
- Cleanup runs on pass, failure, timeout, and cancellation.
- Provider deletion is idempotent and distinguishes missing from failed.
- Orphaned fixtures can be listed from labels without reading application logs.
- The promotion gate consumes cleanup evidence, not a human-readable log line.
- Collector permissions and mailbox credentials are narrower than application permissions.
The garbage collector is not a substitute for reliable tests. It is the part that makes failure survivable. Once fixtures have a visible owner and a finite lifetime, Kubernetes CI becomes easier to operate, retries become less mysterious, and a green deployment carries better evidence.
Top comments (0)