DEV Community

jasonmills94
jasonmills94

Posted on

AWS Test Mailboxes Need a Cost and TTL Budget

An AWS test mailbox is a cloud resource, not just a test detail. Give each run a lease, a TTL, and a cleanup owner before CI starts creating disposable email accounts.

Why mailbox cost belongs in the test design

Email verification tests often begin with a simple assumption: create a temp mail address, wait for a message, and delete it later. That works on a laptop. In a busy CI system, it creates a resource-lifecycle problem.

Parallel jobs can leave mailboxes behind after a cancelled runner. Retries can create another mailbox while the first one is still active. A failed cleanup commmand may not be noticed until the monthly bill or a provider quota makes the problem visible. The test passed, but the cloud system is accumulating state.

I treat every test mailbox as a short-lived lease with four fields:

  • run_id: the pipeline run and attempt that owns it
  • address: the disposable email address assigned to the test
  • expires_at: the hard end of the lease
  • cleanup_state: pending, deleted, or cleanup_failed

This also makes the boundary clearer for privacy. An email verification flow should retain only the message data needed by the test, then remove the rest. The discussion in email verification needs a privacy budget is a useful reminder that test data still deserves an explicit retention policy.

The run-scoped lease

The lease should be created once per CI run or per isolated test group, depending on the provider's limits. I prefer one mailbox per run when parallel tests need separate recipients, and one mailbox per test group when the suite is small and the message cursor is reliable.

The name should include an opaque run identifier, not a user email or a secret. For example:

ci-20261009-1842-attempt-1
Enter fullscreen mode Exit fullscreen mode

The TTL is a budget, not a suggestion. If the normal test completes in two minutes, a five-minute lease may be reasonable. Add room for queue time and one retry, but avoid a twenty-four-hour default because it feels safe. Long leases hide cleanup failures.

The lease record can live beside the test run metadata:

{
  "run_id": "1842-1",
  "address": "ci-1842-1@example.test",
  "expires_at": "2026-10-09T08:30:00Z",
  "cleanup_state": "pending"
}
Enter fullscreen mode Exit fullscreen mode

The mailbox client should reject reads for a different run_id. That check prevents a late message from a previous attempt becoming a false positive in the current test.

A small AWS implementation

For an AWS-based pipeline, I keep the lease metadata in a small DynamoDB table or in the system that already records CI fixtures. The important part is the contract, not the storage product:

  1. The job requests a lease with a bounded TTL.
  2. The service returns the address and a correlation token.
  3. The test filters messages by recipient, run ID, subject, and token.
  4. The job marks the lease consumed or failed.
  5. A scheduled cleanup worker deletes expired leases.

The CI role should have only the permissions needed for its path. The test role may create and consume a lease, while the cleanup role can delete expired records. Separating those permissions makes a bug easier to contain.

I also pass the lease identifier through the pipeline as ordinary metadata, not as a secret. The mailbox address is not a credential. Provider tokens and AWS credentials must stay in the CI secret store and should never be copied into logs.

At the start of a run, write a structured event such as:

mailbox_lease_created run_id=1842-1 ttl_seconds=300 owner=signup-tests
Enter fullscreen mode Exit fullscreen mode

At the end, write the matching state transition. This is more useful than printing the full message body, which adds noise and can expose test data.

Promotion and cleanup gates

A green test result is not enough for a production promotion. I use two gates:

  • Test gate: all expected messages were consumed with the correct run ID.
  • Lifecycle gate: the lease was deleted, or an explicit cleanup task was queued with an owner and deadline.

The lifecycle gate catches a common failure mode: the application test passes, the runner disappears, and the cleanup step never runs. A finally block is still useful, but it is not a recovery plan for a terminated worker.

For deployment evidence, I keep the lease state and test result linked to the pipeline execution. This follows the same operational idea as a verifiable AWS receipt for CI releases: a later operator should be able to see what ran, what it proved, and what resources it left behind.

Set an alert for expired leases older than the allowed grace period. The alert should point to the run ID and cleanup owner, not just say “email test failed.” That difference saves time during an incident.

What to measure in CloudWatch

The minimum useful metrics are:

  • leases created and deleted
  • cleanup failures
  • active leases by environment
  • message wait latency
  • expired leases beyond the grace period
  • test retries caused by missing messages

Do not use only the number of successful tests. A rising success count can hide a growing active-lease count. A small dashboard with these signals gives a better view of reliability and cost.

If a team searches for “temp org mail” while troubleshooting, the runbook should direct them to the lease dashboard and the provider request ID. Search phrases are not a substitute for an ownership record.

Checklist for the next pipeline run

Before enabling this pattern, verify:

  • Every lease has a bounded TTL.
  • The run ID is included in message correlation.
  • Cleanup runs on success, failure, and expiry.
  • The cleanup worker is idempotent.
  • CI permissions are split from cleanup permissions.
  • Message bodies are not copied into normal logs.
  • CloudWatch alarms cover cleanup failures and stale leases.
  • Promotion evidence contains the test and lifecycle states.

One small warning: do not make the cleanup job depend on the same runner that owns the test. Put recovery on a seperate schedule or worker. Otherwise the failure that needs cleanup can also remove the only process capable of doing it.

Questions I ask before shipping

What if the provider delays a message?

Use a bounded polling policy with a clear timeout and record the latency. Do not retry forever, and do not read the newest message without checking its correlation token.

What if the cleanup API is unavailable?

Keep the lease record, mark cleanup_failed, and let the recovery worker retry with backoff. The record should survive long enough to explain the failure, but the mailbox itself should still have a provider-side expiry.

Is a disposable email account safe for every test?

No. Use synthetic or controlled addresses for sensitive workflows and follow the provider's terms. A disposable mailbox is a test fixture, not an identity system.

The practical outcome is simple: a temp mail fixture should have the same operational discipline as any other AWS resource. A short TTL, explicit ownership, and visible cleanup state keep CI predictable without letting forgotten test data become tomorrow's cloud bill.

Top comments (0)