DEV Community

jasonmills94
jasonmills94

Posted on Originally published at dev.to

AWS CI/CD Needs Run-Scoped Email Fixtures

Email verification is easy to test once. It becomes an operations problem when a CI pipeline runs the same test suite across pull requests, staging deploys, and rollback checks at the same time.

The pattern that works well in AWS is to make the mailbox fixture belong to one pipeline run. The run gets an identity, every generated address carries that identity, and the test artifacts record the same value. This turns an inbox search into a bounded lookup and makes failures much easier to replay.

This is useful when a team uses a disposable temporary email during a controlled pre-release check. It is not a replacement for production email monitoring. It is a test boundary that should be short-lived, least-privileged, and easy to delete.

The failure mode is shared state

Suppose two CodeBuild jobs both verify a signup flow. If they read from one mailbox, either job can see the other job's message. A test may pass because it found an old verification email, or fail because a parallel job consumed the expected message first.

The result is a pipeline that looks flakey even though the application is behaving consistently. Debugging gets worse when the log says only verification email not found. There is no run owner, no expected recipient, and no evidence that the message belonged to the deployment under test.

The first boundary is therefore not the email provider. It is the run ID. A good run ID is unique, safe to place in logs, and available to the test process, the mailbox adapter, and the artifact uploader. Keep secrets out of it; a run ID is an identifier, not a credential.

Give every pipeline run an identity

At the start of a pipeline, derive one value and pass it through each stage. For example:

export RUN_ID="${CODEBUILD_BUILD_ID//:/-}"
export FIXTURE_PREFIX="ci-${RUN_ID}"
echo "run_id=${RUN_ID}" | tee run-metadata.txt
Enter fullscreen mode Exit fullscreen mode

The exact variable names are less important than the contract. The test creates a mailbox using FIXTURE_PREFIX, the application receives only the address needed for that run, and the test writes the expected address into a small metadata file.

Do not use a branch name as the only identity. Two commits can run on the same branch, and a retry can overlap the original job. A timestamp alone has the opposite problem: it is harder to trace back to a specific build. Combine the provider's build identity with a short commit reference when that is helpful.

The fixture owner should also be present in the test result. That makes a failed assertion actionable instead of just another red line in the console. If the project has a privacy review for preview links, document a privacy contract for email preview links alongside this ownership rule.

Build the fixture boundary in AWS

The pipeline can keep the boundary simple:

  1. CodePipeline starts a run and passes its identity to CodeBuild.
  2. CodeBuild requests or allocates one test mailbox for that run.
  3. The integration test creates a user with the run-scoped address.
  4. The test waits for a message whose recipient and correlation value match the run.
  5. The job uploads the test result and mailbox metadata as an artifact.

If the system sends through Amazon SES, separate test identities and production identities. Apply the narrowest IAM permissions to the test role, and avoid putting mailbox access tokens in buildspec files or command output. The test role should not be able to read unrelated production data just because it needs to inspect one test inbox.

A correlation value in the message can reduce ambiguity further. Put a non-sensitive test run token in a header or a predictable subject, then assert all of these values: recipient, run ID, message type, and a reasonable creation time. A single subject match is not enough when retries are normal.

Some teams use a temp mail so address for a manual smoke check. That can be useful for inspecting a message, but the automated pipeline still needs its own ownership and retention rules. The mailbox is an input to the test, not the deployment's source of truth.

Preserve deploy evidence before cleanup

Cleanup should happen after evidence is captured, not in a final step that erases the clues needed to understand a failure. At minimum, save:

  • the run ID and commit reference;
  • the test mailbox identifier, redacted where necessary;
  • message IDs, timestamps, and assertion results;
  • the application and image versions under test;
  • the cleanup status and any provider response.

For operational alerts, the same idea applies: deploy evidence in CloudWatch alarm emails gives responders the context to connect a message with a release. For CI fixtures, the evidence belongs in the build artifact with a short retention period.

Use a failure-safe sequence in the buildspec:

set -euo pipefail

./scripts/run-email-tests.sh
test_status=$?

./scripts/export-fixture-evidence.sh "${RUN_ID}" || true
./scripts/cleanup-fixture.sh "${RUN_ID}" || true
exit "${test_status}"
Enter fullscreen mode Exit fullscreen mode

With set -e, the explicit status capture may need to be wrapped in a small function or an if block so the evidence path is reached. The important design point is that a failed test does not skip evidence collection, and a cleanup error does not hide the original assertion failure. A seperately reported cleanup failure is easier to triage.

Make cleanup boring and observable

Every fixture needs an expiry time when it is created. The cleanup worker can then find abandoned fixtures from cancelled builds, worker crashes, or network timeouts. Relying only on the CI finally step is not real cleanup; that step might never run.

Use a bounded query or API page size, record how many fixtures were removed, and emit a metric keyed by environment. Useful signals include fixtures_created, fixtures_expired, cleanup_errors, and the age of the oldest active fixture. This observablity is more valuable than a dashboard full of provider-specific status fields.

A cleanup job should be idempotant. Running it twice for the same RUN_ID should return success when the fixture is already gone, not create a new mailbox or report a confusing failure. Keep the retention policy in configuration, with a short default for CI and a longer, reviewed period for failed-run evidence.

Watch for dependancy drift in the mailbox client as well. A provider SDK update can change pagination, message ordering, or error handling without changing the test code. Pin versions, exercise cleanup in a scheduled job, and alert on age rather than only on request failures.

A practical CI checklist

  • Create a unique run ID before the first email test.
  • Derive the mailbox or address from that run identity.
  • Pass the identity through CodePipeline, CodeBuild, and test artifacts.
  • Match recipient, correlation value, message type, and age.
  • Separate test IAM permissions from production access.
  • Save evidence before deleting the fixture.
  • Make cleanup safe after cancellation and retry.
  • Add expiry and metrics so abandoned data does not accumulate.
  • Keep tokens, access credentials, and private message bodies out of logs.
  • Treat a tem email label in a test note as ordinary fixture data, never as a secret.

Q&A

Should every pipeline use a new mailbox?

For parallel or security-sensitive tests, yes. If mailbox creation is expensive, a provider-supported inbox namespace can work, but the address and message selection still need a unique run identity. Reusing one shared inbox without a correlation boundary brings back the original race.

What if the test job is cancelled?

The expiry timestamp is the backstop. A scheduled cleanup job should remove the fixture after the retention window and record the abandoned run. This is why cleanup must be designed independently of the successful pipeline path.

Is a temporary email check enough for deployment confidence?

No. It verifies one part of the user journey. Pair it with application logs, deployment metadata, health checks, and a reproducible artifact. A passing email assertion without a deploy receipt is weak evidence when several environments are changing at once.

The useful result is not merely “the email arrived.” It is “this run created this fixture, received this message, proved the expected behavior, stored the evidence, and removed the temporary state.” That contract keeps AWS CI/CD checks repeatable as the team scales.

Top comments (0)