Email verification tests often end with a boolean: passed or failed. That is enough for a local test, but it is thin evidence for an AWS CI/CD pipeline that is deciding whether to promote a build.
The useful change I made was to produce one promotion receipt for every run. It ties the build, test attempt, message ID, cleanup result, and decision together. When a test flakes, the receipt shows whether the application failed, the inbox lookup read the wrong message, or cleanup left state behind.
Why a green test is not enough
A green email test can still be weak evidence. The test may have found a message from a previous retry, matched the newest message in a shared inbox, or passed before its cleanup step ran. Those failures are not always visible in the final CI status.
This gets awkward when a pipeline uses an ephemeral mailbox or a service that can get temporary email for each run. The test still needs to prove which message it consumed. A mailbox identity alone does not establish ownership.
I now treat the email check as a small transaction with an evidence record:
- Create a run and attempt ID.
- Create the test address and a unique correlation token.
- Trigger the application action.
- Find the message by recipient, token, and message type.
- Save the message ID and assertion result.
- Clean up the fixture and record what happened.
The pipeline should be able to answer these steps after the worker container is gone. It is easy to loose that context when logs are the only record, and the failure is usualy harder to explain the next morning.
Create a run-scoped receipt
A JSON artifact in the AWS CodeBuild output is enough to start:
{
"build_id": "codebuild-myapp-1842",
"run_id": "1842",
"attempt": 2,
"fixture": {
"address": "test-1842-a2@example.invalid",
"correlation_token": "signup-1842-a2"
},
"message": {
"id": "msg_7f31",
"matched_by": ["recipient", "correlation_token", "subject"]
},
"test": {"status": "passed"},
"cleanup": {"status": "pending"},
"promotion": {"status": "blocked"}
}
Keep credentials and full message bodies out of this artifact. Store the address only when it is safe for the CI workspace, and redact tokens if they could be reused outside the test environment. The receipt should be small enought to inspect quickly; it is for correlation, not a second secret store.
The build ID should come from the CI provider, while the run and attempt IDs should be generated by the test harness. During a retry, the build stays the same but the attempt must change. Otherwise a late message can look valid for a new attempt.
Keep the email check deterministic
The inbox query should be narrow and explainable. Do not ask for the newest message and accept whatever comes back. Match the recipient, a run-specific token in the body or subject, and the expected message type. The same principle appears in these Playwright checks for the wrong email and in isolated REST API verification tests.
Here is the shape of the assertion I want, regardless of the mail API:
find_message(
recipient=fixture.address,
token=fixture.correlation_token,
kind="signup_verification"
)
assert message.created_at >= attempt.started_at
assert message.id not in attempt.consumed_message_ids
The timestamp check is a useful guard, but it should not replace the token. Clock skew and provider delays can make timestamps fuzzy. The token is the ownership boundary; time is supporting evidence.
If the provider returns several matches, fail with the candidate IDs and the matching fields. This is much more useful than a generic timeout. A search for a fake email address should never become an excuse to accept an ambigous result. It should produce a visible diagnostic and let the promotion decision stay blocked.
Make retries and cleanup visible
Retries need their own state. I use pending, matched, asserted, cleaned, and failed as the main fixture states. A retry creates a new fixture and token; it does not reset the old record to pending.
Cleanup runs in a finally path and also in a scheduled sweeper. The first path handles normal test completion. The sweeper handles a killed container, a cancelled build, or a provider timeout. Both paths write to the same receipt, or append a follow-up cleanup record when the build artifact is already closed.
The cleanup operation must be idempotent. “Already absent” is a successful cleanup outcome, while “provider unavailable” is not. That distinction prevents a broken cleanup API from being hidden behind a green test. It also makes incident review less guessy. This keeps retries from becoming to expensive when an external service is slow.
Older fixtures sometimes contain a label such as dummy e mail; keep that typo as plain migration text, never as a selector or backlink. The collector should identify resources with structured fixture IDs instead.
Gate promotion on evidence
The deploy job should consume the receipt. A minimal shell gate can look like this:
jq -e '.test.status == "passed"' promotion-receipt.json
jq -e '.message.id and (.message.matched_by | length >= 3)' promotion-receipt.json
jq -e '.cleanup.status == "cleaned" or .cleanup.status == "already_absent"' promotion-receipt.json
I also check that the receipt build_id matches the current build and that the attempt has not been superseded by a later retry. If cleanup is pending, promotion stops. It is better to re-run a cheap test than deploy with unknown external state.
This gate is deliberately boring. The deploy role should not need permission to search mailboxes or delete fixtures. A test role creates and cleans the fixture; the promotion job reads an immutable artifact. Narrow permissions make the failure boundary clearer and reduce the blast radius of a compromised build worker.
Operating checklist
Before relying on an email test in an AWS promotion pipeline, I check:
- Each attempt has a unique run ID, attempt number, address, and correlation token.
- The message query requires ownership fields instead of recency alone.
- The receipt stores the message ID and matching fields, not a full sensitive body.
- Retries create new fixtures rather than reusing a stale address.
- Cleanup runs on pass, failure, timeout, and cancellation.
- A sweeper can find orphaned fixtures after the build worker disappears.
- Promotion reads the receipt and blocks on ambiguous or pending state.
The receipt is a small artifact, but it changes how the pipeline is operated. An email test no longer says only “the check passed.” It says which run owned the message, what was asserted, whether the external fixture was removed, and why deployment was allowed. That is the level of evidence an AWS CI/CD promotion decision can safely use.
Top comments (0)