Email checks in CI often begin as a tiny convenience: create a user, wait for a verification message, click the link, and move on. A few weeks later, the pipeline has a flaky step, old inboxes, and a failure log that says only “timeout.” It sound small, but this kind of test can quietly become one of the least trusted parts of a delivery workflow.
The useful mental model is to treat an email check like a deployment receipt. The test should prove what happened, preserve enough evidence to debug it, and clean up after itself. This makes Automation feel less like a collection of sleeps and more like a small, reviewable system.
The hidden cost of an email check
An email smoke test crosses several boundaries at once:
- the application creates an event
- a mail provider or test inbox accepts it
- a worker delivers or exposes the message
- the test finds the right message
- the link is checked for freshness and ownership
Each boundary can be healthy while the whole test still fails. A slow worker looks like a missing email. A reused inbox looks like a successful verification. A cleanup failure leaves data that changes the next run. If every test are using the same mailbox, the result is not really isolated.
For a useful failure, record an opaque run ID, the expected recipient, the message query, the number of polls, and the final reason. Do not record the full token or the complete email body. A short receipt is usually enough.
Define the receipt before writing the test
Start with a contract that a reviewer can understand:
email_check:
run_id: ci-4812
recipient: test+ci-4812@example.invalid
subject: Verify your account
result: received
polls: 4
cleanup: completed
The address should be unique to the run, and the message query should include more than a subject. Match a run marker in the recipient or a header, then verify that the link belongs to the expected environment. This prevents an old message from passing a new test.
For a broader perspective on making delivery evidence reviewable, see the Docker release receipt pattern. The same idea works for email: a receipt is not a dump of secrets, it is a compact explanation of the system’s decision.
Bound polling and make cleanup explicit
Avoid an unbounded loop and avoid a single long sleep. A short polling interval with a hard deadline gives both faster success and clearer failure:
deadline=$((SECONDS + 45))
polls=0
while (( SECONDS < deadline )); do
polls=$((polls + 1))
message_id="$(./mail-fixture find --recipient "$RECIPIENT" --marker "$RUN_ID" || true)"
if [[ -n "$message_id" ]]; then
./mail-fixture verify --id "$message_id" --run-id "$RUN_ID"
break
fi
sleep 3
done
if [[ -z "${message_id:-}" ]]; then
echo "email check timed out after $polls polls" >&2
exit 1
fi
./mail-fixture delete --run-id "$RUN_ID"
The cleanup command should run on both success and failure. In a real script, put it in a trap or a test framework’s teardown hook. The job finish is not the same thing as the mailbox being clean. Also make deletion idempotent, because a retry after a partial cleanup should be safe.
If a team is comparing providers or fixtures, a temp mail so can be useful as a disposable boundary for manual checks. Keep it out of production identity flows, and never put access tokens in the pipeline log. A short-lived fixture is a test input, not a security control.
A small GitHub Actions pattern
Expose the receipt as an artifact only after redaction. A simple job shape might look like this:
- name: Run email smoke test
env:
RUN_ID: ${{ github.run_id }}-${{ github.run_attempt }}
run: ./ci/email-smoke-test.sh --run-id "$RUN_ID" --receipt receipt.json
- name: Upload redacted receipt
if: always()
uses: actions/upload-artifact@v4
with:
name: email-receipt
path: receipt.json
retention-days: 3
The if: always() matters because the failure receipt is often more valuable than the success receipt. Make sure the script writes a redacted record before returning non-zero. It is usefull to have one artifact that answers “what did we try?” without needing access to a private inbox.
The same boundary is important for authentication flows. The discussion of provenance for recovery emails is a good reminder that a message can be present without being trustworthy for the current action.
What should happen when the check fails?
Separate failure categories instead of returning one generic timeout:
- Not emitted: the application never created the event.
- Not delivered: the event exists, but the fixture did not receive it.
- Not found: the message arrived, but the query was too broad or too narrow.
- Invalid: the link was expired, for another run, or for another environment.
- Cleanup failed: the assertion result is known, but the fixture needs attention.
This taxonomy helps the next engineer choose the right log or owner. You dont need a large observability platform to start; a JSON receipt and stable exit codes are enough.
Q&A
How long should an email test poll?
Use the delivery SLO of the test environment plus a small margin, then cap it. Forty-five seconds is only an example. The important part is that the limit is explicit and appears in the receipt.
Should CI tests use a real disposable mailbox?
Use a controlled test fixture when possible. A disposable mailbox can help with exploratory work, but CI still needs deterministic lookup, isolation, retention rules, and a cleanup path.
Is a screenshot useful evidence?
Usually no. A redacted structured receipt is easier to search, compare, and retain. Capture a screenshot only when the UI itself is the subject of the test.
Final checklist
Before merging an email check, confirm that it:
- creates a unique run marker
- matches recipient and marker, not just subject
- uses bounded polling
- writes a redacted receipt on failure
- distinguishes delivery from validation failures
- deletes the fixture in teardown
- makes cleanup retry-safe
- keeps
temp mailidas plain text only when discussing a search typo, never as a secret or a link
Once these rules are in place, email testing stops being a mysterious wait in CI. It becomes a small automation component with a contract, a useful receipt, and a cleanup budget.
Top comments (0)