DEV Community

ZorvynGale1729
ZorvynGale1729

Posted on

Transactional Email API Deliverability Setup — Trace Domain Verification and Bounce Alerts

The page says marketplace signups cannot verify their accounts. The queue looks healthy, but the verification links have not reached recipients. Short answer: use a transactional email HTTP API with domain authentication, asynchronous delivery events, and an exportable suppression state; choose the integration that lets your team connect those events to each signup attempt with the least custom plumbing. An accepted API request is not a delivered message. Treat the time between signup and a usable link as the operational signal.

What should a transactional email API deliverability setup reveal about domain verification?

A queue-depth alarm can miss this incident entirely: workers dequeue jobs, the email API accepts requests, and the signup flow still stalls. A provider acceptance response establishes only that the request crossed one boundary. Delivery, bounce, and suppression are later states. The on-call view should show a time window, affected signup attempts, and the oldest attempt without terminal delivery evidence. Keep account identifiers out of metrics labels; use an opaque attempt ID to join detailed events in logs.

Queue depth is zero. The incident isn't.

Start with the user's clock. Record signup_started_at, send_requested_at, api_accepted_at, and the timestamps of subsequent provider events as distinct observations. A dashboard that collapses those into one sent timestamp hides the very interval the page is about. The early warning is a rising age of accepted attempts with no matching delivery or terminal failure event, measured alongside the rate of actual verification completions. A missing event may also mean the event feed is late, so alarm on feed freshness separately before assigning blame to mail delivery.

Where should the earlier signal come from?

For each attempt, persist a durable record before making the API call. Give retries the same application attempt ID and reconcile uncertain responses before resending. Network timeouts are ambiguous. Repeating a request after an acceptance that the client never saw can send two valid links. A short-lived token and a single-use redemption rule limit the harm, but neither makes duplicate emails pleasant for a new user.

Map provider callbacks or polled events into your own states: requested, accepted, delivered, bounced, suppressed, and unknown. Store the provider message ID when returned and keep the raw event identifier for deduplication. Process repeated notifications idempotently. Do not infer delivery from absence of a bounce. When events arrive out of order, preserve their observed timestamps and avoid letting an older acceptance overwrite a later bounce. A polling-only integration needs a documented cursor, retention window, pagination behavior, and recovery procedure; otherwise the integration effort has merely moved from webhook handling into scheduled reconciliation.

Domain setup belongs in this same trace. DKIM signs selected headers and the body; SPF authorizes sending hosts for an envelope domain; DMARC evaluates alignment with the visible From domain. A green DNS check by itself does not prove inbox placement. For a US/EU marketplace, decide which domain sends signup links, who owns its DNS changes, and how the team verifies authentication after deployment. Verify the actual From identity and resulting authentication results with controlled test mail, not just the existence of records.

The trade-off is visible when a request times out: sending again immediately reduces apparent latency but risks two links, while reconciling first delays the retry. Keep the attempt ID stable through both paths. If a callback arrives after the timeout, correlate it with the existing attempt; if a bounce follows acceptance, show the bounce as the current outcome without erasing the original acceptance timestamp. This makes the next page actionable instead of forcing the on-call engineer to reconstruct state from unrelated logs.

How do we change the instrumentation without losing the trail?

Add a synthetic signup-mail attempt to a controlled mailbox and measure its milestones separately from real signups. The synthetic check tests the API-to-mailbox path; the production cohort checks the user's path, including token redemption. Neither replaces the other. Use a bounded age threshold based on measured event lag and the signup experience your team is willing to accept, then review a sample of missing events before paging. No universal minute value follows from the email standards.

Here is the alert's core calculation in Go. It counts accepted attempts that have neither a terminal event nor a redeemed link after the chosen threshold; the caller must supply durable, deduplicated attempt records and a threshold derived from observed latency.

package signup

import "time"

type Attempt struct {
    AcceptedAt time.Time
    TerminalAt time.Time
    RedeemedAt time.Time
}

func UnresolvedOlderThan(attempts []Attempt, now time.Time, limit time.Duration) int {
    count := 0
    for _, a := range attempts {
        if a.AcceptedAt.IsZero() || !a.TerminalAt.IsZero() || !a.RedeemedAt.IsZero() {
            continue
        }
        if now.Sub(a.AcceptedAt) > limit {
            count++
        }
    }
    return count
}
Enter fullscreen mode Exit fullscreen mode

One count isn't a diagnosis.

The integration decision is mostly ownership. A straightforward HTTP send call is only the first piece. During a trial, walk through DNS verification, event ingestion or polling, replay after consumer downtime, suppression inspection, and a DNS rollover in a staging domain. Ask whether exports expose enough identifiers to reconcile an accepted request with a bounce and whether suppression updates can be reflected in your own records. Prefer the path whose failure states the on-call team can actually observe. This is a trade-off, not a vendor ranking.

Test the awkward states before rollout: duplicate callbacks, an API timeout after remote acceptance, an address already suppressed, and an event feed that pauses while sends continue. Keep a small reconciliation job that revisits attempts stuck in unknown state, with a bounded retry policy and an audit trail for manual resolution. If the provider offers only event polling, budget for cursor persistence and backfill tests. If it offers callbacks, budget for signature validation, deduplication, and a retryable ingestion endpoint. Either way, deployment is not finished when the first message arrives.

When does the alert become noise?

An aggressive threshold pages on normal event delay; a loose threshold discovers the failure after the signup cohort has already abandoned the flow. Track the distribution of acceptance-to-event latency and the age of unresolved attempts, and compare both with verification completions. A sustained increase in unresolved age while event-feed freshness is normal deserves investigation. A stale feed with steady verification completions calls for an instrumentation incident first.

Write the runbook around those branches. Check DNS authentication and recent changes, inspect suppression and bounce evidence for affected attempts, compare API acceptance with event-feed freshness, then decide whether to pause retries or reconcile unknown outcomes. Record why an alarm was false and adjust its threshold against observed data, not against the desire for a quiet pager. The cost of a false positive is repeated investigation and eventual distrust of the signal; the cost of a late page is a signup link that never becomes usable.

Further reading

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.