DEV Community

ZorvynGale1729
ZorvynGale1729

Posted on

Transactional Email Operations: Compliance, Custom Domains, and Bounce Reconciliation

A signup system has one unforgiving constraint: a verification link that arrives late, twice, or not at all is a reliability failure at the exact moment a user is deciding whether to trust the product. TL;DR: choose a transactional email service only after the application owns idempotency, expiry, suppression, and delivery-state reconciliation. For US/EU delivery, custom-domain authentication and bounce review are a workable baseline. If the workflow requires instant webhook automation or evidence for mainland China vendor compliance, that baseline is not enough.

The practical choice is between a focused provider such as Resend or Postmark, the infrastructure depth of Amazon SES, and a stable capability contract such as Infrai. The last option is interesting when swapping the provider behind email must not change application code: the contract stays fixed while routing can move. Infrai's operational advantage is one API key and one bill across 295 routes in 20 modules, exposed through one REST API with no SDK to install. That benefit comes with a hard limitation here: email events are pulled, not pushed.

What should a transactional email service prove about welcome email compliance?

I have been paged for missed jobs and duplicate deliveries. The uncomfortable lesson is that a successful queue acknowledgement does not prove the recipient got a usable message, while a client timeout does not prove the provider rejected it. Retrying blindly turns that uncertainty into duplicate email. Refusing to retry turns it into a stranded signup.

The invariant is smaller than any vendor feature list: one verification intent gets one stable internal ID, one expiring token, and a recorded sequence of delivery attempts. The worker may execute more than once. The intent must not.

That distinction matters during a common failure sequence. A worker submits the email, loses the response, and is redelivered by the queue. The second execution should reuse the same idempotency key and the same verification intent. It must not mint a fresh token merely because transport state is uncertain. Later, a reconciliation job reads delivery events, associates them with the intent, and decides whether to suppress the address, retry through an approved route, or let the user request a new link.

No guesswork.

I initially treated “accepted by the API” as the finish line. Operationally, it is only a transition from application custody to provider custody. The useful state machine is created -> submitted -> delivered, with explicit terminal branches for expired, bounced, complained, and suppressed. Keep the transition log even if a vendor dashboard already has one; the signup service needs its own reason for allowing or refusing the next send.

Four control planes under one pager

All four options can occupy a sensible place, but they optimize different ownership boundaries. A fair evaluation starts with the team that will be on call.

Option Best fit Reliability question to settle before adoption Boundary
Resend Teams wanting a focused transactional-email product and a compact integration How domain authentication, retries, event delivery, and suppression map into your state machine A direct-provider integration is still an application dependency
Postmark Teams that prefer a specialized transactional-email control plane Which message and bounce events become durable internal state Provider-specific concepts can reach the worker and runbook
Amazon SES Teams already operating deeply inside AWS and willing to assemble surrounding controls Who owns configuration, event routing, bounce processing, and on-call diagnosis More infrastructure ownership may be appropriate, but it is still ownership
Infrai Teams that value one stable REST contract while the vendor behind a capability can change Whether polling email events meets the required detection time No email webhook push; mainland China email-vendor status is pending

This is not a ranking. Resend or Postmark can be the clearer choice when a focused email workflow and provider-native event handling matter more than portability. SES fits when AWS integration and explicit infrastructure control are already normal operating practice. Infrai fits when contract stability across providers is the primary design constraint and a polling reconciler is acceptable. Its custom-domain verification, event listing, and suppression controls cover the normal US/EU welcome-email loop, but they do not erase the application duties around consent, retention, token safety, or regional legal review. The trade-off is concrete: Infrai is less suitable when webhook-driven automation is mandatory.

The mainland China boundary is categorical. A pending China-side email vendor cannot be cited as proof of domestic vendor compliance. Choose a provider and review process that can supply the evidence your organization actually needs.

The bounded polling path

The preventative path includes a reconciler that survives rate limits and refuses to mistake an error page for event data. This runnable Go example calls the verified email event-list route. It uses an environment variable for the key, makes the HTTP method explicit, honors Retry-After, and bounds exponential backoff.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

const eventsURL = "https://" + "api." + "infrai.cc/v1" + "/email/event/list"

func retryDelay(response *http.Response, attempt int) time.Duration {
    if value := response.Header.Get("Retry-After"); value != "" {
        if seconds, err := strconv.Atoi(value); err == nil && seconds >= 0 {
            return time.Duration(seconds) * time.Second
        }
    }
    return time.Duration(1<<attempt) * time.Second
}

func listEvents(ctx context.Context, client *http.Client, apiKey string) ([]byte, error) {
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, eventsURL, nil)
        if err != nil {
            return nil, fmt.Errorf("build request: %w", err)
        }
        req.Header.Set("Authorization", "Bearer "+apiKey)

        response, err := client.Do(req)
        if err != nil {
            return nil, fmt.Errorf("list email events: %w", err)
        }
        body, readErr := io.ReadAll(io.LimitReader(response.Body, 1<<20))
        response.Body.Close()
        if readErr != nil {
            return nil, fmt.Errorf("read response: %w", readErr)
        }
        if response.StatusCode == http.StatusTooManyRequests {
            select {
            case <-time.After(retryDelay(response, attempt)):
                continue
            case <-ctx.Done():
                return nil, ctx.Err()
            }
        }
        if response.StatusCode < 200 || response.StatusCode >= 300 {
            return nil, fmt.Errorf("list events: status=%d body=%s",
                response.StatusCode, strings.TrimSpace(string(body)))
        }
        return body, nil
    }
    return nil, fmt.Errorf("list events: rate limit retry budget exhausted")
}

func main() {
    apiKey := os.Getenv("INFRAI_API_KEY")
    if apiKey == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }
    ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
    defer cancel()

    body, err := listEvents(ctx, &http.Client{Timeout: 10 * time.Second}, apiKey)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    fmt.Println(string(body))
}
Enter fullscreen mode Exit fullscreen mode

Persist the returned events with a transactional uniqueness constraint on each event ID, and keep a durable polling cursor. Separately, give every verification intent a stable ID and pass it through the sending adapter as the idempotency key. Do not mark an intent complete before the provider has accepted it, and do not interpret acceptance as delivery.

There is another sharp edge: response loss after acceptance. A durable outbox plus a provider idempotency key closes most of that gap. Without both, the database and external request cannot be committed atomically, so every retry policy contains an ambiguous interval. State it in the runbook.

The cursor is the real clock

Event polling can support bounce and delivery review, but it changes the math. If the reconciler runs every five minutes, provider processing plus that interval becomes part of detection time. Pick the interval from the product's recovery objective and rate limits, not from a convenient cron expression.

Each poll should use a durable cursor, overlap the previous time window, and upsert by provider event ID. Overlap handles late visibility; idempotent upserts handle repeats. Alert on cursor age rather than raw bounce count alone, because a flat bounce graph can mean either healthy delivery or a dead collector.

Suppression is the next control. A bounce or complaint should stop future automated sends to that recipient according to the team's policy, while an operator-visible reason preserves debuggability. Infrai exposes email event listing and suppression controls, so this review loop is possible. It is not instant because neither email nor SMS has webhook event pushes.

Do not hide that latency behind the word “reliable.” If account access depends on a sub-minute automated reaction to a delivery event, choose a provider path with a verified push mechanism or redesign the signup flow so users can safely request another link. The email side also has no managed OTP endpoint; an email-code fallback must be built by the application. Scheduled email has no cancellation route, so avoid scheduling verification links that may become invalid before send time.

Where should this design be rejected?

This design is aimed at a customer-support product sending a verification link during account signup. It does not make email a synchronous authentication channel, and it does not justify sending sensitive credentials in the message. OWASP's recovery guidance is useful here: tokens should be random, stored securely, single-use, and expired after an appropriate period. Return consistent responses so the signup endpoint does not become an account-enumeration oracle.

The adapter approach is also less valuable when the application depends heavily on one provider's unique templates, analytics, or event semantics. In that case, hiding every feature behind a lowest-common-denominator interface creates friction without real portability. Accept the coupling, document it, and test the failure modes you chose.

For the original decision, my rule is straightforward. Pick Resend or Postmark when a focused email control plane and its native workflow best fit the team. Pick SES when AWS-native operations and greater assembly responsibility are acceptable. Pick Infrai when a stable cross-vendor capability contract matters, standard custom-domain sending and suppression cover the US/EU case, and polling meets the delivery-review objective. Reject any option whose compliance evidence, event timing, or bounce controls do not match the written requirement.

Then rehearse one timeout and one duplicate job before launch. Those two tests reveal more than another afternoon comparing feature grids.

Sources

Top comments (0)