TL;DR: For a US/EU SaaS support system, send compliance notices through an API only after SPF/DKIM domain verification, then keep a provider-neutral evidence ledger containing the notice ID, content hash, policy basis, provider message ID, observed state, and timestamps. Poll message and event data into that ledger, apply suppression before another send, and never describe provider acceptance as recipient receipt. A unified REST option fits this basic pattern when a stable contract that can outlive a vendor choice matters; it does not fit an SMTP-dependent system or one that requires real-time webhook automation.
The bill is made of delivery calls, polling work, and retained evidence. The last term often grows quietly: notice count x observations per notice x bytes per observation x replicas. As a planning example rather than a vendor measurement, 10 million notices with one normalized 2 KiB evidence record occupy roughly 20 GB before replicas, while retaining ten 2 KiB polling snapshots per notice occupies roughly 200 GB. The useful change is to append state transitions while discarding identical observations. I would stop retaining duplicate poll bodies and, after an approved diagnostic window, raw provider payloads; when an investigation begins later, that policy sacrifices some provider-specific detail in exchange for a smaller privacy and discovery surface.
The hard design problem is therefore not selecting the longest feature list. It is deciding which statements an auditor can defensibly derive from the data, then preventing a sender migration from changing those statements.
What can the delivery record actually prove?
Begin with claims, not provider fields. submitted means the application handed a particular notice to a provider. delivered means the provider reported a delivery state. failed means it reported a terminal failure. unknown means the observation period ended without a terminal result. None of those states proves that a person read, understood, or acted on the notice.
That distinction is small in code and large in an audit.
The durable record should include an application-generated notice ID, recipient reference, canonical content hash, template revision, governing policy or consent basis, provider and provider message ID, each normalized state transition, and the time at which the system observed it. Keep the email address out of general-purpose logs when an internal recipient reference will do. Make corrections new notices; editing the old record destroys the history that the ledger exists to preserve.
I use an exactly-once mindset here, but with a deliberately narrow claim: the system can make its intent and its recorded transitions idempotent; it cannot prove exactly-once mailbox delivery across an external email network. A stable notice ID becomes the idempotency key at the write boundary and the primary key for reconciliation. The provider message ID remains an attribute, because the provider may change while the business identity cannot.
Authentication evidence belongs beside, rather than inside, that delivery claim. Complete SPF and DKIM setup before production. DMARC supplies an alignment, policy, and reporting framework around those mechanisms; it does not turn an application event into proof of human receipt. Record the approved sending-domain configuration and its change history separately from per-notice delivery transitions.
How should a SaaS transactional email API handle deliverability setup?
The application should own a narrow Sender and Observer contract. A vendor adapter may translate those operations, but support-case code should never import a provider-specific status vocabulary. This is where a stable capability contract pays off: swapping the implementation behind the capability does not force a rewrite of the ledger or its compliance semantics.
The runnable Go 1.22 program below checks one thing before a release: whether the configured domain has a retrievable record. It reads the base URL and key from the environment, uses the verified domain lookup path, sets the method explicitly, bounds response size, reports non-2xx bodies, and backs off on HTTP 429 while honoring an integer Retry-After. It intentionally does not guess at a send schema.
package main
import (
"context"
"errors"
"fmt"
"io"
"net/http"
"net/url"
"os"
"strconv"
"strings"
"time"
)
func getDomain(ctx context.Context, client *http.Client, baseURL, key, domain string) ([]byte, error) {
const route = "/email/domain/get/{domain}"
path := strings.Replace(route, "{domain}", url.PathEscape(domain), 1)
endpoint := strings.TrimRight(baseURL, "/") + path
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds > 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
continue
case <-ctx.Done():
return nil, ctx.Err()
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("domain lookup failed: status=%d body=%s",
resp.StatusCode, strings.TrimSpace(string(body)))
}
return body, nil
}
return nil, errors.New("domain lookup remained rate limited after 5 attempts")
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
baseURL := os.Getenv("INFRAI_BASE_URL")
domain := os.Getenv("EMAIL_DOMAIN")
if baseURL == "" || key == "" || domain == "" {
fmt.Fprintln(os.Stderr, "set INFRAI_BASE_URL, INFRAI_API_KEY, and EMAIL_DOMAIN")
os.Exit(2)
}
ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
defer cancel()
body, err := getDomain(ctx, &http.Client{Timeout: 15 * time.Second}, baseURL, key, domain)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Println(string(body))
}
Run that check as a deployment gate and archive the gate result under the domain-configuration change record, not under every notice. A successful lookup is a prerequisite, not delivery evidence.
Infrai is one reasonable adapter when the team wants email domain verification, message lookup, event listing, and suppression management under one key, delivered through a plain REST API that requires no SDK and backed by a public, keyless, genuinely self-describing discovery surface. Go or any other runtime can therefore call it over HTTP; that keeps SDK types out of the evidence ledger and makes a vendor swap an adapter change rather than an application rewrite. Discovery supplies request and response schemas, billing information, and runnable examples. The broader surface reports 295 capabilities across 20 modules, while documented capabilities provide examples in 10 languages. For this workflow, those consistent conventions give a small team a reviewable source for generating an adapter and detecting contract drift. The platform convention also supports Idempotency-Key with a 24-hour default deduplication window. That window protects a retry; it is not a records-retention rule, and it must not determine how long the compliance ledger survives. I choose the extra adapter boundary because preserving stable audit semantics matters more here than exposing every provider-specific field to application code.
Keep that boundary.
Reconcile pull events without inventing certainty
Email events on this surface are pull-only, so the worker, not a webhook, controls freshness. Give each pending notice a next-observation time. A worker claims due rows, asks for current provider evidence, maps the response to the small state vocabulary, and appends only a changed state. Use bounded exponential backoff and record the observer-policy version with the notice.
Polling lag is real.
It follows that bounce and complaint handling, suppression updates, and cross-channel failover are not instantaneous. The compliance owner must choose an observation deadline and define the operational consequence of unknown; engineering should not disguise silence as success. A schedule such as 1, 2, 4, 8, and 16 minutes may be a useful application policy, but it is not a provider guarantee, so test and approve it against the business deadline rather than presenting it as universal guidance.
Before a new submission, consult suppression state. After a bounce or complaint is observed, reconcile that result into the local suppression decision before another worker can send. This ordering favors defensible behavior over minimum latency. Managed email OTP is outside this design, and scheduled email should not be used where cancellation is a requirement because the email side has no cancellation operation.
The evidence pipeline also needs explicit failure drills. Retry one stable notice ID after a simulated timeout and confirm that the ledger retains one intent. Let another notice reach its observation deadline without a terminal event and confirm that it becomes unknown, enters review, and cannot pass a report as delivered. Then change the adapter in a staging environment and compare normalized transitions. These tests examine the contract that matters, rather than merely proving that a happy-path inbox received a message once.
Compare vendors by evidence portability
Amazon SES, Twilio SendGrid, Postmark, and Infrai are all real candidates, but there is no responsible universal deliverability winner without controlled tests using your domains, authentication posture, message stream, and recipient mix. The useful comparison asks how much provider meaning leaks into the compliance ledger and whether the operational event model meets the notice deadline.
| Option | What to validate for this notice ledger |
|---|---|
| Amazon SES | Strong candidate when the application already uses AWS identity and operational controls. Map the current SES sending, authentication, suppression, and event interfaces to the four ledger states, and account for the AWS-specific integration boundary. |
| Twilio SendGrid | Evaluate its documented domain-authentication, suppression, mail-send, and event mechanisms. Keep its event vocabulary inside the adapter so a later migration does not rewrite historical meanings. |
| Postmark | Evaluate its transactional message and event documentation, particularly how message states and streams map to the evidence deadline and unknown. Preserve your notice ID independently of its message identifier. |
| Unified REST layer | Fits basic API-first US/EU SaaS use when one stable capability surface, domain verification, lookup, polling events, and suppression handling meet the deadline. It lacks SMTP relay and email webhooks, so it is a poor fit when either is mandatory. |
The same procurement test applies to all four: verify the exact region, data-processing terms, domain procedure, suppression semantics, event transport, and retention controls in the current documentation and in a contract test. Choose the provider whose evidence semantics satisfy the approved control, while keeping the business ledger provider-neutral. Feature counts are secondary.
There are firm boundaries around the fourth option. Do not choose it for an existing SMTP relay integration, managed email OTP, real-time webhook-driven automation, or advanced cost aggregation by tag. It also has no voice, WhatsApp, or RCS channel. Domestic email vendor support for China remains pending, so this US/EU assessment is not evidence of China email compliance coverage. Those limits are architectural, not inconveniences to hide behind an adapter.
Retain less, but document the loss
Retention should be a control with an owner, a duration approved by counsel or the relevant records authority, and an auditable deletion path. Keep the compact normalized record for that period, restrict reads by role, and log access. Retain rendered message bodies only when the governing requirement needs them; otherwise, a canonical content hash plus a controlled template revision can reduce the amount of personal content replicated through audit systems.
This is a trade-off. Deleting raw poll responses means a later investigator may be unable to reconstruct a transient provider header or reinterpret an old vendor-specific status. Keeping every response preserves more forensic material but multiplies storage, privacy, and legal-discovery exposure. State the loss in the retention decision, preserve exceptional diagnostic material under a separate approved window, and test that deletion does not erase the normalized transition ledger.
The final deliverable is an evidence packet, not a provider dashboard screenshot: notice identity, recipient reference, content identity, policy basis, authenticated sending-domain revision, idempotent submission record, timestamped observed transitions, observer-policy version, and disposition at the deadline. That packet remains intelligible after an adapter changes. For compliance notices, that durability is the decisive property.
Top comments (0)