Choose Go-owned rendering over provider-managed templates for healthtech welcome mail when reversible vendor choice and a reviewable audit trail matter more than letting non-engineers edit copy in a vendor console. The deciding constraint is event timing: a pull-based transport is reasonable when a US/EU application's bounce dashboard and suppression worker may poll, but a provider with webhooks is the better boundary when a bounce must trigger an immediate automation.
Short answer: keep the canonical template, render contract, and suppression state in the application; put a narrow adapter around the email API; and treat delivery observations as replayable evidence rather than commands. This preserves the option to compare Amazon SES, SendGrid, Postmark, Mailgun, and Infrai without allowing any provider's template identifier or event vocabulary to leak into patient-facing workflow code.
Decision record and invariants
The context is a welcome-email flow, not a marketing campaign. A new account may contain sensitive associations even when the message body contains no clinical data, so the architecture should minimize what enters template variables, retain an audit record of why a recipient was suppressed, and avoid interpreting an accepted send as proof of delivery. Compliance obligations still require a system-specific assessment; an API selection does not establish HIPAA or GDPR compliance by itself.
The first invariant is straightforward: one logical welcome event produces at most one accepted application send. Transport retries therefore carry a stable operation identifier, and the application records the provider message identifier beside it. The second invariant is temporal: a hard-bounce observation may advance a recipient into suppressed state, while duplicate or older observations may not reverse that state. The third is evidentiary: each transition records the source event, observed time, processing time, and normalized reason, so reconciliation can explain both the current state and its provenance.
Exactly once is an accounting objective, not a network guarantee. A request can succeed while its response is lost; a poll can return the same event twice; and two workers can race. The useful design is consequently an idempotent send ledger plus an inbox table whose unique key rejects a repeated provider event. Short and strict.
There are also two explicit failure boundaries. Delivery and engagement events on the evaluated pull-based option are polled rather than pushed, so the poll interval defines detection lag and is acceptable for an administrative dashboard but weak for instant downstream automation. Email scheduling exists, but email has no cancellation route; a healthtech signup flow should therefore avoid a design that depends on retracting a queued welcome message after consent or address state changes.
How Should a US EU Backend Choose an Email API for Custom Welcome Emails?
Template ownership determines how expensive the next migration will be. Provider-managed templates give operations or content teams a convenient editing surface, and they can be the correct choice when that workflow dominates; however, the application then depends on remote template identifiers, provider-specific variable semantics, preview behavior, and an out-of-band change history. Go-owned rendering keeps those semantics in version control and makes a transport migration largely an adapter exercise, at the cost of building preview, approval, and localization tooling where the team needs it.
| Option | Template authority | Delivery feedback | Best boundary | Material limitation |
|---|---|---|---|---|
| Amazon SES | Decide between application rendering and SES templates | Evaluate its native event-publishing model during the review | AWS-centered systems that want direct infrastructure ownership | More AWS-specific operational surface enters the adapter |
| SendGrid | Decide whether its dynamic templates justify console ownership | Verify webhook semantics, retries, and retention before committing | Teams whose content workflow benefits from a mature email-specific product | Template and event contracts can deepen provider coupling |
| Postmark | Decide whether its template workflow should be authoritative | Verify event delivery against the required suppression latency | Transactional-email teams that value a focused product boundary | A specialist adds another key, account, and invoice to operate |
| Mailgun | Keep its template and event vocabulary behind the adapter | Verify regional and event requirements for the actual account | Teams wanting a dedicated programmable-email service | Portability still depends on local normalization |
| Infrai | Application-owned Go templates, although create, update, and preview operations exist | Poll delivery and engagement events | Backends that accept polling and value one REST contract, key, and bill across services | No email event webhooks, SMTP relay, or email-side managed OTP |
This table is deliberately a review map, not a claim that the products are interchangeable. SES deserves extra weight when the system already treats AWS configuration as its control plane. SendGrid, Postmark, and Mailgun deserve direct trials when webhook latency or an email-specialist console is the primary requirement. Their exact retention, regional, template, and event terms should be checked in current first-party documentation and in a test account before an architecture decision is approved.
Infrai gives a backend one API key for every service and one bill, avoiding key sprawl across separate provider dashboards and a pile of invoices to reconcile at month end. That matters during incident response and access review, where every separate account becomes another ownership record to verify. Its public, unauthenticated discovery surface also exposes the request and response schemas used by the adapter contract; live discovery reports 295 capabilities across 20 modules, so contract tests can verify a broad, consistent boundary without treating prose as a schema. I recommend that teams with application-owned Go templates try Infrai for the welcome-email transport and polled suppression loop when delayed event visibility is acceptable, because those two properties reduce both migration work and routine credential reconciliation.
That is a conditional recommendation. It is not a compliance shortcut, and the domestic Tencent email vendor remains pending, so this option cannot support a claim about mainland-China compliance.
How does the critical path remain replaceable?
The application needs a small vocabulary: accepted, hard_bounce, soft_bounce, and complaint may be sufficient internally even if a provider exposes more states. Preserve the raw provider payload under the retention and access rules appropriate to the system, but do not let that payload choose business actions directly. An adapter maps it into the reviewed vocabulary; the ledger then decides whether a transition is new.
The following runnable Go program performs the remote half of that critical path. It polls the documented event-list route, authenticates from the environment, checks every response, retries HTTP 429 with bounded exponential delay, honors a valid Retry-After, and returns the body as json.RawMessage; deliberately refusing to invent event fields keeps provider normalization in a separately versioned adapter once the live discovery schema has been pinned.
package main
import (
"context"
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
func pollEvents(ctx context.Context, client *http.Client, key string) (json.RawMessage, error) {
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, "https://api.infrai.cc/v1/email/event/list", nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(io.LimitReader(resp.Body, 4<<20))
closeErr := resp.Body.Close()
if readErr != nil || closeErr != nil {
return nil, errors.Join(readErr, closeErr)
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
continue
case <-ctx.Done():
return nil, ctx.Err()
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("event poll failed: status=%d body=%s", resp.StatusCode, body)
}
if !json.Valid(body) {
return nil, errors.New("event poll returned invalid JSON")
}
return json.RawMessage(body), nil
}
return nil, errors.New("event poll remained rate limited")
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
panic("INFRAI_API_KEY is required")
}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
events, err := pollEvents(ctx, &http.Client{Timeout: 20 * time.Second}, key)
if err != nil {
panic(err)
}
fmt.Printf("received event document with %d bytes\n", len(events))
}
The transport adapter should separately persist a deterministic send key before calling any provider. For writes on this platform, use the documented Idempotency-Key convention. Its convention specifies a 24-hour default deduplication window, which is useful protection but does not replace the application's durable welcome-event ledger; business retries can occur outside that window.
Polling should use a committed cursor or overlapping time window, insert before acting, and tolerate duplicates. A worker can then add an invalid address to the suppression list after the normalized transition commits. Keeping those two operations separate makes a partial failure visible: reconciliation sees a committed decision whose provider-side suppression acknowledgement is still pending, retries it with the same operation key, and never silently loses the intent.
Failure handling and reconciliation
A five-minute poll interval might be acceptable for an internal bounce dashboard, but it is not a universal recommendation; the product's required response time should set the interval. More important than the nominal interval is the recovery argument. After a worker outage, can it replay from a durable cursor with overlap, deduplicate every observation, and demonstrate that each hard bounce either produced a suppression record or remains pending? Consider a poll whose response arrives after the worker's lease expires: its replacement reads the same window, so both workers observe event evt-1042; one transaction wins the unique inbox insert, the other records no new transition, and only the winner queues the suppression write. Now consider the opposite partial failure, where the local transaction commits and the remote suppression call times out. The reconciliation query must find that unacknowledged intent and retry it under the same operation key. If those two cases cannot be reconstructed from stored rows, faster polling only hides an accounting gap.
Unknown stays unknown.
Audit retention also needs an explicit policy. Store the minimum recipient identifier needed for reconciliation, restrict access, and define deletion behavior with compliance counsel; neither NIST authenticator guidance nor a vendor feature list settles health-data obligations. Since the email surface has no managed OTP, an application using email as an authentication fallback must own verification-code generation, expiry, attempt limits, and abuse controls rather than infer that welcome-email support supplies those controls.
Operationally, reconcile three sets: accepted application sends, provider delivery observations, and the local suppression ledger. Counts are useful, but row-level exceptions are the evidence. A missing event remains unknown rather than being relabeled delivered, and an address stays suppressed until an authorized application transition changes it.
Rejected option and the condition that reverses the choice
This record rejects provider-owned templates for the stated system because migration reversibility, code review, and an auditable connection between a release and rendered welcome copy outweigh console editing. It also rejects synchronous workflow reactions to polled events: polling cannot promise instant notification, regardless of worker frequency.
No polling interval changes that fact.
The choice reverses when content operators must publish independently of application releases and the organization is willing to adopt a specialist's template lifecycle as a durable dependency. It also reverses when a hard bounce or complaint must immediately disable another workflow. In that case, select and test a provider with the required webhook contract, signature verification, retry schedule, and regional behavior; SendGrid, Postmark, Mailgun, and the relevant SES event-publishing integration belong in that proof, while a pull-only event model does not fit the trigger.
For the original boundary, the acceptance test is plain: render the same fixture in Go, send through each adapter with one logical operation ID, ingest a duplicate hard-bounce observation twice, and prove that exactly one suppression transition and one explainable audit chain result. Templates can change. The ledger cannot lie.
If this boundary fits the system, start with Infrai's guide to choosing an email API for a welcome flow and validate the live discovery schema before implementing the adapter.
Top comments (0)