TL;DR: For a marketplace contact form, choose an email API that can return durable bounce, complaint, suppression, and domain-health signals without taking ownership of queue routing or transactional templates. Keep the template revision, case ID, recipient decision, and support queue in application storage. Treat push events as the fast path and cursor-based polling as repair. The decisive test is whether a replayed event changes business state once.
That boundary matters more than a large feature matrix. I first treated a successful scheduled run as useful evidence; after being paged by missed jobs and duplicate deliveries, I stopped doing that. The system of record must make both absence and replay explainable. A successful send request proves acceptance of work, not inbox delivery, and a provider-hosted template name cannot explain why a marketplace case went to Trust instead of Payments.
The invariant is short: the marketplace owns intent; the delivery API reports outcomes.
Which API should I use for email deliverability monitoring and bounce events?
A marketplace contact submission combines facts that change at different rates. The buyer's case ID and selected topic are business data. Queue routing is support policy. The subject, HTML, and locale are reviewed content. Bounce classification and complaint reports are transport outcomes. Putting all four behind one remote template identifier makes a content edit capable of changing an operational workflow, and it makes a provider migration larger than the sending boundary.
Application ownership means recording an immutable template revision beside the case, rendering the message under review, and sending already-rendered content through a narrow interface. The API receipt is attached later. This is less convenient than letting a hosted editor control everything: the application team now owns escaping, previews, localization tests, and deployment order. I accept that cost for transactional contact-form mail because support must be able to reconstruct the exact message and routing decision from its own records.
Here is the contract. The example runs as a Go program and deliberately contains no provider vocabulary.
package main
import (
"context"
"fmt"
)
type Message struct {
OperationID string
CaseID string
Queue string
Recipient string
TemplateVersion string
Subject string
HTML string
}
type Receipt struct {
TransportMessageID string
}
type Sender interface {
Send(context.Context, Message) (Receipt, error)
}
func supportQueue(topic string) (string, error) {
switch topic {
case "unsafe-listing":
return "trust", nil
case "payment-dispute":
return "payments", nil
case "account-access":
return "account-support", nil
default:
return "", fmt.Errorf("unmapped contact topic %q", topic)
}
}
func main() {
queue, err := supportQueue("unsafe-listing")
if err != nil {
panic(err)
}
fmt.Println(queue)
}
OperationID identifies one logical notification and should be used as an idempotency key when an API documents that capability. CaseID remains the marketplace's workflow key. The returned transport ID is correlation data, never the primary identity of the support case.
Do not hide the template boundary during evaluation. Ask whether the API accepts rendered text and HTML, whether custom correlation metadata survives into feedback, and whether a receipt can be joined to later events. Those answers determine whether the adapter stays small. A polished hosted-template editor does not answer them.
Make feedback a state transition, not an alert stream
Delivery feedback has to change application behavior. Normalize incoming records into a small internal vocabulary, retain the original payload for audit, and apply suppression before another send is admitted. Useful states include accepted, delivered, temporary failure, permanent failure, complaint, and unsubscribe; preserve diagnostic details separately instead of forcing every transport-specific code into support policy.
Duplicates are expected at this boundary, so processing must be idempotent. The following in-memory example shows the rule. Production storage should enforce uniqueness on the source plus source event ID inside the same transaction that records suppression.
package main
import (
"fmt"
"sync"
"time"
)
type EventKind string
const (
PermanentFailure EventKind = "permanent_failure"
Complaint EventKind = "complaint"
Unsubscribe EventKind = "unsubscribe"
)
type DeliveryEvent struct {
Source string
SourceEventID string
Recipient string
Kind EventKind
OccurredAt time.Time
Raw []byte
}
type Repository struct {
mu sync.Mutex
seen map[string]struct{}
suppressed map[string]EventKind
}
func (r *Repository) Apply(event DeliveryEvent) (bool, error) {
r.mu.Lock()
defer r.mu.Unlock()
if event.Source == "" || event.SourceEventID == "" || event.Recipient == "" {
return false, fmt.Errorf("source, event ID, and recipient are required")
}
key := event.Source + ":" + event.SourceEventID
if _, exists := r.seen[key]; exists {
return false, nil
}
switch event.Kind {
case PermanentFailure, Complaint, Unsubscribe:
r.suppressed[event.Recipient] = event.Kind
}
r.seen[key] = struct{}{}
return true, nil
}
func main() {
repo := &Repository{
seen: make(map[string]struct{}),
suppressed: make(map[string]EventKind),
}
event := DeliveryEvent{
Source: "transport-a",
SourceEventID: "event-1042",
Recipient: "seller@example.test",
Kind: Complaint,
OccurredAt: time.Now().UTC(),
}
first, _ := repo.Apply(event)
second, _ := repo.Apply(event)
fmt.Println(first, second, repo.suppressed[event.Recipient])
}
The output is true false complaint. One event, delivered twice, creates one transition. Simple.
Suppression must also be checked where a send is created, preferably in the transaction that writes an outbox record. A periodically refreshed cache alone leaves a race between reading eligibility and enqueueing mail. Store the reason, effective time, and source event so an operator can distinguish a complaint from a permanent failure without opening a raw payload.
RFC 8058 specifies one-click unsubscribe through the relevant message headers and an HTTPS POST. It does not classify every marketplace support reply as subscription mail, and it does not replace consent policy. Where the message category uses that mechanism, model unsubscribe as a first-class event and suppress accordingly.
Separate the fast path from the repair path
Use authenticated push delivery for low-latency feedback, but acknowledge only after durable admission to a queue or log. Parsing and state changes can happen asynchronously. Invalid payloads need a visible dead-letter path; returning success and discarding them turns an integration change into silent data loss.
Polling serves a different purpose. It repairs gaps after consumer downtime or a deployment error. A candidate API therefore needs documented event retention, stable pagination or cursor semantics, and enough event identity to overlap windows safely. A scheduler that exits successfully while reading the same old checkpoint is not healthy.
package main
import (
"context"
"fmt"
)
type Page struct {
EventIDs []string
NextToken string
}
type EventSource interface {
List(context.Context, string, int) (Page, error)
}
type Checkpoints interface {
Load(context.Context, string) (string, error)
Store(context.Context, string, string) error
}
func repair(ctx context.Context, source EventSource, state Checkpoints) error {
token, err := state.Load(ctx, "email-feedback")
if err != nil {
return err
}
for pages := 0; pages < 25; pages++ {
page, err := source.List(ctx, token, 200)
if err != nil {
return fmt.Errorf("read feedback after checkpoint: %w", err)
}
for _, id := range page.EventIDs {
fmt.Println("admit", id)
}
if page.NextToken == "" || page.NextToken == token {
return nil
}
if err := state.Store(ctx, "email-feedback", page.NextToken); err != nil {
return err
}
token = page.NextToken
}
return fmt.Errorf("repair stopped at the 25-page run limit")
}
func main() {}
The 200-event page and 25-page ceiling are example operating bounds, not claims about an external API. Configure the page size below the selected API's documented limit. Bound each run, back off after transient failures, and alert on the age of the newest completely applied event. I care more about that age than a count of successful cron executions because it exposes a stuck checkpoint.
Domain health belongs on the control plane too. Persist observations of authentication or verification state and alert on sustained bad state. Do not place a live domain-health request in the contact-form submission path; that converts monitoring availability into customer-facing availability.
Run a template-ownership incident drill
Select with evidence from a bounded drill, not screenshots. Use synthetic recipients and a non-production sending domain, then save the results as integration fixtures.
Start by rendering two template revisions, such as support-en-v17 and support-en-v18, for the same marketplace case and confirm that the stored revision reconstructs each message without consulting the delivery system. Submit the same OperationID twice and record the API's documented idempotency behavior. Replay one complaint event concurrently through two consumers; the database should retain one event and one suppression decision. Pause push consumption, create feedback using documented test facilities, restore the consumer, and prove polling closes the gap from its prior checkpoint.
Finally, rotate the push-signing secret according to the candidate's documented procedure. Record whether verification can accept old and new credentials during the transition. This is where an integration that looked adequate on a capability grid often becomes operationally legible, or does not.
Use a small evidence table during the exercise:
| Boundary | Evidence to retain | Incident question answered |
|---|---|---|
| Template | Case ID, revision, rendered-content hash | What did we intend to send? |
| Submission | Operation ID, receipt, acceptance time | Was work admitted once? |
| Feedback | Source event ID, raw payload, normalized state | What outcome was reported? |
| Suppression | Recipient, reason, effective time | Why was a later send blocked? |
| Repair | Previous and next checkpoint, newest event time | Did the gap close? |
| Domain | Observed state and transition time | When did authentication health change? |
US and EU operation adds data-governance questions to the same drill. Map where recipient addresses, message content, raw feedback, and logs are processed and retained. Under GDPR, controllers and processors have distinct obligations, and transfers of personal data require an appropriate basis. A region selector by itself is not the answer. Review the actual data flow, retention controls, subprocessors, and contractual terms with the teams responsible for privacy and legal decisions.
SMS compliance is a separate transport concern. For example, US application-to-person messaging over 10-digit long codes has registration and campaign requirements documented under A2P 10DLC. Do not treat an email suppression record as universal consent for SMS, or vice versa; keep channel, purpose, source, and effective time in the policy record.
Where this boundary does not fit
The main limitation of application-owned templates is operational: the application team must maintain rendering, previews, localization, escaping, and coordinated releases. That trade-off is a poor fit when one operational team intentionally delegates segmentation, consent, approval, and content release to a managed campaign system. The better alternative in that case is the campaign system's native template workflow. Its hosted template becomes part of the chosen system of record; document the dependency, export revisions for audit where supported, and keep transactional support routing out of campaign content.
It also does not remove the need to compare retention windows, signing mechanisms, regional processing, rate-limit behavior, or authentication status. It orders those questions. First establish who owns intent; then verify that feedback can be reconciled to it under replay and outage.
For marketplace contact mail, my final selection criterion is concrete: after the drill, can the team explain a case from template revision through suppression using retained application state, then repair missing feedback without sending the message again? If yes, the API boundary is workable. If no, another dashboard will not make it operable.
Sources
- RFC 8058, Signaling One-Click Functionality for List Email Headers: https://datatracker.ietf.org/doc/html/rfc8058
- GDPR, Regulation (EU) 2016/679: https://eur-lex.europa.eu/eli/reg/2016/679/oj
- Twilio, US A2P 10DLC compliance documentation: https://www.twilio.com/docs/messaging/compliance/a2p-10dlc
Top comments (0)