Route each media contact submission by a versioned policy, and record the evidence for that decision before attempting a welcome email or SMS. Compliance evidence is the deciding constraint: a transport acceptance response cannot explain which consent text a visitor saw, why a suppression rule applied, which support queue received the case, or which domain authenticated the message.
TL;DR: treat routing, consent, and channel eligibility as one deterministic decision whose inputs and policy version are stored in Postgres. Authenticate the email identity with SPF and DKIM, consult suppression before every attempt, keep SMS consent evidence separate from email eligibility, and make queue age plus unexplained decisions the operational signals. The SLO should cover decisions the system can prove, not inbox placement that it cannot control.
For a media site, the risky edge is easy to miss. The same form may receive a subscription question, an advertising request, a correction, or a sensitive newsroom tip, yet a cheerful auto-reply can cross a channel boundary before an operator has even seen the case. Start with evidence, then send.
What should a SaaS welcome email deliverability checklist prove?
An audit question is rarely “did the API return success?” It is closer to: “Given submission ct_01J9Y, why did the system route it to audience support, send one transactional acknowledgement from this domain, and decline SMS?” A useful record answers that without reconstructing mutable configuration from memory.
Capture an opaque submission identifier, receipt time, normalized category, destination queue, policy version, requested channels, consent artifact identifier, suppression decision, message class, template version, and the authentication configuration version expected for the sender. Do not put the contact's address, free-form message, or phone number into metric labels. Those values have high cardinality and may be sensitive; the decision record can refer to access-controlled data by opaque key.
Keep the raw form payload under its own access and retention rules. The routing evidence is a compact explanation, not a second copy of the message.
The distinction matters because channel permission and transport health answer different questions. An address may be eligible for a support acknowledgement while suppressed for promotional mail. A phone number supplied for a callback is not, by itself, evidence that automated text messages were requested. CTIA's messaging principles emphasize consumer choice and consent; the practical design consequence is to store what was presented and what action the visitor took, rather than infer permission later from the mere presence of a number.
I use three invariants for the runbook, because each can be tested without trusting a dashboard:
- Every accepted form submission produces exactly one versioned routing decision.
- No delivery attempt begins until the applicable suppression and consent checks are recorded.
- Every attempt points back to the decision, policy, template, and sender configuration that authorized it.
Exactly one does not mean “insert once and hope.” Put a uniqueness constraint on the submission identifier and decision kind, then make retries return the existing decision. A worker restart should not create a new explanation.
Make the safe decision before touching a transport
The policy should consume facts and return a result. Network calls do not belong inside it. This keeps the highest-value compliance behavior small enough for table tests and lets a deployment replay historical, redacted fixtures against a proposed policy without sending anything.
The Go example below deliberately rejects ambiguous inputs. It also separates a request for SMS from evidence that SMS is allowed; collapsing those fields is the sort of convenient shortcut that survives until the first compliance review.
package routing
import (
"errors"
"time"
)
type Channel string
const (
Email Channel = "email"
SMS Channel = "sms"
)
type Input struct {
SubmissionID string
Category string
EmailRequested bool
SMSRequested bool
SMSConsentID string
EmailSuppressed bool
SMSSuppressed bool
ReceivedAt time.Time
}
type Decision struct {
SubmissionID string
Queue string
Allowed []Channel
Policy string
DecidedAt time.Time
}
func Decide(in Input, now time.Time) (Decision, error) {
if in.SubmissionID == "" || in.ReceivedAt.IsZero() {
return Decision{}, errors.New("missing submission evidence")
}
queues := map[string]string{
"subscription": "audience-support",
"advertising": "commercial-support",
"correction": "editorial-review",
"other": "triage",
}
queue, ok := queues[in.Category]
if !ok {
queue = "triage"
}
allowed := make([]Channel, 0, 2)
if in.EmailRequested && !in.EmailSuppressed {
allowed = append(allowed, Email)
}
if in.SMSRequested && in.SMSConsentID != "" && !in.SMSSuppressed {
allowed = append(allowed, SMS)
}
return Decision{
SubmissionID: in.SubmissionID,
Queue: queue,
Allowed: allowed,
Policy: "contact-routing/v3",
DecidedAt: now.UTC(),
}, nil
}
The example does not claim that four categories fit every newsroom. It shows the boundary: unknown categories fall into human triage, while missing evidence fails closed. A policy release should change the explicit version, and the database write for the routing decision and queued work should be atomic. The later transport call remains outside that transaction; holding a database transaction open during a network request creates a longer failure window without making the database and remote system atomic.
Suppression needs scope and provenance. Store the channel, message class, reason, source event, effective time, and the actor or process that removed it. A single global boolean cannot express the difference between a hard email-address failure, an SMS opt-out, and a preference against promotional updates. It also makes review nearly impossible: an operator sees suppressed=true but cannot say what was blocked or why.
Domain authentication belongs in the evidence chain, although it cannot promise inbox delivery. SPF, defined by RFC 7208, authorizes hosts for an SMTP identity and limits evaluation to ten terms that cause DNS queries. Count the expanded lookup path during change review, not merely the visible top-level record. DKIM supplies a domain signature over selected headers and the body; record the signing domain and selector with the attempt so an operator can connect a message to the intended configuration. Keep the visible From identity stable and review alignment deliberately rather than treating a passing submission as proof of authentication.
For a SaaS contact workflow, the practical deliverability checklist therefore spans the custom sending domain, its DKIM selector, the scoped suppression list, and the durable routing decision. A Node.js application needs the same boundaries even though the policy example here is Go: runtime choice does not change what the evidence must prove. “Best” means the design fits the audit and on-call obligations, not that one transport wins a generic feature contest.
Capacity is part of the compliance design
Evidence that arrives hours after the message is operationally weak. Set an SLO on the part under platform control: for example, 99.9% of accepted contact forms should have a durable routing decision within 60 seconds. This is an engineering target, not a measured claim about any existing system. Define a separate service-level indicator for eligible acknowledgements reaching a recorded terminal transport state; do not label it “delivered to a human.”
Capacity planning starts with bursts, not daily averages. If a breaking-news event produces 300 submissions per second for 10 minutes, that is 180,000 decisions. At a sustained worker rate of 100 decisions per second, arrivals exceed service by 200 per second and backlog grows by 120,000 during the burst; after arrivals return below capacity, three workers at that same assumed rate would need at least 400 seconds to drain it if no new work arrived. These numbers are an explicit planning example, not a benchmark. Replace them with load-test results before setting production concurrency.
Queue age is the better paging signal than queue depth alone. Depth changes with normal traffic; age says a particular contact has waited too long. Page on burn rate against the routing SLO, and retain lower-urgency alerts for increasing suppression-check latency, policy evaluation errors, and events that cannot be correlated to a decision. Count attempts by channel, class, sender domain, and outcome, but keep recipient data out of those dimensions.
Here is the buy-versus-build decision I would put in front of a platform review. The rows are obligations, not product rankings.
| Concern | Build and operate the control | Use a managed control | Evidence required either way |
|---|---|---|---|
| Routing policy | Maximum reviewability; team owns releases and on-call | Faster configuration changes; remote state must be governed | Immutable policy version and decision inputs |
| Suppression | One model across send paths; more reconciliation work | Transport events may feed it directly; portability needs attention | Scoped reason, source, effective time, removal audit |
| Template release | Code review and fixtures share a deployment | Non-engineers may publish independently | Pinned version, required-field contract, promotion history |
| Event retention | Precise access and retention controls; storage is yours | Less operational work; export and deletion boundaries matter | Append-only observations plus derived current state |
| Domain keys | Direct custody; rotation runbook stays with the team | Reduced key handling; dependency boundary expands | Selector inventory, activation time, retirement evidence |
Managed versus self-hosted is not the primary compliance answer. Evidence ownership is. A service is acceptable when the organization can export the relevant history, reconcile callbacks idempotently, constrain access, and explain retention; a self-hosted component is unacceptable when nobody owns patches, rotation, capacity, or the pager.
This design has limits. It is a poor fit for a low-risk form that never sends an acknowledgement, collects no channel consent, and routes every submission to one queue; the policy ledger and replay machinery would add operational cost without useful evidence. A simpler database record and manual queue may be the better choice there. At the other extreme, a newsroom that cannot staff key rotation, event reconciliation, and round-the-clock queue response should use managed transport and identity controls while retaining its own exportable decision record. The trade-off is less infrastructure ownership in exchange for a larger dependency and portability boundary.
Verify the path, then practice rollback
Test the decision layer with category, consent, and suppression combinations, including missing evidence and unknown categories. Test the Postgres uniqueness constraint under concurrent inserts. Feed duplicate and reordered delivery observations to the event consumer and confirm that append-only evidence remains intact while the derived status follows explicit transition rules. Finally, resolve the production-shaped SPF and DKIM records from outside the deployment network and inspect a synthetic message sent to a controlled mailbox.
Do not use a real reader's message for that test. A fixture such as contact-test-2026-10-05 can exercise correlation without copying editorial content into logs.
Deployment should be staged by policy version. First write v3 decisions in shadow mode and compare them with v2 on redacted fixtures; do not enqueue shadow deliveries. Then enable a controlled internal cohort, inspect routing explanations and authentication results, and expand while watching SLO burn, oldest-item age, suppression outcomes, and unmatched events. The rollback switch selects the prior policy for new decisions. Existing decisions keep their original version and should never be rewritten to make the history look consistent.
Roll back independently. A routing-policy error calls for selecting the prior policy; a template defect calls for pinning the prior template; a DNS authentication change calls for restoring the reviewed sender configuration; transport degradation calls for pausing attempts while durable decisions continue to accumulate within the capacity budget. One global rollback couples unrelated failure domains and destroys diagnostic value.
The release gate is blunt: an operator must be able to select one synthetic submission and explain its queue, channel eligibility, suppression result, policy version, template version, sender identity, and final observed outcome from stored evidence. If any answer depends on a mutable dashboard value, the path is not ready.
Top comments (0)