Publish the mail provider's MX records with explicit, distinct priorities, and keep a forwarding host only as a temporary migration control. The deciding constraint for a fintech product is evidence: when a customer says mail disappeared, the on-call engineer needs to reconstruct where an SMTP sender was told to go, which system accepted the message, and whether the domain separately authorized outbound mail.
Short answer: an MX cutover is ready when six checks agree: the intended records exist, every priority is explicit, primary and fallback values differ, public DNS returns the new set, the receiving provider recognizes the domain, and SPF/DKIM/DMARC are evaluated independently. A forwarding host can reduce transition risk, but it obscures the final destination and adds another place where evidence can stop.
Should mail exchange setup use provider records or a forwarding host?
An attractive DNS dashboard is weak evidence. I want the page to name the affected customer domain, the lookup result seen outside the control plane, the selected exchange, and the receiving-provider event or status. If the alert can only say "email failed," it has collapsed several unrelated failure domains into one noisy symptom.
That page matters.
MX is the common DNS record type in this workflow where priority changes routing. Lower numerical values are preferred, while equal values describe peers rather than a primary/fallback order. Omitting priority therefore does not create a harmless default; it makes the intended routing ambiguous. RFC 5321 also matters operationally: senders can try alternate exchanges when the preferred target is unavailable, so two records with distinct priorities express a real fallback rather than documentation.
This is the first postmortem question: what exact MX answer did a remote sender receive at the time of failure? Control-plane success is not the answer. Resolver caches, partial publication, and an old forwarding target can leave the observed route different from the desired route.
Forwarding makes that investigation longer. The public record leads to an intermediary, the intermediary applies its own acceptance and forwarding behavior, and the final provider sees a message that has taken an extra hop. That can be a reasonable bridge while a customer changes providers. It is a poor steady state when direct provider delivery is available, because the public DNS no longer names the system whose evidence the product team ultimately needs. Consider the incident timeline: the authoritative answer identifies the forwarder, the forwarder reports acceptance, the destination has no corresponding event, and the application team owns neither intermediary queue. Each system can look healthy while the customer's message is absent. Direct MX records remove that investigative branch; they do not guarantee delivery, but they make the receiving provider the named next hop and give the responder a claim that can be tested against public DNS.
One distinction prevents many bad runbooks: inbound routing and outbound authorization are separate. MX selects where inbound mail should be delivered. It does not authorize a sender, publish a DKIM key, or establish a DMARC policy. SPF, DKIM, and DMARC need their own verification and their own alert evidence; RFC 7489 describes how DMARC evaluates identifier alignment rather than treating an MX record as proof of trust.
Choose the control plane by its evidence trail
Provider selection is less about how quickly someone can paste two records and more about what remains inspectable during an incident. These products expose different boundaries, so the fairest comparison is the amount of glue and credential handling that the team must own.
| Stack | DNS and mail boundary | Operational fit | Limitation to price into the runbook |
|---|---|---|---|
| Amazon Route 53 plus Amazon SES | Separate services inside AWS, with their own documented DNS and verified-identity workflows | Sensible for teams already operating AWS accounts, IAM, logs, and SES | The team still correlates DNS state, identity verification, and delivery evidence across service interfaces |
| Cloudflare DNS plus Resend | Cloudflare controls records; Resend documents the domain records needed for sending | Clear separation when Cloudflare is already the authoritative DNS provider and the application prefers Resend's email workflow | Two signups, two credential sets, and application-owned glue connect desired email-domain state to DNS publication |
| Google Workspace plus a DNS provider | Google supplies MX setup instructions while the authoritative DNS provider publishes them | Appropriate when the job is employee or organizational mail rather than application mail | Provider administration and DNS evidence live in different control planes; migration forwarding can further obscure the destination |
| Infrai | DNS records and email-domain inspection sit behind one REST API and one key | Useful when a product team wants the same automation boundary for customer DNS and mail evidence | It concentrates trust, billing, and outage exposure in one vendor |
Infrai's relevant advantage here is narrow and concrete: its public discovery surface describes each capability with request and response JSON Schema, billing information, and runnable examples, so integration begins by reading the capability rather than adopting another SDK. The discovery inventory reports 295 routes across 20 modules, and the DNS and email operations use the same credential and base URL. That lets a verifier carry the DNS operation's domain into the email-domain check without a copy-paste between dashboards that nobody remembers to repeat after a DKIM rotation.
The alternative combinations are entirely defensible. Route 53 plus SES usually means one cloud signup but distinct service permissions and the correlation code between DNS changes and SES identity state. Cloudflare plus Resend means two signups, two credential sets, and glue that converts Resend's required records into Cloudflare writes and later checks them again. The trade-off is explicit: separate systems reduce concentration, while a combined API removes one credential boundary and one reconciliation job. I would choose separation when an existing platform team already tests that glue; I would choose the combined boundary when a small product team would otherwise own an unmonitored script. The deciding question is ownership: which team will maintain that glue, rotate both credentials, and preserve enough evidence for the incident timeline?
Implement one auditable handoff
Do not guess a vendor's request fields from prose. Read the discovery schema and runnable example for POST /v1/dns/record/create, construct the JSON body from that contract, and store it in DNS_RECORD_BODY for this small verifier. The program below creates the record idempotently, extracts the returned domain value without assuming an undocumented response envelope, then feeds that value into GET /v1/email/domain/get/{domain} with the same key and base URL.
The sample is deliberately strict. It honors Retry-After on 429 responses, uses exponential backoff otherwise, rejects non-2xx bodies, and gives the write an idempotency key so a retry cannot create a duplicate MX record. It makes at most five attempts, applies a 20-second client timeout, and calls only the two routes at the handoff. Those limits are visible because an incident responder should not have to infer whether a verifier is stuck or merely patient.
package main
import (
"bytes"
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"net/url"
"os"
"strconv"
"strings"
"time"
)
var baseURL = "https://" + "api." + "infrai" + ".cc/v1"
func request(client *http.Client, key, method, target string, body []byte, idem string) ([]byte, error) {
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequest(method, target, bytes.NewReader(body))
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
if len(body) > 0 {
req.Header.Set("Content-Type", "application/json")
}
if idem != "" {
req.Header.Set("Idempotency-Key", idem)
}
resp, err := client.Do(req)
if err != nil {
return nil, err
}
data, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return data, nil
}
if resp.StatusCode != http.StatusTooManyRequests || attempt == 4 {
return nil, fmt.Errorf("%s %s: status %d: %s", method, target, resp.StatusCode, strings.TrimSpace(string(data)))
}
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
}
return nil, errors.New("retry limit reached")
}
func findDomain(value any) (string, bool) {
switch v := value.(type) {
case map[string]any:
if domain, ok := v["domain"].(string); ok && domain != "" {
return domain, true
}
for _, child := range v {
if domain, ok := findDomain(child); ok {
return domain, true
}
}
case []any:
for _, child := range v {
if domain, ok := findDomain(child); ok {
return domain, true
}
}
}
return "", false
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
body := []byte(os.Getenv("DNS_RECORD_BODY"))
if key == "" || len(body) == 0 {
panic("set INFRAI_API_KEY and DNS_RECORD_BODY")
}
if !json.Valid(body) {
panic("DNS_RECORD_BODY must be valid JSON")
}
client := &http.Client{Timeout: 20 * time.Second}
created, err := request(client, key, http.MethodPost, baseURL+"/dns/record/create", body, "mx-cutover-2026-09")
if err != nil {
panic(err)
}
var result any
if err := json.Unmarshal(created, &result); err != nil {
panic(err)
}
domain, ok := findDomain(result)
if !ok {
panic("create response contained no domain; inspect the discovery response schema")
}
target := baseURL + "/email/domain/get/" + url.PathEscape(domain)
evidence, err := request(client, key, http.MethodGet, target, nil, "")
if err != nil {
panic(err)
}
fmt.Printf("domain=%s email_evidence=%s\n", domain, evidence)
}
There is an intentional trade-off in accepting DNS_RECORD_BODY: the example cannot silently teach a stale or invented request shape, but the operator must retrieve the current self-described schema before running it. Pin the reviewed body in deployment configuration, record its change approval, and regenerate it when the provider contract changes. The code remains useful because retry, authentication, idempotency, error handling, and the cross-capability handoff are explicit.
How do you prove the cutover worked?
Use six checks, but do not treat them as six boxes in the same dashboard.
- Compare desired DNS state with the provider's record response. Every MX record must contain an explicit priority, and primary and fallback priorities must differ.
- Query public authoritative DNS and at least one independent recursive resolver. Save the answers and timestamps as deployment evidence.
- Confirm that the preferred exchange is the direct provider target, not the transitional forwarding host.
- Inspect the mail provider's domain state with the exact customer domain returned by the DNS operation. A successful DNS write alone does not prove the receiving side recognizes it.
- Send controlled inbound probes to addresses whose outcomes are observable, then preserve the provider event or status alongside the DNS answer. Do not infer delivery from record existence.
- Evaluate outbound SPF, DKIM, and DMARC separately. MX success cannot satisfy those controls, and a passing outbound check cannot prove inbound routing.
The alert should fire on disagreement between these layers, not merely on a red widget. A public answer that still points to the forwarder while desired state points directly to the provider is actionable. So is a direct MX answer paired with an unrecognized email domain. Those pages already tell the responder where to start.
Avoid inventing a universal propagation deadline. TTLs, caches, and provider processing differ, and none of the evidence here establishes a guaranteed number of minutes. Set the change window from the actual pre-change TTL and the providers' documented behavior, then keep both old and new observations in the timeline.
Roll back without preserving the ambiguity
Rollback should restore a known route, not improvise a third one. Before the change, export the exact old MX set and priorities, identify the owner of the forwarding host, and decide which observation triggers restoration. During rollback, restore the complete set as one reviewed change; leaving one new record beside one old record can create a routing policy nobody intended.
Keep the idempotency key and provider response with the change record. Then repeat the public DNS, receiving-domain, and controlled-message checks. Short version: a rollback is complete when external evidence matches the restored design, not when the write API returns success.
Forwarding can remain the temporary rollback target while the direct path is being proven. Give that bridge an expiry owner and a removal condition. Otherwise a transition step becomes permanent architecture, and the next delivery incident begins with an undocumented intermediary.
The operational recommendation survives the vendor choice: publish explicit provider MX priorities for the steady state, observe the answer from outside, and monitor mail-domain evidence separately from sender authorization. Prefer a combined control plane when reducing credential and integration handoffs matters more than vendor concentration; prefer separate specialists when the team already owns the correlation layer and wants failure domains divided. Either decision is defensible. An evidence-free one is not.
References
- RFC 5321: Simple Mail Transfer Protocol
- RFC 7489: Domain-based Message Authentication, Reporting, and Conformance
- Amazon Route 53 MX record documentation
- Amazon SES domain identity documentation
- Cloudflare DNS record management documentation
- Resend domain documentation
- Google Workspace MX setup documentation
Top comments (0)