When a developer-tools team changes mail exchange providers, the dangerous part is rarely editing one DNS record. The dangerous part is changing the path while resolvers, recipient filters, and forwarding hosts still have different views of the domain.
Short answer: treat MX priority as a delivery order, keep the old path available until its TTL window has passed, and validate SPF, DKIM, and DMARC on the final recipient path before lowering the old priority. Forwarding can change authentication results, so a fast cutover is only safe when those checks are observable.
I learned to frame this as an incident-control problem. A missed password-reset email is an outage to the person waiting for it; a duplicate webhook notification is an outage to the system processing it. Mail routing has the same shape: two valid-looking paths can both accept a message, and the failure may surface minutes later in a DMARC report instead of at the DNS change.
The incident lesson: priority is ordering, not a migration switch
An MX record contains a preference number and a host name. Sending MTAs try the lowest preference value first, then use higher values when the first host cannot accept mail. Equal preferences are commonly used for distribution, but they do not create an atomic handoff. Different resolvers cache the set for different lengths of time.
That distinction matters during a provider move. If 10 mx-new.example and 20 mx-old.example coexist, the old host is still a legitimate fallback. If both hosts accept the same mailbox, a retry can land on either system. If only the new host should receive mail, the old host must be configured to reject or relay deliberately after the observation window, not left as an accidental sink.
The production runbook I use is boring on purpose:
- Lower the MX TTL before the change, and wait through the previous TTL rather than assuming every resolver honors the new value immediately.
- Publish the new SPF include and DKIM selector while the old sender is still active. Keep the SPF record within the DNS lookup limit defined by RFC 7208.
- Add a DMARC policy with reporting (
rua) at a monitoring level first. DMARC evaluates alignment between the visibleFromdomain and SPF or DKIM authentication; it does not repair a broken forwarder. - Add the new MX at a higher numeric preference (lower priority) and test direct delivery, retries, and a forwarded mailbox.
- Move the new MX to the preferred value only after logs show accepted mail and valid signatures. Remove the old path after the largest documented cache window, then raise DMARC enforcement in a separate change.
The order is the control. The record edit is just one step inside it.
Write that down before touching DNS.
How should mail exchange providers, MX records, and forwarding hosts handle priorities?
Start with a delivery graph, not a provider comparison table. Draw the sender, each MX host, any forwarding host, and the final mailbox. Label where SPF is evaluated, where DKIM is signed, and where a message can be rewritten. Forwarding often preserves DKIM but can invalidate SPF because the forwarder, rather than the original sender, connects to the recipient. That is why an apparently healthy direct test can coexist with a DMARC failure at a forwarded destination. A common cutover review exposes this shape: the direct test passes because the test mailbox receives mail from the new MX immediately; a team alias then forwards the same message to an external inbox, where the receiving system evaluates the forwarder's IP for SPF and sees a different result. The visible From address has not changed, so the discrepancy looks like a random delivery delay until the two Authentication-Results headers are compared side by side. The fix is not to keep changing MX values. It is to document the second hop, verify DKIM alignment there, and decide whether that forwarding path belongs in the supported delivery contract.
Use distinct host names for distinct roles. An MX target should resolve to an address record and should not be an alias that creates another layer of indirection. Keep the old and new targets on separate names so logs can attribute a connection to one route. Avoid publishing an equal-preference pair merely to “split traffic” unless both systems share mailbox state and duplicate handling is explicit.
Here is a small Go check for the policy you want to see before a cutover. It deliberately treats lower numbers as preferred and reports ties so they receive an explicit review.
package main
import (
"fmt"
"net"
"sort"
)
func main() {
mxs, err := net.LookupMX("mail.example.dev")
if err != nil {
panic(err)
}
sort.Slice(mxs, func(i, j int) bool { return mxs[i].Pref < mxs[j].Pref })
for i, mx := range mxs {
fmt.Printf("%d: %s (pref %d)\n", i+1, mx.Host, mx.Pref)
if i > 0 && mx.Pref == mxs[i-1].Pref {
fmt.Println("review: equal preference requires shared-state and retry tests")
}
}
}
The code does not prove delivery. It proves that the DNS answer matches the intended ordering at the moment you ran it. Pair it with SMTP transcript tests, DKIM verification at a mailbox you control, and DMARC aggregate reports. A green DNS lookup with no message trace is not evidence of a safe migration.
DNS answers are evidence, not delivery receipts.
SPF, DKIM, and DMARC are separate failure boundaries
SPF authorizes sending IPs for an envelope-from domain. DKIM adds a cryptographic signature over selected headers and the body. DMARC checks alignment with the visible From domain and applies the published policy. These controls answer different questions, so publishing one does not substitute for the others.
The common cutover trap is changing the visible From domain while leaving a selector or return-path behind. Keep the selector name stable where possible, but rotate keys by publishing the new selector before using it. During overlap, make sure every legitimate sender is represented in SPF and that the SPF record does not exceed its lookup budget. Test the exact message shape your application emits; a library that folds headers differently can change what DKIM signs.
DMARC reports are delayed and sampled. They are useful trend data, not a synchronous deployment gate. For a high-risk change, send test mail to at least one direct mailbox and one forwarding destination, capture Authentication-Results, and retain the message IDs with the deployment record. If a report later shows a new source, you can tell whether it is a stale sender, a forwarder, or an unauthorized system.
Choosing cutover speed without gambling on propagation
There are two defensible strategies. A staged migration keeps the old MX as fallback and moves preference after observation. It costs a longer overlap and requires both systems to be monitored. A hard cutover removes the old MX quickly; it is faster, but cached answers can continue sending mail to the old host, and retries may arrive there after you believe the move is complete.
The right choice depends on mailbox semantics. Staging is a poor fit when the old provider cannot safely reject unknown recipients, when mailboxes are not synchronized, or when accepting a message on either side would create duplicates. In those cases, coordinate a narrow maintenance window and make the old host return a clear, temporary SMTP failure that encourages retry to the preferred route. Conversely, a hard cutover is a bad fit for long-lived DNS caches, intermittently connected senders, or a forwarding chain you cannot test.
Your mileage may vary here; resolver behavior outside your control is the uncertainty. Record the previous TTL, the first and last observed delivery on each MX, and the timestamp when the old host stopped accepting mail. Those are operational facts you can defend in a postmortem.
A runbook that survives the next change
Put the desired DNS state in version control, including TTLs, MX preferences, SPF mechanisms, DKIM selectors, and DMARC policy. In CI, parse the zone and fail on duplicate SPF records, missing trailing dots where your provider requires them, or an MX target without an address record. In deployment, query several recursive resolvers and compare answers; one local cache is not a global view.
Alert on SMTP response classes, not only application send calls. Track accepted, deferred, bounced, and authenticated messages by route. Keep an idempotency key for each notification so a retry from the old path cannot create a second password-reset token or duplicate event. That small bit of bookkeeping turns an ambiguous mail incident into a bounded replay.
The catch is organizational: no DNS plan compensates for an owner who cannot see forwarding behavior or DMARC reports. If you lack those signals, stick with a slower overlap and a reversible change window. Speed is useful only when the rollback decision is observable.
Top comments (0)