DEV Community

ZorvynGale1729
ZorvynGale1729

Posted on

DNS Zone Migration Explained: 3 Passes to Enumerate, Diff and Verify Before Cutover

Treat a registrar migration as three passes over the zone — enumerate, diff, verify — and change the nameservers only after the third pass comes back empty. The constraint that decides everything else is propagation: once the delegation moves, you can't un-cache it, so the cost of a missing record is set by TTLs rather than by how fast you can fix the zone. Use the old zone as the source of truth until the new one answers identically.

That ordering is the entire runbook. Apply-then-check is the version that pages someone at 03:00.

For a media company the blast radius lands on mail long before anyone notices the website. A newsroom usually sends from several systems at once — a newsletter platform, a ticketing system, transactional alerts out of the CMS — and each of them publishes its own DKIM selector under _domainkey. Miss one selector while copying the zone and the site stays up while a day of newsletters quietly fails DKIM, then fails DMARC alignment, then lands in spam folders. The first signal is a reader complaint or an aggregate report that arrives the next day, which is far too late to call it a fast rollback.

Enumerate what the zone serves, not what the console shows

The export you get from a provider console is a rendering of its own database, and it's not always the same thing as what its nameservers answer on the wire. AXFR would be the clean way to enumerate a zone, but most hosted DNS refuses zone transfers to arbitrary clients, so in practice you build the record list from the provider's export and then confirm every entry by querying the authoritative servers directly with recursion disabled.

Store that enumerated set in version control before you change anything. It's your rollback material, and it's the only artifact that lets a postmortem say what the zone looked like at 14:00 rather than what somebody remembers it looking like. Enumerate more than the record types you think are in play, too: a zone that has been edited by four teams over five years accumulates TXT records for domain verification, a CNAME for a link-tracking host that a newsletter vendor still signs against, an old _acme-challenge left behind by a certificate renewal, MX records for a subdomain that only ever received bounce mail, and at least one A record whose owner left the company. None of those are interesting individually. Collectively they are the reason a migration that looked complete on Friday produces a support queue on Monday, because the records that break loudly get copied and the records that break quietly get forgotten.

Record class Blast radius if missed What the copy usually drops
MX and SPF TXT All inbound mail, all outbound authentication An include: for a sender nobody owns anymore
DKIM selectors One sending system each, silently Selectors added ad hoc by a marketing team and never documented
Apex CNAME-like records Website and tracking hostnames Provider-specific apex flattening with no standard equivalent

SPF is worth a dedicated look during the copy because of a limit people forget: a receiver must fail evaluation with a permerror after 10 DNS-querying mechanisms, and RFC 7208 also caps void lookups at two. A zone that accumulated includes over five years is often one merger away from that ceiling. Migration is the moment to delete the dead ones — carefully, because deleting an include that some forgotten system still needs is exactly the kind of change whose failure mode is invisible for a week.

How do I diff the old and new zone before changing nameservers?

Query both authoritative sets directly and compare canonical forms, never console screenshot against console screenshot. Recursion off, one server at a time, same question list for both sides. The diff you care about is authoritative-versus-authoritative, because that's the comparison that survives caching.

Most false diffs come from formatting rather than content. A 2048-bit DKIM key doesn't fit in a single DNS character-string — the 255-octet limit in RFC 1035 — so it gets split across several strings inside one TXT record, and two providers will happily split the same key at different offsets. Join the strings before you compare. Trailing dots, owner-name case, and MX preference formatting cause the same class of noise, while SPF mechanism order is semantically meaningful and must never be sorted away.

The question usually arrives asking for a Node.js example. dns.promises.Resolver with setServers() does this job well: it points at one authoritative server at a time, and resolveTxt() hands back string[][], one array of character-strings per record, which you join yourself. I write ops tooling in Go, so that's what these examples are.

// zonediff reads the same names from two authoritative servers and compares
// canonical forms. Read-only, safe to run in a loop while you converge.
package main

import (
    "fmt"
    "sort"
    "strings"

    "github.com/miekg/dns"
)

type key struct{ Name, Type string }

// canonical collapses provider formatting differences: TXT character-strings
// are joined, names are lowercased, and each RRset is sorted so that ordering
// never shows up as a difference. SPF mechanisms inside a string stay put.
func canonical(rrs []dns.RR) []string {
    out := make([]string, 0, len(rrs))
    for _, rr := range rrs {
        switch v := rr.(type) {
        case *dns.TXT:
            out = append(out, strings.Join(v.Txt, ""))
        case *dns.MX:
            out = append(out, fmt.Sprintf("%d %s", v.Preference, strings.ToLower(v.Mx)))
        default:
            rdata := strings.TrimPrefix(rr.String(), rr.Header().String())
            out = append(out, strings.ToLower(strings.TrimSpace(rdata)))
        }
    }
    sort.Strings(out)
    return out
}

func query(server, name string, qtype uint16) ([]dns.RR, error) {
    m := new(dns.Msg)
    m.SetQuestion(dns.Fqdn(name), qtype)
    m.RecursionDesired = false // ask the authority, not somebody's cache
    r, _, err := new(dns.Client).Exchange(m, server+":53")
    if err != nil {
        return nil, err
    }
    if r.Rcode != dns.RcodeSuccess && r.Rcode != dns.RcodeNameError {
        return nil, fmt.Errorf("%s %s: rcode %s", server, name, dns.RcodeToString[r.Rcode])
    }
    return r.Answer, nil
}
Enter fullscreen mode Exit fullscreen mode

The diff itself is boring on purpose, and boring is what you want in front of an irreversible step:

// diff reports what the destination is still missing. An empty result is the
// only state that may proceed to a nameserver change.
func diff(src, dst map[key][]string) []string {
    var changes []string
    for k, want := range src {
        got, ok := dst[k]
        if !ok {
            changes = append(changes, fmt.Sprintf("MISSING %s %s -> %v", k.Name, k.Type, want))
            continue
        }
        if strings.Join(want, "|") != strings.Join(got, "|") {
            changes = append(changes, fmt.Sprintf("DIFF %s %s: want %v got %v", k.Name, k.Type, want, got))
        }
    }
    for k := range dst {
        if _, ok := src[k]; !ok {
            changes = append(changes, fmt.Sprintf("EXTRA %s %s", k.Name, k.Type))
        }
    }
    sort.Strings(changes)
    return changes
}
Enter fullscreen mode Exit fullscreen mode

Join first, compare second.

EXTRA entries deserve a human decision rather than an automatic delete. A staging record on the new provider is harmless; a leftover verification TXT from a service you already migrated is harmless too, until a receiver reads it as an active authorization.

Apply by desired state, so the third run is a no-op

Write the apply step as an upsert keyed on name plus type, then run it until the diff comes back empty. Two runs should produce the same zone as one run. That property is what makes a partially failed apply survivable — you rerun the whole intended set instead of reconstructing which of eleven records made it through before the API call timed out.

Desired-state DNS tooling exists for exactly this shape of work; dnscontrol and octodns both render a zone from a checked-in source file and reconcile the provider toward it. Neither removes the delegation problem, and neither can tell you whether mail still authenticates.

Now the part that governs your schedule. Lower the TTLs on the records you intend to move at least one full TTL before the change window, because a TTL reduction only takes effect after the old, longer TTL has expired everywhere. The delegation TTL is a different story: the parent zone publishes the NS set with its own TTL — 172800 seconds in .com, two full days — and you don't control it. Resolvers that have already cached the old delegation will keep asking the old nameservers after your flip, and some will keep doing it for the full two days.

The practical consequence is that both providers must serve identical answers for the whole overlap window.

Cutover speed is mostly a fiction. What you control is how long the two answers agree.

Verify mail while the old zone is still live, and plan the rollback that isn't a rollback

Verification means sending real messages through each sending system and reading the Authentication-Results header for spf=pass, dkim=pass, and dmarc=pass with the right identifier alignment. Do that against the new nameservers before the delegation moves, by pointing a resolver at them explicitly. Doing it afterwards only tells you what already broke.

DMARC aggregate reports are evidence, not monitoring. RFC 7489 defines a default reporting interval of 86400 seconds, so a report about your cutover typically arrives the day after the cutover. If your rua= address lives on a different domain, the receiving domain must publish the _report._dmarc authorization record, and that record is easy to forget in a migration because it lives in the other zone.

The catch is that a nameserver change has no clean rollback. Flipping the delegation back is another two-day propagation, so recovery comes from the overlap window and not from an undo button. Keep the old zone paid for and serving the identical record set for at least a week after the cutover. The irreversible step is deleting the old zone, not changing the delegation — treat those as two separate change tickets with two separate approvals.

If your zone is DNSSEC-signed, this whole plan needs a different one. Moving a signed zone between operators means either unsigning first, waiting out the DS TTL at the parent, then migrating and re-signing, or running a multi-signer model as described in RFC 8901. Stick with the unsign-wait-migrate sequence unless your team already operates DNSSEC day to day; the multi-signer path is correct and considerably harder to run correctly under time pressure.

Move the registrar and the DNS host as separate changes

Registrar transfer and DNS hosting are two different systems that happen to be sold together. ICANN's Transfer Policy imposes a 60-day lock after a completed inter-registrar transfer, so a rushed combined migration can leave the domain stuck at a registrar you wanted to leave, with mail authentication in flux at the same time.

Move DNS first, prove mail, then move the registration. One variable at a time. I'm not certain that ordering is right for every shop — if your registrar is the party actually losing your business, doing the registration first may be the safer call — but the rule that holds either way is that no two irreversible DNS changes should share a change window.

References

Top comments (0)