Use a read, compare, write, read-back loop around every DNS record change, and skip the write entirely when the stored value already matches. That loop is about eighty lines in any language, and it is the whole difference between a migration job that reports success and one that tells the truth — because an accepted write and a served answer are two different events, separated by a propagation window you do not control.
The rest of this is what each of those four steps has to check, and where the read-back threshold turns into a pager problem.
The 02:40 page and the signal that should have fired at 01:52
The page reads something like tenant-domain-health: CNAME mismatch on 9 customer domains. On-call opens the runbook, and the first three facts all point the wrong way. The zone-move job finished at 01:47. Every write returned a 2xx. The new provider's console lists the records, spelled correctly, with the right targets.
Then someone queries the authoritative nameservers directly and gets the old value back.
Nothing upstream of the resolver was red, which is exactly why the page arrived at 02:40 instead of 01:52. The job had been migrated off a registrar-specific API the week before, and the old API answered a write with a queued-change acknowledgement while the new one answers with the stored record object. Neither response is a claim about what the zone serves. The job treated both as one, so its success log was a log of requests it had sent, not of answers the internet could get. For a B2B SaaS with customer domains delegated to you, that gap is where tenant onboarding quietly stops working.
The DMARC case makes the timing worse. A _dmarc TXT record that got stored with a mangled value doesn't bounce anything and doesn't return an error; it just stops aligning, and the evidence arrives in an aggregate report whose default reporting interval under RFC 7489 is 86400 seconds. A feedback loop one day long is not a feedback loop you can run a cutover on.
The signal that should have fired is cheap: a counter that increments when, some minutes after a write, the authoritative servers still serve the old value. It names the zone and the record. It fires while the old zone is still live and rollback is still a five-minute change rather than a TTL problem.
What does a safe DNS record writer compare before it writes anything?
Five fields, normalized, passed explicitly at every call site: zone, owner name, type, content, TTL. Default arguments are how a writer ends up creating an A record where a CNAME was meant, or writing into the apex because the name was empty.
Normalization is most of the work, and it is where a naive string comparison produces phantom diffs that turn into pointless writes. Owner names match case-insensitively, so lowercase both sides before comparing — RFC 4343 is the clarification worth reading if someone argues about it. Trailing dots differ between providers and between exports of the same zone. A TXT record's content is one or more character-strings of at most 255 octets each, per RFC 1035, so a 2048-bit DKIM key or a long DMARC record gets split, and two providers will split the same value at different offsets; join the strings before you compare, and never sort the tokens inside an SPF value, because mechanism order there is semantic. MX and SRV carry numeric fields that belong in the comparison rather than glued into the content string. A TTL-only difference is a real difference, but it is one you change deliberately as part of the cutover plan, not something a 02:00 reconcile job should quietly fix.
When all five match, do nothing, and record that you did nothing.
Skipping the no-op is not just tidiness. A write that changes nothing can still register as a change on the provider side, bumping the zone serial and filling the change history with identical entries, which is the audit trail you will want to read during the next incident. It also keeps your write rate proportional to actual drift, which matters when the reconcile loop runs every few minutes across a few thousand tenant domains.
The read-then-write pair is not atomic, and it's worth being honest about that. RFC 2136 defines UPDATE prerequisites — an update can be made conditional on an RRset existing, not existing, or matching an exact value — which is a real compare-and-swap at the protocol level. HTTP record APIs generally don't expose anything equivalent, so your compare is optimistic. The practical fix is boring: one writer per zone, enforced with a lease, and humans locked out of the console during a migration. I've been paged for duplicate deliveries caused by two workers thinking they owned the same job; a zone with two writers is the same failure with a longer blast radius.
The loop, with the no-op skip and a bounded read-back
Most examples for this land in Node.js, and its standard library is a good fit for the verification half: dns.promises.Resolver with setServers() aims queries at one authoritative server at a time, and resolveTxt() returns an array of character-string arrays, which is the split problem handed to you directly rather than hidden. I write ops tooling in Go, so that's the version below.
// Package dnswrite applies one record at a time: read, compare, skip or
// write, then read back from the zone's own nameservers before reporting
// success. Upsert is idempotent by design (RFC 9110, section 9.2.2), so a
// retried Apply is safe.
package dnswrite
import (
"context"
"errors"
"fmt"
"strings"
"time"
)
type Record struct {
Zone string // "example.com"
Name string // FQDN without the trailing dot
Type string // "A", "CNAME", "TXT", "MX"
Content string
TTL int
}
// Provider is the seam that makes a registrar-specific API replaceable.
// Two methods are all the writer needs from a vendor.
type Provider interface {
List(ctx context.Context, zone, name, rtype string) ([]Record, error)
Upsert(ctx context.Context, r Record) error
}
// Authoritative queries the zone's own nameservers with recursion disabled,
// so a match means the zone serves the value, not that a cache does.
type Authoritative interface {
Lookup(ctx context.Context, zone, name, rtype string) ([]string, error)
}
type Result struct {
Skipped bool
Applied bool
Verified bool
}
var ErrNotServed = errors.New("write accepted but the authoritative servers still serve the old value")
func Apply(ctx context.Context, p Provider, auth Authoritative, want Record, budget time.Duration) (Result, error) {
want = normalize(want)
current, err := p.List(ctx, want.Zone, want.Name, want.Type)
if err != nil {
return Result{}, fmt.Errorf("read %s %s in %s: %w", want.Type, want.Name, want.Zone, err)
}
for _, have := range current {
if normalize(have) == want {
return Result{Skipped: true, Verified: true}, nil
}
}
if err := p.Upsert(ctx, want); err != nil {
return Result{}, fmt.Errorf("upsert %s %s in %s: %w", want.Type, want.Name, want.Zone, err)
}
deadline := time.Now().Add(budget)
for {
served, err := auth.Lookup(ctx, want.Zone, want.Name, want.Type)
if err == nil && matches(served, want.Content) {
return Result{Applied: true, Verified: true}, nil
}
if time.Now().After(deadline) {
return Result{Applied: true}, fmt.Errorf("%w: %s %s in %s after %s",
ErrNotServed, want.Type, want.Name, want.Zone, budget)
}
select {
case <-ctx.Done():
return Result{Applied: true}, ctx.Err()
case <-time.After(5 * time.Second):
}
}
}
// normalize collapses the differences that are formatting rather than
// content: case, trailing dots, and TXT values that a provider chose to
// split into several character-strings.
func normalize(r Record) Record {
r.Name = strings.ToLower(strings.TrimSuffix(r.Name, "."))
r.Type = strings.ToUpper(r.Type)
r.Content = strings.TrimSpace(r.Content)
if r.Type == "TXT" {
r.Content = strings.ReplaceAll(r.Content, `" "`, "")
r.Content = strings.ReplaceAll(r.Content, `"`, "")
}
if r.Type == "CNAME" || r.Type == "MX" {
r.Content = strings.ToLower(strings.TrimSuffix(r.Content, "."))
}
return r
}
func matches(served []string, want string) bool {
for _, v := range served {
if normalize(Record{Type: "TXT", Content: v}).Content == normalize(Record{Type: "TXT", Content: want}).Content {
return true
}
}
return false
}
Two details in there are load-bearing. Apply returns Skipped and Applied as separate facts, so the caller can emit them as separate counters instead of one ambiguous "success". And the read-back failure is a typed error carrying zone, name and type, which is what makes it findable later — an error string of dns write failed in a log with four thousand tenants in it is a way of not being told anything.
The caller is small, and it is where the instrumentation lives:
res, err := dnswrite.Apply(ctx, prov, auth, want, 10*time.Minute)
switch {
case errors.Is(err, dnswrite.ErrNotServed):
metrics.Inc("dns_readback_mismatch", want.Zone, want.Type)
obs.CaptureError(ctx, err, map[string]string{"zone": want.Zone, "record": want.Name, "type": want.Type})
case err != nil:
metrics.Inc("dns_write_error", want.Zone, want.Type)
obs.CaptureError(ctx, err, map[string]string{"zone": want.Zone, "record": want.Name, "type": want.Type})
case res.Skipped:
metrics.Inc("dns_write_skipped", want.Zone, want.Type)
default:
metrics.Inc("dns_write_applied", want.Zone, want.Type)
}
For the verification side, the cheapest thing that works during a cutover is a direct query per nameserver, recursion off:
for ns in ns1.example-dns.net ns2.example-dns.net; do
dig +norecurse +short @"$ns" app.tenant.example CNAME
done
Propagation delay against cutover speed, and the threshold that pages for nothing
Where you read back decides what a green check means, and the three options are not interchangeable.
| Read-back target | What a match proves | What it costs |
|---|---|---|
| Provider API | The record is stored in their database | Nothing, and it proves nothing about resolution |
| Authoritative nameservers, recursion off | The zone actually serves the value | One query per nameserver, and you must enumerate the NS set first |
| Public recursive resolver | One cache somewhere has caught up | Up to a full TTL of waiting, measuring someone else's cache |
Some APIs make the distinction explicit. A Route 53 change comes back with a status of PENDING and only becomes INSYNC once that service reports the change distributed to its nameservers, which is a propagation statement rather than a storage receipt. Others return the stored record immediately, which is honest about what it is: the database now holds this row.
There's a trap on the read side that bites during creation, not update. If your pre-flight check queried the name through a recursive resolver before the record existed, the NXDOMAIN is now cached, and RFC 2308 sets that negative TTL from the SOA — values of one to three hours are common. Your own read poisons your read-back. Query the authoritative servers directly and the problem disappears.
Then the threshold. The budget has to be at least the provider's published propagation window, and I give it double that before anything pages a human, because the cost of getting this wrong is not a missed alert — it's an alert that fires on normal propagation, twice a week, until nobody reads it. Worse is the automated version: wire a rollback to the first mismatch and a slow propagation turns into a write loop, with the job rewriting a record that was already correct, churning the zone serial and burning provider rate limit while the on-call watches a dashboard flap.
Alert on the counter, not on the record. One mismatch inside the budget is physics; a mismatch rate above zero after the budget expires is a bug.
The TTL plan is the other half of the same axis. Lower the TTL on the records you are moving a full old-TTL window ahead of the cutover — 300 seconds is a common working value — hold it there until read-back is clean on every nameserver, then put it back afterwards. Short TTLs buy cutover speed and cost query volume, which is billable on most hosted DNS. And they are a request rather than a guarantee: recursive resolvers can clamp both ends, unbound's cache-min-ttl being the obvious example, so plan for some share of clients holding the old answer longer than your arithmetic says. How large that share is depends on your users' resolvers, and I'm not sure anyone can tell you in advance.
What the loop doesn't fix
It propagates a wrong desired state perfectly. Compare-before-write protects you from drift and from duplicate writes; it has no opinion about whether the value you decided on is correct, so a review step on the desired state is still the thing standing between you and a confidently applied mistake.
It also gives you no atomicity across records. RFC 2136 makes a single UPDATE message atomic, but an HTTP API applying records one at a time has no such property, so order matters: add the new target before removing the old one, and move MX records last, so the intermediate state is always a serving state.
The single-writer assumption is the fragile one. Console access during a migration will break it, and the loop can't tell a human edit from drift.
If the entire zone is already managed as code, stick with a desired-state tool — octoDNS and DNSControl both build a plan against live records and apply only the diff, which is this same loop at zone granularity with better ergonomics. The catch is that they want to own the whole zone. For a product that writes per-tenant records at runtime, in response to a customer clicking a button, that ownership model isn't a good fit, and you end up back here, with a writer inside your own service and a read-back that means something.
Further reading
- RFC 7489, DMARC, including the
rireporting interval default: https://datatracker.ietf.org/doc/html/rfc7489 - RFC 2136, Dynamic Updates in the DNS, prerequisite section: https://datatracker.ietf.org/doc/html/rfc2136#section-2.4
- RFC 2308, Negative Caching of DNS Queries: https://datatracker.ietf.org/doc/html/rfc2308
- RFC 1035, character-string and TTL semantics: https://datatracker.ietf.org/doc/html/rfc1035#section-3.3
- RFC 4343, DNS Case Insensitivity Clarification: https://datatracker.ietf.org/doc/html/rfc4343
- RFC 9110, idempotent methods: https://datatracker.ietf.org/doc/html/rfc9110#section-9.2.2
- Node.js dns module, Resolver and setServers: https://nodejs.org/api/dns.html
- Unbound configuration, cache-min-ttl and cache-max-ttl: https://unbound.docs.nlnetlabs.nl/en/latest/manpages/unbound.conf.html
- Amazon Route 53 GetChange, PENDING and INSYNC: https://docs.aws.amazon.com/Route53/latest/APIReference/API_GetChange.html
- octoDNS, desired-state zone management: https://github.com/octodns/octodns
Top comments (0)