DEV Community

ZachariahHolloway9058
ZachariahHolloway9058

Posted on

How to Log Every DNS Record Change Before It Propagates (400 Domains)

Write the audit record in the console that took the click, one transaction before the provider API call, and keep it somewhere you can query by zone, actor, and time. Use the zone itself only as a cross-check. A zone holds current state: it will tell you that www resolves to a new load balancer, never who changed it, what the old value was, or which TTL was in effect when the change went out. Provider change logs sit closer to the truth, but from the provider's side the actor is your service credential — so in a property management console that fronts 400 domains, one per community site, every edit is attributed to console-prod and the audit trail dead-ends there.

The actor identity exists in exactly one place: the call site.

That single fact decides the design, and most of the work after it is making sure the record survives retries, out-of-band edits, and the gap between when you write a record and when the internet believes it.

What the zone knows, and what the compliance review asks

The two questions are not the same shape. A zone answers "what is the MX for this domain right now." A review asks "who repointed the MX for 118 Maple Court on the 14th, what did it say before, and who approved it." Current-state reads cannot answer the second one at any price, because the previous value stopped existing the moment it was overwritten.

Propagation delay makes the timestamp ambiguous too. A record edited at 14:02 with a 3600-second TTL can still be served from resolver caches well past 15:00, which means "when did this change take effect" has two honest answers — the authoritative write time and the cache expiry window — and an audit row that stores only the first one will mislead whoever reads it during an incident. Store both the TTL in effect before the write and the one you set. Negative answers cache too, governed by the SOA minimum, so a record that briefly did not exist has its own decay window (RFC 2308 covers the rules, and they surprise people more often than the positive TTL does).

I learned this reflex from queue consumers rather than from DNS: any event that can be delivered twice needs a record that converges instead of accumulating. DNS writes have the same property. A console retry, a double-clicked save button, and a resumed background job all produce the same intent, and if each one appends a fresh audit row, the reviewer now has three edits to explain and no way to tell which one moved the world.

Where should the record of who changed a DNS record live?

Three homes are plausible, and they fail in different places. The provider's change feed is free and requires no code from you, but it records the credential rather than the person, its retention is set by someone else's policy, and it disappears the day you move zones to a different provider. Git-tracked zone config is the strongest audit artifact of the three: with DNSControl or octoDNS, the record of who changed what is the commit, reviewed before it lands. The catch is cutover speed — every change now costs a pull request and a CI run, which is the correct trade for a slow, deliberate estate and the wrong one for a support agent who needs to repoint a property's site during a 20-minute maintenance window.

Where the record lives Knows the human actor Keeps the previous value Survives a provider migration Cost to cutover speed
Provider change feed No — sees the API credential Sometimes, for a fixed window No None
Git-tracked zone config Yes, via commit and review Yes, in full history Yes Minutes to hours per change
Console change table, shipped to logs Yes, from the session Yes, if you read before write Yes Milliseconds

The third row is the one I'd build for an internal admin console, with a caveat that matters: the database row is the authoritative record and the log pipeline is the searchable copy, not the reverse. Logs are where you answer "show me every TXT change across all zones last quarter" in a few seconds; a relational row is where you enforce uniqueness and referential integrity against the ticket that authorized the change. NIST's log management guidance is worth reading on retention and protection before you pick a store, because "we keep it in the app database forever" is a policy statement, not a design.

Write the intent before the change, not after

The ordering is the whole trick. Read the current value, record the intent, then call the provider — so that a crash between steps leaves you with an audit row marked pending rather than a mutated zone nobody can explain.

package dnsaudit

import (
    "context"
    "errors"
    "fmt"
    "time"
)

// Change is the audit row. It is written before the provider call and
// updated once, under the same ID, when the write is confirmed.
type Change struct {
    ID       string // client-generated; also the provider idempotency key
    Zone     string // "118-maple-court.example"
    Name     string // "www"
    Type     string // A, CNAME, MX, TXT
    OldValue string
    NewValue string
    OldTTL   int // seconds in effect before this write
    NewTTL   int
    Actor    string // signed-in console user, never the API credential
    Ticket   string
    At       time.Time
}

var ErrDuplicate = errors.New("change already recorded")

type Store interface {
    // Insert returns ErrDuplicate when the ID is already present.
    Insert(ctx context.Context, c Change) error
    MarkApplied(ctx context.Context, id string, serial uint32) error
    MarkFailed(ctx context.Context, id, reason string) error
}

type Zone interface {
    Get(ctx context.Context, zone, name, rtype string) (value string, ttl int, err error)
    Upsert(ctx context.Context, zone, name, rtype, value string, ttl int, idem string) (uint32, error)
}

func Apply(ctx context.Context, s Store, z Zone, c Change) error {
    cur, ttl, err := z.Get(ctx, c.Zone, c.Name, c.Type)
    if err != nil {
        return fmt.Errorf("read current record: %w", err)
    }
    c.OldValue, c.OldTTL = cur, ttl

    switch err := s.Insert(ctx, c); {
    case errors.Is(err, ErrDuplicate):
        // Retry of an intent we already recorded. Nothing to append; the
        // provider write below carries the same ID and converges.
    case err != nil:
        return fmt.Errorf("record intent: %w", err)
    }

    serial, err := z.Upsert(ctx, c.Zone, c.Name, c.Type, c.NewValue, c.NewTTL, c.ID)
    if err != nil {
        if mErr := s.MarkFailed(ctx, c.ID, err.Error()); mErr != nil {
            return fmt.Errorf("apply %s: %v; audit update also failed: %w", c.ID, err, mErr)
        }
        return fmt.Errorf("apply %s: %w", c.ID, err)
    }
    return s.MarkApplied(ctx, c.ID, serial)
}
Enter fullscreen mode Exit fullscreen mode

Actor is the field people get wrong. It must come from the authenticated console session, not from a request body field and not from the credential the server happens to hold, or you have rebuilt the same dead end the provider feed gave you.

There's a race in the read-then-write, and I'd rather name it than pretend otherwise: another operator can edit the same record between Get and Upsert, which makes OldValue a claim rather than a guarantee. Where the API supports a conditional write against the observed value, use it. Where it doesn't, the reconciler in the next section is what catches the divergence, and the detection delay is your polling interval.

Reconcile the log against the zone on a schedule

An internal console is never the only way a record changes. Someone edits at the registrar, a platform integration adds a verification TXT, a contractor still holds a token from a migration two years back. None of that passes through your call site, so none of it lands in your table.

So diff it. A scheduled job lists every record in every zone, compares the result against the last applied state you recorded, and emits an event for each difference with no matching change row. Run it every 15 minutes if your change volume justifies it, hourly if it doesn't. Keep the job read-only: a reconciler that quietly reverts drift destroys the evidence that a review needs, and it turns a detectable problem into a recurring mystery. Report, alert, let a human decide.

Mail records deserve their own alert rule. A _dmarc TXT record that moves from p=reject to p=none is a one-line change with a policy-sized consequence, and because DMARC failures degrade quietly rather than loudly, nobody pages you — the aggregate reports just start looking different a week later (RFC 7489 defines the record format and the reporting flow). The same goes for SPF and for the CAA records that constrain who can issue certificates for a property's domain.

For a 400-domain estate, the reconciler doubles as your zone inventory. It's the only component that knows about the domain someone registered in 2026 and never told the console about.

Verification and rollback, or why the TTL is the real decision

Cutover speed is bought in advance, not at the moment you need it. Lower the TTL on the record you are about to change — 300 seconds is a common landing spot — and do it at least one full old-TTL period before the cutover, so caches holding the 3600-second copy have expired by the time you flip the value. That TTL reduction is itself a change, and it goes through the same audited path. Two rows, one ticket.

Then verify against more than one vantage point, because the authoritative answer and the cached answer are different facts:

dig +short @ns1.provider.example www.118-maple-court.example A
dig +short @9.9.9.9 www.118-maple-court.example A
dig +noall +answer www.118-maple-court.example A
Enter fullscreen mode Exit fullscreen mode

The third command prints the remaining TTL alongside the value, which tells you how much of the cache window is left. If the authoritative server has the new value and a public resolver still has the old one, nothing is broken — you are simply inside the propagation window you paid for, and the audit row you wrote earlier tells you exactly how long that window is.

Rollback follows the same rule as everything else here: it is a new audited change that restores OldValue and OldTTL, never a quiet fix applied out of band. That is what makes the record of who changed what usable during an incident instead of after it.

A few boundaries worth stating plainly. This design is not a good fit for a team with a dozen domains and monthly changes — git-tracked zone config gives you a better audit trail for less code, and the cutover latency won't hurt you. It doesn't satisfy regulators who require write-once storage with independent retention; for that, ship the same events to a store whose retention you cannot edit from the application. And it can't prove that a record was never changed between two polls, only that it matches now; I'm not sure any application-side design can close that gap without the provider exposing a signed change history. If yours does, that changes the calculus, and I'd revisit the table above.

The postmortem question I'd want answered on day one: for any record in any of those 400 zones, can you name the human, the ticket, the previous value, and the cache window — without reading a single application log line by hand? If the answer needs a grep, the record isn't living in the right place yet.

References

Top comments (0)