DEV Community

tony chen
tony chen

Posted on

DNS Auditing: Application Logs vs Live Zone Reads for Complete History Reconciliation

When an audit asks what a DNS record was yesterday, propagation delay changes the answer. Short answer: use live zone reads as the current-state authority, and use application logs plus versioned snapshots to reconstruct history; neither source alone can prove complete SPF, DKIM, or DMARC coverage.

I once treated a successful change response as the end of a mail cutover. It was only the beginning. The control plane had accepted a new DMARC policy, but recursive resolvers still returned the previous TXT value for the old TTL. A log showed intent and timestamp. A live read showed one resolver's view. The audit question needed both, tied to the same change ID.

That distinction is easy to miss.

What should DNS auditors compare in application logs and live zone reads?

Application logs are an event stream: who requested the change, which validation ran, what the authoritative service returned, and when the write was acknowledged. They are strong evidence of intent and workflow. They are weak evidence of what an external resolver actually served later. A missing log line can mean a logging gap, a rejected request, or an operation that happened outside the application.

Live zone reads answer a different question. Query an authoritative name server for the current RRset and record the resolver, timestamp, response code, TTL, and DNSSEC validation result where relevant. A live read is excellent for present state, but it cannot tell you which value was served last Tuesday after a record was replaced. Caches also make a recursive lookup a measurement of that resolver's cache, not a canonical history.

For reconciliation, normalize each source into an evidence record. Keep the owner name, type, value set, TTL, observed-at time, actor or system identity, and change ID. Store the full RRset rather than one string, because SPF can have multiple TXT fragments and DKIM selectors are separate names. Hashing the normalized value helps detect a change without hiding the original evidence.

How do SPF, DKIM, and DMARC propagation affect audit history completeness?

The records have different operational shapes. SPF is usually a TXT record at the organizational domain and can break when a migration exceeds the DNS lookup limit defined by its standard. DKIM publishes a selector-specific TXT record, so rotating selectors creates a period where old and new keys may both matter. DMARC lives at _dmarc and its policy (p=none, quarantine, or reject) influences enforcement after resolvers refresh their cached answer.

Propagation is not a single global event. Authoritative publication can be complete while receivers continue using a cached RRset until its TTL expires. During a cutover, sample several authoritative servers and several independent recursive resolvers, then record the exact answers. Do not label an absent answer as deletion until the negative-cache window is understood; NXDOMAIN responses have their own caching behavior.

Here is the failure mode I put in an eval harness before approving a mail-domain change. The harness publishes a new DKIM selector, records the authoritative response, and then asks three recursive resolvers at fixed intervals. One resolver returns the new key immediately, another keeps the old answer for the full TTL, and the third returns an empty response because the selector name was mistyped in the probe. If the report stores only the final “pass” result, those three states collapse into one misleading line. A reviewer cannot tell whether the key propagated slowly, whether the probe was wrong, or whether a resolver had cached NXDOMAIN. I keep each raw response, the query name, the resolver address, and the request timestamp, then attach a classification to the observation. That extra data makes the report larger, but it also makes a disputed cutover replayable. The same pattern applies to a DMARC policy flip and to SPF changes that add or remove an include domain.

Short windows need stronger evidence.

The practical audit rule is to separate three timestamps: requested-at, authoritative-published-at, and observed-at. A compliance report that collapses them into “changed at” hides the delay that caused the incident. Your mileage may vary across resolvers, and I'm not sure any fixed polling interval is defensible without looking at the zone's TTL and the receiver population.

A small evidence model for reconciliation

The model below keeps intent and observation separate. It is deliberately plain Python so it can run in a notebook during an investigation and later in a scheduled job.

from dataclasses import dataclass
from datetime import datetime
from hashlib import sha256


@dataclass(frozen=True)
class Evidence:
    source: str  # "application_log" or "live_read"
    name: str
    rrtype: str
    values: tuple[str, ...]
    observed_at: datetime
    change_id: str | None
    ttl: int | None

    @property
    def value_hash(self) -> str:
        normalized = "\n".join(sorted(self.values)).encode("utf-8")
        return sha256(normalized).hexdigest()


def reconcile(log_event: Evidence, live_read: Evidence) -> str:
    if (log_event.name, log_event.rrtype) != (live_read.name, live_read.rrtype):
        return "different_rrsets"
    if log_event.value_hash == live_read.value_hash:
        return "observed_match"
    return "pending_propagation_or_untracked_change"
Enter fullscreen mode Exit fullscreen mode

The final status is not a verdict by itself. pending_propagation_or_untracked_change should trigger another authoritative read, a check of the change ID, and a search for out-of-band edits. Preserve the raw responses alongside the normalized row; a reviewer may need to inspect quoting, TXT chunk boundaries, or DNSSEC status.

Choosing evidence by cutover speed and propagation delay

Fast cutovers favor short, controlled TTLs before the change, a staged selector or policy, and a poller that records observations from named vantage points. That increases query volume and operational work. Slow, planned migrations can tolerate longer TTLs and fewer probes, but they need a longer overlap window for DKIM keys and a documented rollback record.

Audit need Primary evidence Why Limitation
Prove who approved a change Application log and ticket Captures intent and identity Does not prove public visibility
Prove current published RRset Authoritative live read Shows present source of truth No historical record
Prove resolver exposure during cutover Timestamped recursive samples Exposes cache and propagation delay Vantage points are incomplete
Prove a complete timeline Immutable logs plus snapshots and reads Joins intent, publication, and observation Requires retention and clock discipline

The catch is cost in the broad engineering sense: retention, clock synchronization, probe credentials, and a schema that survives record formatting changes. This approach is not suitable when you cannot retain immutable evidence or identify which authoritative servers were queried. In that case, stick with a smaller claim: current-state verification, not historical completeness.

I measure two gaps before declaring a cutover done: the time between authoritative publication and the last stale recursive observation, and the percentage of change IDs that have at least one matching live read. Those metrics expose a fast API with slow propagation, and they keep the audit conversation grounded in evidence rather than a green deployment badge.

For mail delivery, the durable pattern is simple: logs explain intent, authoritative reads establish publication, and resolver samples show propagation. SPF, DKIM, and DMARC each need their own record-aware checks. Reconciliation is complete only when those three views agree within a stated time window.

References

Top comments (0)