DEV Community

daxharrington5274
daxharrington5274

Posted on

Live DNS Zone Reads Expose Gaps in Customer Evidence Ledgers

A healthtech platform cannot treat a successful domain check as permanent evidence. Application logs preserve the product's decisions; live DNS zone reads capture a later observation. Customers keep control of their zones, while the platform still has to explain what it accepted, when it checked, and what the domain publishes now. Those are different questions, and DNS auditing needs both records.

TL;DR: keep application events as the historical ledger, capture live DNS observations as time-stamped evidence, and reconcile the two without overwriting either. An event can prove that the platform requested or accepted a change. A live read can show the state visible during a particular observation. Neither source completes the other, and a mismatch is a result to investigate rather than data to quietly repair.

That choice adds one evidence record and one comparison function to the first implementation. It does not require a sprawling policy engine. Good. Audit code should be boring enough to inspect.

How should application logs and live DNS reads be reconciled?

A live observation answers a narrow question: what did this lookup return at this time? It cannot establish every earlier value, the actor who initiated a platform-side action, or the decision rule that accepted a domain. Reading the zone again during an audit only produces newer evidence. It does not travel backward.

Application logs have the opposite strength. They can preserve intent and workflow transitions: a customer submitted a domain, a verification job ran, and the application changed its own status. Yet a log entry saying verified is not the DNS response itself. It may describe cached input, a parser result, or a decision made under an older policy. Calling it ground truth hides all three boundaries. The reconciliation job therefore joins evidence; it does not promote one source to master and rewrite the other source to match.

No source wins.

The customer-owned zone is the constraint that changes the design. The healthtech platform can record its own actions, but it does not control every later edit in that zone. For a platform-owned zone, the platform may also possess a change trail from the system that writes records. Even then, the current observation and the write trail remain separate evidence. A requested write is not proof of a particular read result.

DMARC makes this distinction concrete. RFC 7489 describes a DNS-published policy record and aggregate feedback; it also discusses external reporting destinations and their authorization. The useful lesson is narrow: published policy, application decisions, and reports are separate artifacts with separate timestamps and provenance.

RFC 7489 is the concrete version marker here. It is not a generic DNS audit specification.

Completeness requires both timelines. It does not mean pretending either timeline is continuous.

The smallest evidence model that works

Start with three records: an application event, a DNS observation, and a reconciliation result. Keep the raw answer beside the normalized comparison value. Normalization helps code compare values; raw material lets an auditor inspect what the code received.

type ApplicationEvent = {
  eventId: string;
  customerId: string;
  domain: string;
  kind: "domain_submitted" | "verification_accepted" | "verification_revoked";
  occurredAt: string;
  policyVersion: string;
};

type DnsObservation = {
  observationId: string;
  domain: string;
  recordType: "TXT";
  observedAt: string;
  resolverLabel: string;
  rawAnswers: string[];
  normalizedAnswers: string[];
  outcome: "answered" | "no_answer" | "indeterminate";
};

type Reconciliation = {
  eventId: string;
  observationId: string;
  result: "consistent" | "drift" | "indeterminate";
  reasons: string[];
};
Enter fullscreen mode Exit fullscreen mode

resolverLabel identifies the observation path without claiming that one resolver saw everything. policyVersion matters because acceptance logic changes; without it, a future reviewer has to guess which rule produced an old decision. The timestamp fields are evidence boundaries, not decoration.

Do not turn no_answer into an empty successful answer. Do not turn a timeout or parser failure into drift. Those shortcuts make dashboards tidy and audits false. An indeterminate observation should stay indeterminate until another observation supplies evidence.

Here is the reconciliation core. It compares a verification token recorded by the application with one observed in DNS. The caller remains responsible for collecting the DNS answer and storing both input records.

type VerificationInput = {
  event: ApplicationEvent;
  observation: DnsObservation;
  expectedToken: string;
};

function reconcileVerification(input: VerificationInput): Reconciliation {
  const { event, observation, expectedToken } = input;

  if (event.domain !== observation.domain) {
    return {
      eventId: event.eventId,
      observationId: observation.observationId,
      result: "indeterminate",
      reasons: ["domain_mismatch"],
    };
  }

  if (observation.outcome === "indeterminate") {
    return {
      eventId: event.eventId,
      observationId: observation.observationId,
      result: "indeterminate",
      reasons: ["observation_failed"],
    };
  }

  const found = observation.normalizedAnswers.includes(expectedToken);
  return {
    eventId: event.eventId,
    observationId: observation.observationId,
    result: found ? "consistent" : "drift",
    reasons: [found ? "expected_token_observed" : "expected_token_absent"],
  };
}
Enter fullscreen mode Exit fullscreen mode

This pure function is cheap to test against reordered answers, duplicates, missing answers, mismatched domains, and failed observations. Rerunning it does not mutate the ledger. A changed policy can create a new reconciliation result while preserving the old one.

There is a trap here. A design that stores only the latest domain status on the customer row looks efficient during the first API call. Later, it cannot distinguish a domain that never verified from one that verified, drifted, and recovered. One field compressed a sequence into a label. That loss is permanent unless another system retained the events.

Reconciliation is a join, not a repair job

Match records by customer scope, normalized domain, evidence purpose, and an explicit time window chosen by policy. Never select “the closest DNS row” without recording why that row qualified. If no observation falls inside the policy window, return an evidence gap. Do not manufacture consistency from an observation taken later.

The output needs enough structure to answer four questions:

  • What did the application decide?
  • What DNS material was observed?
  • Which policy and code path compared them?
  • Was the evidence consistent, drifting, or indeterminate?

That final state is three-valued on purpose. Binary pass/fail logic tends to misclassify collection failures as customer misconfiguration. In a healthtech workflow, that can trigger the wrong operational response: asking a customer to edit a correct zone when the platform merely lacks a usable observation. The trade-off is extra state and a less comforting dashboard. The reason to accept it is specific: indeterminate keeps missing evidence from masquerading as negative evidence.

Reconciliation should avoid claiming ownership it cannot prove. A token match shows that the expected value was observed. It does not identify the human who changed the customer's zone. The application event may identify who requested verification inside the product, but those are two actors and two control planes.

For DMARC-specific evidence, preserve the published record and the interpretation separately. RFC 7489 defines tag syntax and processing behavior, so a future parser or policy change can alter an interpretation without altering what was observed. The same evidence principle applies to a verification TXT value even though the DMARC semantics do not: raw observation first, derived decision second.

What changes when the evidence volume grows

At scale, change storage and collection mechanics before adding policy knobs. Partition immutable events and observations by a stable tenant key and time. Keep reconciliation outputs append-only or versioned. Apply a documented retention rule to each artifact class; an application security event, a raw DNS observation, and a derived comparison may have different operational value. The rule must come from the organization's compliance requirements, not from a generic DNS recipe.

Then benchmark the paths that can distort evidence. Measure collection latency separately from reconciliation throughput. Track counts of answered, no-answer, and indeterminate observations by resolver label and record type. Watch the age of the newest usable observation per active domain. A fast comparator cannot compensate for a stalled collector.

Replay deserves its own test. Feed a fixed set of stored application events and observations into the current comparison code, then assert that every changed result names the new policy or implementation version. The goal is not to force old and new results to match. The goal is to make the difference explainable.

Tamper evidence can be layered on later with restricted write paths, immutable storage controls, hashes, signatures, or external timestamping, depending on the threat model. Do not claim that an append-only TypeScript type provides immutability. It does not. The first build should expose that boundary clearly so storage controls can strengthen it without changing the evidence schema.

Keep configuration small: observation cadence, evidence window, retention class, and policy version expose the real decisions. A dozen per-tenant retry switches will make behavior harder to reconstruct. Operational exceptions belong in recorded policy changes, not mystery flags.

The limitation is operational cost. Dual-source evidence needs scheduled collection, protected storage, retention review, and replay tests. It is a poor fit when the actual requirement is only a one-time setup check with no historical or compliance claim; in that case, a current read plus an explicit timestamp may be sufficient. It is also incomplete when the threat model requires proof against privileged storage operators. Then storage controls or external attestation must carry that part of the argument. More logs cannot solve it.

The decision rule

Use application events to answer who requested what inside the product and which rule the product applied. Use time-stamped DNS observations to answer what the collection path saw. Use reconciliation records to explain the relationship between those two artifacts.

For customer-owned zones, schedule observations around lifecycle events and at the cadence required by the actual control objective. For platform-owned zones, add the platform's DNS change events, but do not discard independent reads. In both cases, preserve gaps as gaps.

This costs more storage than one mutable verified field and demands explicit retention decisions. The return is defensible scope: no source is asked to prove facts it cannot contain. That is the practical standard for complete audit evidence.

Sources

Top comments (0)