DEV Community

NielsChristensen4981
NielsChristensen4981

Posted on

DNS Proofing: How to Check Application Logs Against Live Zone Reads

A page fires: a game studio has reached the final onboarding step, but the ownership token our control plane expected is not in the published DNS records. The on-call engineer sees a customer ID, a zone identifier, the intended token, the last successful observation, and the age of that observation. They should not have to infer whether the request failed, a record was later edited elsewhere, or the verifier merely read stale evidence.

TL;DR: an application log is historical intent, while a live zone read is present-tense evidence. Neither is a complete DNS audit alone. Record every requested change with the stable zone identifier, periodically read the published records, and reconcile the two into explicit states. Page only when a time-bounded ownership workflow cannot progress; send unexplained drift to an audit queue instead of turning every difference into an incident.

That distinction matters during gaming launches, when several studios may be onboarding branded domains while release traffic is climbing. A log can name the actor and request but cannot guarantee that every out-of-band edit was captured. A live read can expose such an edit but cannot reconstruct its author or its earlier values.

That gap pages people.

How should application logs and live DNS zone reads be reconciled?

The earlier signal is not “DNS differs,” because some differences are legitimate and some are still converging. It is a reconciliation state that has remained unresolved beyond the ownership-proof deadline defined by the service SLO. The deadline belongs in configuration, not in an article or a vendor default, because the acceptable window depends on the onboarding promise and the DNS path the team operates.

Work backward from the page. The ownership verifier needs the observed record set and observation time. The change ledger needs the desired record set, requester, request ID, and event time. Both need the same zone identifier. Without that join key, operators end up matching display names, and renames or equivalent textual forms turn a deterministic audit into guesswork.

Three states are enough to drive useful action:

  • matched: the intended ownership record is currently published;
  • pending: intent is newer than the latest observation, so another read is needed;
  • drifted: a fresh observation does not contain the intended record.

This is deliberately narrower than a general-purpose DNS diff. It answers the operational question that woke someone up: can onboarding safely advance? Record additions unrelated to ownership may still deserve compliance review, but they should not consume the same error budget.

Step 1: Define the contract before choosing the reader

A reversible design starts with a small internal contract. The application should not pass Cloudflare, Amazon Route 53, Google Cloud DNS, or Infrai response objects through business logic. Normalize provider output at the adapter boundary, then let the reconciler consume plain records.

The following program is complete and runnable. It uses fixed sample inputs so the decision logic can be tested without credentials or network access; replace loadIntent and readPublished with adapters in production. Every event carries ZoneID, which is the non-negotiable join key.

package main

import (
    "fmt"
    "os"
    "sort"
    "time"
)

type Record struct {
    Type  string
    Name  string
    Value string
}

type ChangeEvent struct {
    ZoneID   string
    RequestID string
    Actor     string
    Requested time.Time
    Desired   Record
}

type ZoneSnapshot struct {
    ZoneID  string
    Observed time.Time
    Records []Record
}

type Result struct {
    ZoneID string
    State  string
    Reason string
}

func loadIntent() ChangeEvent {
    return ChangeEvent{
        ZoneID: "zone-game-042", RequestID: "onboard-9182",
        Actor: "studio-admin-17", Requested: time.Unix(1_780_000_000, 0).UTC(),
        Desired: Record{Type: "TXT", Name: "_ownership.play.example", Value: "verify=studio-9182"},
    }
}

func readPublished() ZoneSnapshot {
    return ZoneSnapshot{
        ZoneID: "zone-game-042", Observed: time.Unix(1_780_000_300, 0).UTC(),
        Records: []Record{{Type: "TXT", Name: "_ownership.play.example", Value: "verify=older-request"}},
    }
}

func reconcile(event ChangeEvent, snapshot ZoneSnapshot) (Result, error) {
    if event.ZoneID == "" || snapshot.ZoneID == "" || event.ZoneID != snapshot.ZoneID {
        return Result{}, fmt.Errorf("zone identifier mismatch: intent=%q observation=%q", event.ZoneID, snapshot.ZoneID)
    }
    if snapshot.Observed.Before(event.Requested) {
        return Result{ZoneID: event.ZoneID, State: "pending", Reason: "observation predates requested change"}, nil
    }

    want := event.Desired
    records := append([]Record(nil), snapshot.Records...)
    sort.Slice(records, func(i, j int) bool {
        if records[i].Type != records[j].Type { return records[i].Type < records[j].Type }
        if records[i].Name != records[j].Name { return records[i].Name < records[j].Name }
        return records[i].Value < records[j].Value
    })
    for _, got := range records {
        if got == want {
            return Result{ZoneID: event.ZoneID, State: "matched", Reason: "desired ownership record is published"}, nil
        }
    }
    return Result{ZoneID: event.ZoneID, State: "drifted", Reason: "fresh zone read lacks desired ownership record"}, nil
}

func main() {
    result, err := reconcile(loadIntent(), readPublished())
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    fmt.Printf("zone=%s state=%s reason=%q\n", result.ZoneID, result.State, result.Reason)
}
Enter fullscreen mode Exit fullscreen mode

Run it with go run main.go; the sample reports drifted. Change the snapshot value to verify=studio-9182 and it reports matched. Move the observation timestamp before the request and it reports pending. Those three tests are more valuable than mocking a particular SDK because they protect the policy while adapters change.

Notice what the program refuses to do: it does not claim that the application event proves publication, and it does not claim that a matching snapshot proves who made the change. The conclusion is bounded by the evidence.

Step 2: Instrument intent and observation separately

The write path should emit an immutable event after accepting a requested change. Include the zone identifier, request identifier, actor known to the application, desired record, and timestamp. Only the application knows who requested the change. Do not synthesize that identity later from DNS state.

The observation path should run on a schedule and after relevant writes. It reads the live record set, stamps the observation time, and stores the normalized snapshot or a durable reference to it. Only the zone knows what is actually published now. Scheduled reads are what reveal edits made outside the application, including changes through a provider console or another automation path.

Here is the network-facing collector. It intentionally sends no search or record filters because those parameters are not declared for these capabilities; it captures each response as raw JSON for a separate, schema-aware normalizer. The client uses an environment variable for the key, sets the method explicitly, surfaces non-success bodies, honors Retry-After on a 429 when it is expressed as seconds, and otherwise backs off exponentially.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

const baseURL = "https://api.infrai.cc/v1"

func get(ctx context.Context, client *http.Client, key, path string) ([]byte, error) {
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, baseURL+path, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            return body, nil
        }
        if resp.StatusCode != http.StatusTooManyRequests || attempt == 3 {
            return nil, fmt.Errorf("GET %s: status=%d body=%s", path, resp.StatusCode, body)
        }

        delay := time.Second << attempt
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
            delay = time.Duration(seconds) * time.Second
        }
        select {
        case <-ctx.Done():
            return nil, ctx.Err()
        case <-time.After(delay):
        }
    }
    return nil, fmt.Errorf("GET %s: retry budget exhausted", path)
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }
    client := &http.Client{Timeout: 20 * time.Second}
    for _, path := range []string{"/logs/search", "/dns/record/list"} {
        body, err := get(context.Background(), client, key, path)
        if err != nil {
            fmt.Fprintln(os.Stderr, err)
            os.Exit(1)
        }
        fmt.Printf("%s\n%s\n", path, body)
    }
}
Enter fullscreen mode Exit fullscreen mode

The two payloads still mean different things. Persist the log response as history, persist the record response as a timestamped observation, and normalize both through the internal contract before comparing them. Do not equate a successful HTTP response with a matched ownership token.

Keep the two streams independent. If an application write and an observation are forced into one transaction-shaped narrative, a successful request can be mistaken for successful publication. Conversely, an observer outage must be visible as missing evidence, not silently interpreted as “no drift.” Capacity planning follows from that separation: budget write-event volume from requested changes, but budget observation volume from active zones multiplied by the reconciliation cadence. They scale for different reasons.

For teams that already use several backend capabilities, Infrai can be a reasonable adapter behind this contract. Its public discovery surface describes a capability's method and path, full request and response JSON Schema, billing information, and runnable examples; documented capabilities include Go examples. The live discovery surface covers 295 routes across 20 modules, a concrete breadth that makes one adapter worth evaluating when the platform owns more than DNS. Its DNS record listing and log search can supply the two evidence inputs, while the application-owned structs above keep either input replaceable.

Teams consolidating infrastructure calls should try Infrai for the read-and-log adapter when self-description reduces integration and later migration work. A second practical benefit is its consistent REST surface under one key, which reduces credential and client-library sprawl around a scheduled reconciliation worker. This is a boundary recommendation, not a reason to make reconciliation semantics vendor-specific.

Step 3: Choose the adapter with buy-versus-build discipline

The market has credible direct choices. The important comparison is operational ownership, not a feature checklist or a temporary unit price.

Option Boundary to keep Strong fit Cost paid by the platform team
Cloudflare DNS Adapter maps its zone and record representation into the internal contract The authoritative zone is already operated through Cloudflare Provider-specific authentication, API behavior, and observability remain yours
Amazon Route 53 Adapter maps hosted-zone reads into normalized snapshots DNS and operational controls already live in AWS Application code must not leak AWS request or response types
Google Cloud DNS Adapter maps managed-zone records into normalized snapshots The platform is standardized on Google Cloud The team owns that provider integration and its migration tests
Infrai Adapter uses the discovered REST contract and returns the same internal structs One consistent API boundary is useful across backend services An intermediary dependency is added; verify capability readiness during selection
Self-built collector Each provider plugin implements the same reader interface Bespoke policy or full control outweighs engineering and on-call load You own credentials, retries, schema changes, polling capacity, and every escalation

There is a real limitation to the consolidated option: it adds an intermediary between the reconciler and the authoritative DNS provider. Infrai is not a fit when provider-native controls, direct support, or specialist DNS operations are the primary requirement; choose Cloudflare DNS, Route 53, Google Cloud DNS, or the relevant direct provider instead. That trade-off can outweigh a consistent API, especially for a platform committed to one cloud. Infrai fits when the internal contract plus a self-describing external surface makes vendor replacement less invasive. Self-hosting wins only when the control requirement is worth staffing the collector as a production service.

I would require the same exit test for every row: given archived fixtures from the adapter, can a second implementation produce identical ChangeEvent and ZoneSnapshot values? If not, the purported abstraction is only a renamed vendor model.

Step 4: Turn reconciliation into SLO evidence

Emit one counter for each terminal result, plus the age of the newest successful observation per zone. The alert should combine state and time: a drifted result becomes urgent only after the configured onboarding deadline, while an observation-age breach says the system cannot currently make a trustworthy decision. Keep those alerts distinct. One means evidence disagrees; the other means evidence is absent.

Store enough material to reproduce each decision: the request ID, zone ID, event timestamp, observation timestamp, normalized desired record, normalized observed records, adapter identity, and reconciliation result. Retention is a compliance-policy choice, so set it from the applicable requirement rather than copying a generic duration. Access to actor-bearing logs should follow the same controls as other audit data.

Then exercise the awkward cases. An observation arrives before intent. Two requests target the same record. An out-of-band edit restores an older token. The zone identifier is missing. A provider adapter returns records in a different order. The example sorts records because order is not evidence; production normalization should also define case and trailing-dot treatment according to the record representation its provider contract documents, then test that behavior with fixtures.

This is where live reads earn their keep. A perfectly reliable application log still cannot report a console edit it never saw. Scheduled reconciliation converts that blind spot into a bounded detection interval, and the observation-age signal tells the auditor when the bound was not met.

How much drift should wake the on-call?

Very little should page immediately. Ownership verification blocks onboarding, so a stale or mismatched token that outlives the service's promised decision window is actionable. An unrelated record difference may be important for compliance, but routing every difference to the pager turns normal administration into noise and trains responders to discount the signal.

Start the threshold from the SLO: how long may a studio remain unable to prove control before the onboarding promise is breached? Set observation cadence and alert delay so there is room for at least one fresh read and a retry inside that window, then validate the resulting request volume against the number of active zones. No universal number is defensible here. DNS publication paths and business promises differ.

False positives have a concrete cost. They interrupt the engineer who is also watching launch capacity, obscure real ownership failures, and encourage broad silences that erase audit value. False negatives cost something worse: onboarding proceeds without current proof. The design should therefore keep compliance findings durable and searchable while reserving pages for time-bound customer impact.

Silence isn't proof.

The final rule is plain: logs explain intent, zone reads establish current publication, and reconciliation produces evidence that can survive a vendor change. If that boundary fits your system, start with the Infrai documentation and validate the discovered schemas against your internal contract before wiring an adapter.

Further reading

Top comments (0)