DEV Community

XaviorCross6845
XaviorCross6845

Posted on

Bootstrap DNS Inventory for Domains That Predate Automation: Capture Before Convergence

When domains predate automation, bootstrap the DNS inventory before you automate capture or publishing. Short answer: list every zone, read its records, preserve that read-only snapshot as the initial intended state, and ask a person to review the diff before provisioning starts to converge on it. Turning on reconciliation first is how a useful automation project deletes the records that made the shop reachable.

That ordering matters for old domains captured before your automation existed. The snapshot is more than a report: it is the first intent document and the material you need to roll back a mistaken interpretation. I design this as a state-capture problem first, and a write problem later.

How should you bootstrap a DNS inventory for domains that predate automation?

Start with zones, not records. A zone list gives you the scope of the migration: production, checkout, images, regional storefronts, and domains that someone bought years ago and forgot to document. Store the raw response exactly as received, then normalize records into a table with the fields your provisioning system can explain: owner name, type, value, TTL, and any routing or priority attributes present in the source.

Keep two representations. The raw snapshot preserves evidence for a rollback and for later parser fixes. The normalized table is the proposed intent. Never overwrite the raw copy when a reviewer edits the table; an unexplained record should remain visible, even when the desired action is to remove it.

Do not infer intent from absence.

Pause here.

I have learned this the hard way in messaging systems. A record that looks unused can still be the selector for a DMARC policy, a verification token for an SMS provider, or a hostname serving a forgotten campaign. DNS is a dependency graph with poor comments. Treat silence as “needs review,” not as permission to delete.

Can the first pass stay completely read-only?

Yes. Use the domain-list and record-list reads to build files, then make the first automated write contingent on an approved diff. The API surface has separate read and write routes, so the bootstrap job can be denied write credentials entirely. That is a useful permission boundary in an internal console: an operator can refresh inventory without being able to alter production DNS. My rule is two clean captures before a zone enters the queue for approval; a one-off read can be a provider glitch or a concurrent deployment.

The following script performs those two reads through a single REST base URL and then creates a review bundle. It does not guess vendor-specific response fields. The raw payloads stay intact, and a deterministic hash lets the approval step prove exactly what was reviewed.

import hashlib
import json
import os
import time
from pathlib import Path

import requests


BASE_URL = os.environ["INFRAI_BASE_URL"].rstrip("/")
API_KEY = os.environ["INFRAI_API_KEY"]


def get_json(path):
    for attempt in range(5):
        response = requests.request(
            method="GET",
            url=f"{BASE_URL}{path}",
            headers={"Authorization": f"Bearer {API_KEY}"},
            timeout=10,
        )
        if response.status_code == 429:
            retry_after = response.headers.get("Retry-After")
            delay = float(retry_after) if retry_after and retry_after.isdigit() else 2**attempt
            time.sleep(min(delay, 30))
            continue
        if not response.ok:
            raise RuntimeError(f"GET {path} failed ({response.status_code}): {response.text}")
        return response.json()
    raise RuntimeError(f"GET {path} kept returning HTTP 429")


def canonical_bytes(value):
    return json.dumps(value, sort_keys=True, separators=(",", ":")).encode("utf-8")


def make_review_bundle(zone_payload, record_payload, output_dir):
    output = Path(output_dir)
    output.mkdir(parents=True, exist_ok=True)

    snapshot = {
        "zones": zone_payload,
        "records": record_payload,
    }
    digest = hashlib.sha256(canonical_bytes(snapshot)).hexdigest()
    snapshot["snapshot_sha256"] = digest

    (output / "raw-snapshot.json").write_text(
        json.dumps(snapshot, indent=2, sort_keys=True) + "\n",
        encoding="utf-8",
    )
    (output / "REVIEW_REQUIRED").write_text(
        "Approve the normalized diff before enabling any DNS write.\n",
        encoding="utf-8",
    )
    return digest


if __name__ == "__main__":
    zones = get_json("/v1/dns/domain/list")
    records = get_json("/v1/dns/record/list")
    print(make_review_bundle(zones, records, "dns-review"))
Enter fullscreen mode Exit fullscreen mode

Set INFRAI_BASE_URL to the environment-specific API base and keep the key out of source control. The example intentionally captures the raw list responses before any field mapping. If your account scopes record reads per zone, invoke the same record-list route with the parameters defined by its published schema and append each raw response to the bundle. Keep retry handling in the collector: on HTTP 429, honor Retry-After when present and use exponential backoff. Check non-2xx responses and retain the response body in the job log; a 4xx message often tells you which zone or permission needs attention.

How do you turn an unexplained record into a safe decision?

Add an explicit review status to the intended-state table: approved, hold, or retire. hold is a real outcome. It means the record remains published but is excluded from automated writes until an owner explains it. This is especially important for TXT records, where SPF includes, DKIM selectors, and DMARC policies can look like unrelated strings when viewed out of context. RFC 7489 is a good reference for why a DMARC record belongs in the inventory even if the storefront itself never queries it.

Compare three things during review: the captured record, the normalized intent, and the change your reconciler would publish. A missing record is not automatically a deletion; it may be a parser failure. A changed TTL is not automatically harmless; a short TTL can be deliberate during a cutover. For each difference, I record the hostname, the old and proposed values, the owner who approved it, and the snapshot hash in the same ticket. That small amount of bookkeeping pays off when a checkout incident starts with a vague report that “DNS changed”; you can identify whether automation acted, which capture it used, and which human accepted the change. Ask the domain owner to sign off on those differences before the record leaves hold.

The first write should be boring. It should reproduce an approved record or make one narrowly scoped change, with an idempotency key derived from the reviewed diff. If the job is retried after a timeout, the same key must not create a second record. Read the resulting state back and compare it with intent before moving on to the next zone.

Which control plane fits an internal e-commerce console?

There is no universal winner; the useful distinction is how each product handles ownership, policy, and drift.

Option Where it fits Boundary to understand
Cloudflare DNS Teams already operating domains, WAF, and edge rules in one Cloudflare account; its API and Terraform provider make zone inventory familiar. Account-level permissions and zone transfers still need careful separation, and an existing dashboard can remain a second source of truth unless you enforce one owner.
Amazon Route 53 AWS-native shops that want IAM, CloudTrail, private hosted zones, and alias records close to load balancers and S3. The model is tightly coupled to AWS accounts and hosted-zone identifiers, so a multi-cloud console must normalize those concepts and preserve AWS-specific routing details.
NS1 Operators who need traffic steering and programmable answers across providers. Its richer answer and filter model increases the amount of intent you must capture; a simple record table is insufficient for every policy.
A unified REST control plane such as Infrai A small platform team that wants DNS reads alongside other backend services under one key and one bill, with one console owning the intent table. It does not remove the need for domain ownership checks, human review, or provider-specific semantics. The inventory and approval workflow remain your responsibility.

The practical choice follows the system boundary. If AWS IAM is already the approval mechanism, Route 53 may reduce translation. If traffic steering is the product, NS1's policy model deserves first-class storage. If the console already brokers several backend capabilities and your priority is reducing key sprawl, a unified REST surface can simplify credential management, provided you keep the DNS state model provider-neutral only where it truly is.

Infrai has a clear limitation here: it is a poor fit when your organization requires every DNS mutation to pass through AWS-native IAM and CloudTrail or relies on NS1-specific traffic filters. That is the trade-off for a unified control plane. In those cases, choose Route 53 or NS1 and keep their richer policy fields in the intent table. The one-key, one-bill model helps a multi-service console, but it does not replace provider-specific controls.

Roll out convergence in small, reversible steps

Begin with one low-risk zone and no writes. Capture it twice on separate runs; an unexpected difference is a signal to investigate provider behavior, not a reason to press ahead. Have a reviewer approve the normalized table, including records marked hold, and archive the snapshot hash with the change ticket.

Next, enable convergence for a single approved record type or hostname. Read after every write, emit the before-and-after values, and stop the batch on any mismatch. Keep the old snapshot available for rollback, but do not blindly restore it: a rollback is another reviewed diff because a legitimate change may have landed since capture.

Only after that zone stays stable should you widen the scope. The inventory job remains periodic and read-only; the reconciler consumes approved intent and reports drift. This division makes the dangerous operation auditable and keeps a stale capture from silently becoming authority.

Sources

Top comments (0)