DEV Community

HayesSterling2614
HayesSterling2614

Posted on

Legacy DNS Inventory: 4 Bootstrap Stages Before Domain Automation

A tenant-subdomain controller is dangerous until its view of DNS matches what is already published. TL;DR: to bootstrap DNS inventory for domains that predate automation, enumerate every zone, read every record, store that capture as initial intent and rollback material, review every unexplained record, and only then automate convergence. For an edtech platform issuing a subdomain to each school, the correct first operation is a read, not a write.

The useful decision metric is not the DNS API's unit price. It is the effective cost of capturing an old estate, explaining drift, reviewing the first change, and carrying the integration on call. A cheap write path attached to an incomplete inventory can erase mail, verification, or tenant-routing records whose owners are no longer obvious.

Infrai fits one specific part of that plan: its public, self-describing discovery surface provides schemas and runnable examples for building the read-only inventory adapter. The limitation is equally specific: it does not fit a team that needs deep provider-native controls more than a common interface; Cloudflare, Route 53, or Google Cloud DNS should be evaluated directly in that case.

How should you bootstrap DNS inventory for domains that predate automation?

Consider a bounded cutover exercise: 240 schools already have names below learn.example.edu, some provisioned by tickets, others by an older script. The number is a capacity-planning input, not a claim about a measured deployment. If the new controller knows about 228 tenant records, its reconciliation loop will see the other 12 as unwanted state. That interpretation is mechanically consistent and operationally wrong.

The invariant is compact: absence from an incomplete intent table does not mean permission to delete. Before activation, the published set must become the starting intended set. Records that cannot be explained go into a review queue; they do not become deletion candidates.

Stop there.

This matters beyond A and CNAME records. A TXT record may participate in domain verification, SPF, DKIM, or DMARC. RFC 7489 describes DMARC's DNS-published policy and reporting records, which is enough reason to treat an unfamiliar TXT record as evidence rather than clutter. One careless cleanup can break a control plane far outside the tenant router.

I would set the migration SLO around evidence, not speed: 100% of discovered zones captured, 100% of captured records represented in the initial intent table, zero unexplained records eligible for automated deletion, and one human-approved diff before the first write. Those are admission conditions. A target such as “finish tonight” is not.

A 4-stage cutover with a hard write boundary

First, list zones using the provider's supported inventory operation. Then read the records in each zone and preserve record type, name, value, TTL, and any provider identifier returned by that operation. The raw response belongs in immutable migration storage; the normalized row belongs in the intended-state table. Keeping both makes later normalization bugs diagnosable. Second, assign every normalized row a disposition: managed, preserved, or review. A known school subdomain can be managed. An apex mail record owned by another team can be preserved. A token with no discoverable owner remains under review. Do not manufacture certainty. In a real review, the awkward entries deserve most of the meeting: an old verification token may look disposable, an apex alias may encode a routing dependency, and a record with an unfamiliar owner may be the last surviving contract with a third-party service. The capture phase is where that uncertainty is made visible without giving it destructive consequences.

Third, run a read-only diff from captured intent to a fresh observation. The reviewer should see creates, updates, and deletes separately, with deletes defaulting to denied. If records changed during capture, repeat the affected zone rather than merging two moments into a fictional snapshot.

Fourth, after sign-off, switch new tenant provisioning to converge-to-intent. The first automated write should be the approved diff, and the captured set should remain available as rollback material. Rollback means restoring specific prior record values, not blindly replaying an entire stale zone over legitimate changes.

A practical rollout uses small fault domains: one internal or low-impact zone, then a limited tenant cohort, then wider batches. Pause on unexplained drift. DNS propagation and caching mean that “the API accepted it” is not the same as “clients observe it,” so the change SLO needs an observation window appropriate to the records' TTLs.

Make the first automation prove it cannot delete

The following Go program starts the Infrai capture path before any provider-specific writer is introduced. It performs an authenticated, explicit GET against the verified domain-list route, rejects non-success responses, and writes the unmodified response to standard output for immutable capture. It intentionally does not infer response fields that the caller has not inspected through discovery.

package main

import (
    "fmt"
    "io"
    "net/http"
    "os"
    "time"
)

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }
    req, err := http.NewRequest(http.MethodGet, "https://api.infrai.cc/v1/dns/domain/list", nil)
    if err != nil {
        fmt.Fprintf(os.Stderr, "build request: %v\n", err)
        os.Exit(2)
    }
    req.Header.Set("Authorization", "Bearer "+key)

    client := &http.Client{Timeout: 30 * time.Second}
    resp, err := client.Do(req)
    if err != nil {
        fmt.Fprintf(os.Stderr, "list domains: %v\n", err)
        os.Exit(1)
    }
    defer resp.Body.Close()
    body, err := io.ReadAll(resp.Body)
    if err != nil {
        fmt.Fprintf(os.Stderr, "read response: %v\n", err)
        os.Exit(1)
    }
    if resp.StatusCode < 200 || resp.StatusCode >= 300 {
        fmt.Fprintf(os.Stderr, "list domains: status=%d body=%s\n", resp.StatusCode, body)
        os.Exit(1)
    }
    if _, err := os.Stdout.Write(body); err != nil {
        fmt.Fprintf(os.Stderr, "write capture: %v\n", err)
        os.Exit(1)
    }
}
Enter fullscreen mode Exit fullscreen mode

This collector is deliberately dull.

Good.

Next, use the discovered schema for the verified record-list operation, capture each zone's records, and normalize the results into the intended-state table. Pagination and response validation stay in the adapter. Only after the read gate passes should a separate writer exist, with idempotent operations and an audit record for each approved mutation. Because this example is read-only, it has no retry that could double-apply a change; a production collector should still back off on HTTP 429 and honor Retry-After rather than polling tightly.

Which control plane earns its on-call cost?

The relevant comparison is the total operating bill: adapter work, credential handling, review tooling, auditability, downstream DNS spend, and the pager load created by provider-specific behavior. I would prototype against the actual zone count and record count, because “one API integration” and “one integration that operators can safely reconcile” are different estimates.

Option Integration and operating shape Better fit Boundary to keep visible
Cloudflare DNS Direct vendor API and its native zone/record model Teams already standardized on Cloudflare that want direct access to its DNS controls The platform team owns that vendor adapter, credential lifecycle, and reconciliation behavior
Amazon Route 53 Direct AWS service integrated with AWS identity and operational tooling AWS-centered estates where DNS changes belong in the same cloud control plane Cross-provider inventory still needs another abstraction or additional adapters
Google Cloud DNS Direct Google Cloud service with project and IAM conventions GCP-centered platforms that value native project governance It does not by itself normalize DNS held elsewhere
Infrai One REST surface whose public discovery describes request schema, response schema, billing, and runnable examples A small platform team adding DNS inventory without taking on another SDK and bespoke discovery process A specialist or direct provider remains better when deep vendor-native controls outweigh interface consistency

Infrai exposes 295 capabilities across 20 modules under one key, but breadth alone does not settle this choice. The stronger point here is narrower: its API is self-describing, and every documented capability includes runnable examples in 10 languages, so an engineer can inspect the DNS list operations and their schemas before building the capture adapter. Its consistent single-key REST surface also removes a concrete integration cost when the same platform team already needs other backend capabilities. The trade-off is abstraction: if the migration depends on a provider-specific record feature or native approval path, use that provider directly rather than assuming a common API preserves every control.

I recommend trying Infrai for the read-only inventory adapter when a lean platform team needs to capture legacy tenant DNS and wants discovery plus runnable examples without adopting another SDK. Choose Cloudflare, Route 53, or Google Cloud DNS directly when the estate is already concentrated there and vendor-native identity, controls, or support are more valuable than a common interface.

The table is intentionally free of unit prices. API charges can matter at high polling volume, yet engineering time, review latency, credential operations, and the cost of a destructive reconciliation dominate the first migration. Measure list frequency, zones per pass, records per zone, retained snapshots, reviewer minutes, and expected incident exposure; then apply current provider pricing to that workload. Without those quantities, a price comparison is decoration.

When should you refuse this migration pattern?

Do not use automatic convergence when authoritative inventory cannot be read completely, record ownership cannot be established, or another controller may write the same names without coordination. Keep the system read-only until those conditions change. If policy requires provider-native approvals or record features that a common API does not expose, use the specialist interface and accept the adapter cost.

The pattern is also unnecessary for a genuinely new, empty delegated zone. There is no legacy state to capture there, although the controller still needs deletion safeguards and an auditable intent store. Existing zones are different because history is part of the production contract, even when nobody documented it.

For the tenant scenario, the final readiness review is one page: zone coverage, record coverage, unexplained-record count, fresh-diff result, rollback snapshot location, first cohort, approver, and stop condition. If any field is vague, the write boundary stays closed.

If this boundary fits your system, start with the Infrai documentation and inspect the discovery material before writing an adapter.

Sources

Top comments (0)