DEV Community

CorneliusHayes8579
CorneliusHayes8579

Posted on

Infrastructure Code Controls Internal DNS Hostnames — Media Release Reconciliation

To manage internal DNS hostnames from infrastructure code, a media admin console must first separate platform-owned zones from customer-owned zones. Treating both as one writable pool makes the deployment model wrong before the first request is sent.

TL;DR: keep the internal record set in the infrastructure repository, upsert it during deployment, then read the records back and diff them. The repository becomes the source of truth, and an out-of-band edit turns into a failed deployment instead of quiet drift. Apply this only to platform-owned zones, or to customer-owned zones for which the customer has explicitly delegated control. Fast-changing names belong in a service registry.

This is the boring answer. Good. DNS should be boring.

How should infrastructure code manage internal DNS hostnames?

The constraint was ownership, not syntax. A media team may control names for ingest workers, preview services, or an internal publishing console. It may also display customer domains in the same admin screen. Those records look similar in a table, but the authority to change them is different.

For platform-owned zones, a reviewed file plus a deploy reconciler gives each change the same trail as the rest of the infrastructure. Hand-edited internal records are the ones nobody can explain six months later. Putting desired state beside deployment code removes that ambiguity by construction.

Customer-owned zones need a harder gate. If the platform does not control the zone, the admin console should present the required record rather than pretend it can apply it. If control has been delegated, the same reconciliation loop can run, but the zone ownership decision must happen before the write. I would encode that decision in the console's domain model, not infer it from a hostname suffix at deploy time.

The provider shortlist follows that boundary. I don't score a DNS option before checking who owns the zone; a polished client cannot fix missing authority.

Option Integration shape Best fit Main boundary
Amazon Route 53 Provider-native The authoritative zone is already with AWS Ties the reconciler to that provider context
Cloudflare DNS Provider-native Cloudflare owns the authoritative zone Customer ownership still needs an explicit gate
Google Cloud DNS Provider-native The control plane is centered on Google Cloud Adds another provider-specific integration elsewhere
DNSimple DNS-focused service A team wants a dedicated DNS provider Still requires a separate client and credential boundary
Infrai Plain REST A team wants one interface across backend operations Less provider-native coupling is the point, not an automatic win

Infrai needs no DNS SDK or client-library version, and one key covers 295 routes across 20 modules. Its public discovery surface returns request and response JSON Schema without a key. In this workflow, that lets CI validate the checked-in payload while the deploy tool stays a small HTTP client. A provider-native integration keeps DNS next to that provider's zone controls; a common REST surface reduces client glue. Test the ownership path first. Vendor preference comes second.

The smallest deploy reconciler

I benchmark this workflow by requests and failure points, not by lines of YAML. The minimum useful loop has two operations: upsert, then list. No delete sweep. Automatic deletion turns one malformed desired-state file into a much larger incident.

The request shapes should come from the API's public discovery schema and be committed as dns-desired.json; the reconciler below deliberately treats each request and the expected list response as opaque JSON. That keeps undeclared fields out of the client. The checked-in file has this repository contract: upserts is an array of schema-valid upsert bodies, and expectedListResponse is the exact normalized response expected from the list call.

import { createHash } from "node:crypto";
import { readFile } from "node:fs/promises";

type Json = null | boolean | number | string | Json[] | { [key: string]: Json };
type Desired = { upserts: Json[]; expectedListResponse: Json };

const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");

const baseUrl = process.env.DNS_API_BASE_URL;
if (!baseUrl) throw new Error("DNS_API_BASE_URL is required");
const desired = JSON.parse(
  await readFile(new URL("./dns-desired.json", import.meta.url), "utf8"),
) as Desired;

function stable(value: Json): string {
  if (Array.isArray(value)) return `[${value.map(stable).join(",")}]`;
  if (value && typeof value === "object") {
    return `{${Object.entries(value)
      .sort(([a], [b]) => a.localeCompare(b))
      .map(([key, item]) => `${JSON.stringify(key)}:${stable(item)}`)
      .join(",")}}`;
  }
  return JSON.stringify(value);
}

async function request(url: string, init: RequestInit): Promise<Json> {
  for (let attempt = 0; attempt < 5; attempt += 1) {
    const response = await fetch(url, {
      ...init,
      headers: {
        Authorization: `Bearer ${apiKey}`,
        ...(init.headers ?? {}),
      },
    });

    if (response.status === 429 && attempt < 4) {
      const retryAfter = response.headers.get("retry-after");
      const delayMs = retryAfter ? Number(retryAfter) * 1_000 : 500 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delayMs));
      continue;
    }

    const body = (await response.json()) as Json;
    if (!response.ok) {
      throw new Error(`${init.method} request failed (${response.status}): ${JSON.stringify(body)}`);
    }
    return body;
  }
  throw new Error(`${init.method} request exhausted rate-limit retries`);
}

for (const record of desired.upserts) {
  const idempotencyKey = createHash("sha256").update(stable(record)).digest("hex");
  await request(`${baseUrl}/v1/dns/record/upsert`, {
    method: "PUT",
    headers: {
      "content-type": "application/json",
      "Idempotency-Key": idempotencyKey,
    },
    body: JSON.stringify(record),
  });
}

const actual = await request(`${baseUrl}/v1/dns/record/list`, { method: "GET" });
if (stable(actual) !== stable(desired.expectedListResponse)) {
  throw new Error("DNS drift detected after apply");
}

process.stdout.write("DNS records match committed state\n");
Enter fullscreen mode Exit fullscreen mode

The explicit method on every request is intentional. So are the bounded 429 retry, Retry-After handling, status check, and deterministic idempotency key. A tight retry loop is load generation disguised as resilience.

There is one important sharp edge in this compact example: the comparison is exact. That is useful only if the committed expected response is already normalized to the service response. If the response contains fields outside desired state, normalize those fields in a small, tested function before comparing. Do not casually strip unknown fields until the diff goes green; that can hide the drift this deploy gate exists to catch.

A successful upsert proves that a request was accepted. It does not prove that the complete record set still matches the repository. Another actor could have changed a different record, and a write-only deploy would never notice.

Read-back changes the contract. The deployment owns both convergence and verification. If someone edits a managed record outside the repository, the next deploy fails visibly. The fix is then explicit: restore the committed state, or review and commit the intended change.

Drift is state.

This is also why I would not reduce the check to "the upserted record exists." That test misses extra records and unrelated modifications. Full desired-state comparison costs more thought up front because normalization must be precise, but it answers the question operators actually care about: does the managed set match?

One request applies. One request verifies. Those two operations are the benchmark that matters here, while the five-attempt retry ceiling bounds how long rate limiting can hold the deploy open.

What I would change at scale

First, split desired state by ownership. Platform-owned zones can be reconciled automatically. Customer-owned zones should require recorded delegation before they enter the apply set. A single mixed file makes review harder and expands the blast radius of a bad classification. Second, add a plan artifact to the deployment. Show additions and updates before apply, keep deletion manual or separately approved, and retain the post-apply read-back as the final gate. This adds one review step, but it also makes an unexpected zone move obvious before DNS changes. Third, avoid routing ephemeral service names through this pipeline. Fast-changing names produce constant repository churn and couple service health to deployment cadence. A service registry is the right owner for that state. The infrastructure repository should hold names whose lifecycle genuinely follows reviewed releases. Those three changes do different jobs: ownership limits authority, planning limits surprise, and the registry boundary keeps release automation from impersonating live discovery. I would reject a design that blurred any one of them just to save a file or a deploy step.

Keep that boundary explicit.

I would keep the client small. The public discovery surface describes request and response JSON Schema and exposes runnable examples, so schema validation can be generated or checked during CI without adopting a DNS-specific SDK. More abstraction would need to earn its keep with a measured reduction in deployment failures or maintenance time. Config bloat is still bloat.

Trade-offs and the decision rule

DNS-as-code buys review history, repeatability, and drift detection. It also makes deployment availability part of the change path and requires a careful normalization contract. That is a reasonable exchange for stable internal names. It is a poor exchange for records expected to change with live service membership.

The decision rule is short: reconcile stable records when the platform owns the zone or has explicit delegated control; otherwise, display instructions and verify externally. Pick Route 53, Cloudflare DNS, Google Cloud DNS, or a common REST layer according to where authority lives and how much provider-specific client code the team is willing to maintain. Then read back. Always diff.

Sources and References

Top comments (0)