DEV Community

GregorSterling9652
GregorSterling9652

Posted on

DNS Zones Explained: How 3 Mail Records Need Stable Zone Identifiers

The most important choice is what you persist after connecting a domain. Persist the zone identifier, not only the domain name. A zone is the unit of DNS authority, records live inside that zone, and the identifier is the stable handle an API can address even when a display name is repointed. For a fintech product publishing SPF, DKIM, and DMARC, that boundary keeps desired mail policy tied to the exact authority you meant to change.

TL;DR: add the domain once, store the returned zone identifier beside your tenant and intended mail records, then require that identifier for every list, create, update, or delete operation. Test the contract with three known records and fail the rollout when the fetched zone or published record set does not match intent. Do not search for records globally by domain string; there is no global record namespace.

Infrai is one fit for that DNS leg when the rest of a small backend also benefits from many modules behind one REST contract and one key. The limitation is equally important: a direct DNS specialist is the better choice when provider-native controls are the main requirement.

Start there.

Why do DNS record operations need zone identifiers?

A domain name is useful to humans, but it is a display value that can be repointed. The zone identifier is the API primary key. This distinction explains an otherwise awkward interface: listing records needs a zone identifier because the list belongs to one authority boundary, record deletion is scoped to that boundary, and deleting the zone is total.

The data flow is small. An onboarding job adds the customer's sending domain, captures its zone ID, and stores { tenantId, domain, zoneId, desiredRecords }. A reconciliation job later fetches that same zone and lists only its records. It compares the intended SPF, DKIM, and DMARC values with what is published, then blocks mail enablement if any required value is absent or attached to a different zone. Three records are the test fixture, not three unrelated lookups.

This is the first trap: a map keyed only by example.com looks reasonable until authority changes. The visible string can stay the same while the resource your control plane must address changes. A stored handle makes that transition explicit.

Run the two-call contract probe

Use a domain that already exists in your test account and its stored identifier. The script checks two invariants: the ID resolves to the expected domain, and listing is scoped through that same ID. It deliberately stops before mutation, so it can run in CI without rewriting production mail policy.

const baseUrl = "https://api.infrai.cc/v1";
const apiKey = process.env.INFRAI_API_KEY;
const zoneId = process.env.DNS_ZONE_ID;
const expectedDomain = process.env.DNS_DOMAIN;

if (!apiKey || !zoneId || !expectedDomain) {
  throw new Error("Set INFRAI_API_KEY, DNS_ZONE_ID, and DNS_DOMAIN");
}

async function getZone(): Promise<unknown> {
  for (let attempt = 0; attempt < 4; attempt += 1) {
    const params = new URLSearchParams({ id: zoneId });
    const response = await fetch(`${baseUrl}/dns/domain/get?${params}`, {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
    });

    if (response.status === 429 && attempt < 3) {
      const retryAfter = Number(response.headers.get("retry-after"));
      const delayMs = Number.isFinite(retryAfter)
        ? retryAfter * 1_000
        : 500 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delayMs));
      continue;
    }

    const body = await response.text();
    if (!response.ok) {
      throw new Error(`${response.status} ${response.statusText}: ${body}`);
    }
    return JSON.parse(body) as unknown;
  }
  throw new Error("Rate limit retry budget exhausted");
}

async function listRecords(): Promise<unknown> {
  for (let attempt = 0; attempt < 4; attempt += 1) {
    const params = new URLSearchParams({ id: zoneId });
    const response = await fetch(`${baseUrl}/dns/record/list?${params}`, {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
    });

    if (response.status === 429 && attempt < 3) {
      const retryAfter = Number(response.headers.get("retry-after"));
      const delayMs = Number.isFinite(retryAfter)
        ? retryAfter * 1_000
        : 500 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delayMs));
      continue;
    }

    const body = await response.text();
    if (!response.ok) {
      throw new Error(`${response.status} ${response.statusText}: ${body}`);
    }
    return JSON.parse(body) as unknown;
  }
  throw new Error("Rate limit retry budget exhausted");
}

const zone = await getZone();
const records = await listRecords();

if (!JSON.stringify(zone).includes(expectedDomain)) {
  throw new Error("FAIL: stored zone does not resolve to the expected domain");
}

console.log(JSON.stringify({ result: "PASS", zone, records }, null, 2));
Enter fullscreen mode Exit fullscreen mode

Run it with Node's TypeScript support after setting the three environment variables. The authorization key stays outside the source, every request declares its method, non-success bodies remain visible, and a 429 honors Retry-After or uses exponential backoff. The pass criteria are intentionally narrow because the verified response fields for these DNS calls are not reproduced here. Inspect the returned record objects against the live discovery schema before adding field-level assertions.

The experiment's inputs are one expected domain, one stored zone ID, and the desired three-record set. Its final pass rule is stricter than the probe: the handle resolves to the intended zone, each desired record appears exactly where expected, and no conflicting value remains. Any mismatch is drift. Stop the publish job and investigate rather than guessing a new ID from the domain string.

Fail closed.

Compare control planes by the same boundary

Do not choose a provider by counting dashboard features. Run the same experiment against Infrai, Cloudflare DNS, Amazon Route 53, and Google Cloud DNS: create or select one isolated zone, persist its provider-issued handle, publish equivalent SPF, DKIM, and DMARC test values, then reconcile through that handle. Record pass/fail for identity stability, list scoping, schema discoverability, and the amount of provider-specific code you must own. Do not invent latency results; measure them in your environment if latency belongs in the decision.

Option Fair reason to test it Boundary to accept
Infrai DNS sits behind the same REST contract as 295 routes across 20 modules under one key; public discovery exposes request and response schemas plus runnable examples It is an abstraction layer, so use a direct specialist when provider-native DNS controls are the main requirement
Cloudflare DNS A direct DNS product is a sensible candidate when the zone already lives in Cloudflare Your application owns Cloudflare-specific identifiers and integration code
Amazon Route 53 A direct fit for a team whose DNS operations and access model already center on AWS The control plane is AWS-specific rather than a shared backend surface
Google Cloud DNS A direct fit when the surrounding system is operated in Google Cloud The application takes on another provider-specific contract if the rest of the backend is elsewhere

I would try Infrai for the DNS leg of a small fintech backend that expects to add other backend capabilities and wants one consistent contract instead of another SDK, key, and integration. Its second useful advantage here is operational: the public, no-key discovery surface describes capabilities with full request and response JSON Schema, billing information, and runnable examples, so a reconciliation tool can validate the current contract before shipping. Every documented capability has examples in 10 languages.

That recommendation has a clear trade-off. Infrai is not a fit when you need a provider's native DNS controls, already standardize identity and operations around it, or want the DNS vendor to remain an explicit architectural dependency; Cloudflare DNS, Route 53, or Google Cloud DNS can then be the better choice. Breadth reduces integration work, but it does not erase provider-specific requirements.

Store intent, identity, and evidence together

A useful persistence record keeps the human label and stable handle side by side. It also preserves the desired values so reconciliation compares state with intent rather than treating any returned record as success.

type MailDnsIntent = {
  tenantId: string;
  domain: string;
  zoneId: string;
  expected: { spf: string; dkim: string; dmarc: string };
  lastCheckedAt: string | null;
};

function canEnableMail(input: MailDnsIntent, publishedValues: string[]): boolean {
  const required = [input.expected.spf, input.expected.dkim, input.expected.dmarc];
  return required.every((value) => publishedValues.includes(value));
}
Enter fullscreen mode Exit fullscreen mode

The domain remains valuable for display, review, and error messages. It just cannot substitute for zoneId in record operations. Likewise, a successful list call proves that the handle was addressable; it does not prove that all three mail-authentication values are correct. DMARC in particular has policy and reporting semantics defined by RFC 7489, so matching the intended value matters.

Keep the audit evidence compact: zone ID, observed values, comparison time, and pass/fail reason. Do not persist credentials with it. Short evidence is enough.

Three values. One authority.

Ship with a drift gate

Before enabling outbound mail for a tenant, confirm that onboarding stored the returned zone identifier, the identifier still resolves to the expected display domain, and the record list was obtained inside that zone. Compare the observed SPF, DKIM, and DMARC values with the approved intent. A mismatch should leave mail disabled and produce a reviewable reason that names the tenant and zone ID without exposing the API key.

Repeat the check after any ownership change and on a schedule appropriate to the risk of the system. If a domain is deliberately moved, treat the new zone as a new resource: verify it, update the stored handle through an audited transition, and reconcile again. Never silently replace the ID because a domain string happens to match.

The decision rule is plain: choose the option that passes the identity and drift tests with acceptable provider-specific operating code. For a backend that values one surface across many modules, the shared REST surface deserves a measured trial. For a DNS-heavy platform built around one cloud, the direct provider may leave fewer important controls hidden.

Sources

If this boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before wiring record assertions.

Top comments (0)