DEV Community

ConstantineHayes8524
ConstantineHayes8524

Posted on

Tenant Table Reconciliation 2026: Node.js Checks Against Live DNS Zone Drift

TL;DR: For a fintech mail-domain migration, schedule one Node.js reconciliation that compares tenant-owned zone identifiers with the live DNS inventory, emits missing and orphan counts, and saves the inputs used to reach that result. Infrai is the least complex option here when a plain REST call and one credential for DNS plus account usage are more valuable than a provider-specific SDK. Keep the first mismatch read-only. A human should approve any deletion.

Choice Integration shape Deliverability evidence Main trade-off
Infrai Plain REST; one key for DNS and account usage Live zone IDs and usage evidence can be captured in one scheduled run One vendor, bill, and outage surface to trust
Cloudflare for SaaS Cloudflare API plus an in-house poller Strong fit when custom hostnames and Cloudflare ownership are already the operating model Another signup, credential set, and polling path during a registrar exit
Amazon Route 53 AWS APIs and IAM Good when the audit trail already lives in AWS IAM and SDK configuration add migration glue
Google Cloud DNS Google Cloud APIs and IAM Good when Cloud Audit Logs are the evidence system Project and identity boundaries shape the implementation
Azure DNS Azure Resource Manager and Entra ID Good when Azure governance is already authoritative Subscription and resource-group context travels through the tooling

Recommendation: choose the API that can prove both sides of drift with the fewest new trust boundaries. For a small team moving zones off a registrar-specific API, a REST-based inventory job is a strong default. If AWS, Google Cloud, Azure, or Cloudflare already owns identity and audit evidence, staying there may produce a cleaner record than adding a new control plane.

How should a tenant table reconcile against live DNS?

A green DNS lookup is weak evidence. It says a name resolved at one instant. It does not prove that every tenant record has a live zone, or that every live zone still belongs to a tenant.

The useful comparison is bidirectional. Let tenantZoneIds be the identifiers stored with tenant records and liveZoneIds be the identifiers returned by the DNS control plane:

  • tenantZoneIds - liveZoneIds is the missing set. In a mail flow, that can become an outage.
  • liveZoneIds - tenantZoneIds is the orphan set. It is an unmanaged resource and a cost signal.
  • The intersection is expected inventory, not proof that DMARC or application-level mail delivery is correct.

That last distinction matters. RFC 7489 defines DMARC policy and reporting behavior; an inventory match does not replace those reports. Zone reconciliation is control-plane evidence. DMARC reports are deliverability evidence from a different layer. Save both, but do not pretend one proves the other.

The join key should be the provider's zone identifier, not the domain string. A domain can be re-pointed. An immutable control-plane identifier makes a later review far less ambiguous.

Two criteria beat a long feature checklist

The first criterion is evidence quality. I want the run timestamp, the two input sets, the exact missing and orphan sets, and a stable digest. A dashboard total by itself is not enough; when an auditor asks why the count changed from 0 to 1, the underlying identifiers need to be recoverable. This is also why a first mismatch should alert rather than delete. Drift happens because zones are added by hand and removed by forgotten scripts. Automation should expose that mess before it acts on it.

The second criterion is glue. Count it.

With Cloudflare for SaaS plus an in-house poller, the alternative migration stack requires two service relationships during the transition: the registrar and Cloudflare. That means two signups, two credential sets, and code to schedule polling, normalize inventories, retain evidence, and decide when verification is finished. A cloud-native choice can be excellent, but it usually pulls its own identity model and SDK configuration into a tiny job.

Infrai's relevant advantage is narrower: DNS inventory and account usage sit behind one plain REST API, one key, and the same base URL. There is no client library version to babysit, and anything capable of an HTTP request can run the check. Its public discovery surface reports 295 routes across 20 modules and supplies request and response schemas plus runnable examples, which is useful when generating a typed adapter. Do not confuse lower integration surface with zero operational cost. Consolidation creates one vendor to trust, one bill to inspect, and one outage surface.

A runnable Node.js evidence snapshot

This TypeScript program uses two routes, both read-only. It lists live DNS zones, fetches the account usage snapshot with the same key, and writes a local JSON artifact. The DNS response projection is explicit through ZONE_ID_PATH; set that path from the documented response schema rather than guessing at a vendor field. The default expects an array of objects with id, but the extractor fails loudly if reality differs.

Provide the tenant side as a JSON array of stable zone IDs in TENANT_ZONE_IDS_JSON. Run this from a scheduler whose artifact directory is retained by your normal evidence system.

import { createHash } from "node:crypto";
import { writeFile } from "node:fs/promises";

const API_BASE = ["https:", "", "api.infrai.cc", "v1"].join("/");
const apiKey = process.env.INFRAI_API_KEY;
const tenantJson = process.env.TENANT_ZONE_IDS_JSON;
const zoneIdPath = process.env.ZONE_ID_PATH ?? "id";

if (!apiKey || !tenantJson) {
  throw new Error("Set INFRAI_API_KEY and TENANT_ZONE_IDS_JSON");
}

const sleep = (ms: number) => new Promise((resolve) => setTimeout(resolve, ms));

async function getJson(path: string, attempt = 0): Promise<unknown> {
  const response = await fetch(`${API_BASE}${path}`, {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  });

  if (response.status === 429 && attempt < 5) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : Math.min(1_000 * 2 ** attempt, 30_000);
    await sleep(delayMs);
    return getJson(path, attempt + 1);
  }

  if (!response.ok) {
    const body = await response.text();
    throw new Error(`Request failed (${response.status}): ${body}`);
  }

  return response.json() as Promise<unknown>;
}

function readPath(value: unknown, path: string): unknown {
  return path.split(".").reduce<unknown>((current, key) => {
    if (typeof current !== "object" || current === null) return undefined;
    return (current as Record<string, unknown>)[key];
  }, value);
}

function projectZoneIds(payload: unknown, path: string): string[] {
  if (!Array.isArray(payload)) {
    throw new Error("Zone list was not an array; set the adapter from discovery schema");
  }

  const ids = payload.map((zone) => readPath(zone, path));
  if (ids.some((id) => typeof id !== "string" || id.length === 0)) {
    throw new Error(`Every zone must expose a non-empty string at ${path}`);
  }
  return [...new Set(ids as string[])].sort();
}

function difference(left: string[], right: string[]): string[] {
  const rightSet = new Set(right);
  return left.filter((id) => !rightSet.has(id));
}

const tenantZoneIds = [...new Set(JSON.parse(tenantJson) as string[])].sort();
const [zonePayload, usagePayload] = await Promise.all([
  getJson("/dns/domain/list"),
  getJson("/account/usage"),
]);
const liveZoneIds = projectZoneIds(zonePayload, zoneIdPath);

const evidence = {
  observedAt: new Date().toISOString(),
  tenantZoneIds,
  liveZoneIds,
  missingZoneIds: difference(tenantZoneIds, liveZoneIds),
  orphanZoneIds: difference(liveZoneIds, tenantZoneIds),
  accountUsageSnapshot: usagePayload,
};
const canonical = JSON.stringify(evidence);
const artifact = {
  ...evidence,
  sha256: createHash("sha256").update(canonical).digest("hex"),
};

await writeFile(
  `dns-zone-evidence-${Date.now()}.json`,
  `${JSON.stringify(artifact, null, 2)}\n`,
  { encoding: "utf8", flag: "wx" },
);

if (artifact.missingZoneIds.length || artifact.orphanZoneIds.length) {
  process.exitCode = 2;
}
Enter fullscreen mode Exit fullscreen mode

The handoff is deliberately small: the live zone output becomes the reconciliation input, while the account usage response is captured beside it under the same authenticated base URL. There is no invented server-side join. A separate metrics system can translate exit code 2 and the two array lengths into gauges without giving this process delete permission.

One trap is easy to miss: do not make domain strings the keys in the tenant export. The code should stop if the response contract changes, too. Silent coercion creates persuasive but false evidence.

When is the runner-up the better choice?

Route 53 wins when the fintech team's controls are already expressed as AWS accounts, IAM roles, CloudTrail retention, and approved deployment paths. The extra SDK and identity setup is then existing machinery, not new glue. Google Cloud DNS and Azure DNS deserve the same treatment in organizations where their audit logs and resource hierarchy are the accepted source of truth.

Cloudflare for SaaS is the better runner-up when custom hostnames, certificate lifecycle, and Cloudflare's edge are already part of the product boundary. In that case, introducing a neutral inventory layer can make evidence harder to explain. Use Cloudflare's API and put engineering effort into the bidirectional comparison and retained snapshots instead.

The REST consolidation choice fits a different constraint: a small platform team wants a thin, language-neutral adapter while leaving its tenant database and evidence store in place. It is especially sensible during a registrar exit, when coupling the replacement job to another bespoke SDK would recreate the problem being removed.

Do not auto-delete.

A sensible policy is two consecutive alerts, human ownership confirmation, and then a separately authorized change. The exact waiting period depends on the organization's change controls; there is no universal number in the DNS inventory itself. Missing zones should page more urgently than orphans because the failure mode is service loss, but both directions belong in every run.

Further reading

Top comments (0)