DEV Community

MidnightEcho794261
MidnightEcho794261

Posted on

How to Use DNS Verification in Node.js (What Domain Control Actually Proves)

TL;DR: A DNS verification token proves that someone could arrange for one value to be published at one DNS name and that your resolver observed it. It does not prove that the person works for the customer, may edit every record in the zone, or should retain access tomorrow. For an internal customer-support console, verify the domain and authorize the operator as two separate decisions. Require both decisions again when a record mutation is requested.

That separation is the result that matters. The tempting implementation is a single verified: true column followed by an enabled Edit DNS button. I would not ship that model: one old observation becomes a permanent permission shared by every staff role, even though the observation said nothing about those roles.

The evaluation constraint is customer ownership. A platform-owned zone can follow the platform's own access policy. A customer-owned zone crosses a trust boundary, so a support agent's authenticated session, an observed TXT token, and approval for the proposed change must remain independent evidence.

What does domain verification actually prove about DNS control?

Suppose Acme Support asks to connect help.acme.example. The console generates an unpredictable token and asks an administrator to publish it at _support-verify.help.acme.example. When the authoritative DNS path returns that token and the application observes it, the narrow conclusion is useful: the claimant could cause that value to appear under the requested name at that time.

Stop there.

The result does not identify who changed DNS. It does not reveal whether they used a registrar, a delegated DNS host, infrastructure automation, or an internal request to another team. It also does not grant a support agent permission to replace MX, TXT, CNAME, or any other record. DNS answers carry record data, TTLs, and protocol semantics; they do not carry employment roles or approval rules.

This distinction matters especially for mail. DMARC publishes policy in DNS and lets a domain owner express requested handling for messages that fail authentication checks. Observing a separate verification token is not authorization to alter that policy. A console that treats the two as equivalent turns a possession signal into a broad administrative capability.

I use this compact evidence model:

Evidence What it supports What it cannot support by itself
Observed challenge token Control of a specific DNS name at an observation time Staff identity or permission to mutate records
Authenticated staff session Identity known to the internal console Customer consent or DNS control
Customer-scoped role Allowed class of actions for that tenant Current control of the DNS name
Change approval Intent for a specific mutation Unrelated or future mutations

The failed/simple approach collapses all four rows into one Boolean. The chosen approach stores them separately because they expire, get revoked, and answer different questions.

Model customer-owned and platform-owned zones differently

Ownership determines the boundary before any code runs. For a platform-owned zone, the platform controls the authoritative configuration and can define which internal service may write which record set. Domain verification adds little because control is already established through that infrastructure boundary. Authorization and change review still matter.

For a customer-owned zone, the console should assume it has no write authority. It can issue a challenge, observe public DNS, and prepare an exact change request. A customer may then apply the change through their own system, or explicitly delegate a constrained automation path. Verification alone must never switch the zone into platform-owned mode.

Represent that decision directly instead of inferring it from a successful lookup:

type ZoneCustody =
  | { kind: "platform"; zoneId: string }
  | { kind: "customer"; tenantId: string };

type Verification = {
  tenantId: string;
  fqdn: string;
  tokenHash: string;
  observedAt: Date | null;
  expiresAt: Date;
};

type ChangeIntent = {
  tenantId: string;
  fqdn: string;
  recordType: "CNAME" | "TXT";
  value: string;
  requestedBy: string;
  approvedBy: string | null;
};
Enter fullscreen mode Exit fullscreen mode

The union prevents a useful category error. A customer value has no platform zone identifier, so generic application code cannot accidentally route it into the platform's writer. This is a small type-level guard, not a complete security boundary, but it makes the dangerous state harder to express.

Implement the narrow verification check

The verifier needs an exact name, an exact token, an expiry, and a fresh DNS observation. It should not accept a token found at a parent name, a sibling, or an arbitrary TXT record elsewhere in the zone. Normalize the requested domain before constructing the challenge name, and store only a hash of the token after issuance.

Here is the focused Node.js example. It checks the public TXT answer and returns evidence; it does not return permission.

import { resolveTxt } from "node:dns/promises";
import { createHash, randomBytes, timingSafeEqual } from "node:crypto";
import { domainToASCII } from "node:url";

function canonicalDomain(input: string): string {
  const ascii = domainToASCII(input.trim().replace(/\.$/, "").toLowerCase());
  if (!ascii || ascii.length > 253) throw new Error("Invalid domain");
  return ascii;
}

function digest(value: string): Buffer {
  return createHash("sha256").update(value, "utf8").digest();
}

export function issueChallenge(domain: string) {
  const fqdn = `_support-verify.${canonicalDomain(domain)}`;
  const token = randomBytes(32).toString("base64url");
  return { fqdn, token, tokenHash: digest(token).toString("hex") };
}

export async function observeChallenge(
  fqdn: string,
  expectedHashHex: string,
  expiresAt: Date,
): Promise<{ observedAt: Date }> {
  if (Date.now() >= expiresAt.getTime()) throw new Error("Challenge expired");

  const answers = await resolveTxt(fqdn);
  const expected = Buffer.from(expectedHashHex, "hex");
  const matched = answers
    .map(parts => parts.join(""))
    .map(digest)
    .some(actual =>
      actual.length === expected.length && timingSafeEqual(actual, expected),
    );

  if (!matched) throw new Error("Challenge token not observed");
  return { observedAt: new Date() };
}
Enter fullscreen mode Exit fullscreen mode

TXT data can be returned as multiple character strings, which is why the example joins each record's parts before comparison. DNS caching also means a newly added or removed answer may not be visible to every resolver immediately. Treat ENOTFOUND, ENODATA, and timeouts as inconclusive observations that can be retried with bounded backoff, not as evidence that the claimant lacks control.

Do not log the raw token. Log the challenge identifier, normalized name, resolver outcome, attempt time, and final state. That is enough to investigate propagation and retry behavior without turning logs into a second credential store.

Require authorization again at mutation time

A valid observation should advance a verification record, not mint an unrestricted DNS session. At the mutation boundary, check tenant membership, role, zone custody, verification freshness, the exact requested name, allowed record type, and approval state. Re-check on every request because roles and approvals can change after verification.

type Actor = { id: string; tenantId: string; roles: string[] };

export function authorizeChange(input: {
  actor: Actor;
  custody: ZoneCustody;
  verification: Verification;
  change: ChangeIntent;
  now: Date;
}): void {
  const { actor, custody, verification, change, now } = input;
  if (actor.tenantId !== change.tenantId) throw new Error("Wrong tenant");
  if (!actor.roles.includes("dns-change-operator")) throw new Error("Role denied");
  if (verification.tenantId !== change.tenantId) throw new Error("Evidence mismatch");
  if (!verification.observedAt || verification.expiresAt <= now) {
    throw new Error("Fresh verification required");
  }
  if (verification.fqdn !== `_support-verify.${change.fqdn}`) {
    throw new Error("Name outside verified scope");
  }
  if (!change.approvedBy || change.approvedBy === change.requestedBy) {
    throw new Error("Independent approval required");
  }
  if (custody.kind === "customer") {
    throw new Error("Return change instructions to the customer");
  }
}
Enter fullscreen mode Exit fullscreen mode

The permitted record types and approval policy are application choices, not DNS facts. In this customer-support console, the policy deliberately rejects direct writes to customer-owned zones. A different organization might accept an explicit delegation, but that delegation needs its own scoped credential and audit trail. The TXT challenge is still not the credential.

This design costs a few extra reads and creates five states for the UI: pending observation, verified, expired, awaiting approval, and authorized. The trade-off is more policy code and more ways for an operator to get stuck between states. I accept that overhead because it keeps the cheap path cheap. DNS lookups happen during verification; routine console page loads read stored evidence, while actual mutations pay for the current authorization checks. There is no reason to poll DNS on every render.

The limitation is important: this split is a poor fit for a platform that needs immediate, fully automated writes to a customer-owned zone but has no explicit delegation from that customer. No verification design can manufacture the missing authority. In that case, keep the workflow customer-applied, or establish a separately revocable and narrowly scoped delegation before enabling automation.

Test the boundaries, not just the happy path. Useful cases include a token at the parent rather than the exact challenge name, two TXT records with only one match, an expired challenge, Unicode input normalized to its ASCII form, a staff role revoked after observation, cross-tenant evidence, self-approval, and a customer-owned zone that must produce instructions instead of a write. Use an authoritative test zone or a controllable DNS server for integration tests; mock resolver results are enough for the pure state transitions.

What should you measure before adopting this split?

Measure observation latency from challenge issuance to the first matching answer, retry counts by resolver outcome, challenge expiry rates, and authorization denials by reason. Also track how often operators request a new challenge after a role or ownership change. Those signals reveal confusing UI and slow DNS propagation without weakening the boundary.

For operations, emit an immutable event for challenge issuance, successful observation, expiry, approval, denial, and mutation result. Include tenant ID, actor ID, normalized FQDN, record type, evidence ID, and timestamps; exclude raw tokens and sensitive record values where they are unnecessary. Alert on repeated cross-tenant denials and unusual approval patterns, not on every ordinary DNS miss.

The decision rule is plain: if the platform owns the zone, authorize writes through the platform's infrastructure policy. If the customer owns it, verification may establish narrow control evidence, but mutation requires separate, explicit delegation and staff authorization. A green verification badge is evidence, not authority.

Further reading

Top comments (0)