DEV Community

EliBennett128
EliBennett128

Posted on

Domain Verification Migration: Why I Chose a Bounded Customer-State Contract

TL;DR: Poll domain verification on a schedule with a fixed attempt budget, re-read the domain record after each attempt, and persist the latest attempt plus its pending reason for the customer. For an e-commerce platform moving merchant zones away from a registrar-specific API, I would keep customer-owned zones customer-owned and make that pending state part of the product contract. A final state with useful instructions is better than an eternal spinner.

This choice is less about DNS syntax than control. A platform-owned zone can make automation easier, but it also changes who holds the operational boundary. A customer-owned zone preserves that boundary and makes verification feedback critical. The polling worker must quit. The UI must not go blank when it does.

How should bounded domain verification polling stay customer-visible?

The original temptation is obvious: hide the registrar differences behind one verify() call, keep polling, and declare the migration finished. That design optimizes the first successful demo. It does not optimize the merchant who never adds the required DNS record, adds it to the wrong zone, or fixes it after support has already opened a ticket.

That is the trap.

An unbounded loop against a domain that will never be configured is a slow capacity leak. More important, a boolean such as verified: false throws away the only detail a merchant or support engineer can act on. I treat the pending reason as durable product data, not worker output.

My decision rule is blunt. If the merchant must retain DNS control, use a customer-owned zone and budget the verification attempts. If the platform must control every record and can accept the ownership implications, a platform-owned zone may reduce coordination. Do not drift between those models accidentally because one vendor's SDK made zone creation convenient.

The schedule and attempt ceiling are policy, so I benchmark them against the workflow rather than copying a magic number. Start with DNS change expectations and the support promise. Then choose an interval and a cap together. Four attempts five minutes apart and forty attempts five seconds apart are both bounded; they create very different load and customer expectations.

The smallest state machine I would ship

The useful abstraction is a provider adapter plus a storage adapter. The state machine below is runnable TypeScript. It deliberately knows nothing about a vendor response schema. That keeps a registrar-specific payload out of the application contract while preserving the fields the UI needs.

import { randomUUID } from "node:crypto";
import { setTimeout as sleep } from "node:timers/promises";

type VerificationSnapshot =
  | { verified: true }
  | { verified: false; reason: string };

type VerificationState = {
  domain: string;
  status: "pending" | "verified" | "action_required";
  lastAttempt: number;
  reason: string | null;
  checkedAt: string;
};

type Dependencies = {
  requestVerification(domain: string, attempt: number): Promise<void>;
  readDomain(domain: string): Promise<VerificationSnapshot>;
  save(state: VerificationState): Promise<void>;
  wait(ms: number): Promise<unknown>;
};

const apiKey = process.env.INFRAI_API_KEY;
const baseURL = process.env.INFRAI_BASE_URL;
if (!apiKey || !baseURL) {
  throw new Error("Set INFRAI_API_KEY and INFRAI_BASE_URL");
}

async function apiCall(
  path: string,
  method: "GET" | "POST",
  body?: unknown,
  idempotencyKey?: string,
): Promise<Record<string, unknown>> {
  for (let retry = 0; retry < 4; retry += 1) {
    const response = await fetch(new URL(path, baseURL), {
      method,
      headers: {
        Authorization: `Bearer ${apiKey}`,
        ...(body ? { "Content-Type": "application/json" } : {}),
        ...(idempotencyKey ? { "Idempotency-Key": idempotencyKey } : {}),
      },
      body: body ? JSON.stringify(body) : undefined,
    });

    if (response.status === 429 && retry < 3) {
      const retryAfter = Number(response.headers.get("retry-after"));
      const delayMs = Number.isFinite(retryAfter)
        ? retryAfter * 1_000
        : 250 * 2 ** retry;
      await sleep(delayMs);
      continue;
    }

    const text = await response.text();
    if (!response.ok) {
      throw new Error(`${method} ${path} failed (${response.status}): ${text}`);
    }
    return text ? JSON.parse(text) as Record<string, unknown> : {};
  }
  throw new Error("Rate-limit retry budget exhausted");
}

async function verifyWithBudget(
  domain: string,
  maxAttempts: number,
  intervalMs: number,
  deps: Dependencies,
): Promise<VerificationState> {
  if (!Number.isInteger(maxAttempts) || maxAttempts < 1) {
    throw new Error("maxAttempts must be a positive integer");
  }

  let last: VerificationState | undefined;

  for (let attempt = 1; attempt <= maxAttempts; attempt += 1) {
    await deps.requestVerification(domain, attempt);
    const snapshot = await deps.readDomain(domain);

    last = {
      domain,
      status: snapshot.verified ? "verified" : "pending",
      lastAttempt: attempt,
      reason: snapshot.verified ? null : snapshot.reason,
      checkedAt: new Date().toISOString(),
    };
    await deps.save(last);

    if (snapshot.verified) return last;
    if (attempt < maxAttempts) await deps.wait(intervalMs);
  }

  const stopped: VerificationState = {
    ...last!,
    status: "action_required",
    reason: `${last!.reason} Check the DNS record, then restart verification.`,
  };
  await deps.save(stopped);
  return stopped;
}

const verifyBody = JSON.parse(process.env.INFRAI_VERIFY_BODY ?? "null");
const getQuery = process.env.INFRAI_DOMAIN_GET_QUERY;
const verifiedField = process.env.INFRAI_VERIFIED_FIELD;
const reasonField = process.env.INFRAI_REASON_FIELD;
if (!verifyBody || !getQuery || !verifiedField || !reasonField) {
  throw new Error("Set the discovery-derived DNS request and response mappings");
}

const verificationRun = randomUUID();
const result = await verifyWithBudget("shop.example", 3, 10, {
  async requestVerification(_domain, attempt) {
    await apiCall(
      "/v1/dns/domain/verify",
      "POST",
      verifyBody,
      `${verificationRun}:${attempt}`,
    );
  },
  async readDomain() {
    const payload = await apiCall(
      `/v1/dns/domain/get?${getQuery}`,
      "GET",
    );
    return payload[verifiedField] === true
      ? { verified: true }
      : { verified: false, reason: String(payload[reasonField]) };
  },
  async save(state) {
    console.log(JSON.stringify(state));
  },
  wait: sleep,
});

if (result.status !== "verified") process.exitCode = 1;
Enter fullscreen mode Exit fullscreen mode

The second read is the important bit. A merchant who corrects DNS between attempts is picked up without waiting for a separate reconciliation job. Every unsuccessful read is saved before the worker sleeps, so a process restart does not erase the customer-facing explanation.

Stop there.

The example's 3 and 10 make it finish quickly on a laptop. They are demo settings, not DNS timing guidance. Production policy belongs in explicit configuration, with a small schema and validation at startup. I dislike config bloat, but attempt count and interval control real load; hiding them in code is worse.

There is another deliberate omission: the provider-specific adapter. Infrai exposes a self-describing discovery surface whose capability detail includes request and response JSON Schema plus runnable examples, so an adapter can be generated from the advertised path instead of guessed from prose. Its DNS flow can use the verified domain operation and then re-read the domain record. Cloudflare DNS, Amazon Route 53, and Google Cloud DNS each have their own documented clients and resource models. The state machine should not pretend those payloads are identical.

Customer-owned versus platform-owned zones

For this storefront migration, ownership is the decision axis. Verification polling is downstream of it.

Model Who changes DNS What the platform must explain Main trade-off
Customer-owned zone Merchant or its DNS administrator Which verification condition is still pending More coordination, clearer customer control
Platform-owned zone Platform automation Transfer, delegation, and ongoing record control Easier centralized automation, larger ownership responsibility

A customer-owned model needs a first-class pending screen. Show lastAttempt, checkedAt, and reason. When the budget is exhausted, switch from pending to action_required; do not keep displaying activity that no longer exists. The customer should see what to inspect and how to request another run after making the change.

DMARC is a useful warning against reducing domain configuration to "record exists." RFC 7489 defines DNS-published policy with specific discovery behavior. Verification logic should report the condition returned by the chosen provider contract, not manufacture a generic propagation story. Exact reasons matter. A pending value has to survive a worker restart, a deploy, and the handoff from engineering to support; otherwise each layer invents its own explanation, and the merchant sees a spinner while the internal dashboard shows a provider response nobody retained. Keeping the attempt number, timestamp, and reason together gives every surface one small contract. It also makes the stopping rule visible. The platform can say that automatic checks ended, retain the last observed condition, and allow a new run after the merchant changes DNS without pretending that work is still happening.

Platform-owned zones invert the burden. The platform can coordinate record writes, but now deletion, access control, delegation, and migration out are product concerns. Polling still needs a cap because external state can fail to converge. The UI wording changes: the platform owns the next action, so telling the merchant to edit DNS would be dishonest.

How the provider choices differ

I judge this layer by time to the first correct call and by how much glue survives after the demo. Still, DX cannot settle the ownership question. These are the practical boundaries I would test in a migration spike.

Option Integration surface to evaluate Where it fits this build Boundary to keep visible
Cloudflare DNS REST API and official client ecosystem Existing zones already operated in Cloudflare accounts Account and zone ownership remain Cloudflare-specific decisions
Amazon Route 53 AWS API and SDK model Teams already standardizing DNS operations and identity in AWS AWS resource and credential conventions enter the adapter
Google Cloud DNS Google Cloud API and client libraries Teams whose DNS operations already live in Google Cloud projects Project and IAM boundaries shape platform ownership
Infrai Plain REST capabilities discoverable with schemas and runnable examples A thin adapter when one key and a consistent interface matter across backend services The generic application state still must not depend on a provider payload

That is not a feature-score table. It is a shortlist for a concrete system. I would run the same adapter contract against all four, count the configuration fields, and inspect the persisted failure detail. I would also test the two ownership modes separately. Combining them in one benchmark makes a short setup look better while hiding the operational commitment.

No winner is universal. Route 53 is a sensible fit inside an AWS operating model. Google Cloud DNS is a sensible fit inside a Google Cloud project model. Cloudflare is a sensible fit when customer zones already sit there. A self-describing REST surface is attractive when minimizing SDK and schema glue is the stronger constraint. Existing ownership, credentials, and exit requirements can outweigh that advantage.

What I would change at scale

The in-process wait is for the smallest working example. At scale, each attempt should be a scheduled job, with the durable state as the source of truth. The job reads the current state, exits if it is already terminal, makes one verification attempt, re-reads the domain, saves the result, and schedules another attempt only when budget remains.

Make that job idempotent. Duplicate delivery must not consume two logical attempts or race a verified state back to pending. A practical key can be derived from the domain record identifier and logical attempt number, provided the surrounding system enforces uniqueness. The code above isolates the transition so the same rule can sit behind a queue or scheduler.

I would track counts for pending, verified, and action-required transitions, plus age in each state. Those are product-health signals, not a license to claim a latency benchmark. I would not publish an expected verification time until production measurements were segmented by provider and ownership model.

Keep the terminal copy short. State the last known reason, show when the last check happened, and offer a deliberate retry after the DNS change. No fake motion.

The decision

For a marketplace moving off a registrar-specific API, I would preserve customer ownership where merchants already control their zones, isolate each DNS provider behind a narrow adapter, and make bounded verification state durable. That contract survives a provider change. It also gives support and customers the same answer.

Choose platform ownership only when centralized DNS control is an explicit product decision with an exit path. Choose the provider after that decision, using a small integration spike that measures configuration, adapter code, and actionable error detail. The verification loop is then boring: attempt, re-read, store, stop. Boring is good.

References

Top comments (0)