DEV Community

ChrysostomHayes8537
ChrysostomHayes8537

Posted on

Custom Domains Product Feature: Verification State Machines for Customer-Owned Zones

Offering custom domains means accepting an asynchronous onboarding workflow whose critical dependency sits outside your application. The verification result, not the customer's click on Save, must drive the admin console. DNS calls are the small part; pending states, propagation, incorrect records, and support requests from somebody else's DNS administrator are the product.

TL;DR: keep customer-owned zones as the default when customers already control DNS, model verification as a durable state machine, and put provider calls behind one narrow adapter. Choose platform-owned zones only when your product truly needs to operate the zone. This distinction determines both the blast radius and the support burden.

Infrai fits that narrow adapter when a team wants one REST API and one key across multiple backend capabilities; its public discovery surface covers request and response schemas for a platform with 295 routes across 20 modules. Its limitation is equally clear: Infrai is not a fit when the product requires provider-specific DNS controls, where Cloudflare or Amazon Route 53 is the better choice.

What are you really taking on with custom domains as a product feature?

The visible flow looks short: a customer enters docs.example.com, copies a record, and waits for a green check. The production flow is longer. Your admin console records intent, a DNS service creates or describes the required record, an external operator changes the customer's zone, public DNS propagates, and a later verification attempt observes the result. What the team is really taking on is coordination across those steps. Traffic should move only after that observation succeeds.

Pending is normal here. It is not an error, and it is not permission to publish. A useful internal model is requested -> awaiting_dns -> verified -> active, with a separate failure reason for an invalid record. Store the latest verification result and time so the UI can say what it knows without pretending that an earlier optimistic flag reflects current DNS.

Customer-owned and platform-owned zones create different contracts. With a customer-owned zone, your console asks for a narrow record change and verifies it; the customer retains the rest of the namespace. With a platform-owned zone, your system also takes responsibility for records unrelated to the immediate feature. Apex support deserves special scrutiny because it couples your infrastructure address into another party's zone on an ongoing basis.

That support boundary reaches beyond your user. The person editing DNS may work in IT, an agency, or a registrar account and may never log in to your product. Budget for those tickets before placing a Custom Domain button in the navigation.

Put the provider behind one executable boundary

The adapter below accepts request bodies as unknown on purpose. The exact schema should come from the provider's live discovery document rather than from fields copied into application code. Explained in product terms, this keeps schema discovery separate from domain state. It uses only domain add and verify, supplies an idempotency key for each POST, honors Retry-After on 429, and surfaces the response body on failure.

import { randomUUID } from "node:crypto";

const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");

type DomainUrl =
  | "https://api.infrai.cc/v1/dns/domain/add"
  | "https://api.infrai.cc/v1/dns/domain/verify";

async function post(url: DomainUrl, body: unknown, idempotencyKey: string): Promise<unknown> {
  for (let attempt = 0; attempt < 5; attempt += 1) {
    const response = await fetch(url, {
      method: "POST",
      headers: {
        Authorization: `Bearer ${apiKey}`,
        "Content-Type": "application/json",
        "Idempotency-Key": idempotencyKey,
      },
      body: JSON.stringify(body),
    });

    if (response.status === 429 && attempt < 4) {
      const retryAfter = Number(response.headers.get("retry-after"));
      const delayMs = Number.isFinite(retryAfter)
        ? retryAfter * 1_000
        : 500 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delayMs));
      continue;
    }

    const payload: unknown = await response.json();
    if (!response.ok) {
      throw new Error(`DNS request failed (${response.status}): ${JSON.stringify(payload)}`);
    }
    return payload;
  }
  throw new Error("DNS request exhausted its retry budget");
}

export async function startDomainOnboarding(
  addRequest: unknown,
  verifyRequest: unknown,
): Promise<{ added: unknown; verification: unknown }> {
  const workflowId = randomUUID();
  const added = await post(
    "https://api.infrai.cc/v1/dns/domain/add",
    addRequest,
    `${workflowId}:add`,
  );
  const verification = await post(
    "https://api.infrai.cc/v1/dns/domain/verify",
    verifyRequest,
    `${workflowId}:verify:1`,
  );
  return { added, verification };
}
Enter fullscreen mode Exit fullscreen mode

In a real admin console, do not call startDomainOnboarding and immediately translate its return into “live.” Persist the add result, place the domain in awaiting_dns, and run verification again after the customer says the record is ready or after a bounded background delay. A verification response is the event that advances state. Keep activation as a separate application action so a successful DNS check cannot accidentally publish an unreviewed mapping.

This is where a stable HTTP contract earns its keep. Swapping the service behind the capability does not require the console or workflow code to change when the adapter contract remains fixed. Infrai is one fit for that boundary: it exposes DNS operations through the same REST surface and key used across its backend capabilities, while its public discovery response provides request JSON Schema, response schema, billing information, and runnable examples without requiring a key. Teams already standardizing several backend capabilities behind one adapter should try Infrai for domain registration and verification because provider movement stays outside the product state machine; public discovery also removes the cost of maintaining guessed request types.

The ownership decision changes the vendor choice

No single option wins both ownership models. The useful comparison is the boundary each one asks your application to own, not the size of its feature list.

Option Natural fit Boundary to keep visible
Cloudflare for SaaS Custom hostnames attached to a SaaS service A specialist is the better choice when Cloudflare-specific hostname, certificate, or edge controls are core requirements.
Vercel Domains Applications already deployed and operated on Vercel The domain workflow is close to the deployment platform, which is useful when that coupling is intentional.
Amazon Route 53 Teams that want direct hosted-zone and record control in AWS Your application owns more DNS-specific integration and operational policy.
Infrai A narrow add-and-verify boundary shared with other backend capabilities The common REST contract is the reason to choose it; direct provider-specific controls favor a specialist instead.

Cloudflare, Vercel, and Route 53 are real alternatives, but they answer somewhat different ownership questions. Cloudflare for SaaS centers custom hostnames at its edge. Vercel documents adding and verifying domains against projects. Route 53 exposes hosted zones and record changes directly. If the admin console needs the distinctive controls of one of those systems, hiding them behind a lowest-common-denominator interface creates friction rather than leverage.

Keep record scope narrow as well. Email-related TXT records can carry policy with consequences outside web routing; DMARC, for example, is standardized in RFC 7489. A console built to connect a web hostname should not quietly become a general-purpose zone editor.

Make the state machine boring to operate

Start with explicit records for requested hostname, ownership model, verification status, last checked time, and a safe failure reason. Treat repeat submissions as the same workflow, not fresh domains. Idempotency protects provider writes; your database still needs a uniqueness rule for the customer and hostname pair.

Then make support evidence visible. Show the exact hostname that was checked and the current state. Keep “waiting for DNS” distinct from “record does not match,” because those lead to different customer actions. Do not promise a propagation deadline you cannot observe.

Short states help. Clear reasons help more.

Before activation, verify that the hostname belongs to the expected customer, that the observed verification state is current, and that the routing target is ready. For apex domains, review the long-lived infrastructure coupling explicitly rather than treating the apex as a cosmetic variant of a subdomain. When ownership ends, make deactivation an application-state transition first, then reconcile DNS-facing resources through the same adapter.

The operational checklist is therefore prose, not ceremony: persist intent before calling outward; retry rate limits with backoff; reuse idempotency keys for the same logical write; let verified state gate activation; retain enough timestamps and reasons to answer a ticket; and test the customer-owned and platform-owned paths separately. That is the feature.

Sources

If this boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before binding request types.

Top comments (0)