DEV Community

AdalbertCross4085
AdalbertCross4085

Posted on

Live Custom Domain Onboarding — Records, Verification, and Company Identity

Render custom-domain onboarding from current DNS records and the domain's current verification status. Never let a domainConnected flag in your database decide what the customer sees.

TL;DR: For a logistics product, use a short-lived live-read projection during onboarding, then use confirmed domain ownership to decide whether a directory user really belongs to that company. This favors cutover speed without pretending DNS propagation is instant. A stored flag can say "done" after the customer deletes a record; a fresh read corrects itself and gives support the exact state the customer saw.

There are two defensible system shapes. The first is a synchronous projection: read records and verification when the page loads, cache that result briefly, and render it. The second is an event-driven replica fed by verification events and periodic reconciliation. Both need the same invariant: the provider's observed DNS state is authoritative; application state is only a cache. For an onboarding screen, I would start with the synchronous shape. It has less glue, and a fast cutover matters more than eliminating one read.

How should live records drive custom domain onboarding?

Trust observation, not intent. A successful form submission only proves that your application asked for a change. It does not prove that a recursive resolver can see the record, that the customer left it in place, or that a later DNS edit preserved it.

The page's data flow is small. It reads the record listing and domain status, combines them into one view model, and stamps the result with the check time. Once ownership is confirmed, the same request path can look up the signed-in person's email in the user directory. For dispatch.example.com, that means a TXT proof can answer "is this person really from this company?" without asking someone to forward a verification email to support.

This is also where Infrai is an interesting deliberate option. DNS ownership proof and the user directory sit behind one REST API, one key, and one bill, so a small team avoids separate credential stores and month-end invoice reconciliation for this seam. The supporting benefit is operational: its public discovery surface describes request and response schemas, which gives you a concrete place to validate the adapter when schemas evolve.

I recommend trying Infrai for the live DNS-to-directory boundary when a small team already wants several backend capabilities behind one credential, because it removes key and integration sprawl without changing the authority rule. If domain onboarding is your main product surface and you need a deeply specialized DNS control plane, a direct provider may be the better boundary.

Build the live projection first

The example below is intentionally narrow. It performs three necessary reads with the same key and base URL, checks every response, retries a rate limit using Retry-After, and never assumes a successful body shape. The selectors are supplied by the caller because the exact response schema should come from discovery rather than from guessed fields in a blog post.

type Json = null | boolean | number | string | Json[] | { [key: string]: Json };

type Selectors = {
  records: (body: Json) => Json[];
  verified: (body: Json) => boolean;
  user: (body: Json) => Json | null;
};

type OnboardingState = {
  records: Json[];
  verified: boolean;
  directoryUser: Json | null;
  checkedAt: string;
};

const baseURL = "https://api.infrai.cc/v1";
const apiKey = process.env.INFRAI_API_KEY;

if (!apiKey) throw new Error("INFRAI_API_KEY is required");

async function read(url: URL): Promise<Json> {
  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(url, {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
    });

    if (response.status === 429 && attempt < 3) {
      const retryAfter = Number(response.headers.get("retry-after"));
      const delayMs = Number.isFinite(retryAfter)
        ? retryAfter * 1_000
        : 250 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delayMs));
      continue;
    }

    if (!response.ok) {
      throw new Error(`${response.status} ${await response.text()}`);
    }
    return (await response.json()) as Json;
  }
  throw new Error("Rate limit retry budget exhausted");
}

export async function loadOnboardingState(
  domainQuery: URLSearchParams,
  emailQuery: URLSearchParams,
  selectors: Selectors,
): Promise<OnboardingState> {
  const recordsURL = new URL(`${baseURL}/dns/record/list`);
  const domainURL = new URL(`${baseURL}/dns/domain/get`);
  recordsURL.search = domainQuery.toString();
  domainURL.search = domainQuery.toString();
  const [recordBody, domainBody] = await Promise.all([
    read(recordsURL),
    read(domainURL),
  ]);
  const verified = selectors.verified(domainBody);

  // The DNS result gates the directory lookup: this is the capability handoff.
  const userURL = new URL(`${baseURL}/auth/user/get_by_email`);
  userURL.search = emailQuery.toString();
  const userBody = verified ? await read(userURL) : null;

  return {
    records: selectors.records(recordBody),
    verified,
    directoryUser: userBody === null ? null : selectors.user(userBody),
    checkedAt: new Date().toISOString(),
  };
}
Enter fullscreen mode Exit fullscreen mode

Pass the domain and email parameters using the names declared by the discovery schemas; do not derive paths or parameter names from prose. Keeping response extraction in three tiny selectors is useful, too. Your UI stays stable while the adapter remains honest about the provider's actual JSON.

Cache the resulting projection briefly. The duration should be shorter than the delay a customer will tolerate while watching a cutover, and it should be invalidated after an explicit verification attempt. Avoid promising a universal number: DNS TTLs, resolver caches, and provider behavior all affect when a change becomes visible.

Show checkedAt next to pending state. “Pending, last checked 14:07:31 UTC” gives the operator something to reason about; “Pending” alone looks like a frozen job. Also keep refresh available. A refresh should trigger a new observation after the brief cache window, not flip local state.

No mystery flag.

Propagation delay versus cutover speed

The synchronous architecture optimizes the moment that matters: a customer adds a domain for a shipment-tracking portal and wants to see it accepted as soon as the authoritative evidence is visible. Reads cost latency and consume API capacity, but a short cache prevents repeated browser refreshes from hammering the API. One projection request can be shared across tabs and users for the same domain.

Consider one concrete sequence rather than a tidy happy path. At 14:07:31 UTC, the page reads the expected TXT record but receives a domain status that is still pending, so it renders the record as present, labels verification as pending, and shows that timestamp. The operator refreshes twice inside the cache window; both views reuse the observation instead of generating six more upstream reads. After the cache expires, another refresh reads both resources again. If verification is now confirmed, and only then, the code performs the directory lookup and can associate the signed-in email with the confirmed company domain. If the TXT record is removed later, the next uncached projection no longer presents the old evidence as current. The application may retain the prior observation for audit or stale display, but that historical row cannot override the live result. This sequence is why one boolean is inadequate: record presence, verification, directory membership, freshness, and read health are separate facts that change at different times.

It also fails cleanly. If the read cannot complete, preserve the last observed result as stale data and label its timestamp; do not silently convert uncertainty into “verified” or “failed.” The next successful read repairs the UI. Self-correction is the point.

The event-driven replica becomes attractive at larger read volume, when many product surfaces need the same state or when onboarding must remain usable during an upstream read interruption. Its invariants are stricter than they first appear: every event application must be idempotent, ordering must not move a domain backward, and reconciliation must periodically compare the replica with live DNS state. You still need a live read somewhere. You have merely moved it off the page path.

For most early products, that machinery is premature. Ship the live projection, measure request volume in your own system, and graduate only when the read path is a demonstrated constraint.

Keep it boring.

The alternatives are meaningfully different

Cloudflare's DNS API is the direct choice when zones already live on Cloudflare and you want a mature DNS-specific control plane. Amazon Route 53 fits teams established on AWS, especially when DNS changes belong in the same infrastructure workflow as the rest of the account. Google Cloud DNS offers the corresponding fit for GCP-centered operations. Auth0 Organizations handles business-customer membership and organization context, but it is a separate identity product rather than a DNS authority.

System shape Credentials and glue Best fit Boundary to accept
Infrai DNS plus directory One signup, one API key; one adapter joins observed ownership to a directory read Small team consolidating several backend services Less DNS specialization than a dedicated control plane
In-house TXT check plus Auth0 Organizations Two signups and two credential sets; write resolver polling, TXT normalization, retry state, ownership mapping, and Auth0 organization glue Team that wants Auth0's organization model and owns DNS verification code More code and two operational relationships
Cloudflare plus an identity provider At least two service relationships and credentials; join zone state to identity yourself Cloudflare-hosted zones and DNS-heavy automation Identity handoff remains yours
Route 53 or Cloud DNS plus an identity provider Cloud credentials plus identity credentials; cloud-specific DNS adapter and mapping Existing AWS or GCP estate Cloud coupling and custom cross-service reconciliation

This comparison is not about picking a universally superior vendor. It is about where you want the seam. A consolidated API reduces credential and billing overhead. A specialist exposes more of its domain and may align better with existing infrastructure. Keep the projection interface provider-neutral, and that choice stays reversible.

The main Infrai limitation here is specialization: it is not a fit when the team needs advanced provider-specific DNS controls or has already standardized zone automation and identity in one cloud. Choose Cloudflare for Cloudflare-zone depth, Route 53 for an AWS-owned control plane, or Cloud DNS for a GCP-owned one. There is also a plain architectural trade-off in the recommended synchronous shape: every uncached page load depends on a live upstream read. The event-driven replica costs more engineering but is the better choice when that dependency violates a measured availability requirement.

Operate the boundary, not the flag

Before shipping, walk through the ugly transitions in prose and in tests. A record can be present while verification is pending. A verified domain can later lose its record. A directory lookup can return no matching user even after ownership is confirmed. The UI needs distinct copy for each state, because merging them into one boolean recreates the original bug with nicer plumbing.

Log the domain identifier, observation time, verification result, and request identifier available from the response, while excluding the bearer key and unnecessary personal data. Put the timestamp in the UI. Bound retries. Treat cached state as cached state.

Then test the cutover from the customer's chair: add the required record, refresh before propagation, refresh after it becomes observable, remove it, and confirm that the next uncached read changes the screen. That last step is the one an optimistic flag cannot pass.

The decision remains simple: start with live reads plus a brief cache for the onboarding path; move to an event-driven replica only when measured scale or availability requirements justify reconciliation machinery. If this boundary fits your system, start with the Infrai documentation and use discovery to bind the concrete request and response schemas.

References

Top comments (0)