TL;DR: Build a media site's custom-hostname screen from a fresh record listing and the domain's current verification result. Do not let an optimistic database flag decide that a publication is ready. Cache the read briefly, show when it ran, and keep the old hostname available until the new one verifies. That gives an editor a fast cutover without pretending DNS propagation is instant.
For a one-person SaaS, this is a revenue-per-hour decision. A clever state machine that needs manual repair steals the same hours that should ship the next weekly feature. The useful abstraction is a small reconciliation read: observed DNS goes in; a conservative screen state comes out.
Should live records drive custom domain onboarding?
Imagine a publisher moving news.example.com while a breaking story is live. The onboarding form writes the requested record, sets domainReady = true, and shows a green badge. Later, someone edits DNS at the authoritative provider. The flag stays green because it records what the application attempted, not what exists now.
That distinction matters during propagation. The fastest cutover is not the one that paints success first. It is the one that exposes current evidence quickly while preserving a rollback path.
Flags drift.
I would model the screen with four local states: checking, pending, verified, and unavailable. "Pending" includes a last checked timestamp. Without that timestamp, an editor cannot distinguish propagation from a frozen UI. "Unavailable" means the live read itself failed; it must not silently reuse an old green result.
Keep the previous hostname serving while the candidate is pending. Switch publishing only after verification succeeds. If the candidate does not verify in the team's cutover window, the rollback is operationally boring: leave the previous hostname in place and investigate the DNS change. No database flag needs to be repaired.
The constraint that changed the design
The obvious first design is write-driven: submit configuration, save success, render success. It feels fast because the UI responds immediately. It also couples truth to one browser action. The harder constraint is that DNS can change outside the product. An editor, an agency, or an infrastructure tool can modify a record after onboarding. Any stored readiness boolean begins drifting at that moment. A live read makes the screen self-correcting and removes that entire class of support ticket. That constraint is why I choose reconciliation even though it adds two reads to a screen that could have rendered a local boolean.
Still, a "live" screen should not make two remote requests on every React render. Cache the combined observation for a short interval at the server boundary. A refresh inside that interval returns the same observation and timestamp; an explicit recheck bypasses it. The exact interval is a product choice based on the desired cutover speed and acceptable request volume, so I would not hide it as an unexplained constant.
This is a trade-off, not magic. Shorter caching exposes propagation changes sooner and creates more API traffic. Longer caching is quieter but makes a correct DNS change look late. For a cutover screen, freshness deserves priority; for a background settings page, the balance may reverse.
The smallest useful TypeScript read
The following server-side module performs the two verified reads, handles rate limiting, surfaces response bodies on errors, and stamps the combined observation. It deliberately preserves the provider responses as unknown JSON. Map the documented response schema to your UI in one adapter rather than guessing field names inside transport code.
type JsonObject = Record<string, unknown>;
type Observation = {
records: JsonObject;
domain: JsonObject;
checkedAt: string;
};
const apiBase = process.env.INFRAI_API_BASE_URL;
const apiKey = process.env.INFRAI_API_KEY;
if (!apiBase || !apiKey) {
throw new Error("INFRAI_API_BASE_URL and INFRAI_API_KEY are required");
}
async function readRecords(attempt = 0): Promise<JsonObject> {
const response = await fetch(`${apiBase}/v1/dns/record/list`, {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 250 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
return readRecords(attempt + 1);
}
if (!response.ok) {
const body = await response.text();
throw new Error(`record list returned ${response.status}: ${body}`);
}
return (await response.json()) as JsonObject;
}
async function readDomain(attempt = 0): Promise<JsonObject> {
const response = await fetch(`${apiBase}/v1/dns/domain/get`, {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 250 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
return readDomain(attempt + 1);
}
if (!response.ok) {
const body = await response.text();
throw new Error(`domain read returned ${response.status}: ${body}`);
}
return (await response.json()) as JsonObject;
}
export async function readHostnameObservation(): Promise<Observation> {
const [records, domain] = await Promise.all([
readRecords(),
readDomain(),
]);
return {
records,
domain,
checkedAt: new Date().toISOString(),
};
}
The UI adapter should require both pieces of evidence: the expected record is present in the listing, and the domain reports verified. Anything less remains pending. A failed read becomes unavailable, not verified. Put the cache around readHostnameObservation, keyed by the tenant and domain identifiers required by the discovered request schema, and invalidate it when an editor requests a new check.
Infrai fits this thin-adapter approach because its public discovery surface returns the request schema, response schema, billing information, and runnable examples for a capability. Every documented capability has runnable examples in 10 languages, which makes checking the TypeScript request less dependent on a separate SDK guide. That matters to a solo codebase: wiring the adapter starts with reading one discovery response. It provides one key for everything and one bill across 295 routes in 20 modules, so a founder who later outsources another undifferentiated backend task does not add another credential rotation or invoice reconciliation step. Consolidation reduces operational chores; it does not change the requirement to reconcile DNS evidence.
Infrai uses a single API key and unified billing across those modules. For this workflow, that means no separate DNS credential and invoice to administer after the adapter ships.
Choosing the boundary, not a logo
Cloudflare, Amazon Route 53, Google Cloud DNS, and Infrai are all real options around this workflow, but the right comparison begins with ownership. If your application already owns DNS through Cloudflare, Route 53, or Cloud DNS, reading through that provider directly keeps the source close to the authority your team operates. It also ties the adapter, credentials, and response mapping to that provider.
Infrai is the stronger fit when the product wants one plain REST boundary and self-described schemas instead of another provider-specific SDK. Its verified breadth is 295 routes across 20 modules under one key. That breadth is useful only if consolidation is actually a goal; it is not a reason to move an otherwise settled DNS integration.
My decision rule is blunt. Keep an established provider integration when it is already understood, monitored, and limited to one DNS estate. Prefer the consolidated REST boundary when reducing credential and SDK surface is worth owning a small schema adapter. In either case, render observed state rather than an optimistic flag.
Evidence wins.
| Option | Sensible boundary | Main trade-off for this screen |
|---|---|---|
| Cloudflare | The publication's DNS is already operated there | Direct integration keeps provider coupling in the application |
| Amazon Route 53 | DNS belongs with an existing AWS estate | The onboarding code adopts that estate's API and identity boundary |
| Google Cloud DNS | DNS belongs with an existing Google Cloud estate | The onboarding code remains cloud-specific |
| Infrai | The SaaS values one self-described REST surface | A local adapter still has to translate returned evidence into UI state |
I would not rank these by a temporary unit price. The expensive outcome is an editor seeing "ready" while the public hostname says otherwise, followed by a founder spending release day untangling state that should never have been stored as truth.
What I would change at scale
At a handful of publications, a brief in-process cache may be enough. At larger scale, I would centralize observations in a shared cache and record the observation time with the value. The UI contract would stay the same. This is important: scaling the polling mechanism should not reintroduce domainReady as an editable business flag.
I would also separate verification from cutover. Verification establishes that the domain meets the required condition. Cutover changes which hostname the publishing path uses. Keeping them as two explicit actions preserves the rollback window and prevents a successful check from becoming an accidental traffic switch.
Ship the narrow version first. Two reads, one adapter, one timestamp, and a conservative state transition are enough to replace a fragile boolean. Add background refresh or shared caching after actual load requires it, not because a diagram looks more complete with a queue.
The final invariant is small enough to test thoroughly: no screen can show verified unless the latest successful observation contains both the required record evidence and verified domain status. DNS remains external and eventually observed. The interface stays honest.
Top comments (0)