TL;DR: To make the whole domain provisioning flow rerunnable, store each tenant's zone identifier and complete intended record set in your database. Let every Node.js run converge on that intent: upsert every record, accept a match as success, and keep verification inside the loop. A fast cutover comes from preparing intent before traffic moves; a low TTL can't repair a half-finished workflow.
| Choice | Best fit | Cutover control | Integration cost |
|---|---|---|---|
| Amazon Route 53 | DNS already governed in AWS | Direct control through its own API | Provider-specific client and model |
| Cloudflare DNS | Zones already operated behind Cloudflare | Direct control through its own API | Provider-specific client and model |
| Google Cloud DNS | DNS already governed in Google Cloud | Direct control through its own API | Provider-specific client and model |
| Infrai | The backing vendor may change while the application contract must stay fixed | One REST API and one key across capabilities | A small adapter; no provider SDK required |
Recommendation: keep reconciliation provider-neutral, then select the control plane that already owns the zone. For a fintech admin console spanning providers, a stable contract is valuable because swapping the service behind the capability does not force a rewrite of the console's provisioning logic. A unified REST control plane fits that boundary, especially when it exposes public discovery for its request schemas. If one cloud already owns every zone, its native DNS product is the cleaner choice.
How do you make the whole domain provisioning flow rerunnable?
There are two clocks, and treating them as one creates bad runbooks.
The first clock is application time: how long the admin console takes to persist intent, reconcile records, and request verification. The second is DNS propagation: how long resolvers continue using cached answers. The provisioning worker controls the first. TTL and resolver behavior constrain the second. That's the distinction I would benchmark, because a single blended duration can't tell an API delay from an expected cached answer.
So optimize them separately. Write the desired record set before starting work. Reconcile ahead of the cutover. Verify in the loop. Then move traffic according to the DNS plan rather than asking an HTTP request to wait for global observation.
No magic here.
A low TTL may shorten how long a cached answer remains useful, but it does not make a sequence of imperative steps resumable. If a process writes three of five records and exits, the next run must discover that the first three are already correct and continue. It must not create duplicates, skip verification, or depend on an in-memory step counter.
That gives four rules:
- Persist the zone identifier with the tenant.
- Persist the complete intended record set, not merely the next operation.
- Upsert every intended record on every run; matching state is success.
- Treat verification as another convergent action, not a one-off epilogue.
The key trade-off is blunt: more reconciliation calls buy simpler recovery. For an internal fintech console, I would take that trade. A provisioning action is rare compared with an ordinary application request, while an ambiguous domain state can stall a scheduled cutover.
Put intent in data, not control flow
An imperative workflow says, "create A, then create B, then verify." Its progress is hidden in the location of the instruction pointer. That state disappears when the process stops.
A reconciler instead asks, "what should be true for this tenant?" The answer must survive a deploy, queue retry, operator click, or worker restart. Store it beside the tenant. A status field may help the UI, but it is not the source of truth; the zone identifier plus desired records are.
Use a narrow domain model. This Node.js example deliberately avoids a provider's request schema. The adapter owns that mapping, while the reconciliation logic owns the invariant.
type DnsRecordIntent = Readonly<{
type: "A" | "AAAA" | "CNAME" | "MX" | "TXT";
name: string;
value: string;
ttl: number;
}>;
type TenantDnsIntent = Readonly<{
tenantId: string;
zoneId: string;
records: readonly DnsRecordIntent[];
}>;
type DnsControlPlane = Readonly<{
upsertRecord(zoneId: string, record: DnsRecordIntent): Promise<void>;
verifyDomain(zoneId: string): Promise<"verified" | "pending">;
}>;
async function convergeTenantDns(
intent: TenantDnsIntent,
dns: DnsControlPlane,
): Promise<"verified" | "pending"> {
for (const record of intent.records) {
await dns.upsertRecord(intent.zoneId, record);
}
return dns.verifyDomain(intent.zoneId);
}
This function is boring. Good. Run it after an admin saves a change, from a queue retry, or from a repair job. The same input should approach the same external state.
Notice what is missing: currentStep, recordCreated, and a branch that assumes verification only happens once. Those fields turn yesterday's execution history into today's behavior. They also multiply recovery paths.
For a record writer, use the provider's upsert primitive where one exists. On the referenced REST surface, the relevant operations are PUT /v1/dns/record/upsert and POST /v1/dns/domain/verify. Keep those calls inside the adapter. The rest of the application should not know their paths or payloads.
Before generating that adapter, check the live contract instead of copying request fields from an article. The discovery endpoint is public, but this runnable probe uses the same bearer-key convention as authenticated calls. It uses an explicit method, honors Retry-After, and surfaces the response body on errors.
type Capability = Readonly<{
id: string;
method: string;
path: string;
available: boolean;
}>;
type Discovery = Readonly<{
version: string;
generated_at: string;
capabilities: readonly Capability[];
}>;
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
const baseUrl = process.env.INFRAI_BASE_URL;
if (!baseUrl) throw new Error("INFRAI_BASE_URL is required");
async function discover(attempt = 0): Promise<Discovery> {
const response = await fetch(`${baseUrl}/discovery`, {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 250 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
return discover(attempt + 1);
}
if (!response.ok) {
throw new Error(`Discovery failed (${response.status}): ${await response.text()}`);
}
return response.json() as Promise<Discovery>;
}
const discovery = await discover();
const required = new Set([
"PUT /v1/dns/record/upsert",
"POST /v1/dns/domain/verify",
]);
const available = discovery.capabilities
.filter((capability) => capability.available)
.map((capability) => `${capability.method} ${capability.path}`)
.filter((operation) => required.has(operation));
if (available.length !== required.size) {
throw new Error("The required DNS contract isn't available");
}
Why is verification part of reconciliation?
Because verification can be pending without making the run a failure.
The console should persist the latest observed result, show pending honestly, and schedule the same reconciler again. On the next run, record upserts settle as matches and verification gets another chance. There is still one path through the code.
Do not hold an admin HTTP request open until DNS caches change. Save intent, start reconciliation, and return a state the UI can poll. This separates cutover coordination from propagation. It also gives operators a useful distinction: the desired records are recorded, the provider has accepted the writes, but external verification is not complete yet.
The verifier is not a rollback oracle. A pending result does not prove that the new records are wrong, and deleting them immediately would work against convergence. Keep the intended state stable unless the operator changes it.
There is one uncomfortable boundary: an upsert loop creates and updates desired records, but it does not define deletion. If the intended set is authoritative, stale-record removal needs its own explicit policy and safety checks. The supplied operations establish upsert and verification; they do not justify inventing deletion semantics for this loop. Scope the first version accordingly.
Native control planes versus a stable contract
Amazon Route 53, Cloudflare DNS, and Google Cloud DNS are sensible runner-ups when the zone already lives in that provider. Their native APIs expose the provider's own resources directly. That reduces conceptual distance for a single-provider system: fewer layers, familiar access controls, and documentation adjacent to the rest of that cloud.
The cost appears when the console must manage zones across control planes. Each native integration gives the application another client, authentication boundary, error model, and request shape. An adapter contains that spread, but the team still maintains each implementation.
A common REST contract changes the ownership line. The business reconciler calls one interface while routing can move behind it. That is the strongest reason to consider the aggregated option here. The aggregated option exposes 295 routes across 20 modules through one REST API and one key; its public discovery surface returns request and response schemas, so an adapter can be generated or checked without installing another SDK.
Do not infer speed from API shape. No comparable runtime latency or propagation benchmark is established here, and vendor choice cannot override recursive resolver caching anyway. Measure the two clocks in your environment: time from saved intent to accepted writes, then time from accepted writes to the observations your cutover requires. Record p50 and p95 separately. Otherwise one slow resolver sample will get blamed on the provisioning client.
Choose native Route 53 when AWS is the durable ownership boundary. Choose Cloudflare DNS when Cloudflare already operates the zones and its native policy surface matters to the team. Choose Google Cloud DNS when Google Cloud is that boundary. Choose a stable cross-provider contract when portability of the admin console is worth the added layer.
A cutover runbook that can be repeated
Before the window, save the target records and zone identifier, then run reconciliation. Record the result without claiming global propagation. At the window, change only the stored intent that actually needs to move and run the same function again.
If the worker stops, rerun it.
Afterward, keep invoking verification through the same path until it returns verified. An operator-triggered retry and an automated retry should execute identical code. That property is more useful than a long list of special recovery buttons.
The decision rule stays small: make the whole flow converge with durable intent and idempotent upserts; manage DNS propagation with a separate cutover plan. Mixing them makes both harder to benchmark and much harder to recover.
Top comments (0)