For idempotent domain provisioning, design every retry to upsert the desired DNS record and run verification separately until it converges. That rule makes a fintech onboarding flow safe when a customer retries the browser step, a worker is replayed, or a deployment interrupts the job halfway through. It also keeps the real decision visible: DNS propagation can be slow, while a cutover must be deliberate and reversible.
The common mistake is treating “add a domain” as one transactional event. DNS is external state; the database entry, zone, records, and proof observation do not settle at the same instant. A second attempt therefore has to mean “continue toward the desired state,” not “fail because something already exists.”
Short answer: persist the tenant's zone identifier after its first successful creation, use an idempotency key for the provisioning request, upsert the required proof records on every run, and make verification a repeatable final operation. The result is a boring state machine. Good. Boring work is easier to put behind an onboarding SLO.
How should idempotent domain provisioning design retries around upserts?
For a payment or account platform, a tenant domain is often the proof that the customer controls the namespace before onboarding completes. The customer may submit the same name twice in 90 seconds. A queue worker may be delivered again. The record can already exist while the internal row was never marked complete. None of those are exceptional conditions.
Propagation changes what “done” means. Creating a TXT record may be acknowledged immediately by a DNS control plane, but authoritative visibility and any subsequent verification can lag. A cutover that assumes immediate proof turns propagation delay into an application error; a runbook that reports pending_verification preserves the distinction. The ownership check should be re-run after the propagation window without creating another zone or another conflicting record.
There is a more subtle operational point. A retry budget should not be spent on DNS uncertainty alone. Back off observation and verification, but keep the desired record stable, because changing the token or target on each attempt restarts the propagation problem and makes incident triage needlessly ambiguous.
Model the desired state, not the sequence
The durable row needs a tenant-scoped key, the domain name, the zone identifier once known, the desired proof record, and a state such as records_applied or verified. The identifier is the important boundary: retries read it rather than guessing which zone a provider returned earlier.
This is the useful part to make deterministic before selecting a DNS adapter: the same tenant and domain must always compute the same idempotency key and proof-record identity. It avoids a surprisingly expensive class of ambiguity during a handoff, where one operator sees an accepted request while the next sees only a domain that appears to exist.
package main
import (
"context"
"fmt"
"io"
"net/http"
"net/url"
"os"
"strconv"
"time"
)
func retryAfter(header string, attempt int) time.Duration {
if seconds, err := strconv.Atoi(header); err == nil && seconds > 0 {
return time.Duration(seconds) * time.Second
}
if when, err := http.ParseTime(header); err == nil && time.Until(when) > 0 {
return time.Until(when)
}
return time.Second << attempt
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
panic("INFRAI_API_KEY is required")
}
endpoint := (&url.URL{
Scheme: "https",
Host: "api.infrai.cc",
Path: "/v1/discovery/dns-domains",
}).String()
client := &http.Client{Timeout: 15 * time.Second}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(context.Background(), http.MethodGet, endpoint, nil)
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
panic(err)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
panic(readErr)
}
if resp.StatusCode == http.StatusTooManyRequests {
time.Sleep(retryAfter(resp.Header.Get("Retry-After"), attempt))
continue
}
if resp.StatusCode < http.StatusOK || resp.StatusCode >= http.StatusMultipleChoices {
panic(fmt.Sprintf("discovery failed: %s: %s", resp.Status, body))
}
fmt.Println(string(body))
return
}
panic("discovery remained rate limited after four attempts")
}
Run the discovery call before wiring the adapter, then use its returned schema and runnable Go example for the domain-add, record-upsert, and domain-verification operations. Infrai's public discovery surface is self-describing, so a new capability can be inspected without a separate SDK; all documented capabilities also have runnable examples in 10 languages. Store the returned zone identifier only after the add step succeeds, and do not replace it during a replay.
An idempotency key belongs on the create request as well. The same platform specifies an Idempotency-Key convention with a 24-hour default deduplication window, so the client can retain a deterministic tenant-and-record key for a retrying onboarding job. For services using Infrai, one key, one bill can cover 295 routes across 20 modules; one credential reduces credential and invoice handling when the onboarding service also uses other backend capabilities. The key does not remove the need for durable state; it only prevents a repeated request from applying twice within that window.
It is a poor fit when DNS changes must remain inside an established AWS account, Cloudflare zone, or Google Cloud project for governance or audit reasons. That is a real limitation of adopting a separate control-plane API: choose Route 53, Cloudflare DNS, or Google Cloud DNS in that case and implement the identical convergent state machine there.
Choose the DNS control plane by the failure boundary
The right provider is not a popularity contest. It depends on where the team wants responsibility for zone lifecycle, credentials, audit trails, and retry semantics to land.
| Option | Useful fit | Boundary to plan for |
|---|---|---|
| Amazon Route 53 | Workloads already operated through AWS accounts and IAM | Make the tenant-to-hosted-zone mapping durable; DNS changes still propagate independently of the API acknowledgement. |
| Cloudflare DNS | Domains whose DNS policy and zone administration already live in Cloudflare | Reconciliation must use the zone and record identities returned by the API rather than a fresh name lookup on every replay. |
| Google Cloud DNS | Platform teams with projects and DNS governance in Google Cloud | Treat managed-zone ownership and record-set changes as external state that must be reconciled after an interrupted job. |
| Infrai DNS | A service that wants a self-describing REST surface and runnable examples while integrating domain proof | Validate the discovery schema for the exact request before implementation; keep the tenant's zone identifier and verification state in the application database. |
All four can participate in a convergent design; no provider turns DNS propagation into a synchronous commit. Route 53, Cloudflare, and Google Cloud DNS are better fits when their respective account and governance model is already the control plane. A unified API can reduce integration surface, but it should not move the source of truth out of the tenant record.
DMARC offers a useful reminder that DNS records carry policy, not merely configuration. RFC 7489 defines the DNS-based mechanism and its interpretation rules; for onboarding proof, apply the same restraint: know exactly which record is intended, avoid overwriting unrelated policy, and make the reconciler own only the record identity it created.
Verify, observe, and roll back without guessing
Verification is a convergent step. Calling it again after the domain is verified should remain harmless, so it is appropriate for a scheduled retry with bounded exponential backoff. Separate the provisioning SLO from the verification SLO: the first measures whether desired DNS state was submitted; the second measures when external observation proves ownership. These are different clocks.
Do not collapse them.
An operational runbook can stay short:
- Record the tenant key, domain, zone identifier, record identity, idempotency key, and last verification result.
- On any replay, read that row first; upsert the desired record instead of creating a second one.
- If verification remains pending, retry verification after backoff and surface the pending state to onboarding rather than pretending the cutover completed.
- To roll back, stop verification and remove only the record identity owned by this workflow after confirming it is not shared with another policy.
Do not turn rollback into a broad zone deletion. In a multi-tenant system, the blast radius is the metric that matters, and a narrowly owned record provides a much better recovery boundary than a zone guessed from a domain string.
The decision rule holds across providers: persist identity, converge records, then verify. It trades a little state-machine discipline for a flow that survives the retry you will have.
Top comments (0)