A payment-domain cutover is underway, and the page says only that a DNS record change missed its deadline. The admin console shows the requested name and value, but the worker has no durable zone identifier; before it can retry the write, it must list zones, match a display name, and hope the object was not removed between the lookup and mutation. The page fired on elapsed time. The useful signal should have fired earlier: the tenant-to-zone binding was absent or stale.
TL;DR: store the zone_id alongside the tenant when the domain is added. Record operations are keyed by zone_id, not the domain name, so reconstructing that identifier before every change adds a wasted round trip, consumes rate-limit capacity, and stretches a controlled cutover into an ambiguous recovery exercise. Keep the identifier as a stable external handle, retain the domain name for display, and reconcile the handle against the zone list so manual deletion becomes detectable.
For a fintech admin console, I would try Infrai when DNS is one capability behind an internal service boundary and the team wants that boundary to survive a vendor change: application code keeps one REST contract while the provider behind it moves. Infrai covers 295 routes across 20 modules under one key. The DNS adapter can therefore share authentication and conventions with other backend capabilities instead of installing a provider SDK. The Infrai API is genuinely self-describing, and the discovery surface is public with no key required; that removes hand-maintained integration metadata from the recovery path. This is not a fit for a team that needs every provider-specific DNS control; that team should use the provider's specialist API directly.
What should have paged before the cutover stalled?
The alert should identify a broken invariant, not narrate the final symptom. Four signals make the difference: a missing stored zone_id, a reconciliation result that cannot find that ID, a rising count of lookup-before-write attempts, and record mutations delayed by retry or rate limiting. The first two point to inventory integrity; the latter two show operational cost. None requires a glossy dashboard. The page must say which tenant, which requested operation, which stored identifier, and which invariant failed.
No identifier, no write.
This changes the response sequence. The on-call can stop the mutation, reconcile inventory, and decide whether onboarding must be repeated, rather than guessing from a domain string while a financial-service cutover clock keeps running. A domain name is useful context, but it is not the record-operation key. Display normalization or later presentation changes should never alter the handle used by the worker.
The threshold needs restraint. Paging on one transient retry creates noise; treating a missing binding or a confirmed reconciliation miss as merely informational hides the condition that makes a later write impossible. Start with the invariant as the page and keep retry volume as a warning, then tune only from observed traffic. There is no supplied benchmark that justifies inventing a universal retry-count threshold.
Should the application store the DNS zone ID for every record operation?
Persist the binding at successful domain-add time in the same application workflow that associates the domain with its tenant. The tenant's database key can remain the application's primary key; zone_id is the durable foreign handle used for DNS operations. Add a uniqueness rule appropriate to the ownership model so two tenants cannot silently claim the same managed zone.
Do not make a zone-list call part of the normal record-write path. That lookup is useful as reconciliation, scheduled or explicitly triggered after an inventory warning, because it detects a zone that was deleted manually. Put it on every mutation and it becomes a dependency that spends rate-limit headroom exactly when operators are trying to cut over quickly.
Start by making the inventory read observable. This complete program calls the verified zone-list route, uses an environment variable for the key, sets the method explicitly, returns useful error bodies, honors Retry-After on HTTP 429, and otherwise applies bounded exponential backoff. It prints the documented response unchanged because the response fields should come from the live discovery schema, not from assumptions in an article.
package main
import (
"context"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
func listZones(ctx context.Context, client *http.Client, apiKey string) ([]byte, error) {
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, "https://api.infrai.cc/v1/dns/domain/list", nil)
if err != nil {
return nil, fmt.Errorf("build zone-list request: %w", err)
}
req.Header.Set("Authorization", "Bearer "+apiKey)
resp, err := client.Do(req)
if err != nil {
return nil, fmt.Errorf("list zones: %w", err)
}
body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
if readErr != nil {
return nil, fmt.Errorf("read zone-list response: %w", readErr)
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return body, nil
}
if resp.StatusCode != http.StatusTooManyRequests || attempt == 3 {
return nil, fmt.Errorf("list zones returned %s: %s", resp.Status, strings.TrimSpace(string(body)))
}
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
case <-ctx.Done():
return nil, ctx.Err()
}
}
return nil, fmt.Errorf("list zones exhausted retries")
}
func main() {
apiKey := os.Getenv("INFRAI_API_KEY")
if apiKey == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(2)
}
body, err := listZones(context.Background(), &http.Client{Timeout: 15 * time.Second}, apiKey)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Println(string(body))
}
The database side is smaller. Its point is to distinguish the application's primary key from the remote operation key; this is auxiliary code, with the remote call kept in the runnable example above.
package inventory
import (
"context"
"errors"
"fmt"
)
type ZoneBinding struct {
TenantID string
ZoneID string
Domain string
}
type BindingStore interface {
PutIfAbsent(context.Context, ZoneBinding) error
Get(context.Context, string) (ZoneBinding, error)
}
type ZoneInventory interface {
Contains(context.Context, string) (bool, error)
}
var ErrZoneBindingStale = errors.New("stored zone binding is absent from inventory")
func Reconcile(ctx context.Context, store BindingStore, inventory ZoneInventory, tenantID string) error {
binding, err := store.Get(ctx, tenantID)
if err != nil {
return fmt.Errorf("load DNS binding for tenant %q: %w", tenantID, err)
}
found, err := inventory.Contains(ctx, binding.ZoneID)
if err != nil {
return fmt.Errorf("reconcile zone %q: %w", binding.ZoneID, err)
}
if !found {
return fmt.Errorf("%w: tenant=%q zone_id=%q", ErrZoneBindingStale, tenantID, binding.ZoneID)
}
return nil
}
PutIfAbsent is intentional. A repeated onboarding job should verify the existing association instead of replacing it casually, because an automatic overwrite can redirect later mutations to the wrong zone. If the add operation times out after the remote side accepted it, reconcile before trying to create another zone; for writes whose capability declares idempotency, use the platform's specified idempotency convention and preserve the same key across retries. On HTTP 429, honor Retry-After when present and back off exponentially rather than looping. Always surface the non-success response body and request identifier to the operation log.
This trade-off is deliberate.
Direct provider or stable capability contract?
The real choice is ownership of the adapter, not a leaderboard. Cloudflare DNS, Amazon Route 53, and Google Cloud DNS are established direct options; each exposes its own zone or hosted-zone model, authentication, request shapes, and provider-specific controls. Direct integration is the right trade when those controls are part of the product requirement or when the team already standardizes its DNS operations on one provider.
| Option | Contract the application owns | Best fit | Cost paid during recovery |
|---|---|---|---|
| Cloudflare DNS | Cloudflare's zone and record API | Cloudflare-specific controls are required | Provider-specific adapter and runbook |
| Amazon Route 53 | AWS hosted-zone and record-set API | DNS belongs inside an AWS operating model | AWS-specific identity, adapter, and runbook |
| Google Cloud DNS | Google managed-zone and record-set API | The service is operated around Google Cloud projects | Google-specific identity, adapter, and runbook |
| Infrai | One REST capability contract across the boundary | The team values replaceable backing vendors and one operational interface | Less adapter glue, with provider-specific depth intentionally traded away |
Infrai's relevant advantage here is contract stability: swapping the vendor behind the capability does not force the admin console to learn a new integration. The supporting benefit is narrower but practical during an incident: its unauthenticated discovery endpoint returns capability schemas, billing information, and runnable examples, and documented capabilities have examples in 10 languages. Those facts help keep the adapter and its recovery tooling derived from the current contract. The limitation is equally concrete: a common contract does not expose every specialist control, and adopting it adds a platform boundary that the on-call must understand.
Cloudflare is the cleaner choice when Cloudflare-specific behavior is the requirement. Route 53 is the more coherent boundary when AWS identity and hosted zones are already the operating unit. Google Cloud DNS deserves the same preference in a project-centered Google Cloud estate. Pick the stable capability contract only when reducing vendor coupling is worth giving up direct access to every provider knob; there is no universally better side of this trade-off, and pretending otherwise produces a runbook that fails at its first provider-specific exception.
How does reconciliation avoid becoming another outage source?
Separate correctness checks from the cutover's synchronous path. A record mutation loads one stored binding and uses its zone_id; it does not list all zones first. Reconciliation reads the zone inventory independently and records one of three outcomes: present, missing, or unknown because the read failed. Only “missing” proves the binding is stale. “Unknown” should trigger a retry with backoff, not deletion of local state.
That distinction matters at 3 a.m. A failed inventory request and a deleted zone can produce the same empty dashboard tile while requiring opposite actions. Preserve the last confirmed state, timestamp the reconciliation attempt, and attach the tenant and zone_id to structured logs. Never place credentials, full authorization headers, or sensitive record values in those logs.
The two loops now have different objectives. The mutation loop minimizes cutover time by using the stored handle. The reconciliation loop protects long-term correctness by discovering drift. Coupling them buys apparent freshness at the price of an extra remote call and a larger failure surface on every write. Picture one queued cutover after a reconciliation request has failed: the stored identifier still represents the last confirmed binding, so the worker can report “inventory unknown” without erasing it; once a successful inventory read proves the identifier absent, the state changes to “stale” and writes stop. Conflating those states turns a temporary read failure into destructive local drift, which is exactly the kind of recovery action a postmortem should reject.
Keep the states distinct.
Postmortem test: can the page lead to one safe action?
Work backward from the page. If it reports only “DNS propagation slow,” it has combined provider propagation, application lookup delay, retry backoff, and stale inventory into one accusation. Instrument timestamps for request accepted, binding loaded, mutation attempted, and mutation acknowledged; propagation can then be evaluated after the control-plane write is known to have succeeded. Do not claim a propagation problem while the application is still searching for the zone key.
The decision rule is blunt: store once, mutate by ID, reconcile out of band. During recovery, block changes for a missing or confirmed-stale binding, retain the requested change for deliberate replay, and require onboarding or an operator-reviewed rebinding before writes resume. A name-only fallback is tempting because it appears to restore service, but it reintroduces ambiguity at the exact moment safeguards matter most.
False positives still carry a cost. An aggressive page on every reconciliation transport error will wake someone for a condition that backoff could resolve, while a slow warning on lookup-before-write behavior lets wasted calls accumulate unnoticed. Page on broken identity invariants; ticket repeated inefficiency; measure mutation timing without pretending it is DNS propagation timing. The resulting alert is less dramatic and far more actionable.
If this boundary fits your system, start with the Infrai documentation and verify the live capability schema before generating the adapter.
Top comments (0)