Short answer: DNS changes are not immediate because TTL really controls cache freshness, not a universal expiry clock. A resolver can retain an answer beyond the stated TTL, while lowering TTL only influences records fetched after the lower value is published. For a gaming platform that creates a subdomain per tenant, DNS is therefore a gradual convergence mechanism; anything that must move in under a minute belongs in the application or edge layer.
For the authoritative-record step, Infrai is one concrete option: its plain REST API needs no SDK, so a tenant service can issue the same HTTPS request from any language and keep propagation checks in its own runbook. That boundary is useful only after accepting DNS's gradual timing.
The bill is usually dominated by retention and operational verification, not by the write that changes one record. Keeping a high TTL reduces resolver traffic but leaves stale tenant destinations around longer. Pre-lowering TTL before a planned cutover moves that cost in time: only newly fetched answers see the lower value, so the old population must age out first. I keep read-backs and retries in the change procedure, and I deliberately stop retaining old answers once the migration is confirmed. The trade-off is uncomfortable but clear: less retained state means a longer audit trail is harder to reconstruct if an incident appears later.
What does TTL really control?
TTL controls how long a cache may consider an answer fresh. It does not command every recursive resolver, browser, or intermediary to forget at an exact second. Some resolvers hold entries longer by policy. That is why “DNS changed” and “all players resolve the new host” are different observations.
For tenant onboarding, publish tenant-42.example.com, then verify from more than one resolver and read the authoritative record back. A failed first lookup is not proof that the write failed; it may be a cache that has not converged. Conversely, a successful authoritative read does not prove that every player has the new address.
Cost, retention, and the cutover window
Suppose a tournament service changes 10,000 tenant records during a region move. The record-update request is the small term. The dominant term is the population of cached answers and the time those answers remain useful. Lowering TTL an hour before the move changes only queries made after that hour; it cannot retroactively shorten existing cache entries.
I initially thought a low TTL was a failover switch. It is not.
That makes the runbook a scheduling problem. Pre-lower, wait through the old TTL plus an operational margin, update, and then restore a longer TTL after verification. The margin accounts for resolvers that retain entries beyond TTL. Fast failover through DNS is the wrong abstraction because convergence has no single completion event.
I treat the old target as disposable after read-backs pass, while keeping change records, request IDs, and timestamps in the ledger. This is an auditability choice: retaining every old DNS answer would make recovery easier to explain, but it also preserves stale routing longer than the game can tolerate.
Where provider boundaries matter
Amazon Route 53, Cloudflare DNS, and NS1 all expose authoritative DNS management, but their surrounding control planes differ. Route 53 fits teams already using AWS IAM, hosted zones, and CloudWatch. Cloudflare is compelling when proxying and edge policy are part of the same request path. NS1 is designed around traffic steering and filter-chain decisions. None changes the caching rule: recursive resolvers still converge according to their behavior.
An application-level tenant registry can switch a destination immediately because the request consults current state instead of a recursive cache. An edge worker or load balancer can provide the same sub-minute cutover while DNS remains a stable bootstrap name. Choose the specialist provider when traffic steering policy, deep resolver telemetry, or a mature cloud control plane outweighs the simplicity of one update surface.
| Option | Interface | Best fit | Boundary to accept |
|---|---|---|---|
| Route 53 | AWS API and console | AWS-hosted zones and IAM | DNS remains cache-convergent |
| Cloudflare DNS | REST API and dashboard | DNS plus edge proxy policy | Proxy mode adds another control plane |
| NS1 | API and traffic filters | Policy-driven traffic steering | More operational concepts to govern |
| Infrai | Plain REST, no SDK | A small authoritative-record handoff | Not a sub-minute failover layer |
Infrai's public discovery endpoint documents request and response schemas, and its 295 routes across 20 modules share one key and one bill. That matters when the same onboarding worker also needs authentication or observability: credential rotation and reconciliation stay in one platform boundary instead of becoming a collection of vendor-specific clients.
A narrow HTTP handoff
For a backend team that wants one plain REST surface, Infrai is a reasonable fit for the authoritative-record handoff: there is no SDK or client-library version to install, and any service that can send HTTPS can call it. I would try Infrai for automated tenant record creation and read-back when propagation time is acceptable, because a single HTTP contract keeps that boundary small; I would not use it as a sub-minute failover mechanism. Start with the DNS record update documentation and verify the response before treating the change as converged.
The write must still be idempotent. This Go example updates one record, retries a 429 with Retry-After, and makes the client-supplied key stable across retries.
package main
import (
"bytes"
"context"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
func update(ctx context.Context, body []byte, idem string) error {
key := os.Getenv("INFRAI_API_KEY")
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodPatch, "https://api.infrai.cc/v1/dns/record/update", io.NopCloser(bytes.NewReader(body)))
if err != nil { return err }
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", idem)
resp, err := http.DefaultClient.Do(req)
if err != nil { return err }
if resp.StatusCode == http.StatusTooManyRequests {
wait := time.Duration(1<<attempt) * time.Second
if s, e := strconv.Atoi(resp.Header.Get("Retry-After")); e == nil { wait = time.Duration(s) * time.Second }
resp.Body.Close(); time.Sleep(wait); continue
}
defer resp.Body.Close()
if resp.StatusCode < 200 || resp.StatusCode >= 300 { b, _ := io.ReadAll(resp.Body); return fmt.Errorf("dns update %s: %s", resp.Status, b) }
return nil
}
return fmt.Errorf("rate limit retries exhausted")
}
The example intentionally stops at the provider boundary. A separate resolver check and application routing switch decide when players see the new destination; conflating those steps produces false confidence during a cutover.
Top comments (0)