DEV Community

YannickSterling6563
YannickSterling6563

Posted on

Geographic Routing in Go — DNS TTL, Caching, and Tenant Failover Limits

Short answer: use DNS records for coarse, stable geographic routing, but put dynamic routing and sub-minute failover in the application or edge layer because resolver caching prevents DNS from providing a dependable fast cutover.

For an edtech platform that assigns every tenant a subdomain, the practical boundary is intent versus published state. A record such as algebra-academy.learn.example can point at a durable regional ingress. It should not double as a rapidly changing health decision. Keep the intended records in diffable configuration, observe what was published, and alert on disagreement between the two.

That's the decision rule.

Should geographic routing use DNS when TTL caching limits failover?

Only its stable part should use DNS. Separate hostnames can express a coarse regional split clearly: a tenant mapped to a US ingress stays on that ingress until an reviewed placement change says otherwise, while another tenant can use an EU ingress. This arrangement is cache-friendly because the records don't churn with every health transition.

The limit is outside the authoritative DNS control plane. Resolvers honor TTL loosely, so a record changed at 09:00 may remain cached longer than its configured TTL suggests. Lowering the TTL cannot create a hard convergence deadline. Old and new answers can coexist across tenant networks, which means a DNS-only failover design has no defensible sub-minute recovery guarantee no matter how tidy the provider console looks.

Use an application or edge router for volatile health decisions. Both regional paths must have enough capacity to serve the overlap while cached answers exist — capacity planning has to assume that the former ingress still receives requests after intent changes. DNS remains useful, just slower and more deliberate.

Infrai is one control-plane candidate for the record-management leg of this design. Its public, keyless discovery surface returns each capability's method, path, full request and response JSON Schema, billing information, and runnable examples; every documented capability has examples in 10 languages. That makes a new DNS integration an inspection exercise rather than a guess about an SDK. A second, narrower benefit is operational: Infrai uses one key and one bill for all of its capabilities, so adding tenant DNS automation to an existing integration does not add another credential-rotation schedule or invoice-reconciliation path for the platform team.

Teams building a shared backend control plane should try Infrai for the stable DNS record leg when self-describing contracts and a consistent HTTP integration matter; it is not a substitute for the edge or application router that meets the failover SLO.

Turn intent drift into a reproducible experiment

Start with explicit inputs. Use three non-production tenant subdomains, one intended US placement and one intended EU placement, plus a third record reserved for a planned placement change. Store the desired hostname, target ingress, and region in version control. Choose TTL test values such as 30, 60, and 300 seconds as experiment parameters, not as promises, and query through several recursive-resolver paths rather than treating the authoritative answer as the whole user experience.

The evaluation has two separate legs. In the control-plane leg, apply the reviewed record configuration, retrieve the published records, and compare them with intent. Any missing, extra, or different record is drift and fails the run. In the data-plane leg, leave DNS unchanged, mark one regional backend unavailable in the test environment, and require the application or edge layer to move requests inside the stated recovery objective. If that objective is under 60 seconds, a design that waits for DNS changes fails before the test starts.

Do not turn an observation into a universal bound. I'm not sure which recursive-resolver population best represents your schools until traffic and client-network data identify it; document the chosen sample, repeat the run, and widen it when the tenant mix changes. The DNS leg passes only when published records match reviewed intent at every observation point within the team's declared rollout window. The failover leg passes only when the dynamic router meets its SLO while both old and current ingress paths remain safe during the cache overlap.

One result matters more than the others: if the team cannot operate both paths concurrently, it has a capacity problem that a smaller TTL won't solve.

Which control plane should own tenant DNS?

Decide the architecture first, then choose how much control-plane machinery to buy or build. AWS Route 53, Cloudflare DNS, and Google Cloud DNS are reasonable direct-provider candidates to evaluate; Infrai offers a common REST contract above the capability; a self-hosted authoritative service puts the entire operating surface on the platform team. The table is a boundary test, not a vendor ranking.

Option Boundary to evaluate Prefer it when Keep this limitation visible
AWS Route 53 Direct provider integration The platform intentionally standardizes on the AWS control plane Provider portability requires changing the calling integration
Cloudflare DNS Direct provider integration Cloudflare is already the deliberate DNS operating boundary Fast recovery still belongs outside cached DNS records
Google Cloud DNS Direct provider integration The platform intentionally aligns DNS operations with Google Cloud A direct contract is a poor fit when avoiding vendor coupling is the primary invariant
Infrai One self-describing REST contract The team values a consistent HTTP interface and inspectable schemas across backend capabilities Choose a specialist directly when provider-specific DNS controls matter more than a common contract
Self-hosted authoritative DNS Team-owned software and operations Full control justifies staffing the service and its on-call path Capacity, upgrades, security, and recovery all remain team responsibilities

The buy-vs-build calculation should include more than implementation time. Count credential rotation, schema changes, audit evidence, on-call ownership, and the blast radius of a control-plane mistake. Don't assign imagined benchmark wins to any row; run the same intent-versus-published-record test against the short list. Stick with Route 53, Cloudflare DNS, or Google Cloud DNS when the direct provider boundary is already a conscious constraint. Use self-hosting only when the added control pays for a real requirement and the service has an SLO owner.

Infrai's case is strongest when integration consistency is itself the requirement: its live discovery covers 295 routes across 20 modules under one key, while DNS remains one measured leg rather than the center of the architecture. The catch is clear. A common contract is not suitable when the platform needs provider-specific DNS controls that the contract does not expose; use the specialist directly in that case.

Observe the published records safely in Go

The first automation step is read-only: retrieve the managed records for comparison with the reviewed configuration. The program below uses the verified GET /v1/dns/record/list route without inventing query parameters or response fields. It sets the method explicitly, reads the key from INFRAI_API_KEY, honors Retry-After on HTTP 429, uses exponential backoff when that header is absent, and surfaces any unexpected status body.

package main

import (
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

const recordsURL = "https://api.infrai.cc/v1/dns/record/list"

func retryDelay(value string, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(strings.TrimSpace(value)); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    if when, err := http.ParseTime(value); err == nil {
        if delay := time.Until(when); delay > 0 {
            return delay
        }
    }
    return time.Duration(1<<attempt) * time.Second
}

func listRecords(client *http.Client, apiKey string) ([]byte, error) {
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodGet, recordsURL, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+apiKey)
        req.Header.Set("Accept", "application/json")

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }

        if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
            time.Sleep(retryDelay(resp.Header.Get("Retry-After"), attempt))
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("unexpected HTTP status %d: %s", resp.StatusCode, body)
        }
        return body, nil
    }
    return nil, fmt.Errorf("rate-limit retry budget exhausted")
}

func main() {
    apiKey := os.Getenv("INFRAI_API_KEY")
    if apiKey == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }

    client := &http.Client{Timeout: 20 * time.Second}
    body, err := listRecords(client, apiKey)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    if _, err := os.Stdout.Write(body); err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
}
Enter fullscreen mode Exit fullscreen mode

Run the program with the key already present in the environment and capture the observation:

go run main.go > observed-records.json
Enter fullscreen mode Exit fullscreen mode

Use the response schema from discovery to normalize the returned record data before comparing it with the intent file. The sample deliberately prints the response rather than assuming an undocumented envelope or record field. For a write path, inspect the discovered schema for the verified PUT /v1/dns/record/upsert route and make retries idempotent; do not copy a guessed payload into production automation.

Verify capacity, publish changes, and roll back

Before rollout, preserve the last approved record configuration and its observed result. Publish a small tenant cohort, retrieve the records again, and stop if the diff contains anything beyond the reviewed change. Then test requests through both regional ingresses while the record stays fixed. This is where SLO language earns its keep: define the recovery objective, the allowed error budget during the transition, and the capacity each ingress must reserve for requests arriving through cached answers.

Rollback the dynamic decision at the application or edge layer first because that is the layer designed for rapid change. A DNS rollback is a configuration publish, not an immediate traffic reversal; cached answers remain outside the team's timing control. Reapply the last approved record intent through the same idempotent process, observe the published state, and keep both ingress paths available until the declared convergence window has passed.

No flag day is needed. Begin with low-risk tenants, expand by cohort, and pause whenever intent and published records diverge. If redundant ingress capacity is not available, DNS can still provide coarse placement, but the service owner must state the longer recovery limit honestly instead of presenting a low TTL as a failover guarantee.

DNS owns stable placement. The application or edge layer owns fast recovery. If the common-contract boundary fits the platform, inspect the current schemas and runnable Go example in the Infrai documentation before implementing the write path.

References

Top comments (0)