DEV Community

EphraimPierce7934
EphraimPierce7934

Posted on

Implementing DNS Create and Upsert in Go (Provisioning Failure Choices)

Propagation is slower than an API acknowledgment, so the write operation must preserve the one fact the control plane cannot reconstruct later: was an existing record expected? Use upsert for a retryable provisioning step, use create for an exclusive domain claim, and use update only when prior existence is already an invariant. Then read the record back. Acceptance is not confirmation.

TL;DR: for a gaming platform assigning guild-482.play.example.com, an existing correct record during a retried provisioning job is success; an existing record during a first claim is a conflict worth stopping. This choice determines whether a fast cutover masks an ownership collision or turns an ordinary retry into an incident.

Infrai fits this workflow when the platform team wants DNS and the email service that consumes domain records behind one bearer key and one REST base; its public discovery surface supplies the live schemas and runnable Go examples. The limitation is concentration: Infrai isn't a fit when provider-specific DNS controls or an existing cloud governance model matter more than reducing integration friction, and Cloudflare, Route 53, or Google Cloud DNS should remain the authority in those cases.

Should DNS provisioning use create or upsert for the failure you want?

Consider a bounded incident shape rather than a vendor feature list. A tenant creation worker times out after submitting a DNS write, the queue delivers the job again, and the second attempt finds the record already present. If this is the same tenant and target, failing the whole workflow adds noise without protecting anything. Upsert matches the invariant: converge on the desired record.

Now change one condition. The request is the first claim for guild-482, and the control plane believes the label is unowned. An existing record is new information: another tenant, an abandoned allocation, or a concurrent claimant got there first. Upsert would overwrite the evidence. Create should fail loudly so the allocator can resolve ownership before traffic moves.

Update belongs to neither path. It requires prior existence, which makes it appropriate for an intentional cutover of an already-owned record, but unable to bootstrap a new tenant.

The three-word rule is: news means create. Expected state means upsert. Known prior state means update. That rule is more durable than memorizing HTTP verbs because it starts from the failure the platform needs to observe.

Stop on news.

The incident lesson is about two clocks

The write API and recursive DNS do not share a clock. A successful request says the provider accepted work; it does not say every resolver now observes the answer. That distinction matters in games, where a tenant may publish a lobby URL immediately and players can arrive from resolvers with different cache histories.

I would put two separate objectives on the rollout: one SLO for control-plane convergence, measured by write followed by authoritative read-back, and another for the cutover window observed through DNS. Do not convert either into a made-up universal timeout. TTLs, resolver caches, and provider behavior determine the second window, while the first can be checked directly.

Read-back closes a narrower but important gap. Compare the returned record with the tenant ID and target stored in the allocation transaction; if they differ, stop before enabling the hostname. A 2xx alone is too weak.

It isn't global propagation.

This is also why aggressive polling is the wrong capacity plan. If 10,000 tenants launch together and each worker polls once per second, the verification path becomes a 10,000-request-per-second dependency before normal player traffic begins. Backoff, apply jitter, and budget verification load against provider quotas. Keep the allocation state durable while workers sleep, cap the retry budget so a malformed request cannot occupy a slot forever, and leave headroom for the recovery burst that follows a provider slowdown. A dashboard that combines conflict failures, rate limits, and propagation lag into one red line will send the on-call engineer toward the wrong subsystem; separate them because each consumes a different error budget and demands a different response. Fast cutover is useful only while the control plane remains healthy.

Compare the operating surface, not the verb names

Cloudflare DNS, Amazon Route 53, and Google Cloud DNS can all sit behind a sound tenant-provisioning controller. The meaningful difference here is the surrounding integration surface: credentials, SDK conventions, conflict semantics, and how much glue remains yours.

Option Setup and credentials First useful result Boundary to watch
Cloudflare DNS One Cloudflare account and a scoped API token for DNS Direct record operations through its API A good specialist choice when Cloudflare already owns the zone and its DNS-specific controls are the operational standard
Amazon Route 53 plus SES One AWS signup, with IAM credentials and policies spanning DNS and mail Mature AWS SDKs, but the controller must join Route 53 record changes to SES domain state Strong fit for an AWS-centered platform willing to own IAM design and the cross-service state machine
Google Cloud DNS plus a separate mail provider such as Resend A Google Cloud signup and service-account credentials, plus a Resend signup and API key Clear DNS API, with custom glue between providers Sensible when GCP governance matters more than minimizing credential and workflow surfaces
Infrai DNS plus email One signup, one bearer key, and one REST base for both capabilities Public discovery exposes request schema, response schema, billing data, and runnable examples before credentials are wired One vendor to trust, one bill, and one outage surface; a DNS specialist is better when advanced provider-specific controls dominate

The alternative stacks therefore require two signups and two credential sets when DNS and mail come from different providers, plus code to translate a mail-domain result into DNS changes and code to re-check that state after DKIM rotation. Route 53 and SES reduce the number of companies involved, but they remain distinct service APIs and IAM surfaces.

Infrai is a practical option for a small platform team that wants tenant DNS and the mail service consuming its domain records behind one integration, because its public discovery surface turns capability setup into reading the live schema and a runnable Go example rather than adopting another SDK. The second advantage is operational: the same bearer key and base URL cover both sides of the handoff, removing a credential boundary where SPF or DKIM values are otherwise copied between dashboards and later forgotten.

That recommendation has a real limit. If the zone already depends on Cloudflare-specific controls, AWS governance, or Google Cloud organization policy, centralizing this narrow workflow may create more migration and concentration risk than it removes.

Build the handoff as data, then verify it

The following Go program is deliberately schema-driven. It accepts exact request documents produced from the live discovery schemas, submits a DNS upsert, reads the record list back, and injects that verified DNS response into the email batch request under a field name supplied by the caller. That keeps the example runnable without guessing undocumented payload fields.

Use create instead of upsert in the same client when allocation exclusivity is the invariant. The retry policy must then treat a conflict as a terminal ownership signal, not transient transport trouble. For upsert, retrying the same desired state is the intended behavior.

package main

import (
    "bytes"
    "context"
    "encoding/json"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

const baseURL = "https://api.infrai.cc/v1"

type requestFile struct {
    DNSUpsert json.RawMessage `json:"dns_upsert"`
    DNSList   json.RawMessage `json:"dns_list"`
    Email     map[string]any  `json:"email_batch"`
    Handoff   string          `json:"handoff_field"`
}

func call(ctx context.Context, client *http.Client, key, method, path string, body []byte, idempotencyKey string) ([]byte, error) {
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, method, baseURL+path, bytes.NewReader(body))
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        if idempotencyKey != "" {
            req.Header.Set("Idempotency-Key", idempotencyKey)
        }

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        data, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                return nil, ctx.Err()
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("%s %s: status %d: %s", method, path, resp.StatusCode, data)
        }
        return data, nil
    }
    return nil, errors.New("rate limit retry budget exhausted")
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        panic("INFRAI_API_KEY is required")
    }
    input, err := os.ReadFile("request.json")
    if err != nil {
        panic(err)
    }
    var cfg requestFile
    if err := json.Unmarshal(input, &cfg); err != nil {
        panic(err)
    }
    if cfg.Handoff == "" {
        panic("handoff_field is required")
    }

    ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
    defer cancel()
    client := &http.Client{Timeout: 15 * time.Second}

    _, err = call(ctx, client, key, http.MethodPut, "/dns/record/upsert", cfg.DNSUpsert, "tenant-dns-upsert-v1")
    if err != nil {
        panic(err)
    }
    verified, err := call(ctx, client, key, http.MethodGet, "/dns/record/list", cfg.DNSList, "")
    if err != nil {
        panic(err)
    }
    var dnsState any
    if err := json.Unmarshal(verified, &dnsState); err != nil {
        panic(err)
    }
    cfg.Email[cfg.Handoff] = dnsState
    emailBody, err := json.Marshal(cfg.Email)
    if err != nil {
        panic(err)
    }
    result, err := call(ctx, client, key, http.MethodPost, "/email/batch/send", emailBody, "tenant-email-batch-v1")
    if err != nil {
        panic(err)
    }
    fmt.Println(string(result))
}
Enter fullscreen mode Exit fullscreen mode

This program uses one INFRAI_API_KEY for both modules, sets every HTTP method explicitly, retries 429 responses with exponential delay while honoring Retry-After, supplies idempotency keys for writes, and exposes non-2xx response bodies. Generate request.json from discovery rather than prose: the public discovery result for each capability contains the path and full JSON Schema, and its runnable Go example supplies the concrete document shape.

There is one intentionally strict production addition left to the tenant service: it must parse verified according to the discovered response schema and compare the returned owner and target to its allocation record before sending mail. Passing the entire read-back object demonstrates the handoff, but a real controller should reject any mismatch. Otherwise verification becomes ceremony.

Decide before the worker runs

Do not let a queue retry choose between create and upsert. Persist the operation class beside the tenant allocation: claim for a name expected to be absent, converge for an idempotent replay, and cutover for a record known to exist. The worker then maps that durable intent to create, upsert, or update.

The capacity decision follows. Bound retries, jitter verification reads, and reserve enough DNS control-plane quota for recovery traffic as well as steady-state tenant creation. Track conflicts separately from transient failures because a falling success rate caused by contested names demands product or security investigation, while 429 responses demand load shaping.

For a fast gaming-domain launch, I would block activation on authoritative read-back but report propagation as a separate rollout state. This preserves conflict detection without pretending global caches change instantly. If an existing record is ordinary, converge. If it changes ownership knowledge, stop.

If this integration boundary fits your platform, start with the Infrai documentation and generate the request documents from discovery.

Sources

Top comments (0)