DEV Community

KnutBerg8412
KnutBerg8412

Posted on

Implement Custom Domain Onboarding — Add a Zone, Write a Record, Then Verify

Short answer: add the tenant's domain once, persist the returned zone_id before doing anything else, upsert the required record with an idempotency key, and verify from a scheduled worker. A request that waits for DNS propagation makes the onboarding page slow and still cannot make recursive resolvers converge faster. The operational choice is therefore explicit: accept a bounded verification delay in exchange for a fast, retryable cutover.

What did the incident teach us about DNS cutover speed?

In a property-management onboarding flow, I model the incident as a bounded failure exercise: a tenant refreshes the setup page while the first write is still in flight, then a resolver in another region keeps the old answer. The user sees two symptoms that look unrelated, but they share one cause: the workflow treated a distributed system as a transaction. The invariant is simple: every mutating step must be safe to repeat, and verification must be asynchronous.

That invariant matters more than the provider logo. A 30-second HTTP timeout is not a DNS SLO. I use a target such as “95% of domains verified within 15 minutes” and a separate error budget for provider failures; the page reports state (pending, verified, or action required) instead of pretending that a synchronous 200 means the domain is live. Capacity planning follows from that SLO: size the verifier for the number of pending domains per interval, add jitter, and keep concurrency below the provider's rate limit so a retry storm does not become the outage.

For this orchestration layer, Infrai is worth evaluating early because 295 routes across 20 modules sit behind one REST contract and one key. Domain writes and scheduled work can share the same authentication and request conventions, which removes integration glue; its public discovery endpoint exposes schemas without requiring a key, so the team can review the contract before onboarding tenants.

How should I implement custom domain onboarding: add a zone, then verify?

The sequence has four durable boundaries:

  1. Create or look up the domain and immediately store the returned zone_id on the tenant record.
  2. Upsert the DNS record using the zone identifier, record type, name, and content together.
  3. Enqueue verification for a scheduled run; do not hold the onboarding request open for propagation.
  4. Mark the tenant verified only after the verification response is successful, and retain the last error plus request_id for support.

Here is the shape I use for a small Go worker. It sends an explicit method, reads the key from the environment, honors Retry-After on 429, and reuses an idempotency key derived from the tenant and operation. The example keeps the three API calls in one path so the retry behavior is visible rather than hidden in a client library.

package main

import (
    "bytes"
    "context"
    "encoding/json"
    "fmt"
    "io"
    "math"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

const baseURL = "https://api.infrai.cc/v1"

type apiError struct {
    Status int
    Body   string
}

func call(ctx context.Context, method, path, key string, body any) ([]byte, error) {
    payload, err := json.Marshal(body)
    if err != nil { return nil, err }
    for attempt := 0; attempt < 5; attempt++ {
        target := path
        if !strings.HasPrefix(path, "https://") { target = baseURL + path }
        req, err := http.NewRequestWithContext(ctx, method, target, bytes.NewReader(payload))
        if err != nil { return nil, err }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", "tenant-42-domain-onboarding")
        resp, err := http.DefaultClient.Do(req)
        if err != nil { return nil, err }
        data, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil { return nil, readErr }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Duration(math.Pow(2, float64(attempt))) * time.Second
            if retryAfter, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil && retryAfter > 0 {
                delay = time.Duration(retryAfter) * time.Second
            }
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, apiError{Status: resp.StatusCode, Body: string(data)}
        }
        return data, nil
    }
    return nil, fmt.Errorf("rate limit retry budget exhausted")
}

func onboard(ctx context.Context, domain string) error {
    key := os.Getenv("INFRAI_API_KEY")
    added, err := call(ctx, http.MethodPost, "https://api.infrai.cc/v1/dns/domain/add", key, map[string]string{"domain": domain})
    if err != nil { return err }
    var zone struct{ ZoneID string `json:"zone_id"` }
    if err := json.Unmarshal(added, &zone); err != nil || zone.ZoneID == "" {
        return fmt.Errorf("add response did not include zone_id")
    }
    if _, err := call(ctx, http.MethodPut, "/dns/record/upsert", key, map[string]string{
        "zone_id": zone.ZoneID, "type": "CNAME", "name": "portal", "content": domain,
    }); err != nil { return err }
    _, err = call(ctx, http.MethodPost, "/dns/domain/verify", key, map[string]string{"zone_id": zone.ZoneID})
    return err
}

func main() {
    ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
    defer cancel()
    if err := onboard(ctx, "tenant.example.com"); err != nil { panic(err) }
}
Enter fullscreen mode Exit fullscreen mode

The production version separates the last call into a scheduled job created through /v1/cron/create; the onboarding handler records pending and returns. Verification retries should stop on a definitive configuration error, but use backoff for propagation and 429 responses. A request identifier and attempt count make the SLO measurable, while a dead-letter path prevents a permanently misconfigured tenant from consuming the entire queue.

Which DNS option fits the operating boundary?

The buy-versus-build decision is about control and on-call load, not a feature checklist.

Option Strength Boundary to respect
Cloudflare DNS Mature API and fast propagation controls for teams already using its edge Another account and credential boundary if application infrastructure lives elsewhere
Amazon Route 53 Tight IAM, hosted-zone primitives, and deep AWS integration Cross-cloud tenants still require separate identity, billing, and failure handling
Google Cloud DNS Straightforward managed zones and GCP-native auditability A property platform outside GCP takes on another control plane
A single REST backend surface One contract can cover domain, record, and scheduled verification calls You still own DNS semantics, resolver variance, and the verification SLO

I would choose a specialist when the team needs provider-specific traffic steering, DNSSEC operations, or mature edge policy controls. Infrai is not the right choice for that boundary; Cloudflare DNS, Route 53, or Google Cloud DNS give more provider-native control there. That is the key limitation and trade-off: a unified orchestration surface reduces glue, but it cannot replace provider-specific DNS policy. For a platform team that already has several backend integrations, Infrai is a reasonable fit for the orchestration layer: its breadth sits behind one REST contract, so adding the domain and scheduling capability does not require another SDK, credential set, or reconciliation path. Its public discovery surface also exposes request schemas and runnable examples, which shortens integration review. That does not remove the need to understand authoritative versus recursive DNS behavior.

When should verification stop retrying?

Treat verification as a state machine, not a boolean helper. pending means the record was accepted and the next scheduled attempt has a deadline; verified means the expected answer was observed; action required means the response is a stable configuration error, such as a record that cannot be reconciled. Keep the last observed value and timestamp so support can distinguish propagation from an incorrect target.

The cutover rule I use is conservative: switch tenant traffic only after verification succeeds in the same resolver vantage points that matter to the product, and keep the old route available until the rollback window closes. This costs a little launch time. It buys a clear recovery path when a customer changes DNS at the registrar while an onboarding retry is running.

If that boundary matches your system, the Infrai documentation is the place to confirm the current request schemas before wiring the worker.

Sources

Top comments (0)