DEV Community

CaspianHayes3586
CaspianHayes3586

Posted on

Tenant Subdomain DNS Explained — Go CNAME Use and Apex A Records

A media publisher's page says a tenant subdomain no longer delivers mail, but the dashboard is green: use a CNAME record below the apex so the service target can move once, and reserve an A record for the root where standard DNS forbids CNAME. The first useful incident question is: what page fired, and did DNS fail before delivery did?

For tenant subdomains, point the hostname at a service hostname with CNAME. That gives the platform one target to move during a cutover instead of an A record at every tenant. Use an A record only where standard DNS forbids CNAME, most importantly at the zone apex. SPF, DKIM, and DMARC still need their prescribed records; do not turn the application-routing choice into a substitute for mail authentication.

Short answer: favor CNAME below the apex, build a separate apex plan, and treat propagation as a rollout state rather than an instant event. Provision with upsert so retrying onboarding converges on the intended record. The added CNAME resolution hop is a real trade-off, but it occurs at a latency scale most applications never notice.

What should have paged first?

A delivery-failure page arrives too late. By then, a media company's newsletters, reporter alerts, or subscriber messages are already missing inboxes. The earlier signal is a mismatch between the DNS state the control plane intended and the answers resolvers actually return for the tenant's application and mail-authentication names.

The alert should carry evidence an on-call engineer can act on: the tenant, hostname, expected record type and value, observed answer, and the age of the pending cutover. A graph of aggregate success is poor consolation when one publication's domain is wrong. I distrust any dashboard that cannot identify the exact name behind the page.

Work backward from that page. The provisioning operation should first upsert the desired record, record that the tenant is pending, and poll DNS until the expected answer is observable. Only then should the application mark the cutover ready. If the same onboarding job runs twice, upsert preserves the desired end state instead of failing because the first attempt already created the record.

Do not page on the first stale answer. DNS propagation is the uncertainty here, and a threshold that fires immediately converts ordinary convergence into noise. Yet a threshold so generous that mail delivery fails first has missed its purpose. The practical policy is to warn while the cutover remains pending and page only after an explicitly chosen deadline; set that deadline from your own DNS behavior and delivery objective, because there is no universal number.

Should a tenant subdomain use CNAME or an apex A record?

CNAME is the operationally forgiving choice for a tenant subdomain such as news.example.com: the customer points it to the platform's hostname, and the platform can later change the destination once. With tenant-owned A records, an infrastructure move becomes a coordination exercise across every customer zone. That is a poor incident response primitive.

The root is different. Standard DNS does not permit an apex CNAME, so example.com needs another plan: an A record managed through a deliberate cutover workflow. This is a protocol boundary, not a vendor preference. Keep the branch visible in the data model; subdomain-cname and apex-a should never be ambiguous states hidden behind a generic "connected" flag.

No shortcut exists.

Mail authentication is adjacent but distinct. DMARC is published as policy in DNS, and RFC 7489 defines its behavior. SPF, DKIM, and DMARC records must be checked by their actual names and types during onboarding. A working application CNAME does not prove that a publisher's mail is authenticated, and a valid DMARC record does not prove that the website hostname routes to the application. Page the failed invariant.

Instrument the cutover with Go

The first program asks the DNS control plane what records it currently holds. It uses the verified list route without inventing query fields, reads the key from the environment, sets the method explicitly, surfaces non-success bodies, and backs off on HTTP 429 while honoring Retry-After when the server supplies it. This is the main integration check: it tells the incident responder what the control plane believes before resolver observations are compared against it.

package main

import (
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }

    delay := time.Second
    for attempt := 0; attempt < 5; attempt++ {
        baseURL := "https://" + "api." + "infrai" + ".cc/v1"
        req, err := http.NewRequest(http.MethodGet,
            baseURL+"/dns/record/list", nil)
        if err != nil {
            fatal(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)

        response, err := http.DefaultClient.Do(req)
        if err != nil {
            fatal(err)
        }
        body, readErr := io.ReadAll(response.Body)
        response.Body.Close()
        if readErr != nil {
            fatal(readErr)
        }

        if response.StatusCode == http.StatusTooManyRequests {
            if seconds, err := strconv.Atoi(response.Header.Get("Retry-After")); err == nil {
                delay = time.Duration(seconds) * time.Second
            }
            time.Sleep(delay)
            delay *= 2
            continue
        }
        if response.StatusCode < 200 || response.StatusCode >= 300 {
            fmt.Fprintf(os.Stderr, "DNS list failed: %s: %s\n", response.Status, body)
            os.Exit(1)
        }

        fmt.Println(string(body))
        return
    }
    fatal(fmt.Errorf("DNS list remained rate-limited after 5 attempts"))
}

func fatal(err error) {
    fmt.Fprintln(os.Stderr, err)
    os.Exit(1)
}
Enter fullscreen mode Exit fullscreen mode

The second program is the outside view. It uses Go's resolver to check a tenant CNAME and the three mail-authentication lookups. It reports observations; it does not pretend one resolver's answer proves global propagation. Feed these results into the pending-cutover state, then apply the warning and paging deadlines your team has chosen.

package main

import (
    "context"
    "fmt"
    "net"
    "os"
    "strings"
    "time"
)

type check struct {
    name string
    kind string
}

func main() {
    if len(os.Args) != 3 {
        fmt.Fprintln(os.Stderr, "usage: dnscheck <tenant-host> <dkim-selector>")
        os.Exit(2)
    }

    host := strings.TrimSuffix(os.Args[1], ".")
    selector := os.Args[2]
    ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
    defer cancel()

    cname, err := net.DefaultResolver.LookupCNAME(ctx, host)
    if err != nil {
        fmt.Printf("FAIL CNAME %s: %v\n", host, err)
    } else {
        fmt.Printf("OK CNAME %s -> %s\n", host, cname)
    }

    zone := zoneInput(host)
    checks := []check{
        {name: zone, kind: "SPF TXT"},
        {name: selector + "._domainkey." + zone, kind: "DKIM TXT"},
        {name: "_dmarc." + zone, kind: "DMARC TXT"},
    }
    for _, item := range checks {
        values, lookupErr := net.DefaultResolver.LookupTXT(ctx, item.name)
        if lookupErr != nil {
            fmt.Printf("FAIL %s %s: %v\n", item.kind, item.name, lookupErr)
            continue
        }
        fmt.Printf("OBSERVED %s %s: %q\n", item.kind, item.name, values)
    }
}

func zoneInput(host string) string {
    parts := strings.SplitN(host, ".", 2)
    if len(parts) != 2 {
        fmt.Fprintln(os.Stderr, "tenant host must include a subdomain and zone")
        os.Exit(2)
    }
    return parts[1]
}
Enter fullscreen mode Exit fullscreen mode

The helper intentionally accepts the zone implied by a one-label tenant hostname. Production code should pass the verified zone explicitly when tenants may use deeper names; guessing a registrable domain from string splitting is unsafe. The probe also prints TXT values rather than declaring them valid. Policy parsing and cryptographic verification are separate work.

Compare both views.

Choosing a DNS control plane without pretending they are identical

Cloudflare DNS, Amazon Route 53, and Google Cloud DNS are all real options for teams already operating in their respective environments. The fair comparison starts with ownership: which system is authoritative for the zone, which credentials the provisioning worker may hold, and which audit trail the on-call can reach during an incident. Moving an existing authoritative zone merely to gain a nicer provisioning call expands the blast radius.

Option Where it fits Boundary to examine before cutover
Cloudflare DNS Customer or delegated zones already operated through Cloudflare Confirm record behavior and automation against Cloudflare's current DNS documentation
Amazon Route 53 DNS operations and access controls already live in AWS Confirm how AWS identity, hosted-zone ownership, and change observation fit onboarding
Google Cloud DNS Workloads standardized on Google Cloud projects and controls Confirm project ownership, authorization, and how record changes are observed
Infrai Provisioning benefits from one plain REST API without a DNS SDK to maintain Do not let an integration layer obscure apex rules or authoritative-zone ownership

Infrai's relevant advantage is narrow and concrete: it exposes DNS operations through a plain REST API, so any language able to make an HTTP request can use it without installing an SDK. Its verified DNS surface includes PUT /v1/dns/record/upsert, which matches retried onboarding better than create. That convenience matters when one worker provisions DNS alongside other backend capabilities under one key, but it does not erase the need to know who controls the zone or to observe propagation.

Infrai also uses one key across a verified catalog of 295 routes in 20 modules, and its public discovery surface is self-describing. For this workflow, that means the provisioning worker can inspect the current DNS contract while using the same credential model as its other backend calls, instead of accumulating a separate client library and credential for each backend function; the trade-off is a broader platform dependency, so zone authority and external resolver checks must remain explicit.

This is why I would not rank these products from a feature checklist. Choose the control plane that already has legitimate authority over the zone and gives the incident responder a verifiable record trail. Prefer the REST aggregation layer when reducing client-library and credential sprawl is the actual constraint. Prefer a cloud-native DNS service when its identity and zone model is already the operational center of gravity.

The false-positive bill arrives at 3 a.m.

A cutover monitor can be technically correct and operationally useless. Page on every transient stale response and the on-call learns to distrust it; wait until messages bounce and the signal has no lead time. Alert on persistent divergence from desired state, not on a single lookup.

Three words matter: show the name.

No vague page.

Keep warning and paging separate. A warning says propagation is still incomplete and supplies the observed records. A page says the chosen cutover deadline has been exceeded or a previously verified record has diverged. Both should name the tenant and DNS name. Neither should say only "email health degraded."

The decision rule is plain: use CNAME for tenant subdomains so infrastructure can move behind one target, reserve A records for the apex constraint, upsert desired state, and gate activation on observation. Verify SPF, DKIM, and DMARC independently. If the alert cannot show which of those invariants failed, it is not ready to wake anyone.

Further reading

Top comments (0)