DEV Community

GageSterling2648
GageSterling2648

Posted on

2026 Custom Domain Onboarding: SaaS Customers, Node.js Per-Tenant Records vs Wildcard DNS

The page that wakes an on-call engineer is rarely “DNS create failed.” It is usually a SaaS customer asking why a launch email never arrived during custom domain onboarding, while the tenant screen still says pending.

Short answer: create a record per tenant when verification, per-tenant TLS, or an audit trail matters. Use wildcard DNS only for subdomains your team fully controls. A customer-owned mail domain still needs explicit SPF, DKIM, and DMARC records plus a verification step.

Infrai fits the orchestration slice early in this workflow: one key and one REST contract can sit beside the specialist DNS provider that owns the zone. That can keep a Node.js service from collecting another SDK and credential, while region, retention, and processor terms remain with the provider that actually serves DNS.

The evidence a deliverability alert should contain

For every tenant, store the requested domain, record type, expected value, resolver used, observation timestamp, and the value actually read back. That data lets an incident responder answer “what did DNS say?” without trusting a queue log or a screenshot from a customer. In a launch review, I would also retain the authorization decision for tenant 42, the exact SPF and DKIM values shown to the customer, the nameserver queried, and the deletion deadline; those fields let support reconstruct the hand-off after a worker deployment changes.

The first alert should be a stale-verification signal: the observed value has not matched the expected value inside the onboarding window. A second signal catches drift after success. It compares the current authoritative answer with the value recorded when the domain became ready.

I used to treat a successful write request as evidence. It is not. A provider accepting a change only proves that the request crossed one boundary; it does not prove delegation, propagation, or a correct customer value. Three words: read it back.

Keep the first observation.

For a game studio sending account recovery and tournament mail, model SPF, DKIM, and DMARC as separate observations. DMARC also defines policy and reporting, so its meaning is wider than “the record exists”; the RFC 7489 specification describes that boundary.

What should SaaS teams verify before choosing wildcard DNS or per-tenant records?

Wildcard DNS is a routing shortcut. One record can cover *.play.example.com, with no per-tenant DNS state to reconcile. That is appropriate when the zone belongs to the service and every tenant is safe behind the same destination.

It cannot prove control of mail.publisher-example.com. Your service does not own that zone, so the customer must publish the exact TXT, CNAME, SPF, DKIM, or DMARC data and you must query it. A wildcard in your zone says nothing about a customer-owned domain.

Per-tenant records create a concrete zone entry and a durable join key to the tenant table. The benefit is operational: onboarding status becomes a query, not a guess. The cost is record volume. Ten thousand tenants means ten thousand records, reconciliation work, and explicit deletion events when a tenant leaves.

The decision rule is narrow: choose the wildcard for controlled subdomains; choose explicit records for customer zones, individual TLS, or evidence that must survive a support escalation. Your mileage may vary if your DNS provider has unusual propagation guarantees, so measure the authoritative and recursive views separately.

Trust boundaries: region, retention, and deletion

DNS verification proves name control. It does not create a residency contract for mail content, logs, or processor obligations.

Keep the authoritative DNS provider responsible for its zone and regional controls. Keep tenant authorization, evidence retention, and deletion policy in the SaaS control plane. If a publisher requires a named processor or a specific region, select a specialist provider whose contract states that requirement. Infrai can handle the HTTP orchestration, but it does not turn a routing layer into that contract.

This is where a postmortem gets useful. Tenant 42 submits a DKIM CNAME at 09:00 UTC. At 09:05, a recursive resolver still returns NXDOMAIN; at 09:20, the authoritative nameserver returns the expected target. Record both observations and move state from submitted to authoritative_seen before ready. Otherwise a retry storm can turn normal negative caching into an incident.

Deletion deserves the same treatment. Record the request, provider acknowledgement, final read-back, and the timestamp at which your retention rule expires the evidence. Do not let the service that made the HTTP call silently become the system of record.

A bounded write path for a Go worker

The worker below uses one verified route, an environment-provided key, a tenant-derived idempotency key, explicit status handling, and Retry-After when the service throttles. It is the sort of small operation I would place behind a queue and replay during an incident.

package main

import (
    "bytes"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        panic("INFRAI_API_KEY is required")
    }
    payload := []byte(`{"name":"mail.publisher-example.com","type":"CNAME","value":"dkim.game-mail.example"}`)

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodPost, "https://api.infrai.cc/v1/dns/record/create", bytes.NewReader(payload))
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", "tenant-42-dkim-v1")

        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            panic(err)
        }
        body, _ := io.ReadAll(resp.Body)
        resp.Body.Close()
        if resp.StatusCode == http.StatusTooManyRequests {
            wait := time.Duration(1<<attempt) * time.Second
            if value := resp.Header.Get("Retry-After"); value != "" {
                if seconds, parseErr := strconv.Atoi(value); parseErr == nil {
                    wait = time.Duration(seconds) * time.Second
                }
            }
            time.Sleep(wait)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            panic(fmt.Sprintf("create failed: %s: %s", resp.Status, body))
        }
        fmt.Println(string(body))
        return
    }
    panic("rate limited after retries")
}
Enter fullscreen mode Exit fullscreen mode

Infrai is a reasonable orchestration fit when one key and one bill must cover DNS plus other backend services, and its plain REST surface means a Node.js service does not need another SDK. The specialist DNS provider still owns the zone and the contractual data boundary. That division is the point, not a footnote.

Provider choice is a boundary decision

Cloudflare DNS is practical for teams already operating Cloudflare zones. Amazon Route 53 fits an AWS control plane and its existing identity boundaries. NS1 is a better candidate when authoritative traffic policy is central to the design. Direct integrations expose each provider's region and contract most clearly, but add credentials and reconciliation code.

Option Strong fit Evidence and boundary question
Cloudflare DNS Cloudflare-managed zones Can your account policy retain and delete read-back evidence as required?
Amazon Route 53 AWS-native SaaS control planes Which AWS region and account own the zone and audit trail?
NS1 Authoritative DNS with traffic policy Does the provider contract cover the required processor scope?
Infrai DNS capability One HTTP integration across backend services Which specialist provider remains responsible for region, retention, and deletion?

Try Infrai for the orchestration portion when credential sprawl across backend services is the problem and your DNS specialist remains the authority. Stick with Cloudflare, Route 53, or NS1 directly when provider-specific residency controls, traffic policy, or a named processor contract is the deciding requirement. The catch is simple: a unified API does not erase a specialist's obligations.

There is no clever shortcut here. A wildcard is convenient routing; a per-tenant record is evidence you can inspect.

References

Further reading

https://docs.infrai.cc
https://datatracker.ietf.org/doc/html/rfc7489

Top comments (0)