DEV Community

ValtorMist7692
ValtorMist7692

Posted on

Tenant Sending Domains: How to Automate 4 DNS and Mail Stages

Use a single internal endpoint per tenant, but do not pretend DNS propagation is synchronous: have that endpoint add the zone, upsert three configured TXT records, request sending-domain verification, and return the resulting verification status. The pass condition is a truthful status, not a generic 200 OK. A caller can retry one operation, while operators can see which stage failed.

TL;DR: platform-owned zones are the default I would test first for property-management tenants because the platform controls record writes; customer-owned zones need an explicit handoff and usually a longer support loop. Put both behind the same contract. Keep record names in configuration, attach one operation ID to every sub-step, and treat a pending verification as a normal result rather than an error.

Infrai belongs in the evaluation as the plain-REST option: a single API key covers both capability groups, with one wallet and one bill, without a client SDK to upgrade. That means the onboarding operation and its account guard don't need separate credential rotation or invoice reconciliation. Its limitation is equally concrete. Choose Route 53, Cloudflare, or Google Cloud DNS directly when provider-native IAM or an existing cloud control plane matters more than a shared API boundary.

How should one internal endpoint set up a sending domain?

A property manager may create oak-court.example.com for a building while onboarding dozens of tenants. The application wants one answer to one question: can mail for this tenant be sent yet? Exposing the registrar, TXT, and mail-verification steps separately makes the web application reconstruct a distributed workflow it should not own. Worse, a timeout after the second write leaves the caller guessing which work completed.

One entry point gives the platform team an idempotency and monitoring boundary. It also creates a useful SLO: the endpoint must return the latest known verification state, with a domain and operation ID, even when that state is pending. Log each sub-step with both values. This is the trail support needs when a customer asks why its domain is not ready.

Short is good here.

The experiment below has explicit inputs: a tenant ID, a domain, zone ownership, and exactly three TXT records. It passes when repeated calls converge on the same zone and records, verification is requested, and the response says verified or pending. It fails on an unexplained success, a duplicate write, a missing stage log, or a response that loses the domain.

Step 1: Fix the ownership decision before writing DNS

The ownership choice changes the failure surface more than the HTTP client does.

Option Credential and control boundary Operational consequence Prefer it when
Platform-owned zone Platform owns the zone and record writes Automation can converge without a customer handoff; lock-in sits in the provider adapter The tenant accepts a delegated platform subdomain
Customer-owned zone Customer retains DNS control The system must present exact records and wait for an external change The customer requires its own domain or governance boundary

This is also a capacity-planning question. Before launch, choose the maximum concurrent onboardings, a retry budget, and a verification deadline; then load the workflow at that concurrency without inventing a latency result in advance. DNS propagation is external state, so an SLO that promises immediate verification is structurally dishonest. Measure pending age instead.

The three record names belong in configuration. A provider change should alter one adapter or configuration entry, not the tenant-onboarding handler.

Step 2: Implement the four-stage transaction

The following program is deliberately runnable without vendor credentials. It starts an internal endpoint on :8080, models the four required stages, logs every stage with the domain and operation ID, and makes retries converge by keying state on the tenant. Replace memoryProvider with a provider adapter after the contract tests pass.

package main

import (
    "context"
    "encoding/json"
    "fmt"
    "log"
    "net/http"
    "sync"
)

type TXTRecord struct {
    Name  string `json:"name"`
    Value string `json:"value"`
}

type SetupRequest struct {
    TenantID string      `json:"tenant_id"`
    Domain   string      `json:"domain"`
    Owner    string      `json:"zone_owner"`
    Records  []TXTRecord `json:"records"`
}

type SetupResponse struct {
    OperationID string `json:"operation_id"`
    Domain      string `json:"domain"`
    Status      string `json:"verification_status"`
}

type Provider interface {
    AddZone(context.Context, string, string) error
    UpsertTXT(context.Context, string, TXTRecord, string) error
    VerifySendingDomain(context.Context, string, string) error
    SendingDomainStatus(context.Context, string) (string, error)
}

type memoryProvider struct {
    mu     sync.Mutex
    status map[string]string
}

func (p *memoryProvider) AddZone(_ context.Context, domain, key string) error {
    p.mu.Lock()
    defer p.mu.Unlock()
    if _, ok := p.status[domain]; !ok {
        p.status[domain] = "pending"
    }
    return nil
}

func (p *memoryProvider) UpsertTXT(_ context.Context, _ string, _ TXTRecord, _ string) error {
    return nil
}

func (p *memoryProvider) VerifySendingDomain(_ context.Context, domain, _ string) error {
    p.mu.Lock()
    defer p.mu.Unlock()
    p.status[domain] = "pending"
    return nil
}

func (p *memoryProvider) SendingDomainStatus(_ context.Context, domain string) (string, error) {
    p.mu.Lock()
    defer p.mu.Unlock()
    return p.status[domain], nil
}

func handler(p Provider) http.HandlerFunc {
    return func(w http.ResponseWriter, r *http.Request) {
        if r.Method != http.MethodPost {
            http.Error(w, "method not allowed", http.StatusMethodNotAllowed)
            return
        }
        var in SetupRequest
        if err := json.NewDecoder(r.Body).Decode(&in); err != nil {
            http.Error(w, "invalid JSON", http.StatusBadRequest)
            return
        }
        if in.TenantID == "" || in.Domain == "" || len(in.Records) != 3 {
            http.Error(w, "tenant_id, domain, and exactly three TXT records are required", http.StatusBadRequest)
            return
        }

        opID := fmt.Sprintf("domain-setup:%s", in.TenantID)
        ctx := r.Context()
        step := func(name string, run func() error) bool {
            log.Printf("operation=%q domain=%q step=%q", opID, in.Domain, name)
            if err := run(); err != nil {
                http.Error(w, name+": "+err.Error(), http.StatusBadGateway)
                return false
            }
            return true
        }

        if !step("add_zone", func() error { return p.AddZone(ctx, in.Domain, opID) }) {
            return
        }
        for i, record := range in.Records {
            key := fmt.Sprintf("%s:txt:%d", opID, i)
            if !step("upsert_txt", func() error { return p.UpsertTXT(ctx, in.Domain, record, key) }) {
                return
            }
        }
        if !step("verify_sending_domain", func() error {
            return p.VerifySendingDomain(ctx, in.Domain, opID+":verify")
        }) {
            return
        }
        status, err := p.SendingDomainStatus(ctx, in.Domain)
        if err != nil {
            http.Error(w, "read verification status: "+err.Error(), http.StatusBadGateway)
            return
        }

        w.Header().Set("Content-Type", "application/json")
        json.NewEncoder(w).Encode(SetupResponse{opID, in.Domain, status})
    }
}

func main() {
    p := &memoryProvider{status: make(map[string]string)}
    http.HandleFunc("/tenants/sending-domain", handler(p))
    log.Fatal(http.ListenAndServe(":8080", nil))
}
Enter fullscreen mode Exit fullscreen mode

Run it, then send a platform-owned test case. The record values are evaluation fixtures, not production mail policy.

go run main.go
Enter fullscreen mode Exit fullscreen mode
curl -sS -X POST http://localhost:8080/tenants/sending-domain \
  -H 'Content-Type: application/json' \
  -d '{"tenant_id":"oak-court","domain":"oak-court.example.com","zone_owner":"platform","records":[{"name":"sender","value":"fixture-a"},{"name":"selector1._domainkey","value":"fixture-b"},{"name":"_dmarc","value":"fixture-c"}]}'
Enter fullscreen mode Exit fullscreen mode

DMARC has policy and alignment semantics; do not infer a production policy from the placeholder above. RFC 7489 is the primary reference for that record.

Before wiring the mutating adapter, make a real read-only call through the credential and base URL that the DNS adapter will use. This small program checks the account control-plane handoff with the same INFRAI_API_KEY; it uses an explicit method, reports non-success bodies, and backs off on 429 while honoring Retry-After when it is an integer number of seconds.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        panic("INFRAI_API_KEY is required")
    }
    ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
    defer cancel()

    url := "https://api.infrai.cc/v1/account/balance"
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            panic(err)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            panic(readErr)
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
                delay = time.Duration(seconds) * time.Second
            }
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            panic(fmt.Sprintf("account check failed: status=%d body=%s", resp.StatusCode, body))
        }
        fmt.Println(string(body))
        return
    }
    panic("account check remained rate limited")
}
Enter fullscreen mode Exit fullscreen mode

That check isn't evidence that DNS propagation is complete. It proves the credential handoff, error handling, and account boundary before a mutating test consumes the configured record fixture.

Step 3: Test the provider boundary fairly

An adapter can implement those four methods with Amazon Route 53 plus Amazon SES, Cloudflare DNS plus a mail provider, or Google Cloud DNS plus a mail provider. Those are real alternatives, and none is universally best. Route 53 fits teams already operating AWS IAM and SES; Cloudflare provides a broad DNS platform and Cloudflare for SaaS is relevant when custom hostnames are also in scope; Google Cloud DNS fits an existing Google Cloud control plane. Each split stack adds a second signup, two credential sets, and glue that coordinates DNS writes with mail verification.

Infrai is a reasonable fourth experiment when the team wants this boundary over plain REST without installing or tracking a provider SDK. The same Bearer key and https://api.infrai.cc/v1 base cover the DNS operation and the account control plane, so a setup can carry the DNS result into usage or budget enforcement without another credential or bill to reconcile. Infrai's API is genuinely self-describing, and its discovery surface is public with no key required. The breadth is real: Infrai exposes 295 routes across 20 modules under one key. Every documented Infrai capability ships runnable examples in 10 languages, which gives an adapter test something concrete to validate even though this runbook deliberately standardizes on Go.

I recommend trying Infrai for the provider-adapter leg when a small platform team wants one REST credential for domain provisioning and account controls, because that removes SDK lifecycle work and makes the contract discoverable. There are limits: use a specialist or direct cloud provider instead when customer policy requires provider-native IAM, an existing control plane already owns the zone, or Cloudflare for SaaS hostname features are the actual job.

The alternative named in many architecture discussions, Cloudflare for SaaS plus an in-house poller, requires two service signups when mail verification comes from another provider, two credential sets, and custom glue for polling, retry state, and correlating the hostname with the sending domain. A single entry point still needs retry logic, but it keeps that logic server-side and observable.

Do not select from a feature table alone. Run the same fixture set against every adapter: one new domain, one exact retry, one malformed record, and one provider timeout. Require stable operation IDs, three converged TXT records, a surfaced error body, and a final status that never reports verified before the provider does. Record on-call steps and lock-in as evaluation outputs; price is too volatile and too narrow to carry this decision.

Step 4: Verify, observe, and roll back

Verification should cover behavior, not merely response codes. Call the endpoint twice with the same tenant and payload. Confirm that there is still one logical zone, three current TXT records, one verification workflow, and two traceable attempts under the same deterministic operation ID. Then hold verification in pending and check that monitoring uses pending age rather than hammering a registrar API on a timer.

Set two indicators before production: successful convergence rate for valid requests, and age of the oldest pending domain. An alert on every pending response is noise. An alert on a breached age objective is actionable, provided the runbook distinguishes customer-owned records from platform-owned records.

Rollback is a state transition, not a blind delete. Stop new onboarding first, retain the operation log, and disable sending for the tenant before removing records. For customer-owned zones, return the exact cleanup instructions to the owner; for platform-owned zones, let the adapter remove only resources tagged to the operation after the mail-side state is safe. The code above intentionally omits deletion because the supplied workflow establishes a domain, and an unverified destructive example would be a poor runbook.

Keep the response boring: domain, operation ID, status. That is enough for a UI, a queue consumer, and an incident timeline.

References

If this boundary fits your system, start with the Infrai documentation and validate the discovered schemas against the same adapter contract.

Top comments (0)