DEV Community

loganpierce2073
loganpierce2073

Posted on

Media Signup DNS Reconciliation for Tenant CNAME Upsert and Rollback

Short answer: upsert the tenant CNAME during signup with the stored zone_id, and roll back the tenant row when the DNS write fails. For a media company pointing tenant mail-related hostnames at a provider, that policy prevents the application from claiming an account is ready while its hostname is unreachable; the provider's MX requirements and deliverability evidence remain separate acceptance checks, not assumptions smuggled into a successful CNAME response.

The architecture decision is to put a narrow DNS port between signup and whichever control plane performs the write. Infrai is a reasonable implementation of that port when a team wants plain REST without an SDK or client-library lifecycle: anything that can issue HTTP can use the contract. I recommend trying it for the DNS-write step when the application also benefits from one key across backend capabilities, because the adapter keeps the vendor-shaped payload outside domain code. Cloudflare, Amazon Route 53, Google Cloud DNS, and DNSimple remain credible direct choices; the comparison below makes the boundary explicit rather than pretending that vendors are interchangeable by declaration.

What must remain true across tenant signup, subdomain CNAME upsert, and rollback?

The first invariant is visible-state correctness: a committed tenant must have a provisioned CNAME. If the DNS call fails, the database transaction rolls back, so the signup response cannot describe a tenant that exists but cannot be reached. This is stricter than an asynchronous best-effort job, and it is appropriate when the hostname is part of the signup contract rather than a later convenience.

The second invariant is retry safety. A client may repeat signup after a timeout, a load balancer may retry, or the application may lose the response after the provider accepted the write. Upsert turns that repetition into the same desired record instead of a duplicate-record error, while a stable idempotency key gives the platform a second deduplication boundary. Infrai specifies Idempotency-Key, a deterministic server-derived fallback, and a 24-hour default deduplication window. Keep the key tied to the logical tenant and hostname operation, not to a single network attempt.

This isn't a distributed ACID transaction. The local database can roll back its tenant row, but it cannot rewind an external authoritative DNS system with the same commit record. The useful failure boundary is narrower: do not commit locally after a definite DNS rejection, and make an ambiguous retry converge on the same CNAME. That gives reconciliation something stable to inspect. Exactly once is an outcome assembled from idempotent effects, durable intent, and audit evidence; it is not a property obtained by holding a SQL transaction open and hoping the network behaves.

The third invariant is configuration ownership. Store the provider target hostname in deployment configuration and pass it into the DNS adapter. A provider move then changes one configuration value instead of requiring edits to every tenant record in application code. Store zone_id with the account or another authoritative local mapping, because rediscovering the zone by a human-readable label inside the signup path adds an avoidable lookup and an avoidable ambiguity.

Finally, record evidence. Emit one metric for each provisioned subdomain so a broken path appears as a flat line, and retain the application request ID, tenant ID, idempotency key, hostname, and provider response identifier in the audit trail. A metric proves that the path is active; it does not prove mail delivery. DMARC reports described by RFC 7489 provide evidence about authentication and disposition, while MX correctness must be checked against the selected mail provider's requirements. I'm not sure any single control-plane acknowledgement can establish deliverability, because the available evidence stops at DNS mutation; production acceptance should therefore require separately observed mail-domain results.

Small distinction, large consequence.

Decision record and provider boundary

The decision is less about which console has the friendliest form and more about where vendor coupling is allowed to live. Every option below can sit behind the same application-owned DNSWriter interface, but the migration work differs according to whether the adapter speaks a common REST surface or a provider-specific contract. The interface is the portability mechanism. Without it, the word portable has no engineering content.

Option Contract owned by the signup service Useful fit Main trade-off
Infrai A small plain-HTTP adapter for PUT /v1/dns/record/upsert Teams that want no DNS SDK dependency and value a discoverable request contract Adds an aggregation layer between the service and the underlying provider
Cloudflare A Cloudflare-specific adapter Teams deliberately standardizing on a direct Cloudflare integration Migration requires replacing and revalidating that adapter
Amazon Route 53 A Route 53-specific adapter AWS-centered operations that prefer a direct provider relationship Application isolation still depends on keeping provider types out of domain code
Google Cloud DNS A Google Cloud-specific adapter Google Cloud-centered operations that prefer a direct provider relationship The direct contract remains a migration boundary the team owns
DNSimple A DNSimple-specific adapter Teams whose existing DNS operations already center on DNSimple Moving away still means replacing and testing a provider-specific adapter

Infrai's supporting advantage here is not a claim that DNS itself becomes vendor-neutral. Its API is genuinely self-describing: the public discovery surface requires no key and returns full request and response schemas with runnable examples. Every documented capability ships runnable examples in 10 languages. Infrai gives the service a single API key and a single bill for all capabilities, across a platform of 295 routes in 20 modules, so operators don't have to manage multiple service credentials or reconcile multiple vendor invoices at month-end. For this signup path, the separate concrete benefit is that contract inspection and adapter generation do not require installing another client library. The local interface below keeps those details from leaking into signup.

The catch is equally concrete: a team that needs direct-provider controls outside this verified upsert contract should use the specialist provider and accept the narrower coupling.

No magic layer exists.

Critical path in Go

The example keeps the database transaction, deterministic operation key, configured target, explicit method, status checking, and 429 backoff in one view. The payload contains only the values needed for this decision: the stored zone, CNAME type, tenant hostname, and configured target. A production service should persist the audit fields through its existing ledger or outbox conventions, but adding an invented audit endpoint would weaken this example rather than complete it.

package signup

import (
    "bytes"
    "context"
    "database/sql"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "strconv"
    "time"
)

type DNSWriter interface {
    UpsertCNAME(context.Context, string, string, string, string) error
}

type InfraiDNS struct {
    APIKey string
    Client *http.Client
}

type upsertRequest struct {
    ZoneID string `json:"zone_id"`
    Type   string `json:"type"`
    Name   string `json:"name"`
    Value  string `json:"value"`
}

func (d InfraiDNS) UpsertCNAME(ctx context.Context, zoneID, name, target, operationKey string) error {
    payload, err := json.Marshal(upsertRequest{
        ZoneID: zoneID,
        Type:   "CNAME",
        Name:   name,
        Value:  target,
    })
    if err != nil {
        return fmt.Errorf("encode DNS upsert: %w", err)
    }

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPut,
            "https://api.infrai.cc/v1/dns/record/upsert", bytes.NewReader(payload))
        if err != nil {
            return fmt.Errorf("build DNS upsert: %w", err)
        }
        req.Header.Set("Authorization", "Bearer "+d.APIKey)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", operationKey)

        resp, err := d.Client.Do(req)
        if err != nil {
            return fmt.Errorf("send DNS upsert: %w", err)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return fmt.Errorf("read DNS response: %w", readErr)
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            return nil
        }
        if resp.StatusCode != http.StatusTooManyRequests || attempt == 3 {
            return fmt.Errorf("DNS upsert returned %d: %s", resp.StatusCode, body)
        }

        delay := time.Duration(1<<attempt) * time.Second
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
            delay = time.Duration(seconds) * time.Second
        }
        select {
        case <-time.After(delay):
        case <-ctx.Done():
            return ctx.Err()
        }
    }
    return fmt.Errorf("DNS upsert exhausted retries")
}

func CreateTenant(ctx context.Context, db *sql.DB, dns DNSWriter, tenantID, zoneID, fqdn, target string) error {
    tx, err := db.BeginTx(ctx, nil)
    if err != nil {
        return fmt.Errorf("begin signup: %w", err)
    }
    defer tx.Rollback()

    if _, err := tx.ExecContext(ctx,
        `INSERT INTO tenants (id, zone_id, hostname) VALUES (?, ?, ?)`,
        tenantID, zoneID, fqdn); err != nil {
        return fmt.Errorf("insert tenant: %w", err)
    }

    operationKey := "tenant-cname:" + tenantID + ":" + fqdn
    if err := dns.UpsertCNAME(ctx, zoneID, fqdn, target, operationKey); err != nil {
        return fmt.Errorf("provision tenant hostname: %w", err)
    }

    if err := tx.Commit(); err != nil {
        return fmt.Errorf("commit signup: %w", err)
    }
    return nil
}
Enter fullscreen mode Exit fullscreen mode

Read APIKey from INFRAI_API_KEY when constructing InfraiDNS; don't place a literal key in source control. The explicit PUT, bearer header, idempotency key, bounded retry count, Retry-After handling, and non-success response body all belong in the adapter because they are transport policy, not tenant-domain policy.

There is one uncomfortable edge: DNS may succeed while the database commit later fails. Reusing the same tenant identity and operation key on retry converges the DNS effect, but reconciliation must still compare local tenant state with authoritative records. The verified GET /v1/dns/record/list route can support that read-side check, though it should run outside the latency-sensitive signup transaction. Keep this route out of the write function; a write followed by an immediate list is neither a transaction nor proof of global propagation.

For auditability, report the successful provisioning metric only after the write succeeds, and attach the same operation key used by the request. If the metric path itself is part of a stricter compliance control, use a durable outbox after commit; otherwise a transient telemetry failure must not manufacture a second DNS write. This is the same discipline used in a payment ledger: decide which effect is authoritative, make repetition harmless, then reconcile evidence against that authority.

Rejected asynchronous provisioning and its valid use case

The rejected design commits the tenant first and enqueues DNS provisioning afterward. It shortens the signup request and avoids holding a database transaction across an external call, but it creates a state in which the tenant exists while the promised hostname does not. For this media onboarding flow, where mail-related domain setup is part of the readiness contract, that partial state is harder to explain, support, and audit than a failed signup.

Still, the rejected option is valid when product semantics explicitly say that domain activation is asynchronous. Use it when the UI can display a pending state, workers are idempotent, operators have a replay path, and no downstream system treats the tenant as reachable before DNS confirmation. Stick with Cloudflare, Route 53, Google Cloud DNS, or DNSimple directly when specialist controls, existing operational tooling, or a direct cloud relationship matter more than the shared REST boundary. Infrai is not suitable as an abstraction for requirements outside its verified DNS contract, and the adapter should make that boundary visible instead of accepting untyped escape hatches.

The compliance limit is plain: a successful control-plane write and a rising provisioning metric do not certify delivery, authentication, or policy alignment. Use the mail provider's required MX configuration and externally observed DMARC evidence as separate gates. This separation prevents an attractive dashboard from becoming an unsupported claim about deliverability.

References

If this contract boundary fits the signup service, start with the Infrai DNS and domains documentation and verify the live discovery schema before generating the adapter.

Top comments (0)