DEV Community

EllisThornton7395
EllisThornton7395

Posted on

Declarative MX Records in Node.js: Priority-Safe Upserts for Mail Routing

Short answer: keep MX intent in versioned configuration, validate priorities and ownership before applying it, then perform an idempotent upsert that treats the complete desired set as the reconciliation target. For a mail-routing onboarding flow, customer-owned zones should require proof of control; platform-owned zones can be managed directly, but they still need the same audit trail.

The bill is rarely dominated by the DNS write itself. It is dominated by the retention around it: stale records, duplicate onboarding attempts, support investigation, and the operational state needed to explain why a message followed one route instead of another. I keep the desired record set, the observed provider response, and the decision that authorized the change. I deliberately stop keeping every polling response after its useful retention window; the trade-off is less forensic detail during a long-running dispute.

What does declarative MX configuration need to prove?

An MX record maps a domain to a mail exchanger and gives that exchanger a preference value. Lower preference values are tried first, while equal values allow a sender to choose among alternatives. That ordering is the mechanism, not a hint for an application-level load balancer.

Configuration should therefore express a set, not an ordered list of API calls. A compact representation might look like this:

type MXIntent struct {
    Name       string
    Exchange   string
    Preference uint16
}

var desired = []MXIntent{
    {Name: "example.customer", Exchange: "mx1.mail.example.net.", Preference: 10},
    {Name: "example.customer", Exchange: "mx2.mail.example.net.", Preference: 20},
}
Enter fullscreen mode Exit fullscreen mode

The trailing dot makes the exchange an absolute DNS name. Normalize case, reject an empty exchange, and reject preferences outside the 0-65535 range before touching a provider. Also reject duplicate (name, exchange, preference) tuples; a duplicate declaration is usually a review mistake, not useful redundancy.

Ownership changes the authorization path. For a customer-owned zone, require a DNS-01 challenge or an equivalent TXT proof and record the challenge identifier before accepting the MX intent. For a platform-owned zone, the control plane can prove authority through its own account boundary, but the resulting change still needs a tenant, actor, and configuration revision in the audit event.

How should Node.js set MX records with priorities from configuration?

The application can be written in Node.js while the reconciliation contract remains language-neutral: parse configuration, canonicalize names, compare the desired set with the observed set, and submit one idempotent operation. The provider adapter is the only layer that knows its wire format. That separation keeps a provider migration from changing payment or onboarding state transitions.

Here is the core policy in Go, using an intentionally generic adapter. It models the operation a Node.js service would call through its own worker or command endpoint without inventing a vendor route.

package dnsreconcile

import (
    "context"
    "fmt"
    "sort"
    "strings"
)

type MXIntent struct {
    Name       string
    Exchange   string
    Preference uint16
}

type DNS interface {
    ListMX(ctx context.Context, zone, name string) ([]MXIntent, error)
    ReplaceMX(ctx context.Context, zone, name string, records []MXIntent, idempotencyKey string) error
}

func canonical(r MXIntent) MXIntent {
    r.Name = strings.ToLower(strings.TrimSuffix(r.Name, "."))
    r.Exchange = strings.ToLower(strings.TrimSuffix(r.Exchange, ".")) + "."
    return r
}

func reconcile(ctx context.Context, dns DNS, zone, name string, desired []MXIntent, revision string) error {
    if zone == "" || name == "" || revision == "" {
        return fmt.Errorf("zone, name, and revision are required")
    }
    want := make([]MXIntent, len(desired))
    for i, record := range desired {
        if record.Preference > 65535 || record.Exchange == "" {
            return fmt.Errorf("invalid MX record at index %d", i)
        }
        want[i] = canonical(record)
    }
    sort.Slice(want, func(i, j int) bool {
        if want[i].Preference != want[j].Preference {
            return want[i].Preference < want[j].Preference
        }
        return want[i].Exchange < want[j].Exchange
    })
    have, err := dns.ListMX(ctx, zone, name)
    if err != nil {
        return fmt.Errorf("read current MX set: %w", err)
    }
    for i := range have {
        have[i] = canonical(have[i])
    }
    sort.Slice(have, func(i, j int) bool {
        if have[i].Preference != have[j].Preference {
            return have[i].Preference < have[j].Preference
        }
        return have[i].Exchange < have[j].Exchange
    })
    if fmt.Sprint(have) == fmt.Sprint(want) {
        return nil
    }
    key := "mx:" + zone + ":" + name + ":" + revision
    return dns.ReplaceMX(ctx, zone, name, want, key)
}
Enter fullscreen mode Exit fullscreen mode

The important detail is replacement of the complete set, rather than appending one record per retry. A retry with the same revision produces the same idempotency key and the same target state. If the provider only exposes record-level changes, the adapter must emulate this set operation with a compare-and-swap or a serialized zone lock; otherwise two onboarding workers can erase each other's priorities.

I once assumed that “upsert” meant a harmless single-row write. DNS taught me otherwise. During a domain onboarding burst, worker A read the primary exchanger, worker B read the same snapshot, and each appended a different backup. A's response arrived last, so the zone looked healthy in both application logs while the second backup vanished from the authoritative set. The support ticket arrived hours later because cached resolvers continued serving the older answer for the full TTL. The fix was to make the set the unit of concurrency and to persist the before-and-after values with the revision; the conflict became a retryable state transition instead of a silent loss. Small change. Big difference.

Write intent first.

What fails when priorities and ownership are treated as incidental?

The first failure is semantic: a migration swaps preference 10 and 20, so backup mail receives traffic before the primary. The second is temporal: a successful write is observed before recursive resolvers have refreshed their caches, and an onboarding screen claims that routing is active too early. DNS TTL controls cache duration; it does not provide a transaction across every resolver.

The third failure is governance. A customer can submit a configuration for a domain they do not control, and a platform-owned zone can be changed by an actor outside the tenant boundary. Keep proof status, actor identity, revision, and intended TTL in the same audit record as the mutation. DMARC alignment does not prove that your application authorized the DNS change; it is a separate mail-authentication policy concern.

Use explicit states such as pending_proof, ready_to_apply, applied, and verification_pending. Make the transition idempotent. A duplicate callback should not create another revision or send another onboarding email.

Cost and retention: what should the system stop keeping?

Retain the configuration revision, the canonical desired set, the provider request identifier, and the observed result for the period your support and compliance obligations require. Drop high-frequency polling samples after they have been summarized into state transitions. That reduces storage and keeps audits readable, but it means a resolver-level timing dispute may need evidence from external DNS measurements rather than your internal log. If an auditor asks why a message was routed to the backup exchanger on a particular afternoon, the revision and authoritative response should answer the question; a thousand identical polling rows do not. The retained record should be compact enough to export, immutable enough to trust, and linked to the onboarding actor without retaining message content.

Decision Keep Give up
Durable revision audit Who changed which MX set and why Every intermediate poll
Short verification window Fresh resolver observations during cutover Long historical cache traces
Full desired-set replacement Deterministic retries and ordering Independent record edits by unrelated workers

The catch is that this approach is not suitable when another team edits the same zone outside your control plane. In that case, use an import-and-merge policy with an explicit ownership boundary, or stick with provider-native change management and make your application verify rather than overwrite. Your mileage may vary because resolver behavior, TTL choices, and mailbox failover policy differ by domain.

A release checklist for mail-routing onboarding

Test canonicalization, duplicate declarations, preference ordering, proof expiry, concurrent revisions, and a retry after a network timeout. Assert that the second application of one revision is a no-op and that a later revision cannot be overwritten by an older worker. Include a resolver check from more than one network; one successful lookup is not global convergence.

Instrument reconciliation latency, proof failures, rejected preferences, stale-read conflicts, and verification age. Alert on a growing verification queue rather than on every individual lookup. That gives the on-call engineer a decision boundary instead of a stream of harmless retries.

For a Node.js implementation, keep parsing and policy in application code, keep the provider SDK behind one adapter, and make the audit event part of the same workflow that advances onboarding. The durable artifact is the declared intent plus its revision, not a screenshot of a DNS dashboard.

References

Further reading

Top comments (0)