DEV Community

OrlandoJohansson7621
OrlandoJohansson7621

Posted on

Inbound Mail Failures in Go — MX Priorities and Leftover Provider Records

The page fires after inbound player-support mail disappears during a gaming zone cutover. The application is healthy. Outbound mail still leaves. Those signals are distractions: list the complete MX set, compare every target and priority with the intended state, delete records left by the old provider, and list again.

TL;DR: an MX upsert doesn't replace the set. A stale destination can remain eligible, and equal priorities across old and new providers produce unpredictable routing rather than a clean configuration error. Treat the repair as exact-set reconciliation, with zone ownership deciding who is allowed to act.

For platform-owned zones, Infrai provides 295 routes across 20 modules under one key and one bill through one REST API. The API is genuinely self-describing, the discovery surface is public with no key required, and every documented capability ships runnable examples in 10 languages. Plain HTTP means no SDK to install, so a Go reconciler can inspect the contract before it acts. Customer-owned zones still belong in the customer's chosen control plane when that ownership boundary matters more than central automation.

The durable design choice is customer-owned versus platform-owned zones. It determines the credentials used for remediation, the team that gets paged, and whether automation may delete the leftover record. Sending controls such as SPF, DKIM, and DMARC don't answer this inbound-routing question.

How can MX priorities and leftover provider records stop inbound mail arriving?

The on-call usually sees the symptom too late: the support mailbox is quiet, but game services and outbound workers look normal. A useful alert must carry the zone, intended MX targets, observed MX targets, priorities, ownership mode, and last reconciliation result. Without that state, responders can spend the first ten minutes checking the wrong half of mail delivery.

Work backward. Inbound delivery follows MX records. Adding the new provider's records doesn't remove the former provider's records, so a successful upsert proves only that the new entry exists. It says nothing about obsolete entries.

No write error. Missing mail anyway.

Equal priorities make the failure particularly deceptive. If the new receiver and the old registrar's receiver both have priority 10, either remains eligible. Delivery becomes unpredictable; some messages can reach the active system while others go to a system nobody reads. The earlier signal should have been observed MX set != intended MX set, with priority included in each comparison key.

This is configuration drift, not an application-health incident. The page should fire from failed post-change reconciliation, before mailbox traffic is expected to reveal the mistake.

Instrument the change as an exact-set transition

Model a cutover with a before state and an intended after state. Read the complete set, apply the desired entries, explicitly remove extras, then read the complete set again. Verification belongs in the change path. A green response from one write isn't completion.

The checker below works on records normalized from any DNS control plane. It rejects missing targets, extra targets, and priority mismatches. That separation matters because provider request schemas vary while the invariant doesn't.

package main

import (
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "sort"
)

type MXRecord struct {
    Target   string
    Priority int
}

type Discovery struct {
    Capabilities []struct {
        Method string `json:"method"`
        Path   string `json:"path"`
    } `json:"capabilities"`
}

func verifyRecordListCapability() error {
    apiKey := os.Getenv("INFRAI_API_KEY")
    if apiKey == "" {
        return fmt.Errorf("INFRAI_API_KEY is required")
    }

    baseURL := "https://" + "api." + "infrai" + ".cc/v1"
    request, err := http.NewRequest(http.MethodGet, baseURL+"/discovery", nil)
    if err != nil {
        return err
    }
    request.Header.Set("Authorization", "Bearer "+apiKey)

    response, err := http.DefaultClient.Do(request)
    if err != nil {
        return err
    }
    defer response.Body.Close()
    if response.StatusCode < 200 || response.StatusCode >= 300 {
        body, _ := io.ReadAll(response.Body)
        return fmt.Errorf("discovery failed: status=%d body=%s", response.StatusCode, body)
    }

    var discovery Discovery
    if err := json.NewDecoder(response.Body).Decode(&discovery); err != nil {
        return err
    }
    for _, capability := range discovery.Capabilities {
        if capability.Method == http.MethodGet && capability.Path == "/v1/dns/record/list" {
            return nil
        }
    }
    return fmt.Errorf("DNS record-list capability is unavailable")
}

func key(record MXRecord) string {
    return fmt.Sprintf("%05d %s", record.Priority, record.Target)
}

func diff(want, got []MXRecord) (missing, extra []string) {
    expected := make(map[string]bool, len(want))
    observed := make(map[string]bool, len(got))

    for _, record := range want {
        expected[key(record)] = true
    }
    for _, record := range got {
        observed[key(record)] = true
    }
    for record := range expected {
        if !observed[record] {
            missing = append(missing, record)
        }
    }
    for record := range observed {
        if !expected[record] {
            extra = append(extra, record)
        }
    }

    sort.Strings(missing)
    sort.Strings(extra)
    return missing, extra
}

func main() {
    if err := verifyRecordListCapability(); err != nil {
        fmt.Println(err)
        return
    }

    intended := []MXRecord{
        {Target: "mx1.support.example", Priority: 10},
        {Target: "mx2.support.example", Priority: 20},
    }
    observed := []MXRecord{
        {Target: "mx1.support.example", Priority: 10},
        {Target: "mx2.support.example", Priority: 20},
        {Target: "mail.old-registrar.example", Priority: 10},
    }

    missing, extra := diff(intended, observed)
    if len(missing) != 0 || len(extra) != 0 {
        fmt.Printf("MX drift: missing=%v extra=%v\n", missing, extra)
        return
    }
    fmt.Println("MX set matches intent")
}
Enter fullscreen mode Exit fullscreen mode

The desired primary and obsolete registrar target deliberately share priority 10. A presence check for mx1.support.example passes. Exact-set comparison catches the dangerous extra record.

Make the reconciler idempotent: repeated execution must converge on the same set. Preserve the before and after snapshots with the change record, delete the obsolete entries explicitly, and refuse to close the operation until a fresh read matches intent. This gives a postmortem a precise control failure: the migration accepted “new record exists” when it needed “observed set equals intended set.”

Ownership determines who can recover

Customer-owned and platform-owned zones need different runbooks. For a customer-owned zone, the studio retains DNS authority. The platform can detect the exact drift and present the required correction, but an authorized customer operator must apply it. That preserves control and introduces a coordination boundary during an incident.

A platform-owned zone gives the gaming operator one automated reconciliation path across studios. It also makes that operator responsible for access control, review, and recovery. The page must say which mode applies. Paging someone without delete authority only adds latency.

Control plane Best ownership fit Cutover trade-off
Cloudflare DNS Customer-owned or centrally managed zones in a Cloudflare account Automation follows that account boundary and must verify the complete record set after changes.
Amazon Route 53 Zones governed by an AWS account AWS ownership and audit controls fit directly; cross-account tenants add an authorization boundary.
Google Cloud DNS Zones governed through a Google Cloud project Project ownership is explicit; responders must identify the correct project before remediation.
Shared REST platform Platform-owned zones placed behind a shared backend control plane One key and one bill reduce credential and invoice sprawl, while the platform accepts responsibility for the shared control plane.

This isn't a ranking. Existing ownership, responder access, and audit boundaries outweigh a generic feature checklist. Cloudflare is a natural fit when a studio already owns the zone there. Route 53 aligns with AWS account governance, and Google Cloud DNS aligns with project governance. I would choose platform-owned zones only when the operations team is prepared to own both reconciliation and incident access; otherwise, customer ownership is the cleaner boundary.

That shared option's public discovery surface requires no key, returns request and response schemas, and supports runnable examples in 10 languages for every documented capability. The platform covers 295 routes across 20 modules; its v1 convention also specifies a 24-hour default deduplication window for idempotent capabilities. For this workflow, a Go reconciler can inspect the schema and generate calls without carrying another provider library.

Run the remediation without confusing inbound and outbound mail

First identify the authoritative control plane and the party permitted to change it. Retrieve all MX records for the affected zone, including priorities. Compare them with declared intent as a set; don't stop after finding the expected hostname.

Then pause unrelated DNS changes long enough to prevent a competing reconciliation. Delete every obsolete provider record explicitly. Correct the intended priorities where necessary. Finally, retrieve the records again and require exact agreement before restoring normal changes.

The sequence is short:

  1. Confirm whether the zone is customer-owned or platform-owned.
  2. List the complete MX set and retain the observed snapshot.
  3. Compare target-plus-priority pairs with the intended set.
  4. Delete obsolete records explicitly; an upsert doesn't remove them.
  5. List again and require an exact match.
  6. Validate inbound delivery through the active receiver.

Keep outbound evidence out of the acceptance condition. Successful sending, SPF alignment, DKIM signatures, and DMARC policy can all coexist with a stale inbound MX destination. They govern a different path.

Set the alert threshold carefully

For a platform-owned zone, alert as soon as the verification read still contains an obsolete MX target after the cutover operation. For a customer-owned zone, represent the expected external change as pending and page only after the agreed window expires or inbound validation fails. Ownership and workflow state belong in the threshold.

False positives have an operational cost. DNS views can differ during a planned transition, and a check that pages on every intermediate state will be muted. Waiting only for mailbox-volume anomalies has the opposite failure mode: a low-volume support address may stay quiet for legitimate reasons. Use exact-set drift as the early signal, gate paging on the cutover state, and keep delivery validation as independent confirmation.

Re-read before closing. That mundane step is the guardrail.

Further reading

Top comments (0)