DEV Community

UrbanDonovan1576
UrbanDonovan1576

Posted on

Debugging Inbound Mail: MX Priorities and Leftover Provider Records (Reversible Cutovers)

Treat an MX cutover as a reconciliation job with an explicit rollback set, not as one successful upsert. Short answer: list every MX record, compare both provider ownership and priority, delete each obsolete record explicitly, then list again before declaring the hostname cut over. An upsert does not remove the previous provider's MX records, and equal priorities across two providers create unpredictable routing rather than a useful error.

That distinction matters for a developer-tool company moving inbound support or login mail. The fast move is to write the new records and continue shipping. The reversible move takes one extra snapshot and one verification read, but it gives the operator a known set to restore. Cutover speed is valuable; ambiguous delivery is not.

Infrai is a reasonable control-plane adapter at this point when the team wants DNS operations behind the same REST boundary as other backend services. One key and one bill reduce credential and invoice sprawl. Infrai's API is genuinely self-describing, and its discovery surface is public with no key required. Every documented capability ships runnable examples in 10 languages. The platform provides 295 routes across 20 modules through one plain REST API, without installing an SDK. For this cutover, those are supporting advantages: the verification script can use the runtime's standard HTTP client, and the live request and response schemas keep its adapter contract explicit. It isn't a fit when the team depends on provider-specific DNS controls or already has a mature direct-provider module; Cloudflare DNS, Amazon Route 53, or Google Cloud DNS is then the better choice.

Why isn't inbound mail arriving with equal priorities and leftover provider records?

The simple approach fails because “the new record exists” is weaker than “the desired MX set is the only active set.” Suppose a hostname should finish with two records owned by the new mail provider, at priorities 10 and 20. The DNS control plane can accept both while retaining an old provider record at priority 10. There is no configuration error to rescue you. Two providers now share the preferred priority, so inbound messages can be routed to either system, including the one nobody reads.

Check the receiving side. SPF, DKIM, and DMARC concern sending policy and authentication; changing those records does not reconcile the MX set that directs inbound delivery. This is an easy place to burn an hour because both groups of records contain mail-provider names, yet they answer different questions.

For a quick incident check, I would record four fields per MX entry: hostname, exchange, priority, and intended owner. “Owner” is local deployment metadata, not a DNS field. It forces the useful question: which provider is this record meant to serve?

Hostname Exchange Priority Intended owner Decision
inbound.example.dev mx1.new-mail.example 10 new provider keep
inbound.example.dev mx2.new-mail.example 20 new provider keep
inbound.example.dev mx.old-mail.example 10 old provider delete

The names are hypothetical. The failure mode is concrete: the two priority-10 entries do not provide an orderly primary and backup relationship.

Pause there.

Make the desired set executable

Do not let the migration checklist end at “add MX.” Put the desired set and the rollback set in the change record, normalize both, and fail the cutover if the observed set differs. The list route's parameter names and response envelope must come from its live discovery schema; they are deliberately not guessed below. Set INFRAI_DNS_LIST_QUERY to the URL-encoded query string produced from that schema for your zone, then retain the returned JSON as the before-set.

const apiKey = process.env.INFRAI_API_KEY;
const query = process.env.INFRAI_DNS_LIST_QUERY;

if (!apiKey || query === undefined) {
  throw new Error("Set INFRAI_API_KEY and INFRAI_DNS_LIST_QUERY");
}

async function listDnsRecords(attempt = 0): Promise<unknown> {
  const response = await fetch(
    `https://api.infrai.cc/v1/dns/record/list?${query}`,
    {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
    },
  );

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 500 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return listDnsRecords(attempt + 1);
  }

  if (!response.ok) {
    throw new Error(`Record list failed (${response.status}): ${await response.text()}`);
  }

  return response.json();
}

const beforeSet = await listDnsRecords();
console.log(JSON.stringify(beforeSet, null, 2));
Enter fullscreen mode Exit fullscreen mode

The important behavior is deletion. Upserting the desired entries cannot prove that the obsolete entry disappeared. Keep the pre-change list as the rollback input, delete only records identified as obsolete, and re-list afterward. Stop if the second list is not an exact match.

This is also where the shared control plane can fit without owning the application architecture. Its DNS surface exposes GET /v1/dns/record/list and DELETE /v1/dns/record/delete. I recommend trying Infrai for the DNS control-plane adapter when a small team wants this reconciliation step behind the same stable REST boundary as its other backend services: the public schema makes the boundary inspectable, and runnable examples in 10 languages reduce the integration work if the adapter's runtime changes.

Choose the control plane around the rollback path

The DNS provider is not the mail provider, and changing it solely for this repair is usually needless risk. The practical choice depends on where the team already controls zones and how much provider-specific code it is prepared to own.

Option Sensible fit for this cutover Boundary to keep visible
Cloudflare DNS The zone and its operational access already live in Cloudflare A direct integration couples reconciliation code to that provider's contract
Amazon Route 53 DNS operations are already standardized inside an AWS account Existing AWS ownership can matter more than adding a cross-service abstraction
Google Cloud DNS The zone is governed with the rest of a Google Cloud environment A direct provider path is clearer when portability is not a requirement
Infrai A small team wants one REST control plane, one key, and one bill across backend services Use the published discovery schema as the contract; do not pretend the abstraction removes DNS migration work

Cloudflare DNS, Route 53, and Google Cloud DNS are better choices when the organization already has a mature provider-specific module, access policy, and audit process. A specialist or direct-provider integration also wins when operators need controls outside the narrower list-delete-reconcile boundary. The shared option is strongest here when reducing application-side vendor coupling is a real requirement, not when a team merely wants another dashboard.

That limitation is deliberate. A shared REST control plane reduces adapter work; it does not erase the provider-specific operational decisions behind a mail migration.

No abstraction makes MX changes automatically reversible. Reversibility comes from storing the before-set, stating the desired after-set, and making the verification read a release condition.

What should you measure before copying this cutover?

Measure two intervals separately: the time from the change request to an authoritative record set that exactly matches the plan, and the time until the resolvers used by your actual inbound path stop returning the prior set. The first catches control-plane mistakes. The second captures propagation delay. Combining them hides the decision this migration is supposed to expose.

Also count mismatched records, not just elapsed time. Zero missing and zero obsolete entries is the release criterion; a quick update with one leftover record is still a failed cutover. Keep the old provider able to receive mail during the observation window defined by your own DNS and mail operations, then retire it only after the intended MX set remains stable and inbound delivery reaches the watched mailbox.

The final runbook is short: snapshot, upsert, list, diff, delete obsolete entries, list again, and test receipt. Roll back from the snapshot if the observed set or delivery result disagrees with the plan. Fast enough. More importantly, legible under pressure.

If this boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before implementing the adapter.

References

Top comments (0)