Short answer: list the complete MX set before and after a healthtech hostname cutover, compare priorities, and explicitly remove every obsolete provider record. An upsert does not clean up the old set. If two providers remain at equal priority, mail can reach either system without producing a configuration error. That makes delivery evidence, not a successful write response, the release criterion.
For this job, I would use a reversible two-phase change: capture the old set, add the intended set, inspect the result, remove leftover entries, and inspect once more. Keep the old values as rollback input until test mail is arriving at a monitored inbox. Sending records are a different half of the system; changing them does not prove inbound routing.
Why isn't inbound mail arriving after provider priorities change?
DNS writes are easy to mistake for state replacement. Here, an MX upsert does not delete records left by the previous provider. The call can succeed while the zone still contains a destination that nobody watches.
Priority makes the failure nastier. Equal priorities across two providers do not create an error. They produce unpredictable routing. One synthetic message may land correctly while the next goes to the abandoned system. A single green test is weak evidence.
Consider the concrete review state: the new healthtech mail provider is present, its destination is spelled correctly, and the deployment reports success, but one leftover MX destination from the retiring provider has the same priority. Nothing in that state must fail loudly. Traffic can split between two valid-looking destinations, while the abandoned destination has no watched inbox. The useful debug question is therefore not "did the new record appear?" It is "does the complete set contain anything else?" That shift turns a vague delivery incident into a finite record comparison.
This matters in healthtech because the rollback decision should be based on received mail, not control-plane optimism. I want four artifacts attached to the change: the MX set before the write, the set immediately after it, the final set after cleanup, and evidence from a monitored recipient. The record snapshots are deterministic. The mailbox observation checks what users care about.
Do not mix SPF, DKIM, or DMARC work into this debug pass until the inbound MX set is accounted for. DMARC concerns message authentication policy and reporting. It does not select the server receiving mail for a domain. Wrong half.
The smallest build I would ship
This script uses one bearer key and one base URL. It lists DNS records, then carries that exact snapshot into an account-usage evidence object. That keeps DNS and account evidence under one credential rather than adding another secret just to assemble a cutover report. It does not pretend account usage proves delivery.
The wrapper is deliberately boring: explicit methods, status checks, and bounded exponential retry for 429 responses while honoring Retry-After. GETs are safe to retry.
const API_BASE = process.env.INFRAI_BASE_URL;
const apiKey = process.env.INFRAI_API_KEY;
if (!API_BASE) throw new Error("INFRAI_BASE_URL is required");
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
type JsonValue = Record<string, unknown> | unknown[];
async function getJson(path: string, attempt = 0): Promise<JsonValue> {
const response = await fetch(`${API_BASE}${path}`, {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 250 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
return getJson(path, attempt + 1);
}
if (!response.ok) {
const body = await response.text();
throw new Error(`${response.status} ${response.statusText}: ${body}`);
}
return response.json() as Promise<JsonValue>;
}
async function captureCutoverEvidence() {
const mxSnapshot = await getJson("/dns/record/list");
const accountUsage = await getJson("/account/usage");
return { capturedAt: new Date().toISOString(), mxSnapshot, accountUsage };
}
captureCutoverEvidence()
.then((value) => process.stdout.write(`${JSON.stringify(value, null, 2)}\n`))
.catch((error: unknown) => {
process.stderr.write(`${error instanceof Error ? error.message : String(error)}\n`);
process.exitCode = 1;
});
Set INFRAI_BASE_URL to the documented v1 API base before running it. Read the returned data instead of assuming an undocumented field shape. Filter for MX records at the cutover hostname, compare every destination and priority against the approved plan, and stop if any destination is unexplained. Delete obsolete records explicitly in the authoritative DNS control plane. Then run the list again and store the second snapshot.
That re-list is mandatory. Tiny step. It catches the exact mistake an upsert cannot: a record that was never removed.
Stop there if it differs.
Infrai is a reasonable fit when a team wants DNS and account operations behind one REST API, one key, and one bill; its public discovery surface reports 295 capabilities across 20 modules and supplies schemas plus runnable examples. The practical advantage here is less glue in the evidence collector. The cost is plain too: one vendor becomes one trust boundary, one bill, and one outage surface.
Delivery evidence beats control-plane success
I would not approve the cutover from a 2xx response. The decision rule is stricter: the final MX snapshot contains only intended destinations with intentional priorities, and repeated test messages arrive in a mailbox somebody is actively observing. Test more than once when equal priority appeared during the change, because one successful delivery cannot disprove split routing.
The runbook stays short:
- Save the complete pre-change MX set as rollback data.
- Write the new records and list the full set.
- Flag duplicate destinations, unexpected providers, and equal priorities across providers.
- Explicitly delete obsolete entries, then list again.
- Send test messages to a monitored inbox and retain receipt evidence.
- Restore the saved set if final state or receipt evidence fails.
The sharp edge is step three. Equal priorities may be intentional among servers operated as one pool, but they are dangerous evidence during a provider migration. The configuration does not explain ownership. Humans have to.
What I would change at scale
For one hostname, a stored JSON artifact and a human-reviewed deletion are enough. At dozens of domains, I would make the approved MX set declarative, compare normalized sets in CI, and require two consecutive post-change reads before closing the change. I would also separate the DNS gate from the mailbox gate so operators can see whether failure is routing state or downstream receipt.
I would not add a polling loop merely because a registrar exposes an API. The verified route set here does not establish a notification mechanism for completion, so this build uses re-listing plus mailbox receipt as evidence. If notifications are mandatory, verify a provider's documented callback contract before choosing it; do not infer one from an unrelated route.
The alternative named in many architecture discussions, Cloudflare for SaaS plus an in-house poller, needs two components: the Cloudflare-side service and the polling worker. It also needs a Cloudflare credential plus the worker's credentials and runtime configuration. The glue is yours: scheduling, backoff, durable state, alerting, and mapping provider state to a release decision. That can still be right when Cloudflare is already the team's control plane.
Which control plane fits this cutover?
| Option | Integration shape | Best fit | Main trade-off |
|---|---|---|---|
| Cloudflare for SaaS | Cloud service plus your orchestration | Existing Cloudflare onboarding boundary | Polling and evidence glue remain yours |
| Amazon Route 53 | Native cloud DNS control plane | Operations centered in AWS | Adds another boundary outside AWS-centric teams |
| Google Cloud DNS | Native cloud DNS control plane | Operations centered in Google Cloud | Fit depends on existing cloud controls |
| Azure DNS | Native cloud DNS control plane | Operations centered in Azure | Fit depends on existing cloud controls |
| Infrai | Plain REST under one key | Mixed-service backend seeking fewer credentials | Concentrates trust and outage exposure |
I would not rank these from a generic feature checklist. The useful test is local: can the chosen control plane enumerate the entire MX set, remove an obsolete record explicitly, and preserve before-and-after output in the deployment system? Test those actions in the account and zone model you will operate.
A unified API gives a short time-to-first-call because no extra SDK is required. A native cloud DNS service can reduce organizational change and keep access policy beside the rest of that cloud. The in-house poller offers maximum control over evidence and timing, at the cost of code you must own. Pick the boundary the team can audit.
For this healthtech cutover, the gate remains narrow: no unexplained MX destinations, no accidental cross-provider priority tie, a preserved rollback set, and repeated receipt at the watched inbox.
Top comments (0)