TL;DR: Keep a marketplace's mail zone under customer control when its team already operates DNS. Before retrying verification, read the expected MX set from durable state, query DNS, normalize both sets, and classify the difference. A matching set is verified. A nonmatching set is a configuration mismatch until repeated observations give you evidence of a transition. Retries alone cannot turn the wrong target into the right one.
| Zone model | Best fit | Verification consequence | Main trade-off |
|---|---|---|---|
| Customer-owned | The marketplace team already controls the domain and its mail changes | Compare observed MX values with the exact expected set issued for that domain | More coordination, but ownership and rollback stay with the customer |
| Platform-owned | The platform owns a delegated zone and can apply every required record | Verification can be coupled to the platform's own desired state | Less customer glue, but the platform owns more DNS operations |
Recommendation: default to customer-owned zones for company mail. Store the expected records first, expose them back to the operator, and make verification a comparison rather than a timer. Choose a platform-owned delegated zone only when the platform is supposed to control the whole namespace and can own the resulting operational duty.
How can a DNS read distinguish propagation delay from a wrong record?
Start with three values: the expected MX set, the observed MX set, and the observation time. The expected set is input, not something reconstructed from the latest DNS response. If that set was never persisted, the verifier has no stable definition of success. It can poll. It cannot reason.
This distinction matters in a marketplace because company mail is usually a shared operational surface. A seller onboarding flow may request a mail-routing change while another team still owns the domain. The verifier should report what it asked for and what it found without pretending that elapsed time explains the difference.
Treat the result as one of four states:
-
verified: normalized expected and observed sets match. -
mismatch: DNS returned data, but at least one priority or exchange differs. -
not_visible: the query did not produce an MX set that can be compared. -
query_error: the lookup failed, so no claim about configuration is justified.
Short states beat vague ones.
Do not label every mismatch as propagation. That word smuggles in a diagnosis the read did not prove. A stale answer and a typo can look identical in one observation. The honest response includes the difference and schedules another read under a bounded retry policy. If the same unexpected set persists, keep calling it a mismatch; do not upgrade hope into evidence.
The comparison also needs set semantics. MX ordering in a response is not a success criterion. Priority and exchange are. Normalize exchange names consistently, preserve priority, remove exact duplicates, and sort only to make equality and logs deterministic. This is a tiny amount of code, which is exactly why it belongs in one shared function rather than three onboarding workers.
Ownership is an operational boundary
Customer-owned and platform-owned zones change who can fix a mismatch. They do not change what a correct verifier needs to know. In both models, verification begins with intended state and compares it with observed state.
With a customer-owned zone, the marketplace generates the required MX values and the customer's operator publishes them. The useful error is concrete: expected priority 10 at one exchange, observed priority 10 at another. "Still propagating" is worse than useless if the record was pasted into the wrong zone or entered with a different exchange. It burns a retry budget while hiding the only actionable clue.
With a platform-owned delegated zone, the write path can produce the desired-state record itself. That reduces the number of hands in the loop. It also expands the platform's job: changes, audit history, rollback, monitoring, and incident ownership now sit on its side of the boundary. Fewer configuration steps are attractive. Config bloat is not. The relevant test is whether the delegated namespace is a real product boundary, not whether an extra form can be removed.
Zone ownership decides who changes DNS; expected-state storage decides whether verification is trustworthy. Keep those choices separate. Otherwise a migration from one ownership model to the other quietly rewrites verification behavior too.
Implement the comparator before the retry loop
The following TypeScript keeps transport out of the decision. It accepts already-resolved records, so the same comparator can run in an API handler, a worker, or a unit test. There is no provider-shaped object and no timer hidden in the function.
type MxRecord = {
priority: number;
exchange: string;
};
type Verification =
| { status: "verified" }
| {
status: "mismatch";
missing: MxRecord[];
unexpected: MxRecord[];
};
function normalizeExchange(exchange: string): string {
return exchange.trim().toLowerCase().replace(/\.$/, "");
}
function key(record: MxRecord): string {
return `${record.priority}:${normalizeExchange(record.exchange)}`;
}
function unique(records: MxRecord[]): MxRecord[] {
return [...new Map(records.map((record) => [key(record), record])).values()];
}
export function compareMx(
expected: MxRecord[],
observed: MxRecord[],
): Verification {
const expectedRecords = unique(expected);
const observedRecords = unique(observed);
const expectedKeys = new Set(expectedRecords.map(key));
const observedKeys = new Set(observedRecords.map(key));
const missing = expectedRecords.filter((record) => !observedKeys.has(key(record)));
const unexpected = observedRecords.filter((record) => !expectedKeys.has(key(record)));
return missing.length === 0 && unexpected.length === 0
? { status: "verified" }
: { status: "mismatch", missing, unexpected };
}
A realistic marketplace fixture should contain more than one record because equality logic often goes wrong when developers test only a singleton. Keep the data fictional and the assertion exact.
import { strict as assert } from "node:assert";
import { compareMx } from "./compare-mx";
const expected = [
{ priority: 10, exchange: "mx1.mail.invalid" },
{ priority: 20, exchange: "mx2.mail.invalid" },
];
const observed = [
{ priority: 20, exchange: "MX2.MAIL.INVALID." },
{ priority: 10, exchange: "mx-wrong.mail.invalid" },
];
assert.deepEqual(compareMx(expected, observed), {
status: "mismatch",
missing: [{ priority: 10, exchange: "mx1.mail.invalid" }],
unexpected: [{ priority: 10, exchange: "mx-wrong.mail.invalid" }],
});
The .invalid names are deliberate test data, so nobody copies an example and accidentally targets a live mail system. The assertion also catches two easy mistakes: treating response order as meaningful and ignoring MX priority.
Put retry policy around the resolver, not inside the comparator. Persist an attempt count, the observation time, a normalized snapshot, and the resolver outcome. Then benchmark the workflow on what users feel: time from publishing the intended records to the first matching read, number of DNS reads per verification, and time spent in an unexplained state. Do not publish a universal propagation number from that benchmark. It describes your observation path and test conditions, not every resolver on the internet.
Retry without erasing the diagnosis
A retry scheduler has two separate jobs. It gives changing DNS data another chance to become observable, and it protects your system from hammering resolvers. Neither job justifies replacing a structured mismatch with a spinner.
Use a bounded schedule with increasing intervals and jitter, then stop automatically. The exact limits are an operational choice and should be measured against the onboarding latency the marketplace accepts. Record each comparison result. If the observed set changes between attempts, the history is useful evidence of a transition. If it never changes, the operator still sees the precise missing and unexpected values.
Errors need a different branch. A timeout or resolver failure says the read was inconclusive; an empty or nonmatching answer says the read completed but did not verify the expected state. Combining those paths makes alerting noisy and support transcripts useless. It also ruins benchmarks because transport availability and configuration accuracy become one metric.
DMARC offers a useful adjacent design lesson for mail-domain tooling: policy discovery is performed through DNS, while reporting provides a separate feedback path. RFC 7489 defines both the DNS-published policy mechanism and aggregate or failure reporting concepts. An MX verifier should keep the same conceptual separation. DNS reads establish what is visible; logs and notifications explain the verification process. Do not make a mail-routing claim from a DMARC record, and do not treat a successful MX comparison as proof that the domain's broader mail authentication policy is correct.
When should the platform own the zone instead?
The runner-up is better when the namespace exists solely for the platform-managed function, delegation is explicit, and the platform is prepared to operate its entire lifecycle. In that case, platform ownership can remove a manual handoff and align the DNS write with desired state. Verification is still valuable because the public read path remains the thing being checked.
It is a poor fit when company mail shares a zone with records owned by several teams, when customers require direct rollback control, or when the platform cannot take responsibility for changes outside its narrow workflow. A convenient setup screen does not settle those questions. Ownership does.
The decision rule is blunt: choose the smallest DNS boundary one team can fully operate. For the usual marketplace company-mail change, that leaves the zone with the customer and gives the platform a strict expected-versus-observed verifier. For a purpose-built delegated namespace, platform ownership may be cleaner. Both paths should produce the same evidence, the same comparison semantics, and the same refusal to call every mismatch propagation.
Further reading
- RFC 7489, Domain-based Message Authentication, Reporting, and Conformance (DMARC): https://datatracker.ietf.org/doc/html/rfc7489
Top comments (0)