Short answer: when email signatures start failing after a DKIM rotation, keep the old public key published, identify which selector each failed message used, and delay the registrar cutover until both selector records resolve correctly through independent recursive resolvers. For an edtech platform moving zones away from a registrar-specific API, the least complex recovery is an overlap window. Do not compensate for uncertain DNS propagation by changing the signer again.
| What you observe | Best next move | Pick this when | Main trade-off |
|---|---|---|---|
| Old selector signs and validates | Keep both keys published | Some mail still leaves through the old signing path | Slower retirement, fastest stabilization |
| New selector signs, but its TXT record is absent from some resolvers | Pause the cutover and measure resolution | Authoritative DNS is correct but recursive answers differ | Slower cutover, lower verification risk |
| Messages carry different selectors across sending paths | Inventory and converge the signers | A queue, region, or service missed the change | More investigation, precise fix |
| DKIM passes but DMARC fails | Check identifier alignment and the From domain | Cryptographic verification is healthy | Avoids rotating a healthy key |
| Neither selector is dependable | Restore the last known signing configuration while retaining both records | The signing rollout cannot finish safely | Stable state, postponed migration |
The decision axis is propagation delay versus cutover speed. Favor speed only after observation proves every active signer and public DNS agree. Half a rotation is two states in production, not one incomplete checkbox.
How should I debug email signatures failing after DKIM rotation?
A DKIM verifier reads the selector and signing domain from the DKIM-Signature field, constructs a DNS name, and retrieves the public key. A signature can fail when the message references an old selector whose record was removed, when a new record is not visible through a resolver yet, or when the signer uses a private key that does not match the published public key. Those cases look similar in a dashboard. Their fixes are different.
DMARC adds another boundary. A message may have a valid DKIM signature and still fail DKIM-based DMARC evaluation when the authenticated signing domain does not align with the domain visible in the From header. That distinction matters during a domain move. It prevents an operator from treating every DMARC failure as proof that the rotated key is wrong. RFC 7489 defines the alignment check and aggregate reporting that exposes results by source.
Think of the path as a diagram in words: signer configuration -> queued message -> receiving resolver -> DKIM verification -> DMARC alignment -> reported result. Put an observation at every arrow. A green answer at the authoritative server proves only one arrow.
That is the split.
Pick overlap when mail still uses both selectors
Overlap is the default recovery choice. Publish the new selector before enabling it on any signer. Keep the old selector available while queued mail and lagging signing processes can still emit signatures that reference it. Retire it only after evidence shows it is no longer used and the operational retention window has passed.
This choice accepts a longer period with two valid public keys. In return, it decouples a DNS migration from a signing deployment. That is useful for an education platform where enrollment notices, password resets, and instructor messages may leave through separate services or queues. One configuration rollout can finish at different times across those paths.
Do not infer completion from the control plane. Sample delivered messages. Record the selector, signing domain, From domain, sending path, and verification result. If even one active path still emits the old selector, removing its record turns a harmless deployment lag into a receiver-visible failure.
Keep both keys.
Pick a propagation hold when recursive answers disagree
A propagation hold fits the case where the new record is present at the authoritative source but public recursive resolvers return different answers. Freeze both DNS and signing configuration. Repeated edits reset the experiment because operators can no longer tell which change produced an answer.
The hold should be driven by measurements, not a slogan such as “wait 48 hours.” DNS caching depends on the published TTL and on when each cache obtained its answer. Negative answers matter too. A resolver that looked up the selector before it existed may retain that negative result according to DNS negative-caching rules.
Fast cutovers are still possible. Lower the relevant TTL ahead of a planned change, confirm that the lower value is being served, then wait for older cached data to age out before switching signers. This costs calendar time before the event, but it buys a shorter and more legible transition. For an emergency recovery, preserve both records and observe. Changing TTL after caches already hold an answer does not rewrite those cached entries.
No edit can recall a cached answer.
Instrument the rotation as a distributed deployment
Start with message evidence. Choose one failed message and one passing message from the same time range. Capture their selector and signing domain from DKIM-Signature, plus the receiver's authentication result. Do not paste message bodies, recipient addresses, or full headers into a shared incident channel; the few authentication fields are enough for this comparison.
Next, query the exact selector names through multiple recursive resolvers from more than one network. Compare presence and record content. Then query the authoritative nameservers to separate publication state from cache state. Consider one concrete split: enrollment mail carries course-new, password-reset mail carries course-old, the authoritative servers publish both records, and one recursive resolver reports the new name as missing. Leave the selectors alone. The password-reset path proves the old signer still exists; the missing new lookup identifies a cache-side risk; and neither observation supports deleting a record. A later sample in which all paths carry course-new resolves signer convergence, but retirement still waits for recursive agreement and zero observed old-selector traffic across the documented window. This sequence answers a crisp question: is the defect at the signer, authoritative DNS, recursive DNS, or alignment layer?
Here is a small TypeScript probe for recursive DNS. It accepts resolver addresses as input instead of baking in a provider. It prints a hash rather than the TXT value, making reports comparable without copying key material through every log pipeline.
import { Resolver } from "node:dns/promises";
import { createHash } from "node:crypto";
type Observation = {
resolver: string;
name: string;
status: "present" | "missing" | "error";
valueHash?: string;
detail?: string;
};
async function observeTxt(resolverAddress: string, name: string): Promise<Observation> {
const resolver = new Resolver();
resolver.setServers([resolverAddress]);
try {
const chunks = await resolver.resolveTxt(name);
const values = chunks.map((parts) => parts.join("")).sort();
const valueHash = createHash("sha256")
.update(JSON.stringify(values))
.digest("hex")
.slice(0, 12);
return { resolver: resolverAddress, name, status: "present", valueHash };
} catch (error) {
const code = error instanceof Error && "code" in error
? String(error.code)
: "UNKNOWN";
return {
resolver: resolverAddress,
name,
status: code === "ENODATA" || code === "ENOTFOUND" ? "missing" : "error",
detail: code,
};
}
}
async function inspectRotation(
domain: string,
selectors: string[],
resolvers: string[],
): Promise<void> {
const checks = resolvers.flatMap((resolver) =>
selectors.map((selector) =>
observeTxt(resolver, `${selector}._domainkey.${domain}`),
),
);
console.table(await Promise.all(checks));
}
await inspectRotation(
"mail.example.edu",
["course-old", "course-new"],
process.argv.slice(2),
);
Run the probe on a schedule during the change and send results into the same telemetry system as mail authentication outcomes. Useful metrics are counts grouped by selector, signing path, and result. Add resolver disagreement as a separate signal. Keep cardinality bounded: domains and selectors are operational dimensions; message IDs and recipient addresses are not.
Alert on sustained user impact, not one isolated lookup. A practical alert requires both meaningful authentication-failure volume and evidence that the rate changed from its established baseline. Route resolver disagreement to the DNS migration owner. Route a selector split to the messaging owner. The page now names an action.
Good alerts teach.
The before-and-after view should be blunt. Before cutover, the old selector carries most signed volume and both records resolve everywhere. During overlap, new-selector volume rises while both remain verifiable. After convergence, old-selector volume reaches zero for the agreed observation period; only then is retirement eligible. Zero observed old-selector use is the gate, not the deployment timestamp.
Cut over the zone without losing the evidence
Registrar independence changes the control path, but the DNS protocol remains the shared contract. Export and review the entire zone before changing delegation. Preserve DKIM TXT records exactly, including selector labels and long string segments. Validate the destination zone from its authoritative servers before updating delegation.
For the change window, keep an immutable manifest of expected record names and hashes. Record the old and new authoritative nameserver sets, the time delegation changed, and observations from recursive resolvers. This gives responders a timeline without binding the runbook to a registrar API. It also makes rollback concrete: restore the last known delegation or signing state, but do not delete either selector during diagnosis.
There is a clear trade-off. Coupling key rotation to zone migration shortens the calendar but multiplies possible causes. Separating them takes another maintenance window, yet each result is easier to interpret. For authentication-sensitive mail, choose the separated sequence unless an external constraint prevents it: establish the destination zone, observe stable resolution, and rotate signing keys afterward.
One boundary at a time.
Use five exit checks: both selector records match the manifest at every authoritative server; independent recursive observations agree; every active sending path uses the intended selector; DKIM verification and DMARC alignment remain healthy at normal volume; and the old selector shows zero use for the team's documented queue and observation window.
Limits of this runbook
This method localizes failures around selector publication, signing rollout, delegation, and DMARC alignment. It does not prove that a receiver will accept a message. Other authentication signals, message content, reputation, and receiver policy remain outside this diagnosis.
It also cannot recover an unknown private key or validate a key pair from DNS observations alone. If the published key hash is consistent but signatures made with that selector fail, test the signer and key pair in a controlled path. Preserve the overlap while doing it.
The operational rule is concise: observe both sides of the DNS boundary before removing compatibility. That produces a slower-looking plan on paper and a faster recovery when one half of the deployment stalls.
Further reading
- RFC 6376, DomainKeys Identified Mail (DKIM) Signatures: https://datatracker.ietf.org/doc/html/rfc6376
- RFC 7489, Domain-based Message Authentication, Reporting, and Conformance (DMARC): https://datatracker.ietf.org/doc/html/rfc7489
- RFC 2308, Negative Caching of DNS Queries: https://datatracker.ietf.org/doc/html/rfc2308
- RFC 1035, Domain Names — Implementation and Specification: https://datatracker.ietf.org/doc/html/rfc1035
Top comments (0)