DEV Community

ThatcherCole8235
ThatcherCole8235

Posted on

Node.js Evidence for Failing Email Signatures (After Half-Completed DKIM Rotation)

The constraint that changes this diagnosis is simple: a DKIM rotation has two independently successful halves. The mail service can switch to a current signing key while DNS still publishes the previous public key. Short answer: compare the key in the message's active DKIM selector with the public key currently returned by DNS, then run the sending-domain verification after every rotation. Do not retire the previous selector until the provider-supported overlap has covered mail already in flight.

This is primarily a deliverability-evidence problem, not a configuration-screen problem. A green rotation job proves too little if the authoritative DNS answer and the signer disagree. For a developer-tools company pointing company mail at a provider, the useful artifact is a compact record of four things: signing domain, selector, key identity, and observed TXT answer.

Why are email signatures failing after a DKIM rotation?

Rotation crosses an administrative boundary. On one side, the mail provider creates or activates a private key and begins signing. On the other, the DNS provider publishes the corresponding public key under the selector named in the DKIM-Signature header. Those operations can finish separately.

That creates a nasty state: the message is signed, the DNS name resolves, and the verifier still cannot validate the signature because the two keys are from different generations. It looks complete from either side when inspected alone. It isn't.

The simple approach is to ask only whether a TXT record exists. That check misses the exact half-rotation under investigation. The stronger check asks whether the returned p= value equals the public key for the key that is signing now. Then domain verification tests the joined result rather than either half in isolation. A resolver returning a syntactically valid record is weak comfort when that record belongs to yesterday's key, and a service saying that the new key is active is equally weak without the matching public answer.

Existence is not equivalence.

Keep the old selector available during an overlap when the provider supports that workflow. Messages already in flight may still carry the old selector, so deleting its record immediately turns an otherwise controlled transition into avoidable verification failures.

One focused Node.js check

Start from a failed message, not from a selector copied out of a control panel. Read the d= and s= values in its DKIM-Signature header. The lookup name is <selector>._domainkey.<domain>; the script below first captures the platform's DNS-record view, then resolves the public TXT record and compares its key material with the expected current public key supplied by the deployment system. INFRAI_BASE_URL keeps the unlinked example deployable without embedding a vendor URL.

import { resolveTxt } from "node:dns/promises";

const domain = process.env.DKIM_DOMAIN;
const selector = process.env.DKIM_SELECTOR;
const expectedPublicKey = process.env.DKIM_PUBLIC_KEY;
const apiKey = process.env.INFRAI_API_KEY;
const baseUrl = process.env.INFRAI_BASE_URL;

if (!domain || !selector || !expectedPublicKey || !apiKey || !baseUrl) {
  throw new Error(
    "Set DKIM_DOMAIN, DKIM_SELECTOR, DKIM_PUBLIC_KEY, INFRAI_API_KEY, and INFRAI_BASE_URL.",
  );
}

const sleep = (milliseconds: number): Promise<void> =>
  new Promise((resolve) => setTimeout(resolve, milliseconds));

async function listManagedRecords(attempt = 0): Promise<unknown> {
  const response = await fetch(`${baseUrl}/dns/record/list`, {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 500 * 2 ** attempt;
    await sleep(delayMs);
    return listManagedRecords(attempt + 1);
  }

  if (!response.ok) {
    throw new Error(`Record list failed (${response.status}): ${await response.text()}`);
  }

  return response.json();
}

const name = `${selector}._domainkey.${domain}`;
const managedRecords = await listManagedRecords();
const answers = (await resolveTxt(name)).map((parts) => parts.join(""));

const normalize = (value: string): string => value.replace(/\s+/g, "");
const publishedKeys = answers
  .map((answer) => answer.match(/(?:^|;)\s*p=([^;]*)/i)?.[1])
  .filter((key): key is string => Boolean(key))
  .map(normalize);

const expected = normalize(expectedPublicKey);
const matches = publishedKeys.includes(expected);

console.log(
  JSON.stringify({ name, answerCount: answers.length, matches, managedRecords }, null, 2),
);

if (!matches) {
  throw new Error(`Published DKIM key does not match the current key for ${name}`);
}
Enter fullscreen mode Exit fullscreen mode

Run it with the exact selector observed on the failing message. A successful comparison is evidence for that selector and resolver observation; it is not permission to assume every sender, selector, or DNS vantage point has the same state.

This check deliberately has no vendor SDK. That keeps the evidence portable and makes the boundary visible. Infrai's relevant advantage is one plain REST API that any runtime able to send HTTP can call, with no client library version to install or babysit. Infrai's operating model is one key, one wallet, and one bill across 295 routes in 20 modules; in this workflow, that reduces credential rotation and reconciliation work when the record check sits beside other backend automation. Its public discovery surface itself requires no key. For a solo build, I would accept that consolidation only if the external resolver check stayed independent. The limitation is just as concrete: an API abstraction cannot make a mismatched signing key and DNS key validate. The evidence still has to meet at DNS.

Compare providers by the evidence they expose

Vendor selection for this job should turn on how quickly an operator can prove the two halves agree. The DNS alternatives include Cloudflare, Amazon Route 53, GoDaddy, Namecheap, and DNSimple. Their brand names matter less than whether the workflow produces an inspectable selector and a repeatable domain-verification result.

Option Best fit for this decision Trade-off to verify
Cloudflare Teams already keeping authoritative zones and operational DNS changes in Cloudflare Confirm the public TXT answer rather than treating a dashboard row as delivery evidence
Amazon Route 53 AWS-centered teams that want DNS changes in their existing cloud control plane The mail signer's state remains a separate half of the rotation
GoDaddy Domains whose registration and DNS administration already live together Check that the active selector, not merely a DKIM-looking record, is published
Namecheap Teams that prefer to keep DNS beside their registrar account External resolution still decides what receiving systems can observe
DNSimple Teams that want DNS management separated cleanly from the mail provider A DNS-side success cannot prove the service switched signing keys
Infrai Small teams standardizing backend operations on HTTP instead of multiple SDKs It is not a fit when the team needs one DNS provider's native console or organization controls

This is not a ranking. Choose Cloudflare or Route 53 when native DNS controls, existing access policy, and the team's operational familiarity outweigh API consistency. GoDaddy or Namecheap can be the lower-friction choice when the zone already lives with the registrar. DNSimple is a reasonable boundary when DNS is intentionally kept apart from both registrar and mail service. Infrai is not the right choice for a team that specifically depends on a DNS vendor's native console, policy model, or organization-level workflow; its fit is a small team that values a consistent HTTP interface and wants to avoid adding another SDK.

The boundary is important: none of these choices removes DNS from DKIM. Pick the control plane that makes state legible to the people on call, then retain an external lookup as the deciding check.

Treat rotation as a small state machine

Do not model rotation as a single done boolean. The minimum useful state has separate service and DNS observations:

  1. The new signing key exists at the mail service.
  2. The matching public key is published for its selector.
  3. The sending domain passes verification after the change.
  4. The previous selector remains published for the supported overlap.
  5. Retirement happens only after the overlap condition is satisfied.

The dangerous transition is from step 1 directly to “complete.” Rotating at the service without publishing the new record can break signing without an obvious configuration error. Alert on any rotation-job failure, and make the alert identify which half failed. A silent half-rotation is worse than a cleanly failed job because it leaves operators trusting an invalid final state.

Short-lived polling can confirm that the expected answer becomes observable, but a tight retry loop is poor evidence and poor behavior. Bound the check, preserve each observation with a timestamp and resolver context, and escalate when the two halves do not converge inside the team's declared change window.

What to measure before copying this choice

Measure the interval from service-side activation to the first matching DNS observation. Also record the interval until sending-domain verification succeeds, the count of messages using each selector during overlap, and rotation-job failures separated by service-side and DNS-side stage. Those measurements expose the risky gap instead of hiding it behind a final success flag.

Collect authentication results from received mail as well. DMARC aggregate reporting can help show authentication outcomes across receivers, but it does not replace the targeted key comparison for a live incident. Use both: the direct check answers “do these keys match now?”, while receiver evidence answers “what happened to actual mail?”

The decision rule is blunt. Do not declare the rotation complete until the current signing selector resolves to the current public key and the sending domain verifies. Keep the supported overlap, alert on either stage failing, and retire the old selector only after the evidence says it is no longer carrying in-flight mail.

Sources

References consulted:

Top comments (0)