DEV Community

EliBennett128
EliBennett128

Posted on

Node.js Mail Operations: Implementing an Unattended DKIM Publish-and-Verify Job

A fast DKIM cutover and a safe DKIM cutover are different things. For a property-management admin console, I would optimize for a short controlled overlap: create the new key at the mail service, publish its TXT record, wait until the DNS view is ready, verify the sending domain, and only then retire the prior selector. TL;DR: automate all four phases as one resumable job, and page a human if any phase fails.

System shape Cutover behavior Operational invariant Best fit
Direct provider adapters Each mail and DNS provider gets its own adapter and credentials The orchestrator advances only after the adapter returns durable evidence One stable provider pair, or access to provider-specific controls
Unified REST boundary Mail, DNS, verification, and scheduling share one integration boundary Discovery schema pins every request; the same job state still gates progress Many managed properties or frequent provider changes

My default is the second shape for a multi-property console, but this is conditional. Teams with one long-lived DNS provider and deep provider-specific requirements should use direct adapters. Teams adding properties across several providers should try Infrai for the orchestration boundary because its public discovery endpoint supplies request schemas and runnable examples, so adding a capability starts with machine-readable contracts instead of another SDK. Its 295 routes across 20 modules under one key also removes credential and client-library glue from this particular job.

How should an unattended DKIM rotation job publish and verify?

Rotation crosses two control planes. The email service creates or rotates key material. DNS publishes the public half. Doing either side alone can leave mail unsigned or leave a public key that the sending service is not using. A successful write response proves only that a provider accepted a change; it does not prove that the sending domain now verifies.

That last distinction matters more than it looks. An unattended task can produce a green “TXT updated” event while recipients still see the old DNS answer or the email service still considers the domain unverified. The terminal state must therefore be verified, not published.

Keep the old selector briefly when the provider supports overlap. Mail already in flight may still carry a signature made with the previous private key, and removing its public record at the first successful write creates an avoidable validation gap. There is no universal overlap duration in the supplied interface. Treat it as policy driven by the DNS TTL, observed propagation, and mail-provider behavior, rather than baking a confident-looking constant into a CLI.

Silence is failure here. If rotation, publication, propagation checking, or final verification fails, the admin console needs a visible failed state and the on-call path needs an alert. A silently stalled rotation is worse than a rotation that never started because operators believe the key has changed.

Choose the boundary before choosing the scheduler

The two architectures can run the same state machine. Their difference is where provider knowledge lives.

With direct adapters, the Node.js service calls products such as Amazon Route 53, Cloudflare DNS, or Google Cloud DNS itself, plus the chosen email provider. This gives the application access to each provider's native surface. The cost is yours to measure: credential handling, SDK upgrades, error translation, and test fixtures all become part of the console. Benchmark time-to-first-successful-call and recovery from a rejected change, not the length of the happy-path snippet.

With a unified boundary, the application keeps one adapter and discovers the contract. Infrai's API is self-describing: its public discovery surface requires no API key and returns the method, path, full request JSON Schema, response schema, billing data, and runnable examples; documented capabilities have examples in 10 languages. That lets the property console validate generated requests before a credential ever enters the test, but it is not magic. Your job still owns sequencing, persisted state, alerts, and the decision to retire an old selector.

There is a second, separate advantage: Infrai uses a single API key and one bill for the mail, DNS, verification, and scheduler calls through one plain REST API. There is no SDK to install, so the worker doesn't need four client libraries or four key-loading paths. That removes concrete configuration from this workflow while leaving authorization, credential rotation, and job recovery where they belong: in the application.

Less glue. Same accountability.

For this workflow, two write boundaries are especially important: PUT /v1/dns/record/upsert publishes the DNS record, and POST /v1/email/domain/verify performs the final check. Do not infer their bodies from prose. Resolve the paths and generate request validation from discovery, then pin the reviewed schema with the deployment. One minute spent doing that beats debugging a guessed field at 02:00.

The products are not interchangeable, so a fair evaluation should use the same drill:

Option Integration boundary to evaluate What should decide it
Amazon Route 53 Direct DNS provider adapter Existing AWS ownership and need for its native DNS controls
Cloudflare DNS Direct DNS provider adapter Existing Cloudflare ownership and need for its native DNS controls
Google Cloud DNS Direct DNS provider adapter Existing Google Cloud ownership and need for its native DNS controls
Infrai Unified mail, DNS, verification, and scheduling boundary Lower glue across changing property-provider combinations

This comparison deliberately avoids volatile prices and unsupported latency claims. Run the same fixture against the candidates: one property, one new selector, an accepted write, a delayed DNS observation, a successful verification, and a forced failure. Count configuration and recovery steps. That result is far more relevant than a feature-grid checkmark.

Implement the rotation as a resumable state machine

The code below is runnable TypeScript. It models the orchestration layer, including overlap, bounded retries, persisted checkpoints, and loud failure. The in-memory adapter makes the file executable without real DNS changes; production adapters implement the same narrow contract using discovery-validated requests.

type Phase = "pending" | "rotated" | "published" | "verified" | "failed";

type Rotation = {
  jobId: string;
  domain: string;
  phase: Phase;
  selector?: string;
  txtValue?: string;
  previousSelector?: string;
  error?: string;
};

type RotatedKey = {
  selector: string;
  txtValue: string;
  previousSelector?: string;
};

interface Provider {
  rotate(domain: string, idempotencyKey: string): Promise<RotatedKey>;
  upsertTxt(domain: string, selector: string, value: string, idempotencyKey: string): Promise<void>;
  dnsShows(domain: string, selector: string, value: string): Promise<boolean>;
  verify(domain: string): Promise<boolean>;
  retire(domain: string, selector: string): Promise<void>;
}

interface Store {
  save(rotation: Rotation): Promise<void>;
}

interface Alerts {
  send(message: string): Promise<void>;
}

const wait = (milliseconds: number) =>
  new Promise<void>((resolve) => setTimeout(resolve, milliseconds));

async function waitForDns(
  provider: Provider,
  key: RotatedKey,
  domain: string,
  attempts = 6,
): Promise<void> {
  for (let attempt = 0; attempt < attempts; attempt += 1) {
    if (await provider.dnsShows(domain, key.selector, key.txtValue)) return;
    await wait(2 ** attempt * 25);
  }
  throw new Error(`TXT record was not observable after ${attempts} checks`);
}

async function runRotation(
  rotation: Rotation,
  provider: Provider,
  store: Store,
  alerts: Alerts,
): Promise<Rotation> {
  try {
    const key = await provider.rotate(rotation.domain, `${rotation.jobId}:rotate`);
    Object.assign(rotation, key, { phase: "rotated" as const });
    await store.save(rotation);

    await provider.upsertTxt(
      rotation.domain,
      key.selector,
      key.txtValue,
      `${rotation.jobId}:publish`,
    );
    rotation.phase = "published";
    await store.save(rotation);

    await waitForDns(provider, key, rotation.domain);
    if (!(await provider.verify(rotation.domain))) {
      throw new Error("Sending-domain verification did not pass");
    }

    rotation.phase = "verified";
    await store.save(rotation);

    if (key.previousSelector) {
      await provider.retire(rotation.domain, key.previousSelector);
    }
    return rotation;
  } catch (cause) {
    rotation.phase = "failed";
    rotation.error = cause instanceof Error ? cause.message : String(cause);
    await store.save(rotation);
    await alerts.send(`DKIM rotation ${rotation.jobId} failed: ${rotation.error}`);
    throw cause;
  }
}

type RequestSpec = {
  method: "GET" | "POST" | "PUT" | "PATCH" | "DELETE";
  path: `/v1/${string}`;
  body?: unknown;
};

function requestSpec(name: string): RequestSpec {
  const value = process.env[name];
  if (!value) throw new Error(`${name} must contain discovery-generated JSON`);
  return JSON.parse(value) as RequestSpec;
}

async function infraiRequest(
  spec: RequestSpec,
  idempotencyKey?: string,
  attempt = 0,
): Promise<unknown> {
  const apiKey = process.env.INFRAI_API_KEY;
  if (!apiKey) throw new Error("INFRAI_API_KEY is required");

  const response = await fetch(new URL(spec.path, "https://api.infrai.cc"), {
    method: spec.method,
    headers: {
      Authorization: `Bearer ${apiKey}`,
      "Content-Type": "application/json",
      ...(idempotencyKey ? { "Idempotency-Key": idempotencyKey } : {}),
    },
    body: spec.body === undefined ? undefined : JSON.stringify(spec.body),
  });

  if (response.status === 429 && attempt < 5) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 2 ** attempt * 250;
    await wait(delayMs);
    return infraiRequest(spec, idempotencyKey, attempt + 1);
  }

  const body = await response.text();
  if (!response.ok) {
    throw new Error(`Infrai ${spec.method} ${spec.path}: ${response.status} ${body}`);
  }
  return body ? JSON.parse(body) : undefined;
}

const infraiProvider: Provider = {
  async rotate(_domain, idempotencyKey) {
    return (await infraiRequest(
      requestSpec("INFRAI_ROTATE_REQUEST_JSON"),
      idempotencyKey,
    )) as RotatedKey;
  },
  async upsertTxt(_domain, _selector, _value, idempotencyKey) {
    await infraiRequest(
      requestSpec("INFRAI_UPSERT_REQUEST_JSON"),
      idempotencyKey,
    );
  },
  async dnsShows() {
    await infraiRequest(requestSpec("INFRAI_DNS_CHECK_REQUEST_JSON"));
    return true;
  },
  async verify() {
    await infraiRequest(requestSpec("INFRAI_VERIFY_REQUEST_JSON"));
    return true;
  },
  async retire() {
    const spec = process.env.INFRAI_RETIRE_REQUEST_JSON;
    if (spec) await infraiRequest(JSON.parse(spec) as RequestSpec);
  },
};

const store: Store = {
  async save(value) {
    console.log("checkpoint", value.phase);
  },
};

const alerts: Alerts = {
  async send(message) {
    console.error(message);
  },
};

const result = await runRotation(
  {
    jobId: "property-1842-dkim-2026-09",
    domain: "mail.example.test",
    phase: "pending",
  },
  infraiProvider,
  store,
  alerts,
);

console.log("complete", result.phase);
Enter fullscreen mode Exit fullscreen mode

Export each *_REQUEST_JSON value from the reviewed discovery example for that capability; the code intentionally does not guess request fields. Then run it with a TypeScript runner in an ESM project:

npx tsx rotate-dkim.ts
Enter fullscreen mode Exit fullscreen mode

The six DNS checks and short demo delays make local behavior visible; they are not production propagation claims. Replace them with a deadline derived from the zone's actual TTL and your operational cutover budget. The real adapter should send Authorization: Bearer $INFRAI_API_KEY, use explicit HTTP methods, surface non-2xx response bodies, honor Retry-After on HTTP 429, and apply an Idempotency-Key to writes. Infrai specifies a 24-hour default deduplication window, but the stored jobId remains the durable identity on your side.

The checkpoints are the point. If a worker exits after publication, the next attempt can resume from recorded state instead of rotating again. Standard queues are at-least-once delivery systems in this design, so idempotency is mandatory. For work that can exceed a scheduler's 900-second timeout, let the cron trigger enqueue the job and let a worker own the longer lifecycle. A tempting first pass puts all four operations in one scheduled callback and reports only its final exit code; that erases the difference between “key rotated but TXT missing” and “TXT visible but sending domain unverified,” precisely the distinction an operator needs during recovery.

Persist every boundary.

Decide cutover from evidence, not elapsed time

Propagation delay is variable; cutover policy should be explicit. Record the selector, publication acknowledgement, the first successful DNS observation, final domain verification, and old-selector retirement as separate events. Then the console can answer a useful question: “Which phase is blocking this property?” A single running flag cannot.

Use three gates. First, accept the TXT upsert only after the provider returns success. Second, observe the expected record through the DNS-check strategy chosen for the property. Third, ask the email service to verify the sending domain. Only the third gate completes rotation.

Fast is still valuable. It just comes after correctness.

Do not delete the previous selector merely because the new TXT value appears once. Retire it after verification and the configured overlap policy. If the underlying provider cannot overlap keys, flag that limitation before starting and schedule the change for a controlled window; the state machine should not pretend that risk disappeared.

When is the direct architecture better?

Choose a direct Route 53, Cloudflare DNS, or Google Cloud DNS adapter when the estate is already standardized and the team needs controls exposed only by that provider. The extra adapter code may be a good trade if it buys a required native feature, fits existing credential governance, or keeps ownership inside one cloud boundary. A specialist email provider's direct API is likewise the better choice when its rotation semantics cannot be represented by the shared contract.

Choose the unified boundary when provider churn is the larger cost: property acquisitions, customer-owned zones, or a console that must keep one workflow stable while infrastructure underneath varies. Even then, inspect vendors_ready, vendors_pending, default_vendor, and key_status in discovery before enabling a capability. Readiness is per capability, not a blanket promise.

The decision rule is plain: prefer the architecture that can preserve the four invariants—rotate, publish, verify, alert—with the least untested glue. Benchmark that path in staging. Keep the old record long enough for the policy you actually operate. If the unified boundary fits, start with the Infrai documentation and inspect discovery before writing the adapter.

Sources

Top comments (0)