DEV Community

Keria
Keria

Posted on

Implementing 3 Scheduled DMARC Policy Gates — Reverify Every Cutover

Progress a DMARC policy through scheduled stages only when two independent checks agree: the organization still controls the rollout, and public DNS returns the exact policy expected from the previous stage. A calendar can decide when to check. It should never decide that a domain is safe to advance.

Short answer: store an expected policy fingerprint beside each scheduled transition, re-check a separate ownership token, query multiple recursive resolvers, and cancel the write if any answer is stale or inconsistent. For an edtech platform moving school email zones away from a registrar-specific API, that makes propagation delay explicit. A fast cutover remains possible, but uncertainty stops the state machine instead of silently promoting p=none to enforcement.

How should a DMARC policy progress through scheduled stages?

DMARC receivers discover policy in a TXT record at _dmarc.<domain>. RFC 7489 defines the requested handling policies none, quarantine, and reject; it also defines pct as the percentage of messages to which the policy applies, with a default of 100. Those are receiver instructions, not proof that a new zone is fully visible everywhere.

That distinction matters during an edtech DNS migration. A registrar control panel may show the new record while recursive resolvers still return an older cached value. The reverse can happen during rollback. If a midnight worker trusts only its database row, it can overwrite a manual correction or advance a zone whose delegation is still settling.

The practical model has three sources of truth: desired rollout state in the application, observed state from public DNS, and authorization represented by a high-entropy token under a separate application-owned label. The token is an application convention, not part of DMARC. Keeping it separate avoids teaching the scheduler that the policy record itself proves permission to change the policy record. Before advancing, reverify all three inputs rather than treating yesterday's successful lookup as durable authorization.

Three states. Two resolvers. One fail-closed gate.

Implement the verification gate first

The example below uses only Node.js built-ins. Save it as rollout.ts, run it with a TypeScript runner, and replace the in-memory writeDmarc function with the narrow DNS adapter used by your infrastructure. The adapter boundary is deliberate: zone migration should not leak a registrar-specific request shape into the state machine.

import { Resolver } from "node:dns/promises";
import { createHash, timingSafeEqual } from "node:crypto";

type Policy = "none" | "quarantine" | "reject";

type Stage = {
  policy: Policy;
  pct: number;
  notBefore: Date;
};

type Rollout = {
  domain: string;
  ownershipTokenHash: string;
  expectedRecordHash: string;
  next: Stage;
};

const resolverAddresses = ["1.1.1.1", "8.8.8.8"];

function normalizeTxt(chunks: string[][]): string[] {
  return chunks.map((parts) => parts.join("").trim()).sort();
}

function sha256(value: string): string {
  return createHash("sha256").update(value).digest("hex");
}

function equalHex(left: string, right: string): boolean {
  const a = Buffer.from(left, "hex");
  const b = Buffer.from(right, "hex");
  return a.length === b.length && timingSafeEqual(a, b);
}

async function txtFrom(server: string, name: string): Promise<string[]> {
  const resolver = new Resolver();
  resolver.setServers([server]);
  return normalizeTxt(await resolver.resolveTxt(name));
}

async function unanimousTxt(name: string): Promise<string[]> {
  const answers = await Promise.all(
    resolverAddresses.map((server) => txtFrom(server, name)),
  );
  const baseline = JSON.stringify(answers[0]);
  if (!answers.every((answer) => JSON.stringify(answer) === baseline)) {
    throw new Error(`DNS answers disagree for ${name}; retry later`);
  }
  return answers[0];
}

function buildRecord(stage: Stage): string {
  if (!Number.isInteger(stage.pct) || stage.pct < 0 || stage.pct > 100) {
    throw new Error("pct must be an integer from 0 through 100");
  }
  return `v=DMARC1; p=${stage.policy}; pct=${stage.pct}`;
}

async function writeDmarc(domain: string, value: string): Promise<void> {
  // Replace this port with an idempotent zone writer in production.
  process.stdout.write(`WRITE _dmarc.${domain} ${JSON.stringify(value)}\n`);
}

async function advance(rollout: Rollout, now = new Date()): Promise<void> {
  if (now < rollout.next.notBefore) return;

  const ownershipName = `_rollout-control.${rollout.domain}`;
  const ownershipAnswers = await unanimousTxt(ownershipName);
  if (ownershipAnswers.length !== 1) {
    throw new Error("Expected exactly one ownership token");
  }
  if (!equalHex(sha256(ownershipAnswers[0]), rollout.ownershipTokenHash)) {
    throw new Error("Ownership verification failed");
  }

  const current = await unanimousTxt(`_dmarc.${rollout.domain}`);
  if (current.length !== 1) {
    throw new Error("Expected exactly one DMARC TXT record");
  }
  if (!equalHex(sha256(current[0]), rollout.expectedRecordHash)) {
    throw new Error("Observed policy does not match the scheduled predecessor");
  }

  await writeDmarc(rollout.domain, buildRecord(rollout.next));
}

const rollout: Rollout = {
  domain: "school.example",
  ownershipTokenHash: sha256("replace-with-the-published-random-token"),
  expectedRecordHash: sha256("v=DMARC1; p=none; pct=100"),
  next: {
    policy: "quarantine",
    pct: 10,
    notBefore: new Date("2026-10-01T02:00:00Z"),
  },
};

advance(rollout).catch((error: unknown) => {
  process.stderr.write(`${error instanceof Error ? error.message : String(error)}\n`);
  process.exitCode = 1;
});
Enter fullscreen mode Exit fullscreen mode

The example deliberately refuses multiple TXT answers. DNS TXT data can be split into character strings inside one record, which is why normalizeTxt joins each record's chunks before comparing it. Multiple DMARC records, however, do not create a useful merge strategy for this worker.

Stop there.

Two public recursive resolvers are a sampling strategy, not a proof of global convergence. Add resolvers that reflect the regions and networks serving your schools, and make the set configuration rather than code. More probes improve confidence but increase query volume and can lengthen the gate when one cache lags. The trade-off is explicit: a broader sample raises confidence in convergence while a smaller sample permits a faster cutover. Neither choice changes what DNS guarantees, so record the resolver set with the transition and choose it from the networks that matter to the schools rather than from a generic global count.

Move through three enforcement states

Use a small state machine: observation at p=none, limited quarantine, then full enforcement. For example, a team could prepare none/100, quarantine/10, and reject/100. Those values are a rollout design, not a universal safety prescription. The evidence required to move between them comes from authentication results and aggregate reports for the domain's legitimate mail streams.

Current state Candidate state Evidence before scheduling
p=none; pct=100 p=quarantine; pct=10 Known learning, billing, and support senders align with SPF or DKIM
p=quarantine; pct=10 p=reject; pct=100 Aggregate reports show no unexplained legitimate stream that would lose delivery
Any state Hold Ownership token, expected policy, or resolver answers disagree

DMARC's pct mechanism lets a domain request that policy be applied to only a percentage of affected messages. It is useful for gradual enforcement, but it is not a deterministic deployment percentage for a particular school, sender, or message class. Do not describe pct=10 as a canary that guarantees exactly one tenth of each traffic segment.

The scheduler should persist the hash of the record it just wrote as the next transition's expected predecessor. It should also record observations per resolver, the candidate value, the decision, and a transition identifier. Never log the raw ownership token. On retry, an idempotency key should make an already completed transition a no-op rather than a second blind write. These details are easy to dismiss as bookkeeping, yet they answer the hardest question after a hold: did the scheduled worker see stale DNS, did someone edit the zone, or did a retry revisit a completed stage? Without the observation set and transition identifier, all three cases can look like the same hash mismatch.

Keep the hold boring.

Treat disagreement as operational data

Propagation waits are cheap; bad enforcement is not. A held transition needs a reason that an operator can act on: token missing, token changed, resolver disagreement, predecessor mismatch, or report review incomplete. Put those labels in metrics and alert on the age of a pending stage, not merely on worker failure counts.

A predecessor mismatch often means a human edited the record, another automation path won a race, or the cutover exposed different authoritative data. The correct response is to reload and review the rollout, not to force the scheduled value. Use compare-and-set behavior in the zone writer when the DNS control plane supports it. Otherwise, perform a final authoritative read immediately before the write and keep the update idempotent.

This costs a few DNS queries per transition. For a solo team, that is a better cost boundary than continuous polling: schedule a check, use bounded retries with jitter for temporary lookup failures, then park the transition for review. Do not compress the interval merely to make a dashboard turn green.

Ship the cutover with an explicit hold rule

Before enabling the worker, publish the ownership token, inventory every legitimate sender, enable aggregate reporting, and capture the exact starting DMARC record. Exercise the adapter against a delegated test zone. Test stale resolver answers, NXDOMAIN, multiple TXT records, a changed token, a manual policy edit, timeout, retry, and rollback. A passing happy path is the small part.

During the registrar-to-independent-zone cutover, keep the old and new authoritative zones aligned before changing delegation. Schedule advancement only after the DNS observations agree and the mail evidence supports the candidate stage. If either check fails, preserve the current policy and surface a hold; do not let elapsed time convert uncertainty into enforcement.

After each write, read the record back through the configured resolver set and store what was observed. The next stage stays locked until that result becomes its expected predecessor. This makes the runbook slower when caches disagree and fast when they converge, which is exactly the behavior the migration needs.

References

Top comments (0)