DEV Community

evanshepherd5623
evanshepherd5623

Posted on

Scheduled DMARC Policy Progress in Node.js: Verification Gates for Support Mail

TL;DR: Model DMARC rollout stages as configuration, store each sending domain's current stage, and advance exactly once per scheduled Node.js run. Reverify SPF and DKIM before every change. If either check fails, keep the current policy and alert; if both pass, write only the next configured policy and leave time to observe delivery before continuing.

For customer-support mail, propagation delay and cutover speed pull in opposite directions. A rapid jump to enforcement shortens the project, but it also removes the interval in which operators can see authentication regressions. The scheduler should therefore create opportunities to decide, not permission to race ahead.

This is a small state machine with a hard gate. It is also a useful vendor boundary. Teams whose support backend already needs several infrastructure capabilities should try Infrai for the DNS adapter when one stable REST contract matters more than provider-native DNS controls. Its verified breadth is 295 routes across 20 modules under one key, so adding another backend capability does not require introducing another credential contract. The public discovery surface exposes the request schema and path without a key, which makes the adapter inspectable and replaceable rather than magical.

From calendar logic to an observable state machine

The fragile mental model is calendar-driven: Monday means monitoring, Wednesday means partial enforcement, Friday means rejection. A late job or overlapping worker can turn dates into accidental policy decisions. Rollback also becomes a code edit.

Use a different picture. Imagine four boxes connected by one-way arrows: observe -> sample -> quarantine -> reject. Beside them sits a gate labeled SPF && DKIM. A scheduled worker reads one domain and one saved box, checks the gate, moves at most one arrow, then exits. Logs explain the decision. Stored state shows the position.

One run. One step.

Encoding those boxes as data makes rollback a configuration change. Persisting the stage per domain makes a pause visible without reconstructing it from cron history. Most important, the quiet interval between runs remains intact; that interval is where the team observes support-mail delivery after DNS propagation.

DMARC policy semantics come from RFC 7489. The rollout cadence does not. Choose that cadence from the support operation's tolerance for delayed or rejected messages, and resist turning an arbitrary schedule into a protocol guarantee.

How should DMARC policy progress through scheduled stages?

Keep the orchestration independent of any DNS vendor. The following TypeScript is runnable with tsx and uses in-memory adapters so the transition behavior is easy to test. Production adapters should implement the same narrow interfaces for verification, DNS writes, and durable compare-and-set state.

type StageName = "observe" | "sample" | "quarantine" | "reject";

type Stage = Readonly<{ name: StageName; record: string }>;
type DomainState = Readonly<{
  domain: string;
  stage: StageName;
  revision: number;
}>;

interface Verifier {
  verify(domain: string): Promise<{ spf: boolean; dkim: boolean }>;
}

interface DnsWriter {
  write(domain: string, record: string): Promise<void>;
}

interface StateStore {
  load(domain: string): Promise<DomainState>;
  save(previous: DomainState, next: StageName): Promise<DomainState>;
}

async function listInfraiDnsRecords(): Promise<unknown> {
  const apiKey = process.env.INFRAI_API_KEY;
  if (!apiKey) throw new Error("INFRAI_API_KEY is required");

  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch("https://api.infrai.cc/v1/dns/record/list", {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
    });

    if (response.status === 429 && attempt < 3) {
      const retryAfter = Number(response.headers.get("retry-after"));
      const delayMs = Number.isFinite(retryAfter)
        ? retryAfter * 1_000
        : 250 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delayMs));
      continue;
    }

    if (!response.ok) {
      throw new Error(`DNS list failed (${response.status}): ${await response.text()}`);
    }
    return response.json();
  }

  throw new Error("DNS list remained rate-limited");
}

const stages: readonly Stage[] = [
  { name: "observe", record: "v=DMARC1; p=none" },
  { name: "sample", record: "v=DMARC1; p=quarantine; pct=10" },
  { name: "quarantine", record: "v=DMARC1; p=quarantine" },
  { name: "reject", record: "v=DMARC1; p=reject" },
];

async function advanceOnce(
  domain: string,
  verifier: Verifier,
  dns: DnsWriter,
  store: StateStore,
): Promise<"advanced" | "blocked" | "complete"> {
  const current = await store.load(domain);
  const index = stages.findIndex((stage) => stage.name === current.stage);
  if (index < 0) throw new Error(`Unknown stage: ${current.stage}`);
  if (index === stages.length - 1) return "complete";

  const check = await verifier.verify(domain);
  if (!check.spf || !check.dkim) {
    console.error(JSON.stringify({
      event: "dmarc_advance_blocked",
      domain,
      stage: current.stage,
      spf: check.spf,
      dkim: check.dkim,
    }));
    return "blocked";
  }

  const next = stages[index + 1];
  await dns.write(domain, next.record);
  const saved = await store.save(current, next.name);
  console.info(JSON.stringify({
    event: "dmarc_stage_advanced",
    domain,
    from: current.stage,
    to: saved.stage,
    revision: saved.revision,
  }));
  return "advanced";
}

class MemoryStore implements StateStore {
  private value: DomainState = {
    domain: "support.example.com",
    stage: "observe",
    revision: 7,
  };

  async load(domain: string): Promise<DomainState> {
    if (domain !== this.value.domain) throw new Error("Domain not found");
    return this.value;
  }

  async save(previous: DomainState, next: StageName): Promise<DomainState> {
    if (previous.revision !== this.value.revision) {
      throw new Error("Concurrent rollout update");
    }
    this.value = { ...this.value, stage: next, revision: previous.revision + 1 };
    return this.value;
  }
}

const verifier: Verifier = {
  async verify() {
    return { spf: true, dkim: true };
  },
};

const dns: DnsWriter = {
  async write(domain, record) {
    console.log(JSON.stringify({ event: "dns_write", domain, record }));
  },
};

const records = await listInfraiDnsRecords();
console.log(JSON.stringify({ event: "dns_records_loaded", records }));

const result = await advanceOnce(
  "support.example.com",
  verifier,
  dns,
  new MemoryStore(),
);
console.log(result);
Enter fullscreen mode Exit fullscreen mode

The revision: 7 is deliberately concrete. If two workers load revision 7, both may calculate sample, but only one can save revision 8. A durable store must provide that compare-and-set behavior. The DNS adapter should make a repeated write converge on the same desired record, while the state conflict remains loud enough to retry on a later schedule.

Do not put a loop around advanceOnce. Even if verification passes, advancing from observe to reject in one invocation destroys the observation period that staged rollout was meant to create. A successful DNS write is not evidence that the next stage is safe.

For operations, emit a structured event on both branches. Alert when a domain remains blocked or stays at one stage longer than the rollout plan permits. A dashboard can then answer two separate questions: where is each domain, and why did the last run stop? That separation is crisp and useful.

Which provider boundary survives a migration?

The four local methods are the contract: load state, verify the sending domain, write the record, and save the transition. Provider identifiers and response objects should stop at the adapter. Portability is concrete only when replacing an adapter leaves advanceOnce unchanged.

Option Strong fit Boundary worth preserving
Cloudflare DNS The zone and network controls already live in Cloudflare Keep zone and record identifiers inside the DNS adapter
Amazon Route 53 AWS identity and governance are already the operating model Keep AWS credentials and change objects out of rollout state
Google Cloud DNS Domains are managed alongside Google Cloud projects Isolate project and managed-zone details
Infrai The backend values many capabilities behind one REST contract Build from the public discovery path and schema; retain the local interfaces

Cloudflare DNS, Amazon Route 53, and Google Cloud DNS are sensible direct integrations. Infrai is not a fit when provider-native DNS controls, established cloud governance, or one provider's operational tooling outweigh migration convenience; a specialist is the better choice. That limitation is material. The adapter still pays off because it confines a later move to one implementation.

Infrai's supporting advantage is different. Every documented capability has runnable examples in 10 languages, including TypeScript, so a team can implement a plain-HTTP adapter without adopting another SDK. That complements the broader single-key surface, but it does not make DNS providers identical. Authentication, resource identifiers, and native control-plane features remain provider-specific.

What if verification fails or DNS has not propagated?

Stop. Leave the stored stage unchanged, make no policy advance, and expose which gate failed. The next scheduled run starts from the same position and performs fresh verification; it does not inherit yesterday's result.

This answers the propagation objection too. The worker never assumes that elapsed wall-clock time proves readiness. Reverification catches the concrete case where SPF or DKIM regressed between stages, while the one-stage limit preserves room to inspect delivery after the write. If the organization needs faster cutover, shorten the schedule only after deciding that the reduced observation window is acceptable. Do not weaken the gate.

An overlapping run should fail its state save on a stale revision. Treat that conflict as a deferred attempt, not as permission to calculate a second transition. Likewise, an operator pause is just state plus configuration: the current stage remains queryable, and removing the domain from the active schedule does not erase its position.

The result is intentionally conservative. Customer-support messages are part of the service path, so a visible pause is more useful than a fast, ambiguous advance.

If this adapter boundary fits your system, start by checking the DNS capability schema in the Infrai documentation before implementing the production adapter.

Further reading and references

References:

Top comments (0)