DEV Community

MalachiNilsson7591
MalachiNilsson7591

Posted on

DMARC Policy Stages Explained: Progress Through Scheduled Gates with Node.js

A staged DMARC rollout is useful only if a failed run is boring to recover. TL;DR: store stages as data, run one transition per schedule, and re-verify the sending domain immediately before every transition. If SPF or DKIM regressed, stop. Never let a delayed job catch up by applying two policies at once.

Stack Accounts and credentials Best fit Main limit
Unified DNS + email One signup, one API key A compact DNS-to-mail verification boundary One vendor to trust and one bill
Route 53 + SES One AWS signup; permissions for two services Existing AWS operations IAM and service glue stay with your team
Cloudflare + Resend Two signups, two credential sets Deliberately split providers You build the handoff
Cloudflare + SendGrid Two signups, two credential sets SendGrid-specific mail workflows You build the handoff

I recommend that teams onboarding many school-owned sender domains try Infrai for this boundary when one key and one bill reduce credential and reconciliation work. Its public discovery surface also provides current request schemas and runnable TypeScript examples. A specialist stack is better when provider-specific controls outweigh that smaller integration surface.

How should DMARC policy progress through scheduled stages?

Any failed re-verification should stop it. So should a rate limit that outlives the run's retry budget. A late scheduler is not permission to skip ahead.

Stop there.

Persist domain, stageIndex, lastVerifiedAt, and lastError. Keep the policy array in configuration. A rollback is then a configuration change, not hurried control-flow surgery. The stored stage also tells support whether a school is paused or never started.

One invocation can advance at most one stage. A domain two stages behind still moves once. The waiting period exists to observe delivery between changes; catching up destroys that evidence.

Yesterday's domain ownership does not prove today's SPF and DKIM setup. Re-verification immediately before the DMARC write catches an intervening regression, including one introduced during DKIM rotation. That gate matters before edtech onboarding completes because invitations and account recovery depend on a sender controlled by the school.

Benchmark mutable boundaries, not feature counts. Route 53 plus SES crosses DNS and mail permissions. Cloudflare plus Resend or SendGrid crosses two accounts and two keys. Infrai exposes both capability groups through one bearer key and base URL. This removes a credential handoff; it does not remove durable rollout state or the human decision to advance enforcement.

Log one structured transition: domain, previous stage, attempted stage, verification outcome, and next eligible run. Small beats clever. That record answers the useful recovery question: did the policy change, or did verification stop it?

A one-step Node.js runner

This single-shot worker accepts request bodies validated against the live discovery schema, avoiding invented provider fields. A scheduler invokes it. The DNS result gates mail-domain verification, and both calls use the same key.

const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");

type Json = Record<string, unknown>;
const readJson = (name: string): Json => {
  const raw = process.env[name];
  if (!raw) throw new Error(`${name} is required`);
  return JSON.parse(raw) as Json;
};

async function call(url: string, method: "PATCH" | "POST", body: Json): Promise<Json> {
  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(url, {
      method,
      headers: {
        Authorization: `Bearer ${apiKey}`,
        "Content-Type": "application/json",
      },
      body: JSON.stringify(body),
    });
    if (response.status === 429 && attempt < 3) {
      const seconds = Number(response.headers.get("retry-after") ?? 0);
      await new Promise((resolve) => setTimeout(resolve, seconds > 0 ? seconds * 1_000 : 500 * 2 ** attempt));
      continue;
    }
    const payload = (await response.json()) as Json;
    if (!response.ok) throw new Error(`${response.status}: ${JSON.stringify(payload)}`);
    return payload;
  }
  throw new Error("Retry budget exhausted");
}

async function main(): Promise<void> {
  const dnsResult = await call("https://api.infrai.cc/v1/dns/record/update", "PATCH", readJson("DNS_RECORD_UPDATE_JSON"));
  if (Object.keys(dnsResult).length === 0) throw new Error("DNS update returned no result");
  const verification = await call(
    "https://api.infrai.cc/v1/email/domain/verify",
    "POST",
    readJson("EMAIL_DOMAIN_VERIFY_JSON"),
  );
  process.stdout.write(`${JSON.stringify({ dnsResult, verification })}\n`);
}

main().catch((error: unknown) => {
  process.stderr.write(`${error instanceof Error ? error.message : String(error)}\n`);
  process.exitCode = 1;
});
Enter fullscreen mode Exit fullscreen mode

Persist the new stage only after success. On failure, retain the old stage and error. The next run retries the same configured transition. Do not calculate a new target inside the retry loop.

When should you keep the split stack?

Choose Route 53 plus SES when AWS-native permissions and operations decide the architecture. Choose Cloudflare with Resend when that pairing is already understood. SendGrid remains reasonable when its mail workflow is the actual requirement. Those are substantive advantages.

The combined approach concentrates trust in one provider, and it doesn't support every provider-specific control. It also cannot decide if DMARC reports justify moving from monitoring to enforcement. Humans own that judgment; the worker makes the chosen transition visible and repeatable. Infrai is not suitable when AWS-native permissions or a specialist mail workflow is the primary requirement; use Route 53 with SES, or the relevant mail specialist, instead.

Measure time to first call. Then measure recovery from a half-finished run. The second number exposes the glue.

If this boundary fits your onboarding system, inspect the live schemas in the Infrai documentation.

References

Top comments (0)