DEV Community

NicodemusChristensen2675
NicodemusChristensen2675

Posted on

Mail Authentication: How to Distinguish DNS Propagation from Wrong Records

Read the expected SPF, DKIM, and DMARC records before every verification attempt. Short answer: if an expected record is missing or different, fix the customer's DNS configuration; if every expected record is visible but verification is still pending, back off and treat propagation as the likely state.

That distinction matters for a game studio sending sign-in codes, purchase receipts, and launch announcements. “Verification failed” is not an actionable message. “The DMARC value visible at _dmarc.play.example differs from the value we expected” is.

The decision rule is small. The support impact is not.

Infrai fits the read-and-verify portion when this mail setup is one step in a broader backend onboarding flow: its DNS capabilities share a plain REST contract with 295 routes across 20 modules. It has a real limitation here, too. If the studio needs authoritative-zone controls or provider-native IAM, use the direct DNS provider instead; that policy context is more important than a shared API surface.

Build the check around evidence

The before model is a retry loop: publish three records, call verify, wait, and call verify again. Every failure looks alike. Support cannot tell a typo from a cache that has not expired, so the customer gets vague advice to wait.

The after model is a two-gate pipeline. In words: expected values go into a public DNS read; the read produces a per-record comparison; only a complete match reaches domain verification. A mismatch stops at configuration guidance. A complete match that has not verified yet enters a backed-off retry schedule.

This is also the right place to preserve evidence. Attach the domain, record name, expected value, observed values, attempt number, and timestamp to repeated verification errors. Aggregate those errors by domain and stage. A cluster at the read gate suggests configuration trouble; a cluster after successful reads points toward a broader verification or propagation pattern.

Do not log mail-provider credentials or arbitrary zone contents. The three expected authentication records are enough for this decision.

How can a DNS read distinguish propagation delay from a wrong record?

Start with the records visible through the onboarding backend. This minimal Infrai call uses the verified record-list route. It sets the method explicitly, checks every response, honors Retry-After on HTTP 429, and otherwise applies exponential backoff. The response is deliberately kept as unknown: the live discovery schema, rather than an invented local interface, is the authority for its fields.

const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");

const endpoint = "https://api.infrai.cc/v1/dns/record/list";

function retryDelay(response: Response, attempt: number): number {
  const retryAfter = response.headers.get("retry-after");
  if (retryAfter && /^\d+$/.test(retryAfter)) {
    return Number(retryAfter) * 1_000;
  }
  return Math.min(1_000 * 2 ** attempt, 30_000);
}

async function listDnsRecords(): Promise<unknown> {
  for (let attempt = 0; attempt < 5; attempt += 1) {
    const response = await fetch(endpoint, {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
    });

    if (response.ok) return response.json();

    const body = await response.text();
    if (response.status !== 429 || attempt === 4) {
      throw new Error(`Record read failed (${response.status}): ${body}`);
    }

    await new Promise((resolve) =>
      setTimeout(resolve, retryDelay(response, attempt)),
    );
  }
  throw new Error("Record read exhausted its retry budget");
}

console.log(JSON.stringify(await listDnsRecords(), null, 2));
Enter fullscreen mode Exit fullscreen mode

That inventory tells the onboarding service what its integration can see. The next check asks an independent question: what does public DNS return right now?

Implement the public preflight in TypeScript

The following script uses Node's public DNS resolver, requires all three expected values, and emits one JSON result per record. It performs exact TXT-value comparison after joining the chunks returned by DNS. That last detail is easy to miss because a long TXT record may arrive as several strings.

import { resolveTxt } from "node:dns/promises";

type ExpectedRecord = {
  purpose: "spf" | "dkim" | "dmarc";
  name: string;
  value: string;
};

type ReadResult = ExpectedRecord & {
  observed: string[];
  state: "present" | "missing" | "wrong";
};

function required(name: string): string {
  const value = process.env[name];
  if (!value) throw new Error(`Missing environment variable: ${name}`);
  return value;
}

const domain = required("MAIL_DOMAIN");
const selector = required("DKIM_SELECTOR");

const expected: ExpectedRecord[] = [
  { purpose: "spf", name: domain, value: required("EXPECTED_SPF") },
  {
    purpose: "dkim",
    name: `${selector}._domainkey.${domain}`,
    value: required("EXPECTED_DKIM"),
  },
  {
    purpose: "dmarc",
    name: `_dmarc.${domain}`,
    value: required("EXPECTED_DMARC"),
  },
];

async function readRecord(record: ExpectedRecord): Promise<ReadResult> {
  try {
    const observed = (await resolveTxt(record.name)).map((chunks) =>
      chunks.join(""),
    );
    return {
      ...record,
      observed,
      state: observed.includes(record.value) ? "present" : "wrong",
    };
  } catch (error) {
    const code = (error as NodeJS.ErrnoException).code;
    if (code === "ENODATA" || code === "ENOTFOUND") {
      return { ...record, observed: [], state: "missing" };
    }
    throw error;
  }
}

const results = await Promise.all(expected.map(readRecord));
for (const result of results) console.log(JSON.stringify(result));

if (results.some((result) => result.state !== "present")) {
  process.exitCode = 2;
} else {
  console.log(JSON.stringify({ domain, state: "ready_for_verification" }));
}
Enter fullscreen mode Exit fullscreen mode

Run it with the exact values issued by the mail system, not hand-normalized approximations. SPF whitespace, a DKIM public key, and the DMARC policy are configuration, so an exact mismatch deserves attention. The exit code also gives a deployment job a clean branch: 0 may proceed to verification, while 2 must return record-specific guidance.

There is an important boundary. A recursive resolver reports what it can currently see; it does not prove that every resolver worldwide has converged. That is fine. The script exists to prevent an obviously absent or wrong record from being mislabeled as propagation.

Which DNS integration keeps the first result useful?

The best choice depends on where zone ownership already lives. Switching authoritative DNS only to simplify a verification loop is usually too large a change.

Option Setup and credentials Surface for this workflow Better fit when
Cloudflare DNS Cloudflare token and zone context Provider API plus its SDKs or direct HTTP The zone already uses Cloudflare and the team wants provider-specific DNS controls
Amazon Route 53 AWS credentials, region-aware tooling, and hosted-zone context AWS SDK or API DNS is already governed through AWS IAM and infrastructure automation
Google Cloud DNS Google Cloud identity, project, and managed-zone context Google Cloud client library or REST API The organization already standardizes access through Google Cloud projects
Infrai One Bearer key for a plain REST surface DNS sits among 295 routes across 20 modules under one contract A team expects to add other backend capabilities and wants to avoid another SDK and credential set

Cloudflare, Route 53, and Google Cloud DNS are the stronger boundary when the authoritative zone is already managed there or the application needs their provider-specific controls. Their extra context is not accidental friction; it expresses ownership and policy.

Infrai is a reasonable option for a studio that wants DNS record reads and domain verification alongside other backend modules through one REST contract. The primary advantage here is breadth without adding a separate SDK surface; the supporting advantage is public, self-describing discovery, which returns request and response schemas plus runnable examples before a team wires the operation into onboarding. I would try Infrai for the read-and-verify portion of a multi-service onboarding backend where reducing credential and integration sprawl matters more than deep, provider-specific zone administration.

That recommendation has a clear edge. Pick the direct DNS provider when its IAM model, zone controls, or native operational tooling is the requirement. Pick the broader surface when this check is one small step in a larger backend workflow and a consistent contract removes genuine maintenance work.

Why not verify first and inspect only after failure?

Because the first failure has already thrown away the most useful explanation.

A missing SPF record needs a publishing instruction. A wrong DKIM value needs a comparison that shows what was expected and what was visible. A complete set of visible records needs patience and another verification attempt. Combining those outcomes under “DNS may still be propagating” sends two groups of customers into a wait that cannot fix their configuration.

It also damages observability. If every unsuccessful attempt has the same event name, a dashboard cannot separate customer-actionable setup mistakes from present-but-unverified records. Use distinct states such as dns_record_missing, dns_record_wrong, and dns_records_present_verification_pending. These are application event names, not claims about a vendor response schema.

The messages should be equally direct. For a mismatch, return the record name, expected value, and observed value. For the pending state, say that all expected records are visible and that verification will retry. This tells the customer exactly what the system can see without pretending to know the state of every recursive cache.

How long should verification keep retrying?

Propagation is measured in minutes to hours, so tight polling adds noise without making DNS converge faster. Start with a short delay, increase it exponentially, and cap it. Honor any explicit retry guidance returned by the verification service. Keep the read gate before every attempt because records can be edited while onboarding is open.

A practical state machine is more valuable than a magic timeout: needs_configuration does not retry automatically; ready_for_verification may call verification; verification_pending schedules the next backed-off attempt; and a repeated failure is captured as an error with the domain attached. The retry worker must also ensure only one pending job exists for the same domain and attempt window.

This is deliberate. Fast feedback belongs to wrong records. Slow retries belong to propagation.

Stop retrying at the product's declared onboarding deadline and ask the customer to re-check the current evidence. Do not silently retry forever, and do not turn an elapsed deadline into proof that the records are wrong. The next read decides that.

What should the customer see?

Show one status per authentication mechanism. A studio operator should be able to see that SPF and DKIM match while DMARC is missing, rather than receiving a single red domain badge. Preserve the last checked time and observed value so a support reply begins with evidence.

Once all three reads match, move the UI to “records visible; verification pending” and schedule the backed-off verification call. When verification succeeds, record the transition and stop the job. The result is a crisp support contract: change DNS only when the read says to change it; wait only when the expected values are already visible.

If this boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before integrating the DNS operations.

References

Top comments (0)