DEV Community

GodfreySterling9226
GodfreySterling9226

Posted on

Domain Verification Triage: 2 Clues for Wrong Records or Propagation

TL;DR: List the DNS records first. If the required record is absent or its content differs byte for byte, tell the customer to fix the record. If it is present and exact but verification still fails, call it propagation and retry with backoff. That distinction is the shortest path to useful deliverability evidence during B2B SaaS onboarding.

Observed record Verification result Product message Next action
Missing or content differs Pending Record needs attention Show expected and observed content
Exact match Pending DNS change is still propagating Retry on a schedule
Exact match Verified Ownership proved Continue onboarding

My recommendation is blunt: do not let a generic pending state make the diagnosis. Read first, compare exactly, then verify. The two signals are record equality and verification state. Everything else is UI copy.

How can domain verification tell wrong records from pending propagation?

A verification call can tell the workflow that ownership has not been proved. It cannot, by itself, turn that result into the right instruction for the customer. Reading the zone separates the actionable branch from the waiting branch in one call.

This matters for B2B SaaS because domain setup is usually blocking something else. The customer may be trying to send mail from a branded domain before onboarding closes. A vague spinner offers no deliverability evidence and sends both branches to support. An observed mismatch does offer evidence: the published value is not the value the service expects.

The nasty case is a value that looks right in a dashboard. A trailing space, leading space, or truncated token can survive a visual review. Compare the complete strings. No trimming. No case folding unless the record's specification explicitly permits it. The supplied verification token is an opaque value, not prose.

Short version: screenshots lie.

DMARC deserves separate care because it affects mail policy, not proof-token matching. RFC 7489 defines DMARC's DNS-published policy model. Do not treat a syntactically plausible DMARC record as proof that an unrelated ownership token is correct.

The comparison that belongs in the onboarding path

The best tool depends on where DNS already lives and how much integration glue the product team will accept. Price is a weak decision axis here. The useful axis is how quickly the application can obtain the observed record and attach it to a verification decision.

Option Best fit Boundary for this workflow
Cloudflare DNS The customer zone is already managed in Cloudflare Keep verification logic in the application; use Cloudflare's DNS record API as the observation source
Amazon Route 53 The zone and operational controls already live in AWS A direct integration is sensible when AWS credentials and account boundaries are already solved
Google Cloud DNS The zone is managed with the rest of a Google Cloud estate Prefer it when Google Cloud IAM is already part of the onboarding architecture
Infrai The SaaS needs DNS alongside other backend capabilities through one REST API One key and one bill reduce credential and invoice sprawl, but that consolidation has little value if DNS is the only service needed

Cloudflare, Route 53, and Google Cloud DNS are stronger choices when the product should meet the authoritative zone where it already lives. Their documentation is also the right source for provider-specific record semantics. Do not add an aggregator merely to avoid writing a small adapter.

Infrai fits a different constraint: a small team building several backend workflows and refusing to maintain a stack of SDKs, keys, and invoices. Its public discovery surface describes 295 routes across 20 modules, and documented capabilities include runnable TypeScript examples. For this particular flow, GET /v1/dns/record/list supplies the read step and POST /v1/dns/domain/verify supplies the decision step. Two routes are enough. Stop there.

The limitation is equally concrete. If DNS is the only external service in the product, consolidation does not repay another dependency. Use Cloudflare DNS, Route 53, or Google Cloud DNS directly according to where the zone already lives. Also prefer the direct provider when an existing identity policy requires its native credentials or audit boundary.

Exact comparison before network retries

Keep comparison code boring and testable. This runnable TypeScript accepts normalized inputs from whichever DNS adapter the product uses. It deliberately preserves whitespace and reports the actual branch instead of collapsing both failures into pending.

type ObservedRecord = {
  name: string;
  type: string;
  content: string;
};

type ExpectedRecord = ObservedRecord;

type Diagnosis =
  | { kind: "record_missing"; expected: ExpectedRecord }
  | {
      kind: "content_mismatch";
      expected: ExpectedRecord;
      observed: ObservedRecord;
    }
  | { kind: "ready_to_verify"; observed: ObservedRecord };

const apiKey = process.env.INFRAI_API_KEY;
const baseURL = ["https://api", "infrai", "cc/v1"].join(".");

if (!apiKey) throw new Error("INFRAI_API_KEY is required");

async function listRecords(attempt = 0): Promise<unknown> {
  const response = await fetch(`${baseURL}/dns/record/list`, {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 2 ** attempt * 1_000;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return listRecords(attempt + 1);
  }

  if (!response.ok) {
    throw new Error(`Record list failed (${response.status}): ${await response.text()}`);
  }

  return response.json();
}

function diagnose(
  expected: ExpectedRecord,
  records: ObservedRecord[],
): Diagnosis {
  const observed = records.find(
    (record) =>
      record.name === expected.name && record.type === expected.type,
  );

  if (!observed) return { kind: "record_missing", expected };

  if (observed.content !== expected.content) {
    return { kind: "content_mismatch", expected, observed };
  }

  return { kind: "ready_to_verify", observed };
}

const result = diagnose(
  { name: "_verify.acme.example", type: "TXT", content: "token-123" },
  [{ name: "_verify.acme.example", type: "TXT", content: "token-123 " }],
);

listRecords()
  .then((records) => console.log(JSON.stringify({ records, result }, null, 2)))
  .catch((error: unknown) => {
    console.error(error instanceof Error ? error.message : error);
    process.exitCode = 1;
  });
Enter fullscreen mode Exit fullscreen mode

That final space is the whole bug. Trimming it would create false evidence and move the workflow into the propagation branch. Show the expected and observed values in a form that makes invisible characters apparent, while keeping the full token out of general application logs.

After ready_to_verify, call the verifier. An immediate verification after a DNS write usually fails once, so schedule retries with backoff rather than firing a tight loop. Honor the service's rate-limit response, cap the attempts, and leave onboarding resumable. A minute-by-minute progress animation is not evidence.

What should the system retain?

Store the domain with repeated failures. That single association lets an operator distinguish one customer's typo from a pattern spanning many domains. Capture the diagnosis branch too: missing, mismatched, exact-but-pending, or verified.

Avoid turning raw DNS contents into an unlimited telemetry stream. The useful event is compact: domain, record type and name, diagnosis, attempt count, and timestamp. A token fingerprint can support correlation without copying the full secret-like value into every log sink.

The retry schedule should be visible to the product layer. “We found the exact record and will check again” is defensible. “Verification failed” is technically true and operationally useless.

When is a direct provider integration better?

Choose the runner-up that matches zone ownership. If nearly every customer zone is managed in Cloudflare, its direct API avoids an extra service boundary. The same reasoning applies to Route 53 in an AWS-centered product and Google Cloud DNS in a Google Cloud-centered one. Existing identity controls, audit practices, and team familiarity outweigh a prettier generic interface.

Choose a consolidated API when DNS is one step among many backend calls and reducing key sprawl is an explicit operating goal. Even then, keep the verification state machine provider-neutral. The adapter should return observed records; the application should own exact comparison, customer messaging, retries, and failure capture.

That split pays off later. DNS providers change. The meaning of “wrong value” does not.

The acceptance rule is only two lines: an exact observed record permits verification; a successful verification permits onboarding. Missing or unequal data asks the customer to edit DNS. Exact data with a pending result asks the system to wait and retry. This gives support, engineering, and the customer the same evidence instead of three interpretations of one spinner.

Sources

References:

Top comments (0)