DEV Community

YatesHolloway6872
YatesHolloway6872

Posted on

Node.js Marketplace Domain Verification with 2 Evidence Paths (Webhooks Need Polling)

Short answer: use webhooks to start verification quickly, but let a slow poller recover completion; neither signal should complete marketplace onboarding until a fresh DNS read produces the deliverability evidence your policy requires.

Design Time to first check Missed-event recovery Operational burden Best fit
Webhook only Fast after an event Weak Low at first, painful later Closed systems with replay guarantees
Polling only Bounded by the interval Built in Predictable background load Low-volume or batch onboarding
Webhook plus polling Fast in the usual path Built in One shared queue and state machine Customer-facing marketplace onboarding

For a one-person SaaS, the third row is the practical default. The webhook buys responsiveness. Polling buys eventual progress. The shared verification job keeps the combination small enough to own while shipping weekly.

The catch is real: two triggers create more moving parts than one. They earn their keep only when they feed the same idempotent transition and store the same evidence. If the application cannot enforce that rule, polling alone is the safer runner-up.

How should Node.js SaaS onboarding combine webhook events and polling?

Treat both mechanisms as hints that say, "look again." A webhook is not proof that the customer's DNS is correct, and a timer firing is not proof either. Each trigger should enqueue a verification attempt for a domain record. The worker then reads DNS, evaluates the result against policy, and writes an observation before it changes onboarding state.

That separation matters for email domains. DMARC is published as a DNS TXT record under the _dmarc label. Its policy works with identifier alignment: the domain visible in the message's From field is compared with identifiers authenticated by SPF and DKIM. RFC 7489 also defines aggregate reporting, which can supply evidence after mail starts flowing. A green onboarding badge cannot promise inbox placement, but it can show exactly which DNS and alignment prerequisites were observed.

Keep the customer-facing states boring: pending, checking, verified, and action_required. Put trigger delivery in a separate log. Then a duplicate webhook, an early callback, or a poll that overlaps a queued event can't manufacture a second state transition.

Fast is useful. Evidence wins.

Build the evidence record before the trigger path

The hard design question is what verified means. Define that before choosing event delivery. For a marketplace that lets sellers use their own sending domain, a useful record contains the domain being checked, the policy version, the observation time, the TXT answers used in the decision, and a reason code. It should also preserve the distinction between "no qualifying observation yet" and "the latest observation conflicts with policy."

Here is a compact policy for the example system. It is a product decision, not a claim that every mail program needs the same gates.

Evidence Onboarding use Why retain it
Expected ownership TXT value is observed Proves control for this workflow Explains why the domain advanced
DMARC TXT record is observed at the expected label Shows the published domain policy Supports later policy review
Policy version and observation timestamp Makes the decision reproducible Prevents an old result from looking current
Machine-readable reason code Drives the next UI action Keeps support text out of worker logic

Do not collapse this into verified: true. A boolean answers yesterday's question forever. An evidence row says what was seen, under which rule, and when. That gives support a useful answer when a seller edits DNS after setup, and it lets a future recheck change the current status without rewriting history.

There is also an important boundary here. DMARC's reporting and policy machinery concerns message authentication and handling; the onboarding check is only a preflight. Actual deliverability depends on evidence the DNS probe cannot produce. I'm not sure how quickly any particular customer's resolver view will reflect a change, so the next scheduled observation, rather than a guessed propagation promise, resolves that uncertainty.

Use one idempotent Node.js transition

The implementation can stay small. The webhook handler authenticates and parses its incoming event at the transport boundary, while the scheduler selects records whose next check is due. Both enqueue the same domain ID. The worker owns DNS access and the state transition.

The following TypeScript sketches that boundary without tying it to a DNS or queue vendor. DnsReader, DomainStore, and JobQueue are application interfaces, so adapters can change without changing the verification policy.

type DomainStatus = "pending" | "checking" | "verified" | "action_required";

type DomainRecord = {
  id: string;
  hostname: string;
  expectedOwnershipValue: string;
  status: DomainStatus;
  policyVersion: 2;
  nextCheckAt: Date | null;
};

type Evidence = {
  domainId: string;
  observedAt: Date;
  policyVersion: 2;
  ownershipAnswers: string[];
  dmarcAnswers: string[];
  reason: "requirements_met" | "ownership_missing" | "dmarc_missing";
};

interface DnsReader {
  txt(hostname: string): Promise<string[]>;
}

interface DomainStore {
  get(id: string): Promise<DomainRecord | null>;
  saveEvidence(evidence: Evidence): Promise<void>;
  transition(id: string, from: DomainStatus[], to: DomainStatus): Promise<boolean>;
  scheduleNextCheck(id: string, at: Date | null): Promise<void>;
}

interface JobQueue {
  enqueueVerification(domainId: string, dedupeKey: string): Promise<void>;
}

const normalizeTxt = (value: string): string => value.trim().replace(/^"|"$/g, "");

async function requestVerification(
  domainId: string,
  queue: JobQueue,
  windowId: string,
): Promise<void> {
  await queue.enqueueVerification(domainId, `domain:${domainId}:${windowId}`);
}

async function verifyDomain(
  domainId: string,
  dns: DnsReader,
  store: DomainStore,
  now: Date,
): Promise<void> {
  const domain = await store.get(domainId);
  if (!domain || domain.status === "verified") return;

  const claimed = await store.transition(
    domain.id,
    ["pending", "action_required"],
    "checking",
  );
  if (!claimed && domain.status !== "checking") return;

  const [ownershipAnswers, dmarcAnswers] = await Promise.all([
    dns.txt(domain.hostname),
    dns.txt(`_dmarc.${domain.hostname}`),
  ]);

  const ownsDomain = ownershipAnswers
    .map(normalizeTxt)
    .includes(domain.expectedOwnershipValue);
  const hasDmarc = dmarcAnswers.some((answer) =>
    normalizeTxt(answer).toUpperCase().startsWith("V=DMARC1;"),
  );
  const reason: Evidence["reason"] = !ownsDomain
    ? "ownership_missing"
    : !hasDmarc
      ? "dmarc_missing"
      : "requirements_met";

  await store.saveEvidence({
    domainId: domain.id,
    observedAt: now,
    policyVersion: domain.policyVersion,
    ownershipAnswers,
    dmarcAnswers,
    reason,
  });

  await store.transition(
    domain.id,
    ["checking"],
    reason === "requirements_met" ? "verified" : "action_required",
  );
  await store.scheduleNextCheck(
    domain.id,
    reason === "requirements_met" ? null : new Date(now.getTime() + 15 * 60_000),
  );
}
Enter fullscreen mode Exit fullscreen mode

The policyVersion: 2 literal is deliberate. A decision made under version 2 should remain explainable after version 3 changes the requirements. The 15-minute delay is merely an example operating choice; use an interval that fits the onboarding promise and DNS workload, then add jitter in the scheduler so many domains do not become due on the same boundary.

There is one subtle race in any implementation of this shape. Two workers can read pending before either claims it. The conditional transition must therefore be an atomic compare-and-set in the backing store, and the evidence insert needs its own idempotency key in production. Those are storage guarantees, not comments that a queue can enforce for you.

Poll for recovery, not for urgency

The poller should inspect durable state, not replay assumptions about which webhook probably arrived. Select unverified domains with nextCheckAt <= now, enqueue them through the same dedupe path, and move the next due time only after the worker has stored its observation. A missed event then costs one polling interval, while an event that arrives normally starts the same work sooner.

Suppose a seller adds shop.example at 10:02 and the application creates a pending row under policy version 2. An event at 10:03 enqueues the domain, but the first DNS read finds neither expected record, so the worker stores ownership_missing, changes the visible state to action_required, and schedules 10:18. The seller finishes the DNS edit at 10:07. No second event is required: the due-row scan enqueues the same domain at 10:18, the worker wins the atomic claim, and the new observation either advances the row or records the next precise reason. If a delayed copy of the 10:03 event arrives at 10:19, its dedupe window and the current state prevent it from manufacturing a second completion. This example is intentionally dull. Every timestamp, answer set, reason, and transition is available from durable records; support does not need to infer an event sequence from scattered logs, and the seller receives instructions tied to the latest observation rather than a generic "DNS is still propagating" message.

One queue. One decision.

Backoff belongs in policy too. A newly added domain can be checked relatively often because a customer is waiting on the screen. Later checks can spread out. Put a ceiling on the onboarding window, then keep the domain in action_required with the latest evidence rather than spinning forever. Don't hide the reason: ownership_missing calls for a different instruction than dmarc_missing.

Observability should follow business transitions. Count domains by current state and reason code; measure the age of the oldest due check; and record trigger source, queue time, DNS observation time, and policy version together. Avoid treating webhook receipt count as a success metric. Ten perfectly delivered duplicate events still represent one domain decision.

This is where the revenue-per-hour lens helps. A custom retry engine, a separate webhook state machine, and a polling subsystem with different rules create three places to debug one seller's setup. One verification worker and two thin triggers outsource the undifferentiated timing problem to a queue and scheduler, leaving the product code responsible for evidence and customer action.

When should polling be the primary completion mechanism?

Stick with polling alone when onboarding volume is low, a completion delay of one interval is acceptable, or the upstream event source has no authenticated delivery contract you are prepared to operate. It is also the cleaner choice when every check is already cheap, rate-limited, and driven by your own DNS reader. Fewer transports mean fewer credentials, replay rules, and dashboards.

Webhook-only can be suitable inside a closed system where the producer and consumer share durable replay, event identity, and ownership. That is a demanding boundary. For public marketplace onboarding, it leaves recovery coupled to a delivery path the customer cannot diagnose, so the burden lands on support.

The hybrid recommendation is not suitable when DNS checks are expensive enough that duplicate triggers materially hurt capacity, yet the datastore cannot provide atomic claims or deduplication. Fix that storage boundary first, or accept polling's slower, simpler behavior. Shipping a weekly product means choosing the smallest system whose failure mode the founder can explain from one evidence row.

References

Top comments (0)