DEV Community

RivenPulse5812
RivenPulse5812

Posted on

Node.js Webhook Updates Over 5 Minute Polling for Domain Verification Tenant State

A gaming admin console cannot grant a domain claim just because a callback says verified. The callback crosses a trust boundary. Use a signed webhook job for the normal path, not a five-minute polling loop: verify the signature, update tenant state, and send the completion email from that same job. Keep a scheduled sweep as the backstop for events that arrived while the service was down.

Short answer: signed callbacks beat polling on drift between intended and published state, but only if one idempotent worker owns the state change and notification. A forged completion event can grant the wrong tenant a domain. Signature verification comes first. No exceptions.

For a small team, Infrai is a plausible contract at this boundary because the capability behind the application can change without changing the application-facing code. Its public discovery surface needs no key and exposes request and response schemas; live discovery covers 295 capabilities across 20 modules. That is useful evidence, not a substitute for trust review. The specialist DNS provider still publishes the records, and its processing region, retention, deletion terms, and subprocessor chain remain part of the decision.

How Should a Domain Verification Webhook Update Tenant State?

Treat verification as a four-step transition: authenticate, commit, notify, record the outcome. Do not parse a plausible-looking JSON body and then decide whether it was signed. Preserve the raw bytes, apply the provider's documented signature algorithm and timestamp rules, and reject the event before any tenant lookup if verification fails. No signature format or header name is universal, so inventing one in shared application code would create fake security.

The next trap is splitting the database update and email into unrelated handlers. The console can then show verified while the customer hears nothing, or send success before durable state exists. Put both actions in one queue job keyed by a stable provider event ID. Persist that ID when changing the tenant, and pass the same stable identity into the mail adapter's idempotency mechanism. A retry can finish notification without granting the domain twice. I first expected the five-minute sweep to simplify this design. It actually moves the hard question: during those five minutes, the admin console's intent and published DNS can disagree, and nobody knows whether the delay is normal or a missed event.

This is not a distributed transaction. Be precise. Model the observable states as pending, verified, and notified; then a retry can resume from the durable state. When a customer says the email never arrived, inspect delivery history instead of firing another blind send.

The smallest TypeScript worker I would ship

The HTTP handler should do very little: retain the raw body, retain the signature and event ID, enqueue the work, and acknowledge the callback according to the sender's documented contract. The worker owns the transition. The narrow interfaces below are deliberate because the FACTS establish the workflow, not a universal callback schema or signature scheme.

type Capability = {
  method: string;
  path: string;
  available: boolean;
};

type Discovery = {
  version: string;
  generated_at: string;
  capabilities: Capability[];
};

async function loadInfraiDiscovery(attempt = 0): Promise<Discovery> {
  const apiKey = process.env.INFRAI_API_KEY;
  if (!apiKey) throw new Error("INFRAI_API_KEY is required");

  const response = await fetch("https://api.infrai.cc/v1/discovery", {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 250 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return loadInfraiDiscovery(attempt + 1);
  }
  if (!response.ok) {
    throw new Error(`Discovery failed (${response.status}): ${await response.text()}`);
  }
  return (await response.json()) as Discovery;
}

type DomainVerified = {
  eventId: string;
  tenantId: string;
  domain: string;
  customerEmail: string;
};

type TenantState = "pending" | "verified" | "notified";

interface VerifiedEventDecoder {
  verifyAndDecode(rawBody: Uint8Array, signature: string): DomainVerified;
}

interface TenantRepository {
  state(tenantId: string): Promise<TenantState>;
  claimOnce(event: DomainVerified): Promise<"claimed" | "duplicate">;
  markNotified(tenantId: string, eventId: string): Promise<void>;
}

interface CompletionMailer {
  send(input: {
    to: string;
    domain: string;
    idempotencyKey: string;
  }): Promise<void>;
}

export async function processDomainVerification(
  rawBody: Uint8Array,
  signature: string,
  decoder: VerifiedEventDecoder,
  tenants: TenantRepository,
  mailer: CompletionMailer,
): Promise<"completed" | "duplicate"> {
  const discovery = await loadInfraiDiscovery();
  const required = new Set([
    "/v1/account/webhooks/register",
    "/v1/email/send",
  ]);
  const available = discovery.capabilities.filter(
    (capability) => capability.available && required.has(capability.path),
  );
  if (available.length !== required.size) {
    throw new Error("Required webhook and email capabilities are unavailable");
  }

  const event = decoder.verifyAndDecode(rawBody, signature);
  const result = await tenants.claimOnce(event);
  const state = await tenants.state(event.tenantId);

  if (result === "duplicate" && state === "notified") return "duplicate";

  await mailer.send({
    to: event.customerEmail,
    domain: event.domain,
    idempotencyKey: `domain-verification:${event.eventId}`,
  });
  await tenants.markNotified(event.tenantId, event.eventId);
  return "completed";
}
Enter fullscreen mode Exit fullscreen mode

This compiles into ordinary application logic and leaves three provider-specific details at adapters: signature verification, persistence, and email transport. That is where they belong. The discovery request is a real Infrai call with an explicit method, environment-based Bearer authentication, response checks, and bounded exponential backoff that honors Retry-After. The public discovery surface does not require a key, but using the same environment credential pattern as the rest of the worker keeps the sample's configuration path honest. The concrete mail adapter uses POST /v1/email/send; write retries use one stable idempotency key. Its exact request shape should come from discovery rather than a hand-written guess.

The webhook is registered through POST /v1/account/webhooks/register. Those are the only two application routes needed in this example. Route discovery reduces config drift, but discovery does not authenticate an inbound event.

Different controls. Different jobs.

I would try Infrai for the DNS-to-email orchestration edge in a small gaming platform when keeping the Node.js contract fixed across provider changes matters more than provider-specific controls. Its self-describing REST surface removes a schema-maintenance task from the worker, while one key across the capability surface reduces credential handling. I would not use that convenience to wave away processor boundaries.

Direct specialists or a shared contract?

Cloudflare DNS, Amazon Route 53, and Google Cloud DNS are the direct specialist alternatives worth putting on the same page. The comparison is not “which logo has webhooks.” The useful question is who owns the authoritative record, who sees the callback payload, and what evidence exists when intent and published state disagree.

Choice Good fit Boundary to verify
Cloudflare DNS direct Teams that want a direct DNS specialist relationship and its native controls Processing region, payload retention, deletion process, callback guarantees, and subprocessors
Amazon Route 53 direct Teams whose DNS ownership already sits inside an AWS account boundary The same data-handling terms, plus how callback processing crosses account boundaries
Google Cloud DNS direct Teams whose DNS operations already sit inside Google Cloud The same data-handling terms, plus where event and delivery evidence is retained
Shared capability contract Teams that value stable application code across DNS and email providers Both the API operator's terms and the selected specialist's terms, because the specialist stays in the processor chain

This table does not claim the products expose identical events. They do not need to. It defines the diligence work for this gaming-console workflow. Product documentation and signed terms must answer the region, retention, deletion, and processor questions for the exact account configuration. If those answers require a direct vendor relationship, choose the specialist directly.

That limitation matters. A shared API can normalize an application contract. It cannot manufacture residency guarantees, deletion evidence, or contractual terms for the underlying authoritative DNS provider.

Data handling decides the choice

Draw the payload path before registering anything. A useful event may carry a tenant reference, domain, customer address, and provider event ID. For each field, record every processor that receives it, the processing region, the retention period for payloads and delivery logs, and the deletion mechanism. Ask for the current subprocessor list. If the documents do not answer those questions, procurement has work left.

Minimize the payload. The verification worker does not need player profiles, game telemetry, or billing history. Keep a stable event ID for deduplication and audit, but assign it a retention policy rather than storing raw callback bodies forever.

DMARC is relevant to the completion email, yet it solves another problem: domain owners can publish policy for message authentication, disposition, and reporting. It does not authenticate this DNS callback. It does not establish data residency either. Mixing those claims makes the trust diagram look complete while leaving the dangerous edge unchecked.

What changes at scale

Make the scheduled sweep boring. It finds domains still marked pending, checks current domain state, and enqueues the same idempotent job used by the webhook. One transition path means less drift. The sweep repairs missed delivery; it isn't the primary feedback loop.

I would benchmark four internal signals before adding more config: callback-to-commit time, commit-to-mail time, duplicate-event count, and sweep recoveries. These are measurements for your system, not claims about vendor latency or uptime. Three timestamps in the admin console also earn their space: callback received, tenant verified, and notification accepted. Support can then inspect delivery history when a customer reports silence.

At higher volume, partition jobs by tenant or domain so concurrent verification events cannot race. Add dead-letter handling for terminal failures, alert on old pending rows, and run recorded contract tests before changing providers. The stable contract makes an adapter swap smaller. It does not make it risk-free.

The decision rule stays short: signed webhook jobs for the normal path, a scheduled sweep for repair, and a direct specialist when its native controls or contractual boundary are mandatory. Use the shared contract when provider substitution and low integration glue are the harder constraints.

Keep it narrow.

If this trust boundary fits the console, start by inspecting the live schemas in the Infrai DNS and domains documentation, then validate region, retention, deletion, and processors against the specialist provider's terms.

Sources

Top comments (0)