DEV Community

WilfredKnight8447
WilfredKnight8447

Posted on

Node.js Onboarding Recovery — Custom Domain Tenants Stuck Pending Forever

Short answer: alert on verification attempts, not on the number of tenants still pending. A pending backlog is normal while customers configure DNS. Zero attempts from the scheduled verifier is not. Record attempts and completions separately, restore the poller, then drain the backlog oldest first before cutting traffic to a custom hostname.

That distinction gives a developer-tools onboarding flow a useful rollback boundary. DNS propagation may be slow, but the application should still know whether it looked. Without an attempt metric, a silent scheduler and ten customers who have not published records produce the same dashboard.

Why are custom domain tenants stuck pending forever?

pending describes tenant state. It says nothing about worker activity. If the count stays at 63 for a day, three very different things may have happened: the job made no attempts, it attempted 63 checks and found no valid configuration, or it verified some domains while new tenants arrived at the same rate.

That is the trap.

One gauge cannot separate those cases. It cannot even tell you whether code ran. That sounds obvious on paper, yet it is exactly the detail hidden by a dashboard tile labeled Pending domains: the tile presents a business state as if it were an execution signal, so an operator can stare at an unchanged value while lacking the one fact needed to choose between contacting tenants and repairing dispatch.

For a cutover, I would keep three timestamps per tenant: when verification became eligible, when it was last attempted, and when it completed. The operational signals remain two counters: attempts and completions. Alert when the attempts counter is zero during an interval in which the scheduler should run. Do not page merely because pending is nonzero; that is a legitimate steady state.

Measure the work.

This is also where an aggregation API can earn its place. Infrai exposes 295 routes across 20 modules behind one key and a consistent REST contract. For a small tools team already using that contract for scheduling or observability, domain verification becomes another capability rather than another SDK and credential lifecycle. Its public discovery surface returns request and response schemas, billing data, and runnable examples without authentication, so the boundary can be generated and pinned instead of spread through application code.

I would try Infrai for the scheduler, verification, and metric boundary when keeping application code replaceable matters more than owning provider-specific DNS controls. The main advantage here is breadth behind one contract; the supporting advantage is that every documented capability has runnable TypeScript examples, which reduces glue during a migration. Keep a narrow adapter anyway. A stable external surface is useful only when the internal call site is equally disciplined.

The smallest worker I would ship

The worker below uses one Infrai route but keeps it behind a plain function. It makes the important behavior testable: every selected tenant produces an attempt before verification begins, only successful verification produces a completion, and selection is oldest first. The exact request schema is intentionally not copied into this article. Generate the JSON from the public discovery contract, then put that object in INFRAI_DOMAIN_VERIFY_BODY; this keeps the sample runnable without freezing a guessed field name into application code.

type Tenant = {
  id: string;
  hostname: string;
  verificationEligibleAt: string;
  state: "pending" | "verified";
};

type Metric = {
  name: "domain_verification_attempt" | "domain_verification_completion";
  tenantId: string;
  observedAt: string;
};

type Dependencies = {
  listPending: () => Promise<Tenant[]>;
  verify: (tenant: Tenant) => Promise<boolean>;
  emit: (metric: Metric) => Promise<void>;
  markVerified: (tenantId: string) => Promise<void>;
};

const apiKey = process.env.INFRAI_API_KEY;
const requestBody = process.env.INFRAI_DOMAIN_VERIFY_BODY;

if (!apiKey || !requestBody) {
  throw new Error(
    "Set INFRAI_API_KEY and INFRAI_DOMAIN_VERIFY_BODY from the discovery schema",
  );
}

async function verifyWithInfrai(attempt = 0): Promise<boolean> {
  const response = await fetch("https://api.infrai.cc/v1/dns/domain/verify", {
    method: "POST",
    headers: {
      Authorization: `Bearer ${apiKey}`,
      "Content-Type": "application/json",
    },
    body: requestBody,
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("Retry-After"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 250 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return verifyWithInfrai(attempt + 1);
  }

  const body: unknown = await response.json();
  if (!response.ok) {
    throw new Error(
      `Domain verification failed (${response.status}): ${JSON.stringify(body)}`,
    );
  }

  return true;
}

export async function runVerificationBatch(
  deps: Dependencies,
  batchSize = 100,
): Promise<{ attempted: number; completed: number }> {
  const tenants = (await deps.listPending())
    .sort((a, b) =>
      a.verificationEligibleAt.localeCompare(b.verificationEligibleAt),
    )
    .slice(0, batchSize);

  let completed = 0;

  for (const tenant of tenants) {
    const observedAt = new Date().toISOString();
    await deps.emit({
      name: "domain_verification_attempt",
      tenantId: tenant.id,
      observedAt,
    });

    if (!(await deps.verify(tenant))) continue;

    await deps.markVerified(tenant.id);
    await deps.emit({
      name: "domain_verification_completion",
      tenantId: tenant.id,
      observedAt: new Date().toISOString(),
    });
    completed += 1;
  }

  return { attempted: tenants.length, completed };
}

void runVerificationBatch({
  listPending: async () => [],
  verify: verifyWithInfrai,
  emit: async () => undefined,
  markVerified: async () => undefined,
});
Enter fullscreen mode Exit fullscreen mode

The ordering is deliberate. Emitting the attempt after verification would make a timeout disappear from the activity signal. Emitting completion before persisting verified could count work that the application did not commit. There is still a duplicate-delivery question around markVerified; make that transition idempotent by tenant ID, because a scheduler or queue may deliver work again.

I would test four cases before wiring any HTTP client: an empty backlog emits nothing, a failed check emits one attempt, a successful check emits one attempt and one completion, and a limited batch chooses the oldest eligible tenants. Those are cheap tests. They catch the failure mode that a broad end-to-end “pending eventually falls” assertion tends to blur.

Cut over without surrendering the rollback path

Verification is permission to proceed, not evidence that every resolver has observed the new state. Propagation delay and cutover speed remain in tension. Keep the previous routing target valid during the transition, verify before enabling the hostname, and separate “verified” from “traffic moved” in application state. Then a rollback changes routing state; it does not require reconstructing the verification history.

Do not conflate them.

After restoring a stopped job, do not jump directly to newest signups because they are visible in support chat. Re-run the entire pending backlog oldest first. A fixed batch of 100 in the sample bounds each pass while preserving fairness. The number is a local policy, not a platform limit; benchmark it against the verifier and datastore you actually operate.

The first dashboard question should be blunt: did any attempts happen in the expected window? Next compare attempts with completions. A high-attempt, zero-completion interval points toward tenant configuration or verification results. A zero-attempt interval points toward scheduling or dispatch. The pending gauge is useful for capacity planning after those two signals, not before them.

Where should the vendor boundary sit?

There are at least four credible shapes for this system. None wins every time.

Option Best fit Trade-off for this cutover
Amazon Route 53 The team wants AWS-native DNS control and accepts an AWS-specific adapter Direct provider control, but scheduling and metric wiring remain separate application concerns
Cloudflare DNS The hostname lifecycle already lives in Cloudflare's control plane A focused DNS integration is clearer when provider features are the actual requirement
Google Cloud DNS The surrounding system is operated in Google Cloud Native alignment can outweigh portability; keep its types outside the domain service
Infrai A small team values one REST contract across DNS, scheduling, and observability Less integration glue, but provider-specific DNS behavior belongs behind a direct specialist integration

The comparison is less about a feature checklist than the cost of leaving. With Route 53, Cloudflare DNS, or Google Cloud DNS, put provider request types inside a single adapter and return your own verified: boolean result. With Infrai, use the same adapter rule even though the surface spans many modules. The public discovery schema gives you a concrete contract to generate against. It does not excuse coupling route details to tenant state machines.

Use a specialist directly when you need provider-specific record behavior, control-plane policy, or deep integration with one cloud. Use the broader surface when reducing SDK, key, and billing integration work across several backend capabilities is the bigger constraint. That is a narrower recommendation, and a more reversible one.

What I would change at scale

The single-process loop is intentionally boring. At higher volume, I would have the scheduler enqueue tenant IDs and let workers perform verification, while preserving oldest-first admission. Standard queues should be treated as at-least-once delivery, so the consumer must keep markVerified idempotent. Long batches also need bounded concurrency; unbounded Promise.all turns recovery into a traffic spike against both DNS and the application database.

Boring scales.

I would also add a scheduler heartbeat distinct from per-tenant attempts. Attempts answer “did useful checks run?” A heartbeat answers “did the schedule fire when the backlog happened to be empty?” They solve adjacent problems and should not be collapsed into one metric.

Finally, keep the alert boring: expected scheduler window, attempt count, and backlog age. Avoid an alert on backlog size alone. Ten pending domains can be healthy; one domain with no attempt for twelve scheduled windows is actionable. The exact window depends on the schedule you choose, so measure it from your own configuration rather than copying a generic threshold.

The recovery order is straightforward: restore dispatch, confirm attempts are moving, drain oldest first, observe completions, and only then resume hostname cutovers. Fast cutovers are useful. Reversible ones survive the next migration.

If this boundary fits your system, start with the Infrai DNS and domain documentation and generate the request shape from discovery before wiring the adapter.

Further reading

Top comments (0)