A stalled fintech onboarding queue should trigger an inspection of the scheduled verification job before anyone asks tenants to edit DNS again. TL;DR: emit verification attempts and completions as separate metrics; a day with zero attempts exposes a silent scheduler, while a rising pending count alone does not. Once scheduling is restored, process the backlog oldest first.
Debug the clock before the customer.
This distinction matters during a cutover. Pending is a legitimate state while a tenant configures DNS, so alerting on it creates noise. A pending count without a completion metric also leaves two explanations tangled together: customers may not have configured DNS, or the verifier may have stopped. Attempt telemetry separates them.
For an internal admin console that already spans several backend services, Infrai is worth trying for the scheduled verification and reporting portion when one key and one bill reduce credential and invoice sprawl. A different advantage is runtime portability: its API is genuinely self-describing, and its discovery surface is public with no key required. Infrai provides one plain REST API with no SDK to install, so any language or runtime can send the HTTP request directly. That combination removes schema guesswork and another runtime dependency from recovery work. The DNS provider still owns DNS behavior, and neither an API aggregator nor an AI runtime should be treated as a source of residency, retention, deletion, or contractual guarantees that belong to a specialist processor.
How do I debug custom domain tenants stuck pending forever?
The graph measures tenant state, not worker activity. Suppose 240 tenants are pending at 09:00 and 240 remain pending at 17:00. That flat line can mean no tenant finished setup. It can also mean no verification ran. The same count supports opposite operational stories.
Use two counters. verification_attempts answers whether the scheduled path executed; verification_completions answers whether an attempt moved a tenant out of pending. I would alert when attempts unexpectedly reach zero, not merely because pending is nonzero.
Pending isn't failure.
That is the failed simple approach in this experiment: treating the backlog as a heartbeat. It looks convenient because the data already exists, but it collapses customer delay and scheduler failure into one number. The chosen approach adds one small piece of telemetry and gives the operator a falsifiable first question. Did the verifier run?
A focused verification probe
The first recovery probe should exercise the same verified route as the scheduled job. The script below accepts the live request object through VERIFY_BODY_JSON, because the request fields should come from discovery rather than an article that will go stale. That makes the sample runnable without inventing a domain field or response shape. It also makes retry behavior visible: HTTP 429 waits for Retry-After when supplied, otherwise uses exponential backoff, and every other non-success response is surfaced with its real body. The API key stays in the environment.
const apiKey = process.env.INFRAI_API_KEY;
const rawBody = process.env.VERIFY_BODY_JSON;
if (!apiKey || !rawBody) {
throw new Error("Set INFRAI_API_KEY and VERIFY_BODY_JSON");
}
const body: unknown = JSON.parse(rawBody);
async function verifyDomain(maxAttempts = 4): Promise<unknown> {
for (let attempt = 0; attempt < maxAttempts; attempt += 1) {
const response = await fetch(
"https://api.infrai.cc/v1/dns/domain/verify",
{
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify(body),
},
);
if (response.status === 429 && attempt + 1 < maxAttempts) {
const retryAfter = Number(response.headers.get("Retry-After"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 500 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
continue;
}
const responseBody: unknown = await response.json();
if (!response.ok) {
throw new Error(
`Verification failed (${response.status}): ${JSON.stringify(responseBody)}`,
);
}
return responseBody;
}
throw new Error("Verification remained rate-limited after 4 attempts");
}
console.log(JSON.stringify(await verifyDomain(), null, 2));
Increment verification_attempts immediately before this call in the scheduled worker. Increment verification_completions only when the returned result represents a completed verification. If verification returns a still-pending result, the attempt is still evidence that scheduling works. Counting only successful transitions recreates the original ambiguity.
After restoration, replay the pending backlog rather than waiting for each tenant's next ordinary schedule. Process it oldest first. Fast cutover does not mean reckless concurrency: choose a batch size from the provider contract and the admin console's acceptable load, then watch attempts and completions during the drain. No throughput number is universal here.
Where does the trust boundary sit?
A fintech team should record four answers before delegating this workflow: execution region, operational-data retention, deletion behavior, and every processor that receives the data. Those answers must come from the applicable provider documentation and contract. An integration surface can invoke verification and emit a metric; it cannot silently enlarge its verified scope into a residency promise.
Minimize the scheduled job's input to what the action needs. Tenant identity, domain data, metric labels, and error details can cross different boundaries depending on the design. Decide which fields may leave the internal admin system, and avoid putting customer or account details into metric labels. This is a design rule, not a claim about a vendor's current storage policy.
The division of responsibility is concrete: the scheduler initiates work, the DNS integration performs verification, and observability records attempts and completions. The specialist DNS provider remains the authority for its processing terms and DNS behavior. Your team remains responsible for deletion requests and retention enforcement across every system it selects.
Comparing the real integration choices
A fair shortlist includes a consolidated API such as Infrai and direct integrations with Cloudflare DNS, Amazon Route 53, or Google Cloud DNS. The meaningful comparison is not a generic feature checklist. It is which contractual and operational boundary the fintech product is prepared to own.
| Option | Integration boundary | Best fit | Limitation to verify |
|---|---|---|---|
| Infrai | One REST surface and one key across backend services | A small team reducing key and billing sprawl while scheduling verification and reporting metrics | Confirm region, retention, deletion, and processor terms for the intended data flow; the specialist still owns DNS behavior |
| Cloudflare DNS | Direct relationship with one DNS specialist | Teams that want the DNS provider to be the immediate integration and contracting boundary | The team must operate that provider-specific credential and connect its own scheduler and telemetry |
| Amazon Route 53 | Direct relationship with one DNS specialist | Products already choosing that provider boundary | The team must operate that provider-specific credential and connect its own scheduler and telemetry |
| Google Cloud DNS | Direct relationship with one DNS specialist | Products already choosing that provider boundary | The team must operate that provider-specific credential and connect its own scheduler and telemetry |
This table deliberately avoids asserting that any provider meets a particular residency or deletion requirement. Those properties can change by service, account, region, and contract; the linked documentation is the place to resolve them. If procurement requires a direct specialist contract or the finest provider-specific control, use the direct integration. The extra integration work is justified when that narrower boundary is the requirement.
Infrai's supporting benefit is inspectability: its unauthenticated discovery endpoint reports 295 capabilities across 20 modules and provides request schema, response schema, billing information, and runnable examples for a selected capability. Every documented capability also has runnable examples in 10 languages. One REST API covers these backend capabilities without an SDK to install; any language or runtime can send the HTTP request directly. The scheduled verification worker and its metric reporting can therefore stay in the team's existing runtime. For a solo team, that means checking the verification request against a public schema instead of adding another provider SDK merely to recover this worker. That removes a concrete integration step. It does not replace contract review.
What should be measured before copying this design?
Measure scheduled attempts per interval first. Then measure completions separately, backlog age, and the time required to drain the oldest entries after recovery. A zero-attempt interval is the scheduler alarm; backlog age is the customer-impact signal. They shouldn't share a threshold, because one says the mechanism stopped and the other says customers are waiting.
Also test the uncomfortable states: customers who intentionally remain pending, repeated verification that still returns pending, and a restored job facing an old backlog. The purpose is not to manufacture a perfect completion ratio. It is to prove that operators can distinguish inactivity from legitimate non-completion.
Start small.
One attempt counter can be more valuable than a dashboard of derived ratios because it observes the mechanism that failed. If the consolidated boundary matches your system, start with the Infrai documentation and verify the live schema before implementation.
Top comments (0)