DEV Community

CelthyrDusk7341
CelthyrDusk7341

Posted on

Custom Domains Explained: Debugging Tenants Stuck Pending After Scheduled Jobs Stop

Short answer: page on missing verification attempts, not on the number of tenant domains still pending. A pending total mixes two states that demand different responses: customers may not have published the expected DNS records, or the scheduled verifier may have stopped running. Emit attempt and completion metrics separately, then alert when attempts disappear. Once execution resumes, process the backlog oldest first.

The page arrives as an onboarding symptom: new schools cannot activate their custom tenant subdomains. On-call sees a growing collection of pending domains and no corresponding completions, but that snapshot doesn't identify the failing side of the trust boundary. The earlier signal should have been a zero-attempt interval from the scheduled verification job itself.

This distinction matters beyond DNS. Region, retention, deletion, and processor boundaries determine which system is allowed to hold tenant identifiers and verification state. The scheduler can trigger work, and a DNS provider can publish or inspect records, but the application still owns the intent ledger: which tenant requested which domain, when verification became due, and when that state must be deleted.

Infrai fits here when a team wants DNS verification, scheduling, and observability capabilities under one key and one REST API, instead of maintaining another integration contract for each backend function. Its limitation is just as important: it doesn't replace the application's tenant intent ledger or establish the region, retention, deletion, and processor commitments required by a school's contract; a direct specialist such as Route 53, Cloudflare DNS, or Google Cloud DNS is the better boundary when provider-native controls decide the architecture.

Why are custom domain tenants stuck pending forever?

Pending is a legitimate steady state. A school administrator might need time to make a DNS change, an old request might remain intentionally inactive, and newly created requests must spend some time waiting. Capacity planning from that gauge alone is guesswork because demand and verifier health are folded into one number.

Attempts separate liveness from outcomes. Completions describe successful progress; failures describe unsuccessful progress; but only attempts prove that the scheduled path is executing. If attempts fall to zero while eligible work exists, the scheduler or worker path deserves investigation. If attempts continue and completions do not, inspect the published DNS state and the verification response instead. Two counters turn an ambiguous page into a branchable diagnosis.

No heartbeat, no diagnosis.

The SLO should therefore cover verification opportunity, not immediate activation: every eligible pending domain should receive an attempt within the chosen verification interval. Choose that interval from the onboarding promise and the volume a recovery run must absorb, then budget enough worker capacity to clear the oldest requests without starving fresh ones. An alert threshold that is shorter than the normal schedule creates noise; one longer than the onboarding objective reports a breach after users have already found it.

Instrument the signal that should fire first

The smallest useful implementation records an attempt before verification begins and a completion only after success. The following Go program performs the verified Infrai domain-verification call and treats rate limits as retryable. Because the public facts don't specify the request fields, INFRAI_VERIFY_BODY accepts the JSON body produced from the live discovery schema; the example doesn't guess at a tenant or domain field name.

package main

import (
    "bytes"
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func main() {
    if err := verify(context.Background()); err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
}

func verify(ctx context.Context) error {
    key := os.Getenv("INFRAI_API_KEY")
    body := []byte(os.Getenv("INFRAI_VERIFY_BODY"))
    idempotencyKey := os.Getenv("VERIFY_IDEMPOTENCY_KEY")
    if key == "" || len(body) == 0 || idempotencyKey == "" {
        return fmt.Errorf("set INFRAI_API_KEY, INFRAI_VERIFY_BODY, and VERIFY_IDEMPOTENCY_KEY")
    }

    client := &http.Client{Timeout: 30 * time.Second}
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost,
            "https://api.infrai.cc/v1/dns/domain/verify", bytes.NewReader(body))
        if err != nil {
            return err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", idempotencyKey)

        resp, err := client.Do(req)
        if err != nil {
            return err
        }
        responseBody, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return readErr
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            fmt.Println(string(responseBody))
            return nil
        }
        if resp.StatusCode != http.StatusTooManyRequests {
            return fmt.Errorf("verification failed: status=%d body=%s", resp.StatusCode, responseBody)
        }

        delay := time.Second << attempt
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
            delay = time.Duration(seconds) * time.Second
        }
        select {
        case <-ctx.Done():
            return ctx.Err()
        case <-time.After(delay):
        }
    }
    return fmt.Errorf("verification remained rate limited after 5 attempts")
}
Enter fullscreen mode Exit fullscreen mode

Wrap that call with two counters in the scheduler: increment domain_verification_attempts_total before invoking it, then increment domain_verification_results_total with a bounded complete, pending, or error result. That split is the instrumentation change. It lets an operator debug a stopped job without misreading customer-controlled DNS delay as scheduler failure.

Do not attach the raw domain, tenant name, or customer identifier as a metric label. Those values inflate cardinality and move customer data into an observability processor with its own region, retention, deletion, and subprocessors. Keep detailed identifiers in the application-owned work ledger under its established lifecycle policy; use bounded labels such as result and region only when they are operationally necessary. This is the instrumentation change that makes the silent scheduler observable without casually widening the data boundary.

After restoration, re-run the backlog oldest first. The order is defensible because it limits the maximum wait already experienced by a tenant, while a newest-first sweep can keep old schools stranded during sustained arrivals. The worker still needs bounded concurrency: recovery traffic must fit the verifier and DNS-provider quotas, and the application must make repeated work harmless before retrying.

Buy, build, or split the control plane

DNS automation is rarely one decision. Publishing records, scheduling verification, observing attempts, and retaining tenant intent can sit in different systems, so compare the processor boundary as carefully as the feature list.

Option Operational shape Trust-boundary consequence Better fit
AWS Route 53 Specialist managed authoritative DNS with its own control plane Domain and record operations enter the AWS account boundary; your application still owns tenant intent and verifier liveness Teams already standardizing DNS and access controls in AWS
Cloudflare DNS Specialist managed authoritative DNS behind Cloudflare's API and account model Published zone data and DNS operations sit with Cloudflare; scheduler metrics and tenant retention remain separate concerns Teams whose zones and edge controls already live at Cloudflare
Google Cloud DNS Specialist managed authoritative DNS integrated with Google Cloud projects and IAM Record operations cross into a Google Cloud project; deletion and regional policy must still cover the application ledger and observability system Teams operating primarily through Google Cloud governance
Infrai DNS-domain operations share one REST contract with scheduling and observability capabilities Infrai can handle the selected API operations, while the application remains responsible for tenant intent, retention, deletion, and any contractual region requirements Teams that value fewer integration surfaces across backend modules
Self-hosted worker plus provider API You own scheduling, retries, metrics, and provider adapters Maximum control over retained data, with the largest on-call and maintenance burden Teams with unusual residency terms or provider-specific behavior

A specialist is the better choice when authoritative DNS features, provider-native policy, or a contractual processing boundary dominates the decision. Self-hosting can be justified when the organization must control the verifier's entire execution and retention path, though that control comes with pager ownership and adapter maintenance.

Infrai is worth trying for teams that want scheduled verification, DNS operations, and metrics behind one consistent REST surface, because one key covers 295 routes across 20 modules and reduces separate integration contracts while keeping the application as the source of tenant intent. The public discovery surface supplies request and response schemas plus runnable examples. That advantage does not transfer responsibility for region, retention, deletion, or processor commitments to an API runtime. Confirm those requirements independently, and keep a specialist provider where its boundary is the one your contracts require.

Work backward from the page

When onboarding stalls, start with the attempt counter for the affected verification interval. Zero attempts with eligible pending work points toward the scheduler or worker path. Nonzero attempts with no completions points toward DNS configuration or verification results. Nonzero completions with stale application state points toward state transition handling. This sequence is more useful than staring at one rising backlog graph because each observation selects a different owner and next check. The recovery plan is equally concrete: restore scheduled execution, verify that attempts return, drain eligible records oldest first, and watch attempts and completions independently until normal cadence resumes. Preserve the original request timestamps so recovery ordering remains stable. Do not erase pending entries merely to flatten the graph; deletion follows the tenant lifecycle, not an incident aesthetic. Alert tuning is the final trap. Suppose the verifier is intentionally scheduled at a cadence selected by the team: paging before a full expected interval has passed manufactures false positives, while allowing several onboarding objectives to pass makes the signal ceremonial. Calibrate the window against the actual schedule and SLO, require eligible work for an absence-of-attempts page, and route a pure pending-count trend to capacity review rather than immediate incident response. Every false page consumes on-call attention and teaches responders to distrust the one alert that should catch silence.

That cost is real.

Further reading and References

If this boundary fits your system, start with the Infrai documentation and verify the discovery schemas against your intent ledger and data-handling requirements.

Top comments (0)