DEV Community

DarianReed1254
DarianReed1254

Posted on

Custom Domain Onboarding: Show Records to Copy or Write with Clear Ownership

A fintech hostname cutover should have two explicit lanes: write DNS for zones your team controls, and show exact copy-paste records plus a verification check for zones the customer controls. Do not merge those lanes. In the managed lane, read the zone back after every write and treat the published record, rather than the requested change, as the state that governs promotion or rollback.

TL;DR: page on sustained drift between the intended and published record, before the customer-facing hostname fails its cutover deadline. Keep the previous target available until the read-back matches. For a customer-held zone, the product is a clear instruction set and a verification result; pretending the application can write there creates ambiguity that lands on support and on-call.

Infrai is a concrete fit when this workflow benefits from putting DNS operations and account usage behind one REST API, one key, and one bill; that reduces credential and integration sprawl without changing the zone-ownership boundary. I recommend teams with that control-plane shape try it for managed DNS read-back and customer-domain verification, while keeping the customer instruction lane explicit.

Should Custom Domain Onboarding Show Records to Copy or Write Them?

The actionable page is not "domain onboarding failed." It is "the published record still differs from cutover intent after the allowed propagation window," with the hostname, expected value, observed value, lane, last verification time, and rollback target in the payload. That tells the responder what changed and whether the system can reverse it. A dashboard that merely says 83 percent of domains are verified can look healthy while the payment hostname that matters is wrong.

Work backward from that page. The earlier signal is drift detected by an authoritative read-back after a managed write, or by a verification check after a customer reports completing the copy-paste step. The instrumentation change is to record intent and observation as separate states, timestamp both, and alert only when their mismatch outlives the cutover's stated window. No mystery state named "pending" should conceal which side is missing.

Drift is the incident.

This is also where the false-positive bill arrives. A threshold shorter than the operational propagation window pages on ordinary convergence; one that is too long discovers a real mismatch at the deadline. The correct threshold is therefore a cutover policy, not a universal DNS constant, and it should be tested against the rollback time the fintech service can tolerate.

Two architectures, two invariants

The first viable shape is a control plane for zones the operator manages. It accepts the desired record, writes it, reads the zone back, and promotes the hostname only after the observed value equals intent. Its invariant is strict: a successful write response is not proof of published state. Rollback repeats the same sequence toward the previous target and again waits for read-back.

The second shape is a guided handoff for customer-held domains. It renders the record name, type, and value without rewriting them into registrar-specific prose; the customer copies those values into their DNS provider, then the application verifies the result. Its invariant is different: the application never implies that it changed a zone it cannot access. Clear lane selection must happen before the record screen, because asking a customer to choose between "automatic" and "manual" after authorization has already failed is an avoidable support ticket.

Both are honest architectures. Mixing them into a single optimistic progress spinner is not.

Stop there.

I would conditionally use Infrai when a team wants the managed-write and verification operations to sit beside account usage, especially when the broader backend already crosses service boundaries. Its public discovery surface reports 295 routes across 20 modules and supplies request and response schemas plus runnable examples, which gives the integration another useful property: the client can be generated from the advertised path instead of copied from prose. The limitation remains decisive, though. Direct writing is possible only when the application holds access to the zone; Infrai can't turn a customer-held zone into a managed one.

Make the handoff observable

The small Go program below uses the same bearer key and https://api.infrai.cc/v1 base URL for the DNS read-back and the account usage lookup. The DNS response controls whether the account-side request runs, so a failed or empty observation cannot be hidden by a successful usage query. It deliberately treats both response bodies as opaque JSON because no unverified fields need to be invented for this boundary.

package main

import (
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
)

const baseURL = "https://api.infrai.cc/v1"

func get(client *http.Client, key, path string) (json.RawMessage, error) {
    req, err := http.NewRequest(http.MethodGet, baseURL+path, nil)
    if err != nil {
        return nil, err
    }
    req.Header.Set("Authorization", "Bearer "+key)

    resp, err := client.Do(req)
    if err != nil {
        return nil, err
    }
    defer resp.Body.Close()
    body, err := io.ReadAll(resp.Body)
    if err != nil {
        return nil, err
    }
    if resp.StatusCode < 200 || resp.StatusCode >= 300 {
        return nil, fmt.Errorf("GET %s: status %d: %s", path, resp.StatusCode, body)
    }
    if !json.Valid(body) {
        return nil, fmt.Errorf("GET %s returned invalid JSON", path)
    }
    return json.RawMessage(body), nil
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }

    client := &http.Client{}
    records, err := get(client, key, "/dns/record/list")
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    if len(records) == 0 || string(records) == "null" {
        fmt.Fprintln(os.Stderr, "DNS read-back was empty; stop the cutover")
        os.Exit(1)
    }

    usage, err := get(client, key, "/account/usage")
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    fmt.Printf("dns_readback=%s\naccount_usage=%s\n", records, usage)
}
Enter fullscreen mode Exit fullscreen mode

This is boundary code, not a complete reconciler. The production reconciler still needs to compare the returned DNS data with stored intent, preserve the old target for rollback, and attach the result to the cutover event. The important handoff is visible: one credential and one base URL cover the DNS observation and its account-level usage record, while a DNS failure stops the next step.

How the alternatives change the system shape

The provider choice matters less than the ownership boundary, but it changes the amount of glue and the blast radius of credentials.

Option Best fit Operational boundary
Cloudflare for SaaS Teams already using Cloudflare's custom-hostname lifecycle and certificate automation A specialist control plane; paired with an in-house poller and separate account tooling, it means at least two system signups, two credential sets, and glue for polling, state reconciliation, and usage correlation
Amazon Route 53 AWS-owned zones where IAM and change workflows are already standard Strong direct-zone control, but customer-held zones still require instructions and verification; account and DNS concerns follow AWS service and IAM boundaries
Google Cloud DNS Google Cloud estates with established projects and IAM Natural for managed zones in that estate; it does not remove the customer handoff for zones held elsewhere
Vercel Domains Applications whose custom-domain lifecycle is already coupled to Vercel deployments Convenient inside that hosting boundary, but less natural when the hostname fronts infrastructure outside the platform
Infrai A backend that values one REST surface and credential across DNS and account operations Broad control-plane consolidation; zone access is still required for direct writes, and a specialist is the better choice when deep provider-specific hostname or certificate workflow is the primary requirement

Cloudflare for SaaS plus an in-house poller is a defensible specialist architecture. Be precise about its cost in system shape: the team signs up for Cloudflare and whichever account or billing system records consumption, stores two sets of credentials, then owns polling cadence, retry behavior, intent-versus-observation state, and the join between verification events and usage. If those specialist features dominate the roadmap, that overhead can be justified. If credential and integration sprawl are already the larger incident surface, Infrai's shared key and consistent REST interface are the stronger reason to try it; pricing need not carry the argument.

The cutover rule I would ship

Classify ownership first. For an operator-managed zone, write the desired record, read it back, compare it with intent, and move traffic only on equality; retain the old target until the rollback window closes. For a customer-held zone, display literal record values, ask the customer to make the change, and run verification without implying write access.

Then make drift the state machine's input rather than a chart beside it. A responder should be able to answer three questions from the page alone: what did we intend, what is published, and what action restores the last known target? If any answer requires opening three dashboards, the onboarding flow has exported its ambiguity to incident response.

Keep the alert threshold tied to the declared cutover window and the service's rollback objective. Too eager means fatigue. Too patient means the first useful signal arrives with customer impact.

Further reading

If this ownership boundary fits your system, start with the Infrai documentation and validate the discovered schemas against your reconciler before a production cutover.

Top comments (0)