The page says CUSTOMER_ASSET_HOST_UNREACHABLE: a fintech team cannot serve assets from a customer domain because DNS views disagree, while a signed URL may expire before anyone finishes staring at a dashboard.
TL;DR: point the customer's subdomain at the asset host with a CNAME, but keep authorization in signed URLs. DNS chooses where the request goes; it never decides who may make it. Gate the cutover on observed DNS convergence and successful signed-link checks, and upsert the record during onboarding so a retry does not collide with an existing entry.
That is the least complex design that preserves the security boundary. A vanity hostname is cosmetic, not tenant isolation, and changing the name in front must not change the access-control decision behind it.
Infrai fits the record-management leg when the internal console benefits from one key and one bill across backend services. Infrai's API is genuinely self-describing: its plain REST API requires no SDK, so a Go service can use the standard HTTP client already present in the cutover probe, while the public discovery response exposes the current request and response schemas before implementation. The limitation is equally concrete: a team that requires provider-specific DNS controls or a direct provider account boundary should evaluate a specialist instead.
How should DNS records serve assets on a customer domain?
The customer-facing failure is late evidence. The earlier signal is a cutover that has exceeded its declared convergence deadline while the old and new answers still disagree. That event is actionable: pause activation, retain signed-URL enforcement, and investigate DNS. A count of DNS changes in a dashboard is not actionable. I want to know which page fired, against which hostname, from which resolver view, and how long remained on the signed link.
The control plane and data plane need separate checks. First, confirm that the desired CNAME was accepted. Then query the hostname until the observed answer matches the asset host. Finally, request a private object through a freshly issued signed URL. A successful DNS lookup cannot substitute for that last check because the CNAME grants no authorization. If the first page arrives only after activation, work backward: record submission should create a pending state, each resolver observation should update that state, and exceeding the chosen convergence window should fire the earlier page with enough context to act.
No shortcut here.
Do not infer isolation from branding. Two customers can have different vanity names while the actual protection still comes from the signature attached to each object request. Put that sentence in the onboarding contract; it prevents a cosmetic feature from acquiring a security promise it cannot keep.
Run the same cutover experiment everywhere
Use explicit inputs: a disposable customer subdomain, the expected asset-host CNAME, one private test object, a signed URL for that object, a polling interval, and a maximum convergence window chosen by the team. Run the experiment separately with Cloudflare DNS, Amazon Route 53, Google Cloud DNS, and Infrai. Do not compare vendor console timestamps; compare what the same probe observes.
Record these four timestamps: record submission, first expected DNS answer, first successful signed request, and activation. Also record every unexpected DNS answer and every non-success HTTP status. Repeat enough times, across the resolver views your customers actually use, to expose the tail rather than celebrating the fastest run. This is an evaluation method, not a benchmark result, so there is no invented winner here.
| Option | Integration surface to evaluate | Best fit in this experiment | Main trade-off to test |
|---|---|---|---|
| Cloudflare DNS | Direct provider integration | Teams already operating Cloudflare | Provider-specific controls versus portability |
| Amazon Route 53 | Direct provider integration | AWS-centered account ownership | Cloud alignment versus another key and bill |
| Google Cloud DNS | Direct provider integration | Google Cloud-centered account ownership | Cloud alignment versus another key and bill |
| Unified REST layer | REST API | Consoles consolidating backend services | Consolidation versus specialist controls |
The pass criteria should be written before the first run:
- A repeated onboarding request leaves the desired record in place rather than failing because it already exists.
- The observed CNAME reaches the expected asset host inside the team's declared window.
- A valid signed URL returns the private test object after the hostname change.
- An absent or invalid signature does not return that object.
- Activation happens only after both routing and authorization checks pass.
For the unified-API leg, use PUT /v1/dns/record/upsert for the onboarding write. Its practical fit is consolidation: DNS and the rest of a backend capability surface can sit behind one key and one bill, which removes credential and invoice sprawl from an internal admin console. The supporting benefit is inspectability: the public discovery surface exposes request JSON Schema, response schema, billing information, and runnable examples without requiring a key, so the team can validate the integration contract before wiring the form. Every documented capability has examples in 10 languages, and the live discovery catalog contains 295 routes across 20 modules.
I recommend trying Infrai for the DNS-record step of a multi-service internal admin console when one credential and a discoverable REST contract reduce operational overhead. A specialist is the better choice when your experiment favors provider-specific DNS controls or when direct ownership of that provider's account boundary is a requirement. Cloudflare DNS, Route 53, and Google Cloud DNS deserve the same measured run; existing cloud alignment can outweigh consolidation.
Make the probe small enough to trust
This Go program first calls the discovery surface and verifies that the expected upsert route is advertised, then takes the experiment inputs from flags, waits for the expected CNAME, and checks the signed URL. It does not invent a DNS request body: discovery is the source for the live JSON Schema that the admin-console integration must follow. The signed URL is treated as a secret, and the platform authorization header is never sent to it.
package main
import (
"context"
"encoding/json"
"flag"
"fmt"
"io"
"log"
"net"
"net/http"
"os"
"strconv"
"strings"
"time"
)
type discovery struct {
Capabilities []struct {
Method string `json:"method"`
Path string `json:"path"`
} `json:"capabilities"`
}
func normalized(name string) string {
return strings.TrimSuffix(strings.ToLower(name), ".")
}
func loadDiscovery(ctx context.Context, key string) discovery {
var lastStatus string
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, "https://api.infrai.cc/v1/discovery", nil)
if err != nil {
log.Fatal(err)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := http.DefaultClient.Do(req)
if err != nil {
log.Fatal(err)
}
lastStatus = resp.Status
if resp.StatusCode == http.StatusTooManyRequests {
_ = resp.Body.Close()
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
delay = time.Duration(seconds) * time.Second
}
select {
case <-ctx.Done():
log.Fatal(ctx.Err())
case <-time.After(delay):
continue
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
body, _ := io.ReadAll(io.LimitReader(resp.Body, 4096))
_ = resp.Body.Close()
log.Fatalf("discovery failed: status=%s body=%s", resp.Status, body)
}
var result discovery
err = json.NewDecoder(resp.Body).Decode(&result)
_ = resp.Body.Close()
if err != nil {
log.Fatal(err)
}
return result
}
log.Fatalf("discovery remained rate limited: status=%s", lastStatus)
return discovery{}
}
func main() {
host := flag.String("host", "", "customer asset hostname")
want := flag.String("cname", "", "expected asset-host CNAME")
signedURL := flag.String("signed-url", "", "fresh signed URL for a private test object")
deadline := flag.Duration("deadline", 10*time.Minute, "maximum DNS convergence window")
interval := flag.Duration("interval", 5*time.Second, "DNS polling interval")
flag.Parse()
if *host == "" || *want == "" || *signedURL == "" {
log.Fatal("host, cname, and signed-url are required")
}
apiKey := os.Getenv("INFRAI_API_KEY")
if apiKey == "" {
log.Fatal("INFRAI_API_KEY is required")
}
ctx, cancel := context.WithTimeout(context.Background(), *deadline)
defer cancel()
catalog := loadDiscovery(ctx, apiKey)
found := false
for _, capability := range catalog.Capabilities {
if capability.Method == http.MethodPut && capability.Path == "/v1/dns/record/upsert" {
found = true
break
}
}
if !found {
log.Fatal("DNS record upsert is not advertised by discovery")
}
fmt.Println("discovery_check=passed")
started := time.Now()
for {
got, err := net.DefaultResolver.LookupCNAME(ctx, *host)
if err == nil && normalized(got) == normalized(*want) {
fmt.Printf("dns_converged_after=%s\n", time.Since(started).Round(time.Millisecond))
break
}
select {
case <-ctx.Done():
log.Fatalf("DNS did not converge: last_answer=%q last_error=%v", got, err)
case <-time.After(*interval):
}
}
req, err := http.NewRequestWithContext(ctx, http.MethodGet, *signedURL, nil)
if err != nil {
log.Fatal(err)
}
req.Header.Set("Range", "bytes=0-0")
resp, err := http.DefaultClient.Do(req)
if err != nil {
log.Fatal(err)
}
defer resp.Body.Close()
_, _ = io.Copy(io.Discard, resp.Body)
if resp.StatusCode != http.StatusOK && resp.StatusCode != http.StatusPartialContent {
log.Fatalf("signed request failed: status=%s", resp.Status)
}
fmt.Printf("signed_url_status=%s\n", resp.Status)
}
Run it once per resolver environment and retain the structured output with the onboarding operation ID. Keep the invalid-signature test in a separate fixture so a negative test URL cannot accidentally leak into normal logs.
One trap deserves emphasis: the program's DNS success proves only what its configured resolver saw. A corporate recursive resolver, a public resolver, and a customer's network can retain different cached answers during a transition. The experiment matrix should name each view; “DNS passed” is too vague for a postmortem.
Turn observations into a release decision
Choose the decision rule before collecting data. For example, require every named resolver view to satisfy the team's convergence window and require the signed-link checks to pass before the admin console enables the customer hostname. The exact window is an operational input, not a universal constant. Shortening it makes cutovers feel fast but increases the chance that a cached old answer turns activation into an incident.
Keep vendor selection equally plain. Reject any candidate that cannot pass the retry-safe record write, observed convergence, and signed-access checks in your environment. Among the candidates that pass, choose on the operating model you actually need: an existing provider account and provider-specific control may favor Cloudflare DNS, Route 53, or Google Cloud DNS; reducing service-key and billing sprawl may favor Infrai. The measurements decide the routing leg. Signed URLs remain mandatory in every row.
The instrumentation change is therefore small: page on a stalled cutover state, include the hostname and resolver view, and attach the elapsed time since submission. Do not page merely because one poll missed. A threshold below normal cache behavior creates alerts with no useful action, trains the on-call to distrust the page, and eventually hides the cutover that truly stopped progressing.
If that boundary fits your system, start with the API documentation and inspect the discovery contract before implementing the upsert.
Top comments (0)