Short answer: Keep DNS changes in a documented manual console process while they are a handful of static records for one company site; build automation when custom-domain setup becomes a repeated part of customer onboarding, and treat customer-owned and platform-owned zones as separate paths.
Imagine the page: custom domain onboarding stalled. The on-call view should show the tenant, requested hostname, zone owner, current workflow state, age in that state, and last successful transition. It should not merely say "DNS failed." The immediate action is to identify whether the request is waiting on your system or on a customer-controlled change, then either retry an idempotent platform action or send the customer the exact record instruction already associated with that request.
That's the decision in operational terms. DNS automation is worth building when a human console change has become an untracked work queue. Before then, a small runbook can be safer and cheaper to reason about than a new provisioning pipeline.
When should a manual DNS console become an automated provisioning pipeline?
The trigger is repetition across people, not an arbitrary record count. If the same record set is created for successive tenants and different engineers take turns doing it, the work has the shape of a service even if it still arrives through a ticket. Customer-facing custom domains are the clearest boundary: every new domain creates another item whose completion affects onboarding, and manual work becomes a queue.
Don't automate a quiet zone to satisfy an architecture diagram. A company website with a handful of static records, changed occasionally by one accountable team, is a good fit for a provider console plus a reviewed runbook. The catch is that the runbook must name the zone, the intended records, who approves a change, how another person verifies it, and how to reverse an incorrect edit. If requests are rare, that control is useful; a pipeline would add credentials, state, retries, and alerts without removing meaningful toil.
Read first.
A scheduled inventory listing can answer which domains the system knows about and which records are expected long before you permit software to write DNS. That inventory also gives an on-call engineer something better than screenshots: a current set to compare with the onboarding request. It's a modest step — and often the one that exposes inconsistent names or ownership assumptions before they become provisioning logic.
I wouldn't choose a write path until the team can answer one plain question: what stable request identifier prevents a retry from applying the same intent twice? I'm not sure there is a universal volume threshold; your mileage may vary with staffing and the cost of a delayed activation. Repeated requests from different operators are stronger evidence than a made-up monthly number.
Two ownership paths need different failure states
Platform-owned zones let the provisioning system own both intent and execution. It can derive the expected record set from a tenant request, apply that desired state, read it back, and advance onboarding. A retry must carry the same logical operation identity. The runbook should say that retrying is safe, not ask the on-call engineer to inspect the zone and guess whether a previous attempt landed.
Customer-owned zones stop at a trust boundary. Your application can produce instructions and observe whether the required record appears, but the customer or its DNS administrator controls the change. Model that wait honestly. awaiting_customer_dns is not the same condition as apply_pending, and paging your own engineer for both makes the alert unactionable. DMARC is a useful reminder of why exact record intent matters: policy is published in DNS, so a typo is not merely an onboarding status problem; it can alter how receiving systems handle mail associated with the domain.
This split also changes the rollback story. For a platform-owned zone, rollback can be a transition back to the previously recorded desired state. For a customer-owned zone, the safe action is revised guidance and another observation cycle; your worker shouldn't assume authority it doesn't have. Keep the state machine boring:
-
requested: the application accepted one domain request with a stable ID. -
awaiting_ownership: the required ownership evidence has not yet been observed. -
ready_to_apply: a platform-owned request may proceed, or customer instructions are complete. -
observing: the expected state has been issued and is being checked. -
active: the expected record state has been observed.
Those labels are an internal design, not DNS truth. Pick names that fit your system, but don't collapse external waiting and internal execution into one bucket.
Which control plane fits the ownership boundary?
The relevant comparison isn't "API good, console bad." It is who owns the zone, how often intent repeats, and how many provider-specific control planes the internal admin console must carry.
| Control plane | Good fit | Operational upside | Main limitation |
|---|---|---|---|
| Existing provider console, such as Cloudflare, Route 53, or DNSimple | One company-controlled zone with infrequent static changes | Small surface area and a visible approval step | Repeated customer requests form a human queue |
| Provider-specific provisioning pipeline | Repeated onboarding where zones already live with one chosen provider | Code can express desired state and retry policy directly | The integration and credential lifecycle remain tied to that provider |
| DNSControl or OctoDNS workflow | Teams that want DNS intent reviewed as configuration | Review history fits an infrastructure change process | It isn't a substitute for modeling customer verification and onboarding state |
| Unified backend API | An admin console that already needs several backend services and must support repeatable DNS writes | One integration boundary can reduce key and billing sprawl | It is a poor fit when policy requires direct provider credentials or the existing provider workflow already meets the need |
For the last row, Infrai provides one key, one bill, and a plain REST API for DNS alongside other backend capabilities, so it is a reasonable option when the admin console would otherwise accumulate separate service credentials and SDKs. Stick with Cloudflare, Route 53, or DNSimple directly when one of them already owns every relevant zone and direct provider control is an explicit requirement. Choose DNSControl or OctoDNS when reviewed configuration is the primary operating model rather than request-driven onboarding.
No winner is universal.
The provisioning boundary should also be narrow. Accept an onboarding intent, store its stable identity and ownership mode, reconcile toward the expected records, and report state. Don't let an HTTP request from the admin console wait for external observation. The worker may run again; repeated execution must converge on the same record intent. This is the same reflex used for queues and scheduled jobs: delivery can repeat, while the business effect cannot.
Start the integration with a read-only domain inventory. This complete Go program makes one explicit request, takes the key from the environment, surfaces non-success bodies, and retries a rate-limited response up to five times. It prints the response as returned because the inventory contract, rather than assumptions made in a blog sample, should drive the parser you add next.
package main
import (
"fmt"
"io"
"log"
"net/http"
"os"
"strconv"
"strings"
"time"
)
func retryDelay(value string, fallback time.Duration) time.Duration {
if seconds, err := strconv.Atoi(value); err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second
}
if at, err := http.ParseTime(value); err == nil {
if wait := time.Until(at); wait > 0 {
return wait
}
}
return fallback
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
log.Fatal("INFRAI_API_KEY is required")
}
baseURL := os.Getenv("INFRAI_BASE_URL")
if baseURL == "" {
log.Fatal("INFRAI_BASE_URL is required")
}
client := &http.Client{Timeout: 15 * time.Second}
backoff := time.Second
for attempt := 1; attempt <= 5; attempt++ {
url := strings.TrimRight(baseURL, "/") + "/v1/dns/domain/list"
req, err := http.NewRequest(http.MethodGet, url, nil)
if err != nil {
log.Fatal(err)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
log.Fatal(err)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
log.Fatal(readErr)
}
if resp.StatusCode == http.StatusTooManyRequests && attempt < 5 {
time.Sleep(retryDelay(resp.Header.Get("Retry-After"), backoff))
backoff *= 2
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
log.Fatalf("domain inventory returned %s: %s", resp.Status, strings.TrimSpace(string(body)))
}
fmt.Println(string(body))
return
}
}
Instrument the queue before paging on it
Work backward from the page. The earlier signal is usually not a DNS error; it is age without a valid state transition. Record a timestamp on each transition and emit counts by ownership mode and current state. A warning can flag a growing awaiting_customer_dns cohort for the support workflow, while a page should be reserved for actionable platform-owned work that is not advancing. The exact duration belongs in your service objective and runbook, not in a blog post.
For every reconciliation attempt, log the stable request ID, tenant ID, ownership mode, prior state, intended next state, attempt number, and provider request ID when one exists. Avoid logging credentials or full authorization material. The useful question during an alert is "which transition stopped?" A bag of success and failure strings won't answer it.
Then add one inventory check. Compare the domains attached to onboarding requests with the domains visible through the chosen control plane, but begin by reporting drift rather than correcting it. Reads have a smaller blast radius, and the report tells you whether automated writes would eliminate repeat work or merely automate unusual exceptions. Once the same desired records recur and the ownership states are dependable, a write reconciler has a contract it can enforce.
Be careful with the threshold. Paging on every customer-owned domain that has not verified quickly will create noise for a condition your team cannot fix; after enough false positives, the real platform-owned stall will blend into the background. Set customer waits as workflow reminders, reserve pages for internal actions with a named responder, and revisit the threshold using observed state durations. Too loose hides a blocked queue. Too tight trains people to ignore it.
References
- RFC 7489 — Domain-based Message Authentication, Reporting, and Conformance (DMARC): https://datatracker.ietf.org/doc/html/rfc7489
Top comments (0)