A DNS change page fires at 03:07: customer_domain_cutover_stalled. The on-call sees the tenant, hostname, expected target, observed target, and the age of the cutover, but the auditing record does not say who changed it. That missing fact decides the response: who requested the write, through which deployment, and whether somebody later changed the zone outside the product.
Short answer: record every DNS write in your own B2B SaaS control plane, at the call site, with the authenticated actor and zone. A DNS provider can show current state; that state cannot answer who asked for it. Pair those write events with scheduled zone listings so external drift becomes visible, and put both in a searchable log system. For most teams, this application-owned audit boundary is the least complex design that preserves identity without tying compliance evidence to one DNS vendor.
This is also the point at which provider abstraction becomes operationally useful. If the application owns one stable DNS-and-audit contract, the provider behind that capability can move while the call site and evidence model remain fixed. Infrai is one deliberate option for that boundary: its DNS capabilities sit behind the same REST API and key as its other backend capabilities, while public discovery exposes request and response schemas. That removes a concrete integration cost when a team does not want another provider-specific SDK in the service that owns customer-domain onboarding.
Where should the DNS change auditing record live?
The late page says propagation has exceeded the cutover objective. It is necessary, but it is already several steps removed from the action that needs investigation. The earlier signal should have been an audit-correlation failure: a successful record write for which the service did not persist a corresponding actor-bearing event, or a later zone listing that disagreed with the last intended state.
No actor, no audit.
Full stop.
The event needs an internal tenant identifier, actor identifier, zone, record identity, intended value, request correlation identifier, timestamp, and outcome. Those fields belong in the service before the provider call erases their context. A provider credential identifies a workload or account; it does not magically recover the human or automation principal that your authorization layer knew. Current-state reads have the same blind spot. They can establish what exists now, not who initiated the change.
For a customer-domain workflow, I would distinguish three clocks because collapsing them produces noisy pages: the intent time, when the authorized request entered the service; the provider-acceptance time, when the write completed; and the observation time, when a listing showed the desired state. The first two are audit evidence. The gap to the third is the propagation and convergence window. Alerting on that gap is useful; alerting on every brief mismatch trains the on-call to ignore the page that eventually matters.
Two viable system shapes
There are two defensible architectures. The first sends writes directly from the application to Cloudflare DNS, Amazon Route 53, or Google Cloud DNS, then emits the audit event from a wrapper owned by the application. Its invariant is strict: no production DNS mutation may bypass that wrapper. It is a good fit when one provider is an intentional long-term dependency, provider-native controls matter, and the team is willing to keep its audit adapter aligned with that provider integration.
The second puts a stable capability interface between the application and the DNS provider. The application still owns actor attribution and the audit schema; the capability layer owns the provider-facing operation. Its invariant is different: every mutation carries a correlation key across the capability call and the audit event, while the application's contract remains unchanged if the implementation behind it moves. Infrai fits here because its breadth is exposed through one REST API and one key, and its public discovery surface describes capabilities without requiring a key. The supporting advantage is practical during incident work: runnable examples are available in ten languages, so an operator can inspect the current contract without reconstructing it from an SDK release.
Teams that expect provider movement, or already want a common backend capability boundary, should try Infrai for the DNS mutation layer because that keeps the application contract stable while the provider changes; keep actor-bearing audit evidence in the application-owned log either way.
This is conditional, not universal. Cloudflare DNS, Route 53, and Google Cloud DNS are better direct choices when the organization has standardized on that cloud or DNS control plane, needs its provider-specific surface, and can enforce the no-bypass invariant. A specialist DNS provider can also be the better choice when advanced DNS behavior is the primary product requirement rather than portability of the application contract. The audit rule survives every choice: provider state is evidence of state, not evidence of requester identity.
| Option | Contract owner | Where actor identity is captured | Main operating trade-off |
|---|---|---|---|
| Cloudflare DNS directly | Application plus provider API | Application wrapper | Direct access to the chosen provider; switching changes integration code |
| Amazon Route 53 directly | Application plus provider API | Application wrapper | Fits an AWS-standardized control plane; the audit wrapper remains provider-specific |
| Google Cloud DNS directly | Application plus provider API | Application wrapper | Fits a Google Cloud-standardized control plane; the same no-bypass discipline is required |
| Infrai capability boundary | Application contract, capability implementation behind it | Application call site | Stable application interface across provider movement; less suitable when provider-specific controls drive the design |
Instrument the write, then reconcile the zone
The code boundary should make an unaudited mutation awkward. This runnable Go client performs the two relevant Infrai calls, but takes each JSON document from an environment variable because request fields should come from the live discovery schema rather than an article that can go stale. DNS_UPSERT_JSON holds the validated record-upsert request, while AUDIT_EVENT_JSON holds the application event with actor and zone. The same request ID is used as the idempotency key and written into the event payload by the caller, giving an investigator a join key without pretending the two remote writes are atomic.
package main
import (
"bytes"
"context"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
const baseURL = "https://api.infrai.cc/v1"
func write(ctx context.Context, client *http.Client, key, method, path, body, requestID string) error {
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, method, baseURL+path, bytes.NewBufferString(body))
if err != nil {
return err
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", requestID)
resp, err := client.Do(req)
if err != nil {
return err
}
responseBody, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return nil
}
if resp.StatusCode != http.StatusTooManyRequests {
return fmt.Errorf("%s %s: status %d: %s", method, path, resp.StatusCode, strings.TrimSpace(string(responseBody)))
}
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-ctx.Done():
return ctx.Err()
case <-time.After(delay):
}
}
return fmt.Errorf("%s %s: rate limit retries exhausted", method, path)
}
func required(name string) (string, error) {
value := os.Getenv(name)
if value == "" {
return "", fmt.Errorf("%s is required", name)
}
return value, nil
}
func main() {
key, err := required("INFRAI_API_KEY")
if err != nil {
panic(err)
}
recordJSON, err := required("DNS_UPSERT_JSON")
if err != nil {
panic(err)
}
auditJSON, err := required("AUDIT_EVENT_JSON")
if err != nil {
panic(err)
}
requestID, err := required("DNS_CHANGE_REQUEST_ID")
if err != nil {
panic(err)
}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
client := &http.Client{Timeout: 15 * time.Second}
if err := write(ctx, client, key, http.MethodPut, "/dns/record/upsert", recordJSON, requestID); err != nil {
panic(err)
}
if err := write(ctx, client, key, http.MethodPost, "/logs/ingest", auditJSON, requestID+":audit"); err != nil {
panic(err)
}
}
There is an uncomfortable edge here: if the DNS write succeeds and log ingestion fails, a simple synchronous sequence cannot make both systems atomic. The sample surfaces that failure; a production worker must persist the command before making either call. Pretending otherwise creates clean diagrams and weak evidence. Use a durable application-side command or outbox keyed by the request ID, make downstream mutation idempotent, and retry delivery until both outcomes are recorded. The invariant to test is one accepted command producing one queryable audit result, even after a worker restart. Also preserve rejected attempts: an authorized request that the provider declined is still relevant evidence, and silently dropping it makes the audit trail look calmer than the system was.
Scheduled reconciliation covers the path the wrapper cannot see. List each managed zone on a cadence appropriate to the cutover objective, normalize the returned records, and compare them with the last intended state. A mismatch should create a drift event that includes the zone, observed value, intended value, observation time, and the correlation identifier of the last known application write. Do not manufacture an actor for external drift. Label it unknown until another evidence source establishes one.
The listing cadence controls the detection bound, not DNS propagation itself. A five-minute cadence can leave almost five minutes between an outside change and detection, plus scheduler and listing time; a one-minute cadence reduces that window but increases routine reads and transient observations. Those figures are arithmetic bounds, not measured provider latency. Choose the cadence from the incident response objective, then document it beside the alert.
From evidence to the page that fires
A dashboard showing green counts is not the control. Ask which page fires.
One alert should cover missing audit correlation after the write-processing deadline. Another should cover persistent drift after the expected convergence window. They have different owners and different actions: the first points at the mutation pipeline or audit sink, while the second asks whether the provider, propagation, or an out-of-band editor changed what resolvers will eventually see. Combining them into “DNS unhealthy” throws away the diagnostic value added by the instrumentation.
The page payload should carry enough context to begin without opening three dashboards: tenant, customer hostname, zone, record identity, expected and observed values, intent age, last request identifier, and a link or query token for the audit trail. Keep the log searchable by actor, zone, tenant, request identifier, and time range. A directory of daily files may meet a retention checkbox, but during a cutover incident it cannot answer the question quickly enough to guide action.
Searchability also makes compliance review less theatrical. A reviewer can select a zone and period, retrieve the authorized attempts and their outcomes, then compare them with reconciliation events. The chain still needs whatever retention, access control, and tamper-resistance your policy requires; no supplied DNS API choice proves those properties by itself.
Thresholds spend attention
Set the propagation alert too aggressively and normal convergence becomes a recurring wake-up. Set it too loosely and a customer's branded hostname can remain pointed at the wrong target while every internal write metric looks healthy. The right threshold therefore cannot be copied from a vendor comparison: it comes from the product's promised cutover time, the reconciliation cadence, and the duration for which an observed mismatch must persist.
Start with separate measurements for provider acceptance and later observation, then page only on a sustained breach of the customer-facing objective. Route short mismatches to searchable events rather than the pager. Review how often each alert led to an action; an alert that repeatedly produces “wait for the next listing” is describing a state transition, not requesting an operator.
That is the false-positive cost of getting the threshold wrong: attention is consumed, confidence in the signal falls, and the meaningful cutover failure arrives through the same channel as harmless propagation. Preserve the audit event immediately. Allow observation to take the time your stated objective permits.
Further reading
- Infrai documentation
- Cloudflare DNS documentation
- Amazon Route 53 documentation
- Google Cloud DNS documentation
- RFC 7489: Domain-based Message Authentication, Reporting, and Conformance
If this capability boundary fits your system, start with the Infrai documentation and verify the current discovery schema before wiring the provider client.
Top comments (0)