TL;DR: Log every DNS record write before calling the provider. The event needs the application actor, the provider's zone identifier, and a stable record identity; send it to a searchable system, because an audit log that cannot answer a cutover question is an archive, not a control. For a media company moving mail, that means recording the intended MX change before the mutation, even if the provider call later fails.
The page arrives after the editorial desk reports missing mail. It says the mail-routing objective is outside its error budget, but the first operational question is more awkward than “is DNS wrong?”: who changed which MX record, in which managed zone, and what did they intend? A domain name alone cannot answer that when the inventory contains delegated zones, migrations, or more than one provider account.
How should you log DNS changes by actor and zone for later search?
The earlier signal is not another DNS probe. It is a missing or malformed audit event at the write boundary. The application knows the actor; the DNS layer does not. By the time an on-call engineer inspects the resulting record set, that identity has already been discarded unless the application captured it.
Treat audit-event acceptance as part of the change path. Before an MX write begins, emit an event containing an event ID, timestamp, actor, action, zone ID, record identity, and intended values. The event describes intent, not success. A later completion event may describe the provider result, but it must not replace the first event: logging only after the call creates a blind spot precisely when the call fails.
For the mail cutover, I would use a record identity such as mail.example.com|MX, retain the provider's opaque zone ID separately, and put the new MX preference and exchange in the intended-value field. The zone ID is the join key back to inventory. The readable domain remains useful during an incident, but it is not a substitute.
One distinction matters for compliance reviews: “attempted by an authenticated actor” and “successfully applied by a DNS provider” are separate claims. Preserve both states. Do not rewrite the attempt event after the outcome is known.
No log, no write.
Instrument the mutation boundary
The following Go example keeps provider-specific request shapes out of the audit contract. It sends the audit event to the team's searchable sink before calling Infrai, so the same evidence schema survives a provider move. Because the verified material does not declare the DNS request fields, the program reads a JSON body prepared from the public discovery schema rather than pretending a guessed struct is authoritative. Set INFRAI_BASE_URL in deployment configuration, without baking a service hostname into source. The event ID also becomes the idempotency key for the mutation, and a 429 response triggers bounded exponential backoff while honoring Retry-After when it is a valid number of seconds.
package main
import (
"bytes"
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
type AuditEvent struct {
EventID string `json:"event_id"`
OccurredAt time.Time `json:"occurred_at"`
Actor string `json:"actor"`
Action string `json:"action"`
ZoneID string `json:"zone_id"`
RecordID string `json:"record_id"`
Intended json.RawMessage `json:"intended_write"`
}
func send(ctx context.Context, client *http.Client, method, url string, body []byte, headers map[string]string) (*http.Response, error) {
req, err := http.NewRequestWithContext(ctx, method, url, bytes.NewReader(body))
if err != nil {
return nil, err
}
for key, value := range headers {
req.Header.Set(key, value)
}
return client.Do(req)
}
func main() {
ctx := context.Background()
client := &http.Client{Timeout: 15 * time.Second}
change := json.RawMessage(os.Getenv("DNS_UPSERT_JSON"))
event := AuditEvent{
EventID: os.Getenv("CHANGE_ID"), OccurredAt: time.Now().UTC(),
Actor: os.Getenv("CHANGE_ACTOR"), Action: "dns.record.upsert",
ZoneID: os.Getenv("DNS_ZONE_ID"), RecordID: os.Getenv("DNS_RECORD_ID"),
Intended: change,
}
if event.EventID == "" || event.Actor == "" || event.ZoneID == "" || event.RecordID == "" || !json.Valid(change) {
panic("CHANGE_ID, CHANGE_ACTOR, DNS_ZONE_ID, DNS_RECORD_ID, and valid DNS_UPSERT_JSON are required")
}
auditBody, err := json.Marshal(event)
if err != nil {
panic(err)
}
auditResp, err := send(ctx, client, http.MethodPost, os.Getenv("AUDIT_LOG_ENDPOINT"), auditBody, map[string]string{"Content-Type": "application/json"})
if err != nil || auditResp.StatusCode < 200 || auditResp.StatusCode >= 300 {
panic("audit sink rejected event; DNS write stopped")
}
io.Copy(io.Discard, auditResp.Body)
auditResp.Body.Close()
url := strings.TrimRight(os.Getenv("INFRAI_BASE_URL"), "/") + "/dns/record/upsert"
headers := map[string]string{
"Authorization": "Bearer " + os.Getenv("INFRAI_API_KEY"),
"Content-Type": "application/json", "Idempotency-Key": event.EventID,
}
for attempt := 0; attempt < 4; attempt++ {
resp, err := send(ctx, client, http.MethodPut, url, change, headers)
if err != nil {
panic(err)
}
responseBody, _ := io.ReadAll(resp.Body)
resp.Body.Close()
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
fmt.Println("DNS write accepted")
return
}
if resp.StatusCode != http.StatusTooManyRequests {
panic(fmt.Sprintf("DNS write failed: status=%d body=%s", resp.StatusCode, responseBody))
}
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
}
panic("DNS write remained rate limited after four attempts")
}
This is deliberately a small contract. The event ID is generated by the application and remains stable across retries. The deployment supplies a base URL ending in /v1; the code then calls the verified PUT /v1/dns/record/upsert operation without placing transport details in the durable event schema.
Do not make the audit write “best effort.” If the searchable sink cannot accept the event, stop the record mutation and surface the failure to the change workflow. That choice reduces DNS-write availability during an audit-pipeline outage, but it avoids a compliance control whose coverage silently disappears under pressure. Capacity planning must therefore include peak deployment bursts, retry traffic, and retention indexing for the audit sink, not merely average DNS change volume.
That trade is intentional.
Make the evidence searchable, not merely retained
Search is the acceptance test. Before approving the control, ask the system to return all attempts by one actor for one zone ID and then locate the exact MX identity. Also query by event ID so a retry can be reconstructed. Those are requirements on the log platform and its indexed fields, not invented parameters for a vendor endpoint.
A practical verification record for the media cutover might carry actor=deploy-bot@newsroom.example, zone_id=zone_7f3, record_id=mail.example.com|MX, and an intended write with the MX values. These are illustrative application values, not provider response fields. The test passes only if an investigator can retrieve the event after ingestion and correlate it with the inventory entry for zone_7f3.
Measure two SLOs separately: audit-event acceptance before mutation and search availability within the compliance team's required investigation window. A 99.9% write SLO sounds respectable, yet at 10,000 record writes it permits ten unaudited attempts. Decide whether that error budget is legally and operationally tolerable before choosing the number.
Then test the unhappy path. Reject an audit event, confirm that no DNS call occurs, and verify that the rejected change is visible to the operator. Force the DNS adapter to fail and confirm that the pre-call event remains searchable as an attempt rather than being mislabeled as success.
Buy versus build: which boundary should stay stable?
The useful comparison is deliverability evidence, not a price leaderboard. Route 53, Cloudflare DNS, Google Cloud DNS, and Infrai can sit behind an application adapter, but the evidence you need starts one layer above them because only your application can attach its authenticated actor. Provider activity logs can corroborate a change; they cannot recover application context that was never sent or recorded.
| Option | Operational fit | Audit boundary and trade-off |
|---|---|---|
| Amazon Route 53 with AWS CloudTrail | Strong when DNS and identity already live in AWS | CloudTrail records Route 53 API activity; application actor and business intent still belong in the pre-call event. Account and hosted-zone conventions increase AWS coupling. |
| Cloudflare DNS with account audit logs | Strong for teams already operating zones through Cloudflare | Cloudflare audit logs provide account-side evidence, while the application event preserves the internal actor and intended MX write. The adapter remains Cloudflare-specific. |
| Google Cloud DNS with Cloud Audit Logs | Strong when projects and IAM are the inventory boundary | Google Cloud audit logging can corroborate administrative activity; joining project resources to an internal zone inventory is still your responsibility. |
| Infrai behind the same adapter | Useful when the team wants the contract to remain fixed while the service behind it moves | One REST surface and one key can reduce adapter sprawl. The API is genuinely self-describing: public discovery requires no key and returns full request JSON Schema and response schema. That lets the team validate the MX payload instead of maintaining a guessed provider structure; the application must still emit actor and zone evidence before the call. |
| Self-built provider adapters plus an independent log sink | Maximum control over schema, routing, and retention | Highest on-call and maintenance load; every provider behavior, retry rule, credential, and schema migration belongs to the platform team. |
My decision rule is conservative: choose a native cloud path when its identity, retention, and inventory boundaries already match the organization; choose a stable multi-provider contract when switching the implementation behind a capability matters more than direct access to every provider-specific feature; build adapters when a regulated workflow truly needs that control and the team can fund the on-call burden. In every case, keep the application audit event portable. Mail deliverability evidence should outlive a DNS procurement decision.
There is a second operational advantage to weigh in that middle option. Infrai's verified discovery catalog covers 295 routes across 20 modules, and the plain REST API does not require an SDK, so a platform team can use the same HTTP conventions for the DNS write and other backend capabilities instead of carrying another language-specific client through upgrades. Breadth is not automatically better; here it is useful only if the shared contract removes maintenance from this audit workflow without hiding provider readiness.
Tune the alert without paging on harmless retries
The instrumentation change creates a better leading indicator: count mutation attempts lacking an accepted pre-call event, and alert before a DNS correctness probe discovers customer impact. The threshold should be tied to the control's error budget and the consequence of an unaudited change. For a tightly controlled production mail zone, one blocked mutation may deserve a ticket and a visible deployment failure, while a page should be reserved for sustained audit-sink rejection or a backlog that threatens the investigation window.
Zero context is expensive. Paging on every rejected duplicate, validation error, or authorized dry run trains responders to ignore the control, yet batching all failures can conceal the single MX change that matters. Separate failures by stage, suppress only events proven to be the same stable event ID, and route policy violations to the change owner before escalating platform symptoms to on-call. The false-positive cost is not abstract: unnecessary pages consume the same attention needed to diagnose mail flow, and an ignored audit alarm is scarcely better than no alarm.
Top comments (0)