TL;DR: for a media hostname cutover, use create semantics to establish the first ownership claim, then treat an existing record whose full intended state differs as a conflict. Permit an idempotent retry only when it describes the same RRset, and make a rollback conditional on the version that was approved. This makes propagation delay visible instead of mistaking a successful write for completed cutover.
The least complex workable plan preserves the old target, publishes one reviewed candidate through a controlled write, and keeps enough evidence to establish which version was authoritative at each decision. A 300-second TTL is a cache-control instruction, not a promise that every viewer changes at the same instant.
The bill is mostly retained change evidence
The bill for a safe DNS change is usually not the 12 RRsets involved in a media cutover. It is the retained evidence around them: the pre-change observation, approved desired record, write result, and post-change observation. Four artifacts for each of 12 hostnames produces 48 reviewable artifacts, so the dominant object count is the change evidence rather than the DNS data itself.
That count should shape the design before it shapes a storage bill. Keep a compact immutable event for every transition, keyed by a change ID and record version, and avoid retaining unbounded recursive-resolver logs whose consent, access-control, and deletion obligations can exceed their operational value. An audit record needs the authoritative observation, timestamp, TTL, actor, approval reference, requested disposition, and resulting disposition; it does not need an indefinite copy of every resolver trace. The trade-off is real: when a regional cache report arrives later, discarding detailed resolver traces reduces the precision of the investigation. It also keeps the retention purpose legible, which matters where compliance rules require a defined reason and duration for operational records.
The change that moves the retained-data term is aggregation. Store one signed or append-only transition record per relevant RRset state, rather than a stream of duplicate observations that cannot alter the cutover decision.
This is intentional loss.
A cutover timer belongs in that record and should be derived from the published TTL and rollout policy, not inferred from a write response. RFC 2181 specifies that cached TTL values decrement, while RFC 2308 covers negative caching; both are reasons to keep an authoritative response distinct from the response a particular resolver may still serve.
What DNS upsert or create failure should provisioning expose?
Create and upsert encode different failures. A create that finds an existing www record can return a conflict, forcing the provisioning workflow to compare the observed record with the desired record. An unconditional upsert can turn the same event into a successful overwrite, losing the signal that a newsroom launch, a certificate workflow, or a separate deployment has already claimed the name.
Choose the conflict when a workflow is establishing ownership of a name or when a retry has no trustworthy record version. Choose an idempotent no-op only after comparing the full record identity and intended data. For one A record, that means at least owner name, type, target value, and TTL. For a set-valued RRset, comparison must cover the complete set, because matching one member does not establish that the intended set is present.
Faster cutover comes from allowing a replacement to proceed. Safer provisioning comes from stopping when current state is no longer the state the change request approved. A conflict is useful operational data.
The following self-contained Go program models the decision. It rejects a mismatched existing record, accepts an exact retry, and requires a version match before rollback changes stored state. The documentation address range in the example is reserved for examples by RFC 5737.
package main
import (
"fmt"
)
type Record struct {
Name string
Type string
Value string
TTL uint32
Version int
}
type Store map[string]Record
func sameIntent(a, b Record) bool {
return a.Name == b.Name && a.Type == b.Type && a.Value == b.Value && a.TTL == b.TTL
}
func provision(s Store, want Record) error {
have, exists := s[want.Name]
if !exists {
s[want.Name] = want
return nil
}
if sameIntent(have, want) {
return nil // An exact retry is idempotent.
}
return fmt.Errorf("conflict for %s: observed version %d differs from approved intent", want.Name, have.Version)
}
func rollback(s Store, prior Record, expectedVersion int) error {
current, exists := s[prior.Name]
if !exists || current.Version != expectedVersion {
return fmt.Errorf("rollback conflict for %s", prior.Name)
}
prior.Version = current.Version + 1
s[prior.Name] = prior
return nil
}
func main() {
before := Record{Name: "video.example.test", Type: "A", Value: "198.51.100.10", TTL: 300, Version: 7}
proposed := Record{Name: "video.example.test", Type: "A", Value: "198.51.100.20", TTL: 300, Version: 8}
records := Store{before.Name: before}
fmt.Println(provision(records, proposed))
fmt.Println(rollback(records, before, 7))
}
The first line reports a conflict, which is the desired result before an approved replacement operation exists. The second line succeeds because the stored record remains version 7. If another actor had changed it first, rollback would stop instead of restoring an obsolete target over a newer decision.
Carry the idempotency key through approval
An idempotency key attached only to a transport request is insufficient. Bind one change ID to the intended RRset, expected current version, approver, and expiry of the cutover window. Persist its final disposition: created, no-op, conflict, replaced, or rolled back. A repeated request with the same ID can then return the prior disposition without issuing another mutation.
The awkward state is a timeout after the authoritative service accepts a write but before the client receives a response. A blind upsert makes that ambiguity disappear on paper and may create a second, unrelated mutation. A retry that first reads and compares intent produces three actionable outcomes: the desired record is already present; the prior version remains present and can be retried under the same precondition; or another value exists and needs a fresh decision. This is the same exactly-once mindset used for ledger mutations, applied to an external state machine where a timeout does not reveal whether the mutation committed.
For DMARC records, the discipline is particularly important because the _dmarc owner name carries policy data receiving systems consume. RFC 7489 defines its record format and policy behavior, so a TXT change should be treated as a complete intended record with the same review and rollback evidence as an application hostname.
Rollback is another controlled write
Prepare rollback before the first mutation. Its payload should be the exact prior RRset, not an instruction such as "put it back." Record the version observed immediately before the forward change, the intended post-change version, the TTL, and the point after which rollback needs explicit re-approval because the operational context may have changed.
During the waiting period, measure authoritative responses from the expected nameservers and separately observe representative recursive resolution where that is permitted and useful. Neither observation establishes universal client visibility. DNS delegation, caching, resolver behavior, and application connection reuse make a hostname cutover a distributed transition.
Do not automate rollback from a single resolver result.
A resolver can still be within its cache horizon, and one probe cannot establish broad reachability. Automate collection and correlation; require an approved decision to replace a currently different record. The resulting runbook is short enough to execute under pressure: create only when the name is absent, compare on every retry, replace only with a reviewed version precondition, and record the disposition before moving to the next hostname. This sacrifices some cutover speed when ownership is ambiguous, which is the correct cost for preserving an audit trail and avoiding an unreviewed deployment.
Top comments (0)