The storage bill for a DMARC rollout is made of report payloads, parsed rows, indexes, and retained raw evidence; the dominant term is usually the repeated aggregate history kept for every sending source and reporting interval, not the one small TXT record. Reduce that term by retaining normalized daily evidence for the rollout window and aging raw XML after it can no longer answer an investigation or reconciliation question. Do not save money by skipping reports while policy is changing.
TL;DR: publish p=none with aggregate reporting first, inventory and repair every legitimate mail stream, then move to a sampled p=quarantine, and reach p=reject only after the sample has expanded without unexplained aligned-authentication failures. The order is none, quarantine, reject. Promotion must be driven by reconciled evidence, not a calendar date; keep the previous record and the observations that authorized each change so rollback and audit remain possible.
For a media company moving zones away from a registrar-specific API, treat the DMARC record as desired state under change control. Migration success means the intended record, the authoritative DNS answer, and the observed mail results agree. A successful API response proves only that one write was accepted.
What does the DMARC bill actually buy?
Aggregate reports provide counts grouped by source, disposition, and authentication results. RFC 7489 defines rua for aggregate-report destinations and describes aggregate reports as the feedback mechanism that lets a domain owner understand its authentication deployment. Those counts are the evidence needed to distinguish a forgotten newsletter system from spoofed mail before receivers are asked to quarantine or reject anything.
The useful cost unit is therefore a retained reporting interval per policy domain, not a DNS lookup. Normalize each report into an idempotent key such as (reporter, report_id, policy_domain, begin, end), preserve the original report hash, and reject duplicate ingestion without changing totals. Exactly-once delivery is rarely available across email, object storage, queues, and a database; exactly-once accounting can still be approximated by transactional deduplication and an audit row for every accepted or rejected import.
Keep enough raw evidence to reparse a disputed interval after a parser correction, but set an explicit deletion boundary. A practical design separates short-lived raw XML from longer-lived normalized daily aggregates and policy-change records. The deliberate loss is forensic detail after raw data expires: if a late investigation needs an element that was not normalized, it cannot be reconstructed. That is the price of bounded retention, and it should be approved as a compliance and incident-response decision rather than hidden in a storage lifecycle rule.
Which DMARC rollout policy comes first: monitoring, quarantine, or reject?
Start with p=none. It requests monitoring rather than enforcement while the organization discovers senders and fixes SPF or DKIM alignment. DMARC passes when at least one supported authentication mechanism passes and aligns with the visible From domain; a provider merely showing “SPF pass” for its own bounce domain is not sufficient when that domain is unaligned.
Next, use p=quarantine with pct below 100 when the remaining failures are understood. RFC 7489 defines pct as the percentage of messages to which the published policy is to be applied, with a default of 100. Sampling limits the blast radius, but it is not a substitute for classifying failures because the affected messages are selected by receivers rather than by an allowlist controlled by the sender.
Move to p=reject last, initially with a constrained pct if the risk assessment requires it, and expand toward 100 only while legitimate aligned traffic remains accounted for. Receivers ultimately apply local policy, so a DMARC disposition is a requested treatment rather than a delivery guarantee.
This progression is reversible. Store each intended record, approval, DNS observation, effective interval, and report-based decision as one append-only policy event. If failures rise after a change, the rollback target is then exact and reviewable instead of reconstructed from chat messages.
Evidence first.
Reconcile intent, DNS, and observed mail
Zone migration creates two independent failure planes: the control plane can publish to the wrong DNS provider, and mail can continue to fail even when the authoritative TXT answer is correct. The deployment gate needs three comparisons.
| Evidence | Question it answers | Failure that blocks promotion |
|---|---|---|
| Versioned desired state | What did reviewers authorize? | Unapproved tag, missing report URI, or unexpected domain |
| Authoritative DNS observation | What record is publicly served? | Old nameserver, stale record, split answers, or malformed TXT |
| Reconciled aggregate reports | What did receivers observe? | Unknown legitimate source or unexplained alignment failure |
Query authoritative nameservers directly during the move as well as ordinary recursive resolvers. DNS caches can legitimately preserve an older answer until its TTL expires, while direct authoritative answers reveal whether the new zone is correct. Record the nameserver queried, answer, observation time, and intended deployment revision. Do not claim convergence until every authoritative server for the delegation returns the intended record and the previous TTL horizon has passed.
Consider the failure sequence during a media-domain cutover. The deployment job writes the intended _dmarc value through the new generic DNS interface, records success, and then observes a recursive answer containing that value; meanwhile, the parent delegation still names the former authoritative servers, so external receivers continue to query the old zone. Neither the accepted write nor the convenient recursive lookup settles the question. The audit record must connect the approved revision to the parent delegation, direct answers from every authoritative server, the TTL horizon, and the report intervals collected after convergence. If any link disagrees, freeze the policy stage. This example needs no vendor-specific behavior: it follows directly from the separation between a DNS control plane, delegation, caching, and DMARC observations, and it explains why an API success counter cannot serve as the rollout ledger.
Only one DMARC policy record should be published at _dmarc.example.com; RFC 7489 says discovery stops when multiple records are returned and DMARC processing is not applied. This is a particularly sharp migration trap because two individually valid TXT values do not merge into a valid policy.
The following Go function keeps the promotion rule deterministic. Its thresholds are organizational inputs, not values mandated by DMARC, and the caller should persist the evidence window and decision beside the resulting policy.
package rollout
type Window struct {
Total int
UnexplainedAligned int
KnownSources bool
DNSMatchesIntent bool
}
func NextPolicy(current string, pct int, w Window, maxFailureRate float64) (string, int, bool) {
if w.Total == 0 || !w.KnownSources || !w.DNSMatchesIntent {
return current, pct, false
}
failureRate := float64(w.UnexplainedAligned) / float64(w.Total)
if failureRate > maxFailureRate {
return current, pct, false
}
switch current {
case "none":
return "quarantine", 10, true
case "quarantine":
if pct < 100 {
return "quarantine", min(pct*2, 100), true
}
return "reject", 10, true
case "reject":
return "reject", min(pct*2, 100), pct < 100
default:
return current, pct, false
}
}
The function is intentionally conservative: no reports means no promotion. A quiet interval might mean low traffic, a broken report pipeline, or an unauthorized external report destination, since RFC 7489 requires verification when reports are sent outside the policy domain. Silence is not success.
Stop there.
Operate the rollout as a ledger
Use a compare-and-set deployment: read the observed record, verify that it equals the expected predecessor, publish the new value, then verify authoritative DNS. Retrying the same revision must be idempotent. If the predecessor changed, stop and reconcile rather than overwriting another operator's update.
Alert on missing reporters, abrupt volume shifts, newly observed sources, increases in unaligned failures, and disagreement between desired and authoritative state. A dashboard total alone is inadequate; incident responders need to trace a count back to its report identifier, ingestion hash, time range, parser version, and policy revision. Access to report data also needs a retention limit and authorization review because report contents expose sending infrastructure and message-flow metadata.
Promotion criteria should be written before the first enforcement change: the observation window, minimum volume, tolerated unexplained-failure rate, required approvers, rollback condition, and maximum time allowed for report-pipeline silence. The standard does not prescribe those organizational limits. Publishing invented universal thresholds would turn local risk tolerance into false protocol guidance.
Do not automate past ambiguity. Unknown traffic should enter a review queue with an owner and disposition; it should not be silently labeled malicious merely because alignment fails. Conversely, an endlessly renewed exception defeats the purpose of enforcement. Expire exceptions, attach evidence, and make renewal an auditable decision.
The final decision rule
Choose none while the inventory or telemetry is incomplete, quarantine when known legitimate streams align and a reversible sample is justified, and reject when reconciled evidence supports enforcement across the full intended scope. During a registrar API migration, freeze promotion whenever desired state and authoritative DNS differ, even if mail metrics look healthy.
This method deliberately stops keeping raw reports after the approved forensic window. The system retains normalized counts, hashes, policy revisions, approvals, and deletion events, but accepts that future parser changes cannot recover discarded fields. That trade-off bounds storage and sensitive-data exposure while leaving a defensible audit trail of why each policy transition occurred.
Further reading
- RFC 7489, Domain-based Message Authentication, Reporting, and Conformance (DMARC): https://datatracker.ietf.org/doc/html/rfc7489
Top comments (0)