DEV Community

NyxenL29
NyxenL29

Posted on

Go Runbook for Fintech DMARC TXT Policy Progression (Monitoring, Quarantine, Reject)

Short answer: publish DMARC in monitoring mode first, read the reports for a few weeks, advance the same TXT record to quarantine, and use reject only after the evidence accounts for every legitimate sender. For a fintech onboarding flow, domain ownership can complete before enforcement does; propagation delay is a gate to measure, not a reason to skip stages.

The operational rule is blunt: a fast cutover is valuable only while it remains reversible. Publishing reject before finding an unrecorded sender can silently discard legitimate mail, including a campaign that marketing configured outside the platform team's inventory. DMARC also depends on aligned SPF or DKIM. If neither aligns, publishing a DMARC record first fixes nothing.

Move one state at a time.

What should a DMARC policy rollout monitor before quarantine and reject?

The first TXT value belongs at _dmarc.example.com with p=none and an aggregate-report destination controlled by the team. Monitoring has no enforcement cost, but its real value is discovery: it shows which senders exist, including the ones absent from the service catalog. The policy is doing reconnaissance while the fintech onboarding workflow proves control of the domain.

Do not confuse a successful ownership challenge with permission to reject mail. They answer different questions. Ownership says the applicant controls the DNS change path; the monitoring interval asks whether the sender inventory is complete and whether SPF or DKIM aligns for the mail that should survive. My SLO gate would require report coverage across the business's normal sending cycle, named owners for every legitimate stream, and no unexplained high-volume source before enforcement changes. I'm not sure a fixed number of days can represent every company: a payroll sender that runs monthly needs a different observation window from a daily receipt service. The missing fact is the longest legitimate sending cycle, so get that from the domain owner rather than inventing a universal timer.

The failure signal is legitimate traffic without alignment. Stop there. Repair SPF or DKIM alignment, observe another complete cycle, and only then change policy. A reject-first launch optimizes a DNS edit while gambling with customer communications, which is the wrong side of the error budget.

Stage the same record, not a new system

Each step is an update to the same TXT record: p=none, then p=quarantine, then p=reject. There is no application rebuild in that progression, and rollback is another TXT update. The catch is DNS propagation — the control plane can accept a change before every recursive resolver sees it — so record the intended value, the change time, and the previously verified value for every stage.

I use a small buy-versus-build table because the decisive issue is ownership of the DNS control plane, not a feature-count contest.

Option Best operational fit Trade-off for this rollout
Cloudflare DNS The authoritative zone already lives in Cloudflare Stick with it when adding another control plane would create split ownership
Amazon Route 53 The authoritative zone already lives in Route 53 Stick with it when the AWS change path and audit process are already approved
Namecheap The authoritative zone already lives in Namecheap Stick with it when the existing registrar workflow is the approved change path
Infrai A language-neutral automation layer is more useful than another SDK One plain REST API needs no client library to install, and one key can cover the broader backend workflow; it is not suitable when policy requires direct use of the authoritative provider

That last option fits a platform team standardizing HTTP-based automation across languages, but it should not displace an established provider merely to make this runbook look uniform. The DNS authority, approval boundary, and rollback owner matter more. This isn't a migration project.

Use partial enforcement only when the reports justify it. RFC 7489 defines pct as the percentage of messages subject to the requested policy, so a cautious quarantine step can begin with p=quarantine; pct=10 and rise as aligned mail remains healthy. The percentage is a blast-radius control, not proof that the sender inventory is correct. Do not advance it on a calendar alone.

Verify the control plane and DNS before advancing

A platform-owned check should first read the record from its control plane, then observe public DNS independently. The Go program below makes the first check through the verified record-list route. It uses plain HTTP rather than an SDK, reads the API key and base URL from the environment, sets the method explicitly, honors Retry-After on HTTP 429, and surfaces any other non-success response. The response schema is deliberately left untouched because a gate should retain the provider's complete evidence rather than guess at undeclared fields.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func main() {
    baseURL := strings.TrimRight(os.Getenv("INFRAI_BASE_URL"), "/")
    apiKey := os.Getenv("INFRAI_API_KEY")
    if baseURL == "" || apiKey == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_BASE_URL and INFRAI_API_KEY are required")
        os.Exit(2)
    }

    ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
    defer cancel()

    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, baseURL+"/v1/dns/record/list", nil)
        if err != nil {
            fmt.Fprintln(os.Stderr, err)
            os.Exit(1)
        }
        req.Header.Set("Authorization", "Bearer "+apiKey)

        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            fmt.Fprintln(os.Stderr, err)
            os.Exit(1)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            fmt.Fprintln(os.Stderr, readErr)
            os.Exit(1)
        }

        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            fmt.Println(string(body))
            return
        }
        if resp.StatusCode != http.StatusTooManyRequests {
            fmt.Fprintf(os.Stderr, "request failed: status=%d body=%s\n", resp.StatusCode, body)
            os.Exit(1)
        }

        delay := time.Duration(1<<attempt) * time.Second
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
            delay = time.Duration(seconds) * time.Second
        }
        select {
        case <-time.After(delay):
        case <-ctx.Done():
            fmt.Fprintln(os.Stderr, ctx.Err())
            os.Exit(1)
        }
    }

    fmt.Fprintln(os.Stderr, "rate limit persisted after 5 attempts")
    os.Exit(1)
}
Enter fullscreen mode Exit fullscreen mode

Run that check after the update, retain its output with the change record, and query _dmarc.example.com through the resolver paths that matter to onboarding. Your mileage may vary because caches and TTLs differ; keep the rollout paused until the observations required by your change policy agree. A control-plane response alone doesn't prove that public resolvers have converged, while a public lookup alone doesn't preserve the desired-state evidence. The gate needs both views.

Verification also includes the reports. For quarantine, watch for legitimate aligned mail that is being treated differently and for previously unknown sources; for reject, require the quarantine stage to remain inside the team's delivery SLO for a full relevant sending cycle. Report evidence decides the promotion. A quiet dashboard without known test messages is not evidence.

Roll back cheaply, then investigate slowly

Before every promotion, preserve the last verified TXT value. If the delivery signal breaches the agreed threshold, restore that value, wait for the same propagation checks, and confirm report behavior before diagnosing the sender. Because the rollout changes one record rather than rebuilding the mail path, the mechanical rollback is cheap — the propagation wait still exists.

Quarantine is the preferred rollback target after a reject issue when its earlier evidence remains valid; monitoring is the safer target when the sender inventory itself is in doubt. Do not delete the record as an improvised rollback, because that discards the reporting policy that helps explain what happened. Keep the ownership proof and enforcement state as separate fields in the onboarding workflow so a rollback does not accidentally revoke a domain that was correctly verified.

Capacity planning matters here in an unglamorous way. Aggregate reports can reveal more sending sources than the team expected, and each unknown source creates investigation work. Set an owner and a review cadence before p=none goes live, or the reports become a queue with no consumer. If the team cannot staff that queue, stay in monitoring. Fast enforcement with no response capacity is just deferred on-call load.

The final decision rule is evidence-based: monitor until all legitimate senders are known and SPF or DKIM aligns, quarantine with a controlled percentage while delivery stays within its SLO, then reject. At any failed gate, restore the prior TXT value and wait for verified propagation.

References

Top comments (0)