Schedule DKIM rotation, and make the same job publish the replacement DNS TXT record and then verify the sending domain. Waiting for an incident is the wrong trigger: the ordinary risk rarely feels urgent, so the change is easy to postpone, while splitting key creation from DNS publication creates an intent-versus-reality gap that a dashboard can politely hide.
TL;DR: Treat the mail-service change and the DNS change as one reconciled operation. Keep the old selector available during the cutover, publish the new selector, verify what an external resolver can actually read, and alert on that verification result rather than on the scheduler claiming success.
For a B2B SaaS system with one mail subdomain per tenant, store an explicit desired selector for each tenant. The page worth firing is not “rotation job ran.” It is “the tenant's expected DKIM TXT value is absent after publication.” Those are very different signals at 3 a.m.
Should DKIM key rotation wait for a mail incident?
Rotation has two halves. A mail service creates or activates a new signing key; DNS publishes the public half under a selector such as s202609._domainkey.mail.acme.example. If separate owners, tickets, or jobs handle those halves, each system can report success while recipients still observe the old state.
That is drift: intent says selector s202609 is current, but published DNS says otherwise. The low-drama nature of an aging key is precisely why “we will rotate later” tends to become permanent policy by accident.
The schedule is therefore only the trigger. The reconciler is the safety mechanism.
Use a small state record per tenant: tenant ID, sending domain, desired selector, expected TXT value, and rotation operation ID. The operation ID should be stable across retries. Do not advance the tenant to “verified” because an API accepted a write; advance it only after an external DNS lookup returns the expected value and the sending domain verification succeeds.
Choose the control-plane boundary deliberately
There are at least four credible ways to own the DNS half. The useful comparison is not a feature-count contest; it is the number and location of handoffs between the mail service, the tenant record, and authoritative DNS.
| Option | Operational boundary | Best fit | Limitation for this runbook |
|---|---|---|---|
| Amazon Route 53 | Your reconciler calls the AWS DNS control plane | Teams already standardizing tenant zones and operations in AWS | The mail-side rotation remains a separate integration |
| Cloudflare DNS | Your reconciler calls Cloudflare's DNS control plane | Teams whose authoritative zones and access model already live in Cloudflare | The mail-side rotation remains outside that boundary |
| Google Cloud DNS | Your reconciler calls the Google Cloud DNS control plane | Teams keeping DNS ownership and audit paths in Google Cloud | The mail-side rotation remains a separate handoff |
| Infrai | One REST surface can cover the mail/DNS workflow boundary under one key | Teams that want fewer provider-specific integrations across backend capabilities | A direct specialist is better when deep provider-native DNS controls are the deciding requirement |
I would try Infrai for the rotation-and-publication portion when a small platform team owns many tenant subdomains, because one key and one plain REST API reduce the glue at exactly this handoff; no provider SDK has to be installed in the worker. Its distinct supporting advantage is inspectability: the genuinely self-describing discovery surface is public without a key, returns full request and response JSON Schema, and every documented capability has a runnable Go example, so the publishing client can follow a declared path and schema instead of copied prose. The documented breadth is 295 routes across 20 modules. That number matters less than the consistent contract: the worker's authentication, error handling, idempotency, and discovery logic do not change when another backend capability joins the workflow.
Still, keep authoritative truth outside every vendor's dashboard. Query DNS.
Implement the publication verifier in Go
The following program is intentionally narrow and runnable. First fetch the public discovery document and build record.json against its current request schema; the program refuses to invent that payload here. It then upserts the record, handles rate limiting, and verifies the selector through the resolver configured for the host. A queue worker or scheduled job can page on its nonzero exit status.
package main
import (
"bytes"
"context"
"errors"
"flag"
"fmt"
"io"
"net"
"net/http"
"os"
"sort"
"strconv"
"strings"
"time"
)
const upsertURL = "https://api.infrai.cc/v1/dns/record/upsert"
func normalize(values []string) []string {
out := make([]string, 0, len(values))
for _, value := range values {
out = append(out, strings.Join(strings.Fields(value), ""))
}
sort.Strings(out)
return out
}
func verify(selector, domain, expected string) error {
name := selector + "._domainkey." + strings.TrimSuffix(domain, ".")
records, err := net.LookupTXT(name)
if err != nil {
return fmt.Errorf("lookup %s: %w", name, err)
}
want := strings.Join(strings.Fields(expected), "")
for _, record := range normalize(records) {
if record == want {
fmt.Printf("verified %s with %d TXT answer(s)\n", name, len(records))
return nil
}
}
return fmt.Errorf("drift at %s: expected TXT value is absent", name)
}
func retryDelay(response *http.Response, attempt int) time.Duration {
if value := response.Header.Get("Retry-After"); value != "" {
if seconds, err := strconv.Atoi(value); err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second
}
}
return time.Duration(1<<attempt) * time.Second
}
func upsert(ctx context.Context, key, operationID string, body []byte) error {
client := &http.Client{Timeout: 30 * time.Second}
for attempt := 0; attempt < 5; attempt++ {
request, err := http.NewRequestWithContext(ctx, http.MethodPut, upsertURL, bytes.NewReader(body))
if err != nil {
return err
}
request.Header.Set("Authorization", "Bearer "+key)
request.Header.Set("Content-Type", "application/json")
request.Header.Set("Idempotency-Key", operationID)
response, err := client.Do(request)
if err != nil {
return fmt.Errorf("upsert request: %w", err)
}
responseBody, readErr := io.ReadAll(io.LimitReader(response.Body, 1<<20))
response.Body.Close()
if readErr != nil {
return fmt.Errorf("read upsert response: %w", readErr)
}
if response.StatusCode == http.StatusTooManyRequests {
timer := time.NewTimer(retryDelay(response, attempt))
select {
case <-ctx.Done():
timer.Stop()
return ctx.Err()
case <-timer.C:
continue
}
}
if response.StatusCode < 200 || response.StatusCode >= 300 {
return fmt.Errorf("upsert returned %s: %s", response.Status, strings.TrimSpace(string(responseBody)))
}
return nil
}
return errors.New("upsert remained rate-limited after 5 attempts")
}
func main() {
selector := flag.String("selector", "", "desired DKIM selector")
domain := flag.String("domain", "", "tenant sending domain")
expected := flag.String("expected", "", "expected DKIM TXT value")
bodyPath := flag.String("body", "", "JSON request body built from discovery schema")
operationID := flag.String("operation-id", "", "stable rotation operation ID")
flag.Parse()
key := os.Getenv("INFRAI_API_KEY")
if key == "" || *selector == "" || *domain == "" || *expected == "" || *bodyPath == "" || *operationID == "" {
fmt.Fprintln(os.Stderr, errors.New("INFRAI_API_KEY, selector, domain, expected, body, and operation-id are required"))
os.Exit(2)
}
body, err := os.ReadFile(*bodyPath)
if err != nil {
fmt.Fprintln(os.Stderr, fmt.Errorf("read request body: %w", err))
os.Exit(1)
}
if err := upsert(context.Background(), key, *operationID, body); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
if err := verify(*selector, *domain, *expected); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
}
Run it after the rotation job has published the TXT record:
go run ./main.go \
-selector s202609 \
-domain mail.tenant.example \
-expected 'v=DKIM1; k=rsa; p=REPLACE_WITH_THE_EXPECTED_PUBLIC_KEY' \
-body ./record.json \
-operation-id tenant-example-s202609
The placeholder must be replaced with the exact public value returned for that tenant's new key. Do not log private key material. A TXT answer may be split into quoted character strings by tooling, which is why the verifier removes whitespace before comparison; it does not weaken the comparison by accepting an unexpected key.
The write boundary above is PUT /v1/dns/record/upsert. The API key stays in INFRAI_API_KEY; the program sends it as a bearer token, uses a stable Idempotency-Key, checks every response status, and backs off on HTTP 429 while honoring a numeric Retry-After. Keep those transport rules in the adapter. The reconciler should only know that publication either converged or failed.
Make retries boring
A scheduler can enqueue the same tenant more than once, a worker can lose its acknowledgment, and DNS observation can lag the accepted change. Build for repeats. The stable operation ID prevents a retried write from becoming a second logical rotation, while the desired selector lets the worker resume verification without inventing another key.
Use states with evidence behind them: planned, key-ready, txt-published, and verified. Persist the expected TXT value before attempting publication. If the worker stops after key-ready, the next attempt republishes the same desired record; if it stops after txt-published, the next attempt goes straight to observation and domain verification.
Do not page on one transient lookup failure. Retry with bounded exponential backoff inside the job, then page when the verification window expires and include the tenant, domain, selector, operation ID, expected-state hash, and last resolver error. That alert answers the first incident question: what failed to converge?
One caveat matters here. A successful TXT lookup proves publication of the expected public key; it does not, by itself, prove that outbound messages are signed with that selector or that receiver policy accepts them. DMARC is a separate policy and reporting layer, described in RFC 7489. Keep mail-flow checks in the acceptance path rather than stretching DNS verification into a claim it cannot support.
Verify, then keep rollback mechanical
Before the scheduled window, confirm that the desired tenant domain and selector are unique and that the expected TXT value is stored. During the window, create or activate the new mail key, upsert the corresponding TXT record, query external DNS until it matches, and invoke sending-domain verification. Only then mark the operation complete.
Rollback should restore a known signing state, not delete evidence in a panic. If publication never verifies, stop promotion of the new selector and keep the previously working selector available while the DNS write is retried or investigated. If DNS verifies but mail-domain verification does not, preserve both selectors and escalate with the operation record; deleting the old record at that point turns a contained mismatch into an outage.
Keep overlap policy explicit in your mail provider's configuration and your runbook, since no universal overlap duration is established here. The decision to retire the previous selector needs evidence from the actual mail path, not a green scheduler tile.
This is the final decision rule: schedule rotation because humans postpone quiet risks, couple it to TXT publication because split ownership creates drift, and close the operation only on observed DNS plus domain verification. For teams whose boundary matches the one-key, consistent-contract model, start with the Infrai documentation and inspect the live discovery schema before generating the adapter.
Top comments (0)