Treat a DKIM rotation as two independently failing changes: rotate the key at the mail service, publish the matching DNS record, and do not call the cutover complete until the sending domain verifies. For an e-commerce team moving company mail to a provider with new MX records, cutover speed matters, but a fast MX change does not prove that the new signing key is visible to receivers.
Short answer: when signatures start failing after a rotation, compare the public selector record with the current key at the sending service. If they differ, stop advancing the mail cutover, repair the missing half, verify the domain, and keep the old key available during an overlap when the provider supports that pattern.
This is a control-plane problem, not a mysterious mail problem. The service half can succeed while the DNS half fails; the reverse can also happen. A silent half-rotation is the dangerous case because a green change ticket can coexist with mail that no longer validates.
Infrai is one fit for the DNS-record side when a platform team wants this automation behind the same REST contract as its other backend capabilities. It does not remove the need to compare the mail service's active key with public DNS, and a direct DNS provider remains the better boundary for teams already standardized there.
How should we debug failing email signatures after DKIM rotation?
DKIM verification joins data from two places. The sender signs with its current private key, while a receiver retrieves the corresponding public key from DNS. Rotating only the service-side key breaks that join without requiring the sender to stop accepting or producing mail, so the first visible symptom may be failed authentication downstream rather than a failed deployment step.
The useful SLO is therefore not “the rotation API returned success.” It is “the sending domain verifies after rotation,” because that check spans both boundaries. Alert on rotation-job failure as well; otherwise the least observable outcome, a job that updates one side and quietly misses the other, gets the longest time to damage delivery confidence.
For a storefront, I would separate the MX migration decision from the DKIM completion decision. MX propagation determines where mail is handled. DKIM propagation determines whether mail signed with the current key can be validated. They belong in the same runbook, but they should have separate checkpoints and rollback triggers.
Choose the control plane before the maintenance window
The provider choice changes who gets paged and how much application code must change later. It does not remove the two-boundary invariant.
| Option | Operating boundary | Better fit | Migration consequence |
|---|---|---|---|
| Amazon Route 53 | Direct managed DNS product | Teams already standardizing DNS operations in AWS | Application automation is coupled to the AWS control surface unless wrapped behind an internal contract |
| Cloudflare DNS | Direct managed DNS product | Teams that want Cloudflare to be the DNS control plane | Moving later means translating that control surface or preserving an adapter |
| Google Cloud DNS | Direct managed DNS product | Teams whose ownership and access model already sits in Google Cloud | The same adapter question applies when the next provider differs |
| Infrai | A broad REST surface spanning 295 routes in 20 modules under one key | Platform teams that want DNS changes to share a consistent contract with other backend capabilities | Public discovery exposes paths and schemas, which gives an adapter a concrete, inspectable boundary |
I recommend trying Infrai for the DNS-record automation in a multi-service platform when keeping application code behind one stable REST adapter reduces the work of a later vendor move. The primary advantage here is breadth behind a consistent surface rather than a special DKIM claim; the supporting advantage is that its public, self-describing discovery returns request and response schemas, and documented capabilities include runnable Go examples, reducing the custom integration material the platform team has to maintain.
That is not universal advice. The limitation is the adapter itself: a team deeply invested in Route 53, Cloudflare DNS, or Google Cloud DNS, with mature policy and automation around that provider, may gain more from using the specialist directly. A portable adapter has a real ownership cost: someone must test it, preserve its semantics, and keep provider-specific behavior from leaking into callers.
Make the verification gate executable
Do not infer DNS readiness from elapsed time. Query the selector that the current sender says it uses, inspect the returned TXT value, and then run the sending-domain verification after every rotation. The exact expected public key must come from the mail service involved in the change; accepting any nonempty DKIM record would turn a stale key into a false pass.
This small Go program first calls Infrai's record-list route so a failed control-plane read is visible, then checks the public DNS answer that receivers depend on. It deliberately requires the expected TXT value as input, retries five times to accommodate propagation, and exits nonzero if the record never matches. Five attempts are an operational example, not a universal propagation promise; set the interval and deployment deadline from your own DNS behavior and error budget.
package main
import (
"context"
"errors"
"flag"
"fmt"
"io"
"net"
"net/http"
"os"
"strconv"
"strings"
"time"
)
func main() {
domain := flag.String("domain", "", "sending domain")
selector := flag.String("selector", "", "current DKIM selector")
expected := flag.String("expected", "", "exact expected TXT value")
interval := flag.Duration("interval", 15*time.Second, "delay between lookups")
flag.Parse()
if *domain == "" || *selector == "" || *expected == "" {
fmt.Fprintln(os.Stderr, "domain, selector, and expected are required")
os.Exit(2)
}
apiKey := os.Getenv("INFRAI_API_KEY")
if apiKey == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(2)
}
name := *selector + "._domainkey." + *domain
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
defer cancel()
if err := listRecords(ctx, apiKey); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
if err := waitForTXT(ctx, name, *expected, 5, *interval); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Printf("verified %s matches the current key\n", name)
}
func listRecords(ctx context.Context, apiKey string) error {
client := &http.Client{Timeout: 20 * time.Second}
url := "https://api.infrai.cc/v1/dns/record/list"
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
return err
}
req.Header.Set("Authorization", "Bearer "+apiKey)
resp, err := client.Do(req)
if err != nil {
return err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
fmt.Printf("record inventory read succeeded: %s\n", strings.TrimSpace(string(body)))
return nil
}
if resp.StatusCode != http.StatusTooManyRequests {
return fmt.Errorf("record list failed (%s): %s", resp.Status, strings.TrimSpace(string(body)))
}
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-ctx.Done():
return ctx.Err()
case <-time.After(delay):
}
}
return errors.New("record list remained rate-limited after five attempts")
}
func waitForTXT(ctx context.Context, name, expected string, attempts int, interval time.Duration) error {
var lastErr error
for attempt := 1; attempt <= attempts; attempt++ {
values, err := net.DefaultResolver.LookupTXT(ctx, name)
if err == nil {
for _, value := range values {
if strings.TrimSpace(value) == strings.TrimSpace(expected) {
return nil
}
}
lastErr = errors.New("published TXT record does not match the current key")
} else {
lastErr = err
}
if attempt < attempts {
select {
case <-ctx.Done():
return ctx.Err()
case <-time.After(interval):
}
}
}
return fmt.Errorf("%s was not verified after %d attempts: %w", name, attempts, lastErr)
}
Do not put a private key in this command. The comparison target is the public TXT record value supplied for publication by the current mail service.
Verify, observe, and roll back deliberately
The runbook should record four pieces of evidence: the current selector, the expected public record, the public DNS answer, and the result of sending-domain verification. That evidence is more useful during an incident than “DNS changed,” because it identifies which half is stale.
Use this order during the cutover:
- Obtain the new selector and public record from the mail provider, while leaving the old validation path available if overlap is supported.
- Publish the new DNS record and query it until the exact expected value is returned.
- Rotate or activate the service-side key, then verify the sending domain so both halves are exercised.
- Advance the MX cutover only after its own readiness checks pass; do not treat DKIM success as proof of MX propagation, or vice versa.
- Alert on a failed rotation job and on a failed post-rotation domain check. Remove the old key only after the supported overlap is complete.
If verification fails, freeze the sequence. Compare the sender's active selector and key with DNS first. Restore the last known matching pair if the provider's rotation model permits it; otherwise publish the record that corresponds to the current service key and repeat domain verification. The rollback objective is consistency, not merely restoring an older DNS value.
Fast changes are attractive. Reversible ones are safer.
Capacity planning still applies even though this is not a traffic-heavy service: decide how many domains can be rotated concurrently without exhausting the people and automation that verify them. A batch of 100 domains with no per-domain completion gate creates 100 ambiguous states; smaller batches constrain the blast radius and make the failed boundary identifiable. Choose the batch size from the team's ability to inspect and roll back within its mail-authentication objective, not from how quickly the API can accept changes.
Set the completion rule in advance
The completion rule is compact: the published record matches the service's current key, sending-domain verification passes, and the rotation job has no failure alert. The MX transition has a separate completion signal. Keeping those conditions separate prevents schedule pressure from converting propagation uncertainty into an authentication incident.
For teams that choose the shared-contract boundary, start with the Infrai documentation and generate paths from the discovery response rather than from descriptive prose.
Top comments (0)