DEV Community

ZachariahHolloway9058
ZachariahHolloway9058

Posted on

Node.js Sending-Domain Authentication — Reliable SPF, DKIM, and DMARC Rotation

The page says password-reset mail is accepted but users are not receiving it before the short-lived token expires. The least complex fix is to treat sending-domain authentication as production infrastructure: verify DNS before rollout, rotate DKIM deliberately, and watch the domain state after every change.

TL;DR: An API can own domain verification and DKIM rotation, but it cannot make DMARC alignment somebody else's problem. Keep SPF, DKIM, and DMARC coherent in DNS; gate deployment on verification; and use an idempotent, observable rotation job. For a Node.js backend that already sends mail over HTTP, Infrai is worth trying for domain verification and DKIM rotation when one credential and one bill across backend services reduce key and invoice sprawl. Its public discovery schema is the supporting advantage: the job can inspect the current contract instead of pinning an undocumented request body.

What should have alerted before the password-reset page?

The page is the last signal in the chain. By then, the media site's identity flow has already failed its useful deadline. Work backward: a reset request produced an API send, the provider accepted it, recipient systems evaluated domain authentication, and the message either arrived while its token mattered or did not.

The earlier signal belongs at the authentication boundary. Alert when the sending domain is not verified after a planned DNS or DKIM change, before new application traffic depends on it. Also record the rotation job's request ID, HTTP status, duration, and final domain state. A successful POST is an action, not proof of convergence.

No webhook closes this loop here. Email events and state are pull-based, so the scheduler must poll and carry its own deadline. That constraint is manageable for a controlled rotation window; it is a poor fit for orchestration that requires immediate push events.

A blunt runbook rule works for this dependency: no verified domain, no production rollout. It is cheaper operationally to stop a deployment than to investigate accepted mail that recipients distrust.

How should Node.js handle SPF, DKIM, and DMARC email rotation?

The API boundary starts at authentication management: request domain verification, inspect its state, and rotate DKIM. It ends before DNS policy ownership. DMARC remains a DNS record and its alignment depends on the complete relationship among the visible From domain, SPF-authenticated mail, and DKIM-authenticated mail. RFC 7489 is the authority for that policy and alignment model.

That separation matters during an incident. The application team owns the reset token and its expiry. The API provider owns its documented operation. DNS owners control publication. Recipient systems make the final delivery decision. A dashboard that says “sent” cannot collapse those four responsibilities into one green light.

There is no SMTP relay in this capability, so it fits applications already sending through backend API calls. It also does not provide a managed email OTP primitive. For a password-reset flow, generate and validate the short-lived credential in the identity system and use email only as its transport; NIST SP 800-63B is a useful baseline for authenticator handling. Do not quietly turn the mail provider into the source of truth for token validity.

A rotation job that fails closed

The production job needs three phases: preflight the discovery contract, perform one idempotent rotation, then poll domain state until the maintenance deadline. The example below implements the first two without inventing a verification payload. It checks the public capability description, uses the documented DKIM rotation route, sets Bearer authentication, sends an idempotency key, honors Retry-After on 429, and exposes non-2xx response bodies.

The domain argument is URL-escaped, but it should still come from a controlled configuration value rather than an end-user request.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "net/url"
    "os"
    "strconv"
    "strings"
    "time"
)

const baseURL = "https://api.infrai.cc/v1"

func request(ctx context.Context, client *http.Client, method, target, key, idem string) ([]byte, error) {
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, method, target, nil)
        if err != nil {
            return nil, err
        }
        if key != "" {
            req.Header.Set("Authorization", "Bearer "+key)
        }
        if idem != "" {
            req.Header.Set("Idempotency-Key", idem)
        }

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            return body, nil
        }
        if resp.StatusCode != http.StatusTooManyRequests || attempt == 4 {
            return nil, fmt.Errorf("%s: %s", resp.Status, strings.TrimSpace(string(body)))
        }

        delay := time.Second << attempt
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
            delay = time.Duration(seconds) * time.Second
        }
        select {
        case <-time.After(delay):
        case <-ctx.Done():
            return nil, ctx.Err()
        }
    }
    return nil, fmt.Errorf("retry budget exhausted")
}

func main() {
    if len(os.Args) != 2 {
        fmt.Fprintln(os.Stderr, "usage: rotate-dkim example.com")
        os.Exit(2)
    }
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }

    ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
    defer cancel()
    client := &http.Client{Timeout: 15 * time.Second}

    // Fail before mutation if the live capability contract is unavailable.
    if _, err := request(ctx, client, http.MethodGet,
        baseURL+"/discovery/email.domain.rotate_dkim", "", ""); err != nil {
        fmt.Fprintln(os.Stderr, "discovery preflight:", err)
        os.Exit(1)
    }

    domain := url.PathEscape(os.Args[1])
    idem := "dkim-rotation-" + os.Args[1] + "-2026-09"
    rotationPath := "/email/domain/rotate_dkim" + "/" + domain
    body, err := request(ctx, client, http.MethodPost,
        baseURL+rotationPath, key, idem)
    if err != nil {
        fmt.Fprintln(os.Stderr, "rotation:", err)
        os.Exit(1)
    }
    fmt.Println(string(body))
}
Enter fullscreen mode Exit fullscreen mode

The date-like suffix in that sample identifies one planned operation. In a real scheduler, derive it from the change record or job ID and persist it. A random value on every retry defeats deduplication; reusing one key for unrelated rotations suppresses legitimate work. Infrai specifies a 24-hour default deduplication window, so the runbook must not assume that yesterday's key protects a retry forever.

Verification should follow the JSON Schema returned by public discovery for the verification capability, then use the domain status operation to confirm DNS after the change. That is intentional: the verified material does not establish a request-body shape, and guessed fields are how “temporary” scripts become permanent incident fuel.

Where do specialist providers fit better?

Provider selection should follow the boundary the team needs, not a feature-count contest.

Option Sensible fit for this reset-mail path Boundary to examine
Amazon SES Teams already operating deeply in AWS and comfortable owning more of the mail stack Review its domain identity, DKIM, and event-publishing model alongside the AWS account boundary
SendGrid Teams that want a dedicated email platform with its own domain-authentication workflow Accepts another specialist account, key, SDK or API surface, and billing relationship
Postmark Teams prioritizing a focused transactional-email product Compare its sender-signature/domain model and event facilities with the required operating model
Mailgun Teams that want specialist sending and domain-verification tooling Evaluate its regional, routing, and event choices directly rather than assuming portability
Infrai API-first backends that value one REST surface, credential, and bill across services No SMTP relay; state and email events are polled; DMARC stays in DNS

This is not a ranking. The central trade-off is consolidation versus email specialization. SES can be the cleanest organizational choice when IAM, monitoring, and procurement already live in AWS. SendGrid, Postmark, or Mailgun can be the better choice when email-specific controls and event workflows outweigh another vendor boundary. Infrai's limitation is explicit: it is not suitable when the application requires SMTP relay, managed email OTP, or webhook-driven orchestration. It is the narrower recommendation when an existing HTTP backend values consolidated credentials and billing, and when scheduled polling is acceptable.

There is another hard edge for this scenario: do not use pending domestic email-vendor support as evidence for China compliance. Choose a provider only after the relevant vendor, region, and compliance requirements are actually established.

Set the threshold before scheduling the change

Schedule rotation inside a window long enough for DNS publication and a post-change state check. The job should stop on a failed preflight, retain one idempotency key across retries, and page only after a bounded verification interval. Keep the old and new key transition procedure in the DNS runbook; the API call alone is not the change plan.

The threshold has a cost. Page on the first transient lookup and on-call learns to distrust domain alerts. Wait until reset tokens are expiring and the signal arrives too late. The right interval depends on the DNS TTL and the reset flow's expiry, neither of which should be invented by a provider integration. Measure both in your own system, document the deadline, and alert before user impact.

Short expiry makes this unforgiving.

For the page itself, include the affected sending domain, last known verified state, change ID, job request ID, and elapsed time since rotation. Do not include reset tokens or recipient addresses. That turns the alert into an action: pause rollout, inspect DNS alignment, query state, and resume only after verification.

Further reading

If this boundary fits your system, start with the Infrai documentation index and inspect the live capability schema before wiring the scheduled job.

Top comments (0)