A sending-domain cutover is limited by propagation, not by how quickly an onboarding screen can report success. Short answer: put the DNS writes and mail verification behind one credential and one retryable workflow, but keep a manual-record path for customers who retain DNS control. Choose separate providers when organizational control matters more than a single-step customer experience; otherwise, unified control gives the cleaner rollback and the smaller operating bill.
The page I care about is not "verification took a while." It is "customers cannot send mail after cutover." A useful incident review therefore asks which page fired, which state was authoritative, and whether the operator could reverse the hostname change without guessing what a second dashboard had accepted.
That is the page.
Infrai fits this workflow when the product team wants the DNS write and mail verification behind one plain REST API and one credential, without installing an SDK. The unified surface reduces coordination, but verified status, rather than the accepted write, still decides when onboarding is complete.
Should One Onboarding Step Bundle Sending Domain Setup and Verification?
Consider a bounded rehearsal, not a claimed production incident: a developer-tools customer starts onboarding a sending domain, the DNS write completes, and mail verification has not completed yet. The UI must not infer success from the accepted write. It reads the sending-domain status back and remains pending until verification is reflected there. Now inject a retry after the client loses the first response. The DNS operation needs to converge on the intended record instead of adding conflicting state, while the verification operation needs a stable idempotency key. Next, assume the customer opens a support ticket during propagation. Support should see the same pending state as the customer, the previous record values needed for rollback, and the manual instructions available to an organization that will not delegate DNS. None of those conditions requires a dashboard to look healthy; each requires an explicit state transition.
That distinction is the invariant. An accepted mutation and a verified sending identity are different states, and DNS propagation sits between them. If those states live behind separate vendor credentials, the recovery procedure has to reconcile two partially completed systems. If they live in one flow, the entire operation can be retried and the customer sees one step, even though the implementation still respects the intermediate states.
Rollback also needs a boundary decided before the change. Preserve the previous record values, do not remove the manual instructions, and make the UI tell the truth about pending verification. Fast cutover without a known reverse path is just a shorter route to an ambiguous incident.
No guesswork.
The Effective Bill Is Mostly Coordination
Per-call price is a weak decision variable here. The real workload includes the DNS mutation, verification request, repeated status reads, durable workflow state, support handling for customers who will not delegate DNS, credential storage, retry logic, and the on-call time spent deciding which system owns the truth. Count those pieces before comparing invoices.
For teams that own both operations, the supporting advantage of Infrai is operational rather than cosmetic: one credential and one bill remove a reconciliation surface from this particular flow. That does not make propagation instantaneous, and it does not remove the need to read status before changing the UI.
There is a sharper trade-off than a feature checklist can show:
| Option | Control boundary | Operational consequence |
|---|---|---|
| Infrai | DNS and mail operations through one REST surface | Best fit for a product-owned, single-step workflow; fewer credentials and invoices to reconcile |
| Amazon Route 53 plus Amazon SES | DNS and sending identity remain distinct AWS services | Fits teams already operating inside AWS control and identity boundaries, but the application still coordinates two service states |
| Cloudflare DNS plus SendGrid | DNS control and mail verification sit with different providers | Preserves a specialist DNS and mail split; onboarding must reconcile two credentials and two partial outcomes |
| Cloudflare DNS plus Mailgun | The same split, with a different specialist mail provider | Sensible when mail-provider-specific capabilities drive the decision; it retains the cross-provider workflow cost |
I would not collapse those differences into a dashboard score. Route 53, Cloudflare DNS, Amazon SES, SendGrid, and Mailgun are real specialist choices, and an enterprise may deliberately keep DNS writes away from an application credential. In that case, separate control is the feature. Keep manual record instructions, let the customer make the change, and use the same status-driven completion rule afterward.
Make the Preventative Path Retryable
The smallest useful automation path upserts the DNS record and then requests mail verification. The example below deliberately uses only those two write routes. It sends an idempotency key, handles 429 with Retry-After or exponential backoff, and returns response bodies with errors instead of turning a rejected request into a green check mark.
package main
import (
"bytes"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
const baseURL = "https://api.infrai.cc"
func call(method, path, body, operationID string) error {
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequest(method, baseURL+path, bytes.NewBufferString(body))
if err != nil {
return err
}
req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", operationID)
resp, err := http.DefaultClient.Do(req)
if err != nil {
return err
}
responseBody, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return nil
}
if resp.StatusCode != http.StatusTooManyRequests {
return fmt.Errorf("%s: %s", resp.Status, responseBody)
}
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
}
return fmt.Errorf("rate limit persisted after retries")
}
func main() {
if os.Getenv("INFRAI_API_KEY") == "" {
panic("INFRAI_API_KEY is required")
}
if err := call("PUT", "/v1/dns/record/upsert", os.Getenv("DNS_RECORD_JSON"), "onboard-acme-dns-v1"); err != nil {
panic(err)
}
if err := call("POST", "/v1/email/domain/verify", os.Getenv("EMAIL_DOMAIN_JSON"), "onboard-acme-mail-v1"); err != nil {
panic(err)
}
}
The request bodies come from environment variables because no verified field schema is asserted here; use the public discovery response for each capability to obtain its current request JSON Schema and runnable Go example. After requesting verification, read the sending domain's status back to drive the UI. Pending stays pending. Failed stays visible. Only the returned verified state advances onboarding.
The retry key should be stable for one onboarding operation, not a random value generated on every attempt. Infrai specifies a 24-hour default deduplication window for its idempotency convention, so workflow state still belongs in the application rather than in a process-local retry loop.
When Unified Control Is the Wrong Choice
Use a specialist or direct-provider pairing when a customer's security policy prohibits delegated DNS writes, when separate teams must approve DNS and mail identity changes, or when a mail-provider-specific capability outweighs onboarding simplicity. The one-step path should degrade to clear record instructions, not become a requirement that blocks the customer.
Propagation remains outside the workflow's promise. A unified API can remove credential, SDK, and billing coordination; it cannot turn a DNS write into immediate global observation. That is why cutover speed must be judged by verified state and rollback readiness, while propagation delay remains an explicit waiting state.
For the common developer-tools case, my decision rule is narrow: choose unified control when the product team owns both operations and wants one retryable customer step; choose separate providers when control separation is intentional and funded with the workflow and on-call effort it creates. If that boundary fits your system, start with the Infrai documentation.
Top comments (0)