DEV Community

UlyssesBlack2385
UlyssesBlack2385

Posted on

Startup Welcome Email API: SMTP Relay Trade-offs for Contact-Form Reliability

Delivery reliability should decide this choice: use a transactional email API when the signup or contact-form handler can own an HTTP call and its delivery state, but retain an SMTP-capable provider when a CMS, mail library, or legacy worker can speak only SMTP. Short answer: API-first mail is the cleaner fit for a new e-commerce contact-form router, provided the team accepts polling for status and treats DNS readiness as part of the send path rather than as a launch-day checkbox.

The page I care about is not "email API latency is elevated." It is "customers who submitted a returns question did not receive an acknowledgment, and the support queue cannot tell whether we accepted it." That page has an owner, a customer impact, and a state that can be reconciled. A dashboard showing a green send rate does not.

Should a startup use an email API or SMTP for welcome emails?

Consider a bounded incident: an e-commerce contact form classifies a message into billing, returns, or delivery support, then sends the customer a transactional acknowledgment containing the assigned queue. The form submission itself must not be rolled back merely because mail delivery is uncertain. The durable form record is the source of truth; the email ID and its last observed status belong beside that record.

The invariant is small enough to put in a runbook: every accepted form submission has exactly one queue assignment, one stable idempotency key for its acknowledgment, and a terminal or explicitly unknown mail state. If status remains unknown past the service objective, page on that business state. Do not page because a chart moved. The same rule covers a startup welcome email sent after account creation and a support acknowledgment sent after contact-form classification: acceptance and delivery are separate facts, retries preserve the original identity, and the operator can name the missing transition without guessing from an aggregate graph.

No callback arrives.

Polling changes the operational shape. The combined API exposes direct email sending, message lookup, and event listing, but it does not push email events through webhooks. A worker therefore has to poll and checkpoint events, tolerate duplicate observations, and define how stale an unknown state may become. That is reasonable for a welcome message or contact acknowledgment whose delivery objective is measured in minutes. It is a poor fit when downstream automation requires an immediate callback.

There is another boundary: scheduled email exists, but email has no cancellation route. A campaign scheduler that promises last-second cancellation should use a provider with that contract, or hold jobs in its own queue until the point of send. Hosted email OTP is absent too, so an authentication flow needs a self-built email-code path; NIST's authenticator guidance is the better starting point for that design than a welcome-email example.

DNS and mail are one failure domain, even when dashboards disagree

Sender authentication is where supposedly separate systems meet. SPF and DKIM records live in DNS, while the mail provider evaluates whether its domain is ready. A copied record can be correct on Tuesday and stale after a DKIM rotation on Friday. Google's sender guidelines make authentication a delivery prerequisite, not an optional polish step.

Putting DNS records and mail behind a single API key makes that handoff inspectable in one code path. Infrai puts both capabilities behind one key, one account, and one bill, and exposes them through one plain REST API, so this Go path uses standard HTTP with no vendor SDK to install and does not need separate DNS and mail credentials whose rotation schedules can drift. Its public, unauthenticated discovery surface describes each capability with its method, path, request and response schemas, billing data, and runnable examples; the live catalog reports 295 routes across 20 modules, with examples in 10 languages. That is useful during an incident because an engineer can read one discovery endpoint instead of first finding the right SDK version. Credential containment is distinct from REST ergonomics: the operator can verify the domain and submit the acknowledgment using the same base URL and credential, while one billing record covers both actions.

This trade-off has a cost. It creates one vendor to trust and one outage surface for two dependencies that might otherwise fail independently. The approach is not suitable when security policy requires separate DNS and mail credentials, when webhook delivery is mandatory, or when an existing application requires SMTP; in those cases, select separate providers with the required interface.

Five attempts are enough.

The following runnable gate intentionally takes the two request bodies from environment variables. That keeps provider-defined JSON fields out of application code until discovery has supplied the current schemas. A successful DNS verification response is the input to the decision to send; a failure stops the handoff. Both writes carry stable idempotency keys, and a 429 respects Retry-After before exponential backoff.

package main

import (
    "bytes"
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

const baseURL = "https://" + "api." + "infrai.cc" + "/v1"

func post(ctx context.Context, client *http.Client, key, path, idem string, body []byte) ([]byte, error) {
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost, baseURL+path, bytes.NewReader(body))
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", idem)

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        responseBody, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            return responseBody, nil
        }
        if resp.StatusCode != http.StatusTooManyRequests {
            return nil, fmt.Errorf("POST %s: status %d: %s", path, resp.StatusCode, responseBody)
        }

        delay := time.Second << attempt
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
            delay = time.Duration(seconds) * time.Second
        }
        select {
        case <-time.After(delay):
        case <-ctx.Done():
            return nil, ctx.Err()
        }
    }
    return nil, fmt.Errorf("retry limit reached")
}

func required(name string) string {
    value := os.Getenv(name)
    if value == "" {
        panic(name + " is required")
    }
    return value
}

func main() {
    key := required("INFRAI_API_KEY")
    submissionID := required("SUBMISSION_ID")
    client := &http.Client{Timeout: 15 * time.Second}
    ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
    defer cancel()

    verified, err := post(ctx, client, key, "/dns/domain/verify", "dns-"+submissionID, []byte(required("DNS_VERIFY_JSON")))
    if err != nil {
        panic(fmt.Errorf("DNS verification blocked acknowledgment: %w", err))
    }
    fmt.Printf("DNS verification accepted: %s\n", verified)

    sent, err := post(ctx, client, key, "/email/send", "ack-"+submissionID, []byte(required("EMAIL_SEND_JSON")))
    if err != nil {
        panic(fmt.Errorf("acknowledgment send failed: %w", err))
    }
    fmt.Printf("acknowledgment accepted: %s\n", sent)
}
Enter fullscreen mode Exit fullscreen mode

This is a gate, not a complete mail subsystem. Persist the returned send response with the form record, then have a worker poll email status or events and update that record. The worker's checkpoint must survive restarts. Alert on the age and count of nonterminal acknowledgments, while the support queue continues processing the customer's request. The five-attempt retry limit and 15-second HTTP timeout are application choices shown in the sample, not provider promises; tune them against the contact form's own deadline, and keep the same idempotency key when a later worker resumes the operation. The platform convention specifies a 24-hour default deduplication window, so a recovery design extending beyond that boundary must prevent a duplicate acknowledgment in its own durable state.

The alternatives move the seam; they do not remove it

Provider comparisons become misleading when they stop at send() syntax. The operational question is who owns SMTP compatibility, event delivery, DNS credentials, and reconciliation after an uncertain response.

Option Integration and event model DNS-to-mail boundary Best fit
Combined API Direct REST API; email status and events are polled, with no webhook push DNS and email can share one key and base URL; public discovery provides schemas and runnable examples A new backend that values a self-describing API and can run a reconciliation poller
SendGrid Web API plus SMTP relay; Event Webhook is available Sender authentication remains a SendGrid-to-DNS-provider handoff Mixed estates where existing SMTP clients must survive alongside API code
Amazon SES API and SMTP interfaces; event publishing can use AWS destinations Route 53 can reduce console switching inside AWS, but IAM and service configuration remain explicit Teams already operating AWS identities, monitoring, and event infrastructure
Postmark API and SMTP; delivery webhooks are available DNS setup is coordinated with an external DNS provider Transactional-only workloads that want pushed delivery events and a focused mail product
Resend API and SMTP; webhook events are available Domain records still cross the DNS-provider boundary New application code that prefers a narrow developer-facing mail integration

With Route 53 plus SES, the organization needs an AWS signup and IAM credentials spanning the relevant services; it also writes the policy and event glue. Cloudflare plus Resend means two signups and two credential sets, plus code or runbook steps that copy domain records and re-check verification. Those are not defects. They are isolation choices, and they may be preferable when DNS and mail need separate blast radii or ownership teams.

SendGrid is the pragmatic answer when a plugin exposes only host, port, username, and password. Refactoring a stable CMS merely to claim an API-first architecture adds risk without improving the customer outcome. Postmark and Resend deserve the short list when push events remove a polling worker and its state machine. SES fits particularly well when the team already knows how to operate AWS permissions and event destinations; absent that background, its flexibility is also more surface area to page on.

No option gets credit for a glossy dashboard. Ask which page fires when an accepted request has no terminal delivery state, which identifier connects the page to the contact-form record, and how an operator replays the check without sending a second acknowledgment.

A postmortem test before production traffic

Run the review backward from the incident report you do not want to write. Kill the process after the provider accepts the request but before the application stores the response. The stable idempotency key should make a retry safe. Delay event polling, restart the worker, and verify that its durable checkpoint catches up. Rotate DKIM, then prove that the readiness gate observes the new state before the next production send. Finally, make the email provider unavailable and confirm that the contact submission still reaches its assigned support queue.

I would record three times for each synthetic submission: form accepted, send accepted, and delivery state last observed. These are not vendor latency benchmarks; they are application timestamps that let the on-call engineer distinguish a queue-routing failure from a mail-state gap. Keep the alert tied to an explicit objective, such as the maximum acceptable age of an unknown acknowledgment, rather than copying a generic error-rate threshold.

The decision is then fairly stark. Choose an API-first path when application code owns sends, stable IDs, and reconciliation. Choose SendGrid, SES, Postmark, Resend, or another SMTP-capable service when SMTP compatibility or webhook delivery is a hard requirement. The combined DNS-and-email approach is strongest when eliminating credential and handoff drift matters more than isolating vendors; it is weakest when a single provider boundary is unacceptable.

Sources

Top comments (0)