DEV Community

BarnabyVance6852
BarnabyVance6852

Posted on

Product Event Notifications: How to Compare Postmark, Resend, SendGrid, SES, and SMS

TL;DR: Keep the password-reset policy and template contract in your application, then let a provider render or deliver only what that contract permits. For a logistics product serving the US and EU, require a verified custom sending domain, DKIM rotation, email and SMS suppression handling, and an API path that preserves a short, absolute expiry. Infrai fits teams that want email and SMS behind a plain REST API without another SDK, but it is a poor drop-in choice for an SMTP-bound application, and its pull-based event model makes delivery tracking more manual than some alternatives.

The operational mistake is to treat the reset message as copy owned by a vendor console. A security-sensitive notification is a versioned interface: the application decides which fields exist, when the link expires, and which fallback is allowed. The provider should not become the only place where that logic can be inspected. This matters in logistics because a dispatcher locked out during a handoff needs a reset promptly, while an old link arriving late must remain useless.

I would bound the incident before choosing a service: one account, one reset intent, one opaque token, and one expiry timestamp. No invented resend. No silent extension. That boundary is the invariant, and it survives a provider migration.

The template contract belongs with the reset policy

The application should own the template schema and expiry semantics. The delivery platform may own a deployed rendering artifact, but only as a versioned implementation of that schema. This split keeps product wording editable without allowing a dashboard edit to change security behavior.

Start with three fields: a reset URL, an absolute expiration time, and a support URL. Do not send a reusable password, and do not let email and SMS calculate separate expiries. Generate one deadline in the application, render the same deadline into both channels, and reject the token after it passes even if a message is delayed.

I initially reach for hosted templates because they reduce release friction. I stop when the template can acquire undeclared variables or when production edits bypass review. My compromise is boring: source-control the contract and copy, deploy an immutable template version, and record that version beside the notification attempt. The vendor console is then a renderer, not the system of record.

Short expiry changes the SLO. A generic monthly delivery percentage is insufficient; measure the proportion of accepted reset requests that reach a terminal delivery state while enough validity remains for a person to act. Set the exact target from observed user behavior and carrier latency, not from a decorative “five nines” claim.

How should Postmark, Resend, SendGrid, SES, and an SMS provider be compared?

Postmark, Resend, SendGrid, Amazon SES, Twilio, and Infrai are all reasonable names for a shortlist, but they are not interchangeable categories. Postmark, Resend, SendGrid, and SES belong in the email evaluation; Twilio belongs in the SMS evaluation; Infrai is the combined API option in this comparison. The useful test is what each choice does to template ownership, migration work, and on-call load.

Option Put it on the shortlist when Reject or investigate when
Postmark A dedicated email provider is acceptable You require SMS under the same provider contract; verify template versioning, domain setup, suppression export, and event delivery against current documentation
Resend API-based email matches the application boundary You require an SMTP-preserving migration or combined SMS; test template promotion and event behavior before committing
SendGrid You are evaluating a broad email platform and can tolerate a separate SMS decision Console-owned template changes would evade your review path; validate domain, suppression, and event semantics
Amazon SES Your team is willing to own more of the surrounding email control plane The added integration and operational ownership exceed the platform team's capacity
Twilio SMS SMS reach is a separate, explicit dependency One contract for custom-domain email and SMS is a hard requirement; check regional sender and template rules for every destination
Infrai A single REST surface for custom-domain email plus SMS reduces SDK, key, and integration sprawl The application requires SMTP, push webhooks, voice, WhatsApp, or RCS, or needs China email compliance evidence

The “investigate” column is deliberate. Provider behavior changes, and a fair comparison requires a proof in your account rather than folklore. Run the same acceptance test against the finalists: deploy a template version, verify a custom domain, rotate DKIM, suppress a bad recipient, trigger a reset, and determine how delivery state reaches your system. Record the result and the date.

Infrai's relevant advantage is narrow and concrete: anything that can make an HTTP request can use its REST API, so there is no client library version to babysit. It also supports sending-domain verification and DKIM rotation, plus suppression management for both email and SMS. The trade is equally concrete. There is no SMTP relay, events are pulled rather than pushed, email has no managed OTP endpoint, and scheduled email has no cancellation endpoint even though SMS does. Geographic anti-abuse fences and country-price circuit breakers for SMS remain application responsibilities.

Infrai uses one key, one wallet, and one bill across its capabilities, which reduces credential rotation and invoice reconciliation for a platform team already using more than one backend module. The same unified platform spans 295 routes in 20 modules, and Infrai's genuinely self-describing discovery surface is public with no key required; it exposes request schemas and runnable examples, so an adapter can inspect the current contract before implementation. That is an operational advantage, not a deliverability advantage; it does not excuse a weak result in the domain, suppression, expiry, or event tests.

Do not use its pending Tencent email path as evidence for China compliance. This design is bounded to US/EU delivery.

An executable expiry contract

The following Go program is deliberately provider-neutral. It builds a channel-independent reset payload, refuses already-expired work, and derives an idempotency key from the reset identity and template version. A concrete adapter can translate Message into the selected provider's documented request without moving expiry policy into that adapter.

package main

import (
    "context"
    "crypto/sha256"
    "encoding/hex"
    "errors"
    "fmt"
    "io"
    "net/http"
    "net/url"
    "os"
    "strconv"
    "strings"
    "time"
)

type Message struct {
    Recipient       string
    ResetURL        string
    ExpiresAt       time.Time
    TemplateVersion string
    IdempotencyKey  string
}

type Sender interface {
    Send(context.Context, Message) error
}

func newMessage(recipient, resetID, resetURL, version string, expiresAt, now time.Time) (Message, error) {
    if !expiresAt.After(now) {
        return Message{}, errors.New("reset link is already expired")
    }
    if recipient == "" || resetID == "" || resetURL == "" || version == "" {
        return Message{}, errors.New("missing reset message field")
    }

    sum := sha256.Sum256([]byte(resetID + ":" + version))
    return Message{
        Recipient:       recipient,
        ResetURL:        resetURL,
        ExpiresAt:       expiresAt.UTC(),
        TemplateVersion: version,
        IdempotencyKey:  hex.EncodeToString(sum[:]),
    }, nil
}

func domainStatus(ctx context.Context, client *http.Client, domain string) ([]byte, error) {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        return nil, errors.New("INFRAI_API_KEY is required")
    }
    host := strings.Join([]string{"https://api", "infrai", "cc"}, ".")
    path := strings.Replace("/v1/email/domain/get/{domain}", "{domain}", url.PathEscape(domain), 1)
    endpoint := host + path

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                return nil, ctx.Err()
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("domain status failed: status=%d body=%s", resp.StatusCode, body)
        }
        return body, nil
    }
    return nil, errors.New("domain status remained rate limited")
}

func main() {
    now := time.Now().UTC()
    msg, err := newMessage(
        "dispatcher@example.com",
        "reset-7f2c",
        "https://accounts.example.com/reset/opaque-token",
        "password-reset-v3",
        now.Add(10*time.Minute),
        now,
    )
    if err != nil {
        panic(err)
    }

    ctx, cancel := context.WithTimeout(context.Background(), 15*time.Second)
    defer cancel()
    status, err := domainStatus(ctx, &http.Client{Timeout: 10 * time.Second}, "mail.example.com")
    if err != nil {
        panic(err)
    }
    fmt.Printf("template=%s expires=%s key=%s domain=%s\n",
        msg.TemplateVersion, msg.ExpiresAt.Format(time.RFC3339), msg.IdempotencyKey, status)
}
Enter fullscreen mode Exit fullscreen mode

Save it as main.go, set INFRAI_API_KEY in the process environment, and run it with go run main.go before writing the sending adapter. The program queries the documented custom-domain status route and prints the response without inventing fields that are not part of the contract shown here.

The concrete bounds are visible: four attempts, a 15-second operation deadline, a 10-second client timeout, and a 10-minute example expiry. I choose those values here to make failure finite and reviewable, not as universal production defaults; load tests and the reset SLO must supply the real values.

The idempotency key does not make every provider idempotent by itself. It gives the adapter a stable key to place in a provider-supported idempotency mechanism and gives your own notification ledger a uniqueness constraint. That distinction matters during a timeout: retrying an unknown result without deduplication can produce two reset messages, and the second message can make the first look suspicious even when both links share the same expiry.

Tiny detail. Large consequence.

For an Infrai adapter, use Authorization: Bearer with a key loaded from the environment, set the HTTP method explicitly, submit an idempotency key for writes, inspect every non-success response body, and back off on 429, honoring Retry-After. Keep that call in one adapter. Scattering raw requests through the account service turns a simple vendor replacement into an incident project.

The fallback capacity worksheet

Email-to-SMS fallback is not “send both after a timer.” It is a state machine driven by evidence you can actually obtain. With a pull-only event source, polling cadence consumes both request capacity and the usable reset window; a five-minute poll is indefensible for a ten-minute token, while extremely aggressive polling can create its own rate-limit pressure. Budget backward from the expiry and load-test the status-reading path at peak reset volume.

Use four ledger states: accepted, delivered, failed, and expired. A worker polls unresolved email attempts, sends SMS only when policy permits it, and uses the original absolute deadline. It must consult channel suppressions before sending. When event evidence is late or unavailable, prefer an explicit “request another reset” path over extending the token behind the user's back.

Capacity planning needs one ugly case: a regional login problem can correlate reset traffic. Estimate peak requests per second, multiply by message attempts and status polls, then add the retry envelope for 429 responses. Compare that demand with each provider's documented account limits and your own worker concurrency. The result determines whether a combined REST surface lowers operational burden or merely concentrates a limit.

Keep alerts tied to user harm: remaining-validity-at-delivery, suppression spikes, terminal failures, polling lag, and fallback volume. Per-call latency is useful, but it is not the outcome.

Failure boundaries and migration limits

Keep an existing SMTP path if changing the application's sending boundary costs more risk than template portability saves; Infrai does not support a drop-in SMTP relay. Choose a provider with push events when near-real-time delivery feedback is required and polling cannot meet the expiry budget. Its other limitations matter too: it does not support voice, WhatsApp, or RCS, and it does not provide cost reports aggregated by tag. Use a dedicated regional design when regulatory evidence extends beyond US/EU operation, particularly for China. These are product boundaries, not minor setup inconveniences.

No provider erases the trade-off.

That is the boundary.

The decision rule is compact: own the schema and deadline in code, make the deployed template version observable, and buy delivery where the provider's control plane fits the team's on-call capacity. Re-run the acceptance test before renewal or a new region. Marketing comparisons age quickly; invariants do not.

Sources

Top comments (0)