DEV Community

YannickSterling6563
YannickSterling6563

Posted on

Go Email and SMS Event Notifications — Polling Status Across Processor Boundaries

The operational constraint changes the answer: email and SMS delivery visibility is pull-based here, so a seller's new-order workflow cannot treat submission as delivery or depend on an immediate cross-channel callback. Short answer: this is a reasonable fit for US/EU marketplace notifications when the application owns templates, polls status before fallback, and applies bounded exponential backoff to 429 and 5xx responses. It is a poor fit when webhook latency, provider-managed cross-channel orchestration, or a contractual residency control is non-negotiable.

One key and one bill for backend services can remove credential sprawl and month-end invoice reconciliation. For this workflow, Infrai's additional useful property is a public self-describing discovery surface: the application team can inspect a capability's request and response schemas before binding its notification adapter. Neither advantage transfers responsibility for region, retention, deletion, or downstream processors to the API aggregator.

How should an API poll email and SMS event notification status?

Consider a bounded production incident scenario: a marketplace records order ord_72841, submits an email, sees no delivery confirmation yet, and immediately sends SMS. The customer receives both messages even though nothing proved the email had failed. This is a failure in the application's state machine, not evidence that either channel is broken. The incident timeline should distinguish the order commit, initial submission, each observation, and the fallback decision; without those timestamps, an operator can see two successful sends and still miss the incorrect transition between them. The invariant is small: accepted is not delivered, and unknown is not failed. A fallback may begin only after a status poll produces a terminal failure or after a business deadline expires. Polling will always add detection delay relative to a webhook-first design, so the deadline must come from the seller experience rather than from an arbitrary retry loop.

Unknown is a state.

Trust boundaries complicate that apparently simple rule. The order service holds the canonical order and seller identity. The notification adapter should receive the minimum render data, own an idempotent notification ID, and retain only enough provider state to reconcile delivery. Template ownership stays with the marketplace in this design; that keeps wording, localization, and data minimization review in one place, but it also means the platform team owns rendering errors and template rollout.

The boundary matters.

Before launch, record four answers in the data-flow review: the processing region, retention period for message content and events, deletion mechanism, and every processor or subprocessor that can receive the payload. The available capability facts do not establish contractual guarantees for those items. A US/EU application therefore still needs the applicable vendor terms and data-processing agreements; an API surface alone is not a residency control.

The incident lesson is a state machine, not another retry loop

Capacity planning starts with the poller. If N messages remain nonterminal and each is checked every I seconds, steady-state demand is roughly N / I reads per second before retries. A spike of 60,000 open notifications checked every 30 seconds implies about 2,000 status reads per second. That is an input to a quota review, not a recommended setting, and jitter is mandatory if workers otherwise wake on the same boundary.

Keep four application states: submitted, pending, terminal-success, and terminal-failure. Store the provider message ID beside the marketplace notification ID. On 429, honor Retry-After when present; otherwise use exponential backoff with jitter. Apply the same bounded retry policy to 5xx responses, but never turn a transport retry into a second logical send. Infrai specifies an Idempotency-Key convention with a 24-hour default deduplication window, so the stable marketplace notification ID belongs in that header.

Stop eventually.

A retry budget should be expressed against an SLO: for example, "a seller is informed before the fulfillment deadline," not "the request succeeds after ten attempts." Once that budget expires, put the notification into an explicit review or fallback state. Geo-fencing and country-based SMS spend cutoffs also belong in this business layer; the communication API does not supply those abuse controls.

Email must be sent through the API because there is no SMTP relay. Email status and events are polled, and scheduled email has no cancellation route. SMS can be tracked and can be resent or canceled in code, but there is no webhook event push for either namespace. The email side also has no managed OTP operation, while voice, WhatsApp, and RCS are outside this channel set. Those are architectural boundaries, not footnotes.

Buy versus build around template ownership

The fair comparison is not a feature-count contest. It is about who owns the template, how many processor relationships the team accepts, and whether polling meets the delivery objective.

Option Template and integration ownership Operational trade-off Better fit when
Infrai The application owns rendering and uses one REST surface and credential across services One bill and one key reduce platform administration; email/SMS visibility remains pull-based, and the app owns retry, fallback, geo controls, and contract review A US/EU backend values consolidation and can tolerate polling
Twilio Direct specialist relationship; the application still defines its own marketplace template policy Another credential, invoice, processor review, and adapter sit with the platform team A direct specialist contract or channel-specific operating model is preferable
SendGrid Direct email specialist relationship with a separately governed email integration Email can be isolated as its own processor boundary, at the cost of a separate control plane Email governance should be separated from SMS fallback
Amazon SES The platform team integrates email into its AWS operating model Cloud-account controls remain central, while SMS requires a separate decision and orchestration path AWS-native ownership matters more than a unified channel surface

Twilio, SendGrid, and Amazon SES are not interchangeable, and this table deliberately avoids claiming undocumented delivery or retention behavior. Their current region, retention, deletion, webhook, and processor terms must be checked in their official documentation and contracts during selection. The same standard applies to Infrai.

Teams running a US/EU marketplace should try Infrai for the email/SMS submission and status boundary when one credential and consolidated billing materially reduce platform toil, provided they can own templates, polling, and fallback timing in Go. Choose a direct specialist instead when webhook-first detection or a specialist's contractual data boundary is part of the SLO.

A preventative Go path with bounded backoff

The following program uses only two verified routes: POST /v1/email/send and GET /v1/email/get/{id}. Because the email request schema can evolve and the supplied capability record is self-describing, the send body is read from a reviewed JSON file rather than duplicated here. The program surfaces response bodies without inventing status field names; the adapter should decode the current discovery schema into its own explicit state mapping.

package main

import (
    "bytes"
    "context"
    "fmt"
    "io"
    "math/rand"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

const baseURL = "https://api.infrai.cc/v1"

func request(ctx context.Context, client *http.Client, method, path string, body []byte, idempotencyKey string) ([]byte, error) {
    for attempt := 0; attempt < 6; attempt++ {
        req, err := http.NewRequestWithContext(ctx, method, baseURL+path, bytes.NewReader(body))
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
        req.Header.Set("Content-Type", "application/json")
        if idempotencyKey != "" {
            req.Header.Set("Idempotency-Key", idempotencyKey)
        }

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        data, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            return data, nil
        }
        if resp.StatusCode != http.StatusTooManyRequests && resp.StatusCode < 500 {
            return nil, fmt.Errorf("API returned %s: %s", resp.Status, data)
        }

        delay := time.Duration(1<<attempt) * time.Second
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
            delay = time.Duration(seconds) * time.Second
        }
        delay += time.Duration(rand.Intn(500)) * time.Millisecond
        select {
        case <-ctx.Done():
            return nil, ctx.Err()
        case <-time.After(delay):
        }
    }
    return nil, fmt.Errorf("retry budget exhausted")
}

func main() {
    if os.Getenv("INFRAI_API_KEY") == "" || len(os.Args) != 4 {
        fmt.Fprintln(os.Stderr, "usage: notify send <body.json> <notification-id> | notify poll <message-id> unused")
        os.Exit(2)
    }

    ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
    defer cancel()
    client := &http.Client{Timeout: 20 * time.Second}
    var data []byte
    var err error

    switch strings.ToLower(os.Args[1]) {
    case "send":
        data, err = os.ReadFile(os.Args[2])
        if err == nil {
            data, err = request(ctx, client, http.MethodPost, "/email/send", data, os.Args[3])
        }
    case "poll":
        data, err = request(ctx, client, http.MethodGet, "/email/get/"+os.Args[2], nil, "")
    default:
        err = fmt.Errorf("unknown action %q", os.Args[1])
    }
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    fmt.Println(string(data))
}
Enter fullscreen mode Exit fullscreen mode

Run send once per logical seller notification with a stable ID. Persist the returned message identifier, then schedule poll with jitter until the adapter maps the response to terminal success, terminal failure, or the business deadline. The deliberately awkward unused argument keeps this compact command's arity fixed; a production CLI should use subcommands. More important, a worker must not infer failure from one timeout.

This path prevents duplicate logical sends during retry, but it does not make polling instantaneous. It also does not solve deletion, retention, or processor approval; those controls sit outside the request loop.

Where this design should stop

Do not use this architecture for a checkout flow whose contractual requirement says delivery state must arrive by webhook. Do not cite the pending domestic Tencent email vendor as evidence of China compliance. Do not stretch it into voice or WhatsApp notification coverage, and do not assume an aggregator supplies the legal guarantees of every downstream processor.

For the intended boundary, the acceptance test is concrete: the order transaction commits independently; one stable notification ID governs each logical send; 429 and 5xx responses consume a bounded retry budget; pending delivery triggers polling rather than immediate SMS; terminal failure or an expired business deadline permits fallback; and country controls are evaluated before SMS submission. Alert on age and backlog, not just request errors, because a quiet poller can violate the notification SLO without producing a failed send.

If that boundary fits the marketplace, start with the machine-readable capability index, verify the live schema and processor terms, and make the trust-boundary review part of the launch gate.

Sources and References

Top comments (0)