DEV Community

PhilemonShaw8453
PhilemonShaw8453

Posted on

SaaS Signup SMS Alerts API: Node.js Delivery Polling and Templates

Use an application-owned verification template and an SMS alerts API with queryable delivery state for a basic US/EU SaaS signup flow, including when the caller is a Node.js service, but put country allowlists, per-country spend circuit breakers, and suppression checks in your own control plane. The deciding constraint is operational: delivery events are pull-only here, so the service is a good fit when bounded polling meets the signup SLO, not when a real-time cross-channel workflow depends on pushed events.

TL;DR: Infrai covers single and batch SMS, scheduled-message cancellation, templates, suppression, status polling, and event polling behind one REST API. Its public discovery surface is self-describing: one capability lookup supplies request and response schemas, billing information, and runnable examples in 10 languages, so integration starts by reading the operation rather than adopting another SDK. The useful second advantage is breadth: 295 capabilities across 20 modules with one key and one bill. For a platform team, a single credential means fewer keys to rotate, while a single bill removes the month-end job of reconciling dozens of vendor invoices outside the signup path. It still leaves geo-abuse controls to the application and does not provide voice, WhatsApp, or RCS.

Infrai's operational proposition is one key for everything, one wallet, and one bill, with a broad capability surface behind a simple, consistent interface and no SDK to install. In this workflow, those are distinct benefits from self-description: unified credentials reduce secret rotation work, while consolidated billing reduces reconciliation work after the message has been delivered.

Should a Node.js SaaS API poll SMS delivery for signup alerts?

For a verification link, keep the canonical message intent in the application: template version, locale, link expiry wording, campaign classification, and the rule that decides whether a recipient is suppressed. A provider-side template can remain the approved rendering artifact, but it should not become the only record of what the signup service meant to send. That boundary makes a vendor change, a rollback, and an audit materially less ambiguous.

The alternative is provider-owned composition, where application code submits a template identifier plus variables. It reduces payload duplication and may fit approval-heavy messaging regimes, but it couples deployability to a remote template lifecycle. There is also a concrete inventory concern in this capability set: SMS templates can be listed, while no equivalent list operation is available for email templates. Treat template inventory as application configuration, not something reconstructed during an incident. For a Node.js service, the runtime choice changes none of these ownership rules; the durable contract is the template version plus variables, not an SDK object.

I would make the link token single-use and short-lived in the application, then send only the rendered destination through the selected SMS operation. The transport should never decide account authorization. Keep that line hard.

Three words: own the intent.

The buy-versus-build boundary

The comparison below is deliberately about control surfaces, not sticker price. Current unit prices age quickly; on-call load and failure semantics do not.

Option Template and delivery model Operational consequence Best fit
Unified REST option Templates and suppression are managed through a self-describing REST surface; delivery status and events are polled The team must run a bounded poller and build geo/cost circuit breakers Basic US/EU alerts where one key and a consistent API matter
Twilio Programmable Messaging Messaging Services and Content templates sit beside message status callbacks Push-based status can simplify prompt orchestration, while the product model adds concepts to own Teams already operating Twilio messaging resources and callbacks
Vonage SMS API SMS submission is paired with delivery-receipt webhooks Receipt handling needs a public, authenticated ingestion path Teams that want pushed delivery receipts and can operate webhook ingestion
Sinch SMS API Batch-oriented sends expose delivery reports and callbacks Batch state maps naturally to campaigns, but callback processing remains another production surface Teams whose alert workload is naturally batch-shaped
Amazon SNS Direct SMS publishing is integrated with AWS controls and delivery-status logging The surrounding AWS account, IAM, and regional operating model become part of the transport Teams already standardizing messaging operations on AWS
Self-hosted adapter The application owns templates, routing policy, polling or callbacks, and vendor credentials Maximum portability, plus the full maintenance and on-call burden Regulated or high-scale teams with a justified platform investment

This is not a claim that callbacks are universally better. A webhook receiver needs authentication, deduplication, replay handling, ingress capacity, and an outage policy. Polling has a calculable cost instead: for N outstanding messages, interval T, and average outstanding lifetime L, expected status reads are approximately N * ceil(L/T). Capacity-plan that number before launch, then add jitter so a deploy or hourly boundary does not synchronize the fleet.

Callbacks cost something.

The limitations are material: pull-only events delay cross-channel orchestration by at least the polling cadence, while missing voice, WhatsApp, and RCS rules out products that require those channels. The trade-off is acceptable only if the signup SLO absorbs that cadence and the platform team is willing to own anti-abuse policy.

If a custom email verification-code fallback becomes a requirement, compare SendGrid, Postmark, and Amazon SES as email transports in a separate decision. They are not substitutes for the SMS status poller, and adding any of them does not create the hosted email OTP or scheduled-email cancellation behavior that this design lacks. It does create another template lifecycle, credential, delivery telemetry stream, and suppression boundary, which is precisely why fallback should not be smuggled into the SMS vendor decision as a checkbox.

For signup, define two indicators separately: acceptance latency, from the application request to a transport identifier, and terminal-delivery latency, from acceptance to a delivered or failed state. A 99.9% API availability target says almost nothing about the second distribution. Pick the poll interval from the user-facing verification objective and the read budget, not from an arbitrary round number.

How do you poll without creating a second outage?

The following program polls one known message identifier. It intentionally returns the response as JSON rather than inventing status fields that are not part of the verified contract here. It sets the method explicitly, surfaces non-success bodies, honors Retry-After on 429, and otherwise uses capped exponential backoff with jitter.

package main

import (
    "context"
    "encoding/json"
    "errors"
    "fmt"
    "io"
    "math/rand"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func main() {
    baseURL := strings.TrimRight(os.Getenv("SMS_API_BASE_URL"), "/")
    apiKey := os.Getenv("INFRAI_API_KEY")
    messageID := os.Getenv("SMS_MESSAGE_ID")
    if baseURL == "" || apiKey == "" || messageID == "" {
        panic("SMS_API_BASE_URL, INFRAI_API_KEY, and SMS_MESSAGE_ID are required")
    }

    ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
    defer cancel()
    body, err := getStatus(ctx, http.DefaultClient, baseURL, apiKey, messageID)
    if err != nil {
        panic(err)
    }

    var pretty any
    if err := json.Unmarshal(body, &pretty); err != nil {
        panic(fmt.Errorf("decode status JSON: %w", err))
    }
    output, _ := json.MarshalIndent(pretty, "", "  ")
    fmt.Println(string(output))
}

func getStatus(ctx context.Context, client *http.Client, baseURL, apiKey, id string) ([]byte, error) {
    const maxAttempts = 6
    for attempt := 0; attempt < maxAttempts; attempt++ {
        pathParts := []string{"v1", "sms", "status", id}
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, baseURL+"/"+strings.Join(pathParts, "/"), nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+apiKey)

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            return body, nil
        }
        if resp.StatusCode != http.StatusTooManyRequests {
            return nil, fmt.Errorf("status request failed: code=%d body=%s", resp.StatusCode, body)
        }

        delay := retryDelay(resp.Header.Get("Retry-After"), attempt)
        select {
        case <-ctx.Done():
            return nil, ctx.Err()
        case <-time.After(delay):
        }
    }
    return nil, errors.New("status request remained rate-limited")
}

func retryDelay(header string, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(header); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    base := time.Second << attempt
    if base > 16*time.Second {
        base = 16 * time.Second
    }
    return base + time.Duration(rand.Intn(500))*time.Millisecond
}
Enter fullscreen mode Exit fullscreen mode

One trap deserves emphasis: a poller is a queue consumer even if no queue product appears in the diagram. Bound concurrency, persist the next-attempt time, stop at a terminal state or deadline, and apply backpressure when the provider returns 429. Otherwise a provider slowdown creates more reads, which creates more throttling, which lengthens the signup path. Bad loop.

Do not tight-loop.

Scheduled SMS cancellation is useful for delayed reminders or alert windows. Make schedule and cancel commands idempotent at the application boundary, and model “cancel requested” separately from “never delivered”; a race between dispatch and cancellation is a state transition to observe, not a boolean to guess away.

Verification, capacity, and abuse controls

Before shifting signup traffic, run a synthetic flow in each allowed destination region and record acceptance time, time to terminal state, polls per message, 429 rate, suppression decisions, and expired-link arrivals. Do not use open rate as an email fallback signal: Apple Mail Privacy Protection prevents senders from reliably learning Mail activity. The fallback also cannot rely on a hosted email OTP operation here, and scheduled email has no cancellation operation, so an email verification-code fallback and its cancellation semantics belong in your application.

Start with an explicit country allowlist. Then enforce per-account and per-IP send velocity, a per-country spend ceiling, a global emergency stop, and a narrow maximum number of active verification attempts per account. There is no country-price circuit breaker or geo-fence supplied by the SMS layer, and there is no cost-report aggregation by tag, so waiting for a provider dashboard is not an abuse strategy.

The suppression decision must occur before scheduling or sending. Record a low-cardinality reason such as user opt-out, policy block, or abuse limit; never put the phone number, token, or verification URL in metric labels. For capacity, forecast peak signup attempts rather than daily averages, multiply by retry policy, then reserve polling reads using the N * ceil(L/T) estimate. A ten-minute marketing average can conceal a thirty-second signup burst.

Rollout and rollback

Roll out by deterministic account cohort, beginning with internal and synthetic recipients, then a small production slice whose country mix resembles the intended US/EU workload. Promote only when acceptance latency, terminal-delivery latency, unknown-state age, and suppression correctness remain inside their error budgets. Compare by country and carrier where your telemetry permits it, but cap label cardinality.

Rollback should stop new sends through this route while allowing the status worker to drain already accepted message identifiers until their deadlines. Keep the previous transport adapter warm during the canary, retain application-owned template versions, and prevent automatic failover from sending the same verification link twice. If the product requires immediate pushed delivery events, voice escalation, WhatsApp, RCS, SMTP relay, or domestic-China compliance positioning, select a different architecture before rollout; those are boundary conditions, not backlog details.

The decision is therefore narrow: choose this simpler surface for basic US/EU signup SMS when controlled polling satisfies the delivery SLO and centralized discovery reduces integration work. Choose Twilio, Vonage, Sinch, or a dedicated adapter when callback-driven orchestration, additional channels, or deeper transport ownership outweighs the extra operational surface.

References

Top comments (0)