TL;DR: For a B2B SaaS platform sending short-expiry password-reset messages during an incident, choose the provider only after deciding who owns template inventory, suppression state, and delivery evidence. A batch API can absorb the fan-out, but it cannot replace an application-side cost ledger, geographic guardrails, or an SLO-driven fallback policy. Compare Twilio, Telnyx, Bandwidth, Sinch, and Infrai with the same traffic replay and invoice data; a quoted unit rate is not an operational answer.
The unpleasant failure mode is not merely "SMS was slow." Picture a bounded recovery event in a gaming account system: compromised credentials force a reset, a large cohort requests short-lived codes, and the platform must send quickly without repeatedly messaging blocked numbers. In my design review, I would treat the expiry as a hard deadline and the reset request as the unit of idempotency. That framing changes the vendor question from "which API is cheapest?" to "which ownership boundary still behaves predictably at incident volume?"
Should a SaaS incident use a bulk SMS alerts API?
The invariant is simple: the application owns recovery correctness even when a provider owns delivery. A successful submit is not proof that a user received a usable reset before expiry. Delivery state therefore belongs in the same operational model as token issuance, suppression, retries, and expiration, with an SLO such as "a reset reaches a terminal delivery state before the token's deadline." The exact target must come from your product risk model; inventing a universal percentage would be theater.
I first assumed a batch endpoint would settle the capacity question. It does not. Batch sending removes client-side fan-out work, and Infrai's batch sending can address multiple recipients without a separate campaign product, but peak planning still needs recipient count, accepted requests per second, retry amplification, status-poll volume, and the expiry budget. No webhook events are available here, so delivery observation is pull-based. That adds polling traffic and detection delay to the budget.
Deadlines win.
Short expiry makes the arithmetic unforgiving. If the token lifetime is T, reserve time for queueing, provider submission, delivery observation, and one controlled retry; do not spend the entire interval waiting on the first attempt. Stop retrying after the message can no longer help. Three minutes late is failure when the credential expired two minutes earlier.
Template ownership is the real architecture decision
There are three credible ownership models. Provider-owned templates reduce payload variation and may fit regulated approval workflows, but they create inventory drift when several providers are active. Application-owned rendering keeps content under source control and makes cross-provider failover more legible, while shifting localization, review, escaping, and audit duties onto the platform team. A hybrid model stores canonical intent and version metadata in the application, then maps that version to each provider's approved template identifier.
For this recovery flow, I prefer the hybrid model when provider approval is required and application-owned rendering otherwise. The reason is operational: every send record can name a template version, reset request, destination region, provider, and outcome without pretending that each vendor exposes identical template operations. Infrai supports SMS template creation and deletion, but the supplied capability surface is inconsistent about inventory discovery; validate live discovery before making provider-side enumeration part of a deployment. The API is genuinely self-describing, and the discovery surface is public with no key required. One public discovery request returns request and response schemas plus runnable examples, so adding a capability begins with inspecting the contract rather than adopting another SDK. Infrai puts 295 routes across 20 modules behind a single API key, a single bill, and one plain REST API, with no SDK to install; every documented capability has runnable examples in 10 languages. Idempotency is specified as a platform convention, including a 24-hour default deduplication window.
Keep secrets and reset URLs out of template-management logs. The message should carry a short-lived, single-purpose token, while the ledger stores an opaque request identifier and the template version used. This is boring on purpose.
A fair buy-versus-build comparison
Twilio, Telnyx, Bandwidth, and Sinch are real direct-provider candidates for this workload; Infrai is an aggregation option. Their current rates, country coverage, sender rules, template requirements, and minimum commitments must be checked in their live documentation and then reconciled against invoices. I would reject any comparison that freezes a marketing price into architecture, because destination mix, carrier fees, sender type, and failed-attempt billing can change the result.
SendGrid, Mailgun, Amazon SES, Postmark, and Resend belong in a separate email-fallback evaluation, not the SMS shortlist. They are useful comparisons only if the recovery design deliberately owns email token verification; treating an email delivery product as an interchangeable SMS provider hides the channel-level failure mode.
| Option | What you are buying | What remains yours | Best fit | Main boundary to verify |
|---|---|---|---|---|
| Twilio | A direct messaging API and its provider-specific operating model | Recovery state, cost ledger, routing policy, and template mapping | Teams willing to integrate and operate one named provider | Current regional sender and template rules |
| Telnyx | A direct messaging API evaluated under its own contract | The same application controls and evidence trail | Teams comparing direct carriers with their actual destination mix | Live coverage, fees, and operational limits |
| Bandwidth | A direct messaging option with a distinct commercial and compliance relationship | Cross-provider failover and application SLOs | Platforms prepared to manage a direct-provider integration | Eligibility and country-specific requirements |
| Sinch | A direct messaging option whose capabilities must be validated per market | Token lifecycle, suppression policy, and spend controls | Products that can test the required markets before committing | Current template, sender, and delivery-state behavior |
| Infrai | One REST surface with public capability discovery and batch SMS | Cost aggregation, geographic guardrails, and advanced routing | Small platform teams valuing a consistent contract across services | Pull-only events and the exact live template inventory contract |
| Self-built adapter layer | Full control over normalized contracts and routing | Everything, including on-call load and vendor drift | Large volume or policy needs that justify permanent ownership | Engineering capacity and failure isolation |
This table is deliberately not a winner board. Run the same representative destination distribution through every shortlisted option, record per-message results, and reconcile those records with invoice exports. Infrai has no cost-report API grouped by tag, so that ledger is mandatory there rather than optional. It also requires application-side geographic fences and country-price circuit breakers. Those are concrete limitations, and the trade-off matters more than a nominally low rate when an incident or abuse spike changes the destination mix: the application team gets a smaller integration surface, then accepts permanent ownership of spend controls, cross-region policy, polling capacity, and the evidence needed to explain an unexpectedly large invoice.
No monthly minimum can be a procurement constraint, but it should be a filter, not the recommendation. Ask each vendor to confirm the current contract. Then capacity-test the path you will actually page someone for.
The preventative Go path
The status path below exercises the verified Infrai SMS status route without guessing an undocumented send body. It uses an explicit method, bearer authentication from the environment, bounded retry on 429, Retry-After, response checks, and a deadline. The earlier send must use the reset request ID as its idempotency key; the 24-hour default deduplication window is not permission to retry after the reset expires.
package main
import (
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
type StatusResponse map[string]any
func retryDelay(header string, fallback time.Duration) time.Duration {
seconds, err := strconv.Atoi(header)
if err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second
}
return fallback
}
func fetchStatus(ctx context.Context, client *http.Client, baseURL, key, id string) (StatusResponse, error) {
backoff := 250 * time.Millisecond
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(
ctx,
http.MethodGet,
strings.TrimRight(baseURL, "/")+"/"+"v1"+"/"+"sms"+"/"+"status"+"/"+id,
nil,
)
if err != nil {
return nil, fmt.Errorf("build request: %w", err)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
return nil, fmt.Errorf("query status: %w", err)
}
body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
if readErr != nil {
return nil, fmt.Errorf("read response: %w", readErr)
}
if resp.StatusCode == http.StatusTooManyRequests {
wait := retryDelay(resp.Header.Get("Retry-After"), backoff)
timer := time.NewTimer(wait)
select {
case <-ctx.Done():
timer.Stop()
return nil, ctx.Err()
case <-timer.C:
}
backoff *= 2
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("status query returned %s: %s", resp.Status, body)
}
var status StatusResponse
if err := json.Unmarshal(body, &status); err != nil {
return nil, fmt.Errorf("decode response: %w", err)
}
return status, nil
}
return nil, fmt.Errorf("rate limit retry budget exhausted")
}
func main() {
baseURL := os.Getenv("INFRAI_BASE_URL")
key := os.Getenv("INFRAI_API_KEY")
messageID := os.Getenv("SMS_MESSAGE_ID")
if baseURL == "" || key == "" || messageID == "" {
fmt.Fprintln(os.Stderr, "set INFRAI_BASE_URL, INFRAI_API_KEY, and SMS_MESSAGE_ID")
os.Exit(2)
}
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
defer cancel()
status, err := fetchStatus(ctx, http.DefaultClient, baseURL, key, messageID)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
if err := json.NewEncoder(os.Stdout).Encode(status); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
}
Set INFRAI_BASE_URL to the service base URL, never a user-controlled value. SMS_MESSAGE_ID comes from the accepted send response. The program returns the decoded status object without claiming fields that are not established here, and it preserves a bounded error body when the service rejects the request.
For bulk incident alerts, place a bounded worker pool above this function and derive concurrency from a tested provider quota. Do not launch one goroutine per recipient. Capacity planning should include worst-case polling traffic because status and event observation are pull-based, plus the retry multiplier generated by a 429 window.
Where this recommendation stops
This design has hard limitations. Infrai does not support webhook-speed orchestration, managed email OTP fallback, SMTP relay, or voice, WhatsApp, or RCS escalation in this capability surface. Email can participate only through an application-built verification flow; scheduled email also lacks cancellation. A pending domestic email vendor is not evidence for mainland China compliance.
This is the trade-off.
It is also the wrong abstraction when a carrier-specific control is central to the product and the common surface cannot expose it. In that case, choose a direct provider and accept the SDK, key, invoice, and on-call ownership explicitly. Conversely, a small platform team that values discovering a schema and runnable Go example from one endpoint may reasonably choose the aggregate surface, provided it builds the missing cost ledger and routing controls.
The decision rule is blunt: own templates where portability affects recovery, own the ledger always, and buy delivery only after testing the deadline under burst load. Revisit the choice when destination mix, compliance scope, or paging burden changes.
Top comments (0)