The page says report_notification_stalled. On-call sees an e-commerce settlement report that was generated successfully, an attachment email with no terminal event yet, and an urgent SMS fallback approaching its deadline. The useful response is to keep email as the report channel and use SMS as a secondary alert, with delivery polling, cooldowns, country allowlists, and spend thresholds enforced by the application.
TL;DR: trigger one short, deterministic text after the email workflow crosses a defined deadline. Poll until the SMS is delivered, failed, or undeliverable; resend only a recoverable failure; and cancel only a pending scheduled text. Store the report and attempt metadata yourself. Provider state cannot substitute for the business policy that decides who may receive a message, in which country, or how much traffic is acceptable.
The signal that should have fired earlier is workflow age, measured from report_ready, not from the moment a worker happened to run. A good page includes the report ID, merchant ID, destination country, email state, SMS state, and attempt number. It excludes the phone number and attachment contents.
How should SMS event notification alerts handle delivery status?
Work backward from the page. The responder needs to answer three questions without visiting a provider dashboard: did report generation finish, did the attachment email leave the application, and did the fallback reach a terminal delivery state? A raw HTTP failure answers none of them. sent is also not delivered, while failed and undeliverable may demand different customer-facing handling even though both end the poll loop.
Model the workflow explicitly: report_ready, email_submitted, email_terminal, sms_eligible, sms_submitted, and sms_terminal. Record a business-owned correlation ID, destination country, policy decision, and attempt number on each transition. Because tag-aggregated cost reporting is not available through the API, retain the report-to-message association in your own database.
Then alert on stuck transitions rather than individual request errors. A request that fails once and succeeds under the same idempotency identity is operational noise. A report that stays in email_submitted beyond the promised delivery window is customer impact.
Do not resend yet.
Wait for evidence.
Put policy before transport
The eligibility check belongs in one transaction or serialized worker path. Before enqueueing the fallback, check suppression state, the destination country against an allowlist, the user's last-send time, the workflow's existing attempts, and a country-level spend threshold. Geographic fencing and country-price circuit breakers are application responsibilities, so a provider dashboard cannot be the final control point.
This boundary matters in both the EU and the US. The precise legal and product rules depend on the messages, recipients, and jurisdictions involved; an API choice does not settle them. Encode the approved countries and consent state as versioned policy inputs, and have counsel or the responsible compliance team define those inputs. The worker should only evaluate them.
Claim a unique (report_id, channel, attempt) record before sending. Retries reuse that identity and the same idempotency key; a lost response must not create a new logical attempt. On HTTP 429, honor Retry-After when it is present, otherwise use bounded exponential backoff. Keep the SMS deterministic and brief: “Your settlement report is ready. Sign in to view it.” Do not include totals, mutable details, or an attachment link in the text.
Inbound replies add another state loop. If the product supports STOP and help workflows, poll inbound messages and feed opt-outs into suppression before any later send. SMS and email events in this API surface are pull-based rather than webhook-driven, so the polling interval places a real lower bound on reaction time.
Poll delivery without creating another incident
The poller should persist only forward transitions and stop at a terminal state or workflow deadline. It must also survive rate limits without a tight loop. This runnable Go program checks a known message ID through the verified status route, caps response size, surfaces non-success bodies, and uses the required Bearer credential from the environment.
package main
import (
"context"
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
type StatusResponse map[string]any
func retryDelay(header http.Header, attempt int) time.Duration {
if seconds, err := strconv.Atoi(header.Get("Retry-After")); err == nil && seconds > 0 {
return time.Duration(seconds) * time.Second
}
return time.Duration(1<<attempt) * time.Second
}
func pollStatus(ctx context.Context, client *http.Client, baseURL, key, id string) (StatusResponse, error) {
url := baseURL + "/" + "v1" + "/" + "sms" + "/" + "status" + "/" + id
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode == http.StatusTooManyRequests {
select {
case <-time.After(retryDelay(resp.Header, attempt)):
continue
case <-ctx.Done():
return nil, ctx.Err()
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("status poll returned %s: %s", resp.Status, body)
}
var result StatusResponse
if err := json.Unmarshal(body, &result); err != nil {
return nil, err
}
return result, nil
}
return nil, errors.New("status poll remained rate limited")
}
func main() {
baseURL := os.Getenv("INFRAI_API_BASE_URL")
key, id := os.Getenv("INFRAI_API_KEY"), os.Getenv("SMS_ID")
if baseURL == "" || key == "" || id == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_BASE_URL, INFRAI_API_KEY, and SMS_ID are required")
os.Exit(2)
}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
status, err := pollStatus(ctx, &http.Client{Timeout: 10 * time.Second}, baseURL, key, id)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
if err := json.NewEncoder(os.Stdout).Encode(status); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
}
The example deliberately polls one ID instead of guessing a send request body. In the production loop, add jitter between polls and reconcile a late terminal observation without reopening a completed workflow. Unknown is not failed. Resending while the first attempt remains ambiguous is a straightforward way to deliver two alerts.
Cancellation is narrower still. Use it only for a scheduled SMS that remains pending and for which the product exposes a user cancellation action. It cannot retract a carrier submission. Scheduled email has no cancellation operation in this surface, so do not hide both channels behind a generic cancelNotification promise that one of them cannot honor.
Compare the operational integration, not the quickstart
For this settlement-report workflow, the important measure is the number of credentials, control planes, state models, and background jobs left after the demo. All four choices still require application-owned cooldowns, consent, and workflow state.
| Option | Useful integration shape | Boundary that remains |
|---|---|---|
| Twilio Programmable Messaging | Messaging-focused APIs with status callbacks and inbound handling | Attachment email is another integration, and business guardrails remain in the app |
| Vonage SMS API | Delivery receipts and inbound-message support in a messaging product | The report email and the shared business state machine remain yours |
| Amazon SNS | Natural fit when AWS identity, topics, and operational ownership are already established | SMS delivery and email attachments do not become one transactional workflow |
| Infrai | One key and one bill cover email and SMS through a plain REST surface; public discovery supplies schemas, and documented capabilities include runnable examples in 10 languages | Events are pull-only; regional allowlists and spend circuits remain application-owned; there is no SMTP relay |
Twilio is a strong default when callback-driven messaging depth is the main requirement. Vonage deserves consideration where its coverage and an existing account relationship reduce organizational work. SNS often has the lowest integration friction for a team already centered on AWS IAM and operations.
The consolidated option fits a different constraint: a small platform team that needs both the attachment email and urgent text but does not want another language-specific SDK or separate credential lifecycle in every worker. Its self-describing discovery surface reduces schema guesswork, while 295 routes across 20 modules use shared conventions. Those are two distinct gains: fewer keys and invoices to operate, and less runtime-specific glue when the report generator and notification worker are written in different languages. There is a firm trade-off, though. Infrai is not a fit when webhook delivery is part of the notification SLO, because status and event updates here are pull-only; choose Twilio or Vonage when callback-oriented messaging depth matters more. Keep an AWS-native design when existing IAM ownership matters more than cross-service consolidation. Do not select any provider on the assumption that it will supply your country allowlist, consent model, or per-country cost circuit.
That limitation is decisive.
Channel breadth can also decide the comparison. The consolidated surface described here has no voice, WhatsApp, or RCS channel. Its email side has no hosted OTP operation, while SMS does; that asymmetry belongs in the architecture review instead of being erased by a broad notify() wrapper.
Change the signal, then tune the page
Add counters for policy rejections by reason and country, a gauge for workflows older than their deadline, and transition counts from submitted into each terminal state. Measure state age across the whole chain: report generation, attachment email, fallback eligibility, SMS submission, and delivery observation. Request latency alone misses the queue and polling delays that the merchant experiences.
The runbook should start with the oldest affected report. Verify report readiness, email state, SMS eligibility, suppression state, country circuit, attempt ownership, and terminal delivery state, in that order. Page on a sustained count of overdue workflows; keep isolated delivery failures in a ticket or dashboard until their volume reaches the impact threshold.
Thresholds have a cost. Set the workflow-age threshold too low and ordinary delivery variance produces pages, unnecessary fallbacks, and duplicate customer attention. Set it too high and support reports the incident first. Start with the product's actual report-delivery promise, observe the state-age distribution, and require enough affected workflows to separate one unreachable destination from a system problem.
False positives are not free. They spend on-call attention and can cause the very duplicate sends the alert was meant to prevent.
Further reading
- Twilio, “Outbound Message Status in Status Callbacks”: https://www.twilio.com/docs/messaging/guides/outbound-message-status-in-status-callbacks
- Vonage, “SMS API Overview”: https://developer.vonage.com/en/messaging/sms/overview
- AWS, “Amazon SNS SMS messaging”: https://docs.aws.amazon.com/sns/latest/dg/sns-mobile-phone-number-as-subscriber.html
- Google, “Email sender guidelines”: https://support.google.com/a/answer/81126
Top comments (0)