Suppress a destination only after a durable, recipient-specific failure; retry temporary delivery failures with a bounded schedule, and keep OTP fallback separate from security-notification delivery. That rule matters more than the provider logo. In an edtech system, one recycled or mistyped guardian number can otherwise produce repeated sends long after the first failure signal.
TL;DR: put a provider-neutral delivery ledger and suppression registry between the application and every SMS API. Normalize numbers to E.164, assign one idempotency key per logical message, ingest signed status callbacks, and make suppression a reviewed state transition. Use a second provider only for narrowly defined temporary failures, never to evade an invalid-number, opt-out, or policy rejection.
How should an SMS provider handle security alerts and OTP notifications?
An accepted request proves that an upstream API received work. It does not prove that a handset received the message. SMS delivery continues through carrier networks, so the final outcome may arrive later through a status callback. Treating the initial 2xx as delivery creates false success; treating every delayed receipt as failure creates duplicates.
Accepted is not delivered.
The useful state machine is small: queued, accepted, delivered, temporary_failure, and permanent_failure. Keep the provider's raw code beside the normalized state because mappings change and operators need evidence during review. A callback may be repeated or arrive after an earlier event, so updates must be idempotent and monotonic: delivered must not slide back to accepted because an old callback was retried.
For this workload, "bounce" is an operational analogy, not an SMS protocol term. The actionable signals are provider and carrier delivery statuses. A syntactically valid E.164 number can still be unreachable, inactive, opted out, or unable to receive the traffic class. Syntax validation is an admission check, not proof of ownership or reachability.
Make suppression an application-owned decision
Store the normalized recipient, scope, reason class, source event, observed time, and review state. Scope matters. An explicit opt-out can require broader blocking than an unreachable handset, while a failed OTP attempt must not silently suppress unrelated account notices. Legal and policy handling varies by jurisdiction and use case, so encode the applicable consent policy rather than inferring it from a +1 or +33 prefix.
The dangerous shortcut is a single Boolean named invalid. It loses the distinction between a permanent address problem and a provider timeout. Use a reasoned record instead:
| Signal | Application action | Cross-provider fallback |
|---|---|---|
| Malformed or impossible destination | Reject before enqueue; record validation reason | No |
| Explicit recipient opt-out | Suppress under the applicable consent scope | No |
| Provider authentication or policy rejection | Stop the route and alert an operator | No |
| Carrier or destination reports a permanent failure | Suppress after mapping the documented code | No |
| Timeout or documented temporary failure | Retry with a cap and jitter | Only if policy permits |
| No final receipt before the OTP expires | End that attempt; require a fresh challenge | No stale resend |
No loopholes. Routing an opt-out or permanent recipient failure through another provider converts a reliability mechanism into repeated unwanted traffic.
Keep OTP data lean. NIST SP 800-63B treats use of the public switched telephone network for out-of-band authentication as a restricted authenticator and calls for considering risks such as SIM change or number porting. SMS fallback should therefore be one bounded authentication path, not evidence that the recipient is safe or that a sensitive account-change notification was read. Do not place secrets, student data, or account-recovery detail in the message body; send the minimum context needed to recognize the event.
There is a real trade-off here. An application-owned suppression registry adds storage, code mappings, privacy review, and an on-call procedure; it is unsuitable when a team cannot maintain callback verification and reason-code mappings. In that case, narrow the launch to one documented route and its provider-managed controls until the operational ownership exists. A second provider increases route diversity, but it also doubles integration tests and can produce conflicting final states. It is the wrong fallback for permanent failures, opt-outs, and OTPs whose validity window is already too short for another delivery attempt.
Implement the send boundary in Go
The provider adapter should return an acceptance identifier, not delivered=true. The application owns the logical message ID and uses it for deduplication before any network call. The following boundary leaves HTTP details and vendor error mappings inside adapters while keeping the scheduling rule testable.
package messaging
import (
"context"
"errors"
"time"
)
type FailureClass string
const (
Temporary FailureClass = "temporary"
Permanent FailureClass = "permanent"
Policy FailureClass = "policy"
)
type Message struct {
ID string // Stable for one logical notification.
Recipient string // Canonical E.164 form after validation.
Body string
ExpiresAt time.Time
}
type Acceptance struct {
ProviderMessageID string
AcceptedAt time.Time
}
type SendError struct {
Class FailureClass
Code string
Err error
}
func (e *SendError) Error() string { return e.Err.Error() }
func (e *SendError) Unwrap() error { return e.Err }
type Sender interface {
Send(context.Context, Message) (Acceptance, error)
}
type Registry interface {
Suppressed(context.Context, string) (bool, error)
Reserve(context.Context, string, time.Time) (bool, error)
}
func Dispatch(ctx context.Context, now time.Time, r Registry, s Sender, m Message) (Acceptance, error) {
if !now.Before(m.ExpiresAt) {
return Acceptance{}, errors.New("message expired before dispatch")
}
blocked, err := r.Suppressed(ctx, m.Recipient)
if err != nil {
return Acceptance{}, err
}
if blocked {
return Acceptance{}, errors.New("recipient suppressed")
}
reserved, err := r.Reserve(ctx, m.ID, m.ExpiresAt)
if err != nil {
return Acceptance{}, err
}
if !reserved {
return Acceptance{}, errors.New("duplicate logical message")
}
return s.Send(ctx, m)
}
Reserve must be atomic in the backing store. A read followed by a write is not enough; two workers can both observe absence and send. The reservation also needs a deliberate lifecycle. Keep the audit record after the OTP expires even though the credential itself should no longer be usable.
For a concrete test fixture, give an OTP a five-minute lifetime, pause the queue for six minutes, and assert that dispatch returns the expiry error without calling either adapter. The five-minute value is test data, not a universal recommendation; production validity must follow the authentication policy. This one case catches a costly scheduling mistake: retry code that checks its attempt count but never checks whether the message still has meaning.
Callbacks need the same discipline. Verify the signature exactly as documented by the selected provider, retain the raw payload under the organization's data-retention policy, and deduplicate on the provider event identifier when one is guaranteed. If no stable event identifier exists, build a documented composite key from immutable fields. Never acknowledge first and hope an in-memory worker persists the event later.
Compare operational contracts, not feature grids
Twilio, Vonage, and Sinch all document outbound messaging and delivery reporting, but their request fields, callback authentication methods, status vocabularies, regional setup, and error taxonomies are provider-specific. Those are adapter boundaries. They are not a reason to let a vendor callback mutate a user profile directly.
A short evaluation should replay the same cases against each candidate: a valid US destination, a valid EU destination, malformed input, an opted-out recipient, a documented permanent failure, callback duplication, callback reordering, and a timeout after acceptance. Record whether the API exposes a stable message identifier, how signatures are verified, which statuses are final, and how long status evidence remains available. Confirm current country, sender-registration, and traffic-class requirements with the provider and applicable authorities before launch; they cannot be inferred from a generic API comparison.
The selection outcome is a routing policy, not a winner. One provider may be eligible for a particular sender type and geography while another route is ineligible. Preserve that decision as configuration with an owner and review date. Do not make fallback automatic merely because two adapters compile. The limitation of a generic adapter becomes visible here: it can unify Send, but it cannot erase country registration, sender identity, consent, callback security, or final-status semantics. Expose those capabilities in route configuration instead of reducing them to a lowest-common-denominator Boolean.
Verify, deploy, and roll back the runbook
Before production, run contract tests against captured, sanitized callback fixtures. Test signature rejection, duplicate delivery, out-of-order transitions, unknown status codes, database unavailability, and an expired OTP waiting in the queue. Property tests are useful for one invariant: no event sequence may move a final state back to a non-final state.
Deploy callback ingestion in observe-only mode first. It should authenticate, parse, and record events without creating suppressions. Compare normalized outcomes with the raw events, review every unknown mapping, then enable suppression for one documented permanent class at a time. Watch queue age, time from acceptance to final status, duplicate reservation attempts, callback verification failures, unknown-code count, suppression additions, and sends blocked by suppression. Alert on rates and sustained age, not on every individual handset failure.
Rollback has two switches: disable new suppression writes, and disable the affected route. Do not delete the registry. Mark questionable entries for review, retain their evidence, and restore a recipient only through an audited transition. If callbacks are failing verification, hold them for bounded reprocessing after the verifier is corrected; do not accept unsigned events to clear the queue.
Stop there.
For an OTP incident, stop retries when the challenge expires. For an account-change alert, preserve the logical message record and move to an approved non-SMS recovery path if one exists. Reliability means reaching a defensible outcome, not maximizing send attempts.
References
- ITU-T Recommendation E.164, The international public telecommunication numbering plan: https://www.itu.int/rec/T-REC-E.164/en
- NIST Special Publication 800-63B, Authentication and Lifecycle Management: https://pages.nist.gov/800-63-3/sp800-63b.html
- Twilio Messaging status callback documentation: https://www.twilio.com/docs/messaging/guides/track-outbound-message-status
- Twilio webhook request validation documentation: https://www.twilio.com/docs/usage/webhooks/webhooks-security
- Vonage SMS API documentation: https://developer.vonage.com/en/messaging/sms/overview
- Vonage signed webhooks documentation: https://developer.vonage.com/en/getting-started/concepts/signed-webhooks
- Sinch SMS delivery report documentation: https://developers.sinch.com/docs/sms/api-reference/sms/tag/Delivery-Reports/
- Sinch webhook security documentation: https://developers.sinch.com/docs/sms/api-reference/authentication/webhook-signing/
Top comments (0)