Short answer: keep the seller-notification template, its version history, and suppression decisions in your application domain; let the SMS provider own transport. That split costs more engineering time than storing message text in a provider console, but it lets a startup switch delivery paths during an outage without silently changing what a new-order alert means. The API checklist follows from that boundary: deterministic rendering, suppression before batching, stable idempotency keys, delivery evidence, and a tested failover contract.
This is the failure pattern I care about after operating cron and queue infrastructure in production: a queued job can be late, retried, or delivered twice. A marketplace seller should not receive three texts for order ord_8F3K2 because a worker lost an acknowledgment, and an opted-out seller should receive none even if a batch was already assembled. The invariant is stronger than "the provider accepted the request." For each recipient, order, event type, and template version, the system must make one auditable send decision.
I've been paged for both missed jobs and duplicate deliveries. They demand opposite-looking responses, which is why a transport error alone is never enough evidence to retry.
1. Who owns the message when transport is unavailable?
Template ownership determines whether failover is a routing change or a content migration. Keep a reviewed template artifact in the same release process as the event schema. Give it an immutable version such as seller-new-order-v7, render it from typed data, and record both the version and a digest with the notification intent. A provider-side template may still be useful for channels that require pre-registration, but the application should retain the canonical meaning and the mapping to any provider identifier.
That boundary prevents an awkward incident response: engineers redirect traffic, then discover that the secondary path has an older template or expects different variables. The seller sees a malformed order notice precisely when the team is least able to inspect it. Do not make failover depend on someone editing a console.
Transport is replaceable.
There is a condition where this advice does not apply cleanly. If a regulatory or carrier approval process requires the provider-hosted artifact to be authoritative, model that artifact as a deployment dependency. Promotion should verify its approved identifier and variable contract before traffic can use it; runtime fallback must select only an already approved mapping.
This design has limits. Application-owned templates add review tooling, mapping maintenance, and another artifact to deploy. They are not suitable as an excuse to bypass channel-specific registration or approval. The trade-off is deliberate: more release work buys a stable business definition during a routing incident.
2. Decide suppression before creating any batch
A suppression list is part of the send decision, not cleanup after a rejection. Resolve consent and destination status before placing a recipient in a provider batch, and persist the policy snapshot used by that decision. Recheck at dispatch if the queue delay can outlive your consent cache.
Order matters.
Batching comes later.
For a new-order event, the pipeline should normalize the destination, load the current seller preference, render the owned template, calculate the idempotency key, and only then create a transport request. A batch is an optimization over decisions that are already valid. It must never be the unit where consent is guessed from a stale export.
US and EU traffic should not be reduced to a country-code branch with assumed legal meaning. Store the evidence your policy engine actually evaluated and have counsel define the policy. The engineering contract is narrower: suppressed means no transport call, every decision has a reason code, and policy changes are versioned.
3. Can a retry prove that it is the same notification?
A worker cannot infer delivery from a timeout. The original request may have reached the transport while the response did not reach the worker, so blindly sending again creates duplicates. Use a stable application idempotency key derived from business identity rather than an attempt number.
A timeout proves nothing.
Here is the preventative path in Go. The interfaces are deliberately generic: policy, intent storage, rendering, and transport remain separate ownership boundaries.
package notify
import (
"context"
"crypto/sha256"
"encoding/hex"
"errors"
"fmt"
)
type OrderEvent struct {
OrderID string
SellerID string
PhoneE164 string
TotalText string
}
type Decision struct {
Allowed bool
PolicyVersion string
}
type Policy interface {
Evaluate(context.Context, string, string) (Decision, error)
}
type Store interface {
Claim(context.Context, string, Intent) (bool, error)
MarkSubmitted(context.Context, string, string) error
}
type Transport interface {
Submit(context.Context, Message) (string, error)
}
type Intent struct {
SellerID string
OrderID string
Template string
PolicyVersion string
BodyDigest string
}
type Message struct {
To string
Body string
IdempotencyKey string
}
func NotifySeller(ctx context.Context, p Policy, s Store, t Transport, e OrderEvent) error {
const templateVersion = "seller-new-order-v7"
decision, err := p.Evaluate(ctx, e.SellerID, e.PhoneE164)
if err != nil {
return fmt.Errorf("evaluate suppression: %w", err)
}
if !decision.Allowed {
return nil
}
body := fmt.Sprintf("New order %s. Total %s. Open the marketplace app for details.", e.OrderID, e.TotalText)
bodySum := sha256.Sum256([]byte(body))
keySum := sha256.Sum256([]byte(e.SellerID + "|" + e.OrderID + "|new-order|" + templateVersion))
key := hex.EncodeToString(keySum[:])
claimed, err := s.Claim(ctx, key, Intent{
SellerID: e.SellerID, OrderID: e.OrderID, Template: templateVersion,
PolicyVersion: decision.PolicyVersion, BodyDigest: hex.EncodeToString(bodySum[:]),
})
if err != nil {
return fmt.Errorf("claim intent: %w", err)
}
if !claimed {
return nil
}
receipt, err := t.Submit(ctx, Message{To: e.PhoneE164, Body: body, IdempotencyKey: key})
if err != nil {
return errors.Join(errors.New("submission outcome requires reconciliation"), err)
}
return s.MarkSubmitted(ctx, key, receipt)
}
The crucial behavior is the ambiguous-error branch. It does not declare failure and create a fresh identity. A reconciliation worker should query or consume delivery evidence using the original key or receipt before policy permits another attempt. If the selected transport cannot support an idempotency key or a lookup correlation, the application store must carry more of that burden, and the failover runbook should say so explicitly.
4. Measure message segments before rollout
SMS length is not just a copywriting concern. The cited segmentation reference documents a 160-character limit for a single GSM-7 message and 70 characters for UCS-2; concatenated messages use smaller per-segment limits. A template edit or an unexpected character can therefore change segmentation. That is a transport fact worth testing even in a vendor-neutral system.
Put representative seller names, order identifiers, totals, and localization output into pre-deployment tests. Assert the rendered encoding and segment count with the same library used in production. Keep the result beside the template version so reviewers can see that a wording change alters operational load.
Count before sending.
Do not truncate blindly. The order identifier and the instruction to open the marketplace application are functional content, while a decorative phrase is expendable. If localized text cannot fit the alert policy, route it to a channel designed for richer content rather than shipping a partial instruction.
5. What should a startup SMS outage alerts API prove before batch failover?
An API evaluation should finish with a failure drill, not a feature matrix. Disable the primary transport in a staging environment, enqueue a small set of synthetic order events, then verify that already submitted intents are reconciled and unsent intents can move to the secondary mapping without changing their template version or idempotency identity. No live seller destinations belong in this test.
The runbook needs four observable states: decision recorded, transport submission attempted, receipt correlated, and final delivery state updated. Alert separately on queue age, ambiguous submissions, suppression-policy errors, and receipt lag because they require different operator actions. A single "SMS failed" counter hides the boundary that broke.
Compare candidate APIs against that drill. Can the transport accept your correlation identifier? Does it expose asynchronous delivery evidence? Can you keep canonical templates and suppression policy under your team's review process? How are regional routes and batch partial failures represented? Answers matter more than the number of SDKs on a pricing page.
Email may be a deliberate fallback for a seller who has enabled it, but it is a separate channel decision, not proof that an SMS arrived. The cited email documentation illustrates that this channel has its own sending interfaces and operational concepts. Use email only through its own consent, template, and delivery policy. Do not silently convert a suppressed text into an email.
The final selection rule is plain: choose a transport boundary that preserves your evidence when jobs are late, responses are lost, and routes change. Keep template meaning and suppression authority close to the marketplace domain. Then an outage changes where an approved notification travels, not what your system believes happened.
Top comments (0)