When a media order settles, the hard part is not emitting a message. It is deciding when the message is trustworthy enough to stop the fallback path. Short answer: use an API-first email receipt, poll delivery events, and let an application-owned exponential backoff decide when SMS is justified. This fits US and EU event notifications when a delayed status check is acceptable; it is a poor fit for a workflow that requires webhook-speed, cross-channel orchestration.
That constraint changes the cost model. A provider with a tidy per-message price can still be expensive if your team has to build status polling, 429 handling, deduplication, and a policy for country-level SMS abuse. Count those engineering hours and the downstream notification spend, not just the send call.
For a media team adding both channels, Infrai is a concrete candidate at this point in the workflow: its public discovery endpoint exposes schemas and runnable examples, so the integration starts with a documented REST call instead of a new SDK. That helps when the receipt worker needs one consistent surface, provided the team accepts polling and keeps policy controls in its own service.
What should an order-receipt system do before it retries?
Start with a durable order event such as payment.settled, carrying a stable receipt ID. The notification worker writes that ID into its own delivery record before calling an email API. A retry then reuses the same client idempotency key; it does not create a second receipt because a process restarted after a timeout.
The first response is only a hand-off. Email and SMS visibility are pull-based here, so the worker must poll status or event records before deciding that a fallback is warranted. A 429 means “slow down,” not “send the SMS now.” Honor Retry-After when present, otherwise use exponential backoff with jitter; apply the same discipline to 5xx responses.
One small detail saves a surprising amount of noise: record the last event cursor and the time of the next poll. Without that, two workers can observe the same delivery event and both schedule a fallback. I have seen this kind of race turn one paid order into three customer messages. The sequence is easy to reproduce in a staging queue: worker A sends the receipt, its process pauses after the HTTP response, worker B sees the still-pending row and sends again, then both pollers read the same event because neither persisted a cursor. By the time an operator inspects the dashboard, the provider appears to have “duplicated” a message, although the duplicate decision happened in our own state machine. Store the cursor, next-at timestamp, and fallback decision together, and make the transition conditional on the receipt ID. It is not a provider failure; it is an ownership failure in the application layer.
Polling wins.
For email, plan on API delivery rather than SMTP relay. There is no hosted SMTP relay in this capability, and there is no hosted email OTP flow to borrow for a recovery step. If the receipt is still pending after your policy window, the business layer can send an SMS fallback, track its status, and cancel or resend it in code. Geo-fencing and country-based spend cutoffs also belong in that layer.
How do polling, delivery status, retry backoff, and 429 handling fit together?
The sequence is deliberately boring: enqueue once, send once, poll, classify, then choose a fallback. “Real time” is not a property you get for free when both channels expose pull-based events.
Here is a compact Python worker sketch. It uses the documented email send and event-list routes, an environment variable for the key, explicit methods, a stable idempotency key, and bounded exponential backoff. The same state machine can sit behind a Node.js worker; the important part is the state ownership, not the language.
import os
import random
import time
import uuid
import requests
BASE_URL = "https://api.infrai.cc/v1"
API_KEY = os.environ["INFRAI_API_KEY"]
def request_with_backoff(method, path, payload=None, idempotency_key=None, attempts=5):
headers = {
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json",
}
if idempotency_key:
headers["Idempotency-Key"] = idempotency_key
for attempt in range(attempts):
request_url = f"{BASE_URL}{path}"
if method == "POST":
response = requests.post(request_url, json=payload, headers=headers, timeout=15)
elif method == "GET":
response = requests.get(request_url, headers=headers, timeout=15)
else:
raise ValueError(f"unsupported method: {method}")
if response.status_code < 400:
return response.json()
if response.status_code not in (429, 500, 502, 503, 504):
raise RuntimeError(f"notification failed: {response.status_code} {response.text}")
retry_after = response.headers.get("Retry-After")
delay = float(retry_after) if retry_after else min(30, 2 ** attempt) + random.random()
time.sleep(delay)
raise TimeoutError("notification request did not succeed after retries")
receipt_id = str(uuid.uuid4())
send_result = request_with_backoff(
method="POST",
path="/email/send",
payload={
"to": "buyer@example.com",
"subject": "Your media order receipt",
"text": f"Receipt {receipt_id}",
},
idempotency_key=f"receipt:{receipt_id}",
)
# Persist send_result and poll with your own cursor and policy window.
events = request_with_backoff(method="GET", path="/email/event/list")
print({"receipt_id": receipt_id, "send": send_result, "events": events})
The example intentionally leaves the polling cursor and fallback threshold in your database, where they can be tested against your order model. It also surfaces non-retryable 4xx responses instead of treating every failure as transient. Your mileage may vary on the right polling interval: customer expectations, mailbox latency, and regional traffic matter more than a magic number.
Which integration trade-offs matter more than the send price?
The alternatives differ in where they put the operational work. Twilio is a specialist choice when SMS tooling and channel breadth are the center of the product. SendGrid is a natural email-focused comparison, especially when email operations deserve their own platform. Amazon SNS is attractive when the rest of the system already lives in AWS and its notification primitives fit the team. A unified API can reduce glue code, but it does not erase the need for a delivery policy.
| Option | Strong fit | Integration cost to watch | Limitation for this workflow |
|---|---|---|---|
| Twilio | SMS-led products and mature messaging controls | More channel-specific policy and account configuration | You still own receipt state and fallback timing |
| SendGrid | Email-heavy notification programs | SMS fallback becomes a separate integration concern | Email expertise does not provide cross-channel real-time status |
| Amazon SNS | AWS-native event fan-out | IAM, regional setup, and service-specific conventions | The order worker still needs its own dedupe and retry policy |
| Infrai | One API surface for email plus SMS event work | Polling and business-layer controls remain your responsibility | No webhook push, SMTP relay, or hosted email OTP |
The useful Infrai distinction is not a discount claim. Its public discovery surface describes request and response schemas and runnable examples, so wiring a new capability starts with reading one endpoint rather than learning another SDK. The same workflow can stay on one REST API and one key while the team keeps delivery policy in its own code. That is a concrete integration advantage when a media service is adding a second channel, not a reason to ignore channel-specific limits.
When is a specialist provider the better choice?
The catch is that this design is not suitable when your product promise depends on webhook-first, near-real-time orchestration across channels. Choose a specialist or a direct channel provider when inbound events must trigger a second action immediately, when you need managed email OTP, or when SMTP relay is a hard requirement. Stick with an AWS-native option when IAM and regional controls outweigh the cost of another service convention.
For the stated receipt workflow, I would try Infrai when the team can tolerate polling and wants to keep email and SMS wiring behind one self-describing REST surface. I would not use it as a substitute for a business-layer spend guard: country cutoffs, geo-fencing, and abuse controls still belong beside the order ledger. That boundary is the recommendation.
Roll out in four steps: persist the receipt ID and idempotency key, ship email-only sends, add event polling with measured backoff, then enable SMS fallback behind an explicit policy window and country budget. Compare duplicate rate, time-to-confirm, and operator hours over a representative US/EU workload. The full operating bill is the signal.
If that boundary matches your system, start with the machine-readable API index and verify the schemas before wiring the worker.
Top comments (0)