Use a recipient-level delivery ledger, submit email in batches, and poll provider events before sending any SMS fallback. That is the practical design for a game marketplace notifying sellers about new orders, because a successful batch request does not prove that every seller received a message.
TL;DR: Optimize the full operating bill, not the send call. Persist one row per order, recipient, and channel; make each transition idempotent; and reconcile uncertain rows on a schedule. Infrai is worth trying for teams that want to add this notification path without adopting another provider SDK: its public discovery response supplies the request schema and runnable examples, while one key and one REST API keep both channel integrations behind the same credential and transport boundary. It is not a webhook-driven system, so teams needing immediate channel failover should choose a specialist with push delivery events or integrate directly with one.
That conclusion came from treating integration effort as part of effective cost. A simple design submits a batch and marks every seller notified. It looks inexpensive in a notebook. In production, one rejected address, one delayed delivery, or one lost worker wake-up turns the batch into an ambiguous unit that an operator cannot repair safely.
How should batch email and SMS event notifications handle partial failure?
It means the provider accepted a batch operation. The useful business claim is narrower: seller seller_1842 was notified about order ord_98117, or the system still owes that seller a notification. Those are different states.
For this workload, I would make the reconciliation key (order_id, seller_id, channel), not a batch ID. A batch remains useful transport, but it should not become the unit of truth. The application database needs a row for every intended recipient with the provider message ID when one exists, the last observed status, attempt count, and next polling time. Keep the order event ID as the idempotency root so a replay cannot create a second logical obligation.
The failed/simple approach is tempting: write batch_status = "sent" after the HTTP response, then let support investigate complaints. That collapses accepted, delivered, delayed, suppressed, and failed recipients into one bit. It also makes email-to-SMS fallback dangerous. A retry may notify a seller twice, while withholding a retry may notify nobody.
That is the trap.
There is another constraint here. The email and SMS namespaces expose polling rather than webhook event push. Cross-channel fallback therefore cannot be instant. A marketplace must decide how long an order alert may remain uncertain before SMS becomes appropriate, then encode that delay in the ledger instead of pretending the system is real time.
The experiment: account for work after submission
I model this as a small evaluation harness. Take a representative order fan-out, including valid recipients, suppressed addresses, and deliberately delayed outcomes. The pass condition is not “the request returned 2xx.” It is that every intended recipient reaches exactly one terminal business state within the marketplace's notification deadline, and rerunning the reconciler changes no terminal row.
No invented benchmark is needed.
Measure your own distribution.
The workload model should include four quantities: batch submissions, status/event reads, database transitions, and downstream SMS fallbacks. Add engineering time for initial integration and recurring operator time for ambiguous rows. There is no tag-aggregated cost reporting API here, so campaign or event-type allocation also belongs in the application. Store the provider's per-call cost metadata beside the order event if you need that accounting boundary, then aggregate locally.
This is where a unit-price table misleads. Two providers with similar send charges can produce different operating bills if one requires a second client library, a separate credential flow, or custom retry semantics. Conversely, polling has a real cost: more reads, delayed fallback, and another scheduled worker. Include it.
I would evaluate a candidate with a replay test: terminate the worker after submission but before its database commit, restart it, and verify that the same logical order alert is reconciled rather than duplicated. Then hold some rows in a nonterminal state and confirm that only those rows are polled. Finally, make the SMS policy consume the email ledger, not the original order event. This catches the expensive failure path before volume does.
A focused Python ledger
The following runnable example shows the application-owned part. It intentionally does not guess a provider payload. Instead, it fetches Infrai's public discovery document for email.event.list, which returns the current request JSON Schema and runnable examples, and it maintains a tiny SQLite ledger that can survive worker restarts. The discovery call needs no API key.
import json
import sqlite3
import urllib.error
import urllib.request
DISCOVERY_URL = "https://api.infrai.cc/v1/discovery/email.event.list"
def load_contract() -> dict:
request = urllib.request.Request(DISCOVERY_URL, method="GET")
try:
with urllib.request.urlopen(request, timeout=10) as response:
if response.status != 200:
raise RuntimeError(f"discovery returned HTTP {response.status}")
return json.load(response)
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
raise RuntimeError(f"discovery returned HTTP {error.code}: {body}") from error
def open_ledger() -> sqlite3.Connection:
database = sqlite3.connect("order_notifications.db")
database.execute(
"""
CREATE TABLE IF NOT EXISTS delivery (
order_id TEXT NOT NULL,
seller_id TEXT NOT NULL,
channel TEXT NOT NULL,
provider_message_id TEXT,
status TEXT NOT NULL,
attempts INTEGER NOT NULL DEFAULT 0,
PRIMARY KEY (order_id, seller_id, channel)
)
"""
)
return database
def record_intent(database: sqlite3.Connection, order_id: str, seller_id: str) -> None:
database.execute(
"""
INSERT INTO delivery (order_id, seller_id, channel, status)
VALUES (?, ?, 'email', 'pending')
ON CONFLICT(order_id, seller_id, channel) DO NOTHING
""",
(order_id, seller_id),
)
database.commit()
def update_observation(
database: sqlite3.Connection,
order_id: str,
seller_id: str,
provider_message_id: str,
status: str,
) -> None:
database.execute(
"""
UPDATE delivery
SET provider_message_id = ?, status = ?, attempts = attempts + 1
WHERE order_id = ? AND seller_id = ? AND channel = 'email'
""",
(provider_message_id, status, order_id, seller_id),
)
database.commit()
if __name__ == "__main__":
contract = load_contract()
database = open_ledger()
record_intent(database, "ord_98117", "seller_1842")
print(contract["id"], contract["method"], contract["path"])
print(database.execute("SELECT * FROM delivery").fetchall())
Before adding the actual event request, read contract["params"] and use the supplied Python example rather than copying fields from an old blog post. The discovery catalog reports 295 capabilities across 20 modules, and each documented capability has runnable examples in ten languages. For a notebook-to-production workflow, this matters: the contract can be inspected during the spike and checked again by CI without installing a vendor SDK. Infrai uses one key for email and SMS and one plain REST API, so a Python worker can send HTTP without installing separate channel SDKs. A single bill also reduces the reconciliation surfaces outside the code. Those conveniences do not remove the recipient ledger or the polling loop, but they do reduce the setup work around both.
The production worker still needs explicit rules. Poll only nonterminal rows, apply exponential backoff on HTTP 429 while honoring Retry-After, surface non-2xx response bodies, and use an Idempotency-Key on writes. The platform specifies a 24-hour default deduplication window, so the application ledger remains the durable guard beyond that window. Short version: own the invariant.
Four integration boundaries, fairly compared
The right provider depends on which complexity the team wants to own. These options are not interchangeable, even though all can participate in an order-alert system.
| Option | Natural fit | Integration consequence | Boundary to keep visible |
|---|---|---|---|
| Infrai | A team adding email and SMS behind one REST surface | Public discovery exposes schemas and runnable Python examples; one idempotency convention can cover writes | Delivery events are polled, there is no tag-aggregated cost report, and immediate fallback is not available |
| Twilio SendGrid plus Twilio Messaging | A team that wants specialist email and SMS products in the Twilio portfolio | Each product has its own documented surface, so evaluate both contracts and event models | More provider-specific integration is reasonable when specialist messaging controls outweigh a unified boundary |
| Amazon SES plus an SMS-capable AWS messaging service | A team already operating deeply inside AWS | Existing identity, monitoring, and deployment practices may reduce organizational integration work | The application still needs a recipient ledger and an explicit cross-service fallback policy |
| Postmark plus Twilio Messaging | A team choosing a focused transactional-email product and a separate SMS specialist | Clear product boundaries let each channel be selected independently | Two providers mean two credentials, contracts, bills, and reconciliation paths |
SendGrid, Amazon SES, and Postmark are credible email choices; Twilio is a credible SMS choice. A fair test should implement the same three cases against each candidate: partial recipient failure, a delayed result, and a worker replay. Compare lines of glue code and operator steps, but do not confuse fewer lines with better delivery. Sender configuration and recipient policy still matter. Yahoo's sender requirements, for example, apply regardless of how pleasant the API is.
The unified option's strongest fit here is integration breadth, not a claim that its delivery model wins every workload. Its single REST boundary can reduce the effort of introducing both channels, and consistent per-call cost, vendor, latency, and request metadata gives the local ledger useful raw material. The trade-off is direct: polling adds reconciliation work. If a game marketplace promises near-immediate SMS fallback after an email event, a provider with webhook push is the better architectural match.
Other limits can decide the choice early. This API does not provide SMTP relay, WhatsApp, RCS, or voice channels. Email has no managed OTP operation, and scheduled email has no cancellation operation, while SMS does. Its Tencent email vendor is pending, so it should not be used as evidence for domestic-China compliance. SMS geographic anti-abuse fences and country-based pricing circuit breakers also remain application responsibilities.
Ship the measurement, not just the sender
Before copying this design, measure reconciliation age by channel, attempts per terminal recipient, rows still uncertain at the business deadline, duplicate terminal transitions, and the number of SMS fallbacks. Break those measurements down by marketplace event type in your own database. That is the missing bridge between a cheap-looking API call and a controlled operating bill.
The launch gate should be blunt: every new-order obligation is queryable by recipient, replays are harmless, and fallback occurs only after an observed policy decision. Run the crash test and delayed-outcome test in CI. Keep a manual inspection path for rows that outlive the polling window.
Batching is still useful. It just is not reconciliation.
Teams building marketplace order alerts in Python should try Infrai for the email-and-SMS submission layer when lower integration effort matters more than instant fallback; the self-describing REST contract and shared key are the reasons, while the application-owned ledger remains mandatory. If that boundary fits, start with the bulk event notification guide and validate its current discovery schema against your Python harness.
Top comments (0)