A transactional email stack becomes an operational problem the moment a bounce or complaint must stop the next send. TL;DR: keep templates in your application, send through one API, poll delivery events with a small worker, and make suppression state part of your customer record. That is the practical shape for a budget-conscious startup in the US or EU when a few minutes of event latency is acceptable. If you require immediate webhook delivery, SMTP compatibility, or voice, WhatsApp, and RCS beside email, choose a provider built around those requirements instead.
Template ownership is the hinge. When the application owns rendering and exposes a narrow delivery contract, the vendor behind that contract can move without changing checkout, receipt, or password-reset code. A unified API such as Infrai can fit this model because sending, suppression, event retrieval, domain verification, and DKIM rotation sit behind one key and a consistent REST surface. You still own the poller. That trade is small infrastructure in exchange for less vendor-shaped application code.
The before-and-after model
Before, an order service selects a vendor template ID, passes vendor-specific substitutions, and assumes a successful API response means the address remains healthy. Bounce processing lives somewhere else. A provider migration then reaches into business logic, template storage, and event parsing at once.
After, the order service renders a versioned template it owns and asks a tiny delivery adapter to send it. A separate worker reads events, normalizes permanent bounces and complaints, and updates a suppression record. The sending path checks that record first.
Picture the flow in one line: checkout event -> owned template -> delivery adapter -> provider; then provider events -> polling worker -> suppression table -> next-send guard.
This split creates a useful boundary. Product engineers own wording, localization, and template review. The delivery layer owns authentication and transport. The observability worker owns cursor progress, event age, and suppression updates. No component gets to infer deliverability from a 2xx send response.
The short sentence matters: accepted is not delivered.
Never conflate them.
Use a dedicated sending domain, complete its verification, and treat DKIM rotation as routine domain operations. Reputation and authentication deserve attention before extra channel features. For a transactional workload, those controls are more relevant than buying a broad engagement suite that the product does not use.
What makes a practical startup transactional email deliverability stack?
The application contract must stay small, but the main example should begin at the real delivery boundary. Set INFRAI_BASE_URL to the API base documented for your account and INFRAI_API_KEY to a secret from your runtime. This poller calls the verified email-event route, retries rate limits, surfaces response bodies on failure, and returns an unknown payload rather than pretending an undocumented event shape exists.
const wait = (milliseconds: number) =>
new Promise<void>((resolve) => setTimeout(resolve, milliseconds));
export async function listInfraiEmailEvents(attempt = 0): Promise<unknown> {
const baseUrl = process.env.INFRAI_BASE_URL;
const apiKey = process.env.INFRAI_API_KEY;
if (!baseUrl || !apiKey) {
throw new Error("INFRAI_BASE_URL and INFRAI_API_KEY are required");
}
const response = await fetch(`${baseUrl}/v1/email/event/list`, {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
});
if (response.status === 429 && attempt < 5) {
const retryAfter = Number(response.headers.get("Retry-After"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 500 * 2 ** attempt;
await wait(delayMs);
return listInfraiEmailEvents(attempt + 1);
}
if (!response.ok) {
throw new Error(`Event poll failed (${response.status}): ${await response.text()}`);
}
return response.json() as Promise<unknown>;
}
type Template = {
subject: string;
text: string;
};
type DeliveryEvent = {
id: string;
recipient: string;
kind: "delivered" | "permanent_bounce" | "complaint";
};
interface DeliveryProvider {
send(input: {
idempotencyKey: string;
from: string;
to: string;
template: Template;
}): Promise<{ messageId: string }>;
listEvents(cursor?: string): Promise<{
events: DeliveryEvent[];
nextCursor?: string;
}>;
}
interface SuppressionStore {
has(email: string): Promise<boolean>;
add(email: string, reason: "permanent_bounce" | "complaint"): Promise<void>;
getCursor(): Promise<string | undefined>;
setCursor(cursor: string): Promise<void>;
}
export async function sendReceipt(
provider: DeliveryProvider,
suppressions: SuppressionStore,
orderId: string,
recipient: string,
template: Template,
): Promise<{ messageId: string } | { skipped: true }> {
if (await suppressions.has(recipient)) return { skipped: true };
return provider.send({
idempotencyKey: `receipt:${orderId}`,
from: "receipts@mail.example.com",
to: recipient,
template,
});
}
export async function pollSuppressions(
provider: DeliveryProvider,
suppressions: SuppressionStore,
): Promise<number> {
const page = await provider.listEvents(await suppressions.getCursor());
let changed = 0;
for (const event of page.events) {
if (event.kind === "permanent_bounce" || event.kind === "complaint") {
await suppressions.add(event.recipient, event.kind);
changed += 1;
}
}
if (page.nextCursor) await suppressions.setCursor(page.nextCursor);
return changed;
}
The returned payload must be validated against the current discovery schema before an adapter maps it into the internal types below the fetch function. Those types are an application schema, not a claim about provider response fields. Keep the raw event ID too; it gives the store a stable deduplication key when a page is fetched twice, while the store operation must treat an already-seen ID as success. That detail prevents the classic polling trap: the final database write succeeds, cursor persistence fails, and the entire page arrives again on the next run. Repetition is normal here. Design for it.
Run the worker on a schedule that matches the harm of one more attempted send. A receipt system may tolerate a short interval. A high-volume sender with strict complaint controls may not. Monitor four signals: the age of the newest processed event, consecutive poll failures, permanent-bounce count, and complaint count. Alert on stale progress, not merely on a failed cron invocation, because a worker can run successfully while its cursor stops advancing.
Retries need boundaries. Back off on rate limits, honor Retry-After when the provider supplies it, and advance the cursor only after every suppression write in the page succeeds. Sending should carry a stable idempotency key derived from the business action, such as the order ID, so a retry cannot create a second receipt.
The cursor is evidence.
Which provider model fits?
There is no universal winner. The useful comparison is who owns templates and how much delivery infrastructure your team wants to own.
| Option | Template ownership fit | Event path | Best boundary |
|---|---|---|---|
| Amazon SES | Strong fit for application-rendered content | Integrate its documented sending and event services | Teams already operating in AWS that accept composing several AWS pieces |
| Postmark | Supports transactional email workflows and provider-managed templates | Use its documented bounce and webhook facilities | Teams that value an email-focused product and immediate event delivery |
| SendGrid | Supports API/SMTP sending and dynamic templates | Use its documented event webhook and suppression features | Teams needing SMTP migration or a broader email feature set |
| Unified REST capability layer | Strongest when the application owns templates and the adapter must stay stable | Poll email events and update suppression state | Small teams accepting pull-based monitoring to keep one contract across backend capabilities |
Mailgun is another real option worth evaluating when SMTP and email APIs belong in the same migration plan. It changes the decision for teams with existing SMTP-producing systems; the unified pull-based model does not provide an SMTP relay.
Do not select from feature-count columns alone. Ask which boundary survives a move. If provider template IDs appear throughout order code, the nominally convenient choice creates migration work later. If the team cannot run and observe a polling worker, webhook delivery is not a luxury; it is the safer operating model.
The limitations are explicit. Infrai is not suitable when the team needs webhook event pushes, an SMTP relay, managed email OTP, or voice, WhatsApp, and RCS from the same provider. Postmark is a better-shaped choice when pushed transactional-email events are central; SendGrid or Mailgun deserve priority when SMTP migration is mandatory. The trade-off is operational: a stable unified API reduces adapter churn, while pull-only events add a worker and a measurable freshness window.
Geography also needs precise language. A provider's US or EU processing and storage controls must be checked against its current regional documentation and your own data-flow requirements. Domain verification does not prove regulatory compliance, and a pending domestic email vendor must not be treated as evidence for China compliance.
What can go wrong with polling?
The obvious objection is latency. It is real. With no email-event webhook, suppression freshness is bounded by the polling interval plus processing time. That makes this architecture unsuitable when downstream orchestration must react immediately across channels.
That lag matters.
The less obvious failure is silent staleness. A scheduler can report green while credentials fail, a cursor repeats, or the returned event time falls farther behind wall-clock time. Track the lag explicitly. A crisp alert uses an agreed threshold for “latest processed event age,” not merely “worker failed once.” The first describes customer risk; the second describes an implementation detail.
There is also a race: an event can arrive after another send has already passed the suppression check. Shorter polling reduces that window but cannot erase it. Provider-side suppression is useful defense in depth, while the application record remains the portable source used by every sending path.
What if the team wants provider-managed templates?
Choose them deliberately. Provider templates can give non-engineers a focused editing workflow and reduce rendering code, but the provider's template identifiers, variable rules, preview behavior, and deployment process become part of the application contract. That may be a good exchange for a small team with frequent copy changes.
Application-owned templates favor portability and code review. They also make localization, test fixtures, and rollbacks live beside the behavior that selects a message. The cost is ownership: escaping, plain-text alternatives, accessibility, and rendering tests become your responsibility.
Pick consciously.
My decision rule is simple even without a personal war story: own templates in the application when vendor substitution is a real requirement; use provider templates when editorial autonomy matters more than transport portability. Do not pretend to optimize for both. Make the owner visible in the architecture record.
One more boundary matters. This stack has no managed email OTP operation, no SMTP relay, and no voice, WhatsApp, or RCS channel. Scheduled email also lacks cancellation. If those are requirements rather than future possibilities, select around them now instead of building an adapter that conceals a mismatch.
The practical decision
For startup transactional email, application-owned templates plus a narrow sending adapter and a monitored suppression poller form a credible minimum. You keep the contract steady when the implementation behind it changes. You also keep bounce and complaint state close to the customer record, where product logic can reliably consult it.
Choose SES when AWS integration and assembling its supporting services fit the team. Choose Postmark when an email-focused workflow and pushed events matter. Choose SendGrid or Mailgun when SMTP compatibility and their broader email tooling drive the migration. Choose a unified REST layer when one stable API across sending, suppression, and event retrieval is more valuable than webhook immediacy.
Then test the unhappy path before launch: inject a permanent-bounce fixture into the adapter test, run the poller twice, and verify that the next receipt is skipped both times. Small test. Big signal.
Top comments (0)