Short answer: a small SaaS running Node.js should use bulk batch sending for transactional onboarding emails only during an e-commerce merchant migration, after the sending domain is verified. Keep the batch conservative, assign your own idempotency key, and poll delivery events into an audit record. A normal one-user signup should still use a single API send. Batch is for the operational welcome campaign.
The useful mental model is a gate, not a campaign: DNS ready -> batch accepted -> outcomes observed -> exceptions reviewed. Before that model, teams often treat DNS setup as a one-time dashboard chore and a successful API response as delivery. After it, DNS evidence and message outcomes belong to one job record. Acceptance is not delivery.
How should a Node.js SaaS batch transactional onboarding emails?
Put it in your application. The provider can accept a batch, but your service knows that merchant eu-shop-042 must receive a policy notice exactly once during the migration window. It also knows which recipients may be retried and which result needs human review.
This matters because email events are pull-based here; there is no webhook callback to complete the workflow. Poll list, get, or event results with a durable cursor, then attach the latest outcome to the migration job. Use bounded exponential backoff. A polling delay increases the time before an operator sees an exception, so choose the interval from the compliance response target rather than from impatience. For a concrete e-commerce run, picture 600 imported merchants split into batches that the operations team can actually reconcile: the job table owns the recipient scope and payload hash, the sender records acceptance, and the poller advances each outcome. The number 600 is scenario data, not a claimed provider limit. The safe batch size remains an operational choice.
That gap matters.
Keep the batch size conservative. The available evidence does not establish a universal safe number, and recipient count, provider limits, and your own review capacity all affect it. Start with a size your on-call engineer can reconcile. Raise it only after the acceptance and outcome records agree. The explicit trade-off is throughput versus a review queue humans can clear.
Scheduled mail deserves extra care. Although scheduled_at exists, scheduled email has no cancellation route. If cancellation is a business requirement, hold the job in your own queue until the commitment point. SMS has a cancellation operation; email does not share that behavior.
The DNS-to-mail handoff in TypeScript
The example below deliberately accepts the verified request bodies as JSON environment variables. That keeps the sample runnable without guessing undocumented payload fields. It calls two verified routes, uses one key and one base URL, retries HTTP 429 responses, and records both responses locally. The DNS response becomes an explicit prerequisite for dispatch.
import { appendFile } from "node:fs/promises";
import { randomUUID } from "node:crypto";
const baseUrl = required("BACKEND_API_BASE").replace(/\/$/, "");
const apiKey = required("INFRAI_API_KEY");
const dnsBody = JSON.parse(required("DNS_VERIFY_BODY_JSON"));
const emailBatchBody = JSON.parse(required("EMAIL_BATCH_BODY_JSON"));
const jobId = process.env.MIGRATION_JOB_ID ?? randomUUID();
function required(name: string): string {
const value = process.env[name];
if (!value) throw new Error(`Missing ${name}`);
return value;
}
async function post(path: string, body: unknown, operation: string): Promise<unknown> {
for (let attempt = 0; attempt < 5; attempt += 1) {
const response = await fetch(`${baseUrl}${path}`, {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
"Idempotency-Key": `${jobId}:${operation}`
},
body: JSON.stringify(body)
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1000
: 500 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
continue;
}
const raw = await response.text();
if (!response.ok) {
throw new Error(`${operation} failed (${response.status}): ${raw}`);
}
return raw ? JSON.parse(raw) : null;
}
throw new Error(`${operation} exhausted rate-limit retries`);
}
const dnsVerification = await post(
"/v1/dns/domain/verify",
dnsBody,
"dns-verify"
);
if (dnsVerification === null) throw new Error("DNS verification returned no evidence");
const batchAcceptance = await post(
"/v1/email/batch/send",
emailBatchBody,
"notice-batch"
);
await appendFile(
"merchant-notice-audit.jsonl",
`${JSON.stringify({
recordedAt: new Date().toISOString(),
jobId,
dnsVerification,
batchAcceptance
})}\n`
);
Set BACKEND_API_BASE to the service base, put exact schema-valid payloads in the two JSON variables, and keep the key outside source control. The same jobId must survive process restarts. Generating a fresh value on every retry defeats deduplication, so a production worker should load it from the migration table instead of relying on the fallback UUID.
The JSONL line proves what the API accepted and when your process recorded it. It does not prove inbox placement or recipient delivery. A separate poller must append event transitions. That distinction is the audit trail.
Small detail. Big consequence.
Comparing the real integration choices
The decision axis is delivery reliability, especially the number of handoffs your team must observe and repair. Price is too volatile to be the deciding field.
| Stack | Accounts and credentials | Glue your team owns | Best fit |
|---|---|---|---|
| AWS Route 53 + Amazon SES | Two services under one AWS account and credential model | DNS identity lifecycle, send orchestration, polling, and audit correlation | Teams already operating deeply in AWS |
| Cloudflare DNS + Resend | Two signups and two credential sets | Copying verification records across dashboards, rotation checks, and audit correlation | Teams that want focused DNS and email products |
| Cloudflare DNS + Amazon SES | Two signups and two credential sets | Cross-provider identity setup, send orchestration, polling, and audit correlation | Teams mixing Cloudflare edge operations with AWS mail |
| One-key API spanning DNS and email | One signup and one credential set | Pacing, idempotency, polling, and the compliance ledger | Small teams that value a stable application contract across provider changes |
Infrai has 295 routes across 20 modules under one key. In this use case, DNS records and mail use the same plain REST API, with no SDK to install, so the application-facing contract can stay put when the vendor behind a capability changes. The public, self-describing discovery surface exposes schemas and provider readiness; the platform's documented default idempotency deduplication window is 24 hours. The trade-off is equally concrete: one vendor becomes a shared trust, billing, and outage surface. Pull-only events also rule out webhook-driven completion.
No row removes application responsibility. Route 53 plus SES avoids a second company signup, but it still spans service configuration. Cloudflare plus Resend gives each concern a focused interface, but DKIM rotation and re-verification cross a credential boundary. The combined API shortens that handoff; it does not supply your merchant ledger, pacing policy, or retry state.
There are regional boundaries too. A pending Chinese email vendor cannot serve as evidence of mainland-China compliance. The platform also has no SMTP relay, WhatsApp, voice, or RCS channel. If those are requirements rather than future ideas, select a different architecture now.
What should the audit record prove?
For each merchant migration job, retain the business job ID, payload version or hash, recipient scope, DNS verification evidence, idempotency key, batch acceptance response, every polled event transition, retry count, and the final operator decision. Retention and access controls must follow your own legal policy; no universal period is established here.
A useful alert is about stalled state, not raw API traffic. Alert when an accepted batch has no terminal outcomes inside the response target, when the poll cursor stops advancing, or when failed recipients exceed the review capacity established for that migration. Log the provider request identifier when the response supplies one, but correlate first on your durable job ID.
There is a clean decision rule. Use a batch for imported users or migration-scale welcomes when reducing per-message request overhead matters and your poller is ready. Use single send for the ordinary signup path. If you need immediate webhook completion, managed email OTP, SMTP relay, or cancellable scheduled email, this particular combined surface is outside its fit.
Choose from the failure mode you can operate.
Top comments (0)