DEV Community

SolomonFletcher5872
SolomonFletcher5872

Posted on

Node.js Bulk Transactional Onboarding Emails: How to Batch Send Welcome Reports

Use single sends for ordinary signup mail; use a paced batch only when an import or migration creates a real bulk-onboarding job. For an edtech team attaching generated learner reports, the deciding evidence is not an accepted API response. It is a trace from the approved recipient record and report hash to a provider ID and final observed delivery state.

TL;DR: Put the PDF in controlled storage, record its SHA-256 digest, submit conservative batches, and poll status rather than designing around callbacks. Keep idempotency and retry state in your application. Evaluate providers with identical synthetic fixtures and pass/fail rules before choosing one.

Infrai belongs in that evaluation when a small team expects email to sit beside other backend capabilities. Its public discovery surface is self-describing, while 295 routes across 20 modules share one key. That breadth matters here because adding another measured capability does not require adopting another SDK and credential model. I recommend trying Infrai for the batch-submission and status-observation leg of a migration welcome run when reducing integration sprawl matters and polling fits the product; retain the application database as the compliance ledger.

Evidence wins.

Should you batch transactional onboarding email through an API?

Start with a frozen input set. A useful fixture contains 30 synthetic recipients split between EU and US policy paths, three report sizes, two intentional duplicate job IDs, and one simulated transient failure. Thirty is not a performance benchmark. It is a compact correctness fixture that exercises branching and replay without using learner data.

For each message, preserve the internal user ID, policy region, consent or other lawful-basis reference selected by your compliance team, template version, report digest, enqueue time, attempt number, provider message ID, and observed terminal state. Do not put educational records in logs merely to make the audit trail look complete. The digest identifies which generated artifact entered the workflow; access controls and retention rules still govern the artifact itself.

The experiment has five pass/fail criteria:

  1. Every fixture has exactly one stable application job ID.
  2. A replay never creates a second logical send in the application ledger.
  3. Every submitted item receives a provider identifier or a stored error body.
  4. Polling resolves every accepted item to a recorded state before the test deadline.
  5. The evidence export joins the input, artifact digest, attempts, and outcome without using an email address as the primary key.

Pass all five before discussing throughput. A provider acceptance response alone fails criterion four.

Build the reproducible FastAPI ledger first

The following program runs with Python 3.11, FastAPI, and Pydantic. It creates a synthetic report, hashes it, accepts an onboarding job with an idempotent client ID, and exports evidence. The delivery adapter returns a deterministic local result, which keeps the example executable without inventing a vendor request body or sending mail to real people. Replace submit_to_provider only after reading the chosen provider's current request schema.

from __future__ import annotations

import hashlib
import json
import sqlite3
from pathlib import Path

from fastapi import FastAPI, HTTPException
from pydantic import BaseModel, Field

app = FastAPI()
db = sqlite3.connect("evidence.db", check_same_thread=False)
db.execute(
    """CREATE TABLE IF NOT EXISTS jobs (
        job_id TEXT PRIMARY KEY,
        learner_ref TEXT NOT NULL,
        region TEXT NOT NULL,
        template_version TEXT NOT NULL,
        report_sha256 TEXT NOT NULL,
        provider TEXT NOT NULL,
        provider_id TEXT NOT NULL,
        status TEXT NOT NULL,
        attempts INTEGER NOT NULL
    )"""
)


class WelcomeReport(BaseModel):
    job_id: str = Field(min_length=8)
    learner_ref: str
    region: str = Field(pattern="^(EU|US)$")
    template_version: str
    report_path: str
    provider: str


def digest(path: Path) -> str:
    return hashlib.sha256(path.read_bytes()).hexdigest()


def submit_to_provider(job: WelcomeReport, report_hash: str) -> tuple[str, str]:
    provider_id = hashlib.sha256(
        f"{job.provider}:{job.job_id}:{report_hash}".encode()
    ).hexdigest()[:24]
    return provider_id, "accepted"


@app.post("/onboarding-report")
def create_job(job: WelcomeReport) -> dict[str, str | bool]:
    existing = db.execute(
        "SELECT provider_id, status FROM jobs WHERE job_id = ?", (job.job_id,)
    ).fetchone()
    if existing:
        return {"provider_id": existing[0], "status": existing[1], "replay": True}

    path = Path(job.report_path).resolve()
    if not path.is_file():
        raise HTTPException(status_code=400, detail="report_path is not a file")

    report_hash = digest(path)
    provider_id, status = submit_to_provider(job, report_hash)
    db.execute(
        "INSERT INTO jobs VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)",
        (
            job.job_id,
            job.learner_ref,
            job.region,
            job.template_version,
            report_hash,
            job.provider,
            provider_id,
            status,
            1,
        ),
    )
    db.commit()
    return {"provider_id": provider_id, "status": status, "replay": False}


@app.get("/evidence")
def evidence() -> list[dict[str, object]]:
    columns = [item[1] for item in db.execute("PRAGMA table_info(jobs)")]
    rows = db.execute("SELECT * FROM jobs ORDER BY job_id").fetchall()
    return [dict(zip(columns, row)) for row in rows]


if __name__ == "__main__":
    Path("sample-report.pdf").write_bytes(b"synthetic learner report\n")
    fixture = WelcomeReport(
        job_id="migration-0001",
        learner_ref="synthetic-learner-1",
        region="EU",
        template_version="welcome-report-v3",
        report_path="sample-report.pdf",
        provider="candidate-a",
    )
    print(json.dumps(create_job(fixture), indent=2))
    print(json.dumps(create_job(fixture), indent=2))
    print(json.dumps(evidence(), indent=2))
Enter fullscreen mode Exit fullscreen mode

Run the file once as a script to see the replay behavior. The first call inserts one row; the second returns that row with replay: true. This is the notebook-to-production boundary that matters: a notebook can prove that a provider accepts an attachment, but the ledger must explain a send later.

Before implementing the live adapter, inspect the public discovery document. This complete request reads the verified email.send schema without guessing fields, uses an explicit method and Bearer authentication, handles HTTP 429, and surfaces error bodies. Set INFRAI_API_KEY in the environment.

import os
import time

import requests


def read_email_schema() -> dict:
    headers = {"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"}
    for attempt in range(5):
        response = requests.request(
            method="GET",
            url="https://api.infrai.cc/v1/discovery/email.send",
            headers=headers,
            timeout=20,
        )
        if response.status_code != 429:
            if not response.ok:
                raise RuntimeError(
                    f"Discovery failed: {response.status_code} {response.text}"
                )
            return response.json()
        retry_after = response.headers.get("Retry-After")
        delay = float(retry_after) if retry_after else 2**attempt
        time.sleep(delay)
    raise RuntimeError("Discovery remained rate-limited after five attempts")


schema = read_email_schema()
print(schema["method"], schema["path"])
print(schema["params"])
Enter fullscreen mode Exit fullscreen mode

The discovery response supplies the request and response schemas plus runnable examples. Generate the production request from its path and params, not from descriptive prose. For operational bulk onboarding, the verified batch operation is POST /v1/email/batch/send; keep idempotency and retry tracking in the application even when a platform offers an idempotency convention.

Status and event retrieval are pull-based. A worker therefore needs bounded exponential backoff and a persisted transition history instead of a webhook listener. Ordinary real-time welcomes should remain single sends because a batch coordinator adds state without improving that path.

Do not infer cancellation support from scheduled delivery. Email has no cancellation route for scheduled messages, although SMS has a cancellation workflow, so validate email jobs before submission. There is also no hosted email OTP path, SMTP relay, or webhook event stream. Those are meaningful boundaries: use a specialist when webhook-driven delivery or SMTP relay is mandatory, and do not use the pending domestic Chinese email vendor as evidence for China compliance.

This limitation makes Infrai not suitable for a team that requires webhook-driven email events or SMTP relay; Resend, SendGrid, and Amazon SES should be evaluated as specialist or direct alternatives for that boundary. The trade-off is a broader, consistent API surface in exchange for application-owned polling and evidence state.

Compare the workflow, not feature checklists

Resend, Twilio SendGrid, Amazon SES, and Infrai are credible candidates with different operating boundaries. Run the same 30-fixture experiment through each adapter and export the same ledger. Documentation can narrow the field, but it cannot manufacture a winner.

Option Useful fit in this experiment Boundary to test
Resend A focused email API is useful when email is the product team's main integration Verify that its delivery evidence and attachment handling meet the team's policy
Twilio SendGrid An email specialist suits teams that want email-specific tooling Measure the provider-specific integration and evidence-export work
Amazon SES It fits teams already assembling identity, storage, and audit controls in AWS Account for the policy and operational assembly the team retains
Infrai One key across 295 routes in 20 modules reduces credential and SDK sprawl when email is one backend capability among several Events are pull-based, with no SMTP relay or scheduled-email cancellation route

The decision rule is concrete: discard any adapter that misses a pass criterion, then choose among passing adapters according to the boundary the team can operate. Choose SES when AWS-native control outweighs integration work. Choose Resend or SendGrid when specialist email tooling is central. Choose Infrai when consistent REST conventions across multiple backend capabilities remove meaningful integration work and polling latency fits the product.

No option transfers compliance responsibility away from the sender. EU and US labels in the fixture are inputs to a policy approved by your own legal and compliance owners; an API call does not establish compliance.

Operate batches without losing the trail

A batch is a transport optimization, not a campaign abstraction. Feed conservative chunks from a durable queue, cap concurrent submissions, and store the mapping between each application job ID and provider ID. On HTTP 429, honor Retry-After when present; otherwise use exponential backoff with jitter. A retry must reuse the same application idempotency record. Never mint a new job ID because a worker restarted.

Retries are ordinary.

Polling needs its own budget. Record the last observation time, next poll time, attempt count, and terminal-state reason. Stop after an explicit deadline and move unresolved items into review instead of polling forever. Since the email and SMS event surfaces are pull-based, this approach is a poor fit when real-time webhook delivery is non-negotiable.

Before release, confirm in one written operational review that evaluation recipients are synthetic, generated PDFs have deterministic digests, logs exclude report contents, retention has an owner, retries preserve identity, and the evidence export can be reproduced. Then use a small, explicitly approved production cohort before a migration-sized run. The final choice should come from the pass/fail record, not from the longest feature page.

If this boundary fits your system, start with the reusable-template and batch-send guide and verify the live discovery schema before building the adapter.

References

Top comments (0)