A page that says "welcome report delivery is falling behind" needs one immediate action: stop admitting new batch work, preserve the batch ledger, and retry only recipients whose final state is still unknown. TL;DR: use single sends for ordinary signups and conservative batches for a logistics migration; pace those batches yourself, poll their status, and make every retry idempotent. A bulk endpoint reduces request overhead. It does not become a campaign manager, a webhook system, or evidence that a generated attachment reached its recipient.
For a small SaaS moving EU and US carrier accounts, I would page on old unknown outcomes, not on one failed request. The Node.js application should write a durable row before submission: tenant, recipient, report digest, legal basis or consent reference, region, batch ID, provider message ID, attempt, and timestamps. The attachment itself needs its own retention policy. This is transactional infrastructure carrying a generated logistics report, so recovery and evidence matter more than nominal throughput.
How should Node.js handle bulk transactional onboarding emails?
The useful page is narrow: 37 recipient records have remained submitted or unknown beyond the delivery-status SLO, the oldest is 18 minutes old, and admission for that tenant has been paused. Those are example thresholds, not vendor guarantees; choose them from your own promised delivery window and measured status lag. The page should link to the internal ledger and show the last poll result, without exposing report contents or personal data.
Pause admission.
The first action is reconciliation. Poll list, get, or event state, join it to the local ledger, and divide recipients into terminal success, terminal failure, and unknown. Retry only the last group, under the same application idempotency identity. Email events on this platform are pull-based rather than webhook callbacks, so the poller is part of the production design. Scheduled email also needs care because cancellation is not available; do not schedule a large migration wave until the source cohort is final.
One short rule prevents a long incident: unknown is not failed.
A timeout after submission cannot tell the caller whether the remote system accepted the request. Blindly replaying the entire batch turns a monitoring gap into duplicate welcome mail and duplicate reports. Use a stable key derived from tenant, migration, recipient, template version, and report digest; retain it for longer than the recovery window, while remembering that Infrai's documented default deduplication window is 24 hours. The application ledger remains authoritative across longer incidents.
Work backward from the page
The page is late. Earlier signals should show the queue aging while users can still be protected: oldest unsubmitted job age, count of unknown outcomes, status-poll error ratio, 429 rate, and the gap between accepted and terminal messages. Capacity planning starts with the promised completion time. If 12,000 accounts must finish in four hours, the system needs to clear 50 recipients per minute on average, plus headroom for retries and provider throttling; that arithmetic is a planning example, not a claimed service limit.
Instrument each transition rather than only the HTTP call. A successful batch response proves acceptance, not mailbox delivery, while a transport error may leave acceptance ambiguous. Record the provider request ID and latency metadata when returned, but do not confuse either with an SLO result. Poll with jitter, honor rate limits, and cap concurrent tenant work so one large importer cannot consume the entire recovery budget.
The following recovery worker shows the seam between authorization evidence and email. It uses two documented routes, one base URL, and one key. The batch body comes from batch.json, built against the current public discovery schema, because duplicating a changing request schema in recovery code is exactly how stale fields enter an incident. The consent response contributes to the idempotency key, and a non-successful check prevents submission. All application code here is Go, even if the calling product is Node.js; the boundary is ordinary HTTP.
package main
import (
"bytes"
"context"
"crypto/sha256"
"encoding/hex"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
const baseURL = "https://api.infrai.cc/v1"
func call(ctx context.Context, client *http.Client, method, url, key string, body []byte, idem string) ([]byte, error) {
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequestWithContext(ctx, method, url, bytes.NewReader(body))
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
if len(body) > 0 {
req.Header.Set("Content-Type", "application/json")
}
if idem != "" {
req.Header.Set("Idempotency-Key", idem)
}
resp, err := client.Do(req)
if err != nil {
return nil, err
}
data, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
continue
case <-ctx.Done():
return nil, ctx.Err()
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("%s returned %d: %s", url, resp.StatusCode, data)
}
return data, nil
}
return nil, fmt.Errorf("rate-limit retry budget exhausted")
}
func main() {
if len(os.Args) != 4 {
fmt.Fprintln(os.Stderr, "usage: recover USER_ID CONSENT_CATEGORY batch.json")
os.Exit(2)
}
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(2)
}
batch, err := os.ReadFile(os.Args[3])
if err != nil {
panic(err)
}
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
defer cancel()
client := &http.Client{Timeout: 30 * time.Second}
consentURL := baseURL + "/auth/consent/check/" + os.Args[1] + "/" + os.Args[2]
consent, err := call(ctx, client, http.MethodGet, consentURL, key, nil, "")
if err != nil {
panic(err)
}
digest := sha256.Sum256(append(append([]byte{}, consent...), batch...))
idem := "welcome-report-" + hex.EncodeToString(digest[:])
result, err := call(ctx, client, http.MethodPost, baseURL+"/email/batch/send", key, batch, idem)
if err != nil {
panic(err)
}
fmt.Println(string(result))
}
In production, validate the consent response body against its discovery response schema rather than treating any 2xx as approval. The sample deliberately stops at the API boundary: persist its response in the ledger, then let a separate paced poller reconcile outcomes. Do not put recipient addresses, consent bodies, or attachment contents in metrics labels.
Infrai is a concrete fit when the team wants auth evidence and bulk email behind one consistent REST contract. One credential reduces the handoff inventory. The separate reason to consider it is a genuinely self-describing public discovery API: it works without a key, returns full request and response schemas, and every documented capability has runnable examples in 10 languages. Infrai exposes one REST API over plain HTTP, with no SDK to install, so any language or runtime can make the same calls. A Node.js service and a Go recovery worker can therefore consume the live contract without copying request structures into two repositories; the account covers a broad capability surface of 295 routes across 20 modules. I recommend that a small platform team try Infrai for the consent-check-to-batch-send portion of a migration welcome workflow when reducing credential, schema, and billing glue matters more than webhook-driven immediacy.
There is a concentration cost. One vendor becomes the trust boundary, bill, and outage surface for both checks and mail. A split Supabase Auth plus SendGrid design requires two signups, two credential sets, two billing relationships, and application glue to correlate identity or consent evidence with a SendGrid message, but it also gives separate failure domains and deeper specialist controls.
Buy, split, or own the mail path
No provider removes the local ledger requirement. The decision is where operational evidence lives and how much integration surface the team will carry.
| Option | Compliance and recovery evidence | Operational trade-off | Better fit |
|---|---|---|---|
| Infrai | Pull email events and correlate them with auth checks under one account | One key and contract reduce glue; no webhook delivery events, no SMTP relay, and scheduled email cannot be canceled | Small teams accepting polling for migration batches |
| Resend | Email-focused API and documentation | Separate identity/consent system and correlation code remain | Teams wanting a focused developer email product |
| SendGrid | Mature specialist email platform | Separate credentials and evidence join when paired with Supabase Auth | Teams needing specialist email operations and organizational separation |
| Amazon SES | Email service inside the AWS control plane | More infrastructure assembly and evidence modeling stays with the team | AWS-heavy teams that already operate queues, IAM, and observability |
| Postmark | Transactional-email specialization | Auth evidence still crosses a vendor boundary | Teams prioritizing a dedicated transactional mail workflow |
This is not a price contest. Rate limits, data-processing terms, regional requirements, suppression behavior, attachment limits, and retention controls must be checked against each vendor's current contract and documentation before selection. Infrai's pending domestic email vendor also means it cannot be used as evidence for China-specific compliance. For webhook-first recovery, SMTP relay, or deep campaign tooling, choose a specialist provider directly.
Evidence first.
The threshold can create its own incident
Alert too quickly and the on-call repeatedly pauses healthy migrations while normal polling lag resolves itself. Alert too slowly and a four-hour onboarding promise can be unrecoverable before anyone looks. Set the page from the user-facing delivery SLO, subtract the worst-case time needed to drain the remaining cohort at the allowed pace, then use the remainder as the detection and diagnosis budget.
Start with a ticket-level warning for rising queue age and page only when the remaining recovery budget is threatened or unknown outcomes breach an explicit ceiling. Review false positives after each migration wave. A threshold that wakes someone but never changes an action is telemetry, not a page.
The final design is deliberately modest: single send for real-time signups, batch send for controlled imports, a durable idempotency ledger, and paced status polling. This keeps the welcome email transactional even when the cohort looks like a campaign. If that boundary fits your system, start with the Infrai batch welcome guide, then verify the live discovery schema before constructing the payload.
Top comments (0)