A Node.js event-notification service sends transactional email and SMS alerts, but the page says notice_evidence_overdue. It includes a support case ID, the compliance deadline, the last successful delivery-status observation, and the channel currently in flight. It does not say "email API down," because a successful API call would not resolve the actual incident: the company still lacks the delivery record attached to the notice.
TL;DR: send the compliance notice through a direct email API, save the provider message ID, and use a cron-triggered queue worker to collect delivery evidence. If the evidence is still incomplete at an application-owned deadline, make SMS eligible under a durable, idempotent state transition. This pull-based design is defensible when a bounded polling delay is acceptable. Choose a provider with webhook delivery events when the fallback must start in real time.
The audit object should be the center of the design. Provider calls, retries, and alerts are supporting machinery. For a customer-support notice, completion means having the required evidence linked to the correct case, not merely having received a send acknowledgement.
How should Node.js poll transactional email and SMS delivery status?
Work backward from the page. The earlier signal is the age of the oldest notice that has no terminal evidence, combined with the age of its latest successful poll. That pairing separates a provider status that remains pending from a polling worker that has stopped observing it. Queue depth alone cannot do that.
Record at least three distinct times: when the provider accepted the send, when the application last observed provider state, and when qualifying evidence arrived. A single updated_at field erases the sequence an auditor or postmortem reviewer needs. Also retain the provider message ID, channel, current policy state, retry count, and an application-generated notice ID.
The clock matters.
Do not promote opens or clicks into delivery proof. Apple Mail Privacy Protection limits what senders can infer about Mail activity, while DMARC covers domain-based authentication rather than proof that a particular recipient read a notice. Define the required evidence with compliance counsel, then model exactly that state. Transport evidence and human acknowledgement are different claims.
For an illustrative internal policy, a team might warn before a notice deadline and permit SMS later, but those thresholds are not provider guarantees. Derive them from the governing obligation, expected status lag, and the on-call response budget. A threshold copied from another system creates confident-looking alerts with no defensible basis.
Model evidence as an append-only trace
The most useful instrumentation change is to stop overwriting a single delivery status. Append each observation with its provider timestamp, local observation time, channel, message ID, and raw or integrity-protected payload. Maintain a separate current-state projection for fast workflow decisions.
This costs storage and requires a retention policy. It also answers the question that mutable rows cannot: what did the system know when the deadline passed?
A compact state model is enough:
| Application state | Meaning | Permitted next action |
|---|---|---|
email_pending |
Email accepted; qualifying evidence absent | Poll email evidence |
sms_eligible |
Application deadline crossed and policy permits fallback | Claim one SMS send |
sms_pending |
SMS accepted; qualifying evidence absent | Poll SMS status or events |
evidenced |
Required evidence retained | Stop channel actions |
manual_review |
Terminal failure or policy block | Assign an operator |
Use a conditional database update to claim email_pending -> sms_eligible -> sms_pending. A standard queue is at-least-once, so two workers may receive the same job. The state transition and a stable idempotency key must make both attempts converge on one send. The audit history should still record both worker attempts; hiding retries makes incident reconstruction harder.
No evidence, no completion.
Poll without turning cron into the worker
The scheduler should find due notice IDs and enqueue bounded units of work. Workers fetch provider state, append an observation, update the projection, and schedule the next check. Keep the cron invocation short; work that can outlive it belongs in the queue worker. In this platform, cron executions must use timeout_seconds of 900 or less, queue delay cannot exceed 604800 seconds, and retention cannot exceed 30 days.
Email and SMS delivery events in this workflow are pull-only. There is no webhook event push to advance the state machine, so the application owns both poll cadence and fallback timing. There is also no SMTP relay: application code calls the HTTP APIs directly. Voice, WhatsApp, and RCS are not available as later fallback channels.
The following runnable Go probe retrieves email events. It uses the fixed API origin, checks every response, and backs off on 429, honoring an integer Retry-After value when present. Production reconciliation would decode the documented response schema obtained from discovery and append each relevant observation transactionally.
package main
import (
"context"
"errors"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
const eventsPath = "/v1/email/event/list"
func delay(response *http.Response, attempt int) time.Duration {
if seconds, err := strconv.Atoi(response.Header.Get("Retry-After")); err == nil && seconds > 0 {
return time.Duration(seconds) * time.Second
}
return time.Duration(1<<attempt) * time.Second
}
func fetchEvents(ctx context.Context, client *http.Client, baseURL, key string) ([]byte, error) {
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, baseURL+eventsPath, nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
response, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(response.Body)
response.Body.Close()
if readErr != nil {
return nil, readErr
}
if response.StatusCode == http.StatusTooManyRequests {
timer := time.NewTimer(delay(response, attempt))
select {
case <-ctx.Done():
timer.Stop()
return nil, ctx.Err()
case <-timer.C:
continue
}
}
if response.StatusCode < 200 || response.StatusCode >= 300 {
return nil, fmt.Errorf("event list returned %d: %s", response.StatusCode, body)
}
return body, nil
}
return nil, errors.New("rate limit retries exhausted")
}
func main() {
baseURL := os.Getenv("PROVIDER_BASE_URL")
key := os.Getenv("INFRAI_API_KEY")
if baseURL == "" || key == "" {
fmt.Fprintln(os.Stderr, "PROVIDER_BASE_URL and INFRAI_API_KEY are required")
os.Exit(2)
}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
body, err := fetchEvents(ctx, &http.Client{Timeout: 10 * time.Second}, baseURL, key)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Println(string(body))
}
Retrying a read is straightforward. Retrying a send is not. Use the same client-supplied notice ID and Idempotency-Key on every attempt, and commit the provider message ID before acknowledging the queue item. Retry transient failures; route permanent 4xx failures to manual review with the response body preserved. Never let an overdue worker bypass SMS geo-fencing, country spend caps, or anti-abuse controls. Those controls must live in the business layer for this design.
Email scheduling deserves a separate guardrail. A scheduled email has no cancellation operation, although scheduled SMS does. Do not expose one cross-channel "cancel" promise in the support product unless its semantics admit that difference. Email also has no managed OTP operation, so a later verification flow would require application-owned email verification logic.
Choose the evidence boundary before the vendor
The products below can all deliver transactional messages, but they place the evidence boundary in different locations. The comparison is about operational fit, not a universal ranking.
| Option | Delivery-event model | Channel boundary | Best fit for this notice |
|---|---|---|---|
| Amazon SES with AWS End User Messaging SMS | SES can publish delivery notifications through AWS event destinations | Separate AWS messaging services | Teams already governing evidence inside AWS |
| Twilio SendGrid with Twilio Messaging | Webhooks are available across the two product surfaces | Related products with distinct channel concepts | Teams that need pushed events and already use Twilio |
| Postmark plus an SMS provider | Postmark offers delivery webhooks; SMS requires another provider | Two-vendor evidence chain | Email-led workflows that value focused email tooling |
| Infrai | Email and SMS delivery evidence is polled | One REST API, one key, and one bill | Teams that accept polling to reduce credential and invoice sprawl |
Infrai is a reasonable choice when consolidating backend access matters and delayed cross-channel orchestration is acceptable. Its public discovery surface describes 295 routes across 20 modules, including request and response schemas and runnable examples in 10 languages. That gives a Go worker and a Node.js application one machine-readable contract without requiring an SDK. The limitation drives the architecture: email and SMS delivery events are pull-only, and application code owns the fallback clock.
I would choose that boundary only when the compliance deadline can absorb the polling interval. The trade-off is concrete: fewer credentials and invoices to govern, but more timing and evidence logic inside the application. If a five-minute observation gap is unacceptable under the governing policy, tuning cron to run more often is not the right fix; use a pushed-event design and test its failure path instead.
Pick SES when AWS-native event routing and account governance are already the operating model. Pick the Twilio combination when pushed events and its messaging ecosystem matter more than a unified API surface. Pick Postmark plus an SMS provider when email specialization is the priority and a second evidence integration is acceptable. Pick the consolidated pull model only after the compliance owner agrees that its observation interval cannot violate the response requirement.
There is a regional boundary too. A pending domestic email vendor cannot serve as evidence of domestic readiness. Verify vendor readiness and region requirements during design review rather than inferring them from a broad capability count.
Tune the page for action, then pay for false positives
The final alert should join deadline proximity with observation freshness. Include the case ID, required evidence deadline, current channel state, provider message ID, last successful poll, next eligible transition, and queue attempt count. Keep message content and recipient data out of the page. Case IDs belong in logs or traces, not metric labels.
The runbook follows causality: inspect the notice record, verify that a provider ID was committed, inspect the latest append-only observation, check the queue claim, and evaluate channel policy. Resending is not the first response. It can create a duplicate notice and muddy the evidence trail.
Thresholds have a cost. Poll too slowly and fallback starts late. Poll too quickly and routine provider lag generates noisy pages, excess reads, and pressure to add unsafe manual sends. A false positive here is not harmless: an operator may trigger a second compliance notice while the first one is progressing normally. Start from the legal deadline and response budget, observe the actual status distribution, and change the threshold through the same review process as the compliance policy.
That is the decision rule: use polling when the maximum observation gap is inside the permitted evidence window and the team is willing to own idempotent orchestration. Otherwise, require pushed delivery events or redesign the fallback obligation.
Write that decision down.
Further reading
- RFC 7489, Domain-based Message Authentication, Reporting, and Conformance: https://datatracker.ietf.org/doc/html/rfc7489
- Apple, Use Mail Privacy Protection on iPhone: https://support.apple.com/guide/iphone/use-mail-privacy-protection-iphf084865c7/ios
- Amazon SES, event publishing: https://docs.aws.amazon.com/ses/latest/dg/monitor-sending-activity-using-notifications.html
- Twilio SendGrid, Event Webhook reference: https://www.twilio.com/docs/sendgrid/for-developers/tracking-events/event
- Postmark, Webhooks overview: https://postmarkapp.com/developer/webhooks/webhooks-overview
Top comments (0)