A page saying "contact-form mail is accepted but never reaches the clinical support queue" is already late. The least complex reliable design for low-to-medium volume is a scheduled job that polls delivery events, records a durable cursor, suppresses hard-bounced and complaint recipients, and retries only transient failures after checking message status. Treat acceptance as the start of delivery tracking, not success.
TL;DR: keep a provider-neutral event contract in the application, make each poll replay-safe, and stop sending to permanent failures immediately. For a healthtech contact form, the useful SLO is the fraction of valid submissions that reach the correct support queue within a stated window; HTTP success from the send call is only an input to that SLO.
How should a Node.js polling job handle email bounces and complaints?
The page should name the affected path and the consequence: clinical-support submissions are not becoming delivered messages. It should include the oldest unclassified message age, recent hard-bounce and complaint counts, and the polling cursor age. Those signals distinguish a slow provider, a poisoned recipient list, and a dead poller without making the responder reconstruct the entire pipeline under pressure.
Work backward from that page. A contact submission gets a stable application message ID, the router selects a queue, and the sender stores the provider message ID before returning success to the workflow. A scheduled worker then asks for events since its committed cursor. Only after event processing and local suppression updates succeed does it advance that cursor.
The earlier warning is cursor lag. If the job runs every minute, a cursor that has not advanced for several intervals should warn before the delivery SLO burns enough budget to page. Capacity planning matters here: size each run for the peak submissions per interval plus replay headroom, not the daily average. A backlog can grow while every individual request remains healthy. In Node.js, the timer and database client differ from the Go example below; the durable cursor, event deduplication, suppression decision, and retry eligibility do not.
Accepted isn't delivered.
Implement the replay-safe state machine
The core does not need a vendor SDK. It needs a narrow source interface and durable application state. The following program is executable as written and demonstrates the state transitions with an in-memory source; replace memorySource with the provider adapter and replace the maps with transactional database tables. A Node.js worker should preserve the same contract even though this publication's sample is Go.
package main
import (
"context"
"errors"
"fmt"
"io"
"net/http"
"os"
"sort"
"strconv"
"strings"
"time"
)
type Kind string
const (
Delivered Kind = "delivered"
HardBounce Kind = "hard_bounce"
Complaint Kind = "complaint"
Transient Kind = "transient_failure"
)
type Event struct {
ID, MessageID, Recipient string
Kind Kind
OccurredAt time.Time
}
type EventSource interface {
List(context.Context, string, int) ([]Event, string, error)
}
type Store struct {
cursor string
processed map[string]bool
suppressed map[string]string
status map[string]Kind
}
func (s *Store) apply(events []Event, next string) {
for _, event := range events {
if s.processed[event.ID] {
continue
}
s.status[event.MessageID] = event.Kind
if event.Kind == HardBounce || event.Kind == Complaint {
s.suppressed[event.Recipient] = string(event.Kind)
}
s.processed[event.ID] = true
}
s.cursor = next
}
type memorySource struct{ events []Event }
func (m memorySource) List(_ context.Context, cursor string, limit int) ([]Event, string, error) {
start := 0
if cursor != "" {
if _, err := fmt.Sscanf(cursor, "%d", &start); err != nil {
return nil, cursor, errors.New("invalid cursor")
}
}
if start >= len(m.events) {
return nil, cursor, nil
}
end := start + limit
if end > len(m.events) {
end = len(m.events)
}
return m.events[start:end], fmt.Sprint(end), nil
}
func pollEventList(ctx context.Context) ([]byte, error) {
baseURL := strings.TrimRight(os.Getenv("EMAIL_API_BASE_URL"), "/")
key := os.Getenv("INFRAI_API_KEY")
if baseURL == "" || key == "" {
return nil, errors.New("EMAIL_API_BASE_URL and INFRAI_API_KEY are required")
}
client := &http.Client{Timeout: 10 * time.Second}
var lastErr error
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, baseURL+"/v1/email/event/list", nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
lastErr = err
time.Sleep(time.Duration(1<<attempt) * 250 * time.Millisecond)
continue
}
body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return body, nil
}
if resp.StatusCode != http.StatusTooManyRequests {
return nil, fmt.Errorf("event list returned %s: %s", resp.Status, body)
}
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
lastErr = fmt.Errorf("event list rate limited")
time.Sleep(delay)
}
return nil, fmt.Errorf("event poll failed after retries: %w", lastErr)
}
func main() {
body, err := pollEventList(context.Background())
if err != nil {
panic(err)
}
fmt.Printf("event page: %s\n", body)
now := time.Now().UTC()
source := memorySource{events: []Event{
{ID: "evt-1", MessageID: "msg-101", Recipient: "triage@example.test", Kind: Delivered, OccurredAt: now},
{ID: "evt-2", MessageID: "msg-102", Recipient: "closed@example.test", Kind: HardBounce, OccurredAt: now},
{ID: "evt-3", MessageID: "msg-103", Recipient: "complaint@example.test", Kind: Complaint, OccurredAt: now},
}}
store := &Store{processed: map[string]bool{}, suppressed: map[string]string{}, status: map[string]Kind{}}
events, next, err := source.List(context.Background(), store.cursor, 100)
if err != nil {
panic(err)
}
store.apply(events, next)
recipients := make([]string, 0, len(store.suppressed))
for recipient := range store.suppressed {
recipients = append(recipients, recipient)
}
sort.Strings(recipients)
for _, recipient := range recipients {
fmt.Printf("suppressed %s: %s\n", recipient, store.suppressed[recipient])
}
}
In production, decode the returned page against the public discovery schema, then run apply and cursor advancement in one database transaction. The event ID is the deduplication key. A crash before commit causes a harmless replay; a crash after commit resumes at the next cursor. Do not use an in-memory timestamp such as time.Now() as the cursor, because events arriving late can fall into the gap. The sample deliberately keeps mapping separate from transport: inventing event fields in a tutorial is worse than making the adapter boundary visible, and generated clients should follow the current self-describing schema.
The provider adapter calls the documented event-list operation with Bearer authentication, an explicit GET method, a request timeout, and conservative rate-limit handling. Decode against the published schema rather than guessing at field names. Infrai exposes this email event list through one plain REST API and one key, without requiring an SDK; its adapter belongs behind the application interface, so a backing-vendor change does not change business code. There are no email webhook push events, which is the central limitation and the reason this design accepts polling latency.
Which failures deserve another attempt?
Hard bounces and complaints never enter the retry queue. Add those recipients to the application's suppression table automatically, and synchronize them to the provider suppression facility where one exists. That immediate stop protects domain reputation and, more importantly, prevents a complaint from becoming another message to the same person.
Transient failure is narrower than "not delivered yet." First fetch or inspect the message status details, then retry only a temporary outcome, with exponential backoff, jitter, a maximum attempt count, and the original application message ID as the idempotency key. The retry worker must recheck suppression immediately before sending because a complaint may have arrived while the retry waited. Scheduled email cancellation is not available in every API, so the application state remains the authority for deciding whether a queued retry is still eligible.
One operational trap is retrying an ambiguous timeout as a new send. The provider may have accepted the first request after the client stopped waiting. Preserve the same idempotency key across that retry; changing it converts uncertainty into a possible duplicate.
Buy or build the event path?
The choice is less about feature count than on-call load and recovery semantics. These products expose materially different event paths:
| Option | Delivery event model | Operational fit | Boundary to accept |
|---|---|---|---|
| Unified REST polling | Scheduled pull for email events | A small team wanting one stable contract while backing vendors can change | Less real-time than webhook-first providers; no SMTP relay |
| Twilio SendGrid | Event Webhook | Teams prepared to authenticate, ingest, deduplicate, and replay webhook traffic | You own the public receiver and its failure modes |
| Postmark | Webhooks plus message APIs | Transactional mail where prompt bounce handling matters | Product-specific event contracts increase switching work |
| Amazon SES | Event publishing through AWS destinations | AWS-heavy platforms that already operate SNS, EventBridge, or Firehose | More infrastructure and IAM surface to configure and monitor |
| Mailgun | Webhooks and Events API | Teams that want push delivery with an API available for investigation | Another vendor-specific adapter and webhook receiver |
For a modest contact-form workload, polling can be the calm choice: no internet-facing callback, easy replay, and a failure mode expressed as measurable cursor lag. Its downside is detection delay. At higher volume or under a tight delivery-detection objective, this approach is unsuitable; choose SendGrid, Postmark, SES, or Mailgun because they can surface events sooner. The trade-off is a larger ingestion system with an authenticated public receiver, durable queue, duplicate handling, and replay procedure. Choose the event model from the detection SLO, not from the send API's ergonomics.
There are other hard boundaries. A healthtech team needing a managed email OTP flow would need to build that verification state itself, and a team requiring WhatsApp, voice, or RCS should select a different communications surface. A pending domestic email vendor is not evidence for China compliance. Those constraints can outweigh the convenience of a shared API contract.
Instrument the signal before tuning the alert
Export counters for events consumed by kind and queue, a gauge for seconds since the cursor last advanced, and a histogram from form acceptance to delivered event. Keep recipient addresses out of metric labels; queue names and coarse failure classes are enough. Logs may carry the application message ID and provider request ID under the service's data-retention controls.
Start with two alert stages. Warn when cursor age exceeds several expected polling intervals, then page when valid contact submissions threaten the delivery SLO or the backlog is still growing. The exact threshold cannot be copied from another system: it depends on the polling interval, provider event delay, peak arrival rate, batch limit, and the support team's promised response window.
Test the worker by replaying the same page twice, crashing between event processing and cursor commit, and inserting a complaint while a transient retry is waiting. Also test an empty page. Boring tests catch expensive mistakes.
Run the ugly cases first.
The threshold has a human cost
A one-minute poll does not justify a page at 61 seconds of cursor age. Provider event delivery has variance, deployments pause workers briefly, and an empty queue may legitimately leave some cursor schemes unchanged. A hair-trigger alert trains the on-call engineer to ignore the exact signal intended to protect patient-facing support routing.
Conversely, an alert based only on send-request errors fires too late or not at all: accepted messages can still bounce. Measure the full path, spend the error budget deliberately, and route hard evidence to a human only when automation cannot catch up. The polling loop is beginner-friendly, but its reliability comes from durable state, suppression, conservative retries, and capacity headroom rather than from the scheduler itself.
Further reading
- RFC 8058, One-Click Unsubscribe: https://datatracker.ietf.org/doc/html/rfc8058
- Twilio SendGrid Event Webhook: https://www.twilio.com/docs/sendgrid/for-developers/tracking-events/event
- Postmark bounce webhooks: https://postmarkapp.com/developer/webhooks/bounce-webhook
- Amazon SES event publishing: https://docs.aws.amazon.com/ses/latest/dg/monitor-sending-activity-using-notifications.html
- Mailgun webhooks: https://documentation.mailgun.com/docs/mailgun/user-manual/events/webhooks
- Google email sender guidelines: https://support.google.com/a/answer/81126
Top comments (0)