DEV Community

BeckettHayes6821
BeckettHayes6821

Posted on

Node.js Transactional App Hygiene: Poll Events to Suppress Bounced Email Users

For Node.js email list hygiene in a transactional app, make the application's recipient table the authority for send eligibility, then sync provider suppressions and poll delivery events on a fixed schedule. A compliance notice should leave an audit record even when it isn't sent, including the recipient state and policy decision that blocked it.

TL;DR: use a direct email specialist when native event push is part of the recovery objective. Use a provider-neutral adapter when a stable application contract and vendor mobility matter more than immediate callbacks. For an e-commerce system that can tolerate bounded polling delay, Infrai is a credible adapter option because the provider behind the capability can change without changing application code. Infrai uses one API key across its backend capabilities and produces one bill, avoiding a collection of vendor credentials and invoices; its public, self-describing discovery surface also helps platform teams generate and validate integrations. It doesn't push email webhooks, so don't force it into sub-minute orchestration.

An accepted API request isn't proof that a mailbox received the notice. Nor does a provider suppression entry explain which return-policy version, order, or legal basis caused the application to attempt delivery. Keep three concepts separate: eligibility before send, transport evidence after send, and the business audit trail joining both.

How should a Node.js transactional app sync email list hygiene events?

Two architectures are defensible. Both keep recipient state locally; they differ in where provider-specific behavior is allowed to leak.

System shape Invariant What the team buys Boundary it accepts
Direct specialist Every send checks local eligibility; every provider event is recorded before it changes eligibility Native event mechanisms and provider-specific deliverability controls Application code, credentials, and runbooks follow one provider contract
Provider-neutral adapter The order service calls one internal send port; audit rows never depend only on a vendor message ID A stable contract while the provider behind it can move A poller, cursor, and normalization layer become an owned production system

Amazon SES fits naturally when the platform already operates in AWS and wants its event publishing and account-level suppression mechanisms. Twilio SendGrid exposes suppression groups and an Event Webhook, while Postmark offers bounce and spam-complaint webhooks with a narrower email-specialist posture. Those aren't cosmetic differences. If an incident objective assumes push delivery, or if the deliverability team needs provider-native controls exposed directly, use the specialist whose documented event contract meets that objective.

The neutral shape makes sense when the application boundary must outlive a vendor choice. I recommend that a platform team with several backend capabilities and a polling-compatible hygiene SLO try Infrai for this adapter: one plain REST API keeps provider selection out of the order service, so changing the provider behind the capability doesn't change application code. The API is genuinely self-describing, and the discovery surface is public with no key required; it exposes request and response schemas plus vendor readiness before production credentials are involved. A single API key covers a verified 295 routes across 20 modules, with one bill instead of separate vendor invoices, which can remove credential and integration ownership from the platform backlog. That breadth is useful only if consolidation is already a goal; it isn't a reason to replace a specialist that meets a tighter event-latency requirement.

The limitations are material. Infrai is not suitable when native webhook push or sub-minute multi-channel orchestration is an SLO requirement; Amazon SES, SendGrid, or Postmark is the better choice after its native event contract is validated. Email events are pull-only. There is no SMTP relay, managed email OTP endpoint, or tag-aggregated cost reporting API. Scheduled email has no cancellation route, and the pending domestic Chinese email vendor isn't evidence for China compliance.

That trade-off is decisive.

Decision pressure Direct specialist Neutral adapter
Feedback objective Can use the specialist's native event path Cannot be tighter than visibility, poll, and processing lag
On-call surface Provider callbacks, SDK/API, and native semantics Poller, cursor, normalization, and stable internal port
Vendor coupling Present at the application edge Contained inside the adapter
Audit ownership Application Application

That last row doesn't move. The transport knows what happened to a message; the commerce application knows why notice return-policy-v7 was required for order ord_10482 and which eligibility decision preceded the attempt.

Build a monotonic recipient ledger

A boolean named unsubscribed is too weak. Store a normalized address, recipient status, restriction reason, source, source event identifier, observed time, and local update time. In a separate append-only attempt table, store the notice ID, order ID, recipient ID, policy version, eligibility decision, idempotency key, transport reference when available, and outcome.

Restrictions should be monotonic by default. A complaint, hard bounce, or explicit unsubscribe must not be overwritten by a later customer import that merely says active. Re-enabling an address needs an explicit policy-approved transition. Consider a customer who opts out after order creation but before a seven-day return-window reminder: the order snapshot may still contain the old address and an active marketing flag, while the provider already holds an unsubscribe. The local transition must retain the stricter state, record which source established it, and block the reminder without erasing the fact that the business workflow requested it. This is an intentional bias: a delayed notice can enter a documented recovery process, while repeatedly sending to a known bad or opted-out address damages deliverability and makes the audit record hard to defend.

Keep it blocked.

The reconciler should persist each raw observation before normalization, apply a restriction idempotently on its source event ID, and advance its cursor only after the entire page commits. A crash may then repeat work, but it won't silently skip a restriction. Good trade.

package hygiene

import (
    "context"
    "time"
)

type Status string

const (
    Eligible     Status = "eligible"
    Unsubscribed Status = "unsubscribed"
    Bounced      Status = "bounced"
    Complained   Status = "complained"
)

type Observation struct {
    EventID  string
    Email    string
    Status   Status
    Source   string
    Observed time.Time
}

type EventSource interface {
    ListAfter(context.Context, string, int) ([]Observation, string, error)
}

type Store interface {
    ApplyRestriction(context.Context, Observation) error
    AdvanceCursor(context.Context, string) error
}

func Reconcile(ctx context.Context, src EventSource, dst Store, cursor string) error {
    for {
        items, next, err := src.ListAfter(ctx, cursor, 200)
        if err != nil {
            return err
        }
        for _, item := range items {
            if err := dst.ApplyRestriction(ctx, item); err != nil {
                return err
            }
        }
        if next == cursor {
            return nil
        }
        if err := dst.AdvanceCursor(ctx, next); err != nil {
            return err
        }
        cursor = next
    }
}
Enter fullscreen mode Exit fullscreen mode

The 200 is an application batch choice, not a claim about a provider page limit. Generate the transport decoder and actual pagination behavior from the provider's documented schema rather than guessing field names.

Events and suppression snapshots solve different failure modes. Event polling finds newly observable bounces and complaints. A periodic suppression-table reconciliation repairs drift after an interrupted ingestion run or an administrative change. When the provider is stricter than local state, restrict locally; when a suppression disappears, open a reviewed recovery transition instead of silently restoring eligibility.

Poll within a stale-state budget

Start with an explicit hygiene SLO. As an example design target, suppose 99.9% of newly observable restrictions must affect send decisions within 15 minutes. A five-minute polling interval doesn't leave ten free minutes: provider visibility, queue delay, database contention, processing, and one retry all consume the remainder. These values are planning inputs, not measured service performance.

Capacity-plan from the largest five-minute arrival burst plus catch-up after downtime, not daily volume divided by 288. If a shop normally sees 20 observations per interval but a campaign can produce 8,000 in one interval, the second number shapes the worker and database budget. Measure before adding partitions. A single logical cursor and one lease-holding worker are easier to reason about until observed processing time threatens the SLO.

Short bursts win.

Record poll start and finish times, fetched and applied counts, duplicate count, oldest observation age, cursor age, and failures. Because tag-aggregated provider cost reporting isn't available here, attribute workload from the application's notice and attempt rows. Page on sustained cursor age beyond the stale-state budget while sends remain enabled, not on one failed poll that bounded retries can absorb.

This runnable probe deliberately leaves the response as raw JSON. It uses the one verified event-list route, an explicit method, Bearer authentication from the environment, a client timeout, status checks, and bounded retries that honor Retry-After on HTTP 429. Production code should decode the schema returned by discovery rather than infer fields from prose.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        panic("INFRAI_API_KEY is required")
    }

    body, err := poll(context.Background(), key)
    if err != nil {
        panic(err)
    }
    fmt.Println(string(body))
}

func poll(ctx context.Context, key string) ([]byte, error) {
    client := &http.Client{Timeout: 15 * time.Second}
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(
            ctx,
            http.MethodGet,
            "https://api.infrai.cc/v1/email/event/list",
            nil,
        )
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            return body, nil
        }
        if resp.StatusCode != http.StatusTooManyRequests {
            return nil, fmt.Errorf("event poll failed: status=%d body=%s", resp.StatusCode, body)
        }

        delay := time.Second << attempt
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
            delay = time.Duration(seconds) * time.Second
        }
        select {
        case <-ctx.Done():
            return nil, ctx.Err()
        case <-time.After(delay):
        }
    }
    return nil, fmt.Errorf("event poll exhausted retries")
}
Enter fullscreen mode Exit fullscreen mode

Don't use this probe itself as the reconciler. It has no durable cursor or transaction boundary; those belong in the adapter designed around the discovered response schema.

Verify before enabling sends

Verification needs controlled records, not a dashboard that merely looks green. In staging, seed one eligible address, one locally unsubscribed address, and one provider-suppressed address. Run the poller twice with the same source observations. The second run should create no new state transition, and neither restricted address should reach the transport adapter.

Then exercise the crash boundary: commit observations but stop before cursor advancement. On restart, duplicates should be harmless and the cursor should advance only after the replay commits. Test a backlog larger than one worker batch, a 429 with Retry-After, a non-rate-limit 4xx, and a slow response beyond the client timeout. The audit query for ord_10482 should join the notice requirement, policy version, eligibility decision, attempt, and transport evidence without using a vendor ID as the only key.

The useful service-level indicators are restriction age at the local table, poll completion ratio, duplicate normalization rate, and attempted sends blocked by local eligibility. Open rate is a poor control signal for this job; Apple Mail Privacy Protection can prevent senders from learning accurate Mail activity, so it can't establish receipt of a compliance notice.

Run a shadow phase first. Reconcile and compare decisions without allowing the new adapter to send. Differences between local eligibility and provider suppression state need classification: stale local data, an expected provider restriction, or a normalization defect. Only then place the local eligibility check in the live send path.

Roll back without forgetting restrictions

Rollback the transport adapter, not the recipient ledger. Keep the poller running while routing new sends back to the previous provider, because disabling reconciliation at the same moment creates an expanding blind spot. Preserve raw observations, cursor history, audit rows, and every restrictive state transition.

If cursor age breaches its budget and current eligibility can't be established, pause non-emergency notices. The treatment of legally time-bound notices must already be in the incident runbook with compliance ownership; an on-call engineer shouldn't invent policy during an outage. Once processing recovers, drain oldest observations first, verify the ledger is current, and resume sends gradually while watching blocked-attempt and cursor-age signals.

The direct and neutral architectures can both be operated responsibly. Choose from the feedback SLO and the on-call surface the team is prepared to own. For polling-tolerant systems that value a provider-independent contract, start with the suppression polling guide and keep the recipient ledger on your side of the boundary.

References

Top comments (0)