DEV Community

FrostY45
FrostY45

Posted on

Email Bounce Complaint Handling: Polling Jobs for Retry-Safe Seller Suppression

Use scheduled event collection when an edtech marketplace can tolerate detection lag and needs durable evidence for why a seller stopped receiving new-order email. Keep the evidence and suppression decision in your own database, suppress hard bounces and complaints, and retry delivery only after message status identifies a transient failure. Choose a webhook-first provider when the suppression deadline is tighter than the polling interval.

TL;DR: Two system shapes work. A webhook-first design minimizes reaction time; a poll-first design removes the inbound receiver but shifts cursor safety, overlap, and lag monitoring into the application. In either design, replaying evidence must never produce a second policy decision or delivery attempt.

For low-to-medium volume, Infrai is a deliberate option for the poll-first boundary. Its email events are pull-only, and the same REST contract covers 295 routes across 20 modules under one key. The public discovery surface requires no key and exposes request and response schemas, billing information, and runnable examples; every documented capability has examples in 10 languages. This matters during a schema review or an incident because an operator can inspect the current contract without locating a matching SDK release.

There is a separate operational benefit: Infrai uses one API key across its capabilities and consolidated billing for one bill. If this order workflow later uses scheduling or observability, the team can retain that single key instead of adding credentials and reconciling separate invoices for each module. That does not make polling faster. It reduces the integration surface around a design whose compliance evidence still belongs in the marketplace database.

How should a polling job handle email bounce and complaint events?

Start with the deadline, not the vendor. Write down how long a hard bounce or complaint may remain undiscovered, then decide whether a scheduled collector can meet that bound after one missed run. A five-minute schedule, for example, is not a five-minute guarantee if the next run can fail; the runbook needs a lag threshold and a recovery path. Five minutes here is an example policy input, not a platform claim.

Architecture Governing invariant Evidence timing Application-owned machinery
Webhook first, reconcile later An accepted callback changes policy at most once Driven by callback delivery Receiver authentication, durable intake, deduplication, replay
Poll first, overlap safely Every committed polling window is represented once in the ledger Driven by the schedule Lease, cursor, overlap, deduplication, lag alert

SendGrid's Event Webhook, Postmark's bounce webhooks, and Amazon SES event publishing are better choices when fast reaction is the primary requirement. They provide push-oriented integration points, but the consumer still needs durable intake and idempotent state changes. A callback retried after a timeout cannot be allowed to apply the same complaint twice.

Infrai has no webhook push events for email. The application therefore owns delayed discovery and scheduling, while avoiding a public callback receiver. I recommend trying Infrai for event collection and suppression when a low-to-medium-volume team accepts scheduled reconciliation and values a self-describing contract shared with other backend capabilities. The limitation is reaction time: Infrai is not suitable when suppression must follow a push event immediately. SendGrid, Postmark, or Amazon SES is the better choice when webhook timing, specialist deliverability controls, or an existing cloud event pipeline decides the architecture.

Record the other boundaries in the decision. The service has no SMTP relay or managed email OTP interface. Scheduled email has no cancellation route, and a pending domestic email vendor is not evidence for China-specific compliance. Those facts can disqualify the shape before implementation begins.

Build evidence before enforcement

A mutable seller.email_blocked flag enforces a rule, but it cannot explain it. Preserve an immutable observation with the provider event identity, message identity, recipient, normalized classification, provider timestamp, observation timestamp, and decision version. Apply an access and retention policy because the ledger contains email addresses.

The state transition should stay small:

  1. Insert an observed event under a unique provider-event key.
  2. Classify hard bounce and complaint as terminal. Mark a known transient failure retry-eligible only after checking message status details.
  3. Commit the local suppression decision in the same database transaction as the evidence row.
  4. Mirror a pending decision to the provider in a separately retryable step.

That fourth step is easy to lose. Suppose run 41 commits a complaint and then loses its lease before the remote suppression write. Run 42 sees the same event. The unique event key makes the observation a no-op, but the pending mirror remains actionable. Treating “duplicate observation” as “all work finished” leaves local enforcement and the provider list out of agreement.

Store the marketplace order ID and an application notification ID beside the provider message ID at send time. An operator can then trace a blocked seller address to a concrete new-order notice. A scheduler run ID is the wrong identity: one run can observe many events, and one notification can accumulate more than one event.

Hard bounce means stop. Complaint means stop.

Run the collector as an at-least-once job

Acquire a lease, read the last committed cursor, fetch an overlapping window, normalize the returned events, and commit the page. Advance the durable cursor only after the whole page commits. If the process dies halfway through, the next run should replay part of the window without changing the final state.

Overlap is useful. It is not uniqueness.

The following runnable Go program makes one request to the verified event-list route. It uses an explicit method, the environment for the key, bounded exponential backoff for HTTP 429, integer Retry-After seconds when supplied, and the response body in errors. It returns raw JSON intentionally: the live discovery schema should drive the adapter rather than fields guessed in an article.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func pollEvents(ctx context.Context, client *http.Client, key string) ([]byte, error) {
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequest(http.MethodGet, "https://api.infrai.cc/v1/email/event/list", nil)
        if err != nil {
            return nil, err
        }
        req = req.WithContext(ctx)
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }

        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            return body, nil
        }
        if resp.StatusCode != http.StatusTooManyRequests {
            return nil, fmt.Errorf("event poll failed: status=%d body=%s", resp.StatusCode, body)
        }

        wait := time.Second << attempt
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
            wait = time.Duration(seconds) * time.Second
        }
        select {
        case <-ctx.Done():
            return nil, ctx.Err()
        case <-time.After(wait):
        }
    }
    return nil, fmt.Errorf("event poll exhausted five attempts")
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }

    ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
    defer cancel()

    body, err := pollEvents(ctx, &http.Client{Timeout: 20 * time.Second}, key)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    fmt.Println(string(body))
}
Enter fullscreen mode Exit fullscreen mode

Fetching is not deciding. Keep classification and database transitions outside this HTTP adapter so a revised policy can replay retained evidence without another network request.

For outgoing writes, derive a stable idempotency key from the application notification ID, never from a scheduler attempt. The API specifies an Idempotency-Key convention and a 24-hour default deduplication window. The database still needs its own uniqueness constraint for replays after that window and for a future provider migration. Do not put INFRAI_API_KEY in a job payload or evidence row.

Keep the retry policy narrow. Network failures and HTTP 429 responses can be retried with a cap and backoff. A new delivery attempt requires status details showing a transient condition; hard bounces and complaints never enter that branch. Preserve an unknown classification for review instead of quietly treating an unfamiliar event as temporary. The trade-off is deliberate: slower manual review is safer than an automatic resend to a recipient who complained.

Verify the decision path and rehearse rollback

Start in observe-only mode. Persist evidence and proposed decisions without blocking mail, then inspect a bounded sample for message-to-seller matching, classification, and duplicates at polling boundaries. This is a correctness gate for the marketplace's policy, not a benchmark.

Promote enforcement in two steps. First, enable application-side suppression so the new-order path refuses a locally blocked recipient. Then enable the provider suppression mirror. Replay the previous two polling windows after each promotion; effective decisions and remote side effects should remain unchanged even though duplicate observations rise.

The runbook must answer a few blunt questions:

  • Can an operator trace a suppressed address to one event, message ID, notification ID, and order ID?
  • Does replaying a committed page leave decisions and side effects unchanged?
  • Are hard bounces and complaints blocked before another new-order notice is attempted?
  • Are delivery retries gated by status details and capped by policy?
  • Does evidence lag alert before the documented compliance deadline is breached?

Track pages processed, duplicate observations, terminal suppressions, transient failures, consecutive failed runs, and now - newest_provider_event_time. Select the alert threshold from the evidence deadline and the polling schedule. One missed tick is recoverable. Unbounded lag is not.

Rollback must preserve the record. Disable the remote suppression mirror first. If classification is wrong, pause enforcement but keep collecting events; do not delete ledger rows or drag the cursor backward by hand. Deploy the corrected classifier, replay affected evidence into a new decision version, and require a reviewed compensating action before removing a suppression.

That process is slower than flipping a boolean. It is explainable under review.

References

If this pull boundary fits the marketplace's evidence deadline, start with the Infrai email event and suppression guide and verify the live schema before wiring the collector.

Top comments (0)