DEV Community

SaxonFletcher2366
SaxonFletcher2366

Posted on

Gaming Email API Evidence — Bounce Handling, Complaint Suppression, Delivery Monitoring

For a gaming password-reset flow, choose an email API by the evidence it lets your team retain, not by the length of its feature list. TL;DR: require durable bounce and complaint events, an inspectable suppression state, and a poll interface with stable cursors or time windows. Then store an append-only local event ledger and update suppression idempotently. A webhook can reduce detection latency, but it should not be the only place where delivery evidence exists.

The reset token itself should expire quickly, while the delivery record should live according to a documented retention policy. Those are separate clocks. Mixing them makes both security review and incident reconstruction harder.

Pick Use it when Evidence requirement Main failure to test
Polling Inbound HTTP is restricted, or scheduled reconciliation is required Stable event IDs, server-side timestamps, pagination, and a documented retention window A page is replayed or the worker stops longer than the provider retains events
Webhook plus polling Fast reaction and later reconciliation both matter Signed pushes and a queryable event history A valid push arrives twice, late, or out of order
Export or log stream Audit and analytics pipelines already ingest immutable batches Versioned schema, delivery manifest, and deterministic object names A partial batch looks complete

That table is the first filter. If an API cannot explain event retention, pagination, replay behavior, and the boundary between a bounce and suppression, a polished dashboard does not repair the evidence gap.

Should an email API poll bounce and complaint events instead of webhooks?

Polling is a serious option when the security boundary does not permit public callbacks. Pick it when a five-minute detection interval is acceptable and the API exposes events long enough to cover the longest credible worker outage. The consumer must advance a cursor only after its database transaction commits. Otherwise a crash can create a silent hole.

Webhook plus polling is the stronger fit when operations needs prompt signals but compliance also needs reconciliation. Treat the webhook as an early notification. Treat the poller as the completeness check. Both paths feed the same idempotent event handler, keyed by a provider event ID or another documented immutable identifier.

An export or log stream fits teams with an established evidence pipeline. It trades freshness for clean batch boundaries. Require a manifest or equivalent completion marker; file arrival alone does not prove that a batch is complete.

The key distinction is easy to miss: an SMTP delivery status and a recipient complaint are different signals. Enhanced SMTP status codes classify delivery outcomes, while Abuse Reporting Format defines a machine-readable feedback report format. The application should preserve the original category and raw reference instead of flattening every negative result into failed.

The trade-off is latency versus recoverability.

Build the evidence ledger before the dashboard

Picture the flow in words: reset request, token record, send attempt, provider acceptance, delivery event, suppression decision, audit query. Each arrow creates a record. The dashboard is merely a view over those records.

Start with three identities: an internal message ID, the provider's message ID, and the provider event ID. Keep the reset token out of logs and event payloads. Store a one-way lookup value for the recipient if staff do not need to see the address during routine investigations, and restrict any reversible mapping separately.

The following TypeScript sketch uses generic interfaces. The poll boundary overlaps by two minutes on purpose. Re-reading is acceptable because the insert and suppression update are idempotent; missing an event is not.

type DeliveryKind = "delivered" | "transient_bounce" | "permanent_bounce" | "complaint";

type DeliveryEvent = {
  id: string;
  messageId: string;
  recipientKey: string;
  occurredAt: string;
  kind: DeliveryKind;
  diagnosticCode?: string;
};

interface EventSource {
  list(input: {
    after: string;
    cursor?: string;
    limit: number;
  }): Promise<{ events: DeliveryEvent[]; nextCursor?: string }>;
}

interface EvidenceStore {
  transaction<T>(work: (tx: EvidenceTransaction) => Promise<T>): Promise<T>;
  checkpoint(): Promise<{ after: string; cursor?: string }>;
}

interface EvidenceTransaction {
  insertEventIfAbsent(event: DeliveryEvent): Promise<boolean>;
  suppress(recipientKey: string, reason: "permanent_bounce" | "complaint"): Promise<void>;
  saveCheckpoint(checkpoint: { after: string; cursor?: string }): Promise<void>;
}

const overlapMs = 2 * 60 * 1000;

export async function reconcile(source: EventSource, store: EvidenceStore): Promise<void> {
  let checkpoint = await store.checkpoint();

  for (;;) {
    const page = await source.list({ ...checkpoint, limit: 250 });

    await store.transaction(async (tx) => {
      for (const event of page.events) {
        const inserted = await tx.insertEventIfAbsent(event);
        if (!inserted) continue;

        if (event.kind === "permanent_bounce" || event.kind === "complaint") {
          await tx.suppress(event.recipientKey, event.kind);
        }
      }

      const latest = page.events.at(-1)?.occurredAt ?? checkpoint.after;
      checkpoint = page.nextCursor
        ? { after: checkpoint.after, cursor: page.nextCursor }
        : { after: new Date(Date.parse(latest) - overlapMs).toISOString() };
      await tx.saveCheckpoint(checkpoint);
    });

    if (!page.nextCursor) return;
  }
}
Enter fullscreen mode Exit fullscreen mode

This is at-least-once ingestion. Good. A unique constraint on the event ID turns duplicates into harmless replays, and placing the checkpoint update in the same transaction prevents advancement past uncommitted evidence. The overlap is an engineering choice, not a universal constant; set it from the source's documented ordering behavior and measured arrival delay.

Walk through one concrete replay before trusting the design. Suppose a page contains 250 events and event 173 is a permanent bounce for a reset message. The transaction inserts all 250 rows, activates suppression for that recipient key, and saves the next cursor. If the process loses its connection before commit, none of those changes survive and the same page is requested again. If the commit succeeds but the process exits before starting the next request, the saved cursor resumes at the following page. If the source repeats event 173 in a later time-window query, the unique event ID makes the insert a no-op, so the old suppression evidence remains linked to one immutable event rather than acquiring a second invented cause. This walkthrough exposes the contract the API must support: repeated reads, stable identity, and a position that is meaningful after restart. A source that offers only a dashboard counter cannot support it. Neither can a poll endpoint whose history expires before the team can restore a failed consumer.

Replays are normal.

Suppression deserves its own state machine. A permanent bounce or complaint can create an active suppression record with the source event ID, reason, and timestamp. A transient bounce should remain evidence without automatically becoming permanent suppression. Any release needs an authorized action and a new audit record rather than deletion of the old one.

Test the gaps, not just the happy path

A successful API response proves acceptance of a request, not inbox placement. Tests should inject duplicate pages, reverse event order, repeat the same complaint, and stop the worker between inserting an event and saving its checkpoint. Also test the ugly boundary: the worker is unavailable longer than the upstream event-retention window. That condition must alert before evidence can expire.

Use three metrics with distinct meanings: age of the newest successfully committed poll, age of the newest event observed, and distance from the known retention boundary. A worker can poll successfully while receiving no data, so one green last_poll gauge is weak evidence. Alert on sustained checkpoint age and repeated page failures; route a retention-boundary warning early enough for recovery.

For the password-reset path, record token issuance and expiry without recording the token. RFC 6238 is about time-based one-time passwords and defines a default 30-second time step, but that value is not a general password-reset-email expiry rule. Choose the reset expiry from the threat model and user journey, document it, and test that an expired link fails even if the email arrives later.

One more trap: unsubscribe requirements are message-class dependent. RFC 8058 specifies one-click unsubscribe for list email through List-Unsubscribe and List-Unsubscribe-Post; it does not turn a password-reset security message into marketing mail. Keep transactional and promotional policy explicit, and do not use a complaint signal as permission to erase historical evidence.

What should the selection review demand?

Ask each candidate for a test account and verify the behavior yourself. A credible review captures the returned schema and answers a short set of questions:

  1. Can events be queried after a consumer outage, and for exactly how long?
  2. Are pagination and ordering guarantees documented? Is there a stable event ID?
  3. Can the team retrieve bounce class, diagnostic code, complaint signal, and suppression reason without scraping a dashboard?
  4. Can suppression changes be exported with timestamps and provenance?
  5. What authentication scopes separate sending, event reading, and suppression administration?
  6. How are schema changes announced, and can raw payloads be retained under the organization's data policy?

Run the proof with synthetic addresses and controlled fixtures, not real complaints. Save request IDs, timestamps, page cursors, and database rows from the exercise. The result is a before-and-after trail: before ingestion, the upstream event exists; after one run, one ledger row and one justified state transition exist; after replay, the counts do not change.

That replay check catches a surprising amount of trouble.

Limits and decision rule

Polling cannot make an upstream history more durable than its published retention window, and it adds detection delay. It is not suitable when suppression must react faster than the chosen poll interval or when the source exposes no replayable history. Webhooks cannot guarantee completeness merely because the endpoint returned success, but signed webhooks are the better primary signal when seconds matter. Exports may arrive too slowly for operational response; they are the better fit only when batch evidence is acceptable. These limitations are boundaries to measure, not labels that select a winner.

Choose the option that can prove completeness across your longest planned outage and can replay every state transition without changing the result. For a short-lived gaming reset link, separately prove that late delivery never extends token validity. If no candidate can provide those two proofs, the system is not ready for a compliance claim.

Sources

Top comments (0)