DEV Community

IversonBlake8417
IversonBlake8417

Posted on

Marketplace Event Notifications Explained: 3 Ledgers for Batch Email and SMS Partial Failure

TL;DR: For marketplace event notifications, use one durable row per recipient and channel, then let callbacks or polling advance that row through a monotonic state machine. The least complex useful option is a recipient ledger with immutable attempt records. A batch is only a grouping key; it must never be the source of truth for delivery.

Ledger shape Pick this when Main trade-off Compliance evidence
Batch row plus counters Every recipient can be retried together and individual proof is unnecessary Small schema, but partial failure is ambiguous Weak: totals do not identify the affected destination
Recipient row with current state You need suppression and support lookup per destination Easy to query; state changes can erase history Moderate: retain timestamps and reason categories
Recipient row plus immutable attempts You must show what was attempted, observed, suppressed, and retried More writes and retention work Strong: each decision has an ordered record

For a marketplace, choose the third shape when a seller asks why 37 order updates produced 34 deliveries, two permanent email bounces, and one unresolved SMS receipt. It gives support a defensible answer without replaying the entire batch. It also keeps one bad destination from holding 36 healthy ones hostage.

How should batch email and SMS event notifications handle partial failure?

The usual culprit is a mismatch between the unit of work and the unit of evidence. The queue carries batch-814, so the worker waits for one batch-level success. Delivery systems report outcomes for individual destinations. One email can be accepted while another receives a permanent SMTP failure; an SMS submission can be accepted before its final receipt arrives. A lone unresolved recipient then leaves the batch labeled processing, even though most useful work finished.

Picture the flow as words: marketplace event to recipient expansion to suppression check to channel submission to provider observation to normalized recipient state. The batch sits above that line as a correlation ID. It does not move through the line itself.

This distinction matters for compliance evidence. sent: 36 is an aggregate, not an explanation. Evidence needs the internal destination identifier, template revision, attempt time, normalized outcome, provider correlation value, and rule that allowed or suppressed the attempt. Store sensitive destinations in the form your retention and access policy permits.

A bounce is not permission to retry blindly. SMTP delivery status notifications define structured status information, including whether a failure is persistent or temporary. Yahoo's sender guidance says to remove invalid recipients and keep complaint rates low. The operational translation is direct: classify observations, suppress permanent failures, and bound retries for transient ones.

Fast retries are not progress.

One row wins.

Pick the ledger that matches the proof you owe

Batch counters fit low-stakes fan-out where the audience can be regenerated and no per-recipient claim will be investigated. They are useful dashboard material. Do not let them drive retries, because failed = 2 cannot tell a worker which two destinations are safe to try again.

Current recipient state is the practical middle ground. Pick it when support needs a quick answer and the audit requirement covers only the latest decision. Add compare-and-set updates so a delayed delivered callback cannot be overwritten by an older submitted poll result. This model is compact, but debugging a disputed transition becomes hard after the old value disappears.

Recipient state plus attempts fits the marketplace case. Each attempt is append-only; the recipient row is a projection for fast reads. A callback and poll may describe the same provider operation, so deduplicate observations by provider event ID when one exists, or by a narrowly defined composite key. Keep raw payloads under a controlled retention policy and expose a normalized reason to application code. The raw event is evidence. The normalized state drives behavior.

There is a real trade-off: more storage, a reconciliation job, and a schema that must survive new reason codes. I would still pick the attempt ledger when the answer controls suppression or demonstrates that a notification was handled according to policy. The reason is practical. Losing an old state saves a write but also removes the evidence needed to explain the next decision.

Implement monotonic recipient transitions

The state machine should separate submission from outcome. accepted means the channel accepted a request for processing; it does not prove delivery. Terminal outcomes close one attempt. suppressed means no send should occur until an explicit policy action changes eligibility.

Here is a small TypeScript core. It has no vendor SDK and no HTTP route. Adapters translate callbacks and poll results into Observation, while the ledger owns transitions.

type Channel = "email" | "sms";
type State =
  | "queued"
  | "accepted"
  | "delivered"
  | "transient_failure"
  | "permanent_failure"
  | "suppressed";

type Recipient = {
  batchId: string;
  recipientId: string;
  channel: Channel;
  state: State;
  version: number;
  attempts: number;
  nextCheckAt?: Date;
};

type Observation = {
  operationId: string;
  observedAt: Date;
  state: Exclude<State, "queued" | "suppressed">;
  reason: string;
};

const terminal = new Set<State>([
  "delivered",
  "permanent_failure",
  "suppressed",
]);

function applyObservation(
  recipient: Recipient,
  observation: Observation,
  operationAlreadyRecorded: boolean,
): Recipient | undefined {
  if (operationAlreadyRecorded || terminal.has(recipient.state)) return undefined;

  if (observation.state === "permanent_failure") {
    return {
      ...recipient,
      state: "suppressed",
      version: recipient.version + 1,
      nextCheckAt: undefined,
    };
  }

  const shouldRecheck =
    observation.state === "accepted" ||
    observation.state === "transient_failure";

  return {
    ...recipient,
    state: observation.state,
    version: recipient.version + 1,
    nextCheckAt: shouldRecheck
      ? new Date(observation.observedAt.getTime() + 5 * 60_000)
      : undefined,
  };
}
Enter fullscreen mode Exit fullscreen mode

The five-minute value is an example policy, not a delivery guarantee. Set it from channel behavior, published limits, and the notification deadline. Add jitter so a large batch does not wake every pending row simultaneously. Cap attempts and elapsed age separately: three rapid attempts and one attempt waiting for a delayed receipt are different operational stories.

Persist the attempt and update the projection in one database transaction. Use version in the update predicate. If another worker wins, reload the row, record the observation only if it is still new, and evaluate again. Callback-versus-poller races become ordinary concurrency.

Templates belong in the evidence chain too. Record a content revision or immutable template hash with each attempt, not rendered message bodies by default. Mustache escapes HTML variables by default; triple braces or an ampersand render unescaped content, so template review must treat those forms as a deliberate security decision. Channel adapters can render from the same approved event data while allowing email and SMS copy to differ.

Observe recipients, alert on age

A useful dashboard has two levels. Batch counters explain scope. Recipient age explains action. Track counts by normalized state and channel, plus the age of the oldest nonterminal recipient. Alerting on generic queue depth confuses a planned marketplace spike with a stalled reconciliation path. Use structured logs keyed by batchId, recipientId, channel, attemptId, and operationId, but do not put a raw email address, phone number, message body, or callback payload in general-purpose logs. The controlled evidence store and operational log serve different readers and usually need different retention and access rules. Start an investigation with one recipient: read its ordered attempts, compare the latest observation time with nextCheckAt, and inspect the suppression decision. Then widen to peers with the same channel, template revision, reason category, and submission window. This sequence finds a shared failure without turning every partial batch into a mass retry. It also exposes the useful contrast quickly: one aging row suggests a recipient-specific path, while many rows with the same template revision or submission window point toward shared processing.

Test the awkward orders. Deliver a duplicate callback. Apply a poll result after a terminal callback. Return one permanent failure among 99 accepted recipients. Crash after the channel accepts a request but before the local transaction commits. The last case requires an idempotency strategy at the adapter boundary or reconciliation by operation ID; the state machine alone cannot prove whether an unrecorded submission occurred.

Limits and decision rule

No generic ledger makes email and SMS semantics identical. SMTP status reports and mobile delivery receipts have different vocabularies, timing, and coverage. Normalize only states the application can act on, preserve the original reason as evidence, and avoid inventing precision an upstream channel did not provide.

Use batch counters for display, current recipient rows for straightforward support, and immutable attempts when suppression or compliance questions require a timeline. For marketplace bounces, retry the recipient, never the counter; suppress on a classified permanent failure; and let an unresolved receipt age into reconciliation without blocking completed peers.

Further reading

Top comments (0)