DEV Community

RhysFalconer159
RhysFalconer159

Posted on

Nodejs Event Notifications for Transactional Email and SMS Status Polling

For a media service, Nodejs event notifications become tricky when transactional email and SMS alerts need delivery status but can't receive webhooks. That constraint changes the design: send each alert once, poll for its result, and make suppression a durable decision before the next edition or breaking-news batch starts.

TL;DR: use a small queue with separate send and receipt-check jobs. Give every logical alert a stable key, store the remote message identifier returned by the transport, and poll only messages that still have a nonterminal status. A hard email bounce or explicit SMS invalid-recipient result suppresses that channel. A timeout, temporary rejection, or unavailable status endpoint schedules another check with bounded backoff. Cron should wake workers, not contain the workflow.

This is less glue than pretending a scheduled script can remember everything. It also works when inbound webhooks are unavailable or prohibited by the deployment boundary. The important metric isn't how quickly the cron process exits. It is how many recipient-channel pairs reach one unambiguous terminal state without receiving a duplicate alert. The trade-off is delayed knowledge: polling spends read capacity and can only discover a transition on the next scheduled check, so this design is a poor fit when a verified status must trigger another action within seconds.

How should Nodejs poll transactional email and SMS delivery status?

The tempting implementation is one loop: select subscribers, send email and SMS, sleep, ask for status, retry failures. It looks compact. It collapses four different clocks into one process.

It won't stay compact.

Publishing controls when an alert becomes useful. A transport controls when it accepts a message. Delivery can settle later. Suppression must survive every worker restart and affect future campaigns. If those facts live in local variables, a crash between the send call and the database update can produce a duplicate. If the same cron run retries every non-success result, it can also treat “still processing” as “send again.” Those are different states.

The smallest useful state machine is boring:

State Meaning Next action
queued No accepted send has been recorded Attempt one send
accepted Transport returned a message ID Schedule a receipt check
pending Receipt is not terminal Poll later
delivered Terminal success Stop
failed Terminal failure that is not an invalid address/number Stop or review by policy
suppressed Recipient-channel pair is invalid Stop and block future sends

Do not map email opens to delivery. Apple documents that Mail Privacy Protection can download remote content in the background, which prevents a sender from learning reliably whether the recipient opened a message. Opens are an engagement signal with a damaged measuring instrument, not a suppression input.

Email authentication is a separate layer too. DMARC defines domain-level policy and reporting around SPF and DKIM alignment. It does not replace per-recipient bounce handling. Keep authentication failures visible in operations, but do not confuse a domain policy problem with proof that one mailbox is invalid.

Different evidence. Different action.

The smallest working pull loop

I judge this interface by time-to-first-call and the amount of adapter code it forces into the application. Two transport methods are enough: one submits an alert, and one reads a receipt. Provider-specific payloads stop at that boundary.

type Channel = "email" | "sms";
type JobState =
  | "queued"
  | "accepted"
  | "pending"
  | "delivered"
  | "failed"
  | "suppressed";

type AlertJob = {
  id: string;
  editionId: string;
  recipientId: string;
  channel: Channel;
  destination: string;
  state: JobState;
  remoteId?: string;
  attempts: number;
  nextRunAt: Date;
};

type Receipt =
  | { state: "pending" }
  | { state: "delivered" }
  | { state: "failed"; invalidRecipient: boolean; reason: string };

interface Transport {
  send(input: {
    destination: string;
    body: string;
    idempotencyKey: string;
  }): Promise<{ remoteId: string }>;
  receipt(remoteId: string): Promise<Receipt>;
}
Enter fullscreen mode Exit fullscreen mode

The database needs a unique constraint on (editionId, recipientId, channel). That stable tuple is the logical alert identity. It prevents two schedulers from creating two jobs, and it gives the transport adapter an idempotency key where that capability exists. The local constraint is still required; a remote guarantee cannot protect inserts that never reached the remote system.

Claim due rows with a lease. The exact database syntax varies, so the repository below exposes the atomic operations instead of hiding dubious SQL in a tutorial.

interface JobRepository {
  claimDue(limit: number, leaseUntil: Date): Promise<AlertJob[]>;
  markAccepted(id: string, remoteId: string, nextRunAt: Date): Promise<void>;
  markPending(id: string, nextRunAt: Date): Promise<void>;
  markDelivered(id: string): Promise<void>;
  markFailed(id: string, reason: string): Promise<void>;
  suppress(id: string, recipientId: string, channel: Channel, reason: string): Promise<void>;
  release(id: string, nextRunAt: Date): Promise<void>;
}

const delayMs = (attempt: number): number => {
  const cappedAttempt = Math.min(attempt, 6);
  return 15_000 * 2 ** cappedAttempt;
};

async function processJob(
  job: AlertJob,
  transport: Transport,
  jobs: JobRepository,
  now: Date,
): Promise<void> {
  try {
    if (!job.remoteId) {
      const sent = await transport.send({
        destination: job.destination,
        body: "A new edition is ready.",
        idempotencyKey: `${job.editionId}:${job.recipientId}:${job.channel}`,
      });
      await jobs.markAccepted(
        job.id,
        sent.remoteId,
        new Date(now.getTime() + delayMs(job.attempts)),
      );
      return;
    }

    const receipt = await transport.receipt(job.remoteId);
    if (receipt.state === "delivered") {
      await jobs.markDelivered(job.id);
      return;
    }
    if (receipt.state === "failed" && receipt.invalidRecipient) {
      await jobs.suppress(
        job.id,
        job.recipientId,
        job.channel,
        receipt.reason,
      );
      return;
    }
    if (receipt.state === "failed") {
      await jobs.markFailed(job.id, receipt.reason);
      return;
    }

    await jobs.markPending(
      job.id,
      new Date(now.getTime() + delayMs(job.attempts)),
    );
  } catch {
    await jobs.release(
      job.id,
      new Date(now.getTime() + delayMs(job.attempts)),
    );
  }
}
Enter fullscreen mode Exit fullscreen mode

There is a deliberate omission: the catch block does not send again when remoteId exists. It releases the same receipt-check job. Once a send has been accepted and recorded, retrying delivery means polling the existing message, not creating another one.

No second send.

The uncomfortable gap is a crash after the remote service accepts a send but before markAccepted commits. A stable idempotency key can close it if the transport supports deduplication. Without that contract, exactly-once delivery is not available. Record this as an explicit integration requirement, because no retry library can manufacture atomicity across your database and an external API. Pulling also has a hard operational limit: if the source doesn't retain receipts long enough for the worst queue delay, the worker can lose the chance to observe a terminal result. In that case, use a push callback or a transport that offers a sufficient status-retention contract rather than increasing retries forever.

Cron is only the metronome

Run one short scheduler tick for both regions, using region-local storage and workers when data residency requires it. The tick claims a bounded number of due rows with a lease, processes them with limited concurrency, and exits. Another worker can reclaim a job after its lease expires.

async function tick(
  jobs: JobRepository,
  transportFor: (channel: Channel) => Transport,
  now = new Date(),
): Promise<void> {
  const leaseUntil = new Date(now.getTime() + 60_000);
  const due = await jobs.claimDue(100, leaseUntil);

  for (let offset = 0; offset < due.length; offset += 10) {
    const batch = due.slice(offset, offset + 10);
    await Promise.all(
      batch.map((job) => processJob(job, transportFor(job.channel), jobs, now)),
    );
  }
}
Enter fullscreen mode Exit fullscreen mode

The numbers here are starting bounds, not benchmarks. A 60-second lease, 100 claimed rows, batches of 10, and six backoff doublings make the behavior concrete enough to test. Measure receipt latency, rate-limit responses, queue age, and lease expirations, then tune from observed distributions. I would reject any SDK evaluation that required those four controls to be scattered across callback hooks and config files.

Test the failure edges with a fake transport: accepted then pending then delivered; invalid recipient; transient exception during polling; two workers claiming simultaneously; and a process exit immediately after send acceptance. Also test suppression at enqueue time. Blocking only inside the worker leaves suppressed jobs accumulating in the queue and makes queue depth lie.

What I would change at scale

First, split submission and receipt polling into separate queues. They have different rate limits, latency, and urgency. A breaking-news send should not sit behind thousands of slow receipt checks. Keep one persisted state machine, though. Duplicating state across two queue products turns reconciliation into the real system.

Second, store suppression as (recipientId, channel) plus reason, source message, and timestamp. Email and SMS destinations fail independently. A bad phone number should not silence a valid mailbox. Changes to a destination should create a fresh validation path rather than silently clearing history.

Third, add a terminal deadline. Polling forever is config bloat disguised as reliability. After the documented receipt window has elapsed, move the job to an explicit unresolved state for review and metrics; do not label it delivered, and do not resend it automatically.

The regional design needs one owner for each logical job. Active-active schedulers pointed at replicated but independently writable queues can both send before replication converges. Partition ownership by recipient or edition, or route writes to one home region. Failover should transfer leases and ownership, not replay every recent alert.

Observability stays small: queue age by channel and region, time from acceptance to terminal state, suppression counts by normalized reason, lease recoveries, and jobs past deadline. Log the job ID and remote ID, but avoid putting email addresses, phone numbers, or message bodies in routine logs. Those fields do not help a latency graph.

The integration decision rule

Choose an integration only after proving three contracts in a disposable adapter: submission returns a durable message identifier; status lookup distinguishes pending from terminal outcomes; and invalid-recipient results can be normalized without parsing prose. Then verify idempotency behavior and status retention against the transport's current documentation.

I would benchmark the adapter by code and operations, not brochure breadth: lines of mapping code, calls required for one terminal result, configuration keys, behavior after a timeout, and the number of states that cannot be mapped. A thin SDK can still be expensive if every error arrives as an unstructured string. A larger SDK can be tolerable if its types make terminality explicit. The boundary matters more than package size.

This design does not promise exactly once. It makes duplicates bounded, suppression durable, and uncertainty visible. For media alerts without webhooks, that is the honest target: one logical job, one recorded submission, repeated reads of the same receipt, and no future contact after a verified invalid-recipient result.

Sources

Top comments (0)