TL;DR: Treat an email or SMS request timeout as an unknown submission result, not a failed notification. Persist the external message ID when one exists, poll for a terminal status on a bounded schedule, and make recipient suppression a separate, idempotent decision. Keep clinical template content owned by the application; let the delivery adapter own transport rendering details and status translation.
The crucial distinction is small: request outcome is not delivery outcome. A timeout says the caller stopped waiting. It does not prove rejection, acceptance, or delivery. Blind retrying can create duplicate appointment reminders. Blindly marking failure can hide a bad address that should be suppressed.
Replace the send result with an evidence ledger
The before model is tempting: call send, wait, then store sent or failed. It compresses several independent events into one boolean. SMTP can accept a message before a later non-delivery report, while an SMS submission can receive an identifier before its delivery state changes. DKIM authenticates domain-level responsibility for an email message; it is not delivery proof.
The after model is a ledger. In words: application event to immutable notification intent; intent to channel attempt; attempt to external ID; external ID to observations; terminal negative observation to suppression decision. Each arrow can happen at a different time. Store enough data to answer four questions without opening the body: which recipient key was targeted, which template revision produced the content, which attempt was submitted, and what evidence changed its state. In a healthtech system, logs can use an opaque recipient key rather than repeating an address or phone number. Template ownership matters here: the application should own appointment-reminder@17, its approved wording, and its rendering data contract, while a channel adapter owns MIME construction, SMS segment handling, and translation from external states. If the transport owns the only template copy, investigators cannot reliably reproduce which revision was attempted; changing transport also changes a content-governance boundary that has nothing to do with polling.
One intent may have several attempts, but one attempt must never pretend to be another. Use a stable idempotency key derived from intent and channel, not mutable template text.
How should a Node.js cron worker handle email and SMS event notification timeouts?
First, look for durable evidence from the original call. If an external ID was received before the client timed out, store it and enqueue reconciliation. With no ID, keep the attempt in submission_unknown; do not label it failed. Retrying an ambiguous submission without an idempotency guarantee may send twice. Refusing ever to retry may omit a time-sensitive reminder. That policy belongs beside the notification intent, not inside a generic HTTP retry helper.
Use a small vocabulary: submission_unknown, accepted, delivered, transient_failure, permanent_failure, and review_required. Keep suppressed on the recipient-channel relationship rather than pretending it is a delivery state.
A polling schedule needs bounds. For example, poll after 30 seconds, 2 minutes, 10 minutes, and 30 minutes, then require review if no terminal evidence appears. Those intervals are an example policy, not a standard. Choose them from notification urgency, status retention, and polling limits.
A copyable TypeScript reconciliation worker
This generic adapter separates lookup, ledger updates, and suppression. The database implementation should use row locking or a lease so two cron processes cannot reconcile the same attempt concurrently.
type State = "accepted" | "delivered" | "transient_failure" | "permanent_failure";
type Attempt = {
id: string;
recipientKey: string;
channel: "email" | "sms";
externalId: string | null;
pollCount: number;
};
type Observation = { state: State; reasonCode?: string; observedAt: Date };
interface Adapter {
lookup(externalId: string): Promise<Observation>;
}
interface Ledger {
claimDue(limit: number): Promise<Attempt[]>;
apply(id: string, observation: Observation): Promise<void>;
defer(id: string, at: Date): Promise<void>;
requireReview(id: string, reason: string): Promise<void>;
}
interface Suppressions {
add(input: {
recipientKey: string;
channel: Attempt["channel"];
reasonCode: string;
sourceAttemptId: string;
}): Promise<void>;
}
const delays = [30_000, 120_000, 600_000, 1_800_000] as const;
export async function reconcile(
ledger: Ledger,
suppressions: Suppressions,
adapters: Record<Attempt["channel"], Adapter>,
now = new Date(),
): Promise<void> {
for (const attempt of await ledger.claimDue(100)) {
if (!attempt.externalId) {
await ledger.requireReview(attempt.id, "missing_external_id");
continue;
}
try {
const observed = await adapters[attempt.channel].lookup(attempt.externalId);
await ledger.apply(attempt.id, observed);
if (observed.state === "permanent_failure" && observed.reasonCode) {
await suppressions.add({
recipientKey: attempt.recipientKey,
channel: attempt.channel,
reasonCode: observed.reasonCode,
sourceAttemptId: attempt.id,
});
}
} catch {
const delay = delays[attempt.pollCount];
if (delay === undefined) {
await ledger.requireReview(attempt.id, "poll_budget_exhausted");
} else {
await ledger.defer(attempt.id, new Date(now.getTime() + delay));
}
}
}
}
The 100 cap keeps one cron run finite; it is not a throughput claim. Make state changes monotonic so an older accepted observation cannot overwrite delivered or permanent_failure. Make suppression insertion idempotent too.
Do not suppress on every failure-shaped string. A malformed template, expired credential, throttled lookup, or temporary remote error says nothing about recipient validity. Normalize external codes into a small internal taxonomy and approve which terminal categories create suppression. Preserve the raw code in restricted diagnostics.
Count transitions by channel and state. Measure the oldest nonterminal attempt, not only queue depth. Alert on exhausted poll budgets and abrupt permanent-failure changes. Logs need attempt ID, opaque recipient key, template revision, old state, new state, poll count, and normalized reason. They do not need rendered content.
Two objections change the design
Why not mark accepted as delivered? Acceptance answers a narrower question. SMTP success means a receiver accepted responsibility under the protocol; delivery status notifications are a separate reporting mechanism. SMS state can also evolve after message creation. Combining the states removes evidence needed for a later permanent failure.
Why not let the delivery service own templates? That can support an external editing workflow, but the ledger still needs a stable template identifier and revision. The application must validate the data contract, and tests must pin a revision. A useful middle path keeps canonical content and its schema under application governance, then publishes a versioned rendering artifact to the transport layer.
This polling design has a real limitation: it is a poor fit when a transport offers no stable status lookup, retains status for less time than the polling window, or charges operationally significant capacity for frequent reads. It also trades faster state updates for scheduled load. In those cases, use a standards-based delivery report or authenticated callback when available, while preserving the same ledger and suppression policy. The callback changes evidence ingestion, not template ownership.
Test accepted, delivered, transient failure, permanent invalid-recipient failure, an unknown code, a lookup timeout, and an out-of-order observation. Then run one concurrency test with two worker instances claiming the same record. Expect one transition and one idempotent suppression entry.
Short rule: kill the duplicate first.
A timeout creates uncertainty. Polling reduces it; polling does not justify guessing. Keep content revisions under explicit application ownership, transport observations in an append-friendly ledger, and suppression behind a mapped terminal signal. This makes bounces actionable without turning a cron worker into a second template system.
Top comments (0)