A compliance notice changes the design constraint: "accepted by an API" is not the evidence the business needs. The useful choice is a small state machine that records every submission and status observation, schedules bounded retries after throttling, and preserves channel-specific evidence without pretending that email and SMS have identical outcomes.
TL;DR: create one internal notice ID, one attempt record per channel submission, and an append-only observation log. Poll only nonterminal attempts. Treat 429 as a scheduling signal, add jitter to exponential backoff, and cap both elapsed time and attempt count. The result is an auditable record even when the final state remains unknown.
How should email and SMS API event notifications handle polling?
The send response proves one narrow event: a remote system accepted a request. It does not prove that a mailbox accepted the message, that a handset received it, or that a person read it. Those are different claims. A defensible record keeps them separate.
For this build, the evidence unit is a compliance notice, not an HTTP request. It contains the notice version, recipient reference, channel, timestamps, a stable idempotency key, the remote message reference, and every state transition observed. Store a digest of the rendered content rather than duplicating sensitive message text in operational logs. Access to the evidence store should be narrower than access to ordinary application telemetry.
Email adds another boundary. DMARC defines domain-level handling and reporting based on SPF and DKIM alignment; it does not turn a delivery callback into proof that a human read a notice. Authentication guidance also warns against treating email as an out-of-band authenticator. Those standards support a conservative evidence vocabulary: submitted, accepted, delivered, failed, and unknown, with the exact source attached.
Unknown is honest.
Consider one concrete timeline. At 09:00:00, the application commits notice N-1042, two channel attempts, and the rendered-content digest in one transaction. The email submission returns a remote reference; the SMS submission times out after the remote side may have accepted it. The email attempt can move to accepted. The SMS attempt cannot safely be submitted under a new key, so the worker retries with the original idempotency key and records that this was a transport retry. At 09:00:02, the first email status poll is throttled. That observation belongs in the ledger, but it says nothing about delivery. The scheduler calculates a later due time. Meanwhile, the SMS adapter recovers the original remote reference and begins status polling. If email eventually reports delivered while SMS stays nonterminal until the 30-minute local budget closes, the notice record contains one delivered channel and one unknown channel. It does not collapse them into a comforting overall success. This is the practical trade-off: the record is less tidy, but every claim can be traced to a source and timestamp.
Build the smallest evidence ledger
I would keep the transport adapter deliberately boring. The application owns normalized states and evidence. An adapter owns remote field names. This avoids spraying provider-shaped objects across the CLI, worker, and audit export.
type Channel = "email" | "sms";
type DeliveryState =
| "queued"
| "submitted"
| "accepted"
| "delivered"
| "failed"
| "unknown";
type Observation = {
noticeId: string;
attemptId: string;
channel: Channel;
state: DeliveryState;
observedAt: string;
source: "submit-response" | "status-poll" | "timeout";
remoteRef?: string;
detailCode?: string;
};
interface DeliveryTransport {
submit(input: {
channel: Channel;
recipientRef: string;
contentDigest: string;
idempotencyKey: string;
}): Promise<{ remoteRef: string; state: "submitted" | "accepted" }>;
status(remoteRef: string): Promise<{
state: "accepted" | "delivered" | "failed" | "unknown";
detailCode?: string;
}>;
}
The ledger should reject mutation of old observations. Corrections become new observations. That gives an export a traceable order and stops a late poll from silently rewriting what the system knew earlier.
The content digest needs a versioned canonical input. For example, hash the template version, locale, subject, and rendered body in a documented order. Do not hash an arbitrary object serialization and hope every runtime orders it the same way.
Poll with a budget, not a loop
An unbounded poller is a rate-limit amplifier. A fixed interval is little better: thousands of notices created together will wake together. The scheduling policy below uses exponential growth, bounded random jitter, and a hard ceiling. The numbers are policy choices for this example, not claims about an external API.
const policy = {
baseDelayMs: 2_000,
maxDelayMs: 60_000,
maxAttempts: 10,
maxElapsedMs: 30 * 60_000,
};
function nextDelayMs(attempt: number, retryAfterMs?: number): number {
if (retryAfterMs !== undefined) {
return Math.min(Math.max(retryAfterMs, 0), policy.maxDelayMs);
}
const ceiling = Math.min(
policy.baseDelayMs * 2 ** Math.max(0, attempt - 1),
policy.maxDelayMs,
);
return Math.floor(Math.random() * (ceiling + 1));
}
function mayPoll(input: {
attempt: number;
startedAtMs: number;
nowMs: number;
}): boolean {
return (
input.attempt < policy.maxAttempts &&
input.nowMs - input.startedAtMs < policy.maxElapsedMs
);
}
When a status request returns 429, the worker records that observation, honors a parsed retry delay when the adapter supplies one, and reschedules. It does not mark delivery failed. If the retry budget expires, append unknown with source timeout; never manufacture delivered to make a dashboard green.
The important retry boundary is the operation. Retrying a status read is normally distinct from resubmitting a notice. A submission retry must reuse the same idempotency key, and the ledger still records each local attempt. Otherwise a network timeout can become two compliance notices.
type PollResult =
| { kind: "state"; state: DeliveryState; detailCode?: string }
| { kind: "throttled"; retryAfterMs?: number };
function planNext(result: PollResult, attempt: number) {
if (result.kind === "state") {
const terminal = result.state === "delivered" || result.state === "failed";
return terminal
? { action: "stop" as const }
: { action: "schedule" as const, delayMs: nextDelayMs(attempt) };
}
return {
action: "schedule" as const,
delayMs: nextDelayMs(attempt, result.retryAfterMs),
};
}
I benchmark this scheduler by behavior, not requests per second alone. The useful checks are maximum concurrent polls after a burst, time spent in nonterminal states, duplicate submission count, and the proportion that ends as unknown. A fast poller that creates duplicate notices has failed the job.
The happy path needs one fixture. The awkward paths deserve most of the suite: submission accepted and response lost; three consecutive throttles; a terminal failure after an earlier accepted state; a late delivered observation after the local timeout; and two workers claiming the same due poll.
Use a fake clock and seeded randomness so delay assertions stay deterministic. Test that terminal attempts are never polled again. Test that a stale worker cannot overwrite a newer lease. Test that an audit export includes the policy version used for each decision.
Keep personally identifying recipient data out of fixtures. A stable pseudonymous reference is enough to prove correlation, while the production lookup can remain behind its own access boundary.
What I would change at scale
At low volume, a database table with next_poll_at and a worker using atomic claims is enough. At higher volume, partition due work by time bucket and channel, then place concurrency limits at both the global and transport-adapter levels. The ledger remains the source of evidence; the queue is only a delivery mechanism for work.
There is a real trade-off here. Polling gives the application explicit control over cadence and recovery, but it consumes calls while nothing changes. Push callbacks reduce idle polling, yet they introduce signature verification, replay handling, public ingress, and out-of-order events. A mixed design can ingest callbacks and keep sparse reconciliation polls for attempts that remain nonterminal. Pick from measured state lag and evidence requirements, not from config fashion.
The operational view should answer four questions without opening raw logs: which notices lack a terminal observation, which are waiting because of throttling, which submissions may be duplicates, and which evidence exports failed validation. Alert on age and backlog, not every individual 429.
For the final audit package, export the notice ID, content digest and version, recipient reference, channel, idempotency key, remote reference, ordered observations, timestamps, and normalization-policy version. Sign or otherwise protect the export according to the surrounding compliance system. The core rule stays modest: report what each source observed, and no more.
Sources
- RFC 7489: Domain-based Message Authentication, Reporting, and Conformance (DMARC): https://datatracker.ietf.org/doc/html/rfc7489
- NIST SP 800-63B, Digital Identity Guidelines: Authentication and Lifecycle Management: https://pages.nist.gov/800-63-3/sp800-63b.html
Top comments (0)