A healthtech signup system has one operational constraint that changes the design: after an email or SMS verification attempt, the team must be able to reconstruct what the system knew, when it knew it, and what action followed. Choose an append-only event ledger over a mutable delivery-status snapshot when compliance evidence matters. A snapshot is acceptable only when transport history has no audit value and the application needs nothing beyond the current state.
TL;DR: write the attempt before dispatch, append provider observations from either callbacks or polling, record HTTP 429 scheduling decisions, and keep link consumption as a separate application event. Polling versus callbacks is a collection detail. The durable comparison is between retaining evidence and repeatedly overwriting it.
What must the incident record answer?
Start at the postmortem, not at the messaging API. For one signup attempt, an investigator should be able to establish that the application created a verification challenge, requested delivery through a named channel, received a transport observation, and either accepted or did not accept the challenge before its expiry. Those are separate claims. Compressing them into sent=true loses the distinctions before an incident even begins.
This is the invariant: each observation is immutable, attributable to a source, timestamped both by that source and by the receiving system when those values exist, and correlated to an internal attempt identifier. A later observation may change the derived view, but it must not erase an earlier one.
The distinction matters because email authentication and message delivery answer different questions. DMARC defines policy and reporting based on alignment with SPF and DKIM identifiers; it does not prove that a particular recipient opened a mailbox or used a verification link. NIST SP 800-63B also treats PSTN out-of-band authentication as restricted and discusses risks including SIM change and number porting. A carrier delivery observation cannot be promoted into proof of identity. The application event that consumes the verification token is the relevant evidence for that step.
What page fired? If the answer is “the provider dashboard turned red,” the alert has skipped the user objective and the local evidence gap. A useful page identifies either an impaired verification outcome or a reconciliation queue whose oldest unresolved item is approaching the link deadline. A chart of raw send failures may help diagnosis, but it is not the incident definition.
Should Email and SMS API Event Notifications Rely on Polling?
A mutable row is attractive because reads are easy. One message ID, one status, one timestamp. It also makes the hardest failure invisible: an older observation can arrive after a newer one and overwrite it, leaving no record that the regression came from reordering rather than from the transport itself.
An event ledger costs more storage and requires a projection rule. In exchange, it preserves duplicates, reordering, the source of each observation, and the scheduler's response to rate limiting. That is evidence an incident responder can replay instead of interpreting whatever value survived the last write. Its main limitation is operational complexity: schema evolution, retention enforcement, replay tooling, projection repair, and access review all become owned production work. If the team cannot operate those controls, a carefully guarded snapshot is the more honest choice.
| Decision | Append-only event ledger | Mutable status snapshot |
|---|---|---|
| Audit question | Reconstructs observations and decisions | Shows the last stored interpretation |
| Late data | Retained without forcing state regression | Can overwrite newer state unless guarded |
| Duplicate input | Visible and deduplicated by policy | Often disappears into the current value |
| Operational burden | Projection, retention, access control | Simpler reads and less stored history |
| Appropriate boundary | Regulated or disputed workflows | Low-consequence, current-state-only workflows |
The ledger does not require one collection mechanism. Authenticated callbacks can append observations as they arrive. Polling can append snapshots obtained later. Either can be delayed, duplicated, or unavailable, so the projection should apply explicit monotonic transition rules while retaining the original provider vocabulary. Do not make a normalized delivered label carry more certainty than its source provides.
Run a single attempt through the ugly order before approving the schema. At 10:00:00 the signup service commits challenge creation, then dispatches an email and stores the provider message ID; at 10:00:04 a polling worker reads an accepted state, but stalls before its local write; at 10:00:06 an authenticated callback reports a later transport state and commits first; at 10:00:08 the stalled worker resumes with its older observation; at 10:00:10 the user consumes the link. A snapshot implementation needs conditional updates precise enough to reject the stale write, yet after rejection it may retain no evidence that polling ran or what it saw. The ledger appends both transport observations with their source times and local receipt times, refuses to regress the derived state, and records token consumption on the application timeline. No claim here depends on those illustrative times being a measured provider latency; they expose an ordering that concurrency permits. Now add a duplicate callback and a worker restart. If the audit view still produces one intelligible timeline without deleting inconvenient inputs, the model is doing useful work.
Reordering is normal.
Store an internal attempt ID, channel, pseudonymous destination reference, template revision, provider message ID, original status, normalized state, observation source, provider event time when supplied, local receipt time, and a deduplication key. Keep full email addresses, phone numbers, tokens, and verification URLs out of routine logs and metric labels. Retention and access rules should apply to the ledger itself; “append-only” is a state-transition rule, not permission to keep personal data forever.
Make the scheduler produce evidence
The polling worker should not be an unbounded loop attached to a request handler. It should claim unresolved attempts from durable storage, read status with capped concurrency, append the result, and persist the next eligible check. If the process stops after the remote read, another worker may repeat it; the append path therefore needs a deterministic deduplication rule.
HTTP 429 is specific: RFC 6585 defines it as “Too Many Requests” and says the response may include Retry-After. RFC 9110 allows that header to contain either a delay in seconds or an HTTP date. Honor a valid value. When it is absent, exponential backoff with jitter prevents every worker from retrying on the same cadence, while an upper bound prevents reconciliation from going dormant.
Stop conditions matter more than clever backoff. Stop after a terminal transport observation, a permanent request error, or expiration of the verification challenge. Never let a reconciliation job outlive the action it is meant to recover.
The Go example below makes each scheduling decision explicit enough to record as an event. It uses math/rand/v2, which is available beginning with Go 1.22; projects on an older Go release should inject a compatible random source rather than silently changing the algorithm.
package reconcile
import (
"math/rand/v2"
"time"
)
type Decision struct {
RetryAt time.Time
Reason string
Retry bool
}
func Schedule(attempt int, retryAfter time.Duration, now, expiresAt time.Time) Decision {
if !now.Before(expiresAt) {
return Decision{Reason: "verification_expired"}
}
var delay time.Duration
var reason string
if retryAfter > 0 {
delay = retryAfter
reason = "retry_after"
} else {
const base = time.Second
const ceiling = 2 * time.Minute
exponent := attempt
if exponent < 0 {
exponent = 0
}
if exponent > 7 {
exponent = 7
}
window := base * time.Duration(1<<exponent)
if window > ceiling {
window = ceiling
}
delay = time.Duration(rand.Int64N(int64(window) + 1))
reason = "jittered_backoff"
}
retryAt := now.Add(delay)
if !retryAt.Before(expiresAt) {
return Decision{Reason: "retry_exceeds_expiry"}
}
return Decision{RetryAt: retryAt, Reason: reason, Retry: true}
}
Parse the HTTP-date form at the HTTP boundary, where the response headers and current clock are available, then pass the resulting duration into the scheduler. Persist attempt, RetryAt, and Reason in the same durable workflow as the unresolved record. That gives the postmortem a concrete answer to “why did we wait?” rather than an inferred answer based on a worker's current configuration.
Short code; long consequences.
Test the timeline, then decide what pages
Happy-path delivery proves little. Test sequences that damage the evidence model: the same callback twice; an older polling result after a newer callback; a valid Retry-After delay; a 429 without that header; a worker crash after reading remote state but before its local commit; and an observation received after the verification link expires. The assertion is not merely the final label. Assert the retained events, their ordering metadata, the projection, and the next scheduling decision.
Deployment should begin with constrained worker concurrency and a bounded reconciliation cohort. Watch the age of the oldest unresolved attempt, verification completion latency, attempts grouped by outcome, 429 frequency, callback authentication failures, and projection rejections caused by stale transitions. Correlation identifiers belong in traces and controlled logs, not in low-cardinality aggregate labels.
A synthetic signup can exercise dispatch through token consumption if its destination is controlled and its records are excluded from compliance reporting. Sampling complete traces is also useful because dashboards flatten timing and causality; distrust a healthy aggregate until a few individual timelines can be reconstructed from durable records.
Page on the user-facing objective when enough traffic exists to evaluate it. Separately page on a reconciliation backlog that threatens the evidence deadline, since verification might still succeed while the audit trail quietly stops updating. Do not page on every retry. That only trains the person carrying the pager to ignore the mechanism before the mechanism becomes consequential.
Where the simpler choice is correct
Use a mutable snapshot when the workflow is low consequence, no historical transport evidence is required, disputes do not depend on intermediate states, and rebuilding the current view from events would add operational weight without reducing risk. Even then, guard updates against stale observations and retain the application's verification result separately.
Use the ledger for the healthtech signup case. It creates a reviewable chain from challenge creation to transport observations to token consumption, regardless of whether a callback or a polling worker collected a particular status. The choice is deliberately about compliance evidence, not about declaring one transport or collection method universally superior.
Sources and References
- RFC 7489, Domain-based Message Authentication, Reporting, and Conformance: https://datatracker.ietf.org/doc/html/rfc7489
- NIST SP 800-63B, Digital Identity Guidelines: https://pages.nist.gov/800-63-3/sp800-63b.html
- RFC 6585, Additional HTTP Status Codes: https://www.rfc-editor.org/rfc/rfc6585
- RFC 9110, HTTP Semantics: https://www.rfc-editor.org/rfc/rfc9110
- Go 1.22 release notes (
math/rand/v2): https://go.dev/doc/go1.22
Top comments (0)