TL;DR: A post-import compliance notice is a ledger event before it is an email. Give every player and message template a deterministic key, persist an outbox row in the same transaction as the import, and let a Node.js worker claim small batches under an explicit provider limit. Retry only transient failures, record every attempt, and make a replay produce the same decision without sending a duplicate.
That framing matters in gaming. A regulator, a support agent, or a player may ask which notice was generated, when it was attempted, and why it was or was not delivered. A loop that calls an email API for each imported row can send a plausible message, yet still fail that audit question after a timeout.
Can a bulk welcome email after user import remain auditable?
Start with the evidence, not the transport. For each imported account, the system should be able to answer five questions: which legal or policy version selected the notice, which destination was used, which content digest was rendered, how many delivery attempts occurred, and what terminal state was reached. Those facts belong to your database; a provider dashboard is only an observation of one hop.
I use an append-only delivery record because mutable status fields erase useful history. The current state can still be materialized for queries, but each transition carries an event time, an attempt number, and a reason code. Keep recipient addresses protected; an audit trail does not require making personal data broadly readable.
A compact schema can look like this:
| Record | Required fields | Why it exists |
|---|---|---|
import_account |
account_id, import_batch, consent_state
|
Establishes the source row and eligibility decision |
notice_outbox |
notice_id, dedupe_key, channel, template_digest, state
|
Durable work item created with the account |
delivery_attempt |
notice_id, attempt_no, started_at, result_code
|
Append-only evidence for each try |
delivery_receipt |
notice_id, provider_id, accepted_at
|
Correlates an accepted handoff without claiming final inbox delivery |
The distinction between accepted and delivered is important. SMTP acceptance is not proof that a person read a message, and an SMS submission is not proof that a handset displayed it. Name those states precisely so an operations report does not overstate compliance.
The trade-off is operational weight. An outbox, lease table, and append-only attempts take more storage and migration work than a direct API call. That cost is justified for regulated notices; it is a poor fit for a one-off internal announcement where replay evidence has no business value. A team with no durable database should delay the send until it has one, because a queue alone cannot prove which policy version was rendered.
Model the import as a ledger, not a loop
The import transaction creates the account and its notification intent together. A unique constraint on (account_id, notice_kind, policy_version) is the dedupe gate. If the import job is replayed, the insert becomes a no-op and no second notice is born. This is the exactly-once mindset applied where it is achievable: the intent is exactly once; external delivery remains at-least-once and must be reconciled.
Here is the state transition core in Go. It is deliberately transport-neutral; the same states can be driven by a Node.js queue consumer.
type NoticeState string
const (
Pending NoticeState = "pending"
InFlight NoticeState = "in_flight"
Accepted NoticeState = "accepted"
Retryable NoticeState = "retryable"
Failed NoticeState = "failed"
Suppressed NoticeState = "suppressed"
)
type Notice struct {
ID string
AccountID string
DedupeKey string
Channel string
PolicyVersion string
TemplateDigest string
AttemptCount int
State NoticeState
}
func nextState(code int, attempts, maxAttempts int) NoticeState {
switch {
case code >= 200 && code < 300:
return Accepted
case code == 408 || code == 429 || code >= 500:
if attempts < maxAttempts {
return Retryable
}
return Failed
default:
return Failed
}
}
The 429 branch is not a license to hammer the endpoint. Persist the server's retry hint when one is supplied, cap the delay, and add jitter so a whole import batch does not wake at the same millisecond. Permanent failures, such as an invalid destination, move to failed and require an explicit correction-and-replay operation.
How should a Node.js worker pace batches and retries?
Treat rate limits as a scheduling input. Suppose a channel contract permits 10 submissions per second and your import contains 2,400 eligible accounts. A batch size of 20 with a 2-second interval is easy to reason about, while a concurrency setting of 200 is not: it creates a burst, obscures which attempt consumed the quota, and makes a timeout ambiguous. The numbers here are configuration examples, not a claim about any provider's limit.
A worker claims rows with a lease, sends at most the claimed batch, and writes an attempt before acknowledging the queue message. The lease timeout must exceed the request deadline, otherwise a slow response can cause two workers to send the same notice. An idempotency key should travel with the request when the transport supports it; still keep your own dedupe key because a remote system may forget keys after its retention window.
type Sender interface {
Send(ctx context.Context, n Notice, idempotencyKey string) (int, string, error)
}
func backoff(attempt int, base, cap time.Duration, jitter time.Duration) time.Duration {
d := base << (attempt - 1)
if d > cap {
d = cap
}
return d + time.Duration(rand.Int63n(int64(jitter)))
}
func deliver(ctx context.Context, s Sender, n Notice, maxAttempts int) NoticeState {
for attempt := n.AttemptCount + 1; attempt <= maxAttempts; attempt++ {
code, receipt, err := s.Send(ctx, n, n.DedupeKey)
appendAttempt(n.ID, attempt, code, receipt, err)
if err == nil && code >= 200 && code < 300 {
return Accepted
}
state := nextState(code, attempt, maxAttempts)
if state != Retryable {
return state
}
delay := backoff(attempt, 500*time.Millisecond, 30*time.Second, 250*time.Millisecond)
if err := sleepContext(ctx, delay); err != nil {
return Retryable
}
}
return Failed
}
In production, appendAttempt and the final state update must be conditional on the lease owner. If the process dies after the remote accepts the message but before the database commit, the notice is replayed; the idempotency key and provider receipt correlation are what turn that uncertain interval into a reconcilable case. Never infer success from a network timeout.
Metrics should expose both progress and ambiguity: eligible rows, suppressed rows, attempts by result code, lease expirations, acceptance latency, and notices awaiting reconciliation. Alert on a growing retryable set, not merely on HTTP errors. A quiet queue with a rising unknown reconciliation count is a delivery incident too.
Email or SMS: which evidence survives the audit?
Use email for a notice that needs durable explanatory text and a link to the policy version; use SMS only when the player has a verified mobile destination and the message can survive segmentation. GSM-7 and UCS-2 encoding change SMS character capacity and can split one logical notice into multiple billable segments, so store the encoded length and segment count in the attempt record. The transport decision is subordinate to the evidence contract.
For email, authenticate the sending domain, publish SPF and DKIM, and monitor the sender requirements described by Google's guidance. Keep unsubscribe and consent rules separate from the compliance-notice eligibility rule; a legal notice can have a different basis, but that decision must be explicit in consent_state and reviewable.
A useful decision table is small:
| Condition | Action | Recorded reason |
|---|---|---|
| Destination unverified | Suppress | destination_unverified |
| Policy version missing | Hold import row | policy_unresolved |
| Provider says temporary throttle | Retry with hint | rate_limited |
| Provider accepts handoff | Mark accepted, reconcile later | accepted_pending_receipt |
Roll out with replay, not hope
Run a dry import that creates outbox rows but invokes no sender. Sample 100 rows across jurisdictions, compare the rendered digest with the approved template, and have compliance sign the policy-version mapping. Then enable one shard, watch lease expiry and reconciliation metrics for at least one full retry horizon, and expand in fixed increments.
A replay command should accept an import batch and a reason, select only failed, retryable, or unknown rows, and preserve the original notice ID. That gives support a reversible action: correcting an address does not rewrite history; it creates a new attempt linked to the same notice.
The practical rule is terse: persist intent with the import, pace work in the worker, and treat every uncertain handoff as evidence to reconcile. That is how a Node.js batch process becomes an auditable gaming control rather than a fast loop with an attractive success log.
Sources
References:
- Google, Email sender guidelines: https://support.google.com/a/answer/81126
- Twilio, SMS character limits and segmentation (GSM-7/UCS-2): https://www.twilio.com/docs/glossary/what-sms-character-limit
- RFC 5321, Simple Mail Transfer Protocol: https://www.rfc-editor.org/rfc/rfc5321
- RFC 6648, Deprecating the X- prefix in protocol parameters: https://www.rfc-editor.org/rfc/rfc6648
Top comments (0)