TL;DR: Keep the password-reset template and its security-critical fields in the authentication service, then let channel adapters render constrained SMS and email variants. SMS can be the first delivery attempt and email the fallback, but neither channel should decide the token, expiry, eligibility, or recovery policy. Poll delivery events into an internal ledger; never treat a provider's accepted status as proof that a person received the message.
For a logistics operation, a password reset often lands at the worst possible moment: a dispatcher is locked out while assigning a time-sensitive load. The tempting beginner design is one template in each provider dashboard, a callback route, and a retry loop. The operational problem is ownership. Once security wording, expiry text, and routing rules live in several consoles, a routine copy edit can become an authentication-policy change with no code review.
I've been paged by missed jobs and duplicate deliveries. That history leaves me with one reflex here: retries must be idempotent, and delivery must be evidence rather than authority. A reset request gets one challenge record and one stable message intent. Every attempt hangs off those records.
How should a beginner OTP login architecture own SMS templates?
The application team should own the canonical message contract in version-controlled code. That contract contains purpose, locale, expiry time, support wording, and the variables each renderer may consume. The SMS and email adapters own channel mechanics: encoding, subject lines, HTML escaping, and transport-specific identifiers. Operations owns alert thresholds and the replay procedure. Security owns the reset policy.
This boundary is deliberate. Provider-hosted templates can make copy changes convenient, but they split the review trail from the code that creates the challenge. Fully rendering arbitrary strings in the authentication handler creates the opposite problem: transport details leak into policy code. A small typed view model between the two keeps the split visible.
For example, the application can produce a channel-neutral intent and require adapters to return their external receipt without granting that receipt any security meaning:
package reset
import (
"context"
"fmt"
"time"
)
type Intent struct {
MessageID string
UserID string
ResetURL string
ExpiresAt time.Time
Locale string
}
type Receipt struct {
ProviderMessageID string
AcceptedAt time.Time
}
type Sender interface {
Send(context.Context, Intent) (Receipt, error)
}
func SMSCopy(i Intent, now time.Time) (string, error) {
remaining := i.ExpiresAt.Sub(now)
if remaining <= 0 {
return "", fmt.Errorf("reset intent %s expired", i.MessageID)
}
minutes := int((remaining + time.Minute - 1) / time.Minute)
return fmt.Sprintf("Reset dispatch access: %s Expires in %d min. Ignore if unrequested.", i.ResetURL, minutes), nil
}
The useful log fields are the internal message ID, template version, channel, region, attempt number, transition, and a redacted destination key. Do not log the reset URL.
Short SMS copy deserves its own test. A single SMS segment allows 160 GSM-7 characters or 70 UCS-2 characters; concatenated messages have lower per-segment limits. A translated character or pasted punctuation mark may therefore change transport behavior even when the sentence still looks short. Test rendered fixtures for the locales you ship, and record segment count as telemetry rather than promising that every locale fits the same envelope.
Count the rendered encoding, not the words.
One challenge, several delivery attempts
Model the reset challenge separately from delivery. The challenge has a digest of the secret, a subject, an expiry, and a consumed timestamp. The message intent references that challenge and carries a unique ID. Attempts then record SMS or email, provider receipt, timestamps, and terminal status.
Readers often describe this as a cheap Node.js 2FA stack with primary SMS, but the language runtime does not settle the ownership question. The same boundary applies when login uses an OTP and when the concrete logistics flow is password recovery: the authentication service decides what may be redeemed, while channel code decides how approved fields are rendered. The Go sample emphasizes the interface because all code in this article is Go; a Node.js worker should preserve the same state machine and idempotency rules rather than treating a framework retry as a delivery guarantee.
This separation prevents a common mistake: generating a fresh secret when SMS is delayed and email fallback begins. Two live secrets make the user's next action ambiguous and complicate incident response. Reusing one challenge across its approved delivery attempts gives the verifier one atomic rule: accept an unexpired, unconsumed challenge once, then mark it consumed in the same transaction.
The sender can use an outbox row committed with the challenge. A worker claims that row, sends it, and stores the receipt. If the worker stops between the network call and the database update, the state is uncertain. Retrying with the stable message ID as an idempotency key is the clean path when a transport honors such keys. When it does not, the ledger must admit the ambiguity; suppressing every retry risks a missed reset, while retrying may produce a duplicate. Pick the behavior explicitly and put it in the runbook.
Do not guess quietly.
Email fallback should be triggered by a policy-owned timer or a terminal delivery event, not by a request handler waiting on the SMS network call. Keep the fallback copy honest: it is another route to the same expiring challenge, not a new challenge and not evidence that SMS failed to reach the handset.
Polling is reconciliation, not authentication
A small team can begin with polling instead of exposing a public event receiver. Store the provider receipt, poll with bounded concurrency, add jitter to backoff, and stop after either a terminal state or a retention deadline. The poller updates only the delivery attempt. It cannot extend expiry, mint another secret, mark the user authenticated, or consume the challenge.
That rule matters because delivery vocabulary is transport-specific and incomplete. Normalize external states into a compact internal set such as pending, delivered, failed, and unknown, while retaining the original value for diagnosis. unknown is real operational state. It should age into a reconciliation queue, not be rewritten as success to make a dashboard green.
Alert on outcomes the team can act on: old pending attempts, a sustained change in terminal failures by channel and region, poller lag, and duplicate-attempt counts. Avoid paging on a single reset. The runbook should let an operator find every attempt for one message ID without exposing the token or full address.
For US and EU traffic, keep the architecture region-aware without pretending geography is merely a provider option. The routing table, permitted data fields, retention, sender identity, and escalation owner should be configuration reviewed by the teams responsible for those regions. The message contract stays common; localized copy and routing policy can differ.
Test the ownership boundary before launch
Start with tests that would have prevented an ugly postmortem. Freeze time immediately before, at, and after expiry. Race two redemption requests and require exactly one successful consume. Replay the same outbox item. Return an accepted receipt and then an unknown poll result. Make SMS stall until the fallback timer fires, then let both channels report delivery. The user may receive two messages, but both must point to the same challenge and neither may extend it.
Template tests need the same seriousness. Review rendered fixtures, not just source strings. Assert that required security wording and the calculated expiry are present in every locale. Check links against an allowlisted application origin. Measure SMS encoding and segment count. Render the email as text and HTML, and ensure authentication headers are configured for the sending domain; Google's sender guidance documents SPF, DKIM, and DMARC expectations along with other delivery requirements.
Deployment should separate policy changes from prose changes even when they share a repository. Give every template a version, include that version in the message ledger, canary a new version, and retain enough metadata to answer which copy a user was sent. Rollback then means selecting the previous reviewed template, not editing a live dashboard while an incident is open.
The main limitation is operational weight. This architecture needs durable state, a worker, reconciliation, expiry tests, and someone who owns the alerts; that is a real trade-off, not free reliability. It is warranted when resets cross channels, workers, or regions and when the team needs an audit trail. It is not a fit for an internal tool with one email transport, low consequence, and no fallback, where a transactional outbox plus one versioned email template may be enough. A tiny team without on-call coverage should also avoid provider-status polling until it can define what action each terminal state triggers.
Complexity spends attention.
The durable decision is about authority: authentication code owns the challenge, the repository owns reviewed message meaning, adapters own rendering constraints, and the ledger owns delivery evidence. With those lines intact, transports can change without silently changing reset policy.
Top comments (0)