TL;DR: For an HR SaaS 2FA login or password-reset flow near a shift boundary, choose SMS OTP or email OTP at runtime according to the employee's verified, reachable destination. Allow one controlled fallback, and keep code validity independent from message delivery. The operational target is a completed login or reset inside the expiry window; a provider acceptance event, an email open, or an SMS handoff is only an intermediate signal.
Integration effort should be judged against that whole path. A channel that takes one afternoon to wire up but gives the on-call engineer no trustworthy state transitions is unfinished integration work.
Read the shift-boundary failure signal first
The user story sounds small: deliver a reset code with a short expiry so an employee can see shift reminders. The system behind it has at least four separately observable stages: create one challenge, enqueue one message attempt, record the channel outcome, and accept the code once. Keep those identities separate. Retries then belong to the delivery attempt, while expiry, attempt limits, and one-time use belong to the challenge.
This distinction matters under deadline pressure. If an SMS attempt stalls, issuing a fresh code by email leaves two valid secrets in circulation and makes the audit trail harder to explain. A safer fallback reuses the same challenge while rendering the message for a different verified destination. The displayed code may stay the same; the delivery-attempt identifier must not.
Use a capacity-planning reflex before choosing defaults. Estimate peak reset starts per minute around shift changes, multiply by the maximum permitted delivery attempts, and provision the queue for that upper bound rather than the daily average. A policy allowing one initial send and one fallback has a worst-case delivery fan-out of 2 per challenge. That is a planning bound, not a prediction.
Deadlines expose ambiguity.
Should an HR SaaS login use SMS OTP or email OTP?
Email and SMS expose different evidence, but neither can prove that the intended employee read the code. For email, domain authentication is part of deliverability practice: DMARC defines a policy and reporting mechanism built on SPF and DKIM alignment. It does not turn an accepted or opened message into proof of possession. Apple Mail Privacy Protection can privately download remote content in the background, so an image load is particularly weak evidence that a person opened a reset message. SMS has its own gap between submission and human receipt. Model queued, submitted, delivered, temporarily failed, and permanently failed as delivery states only; do not let any of them consume or validate the challenge. Likewise, avoid using a fixed timer such as “fallback after 10 seconds” as if every carrier and region behaved alike. Set the timer from the reset SLO and observed latency distribution for the destination cohort, with enough time left for the employee to type the code. The security comparison is therefore contextual. SMS reaches a phone number; email reaches a mailbox session. Either destination may be shared, forwarded, inaccessible on a work floor, or controlled by someone other than the employee. For this password-reset path, the correct question is narrower: which already-verified destination is reachable now, and what independent controls limit damage if it is not? Rate limits, a short expiry, one-time consumption, attempt limits, and a visible account-recovery event carry more weight than arguing that one transport is universally secure. Regional labels such as US or EU do not answer that question by themselves; measure deliverability and latency for the actual workforce cohorts, then keep the security controls consistent across both channels.
Unknown is a valid state.
Implement one challenge across both delivery attempts
Here is the core boundary I would insist on during design review. The example chooses a five-minute lifetime and six verification attempts as explicit local policy values, not universal best practices. Change them through configuration and review them against the reset SLO and support process.
package reset
import (
"crypto/rand"
"crypto/sha256"
"crypto/subtle"
"encoding/binary"
"errors"
"fmt"
"time"
)
type Challenge struct {
ID string
CodeHash [32]byte
ExpiresAt time.Time
Attempts int
UsedAt *time.Time
}
func NewChallenge(id string, now time.Time) (Challenge, string, error) {
var raw [8]byte
if _, err := rand.Read(raw[:]); err != nil {
return Challenge{}, "", err
}
code := fmt.Sprintf("%06d", binary.BigEndian.Uint64(raw[:])%1000000)
return Challenge{
ID: id,
CodeHash: sha256.Sum256([]byte(code)),
ExpiresAt: now.Add(5 * time.Minute),
}, code, nil
}
func (c *Challenge) Verify(code string, now time.Time) error {
if c.UsedAt != nil || !now.Before(c.ExpiresAt) || c.Attempts >= 6 {
return errors.New("challenge unavailable")
}
c.Attempts++
candidate := sha256.Sum256([]byte(code))
if subtle.ConstantTimeCompare(candidate[:], c.CodeHash[:]) != 1 {
return errors.New("challenge unavailable")
}
used := now
c.UsedAt = &used
return nil
}
Persist the attempt increment and one-time consumption atomically. Do not log the plaintext code, and do not put it in metrics labels. The queue payload can carry the challenge ID, delivery-attempt ID, channel, and a reference to a verified destination; the renderer should receive the code through a narrowly scoped path and discard it after sending.
Fallback needs an idempotency boundary. A worker may retry the same delivery-attempt ID after a temporary failure, but the policy engine creates a new attempt ID when it changes channel. This gives an operator a coherent timeline without coupling the authentication decision to any provider-specific callback vocabulary.
Budget the integration ownership before choosing a boundary
A buy-versus-build decision should include on-call work. The smallest useful comparison is not “one API versus two APIs”; it is the ownership boundary around routing, evidence normalization, and failure recovery.
| Boundary | Managed delivery integration | Self-operated delivery components |
|---|---|---|
| Channel adapters | Less adapter code; normalize external states locally | More protocol and carrier/mail-system work |
| Routing policy | Keep policy in your service to limit lock-in | Full control, plus full testing burden |
| Queue and retries | Verify idempotency and retention behavior | Own capacity, backoff, and dead-letter handling |
| Observability | Export events into your SLO model | Build and operate the event pipeline |
| Exit cost | Preserve generic challenge and attempt records | Preserve runbooks and specialist knowledge |
For a small platform team, I would build the challenge state machine and policy layer because they encode product risk, then treat delivery adapters as replaceable infrastructure. That is a trade-off, not a recommendation for a product category. Self-operation may reduce an external dependency, but it also places mail reputation, telecom behavior, queue capacity, and incident response on the same on-call rotation. Price is rarely the deciding line item once that load is counted.
Keep the adapter contract deliberately boring: submit, observe status, and classify retryability. If a channel cannot provide a final signal, represent “unknown” rather than manufacturing “delivered.” Skepticism is useful here.
Canary the policy and preserve in-flight challenges on rollback
Test with clocks and state transitions, not real inbox optimism. Before deployment, cover expiry at the exact boundary, duplicate queue delivery, concurrent correct submissions, wrong-code exhaustion, late status events, and fallback racing with an initial success. Inject temporary and permanent adapter errors. Confirm that every path leaves one challenge record and an ordered set of attempts.
The primary service-level indicator should measure successful reset completion before challenge expiry, segmented by initial channel, fallback path, destination region, and client surface. Supporting indicators include queue age, time to first submission, classified permanent failures, unknown outcomes, fallback rate, and verification exhaustion. Do not use email opens as the success SLI; privacy features make that metric unsuitable, and it never represented reset completion anyway.
Deploy policy changes behind configuration with a narrow cohort first. Roll back by stopping new fallback decisions or restoring the earlier routing policy, not by deleting challenges already issued. In-flight attempts should finish under the policy version recorded with them, while verification continues to enforce the original expiry and one-time-use rules.
The decision rule stays plain: choose the verified channel most likely to be reachable during the employee's shift, preserve one challenge across a tightly bounded fallback, and judge the integration by completed resets within the SLO. Everything else is supporting evidence.
References
- RFC 7489, Domain-based Message Authentication, Reporting, and Conformance (DMARC): https://datatracker.ietf.org/doc/html/rfc7489
- Apple, Use Mail Privacy Protection on iPhone: https://support.apple.com/guide/iphone/use-mail-privacy-protection-iphf084865c7/ios
Top comments (0)