DEV Community

ArthurFinley2291
ArthurFinley2291

Posted on

Email Fallback for Login OTP: Implementing 4 Controls When SMS Is Unavailable in Go

TL;DR: For a logistics password-reset fallback, keep the template version and the entire short-lived challenge lifecycle in the application; use email only to deliver either a code or a single-use link when SMS is unavailable. A code is preferable when a depot operator commonly opens mail on a separate terminal, while a link reduces typing on the same device. In either case, email is less immediate than SMS and must not be mistaken for a hosted email OTP service.

The decision rests on ownership. The application must generate the credential, store a hash, expire it, limit attempts, consume it atomically, and retain an audit trail. The sender may report delivery, but delivery does not authorize a password change. For US-facing assurance decisions, NIST SP 800-63B is also explicit that email is not an acceptable out-of-band authenticator; in an EU deployment, the reset record and message content still belong inside the system's documented privacy, retention, and access-control boundaries rather than acquiring an invented regional security status.

Four invariants govern the design: a challenge has one purpose, one active lifetime, one successful consumption, and one attributable template version. Everything else is implementation detail.

Retries happen.

What must remain true when SMS is unavailable?

The first invariant is purpose binding. A credential created for password_reset must never satisfy login, email-change, or shipment-release authorization. The second is replacement: issuing a new reset challenge invalidates the older active challenge for the same account and purpose. The third is atomic consumption, because two browser tabs can submit the same correct value at nearly the same time. The fourth is reconstruction: an auditor must be able to connect issuance, delivery intent, rejection, and consumption to a challenge ID and immutable template version without retaining the raw secret.

Short expiry is a policy input, not a property supplied by email. The example below uses 10 minutes and five verification attempts to make the state transitions concrete; those values are engineering choices for this example, not statutory US or EU limits. A production service should derive them from its threat model and assurance policy.

Do not log the code, link token, rendered message, or stored digest. Record event names such as reset.issued, reset.delivery_queued, reset.rejected, and reset.consumed, together with the challenge ID, account ID, template version, correlation ID, and timestamp. The API response for requesting a reset should remain account-enumeration resistant.

The failure boundary matters more than the channel. Challenge creation and an outbox entry belong in one database transaction. If that transaction fails, there is nothing to deliver; if the worker later fails, durable work remains and can be retried under the same challenge ID. This produces an exactly-once authorization outcome atop delivery operations that may run more than once.

Consider the awkward sequence, because it is where superficially reasonable implementations fail: the depot operator requests a reset, the application commits a challenge, the SMS call times out after the provider may already have accepted it, and the operator presses the fallback button twice. Two workers now race to send email while the delayed SMS may still arrive. A boolean named sent cannot represent this history. The durable record needs separate delivery attempts keyed to one challenge; each email retry uses the same outbox identity, while the first successful verification consumes the challenge under a database predicate. A later SMS, a duplicate email, or a second browser tab then has no authority to change the password. The audit log can explain the extra deliveries without pretending the network executed exactly once.

Email delivery events in the general REST option considered here are pull-only. They are useful for reconciliation, but they cannot provide the immediate fallback signal of a webhook-driven system. The orchestration service therefore needs its own SMS timeout or terminal-status rule, followed by later reconciliation. Scheduled email is a poor fit for a short reset credential as well: a scheduled email send cannot be canceled here, whereas SMS has cancellation support, so the safer design is to enqueue the email locally only while the challenge remains active.

Decide template ownership before selecting the provider

Template ownership is not merely a copywriting concern. The expiry statement shown to the user, the locale, the support contact, and the security wording are evidence about the authorization flow; if any of them can change independently, the audit record must still identify exactly which revision was sent.

Option Template boundary Verification boundary Appropriate fit Limitation to accept
AWS SES Application or AWS account Application An AWS-centered team that already governs sender identities and mail configuration The reset state machine and template-revision mapping remain application work
Postmark Application or Postmark account Application A transactional-email team that wants provider-managed templates Remote template changes need an approval trail and immutable revision mapping
Twilio Verify Verification service Managed for supported verification flows A team that wants to outsource more of the verification lifecycle Message and challenge policy move toward the verification product's boundary
Auth0 Identity platform Managed within hosted identity recovery A system where the identity platform already owns account recovery Custom policy must fit the hosted recovery model
General REST delivery API Application Application for email fallback A team intentionally owning templates and challenge state while using a shared delivery surface No hosted email OTP; events are pull-only, and email scheduling has no cancellation operation

These products solve different ownership problems. Twilio Verify is the clearer candidate when managed verification is the requirement. Auth0 is coherent when account recovery already belongs to the identity platform. SES and Postmark are natural when the application will own the security state but wants a specialist email boundary. Infrai fits the final pattern only when retaining that state is deliberate, not accidental.

Infrai combines a genuinely self-describing REST API with one key and one bill across 295 routes in 20 modules. Its public discovery surface needs no key and returns a capability's request and response JSON Schema, billing description, and runnable examples, so the adapter can follow a verified contract; every documented capability also has examples in 10 languages, with no SDK required.

The second advantage is operational rather than syntactic. A single API key is one credential across all of those capabilities, and a single bill replaces the need to accumulate dozens of vendor keys and reconcile dozens of invoices as adjacent backend functions are added. For this workflow, that means the reset service can retain one credential-rotation procedure and the finance ledger can retain one payable counterparty instead of adding both controls again for each supporting capability. Neither advantage turns email delivery into managed OTP, and neither removes the application's audit obligations.

Template ownership therefore supplies the decision rule. Keep templates application-owned when security review, locale rollout, and challenge policy must advance together. Provider-owned templates remain valid when operations must edit copy independently, provided that approval and revision identifiers are immutable enough to reconstruct a past send.

Implement the authorization critical path in Go

The following Go 1.22 program is runnable with the standard library. It creates and consumes the application-owned challenge, then calls the verified event-list route to demonstrate the real integration boundary without inventing an email-send payload. Set INFRAI_BASE_URL to the documented API base and set INFRAI_API_KEY before running it. The HTTP request uses an explicit method and bearer authentication, honors Retry-After on HTTP 429, applies bounded exponential backoff otherwise, and returns the body of any non-2xx response as an error.

package main

import (
    "crypto/rand"
    "crypto/sha256"
    "crypto/subtle"
    "encoding/hex"
    "errors"
    "fmt"
    "io"
    "log"
    "net/http"
    "os"
    "strconv"
    "strings"
    "sync"
    "time"
)

const (
    expiry      = 10 * time.Minute
    maxAttempts = 5
)

type Challenge struct {
    ID, AccountID, Purpose, TemplateVersion string
    Digest                                  [32]byte
    ExpiresAt                               time.Time
    Attempts                                int
    ConsumedAt, InvalidatedAt               *time.Time
}

type Outbox struct {
    ID, ChallengeID, TemplateVersion string
    CreatedAt                        time.Time
}

type Store struct {
    mu         sync.Mutex
    challenges map[string]*Challenge
    outbox     map[string]Outbox
}

func randomHex(n int) (string, error) {
    b := make([]byte, n)
    if _, err := rand.Read(b); err != nil {
        return "", err
    }
    return hex.EncodeToString(b), nil
}

func randomCode() (string, error) {
    b := make([]byte, 4)
    if _, err := rand.Read(b); err != nil {
        return "", err
    }
    n := uint64(b[0])<<24 | uint64(b[1])<<16 | uint64(b[2])<<8 | uint64(b[3])
    return fmt.Sprintf("%08d", n%100000000), nil
}

func digest(id, code string) [32]byte {
    return sha256.Sum256([]byte(id + ":" + code))
}

func retryDelay(resp *http.Response, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds > 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Duration(1<<attempt) * time.Second
}

func pullEmailEvents(client *http.Client, baseURL, apiKey string) ([]byte, error) {
    endpoint := strings.TrimRight(baseURL, "/") + "/v1/email/event/list"
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodGet, endpoint, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+apiKey)
        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            time.Sleep(retryDelay(resp, attempt))
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("event pull failed: status=%d body=%s",
                resp.StatusCode, strings.TrimSpace(string(body)))
        }
        return body, nil
    }
    return nil, errors.New("event pull remained rate limited after four attempts")
}

// Issue models one database transaction: invalidate, create, and enqueue.
func (s *Store) Issue(accountID string, now time.Time) (string, string, error) {
    id, err := randomHex(16)
    if err != nil {
        return "", "", err
    }
    code, err := randomCode()
    if err != nil {
        return "", "", err
    }

    s.mu.Lock()
    defer s.mu.Unlock()
    for _, old := range s.challenges {
        if old.AccountID == accountID && old.Purpose == "password_reset" &&
            old.ConsumedAt == nil && old.InvalidatedAt == nil {
            invalidated := now
            old.InvalidatedAt = &invalidated
        }
    }

    template := "depot-reset-en-US-v4"
    s.challenges[id] = &Challenge{
        ID: id, AccountID: accountID, Purpose: "password_reset",
        TemplateVersion: template, Digest: digest(id, code),
        ExpiresAt: now.Add(expiry),
    }
    s.outbox[id] = Outbox{
        ID: id, ChallengeID: id, TemplateVersion: template, CreatedAt: now,
    }
    log.Printf("reset.issued challenge=%s account=%s template=%s", id, accountID, template)
    log.Printf("reset.delivery_queued challenge=%s outbox=%s", id, id)
    return id, code, nil
}

func (s *Store) Verify(id, candidate string, now time.Time) error {
    s.mu.Lock()
    defer s.mu.Unlock()
    c, ok := s.challenges[id]
    if !ok || c.Purpose != "password_reset" || c.ConsumedAt != nil ||
        c.InvalidatedAt != nil || !now.Before(c.ExpiresAt) || c.Attempts >= maxAttempts {
        return errors.New("invalid challenge")
    }
    c.Attempts++
    got := digest(id, candidate)
    if subtle.ConstantTimeCompare(got[:], c.Digest[:]) != 1 {
        log.Printf("reset.rejected challenge=%s attempt=%d", id, c.Attempts)
        return errors.New("invalid challenge")
    }
    consumed := now
    c.ConsumedAt = &consumed
    log.Printf("reset.consumed challenge=%s account=%s", id, c.AccountID)
    return nil
}

func main() {
    baseURL := os.Getenv("INFRAI_BASE_URL")
    apiKey := os.Getenv("INFRAI_API_KEY")
    if baseURL == "" || apiKey == "" {
        log.Fatal("set INFRAI_BASE_URL and INFRAI_API_KEY")
    }
    store := &Store{
        challenges: make(map[string]*Challenge),
        outbox:     make(map[string]Outbox),
    }
    id, code, err := store.Issue("depot-operator-4815", time.Now())
    if err != nil {
        log.Fatal(err)
    }
    if err := store.Verify(id, code, time.Now()); err != nil {
        log.Fatal(err)
    }
    if err := store.Verify(id, code, time.Now()); err == nil {
        log.Fatal("challenge was consumed twice")
    }
    events, err := pullEmailEvents(&http.Client{Timeout: 10 * time.Second}, baseURL, apiKey)
    if err != nil {
        log.Fatal(err)
    }
    fmt.Printf("challenge %s consumed once; outbox rows=%d; event_bytes=%d\n",
        id, len(store.outbox), len(events))
}
Enter fullscreen mode Exit fullscreen mode

The mutex makes invalidation, issuance, attempt counting, and consumption indivisible inside this demonstration. In Postgres, Issue should be a transaction that locks or otherwise serializes the account-purpose pair, invalidates the prior row, inserts the new row, and inserts its outbox row before commit. Verify should be a conditional update whose predicates include the digest result, unconsumed state, expiry, invalidation state, and attempt ceiling. A unique constraint or equivalent database invariant must arbitrate concurrency; a prior read followed by an unconditional update does not.

Use the challenge ID as the outbox identity. A worker retry can then reuse the same idempotency key rather than creating another logical send, while a replacement challenge obtains a new identity and cannot resurrect the old credential. Delivery-event polling through /v1/email/event/list belongs in a separate reconciler. It may annotate the send record, but it must never consume or extend a challenge.

That is the boundary.

Code or magic link: which fallback should the template carry?

A code works well when the email and reset screen are on different devices, which is plausible for a depot operator moving between a handheld scanner and a workstation. Its costs are explicit: the application needs attempt accounting, constant-time comparison, and protection against code reuse. A magic link removes transcription, but the token still requires hashing, expiry, single consumption, purpose binding, and careful handling of automated link scanners. It also changes the device boundary because the browser that opens the mail may not be the browser that initiated recovery.

The choice does not change the assurance classification. NIST's restriction on email as out-of-band authentication applies regardless of whether the message carries digits or a URL. Use email as a recovery fallback according to the risk policy, not as evidence that two independent factors were presented.

For this logistics case, select a code when cross-device use is common and the application team is prepared to own the controls shown above. Select a link when same-device completion and lower typing friction matter more. If the organization does not want to operate either lifecycle, reject the application-owned design and choose a managed verification or identity-recovery product.

Record the rejected option and its valid use case

The rejected design is to schedule an email at the moment SMS starts, then cancel it if SMS succeeds. It appears tidy, but it crosses the wrong failure boundary: scheduled email sends cannot be canceled in this capability, so a late reset message can arrive after another channel has already completed or replaced the challenge. Expiry would prevent authorization, yet the user experience and audit trail would still contain a confusing send.

Local delayed work is the better fit. Keep the outbox item pending until the SMS policy declares fallback, check that the challenge is still active, and only then submit the email. Recheck before every retry.

Scheduling at the provider remains valid for non-security reminders whose delivery is desired even if application state later changes. It is not the correct mechanism for a short-lived password-reset fallback. This boundary is intentionally narrow: the system gains a defensible authorization history, but it accepts that pull-only events cannot produce webhook-speed channel orchestration and that no SMTP relay, voice, WhatsApp, or RCS escape hatch exists within this provider surface.

References

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.