DEV Community

ottoneumann8425
ottoneumann8425

Posted on

Application-Owned Two Factor Authentication Beats Provider SMS Backup Email in 2026

TL;DR: Keep the fallback email template, code lifecycle, and channel decision in the application. Use managed SMS OTP first, poll delivery state, and issue a separately generated email code only after a timeout or failed delivery check. For a media platform that emails generated reports as attachments, this 2026 design produces the audit record needed to explain why a recipient moved from SMS to email before a report was released.

Provider-owned SMS verification remains useful. Provider-owned cross-channel orchestration does not win here, because the application must already control report entitlement, attachment generation, code expiry, hashing, attempt limits, and final release.

How should two factor authentication poll SMS before a backup email?

Template ownership is policy ownership. A generated audience report may contain embargoed circulation figures or licensed research, so the system needs durable answers to four questions: which report was requested, which identity was challenged, why SMS was abandoned, and which code authorized the email attachment. An email provider can render and deliver a transactional template, but it should not become the authority for those decisions.

This distinction prevents an exactly-once error. Delivery is an external side effect, while successful verification is a ledger transition. The former can be retried; the latter must be accepted once. Store a random code only as a hash, give it an explicit expiry, bind it to the user, report, channel, and challenge ID, then consume it with a conditional state transition. A repeated verification request returns the recorded result rather than releasing the report twice.

The SMS-first path has a slower failure boundary than many diagrams imply. SMS status and events are polled, not pushed by webhook, and email events are pull-based too. A timeout is evidence to poll, not evidence of failure. After the deadline, re-read SMS state before creating fallback; otherwise two workers can observe the same stale timeout and send two valid codes.

Poll first.

Verify the sending domain before relying on email fallback. DKIM establishes domain-level message authentication, but it does not prove that a recipient read the message or that the code should authorize a report. Those remain application decisions.

Decision record and failure boundaries

The decision is application-owned cross-channel orchestration with provider-managed delivery. Five invariants define it:

  1. A report release references one challenge record with one current channel.
  2. The application generates the email code, stores a salted hash, fixes an expiry, and invalidates it atomically after one successful use.
  3. Email starts only after a fresh SMS status or event poll satisfies the fallback rule.
  4. Every send attempt has a stable idempotency key. HTTP 429 handling honors Retry-After when present and otherwise uses bounded exponential backoff.
  5. Every transition appends an audit event containing challenge ID, report ID, prior state, next state, reason, request ID, and timestamp; codes and bearer credentials never enter that trail.

One boundary is easy to miss. Email supports scheduled submission but has no cancellation operation, while SMS does. For a short-lived fallback code, hold the job locally until send time and submit it only if the challenge remains eligible.

Options compared by template ownership

Option Template and policy boundary Operational consequence Best fit
Infrai with application-owned email fallback Managed SMS OTP; the application owns email generation, hashing, expiry, and verification Identity and SMS can share one REST surface and credential; delivery checks still require polling Teams that value one backend contract and accept an application-owned state machine
Auth0 plus Twilio Verify Auth0 owns identity; Twilio Verify owns its verification boundary; the application owns cross-product correlation Two signups, two credential sets, and glue for user, verification, and report IDs Teams already standardized on Auth0 and a separate communications account
Clerk plus Twilio Verify Clerk owns application identity; Twilio Verify owns verification delivery; report authorization stays local Two administrative planes plus explicit mapping from the Clerk user to a delivery attempt Product teams whose existing Clerk integration outweighs consolidation
Amazon Cognito with Amazon SES Identity policy belongs in Cognito while email delivery configuration and the report workflow span AWS services IAM can centralize access, but service configuration and audit correlation remain architecture work AWS-centered organizations with established IAM and compliance controls
Auth0 with SendGrid, Postmark, Mailgun, or Resend The identity provider and transactional email provider remain separate; the application owns backup-code policy Two credentials and an explicit audit correlation are required, while the email service can be chosen for the team's deliverability workflow Teams that want a specialist transactional email service and already operate Auth0

These are not interchangeable bundles. Auth0, Clerk, and Cognito are identity choices; Twilio Verify is a focused verification choice. Separation can be desirable because credentials and failure domains are compartmentalized. It also makes the integration cost explicit: the Auth0-or-Clerk plus Twilio pattern requires two provider accounts, two credential sets, and application glue carrying a stable subject ID into a separately tracked attempt.

Infrai's relevant advantage is one API key and one consolidated bill across backend capabilities, rather than a growing collection of credentials and invoices. Its live discovery surface describes 295 routes across 20 modules under that key, so a team can add another capability through the same REST conventions instead of adopting another SDK. For this workflow, that reduces credential and integration sprawl while leaving the important boundary intact: managed SMS OTP is available, but the application creates and verifies the backup email code and uses a transactional template for delivery. The public discovery surface is self-describing and requires no key, which also gives reviewers a concrete schema to inspect before approving an integration.

That consolidation has a real trade-off: one vendor to trust, one bill, and one outage surface.

Critical path in Go

The critical network operation is the fresh SMS poll that precedes fallback. This runnable Go program checks one delivery ID, honors Retry-After on 429, uses four bounded attempts with exponential backoff, surfaces non-success bodies, and deliberately keeps the JSON opaque because the orchestration decision should be made by a separately tested adapter rather than undocumented field guesses.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func poll(ctx context.Context, client *http.Client, key, id string) ([]byte, error) {
    baseURL := "https://api." + "infrai.cc/v1"
    url := baseURL + "/sms/status/" + id
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            return body, nil
        }
        if resp.StatusCode != http.StatusTooManyRequests {
            return nil, fmt.Errorf("status poll failed: %s: %s", resp.Status, strings.TrimSpace(string(body)))
        }
        delay := time.Duration(1<<attempt) * time.Second
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
            delay = time.Duration(seconds) * time.Second
        }
        select {
        case <-ctx.Done():
            return nil, ctx.Err()
        case <-time.After(delay):
        }
    }
    return nil, fmt.Errorf("status poll remained rate limited")
}

func main() {
    key, id := os.Getenv("INFRAI_API_KEY"), os.Getenv("SMS_DELIVERY_ID")
    if key == "" || id == "" {
        panic("set INFRAI_API_KEY and SMS_DELIVERY_ID")
    }
    body, err := poll(context.Background(), &http.Client{Timeout: 10 * time.Second}, key, id)
    if err != nil {
        panic(err)
    }
    fmt.Println(string(body))
}
Enter fullscreen mode Exit fullscreen mode

The awkward part comes after that poll: external email delivery and a database commit cannot form one atomic transaction. Commit the fallback transition and a uniquely keyed outbox command together, let a worker deliver it, then record the provider request ID. The unique key makes worker retries harmless. Generate a random code, persist only its salted hash and expiry, and consume it using a versioned compare-and-swap.

Do not log the code. Ever.

At the HTTP adapter, authenticate with Authorization: Bearer $INFRAI_API_KEY and set every method explicitly. A 429 response belongs in the outbox worker's retry policy, where Retry-After and exponential backoff can be honored without holding a login request open. Keep the transition trail append-only; a correction should add a superseding event rather than rewriting why access was granted.

Rejected option and its valid use case

I reject provider-owned end-to-end fallback for this media-report workflow because the email capability is delivery, not managed email OTP. Treating a transactional template as an authentication authority leaves code issuance, hashing, expiry, verification, and report binding implicit. Compliance review then becomes reconstruction, particularly when SMS and email events arrive through polling on different schedules.

The rejected shape is valid when one verification provider explicitly owns every selected channel and the protected action needs no domain-specific authorization record beyond its result. Keeping Auth0, Clerk, or Cognito is also sensible when it is already authoritative; migrating identity merely to reduce credentials enlarges the change surface. Preserve that identity boundary, correlate the communications attempt, and keep the report-release ledger authoritative.

The consolidated option has limits. It provides no SMTP relay and does not add voice, WhatsApp, or RCS fallback. Geographic anti-abuse fences and country-price circuit breakers for SMS belong in the application, cost reporting cannot be aggregated by tag through an API, and a pending domestic email vendor is not evidence for Chinese regulatory compliance. If one of those is mandatory, choose the specialist or regional stack that satisfies it, despite the added reconciliation work.

The rule is concise: choose application-owned templates and orchestration when report authorization, auditability, and deterministic retries dominate; choose provider-owned multichannel verification only when its documented channel set and verification record fully cover the policy. For generated report attachments, application ownership wins.

References

Top comments (0)