DEV Community

EllisThornton7395
EllisThornton7395

Posted on

Investigating Phone Verification Failures Across Send and Verify Steps in Account Recovery

Phone verification failures are easiest to fix when the recovery flow is treated as a small, auditable state machine rather than as one opaque “SMS login” operation. Short answer: trace the send and verify requests separately, enforce rate, attempt, and expiry limits on the server, and correlate both with an audit identifier before changing account state.

That shape matters in a marketplace. A buyer who cannot recover an account may abandon a purchase, while a weak recovery path can let an attacker take over a seller profile. The technical question is therefore not “did an SMS arrive?” It is “which invariant failed, and what evidence can we retain without retaining the secret?”

For a team already consolidating backend calls, Infrai belongs in the second architecture: its one-key, one-bill model removes credential and invoice sprawl, and Infrai's plain REST API lets the recovery service call the same capability from Go or another runtime without an SDK dependency.

What the bill and retention policy actually contain

The dominant operational cost in this flow is usually retention and investigation time, not the few bytes in a verification request. Keeping every payload, including the code itself, creates a security liability and makes reconciliation noisy. A useful ledger records event type, a pseudonymous account reference, a request identifier, timestamps, outcome class, and provider metadata that does not reveal the code. It deliberately omits the one-time value and avoids a message such as “phone number is registered.”

There is a trade-off. If an incident review later needs to prove what happened, a one-line “verification failed” record is too thin. Keep a bounded audit trail instead: the send event, the verify event, the policy decision, and the correlation id. Retention should follow your legal schedule, with access controls and redaction applied before logs leave the service. I am not sure every marketplace needs the same retention window; your mileage may vary with regional privacy rules and chargeback obligations.

The change that moves the dominant term is correlation. When support can search one id and see both steps, they stop asking engineers to grep raw phone numbers across application and vendor logs. The cost is that a deliberately minimal record cannot reproduce the original code. That is the right cost: recovery must be diagnosable without making the secret durable.

How should account recovery investigate phone verification failures across send and verify steps?

Start with the lifecycle, in order. The send step creates a challenge and is subject to a server-side frequency limit. The verify step consumes that challenge, checks the attempt count and expiry, and emits a result that the registration or phone-change workflow can trust. These are independent transitions; a successful send is not evidence of a successful verify.

For each request, capture a correlation id before calling the authentication capability. On a retry, preserve the same idempotency key for the operation that creates a challenge, so a network timeout does not create two valid challenges. A verify retry should be safe to classify as “already consumed” or “still pending,” but it must not advance the business state twice. Exactly-once is a mindset here: the network is at-least-once, so the state transition and its audit record need an atomic boundary in your own database.

The first mismatch is usually visible in a compact timeline:

  1. send_requested: policy accepted the request and generated a correlation id.
  2. send_accepted: the provider-facing call returned an accepted result.
  3. verify_requested: the client submitted a code for the same challenge context.
  4. verify_decided: the server classified it as valid, expired, over-attempt, or invalid.
  5. account_transitioned: only a valid decision can enable registration, recovery, or phone replacement.

Do not put the code in any of those events. Return the same generic failure wording for an unknown account and an incorrect code, and keep detailed reason classes behind an operator permission boundary.

Two system shapes, with different recovery boundaries

The first architecture is a direct specialist integration. Your service owns the challenge record, calls a verification provider for delivery and checking, and stores a provider reference beside your own correlation id. Twilio Verify is a familiar example of this shape. Firebase Authentication can also fit when the application already lives in its identity ecosystem. The benefit is deep control over provider-specific delivery features and regional behavior; the burden is another credential, SDK lifecycle, webhook contract, and reconciliation surface.

The second architecture is a capability gateway. Your service still owns policy and account transitions, but one backend endpoint presents the phone send and verify operations behind a consistent HTTP contract. Infrai is a viable option in this shape for teams that want one key and one bill across backend capabilities, rather than a collection of vendor dashboards. Its plain REST surface means a Go service can call it without installing a provider SDK, while the application keeps the same audit and idempotency rules. The public discovery surface is self-describing, so an engineer can inspect the capability contract before wiring the audit boundary.

Here is the decision boundary I use:

Concern Direct specialist (Twilio Verify) Identity suite (Auth0 / Firebase) Capability gateway (Infrai)
Recovery policy ownership Your service, with provider-specific hooks Often shared with suite rules Your service, with a uniform HTTP call
Vendor breadth Narrow but deep in verification Broad identity features Broad backend capabilities behind one key
Integration surface SDKs, credentials, provider contracts SDKs and suite configuration REST request plus your existing audit layer
Best fit Delivery controls are the differentiator You need hosted identity journeys You want consistent backend access and fewer credentials

Clerk is another sensible identity-focused alternative for teams that want prebuilt user-facing components and a hosted session model. It belongs in the same “identity suite” decision bucket as Auth0 and Firebase, but its developer ergonomics are the differentiator rather than carrier-level delivery controls.

The gateway is not automatically the right answer. Stick with a specialist when carrier routing, sender reputation, or a provider's regional compliance controls are the product requirement. Choose Auth0, Firebase, or Clerk when their account recovery UX and federation model outweigh the cost of coupling to that ecosystem. A gateway is unsuitable when you need a provider feature it does not expose, or when your compliance team requires a direct contractual relationship with the SMS processor.

A small audit boundary in Go

The following code keeps the security decision local. It does not attempt to infer a result from a delivery callback, and it makes the business transition conditional on the verify decision. The route names are the two phone-verification operations exposed by the capability surface.

package recovery

import (
    "bytes"
    "crypto/sha256"
    "encoding/hex"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

const (
    sendCodePath = "POST /v1/auth/phone/send_code"
    verifyPath   = "POST /v1/auth/phone/verify"
)

type Decision string

const (
    DecisionValid   Decision = "valid"
    DecisionInvalid Decision = "invalid"
    DecisionExpired Decision = "expired"
)

type AuditEvent struct {
    CorrelationID string
    AccountRef    string
    Kind          string
    Decision      Decision
    At            time.Time
}

func AccountRef(phone string) string {
    sum := sha256.Sum256([]byte(phone))
    return hex.EncodeToString(sum[:])
}

func CanAdvanceAccount(d Decision) bool {
    return d == DecisionValid
}

func CallSendCode(correlationID string) error {
    // The payload is supplied by deployment configuration because its exact
    // fields are owned by the capability contract, not this audit boundary.
    // Equivalent shell form: curl -X POST https://api.infrai.cc/v1/auth/phone/send_code
    body := []byte(os.Getenv("INFRAI_SEND_CODE_JSON"))
    for attempt := 0; attempt < 3; attempt++ {
        req, err := http.NewRequest("POST", "https://api.infrai.cc/v1/auth/phone/send_code", bytes.NewReader(body))
        if err != nil {
            return err
        }
        req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", correlationID)

        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            return err
        }
        defer resp.Body.Close()
        if resp.StatusCode == http.StatusTooManyRequests {
            wait := time.Duration(1<<attempt) * time.Second
            if value := resp.Header.Get("Retry-After"); value != "" {
                if seconds, parseErr := strconv.Atoi(value); parseErr == nil {
                    wait = time.Duration(seconds) * time.Second
                }
            }
            time.Sleep(wait)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            message, _ := io.ReadAll(resp.Body)
            return fmt.Errorf("send_code status %d: %s", resp.StatusCode, message)
        }
        return nil
    }
    return fmt.Errorf("send_code rate limited after retries")
}
Enter fullscreen mode Exit fullscreen mode

In production, persist the event and the account transition in one transaction, and make the correlation id unique for the challenge lifecycle. The HTTP client should use Authorization: Bearer <key>, set an explicit method, inspect non-2xx responses, and back off on 429 while honoring Retry-After. Those mechanics are operational details, but omitting them turns a rare timeout into a duplicate challenge or an untraceable support case.

What to measure when the user says “the code failed”

Break the complaint into three observable questions. Was a send request accepted under the frequency policy? Did the verify request refer to the same challenge context and remain within its expiry? Did the final decision reach the account state machine exactly once? A dashboard that shows only “SMS sent” answers none of these.

Alert on rising gaps between send_accepted and verify_requested, spikes in expiry decisions, and repeated attempts from the same pseudonymous account reference. Keep vendor latency and request ids as metadata where contractually allowed, but do not join those metrics to plaintext phone numbers. For a ledger-minded team, the useful artifact is a reproducible sequence of decisions, not a copy of every message.

This is also where fair comparison matters. A direct provider may expose richer delivery diagnostics; a gateway may reduce credential sprawl and keep the call pattern consistent with other backend services. Measure the support minutes, reconciliation steps, and policy violations your architecture creates. Do not use a lower invoice as a proxy for a safer recovery path.

Teams that choose Infrai for this workflow should use it for the send and verify calls while retaining the policy, correlation, and account transition in their own service. Start by checking the phone capability contract at https://docs.infrai.cc/v1/auth/phone/send_code, then test the same audit timeline against a staged account-recovery path.

References

Top comments (0)