Short answer: for a B2B marketplace using SMS OTP login, choose a provider boundary that emits the verification evidence your reviewers need. When there are no webhooks, polling can enrich the delivery-status trail, but a delivered receipt must never decide that the seller proved possession. A hosted verification service is the smallest operational surface; owning the code lifecycle gives finer evidence and policy control.
| Boundary | Pick this when | Evidence retained | Main burden |
|---|---|---|---|
| Hosted verification | Standard issue and check semantics fit policy | Challenge ID, outcome, timestamps, policy version | Reconcile external outcomes with the login audit record |
| Application-owned challenge over SMS | Reviewers require control of expiry, attempts, or resend lineage | Challenge transitions plus a separate delivery trail | Secure generation, race handling, throttling, support tools |
| Phishing-resistant primary method with SMS recovery | Seller accounts can change payouts or expose sensitive orders | Primary-authenticator result and governed recovery events | Enrollment, recovery, and account lifecycle |
The decisive split is small: a delivery system answers what happened to a message, while an authentication system answers whether a presented code is valid for one challenge. Keep those answers separate.
Pick the evidence boundary first
Pick hosted verification when its lifecycle already matches the policy you can defend. The application creates a challenge, stores the opaque identifier, and later records the verification result beside its own seller, order-view session, policy version, and correlation ID. Code generation, expiry, and comparison stay behind one boundary. Confirm the documented status model, retention behavior, regional handling, and rate controls before treating those records as compliance evidence.
Pick an application-owned challenge when the audit question is more specific: Which policy was active? Did this resend supersede an earlier code? Was denial caused by expiry, an attempt ceiling, or replay? The control is real. So is the responsibility. Generate codes with a cryptographically secure random source, store a keyed digest rather than the code, make verification atomic, and ensure one successful comparison consumes the challenge. A transaction must settle two simultaneous correct submissions deterministically.
Pick a phishing-resistant primary authenticator with separately governed SMS recovery when the account's power justifies it. NIST SP 800-63B treats the public switched telephone network for out-of-band authentication as restricted and asks verifiers to consider risks such as SIM change and number porting. This does not make every texted code unusable. It means a marketplace should document why SMS is acceptable for this action, watch for risk signals, and avoid letting a weaker recovery path silently erase the primary path's strength.
These are architecture choices, not a product ranking. Choose the boundary whose records can reconstruct the access decision without pretending transport telemetry proves identity.
What should OTP login provider polling status mean without webhooks?
Polling proves only what the queried status contract defines. A messaging API may expose accepted, queued, sent, delivered, undeliverable, or another provider-specific state. Those names are not universal. A terminal receipt does not show that the intended human read the message, much less prove that the submitted code matched.
No receipt can prove possession.
No webhook means the application owns the observation schedule. Start after a short delay, back off, stop at a documented terminal state or fixed deadline, and preserve the last response time. Do not poll from the browser. A browser loop stops when the tab closes and can multiply traffic during refreshes; a worker keyed by message ID is easier to bound and observe.
Picture two rails. On the upper rail, challenge_created -> code_submitted -> verified | denied | expired. On the lower rail, message_accepted -> status_observed -> terminal | observation_ended. They share a correlation ID, but there is no arrow from delivered to verified.
That missing arrow matters.
Polling changes evidence timeliness too. The database knows only the latest state it observed, at the time it observed it. Store both the provider event time when the contract supplies one and your observation time. If no trustworthy event time exists, do not manufacture one. Record the limitation.
Build one authoritative challenge state machine
This TypeScript sketch uses generic ports rather than a vendor SDK. Its five observation delays are an example marketplace policy, not universal security constants; calibrate them against threat modeling, support load, local rules, and the assurance required for opening an order.
type ChallengeState = "pending" | "verified" | "denied" | "expired";
type DeliveryState =
| "unobserved"
| "accepted"
| "in_transit"
| "delivered"
| "undeliverable"
| "unknown";
type Challenge = {
id: string;
sellerId: string;
orderAccessSessionId: string;
codeDigest: string;
state: ChallengeState;
attempts: number;
expiresAt: Date;
resendAvailableAt: Date;
generation: number;
messageId: string;
policyVersion: string;
};
interface ChallengeStore {
compareAndConsume(input: {
challengeId: string;
expectedGeneration: number;
presentedDigest: string;
now: Date;
maximumAttempts: number;
}): Promise<"verified" | "mismatch" | "expired" | "consumed" | "limited">;
recordDeliveryObservation(input: {
challengeId: string;
messageId: string;
observedAt: Date;
providerEventAt?: Date;
state: DeliveryState;
}): Promise<void>;
}
compareAndConsume is intentionally one store operation. Reading a row, comparing in application memory, and then updating creates a race in which concurrent requests can both appear successful. The store should compare the keyed digest in constant time within the trusted boundary, reject expired or superseded generations, increment failed attempts, and consume a valid challenge atomically. Never write the submitted code, its digest, or the full phone number to logs.
A resend creates a new generation. It does not extend the old code or create a second valid code. Preserve lineage so support staff can say generation 2 replaced generation 1, but permit only the current generation to verify. The UI can offer resend after the server-provided cooldown; the server remains authoritative if a client bypasses the disabled button.
Here is a bounded delivery observer. It ends when delivery telemetry is unavailable because login verification remains on the other rail.
type DeliverySnapshot = {
state: DeliveryState;
terminal: boolean;
providerEventAt?: Date;
};
interface MessageStatusReader {
read(messageId: string): Promise<DeliverySnapshot>;
}
const delaysMs = [2_000, 5_000, 10_000, 20_000, 40_000] as const;
async function observeDelivery(
challenge: Challenge,
reader: MessageStatusReader,
store: ChallengeStore,
wait: (milliseconds: number) => Promise<void>,
): Promise<void> {
for (const delay of delaysMs) {
await wait(delay);
const snapshot = await reader.read(challenge.messageId);
await store.recordDeliveryObservation({
challengeId: challenge.id,
messageId: challenge.messageId,
observedAt: new Date(),
providerEventAt: snapshot.providerEventAt,
state: snapshot.state,
});
if (snapshot.terminal) return;
}
}
In production, the wait belongs in a durable queue rather than an in-process timer. Give each observation job an idempotency key derived from message ID and attempt number. Bound concurrency, apply jitter, and classify authentication failures separately from transport failures. A status API timeout should affect an observability metric; it should not consume one of the seller's code-entry attempts.
Make retry UX and abuse controls agree
A calm verification screen needs a few states: waiting for input, checking, incorrect or expired, temporarily limited, and verified. Keep error copy generic enough to avoid account enumeration. Preserve phone-number masking consistently. Support platform one-time-code autofill where available; automatic entry reduces transcription friction without changing the server's rule.
Retry and resend are different actions. Retry submits another candidate against the same challenge and spends an attempt. Resend asks the server to authorize a new generation, subject to cooldown and rate policy. Delivery polling spends neither. Mixing these counters creates ugly behavior: a delayed receipt can lock out a real seller, or repeated sends can continue while the UI claims only code attempts are limited.
Count them separately.
Apply limits across several keys because each catches a different abuse shape: challenge, seller account, destination, network source, and a privacy-preserving device signal where lawful. NIST requires rate limiting for authentication secrets with less than 64 bits of entropy, but it does not supply one universal number for every application. Make limits policy data. Version them. Return retry timing where appropriate, and record the decision reason internally while keeping the public response nondisclosing.
Metrics should mirror the two rails. For authentication, count challenges created, verification outcomes by reason, time to verification, resends per challenge, and limit decisions. For transport, count status observations, terminal delivery categories, observation age, and status-read errors. Alert on ratios and sustained shifts with minimum-volume guards; raw failure counts mostly track traffic. Keep phone numbers and codes out of metric labels. High-cardinality correlation IDs belong in traces or controlled audit records, not metric dimensions.
The audit record should answer one sentence: seller S gained access to order-view session O because challenge C, under policy P, moved from pending to verified at time T. Delivery observations remain attached evidence rather than the authorization reason.
Test transitions, not just the happy endpoint. Use a fake clock around expiry boundaries. Race two correct submissions and require exactly one success. Submit an old generation after resend. Repeat the same observation job. Simulate a status reader that times out forever and confirm that a correct code can still verify. Finally, test logging with trap values that must never appear.
Know the limits of the evidence
SMS delivery evidence has a hard ceiling. A status can describe message handling, not who controlled the handset, whether a number was reassigned, or whether social engineering moved service to another SIM. Polling cannot raise that ceiling; it only changes how the application collects transport state.
Keep the final rule crisp: authorize order access from the atomic challenge result, retain the minimum evidence required by policy, and treat message status as diagnostic context. For higher-risk seller actions, use an authenticator and recovery design whose phishing resistance and lifecycle match the consequence of account takeover.
Top comments (0)