TL;DR: Treat a missing password reset email as a traceable delivery-state problem, not as a prompt to resend blindly. The least complex useful design gives every reset attempt an opaque message ID, records submission and downstream disposition separately, checks SPF, DKIM, and DMARC alignment at the exact sending domain, and pages only when a seller-impacting symptom burns a defined reliability budget. For a fintech marketplace, this matters because a seller locked out of the account may also miss the new-order workflow waiting behind login.
The page fires: seller_recovery_delivery_degraded. On-call sees that password recovery requests are succeeding, but an abnormal share has no accepted, bounced, or deferred disposition within the observation window. The dashboard can split the symptom by recipient domain, sending domain, template version, and transport route without exposing addresses or reset tokens. It cannot yet say "spam folder." No sender receives a universal mailbox-placement receipt, so claiming inbox placement from an accepted SMTP handoff would be false confidence.
That distinction controls the response. An SMTP acceptance does not prove that the person saw the message. DMARC can establish domain-aligned authentication and publish handling policy, but it does not promise delivery. Start with the evidence chain, preserve account-recovery security, and use a second channel only as a deliberately governed fallback.
This is transactional email troubleshooting, not a hunt for one magic record.
What if the password reset email was delivered to the spam folder?
First decide whether this is submission failure, recipient rejection, delayed delivery, or suspected filtering. Those states need different owners and different remediation. If the application failed before enqueueing, DNS work is noise. If one recipient domain is deferring traffic while others accept it, changing the reset flow adds risk without addressing the boundary that moved.
Work backward from the seller-visible symptom. The order event exists. The seller tries to sign in, requests recovery, and receives no message. The recovery API gives a generic response so account existence is not disclosed, while the internal event stream records an attempt under a random correlation ID. A dispatcher submits the message. A later processor attaches accepted, deferred, or bounced evidence to that ID. The alert evaluates the aging population of attempts rather than a raw count of support complaints.
The signal that should have fired earlier is growth in recovery messages stuck between submission and terminal transport evidence, segmented by domain and route. That is earlier and more actionable than "users say mail is in spam," but it must remain a symptom signal: accepted messages can still be filtered, and missing feedback can be an instrumentation defect. A useful page therefore carries numerator, denominator, observation window, event freshness, and the largest affected segment. Without the denominator, a quiet period and a real outage can look identical.
Stop there.
Keep the state model small:
-
requested: the application accepted a recovery request. -
queued: an immutable notification job exists. -
submitted: the sender handed it to the transport boundary. -
accepted,deferred, orbounced: downstream evidence arrived. -
consumed: the reset token was used, recorded separately from transport success.
Do not collapse accepted and consumed. Consumption is affected by user behavior, token expiry, and product UX; transport disposition is affected by sending infrastructure and recipient policy. Combining them produces an alert that can page the wrong team.
The limitation is deliberate: this state machine describes evidence the SaaS platform can observe, not the private decision inside every mailbox. If the support report says "spam folder" while the trace ends at accepted, keep both facts. Do not promote a guess to a delivery state.
Instrument the boundary without logging the secret
A reset URL is a credential. It does not belong in logs, traces, metrics labels, or delivery metadata. The same applies to the raw recipient address. Correlate with a random message ID, retain a normalized recipient domain only where policy permits, and make identifiers useless for reconstructing the link.
The instrumentation change is modest: emit one structured event at every ownership transfer and reject incomplete events. This Go shape is transport-neutral. It separates message identity from secret-bearing content and keeps metric dimensions bounded.
package delivery
import (
"crypto/rand"
"encoding/hex"
"errors"
"time"
)
type State string
const (
Requested State = "requested"
Queued State = "queued"
Submitted State = "submitted"
Accepted State = "accepted"
Deferred State = "deferred"
Bounced State = "bounced"
Consumed State = "consumed"
)
type Event struct {
MessageID string `json:"message_id"`
State State `json:"state"`
RecipientDomain string `json:"recipient_domain"`
SendingDomain string `json:"sending_domain"`
TemplateVersion string `json:"template_version"`
TransportRoute string `json:"transport_route"`
ObservedAt time.Time `json:"observed_at"`
}
func NewMessageID() (string, error) {
b := make([]byte, 16)
if _, err := rand.Read(b); err != nil {
return "", err
}
return hex.EncodeToString(b), nil
}
func (e Event) Validate() error {
if e.MessageID == "" || e.ObservedAt.IsZero() {
return errors.New("message ID and observation time are required")
}
if e.RecipientDomain == "" || e.SendingDomain == "" {
return errors.New("bounded domain dimensions are required")
}
return nil
}
A 16-byte random identifier gives correlation without embedding account data. Template version deserves a dimension because a content or rendering change can coincide with filtering, while transport route tells on-call whether the symptom follows an infrastructure boundary. Never use a full error string as a metric label; classify it into a reviewed, finite reason set and keep the original response in access-controlled logs under an appropriate retention policy.
Retries need similar discipline. Retry transient deferrals through a durable queue with bounded attempts and backoff, but do not retry a permanent rejection forever. Make submission idempotent against the notification job ID, because a timeout at the transport boundary leaves the sender uncertain whether the remote side accepted the first attempt. Blind resends can turn one recovery request into several valid-looking messages, train recipients to distrust the channel, and create unnecessary load exactly when the system is degraded.
Authentication narrows the search; it does not close the incident
Once the trace proves that messages reach the transport boundary, inspect authentication at the domain that appears to recipients. SPF authorizes hosts to use a domain in SMTP identity. DKIM adds a domain signature to selected headers and the body. DMARC evaluates alignment between the visible From domain and an authenticated SPF or DKIM domain, then lets the domain owner publish a requested handling policy and receive reports. RFC 7489 defines that relationship.
The operational trap is checking for the mere presence of three DNS records. Presence is not alignment. A message can pass SPF for a transport domain that is not aligned with the visible sender, or carry a valid DKIM signature for a different organizational domain. Inspect the authentication results on a received sample, the exact domains involved, and the DMARC reports for aggregate patterns. Avoid putting addresses, tokens, or message bodies into a ticket while collecting that evidence.
Forwarding complicates the picture because the forwarding system may no longer be authorized by the original SPF policy; content modification can also invalidate a DKIM signature. DMARC is designed to work when at least one aligned mechanism succeeds, which is why the investigation should preserve both mechanisms rather than treating SPF as a universal verdict. Reports help locate patterns. They do not identify every individual mailbox decision.
That is the trade-off.
Roll policy changes in stages backed by report review. A syntactically valid but operationally incomplete authorization change can reject legitimate streams that were omitted from the inventory. Inventory every sender using the domain, separate recovery mail from bulk traffic operationally, verify aligned authentication, and then choose policy based on observed legitimate traffic and the organization's abuse posture. This is change management, not a DNS checkbox.
SMS can be a recovery contingency, but it creates a second regulated delivery system rather than a free reliability multiplier. In the United States, application-to-person traffic over ten-digit long codes has registration and compliance requirements documented under A2P 10DLC. A fallback design therefore needs consent, eligibility, sender registration, opt-out handling, and its own disposition telemetry before it belongs in an incident runbook. Do not silently switch channels because email is slow.
Capacity and ownership decide whether the design survives on-call
The buy-or-build question is narrower than it first appears. A team may delegate transport while retaining message orchestration, security, telemetry, and incident ownership. Or it may operate more of the mail path and accept the corresponding queueing, reputation, abuse, authentication, and feedback-processing burden. The correct boundary follows staffing and SLO obligations, not feature-count enthusiasm.
Self-operation is not suitable for a team that cannot staff abuse response and transport operations. Managed transport has a different limitation: normalized events can hide useful raw detail, and changing that event contract later can be expensive. A second channel is an alternative only when the team can own its consent and compliance path. These are operational boundaries, not product rankings.
| Boundary | Team owns | External boundary owns | Primary on-call cost | Lock-in pressure |
|---|---|---|---|---|
| Managed transport | Recovery policy, templates, IDs, event normalization, SLOs | SMTP delivery infrastructure and raw dispositions | Reconciling events and escalating ambiguous states | Event schemas and suppression behavior |
| Self-operated transport | All of the above plus queues, domain operations, feedback ingestion, and abuse controls | Recipient networks | Continuous transport and reputation operations | Lower API coupling, higher operational coupling to mail standards |
| Managed primary plus governed alternate channel | Channel policy, consent, orchestration, cross-channel SLOs | Delivery infrastructure per channel | Correlated incidents and compliance drift across two systems | Multiple event and policy models |
Capacity planning starts at bursts, not daily averages. New-order traffic and account-recovery demand can be correlated: a marketplace promotion increases orders, sellers sign in after a gap, and recovery requests rise at the same time. Queue age, worker saturation, disposition lag, and retry volume should therefore be tested under a declared burst model. Pick the numbers from actual order and recovery distributions; inventing a fashionable multiplier would only disguise missing demand data.
The same skepticism applies to the SLO. Define the user journey and the measurable boundary. An internal service-level indicator can measure the fraction of eligible recovery jobs that receive non-deferred transport evidence within a chosen window, while a separate product indicator measures successful token consumption. The target and window must come from business tolerance and observed baselines. Neither indicator proves inbox placement, so the dashboard should say exactly what it measures.
No universal threshold exists.
Release changes by cohort: sending domain, template version, route, or a stable recipient-domain slice. Compare disposition distributions and trace completeness before widening. A synthetic mailbox set can validate that messages are generated, authenticated, and observable, but it cannot represent every recipient's filtering policy. Keep support reports as a correlated signal, not as the only monitor.
The false-positive budget is part of delivery reliability
A threshold that pages on every short disposition gap will be noisy during normal deferrals and telemetry lag. A threshold so broad that it waits for many locked-out sellers is operational theater. The useful middle ground is a multi-window symptom alert that combines an affected fraction, a minimum event volume, and event-pipeline freshness, then routes low-volume domain-specific anomalies to a ticket unless seller impact is already visible.
This is where the earlier instrumentation pays for itself. On-call can distinguish a stale feedback consumer from a transport rejection surge, identify whether one domain or every route is affected, and decide whether to pause a rollout without touching reset-token logic. The alert should also suppress itself when the denominator is too small to support the chosen inference; low traffic is a reason to gather evidence, not a reason to manufacture certainty.
Every page consumes attention. Repeated false positives teach responders to delay, while a threshold based only on accepted handoffs hides mailbox-placement complaints. Review alert precision after incidents and non-incidents, record which signal led to action, and adjust the window only with enough retained evidence to understand the consequence. The goal is not zero pages. It is a page that names a seller-impacting risk, presents the trace break, and gives the responder one defensible next check.
This alerting approach is also a poor fit for a tiny stream whose denominator rarely supports a stable rate. Use case-level inspection and a ticket there; reserve paging for a symptom with enough volume or direct seller impact to justify interruption. The explicit choice is lower detection speed in exchange for fewer unactionable pages.
For the marketplace seller waiting behind a password reset, reliable notification is the combined result of secure recovery design, aligned authentication, observable ownership transfers, controlled retries, and an alert calibrated to evidence. No single DNS record or fallback channel can substitute for that chain.
Further reading
- RFC 7489, Domain-based Message Authentication, Reporting, and Conformance (DMARC): https://datatracker.ietf.org/doc/html/rfc7489
- US A2P 10DLC compliance documentation: https://www.twilio.com/docs/messaging/compliance/a2p-10dlc
Top comments (0)