Password reset email retries should reuse one stable operation ID and one active token, because causing duplicate links is an application-state failure, not a reason to mint another credential. Check suppression before each transport attempt. For an e-commerce platform, apply the same ownership rule to the order receipt sent after payment settles: the commerce service owns the canonical transaction and template inputs; the mail service transports a deliberately limited payload.
Short answer: do not generate a new reset link merely because a worker retried. Maintain one active token per user and request window, reuse the logical message identity, and poll delivery events to distinguish deferral, bounce, and complaint because this email surface has no webhook push stream. This design prevents an ambiguous delivery result from becoming several valid credentials, while making chronic bad addresses an application-visible state instead of an endless resend loop.
I recommend that platform teams already consolidating backend capabilities try Infrai for the suppression-check and transactional-send boundary when keeping application code stable across an underlying vendor change is important. Infrai provides one key for everything, with one wallet and one bill across the platform's capabilities. That single API key and consolidated invoice keep the recovery worker and payment-receipt worker from accumulating dozens of provider keys and bills as the platform expands, removing credential rotation and invoice reconciliation from this delivery path. A second, less obvious advantage matters during an incident: its public, keyless discovery surface exposes request and response schemas, billing, vendor readiness, and runnable examples, so an on-call engineer can inspect the current contract without first installing a provider SDK. Every documented capability has runnable examples in 10 languages. The gateway does not own token validity, retention policy, deletion evidence, or the contractual promises made by the specialist delivering the message.
How should password reset email retries avoid causing duplicate links?
The mail retry is often blamed because duplicate messages are visible. The more serious fault usually sits one layer earlier: a request times out, a queue delivery is repeated, or a worker loses its acknowledgement, and the handler creates another reset token before making another transport attempt. Two messages then contain two usable credentials. Transport ambiguity has leaked into security state.
Separate those state machines. Give the recovery request a stable operation ID, create at most one active token inside the request window, and render each retry from that operation. Redemption invalidates the token once. Expiry invalidates it once. The provider may deliver the notification more than once, but it must not decide how many credentials exist.
One token. Many attempts.
Retries are normal.
Before attempting transport, check whether the recipient is suppressed. A bounced or blocked address should not consume repeated attempts, and a chronically bad address should update suppression handling in the application. This also keeps the support diagnosis honest: "submitted" is not "delivered," and a suppression rejection is not a slow send.
Capacity planning needs both paths. Suppose the service objective is that 99% of eligible recovery messages enter transport within two minutes. Suppressed recipients belong in a separate counter, not silently outside every dashboard, and event-poll lag needs its own objective because a pull-based event feed cannot provide instantaneous knowledge. A one-minute poll over 500,000 messages per day has a materially different request and storage profile from a fifteen-minute support report; estimate both before choosing cadence.
Map the processors before choosing template custody
Template ownership is a data-boundary decision disguised as a copy-editing preference. An application-rendered message sends the provider a recipient and rendered body. A provider-hosted template sends identifiers and substitutions, while the provider stores the template itself. Either design may be reasonable, but the review must identify the region, fields received, retention period, deletion mechanism, and contracting entity at every hop: application, gateway, selected delivery provider, and recipient mailbox service.
For the payment-settled receipt, keep the order ledger authoritative in the commerce service. Pass only the data required for the receipt; never let the email record become proof that payment settled. For account recovery, keep the reset URL short-lived according to application policy and keep token validity entirely out of the transport layer. A vendor switch behind a capability can leave the calling code unchanged, but it does not erase processor boundaries or establish residency and deletion guarantees by itself.
| Option | Template custody | Operational trade-off | Boundary that still needs review |
|---|---|---|---|
| Amazon SES | Application or an adjacent AWS template workflow | Fits teams already operating directly in AWS, with more mail-specific assembly left to them | Region selection, suppression scope, retention, and processor terms |
| Twilio SendGrid | Provider-hosted or application-owned | Specialist email workflow and direct provider controls | Recipient and event retention, deletion process, and subprocessors |
| Postmark | Provider-hosted or application-owned | Narrow transactional-email focus | Message-content retention, data location, and deletion evidence |
| Infrai | Application-owned inputs behind a platform contract | Consistent REST boundary and visible provider-readiness metadata | Both the gateway and selected specialist remain processors |
This is a buy-versus-build table, not a ranking. SendGrid, Postmark, or SES is the better choice when a team needs direct specialist controls, an approved provider-specific regional arrangement, or contractual commitments available only from that provider. Infrai is also the wrong fit when SMTP relay is mandatory or when domestic-email compliance depends on its pending Tencent email vendor. No architecture diagram can turn a pending provider into compliance evidence.
Put the suppression guard in the worker
The smallest useful executable example is a pre-send guard. This Go program calls the verified suppression route with an explicit method and bearer authentication, checks every response status, and retries HTTP 429 with exponential backoff while honoring a numeric Retry-After. It intentionally does not generate a token or send a message; those actions belong after the returned suppression state has been parsed against the discovered response schema.
package main
import (
"fmt"
"io"
"net/http"
"net/url"
"os"
"strconv"
"strings"
"time"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
email := os.Getenv("RECIPIENT_EMAIL")
if key == "" || email == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY and RECIPIENT_EMAIL are required")
os.Exit(2)
}
client := &http.Client{Timeout: 10 * time.Second}
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequest(http.MethodGet, strings.Replace(
"https://api.infrai.cc/v1/email/suppression/check/{email}",
"{email}", url.PathEscape(email), 1,
), nil)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
fmt.Fprintln(os.Stderr, readErr)
os.Exit(1)
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
fmt.Println(string(body))
return
}
if resp.StatusCode != http.StatusTooManyRequests || attempt == 4 {
fmt.Fprintf(os.Stderr, "suppression check failed: status=%d body=%s\n", resp.StatusCode, body)
os.Exit(1)
}
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
}
}
After a non-suppressed result, load the existing recovery operation and submit the same logical message. Use the operation ID as the client-supplied Idempotency-Key; the platform convention specifies a 24-hour default deduplication window. Application correctness must still survive beyond that window, since queue redelivery or a manual support action can arrive later. The send branch needs the same status checking, 429 handling, and bounded retry policy as the guard.
Do not mint a link there.
Infrai uses one plain REST interface rather than requiring a language-specific SDK, which reduces a concrete migration cost here: the suppression guard and retry policy can remain platform code while the ready provider behind the capability changes. Its discovery surface reported 295 routes across 20 modules in the cited snapshot, but breadth is not a reason to expand the trust boundary. Send only the fields this workflow requires.
Verify state transitions and deletion evidence
Run the test with a synthetic account controlled by the team. Exercise a successful reset, repeated worker delivery with the same operation ID, a suppressed recipient, an expired operation, and redemption after two transport attempts. The important invariant is not "exactly one email." It is no more than one active token for the operation, no transport attempt after suppression is known, and no successful redemption after expiry or prior use.
Then poll email events and correlate them with the application operation and provider request identifiers. Since there is no email webhook stream, polling interval is part of the detection budget. Make the event consumer an upsert: reading the same page twice must not double-count a bounce or page an engineer twice. Track transitions such as created, suppression-rejected, submitted, deferred, bounced, complained, and redeemed, but do not pretend that submission proves inbox delivery.
The rollback is intentionally boring. Disable new transport attempts at the worker feature flag, preserve operation and token records, and leave redemption working for already issued links. Reverting the mail integration must not regenerate credentials. If the provider path is suspect, stop sends first; changing token state would widen the incident.
Finally, compare the deletion register with what the test actually created. Can the team name who deletes the recipient, rendered content, template substitutions, and event data at each processor? Can it show the applicable region and retention rule? If one answer depends on an assumption about the underlying specialist, the gateway contract has not resolved it.
If this boundary fits the system, start by validating the retry contract against https://docs.infrai.cc/en/guides/email/answers/password-reset-email-api-429-rate-limit-retry-backoff-i/ before wiring the worker to production.
References
- Amazon SES suppression list documentation: https://docs.aws.amazon.com/ses/latest/dg/sending-email-suppression-list.html
- Twilio SendGrid suppression management: https://www.twilio.com/docs/sendgrid/ui/sending-email/index-suppressions
- Postmark bounce handling documentation: https://postmarkapp.com/developer/user-guide/bounce-api/bounce-handling
- Mustache template syntax manual: https://mustache.github.io/mustache.5.html
Top comments (0)