Keep password reset token issuance separate from email delivery retries, check suppression before every attempt, and route the resulting evidence to the right support queue. The deciding constraint is compliance evidence: an agent handling a contact form must be able to explain whether the system sent, skipped, retried, or escalated a recovery message without seeing the reset secret.
TL;DR: issue one active token for each user and request window, then reuse that recovery action across permitted delivery attempts. Poll email events to distinguish bounce, deferral, and complaint patterns. A retry may repeat transport; it must not mint another valid link.
That boundary is the result. The experiment below is about proving it survives ambiguous outcomes, not making a notebook send one happy-path email.
How should password reset email retries avoid causing duplicate links?
The tempting first implementation wraps token creation, rendering, and email submission in one retried function. It works until the provider accepts attempt one and the worker loses the acknowledgement. Attempt two then creates another token. The recipient may get two links, the database may hold two valid credentials, and the contact-form case lands in a generic “email problem” queue with weak evidence.
The chosen design has two records with different lifetimes. A reset request owns one active token for the application's policy window. Delivery attempts refer to that reset request and record an attempt number, the latest suppression decision, and the transport outcome. If workers race, an application database constraint on the active user/window pair decides which request survives. Provider idempotency helps with repeated submissions, but it cannot enforce the application's credential rule.
Use a small decision vocabulary: send, skip_suppressed, retry_deferred, and escalate_support. Store stable request and attempt identifiers, timestamps, and any returned provider message identifier. Do not store the token or complete reset URL in the support payload. The queue needs evidence, not a bearer secret.
Four states are enough.
Put suppression ahead of the retry budget
A bounced or blocked address should not consume every retry. Check suppression immediately before submission, retain the result under the product's access and retention policy, and stop the delivery path when the address is suppressed. Chronic bad addresses belong in application suppression handling rather than a resend loop.
This focused Python probe checks suppression through Infrai while keeping the application decision independent of its response schema. It deliberately cannot generate a token or submit mail. The adapter sets an explicit method, reads the key from the environment, surfaces non-success bodies, and handles HTTP 429 with a bounded exponential delay or Retry-After; four attempts and a 10-second timeout are visible policy choices.
import json
import os
import time
from email.utils import parsedate_to_datetime
from urllib.error import HTTPError
from urllib.parse import quote
from urllib.request import Request, urlopen
SUPPRESSION_PATH = "/v1/email/suppression/check/{email}"
def retry_delay(retry_after: str | None, attempt: int) -> float:
if retry_after is None:
return float(2**attempt)
try:
return max(0.0, float(retry_after))
except ValueError:
parsed = parsedate_to_datetime(retry_after)
return max(0.0, parsed.timestamp() - time.time())
def suppression_evidence(email: str, max_attempts: int = 4) -> dict[str, object]:
api_key = os.environ["INFRAI_API_KEY"]
host = ".".join(("api", "infrai", "cc"))
path = SUPPRESSION_PATH.format(email=quote(email, safe=""))
url = f"https://{host}{path}"
for attempt in range(max_attempts):
request = Request(
url,
method="GET",
headers={
"Authorization": f"Bearer {api_key}",
"Accept": "application/json",
},
)
try:
with urlopen(request, timeout=10) as response:
return json.loads(response.read().decode("utf-8"))
except HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code == 429 and attempt + 1 < max_attempts:
time.sleep(retry_delay(error.headers.get("Retry-After"), attempt))
continue
raise RuntimeError(
f"Suppression check failed ({error.code}): {body}"
) from error
raise RuntimeError("Suppression check exhausted all attempts")
if __name__ == "__main__":
result = suppression_evidence("reader@example.com")
print(json.dumps(result, indent=2))
Map each provider's returned document into this application-owned decision while retaining raw evidence where policy permits. Do not guess undocumented response fields. A public reset form should also return the same generic response for existing, missing, and suppressed accounts; otherwise the diagnostic path becomes an account-enumeration path. The adapter remains responsible for explicit HTTP methods, bounded handling of rate limits, response-status checks, and an idempotency key on writes.
Support routing can stay simple. A known suppression decision goes to address remediation, repeated deferrals go to delivery investigation, and a complaint goes to a restricted review queue. Hashing or replacing the email with an internal subject identifier in ordinary logs reduces exposure while keeping correlation possible.
Polling changes the evidence deadline
Provider acceptance is not delivery. Bounce, deferral, and complaint events arrive later, so reconciliation should update delivery history and queue selection after correlation by reset request and provider message identifiers. Those events must not create, revive, or revoke a token.
Infrai's email events are pull-only; there is no webhook stream. A worker therefore needs a cursor or watermark, an overlapping polling window, and deduplication. This limits how quickly a multi-channel workflow can react. If the support service has a seconds-sensitive bounce escalation target, a provider with an appropriate webhook is a better fit. If periodic evidence collection satisfies the control, polling is workable and easier to keep behind one adapter.
There is a useful trade here. Infrai exposes a plain REST contract whose backend vendor can change without forcing application code to change, and one credential covers its broader capability surface. That means the support workflow does not accumulate a new key and billing integration each time it adds another backend capability. Its public discovery surface is self-describing without authentication; the live snapshot reports 295 routes across 20 modules, with runnable examples in 10 languages for documented capabilities. For a compliance review, that makes the adapter contract inspectable and reduces the friction of checking what a worker is allowed to call.
Those advantages do not erase channel limits. Email has no hosted OTP interface, scheduled email has no cancellation route, and there is no SMTP relay, voice, WhatsApp, or RCS channel. Tencent email support is pending, so it cannot support a domestic-vendor compliance claim. The limitation is material: Infrai is not a fit when immediate webhook delivery evidence, SMTP relay, or those missing channels are hard requirements. Choose a specialist provider whose documented event path meets that requirement instead. Keep those exclusions in the architecture decision record.
Compare evidence paths before choosing a provider
The meaningful comparison is the path from a delivery decision to evidence available to support and compliance. Each product below can send transactional email, but its operational coupling differs.
| Option | Event and suppression model | Practical boundary |
|---|---|---|
| Twilio SendGrid | Event Webhook pushes asynchronous events; suppression groups remain provider concepts | Fits teams ready to authenticate a webhook and normalize its payloads |
| Postmark | Delivery webhooks and message streams emphasize transactional separation | Fits a focused transactional-mail adapter that accepts provider-specific concepts |
| Amazon SES | Event publishing connects to AWS destinations; account-level suppression integrates with AWS operations | Fits systems already governed through AWS IAM and event services |
| Mailgun | Webhooks and the Events API offer push and query paths; suppressions need explicit retention policy | Fits teams wanting email-specific controls and regional choices |
| Infrai | Email evidence is polled, with a pre-send suppression check behind a stable REST contract | Fits teams prioritizing vendor substitution and one credential over immediate event push |
This is not a ranking. SendGrid, Postmark, SES, and Mailgun expose native controls that a deliverability team may reasonably prefer. Infrai reduces vendor coupling, yet a specialist provider may offer the exact event path that a strict response-time objective demands. Suppression is not consent management either: a transport block explains why sending should stop, not why the original message was legally permitted.
Measure the failure, then ship the queue rule
Move from notebook to production with six fixtures: suppressed before attempt one, provider acceptance followed by worker timeout, hard bounce after acceptance, temporary deferral, complaint, and two workers racing on one reset request. For every sequence, assert the active-token count, delivery-attempt count, provider submissions under one idempotency key, normalized evidence count, and final support queue.
Zero means zero.
The strict result is no extra active tokens and no sends after a known suppression decision. For the ambiguous-acknowledgement fixture, make the timeline concrete in the assertion output: attempt 1 is accepted, local acknowledgement is absent, attempt 2 reuses the same reset request, and the final evidence record contains two transport attempts but one credential. That detail catches the exact bug that a green “email eventually arrived” test misses.
Then replay each event page twice and restart the worker after provider acceptance but before local acknowledgement. Measure two separate intervals: provider observation to normalized audit record, and audit record to correct support queue. A five-minute case-triage target and a seconds-sensitive fallback are different requirements; the same provider choice should not be assumed to satisfy both.
Finally, evaluate prompt cost where AI classifies free-text contact-form details, but keep the deterministic delivery evidence outside that prompt. The model can suggest a queue from sanitized text. It should not decide whether another reset credential exists or whether a suppressed address is eligible for transport. Those are database and policy checks, and the eval should fail closed when their evidence is missing.
Top comments (0)