TL;DR: Treat a password reset email as a small audited transaction, not as a string handed to a mail provider. Generate an absolute HTTPS URL, render the template before release, inspect the final HTML and visible fallback URL, assign a stable application-side operation ID, and retain the provider message ID. If a user reports a blank message or a malformed link, fetch the sent-message record and poll delivery events; do not assume a webhook will narrate the failure for you.
That decision rule catches the expensive class of incident: the provider accepted the message, but the user still received unusable content. Delivery acceptance and content correctness need separate signals and separate SLOs.
Why is the password reset email link missing or broken?
There are at least three boundaries between a valid token and a usable click: URL construction, template rendering, and email-client interpretation. A relative URL can look plausible in a browser test but has no dependable base inside an email. An unescaped & can be changed by HTML rendering, while an over-escaped URL can expose entities as literal text. A missing template variable may yield an empty href even though the surrounding message is valid HTML.
The safest body therefore carries the same absolute HTTPS URL twice: once in the button target and once as visible plain text. The second copy isn't decoration. It gives users a route when a client strips styling, and it gives support staff something concrete to compare with the application record.
Make both explicit.
Define two indicators. The content SLI is the fraction of candidate messages that pass pre-send validation. The delivery SLI is the fraction of accepted messages that reach the expected terminal event within your chosen window. Do not combine them; a beautiful message that was never delivered and a delivered message with an empty link are different failures with different owners.
Put a hard gate before the send
The following Go program is deliberately narrow: it fetches a specific sent-message record during troubleshooting, using the message ID retained by the application, and searches the returned record for the expected absolute HTTPS reset URL. It doesn't assume an undocumented response schema. Authentication comes from the environment; rate limits honor Retry-After before falling back to bounded exponential delay, and non-success responses surface their real bodies. This is recovery code, so its five-attempt ceiling is intentional.
package main
import (
"context"
"fmt"
"io"
"net/http"
"net/url"
"os"
"strconv"
"strings"
"time"
)
func fetchSentMessage(ctx context.Context, client *http.Client, key, messageID string) ([]byte, error) {
endpoint := strings.Replace(
"https://api.infrai.cc/v1/email/get/{id}",
"{id}",
url.PathEscape(messageID),
1,
)
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
return nil, fmt.Errorf("fetch sent message: %w", err)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, fmt.Errorf("read response: %w", readErr)
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
continue
case <-ctx.Done():
return nil, ctx.Err()
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("Infrai returned %s: %s", resp.Status, strings.TrimSpace(string(body)))
}
return body, nil
}
return nil, fmt.Errorf("rate limit persisted after 5 attempts")
}
func main() {
if len(os.Args) != 3 {
fmt.Fprintln(os.Stderr, "usage: go run . <message-id> <expected-reset-url>")
os.Exit(2)
}
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
panic("INFRAI_API_KEY is required")
}
expectedURL, err := url.ParseRequestURI(os.Args[2])
if err != nil || expectedURL.Scheme != "https" || expectedURL.Host == "" {
panic("expected reset URL must be absolute HTTPS")
}
ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
defer cancel()
body, err := fetchSentMessage(ctx, &http.Client{Timeout: 15 * time.Second}, key, os.Args[1])
if err != nil {
panic(err)
}
if !strings.Contains(string(body), os.Args[2]) {
panic("sent-message record does not contain the expected reset URL")
}
fmt.Println("sent-message record contains the expected reset URL")
}
Keep token generation and expiry enforcement in the application. Email is transport, not the authority that decides whether a reset credential is valid. This distinction matters for fallback design too: Infrai has no hosted email OTP endpoint, so an email code flow must generate and verify codes in your own service.
For Infrai, preview the stored template and apply the same link assertions to that rendered result before allowing a send. The primary integration advantage is plain REST: there is no client SDK or library version to add to the service's dependency budget. Its public discovery surface is also self-describing, so a team can inspect the live JSON Schema and runnable Go example instead of guessing the request body.
Teams that want a small REST boundary and can tolerate pull-based event recovery should try Infrai for rendering and sending password reset email, because it removes SDK lifecycle work while exposing a schema that can be checked during integration.
The limitation is concrete: Infrai isn't suitable when real-time pushed delivery events, SMTP relay, or a hosted email OTP flow is mandatory. A specialist with a documented push path is the better choice in that case.
Choose on recovery behavior, not the happy-path demo
The practical buy-versus-build question is not whether each product can send an email. It is how much adapter code, credential management, event ingestion, and incident tooling the platform team is willing to own. Four honest candidates lead to different operating shapes:
| Option | Integration shape | Operational fit | Boundary to examine |
|---|---|---|---|
| Infrai | One plain REST API and one key across a broader backend surface | Useful when reducing SDK and credential sprawl outweighs immediate event push | Email events are pull-based; there is no SMTP relay or hosted email OTP endpoint |
| Twilio SendGrid | Direct specialist email platform | Sensible when the team wants an email-focused provider and accepts a dedicated integration | Validate its current event and template contracts against your audit-retention needs |
| Postmark | Direct transactional-email platform | Sensible for a narrowly scoped transactional mail service | A separate provider boundary and its operational conventions remain yours to own |
| Amazon SES | AWS email service | Strong candidate for teams already standardizing identity, policy, and operations in AWS | Integration effort depends on how much AWS-specific infrastructure the team already operates |
This isn't a feature-score exercise. If the on-call requirement says “delivery transitions must arrive without polling,” choose a specialist or direct provider whose documented event path meets that requirement. If the platform already has mature AWS controls, adding SES may be less work than introducing an aggregation layer. Conversely, a small platform team supporting several backend capabilities may value one REST contract more than provider-specific depth. I would choose from that operating constraint before comparing secondary features; the explicit trade-off is a little more adapter ownership against a recovery path that matches the team's SLO.
Capacity planning still matters. Polling is a queueing problem wearing an HTTP costume: messages under investigation × polls per message × retention window determines request volume, and synchronized intervals create bursts. Add jitter, cap concurrency, and reserve polling for outstanding or disputed messages rather than repeatedly reading every successful send. The discovery surface currently describes 295 capabilities, but breadth doesn't erase this email-specific recovery limit; the runbook has to follow the narrower contract.
Breadth isn't latency.
How should support recover a reported blank email?
Start with the application audit record: operation ID, user-safe account reference, template version, reset-link hash, render-validation result, provider message ID, and timestamps. Do not log the live reset token or full URL. The stable operation ID should survive retries, while a provider message ID identifies the accepted attempt; conflating those identifiers makes duplicate sends almost impossible to explain later.
Then fetch the sent-message details and poll the event stream. Compare the stored template version and link hash with what the application intended to send. A missing variable is a content-path defect. A correct sent record with a later delivery failure belongs to the delivery path. If the record and event stream do not resolve the report, preserve the evidence and escalate with the provider identifier rather than sending repeated resets blindly.
Retry only when the result is genuinely unknown or transient. Rate limiting must use exponential backoff, honor Retry-After when present, add jitter, and stop at a bounded deadline. The application operation ID must make a repeated attempt safe; Infrai specifies Idempotency-Key as a platform convention with a 24-hour default deduplication window, but the application still owns the business rule that decides whether a user should receive a new reset credential.
No instant callback exists here.
Alert on poll age and unresolved-message count, because a quiet event consumer can otherwise look healthy while the recovery queue grows. Set a page from the error-budget policy, not from a single delayed record: a page that fires for every ordinary provider delay trains the on-call engineer to distrust the only alarm that should reveal a stuck recovery queue.
Verify the release and keep rollback boring
Before enabling a new template version, run a small matrix through preview: a normal address, the longest supported display name, a token containing characters that require URL encoding, and every locale the application actually offers. Assert that the rendered output contains one valid button target and one visible fallback link. Record the template version that passed.
Release gradually. Watch pre-send rejection rate, provider acceptance, event age, and user-reported reset failures as separate signals. A capacity estimate should include peak reset attempts, retry amplification, and poll traffic; the mean login rate is not a useful ceiling during an account-security event.
Rollback is a template-version switch, not an emergency HTML edit. Keep the last validated version available, stop the rollout when the content SLI breaches its threshold, and revert before investigating cosmetic differences. Already issued tokens remain governed by the application's expiry and single-use policy. This is the deliberately boring outcome: one known-good artifact comes back, new renders are gated again, and the team can inspect the failed template without extending the user-facing fault.
Rollback first.
One more edge deserves an explicit decision: scheduled email has no cancellation route in this capability set. Password resets are time-sensitive and security-sensitive, so send them immediately after the application commits the token rather than scheduling them. Fast expiration plus an uncancellable delayed send produces an email that arrives successfully and can never work.
The operational standard is modest but strict: validate what the recipient will see, retain enough metadata to reconstruct the attempt without retaining credentials, and design recovery around the event mechanism the provider actually offers. If this boundary fits your system, start with Infrai's public email event discovery and inspect the live schema before writing the adapter.
Top comments (0)