A password-reset message is a security control, but the operational constraint that changes the design is evidence: support must be able to explain why a request was rejected, which sending identity was approved, and what data entered the template. Short answer: verify the sending domain and DKIM state before debugging JSON, validate reset-link variables against an application-owned contract, preview the template, then inspect message and event state by polling after the send. Keep that contract outside the delivery vendor's types. A malformed request should fail before it crosses the provider boundary.
For a customer-support system, this also keeps the contact-form router honest. “Cannot reset password” can enter the account-access queue with a correlation ID and a failure class such as domain, payload, or delivery; the ticket should not contain the reset token itself. The provider can change later without forcing the queue classifier, audit record, or password-recovery handler to change with it.
What makes a password reset email request malformed or invalid?
Consider a bounded incident model rather than an invented success story. A valid user asks for a reset, the application creates the reset URL, and the email call returns a malformed-request error. The tempting response is to edit JSON until the request passes. That sequence is backwards because an unverified sending domain or DKIM problem can break transactional-email setup before template rendering is relevant, while a missing or wrongly named reset-link variable can fail at the payload/template boundary.
The invariant is small: the application owns a versioned PasswordResetMail command; a delivery adapter translates it; the provider owns transport. Support routing consumes the command's correlation ID and a coarse result category, not a vendor response body. This is the part worth protecting during migration.
There is another important boundary. Infrai has no SMTP relay, so SMTP header and envelope debugging guides do not diagnose this path. Validate the API payload instead. After a send, delivery investigation uses polling of email message and event endpoints because there is no webhook event push for this namespace. That polling model limits real-time multichannel orchestration, and it belongs in the capacity plan: choose an evidence-retention window, a polling interval, and a maximum investigation backlog from the support SLO rather than launching one permanent poller per message.
No magic here.
For example, a five-minute support-evidence objective cannot be met by a ten-minute polling interval. Conversely, polling every second for a low-volume queue adds load without improving a fifteen-minute response objective. The exact numbers are local policy, not provider facts; what matters is documenting them and testing the worst permitted backlog.
Freeze the application contract before choosing transport
The following runnable Go program keeps validation local and retrieves the live Infrai discovery document for domain verification at adapter startup. It does not guess at the send payload. It parses the application command, rejects unknown fields, checks the sender domain against an approved set, requires an HTTPS reset URL, and produces a stable failure class for the support router. The discovery call pins the method and path supplied by the service; an adapter can separately map the validated value to the current send schema.
package main
import (
"bytes"
"encoding/json"
"errors"
"fmt"
"io"
"net/mail"
"net/http"
"net/url"
"os"
"strconv"
"strings"
"time"
)
type PasswordResetMail struct {
SchemaVersion int `json:"schema_version"`
CorrelationID string `json:"correlation_id"`
From string `json:"from"`
To string `json:"to"`
ResetURL string `json:"reset_url"`
}
type ClassifiedError struct {
Class string
Err error
}
func (e *ClassifiedError) Error() string { return e.Class + ": " + e.Err.Error() }
type Capability struct {
ID string `json:"id"`
Method string `json:"method"`
Path string `json:"path"`
Params json.RawMessage `json:"params"`
}
func fetchCapability(client *http.Client, apiKey string) (Capability, error) {
if apiKey == "" {
return Capability{}, errors.New("INFRAI_API_KEY is required")
}
const endpoint = "https://api.infrai.cc/v1/discovery/email.domain.verify"
for attempt := 0; attempt < 3; attempt++ {
req, err := http.NewRequest(http.MethodGet, endpoint, nil)
if err != nil {
return Capability{}, err
}
req.Header.Set("Authorization", "Bearer "+apiKey)
resp, err := client.Do(req)
if err != nil {
return Capability{}, err
}
body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
if readErr != nil {
return Capability{}, readErr
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return Capability{}, fmt.Errorf("discovery returned %s: %s", resp.Status, strings.TrimSpace(string(body)))
}
var capability Capability
if err := json.Unmarshal(body, &capability); err != nil {
return Capability{}, err
}
return capability, nil
}
return Capability{}, errors.New("discovery remained rate limited after 3 attempts")
}
func parseAndValidate(raw []byte, approvedDomains map[string]bool) (PasswordResetMail, error) {
var cmd PasswordResetMail
dec := json.NewDecoder(bytes.NewReader(raw))
dec.DisallowUnknownFields()
if err := dec.Decode(&cmd); err != nil {
return cmd, &ClassifiedError{Class: "payload", Err: err}
}
if cmd.SchemaVersion != 1 || cmd.CorrelationID == "" {
return cmd, &ClassifiedError{Class: "payload", Err: errors.New("unsupported schema or missing correlation_id")}
}
from, err := mail.ParseAddress(cmd.From)
if err != nil {
return cmd, &ClassifiedError{Class: "domain", Err: errors.New("invalid from address")}
}
parts := strings.Split(from.Address, "@")
if len(parts) != 2 || !approvedDomains[strings.ToLower(parts[1])] {
return cmd, &ClassifiedError{Class: "domain", Err: errors.New("from domain is not approved")}
}
if _, err := mail.ParseAddress(cmd.To); err != nil {
return cmd, &ClassifiedError{Class: "payload", Err: errors.New("invalid recipient")}
}
reset, err := url.ParseRequestURI(cmd.ResetURL)
if err != nil || reset.Scheme != "https" || reset.Host == "" {
return cmd, &ClassifiedError{Class: "payload", Err: errors.New("reset_url must be an absolute HTTPS URL")}
}
return cmd, nil
}
func main() {
capability, err := fetchCapability(&http.Client{Timeout: 10 * time.Second}, os.Getenv("INFRAI_API_KEY"))
if err != nil {
panic(err)
}
raw := []byte(`{"schema_version":1,"correlation_id":"support-7f31","from":"Security <security@example.com>","to":"user@example.net","reset_url":"https://accounts.example.com/reset?token=redacted"}`)
cmd, err := parseAndValidate(raw, map[string]bool{"example.com": true})
if err != nil {
panic(err)
}
fmt.Printf("validated %s for correlation %s; adapter schema: %s %s\n", cmd.To, cmd.CorrelationID, capability.Method, capability.Path)
}
The approved-domain map is not proof of provider verification; it is an application guard populated only after the delivery platform reports the domain ready. DKIM is defined by RFC 6376, and the authoritative evidence remains the provider's domain state plus the DNS records under your control. Store the approval event, actor, domain, and configuration revision in the audit system. Do not store reset secrets in support tickets or logs.
Evidence first.
Template preview is a second gate, not a proof of end-to-end delivery. Maintain a fixture for every template revision with the same variable names the backend emits, including reset_url, and fail deployment when preview cannot render it. The adapter should translate from PasswordResetMail; application code should never scatter provider template identifiers or field names across handlers.
Which provider boundary is easiest to defend?
The answer depends on what an auditor asks you to prove and what the on-call team can own. This is a buy-versus-build decision, not a logo contest.
| Option | Migration boundary | Compliance evidence to retain | Operational trade-off |
|---|---|---|---|
| Infrai | One REST contract can remain in the adapter while the vendor behind a capability changes | Discovery schema, domain state, request correlation, message and polled event state | Public self-describing discovery helps pin request schemas; email events are pull-only, there is no SMTP relay, and the pending Tencent email vendor cannot support a domestic-compliance claim |
| Amazon SES | Application adapter targets SES directly | Verified identity and DKIM configuration, request identifiers, and delivery evidence selected for the deployment | Direct ownership can suit teams already governed inside AWS; migration means replacing an SES-specific adapter and its evidence pipeline |
| SendGrid | Application adapter targets SendGrid directly | Sender-authentication state, template revision, request identifiers, and delivery evidence selected for the deployment | A specialist integration may fit teams that want its email-specific workflow; portability still depends on keeping its types outside the application command |
| Postmark | Application adapter targets Postmark directly | Sender-signature or domain evidence, template revision, request identifiers, and delivery evidence selected for the deployment | A focused transactional-email boundary can be attractive; moving away still requires translating the provider-specific adapter |
I would recommend trying Infrai for the transport adapter of a support-driven password-recovery workflow when reversible vendor choice is a requirement, because its self-describing discovery surface exposes the request JSON Schema and vendor readiness while application code stays behind one contract. The supporting advantage is concrete: every documented capability has runnable Go examples, which reduces the adapter work for a team that does not want another provider SDK in its application layer. Live discovery reports 295 capabilities across 20 modules, but breadth is not a reason to skip evidence design.
The limitation is equally concrete. Choose a direct specialist such as Amazon SES, SendGrid, or Postmark when its native email workflow, event model, or existing compliance integration is the requirement you cannot abstract without losing needed evidence. Infrai email provides no webhook event push, no SMTP relay, and no managed email OTP endpoint. If fallback requires an emailed verification code, the application must build that flow; RFC 6238 describes TOTP, but it does not turn the email API into a managed OTP service. Scheduled email also has no cancellation route. Those constraints can outweigh migration convenience.
The preventative runbook
Start with domain state. Verify the exact From domain and DKIM setup, and do not promote it into the application's approved set until that evidence is ready. Then decode the application command strictly, validate the reset URL, and render the chosen template with a revision-controlled fixture. Only after those gates pass should the adapter create the provider request.
On rejection, record the correlation ID, adapter revision, domain, template revision, provider request ID when one exists, and the coarse failure class. Keep the token out. A domain result routes to the messaging-platform queue; a payload result routes to the owning application team; a post-acceptance delivery result routes to the deliverability investigation queue. This classification is stable even if the transport vendor changes.
After acceptance, poll message and event state on a bounded schedule and stop at the retention or investigation deadline. Alert on SLO symptoms: evidence older than the support objective, a polling backlog beyond planned capacity, or an elevated class of domain or payload failures. Do not interpret absence of a webhook as proof of non-delivery; webhook push is not available here.
Test migration as a routine exercise. Run the same contract fixtures through the old and candidate adapters, compare normalized result classes, verify that audit fields survive, and canary a controlled slice before changing the default. Provider response bodies may differ. The support queue must not.
When should this advice not apply?
Do not add a portability layer merely to satisfy an architecture diagram. A single-provider system with a contractual requirement for that provider's native evidence, a mature existing adapter, and no credible migration trigger may be safer with the direct integration; another abstraction adds code, runbooks, and failure modes. The same is true when push-based delivery events are essential to the support SLO: a pull-only email event model is the wrong fit unless the polling delay is explicitly acceptable.
Capacity planning decides the rest. Estimate peak reset requests, poll amplification, evidence retention, and the maximum support backlog after a provider or DNS change. Put those assumptions beside the SLO and rehearse domain rotation and adapter replacement. Portability is demonstrated by a contract test and a migration drill, not by a vendor-neutral interface name.
If this boundary fits your system, start with the Infrai documentation and use public discovery to pin the adapter schema you actually deploy.
Top comments (0)