TL;DR: For a fintech password-reset message with a short expiry, choose between an API-first email service and an SMTP relay by counting the integration work required to make delivery correct, observable, and replaceable. An HTTPS API often exposes request identifiers and structured errors more directly; SMTP remains a stable, broadly understood protocol and can fit an existing mail pipeline. Neither transport proves that a person received the message. Put the reset state, expiry, idempotency, retry policy, and audit trail in your application boundary, then compare providers on how little translation that boundary requires.
I have been paged by missed jobs and duplicate deliveries. Those pages taught the same lesson from opposite directions: a send call is not the unit of correctness. The durable operation is one reset request moving through a state machine, possibly across retries, while one short-lived credential remains valid. For this workload, a quick provider SDK demo is a weak integration test. The useful test is what happens after the process dies between accepting the reset request and recording the provider response.
Should a startup use a transactional email API for welcome emails?
An API-first service and an SMTP relay can both hand mail into the delivery network. The meaningful difference is the adapter your team must own. With an HTTP API, the application commonly sends a structured request and receives a structured status plus a request identifier. With SMTP, the application conducts a protocol exchange and receives SMTP reply codes. A relay accepting a message means it accepted responsibility for the next step; it does not mean the destination mailbox displayed it. An HTTP success has the same limitation unless the service contract explicitly says otherwise.
That boundary matters more for a password reset than for a welcome message. A delayed welcome is annoying. A delayed reset can arrive after the credential has expired, and a second message can make the user unsure which link is current. Keep those streams separate in queues, dashboards, and rate controls even if they share a transport. Security mail should not wait behind a promotional burst.
For a startup already operating SMTP libraries, relay credentials, bounce processing, and trace correlation, changing transports may add work without removing much risk. A team starting from a small web service may find a JSON API easier because authentication, error parsing, and provider identifiers fit its existing HTTP tooling. This is an integration-effort decision, not a protocol morality play.
| Concern | API-first adapter | SMTP relay adapter | Application must still own |
|---|---|---|---|
| Submission result | HTTP status and structured body | SMTP reply code and enhanced status when available | Retry classification |
| Correlation | Provider request or message identifier | Message-ID plus relay response and local trace | Mapping to reset request |
| Authentication | Usually an API credential | Usually relay credentials and transport security | Secret rotation and least privilege |
| Portability | Provider schema must be translated | Message format is standardized, relay behavior still varies | A narrow internal interface |
| Delivery evidence | Events may be available through callbacks or polling | Bounce and delivery status handling depends on the mail path | State transitions and retention |
Count the glue.
Include credential rotation, timeout behavior, webhook or bounce authentication, event normalization, local test fixtures, dashboards, and the deletion path for sensitive data. The shortest send function often hides the longest runbook.
The invariant from duplicate and missed jobs
A reset endpoint should reveal as little as practical about whether an account exists. After accepting an eligible request, create the reset record and an outbox item in the same database transaction. A worker claims that item, renders the message, submits it through a narrow mail interface, and records the outcome. This transactional-outbox shape closes the dangerous gap between committing security state and enqueueing work in a separate system.
Retries are normal.
Design for them. Give each logical reset an idempotency key generated by your system, and keep it stable across transport retries. If an API supports an idempotency facility, the adapter can pass that key through, but local deduplication remains necessary because remote semantics differ. SMTP has no general request-idempotency command, so the sender must prevent concurrent claims and avoid regenerating a fresh message for the same outbox row.
The token needs a separate rule: store a verifier rather than a reusable plaintext credential, enforce a short configured expiry, invalidate it after successful use, and make the action single-use. The email carries the opaque value needed for verification, while logs carry only safe correlation identifiers. Do not put the token in provider metadata, metric labels, traces, or support screenshots. NIST's digital identity guidance is useful here because it treats recovery as part of the authenticator lifecycle, not as ordinary marketing messaging.
One subtle failure deserves a runbook entry. If every retry creates a new reset token, an uncertain submission result can produce multiple valid messages. If every new user request invalidates the prior token immediately, delivery reordering can make the newest-looking email contain an invalid link.
Ambiguity is the bug.
Pick a policy deliberately. One defensible policy is to keep one active challenge for a bounded window and resend a message referring to that same challenge, while rate-limiting both the account and request origin. The exact window is a product and threat-model decision, not a mail-provider default.
A preventative Go boundary
The interface below keeps provider-specific payloads away from the reset workflow. The worker assigns the durable identifier; the adapter reports whether an error is safe to retry and preserves a remote identifier only for correlation. In production, Claim must use a database lease or equivalent atomic transition so two workers cannot submit the same row concurrently.
package resetmail
import (
"context"
"errors"
"time"
)
type Message struct {
OperationID string
Recipient string
Subject string
TextBody string
HTMLBody string
}
type Receipt struct {
RemoteID string
}
type Sender interface {
Send(ctx context.Context, msg Message) (Receipt, error)
}
type SendError interface {
error
Retryable() bool
}
type Job struct {
ID string
Recipient string
ResetURL string
ExpiresAt time.Time
Attempts int
}
type Store interface {
Claim(ctx context.Context, jobID string, lease time.Duration) (Job, error)
MarkSent(ctx context.Context, jobID, remoteID string, at time.Time) error
RetryLater(ctx context.Context, jobID string, at time.Time, reason string) error
MarkFailed(ctx context.Context, jobID, reason string) error
}
func Deliver(ctx context.Context, store Store, sender Sender, jobID string, now time.Time) error {
job, err := store.Claim(ctx, jobID, 30*time.Second)
if err != nil {
return err
}
// Expired credentials are not useful messages; close the job without submitting it.
if !now.Before(job.ExpiresAt) {
return store.MarkFailed(ctx, job.ID, "reset expired before submission")
}
msg := Message{
OperationID: job.ID,
Recipient: job.Recipient,
Subject: "Reset your password",
TextBody: "Use this link before it expires: " + job.ResetURL,
HTMLBody: "", // Render with an escaping template in the real adapter.
}
receipt, err := sender.Send(ctx, msg)
if err == nil {
return store.MarkSent(ctx, job.ID, receipt.RemoteID, now)
}
var sendErr SendError
if errors.As(err, &sendErr) && sendErr.Retryable() && job.Attempts < 5 {
delay := time.Duration(1<<job.Attempts) * time.Second
next := now.Add(delay)
if next.Before(job.ExpiresAt) {
return store.RetryLater(ctx, job.ID, next, "temporary submission failure")
}
}
return store.MarkFailed(ctx, job.ID, "submission failed")
}
The numbers here are example policy, not universal recommendations: a 30-second worker lease, no more than five attempts, and exponential delays bounded by the credential expiry. Make them configuration with tests around the boundary. In particular, ensure the lease exceeds the sender timeout or is renewed safely. Otherwise a second worker can claim the job while the first call is still in flight.
There is no provider route in the workflow. An HTTP adapter can map Message to a vendor schema and classify timeouts, throttling, and permanent validation errors. An SMTP adapter can build the MIME message, issue the protocol commands, and classify transient 4xx replies separately from permanent 5xx replies. Both adapters must apply a hard timeout and return errors without leaking the reset URL.
How should the integration be tested and operated?
Start below the happy path. Kill the worker after submission but before MarkSent, run two workers against one job, delay delivery beyond expiry, and return malformed or unauthenticated event callbacks. Verify that retries retain the same operation ID, the queue does not grow without an alert, and no reset credential appears in logs. A fake Sender is enough for state-machine tests; use a controlled mailbox and the real adapter in preproduction to inspect headers and rendering.
The minimum operational signals are accepted reset requests, outbox age, attempts, submission latency, retry classifications, terminal failures, and expiry-before-send count. Keep dimensions bounded. A recipient address, token, URL, or provider response body must never become a metric label. Trace with the internal operation ID and map it to the remote identifier in access-controlled storage.
Email authentication is part of the deployment, not an item to postpone until volume grows. Google's sender guidance calls for authentication and valid message formatting, with additional requirements for bulk senders. Configure the sending domain, SPF, DKIM, DMARC policy, forward and reverse DNS where applicable, and TLS according to the receivers and infrastructure in scope. Test alignment after DNS changes. Separate transactional traffic operationally so its reputation and queue are visible, but do not assume separation cures poor consent or unsafe sending practices.
On call, the first question is where progress stopped: application acceptance, outbox claim, transport submission, downstream delivery event, or user redemption. A single “email failed” counter cannot answer it. The runbook should also say what operators may replay, what has already expired, and how to avoid manually sending a credential outside the audited path.
When does this advice not apply?
A synchronous internal tool with no durable user account and no security-bearing link may not justify an outbox. A mature mail platform may already supply queueing, deduplication, standardized events, and audited templates behind an internal SMTP relay; adding a direct API adapter could duplicate controls. Conversely, a small service with no mail operations may reduce integration work through an API-first adapter, provided it still owns reset state and failure handling.
Do not use email as proof that the requester controls an active authenticated session, and do not let transport acceptance restore account access. The reset completes only when the server verifies the presented credential, confirms it is unexpired and unused, applies the account change atomically, and invalidates the credential.
The decision rule is compact: choose the transport whose adapter leaves the fewest unowned failure states in your system. Preserve a narrow sender interface, a durable outbox, and provider-independent security state. Then a later transport change is an adapter migration rather than a rewrite of account recovery.
Top comments (0)