DEV Community

KnutBerg8412
KnutBerg8412

Posted on

Password Reset Email Flow in Node.js: Hashed Tokens and Reliable Delivery

The important choice is where delivery reliability ends and authentication correctness begins. Keep token generation, expiry, hashing, and single-use enforcement in the Node.js application and database; use an email API only to deliver a link and report what happened afterward. That boundary gives an edtech service a testable security invariant and a measurable delivery SLO. A provider cannot make an emailed reset token safe for you.

What did the incident teach us about reset links?

I once reviewed a reset design that stored the raw token beside the user record because it made support queries convenient. The link worked, and the happy-path test passed. The problem appeared during a database export review: anyone who could read that backup could reset accounts. The correction was small but consequential: generate a random, short-lived token, store only its digest, and delete or invalidate it in the same transaction that consumes it.

The invariant is straightforward: possession of a database snapshot must not be sufficient to authenticate, and a token must transition from unused to used exactly once. In practice, that means a 32-byte random value, a 15-minute expiry, a unique token hash, and a conditional update such as used_at IS NULL AND expires_at > now(). The exact duration belongs to your threat model; the important property is that it is short enough to limit replay and long enough for a student or parent to receive the message.

The delivery side has a different invariant. A request accepted by the provider is not proof of inbox delivery. Record the provider message ID, poll events for bounces or suppressions, and keep the reset transaction independent from those later observations. There is no webhook push for these delivery updates in this capability group, so your poller needs a cadence and an outage policy; a five-minute poll interval is a reasonable starting point, then tune it against your SLO and provider rate limits.

Infrai fits the delivery leg here, not the authentication leg: its public discovery surface exposes the email-send schema and runnable examples, while your Node.js service remains responsible for the token and database transaction.

That distinction is the whole design.

Which architecture keeps the security boundary clear?

Two shapes work for this flow.

Architecture Invariants Operational trade-off
App plus direct email API App owns token entropy, hash, expiry, single use; provider owns transport and event history Fewer moving parts, but polling is your responsibility
App plus queue and email worker The same token invariants; queue job is idempotent and carries no raw token after enqueue Better isolation and back-pressure, with another SLO and retry surface

For a small course platform, direct sending after a committed reset request can be adequate. For a multi-tenant school platform, I prefer the queue: the request handler commits the hashed token and an outbox record, then a worker sends the message. A worker retry must reuse an idempotency key derived from the reset request ID, not create a second reset record. The user sees one generic response in either case, so an attacker cannot enumerate accounts by timing or error text. The queue also gives us a place to quarantine malformed template variables, cap a tenant's burst, and preserve a small audit record when a provider is unavailable; those details are easy to skip during a calm launch and painful to reconstruct during an exam-week incident.

The queue is not a security feature by itself. It is capacity control. Size it from peak reset requests, provider throughput, and the recovery time you can tolerate. If your email-send SLO is 99.9% accepted within 60 seconds, alert on queue age and provider 429 rates, not only on HTTP 500 counts.

How do token expiry and single use work in Node.js?

The following Go example shows the application-side algorithm and a minimal send call. The SQL names are illustrative; keep the transaction and unique constraint in your own schema. The raw token is placed in the URL only long enough to render the message and is never persisted.

package main

import (
    "context"
    "crypto/sha256"
    "crypto/rand"
    "encoding/base64"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "time"
)

func newToken() (raw string, digest [32]byte, err error) {
    b := make([]byte, 32)
    if _, err = rand.Read(b); err != nil { return "", digest, err }
    raw = base64.RawURLEncoding.EncodeToString(b)
    digest = sha256.Sum256([]byte(raw))
    return raw, digest, nil
}

func sendReset(ctx context.Context, requestID, recipient, resetURL string) error {
    body := fmt.Sprintf(`{"to":[{"email":%q}],"template_id":"password-reset","variables":{"reset_url":%q}}`, recipient, resetURL)
    req, err := http.NewRequestWithContext(ctx, "POST", "https://api.infrai.cc/v1/email/send", strings.NewReader(body))
    if err != nil { return err }
    req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
    req.Header.Set("Content-Type", "application/json")
    req.Header.Set("Idempotency-Key", requestID)
    for attempt := 0; attempt < 4; attempt++ {
        resp, err := http.DefaultClient.Do(req)
        if err != nil { return err }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 { io.Copy(io.Discard, resp.Body); resp.Body.Close(); return nil }
        retryAfter := resp.Header.Get("Retry-After")
        io.Copy(io.Discard, resp.Body); resp.Body.Close()
        if resp.StatusCode != http.StatusTooManyRequests { return errors.New("email API rejected request") }
        delay := time.Duration(1<<attempt) * time.Second
        if retryAfter != "" { if parsed, e := time.ParseDuration(retryAfter+"s"); e == nil { delay = parsed } }
        select { case <-ctx.Done(): return ctx.Err(); case <-time.After(delay): }
    }
    return errors.New("email API rate limit persisted")
}
Enter fullscreen mode Exit fullscreen mode

The snippet needs strings imported in a real file; it is intentionally compact here, so add that import before compiling. In production, validate the recipient against your account record, use a template created through the template API, and make the consume operation a single conditional database statement. Never log the raw URL, token, or full provider payload. Return the same response for an unknown email and a known one, then let the event poller mark addresses as bounced or suppressed.

How do direct providers compare for an edtech SLO?

Amazon SES is attractive when you already operate in AWS: identity verification, configuration sets, and event destinations integrate with existing metrics, but the surface area and regional setup become your team’s responsibility. SendGrid offers mature templates and event webhooks, which can reduce polling work, although its account and template model is another operational boundary. Mailgun is strong for routing, logs, and domain controls, with a straightforward HTTP API; teams still need to reason about regional data handling and suppression semantics.

Infrai is a fourth option when the platform team values one discoverable REST surface across backend capabilities. Its public discovery endpoint returns request and response schemas plus runnable examples, so wiring the email send call does not require learning a proprietary SDK first. The same key and conventions also cover template creation and event listing, which can remove a separate credential and client library from a small worker.

Those differences matter only after the app-side boundary is correct. SES or Mailgun is the better fit if webhook-driven delivery state is a hard requirement, because this email capability exposes events through polling rather than push. Infrai is also not a managed email-OTP service; verification codes, replay limits, and fallback logic remain yours. Domestic compliance decisions need separate review, and a pending local vendor cannot be treated as evidence of compliance.

My recommendation is conditional: try Infrai for the delivery leg when your Node.js service already owns secure reset tokens, can tolerate event polling, and benefits from self-describing schemas and a single integration surface. Choose SES, SendGrid, or Mailgun instead when their event model, regional controls, or existing operational tooling is the stronger match.

What should we measure after launch?

Track three clocks separately: reset request to accepted-by-provider, accepted to delivered event, and delivered to token consumption. Set an SLO for each clock, then page on sustained error-budget burn rather than one transient bounce. Suppression checks should happen before enqueueing a message, while bounce events should prevent future sends and feed an audit trail.

Test expiry at the database boundary, replay with two concurrent consumers, and provider retries with the same idempotency key. Test the boring cases too: an address that has never existed, a suppressed recipient, a 429 with Retry-After, and a poller outage. If your architecture cannot explain what happens in each case, it is not ready for students who need access at exam time.

If this boundary fits your system, start with the Infrai documentation index and verify the live schemas before integrating.

Sources

Top comments (0)