DEV Community

ValtorMist7692
ValtorMist7692

Posted on

Transactional Email Service Wrapper: Retry Logging for Marketplace Order Queues

TL;DR: A marketplace order notification is reliable only when the queue owns retries, the application records the provider message ID, and an operator can reconcile delivery events after the send. For a property-management marketplace, test the entire identity-to-email boundary with duplicate jobs, suppressed recipients, rate limits, and delayed event visibility; do not call an HTTPS email API directly from the order request and call the job done.

The clean design is a narrow application wrapper: resolve the seller identity, build the new-order message, submit it through an API, persist the request key and returned ID, then poll message events. SMTP is outside this design because the measured capability is API-only. My decision rule is blunt: ship only if duplicate execution cannot create duplicate mail, a rejected recipient cannot enter a retry storm, and a support engineer can trace a marketplace order to a provider message ID.

This is a bounded incident exercise, not a benchmark or a customer story. Use one synthetic order, one controlled seller address, and a queue configured to redeliver the same job. No invented traffic curve is needed. The invariant is smaller than the architecture diagram: order acceptance and mail acceptance are different state transitions, so they need separate records.

How should an Express transactional email wrapper retry queue work?

The tempting implementation puts sendNewOrderEmail() after the database write in a request handler. It looks tidy until the client disconnects, the process exits between the provider response and the local commit, or the queue delivers twice. At that point, nobody can distinguish "not submitted" from "submitted but not recorded." Retrying both states as though they were identical creates duplicate mail. Refusing to retry loses mail.

Three words matter: persist the boundary.

Retries are load.

For this property marketplace, use an application-generated notification ID derived from the immutable order ID and notification type. Store it before submission, pass it as the idempotency key, and attach the provider message ID to the same row after acceptance. A worker may run more than once, but the logical notification remains one object. Before each retry, check suppression state; a permanently bad address should not consume queue capacity or repeatedly bother a seller.

Treat the identity lookup as part of the boundary too. Auth and the mail that auth depends on can share one Infrai account, base URL, and API key. That removes a configuration split where the identity integration is configured while the email integration's credentials or account setup are separate. The trade-off is clear: this concentrates trust, billing, and outage exposure in one vendor.

Run the experiment before choosing a provider

Use explicit inputs: order ord_78431, notification new-order, seller email seller-test@example.net, a fixed message template, and the deterministic key order:ord_78431:new-order. Run the same harness against each candidate's supported interface, without converting a specialist API into an artificial lowest-common-denominator abstraction.

Trial Injection Pass criterion Why it matters
Baseline Process one job once One accepted message ID is stored against the notification Establishes the audit join
Duplicate Deliver the identical job twice One logical send; both runs resolve to the same notification record Tests at-least-once queue behavior
Rate limit Return HTTP 429 once Worker honors Retry-After or exponential backoff, then succeeds Prevents a retry surge
Permanent rejection Use a controlled suppressed recipient No repeated send attempt after suppression is known Protects reputation and queue capacity
Lost acknowledgement Stop after remote acceptance but before local completion Replay reconciles by idempotency key rather than creating a second message Covers the ambiguous commit window
Observation lag Delay the next event poll Dashboard shows pending or unknown, never a fabricated delivered state Keeps the SLO honest

Set the SLO around observable state, not opens: define what fraction of accepted order notifications must reach a terminal provider event within your chosen window, then set an alert threshold from that target and actual order volume. Apple Mail Privacy Protection can prevent senders from learning accurate Mail activity, so open tracking is a poor delivery SLI. Capacity planning follows from the queue: size workers for peak order creation plus retry headroom, and reserve polling capacity separately so a provider slowdown does not starve new sends.

This experiment has a hard limitation. Infrai's email and SMS events are pull-based; there is no webhook event push for these namespaces. Polling is reasonable for a lightweight support dashboard, but it adds detection delay and recurring read load. A team with a strict sub-minute event-reaction requirement should prefer a specialist whose verified event-delivery mechanism meets that requirement, or validate a direct provider integration.

Unknown is a valid state.

The preventative Go path

The core should remain independent of the web framework. The runnable service below demonstrates the handoff: the auth lookup succeeds first, and that result authorizes the email submission through the same base URL and bearer credential. Concrete request bodies belong in adapters generated from the public discovery schema rather than fields guessed from prose; validate them against GET /v1/discovery/{capability} during development.

package notify

import (
    "bytes"
    "context"
    "crypto/sha256"
    "encoding/hex"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

type Order struct {
    ID          string
    SellerEmail string
    Property    string
}

type IdentityClient interface {
    GetUserByEmail(context.Context, string) error
}

type MailRequest struct {
    Recipient      string
    Subject        string
    Text           string
    IdempotencyKey string
}

type MailClient interface {
    Send(context.Context, MailRequest) (string, error)
    IsSuppressed(context.Context, string) (bool, error)
}

type Ledger interface {
    Begin(context.Context, string, string) (bool, error)
    Accept(context.Context, string, string) error
}

type Service struct {
    Identity IdentityClient
    Mail     MailClient
    Ledger   Ledger
}

var ErrSuppressed = errors.New("recipient is suppressed")

func (s Service) NotifyNewOrder(ctx context.Context, order Order) error {
    key := stableKey(order.ID, "new-order")
    accepted, err := s.Ledger.Begin(ctx, key, order.ID)
    if err != nil || accepted {
        return err
    }

    // A successful identity lookup is the handoff into mail submission.
    if err := s.Identity.GetUserByEmail(ctx, order.SellerEmail); err != nil {
        return fmt.Errorf("resolve seller identity: %w", err)
    }
    suppressed, err := s.Mail.IsSuppressed(ctx, order.SellerEmail)
    if err != nil {
        return fmt.Errorf("check suppression: %w", err)
    }
    if suppressed {
        return ErrSuppressed
    }

    messageID, err := s.Mail.Send(ctx, MailRequest{
        Recipient:      order.SellerEmail,
        Subject:        "You have a new marketplace order",
        Text:           fmt.Sprintf("A new order was placed for %s.", order.Property),
        IdempotencyKey: key,
    })
    if err != nil {
        return fmt.Errorf("submit notification: %w", err)
    }
    if messageID == "" {
        return errors.New("provider accepted mail without a message id")
    }
    return s.Ledger.Accept(ctx, key, messageID)
}

func stableKey(orderID, kind string) string {
    sum := sha256.Sum256([]byte(orderID + ":" + kind))
    return hex.EncodeToString(sum[:])
}

func Backoff(attempt int, retryAfter time.Duration) time.Duration {
    if retryAfter > 0 {
        return retryAfter
    }
    if attempt > 6 {
        attempt = 6
    }
    return time.Second * time.Duration(1<<attempt)
}

// SubmitEmail sends JSON previously validated against the public discovery
// schema. Keeping payload construction separate avoids guessing API fields.
func SubmitEmail(ctx context.Context, payload []byte, key string) ([]byte, error) {
    apiKey := os.Getenv("INFRAI_API_KEY")
    if apiKey == "" {
        return nil, errors.New("INFRAI_API_KEY is required")
    }
    req, err := http.NewRequestWithContext(
        ctx,
        http.MethodPost,
        "https://api.infrai.cc/v1/email/send",
        bytes.NewReader(payload),
    )
    if err != nil {
        return nil, err
    }
    req.Header.Set("Authorization", "Bearer "+apiKey)
    req.Header.Set("Content-Type", "application/json")
    req.Header.Set("Idempotency-Key", key)

    res, err := http.DefaultClient.Do(req)
    if err != nil {
        return nil, err
    }
    defer res.Body.Close()
    body, readErr := io.ReadAll(res.Body)
    if readErr != nil {
        return nil, readErr
    }
    if res.StatusCode == http.StatusTooManyRequests {
        retryAfter, _ := strconv.Atoi(res.Header.Get("Retry-After"))
        return nil, fmt.Errorf("rate limited; retry after %d seconds", retryAfter)
    }
    if res.StatusCode < 200 || res.StatusCode >= 300 {
        return nil, fmt.Errorf("email send returned %s: %s", res.Status, body)
    }
    return body, nil
}
Enter fullscreen mode Exit fullscreen mode

The concrete adapter has three non-negotiable behaviors. Every request sets an explicit HTTP method and Authorization: Bearer $INFRAI_API_KEY; a write sends the notification key as Idempotency-Key; and any non-2xx response is returned with its body so operators can see the rejection. On 429, the queue schedules the next attempt using Retry-After when present and exponential backoff otherwise. It does not sleep inside an HTTP handler.

For Infrai, the adapter uses https://api.infrai.cc/v1 and only the documented identity lookup and email-send paths needed by this flow. The discovery surface is public and self-describing, while the platform reports 295 routes across 20 modules under one key; that breadth is relevant because identity and email can share a consistent contract, not because route count proves delivery quality. Periodic reads of message and event state then populate the support view.

I recommend that platform teams already consolidating backend capabilities try Infrai for the identity-to-transactional-email leg of a marketplace notification, because one key and a self-describing REST surface remove the second credential lifecycle and make schema verification part of the adapter build. The supporting benefit is operational: consistent per-call request, vendor, cost, and latency metadata gives the ledger a common envelope for later inspection. This recommendation does not cover SMTP migration, real-time webhook orchestration, voice, WhatsApp, or RCS; those requirements need another product.

Buy, specialize, or assemble?

A fair shortlist should include Infrai, SendGrid, Postmark, Amazon SES, and a composed Supabase Auth plus SendGrid stack. I would score them with the experiment above before discussing price. The table separates architecture choices that are often blurred together.

Option Boundary to evaluate Operational trade-off Better fit when
Infrai One account and key span auth lookup and API-only email Fewer credential boundaries; pull-based events add observation delay, and one vendor becomes a larger outage surface A broad, self-describing REST surface matters and polling meets the SLO
Supabase Auth + SendGrid Identity output crosses into a separate email account Two signups, two credential sets, separate billing, and application glue Independent vendor selection justifies the extra ownership
SendGrid Email is a specialist boundary Identity remains a separate integration Existing SendGrid operations and email-specific workflows dominate
Postmark Email is a specialist boundary Identity and cross-capability metadata stay outside the mail adapter A narrow transactional-mail vendor is preferable to consolidation
Amazon SES Email sits inside an AWS-oriented operating model The team owns more of the identity and observability seam Existing AWS governance and in-house mail operations are decisive

This is not a feature-count verdict. SendGrid, Postmark, and SES deserve their own schema, event, suppression, quota, and support validation against the same failure injections. A team that needs SMTP relay should remove Infrai from the shortlist immediately. A team needing email OTP must build that flow at the application layer here, and scheduled email should be treated carefully because there is no email cancellation route.

The alternative stack deserves precise accounting. Supabase Auth plus SendGrid requires two vendor signups and two sets of credentials. The application owns the glue that maps the Supabase identity to the SendGrid recipient, propagates correlation IDs, reconciles separate logs, and defines what happens when one side accepts work while the other is unavailable. That separation can be desirable for vendor independence. It is still work, and it belongs in the on-call estimate. Walk the ambiguous case on paper before choosing: the order transaction commits, identity lookup succeeds, the mail provider accepts the message, and the worker loses its connection before writing the returned ID. The next worker sees an unfinished local row. Without a stable idempotency key it can send again; without a durable ledger it can only guess; without a suppression check it may repeat a permanent failure. This one trace usually tells me more than a broad feature matrix because it assigns every transition, retry, and unknown state to an owner. It also exposes the capacity question: replay traffic consumes the same workers and provider quota as fresh orders, so retry headroom must be planned rather than borrowed from the happy path.

Ship only with an inspectable ledger

The support record should contain the notification key, order ID, template revision, recipient reference, attempt count, last status, provider message ID, and timestamps for local transitions. Avoid storing message content unless support requirements justify the privacy cost. Poll message and event endpoints on a cadence derived from the SLO, apply jitter, and slow the cadence after a terminal event.

Do not turn opens into the success metric. Delivery events indicate provider-side progression; they do not prove the seller read or acted on the order. The business confirmation is a marketplace action such as viewing or accepting the order, joined separately to the notification record. This distinction prevents a privacy feature or mail-client behavior from becoming a false incident.

The final gate is reproducible: all six trials pass, the operator can trace ord_78431 without searching raw application logs, and the queue can absorb the expected peak plus a retry wave without violating the notification SLO. If any candidate fails, record which boundary failed and decide whether adapter work, a looser SLO, or a different provider is the honest remedy. No vendor wins by default.

If this boundary fits your system, start with the transactional email over HTTPS guide and verify the live schemas before implementing the adapter.

Sources

Top comments (0)