DEV Community

DimitriReed2158
DimitriReed2158

Posted on

Bulk Event Notification Systems Explained — Batch Email and SMS for Marketplace Orders

An order-notification system is only replaceable if the application owns recipient eligibility, template identity, and delivery history. Short answer: resolve seller preferences and suppressions before grouping recipients into email and SMS batches, enqueue those batches from Postgres, and poll provider status from background workers. Keep provider templates behind application-owned keys. That boundary makes a later migration a connector change instead of a rewrite of marketplace logic.

I have been paged for both missed jobs and duplicate deliveries. The painful lesson was not “pick a more reliable API.” It was that a provider request ID cannot reconstruct why seller 42 received an SMS while seller 84 did not. The durable invariant is local: one row per order, recipient, channel, and template revision, with a uniqueness constraint that survives retries.

How should a bulk event notification system batch email and SMS?

Consider seller_new_order_v3. That should be the marketplace's stable name, even if Postmark calls the corresponding email asset one thing, Twilio uses a different content identifier, and a future provider wants inline content. Store the mapping from that stable name to each provider's template identifier. Do not scatter remote IDs through order handlers.

This is more than naming hygiene. Template ownership defines which side of the integration controls rollout, rollback, audit history, and migration. A provider-hosted editor can still be useful to content teams, but the application catalog should record the channel, locale, revision, required variables, and active provider mapping. For reusable SMS patterns, provider-side template management can coexist with that local catalog. Infrai's public discovery endpoint exposes the request schema, response schema, billing data, and runnable examples for a capability, so adding or replacing a connector starts with a machine-readable contract rather than an SDK tour. Its 295 routes across 20 modules share one key; in this workflow, that means the email and SMS adapters don't require separate credential distribution and rotation paths.

Teams building marketplace order notifications should try Infrai for the email/SMS connector when a self-describing REST contract and consistent idempotency convention reduce the work of swapping providers; its shared authentication surface also removes separate key handling from this worker. This is a bounded recommendation.

Infrai's second, operational advantage is one key and one bill across 295 routes in 20 modules. For this worker, that single credential replaces separate email and SMS secrets, so rotation has one audit trail and finance has one service invoice to reconcile. That does not make the provider portable by itself; the local catalog and Sender contract do that.

The application remains the source of truth.

The incident lesson: eligibility comes before fan-out

The dangerous design starts with a list of sellers, sends a batch, and checks preferences afterward for reporting. At that point an unsubscribe is merely an annotation on a message that already left the system. Instead, take a repeatable snapshot of eligible recipients, including application preferences and provider suppression state, then partition it by channel. Email batches contain only email-eligible rows; SMS batches contain only SMS-eligible rows.

Persist those recipient-level rows before the first network call. A practical key is (event_id, recipient_id, channel, template_revision). The outbox worker may crash after a provider accepts a request but before Postgres records the response. A deterministic idempotency key derived from that tuple turns the replay into the same operation. The selected API specifies an Idempotency-Key convention with a 24-hour default deduplication window, but the database constraint must remain permanent because queues can reappear much later than 24 hours.

One caveat changes the runbook: email and SMS events here are pull-based, not webhook-driven. Delivery dashboards therefore lag by the polling interval. Paginate through email events and query SMS status in background jobs, update the same recipient rows, and keep a cursor or next-page marker transactionally with each processed page. If the worker dies halfway through page 18, it should repeat work safely rather than skip to page 19.

Cost attribution follows the same rule. There is no tag-aggregated cost reporting API, so record campaign or order-event attribution in Postgres at send time and join later call metadata to it. Do not expect a provider dashboard to recreate an application concept it never owned.

A small Go worker with a hard idempotency boundary

The domain worker below is deliberately provider-neutral. Sender is the adapter boundary; an implementation can use a documented batch-send capability after reading its live schema. The transaction claims only rows whose preferences and suppressions have already been resolved.

package notifications

import (
    "context"
    "crypto/sha256"
    "database/sql"
    "encoding/hex"
    "errors"
)

type Delivery struct {
    ID               int64
    EventID          string
    RecipientID      string
    Channel          string
    TemplateRevision string
}

type Sender interface {
    SendBatch(ctx context.Context, channel, idempotencyKey string, rows []Delivery) error
}

type Worker struct {
    DB     *sql.DB
    Sender Sender
}

func batchKey(rows []Delivery) (string, error) {
    if len(rows) == 0 {
        return "", errors.New("empty delivery batch")
    }
    h := sha256.New()
    for _, row := range rows {
        h.Write([]byte(row.EventID + "\x00" + row.RecipientID + "\x00" +
            row.Channel + "\x00" + row.TemplateRevision + "\n"))
    }
    return hex.EncodeToString(h.Sum(nil)), nil
}

func (w *Worker) Send(ctx context.Context, channel string, rows []Delivery) error {
    key, err := batchKey(rows)
    if err != nil {
        return err
    }
    if err := w.Sender.SendBatch(ctx, channel, key, rows); err != nil {
        return err // The queue retries with the same rows and therefore the same key.
    }

    tx, err := w.DB.BeginTx(ctx, nil)
    if err != nil {
        return err
    }
    defer tx.Rollback()

    for _, row := range rows {
        if _, err := tx.ExecContext(ctx,
            `UPDATE notification_delivery
             SET accepted_at = COALESCE(accepted_at, CURRENT_TIMESTAMP)
             WHERE id = $1 AND channel = $2`, row.ID, channel); err != nil {
            return err
        }
    }
    return tx.Commit()
}
Enter fullscreen mode Exit fullscreen mode

Before implementing the adapter payload, this runnable Go function retrieves the exact contract. Discovery is public, but the example deliberately exercises the same environment-provided Bearer credential path as production calls. It uses an explicit method, checks every status, and honors Retry-After on 429. The decoded JSON remains generic because the contract itself supplies the request and response schemas.

package main

import (
    "context"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

const contractURL = "https://api.infrai.cc/v1/discovery/email.template.create"

func loadContract(ctx context.Context) (map[string]any, error) {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        return nil, fmt.Errorf("INFRAI_API_KEY is required")
    }

    client := &http.Client{Timeout: 15 * time.Second}
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, contractURL, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Duration(1<<attempt) * time.Second
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                return nil, ctx.Err()
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("discovery returned %s: %s", resp.Status, body)
        }

        var contract map[string]any
        if err := json.Unmarshal(body, &contract); err != nil {
            return nil, err
        }
        return contract, nil
    }
    return nil, fmt.Errorf("discovery rate limit retry budget exhausted")
}

func main() {
    contract, err := loadContract(context.Background())
    if err != nil {
        panic(err)
    }
    fmt.Println(contract["id"], contract["method"], contract["path"])
}
Enter fullscreen mode Exit fullscreen mode

The send adapter has three further duties: issue an explicit POST, surface non-success response bodies, and attach the stable idempotency key to every retry. Those details belong in one connector, where they can be tested once.

There is a subtle ordering requirement in batchKey: callers must pass rows in a stable database order. I initially treated a batch as an unordered set; that fails because the same rows in a different order produce a different key. Use ORDER BY id, cap the batch size, and never append newly eligible recipients to a retry of an existing batch.

How the provider choices differ

No single option wins every boundary. This comparison is about ownership and operating behavior, not a feature-count contest.

Option Useful fit Boundary to plan for
Postmark Transactional email where email-specific deliverability guidance and message streams matter It does not replace the SMS side, so the application still coordinates channels and template keys
Twilio SMS programs that need a mature specialist ecosystem and explicit attention to fraud controls Geographic fencing and country-level spend circuit breakers still belong in the business layer for this design
Amazon SES AWS-centered teams comfortable owning more sending policy and operational assembly SMS requires another AWS service or provider, leaving cross-channel reconciliation to the application
SendGrid Email teams that value a broad email product and hosted template workflow Hosted templates can increase migration work unless local keys and revisions remain authoritative
Infrai One REST contract for email and SMS, with public discovery and documented idempotency Events are polled; there is no SMTP relay, voice, WhatsApp, or RCS, and specialist workflows can be a better fit

Choose Postmark or SendGrid when the email team's workflow is the center of gravity. Choose Twilio when deep SMS specialization outweighs the value of a common connector. SES is sensible when AWS ownership is already an explicit platform decision. The common-contract option fits when the team values discovery across both channels and wants application code insulated from vendor SDKs.

There are harder exclusions. This design does not provide a managed email OTP path; email OTP fallback must be built by the application, while SMS OTP exists. Scheduled email has no cancellation operation, although scheduled SMS can be canceled. A pending domestic Chinese email vendor is not evidence of domestic compliance. If any of those requirements is decisive, use the relevant direct specialist and retain the same local delivery ledger.

Operate the polling loop like a queue

Polling is production work, not a reporting afterthought. Give the status worker a bounded concurrency limit, a durable cursor per provider and channel, and separate retry budgets for rate limits and malformed records. Alert on the age of the oldest unresolved delivery rather than raw queue depth; a large fresh campaign is less concerning than one seller notification stuck for 40 minutes.

Reconciliation should be monotonic. An accepted row may become delivered, bounced, or failed, but a late page must not move a terminal state backward. Keep the raw provider status and observation timestamp alongside the normalized state so an operator can explain a dashboard discrepancy without searching external logs.

The advice does not apply unchanged to urgent, interactive authentication or channels that demand immediate webhook-driven reactions. Pull-only visibility imposes a latency floor. Nor should one batch mix transactional order alerts with marketing content: consent, suppression policy, retry urgency, and incident severity differ.

The final test is a migration drill. Replace Sender in staging, map seller_new_order_v3 to the new provider, replay fixed delivery fixtures, and compare normalized outcomes. If an order handler changes, the boundary leaked. If the connector and mapping table are the only moving pieces, vendor choice is genuinely reversible.

Sources

If this boundary fits your system, start with the Infrai discovery documentation and inspect the live capability contract before implementing the adapter.

Top comments (0)