DEV Community

thomasmoore5082
thomasmoore5082

Posted on

Media Report Event Notifications: A Transactional Email and SMS API Guide

The transactional email and SMS API page says a generated media report event is ready, yet the recipient still has no useful notification. On-call can see that email submission succeeded, an SMS fallback is eligible, and the current template revision changed shortly before the alert; what the page can't establish is whether delivery observation is late, the attachment contract is wrong, or the fallback clock was started from the wrong event.

TL;DR: keep the report attachment contract and the fallback state machine in application code, then let a provider own presentation only when the team changing copy can also test and roll back that template. For this workflow, the best transactional email and SMS API is the one that makes template revisions, submission state, and delivery evidence distinguishable inside the SLO. A unified REST option is practical when polling is acceptable; require a different event boundary when fallback must react to webhooks, or when SMTP relay, managed email OTP, voice, WhatsApp, or RCS is mandatory.

Do not select on the advertised message rate first. The recurring cost of polling, geo-fencing, country-level SMS spend guards, contract testing, and on-call ambiguity can dominate a small difference in provider charges, and none of those responsibilities disappears because a dashboard says “sent.”

How should a transactional email and SMS API handle event notifications?

The earlier signal is not an open. Apple Mail Privacy Protection can download remote content without a deliberate reader action, so an open event cannot be the sole trigger for an urgent SMS. DMARC answers a separate question about authentication and receiver policy; it does not prove that a person received or read the report.

Page on stale delivery knowledge while actionable work is pending. Record report_ready_at, email_submitted_at, delivery_observed_at, fallback_due_at, sms_submitted_at, and template_revision under one stable notification ID. Then alert when a pending notification's observation age consumes the part of the objective reserved for the poller. An old poll timestamp with no pending notifications is background noise. The same age while hundreds of reports await a decision is a capacity signal.

Short fields matter here. One generic updated_at does not.

Both email and SMS events in the unified capability are pull-based rather than webhook-driven, which fixes the shape of the capacity calculation. If P notifications are pending, the verified page size is B, and the target observation interval is T, the poller must finish at least ceil(P/B) reads within T, plus retry headroom. Measure B, request duration, and rate limits in the deployed environment; no supplied evidence supports a universal concurrency value or latency promise. That's a hard limitation, not an implementation detail.

This is also where the false-positive bill starts. A threshold below the normal poll-cycle envelope pages on expected work, while a threshold above the fallback budget quietly spends the user's objective. Reserve time for the SMS attempt first, use the remaining budget for observation, and make pending volume part of the alert condition.

Make template publication an observable production change

A media report email is an interface. The subject, attachment filename and media type, localization inputs, suppression behavior, legal footer, and fallback copy form a contract between the report generator and the communications system. Moving HTML into a provider does not move that contract; it creates another release surface.

Use a release gate based on rollback authority:

Template boundary Release gate Evidence on-call needs
Application-owned Code review and a report fixture pass together Commit, fixture result, and deploy revision
Provider-owned A named publisher promotes an immutable revision Revision, approver, preview result, and rollback target
Hybrid Both sides validate one versioned field contract Contract version and results from both release paths

For generated attachments, application ownership is the conservative default because the producer already controls the file, recipient, locale, and data classification. Provider ownership makes sense when a communications team has genuine release authority, but only if the template revision is emitted with every notification and the pager holder can restore a known revision without guessing which copy changed.

The key instrumentation change is therefore small: treat template_revision as a deployment dimension, not decorative metadata. Compare failure and stale-observation counts before and after a revision, but do not alert on the revision itself. A rollback drill should cover one representative attachment, one localization change, one suppressed recipient, and one forced SMS fallback. Four cases are enough to expose an unclear owner; they are not a claim of exhaustive testing.

Test provider boundaries with one report fixture

Resend, Postmark, SendGrid, Twilio, and Bird (formerly MessageBird) are real candidates, but they should not be assigned capabilities by reputation. Run the same fixture through each shortlisted boundary and consult the current vendor documentation for the features you require. This keeps the comparison fair as products change.

Candidate Role worth evaluating Boundary question that decides the test
Resend Focused report-email submission Will a separate SMS system leave revision correlation and fallback ownership with a team prepared to operate it?
Postmark Transactional report email Can its delivery evidence be joined to the chosen SMS state inside the notification objective?
SendGrid Email templates and email operations Who promotes and rolls back the template while preserving the attachment contract?
Twilio Urgent SMS delivery Can the application correlate SMS status with the email provider without creating a second paging boundary?
Bird (MessageBird) Broader messaging evaluation Do the currently documented channel and event controls match the release and response process?
Unified REST API One application boundary for email and SMS Can the product accept polling and application-owned geographic safeguards?

The table deliberately identifies tests, not winners. A channel-native pairing is often easier to reason about when separate teams already own email and SMS. It also means operating correlation, credentials, and fallback behavior across two systems. A broader provider can reduce those handoffs, but breadth is irrelevant if its event model cannot meet the response budget.

Infrai is one credible unified option for the narrow case described here because it provides one key, one wallet, and one bill across one REST API, with no SDK to install. Its public discovery surface is self-describing: a capability description provides request and response schemas, billing information, and runnable examples, so adding a capability begins by reading one contract rather than adopting another language SDK. The live discovery snapshot reports 295 routes across 20 modules and examples in 10 languages. For this workflow, the shared credential and bill mean the report worker and poller don't need separate credential rotation or invoice correlation just because they run in different languages; plain HTTP also keeps either runtime from owning another vendor client. Idempotency is specified on 171 of 294 documented capabilities, with a 24-hour default deduplication window where the convention applies.

The trade-off is material. Infrai is not a fit when webhook-speed fallback, SMTP relay, managed email OTP, or voice, WhatsApp, and RCS are requirements; choose a channel-native provider or a broader messaging candidate whose current documentation proves the missing behavior. Its email and SMS delivery events require polling. Email scheduling has no cancellation operation, tag-aggregated cost reporting is unavailable through an API, and domestic email vendor support is pending, so it cannot support a China-compliance claim. SMS geographic anti-abuse controls and per-country spend circuit breakers remain application responsibilities.

Reject any candidate that cannot pass the fixture and the alert trace. That rule is stricter, and more useful, than counting catalog entries.

Instrument the decision, not a vendor response

The first runnable step is to retrieve the exact email-send contract rather than guess its attachment fields. This Go program calls the public capability discovery surface, sets an explicit method and environment-based authorization, checks status bodies, and handles HTTP 429 with bounded exponential backoff plus an integer Retry-After. Pin the returned schema in a contract test before the report worker constructs a write payload.

package main

import (
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

const (
    scheme         = "https://"
    host           = "api." + "infrai" + ".cc"
    discoveryRoute = "/v1/discovery/email.send"
)

func main() {
    body, err := fetchContract()
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    fmt.Println(string(body))
}

func fetchContract() ([]byte, error) {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        return nil, fmt.Errorf("INFRAI_API_KEY is required")
    }
    client := &http.Client{Timeout: 10 * time.Second}
    var lastErr error

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodGet, scheme+host+discoveryRoute, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        resp, err := client.Do(req)
        if err != nil {
            lastErr = err
            time.Sleep(time.Duration(1<<attempt) * time.Second)
            continue
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Duration(1<<attempt) * time.Second
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
                delay = time.Duration(seconds) * time.Second
            }
            lastErr = fmt.Errorf("rate limited: %s", body)
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("discovery failed with %s: %s", resp.Status, body)
        }
        return body, nil
    }
    return nil, fmt.Errorf("discovery failed after retries: %w", lastErr)
}
Enter fullscreen mode Exit fullscreen mode

The 10-second timeout and four attempts are safety bounds in this example, not measured provider latency or a recommended production retry budget. The actual write must attach an idempotency key so a retry doesn't duplicate a notification. Pollers still need a durable processed-event key because repeated observations are expected.

This division also makes migration less dramatic. Swap the adapter, replay the same fixture, and keep the alert rule stable. Template ownership is then a conscious release decision rather than an accidental property of whichever dashboard somebody opened first.

Set the threshold with the interruption cost visible

The final choice is conditional. Use separate email and SMS specialists when their verified event behavior meets the SLO and the organization can carry the correlation and on-call split. Use a unified option when one team owns both channels, one contract removes meaningful integration work, and polling latency fits inside the fallback budget. Choose neither path until the report fixture proves attachment rendering, suppression, revision rollback, and idempotent fallback.

False positives are not free. They interrupt the same engineers needed to restore the delivery path, and repeated pages train responders to discount the signal. Yet a comfortable threshold that leaves no time for SMS is merely quiet failure. Start from the user-facing objective, subtract the worst acceptable fallback allowance, and capacity-plan the poller against the remainder. The cheapest acceptable system is the one whose complete operating boundary stays inside that budget, not the one with the smallest number on a pricing page.

Further reading and References

Top comments (0)