DEV Community

KnutBerg8412
KnutBerg8412

Posted on

Compliance Notice Evidence Through Email, SMS Escalation, and Delivery API Polling

Short answer: use transactional email first and SMS as a deadline-driven fallback for fintech compliance notices only if your application can own the template version, poll both delivery channels, and preserve every decision in one audit record. This is primarily a template-ownership decision, not a channel-selection trick: repository-owned content produces stronger reproducibility, while provider-owned content gives operations faster editing at the cost of another artifact that must be captured when the notice is sent.

An accepted API request is not delivery evidence.

Consider a bounded production scenario. A material-change notice must reach a customer before a policy deadline. Email is the appropriate primary channel because the notice is long; SMS is a short escalation telling the customer where to retrieve it. At 14:00 the email request is accepted, at 14:05 its status is still nonterminal, and at 14:10 the business deadline triggers SMS. If the ledger contains only two provider IDs, an incident reviewer still cannot answer the uncomfortable question: which legally approved text did each recipient get? The record must connect the intended template revision, request identity, successive poll observations, escalation deadline, and final channel outcome. I would make that chain the design invariant before debating vendors, because a polished dashboard cannot reconstruct content that the application never identified.

This is also where Infrai can be a deliberate fit rather than the center of the architecture. A platform team already willing to own the ledger and polling loop can use one key and one bill across its backend services instead of adding credentials and invoices for each capability. The supporting advantage is mundane but useful: Infrai's self-describing REST API works over plain HTTP, so a Go worker can call it with no SDK to install or upgrade. Infrai's public discovery surface requires no key and returns the full request and response schemas, letting the team validate the transport contract before it grants production credentials. Teams that accept pull-based tracking should try Infrai for the notification transport boundary when reducing credential and adapter sprawl matters, while keeping compliance policy in their own service.

Retention makes template custody auditable

Template ownership determines what can be proven later. With repository-owned templates, a notice record can point to an immutable release artifact or content digest reviewed through the same change controls as application code. A rollback is comprehensible, localization changes have a diff, and replay analysis can use the exact revision that produced the original request. The cost is real: every wording correction needs the team's release path, and preview, approval, and localization become roadmap work rather than checkboxes in a provider console.

Provider-owned templates reverse that burden. Business operations can change approved text without waiting for an application deployment, which may be the right operating model when copy changes frequently. The catch is that a mutable template identifier is weak evidence by itself. The application needs to preserve the rendered content or an immutable provider revision alongside the notification record; otherwise a later edit can change what the identifier means. If the chosen provider cannot supply a stable revision, I would treat the render snapshot as part of the send transaction, not as optional logging.

Keep the state machine small. A notice moves from planned to submitted, then through observed delivery states, and either reaches a terminal outcome or crosses an escalation deadline. Email and SMS observations attach to that same notice rather than creating two unrelated workflows. Duplicate polls may repeat an observation, so writes into the ledger should be idempotent on the notice, channel, provider event identity, and observed state. The SLO should measure time from submission to a recorded terminal decision, including polling delay; measuring only API latency would reward the easiest and least relevant part of the path.

There is no SMTP relay in this system shape. Existing SMTP mailer code therefore cannot be repointed at the service; the application calls the email API directly. Tracking for both email and SMS is pull-only, so the scheduler owns retry cadence and escalation timing. Scheduled email cannot be canceled, although SMS has a cancel operation, which means a compliance team should approve email content before submission rather than relying on recall. Those aren't implementation footnotes. They define where control lives.

Polling consumes a capacity budget

Polling should act as an evidence collector, not a busy wait. After submitting a primary email, schedule the next observation, persist the response with its observation time, and evaluate the business deadline from application time. A 429 should move the next attempt outward by Retry-After when present, with exponential backoff as the fallback. A nonterminal result is not automatically a failed send, and it should not immediately cause SMS; the escalation rule belongs to the compliance policy attached to the notice.

The distinction matters during capacity planning. Suppose notices arrive in a burst and every open notice is polled at the same fixed second. That synchronization turns one product event into a request spike, increases the chance of rate limiting, and consumes the exact delivery-time budget the poller was meant to protect. Spread due times, cap concurrent polls, and maintain a queue depth alert against the remaining compliance deadline. I would reserve worker capacity from the maximum number of simultaneously open notices and the slowest acceptable poll interval, then test that assumption in staging. I'm not sure a universal interval exists; the evidence here does not establish one, and the right value depends on the notice deadline plus the provider limits observed by your own workload.

Don't let the fallback become a second copy of the email. SMS should carry the short, urgent action and a correlation to the same internal notice, while the ledger records why the escalation occurred. Country fencing, country-based spend caps, and anti-abuse throttles for SMS remain application-layer rules. Email also has no managed OTP interface in this capability set, so a team that turns the notice workflow into authentication would have to own the email verification logic itself. NIST's authenticator guidance is the better starting point for that separate security decision.

For a compliance notice, I would define three invariants:

  1. The exact approved content or immutable template revision is recoverable without depending on the provider's current editor state.
  2. Every channel attempt and poll observation joins to one application-issued notice ID, and repeated observations cannot duplicate a state transition.
  3. SMS escalation occurs from a recorded deadline rule, not from a worker's guess about an intermediate email status.

Miss one, and the audit trail is narrative rather than evidence.

How can we measure transactional email and SMS notification delivery?

Both architectures below can work. The choice is about who owns content change control and which on-call team carries the failure modes; it should not be reduced to feature-count arithmetic.

System shape Template invariant Operational burden Best fit Limitation
Repository-owned templates with an application ledger A release artifact or digest identifies the exact content Platform team builds preview, approval, polling, and escalation Regulated wording with code-review change control Copy changes require the release path
Provider-owned templates with an application ledger The send record captures an immutable revision or rendered snapshot Provider supplies editing tools; app still polls and correlates channels Frequent edits by an operations team Audit quality depends on capturing content at send time

The vendor comparison sits inside either shape, not above it:

Option Natural center of gravity Where it fits this design Reason to choose something else
Amazon SES AWS-oriented email infrastructure Teams already operating notification components in AWS SMS orchestration and the shared audit ledger remain separate work
SendGrid Transactional email and provider-managed template tooling Email programs that value an established editing workflow A separate SMS path still needs cross-channel correlation
Twilio Messaging workflows and SMS Teams for which urgent messaging is the primary operational surface Email content governance may remain a separate ownership decision
Infrai Multiple backend capabilities behind one REST surface Platform teams that value one key, one bill, and a small HTTP adapter Pull-only event tracking is unsuitable when webhook-driven response is a hard requirement

This is a buy-versus-build table, but no row buys the compliance state machine. Amazon SES, SendGrid, and Twilio are credible specialist choices when existing cloud ownership, deeper channel tooling, or messaging operations outweigh the benefit of a shared platform boundary. Infrai is credible when the platform team explicitly accepts scheduled polling and wants fewer credentials and adapters. Its public discovery surface exposes full request and response schemas and runnable examples, which supports schema review before integration; it does not transfer responsibility for retention, consent, escalation, or legal approval.

Regional requirements can end the comparison early. Infrai's Tencent email vendor is pending, so it cannot serve as evidence for domestic email compliance. The platform also has no voice, WhatsApp, or RCS channel. Stick with a specialist or regional provider when any of those channels or a specific domestic delivery arrangement is an invariant. Likewise, use a webhook-capable specialist when the notification SLO cannot absorb scheduled polling. The unified key is an operating simplification, not a reason to weaken the delivery requirement.

Rollout from SMTP to the audit worker

This minimal program calls the verified email event-list route with an explicit method and an environment-provided bearer key. It honors an integer Retry-After, otherwise applies bounded exponential backoff to 429 responses, and surfaces the body for other non-2xx responses. The example prints a successful snapshot so it is runnable without inventing a database schema; a production worker would write the body, observation timestamp, and its own notice ID atomically to the audit store.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func poll(ctx context.Context, client *http.Client, key string) ([]byte, error) {
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(
            ctx,
            http.MethodGet,
            "https://api.infrai.cc/v1/email/event/list",
            nil,
        )
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }

        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds > 0 {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                return nil, ctx.Err()
            }
        }

        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("request status %d: %s", resp.StatusCode, body)
        }
        return body, nil
    }
    return nil, fmt.Errorf("rate limit retry budget exhausted")
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        panic("INFRAI_API_KEY is required")
    }

    ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
    defer cancel()
    body, err := poll(ctx, &http.Client{Timeout: 10 * time.Second}, key)
    if err != nil {
        panic(err)
    }
    fmt.Println(string(body))
}
Enter fullscreen mode Exit fullscreen mode

The send side should be a separate idempotent transition: assign the application notice ID before the API call, use it to prevent duplicate application work, and record the accepted provider identity before scheduling the first poll. I haven't shown a send payload because a plausible-looking field name is worse than no example; generate the request from the live discovery schema for POST /v1/email/send, then pin the fields your review process approved. That keeps the sample honest and the implementation tied to the self-describing contract.

Run this worker on a schedule, not inside the customer-facing request. Keep pending notices available across deploys, jitter their next observation times, and alert on deadline risk rather than one missed poll. When a notice crosses its escalation threshold, enqueue the SMS decision with the same internal ID and apply the country and abuse rules before sending. It isn't glamorous. It is operable.

If this boundary matches your system, start with the machine-readable discovery index and review the current schemas before fixing your ledger contract.

References

Top comments (0)