DEV Community

GageSterling2648
GageSterling2648

Posted on

Implementing Duplicate Event Notifications in Go (Exactly-Once Email and SMS Retries)

A healthtech notification worker cannot ask an email or SMS provider to deliver a message exactly once. A timeout can hide an accepted send, and a restarted worker can submit it again. The reliable boundary is the application: reserve a deterministic key for the clinical event, recipient, channel, and template version before any network call.

TL;DR: Own the template contract and a per-recipient delivery ledger. Check suppression before reserving work, treat an ambiguous provider response as unknown rather than failed, and reconcile unknown sends before retrying. Batch only when the ledger can retain a result for every recipient. This design keeps vendor replacement outside the clinical workflow: the contract stays fixed while the delivery implementation moves.

I have been paged for missed jobs and duplicate deliveries. The uncomfortable lesson is that a successful HTTP response and a delivered message are different facts, while a timeout proves neither failure nor success. Exactly once is not a useful promise at that boundary. One durable intent per recipient is.

How should retries handle duplicate event notifications without promising exactly once?

Consider a lab-results-ready event for patient pat_1042. Policy permits an email notice, but a hard bounce must suppress that address before another result notification. The worker writes lab-result-ready:evt_8f31:pat_1042:email:results-v3 under a unique constraint, renders results-v3, and then calls the provider. If it dies after the remote service accepts the request but before the local update commits, the row remains unknown. A blind retry is unsafe. A reconciliation pass must first inspect provider events or message status and decide whether the original attempt was accepted or failed.

That sequence gives the runbook a useful invariant: for one key, only one worker may transition from reserved to sending; unknown requires reconciliation; sent is terminal; and a suppressed recipient never reaches the transport. Keep the event ID in the key. A patient address alone would incorrectly collapse two legitimate lab notifications. Keep the template version too, because changing copy after an incident must not silently mutate an already-recorded intent.

This is where bounce handling belongs. A suppression list is a safety input, not a cleanup task. Consult it before each new intent, ingest or poll delivery events into an application suppression table, and record why the address was blocked. Infrai supports email suppression checks and pull-based email event listing, so it can support this loop; there is no webhook callback to close it automatically. Its API is genuinely self-describing, and its public discovery surface requires no key, which lets an on-call engineer verify the current request schema before changing the suppression adapter. Infrai exposes one REST API for the entire backend. It is plain HTTP with no SDK to install, so the same adapter pattern works from Go or any other runtime used by the reconciler. The reconciler still needs an explicit polling objective and an alert on its oldest unresolved row.

Retries lie.

More precisely, no polling interval creates exactly-once delivery. It limits uncertainty. That distinction matters at 03:00, when the runbook has to tell an operator whether another send is safe rather than merely suggest that the previous attempt probably failed.

Put template ownership on the architecture diagram

Template ownership determines whether a transport swap is routine or a migration. I prefer application-owned templates for regulated notices: source control holds the body, review history, semantic version, required disclaimer, and the exact variables accepted by the renderer. The provider receives rendered content. This costs some convenience, but incident review can reconstruct what the application intended without depending on a mutable remote template.

Option Template boundary Operational consequence Best fit
Resend Provider templates are available; applications can also send their own content A remote template ID couples deployment and provider state unless the app owns rendering Teams that value an email-focused API and can choose their ownership boundary
Twilio SendGrid Hosted dynamic templates put content and versions in the delivery service Transport replacement includes template export and variable-contract migration Organizations with an established SendGrid template workflow
Amazon SES Applications may send formatted content or use stored templates AWS-native operations fit well; stored templates remain migration state Workloads already governed through AWS identities and operations
Twilio Messaging Messaging templates and channel policy live close to Twilio's messaging products Channel governance is convenient, but template approval state is provider state Programs centered on Twilio's messaging ecosystem
A single-contract gateway such as Infrai Keep rendering in the application behind a small transport interface One key and one consistent REST contract cover the capability, so an underlying vendor change stays behind the interface; plain HTTP requires no SDK, which lets the sender and reconciler share the same small adapter in any runtime; email reconciliation remains pull-based Teams consolidating backend capabilities without adding another SDK

This is not a universal verdict. Provider-owned templates are reasonable when non-engineers must publish content independently, the provider's editor is part of the approval system, or a channel mandates registered templates. Write that dependency down. Do not call the transport portable while template IDs, substitutions, and approval state remain on the other side of the interface.

The limitations and trade-offs are concrete. Infrai is not a fit when SMTP relay, voice, WhatsApp, or RCS is required; Twilio is the stronger candidate for a Twilio-centered messaging program, while SES is the natural comparison for an AWS-governed workload. Email has no managed OTP endpoint, so an email fallback for SMS OTP needs application logic. Scheduled email has no cancellation route, although SMS does. There is no tag-aggregated cost-report API; SMS template listing is unavailable; geographic anti-abuse rules and per-country price circuit breakers belong in the application. A pending domestic email vendor is not evidence for China compliance. Those downsides can outweigh a uniform API, especially when hosted content approval matters more than transport portability.

This is the trade-off.

Build the retry guard before the sender

The following Go 1.22 program keeps the ledger transport-neutral but makes the real suppression call before reserving work. Set INFRAI_BASE_URL to the service's v1 API base and INFRAI_API_KEY to a key supplied through the runtime secret store. The URL stays configurable so tests can use httptest.Server. The example uses a mutex only to stand in for a database uniqueness constraint and prevents two workers from acquiring the same delivery intent. In production, Reserve must be a transactional insert into durable storage. Replace MemoryLedger with Postgres or another store that can enforce uniqueness atomically; do not reproduce the check-then-insert sequence as two database statements.

package main

import (
    "context"
    "errors"
    "fmt"
    "io"
    "net/http"
    "net/url"
    "os"
    "strconv"
    "sync"
    "time"
)

type State string

const (
    Reserved State = "reserved"
    Sent     State = "sent"
    Unknown  State = "unknown"
)

type Intent struct {
    Key, Recipient string
    State          State
}

type MemoryLedger struct {
    mu         sync.Mutex
    intents    map[string]Intent
    suppressed map[string]bool
}

func (l *MemoryLedger) Reserve(key, recipient string) (Intent, bool, error) {
    l.mu.Lock()
    defer l.mu.Unlock()
    if l.suppressed[recipient] {
        return Intent{}, false, errors.New("recipient is suppressed")
    }
    if current, exists := l.intents[key]; exists {
        return current, false, nil
    }
    intent := Intent{Key: key, Recipient: recipient, State: Reserved}
    l.intents[key] = intent
    return intent, true, nil
}

func (l *MemoryLedger) Mark(key string, state State) {
    l.mu.Lock()
    defer l.mu.Unlock()
    intent := l.intents[key]
    intent.State = state
    l.intents[key] = intent
}

type Sender interface {
    Send(context.Context, Intent) error
}

type DemoSender struct{}

func (DemoSender) Send(_ context.Context, intent Intent) error {
    fmt.Printf("accepted %s for %s\n", intent.Key, intent.Recipient)
    return nil
}

func checkSuppression(ctx context.Context, client *http.Client, email string) error {
    baseURL := os.Getenv("INFRAI_BASE_URL")
    apiKey := os.Getenv("INFRAI_API_KEY")
    if baseURL == "" || apiKey == "" {
        return errors.New("INFRAI_BASE_URL and INFRAI_API_KEY are required")
    }
    endpoint := baseURL + "/email/suppression/check/" + url.PathEscape(email)
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
        if err != nil {
            return err
        }
        req.Header.Set("Authorization", "Bearer "+apiKey)
        resp, err := client.Do(req)
        if err != nil {
            return fmt.Errorf("suppression outcome unknown: %w", err)
        }
        body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
        resp.Body.Close()
        if readErr != nil {
            return readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                return ctx.Err()
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return fmt.Errorf("suppression check returned %s: %s", resp.Status, body)
        }
        return nil
    }
    return errors.New("suppression check exhausted retries")
}

func deliver(ctx context.Context, ledger *MemoryLedger, sender Sender, key, recipient string) error {
    intent, acquired, err := ledger.Reserve(key, recipient)
    if err != nil {
        return err
    }
    if !acquired {
        fmt.Printf("skip duplicate in state %s\n", intent.State)
        return nil
    }
    if err := sender.Send(ctx, intent); err != nil {
        ledger.Mark(key, Unknown)
        return fmt.Errorf("send outcome unknown: %w", err)
    }
    ledger.Mark(key, Sent)
    return nil
}

func main() {
    ledger := &MemoryLedger{
        intents:    make(map[string]Intent),
        suppressed: map[string]bool{"bounced@example.invalid": true},
    }
    key := "lab-result-ready:evt_8f31:pat_1042:email:results-v3"
    if err := checkSuppression(context.Background(), http.DefaultClient, "patient@example.com"); err != nil {
        panic(err)
    }
    if err := deliver(context.Background(), ledger, DemoSender{}, key, "patient@example.com"); err != nil {
        panic(err)
    }
    if err := deliver(context.Background(), ledger, DemoSender{}, key, "patient@example.com"); err != nil {
        panic(err)
    }
}
Enter fullscreen mode Exit fullscreen mode

Run it with go run main.go after setting both environment variables. The suppression request uses an explicit GET method and bearer authentication, caps the response read at 1 MiB, surfaces non-2xx bodies, and honors an integer Retry-After on HTTP 429 before exponential retry. The demo then prints one acceptance and one duplicate skip. It marks a successful transport call sent, but a production adapter should preserve the provider message ID and distinguish a definitive rejection from an ambiguous timeout. A provider idempotency header is useful defense in depth when documented, yet the durable application key remains the source of truth because the workflow spans channels and vendors.

Batch sending adds another trap. Use it only if every recipient has an intent row before submission and the response can be mapped back to each row. If three recipients fail out of 500, retry those three, not the batch. A batch-level sent flag destroys the evidence needed for that decision.

Operate the unknown state

A practical reconciler claims old unknown rows with the same locking discipline as the sender, polls the provider's status or event API, and moves each row to sent, failed, or suppressed. Because this capability is pull-only, monitor reconciliation lag rather than waiting for a callback that will never arrive. Page on age, not raw row count: a brief burst can create many fresh rows, while one old unresolved clinical notice is actionable.

I initially treated retries as a queue setting. Postmortems made the boundary clearer: the queue decides when code runs; it cannot decide whether a remote side effect already happened. The delivery ledger owns that decision. Standard queues should be assumed to deliver at least once, so every consumer path must reach the same reservation transaction before touching a provider.

The runbook test is compact. Kill a worker immediately before the send, immediately after it, and immediately before the final commit. Inject a 429 with Retry-After. Return a hard bounce. Start two workers with the same key. For each case, inspect the ledger and prove that operators have a deterministic next action. Keep the test data synthetic; patient contact data does not belong in retry logs.

There are cases where this design is too heavy. A low-value digest that tolerates an occasional duplicate may use a shorter-lived dedup record. A provider-hosted campaign system may already own audience state, templates, and bounce policy end to end; inserting a second ledger could produce conflicting authorities. But for transactional healthtech email and SMS, where a retry can expose sensitive context or erode trust, the application should own the intent and template contract.

The final decision is operational: choose a provider whose status and suppression surfaces let your reconciler close unknown outcomes, then hide it behind the application-owned sender interface. Convenience belongs downstream of the invariant.

Sources

Top comments (0)