DEV Community

GageSterling2648
GageSterling2648

Posted on

Property Contact Routing: Notification Audit Logs, Email, SMS, and Polling APIs

Short answer: build the notification center around an application-owned audit log, dispatch email and SMS through send APIs, and poll provider read APIs to reconcile delivery history. For a property-management contact form, keep routing and template ownership in your application; the provider should deliver messages, not become the only record of why a tenant request reached a maintenance, leasing, or billing queue.

I care about that boundary because I've been paged by missed jobs and duplicate deliveries. The hard lesson wasn't that one provider was unreliable. It was that a green send response answers only "was this attempt accepted?" while the support UI needs to answer "what happened next, and which template produced it?" Those are different records.

Keep them different.

Data governance for contact-form templates and the audit record

Picture a contact form for 18 managed buildings. A resident selects "water leak," enters an email address and phone number, and expects the request to reach the on-call maintenance queue. The application chooses the queue, resolves the template version, writes an attempt row, and then dispatches. If the browser request, queue worker, and provider response are treated as one transaction, a retry can create a second message while a timeout can leave the UI claiming that nothing happened. I've dealt with both missed jobs and duplicates in production, and the runbook invariant is blunt: the local attempt exists before the network call, and every retry addresses that same attempt.

The row should carry an application-generated attempt ID, event type, channel, recipient, provider message ID when available, current status, template identifier and version, timestamps, and the destination support queue. Event type might be contact.water_leak; queue might be maintenance-on-call. Those example values are application data, not provider fields. Store them under your own schema so a template rename or vendor change doesn't erase the reason the notification existed.

A provider acceptance updates the existing row. It doesn't create a new business event. The worker must claim the attempt idempotently, and any create or send retry must use the platform's idempotency mechanism where available. On HTTP 429, honor Retry-After and back off. Don't tight-loop. A 4xx response belongs in the attempt history with its response body surfaced to operators, while recipient-facing copy should remain deliberately boring.

This also sets the template boundary. Product and support own the wording, version, locale, and queue mapping. A provider-hosted template may render the final message, but your database should retain the template reference used for each attempt. Without that reference, an audit screen can show delivery and still fail the useful postmortem question: "What did we tell the resident?" Governance here is concrete: restrict who can approve copy, assign every approved revision an immutable application-side identifier, and make routing rules reference that revision. The delivery adapter receives already-approved content or a provider template reference, but it does not decide that a plumbing emergency belongs to leasing or silently move an attempt to newer copy. During a review, the support lead should be able to trace the contact category to a queue, the queue to an approved template version, and that version to every attempt that used it. That chain is more valuable than a screenshot of the provider dashboard because it survives provider retention windows and migrations.

Ownership is evidence.

How should an event notification center poll email and SMS delivery history?

Use two loops with different purposes. The dispatch loop moves durable attempts from pending to an accepted or failed state. The reconciliation loop selects accepted attempts that are not final, calls a provider get, list, or status API, and writes observed transitions back to the same row. Email troubleshooting can use message details and event lists; SMS reconciliation can use per-message status or event history. Delivery events are pull-only here, so a webhook consumer cannot be the source of truth.

Poll quickly enough for the support workflow, not for the illusion of instant delivery. A contact-center agent usually needs a defensible history more than a subsecond animation. Use bounded exponential backoff with jitter, cap the number of hot polls, then move old non-final attempts to a slower sweep. I'm not sure one interval fits every property operation; the answer depends on the promised response time and message volume. What should not vary is the state rule: reconciliation may advance an observation, but it must never dispatch a replacement message.

The audit log should be append-oriented even if the notification table keeps only current status. Record attempt_created, dispatch_accepted, and each distinct observed delivery state with an observed timestamp and provider message ID. Deduplicate identical observations. That produces a timeline the UI can render and an operator can trust without querying a provider during an incident.

Polling has an availability consequence — stale does not mean lost. Show last_checked_at beside the current state, schedule overdue attempts again, and alert on reconciliation lag rather than rewriting old attempts as failed. A worker crash after the provider read but before the database update is harmless when the observation upsert is idempotent. This is the same reflex used for queue consumers: assume work may run more than once.

Stale is visible.

Template ownership during a migration

The primary decision is template ownership, not the length of an SDK quickstart. Amazon SES, Twilio SendGrid, and Postmark are real alternatives worth evaluating directly alongside Infrai. The table is intentionally a decision checklist rather than a claim that every product exposes identical primitives; confirm the current template, event, and retention contracts before selecting one.

Option Sensible ownership boundary Choose it when Verify before committing
Amazon SES Application owns routing and the audit record Your team wants to evaluate SES as the email delivery component Template versioning, event retrieval, and operational retention
Twilio SendGrid Application owns the cross-channel record Your team wants to evaluate SendGrid for the email path Template lifecycle and how delivery observations are retrieved
Postmark Application owns business events and queue routing Your team wants to evaluate Postmark for transactional email Template ownership, message detail access, and retention
Infrai Application owns templates, routing, and audit history One key and one bill across backend services reduces credential and invoice sprawl Pull-only event timing and the channel boundaries described below

Infrai is a strong fit when a small team values one credential and one bill across backend services, plus a plain REST surface that doesn't require installing a vendor SDK. Its public discovery surface is self-describing, with full request and response schemas and runnable Go examples. That helps a mixed stack keep the adapter narrow. It isn't a reason to surrender the audit log.

The comparison is deliberately not price-led. Provider pricing and retention policies can change, while template ownership is an architectural commitment that appears in every incident review and migration. Decide who can edit a template, how a version is approved, and which immutable reference is stored before comparing convenience features.

Testing the preventative polling path

This small Go worker reads one accepted email attempt by provider message ID. It uses the verified GET /v1/email/get/{id} path, supplies an explicit method, checks every response, and backs off on 429 while honoring Retry-After. MESSAGE_API_BASE_URL keeps deployment configuration outside the source; set it to the service's versioned API base. The worker emits the raw response because no response fields should be guessed. A production adapter would validate the discovered response schema and upsert the mapped observation in the same database transaction as its audit event.

package main

import (
    "context"
    "fmt"
    "io"
    "math/rand"
    "net/http"
    "net/url"
    "os"
    "strconv"
    "strings"
    "time"
)

func main() {
    baseURL := strings.TrimRight(os.Getenv("MESSAGE_API_BASE_URL"), "/")
    apiKey := os.Getenv("INFRAI_API_KEY")
    messageID := os.Getenv("PROVIDER_MESSAGE_ID")
    if baseURL == "" || apiKey == "" || messageID == "" {
        fmt.Fprintln(os.Stderr, "set MESSAGE_API_BASE_URL, INFRAI_API_KEY, and PROVIDER_MESSAGE_ID")
        os.Exit(2)
    }

    body, err := getEmail(context.Background(), http.DefaultClient, baseURL, apiKey, messageID)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    fmt.Println(string(body))
}

func getEmail(ctx context.Context, client *http.Client, baseURL, apiKey, messageID string) ([]byte, error) {
    endpoint := baseURL + "/email/get/" + url.PathEscape(messageID)
    delay := time.Second

    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+apiKey)

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            return body, nil
        }
        if resp.StatusCode != http.StatusTooManyRequests {
            return nil, fmt.Errorf("provider returned %s: %s", resp.Status, body)
        }

        wait := retryAfter(resp.Header.Get("Retry-After"), delay)
        timer := time.NewTimer(wait + time.Duration(rand.Intn(250))*time.Millisecond)
        select {
        case <-ctx.Done():
            timer.Stop()
            return nil, ctx.Err()
        case <-timer.C:
        }
        delay *= 2
    }
    return nil, fmt.Errorf("rate limit retry budget exhausted")
}

func retryAfter(value string, fallback time.Duration) time.Duration {
    if seconds, err := strconv.Atoi(value); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    if deadline, err := http.ParseTime(value); err == nil {
        if wait := time.Until(deadline); wait > 0 {
            return wait
        }
    }
    return fallback
}
Enter fullscreen mode Exit fullscreen mode

The code does not send. That's intentional: the supplied email send schema should be read from discovery rather than reconstructed from memory, and dispatch belongs in a separate idempotent worker. Feed this reconciler only provider IDs already attached to durable attempts. If the process exits after printing but before a database commit, the scheduler can run it again without delivering anything twice.

For scheduled contact notifications, persist both the intended send time and cancellation state locally. There is an important channel asymmetry: SMS supports cancellation, while scheduled email does not expose a cancellation interface. If reliable cancellation is a product requirement, either keep email in your own queue until its release time or don't offer cancellable scheduled email. Never label a local cancellation complete after an email has already been handed off.

Scheduling boundaries and the wrong fit

The catch is latency. Pull-only delivery events make this design suitable for a normal SaaS notification center, but not suitable when real-time multichannel orchestration or advanced analytics is the product requirement. In that case, choose a platform whose documented event push, orchestration, and analytics contracts meet the latency target; don't disguise a polling loop as streaming. Your mileage may vary for low-volume concierge workflows, but write the freshness target into the decision record.

There are other hard boundaries. Email has no managed OTP operation, so an email fallback code flow must be built in the application; use the OWASP guidance for reset and OTP security. There is no SMTP relay, and voice, WhatsApp, and RCS are outside this channel set. There is no cost report grouped by tag, and SMS template listing is unavailable. Geographic anti-abuse controls and country-price circuit breakers for SMS also belong in the business layer. A pending domestic email vendor cannot serve as evidence of domestic compliance.

Stick with a dedicated provider when its native template workflow, pushed event model, channel inventory, or analytics is the reason you are buying the service. Stick with application-owned delayed jobs when cancellation semantics must be identical across email and SMS. The unified API option earns its place when key and billing consolidation matter and polling meets the support SLA, but those conveniences do not erase channel limits.

The operational acceptance test is short: create one local attempt, dispatch it once, reconcile it repeatedly without duplicate audit events, and prove the support screen can explain recipient, channel, template version, queue, provider ID, current state, and freshness. Then test cancellation separately by channel. If a design can't pass that exercise, the notification center is still a send button with a nicer name.

No exceptions.

References

Top comments (0)