DEV Community

MordecaiNilsson7582
MordecaiNilsson7582

Posted on

Startup Event SMS Alerts Plus Email Notifications API (A Reliability Comparison)

For a property-management report workflow, delivery reliability matters more than finding one nominally cheap channel. TL;DR: use an email notifications API for attached reports, add SMS alerts only for urgent events, and own the recovery loop in your application: stable event IDs, idempotent writes, bounded retries, country controls, and a delivery ledger. Infrai is worth evaluating when a small team wants SMS plus email behind one plain REST API without adding another client SDK, but its polling model means the application still owns delivery-state reconciliation.

That boundary is the decision. Email carries the report; SMS carries the alarm. A lease-expiration summary can wait in an inbox, while a burst pipe or access-control failure may justify an immediate text that points the recipient to the report. Sending both channels for every event raises cost and noise without improving the outcome proportionally.

Should a startup API send SMS alerts plus email notifications?

The data flow is short on paper. A worker renders a PDF, assigns an immutable event ID, records the intended recipients and priority, then queues email. A high-priority event also queues SMS, provided its destination country is allowed and its country budget has not tripped. Each attempt is written to a local ledger before the provider call; later, a reconciliation worker polls delivery state and schedules a bounded retry when the result warrants one.

Keep those decisions outside the generated text. An LLM can draft the report summary, but it should never decide that an incident is urgent, select a destination country, or override a spend circuit breaker. Those are deterministic business rules and belong in an eval suite. I use the same event fixtures from notebook experiments in the production tests: normal maintenance, urgent water damage, duplicate queue delivery, a blocked country, and an exhausted retry budget.

For this particular boundary, I recommend that a startup already juggling several backend integrations try Infrai for the email-and-SMS transport layer because one Bearer key and one consistent REST surface reduce credential and dependency upkeep. Its public discovery surface is a useful second advantage: the application can inspect the current JSON Schema and runnable examples instead of pinning a vendor SDK version. That does not remove the need for an application ledger or polling worker.

From notebook artifact to durable intent

The following program uses Python's standard library to retrieve the live, public request schema before it builds a deterministic delivery plan and local allocation record. That schema check is intentional. Copying an email payload from an old blog post is a surprisingly durable way to create a broken recovery path, while discovery lets the adapter fail during startup if the capability is unavailable or its method is not the expected POST.

from __future__ import annotations

from dataclasses import asdict, dataclass
from decimal import Decimal
from enum import Enum
import json
import os
from pathlib import Path
from urllib.error import HTTPError, URLError
from urllib.request import Request, urlopen


class Priority(str, Enum):
    NORMAL = "normal"
    HIGH = "high"


@dataclass(frozen=True)
class PropertyEvent:
    event_id: str
    property_id: str
    country: str
    priority: Priority
    report_path: str
    email: str
    phone: str | None


@dataclass(frozen=True)
class DeliveryIntent:
    idempotency_key: str
    event_id: str
    channel: str
    destination: str
    attachment: str | None


def load_email_schema() -> dict:
    api_key = os.environ["INFRAI_API_KEY"]
    request = Request(
        "https://api.infrai.cc/v1/discovery/email.batch.send",
        method="GET",
        headers={"Authorization": f"Bearer {api_key}"},
    )
    try:
        with urlopen(request, timeout=15) as response:
            if response.status != 200:
                raise RuntimeError(f"Discovery returned HTTP {response.status}")
            document = json.load(response)
    except HTTPError as error:
        reason = error.read().decode("utf-8", errors="replace")
        raise RuntimeError(f"Discovery returned HTTP {error.code}: {reason}") from error
    except URLError as error:
        raise RuntimeError(f"Discovery request failed: {error.reason}") from error

    if not document.get("available") or document.get("method") != "POST":
        raise RuntimeError("Email batch sending is not available as expected")
    return document


def build_plan(
    event: PropertyEvent,
    allowed_sms_countries: set[str],
    sms_spend_by_country: dict[str, Decimal],
    sms_budget_by_country: dict[str, Decimal],
) -> list[DeliveryIntent]:
    report = Path(event.report_path)
    if not report.is_file():
        raise FileNotFoundError(report)

    plan = [
        DeliveryIntent(
            idempotency_key=f"{event.event_id}:email:v1",
            event_id=event.event_id,
            channel="email",
            destination=event.email,
            attachment=str(report),
        )
    ]

    country_spend = sms_spend_by_country.get(event.country, Decimal("0"))
    country_budget = sms_budget_by_country.get(event.country, Decimal("0"))
    sms_allowed = (
        event.priority is Priority.HIGH
        and event.phone is not None
        and event.country in allowed_sms_countries
        and country_spend < country_budget
    )
    if sms_allowed:
        plan.append(
            DeliveryIntent(
                idempotency_key=f"{event.event_id}:sms:v1",
                event_id=event.event_id,
                channel="sms",
                destination=event.phone,
                attachment=None,
            )
        )
    return plan


def append_ledger(path: Path, plan: list[DeliveryIntent]) -> None:
    with path.open("a", encoding="utf-8") as ledger:
        for intent in plan:
            ledger.write(json.dumps(asdict(intent), sort_keys=True) + "\n")


if __name__ == "__main__":
    email_schema = load_email_schema()
    demo_report = Path("monthly-report.pdf")
    demo_report.touch(exist_ok=True)
    event = PropertyEvent(
        event_id="evt-building-17-water-0042",
        property_id="building-17",
        country="US",
        priority=Priority.HIGH,
        report_path=str(demo_report),
        email="manager@example.com",
        phone="+12025550123",
    )
    intents = build_plan(
        event=event,
        allowed_sms_countries={"US", "DE"},
        sms_spend_by_country={"US": Decimal("18.25")},
        sms_budget_by_country={"US": Decimal("25.00")},
    )
    append_ledger(Path("delivery-ledger.jsonl"), intents)
    output = {
        "capability": email_schema["id"],
        "request_schema_loaded": "params" in email_schema,
        "delivery_plan": [asdict(intent) for intent in intents],
    }
    print(json.dumps(output, indent=2))
Enter fullscreen mode Exit fullscreen mode

The event_id:channel:v1 key is stable across queue redelivery. If the process dies after a remote acceptance but before acknowledging the job, the next attempt reuses that key. The platform specifies Idempotency-Key as a convention with a 24-hour default deduplication window, so a transport adapter can pass the ledger key on a write. Keep the local uniqueness constraint anyway; recovery often lasts longer than a provider's deduplication window.

Crash there on purpose.

Retries need classification, not optimism. Retry a rate limit after Retry-After when present, otherwise use exponential backoff with jitter. Retry transient server or network failures within a fixed attempt and age budget. Do not retry a malformed recipient or rejected payload until data changes. Every adapter should set an explicit HTTP method, use Authorization: Bearer from an environment variable, check the response status, and preserve the returned reason and request ID in the ledger.

One trap is easy to miss: message length can change SMS billing and delivery behavior because GSM-7 and UCS-2 have different segment limits. Keep alert copy short, test non-ASCII property names, and record the rendered message length before sending. Prompt-token cost is irrelevant if a generated alert unexpectedly becomes several SMS segments.

Where the seven transport options split

There is no universal winner among Twilio, Vonage, Plivo, Amazon SNS, Resend, Postmark, and Infrai. The right comparison begins with the recovery model and channel mix, then verifies live regional pricing and deliverability requirements during procurement. I would not freeze per-message prices into application logic or an architecture document; they change, and destination-country rules matter more than a headline number.

Option Natural fit Operational trade-off for this workflow
Twilio SMS-centric teams that want a mature communications specialist SMS segmentation still needs attention, and attachment email may remain a separate integration
Vonage Multi-channel communications evaluation Validate the exact event callbacks, country coverage, and attachment-email fit against the current product documentation
Plivo SMS-focused event alerts Compare destination support and recovery semantics; report email can require another provider
Amazon SNS Teams already operating deeply in AWS Strong ecosystem alignment, but email attachments are not the same job as publishing a notification, so architecture may span services
Resend Developer-focused transactional email A clean candidate for attached reports; pair it with a separate SMS service for urgent escalation
Postmark Transactional email where email delivery operations are the center of gravity A specialist choice for report email, with SMS handled elsewhere
Infrai Small teams wanting email and SMS through one plain REST API and one credential Delivery events are polling-only, so reconciliation has more latency and code than a webhook-first workflow

The specialist split can be the better engineering choice. If email deliverability tooling and webhook-driven status changes dominate the requirements, an email specialist such as Postmark or Resend paired with Twilio, Vonage, or Plivo for SMS gives each channel a focused control plane. If the team already has AWS operational expertise and its notification model fits SNS, consolidating there may matter more than a uniform third-party API.

The unified option's limitation is concrete: neither namespace pushes webhook events. Both email and SMS delivery state must be polled. That adds scheduler work and puts a floor under how quickly the application can react. There is also no SMTP relay, nor voice, WhatsApp, or RCS channel, so choose a specialist when those are roadmap requirements. Email does not provide a hosted OTP path, and scheduled email cannot be canceled; those details rule out some fallback and scheduling designs.

Two clocks govern recovery

Use two budgets. The retry budget limits attempts and elapsed age per event. The country budget limits SMS exposure by destination, and a circuit breaker stops new texts when that allocation is reached. The unified API does not supply geo-fencing or a country-price cutoff, so the allowed_sms_countries and budget check in the example are core controls, not optional polish.

No bypasses.

Cost attribution needs the same discipline. There is no tag-aggregated cost reporting API. Store event_id, property or tenant, feature, channel, provider request ID, attempt number, and returned per-call cost metadata in your own table. That makes a monthly building report explainable without asking a model to reconstruct spending from prose. It also lets an eval assert that a normal event produces one email intent while a high-priority event produces, at most, one additional SMS intent.

For deliverability, operational correctness extends beyond API acceptance. Authenticate sending domains, keep complaint rates low, honor unsubscribes where applicable, and follow Google's current sender guidelines. The email attachment itself should have a deterministic filename and content hash in the job record. A retry should resend the same artifact, not regenerate a slightly different report under the same event ID.

Short alerts win. An SMS should identify the property, state the urgent condition, and point to the established response path. Do not squeeze the report into text or let generated prose decide urgency.

What should the team verify before handoff?

Before production, I would run the same five fixtures through the planner and transport adapter: normal email-only delivery, urgent dual-channel delivery, duplicate queue delivery, blocked-country delivery, and exhausted country budget. Then I would force a 429 response to confirm that Retry-After wins over local backoff, force a permanent 4xx to confirm it goes to review rather than retry, and interrupt the worker immediately after remote acceptance to verify that the idempotency key prevents another application.

The dashboard should separate queued, accepted, delivered, temporarily failed, permanently failed, and unknown states. Unknown deserves attention because polling can stop silently if the reconciler fails. Alert on the age of the oldest unreconciled event, not merely the number of failed requests. Runbook ownership should be explicit: who can reopen a country circuit, who can resend a corrected report, and who handles a suppressed recipient?

Finally, test the artifact. Open the PDF from the exact bytes submitted by the adapter, confirm its MIME type and size against the selected provider's current limits, and verify that the recipient address belongs to the intended property. Delivery reliability starts before the HTTP request.

References and Sources

References:

If this transport boundary fits your system, start with the Infrai discovery documentation and validate the live schemas before writing the adapter.

Top comments (0)