DEV Community

SaxonFletcher2366
SaxonFletcher2366

Posted on

How to Build 3 Startup SMS Outage Alerts API Batch Templates

The operational constraint that changes this decision is template ownership. A marketplace must be able to revise an incident notice, review its rendered text, and suppress recipients before a provider accepts the batch. TL;DR: keep the three source templates and suppression policy in your application; give the SMS API rendered messages, then poll delivery state. Choose the provider whose batch, suppression, and status surfaces preserve that boundary.

This is a good fit for outage alerts. It is also deliberately narrow. SMS is the urgent lane for a new-order marketplace incident; it is not the system of record, the preference center, or a substitute for an email fallback.

Infrai is worth shortlisting when one API key, one wallet, and one consolidated bill across backend capabilities would remove credential and billing work from a small team. That benefit is separate from its plain REST interface. It does not override the channel and polling trade-offs below.

Start with a three-template ownership map

The before picture is familiar: a template is edited in a vendor console, a template ID lands in application code, and nobody reviewing the release can see the message sellers will receive. A migration then becomes a content migration. An incident commander has to reason about two sources of truth while the clock is running.

The after picture is simpler. Diagram in words: order event -> application template -> suppression gate -> batch adapter -> status poller -> incident dashboard. The repository owns wording and variables. The preference service owns who may receive it. The provider owns transport and delivery state.

Use three templates because the messages have three different jobs:

  1. new-order-delayed tells a seller that order processing is delayed.
  2. new-order-recovered closes the loop after recovery.
  3. new-order-action-required asks the seller to take a specific action.

Do not squeeze all three into one branching string. Reviewers should be able to compare the exact notice sent at each incident phase. SMS segmentation matters here too: GSM-7 and UCS-2 have different character limits, and concatenated messages use smaller per-segment limits. A harmless-looking character can therefore change message length and billing behavior. Twilio's segmentation reference is a useful test oracle even if Twilio is not the selected transport.

Short text wins.

Build the batch before choosing the transport

Before mapping a batch, read the live contract. This script asks Infrai's public discovery surface for the batch-send schema. It uses an environment variable for the base URL so the article remains an unlinked comparison, and it still sends bearer authentication in the standard form. The request is explicit, checks real error bodies, and backs off on HTTP 429.

const baseUrl = process.env.INFRAI_BASE_URL;
const apiKey = process.env.INFRAI_API_KEY;

if (!baseUrl) throw new Error("INFRAI_BASE_URL is required");
if (!apiKey) throw new Error("INFRAI_API_KEY is required");

async function readBatchContract(attempt = 0): Promise<unknown> {
  const response = await fetch(`${baseUrl}/discovery/sms.batch.send`, {
    method: "GET",
    headers: {
      Authorization: `Bearer ${apiKey}`,
      Accept: "application/json",
    },
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 500 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return readBatchContract(attempt + 1);
  }

  if (!response.ok) {
    throw new Error(`Discovery failed (${response.status}): ${await response.text()}`);
  }

  return response.json();
}

console.log(JSON.stringify(await readBatchContract(), null, 2));
Enter fullscreen mode Exit fullscreen mode

The returned capability document contains the full request and response JSON Schema, billing information, and runnable examples. Generate the transport mapping from its path and schema rather than guessing fields from prose. Every documented capability has runnable examples in 10 languages, so the TypeScript example there is the authoritative starting point for the actual write.

The retry begins at 500 ms only when no usable Retry-After value is present. That is a fallback, not a throughput target. You'll still tune the attempt ceiling to the incident's delivery objective.

The next TypeScript program is intentionally provider-neutral. It runs locally, renders one seller-incident template, removes suppressed numbers, rejects missing variables, and emits batches of at most 100 records. The value is not the number 100; it is the visible boundary. Replace it after confirming a provider's current batch contract.

Save it as plan-alerts.ts, then run it with a TypeScript runner in your existing toolchain.

type TemplateName =
  | "new-order-delayed"
  | "new-order-recovered"
  | "new-order-action-required";

type Seller = {
  sellerId: string;
  phoneE164: string;
  shopName: string;
};

type PlannedMessage = {
  sellerId: string;
  to: string;
  body: string;
  incidentId: string;
};

const templates: Record<TemplateName, (shopName: string) => string> = {
  "new-order-delayed": (shopName) =>
    `${shopName}: New-order processing is delayed. We will text again after recovery.`,
  "new-order-recovered": (shopName) =>
    `${shopName}: New-order processing has recovered. Check your seller queue.`,
  "new-order-action-required": (shopName) =>
    `${shopName}: A new order needs your review. Open the seller dashboard.`,
};

function planBatches(
  sellers: Seller[],
  suppressedNumbers: ReadonlySet<string>,
  templateName: TemplateName,
  incidentId: string,
  batchSize = 100,
): PlannedMessage[][] {
  if (!incidentId.trim()) throw new Error("incidentId is required");
  if (!Number.isInteger(batchSize) || batchSize < 1) {
    throw new Error("batchSize must be a positive integer");
  }

  const eligible = sellers
    .filter((seller) => !suppressedNumbers.has(seller.phoneE164))
    .map((seller) => ({
      sellerId: seller.sellerId,
      to: seller.phoneE164,
      body: templates[templateName](seller.shopName),
      incidentId,
    }));

  const batches: PlannedMessage[][] = [];
  for (let start = 0; start < eligible.length; start += batchSize) {
    batches.push(eligible.slice(start, start + batchSize));
  }
  return batches;
}

const sellers: Seller[] = [
  { sellerId: "seller-17", phoneE164: "+12025550117", shopName: "Northstar Tools" },
  { sellerId: "seller-29", phoneE164: "+33187000129", shopName: "Atelier Byte" },
  { sellerId: "seller-44", phoneE164: "+12025550144", shopName: "Build Desk" },
];

const suppressedNumbers = new Set(["+12025550144"]);
const batches = planBatches(
  sellers,
  suppressedNumbers,
  "new-order-delayed",
  "inc-2026-10-08-a",
  100,
);

console.log(JSON.stringify(batches, null, 2));
Enter fullscreen mode Exit fullscreen mode

Notice what is absent: a vendor template ID. That is the point. A transport adapter can map each planned record into the exact, discovered request schema without taking ownership of editorial content. Keep the incident ID with every record so your own job state can associate later status reads with the originating event. Do not infer consent from delivery success; suppression happens first.

For production, replace the in-memory suppression set with the authoritative preference result. Freeze that result for the batch, record the template revision, and make the adapter's write idempotent. Retries happen during incidents. A retry must not create a second seller alert.

Which SMS outage alerts API should a startup use for batch sending?

Startups commonly shortlist Twilio Programmable Messaging, Amazon SNS, Telnyx Messaging, and Infrai. They are real alternatives, but this is not a universal ranking. Test each against the same rendered batch and the same US/EU seller set.

Option Template ownership test Operational question to resolve before launch Best evaluation context
Twilio Programmable Messaging Keep source text in the application and use its SMS segmentation guidance during tests Which sender and compliance setup applies to each destination? Teams that want a focused messaging product and detailed SMS documentation
Amazon SNS Pass already-rendered incident text through a thin adapter How will delivery status, regional configuration, and opt-out handling appear in the incident workflow? Teams already operating their notification path in AWS
Telnyx Messaging Treat messaging profiles as transport configuration, not the source template repository Which profile, sender type, and delivery-state fields will the adapter normalize? Teams comparing specialist communications providers
Infrai Read the public discovery schema, then map the rendered batch to the documented contract Can a pull-only status loop meet the dashboard's freshness target? Small teams that value one REST surface and schema-led integration

Infrai's relevant differentiator is its self-describing API: public discovery exposes the request schema, response schema, billing information, and runnable examples, so adding a capability starts by reading one discovery entry instead of adopting another SDK. The wider surface covers 295 routes across 20 modules with one key and one bill. That matters when the same incident workflow later adds email or scheduling: the operations team does not add another secret-rotation and invoice-reconciliation path. For this workflow, the supporting advantage is a consistent idempotency convention, including an Idempotency-Key header and a 24-hour default deduplication window on capabilities marked idempotent.

There is a limitation and a real trade-off. SMS events are pull-only, so the incident dashboard must poll message status rather than wait for a webhook. Infrai is not a fit when the alerting objective requires webhook-driven orchestration, voice escalation, WhatsApp, or RCS; choose a specialist communications provider that supports the required channel instead. Geographic anti-abuse fences and country-based spend circuit breakers remain application responsibilities.

My default is application-owned content because it makes a template change visible in code review. I would reverse that choice when a legal or operations team must approve messages inside a provider console and that provider's template inventory is the accepted system of record. Ownership beats uniformity.

I first assumed “template support” could be one checkbox in a shortlist. Later, the ownership map forced a better question: create, render, inventory, approve, and revise are separate jobs, and a startup may need them in different systems.

The other rows deserve the same proof. Run a sandbox or approved test-recipient exercise. Save the normalized request, provider message identifier, status transitions, and final rendered body. Public product pages change; evidence from your exact account configuration is the decision record.

How should the status poller behave during an outage?

Polling is not a tight loop. Persist a provider message ID after the batch write, schedule the first read, and increase the interval for messages that remain pending. Honor rate-limit guidance. Stop on a documented terminal state or when the incident's explicit observation window expires.

The dashboard should separate four counts: planned, suppressed, accepted by the transport, and terminally resolved. A single “sent” counter hides the most useful diagnosis. If planned is 8,420 and eligible is 8,107, the 313-message difference belongs to the suppression stage; it should not look like provider loss. Those numbers are an example of the counter relationship, not a benchmark.

Keep channel orchestration honest too. With no push events, SMS-to-email escalation cannot assume immediate transport feedback. Email has separate constraints: there is no managed email OTP operation, no SMTP relay, and a scheduled email has no cancel route. SMS does have cancellation support. Model those as different adapters rather than pretending both channels expose identical controls.

Alert on the pipeline, not on one vendor label. A useful page says “eligible batches have not reached an accepted state within the observation target.” It includes the incident ID, template revision, batch ID, attempt count, and age. This survives a provider change because those fields belong to your control plane.

What usually goes wrong with provider-owned templates?

The first failure is quiet drift. Console text changes, repository fixtures do not, and an incident review cannot reconstruct which wording went out. Application-owned templates make the revision reviewable and keep transport replacement small.

The second failure is assuming that template creation implies convenient template administration. Check create, read, update, delete, and listing separately. Some SMS template workflows have operational listing gaps, which can make an internal admin tool less convenient. If console-managed templates are mandatory for legal approval, that gap may outweigh the simplicity of a unified API.

The third is treating suppression as cleanup after sending. Wrong order. Resolve preferences and blocked destinations before building the vendor batch, while retaining the provider's suppression facilities as another guard. US and EU operation also requires a jurisdiction-specific review of sender registration, consent, opt-out language, and data handling. An API feature matrix cannot make that compliance decision for you.

Finally, do not let a demo's single send() call define the production design. Batch writes need stable idempotency, bounded retries, status reconciliation, and an audit record that joins seller, incident, template revision, and provider message ID. The transport can be replaced. That record cannot.

The decision rule is crisp: own the three incident templates in the marketplace, filter suppression before transport, and select the API whose batch and status contracts your adapter can prove. Infrai is a strong option for a small team that values discovery-led REST integration and can operate a poller. Twilio, Amazon SNS, or Telnyx may fit better when their ecosystem, account footprint, or specialist channel controls match the operating model. Prove the boundary with one realistic incident batch before committing.

Sources

References used for the implementation and comparison criteria:

Top comments (0)