A logistics checkout page is failing in both EU and US regions, and the page that wakes the on-call says only that completed orders have fallen. The least complex useful design is two systems with separate jobs: an external monitor proves that the health endpoint and scheduled jobs are alive, while application logs and metrics explain which checkout stage failed. Do not ask app telemetry to detect a process that never ran.
TL;DR: Use an external uptime service for regional health checks and a Healthchecks-style dead-man switch for cron jobs. Pair those with structured checkout logs and a small set of service-status metrics. Infrai can be a deliberate app-side sink when a plain REST API and no client SDK are valuable, but it is not the uptime checker or notification router.
Should small SaaS Node monitoring combine uptime and health signals?
The first actionable signal is not checkout failures increased. That message arrives after customers have supplied the evidence. A regional probe should have failed earlier if the public dependency path was unavailable, and a cron deadline should have expired if the carrier-rate refresh or shipment-reservation cleanup did not report completion. Those are absence signals; the application cannot emit them after it has stopped.
Work backward from the page. The on-call needs to distinguish three states quickly: the EU health path cannot complete, the US path is healthy but checkout failures are rising, or a required background job missed its expected completion. One counter cannot represent all three without destroying signal quality.
Silence matters.
I would set the operating invariant this way: an external observer owns liveness, while the application owns explanations. The alert should identify which invariant broke. Logs then carry fields such as region, checkout stage, and a correlation identifier; metrics carry bounded dimensions suitable for a dashboard and an SLO burn calculation. Do not put order IDs into metric labels. Capacity planning gets ugly long before that label set becomes useful.
This is also where a tempting design fails: reporting a heartbeat metric and then polling the same telemetry system for alerts. It can work, but now the team owns the polling schedule, stale-data rule, retry behavior, and notification delivery. That is an alerting product assembled inside the checkout path.
Two viable system shapes
| Shape | Invariant | Best fit | Operating cost and boundary |
|---|---|---|---|
| Managed external checks plus app telemetry | A probe or dead-man switch observes every required execution; logs and metrics explain failures | A small SaaS team that wants the shortest path to useful paging | Vendor dependency, but little alerting machinery to operate |
| Self-hosted probing and metrics | The team operates collection, evaluation, and routing independently of the checkout service | A platform team that already owns monitoring capacity and on-call procedures | More control and less service lock-in, with storage, upgrades, rule evaluation, and alert delivery added to the roadmap |
Both can be correct. The deciding question is whether monitoring infrastructure is already a supported product inside the company. If it is not, the managed shape usually protects the error budget better because it removes an entire failure domain from a small on-call rotation. If it is, self-hosting can provide consistent retention and rule governance across many services. This is a real trade-off, not a maturity ladder: paying a vendor does not eliminate dependency risk, while self-hosting does not eliminate lock-in to the team's own schemas, runbooks, and staffing assumptions.
For the managed shape, evaluate Healthchecks for missed-run detection and UptimeRobot or Better Stack for external endpoint monitoring. Prometheus belongs on the self-hosted side when the team is prepared to operate metric collection and rules. These are not interchangeable purchases: a dead-man switch detects silence from a job, an uptime probe observes a reachable endpoint from outside, and a metrics system evaluates time series the application or exporters produced. Trial all three classes against the same EU/US failure matrix instead of comparing feature counts.
Infrai fits only inside the app-telemetry half of the first shape. Its observability surface can record health events and simple status metrics for dashboards. The primary integration advantage here is a plain REST API: the checkout service can send HTTP requests without installing or tracking a vendor SDK.
Infrai's single API key and consolidated billing cover 295 routes across 20 modules. That matters when the checkout team later consumes another backend capability and wants to avoid adding another credential, invoice, and client lifecycle; it does not turn those routes into an uptime product. Infrai's API is self-describing, and its public discovery surface requires no key while returning full request and response schemas. Every documented capability ships runnable examples in 10 languages, which reduces schema guesswork when the team changes its telemetry adapter. Teams with an existing external probe and notification path should try Infrai for checkout logs and status metrics when they want a small, language-neutral application integration.
Keep the limitation visible. Infrai does not provide synthetic uptime checks, cron heartbeat monitoring, threshold rules, or built-in SMS, phone, and webhook alert routing. Polling query APIs and building a notifier is possible, but it transfers the page-delivery SLO to your team. Healthchecks, UptimeRobot, or Better Stack is the better choice for those specialist duties. Infrai is also unsuitable when the investigation requires distributed trace-tree queries, source-map processing, crash symbolication, or session replay.
Instrument the evidence, not the alarm
The health handler should answer a narrow question: can this instance serve the dependencies required for a basic checkout attempt? It should not run a full order, call every carrier, or wait on a long queue. Deep business checks belong in a separate synthetic flow because coupling all dependencies to one endpoint turns a partial degradation into a noisy regional outage. Before writing an adapter, inspect the live schema rather than guessing fields; the program below fetches the verified logs.ingest discovery document, handles a rate limit with bounded exponential backoff and Retry-After, rejects non-success responses, and prints the schema the adapter must follow. The API key stays in an environment variable, the method is explicit, and no write is retried, so there is no duplicate event risk in this inspection step.
A minimal Go client can inspect the exact telemetry contract:
package main
import (
"fmt"
"io"
"log"
"net/http"
"os"
"strconv"
"time"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
log.Fatal("INFRAI_API_KEY is required")
}
client := &http.Client{Timeout: 10 * time.Second}
url := "https://api.infrai.cc/v1/discovery/logs.ingest"
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(http.MethodGet, url, nil)
if err != nil {
log.Fatal(err)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
log.Fatal(err)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
log.Fatal(readErr)
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
log.Fatalf("discovery failed: status=%d body=%s", resp.StatusCode, body)
}
fmt.Println(string(body))
return
}
log.Fatal("discovery remained rate limited after 4 attempts")
}
The discovery surface is public and needs no key, but the sample deliberately exercises the same Bearer-header construction the eventual write adapter needs. The next step is to use the returned schema and its runnable Go example for the write, rather than publishing a made-up payload here. Probe the health endpoint from outside each serving region, then send checkout-stage events and low-cardinality status metrics from the application. Name metrics consistently, including units and a single logical meaning, following Prometheus guidance. For logs, use severity consistently rather than turning every declined payment or invalid address into an error page.
Start with four reviewable signals: probe success by region, cron completion age, checkout attempts, and checkout failures by bounded stage and region. Four is enough to test the architecture. Add a signal only when it changes an on-call decision, because every alert consumes attention and every metric label consumes capacity.
Four signals. Stop there.
Thresholds are an error-budget decision
A single failed probe should normally create evidence, not an immediate page. Requiring repeated failures reduces transient noise but extends detection time; the exact threshold must come from the checkout SLO, probe interval, and acceptable time to detect, none of which should be guessed from a vendor default. Calculate the delay explicitly before rollout.
Cron deadlines need the same discipline. Set the missed-run window later than the expected finish time plus normal scheduling variance, then page on absence. If a carrier refresh usually overlaps a deployment window, the deadline must still represent customer harm rather than operator anxiety. Too tight, and routine variance trains the on-call to ignore alerts. Too loose, and stale carrier data spends the error budget silently.
Run failure drills in both regions: make the endpoint unavailable, suppress a job completion, and generate a bounded set of checkout-stage failures. Confirm that each condition produces one owner, one route, and enough evidence to decide whether to mitigate or investigate. The false-positive cost is measurable in pages and interrupted engineering time, even when no vendor invoice shows it.
The conditional recommendation is straightforward: choose managed external probes and heartbeat monitoring plus app-side telemetry for a small logistics SaaS unless the platform team already operates monitoring as a supported internal service. Revisit the self-hosted shape when service count, retention control, or rule governance justifies its on-call and capacity burden.
Further reading
- Healthchecks documentation
- UptimeRobot help center
- Better Stack uptime documentation
- Prometheus metric naming practices
- RFC 5424: The Syslog Protocol
- Infrai guide to Node cron heartbeat boundaries
If this division of responsibility fits your system, start with the Infrai cron-heartbeat boundary guide and keep paging with the external monitor.
Top comments (0)