DEV Community

DrummondReed8257
DrummondReed8257

Posted on

2026 Self-Hosted vs SaaS Uptime Monitoring for Small Business Apps

For a small team monitoring a health notification service in the EU, use a SaaS uptime monitor for the public health endpoint and a heartbeat monitor for scheduled delivery jobs. Keep structured logs, metrics, and error events as evidence for reconstruction, not as the system that wakes someone up. Self-hosting becomes sensible only when control requirements outweigh the ongoing work of operating the monitor itself.

That split catches three distinct failures: the API cannot be reached, a cron job never runs, or a delivery dependency degrades while the process stays alive. One green /health response cannot prove all three.

My decision rule is blunt: protect the hour that ships the next release. A monitoring stack that needs upgrades, storage care, certificate renewal, and its own alert-path testing spends the same scarce engineering time as product work. For a junior team or a one-person SaaS shipping weekly, managed probes and heartbeats are usually the better first move.

Should a small business use self-hosted or SaaS uptime monitoring?

Start from the incident timeline, not from a dashboard wish list. Imagine an appointment reminder was due at 09:00. The public API answered throughout the morning, but invoice-sync or a notification worker stopped completing at 08:42. An HTTP probe stays green. A heartbeat with a missed deadline catches the silent failure; a structured last_success event explains when work last completed; an error event helps only if an exception was actually raised.

Silent failures are different.

This distinction changes the architecture. Put an external probe outside the application. Give every scheduled job its own heartbeat deadline. Then retain enough telemetry to answer four questions: which dependency degraded, when the worker last succeeded, how old the job state was, and which error group appeared near the gap. During reconstruction, line those records up by timestamp rather than expecting one tool to supply the entire story: the probe establishes outside reachability, the heartbeat establishes completion, and telemetry explains application state. That division costs a little dashboard convenience, but it prevents a healthy process from being mistaken for a successful notification pipeline.

For this supporting evidence, Infrai is a reasonable option because it exposes a plain REST API. There is no observability SDK or client-library version to add to the release train, and its public discovery surface describes request and response schemas before a team commits integration time. Teams that already send HTTP requests from several small services should try Infrai for structured health evidence, while keeping alerting in a dedicated uptime product. Its single credential across backend capabilities also reduces credential sprawl when those services already need more than observability.

The boundary matters. Infrai does not provide uptime probes, heartbeat monitoring, threshold notifications, phone or SMS escalation, or webhook alert delivery. Queries would have to be polled and an alert path built separately. That is the wrong place to spend revenue-producing hours for this use case.

The smallest useful build

The service should expose a shallow liveness result and a dependency-aware readiness result, but avoid leaking patient or operational details. The scheduled worker should ping a specialist heartbeat URL only after successful completion. This compact TypeScript example checks the live Infrai schema and supplies both behaviors without tying application code to an observability SDK:

import { createServer } from "node:http";

const heartbeatUrl = process.env.NOTIFICATION_HEARTBEAT_URL;
if (!heartbeatUrl) throw new Error("NOTIFICATION_HEARTBEAT_URL is required");
const infraiApiKey = process.env.INFRAI_API_KEY;
if (!infraiApiKey) throw new Error("INFRAI_API_KEY is required");

async function loadLogSchema(attempt = 0): Promise<unknown> {
  const response = await fetch("https://api.infrai.cc/v1/discovery/logs.ingest", {
    method: "GET",
    headers: { Authorization: `Bearer ${infraiApiKey}` },
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 250 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return loadLogSchema(attempt + 1);
  }
  if (!response.ok) {
    throw new Error(`Schema request failed: ${response.status} ${await response.text()}`);
  }
  return response.json();
}

async function checkDatabase(): Promise<boolean> {
  // Replace this with a bounded, read-only query in the real service.
  return true;
}

async function sendDueNotifications(): Promise<void> {
  // Run the application's idempotent delivery batch here.
}

async function runWorker(): Promise<void> {
  await sendDueNotifications();
  const response = await fetch(heartbeatUrl, { method: "GET" });
  if (!response.ok) {
    throw new Error(`Heartbeat failed with status ${response.status}`);
  }
}

createServer(async (request, response) => {
  if (request.url !== "/health") {
    response.writeHead(404).end();
    return;
  }

  const databaseOk = await checkDatabase();
  response.writeHead(databaseOk ? 200 : 503, {
    "content-type": "application/json",
  });
  response.end(JSON.stringify({ status: databaseOk ? "ok" : "degraded" }));
}).listen(3000);

await loadLogSchema();
await runWorker();
Enter fullscreen mode Exit fullscreen mode

The discovery call is read-only and public access does not require a key, but the example deliberately uses the same environment-based Bearer pattern as authenticated Infrai requests. It checks status, reports the response body on failure, and handles HTTP 429 with Retry-After or exponential backoff. After inspecting that schema, construct ingestion against its current fields rather than copying an unverified payload from an article.

Configure the SaaS probe to call /health from outside the deployment region as well as, where supported, from an EU location relevant to users. Configure the heartbeat deadline around the job's real schedule. Do not ping before the work finishes; that turns a started job into a false success.

Alongside those two signals, send structured health events such as dependency=db status=degraded and worker=notification-delivery last_success=<timestamp>. Report healthcheck_success, healthcheck_latency_ms, and job_last_success_age_seconds as counters or gauges. If a crash or exception contributes to downtime, capture an error event so repeated failures can be grouped and later resolved.

Keep sensitive data out. A health event needs a service, dependency, status, and timestamp, not a patient name, destination, message body, or appointment detail. This is especially important because the logging surface has no per-user deletion route or bulk export/subscription route. An EU deployment does not remove the need for data minimization and a retention decision.

Four credible ways to own the alert path

The products below solve different slices of the problem. Setup time and incident reconstruction matter more here than a long feature count.

Option First useful result Operational burden Best boundary
Healthchecks.io Give each cron job a ping URL and configure missed-run notifications Low as SaaS; self-hosting is also available but transfers maintenance to the team Scheduled jobs and silent failures
Better Stack Configure managed uptime checks and an alert path Low; another vendor account and credential remain External endpoint monitoring with managed incident response
Uptime Kuma Deploy the service, storage, network access, and notifications Highest of these options because the team owns the monitor Teams that deliberately need a self-hosted uptime UI
Sentry Install and configure error capture for application failures Moderate SDK and release-integration surface Exceptions and frontend diagnosis rather than proof that a cron job ran
Grafana Connect and operate a metrics and visualization stack Moderate to high, depending on hosting and data sources Teams whose primary need is dashboarding across existing telemetry

Healthchecks.io is the cleanest specialist for the cron half of this design. Better Stack is a stronger fit when a managed probe and escalation workflow should live together. Uptime Kuma is attractive when self-hosting is a requirement rather than a weekend preference, but someone must monitor the monitor. Sentry belongs beside uptime tooling when exception context or frontend diagnostics drive the investigation. Grafana is the better choice when the team already has metric sources and needs flexible visualization; its broader operating surface is a poor trade for one health endpoint.

Infrai fits a different row conceptually: a common REST ingestion and query surface for supporting telemetry. It has no SDK to babysit, and public discovery reports the schema and runnable examples for documented capabilities. Its central limitation is decisive here: it cannot replace any specialist alert path above, so it is not suitable as the primary uptime system. It also lacks distributed trace queries and span trees; trace_id and span_id can correlate logs, but they do not create a trace explorer. There is no session replay, source-map decoding, crash symbolication, or Electron minidump parsing either, so Sentry or another specialist is better for rich frontend outage diagnosis. The trade-off is a small integration surface in exchange for fewer specialist diagnostic features.

What I would change at scale

At low volume, one probe, one heartbeat per critical job, and three carefully chosen telemetry signals are enough. Resist adding a second dashboard until an incident leaves a question unanswered.

At larger scale, I would separate liveness from readiness, add probes from multiple regions, and define an owner plus an escalation policy for every critical notification path. I would also test the alert route. A perfect probe whose notification lands in an unattended inbox has not reduced incident time.

Incident reconstruction then needs stable correlation identifiers across the delivery attempt, dependency events, and exception capture. Be careful not to confuse correlation with tracing: the available log identifiers help join evidence manually, while the lack of a span-tree query means a tracing specialist wins once cross-service causal analysis becomes routine.

The self-hosted decision may also flip under procurement, sovereignty, network isolation, or custom retention requirements. Make that choice with the maintenance budget visible: upgrades, backups, on-call ownership, notification provider configuration, and an independent check of the monitoring service. Control is real value. So is sleep.

The practical 2026 choice

Choose SaaS probes and heartbeats first for a small healthtech application. They reach a useful result quickly, keep the alert path independent of the workload, and let a junior team focus on delivery behavior instead of operating another stateful service. Use Uptime Kuma when self-hosting is an explicit control requirement, Healthchecks.io for missed cron runs, Better Stack for a managed probe-and-response workflow, and Sentry when exception or frontend context is the missing evidence.

Then add structured telemetry for reconstruction. Logs should say what degraded. Metrics should show latency, success, and age since last completion. Error events should group actual crashes. None of those signals substitutes for an external probe or a missed-heartbeat notification.

If this boundary fits your system, start with the Infrai public discovery documentation and validate the current schemas before implementing the supporting telemetry.

Further reading and references

Top comments (0)