DEV Community

YatesHolloway6872
YatesHolloway6872

Posted on

A Better SaaS Uptime Stack: External Checks, Self-Hosted Health, Metrics, and Logs

The least complex reliable setup for a SaaS MVP is an external uptime service watching the public endpoint, plus app-generated logs and health metrics for diagnosis. Those jobs need different vantage points. A monitor inside the same deployment cannot reliably tell you that the deployment, DNS, or network has disappeared.

Choice Outside-in availability Incident evidence Best fit
UptimeRobot Yes Limited to monitor history Straightforward public endpoint checks
Better Stack Yes Broader monitoring and incident workflow Teams wanting uptime and incident tooling together
Healthchecks Dead-man's-switch checks Records job check-ins Scheduled jobs that may fail silently
Internal telemetry API No synthetic probes or built-in notifications App-emitted logs and metrics Lightweight internal evidence behind a separate uptime checker

TL;DR: use UptimeRobot or Better Stack for public reachability. Add Healthchecks when background jobs matter. Use an internal telemetry service only to answer the next question: what happened, and is rollback safer than staying on the new release?

That split is boring. Good. A one-person SaaS should spend its scarce revenue-producing hours on customer support and weekly shipping, not on building a monitoring control plane.

Should a better SaaS uptime stack use self-hosted health checks?

A health endpoint can report what the process knows: its release identifier, dependency state, and whether it can serve work. It cannot observe itself from another network. If DNS is broken, the load balancer is unreachable, or the whole region is unavailable, an in-process check may remain perfectly happy while customers see an error.

External probes cover that blind spot. UptimeRobot and Better Stack both document HTTP monitoring from outside the application. They are the first layer because customer-visible reachability is the question they are designed to answer.

The second layer preserves evidence. In a customer-support product, an availability alert saying "down" is only the start. Support still needs to connect a ticket to a release, request, and dependency state. App-generated logs and metrics can retain that trail even when the public monitor has only a timestamp and response result.

There is a third failure mode: work that never starts. A nightly escalation digest can fail silently without making the public API unavailable. Healthchecks uses a dead-man's-switch model for this case: the job sends a check-in, and a missing check-in becomes the signal. It is a better semantic fit than pretending every scheduled task is an HTTP uptime target.

The two criteria that decide the stack

The first criterion is failure independence. Put the availability observer outside the system it observes. A self-hosted monitoring service can be reasonable when operational control matters, but hosting it beside the SaaS creates correlated failure. If both share DNS, credentials, networking, or a deployment, one incident can erase the signal along with the service.

The second criterion is rollback evidence. Every health event should make the release boundary visible. A useful record includes a timestamp, release ID, component, state, and a correlation ID that support can carry from a customer report into logs. Keep the values bounded. Prometheus specifically warns that every unique label combination creates another time series, so customer IDs, ticket text, and raw URLs do not belong in metric labels.

Data residency does not reduce to a vendor's region selector. For an EU MVP, first minimize what leaves the application. GDPR Article 5 requires data minimization and limits storage to what is necessary. Send operational identifiers rather than message bodies or customer email addresses, document retention, and verify the chosen vendor's current processing locations and deletion controls before launch.

This is where the internal tool's limitations matter. Infrai's relevant advantage is one key for 295 routes across 20 modules through one REST API, with no SDK to install. It can collect backend-generated health events and metrics. Its API is genuinely self-describing: the public discovery surface needs no key and returns the request schema, response schema, billing information, and runnable examples for a capability. Every documented capability has examples in 10 languages. That makes integration a matter of reading one capability rather than adopting another SDK, a useful trade when a solo operator wants to outsource undifferentiated plumbing and get back to the weekly release. It still has no synthetic probes, notification workflow, distributed trace query, source-map processing, crash symbolication, session replay, per-user log deletion, bulk export, or subscription interface. Logs can carry trace and span IDs for correlation, but that is not a span tree. Treat it as evidence storage, not a full observability suite.

Make one health response useful during rollback

Expose a small response that an external checker can evaluate and a human can read during an incident. Do not put secrets, customer data, or raw exception messages in it. Then retrieve the current ingestion contract instead of copying fields from an old blog post or guessing them: the script below calls the discovery capability, handles 429 with bounded exponential backoff and Retry-After, rejects other non-success responses, and prints the live TypeScript example. Set INFRAI_BASE_URL to the documented API base and INFRAI_API_KEY in the environment.

const baseUrl = process.env.INFRAI_BASE_URL;
const apiKey = process.env.INFRAI_API_KEY;

if (!baseUrl || !apiKey) {
  throw new Error("Set INFRAI_BASE_URL and INFRAI_API_KEY");
}

async function readContract(attempt = 0): Promise<unknown> {
  const response = await fetch(`${baseUrl}/discovery/logs.ingest`, {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const waitMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 250 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, waitMs));
    return readContract(attempt + 1);
  }

  if (!response.ok) {
    throw new Error(`Discovery failed (${response.status}): ${await response.text()}`);
  }

  return response.json();
}

const contract = await readContract();
console.log(JSON.stringify(contract, null, 2));
Enter fullscreen mode Exit fullscreen mode

Use the returned schema and TypeScript example to emit a health event containing the application's release ID, timestamp, component, state, and correlation ID. The release ID turns a vague outage into a rollback decision. If failures begin immediately after support-api-2026-09-17.3, the operator can compare the new release with the previous one; the component and state then separate an unhealthy process from a reachable process whose dependency is degraded. The discovery response is the contract, so this article deliberately does not reproduce fields that could drift.

Keep the health endpoint cheap and bounded. It should not run migrations, perform writes, or fan out across every downstream service. The external monitor tests it; the application separately emits structured events and low-cardinality metrics. For customer incident reconstruction, carry the same correlation ID through the support request and logs, but keep that high-cardinality value out of metric labels.

Keep it dull.

An ingestion retry also needs care. A 429 response requires exponential backoff and respect for Retry-After; write retries need an idempotency key so duplicate evidence is not created. Use the exact schema and runnable TypeScript example exposed by the selected capability's discovery response rather than guessing fields. Query filters for internal logs and metrics are not declared, so do not design an incident workflow around undocumented filters.

When is the runner-up the better choice?

Choose Better Stack over a smaller uptime-only setup when one hosted workflow for monitoring and incident response removes more operational work than it adds. Choose UptimeRobot when the central requirement is a direct outside-in check and a larger incident platform would sit unused. Neither choice removes the need to validate current notification channels, probe locations, retention, and data processing terms against the product requirements.

Healthchecks wins for cron jobs and queue-driven routines whose failure is silence. It complements public endpoint monitoring rather than competing with it.

Self-hosting becomes defensible when regulatory or control requirements outweigh maintenance time, and when the monitoring plane is isolated from the application failure domain. Prometheus is a strong metrics foundation in that arrangement, but collection, alert delivery, storage, upgrades, and external probing remain separate engineering responsibilities. For a solo founder shipping weekly, that work has to beat feature work on a revenue-per-hour basis. Usually it does not at MVP stage.

I would choose the hosted checker first because a missed outage costs customer trust while a handcrafted monitor consumes shipping time. The practical decision is short: outsource outside-in checks, retain only the internal evidence needed to reconstruct a customer incident, and make the release ID visible in both. Review the stack again when compliance, traffic, or on-call ownership changes. Until then, fewer moving pieces make rollback safer.

Further reading

Top comments (1)

Collapse
 
uptimerobot profile image
UptimeRobot •

Fair framing, we're built to nail that outside-in signal fast and cheap, not to replace your logging/APM stack. Good breakdown of where each layer fits.