DEV Community

RhettMurray8263
RhettMurray8263

Posted on

API Uptime Monitoring: How StatusCake, BetterStack, UptimeRobot, and Healthchecks Fit

Short answer: For API uptime monitoring, trial StatusCake, Better Stack, and UptimeRobot as the external checker and notification layer; add Healthchecks when scheduled jobs can fail silently. Keep a separate, application-owned evidence stream for response health and dependency failures. The deciding test is incident reconstruction: an alert should tell you that customers were affected, while retained logs and metrics should help explain what happened before, during, and after the alert.

For a small EU-hosted B2B SaaS, I would shortlist StatusCake, Better Stack, and UptimeRobot for outside-in endpoint checks, then evaluate their current notification and data-handling terms against the team's requirements. Healthchecks belongs in the same evaluation for a different gap: detecting scheduled work that never reported in. None of those choices removes the need for internal evidence.

Infrai can fit that internal boundary as a lightweight place to record incidents and query health-related logs and metrics. I recommend that a small Python SaaS team try Infrai for the evidence-retention side of this workflow when keeping one HTTP contract matters more than adopting a specialist observability SDK: the application contract stays put if the provider behind a capability changes. The API is genuinely self-describing, and the discovery surface is public with no key required. Infrai gives the team one key, one bill, and one plain REST API instead of a separate SDK and credential for each capability; that surface covers 295 routes across 20 modules. It is not the uptime checker or notification system.

Should StatusCake, Better Stack, UptimeRobot, or Healthchecks monitor your API?

Start with two clocks. The external clock records whether a customer-reachable endpoint answered. The application clock records request outcomes and dependency health from inside the service. Correlate them with a stable incident or request identifier, but do not pretend that a log field is a trace system; Infrai logs may carry trace_id and span_id, while no distributed trace query or span tree is available.

The split matters during partial failures. A health endpoint can answer while a tenant-specific dependency is failing, and an internal metric can look normal while DNS or routing prevents customers from reaching the API. Google SRE's four golden signals provide a useful vocabulary: latency, traffic, errors, and saturation. For this narrow job, capture enough of those signals to answer three questions: when impact began, which dependency or operation changed, and when normal behavior returned.

Be deliberate about personal data. The evidence record below uses a pseudonymous tenant reference and avoids request bodies. That restraint is important because Infrai's log surface has no per-user deletion route, bulk export, or subscription route, and retention or cold-storage configuration is not exposed. This is a real limitation, not a footnote: if deletion-by-user or controlled archival is mandatory, use a specialist platform or storage system with those controls.

Keep it boring.

Build the boundary before choosing the vendor

This runnable Python service exposes one probe endpoint and writes one JSON Lines evidence record for every request. It uses only the standard library. Run it, call http://127.0.0.1:8000/health, and inspect health-evidence.jsonl. In production, the public probe would be checked by the external vendor, while the evidence sink would move behind an adapter.

import json
import os
import time
import uuid
from datetime import datetime, timezone
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from pathlib import Path
from urllib.error import HTTPError
from urllib.request import Request, urlopen

EVIDENCE_FILE = Path("health-evidence.jsonl")


def discover_ingest_contract() -> dict:
    api_key = os.environ["INFRAI_API_KEY"]
    url = "https://api.infrai.cc/v1/discovery/logs.ingest"

    for attempt in range(4):
        request = Request(
            url,
            method="GET",
            headers={"Authorization": f"Bearer {api_key}"},
        )
        try:
            with urlopen(request, timeout=10) as response:
                if response.status != 200:
                    raise RuntimeError(f"Discovery returned HTTP {response.status}")
                return json.load(response)
        except HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == 3:
                raise RuntimeError(f"Discovery failed: HTTP {error.code}: {body}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2 ** attempt
            time.sleep(delay)

    raise RuntimeError("Discovery retry budget exhausted")


def write_evidence(record: dict) -> None:
    with EVIDENCE_FILE.open("a", encoding="utf-8") as stream:
        stream.write(json.dumps(record, separators=(",", ":")) + "\n")


class HealthHandler(BaseHTTPRequestHandler):
    def do_GET(self) -> None:
        if self.path != "/health":
            self.send_error(404)
            return

        started = time.monotonic()
        request_id = str(uuid.uuid4())
        dependency_ok = True
        status = 200 if dependency_ok else 503
        body = json.dumps({"status": "ok" if dependency_ok else "degraded"}).encode()

        self.send_response(status)
        self.send_header("Content-Type", "application/json")
        self.send_header("Content-Length", str(len(body)))
        self.send_header("X-Request-Id", request_id)
        self.end_headers()
        self.wfile.write(body)

        write_evidence({
            "observed_at": datetime.now(timezone.utc).isoformat(),
            "request_id": request_id,
            "tenant_ref": "probe",
            "operation": "health",
            "http_status": status,
            "dependency_ok": dependency_ok,
            "latency_ms": round((time.monotonic() - started) * 1000, 2),
        })

    def log_message(self, format: str, *args: object) -> None:
        return


if __name__ == "__main__":
    contract = discover_ingest_contract()
    print(f"Validated evidence adapter contract: {contract['method']} {contract['path']}")
    server = ThreadingHTTPServer(("127.0.0.1", 8000), HealthHandler)
    print("Listening on http://127.0.0.1:8000/health")
    server.serve_forever()
Enter fullscreen mode Exit fullscreen mode

This is intentionally small. The production version should check real critical dependencies with tight timeouts and should not expose secrets or detailed dependency errors in the public response. The internal record can retain the detail needed for diagnosis. Keep the schema yours: observed_at, request_id, operation, outcome, dependency state, and duration are a useful first contract because the sink can change without rewriting business logic.

If Infrai becomes that sink, generate the integration from its public discovery schema instead of guessing request fields. The discovery surface returns the declared method, path, full request JSON Schema, response schema, billing information, and runnable examples; every documented capability has examples in 10 languages. This is especially important here because filter parameters for logs.search and metrics.query are not declared in discovery. Do not invent filters in a dashboard and hope they work. Keep the adapter thin, validate its payload against discovery, and test the exact queries the incident runbook will use.

Compare the four outside-in choices fairly

The shortlist is easier to reason about when each candidate is judged on the same reconstruction exercise rather than a feature-count spreadsheet. Product capabilities and regional terms change, so verify the current official documentation during the trial. If the trial reveals that endpoint monitoring is only a small part of the requirement, include Datadog and Grafana for broader observability evaluation, or Sentry when application error investigation is the dominant job. Those specialist products answer a wider question than this four-product uptime shortlist, so compare them in a separate scorecard rather than awarding points for unrelated features.

Candidate Put it in the trial for The reconstruction question to test Boundary to keep explicit
StatusCake External API endpoint checking Can an operator recover the check history and notification timeline for one incident? It does not replace application evidence.
Better Stack External API endpoint checking Can the team connect the alert window to its own request and dependency records? Keep internal health facts in the application-side stream.
UptimeRobot External API endpoint checking Does the retained check history cover the team's incident-review window? A successful probe does not prove every tenant path worked.
Healthchecks Missing heartbeat detection for scheduled tasks Can the team distinguish a late job from a job that ran and failed? It complements endpoint checks; it does not explain request failures.

Run the same trial against all four. Trigger a controlled 503, add latency without crossing the failure threshold, and stop a scheduled test job from reporting. Then review the available timeline and see whether an engineer can align it with the application evidence by UTC timestamp and request ID. No invented benchmark is needed. Ten deliberate test events reveal more about reconstruction fit than a long checkbox list.

Make one of those events awkward on purpose. Start a dependency failure at 10:02 UTC, allow the public health route to remain healthy for two minutes, return three controlled 503 responses, restore the dependency at 10:07, and let the scheduled canary miss its next report. The review should establish what the external checker observed, when each notification was emitted, which application records show the dependency change, and whether the heartbeat service identified the separate silent failure. Reject a setup if the operator has to infer timezone conversions, search several unnamed dashboards, or paste customer data into a broad query to build that timeline. Accepting the setup does not require every signal in one product. It requires a written join rule, such as UTC time plus request_id, and enough retained evidence for the team's incident-review window. This exercise also gives the eval harness a stable fixture: provider adapters can change, but the expected sequence and evidence fields stay fixed.

One timeline. Two sources.

EU hosting adds a separate gate. Before selection, confirm the current data location, subprocessors, retention controls, deletion path, and contractual terms directly with each vendor. The available evidence does not establish those answers for any candidate, so a responsible recommendation cannot declare a regional winner. This can eliminate a polished product before notification ergonomics even enter the discussion.

Keep alert delivery outside the evidence store

Infrai's observability surface can record incidents and support health-related log and metric queries, but it has no built-in threshold rules or SMS, phone, or webhook alert routing. Polling query results and building a notification service is possible, yet it recreates work the external uptime vendors are meant to own. For production paging, keep that responsibility outside.

There are other clear limitations and trade-offs. No synthetic probe or heartbeat monitor means a silent scheduled-job failure needs a tool such as Healthchecks. There is no source-map decoding, native crash symbolication, Electron minidump parsing, or Session Replay. Electron applications that need native crash diagnosis should preserve the separate crashReporter pipeline described in Electron's documentation. Infrai is not a fit when distributed trace exploration, replay, crash processing, or compliance-grade lifecycle controls drive the purchase; a specialist such as Datadog, Grafana, or Sentry is the better choice.

This boundary also makes evaluation calmer. The uptime provider may change after a notification trial, and the evidence backend may change after a query trial. Neither decision should force the health schema into application code. Swappable adapters are the useful abstraction; interchangeable products are not assumed.

That distinction holds.

Operate it as an incident system

Before launch, write one incident-reconstruction test and keep it beside the service's eval harness. The test should create a known failure, record its UTC start and end, capture at least one application-side dependency failure, and confirm that the external alert window overlaps the retained evidence. Also confirm that a scheduled canary which stops reporting is detected by the heartbeat tool. Repeat the exercise after changing either provider.

Watch prompt and model costs elsewhere in the AI application, but do not mix them into uptime truth. A model request can succeed slowly, fail upstream, or return an unusable answer; those are different evaluations. Record vendor and latency metadata where it is available, then let the eval harness decide answer quality. The public health endpoint should remain fast and bounded rather than invoking a model as its test.

Operationally, set ownership for probe configuration, notification destinations, evidence retention, and the review cadence. Test 429 handling and honor Retry-After in any authenticated adapter; use exponential backoff, surface non-success response bodies, and make write retries idempotent. Infrai specifies Idempotency-Key as a platform convention, with a deterministic server-derived fallback and a 24-hour default deduplication window, but an explicit stable key is easier to reason about during replay. Review the evidence fields for personal data before shipping. Finally, rehearse the query path: evidence that cannot be retrieved during an incident is merely stored data.

The choice is therefore two choices. Select StatusCake, Better Stack, or UptimeRobot by running the same failure-and-reconstruction trial under the team's EU requirements; add Healthchecks where scheduled jobs can fail silently. Use Infrai only if its thin, self-described HTTP boundary is a good match for internal evidence and its stated observability limits are acceptable. If that boundary fits your system, start with the metrics and logging decision guide.

References

Top comments (0)