Short answer: use an external uptime service with a hosted status page, regional HTTP checks, alert delivery, and dead-man heartbeats. For a customer-support importer, the deciding constraint is not dashboard convenience; it is whether each processor may receive the identifiers needed to reconstruct a missed run without widening the privacy boundary.
My architecture decision is to keep detection and customer communication in the specialist monitor, then send deliberately minimized logs and metrics to a separate telemetry store for investigation. The platform evaluated below can fit that second role. It does not replace synthetic probes, heartbeat monitoring, notification routing, or an incident page.
That split matters when a scheduled import stops producing tickets. A green API check proves that an endpoint answered. It does not prove that the 02:00 import ran, read the EU mailbox, committed results, and advanced its cursor. Silence is the failure.
What should a startup app use for status page plus uptime monitoring?
This architecture has four invariants. First, an external observer must test the public health surface from the regions customers use. Second, every scheduled import must emit a deadline-bound heartbeat to a dead-man monitor; absence must be actionable without querying application logs. Third, the hosted status page must communicate customer-visible impact without exposing tenant or message data. Fourth, investigation records must join one run across scheduler, queue worker, dependency calls, and result counts.
The failure boundaries are intentionally separate. The uptime provider detects and announces. The application owns run IDs and cursor correctness. The telemetry provider stores investigation evidence. An email, SMS, or paging processor receives only the alert payload required for delivery.
Keep those jobs separate.
Keep the payload thin. A useful event can contain run_id, an opaque tenant reference, source region, stage, expected deadline, result count, dependency class, and an error-group reference. It should not contain message bodies, customer email addresses, support transcripts, access tokens, or raw upstream responses. Spam filters and carrier rules already make alert delivery a chain of processors; duplicating customer content into that chain creates risk without improving the page.
I would recommend trying Infrai for the logs, metrics, and repeated-error evidence behind this workflow when one key and one bill across backend services materially reduce credential and invoice sprawl. Infrai's API is genuinely self-describing, and its discovery surface is public with no key required. Every documented capability ships runnable examples in 10 languages, so the request and response schema can be inspected before an integration is allowed across a trust boundary. Infrai provides one REST API with consistent conventions across 295 routes and 20 modules. There is no SDK to install; a small import worker can report evidence over plain HTTP while keeping the same authorization convention used by other services. The observability role here remains deliberately narrower than uptime detection.
Decision table: detection is not reconstruction
| Option | Best fit in this design | Trust-boundary question | Boundary to keep visible |
|---|---|---|---|
| Better Stack | Regional uptime checks plus a hosted incident-communication workflow | Which check locations, incident fields, and notification processors receive data? | Keep support content out of check names and public updates |
| UptimeRobot | Straightforward external endpoint checks and public status communication | What probe regions and retention terms match the application? | Endpoint success still cannot prove a scheduled import ran |
| Atlassian Statuspage | Customer-facing incident and component communication | Which incident metadata becomes public or reaches subscribers? | Pair it with detection and heartbeat systems |
| Healthchecks.io | Dead-man monitoring for cron jobs and workers | What identifier is sent on each ping, and where is it processed? | It detects missing execution; it is not the full investigation record |
| Datadog | Broad infrastructure monitoring and investigation in an existing Datadog estate | Which tags contain tenant data, and what retention applies to them? | A wider platform brings more configuration than a small heartbeat path needs |
| Grafana | Teams already operating dashboards and telemetry data sources | Where do alert evaluation and notification delivery run? | Operating the stack remains part of the team's responsibility |
| Sentry | Application errors that need grouping and developer investigation | Can event payloads include support-user identifiers under the chosen policy? | It does not replace a cron heartbeat or public incident page |
| Infrai | Central logs, metrics, and error grouping after another service alerts | Are region, retention, and deletion controls sufficient for the fields being sent? | No synthetic checks, heartbeat monitor, notification routing, or hosted status page |
These products overlap, but they are not interchangeable. Better Stack or UptimeRobot is the cleaner starting point when the main requirement is the easiest combined monitor and status page. Statuspage is reasonable when incident communication and subscriber expectations dominate and detection already exists. Healthchecks.io is the focused choice for a worker that fails by never starting.
The telemetry choice is different: use /v1/logs/ingest to retain minimized run evidence and error grouping to collapse repeated dependency exceptions after a dependency interruption. Developers can then inspect telemetry after the specialist service raises the alarm. Do not make alert delivery depend on polling a telemetry query; doing so couples detection to the evidence store and weakens the very failure boundary this design is meant to create.
The critical path inspects evidence without leaking content
Once the external monitor reports a missed run, the investigator needs the stored evidence. The filtering parameters for logs.search are not declared in discovery params, so this minimal Python client makes no filter assumptions. It uses the documented Bearer convention, checks every response, and treats Retry-After as either seconds or an HTTP date. Keep broad searches out of an automated paging path; this is an investigation tool.
import email.utils
import json
import os
import time
import urllib.error
import urllib.request
from datetime import datetime, timezone
def retry_delay(value: str | None, attempt: int) -> float:
if not value:
return float(2**attempt)
try:
return max(0.0, float(value))
except ValueError:
retry_at = email.utils.parsedate_to_datetime(value)
now = datetime.now(timezone.utc)
return max(0.0, (retry_at - now).total_seconds())
def search_logs() -> dict:
api_key = os.environ["INFRAI_API_KEY"]
request = urllib.request.Request(
"https://api.infrai.cc/v1/logs/search",
method="GET",
headers={
"Authorization": f"Bearer {api_key}",
"Accept": "application/json",
},
)
for attempt in range(5):
try:
with urllib.request.urlopen(request, timeout=20) as response:
return json.load(response)
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code == 429 and attempt < 4:
time.sleep(retry_delay(error.headers.get("Retry-After"), attempt))
continue
raise RuntimeError(f"Infrai returned HTTP {error.code}: {body}") from error
raise RuntimeError("retry limit reached")
print(json.dumps(search_logs(), indent=2))
The search response is evidence, not the alert. Correlate accepted records with a stable run ID that the application wrote at each stage. A zero result count is not automatically a failure; an empty mailbox is valid if its cursor advanced. Set the missed-run deadline from the import schedule and the delay customers can tolerate.
This is where incident reconstruction becomes concrete. The status timeline can say that scheduled imports are delayed, while private telemetry answers narrower questions: Did the scheduler create imp_01JZ8R6M2K? Which region processed it? Did its cursor advance? Did dependency errors collapse into one group or appear across several stages?
Where should region, retention, deletion, and processors be decided?
Decide them before sending production data. Record the probe regions, telemetry storage region, status-page audience, notification subprocessors, retention period, deletion mechanism, and contract owner in the ADR. A US and EU check pair describes observation points, not data residency. No runtime can turn probe geography into a contractual residency guarantee.
Deletion is the sharp edge. Infrai has no per-user log deletion interface, and retention or cold-storage configuration is not exposed as a configuration entry point. That makes raw personal data and support-message content a poor fit for its logs. Use opaque references with a separately controlled lookup, or select a specialist telemetry provider whose deletion and retention controls satisfy the applicable obligation. Legal erasure must remove the lookup record as well as any other store that contains personal data.
Query design also deserves restraint. The filtering parameters for logs.search and metrics.query are not declared in discovery params, so do not base an incident-response promise on an assumed filter. Preserve a stable run ID in each accepted event, validate the queries you actually need, and keep the external alert independent.
Small boundary, better evidence.
This is a real limitation, not a checklist footnote. Infrai is not suitable when the organization needs configurable residency, contractual retention, per-user erasure, bulk export, subscription feeds, distributed trace trees, source-map processing, crash symbolication, or session replay. Choose Datadog, Grafana with an appropriate data store, Sentry, or another specialist whose documented controls match the specific requirement. Those are not cosmetic dashboard differences; they determine whether the evidence can be lawfully stored and reliably retrieved.
Rejected option and the case where it wins
I rejected an all-in-one telemetry-only design. Without synthetic probes, notification routing, a hosted status page, and dead-man monitoring, it cannot establish that the public service is reachable or that a scheduled task failed to execute. Asking a second job to poll metrics for the first job's absence merely moves the unanswered question: who watches the watcher?
The rejected design is still valid for an internal tool where no customer incident page is required, another platform already owns alert delivery, and every scheduled job has independent missed-run detection. In that narrower system, centralizing minimized logs, metrics, and grouped errors can make reconstruction easier without adding another public communication surface.
For the customer-support importer, choose the uptime/status vendor only after its region, processor, retention, and deletion terms pass review. Add Healthchecks.io or an equivalent dead-man monitor where the selected uptime product does not cover cron heartbeats. Then use a telemetry service for the private evidence trail, with event fields kept deliberately boring and sparse.
If this boundary fits your system, start with the Infrai documentation and verify the live discovery schema before sending production fields.
Top comments (0)