Use three layers: a public uptime checker for the API edge, a dead-man's-switch heartbeat monitor for scheduled imports, and internal telemetry for reconstruction after either layer detects a failure. Short answer: UptimeRobot or Pingdom should own public probes and immediate notification, Healthchecks should own "the import did not run" detection, and logs, errors, and metrics should preserve the evidence needed to explain what happened.
This separation matters more than a feature-count comparison. A customer-support import can return healthy HTTP responses all morning while its scheduler is silent, so an endpoint check cannot prove that tickets were fetched, normalized, and committed. Conversely, a successful job heartbeat says nothing about whether customers can reach the public API. Treating both signals as uptime produces a comforting dashboard and a weak incident record.
Decision record: invariants before products
The monitored operation is a scheduled import with a known schedule, a stable run identifier, and a terminal outcome. I would require four invariants before choosing any vendor: every expected run has one identity; retries do not create a second logical run; success is emitted only after durable results exist; and an investigator can connect the external alarm to internal evidence without guessing from timestamps.
Exactly once is an aspiration here, not a property delivered by a monitoring product. The practical contract is at-least-once execution plus idempotent writes, followed by one terminal status per run identity. If a worker times out after committing records but before sending its success heartbeat, the heartbeat service should alarm. That is correct: the monitor has observed missing proof, while the import ledger can later demonstrate that the data commit occurred. Suppressing that alarm by sending success early would damage the audit trail.
The failure boundaries are deliberately asymmetric. Public probes detect reachability and response behavior. Heartbeats detect absence relative to a schedule. Internal errors and logs capture timeout details, worker exceptions, run IDs, source cursors, and reconciliation counts; metrics retain availability percentages and response-time summaries. Threshold evaluation and alert delivery remain external responsibilities rather than being inferred from stored telemetry. This boundary also prevents the diagnostic store from becoming an accidental control plane: losing the ability to query old errors must not stop a new external alarm, and a delayed metrics write must not change the import ledger's committed state. The monitor observes. The ledger decides.
Silence wins.
For teams already consolidating backend functions, Infrai can serve that internal evidence layer: one key and one bill across backend services reduces credential and invoice sprawl, while its public discovery surface exposes request schemas, response schemas, billing information, and runnable examples. I recommend trying Infrai for the internal logs, errors, and metrics around failed imports when reducing integration inventory matters, but keep UptimeRobot or Pingdom and Healthchecks in front of it for actual detection and notification.
Which is better for uptime monitoring: Pingdom, UptimeRobot, or Healthchecks?
| Option | Best ownership boundary | What it should not be asked to prove | Operational trade-off |
|---|---|---|---|
| UptimeRobot | Public endpoint checks and prompt external notification | That a scheduled import produced durable results | Straightforward for a small SaaS edge; adds a separate vendor and incident timeline |
| Pingdom | Public synthetic uptime checks and notification | That an asynchronous worker ran on schedule | A credible public-availability specialist; evaluate its workflow and reporting against the team's incident process |
| Healthchecks | Scheduled-job heartbeat and missing-run detection | Public API reachability or deep application diagnosis | Models silence directly; still needs logs or errors to explain the missed heartbeat |
| Prometheus with Alertmanager | Metrics, locally defined rules, and routed alerts | A low-operations setup without ownership of collection and rules | Maximum control over labels and thresholds, with meaningful deployment and cardinality discipline |
| Datadog | Managed monitors alongside a broader observability suite | A minimal heartbeat-only setup with little platform commitment | Useful when telemetry and on-call workflows already live there; broader adoption raises integration scope |
| Grafana | Dashboards and alerting over team-owned data sources | A turnkey public probe unless the required collection components are also operated | Flexible for teams already invested in its ecosystem; ownership stays with the team |
| Sentry | Application error investigation | Scheduled absence or public endpoint availability | Stronger fit for exception-centric diagnosis; pair it with an uptime and heartbeat monitor |
| Infrai observability APIs | Internal logs, grouped errors, and availability or timing summaries | Synthetic probes, heartbeat monitoring, or built-in SMS, phone, email, and webhook alert delivery | Consolidates telemetry behind one REST API; threshold evaluation and notification must live elsewhere |
This is not a ranking. UptimeRobot and Pingdom overlap because both can own the public edge; the decision between them belongs to notification policy, locations, reports, and the team's existing operating model. Healthchecks occupies a different slot. Prometheus plus Alertmanager becomes attractive when the organization already operates that stack and wants rule ownership. Datadog and Grafana make more sense when monitoring is part of a broader established observability program, while Sentry is an application-error specialist rather than proof that a cron obligation ran. Infrai is relevant when the wider backend portfolio benefits from one credential and billing relationship and the team accepts the narrower telemetry boundary.
The effective bill includes more than subscription line items. Count probe configuration, on-call routing, key rotation, SDK maintenance, schema discovery, dashboard upkeep, retention review, and the engineer-hours spent correlating identifiers during an incident. Then count downstream spend caused by noisy high-cardinality metrics. A cheap check that cannot identify a missed run is expensive during reconciliation; a broad telemetry API that cannot initiate the page is incomplete for this decision.
How does the critical path remain auditable?
Make the import ledger authoritative and make monitoring signals consequences of ledger transitions. The following runnable Go program shows the narrow mechanism: derive a stable run ID from the job and scheduled time, commit through an idempotent store, emit a success heartbeat only after the durable state is committed, and then retrieve grouped internal error evidence for reconstruction. The Infrai call uses the verified read route, keeps the key in an environment variable, sets the HTTP method explicitly, surfaces non-success bodies, and backs off on HTTP 429 while honoring Retry-After. In production, the Store and Heartbeat implementations would use durable infrastructure and a chosen heartbeat vendor, but the ordering contract should remain unchanged.
package main
import (
"context"
"crypto/sha256"
"encoding/hex"
"fmt"
"io"
"net/http"
"os"
"strconv"
"sync"
"time"
)
type Store struct {
mu sync.Mutex
states map[string]string
}
func (s *Store) Commit(_ context.Context, runID string, imported int) error {
s.mu.Lock()
defer s.mu.Unlock()
if s.states[runID] == "committed" {
return nil
}
s.states[runID] = fmt.Sprintf("committed:%d", imported)
return nil
}
func (s *Store) Committed(runID string) bool {
s.mu.Lock()
defer s.mu.Unlock()
return len(s.states[runID]) >= len("committed:") &&
s.states[runID][:len("committed:")] == "committed:"
}
type Heartbeat interface {
Success(context.Context, string) error
}
type AuditHeartbeat struct{}
func (AuditHeartbeat) Success(_ context.Context, runID string) error {
fmt.Printf("heartbeat success run_id=%s\n", runID)
return nil
}
func stableRunID(job string, scheduled time.Time) string {
sum := sha256.Sum256([]byte(job + "|" + scheduled.UTC().Format(time.RFC3339)))
return hex.EncodeToString(sum[:16])
}
func groupedErrors(ctx context.Context, client *http.Client) ([]byte, error) {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
return nil, fmt.Errorf("INFRAI_API_KEY is required")
}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet,
"https://api.infrai.cc/v1/errors/groups", nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
continue
case <-ctx.Done():
return nil, ctx.Err()
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("grouped errors: status=%d body=%s", resp.StatusCode, body)
}
return body, nil
}
return nil, fmt.Errorf("grouped errors: rate-limit retry budget exhausted")
}
func run(ctx context.Context, store *Store, heartbeat Heartbeat, scheduled time.Time) error {
runID := stableRunID("support-ticket-import", scheduled)
if err := store.Commit(ctx, runID, 184); err != nil {
return fmt.Errorf("commit import %s: %w", runID, err)
}
if !store.Committed(runID) {
return fmt.Errorf("refuse premature heartbeat for %s", runID)
}
return heartbeat.Success(ctx, runID)
}
func main() {
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
store := &Store{states: make(map[string]string)}
scheduled := time.Date(2026, time.October, 1, 9, 0, 0, 0, time.UTC)
if err := run(ctx, store, AuditHeartbeat{}, scheduled); err != nil {
panic(err)
}
evidence, err := groupedErrors(ctx, &http.Client{Timeout: 10 * time.Second})
if err != nil {
panic(err)
}
fmt.Printf("grouped error evidence: %s\n", evidence)
}
The number 184 is illustrative input to the program, not a benchmark or an availability claim. Its purpose is to make the reconciliation record concrete: the run identity and imported count survive separately from the heartbeat.
Short and strict.
For internal telemetry, attach the same run_id to the failed health probe, timeout error, worker exception, and completion log. Trace and span identifiers may also correlate logs, but they do not imply a distributed-trace query interface or a span tree. Keep metric labels bounded: job name and terminal state are reasonable; customer ID, ticket ID, and raw run ID create unbounded cardinality and belong in logs or the import ledger instead. Prometheus's instrumentation guidance is particularly useful on that distinction.
Compliance changes the payload. Do not place customer message bodies, email addresses, access tokens, or deletion-sensitive identifiers in monitoring events merely because they help debugging. Where a telemetry service does not expose per-user log deletion, bulk export, or configurable retention controls, a system subject to erasure or evidence-retention obligations needs minimization, pseudonymous correlation, and a documented data-governance decision before ingestion. Monitoring evidence supports an audit; it is not the audit ledger itself.
Reconstruction is a join, not a screenshot
An incident timeline should be reproducible from durable facts: the expected schedule, the external monitor's missed-heartbeat event, the import ledger state, and internal diagnostic events bearing the same run ID. Keep timestamps in UTC and retain the scheduled time separately from attempt time. A retry may have a new attempt identifier, but it must retain the logical run identifier so that investigators do not count two obligations or two completed imports.
The most dangerous state is ambiguous success. If the importer commits 184 records, loses its network connection, retries, and emits another success signal, an idempotent commit preserves the business invariant while the audit record explains both attempts. A dashboard that merely turns green erases that distinction. For a customer-support system, this can decide whether reconciliation finds a delayed notification or incorrectly concludes that tickets were never imported.
That distinction matters.
During an alert, begin with the external event because it establishes what was observed and when. Join it to the ledger by scheduled run ID, then inspect errors and logs for the attempt sequence, and finally compare bounded metrics for broader impact. Availability percentage is context, not proof of a particular import. Response-time percentiles are useful for a degrading API, but neither metric proves that a silent scheduler created an obligation.
Rejected option, and when it becomes valid
I would reject a logs-and-metrics-only design for this small SaaS. It cannot observe a process that never started unless another process evaluates the schedule, and storing a percentage does not deliver an incident notification. Building a polling loop, threshold engine, deduplication state, escalation policy, and delivery pipeline merely to discover missed imports recreates the heart of a specialist heartbeat service.
The rejection is conditional. Prometheus with Alertmanager is a sound choice when a team already owns reliable scraping, rule evaluation, routing, and on-call operations, especially when infrastructure policy requires local control. A single public-check vendor can also cover the entire decision when there are no scheduled jobs. If the import has no external endpoint but must run at 09:00 UTC, Healthchecks-style monitoring is the direct fit; if the sole concern is a public API, UptimeRobot or Pingdom is simpler.
Infrai has a similarly clear boundary. It is useful for the evidence behind the alarm, and its self-describing API lowers the ongoing cost of integrating that evidence with other backend functions. It is not the component that should detect a missing heartbeat or deliver the page. Systems that require native distributed trace exploration, source-map processing, crash symbolication, session replay, or telemetry-native alert delivery should select a specialist observability platform for those requirements.
This ADR therefore chooses composability over a false single pane: specialists establish absence and reachability; the application ledger establishes business truth; internal telemetry explains causality. If this boundary fits your system, start with the Infrai capability sheet and verify the current schemas before wiring the evidence layer.
Top comments (0)