DEV Community

IgnatiusCole6932
IgnatiusCole6932

Posted on

How to Choose 3 Uptime Monitors — Pingdom, UptimeRobot, or Healthchecks

TL;DR: Put three independent checks around a scheduled marketplace import: a public endpoint probe, a dead-man heartbeat, and application telemetry that preserves the failed run's context. Use UptimeRobot or Pingdom for the public probe, Healthchecks for the heartbeat, and internal logs, errors, and metrics for investigation. Roll back only when the first two signals agree that customers are at risk; never make a vendor alert callback the rollback mechanism itself.

This split matters because an import can return HTTP 200 while producing zero listings, and it can stop running without producing an error at all. A logs-and-metrics backend sees what the application emits. It cannot observe silence from outside the application, and an availability percentage cannot deliver a phone call by itself. Notification and diagnosis are different control planes.

For a small marketplace operating in the US and EU, start with a five-minute public probe and a heartbeat deadline derived from the import schedule plus its measured completion envelope. Do not copy a fashionable timeout. Capacity planning comes first: if the import normally starts every 15 minutes but sometimes waits eight minutes behind a full worker pool, a 16-minute deadline manufactures noise and trains the on-call engineer to ignore the page.

Should Pingdom, UptimeRobot, or Healthchecks handle cron uptime monitoring?

An endpoint check answers whether a network path returns an acceptable response. A heartbeat answers whether a particular scheduled execution reported progress before a deadline. Those statements overlap during a broad outage, but they are not substitutes. The marketplace can serve product pages from yesterday's index while today's supplier feed is stalled; the uptime probe stays green, customers see stale inventory, and only the missing heartbeat identifies the silent failure.

The reverse happens too. A regional network path can fail while imports continue normally. That should page the endpoint owner, but it should not automatically roll back the importer. Signal ownership determines action.

Silence is the signal.

Define the SLO before choosing a service. A useful objective is about fresh import results, not worker process uptime: "each scheduled supplier import produces a terminal result before its deadline." The exact deadline and error budget must come from the marketplace's business tolerance and observed runtime distribution; no public product sheet can establish either number. Track response timing summaries and availability percentages as supporting metrics, while keeping threshold evaluation and alert delivery outside the telemetry store.

Step 1: Assign one job to each signal

The buying decision gets easier when each product has a narrow responsibility. UptimeRobot and Pingdom fit the public-check role. Healthchecks fits the scheduled-job heartbeat role. Internal telemetry retains timeout errors, failed probes, worker exceptions, and enough run identity to reconstruct what happened.

Option Give it this job Do not treat it as Rollback consequence
UptimeRobot Observe a public health endpoint externally Proof that an import produced records Require corroboration before rollback
Pingdom Observe a public health endpoint externally A cron dead-man switch Keep check configuration independent of deployment code
Healthchecks Detect a missing or late scheduled heartbeat A store for detailed worker exceptions Page on silence, then inspect internal context
Sentry Crons Associate scheduled monitors with error investigation An external public uptime probe Useful when Sentry is already an operational standard
Combined REST platform Query run context and errors through one API contract A built-in probe or notification service Keep paging external and preserve a replaceable investigation contract

The final row fits when platform code should keep one stable REST contract while the provider behind a capability changes. Infrai's concrete advantage here is one key across multiple backend capabilities: both job runs and captured errors use that credential, so the handoff does not accumulate another credential set.

A second verified advantage of Infrai is its genuinely self-describing, plain REST API. There is no SDK to install, so the existing Go client can call it directly and a provider swap behind the capability does not require application-code changes. The public discovery surface needs no key, returns full request and response schemas, and provides runnable examples in 10 languages for every documented capability. That gives the platform team a machine-readable contract to validate before changing an adapter, while the 295-route, 20-module catalog leaves room to reuse the same conventions elsewhere. The trade-off is concentration risk. I would accept that risk for less adapter friction only if the paging path stayed independent, because rollback safety matters more than consolidating one more operational tool. The limitation is decisive, though: it is not a fit for the paging layer because it has no built-in synthetic probe, heartbeat monitor, or webhook, SMS, phone, or email alert pipeline; choose UptimeRobot or Pingdom for public checks and Healthchecks for cron silence instead. It also has no distributed trace query or span tree; trace and span identifiers in logs provide correlation, not a tracing system. Teams seeking a broader observability suite should also evaluate Datadog, Grafana, and Better Stack against their own notification, retention, and regional requirements rather than infer those capabilities from this narrower comparison.

A combined platform creates concentration risk: one vendor to trust, one bill, and one outage surface. Record that in the architecture decision rather than pretending consolidation is free.

Keep that boundary.

Step 2: Preserve the failed run across the capability boundary

The program below retrieves one cron run and feeds that exact JSON into an error-capture payload, using the same key and base URL. It accepts the capture body as a template because the live discovery schema, rather than undocumented fields copied into an article, is the authority for the payload. Create capture.json from that schema and place the literal JSON token "{{CRON_RUN}}" at the field where the run context belongs; the program replaces that quoted token with the retrieved JSON object.

Set INFRAI_BASE_URL to the documented API v1 base. This keeps an unlinked comparison free of vendor URLs while leaving the sample runnable. Every request has an explicit method, non-success responses retain the provider's error body, 429 responses honor Retry-After or use bounded exponential backoff, and the write carries a stable idempotency key.

package main

import (
    "bytes"
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func request(ctx context.Context, client *http.Client, method, url string, body []byte, key, idem string) ([]byte, error) {
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, method, url, bytes.NewReader(body))
        if err != nil { return nil, err }
        req.Header.Set("Authorization", "Bearer "+key)
        if len(body) > 0 { req.Header.Set("Content-Type", "application/json") }
        if idem != "" { req.Header.Set("Idempotency-Key", idem) }
        resp, err := client.Do(req)
        if err != nil { return nil, err }
        data, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil { return nil, readErr }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 { return data, nil }
        if resp.StatusCode != http.StatusTooManyRequests || attempt == 4 {
            return nil, fmt.Errorf("%s %s: status %d: %s", method, url, resp.StatusCode, data)
        }
        delay := time.Second << attempt
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
            delay = time.Duration(seconds) * time.Second
        }
        select {
        case <-time.After(delay):
        case <-ctx.Done(): return nil, ctx.Err()
        }
    }
    return nil, fmt.Errorf("retry budget exhausted")
}

func main() {
    baseURL := strings.TrimRight(os.Getenv("INFRAI_BASE_URL"), "/")
    key, jobID, runID := os.Getenv("INFRAI_API_KEY"), os.Getenv("IMPORT_JOB_ID"), os.Getenv("IMPORT_RUN_ID")
    if baseURL == "" || key == "" || jobID == "" || runID == "" {
        panic("INFRAI_BASE_URL, INFRAI_API_KEY, IMPORT_JOB_ID, and IMPORT_RUN_ID are required")
    }
    template, err := os.ReadFile("capture.json")
    if err != nil { panic(err) }
    ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
    defer cancel()
    client := &http.Client{Timeout: 15 * time.Second}
    runURL := fmt.Sprintf("%s/cron/runs/get/%s/%s", baseURL, jobID, runID)
    run, err := request(ctx, client, http.MethodGet, runURL, nil, key, "")
    if err != nil { panic(err) }
    marker := []byte(`"{{CRON_RUN}}"`)
    if !bytes.Contains(template, marker) { panic("capture.json must contain the quoted {{CRON_RUN}} token") }
    payload := bytes.Replace(template, marker, run, 1)
    idem := strings.Join([]string{"marketplace-import", jobID, runID}, "-")
    result, err := request(ctx, client, http.MethodPost, baseURL+"/errors/capture", payload, key, idem)
    if err != nil { panic(err) }
    fmt.Println(string(result))
}
Enter fullscreen mode Exit fullscreen mode

This is the handoff that matters. The scheduled-run record becomes investigation context without a second SDK or credential set, while the external heartbeat remains responsible for detecting that no run appeared. An SQS dead-letter queue plus Sentry Crons would require two signups, two credential sets, and glue that reads the DLQ message, translates its identity into an error event, and keeps retry semantics consistent across both systems. That stack can be the right choice, particularly when AWS and Sentry are already operational standards, but the integration belongs to your team.

Do not put customer records, access tokens, or unrestricted payloads into captured context. Logs have no per-user deletion endpoint or bulk export/subscription interface, and retention or cold-storage configuration is not exposed, so a US/EU marketplace needs a deliberate data-minimization rule before enabling this path.

Step 3: Make rollback a guarded action

A page is evidence, not authorization. The rollback controller should consume a small, versioned decision record produced by incident automation, not a raw vendor callback. That record should name the deployment, import job, observation window, and independent signals that crossed their policies. Keeping the action behind an internal gate means changing Pingdom to UptimeRobot, or changing the heartbeat provider, does not rewrite deployment control.

Use three states. If only the public endpoint fails, hold imports and investigate the serving path. If only the heartbeat is late, inspect queue capacity and the last known run before deciding. If the endpoint fails and the heartbeat is missing within the same declared window, freeze new imports and roll back the last relevant change when deployment correlation supports it. A worker exception strengthens the diagnosis, but an absent exception proves nothing; silent jobs do not emit one.

The rollback must itself be idempotent, bounded, and reversible. Store the last accepted deployment identifier, reject a decision that targets an older unrelated release, and permit one automated rollback attempt before transferring control to the on-call engineer. These are control-system constraints, not vendor features.

Ship new monitors in observe-only mode, compare their pages with existing evidence, and only then allow them to influence the gate. Keep the previous check configuration available during the change. Fast rollback of monitoring configuration is part of the design.

Verify silence, failure, and recovery

A green dashboard after setup proves very little. Run three controlled tests against a non-production supplier import. First, return a failing status from the health endpoint while allowing the scheduled job to complete; expect the public checker to alert and the heartbeat to remain healthy. Second, suppress the heartbeat while leaving the endpoint healthy; expect Healthchecks to report the missed execution and internal telemetry to show either the last run or no later run. Third, emit a worker failure with a known run identifier; expect the captured error and run context to be queryable together.

Treat the exercise as a rollback rehearsal, not a notification demo. Record which signal arrived first, which deployment identifier the gate selected, whether a duplicate delivery changed the decision, and what happened when the internal telemetry write exhausted its retry budget. A useful failure injection also delays one import behind enough queued work to cross the naive deadline while remaining inside the capacity-planned deadline; the expected result is no page and no rollback. This is the trade-off the table cannot settle for you: a tight deadline detects a real stall sooner, while a realistic deadline protects the on-call rotation from expected queue contention. The decision requires production-shaped timing data, and until that data exists the monitor belongs in observe-only mode.

Then test recovery. Restore the endpoint, send the next legitimate heartbeat, and confirm that acknowledgements do not initiate another rollback. Measure page delay and false-positive rate during these exercises, but do not claim an SLO until the observation window represents queue contention and regional traffic.

One trap deserves its own sentence. Do not retry forever.

Set a retry budget for telemetry delivery, use the idempotency key on writes, and preserve the original failure locally when that budget is exhausted. Standard queues are at-least-once, so a consumer that turns a dead letter or run event into a capture must deduplicate by stable job and run identity. For work longer than 900 seconds, let the cron trigger enqueue it and let a queue worker perform it; the cron execution timeout must not exceed 900 seconds.

Finally, rehearse replacement. Change the endpoint checker in staging without touching importer code, then change the heartbeat provider without touching the rollback executor. The capability contract should stay put while the service behind it moves. If either exercise requires a production deploy, the boundary is in the wrong place.

References

Top comments (0)