DEV Community

SeraphinaLyn7139
SeraphinaLyn7139

Posted on

Self-Hosted vs SaaS Uptime Monitoring for Small Go Apps (EU Cron Alerting)

TL;DR: Use SaaS endpoint monitoring plus a managed cron heartbeat for a small edtech application in an EU region. That is the least complex way for a junior team to detect both a dead service and a nightly search pipeline that never finished. Keep structured logs, metrics, and grouped errors as evidence for incident reconstruction; they support the page, but they are not the paging system.

The page arrives at 02:17. The API health check is red, yet the process responds, and no obvious exception explains why tomorrow's course-search index is stale. The on-call engineer needs answers to three different questions: can an outside client reach the application, did the scheduled pipeline complete, and where did its last run stop? One green /health response cannot answer all three.

Three jobs. Three signals.

An external probe establishes reachability. A heartbeat catches the silent absence of an expected job. Application telemetry reconstructs the run. Give paging to the first two and explanation to the third, with thresholds tied to the freshness SLO rather than to whichever metric happens to be easy to collect.

What should have fired before the page?

The public endpoint page is the last signal in this chain, not the first. Work backward. A nightly import can stop producing fresh search data while the web process remains healthy, so an HTTP probe has no evidence that scheduled work happened. A heartbeat deadline should fire when the expected completion signal is absent; Healthchecks.io is designed around that cron-monitoring pattern and documents the ping lifecycle directly.

Before that deadline, a dependency event might say dependency=db status=degraded. The worker should also report job_last_success_age_seconds. A rising gauge shows that freshness is eroding while the application still answers requests. Counters and gauges such as healthcheck_success and healthcheck_latency_ms add trend context, while an error event gives a crashing worker a group that can later be resolved.

Do not compress this into a synthetic "healthy" bit. The service-level objective should name the user outcome: searchable course data is fresh after the nightly processing window. Endpoint availability is related, but it is a different promise with a different failure mode.

Absence matters.

The capacity-planning reflex matters here. A monitor consumes storage, network paths, upgrades, certificates, and on-call attention even when its own CPU graph looks trivial. For a two-engineer rotation, that human capacity is usually the binding limit. I would not add another stateful service to that rotation without a placement or control requirement strong enough to justify every upgrade, backup check, and monitor-the-monitor decision.

No signal is enough.

Instrument the transitions, not the happy ending

Emit one structured event at each pipeline state transition, preserve a scheduled run identifier across the events, and calculate job age from the last confirmed success. The Go sender below accepts a health event that the application has already serialized and validated against the current discovery schema. That avoids duplicating a remote schema in handwritten structs.

Set OBSERVABILITY_BASE_URL, INFRAI_API_KEY, PIPELINE_RUN_ID, and HEALTH_EVENT_JSON in the environment. The example uses one verified route, sets the HTTP method explicitly, checks every response, supplies an idempotency key, and honors Retry-After on a 429. It also makes two deliberately conservative client choices: a 15-second timeout and no more than five rate-limit attempts.

package main

import (
    "bytes"
    "fmt"
    "io"
    "log"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func main() {
    baseURL := strings.TrimRight(os.Getenv("OBSERVABILITY_BASE_URL"), "/")
    key := os.Getenv("INFRAI_API_KEY")
    runID := os.Getenv("PIPELINE_RUN_ID")
    payload := []byte(os.Getenv("HEALTH_EVENT_JSON"))
    if baseURL == "" || key == "" || runID == "" || len(payload) == 0 {
        log.Fatal("OBSERVABILITY_BASE_URL, INFRAI_API_KEY, PIPELINE_RUN_ID, and HEALTH_EVENT_JSON are required")
    }

    client := &http.Client{Timeout: 15 * time.Second}
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequest(
            http.MethodPost,
            baseURL+"/v1/logs/ingest",
            bytes.NewReader(payload),
        )
        if err != nil {
            log.Fatal(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", runID)

        resp, err := client.Do(req)
        if err != nil {
            log.Fatal(err)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            log.Fatal(readErr)
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            fmt.Println(string(body))
            return
        }
        if resp.StatusCode != http.StatusTooManyRequests {
            log.Fatalf("ingest failed: status=%d body=%s", resp.StatusCode, body)
        }

        delay := time.Second << attempt
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
            delay = time.Duration(seconds) * time.Second
        }
        time.Sleep(delay)
    }
    log.Fatal("ingest failed after rate-limit retries")
}
Enter fullscreen mode Exit fullscreen mode

Keep the event boring. Stable field names are more useful at 02:17 than a clever message, and the run identifier correlates scheduler, worker, dependency, and publish events without pretending that a log field is a distributed trace. Logs can carry trace_id and span_id, but this surface has no trace query or span tree.

Infrai is one candidate for this supporting telemetry layer because its self-describing discovery surface covers 295 routes across 20 modules under one key, with full request and response schemas available without authentication. The relevant architectural advantage is contract stability: application code can keep the same plain REST contract while the vendor behind a capability changes. That reduces integration churn, and the public schema makes validation practical.

The limitation is decisive: Infrai is not suitable as the primary uptime monitor. There are no probes, heartbeat monitoring, or notification routes in this API surface, so replacing a managed detector would mean building and operating a polling alert loop. The trade-off is a stable, broad telemetry contract without the independent detector or paging path; choose Healthchecks.io for missing cron completions and a managed uptime product for endpoint checks.

Should a Small Business Use Self-Hosted or SaaS Uptime Monitoring?

Option Operating model Best fit here Boundary to plan for
Healthchecks.io Managed cron heartbeat service Directly detects a missing nightly completion Pair with an external HTTP probe and richer telemetry
Better Stack Uptime Managed uptime product Hosted endpoint checks and alerting without monitor maintenance Verify current EU region, retention, and notification terms
Grafana Cloud Synthetic Monitoring Managed synthetic monitoring integrated with Grafana Synthetic results beside an existing Grafana estate More platform surface than one health endpoint may justify
Uptime Kuma Self-hosted uptime monitor Firm self-hosting requirement with staff to own lifecycle work Its availability, upgrades, storage, and notifications join the on-call load

This is a buy-versus-build decision, not a feature-count contest. Healthchecks.io supplies the direct missing-run signal. Better Stack Uptime and Grafana Cloud Synthetic Monitoring are credible managed choices when endpoint coverage or an existing observability estate matters more. Uptime Kuma supplies control, but control consumes operating capacity.

Managed services introduce vendor dependency and require a current review of region, retention, and notification terms. Self-hosting can be the correct boundary when policy forbids relevant external processing or a staffed platform team needs control over placement. Neither condition follows merely from having a small application.

My decision rule is stricter: require mandated placement, a necessary integration the managed candidates cannot supply, or enough platform capacity to own another reliability system before self-hosting. Otherwise, select the SaaS product whose documented data region and escalation path meet the SLO, and test it from the actual EU deployment context. Do not infer regional behavior from a marketing label.

How does the on-call reconstruct the incident?

Start with the detector and timestamp. If the external probe failed, compare it with application health latency and dependency status. If the heartbeat deadline failed while the endpoint stayed green, search the structured events by scheduled run identifier, locate the last pipeline transition, and inspect job_last_success_age_seconds. A captured exception can explain a crash. No exception and no completion heartbeat points instead toward silent scheduling or execution failure. This separation prevents a common category error: a log search can prove what an instrumented process reported, but it cannot prove that a process which emitted nothing was ever scheduled. A heartbeat monitor exists to detect that absence, which is why converting a query result into an alert is not equivalent to monitoring the schedule from outside the worker.

Do not guess.

There are other limits to name before an incident. The telemetry surface has no distributed trace query or span tree. For frontend diagnosis, there is no session replay, source-map decoding, or crash symbolication, so a separate browser error product may be necessary. Logs also have no per-user deletion route or bulk export/subscription interface, while retention and cold-storage configuration are not exposed.

For an edtech workload, keep student identifiers out of health events unless they are strictly required. Operational fields such as run ID, stage, dependency, status, and last-success timestamp are enough for this reconstruction path. Data minimization is easier than designing a deletion workflow around an interface that does not provide one.

A bad threshold charges interest

Set the heartbeat deadline from the expected processing window plus a defensible margin, then connect the page to the freshness SLO. A threshold that is too tight wakes the team for ordinary runtime variance; repeated harmless pages slow response and consume attention. A threshold that is too loose spends the freshness error budget before anyone acts.

The same caution applies to endpoint latency. Paging on every slow sample mistakes noise for user impact, while waiting for total failure throws away an earlier signal. Review alert counts, acknowledged causes, and missed user-impact windows after each operating cycle. Change the threshold only when that evidence supports the move.

False positives cost capacity.

They also train people to wait.

For this small EU-hosted edtech application, buy the basic detection path, keep the Go event contract portable, and spend engineering effort on reconstruction data that an external probe cannot infer. SaaS owns detection and delivery; structured telemetry owns the explanation. The boundary is plain enough to test during a game day and narrow enough that the junior team can operate it.

Further reading

Top comments (0)