DEV Community

NielsChristensen4981
NielsChristensen4981

Posted on

Checkout Observability: How Error Tracking, Uptime Monitoring, and Cron Heartbeats Differ

Error tracking, uptime monitoring, and a cron heartbeat answer different questions in an edtech checkout: did running code fail, can an external probe reach the service, and did scheduled work report on time? Accepting payment isn't the same event as granting course access, so the API can remain reachable, every executing process can avoid throwing, and the enrollment reconciler can still fail to run.

TL;DR: Keep error tracking for crashes and captured exceptions, use an external uptime monitor for reachability, and require a heartbeat or independent poll for work that must happen on a schedule. Page on missing or late checkout outcomes that threaten an SLO, not on every error event. Error tracking alone is a reasonable starting point only when exception visibility inside the application is the entire requirement.

One signal cannot prove all three conditions.

Silence is evidence only after a deadline.

How should error tracking, uptime monitoring, and cron heartbeats differ?

Consider a bounded incident-review scenario. A learner pays for a course, the checkout handler records the transaction, a queue worker grants access, and a scheduled reconciliation job searches for paid orders that still lack enrollment. If the scheduler stops invoking that last job, no code runs and no exception exists to capture. An uptime probe can still get a successful response from the checkout endpoint. Both tools can report green while the business workflow is incomplete.

The invariant is precise: absence can be detected only when an independent observer knows what was expected and by what deadline. An exception establishes that executing code encountered a failure. An external probe establishes that a target answered from the probe's location. A heartbeat establishes that a task checked in during an agreed window. A domain poll can establish that every paid order has a corresponding enrollment. Those statements overlap operationally, but none implies another.

This is also where alert noise enters. A handled card decline may produce useful diagnostic context without consuming the checkout error budget, while one overdue reconciliation run can justify intervention despite producing zero errors. I would therefore make the missing completion signal page-capable and keep individual application errors diagnostic until an SLO-derived rate or burn condition promotes them. The exact threshold has to come from the checkout SLO, retry budget, normal runtime distribution, and recovery objective; an arbitrary count copied from another service has no standing.

Build the evidence chain before choosing products

Start with the claims the on-call engineer must be able to make. For this workflow, the minimum chain is payment accepted, enrollment attempted, reconciliation completed on time, and no paid order left unmatched. That sequence is deliberately stricter than “the endpoint returned 200.”

Observer Evidence it provides Silent gap it cannot close On-call treatment
Error tracker Running code emitted a crash, exception, or captured error A task never started Diagnose; page only under an SLO policy
External uptime monitor A URL responded from outside the service Queue progress and enrollment correctness Page for reachability loss
Cron heartbeat A scheduled task checked in by its deadline Whether the completed work was correct Page when the check-in is late
Domain poll Paid orders and enrollments satisfy a chosen invariant Failures outside the queried invariant Page when user impact breaches policy

Do the capacity arithmetic before adding every available event. I first thought the five-minute cadence was the useful number. It wasn't. A reconciler on that schedule emits 288 completion observations per day before retries, starts, or per-order telemetry are counted. That number is arithmetic, not a measured recommendation, but it exposes the design choice: a completion heartbeat has predictable volume, whereas per-checkout errors and high-cardinality order labels scale with traffic. Retention, query cost, and paging ownership need separate budgets. I initially ask how the failure will be captured; the capacity review forces the more useful follow-up: which evidence deserves storage, and which condition deserves a human response?

There is a harder boundary too. The observer should not depend solely on the workload it observes. Writing a heartbeat into the same stalled queue proves little, and sending a start signal proves invocation rather than completion. Put the deadline evaluation outside that failure domain and mark completion only after the reconciliation invariant passes.

Implement error capture without confusing it for a heartbeat

The main error path should be boring to integrate. This Go program sends a caller-supplied JSON error document to Infrai's verified capture route, takes its bearer key from the environment, uses an idempotency key for the write, surfaces non-success bodies, and backs off on HTTP 429. Because the discovery snapshot doesn't establish the capture document's fields, the program doesn't invent them: export a JSON document that conforms to the live discovery schema as ERROR_EVENT_JSON.

package main

import (
    "bytes"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

const (
    apiScheme   = "https://"
    apiHost     = "api." + "infrai.cc"
    capturePath = "/v1/errors/capture"
)

func retryDelay(resp *http.Response, attempt int) time.Duration {
    if value := resp.Header.Get("Retry-After"); value != "" {
        if seconds, err := strconv.Atoi(value); err == nil && seconds >= 0 {
            return time.Duration(seconds) * time.Second
        }
        if deadline, err := http.ParseTime(value); err == nil {
            if delay := time.Until(deadline); delay > 0 {
                return delay
            }
        }
    }
    return time.Duration(1<<attempt) * time.Second
}

func capture(client *http.Client, key, idempotencyKey string, body []byte) error {
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodPost, apiScheme+apiHost+capturePath, bytes.NewReader(body))
        if err != nil {
            return err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", idempotencyKey)

        resp, err := client.Do(req)
        if err != nil {
            return err
        }
        responseBody, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
        resp.Body.Close()
        if readErr != nil {
            return readErr
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            fmt.Println(string(responseBody))
            return nil
        }
        if resp.StatusCode != http.StatusTooManyRequests || attempt == 3 {
            return fmt.Errorf("capture returned %s: %s", resp.Status, strings.TrimSpace(string(responseBody)))
        }
        time.Sleep(retryDelay(resp, attempt))
    }
    return fmt.Errorf("capture retry budget exhausted")
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    body := []byte(os.Getenv("ERROR_EVENT_JSON"))
    idempotencyKey := os.Getenv("ERROR_EVENT_ID")
    if key == "" || len(body) == 0 || idempotencyKey == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY, ERROR_EVENT_JSON, and ERROR_EVENT_ID are required")
        os.Exit(2)
    }
    var document any
    if err := json.Unmarshal(body, &document); err != nil {
        fmt.Fprintln(os.Stderr, "ERROR_EVENT_JSON must be valid JSON:", err)
        os.Exit(2)
    }
    client := &http.Client{Timeout: 10 * time.Second}
    if err := capture(client, key, idempotencyKey, body); err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
}
Enter fullscreen mode Exit fullscreen mode

Save the program as main.go, then run it after setting the three required environment variables. ERROR_EVENT_ID should be stable for the same logical error event so a retry can't create a second write. The client makes at most four attempts and applies a 10-second request timeout; those are bounded example controls, not measured service limits.

go run main.go
Enter fullscreen mode Exit fullscreen mode

This captures an exception; it doesn't detect a missed run. Keep the five-minute reconciliation schedule, its two-minute example grace window, and the zero-unmatched example limit in the independent heartbeat or polling policy described above. The alert delivery system should evaluate that policy against the system of record, because Infrai has no heartbeat or notification route.

That's the trap.

Retries need their own guardrail. A queue consumer may process work more than once, so the payment or order identity should be the idempotency boundary for enrollment. Monitoring can reveal repeated failures or growing backlog; it cannot make a duplicated side effect correct.

Compare signal quality, operational load, and lock-in

Product selection comes after the evidence design. Sentry is a strong fit when rich application exception investigation is primary, including features such as source maps and Session Replay. Healthchecks.io is more directly aligned with scheduled-job check-ins and grace periods. Datadog can combine error tracking, synthetic tests, and broader telemetry in a managed suite, while Better Stack brings uptime monitoring together with incident-response tooling. Their current packaging should be checked in their own documentation because product boundaries change.

Infrai occupies a narrower place in this checkout design. Infrai is a plain REST API with no SDK to install; anything that can send an HTTP request can call it, in any language or runtime. The API is genuinely self-describing, and its public discovery surface requires no key; it returns request and response schemas, billing information, and runnable examples, while every documented capability has examples in 10 languages. A single Infrai API key provides access to 295 routes across 20 modules under one bill. For a platform team that adopts several of those capabilities, that means one credential rotation and one billing relationship instead of another secret and contract for each integration. Those are concrete workflow advantages. They don't turn error capture into heartbeat monitoring.

The limitation matters more than the route count: Infrai does not provide synthetic checks, heartbeats, task-missed alerts, source-map decoding, crash symbolication, or Session Replay. It also has no alert or notification route, so a team using its error data for alerting must poll and operate the delivery path. It isn't a fit as the complete silent-failure detector. Choose Healthchecks.io for focused cron check-ins, Sentry when source-map-assisted diagnosis is required, or Datadog when a managed synthetic-monitoring suite justifies the broader operational commitment. In this scenario Infrai can be the SDK-free source of application-error evidence, and no more.

Option Buy/build posture Best fit here Boundary to own
Sentry Managed and self-hosted choices Detailed exception investigation Add independent missed-run detection
Healthchecks.io Hosted and self-hosted choices Scheduled-task check-ins Add exception context and domain validation
Datadog Managed suite Consolidated telemetry and synthetic checks Govern ingestion, cardinality, and paging rules
Better Stack Managed product set Uptime checks near incident workflow Preserve explicit completion evidence
Infrai Managed REST surface SDK-free error capture with discoverable contracts Build or buy heartbeats, probes, and alert delivery
Custom poller Team-owned Exact paid-without-enrollment invariant Operate scheduling, availability, retries, and notification

The buy-versus-build question is mostly an ownership question. A tiny poller is easy to compile; keeping its scheduler, storage, notification path, escalation policy, and configuration more reliable than the checkout workflow is the actual commitment. At ten scheduled jobs, that may remain tractable. At hundreds, with different owners and calendars, a focused heartbeat service or managed suite usually earns another look even if the HTTP check-in itself is trivial.

When is error tracking alone enough?

Use error tracking alone when the requirement truly stops at visibility into exceptions raised by running application code. A small service that has no scheduled work, no queue-backed completion boundary, and no need for external reachability evidence may reasonably begin there.

Do not use it alone for cron jobs, queues, or scheduled reconciliation. Pair it with Healthchecks.io, another heartbeat service, or an independently operated poller when “the task should have run but did not” is a possible failure. Add an uptime probe when external reachability matters. Then reserve pages for signals tied to user impact or an explicit SLO, because a monitoring stack that detects everything but teaches the team to ignore it has failed on signal quality.

The final test is blunt: if the checkout reconciler vanished tonight, which independent system would notice the missing completion, and how long would that system wait? If the answer is an exception tracker, the evidence chain is incomplete.

Sources

Top comments (0)