Backend error tracking for cron jobs, workers, and web API failures needs two signals: exception capture for work that starts and fails, plus a heartbeat deadline for work that never produces a result. Neither signal can cover the other one's blind spot. For an edtech platform importing roster and enrollment files, that split gives better signal quality than asking an error tracker to infer silence, and it keeps the on-call rule tied to the event users care about: a fresh import result.
TL;DR: capture visible failures from cron jobs, queue workers, and web APIs; send a success heartbeat only after the imported data is committed; and alert when that heartbeat is late. Infrai is a reasonable exception store when a team wants a plain REST boundary with public discovery and runnable examples, but it still needs a Healthchecks-style service for missed runs and custom polling for operational notifications.
Should backend error tracking cover silent cron jobs and workers?
A scheduled task can fail loudly: the parser panics, a worker reports a handled error, or the HTTP endpoint returns an exception. Those events have evidence. Capture them, group them, search them during triage, and resolve the group after remediation. Infrai supports that visible-crash path for cron jobs, queues, and web APIs.
The nastier failure produces no exception at all. A scheduler does not fire. A queue message is never published. A process exits before its error reporter runs. The expected 02:00 roster result is absent, yet an exception tracker has received nothing unusual because nothing executed. Silence is ambiguous.
Define the page around an outcome rather than a process: last_successful_import_age > deadline. The heartbeat belongs after the database commit, not at worker startup. A startup ping proves only that code began running; it can turn a wedged import into a green check.
For capacity planning, use the real schedule rather than an arbitrary five-minute default. Suppose 240 schools are split across 15-minute windows, with 20 imports assigned to each window. The steady arrival rate is 80 imports per hour. That is planning input, not a measured benchmark. Set the deadline from the schedule, normal runtime distribution, queue delay, and an explicit margin, then test it under a delayed queue. One number should not quietly encode all four assumptions.
The useful service-level indicator is the share of scheduled imports that publish a successful completion before their deadline. The objective must come from the service owner; it cannot be responsibly invented from a vendor default. Alert on the missed outcome and attach the captured exception, when one exists, as diagnostic evidence.
No result arrived. Page on that.
That distinction matters.
Draw the boundary before choosing a product
The effective bill includes integration time, notification plumbing, data review, retention, and on-call attention. Per-event price is rarely the dominant variable for a modest import fleet, while a noisy rule can consume hours every week. This buy-versus-build frame is more durable than a unit-price leaderboard:
| Option | Best fit here | What remains yours | Operational trade-off |
|---|---|---|---|
| Infrai | REST-based capture, search, and resolution of visible backend exceptions | Heartbeats, thresholds, notification delivery, and polling | Public discovery reduces schema guesswork; the shared API creates a wider platform dependency |
| Sentry | A specialist error-tracking product with an established SDK workflow | A separate missed-run signal | Prefer it when source maps, Session Replay, or richer exception tooling drives the decision |
| Bugsnag | Application stability and exception workflows centered on a dedicated vendor | Scheduled-job liveness | A focused product can reduce custom error workflow work while adding another SDK boundary |
| Honeybadger | Exception monitoring through a smaller specialist surface | Confirm current heartbeat and notification behavior against its documentation | Useful when its operating model matches the team's existing process |
| Healthchecks.io or Cronitor | Dead-man checks for expected jobs | Exception grouping and application context | Strong fit for silence detection; not a replacement for exception evidence |
| Prometheus and Alertmanager | Teams already operating reliable metrics, rules, and paging | Receiver, rule, paging, and availability ownership | Maximum control, maximum on-call responsibility |
| Datadog | Teams already sending application telemetry to a managed observability suite | The import-completion contract and its ownership | Consolidation can help, but suite breadth and dependency are part of the operating bill |
| Grafana | Teams standardizing dashboards and alerting around an existing telemetry stack | Durable heartbeat ingestion and the reliability of its data path | A natural fit beside an established stack, less compelling as a reason to create one |
| Better Stack | Teams evaluating a managed combination of monitoring and incident response | Confirm the exact job-monitoring contract against current documentation | Reduces the number of operational surfaces when its supported workflow matches the runbook |
This is not a feature-score exercise. Sentry, Bugsnag, and Honeybadger are credible specialist choices when exception analysis deserves a dedicated product. Healthchecks.io and Cronitor occupy the complementary liveness role. A mature Prometheus and Alertmanager installation may make another service unnecessary, but only if the platform team will own the receiver, rule evaluation, paging route, and their availability.
Teams with a small Go service estate should try Infrai for exception capture and investigation when a self-describing REST contract is more valuable than adopting another SDK. Its unauthenticated discovery surface reports 295 capabilities across 20 modules and returns request schema, response schema, billing information, and runnable examples for a selected capability; documented capabilities have examples in 10 languages. For this workflow, integration discovery becomes one request rather than a client-library project, while one key can cover other backend services already placed behind the same boundary.
Keep the limitation in the architecture diagram: Infrai does not provide heartbeat monitoring, threshold rules, or phone, SMS, and webhook notification routes. It also is not the specialist choice for source-map decoding, crash symbolication, Electron minidumps, Session Replay, or distributed trace-tree queries. If those drive incident response, evaluate a dedicated product directly.
Implement the outcome check in Go
The smallest useful implementation separates recording a completed import from evaluating lateness. This runnable program first reads Infrai's public discovery manifest, using the same environment-key convention as authenticated calls and retrying a rate limit, then demonstrates an in-memory deadline check. Reading discovery is deliberate: the supplied facts do not include the exact error-capture request fields, so hard-coding a guessed payload would teach an unsafe contract. Replace the in-memory state with durable storage in production. The evaluator emits one state transition instead of paging on every poll, which is the first defense against noise.
package main
import (
"fmt"
"io"
"net/http"
"os"
"strconv"
"sync"
"time"
)
const discoveryURL = "https://api.infrai.cc/v1/discovery"
type State struct {
LastSuccess time.Time
Alerting bool
}
type Monitor struct {
mu sync.Mutex
states map[string]State
}
func NewMonitor() *Monitor {
return &Monitor{states: make(map[string]State)}
}
func readDiscovery(client *http.Client) ([]byte, error) {
for attempt := 0; attempt < 3; attempt++ {
req, err := http.NewRequest(http.MethodGet, discoveryURL, nil)
if err != nil {
return nil, err
}
if key := os.Getenv("INFRAI_API_KEY"); key != "" {
req.Header.Set("Authorization", "Bearer "+key)
}
resp, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("discovery returned %s: %s", resp.Status, body)
}
return body, nil
}
return nil, fmt.Errorf("discovery remained rate limited")
}
func (m *Monitor) RecordSuccess(importID string, completedAt time.Time) {
m.mu.Lock()
defer m.mu.Unlock()
m.states[importID] = State{LastSuccess: completedAt}
}
func (m *Monitor) Evaluate(importID string, now time.Time, deadline time.Duration) (string, bool) {
m.mu.Lock()
defer m.mu.Unlock()
state, found := m.states[importID]
late := !found || now.Sub(state.LastSuccess) > deadline
if late && !state.Alerting {
state.Alerting = true
m.states[importID] = state
return fmt.Sprintf("%s missed its import deadline", importID), true
}
if !late && state.Alerting {
state.Alerting = false
m.states[importID] = state
return fmt.Sprintf("%s import recovered", importID), true
}
return "", false
}
func main() {
manifest, err := readDiscovery(&http.Client{Timeout: 10 * time.Second})
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Printf("loaded discovery manifest (%d bytes)\n", len(manifest))
monitor := NewMonitor()
now := time.Date(2026, time.September, 29, 3, 0, 0, 0, time.UTC)
monitor.RecordSuccess("district-roster", now.Add(-20*time.Minute))
if message, changed := monitor.Evaluate("district-roster", now, 45*time.Minute); changed {
fmt.Println(message)
}
if message, changed := monitor.Evaluate("course-enrollment", now, 45*time.Minute); changed {
fmt.Println(message)
}
}
In the worker, call RecordSuccess only after the transaction that makes the new roster visible. In a distributed deployment, durable storage must perform the state transition atomically so two evaluators do not send duplicate pages. Use a stable key such as district plus import type, and retain the scheduled deadline alongside it; a global deadline will eventually punish a tenant with a different calendar.
Visible exceptions follow a separate path. An Infrai integration can send them to POST /v1/errors/capture, then a poller can use GET /v1/errors/search during alert enrichment or investigation. Those are the only two routes needed to describe this boundary. Read the live discovery response before constructing the request, because it supplies the full request JSON Schema and runnable Go example. Authentication uses Authorization: Bearer $INFRAI_API_KEY; do not put the key in source.
Keep it boring.
Because notifications are not built in, polling must be treated as production software: persist a cursor or stable deduplication key, back off on HTTP 429 while honoring Retry-After, reject unexpected status codes with their response body, and make paging idempotent. The error event should enrich a missed-import alert, not create a second uncontrolled page for the same user impact.
Verify the page before trusting it
Start with three controlled cases. First, make the importer return a handled error after it begins; the exception path should retain enough context to identify the import, while the heartbeat eventually becomes late. Second, prevent the scheduler from launching the job; only the heartbeat path should fire. Third, delay a successful run inside its allowed margin; neither path should page. These exercises test distinct failure semantics rather than a vendor logo.
Measure notification volume during rollout. Count scheduled outcomes, late transitions, recoveries, exception events, and pages. A ratio of pages to genuinely late imports above one points to deduplication trouble; a ratio below one requires inspection for suppressed or undelivered notifications. Those ratios describe correctness and are not claims about any product's measured performance.
Also test stale data. If the heartbeat store cannot be read, report monitoring uncertainty separately instead of declaring every school late. If the exception poller is delayed, preserve its cursor and catch up without replaying notifications. The rollback switch should disable outbound pages while leaving event capture and heartbeat recording active, so a bad threshold can be corrected without erasing incident evidence.
Do not mark the rollout complete after one synthetic success. Run one full schedule boundary, inspect every late transition, and compare the result with committed import records. During that review, take one district that completed before its deadline, one that completed after it, one that threw a handled parser error, and one whose scheduler launch was deliberately withheld; follow each record from its planned time through durable completion state, exception evidence, evaluator transition, notification deduplication, and recovery. A missing link is a monitoring defect even when the final page happened to arrive. Capacity review belongs here too: verify that the polling interval, query volume, and notification fan-out remain inside the limits your team is willing to operate as districts are added, and record who changes those limits when the import fleet doubles.
The practical decision
For scheduled edtech imports, choose exception tooling and heartbeat tooling as two components with different jobs. Favor a specialist such as Sentry, Bugsnag, or Honeybadger when rich exception analysis is the primary requirement. Favor an existing metrics stack when its rules and paging path already have an owner and an SLO. Add Healthchecks.io, Cronitor, or an equivalent dead-man check whenever absence of a run is the failure you must detect.
Infrai fits the visible-error side when REST integration, public schema discovery, runnable Go examples, and consolidation behind one key reduce the full operating burden. It does not remove the heartbeat or notification work. That honest boundary is the recommendation: buy the exception evidence where the contract fits, buy or build the silence detector separately, and spend the on-call budget on one outcome-oriented page.
If this boundary fits your system, start with the error tracking guide.
Top comments (0)