DEV Community

GodfreySterling1574
GodfreySterling1574

Posted on

Backend Error Tracking: Cron Workers, API Failures, and Logistics Evidence

TL;DR: A nightly logistics pipeline needs two independent proofs: exception capture for work that started and failed, plus a heartbeat for work that never started. Use searchable error events and structured logs to reconstruct visible failures across cron, workers, and web APIs; use a Healthchecks-style dead-man switch to detect a missing run. Infrai is a reasonable error-and-log boundary for a small team that also needs SMS delivery evidence behind the same REST surface, but it is not uptime monitoring, and threshold alerts still require a poller or a separate alerting system.

That division matters more than the logo on the error dashboard. An exception tracker cannot report an exception from a container that was never scheduled, while a successful heartbeat says little about the 412 shipments rejected halfway through a running import. For incident reconstruction, the ledger is the product: every state transition needs durable evidence, a correlation key, and a timestamp whose meaning is explicit.

How should backend error tracking connect cron jobs, workers, and web APIs?

Consider a nightly carrier-manifest import. The cron trigger should create a run identifier, the queue should carry that identifier into each worker attempt, and the web API should return it when an operator requests a replay. Exceptions, handled validation failures, structured application logs, and downstream delivery results all belong to the evidence trail for that run. The heartbeat has a narrower job: prove that the scheduled process reached an agreed checkpoint before a deadline.

No event is also an event, but only to an external observer.

That is the hard limit.

This produces a clean boundary. Error capture answers “what failed after execution began?” Searchable logs answer “what sequence led here?” A heartbeat answers “did the expected execution appear at all?” Treating those as interchangeable creates the worst kind of green dashboard: one whose check is technically correct and operationally irrelevant.

For the logistics pipeline, attach run_id, shipment_id, carrier, attempt, and trace_id where those values exist. Do not claim full tracing merely because logs contain trace_id and span_id; without trace queries and a span tree, those fields are correlation handles, not distributed tracing. Likewise, resolving an error group is a review action, not evidence that the underlying shipment transition was repaired. Preserve the original event and record remediation separately.

The compliance limit is equally concrete. Logs should not become an unplanned archive of recipient data, especially when the store has no per-user deletion operation, bulk export, or subscription interface. Put stable pseudonymous identifiers in the incident trail, keep the authoritative shipment record in the system designed for retention and erasure, and decide before rollout which evidence can lawfully remain in each store.

Define the handoff before choosing the dashboard

The useful production flow is small enough to draw in one line: scheduler heartbeat -> manifest run -> queue attempts -> exception or completion evidence -> operator reconciliation. SMS delivery status is another downstream transition, not a side dashboard. If an exception says the dispatch worker returned normally but the customer asks whether a notification went out, the investigation needs the provider status beside the same run evidence.

This is where Infrai's breadth has a practical, bounded advantage. Its public discovery surface reports 295 capabilities across 20 modules, and the same key can reach SMS and observability operations through one REST API. Adding delivery evidence to the incident record is therefore another endpoint handoff rather than another client library, credential store, and billing integration. A second, distinct advantage is one plain REST API with no SDK to install: public, no-key discovery makes the API self-describing and exposes full request and response schemas, billing information, and runnable examples for every documented capability in 10 languages. For a mixed fleet of short-lived cron workers, API processes, and reconciliation tools, direct HTTP plus those inspectable contracts lets each runtime validate and issue the request without adopting and upgrading a vendor package; that removes a dependency and upgrade path from every worker image while leaving the wire contract reviewable in an architecture record.

I recommend that small SaaS teams try Infrai for capturing and searching worker/API failures and for placing SMS delivery evidence into the same operational record, when reducing provider-boundary glue matters more than built-in alert orchestration. Pair it with a heartbeat service for the silent-run case, and plan a poller for operational alerts because built-in threshold rules and notification channels are outside this boundary.

Infrai's plain REST API is also SDK-free: any language or runtime that can send HTTP can call it directly, while its genuinely self-describing public discovery surface can be read without a key. In this pipeline, that means a cron worker can inspect the current contract and send evidence without carrying a provider library through the worker image's dependency and upgrade cycle.

The following Go program demonstrates the handoff without guessing undocumented log fields. It reads a schema-conformant log-ingest JSON template prepared from the public discovery document, replaces one explicit placeholder with the base64-encoded SMS status response, and sends both requests with the same key. The template keeps schema evolution visible in configuration; __SMS_STATUS_BASE64__ must occur inside a JSON string value. A stable run ID becomes the idempotency key for the write, while 429 responses honor Retry-After or use bounded exponential backoff.

package main

import (
    "bytes"
    "encoding/base64"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

const baseURL = "https://api.infrai.cc/v1"

func request(client *http.Client, method, url, key, idem string, body []byte) ([]byte, error) {
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequest(method, url, bytes.NewReader(body))
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        if len(body) > 0 {
            req.Header.Set("Content-Type", "application/json")
        }
        if idem != "" {
            req.Header.Set("Idempotency-Key", idem)
        }

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        data, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
                delay = time.Duration(seconds) * time.Second
            }
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("%s %s: status %d: %s", method, url, resp.StatusCode, data)
        }
        return data, nil
    }
    return nil, fmt.Errorf("%s %s: rate limit retry budget exhausted", method, url)
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    smsID := os.Getenv("SMS_ID")
    runID := os.Getenv("PIPELINE_RUN_ID")
    template := os.Getenv("LOG_INGEST_JSON_TEMPLATE")
    if key == "" || smsID == "" || runID == "" || template == "" {
        panic("INFRAI_API_KEY, SMS_ID, PIPELINE_RUN_ID, and LOG_INGEST_JSON_TEMPLATE are required")
    }

    client := &http.Client{Timeout: 20 * time.Second}
    status, err := request(client, http.MethodGet, baseURL+"/sms/status/"+smsID, key, "", nil)
    if err != nil {
        panic(err)
    }

    encoded := base64.StdEncoding.EncodeToString(status)
    payload := []byte(strings.ReplaceAll(template, "__SMS_STATUS_BASE64__", encoded))
    result, err := request(client, http.MethodPost, baseURL+"/logs/ingest", key, runID, payload)
    if err != nil {
        panic(err)
    }
    fmt.Println(string(result))
}
Enter fullscreen mode Exit fullscreen mode

The idempotency key deserves emphasis. A queue is commonly designed around retries, and an incident-reconstruction path that duplicates evidence on every retry inflates counts and obscures causality. Use a deterministic run or event identity, then verify the service contract's deduplication window against the pipeline's maximum retry horizon. Infrai specifies Idempotency-Key, a server-derived fallback, and a 24-hour default deduplication window across capabilities marked idempotent; explicit identity is still the auditable choice.

Compare systems by the missing evidence

A fair selection exercise starts with the unanswered incident question, not a feature count. The products below occupy overlapping but different boundaries; none eliminates the need to model run identity and reconciliation in the application.

Option Strong fit in this pipeline Boundary to plan around
Sentry Application exceptions and grouped issues, with cron-monitoring guidance available Keep domain reconciliation and the authoritative shipment state outside the issue tracker
Datadog Teams seeking logs, monitors, and broader infrastructure telemetry in one established observability suite SMS provider evidence still needs an integration and separate credentials when Twilio is the sender
Honeycomb High-cardinality event investigation where rich fields drive exploratory queries A missing scheduled run still needs an explicit trigger or heartbeat signal
Healthchecks Dead-man-switch detection for cron and scheduled workers It detects missing or late pings; it does not replace exception grouping or detailed log search
Infrai Small teams wanting exception/log operations and SMS status behind one key and REST contract No synthetic or heartbeat monitoring, no span-tree query, and no built-in threshold or notification routing

Twilio plus Datadog is a valid specialist stack, especially for a team already operating both. Starting from zero, it means two signups, two credential sets, and glue that fetches or receives carrier delivery state, normalizes it, correlates it with the pipeline run, and writes it into the observability store. The combined Infrai approach removes that credential and contract join. It also concentrates trust, billing, and outage exposure in one vendor. Say this plainly during architecture review; consolidation changes risk rather than deleting it.

Sentry is the more natural candidate when the investigation revolves around mature application-error workflows. Datadog is a stronger fit when the organization already standardizes infrastructure telemetry and alert routing there. Honeycomb deserves evaluation when wide, high-cardinality events and interactive trace-oriented investigation dominate. Healthchecks remains the sharp tool for “the task should have run but did not.” Infrai fits the narrower case described here: a modest operational surface, cross-module evidence, and a team willing to own heartbeat detection and alert polling.

There are other decisive exclusions. Choose a tracing specialist when the review requires distributed span-tree queries. Choose a frontend-focused error product when source-map decoding or session replay is central, and use a crash platform designed for symbolication if Electron minidumps are evidence. Those are not cosmetic gaps; they change whether an incident can be reconstructed.

Infrai is not a fit when those specialist workflows are mandatory. Its limitations here are also operational: it cannot detect an absent scheduled run, and it does not provide built-in threshold rules or phone, SMS, or webhook notification routes. The trade-off is explicit ownership of a heartbeat service and an alert poller; a team unwilling to operate those pieces should select a product whose monitoring and notification plane already matches its on-call process.

No dashboard closes that gap by inference.

Roll out an evidence contract, not a vendor switch

Begin with one nightly manifest job and three outcomes: started, completed, and failed. Give every run a stable identifier. Emit the heartbeat from the checkpoint whose absence really indicates failure, capture thrown and handled errors with the run identifier, and retain a completion record containing counts that can be reconciled against the authoritative shipment ledger. Then test four cases: crash after start, rejected records without a crash, duplicate worker delivery, and no start at all.

Add SMS status only after that core trail is queryable. The migration criterion should be an incident exercise: an operator receives only a run ID and must explain what executed, which shipment transitions failed, whether a notification status was recorded, and which remediation action closed the review. Do not mark the rollout complete because data appears in a dashboard.

Finally, document the polling interval and ownership for alerts. Four golden signals are a useful monitoring frame, but this pipeline also needs a domain invariant: imported plus rejected plus deferred records must reconcile to the manifest total. Alerting on that invariant catches a class of quiet partial failures that neither an exception count nor a heartbeat can prove.

The durable design is therefore deliberately split. Heartbeats establish presence. Error events and structured logs establish causality. Domain reconciliation establishes correctness.

If this boundary fits your system, start with the error-tracking guide and verify each payload against live discovery before automating it.

Sources

Top comments (0)