Backend error tracking for cron jobs, workers, and a web API should pair exception capture with deadline and outcome signals, or silent failures will remain invisible. In an e-commerce AI agent loop, the largest observability charge is determined by multiplication: turns create events, events carry bytes, and retention keeps those bytes queryable. Instrumentation breadth matters less than what is repeated on every model call, tool attempt, and retry. The useful design is a compact operational record for every turn, plus richer evidence only for failures and a small, declared sample of successes.
TL;DR: keep durable outcome, latency, cost-allocation, deployment, and correlation fields; retain bounded failure evidence long enough to cover the rollback window; and treat missed schedules, stuck workers, empty successes, and policy rejections as explicit outcomes. A rollback decision should work without replaying prompts, responses, customer messages, or checkout data. Exception capture and heartbeats remain separate signals because either one can be green while the customer journey is broken.
What is the bill actually made of?
Start with an accounting identity, not a feature matrix. For one retention tier, stored volume is approximately turns per day * events per turn * average event bytes * retained days. Query scans, indexing, replicas, and outbound transfer can add work, but reducing a repeated high-cardinality payload attacks the dominant term before those secondary costs appear. Consider a capacity-planning example, not a production benchmark. Assume an agent handles 20,000 shopping-assistant turns per day. Each turn makes two model calls and three tool attempts, producing six operational events after the final outcome is included. At 4 KB per event and 30 days of searchable retention, the raw input is about 14.4 GB. Putting a 40 KB prompt-and-response body on each event changes the same arithmetic to about 144 GB. The tenfold difference comes from payload policy, not from choosing a different exception UI. The change that moves the term is straightforward: store identifiers and measurements on the hot path, then attach a bounded diagnostic envelope when the outcome warrants it. A normal event may contain timestamps, duration, attempt count, normalized outcome, estimated usage units supplied by the application, release ID, workflow version, region, and opaque correlation IDs. It does not need a product description copied six times, a full conversation, or an unbounded stack of tool arguments.
Payload policy wins.
Be strict here. E-commerce payloads can cross awkward boundaries: a delivery-status tool may expose an address, while an OTP step may carry a phone number or a one-time code. Repeating either value in searchable telemetry creates a compliance and access-control problem without improving rollback math. Redaction after ingestion is late; the safer boundary is the event constructor.
from dataclasses import dataclass
from typing import Literal
Outcome = Literal[
"ok", "exception", "timeout", "missed_deadline",
"empty_result", "policy_rejected", "cancelled",
]
@dataclass(frozen=True)
class TurnEvent:
occurred_at: str
trace_id: str
turn_id: str
release_id: str
workflow_version: str
operation: str
outcome: Outcome
duration_ms: int
attempt: int
input_units: int | None
output_units: int | None
diagnostic_ref: str | None = None
def retention_class(event: TurnEvent) -> str:
if event.outcome != "ok":
return "failure_evidence"
if stable_sample(event.turn_id, per_thousand=10):
return "success_sample"
return "aggregate_only"
def stable_sample(turn_id: str, per_thousand: int) -> bool:
import hashlib
bucket = int.from_bytes(
hashlib.sha256(turn_id.encode("utf-8")).digest()[:4], "big"
) % 1000
return bucket < per_thousand
The one-percent success sample above is an explicit planning choice, not a universal optimum. Stable sampling matters because a release comparison should not change merely because a process restarted. Teams should choose the rate from their traffic, rare-path coverage, investigation window, and privacy constraints, then record that policy beside the data.
What should backend error tracking cover for cron jobs and workers?
More than exceptions. A heartbeat answers “did something report near the expected time?” An exception answers “did executed code raise and report a failure?” Neither proves that useful work completed.
A cron process can start, emit its heartbeat, catch a downstream timeout, and exit with status zero. A worker can stay alive while messages accumulate beyond their useful deadline. A web API can return a syntactically valid success with an empty recommendation set. An agent can finish normally after every inventory tool call was rejected by policy. No exception is required for any of these cases.
This is why “healthchecks alternative” is the wrong selection frame. The signals form a small evidence chain:
| Signal | Question it answers | Failure it can miss |
|---|---|---|
| Schedule assertion | Did the job begin in its allowed window? | Work that began but never completed |
| Progress or queue age | Is useful work moving before its deadline? | Semantically wrong output |
| Exception event | Where did executed code fail? | Caught errors and empty success |
| Outcome event | Did the business operation reach a declared result? | A missing job unless absence is checked |
| Latency and traffic | Did behavior change by release or workflow? | A rare failure hidden by aggregation |
Google's SRE discussion of monitoring distributed systems separates latency, traffic, errors, and saturation as four golden signals. That separation is useful here: an exception stream covers only part of “errors,” and error evidence alone says little about rising queue age or slow model calls. For an agent loop, record latency around each model and tool boundary, traffic as completed turns and attempts, errors as explicit outcomes, and saturation at the constrained worker or queue.
Outcome events also need a deadline. Without one, a missing completion is indistinguishable from a slow completion until somebody searches manually. A scheduler can evaluate (workflow, expected_window); a worker monitor can evaluate the age of the oldest eligible item; an API monitor can compare request acceptance with terminal outcomes. Each assertion should create the same searchable outcome shape as an exception, including release and correlation IDs.
Design search around rollback questions
During a risky deployment, broad full-text search is less important than fast answers to constrained questions. Did the new release increase terminal failures? Which operation changed? Are failures isolated to one workflow version or region? Did p95 latency move while traffic stayed comparable? Can the on-call engineer find representative evidence without opening customer content?
Those questions define the event schema. Use low-cardinality fields for aggregation and exact filters: release_id, workflow_version, operation, outcome, and region. Keep trace_id and turn_id for point lookup, not group-by charts. Separate attempt outcome from final turn outcome, or retries will inflate both traffic and failure counts and make a recovered timeout look like a lost purchase.
Cost needs the same discipline. If the application receives usage measurements from a model boundary, store those measurements with the operation and release. If cost allocation depends on a mutable rate card, keep the usage units and rate-card version rather than baking an unexplained currency number into every event. The rollback question is comparative: did this release change units per successful turn, retry count, or latency enough to violate its guardrail?
Never make rollback depend on a field that is sampled away. Terminal outcome counters, schedule assertions, and release labels belong in the durable compact record. Large diagnostic bodies can have shorter retention and stricter access because they support explanation, not detection.
A practical release view needs four comparable slices: the candidate release, its baseline, the same workflow version, and a time window that contains enough completed turns to evaluate the guardrail. Do not compare a fresh deployment's partial turns with yesterday's completed cohort. Late completions should be assigned by their start release and shown separately when they cross the decision window.
Capture exceptions without making them the data model
Exception capture still earns its place. Preserve exception type, a normalized message or fingerprint, stack frames, operation, attempt, release, and trace correlation. Bound message length, strip secrets, and avoid using raw messages as grouping keys; identifiers, URLs, SKU values, and timestamps can fragment one defect into thousands of groups.
The transport should fail quietly from the application's point of view while exposing its own drop count. Logging libraries commonly support appenders, and the Logback documentation describes the appender contract and custom appenders. The architectural lesson crosses runtimes: adapt structured records at the logging boundary, buffer within a fixed limit, and define what happens under backpressure. Blocking checkout because the telemetry sink is slow is a bad dependency direction.
There is a sharp edge. If the process crashes, an in-memory buffer may vanish; if every emit forces synchronous delivery, tail latency and failure coupling rise. Choose the boundary deliberately. Critical terminal outcomes can go through a durable local or queue-backed path, while verbose diagnostics can accept bounded loss. Monitor dropped-event counts and test shutdown flushing, process termination, queue unavailability, duplicate delivery, and malformed payloads.
I would also test the non-exception cases first: a cron invocation that never starts, a handler that catches a timeout, a worker that renews its lease without advancing, and an API response that is valid but empty. This is a design preference, not a claim about a past incident. Those cases reveal whether the system models work or merely models crashes.
Retention is part of rollback safety
Retention should be derived from decisions. Keep compact aggregates and terminal outcomes across the period used for release baselines and capacity trends. Keep searchable failure evidence through the longest plausible detection, triage, and rollback interval. Keep raw prompt, response, and tool bodies only when a separately justified investigation requires them, with narrower access and an explicit deletion schedule.
This approach has limits. It is a poor fit when the actual requirement is a legally authorized, byte-for-byte business archive or deterministic replay of every agent turn; in that case, use a separately governed system of record and keep observability as an index into it. The trade-off is extra correlation and access-control work, but an operational event store should not quietly become the permanent home for customer conversations.
This means deliberately giving something up. With aggregates plus a stable success sample, an engineer may be unable to reconstruct the exact wording of an otherwise successful recommendation after the detailed body expires. That is a real forensic cost. It is acceptable only if the retained record can still establish which release ran, what tools were attempted, how long they took, what resource units were reported, and which terminal outcome occurred.
Some detail is gone.
Run a rollback drill before relying on the design. Deploy two identifiable releases in a test environment, inject a timeout and an empty-success path, omit one scheduled run, and verify that the release view separates all three. Then expire the diagnostic tier and repeat the decision using only durable fields. If the team cannot decide which release to stop, the schema is missing evidence; if it can decide but cannot inspect one representative failure, the retention window is too short.
The selection criterion is therefore concrete: choose an implementation that can ingest structured outcomes from cron jobs, workers, and APIs; correlate attempts with terminal results; assert absence and deadlines; search by release without exposing sensitive payloads; report dropped telemetry; and apply different retention and access policies to metrics, compact events, and diagnostic bodies. The best fit is the one that makes a rollback decision with the smallest sufficient evidence set.
Top comments (0)