The page says a marketplace experiment is breaching its error-budget policy, but the on-call view contains a pile of failed work with no reliable way to separate treatment tenants from control tenants. For Postgres-backed cron worker error tracking, including a Node.js background job runner, the least complex useful fix is to capture one immutable record per attempt, keyed by a stable job ID and carrying the experiment cohort, then alert when terminal outcomes diverge between cohorts.
TL;DR: Capture the first attempt, every retry decision, and the terminal outcome as separate events. Put cohort identity and a sanitized error class beside them, not inside an unstructured message. Alert on a sustained cohort delta with enough completed work in both groups; page on terminal impact, while retry spikes remain a warning. This preserves the evidence needed to reconstruct an incident after a worker has retried and overwritten its immediate error state.
How should a Postgres cron worker capture background job errors?
A useful page should answer three questions before anyone opens a trace: which cohort is affected, how many jobs reached a terminal failure, and what deploy or experiment version bounds the change. A raw exception count cannot do that. One customer action may create several attempts, so counting exceptions can make a retry policy look like user impact.
Start there.
For a marketplace, take a catalog-enrichment experiment split across tenant cohorts. A task reads an item, calls a downstream classifier, and writes the result. The alert should compare terminal outcomes per completed job in treatment and control, while retaining attempt counts for diagnosis. If treatment has 40 terminal failures among 2,000 completed jobs and control has 5 among 2,100, those are example inputs to the decision rule, not a claim about a production incident.
The page payload needs the observation window, cohort counts, experiment version, and a link or query key for the underlying attempts. It should not contain item descriptions, access tokens, connection strings, or raw request bodies. OWASP's logging guidance explicitly warns against recording secrets and sensitive personal data; useful reconstruction depends on disciplined fields, not indiscriminate capture.
Work backward from impact
The late signal is terminal failure rate by cohort. The earlier signal is retry amplification: attempts per completed job climbing in treatment while terminal outcomes still look normal. That signal buys investigation time, but it is a poor paging condition on its own because successful retries have repaired the immediate operation.
This distinction matters. Page on exhausted work that threatens the service objective; ticket or warn on unusual retry cost. Otherwise a transient dependency slowdown wakes someone even though the marketplace completes the work inside its stated deadline.
The reconstruction chain should be explicit:
- The alert groups terminal outcomes by
experiment_versionandcohort. - A stable
job_idfinds every attempt without using a tenant name or payload as a search key. - Each attempt records its ordinal, start time, finish time, outcome, sanitized error class, and the retry decision.
- The terminal row records success, permanent failure, or cancellation once. Later retries do not mutate earlier evidence.
That last property is the trap. Suppose job j-1042 fails validation on attempt one, is retried after a dependency response on attempt two, and then exhausts its policy on attempt three. If the worker updates one jobs.last_error column, responders see only the third error; they cannot tell whether treatment changed the input, the dependency became slow, or the retry policy merely repeated damage that began earlier. An append-only attempt table keeps all three observations beside one job identity, while the experiment version and cohort let the responder compare that sequence with control work in the same window. A uniqueness constraint on (job_id, attempt) also turns duplicate instrumentation into an observable write conflict instead of silently double-counting an attempt. The trade-off is blunt: more writes and retention pressure buy a timeline that can survive the worker process.
Retries distort the picture.
Instrument the attempt boundary
Instrumentation belongs around the unit that the retry mechanism invokes, rather than around the scheduler process. Cron starts processes, and queue libraries schedule retries, but the attempt boundary is where outcome, duration, and retry intent can be recorded consistently.
The following Go example uses a generic recorder so the same contract can sit behind a cron handler or a queue consumer. The fields are deliberately narrow. ErrorClass should come from a small reviewed taxonomy; it must not be the raw error text.
package jobs
import (
"context"
"errors"
"time"
)
type Attempt struct {
JobID string
TenantKey string
Cohort string
ExperimentVersion string
Number int
StartedAt time.Time
FinishedAt time.Time
Outcome string
ErrorClass string
WillRetry bool
}
type Recorder interface {
AppendAttempt(context.Context, Attempt) error
}
type Work func(context.Context) error
func RunAttempt(ctx context.Context, r Recorder, a Attempt, work Work, retryable func(error) bool) error {
a.StartedAt = time.Now().UTC()
err := work(ctx)
a.FinishedAt = time.Now().UTC()
if err == nil {
a.Outcome = "succeeded"
} else {
a.Outcome = "failed"
a.ErrorClass = classify(err)
a.WillRetry = retryable(err)
}
if recordErr := r.AppendAttempt(ctx, a); recordErr != nil {
return errors.Join(err, recordErr)
}
return err
}
func classify(err error) string {
if errors.Is(err, context.DeadlineExceeded) {
return "dependency_timeout"
}
return "unclassified"
}
The uneasy part is the recorder failure. Returning the instrumentation error makes evidence loss visible, but it may cause the job framework to retry work that actually succeeded. Ignoring it protects business processing while accepting a hole in the incident record. There is no universal answer: define which write owns correctness, make the choice explicit, and test it. For high-consequence tasks, persisting the business mutation and its attempt event in one PostgreSQL transaction can remove that ambiguity when both records live in the same database.
Do not confuse a job ID with a trace ID. A trace can describe one execution path; a job ID joins attempts across time. Keep both when tracing exists, but build the incident query on the durable identity.
Set thresholds from an SLO, not anxiety
Capacity planning starts before threshold selection. Estimate completed jobs per tenant cohort in the alert window, expected attempts per job, event retention, and the write rate during a dependency failure. A design that handles the average but drops attempt records during the exact burst being investigated has failed its primary job.
For paging, require a minimum denominator in both cohorts and compare terminal-failure proportions over more than one evaluation interval. The minimum denominator avoids treating one failure among two jobs as equivalent evidence to hundreds of outcomes. Repeated evaluation prevents a single late batch from defining the incident. Those parameters must come from the marketplace's traffic shape and error-budget policy; copying someone else's numeric threshold creates false precision.
Use a separate warning for attempts per completed job. It can reveal a cohort-specific regression earlier, but it should lead to inspection rather than immediate escalation unless retry load itself threatens database or worker capacity.
Wrong thresholds have a measurable operational cost: pages interrupt on-call work, and repeated non-actionable pages teach responders to distrust the channel. The opposite error is expensive too. A threshold that waits for the whole cohort to fail preserves quiet dashboards while consuming the error budget. Review both false positives and missed detections after each experiment, then record why the threshold changed.
Buy, build, or combine?
The durable event model matters more than the screen used to inspect it. A managed error tracker, a self-hosted observability stack, and direct PostgreSQL queries can all present the same evidence if the application emits stable dimensions and preserves attempt history.
| Approach | Incident reconstruction | On-call load | Lock-in boundary | Capacity concern |
|---|---|---|---|---|
| PostgreSQL attempt ledger | Direct joins to experiment and tenant keys | Team owns schema, retention, queries, and alerts | SQL schema and event contract | Primary database write and storage growth |
| Self-hosted telemetry pipeline | Flexible correlation and storage choices | Team owns upgrades, scaling, and failure recovery | Collector schema plus chosen storage | Ingest bursts, queues, and retention |
| Managed error tracking | Fast exception grouping and alert delivery | Provider operates the service; team owns instrumentation policy | Provider event model, queries, and export path | Ingest quotas and retained event volume |
| Combined ledger and telemetry | Durable job history plus faster cross-signal investigation | Two paths must be tested and reconciled | Application event contract becomes the portability layer | Duplicate storage and mismatched retention |
I would make the decision with a replay test, not a feature checklist: given only retained data, can an engineer reconstruct one job's attempts, compare cohort terminal rates, identify the experiment version, and explain why the page fired? Then test deletion. Tenant-linked observability data needs an erasure path because GDPR Article 17 defines a right to erasure under stated conditions; pseudonymous identifiers reduce exposure, but retention and deletion still need deliberate ownership.
The decision rule is plain: keep the attempt ledger close to the transactional system when reconstruction requires database truth, add telemetry when cross-service navigation materially reduces response time, and pay for managed operation when its reduction in on-call burden outweighs export constraints and recurring ingest commitments. Revisit the choice at projected failure-burst volume, not only today's median load.
Further reading
- OWASP Logging Cheat Sheet: https://cheatsheetseries.owasp.org/cheatsheets/Logging_Cheat_Sheet.html
- GDPR Article 17, Right to erasure: https://gdpr-info.eu/art-17-gdpr/
Top comments (0)