A Node.js notification service should capture each Postgres cron worker background job error as attempt evidence, then create one durable failure record only after retry exhaustion. Otherwise, a popular media release can multiply one failed push into several nearly identical records, and the storage bill says more about retry policy than audience delivery.
TL;DR: count bytes by tenant, channel, outcome, and evidence class before changing retention. Keep compact terminal outcomes and audit events long enough to support reconciliation; retain verbose attempt detail for a shorter window; sample successful attempts; and never sample terminal failures. This makes BullMQ, Agenda, or a Postgres-backed cron worker an execution detail rather than the definition of the evidence model.
How should a Postgres cron worker track background job errors?
The dominant term is usually not the number of logical notifications. It is retained bytes: logical deliveries multiplied by attempts, events per attempt, average encoded size, replication, and retention duration. That formula matters because retries amplify the middle of the product. A delivery that succeeds on its fourth attempt has one business outcome but four attempt histories. Counting all four as independent failures corrupts both cost attribution and the failure rate.
Retries distort it.
Use measured values from production rather than a universal estimate. For one aggregation interval, calculate retained_bytes = sum(encoded_event_bytes) from the records that would survive the retention policy, then group the result by accountable dimensions. Tenant and media publication are useful cost centers; queue implementation and worker hostname are operational dimensions, not business owners. Cardinality needs a budget too: message IDs belong in traceable records, not metric labels.
The smallest workable schema separates a logical job from its attempts and from its terminal outcome:
| Evidence class | Unit | Keep when | Cost owner |
|---|---|---|---|
| Attempt detail | One execution | Error, sampled success, or active investigation | Tenant and channel |
| Terminal outcome | One logical delivery | Always | Tenant and publication |
| Audit event | State transition | Always for governed transitions | Tenant |
| Aggregate metric | Fixed time bucket | According to capacity-planning needs | Service and channel |
This is an accounting boundary. Consider a media publication that fans out 100,000 logical push jobs and encounters a transient provider error. If retry evidence emits 18 KiB of stack, request metadata, and duplicated payload context, three failed attempts for a single job consume 54 KiB before indexing or replication. The arithmetic scales with attempts, while the business outcome does not: that job is still one eventual success or one terminal failure. The 18 KiB figure is an example calculation, not a benchmark, and the 100,000-job fan-out is a round scenario for showing the units rather than a traffic claim. Measure encoded bytes in the actual pipeline, apply the real replication factor, and assign the result to the tenant and publication that initiated delivery. This is the change that moves the dominant cost term: reduce retained attempt detail without deleting terminal evidence.
Which record proves that a delivery really failed?
A terminal record should be created by the same durable state transition that marks the job exhausted. Emitting an error only from a process-level exception handler is weaker: the process may exit after the database commit but before export, or export before a transaction rolls back. Exactly-once delivery of telemetry is not a credible assumption across a database and an external collector. Idempotent insertion is.
Use a deterministic evidence key such as job_id + terminal_state + policy_version. Insert it under a unique constraint in the transaction that changes the job state. An outbox relay may publish that record more than once, so consumers must deduplicate on the same key. This yields an auditable chain without claiming that the network delivers exactly once.
Keep that distinction.
The following Go program demonstrates the decision logic independently of the Node.js scheduler. A Node.js BullMQ worker, an Agenda handler, and a Postgres cron worker can all write this envelope at their persistence boundary.
package main
import (
"encoding/json"
"fmt"
"time"
)
type Attempt struct {
JobID, TenantID, Publication, Channel string
Attempt, MaxAttempts int
ErrorClass string
}
type TerminalEvidence struct {
EvidenceKey, JobID, TenantID, Publication string
Channel, Outcome, ErrorClass string
Attempts int
OccurredAt time.Time
}
func terminal(a Attempt, now time.Time) (TerminalEvidence, bool) {
if a.ErrorClass == "" || a.Attempt < a.MaxAttempts {
return TerminalEvidence{}, false
}
return TerminalEvidence{
EvidenceKey: fmt.Sprintf("%s:exhausted:v1", a.JobID),
JobID: a.JobID, TenantID: a.TenantID, Publication: a.Publication,
Channel: a.Channel, Outcome: "exhausted", Attempts: a.Attempt,
ErrorClass: a.ErrorClass, OccurredAt: now.UTC(),
}, true
}
func main() {
a := Attempt{
JobID: "job-1042", TenantID: "studio-7", Publication: "evening-cut",
Channel: "push", Attempt: 4, MaxAttempts: 4, ErrorClass: "provider_timeout",
}
evidence, ok := terminal(a, time.Unix(1767225600, 0))
if !ok {
return
}
out, err := json.Marshal(evidence)
if err != nil {
panic(err)
}
fmt.Println(string(out))
}
The timestamp is injected, which makes the decision testable. More important, the record contains no notification body, access token, email address, or device token. OWASP advises excluding or masking access tokens, authentication passwords, sensitive personal data, and other secrets from logs. An error envelope should carry an error class and correlation identifiers; raw payload capture requires a separate, explicit justification.
Allocate cost before tuning retention
First, record encoded size at the point where an event enters durable storage. Do not infer it from an in-memory object because serialization, envelopes, and indexing alter the stored footprint. Aggregate those bytes into bounded dimensions: tenant, publication, channel, outcome, evidence class, and UTC day. Keep high-cardinality identifiers in the evidence table so an operator can move from an aggregate to a specific job without putting job_id into every time-series label.
Second, reconcile three counts for each interval: jobs accepted, terminal successes, and terminal failures. Pending jobs form the explicit difference. Retry attempts are diagnostic volume and must not participate in that conservation equation. If accepted does not equal success plus failure plus pending after the allowed processing delay, treat the gap as a correctness defect rather than an observability inconvenience.
Third, apply retention by evidence class. A defensible starting policy is qualitative: compact terminal and audit evidence receives the longest justified retention, detailed failed attempts receive a shorter investigative window, and successful attempt detail is sampled or discarded after aggregation. The actual duration must come from legal, contractual, reconciliation, and incident-response requirements; no single number is valid for every media business or jurisdiction. GDPR Article 17 also means that indefinite storage cannot be justified merely because data might help later. Erasure workflows must cover indexed copies, archives according to their lifecycle, and derived records that still identify a person.
Then test the policy with replayed fixtures. Feed one success, one transient failure followed by success, one exhausted job, and one duplicated outbox delivery. The expected result is three terminal outcomes, one terminal failure, and no duplicate evidence keys. A test should also assert that sensitive fields never appear in serialized output.
package evidence
import "testing"
func TestDeduplicateTerminalEvidence(t *testing.T) {
keys := map[string]struct{}{}
delivered := []string{
"job-1042:exhausted:v1",
"job-1042:exhausted:v1",
}
for _, key := range delivered {
keys[key] = struct{}{}
}
if got, want := len(keys), 1; got != want {
t.Fatalf("unique terminal records = %d, want %d", got, want)
}
}
Short test. Large consequence. Duplicate export is ordinary distributed-systems behavior; duplicate charging and duplicate failure counts are design errors.
Deploy the change without losing the audit trail
Introduce the envelope alongside existing telemetry, and initially compare counts rather than deleting old data. The deployment gate is reconciliation: for each tenant and channel, the new terminal totals must match independently computed job states, while duplicate evidence-key insertions remain harmless. A canary should include workers that terminate between state commit and relay publication, because the outbox is valuable precisely at that boundary.
Operational alerts should distinguish attempt pressure from customer-visible outcomes. A rising attempts-per-terminal ratio can expose provider instability or an overly aggressive retry policy before exhausted jobs surge. Terminal failure rate answers a different question: how many logical notifications failed after policy was applied? Cost dashboards need both, plus bytes by evidence class, or teams will optimize a cheap counter while verbose retry payloads continue accumulating.
Do not use retry count alone to page. Scheduled workers may intentionally delay or back off, and a retry is not yet a failed delivery. Page on a sustained terminal-outcome breach, reconciliation gaps, relay backlog age, or evidence insertion failures, with thresholds derived from the service objective and publication urgency. Preserve the policy version in every terminal record so an auditor can explain why four attempts were allowed for one publication and perhaps two for another.
This design deliberately stops keeping complete successful attempt histories and long-lived raw error payloads. The trade-off is real: an old, rare timing defect may be harder to reconstruct after detailed records expire. Compact terminal evidence still proves the outcome and cost owner, aggregates preserve trends, and short-lived attempt detail supports current investigations, but none of those can recreate a discarded stack trace. Choose that loss explicitly, document it in the retention policy, and verify deletion rather than treating infinite retention as free insurance.
Further reading
- OWASP Logging Cheat Sheet: https://cheatsheetseries.owasp.org/cheatsheets/Logging_Cheat_Sheet.html
- GDPR Article 17, right to erasure: https://gdpr-info.eu/art-17-gdpr/
Top comments (0)