Short answer: have the Express Node.js service send structured logs for every completed AI-agent step, then poll the log search for new error events and trigger deduplicated Slack notifications. Attach stable tenant, run, model, token, latency, and outcome fields. For a B2B SaaS agent, keep the small attribution fields longer than bulky request and response bodies. This is the least complex design that can answer both urgent questions: “Which run failed?” and “Which tenant, workflow, and model caused the latency and cost?”
The storage bill is made from event volume, bytes per event, retention time, replication, indexes, and query work. Payload bytes usually deserve attention first because prompts, tool results, and model responses can dwarf the dimensions needed for an alert or a cost rollup. A 12 KB payload retained with every step consumes twelve times the raw capacity of a 1 KB event at the same event rate and retention; replicas and indexes then multiply their own portions of the bill. Before changing a logging backend, remove or separately retain the large fields that do not participate in routing, attribution, or diagnosis.
What does the bill actually preserve?
Start with an inventory, not a product comparison. An agent loop has several boundaries: the inbound request, each model call, each tool call, retries, and the final response. Recording only the final Express error makes Slack easy but destroys cost attribution. Recording every byte forever preserves evidence, yet makes the highest-volume and most sensitive material the default dataset.
For each completed step, the durable event should carry a timestamp, severity, event name, schema version, tenant ID, workflow ID, run ID, step ID, attempt number, model identifier, input and output token counts when the model reports them, elapsed milliseconds measured by the application, and an outcome. Add an error class and a sanitized message on failure. Do not make a free-form message the primary schema; parsers and alerts become dependent on punctuation.
The cost equation is intentionally plain:
def estimated_storage_bytes(events_per_day, average_event_bytes, retention_days, copies=1):
return events_per_day * average_event_bytes * retention_days * copies
This estimate excludes index overhead, compression, query scans, and network transfer, so it is a sizing floor rather than a quote. Its value is diagnostic: if payload removal cuts average_event_bytes, that change affects every retained day and copy. By contrast, shaving a few labels from a compact event is unlikely to move the dominant term.
I would retain three logical classes, even if one storage system implements all three. Attribution events are compact and live long enough for invoice reconciliation and trend comparisons. Diagnostic excerpts have shorter retention and stricter access because they may contain customer data. Alert-delivery records retain the event fingerprint, attempt state, and destination response category so notification failures do not masquerade as application recovery.
| Data class | Needed for | Retention decision | Failure if omitted |
|---|---|---|---|
| Step dimensions and usage | Tenant and workflow allocation | Longest of the three | Costs collapse into an unowned total |
| Sanitized failure detail | Debugging a specific run | Short, access-controlled window | Engineers know a run failed but not why |
| Full prompt or tool body | Rare forensic replay | Off by default or separately approved | Exact reconstruction may be impossible |
| Alert delivery ledger | Deduplication and retry | At least through the retry horizon | Duplicate or silently missing Slack messages |
That table encodes a deliberate loss: full conversational content is not the accounting record.
How should an Express Node.js service send structured logs for a poll?
The Express process should emit after a step finishes, because duration, outcome, and reported usage are then known. The transport may be stdout collected by an agent, a queue, or a batch HTTP exporter; the contract matters more than the pipe. Since this article’s code is constrained to Python, the example below specifies the event builder used by the Node.js service rather than pretending Python middleware is Express middleware.
Keep it boring.
from datetime import datetime, timezone
def agent_step_event(*, tenant_id, workflow_id, run_id, step_id,
model, input_tokens, output_tokens,
elapsed_ms, error=None):
failed = error is not None
event = {
"timestamp": datetime.now(timezone.utc).isoformat(),
"schema_version": 1,
"event_name": "agent.step.completed",
"level": "error" if failed else "info",
"tenant_id": tenant_id,
"workflow_id": workflow_id,
"run_id": run_id,
"step_id": step_id,
"model": model,
"input_tokens": input_tokens,
"output_tokens": output_tokens,
"elapsed_ms": elapsed_ms,
"outcome": "failed" if failed else "succeeded",
}
if failed:
event["error_class"] = type(error).__name__
event["error_message"] = str(error)[:500]
return event
Token counts are measurements supplied by a model response, not universal money values. Convert usage to allocated cost in a versioned rate table outside the raw event, recording which rate-table version produced the result. Otherwise a historical recomputation can silently apply a later rate to an earlier call. For providers or models that do not report comparable usage, preserve the native measurement and mark allocation as incomplete instead of manufacturing precision.
Latency needs the same discipline. Measure the entire agent run for the customer-facing service objective and measure individual steps for diagnosis. Do not add step durations and call the sum wall-clock latency when tools execute concurrently. That arithmetic overstates elapsed time.
Names should be stable and units explicit. Prometheus naming guidance recommends base units and unit suffixes for metrics; the same habit makes log-to-metric transformations less ambiguous. A field named elapsed_ms states its unit, while an eventual metric should follow the conventions of its metrics system rather than inherit an arbitrary log label.
How should polling avoid duplicate Slack alerts?
A poller needs a durable cursor and a delivery ledger. “Query the last five minutes” is not a cursor: overlapping windows duplicate alerts, non-overlapping windows lose late-arriving events, and a worker restart forgets what it has sent. Search by a stable ordered pair such as (timestamp, event_id), include a small overlap for late ingestion, and treat the ledger’s unique fingerprint as the final duplicate barrier. The trade-off is delay: a scheduled search cannot react before the next poll and cannot prove that an event is searchable merely because the application emitted it. This design does not fit a sub-second paging requirement or a stream with sustained volume too high to scan incrementally; use a push-based event consumer for that boundary, while retaining the same event schema and delivery ledger.
The alert path should be at-least-once until the destination accepts the message. Exactly-once delivery across a search service, a local state store, and Slack is not available merely because the loop runs every minute. A crash after Slack accepts a request but before the ledger commits can still produce a duplicate on retry. Make that visible in the design.
import hashlib
import json
import urllib.request
def fingerprint(event):
material = ":".join([
event["tenant_id"],
event["run_id"],
event["step_id"],
str(event.get("attempt", 1)),
event["event_name"],
])
return hashlib.sha256(material.encode("utf-8")).hexdigest()
def post_slack(webhook_url, event):
body = {
"text": (
f"Agent step failed: tenant={event['tenant_id']} "
f"workflow={event['workflow_id']} run={event['run_id']} "
f"step={event['step_id']} latency_ms={event['elapsed_ms']}"
)
}
request = urllib.request.Request(
webhook_url,
data=json.dumps(body).encode("utf-8"),
headers={"Content-Type": "application/json"},
method="POST",
)
with urllib.request.urlopen(request, timeout=10) as response:
if not 200 <= response.status < 300:
raise RuntimeError(f"notification rejected: {response.status}")
def process_page(events, ledger, webhook_url):
for event in events:
key = fingerprint(event)
if ledger.was_delivered(key):
continue
post_slack(webhook_url, event)
ledger.mark_delivered(key)
Keep the webhook secret out of event fields and alert text. Escape or truncate customer-controlled strings before notification, and put identifiers rather than prompts into Slack. The operational message should link an authorized engineer back to internal investigation context where possible; copying the context into a broad chat channel weakens access control.
The query predicate should require both level = error and the event names that represent actionable failures. Severity alone is too broad: a handled tool failure followed by a successful retry may deserve an error event for diagnosis but not a page. Encode alert eligibility as a separate, versioned policy. Then test it with successful retries, permanent failures, late events, duplicate search results, an unavailable destination, and a crash between notification and ledger commit.
Sampling can corrupt the accounting record
OpenTelemetry distinguishes head sampling, decided before a trace completes, from tail sampling, decided after all or part of the trace is available. Either can reduce retained trace volume, but sampled traces cannot be the sole source of tenant billing allocation unless the accounting method explicitly models missing data. An error-biased tail policy also produces a distorted cost dataset because failed and slow executions are retained at a different rate.
Keep compact usage events unsampled. Sample bulky diagnostic traces separately, and propagate run_id and trace_id so retained detail can be joined to the accounting spine. This split also prevents an observability tuning change from rewriting the financial meaning of historical data.
Metrics have another boundary. Counters and histograms are good for rates and latency distributions, while a log event carries the dimensions needed to investigate a run. Putting unbounded tenant_id or run_id values into metric labels creates a growing series set; preserve those identifiers in events and aggregate metrics over bounded dimensions such as workflow class or outcome. The metric should tell an operator that failures rose. The event should identify the affected runs.
Fast alerts can still be wrong. Polling delay, ingestion delay, query timeout, destination throttling, and clock skew each create a distinct failure mode, so measure them separately: newest searchable event age, poll duration, pages processed, cursor lag, delivery attempts, and undelivered ledger entries. A healthy Express service does not prove that its alert worker is healthy. Nor does a successful Slack response prove that a human acted. Alert delivery and incident response are separate states, and combining them hides the precise handoff that operators need to audit.
Watch the cursor.
Deploy the schema before relying on it
Treat the event schema and poller as a compatibility boundary. Add fields before making them required, deploy consumers that tolerate unknown fields, then deploy producers. A schema_version helps migrations, but it does not replace tolerant parsing. Reject malformed events into a bounded quarantine with a counter and an owner; an infinite retry loop turns one bad record into a stalled alert stream.
The useful tests are failure-oriented. Feed the poller two events with the same timestamp and different IDs. Deliver the same page twice. Move the cursor backward. Return a destination timeout after accepting a request. Run an old producer beside a new consumer. Verify that one tenant cannot appear in another tenant’s alert or cost report. These cases expose more than a happy-path screenshot.
Reconcile daily aggregates against the raw compact events before using them for customer-facing allocation. Compare per-tenant event counts, token totals, incomplete-usage counts, and the number of quarantined records. Set an explicit lateness window before finalizing a period because ingestion can arrive after the first rollup. The acceptable window is a business decision tied to reporting latency, not a universal constant.
The final retention policy should be blunt: keep compact, versioned attribution events for the reconciliation period; keep alert delivery state through all retries and audits; expire diagnostic excerpts sooner; and do not retain full prompts and tool bodies by default. When a rare incident requires an exact replay, that last choice costs evidence. It also limits routine storage growth and the amount of customer content exposed to every observability query. For a multi-tenant agent service, that is a defensible trade rather than an accidental deletion policy.
Further reading
- Prometheus, “Metric and label naming”: https://prometheus.io/docs/practices/naming/
- OpenTelemetry, “Sampling”: https://opentelemetry.io/docs/concepts/sampling/
- Slack, “Sending messages using incoming webhooks”: https://api.slack.com/messaging/webhooks
Top comments (0)