For a nightly customer-support data pipeline, use a metrics dashboard to detect that a run is unhealthy, but keep structured logs as the admin's investigation and rollback evidence. Rollback safety is the deciding constraint: a chart can show that failures rose; it cannot, by itself, identify every ticket mutation belonging to one run. Searchable logs can carry the run ID, source record ID, attempted operation, outcome, and schema version needed to reconstruct that boundary.
TL;DR: do not force one telemetry shape to serve two jobs. Alert on bounded, low-cardinality metrics such as run counts, processed records, failures, duplicates, and duration. Investigate with structured logs keyed by run_id and record_id. Keep the business database, not either observability system, as the authority for rollback state.
Should admin analytics use a metrics dashboard or logs search?
I have been paged for missed scheduled jobs and duplicate deliveries. The uncomfortable part is rarely noticing that something is wrong. It is proving the affected set before taking a second action that can make the incident larger.
Consider a nightly job that imports customer-support records, normalizes fields, and updates an admin view. An operator sees the failure counter rise after deployment. The immediate questions are concrete: Did the run start? Which input records were attempted? Which writes succeeded? Did a retry repeat an already completed operation? Which application version emitted the result?
A dashboard answers the first question quickly. Properly structured logs help answer the remaining questions. Neither replaces an idempotency key or a durable mutation ledger in the application database.
That distinction is the invariant: detection may be aggregate, but reversal needs record-level provenance. If a telemetry backend disappears, the application should still prevent a repeated import from applying the same logical mutation twice. Observability explains the decision; transactional state enforces it.
The split that survives an incident
The simplest useful architecture emits both signals from the same processing path. A counter increments for the aggregate outcome, while one structured event records the identifiers an operator will search later. The two representations share semantics, not storage assumptions.
| Operator task | Primary signal | Reason |
|---|---|---|
| Confirm the scheduled run started | Metric | A run count is fast to scan and alert on |
| Detect abnormal failures or duration | Metric | Aggregation exposes change without reading every event |
| Find all records touched by one run | Structured log | The query needs a run_id boundary |
| Explain a duplicate attempt | Structured log plus durable state | Attempt history needs context; prevention needs an idempotency record |
| Approve a rollback | Durable state reconciled with logs | The rollback must act on authoritative business state |
This division also limits a common operational mistake: putting record IDs into metric labels. A pipeline may process an open-ended set of records, so those identifiers belong in events intended for search, not in an aggregate series. Keep the dashboard legible. Keep the evidence specific.
The trace sampling decision is separate. OpenTelemetry documents head sampling as a decision made early and tail sampling as a decision made after more of a trace is available. Sampling can control trace volume, but a sampled trace is a poor sole record of rollback scope: the exact event an operator needs may not be retained. Required mutation evidence needs an explicit retention policy independent of trace sampling.
Encode the rollback boundary in the write path
The preventative path starts before telemetry. Give each scheduled execution a stable run ID, give each logical import operation a stable idempotency key, and commit the business mutation together with its durable processed marker according to the datastore's transaction model. Emit telemetry after the result is known.
The following Go example shows the event shape and the control flow. Store.ApplyOnce is deliberately a domain interface: its implementation must atomically decide whether an operation was already applied and perform the mutation when it was not.
package pipeline
import (
"context"
"log/slog"
)
type Ticket struct {
SourceID string
Status string
}
type Store interface {
ApplyOnce(ctx context.Context, idempotencyKey string, ticket Ticket) (applied bool, err error)
}
type Counter interface {
Add(ctx context.Context, value int64, outcome string)
}
type Importer struct {
store Store
results Counter
logger *slog.Logger
}
func (i *Importer) Import(ctx context.Context, runID string, ticket Ticket) error {
key := runID + ":" + ticket.SourceID
applied, err := i.store.ApplyOnce(ctx, key, ticket)
if err != nil {
i.results.Add(ctx, 1, "failed")
i.logger.ErrorContext(ctx, "ticket import failed",
"run_id", runID,
"record_id", ticket.SourceID,
"idempotency_key", key,
"operation", "ticket_import",
"outcome", "failed",
"schema_version", 1,
"error", err,
)
return err
}
outcome := "applied"
if !applied {
outcome = "duplicate_skipped"
}
i.results.Add(ctx, 1, outcome)
i.logger.InfoContext(ctx, "ticket import completed",
"run_id", runID,
"record_id", ticket.SourceID,
"idempotency_key", key,
"operation", "ticket_import",
"outcome", outcome,
"schema_version", 1,
)
return nil
}
One detail matters: the success event is written after ApplyOnce returns. Logging an intention before the transaction and later treating that event as proof of a committed mutation creates a false rollback set. Intent events can still be useful, but their names and outcomes must make the distinction explicit.
Errors also need care. A searchable error string helps diagnosis, but rollback queries should use stable fields such as operation and outcome. Free-form messages change during refactors. IDs and versioned event schemas form the durable query contract.
Test recovery, not just charts
A dashboard screenshot is not a rollback test. Before a release, run a fixture containing a new record, a record that will fail validation, and the same logical record delivered twice. Then verify the aggregate counts, search the emitted events by run ID, and reconcile every reported success with durable state. Next, rerun the identical fixture. The second execution should not create another logical mutation. This is where an idempotency reflex pays for itself: the retry becomes a normal state transition with a visible duplicate_skipped outcome instead of an emergency procedure. Now exercise the reversal in a non-production environment: build the candidate set from the run events, compare each candidate with current business state, and refuse the reversal when a later run owns the current value. This rehearsal tests the dangerous gap between finding an event and proving that its corresponding mutation is still safe to undo. It also leaves an artifact the reviewer can inspect instead of asking them to trust that the rollback query looks plausible.
Short test, long value.
Deployment should be reversible in two independent senses. The application release can roll back without losing the ability to read events emitted by the newer schema, and the data operation can reverse only the records proven to belong to the affected run. Adding fields is easier to tolerate than silently changing their meaning. A schema_version field gives consumers an explicit branch when semantics must change.
The runbook should require a three-way check before reversal: identify the suspect deployment or configuration change, derive the candidate record set from structured events, and reconcile that set against authoritative state. If those sets disagree, stop. An incomplete rollback is bad; deleting a correct update from another run is worse.
When this split is unnecessary
The limitation of the split is extra operational surface. For a tiny internal job where an operator only needs to know whether the last execution succeeded, a metric and a link to the job result may be enough. Record-level search adds retention, access control, and schema ownership. The trade-off is unjustified when nobody needs record-level reconstruction, so do not collect fields nobody will use.
No search layer fixes missing provenance.
The opposite case also exists. If administrators need audited historical reporting, complex business filters, or guaranteed reconstruction over a defined retention period, logs alone are not the analytics database. Build a durable reporting model or mutation ledger and treat telemetry as navigation into it.
The practical decision rule is narrow: choose metrics for bounded questions about rate, count, duration, and health; choose structured logs for high-detail investigation by execution and record identifiers; choose durable application data for rollback authority. For a nightly support pipeline, that combination is easier to operate than asking a dashboard or a log index to impersonate all three.
Top comments (0)