Short answer: for a small Node.js application, keep failure alerting boring. Poll one completion metric after the nightly e-commerce pipeline's deadline, require the failure state to persist across two observations, and send one deduplicated webhook containing a run identifier. Retain a compact failure record longer than the searchable logs. This gives an operator enough evidence to act without paying to keep every successful debug line searchable.
The dominant bill is usually the data kept and searched, not the tiny polling function. Quantify that term first. If a pipeline emits 2 GB of structured logs per night, 30 days means 60 GB retained before replicas, indexes, or compression alter the storage system's actual footprint. Reducing a poll from five minutes to one minute changes request volume from 288 to 1,440 queries per day, but it does not change those 60 GB. Retention and log volume are therefore the first levers to inspect; poll frequency follows the operational deadline.
This design deliberately stops keeping verbose success logs after a short diagnostic window. When a delayed defect is discovered later, the trade-off is real: the team may retain the run outcome and alert history but lose line-by-line reconstruction of an old successful run.
That loss is intentional.
Should a small app poll a metrics API for failure alerting?
A nightly catalog pipeline can fail loudly, stall without exiting, or finish while producing an incomplete result. A raw process-error counter covers only the first case. The alerting signal needs to describe the job outcome that the business cares about: one expected run, by a known deadline, with a terminal state and a plausible output count.
That makes the useful record small. Give each run an immutable run_id; record started_at, finished_at, status, records_read, records_written, and a bounded error_class. Avoid putting customer email addresses, phone numbers, order notes, or raw payloads in either the metric labels or the webhook. Those values increase cardinality, leak regulated or sensitive context into notification systems, and still do not answer the first operational question: which run failed?
Google's SRE guidance separates symptoms from causes and identifies latency, traffic, errors, and saturation as four useful signals. For this batch job, the alert should begin with the symptom: the expected run is late or failed. CPU saturation and queue depth can help diagnosis, but paging on every causal indicator creates duplicate noise.
Two consecutive observations are a modest noise filter, not a universal constant. Use them only when the extra polling interval fits inside the response budget. A ten-minute confirmation delay is unacceptable if checkout inventory must be corrected in five minutes; it may be reasonable when the catalog publication deadline has an hour of slack.
For a small app with one nightly job, polling a stable query endpoint is a defensible alternative to adopting a complete on-call suite. A Lambda-style scheduled function can own the poll without living inside the Node.js process being monitored. Keep the scheduler independent: an application outage should not stop the code responsible for detecting that outage.
The retention model comes before the poller
Split the data by purpose instead of assigning one retention period to everything.
| Data class | Example | Retention decision | What is lost when removed |
|---|---|---|---|
| Run outcome | status, timestamps, counts, run ID | Keep through the audit and trend window | Long-range reliability and completeness analysis |
| Failure evidence | bounded error class, failed stage, code revision | Keep long enough for investigation and recurrence analysis | Comparison with an older failure |
| Verbose diagnostics | per-item debug events and stack context | Keep for a short operational window | Late, line-by-line reconstruction |
| Alert state | first seen, confirmations, notification ID | Keep through deduplication and review | Proof that notification logic behaved as intended |
The table is also a compliance boundary. A run outcome can often remain useful after detailed event content has expired. Redaction should happen before ingestion, because shortening retention later does not undo disclosure to an index, webhook receiver, or notification transcript.
Redact first.
Estimate daily volume from measured event bytes and event counts, then multiply by the proposed searchable window. Do the same for failure-only evidence. A representative calculation is more useful than a provider price comparison: 2 GB/night times 30 nights is 60 GB, while a 10 KB outcome document per night is about 300 KB across the same period. These are worked inputs, not benchmark claims. Replace both with production measurements, including index overhead and replication reported by the chosen storage layer.
There is a delivery lesson here. An alert accepted by a webhook is not the same as a human notification delivered, just as an OTP accepted by a messaging gateway is not proof that it reached a handset. Preserve the receiver's correlation identifier and response class, retry transient delivery failures with bounded backoff, and route repeated permanent failures to a separately observable sink. Never recursively alert through the same broken path.
A minimal poll and deduplication boundary
Keep metric evaluation separate from notification delivery. The poller reads an aggregate result; a small state store tracks the last observation and the last notified run. The webhook receives a compact event, not the full logs. This Python example shows the boundary without binding it to a metrics vendor or incident product:
import json
import os
import time
import urllib.request
from dataclasses import dataclass
from datetime import datetime, timezone
@dataclass(frozen=True)
class Observation:
run_id: str
status: str
deadline_missed: bool
def get_json(url: str, token: str) -> dict:
request = urllib.request.Request(
url,
headers={"Authorization": f"Bearer {token}", "Accept": "application/json"},
)
with urllib.request.urlopen(request, timeout=10) as response:
return json.load(response)
def post_json(url: str, payload: dict) -> None:
body = json.dumps(payload, separators=(",", ":")).encode("utf-8")
request = urllib.request.Request(
url,
data=body,
method="POST",
headers={"Content-Type": "application/json"},
)
with urllib.request.urlopen(request, timeout=10) as response:
if response.status < 200 or response.status >= 300:
raise RuntimeError(f"webhook returned status {response.status}")
def evaluate(raw: dict) -> Observation:
return Observation(
run_id=str(raw["run_id"]),
status=str(raw["status"]),
deadline_missed=bool(raw["deadline_missed"]),
)
def handler(event, context):
current = evaluate(get_json(os.environ["METRIC_QUERY_URL"], os.environ["METRIC_TOKEN"]))
previous = event.get("previous", {})
failed = current.status == "failed" or current.deadline_missed
confirmed = failed and previous.get("run_id") == current.run_id and previous.get("failed") is True
already_sent = event.get("last_notified_run_id") == current.run_id
if confirmed and not already_sent:
post_json(
os.environ["ALERT_WEBHOOK_URL"],
{
"event_type": "nightly_pipeline_failure",
"run_id": current.run_id,
"observed_at": datetime.now(timezone.utc).isoformat(),
"status": current.status,
"deadline_missed": current.deadline_missed,
},
)
return {
"previous": {"run_id": current.run_id, "failed": failed},
"last_notified_run_id": current.run_id if confirmed else event.get("last_notified_run_id"),
}
The surrounding runtime must persist the returned state atomically; passing it back as the next event is only a generic representation of that contract. Concurrent invocations need a conditional write keyed by run_id, or both can observe “not sent” and emit duplicates. Secrets belong in the runtime's secret facility, and logs must not print tokens or complete webhook bodies.
The poller's own failure also needs treatment. A timeout or malformed response is an unknown monitoring state, not evidence that the pipeline failed. Count consecutive poll errors separately and notify through an independent dead-man path when the monitor has not completed successfully within its own deadline.
Unknown is not healthy.
Signal quality requires boring tests
Start with table-driven cases for failed, deadline_missed, and successful completion. Then test the sequence, because most alert bugs live between observations: failure then failure sends once; failure then success sends nothing; two concurrent confirmations still produce one notification; a new failed run_id can notify even if yesterday's run already did.
Inject three boundary failures in staging: a query timeout, a webhook timeout after the receiver accepted the request, and a state-store write conflict. The second case is awkward. Retrying without an idempotency key may duplicate the alert, while refusing to retry may lose it. Put a stable event identifier derived from the event type and run_id in the payload, and require the receiver or an adapter to deduplicate it.
Deployment should begin in record-only mode. Compare would-have-alerted events with known pipeline outcomes for several normal cycles, then enable notification with a long repeat interval. Review false positives, missed outcomes, time-to-detection, and notification delivery evidence after any pipeline schedule change.
No page should be unactionable. The message should name the failed run, the missed condition, the observation time, and the internal log-search key. It should not paste the entire exception or a customer record into a channel with broader membership.
Keep it dull.
Choosing the least complex operating model
A scheduled function plus a durable state row is enough when there is one pipeline, one responder group, and a forgiving response window. It has a small surface area, but the team owns scheduling, authentication, retries, deduplication, escalation, and the health of the monitor itself.
Move to a fuller on-call system when acknowledgment, rotations, escalation policies, delivery-channel redundancy, and audit history become requirements. That is a capability threshold, not a company-size contest. Keep the metric query and alert event schema portable so notification routing can change without rewriting pipeline instrumentation.
The final decision rule is simple: retain outcomes for analysis, retain bounded failure evidence for investigation, and expire verbose successful logs as soon as their diagnostic value falls below their storage and disclosure cost. Poll only as fast as the response objective demands. The cheapest alert is useless if it is noisy, late, or impossible to reconstruct; the largest log archive is equally useless if responders cannot identify one failed run.
Top comments (0)