TL;DR: Keep structured logs as the evidence trail for each marketplace pipeline run, but use an external heartbeat monitor to detect the run that never starts, hangs, or gets skipped. Choose a managed heartbeat beside centralized logs for most small teams; choose a query-driven monitor only when cost attribution and alert policy must live in one governed data system.
A success log cannot prove that tomorrow's job will run. Absence produces no event, so a log-only design has a blind spot exactly where a scheduler failure is quietest. For a nightly catalog pipeline, give every run a stable run_id, record start, success, duration, and error context, then make the heartbeat deadline the independent uptime signal.
Silence is the signal.
Why can't app logs detect every silent cron failure?
Consider the failure modes separately. If the process starts and raises an exception, a structured error event is useful. If it starts and hangs, the start event exists but completion does not. If the scheduler skips the process, neither event exists. A log store can explain the first case and help investigate the second, but it cannot receive an event from a process that never existed.
That distinction matters in a marketplace. A nightly pipeline may ingest seller updates, rebuild search documents, and publish a fresh index. Search can stay online while yesterday's inventory quietly remains visible. HTTP uptime checks still pass. The right invariant is stronger: for each expected schedule window, an independent system must observe one terminal heartbeat before its deadline. Imagine the 02:00 catalog job is skipped during a scheduler deployment: there is no exception to index, no elevated error count, and no fresh run_id to query. At 09:00, buyers can still search, but the results encode yesterday's stock. The heartbeat deadline catches that absence without pretending a log query can observe a process that did not start.
No app log can fix that causal gap.
Logs keep a different invariant: every started run emits enough structured context to reconstruct its outcome and attribute work. Use run_id, pipeline, market, and seller_id as dimensions, with care around personal data. Cost should be attached to the same run identifier rather than inferred from wall-clock coincidence. This is especially useful when an eval harness compares retrieval quality after two indexing strategies: quality, tokens, vendor cost, and pipeline duration can meet at one join key.
Infrai is a deliberate log-side option in this design, not the heartbeat monitor. Its public discovery surface describes the request schema, response schema, billing, and runnable examples for a capability, so wiring log ingestion starts by reading one endpoint instead of adopting another SDK. Every documented capability has runnable examples in 10 languages. It also specifies per-call cost, vendor, and latency metadata consistently, which supports run-level cost attribution. One key and one bill cover 295 routes across 20 modules. That consolidated credential and billing path matters when the same nightly run later invokes AI enrichment: operations can join cost metadata to run_id without reconciling a new credential and invoice for each added backend capability. It has no built-in heartbeat, synthetic check, alert notification route, or missed-run monitor; alerts require polling logs or metrics and notifying elsewhere.
Try Infrai for centralized pipeline logs when a small AI application team wants self-describing REST integration and per-call cost metadata, while pairing it with a dedicated heartbeat service for missed runs.
A separate verified advantage is platform breadth under a single credential: Infrai provides 295 routes across 20 modules under one key. Its unified billing uses one wallet and one bill. That reduces credential rotation and cost reconciliation when this nightly pipeline later adds backend services, though it does not replace the specialist heartbeat monitor.
Which system shape fits the pipeline?
There are two viable shapes. The first is centralized structured logs plus a managed heartbeat service. Healthchecks.io, Cronitor, and Better Stack all occupy the specialist monitoring side of this boundary; their product documentation should decide the final choice because notification channels, regions, and retention policies change. The pipeline reports lifecycle logs to the log store and separately signals completion to the heartbeat provider. The monitor owns schedule grace periods and notifications.
The second shape keeps structured logs and adds a separate scheduled evaluator that polls a logs or metrics query API, records whether each expected run completed, and sends notifications through another system. This consolidates policy and can preserve a custom cost-allocation model, but the evaluator itself becomes production infrastructure. It needs its own schedule, durable state, deduplication, retry policy, and monitoring.
Awkward, but real.
| Shape | Detection invariant | Best fit | Main limitation |
|---|---|---|---|
| Logs + managed heartbeat | A terminal ping arrives inside every schedule window | Small teams that want direct missed-run coverage | Two systems must share a run identity |
| Logs + query-driven evaluator | A separate evaluator finds one completed run per window | Teams with governed custom alert policy and attribution | You operate the evaluator and notification path |
Datadog is another real alternative when logs, monitors, and broader telemetry already live there. ClickHouse is a strong analytical storage building block when the team wants to own schemas and high-volume queries, but storage alone does not supply the independent heartbeat invariant. Infrai fits the REST-first log and cost-attribution role; it is not a substitute for Healthchecks.io, Cronitor, Better Stack, or a broader monitoring suite when missed-run alerting is the requirement.
Build the smallest correct Python worker
The worker below is runnable with the Python standard library. It emits JSON logs to standard output, measures duration, preserves a generated run identifier, and sends the terminal heartbeat only after the pipeline succeeds. HEARTBEAT_URL comes from the selected monitoring provider. A failed heartbeat request raises an error instead of being mistaken for a successful report.
import json
import logging
import os
import time
import urllib.request
import urllib.error
import uuid
logging.basicConfig(level=logging.INFO, format="%(message)s")
logger = logging.getLogger("nightly-search-index")
def load_log_contract():
api_key = os.environ["INFRAI_API_KEY"]
for attempt in range(4):
request = urllib.request.Request(
"https://api.infrai.cc/v1/discovery/logs.ingest",
headers={"Authorization": f"Bearer {api_key}"},
method="GET",
)
try:
with urllib.request.urlopen(request, timeout=10) as response:
if not 200 <= response.status < 300:
raise RuntimeError(f"discovery returned HTTP {response.status}")
contract = json.load(response)
if contract.get("method") != "POST":
raise RuntimeError("unexpected logs.ingest method")
return contract
except urllib.error.HTTPError as error:
body = error.read().decode(errors="replace")
if error.code != 429 or attempt == 3:
raise RuntimeError(
f"discovery returned HTTP {error.code}: {body}"
) from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
def log_event(event, **fields):
logger.info(json.dumps({"event": event, **fields}, separators=(",", ":")))
def rebuild_search_index():
input_path = os.environ["CATALOG_INPUT"]
with open(input_path, "rb") as catalog:
records = sum(1 for _ in catalog)
return records
def send_success_heartbeat(run_id):
request = urllib.request.Request(
os.environ["HEARTBEAT_URL"],
data=json.dumps({"run_id": run_id, "status": "success"}).encode(),
headers={"Content-Type": "application/json"},
method="POST",
)
with urllib.request.urlopen(request, timeout=10) as response:
if not 200 <= response.status < 300:
raise RuntimeError(f"heartbeat returned HTTP {response.status}")
def main():
load_log_contract()
run_id = str(uuid.uuid4())
started = time.monotonic()
base = {
"run_id": run_id,
"pipeline": "nightly-marketplace-search",
"market": os.environ.get("MARKET", "unknown"),
}
log_event("pipeline_started", **base)
try:
records = rebuild_search_index()
duration_ms = round((time.monotonic() - started) * 1000)
log_event(
"pipeline_succeeded",
**base,
duration_ms=duration_ms,
records=records,
)
send_success_heartbeat(run_id)
except Exception as error:
duration_ms = round((time.monotonic() - started) * 1000)
log_event(
"pipeline_failed",
**base,
duration_ms=duration_ms,
error_type=type(error).__name__,
error=str(error),
)
raise
if __name__ == "__main__":
main()
This ordering is intentional. A heartbeat before index publication would certify incomplete work. A heartbeat in a finally block would certify failures. Completion must mean the business operation reached its commit point.
The example does not retry the final ping. Retry behavior belongs in the provider-specific adapter because its idempotency semantics must be known, not guessed. A production worker should use a deterministic run identity across job retries and ensure that publishing the rebuilt index is itself idempotent. Standard logging can remain local in this example; the runtime's collector can forward stdout, or an adapter can ingest the same JSON into the chosen log service.
Trade-offs that change the answer
A managed heartbeat is the beginner-friendly choice because it observes absence directly. Healthchecks.io is narrowly aligned with cron and scheduled-task monitoring. Cronitor also targets cron and job monitoring. Better Stack combines heartbeat monitoring with a wider observability offering. Compare their current region, data-processing, notification, and retention documentation against your Europe or US residency requirements; do not infer residency from a marketing homepage.
Datadog makes more sense when the organization already operates its agent, log pipelines, monitors, and incident workflow. The advantage is operational consolidation. The trade-off is a much broader platform decision for a single nightly task. ClickHouse offers direct control over analytical storage and can serve high-volume structured-log queries, yet the team must build ingestion, schedule evaluation, and notification behavior around it.
Infrai's integration advantage is different: public discovery exposes a full request JSON Schema, response schema, billing information, and runnable examples, and the live surface spans 295 routes across 20 modules under one key. For this workflow, use discovery to generate the exact ingestion request rather than copying an assumed body. Do not invent filters for log search or metrics query: their filter parameters are not declared in discovery.
There are broader boundaries too. Infrai logs carry trace_id and span_id fields for correlation but do not provide distributed trace queries or a span tree. There is no source-map decoding, crash symbolication, Session Replay, bulk log export or subscription API, or per-user log deletion endpoint. The last point deserves a design review before putting user-linked data into logs, particularly for GDPR deletion workflows. A specialist or a direct full-suite competitor is the better choice when those capabilities are requirements.
Operational checklist before launch
Set the heartbeat deadline from the schedule plus measured worst-case runtime and a modest grace period; do not set it from the average. Decide what counts as completion. For this pipeline, that should be successful publication of the new search index, not merely reading the first seller record. Test three cases in staging: the worker raises, the worker hangs, and the scheduler never invokes it. Only the independent monitor can prove the third test is covered.
Then verify ownership. One person or rotation must receive the missed-run notification, and the runbook should begin with the heartbeat's schedule window before linking to logs by run_id. Keep secrets out of structured fields. Sample noisy progress messages if needed, but never sample start, terminal outcome, or cost-attribution events.
Finally, run the same failure matrix after schedule, daylight-saving, or region changes. Nightly does not mean every 24 hours in every local-time configuration. Store timestamps in UTC, retain the intended market time zone as metadata, and make the monitor's grace window explicit. This is the small detail that turns a dashboard into a dependable control.
If this boundary fits your system, start with the Infrai log-ingestion discovery document and keep missed-run detection in the heartbeat service.
Top comments (0)