A cheap log store is still expensive if every delayed truck manifest wakes someone up. For a startup app, centralized logging should keep only the logs needed to decide whether work happened: schedule due, run completion, and result count. Alert only when a due import has neither completed nor produced results after its allowed lateness. This favors signal quality over raw volume, keeps the storage layer replaceable, and answers the operational question directly.
TL;DR: emit one event when a scheduled logistics import starts and one when it finishes; include a stable job name, scheduled time, status, result count, duration, and run ID. A separate evaluator tracks expected schedules and raises an alert for missing completion, zero results, or an explicit failure. Do not page on the absence of arbitrary application log lines.
What should centralized logging capture from a startup app's logs?
Silence is ambiguous. A carrier feed may have no shipments for a window, the scheduler may not have launched the task, the worker may still be processing a large file, or log delivery may be delayed. A rule that says "no logs for ten minutes" collapses those states into one page. It is easy to configure and hard to trust.
Noise wins otherwise.
The event contract should instead describe business progress. In a Next.js front end with a Node.js scheduler and a separate worker, import_started proves dispatch, while import_finished records the terminal status and number of accepted records. The monitor joins those events to the schedule. Centralization matters because each process may emit a different part of the story, but the detection rule should not depend on which one happened to write last. The same boundary also works when the worker is Python, which is useful when a notebook prototype grows into a scheduled production job.
Keep the payload narrow. A useful event might look like this:
import json
event = {
"event": "import_finished",
"job": "carrier_manifest",
"scheduled_at": "2026-09-23T01:00:00Z",
"run_id": "run-01842",
"status": "ok",
"result_count": 317,
"duration_ms": 42810,
}
print(json.dumps(event, separators=(",", ":")))
There is no manifest body, customer address, or carrier credential in that event. GDPR Article 5 includes data minimization: personal data should be adequate, relevant, and limited to what is necessary. For this alert, counts and identifiers are enough; shipment contents are not.
Build the evaluator before choosing storage
The next example is deliberately plain Python. It consumes normalized events that could have come from files, a queue, or a centralized log API. The storage adapter is outside the decision rule, so a notebook test and the production job exercise the same function.
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
from typing import Iterable
@dataclass(frozen=True)
class ImportEvent:
event: str
job: str
scheduled_at: datetime
status: str | None = None
result_count: int | None = None
def parse_event(row: dict) -> ImportEvent:
return ImportEvent(
event=row["event"],
job=row["job"],
scheduled_at=datetime.fromisoformat(
row["scheduled_at"].replace("Z", "+00:00")
),
status=row.get("status"),
result_count=row.get("result_count"),
)
def evaluate_import(
job: str,
scheduled_at: datetime,
now: datetime,
allowed_lateness: timedelta,
events: Iterable[ImportEvent],
) -> str | None:
relevant = [
item
for item in events
if item.job == job and item.scheduled_at == scheduled_at
]
finishes = [item for item in relevant if item.event == "import_finished"]
if finishes:
latest = finishes[-1]
if latest.status != "ok":
return f"{job}: import finished with status={latest.status}"
if latest.result_count == 0:
return f"{job}: import completed with zero results"
return None
if now > scheduled_at + allowed_lateness:
return f"{job}: no completion after allowed lateness"
return None
if __name__ == "__main__":
due = datetime(2026, 9, 23, 1, 0, tzinfo=timezone.utc)
rows = [
{
"event": "import_started",
"job": "carrier_manifest",
"scheduled_at": "2026-09-23T01:00:00Z",
}
]
alert = evaluate_import(
job="carrier_manifest",
scheduled_at=due,
now=datetime(2026, 9, 23, 1, 21, tzinfo=timezone.utc),
allowed_lateness=timedelta(minutes=20),
events=[parse_event(row) for row in rows],
)
print(alert)
This example separates three outcomes on purpose. An explicit failed finish is actionable immediately. A successful finish with zero results may indicate an upstream empty feed or a parsing problem. A missing finish becomes actionable only after the schedule-specific lateness window. Short rule, real distinctions.
It has a real limitation. This rule is not appropriate for an unscheduled, continuously flowing source because there is no due time to evaluate; that system needs a throughput or freshness objective instead. It also cannot prove that every imported record is correct. The trade-off is deliberate: detect stopped production with a small, explainable signal, then leave record-level quality to a separate check.
For production, make repeated evaluation idempotent by deriving a deduplication key from the job, scheduled time, and alert reason. Send a recovery notification when a later successful completion arrives. If imports can legitimately be empty, replace the universal zero rule with an expectation owned by that feed: for example, allow zero on a documented non-operating day. The exception belongs in configuration and in eval cases, not in a prompt or an operator's memory.
Tune signal quality with an eval set
I would choose the lateness window from the workflow contract, then test it against representative event sequences before connecting paging. This is the same habit that keeps an AI feature honest: write the eval cases first, inspect false positives and false negatives, and change one rule at a time. No model call is needed here. Deterministic logic is cheaper to operate, easier to replay, and easier to explain during an incident.
That separation matters.
Start with at least these cases: an on-time successful import with records, a late success inside the grace window, an explicit failure, a zero-result success, a start with no finish, a duplicated finish, and an event delivered out of order. Then add the logistics exceptions that actually exist, such as a feed that does not run on a regional holiday.
A single global threshold looks tidy in a dashboard and performs poorly when a five-minute inventory delta and a two-hour carrier manifest share it. Put the expected cadence and allowed lateness next to each job's configuration. Keep alert policy separate from retention policy; deleting old debug records should not silently change whether today's scheduled run is evaluated.
The result count is also a candidate for trend detection, but avoid turning the first version into anomaly detection. A zero is crisp. A drop from 317 records to 41 may be suspicious, yet it needs a baseline, seasonality assumptions, and a review path. Page on conditions with an agreed response; route weaker signals to a non-paging review queue until the eval set shows they are reliable.
What should the centralized logging layer actually provide?
The application needs a small interface, not a permanent marriage to a backend. It must accept structured events, preserve timestamps, filter by stable fields, enforce retention, and let the evaluator retrieve the relevant time window. EU data location may be a deployment requirement, so verify the actual storage and processing locations in the service contract or self-hosted topology rather than inferring them from a marketing label.
For beginners seeking an alternative to Datadog, the better first comparison is not a long vendor matrix. Run the same fixture against every candidate. A hosted service, OpenSearch, and Grafana Loki can all be evaluated through that neutral test without assuming that one deployment model fits every startup. The application contract stays the same while the team checks its own EU region, access, retention, and operational requirements.
Compare candidates by running the same fixture through each path. Can a beginner find one run by job, scheduled_at, and run_id? Are delayed and duplicated events handled predictably? Can access be restricted, retention configured, and deletion verified? Does export preserve the structured fields? Those tests reveal more than a feature grid.
Price still belongs in the decision, just not as the detection design. Estimate daily structured event volume, retention duration, query frequency, and expected alert evaluations. Keep verbose framework logs on a shorter tier if they are useful for debugging, while retaining the compact lifecycle events for the operational window. Do not quote a headline entry price as though ingestion, storage, querying, transfer, and support form one stable number.
Field naming deserves discipline too. Prometheus recommends a naming scheme in which names have a single unit and should represent the same logical thing across label dimensions. Although these examples are log events rather than metrics, the underlying discipline transfers cleanly: use one meaning for duration_ms, do not sometimes put seconds into it, and do not overload result_count with bytes. If the team later derives metrics from logs, consistent semantics prevent avoidable cleanup.
Operate the rule, not just the log pipeline
Before deployment, replay the eval fixtures through the exact parser and evaluator that will run in production. Deploy the event contract first, observe that starts and finishes pair correctly, and only then enable paging. Keep the first alert destination attended during the rollout, because an alert with no clear owner is merely another stored event.
The operational checklist is brief in prose. Confirm clocks are UTC, run IDs are unique for an attempt, schedule changes update the monitor, and retry behavior cannot turn an old failure into a fresh page. Confirm malformed events go to a visible error counter rather than disappearing. Review a sample event for personal data, set retention intentionally, and test alert deduplication plus recovery. Finally, schedule a synthetic import that produces a known harmless result; it checks the scheduler, worker, event path, evaluator, and notification path together.
The final choice should be the simplest logging path that passes those tests in the required region. The durable investment is the event contract and evaluated decision rule. With those in place, centralized logs help answer why a logistics import stopped producing results without making every quiet interval look like an emergency.
Top comments (0)