An error-log poller can alert on a failed import, but it cannot tell you that an import produced nothing when no error was emitted. For scheduled media imports, record each run's start, completion, and result count; poll those records alongside errors, then send one webhook notification per missed deadline or failed run. Keep the run ID and region in the alert so the incident can be reconstructed later.
TL;DR: Treat silence as a missing expected result, not as proof that the backend is healthy. A scheduled import that never started, one that crashed before logging an error, and one that completed with zero items need different diagnoses. A poller can make those distinctions only if the import writes explicit lifecycle records.
How can Node.js detect backend failures from error logs?
Imagine an import scheduled for 02:00 in the US region. The scheduler fails before it calls the worker. No worker exception exists to query. A search for level=error returns an empty page, which looks exactly like a clean run to a naive watcher. That's the trap.
The useful unit of observation is the scheduled run, not the individual exception. Assign an ID derived from the job, region, and scheduled slot; record started and completed events under that ID. Completion should include the count of accepted results. An explicit zero is different from no completion record: zero might be a legitimate empty feed, while no completion means the run is unresolved. Define the expected deadline and the policy for zero results with the people who own the feed. Neither can be inferred from a generic error message. For example, if the worker records started and then dies, the missing completion points toward execution; if neither event exists, start with the scheduler or the event pipeline. If completion exists with zero accepted items, investigate the upstream feed and validation policy. Those are three different paths through an incident, even though a generic error search could return the same empty result for all three.
Empty logs prove little.
This shifts the decision axis from alert volume to incident reconstruction. An alert saying "five errors" gives little context. An alert identifying the missing scheduled slot, last observed stage, region, and query window gives an on-call engineer a starting point. Keep the source timestamps as well as the time the poller saw each record; ingestion can lag, and sorting only by arrival time can disguise the order of a failure.
The constraint that changes the design
Polling a logs API looks like a small integration: fetch errors, POST to a webhook, sleep. The schedule makes it a state problem. Each poll needs an overlapping lookback window to tolerate delayed ingestion, yet overlapping windows replay records. If the watcher restarts, an in-memory "already alerted" set vanishes. Store the last observed state and an alert key in durable storage, or accept duplicate notifications as an explicit trade-off. A deployment with two watcher replicas adds another race: both might read the same unfinished slot before either records a claim. The claim therefore needs atomicity across replicas, not merely persistence. Polling frequency, grace interval, and log retention are separate settings; coupling them into one global timeout makes a slow regional feed look broken while a truly stopped fast feed waits too long.
Keep those knobs few.
For a US/EU deployment, keep region in every record and query the appropriate region's log store. Do not merge records by timestamp alone; a run ID must be scoped to its region. Log retention, access controls, and deletion rules matter too: records containing personal data require a deletion path, including derived alert payloads where applicable. GDPR Article 17 describes the right to erasure and its exceptions; it does not grant a blanket exemption to operational logs. Avoid copying article titles, author details, or raw exception payloads into the webhook when a run ID and state suffice.
I would benchmark detection lag before debating polling frequency: measure scheduled deadline to log availability, then log availability to alert delivery, separately. There are no measured timings for this example. Pick a grace interval from your own ingestion-delay distribution and feed behavior, and test the choice against a delayed but successful run. Faster polling cannot repair missing lifecycle events.
A minimal watcher for the missing result
This TypeScript sketch assumes a logs endpoint returning structured import lifecycle events and a durable key-value store supplied by the host application. The URLs and interface are placeholders, not claims about a particular service. A production worker should persist its started and completed events before relying on this watcher. The scheduled slots come from your scheduler's source of truth, rather than being guessed from whatever happens to appear in logs.
type ImportEvent = {
runId: string;
region: "us" | "eu";
stage: "started" | "completed" | "failed";
at: string;
resultCount?: number;
};
type Slot = {
runId: string;
region: "us" | "eu";
deadline: string;
};
interface AlertState {
// Must atomically claim a key across watcher restarts and replicas.
claim(key: string): Promise<boolean>;
}
async function checkSlot(slot: Slot, state: AlertState, webhookUrl: string) {
if (Date.now() < Date.parse(slot.deadline)) return;
const url = new URL("/events", "https://logs.example.invalid");
url.searchParams.set("region", slot.region);
url.searchParams.set("runId", slot.runId);
const response = await fetch(url);
if (!response.ok) throw new Error(`Log query failed: ${response.status}`);
const events = (await response.json()) as ImportEvent[];
const run = events.filter(e => e.runId === slot.runId && e.region === slot.region);
const completed = run.find(e => e.stage === "completed");
const failed = run.find(e => e.stage === "failed");
const status = failed ? "failed" : completed
? completed.resultCount === 0 ? "empty" : "complete"
: run.some(e => e.stage === "started") ? "unfinished" : "not started";
if (status === "complete") return;
const key = `${slot.region}:${slot.runId}:${status}`;
if (!(await state.claim(key))) return;
const alert = await fetch(webhookUrl, {
method: "POST",
headers: { "content-type": "application/json" },
body: JSON.stringify({ text: `Import ${status}: ${slot.region} ${slot.runId}; deadline ${slot.deadline}` }),
});
if (!alert.ok) throw new Error(`Alert delivery failed: ${alert.status}`);
}
There is a delivery gap between claim and the webhook POST: a crash there loses an alert. For an operational implementation, write an outbox item transactionally with the claim, retry delivery with backoff, and mark it sent only after a successful response. Webhook receivers may still see duplicates after an ambiguous timeout, so include the same alert key in the message and make repeated notifications recognizable. Never log the webhook URL; treat it as a secret.
The query above also assumes the API returns all records for one run. If results are paginated, follow every page or the watcher may misclassify a completed run as unfinished. Validate the response schema, bound the query window, and handle API throttling and transient failures without converting "could not check" into "import failed." A broken watcher deserves its own health signal. Check that a response's region matches the request even if the server claims to filter it, and reject malformed timestamps rather than silently sorting them. A truncated page can be especially deceptive: a started event may arrive early while the completed event sits on the next page, creating a false incident precisely when the importer is healthy.
What changes at scale?
First test the states, not the happy-path fetch. Use synthetic slots for a run that never starts, fails, starts without completion, completes with zero results, and completes after the grace window. Replay the same log page twice and restart the watcher between polls. If that produces a second incident notification, the deduplication boundary is wrong. A staging test should also delay event ingestion on purpose; this exposes whether your grace period reflects reality.
Then separate detection from delivery. A single process and a small persistent store are sufficient for a modest number of scheduled feeds. More jobs or longer log delays call for batched queries and explicit per-region capacity limits; otherwise the monitor's query load can obscure the very incident it is meant to detect. Track query failures, oldest unchecked slot, and alert delivery failures. Benchmark time-to-first-use too: count the lifecycle events and configuration fields a new job must add, not just the poller's request latency. Config bloat is an incident risk when each feed has a slightly different undocumented deadline.
Error grouping can help triage repeated exceptions, but a grouping fingerprint is not a substitute for a scheduled run ID: several failures can belong to one run, and one silent run can have no errors at all. The decision rule stays simple: page for a missing required result after its agreed deadline, retain enough structured context to reconstruct the run, and keep failure of the monitoring path visible separately.
References
- Sentry, event grouping and fingerprint mechanics: https://docs.sentry.io/concepts/data-management/event-grouping/
- GDPR, Article 17, right to erasure: https://gdpr-info.eu/art-17-gdpr/
Top comments (0)