DEV Community

jaxmonroe3187
jaxmonroe3187

Posted on

Silent Cron Runs: A Backend Metrics Dashboard Beyond API Failures

An agent dashboard has one awkward constraint: the absence of a failure event does not mean the scheduled loop ran. TL;DR: use metrics for latency, cost, completions, failures, and backlog trends, then pair them with a separate heartbeat monitor for missed cron runs. Keep error details beside the charts, not inside every metric label. This gives a solo operator useful signal without pretending to provide complete monitoring.

The tempting first design is one giant event stream. Every tool call, exception, token count, and scheduler tick goes into it, then the dashboard tries to reconstruct reality. That approach is simple to ship but noisy to operate. High-cardinality error messages fragment charts, while a job that never starts emits nothing at all.

My first-pass design would instead define a small operational contract: each completed loop reports duration and cost; each terminal outcome increments one success or failure counter; a queue reports backlog size; and the scheduler pings a heartbeat service independently. Four signal families are enough for the first useful screen. The trade-off is deliberate: this screen loses event-level detail, but an operator can scan it before opening the error record that preserves that detail.

How should a backend metrics dashboard cover cron jobs and API failures?

Start with decisions, not available fields. If the question is "Did the agent become slower?", chart a duration distribution or stable time buckets. If it is "Are we spending more per successful outcome?", aggregate cost beside the success count, then compute the ratio in the dashboard. A raw total cost line can rise because traffic rose, which is healthy, or because each loop became more expensive, which is not. Those cases need separate widgets.

For an agent loop, I would keep these dashboard signals:

  • completed and failed loops per interval;
  • duration for the whole loop, plus a deliberately small set of stage names;
  • cost per completed loop and total cost per interval;
  • queue backlog at sampling time;
  • grouped error counts, with links from a spike to the relevant error view;
  • heartbeat state for every scheduled producer.

Do not put prompt text, user IDs, exception messages, or run IDs into metric dimensions. They create noise and can turn a bounded chart into an unbounded set of series. Preserve those details in errors or logs and correlate them with a trace ID or span ID where available. Correlation fields are useful, but they are not a distributed trace query or a span tree.

One ratio deserves special care. A cost-per-success widget should divide interval cost by successful outcomes only when that interval has at least one success. Rendering zero during an empty interval lies; render no value instead. Business events need the same restraint: promote a stable event such as an agent task completing into a count, but keep arbitrary event properties out of dimensions until a concrete dashboard decision needs them.

Quiet is ambiguous.

A focused TypeScript aggregation

The following runnable TypeScript retrieves error records for the failure-detail side of the dashboard. It uses one verified route, checks every response, and backs off on rate limits. Set INFRAI_BASE_URL to the documented API base and keep the key outside source control. The response stays unknown because inventing a convenient vendor response type would make the sample look safer than it is.

const apiKey = process.env.INFRAI_API_KEY;
const baseUrl = process.env.INFRAI_BASE_URL;

if (!apiKey || !baseUrl) {
  throw new Error("Set INFRAI_API_KEY and INFRAI_BASE_URL");
}

async function listErrors(attempt = 0): Promise<unknown> {
  const response = await fetch(`${baseUrl}/errors/list`, {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 500 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return listErrors(attempt + 1);
  }

  if (!response.ok) {
    const body = await response.text();
    throw new Error(`Error API returned ${response.status}: ${body}`);
  }

  return response.json() as Promise<unknown>;
}

listErrors().then((records) => console.log(JSON.stringify(records, null, 2)));
Enter fullscreen mode Exit fullscreen mode

Keep timeseries work on the metrics side and use this error response for investigation, rather than turning every error message into a metric dimension. The example intentionally does not add query filters: the discovery parameters for metrics queries and log searches are undeclared, so guessed filters would be brittle. For a real latency chart, prefer a percentile over an average once you have enough samples; a few slow loops can disappear inside a friendly-looking mean.

Metrics reporting should also be off the critical path. Buffer a small batch, flush after the user-visible result, and cap the buffer so observability cannot consume unlimited memory. Lost telemetry is undesirable. A delayed agent response is worse.

Why can't metrics detect a missed cron run?

A metrics API can show success counts, failure counts, durations, backlog, and error-rate trends. It cannot report an event that was never emitted. If the host is down, the scheduler is paused, or deployment configuration stops invoking the task, both the success and failure series remain quiet. A dashboard may interpret that silence as a calm night.

Use a Healthchecks-style dead man's switch. The scheduler pings a distinct heartbeat on success, and the heartbeat service owns the expected cadence and grace period. For longer work, separate "started" from "completed" only if the chosen heartbeat product supports that lifecycle clearly. The core rule stays simple: metrics describe work that happened; heartbeats detect work that should have happened but did not.

This split also keeps alerts honest. Infrai can report and query metrics through a plain REST API with no client SDK to install, supply failure counts through its error APIs, and use one key for 295 routes across 20 modules with one bill. That reduces credential and billing work when the same small backend later needs another capability. The API is genuinely self-describing, and its public discovery surface requires no key while exposing full request and response schemas, billing details, and runnable examples. Every documented capability ships runnable examples in 10 languages. It does not provide heartbeat checks or notification routes, so polling-based thresholds remain your responsibility and silent cron failures need a separate heartbeat tool.

Comparing the realistic tool choices

The right comparison is about signal ownership, not the longest feature list.

Option Best fit in this design Boundary to account for
Healthchecks Detecting a scheduled job that misses its expected ping Complements rather than replaces latency, cost, backlog, and error charts
Grafana Building dashboards across metrics sources with flexible visualization Collection, storage, and heartbeat semantics still need to be chosen
Sentry Investigating grouped application errors and rich failure context Error investigation is different from proving that a cron job ran
Datadog A broader hosted observability stack when one platform should cover many operational signals Scope and operating model may be more than a small SaaS dashboard needs
Infrai Small backends that want metrics and error access through one REST surface No heartbeat monitoring, notification route, distributed trace query, span tree, source-map decoding, minidump symbolication, or session replay

The last row is a narrow recommendation, not a universal one. Its limitations are material: pick it when REST simplicity and a compact integration surface outweigh the missing monitoring layers. It is not suitable as the sole system when on-call alert delivery, distributed tracing, or crash-symbol processing is part of the requirement; pick a broader suite then. Electron teams in particular should treat native crash minidumps as a separate pipeline because collecting metric failures cannot decode those dumps.

Grafana plus a metrics store gives the most control over dashboard composition. Sentry gives deeper error investigation. Datadog consolidates more of the operational stack. Healthchecks solves the silence problem directly. Combining specialized tools means more credentials and integrations, but it also lets each signal retain clear semantics.

Measure this before copying the design

Run the dashboard for a representative workload before committing to its labels and thresholds. Measure the count of unique series, telemetry delay, missing-cost records, the gap between scheduler pings, and the fraction of failed loops that can be opened in an error view. Also compare mean latency with p95; if they tell different stories, the percentile belongs on the main screen.

Then perform two controlled tests. Cause one loop to return a terminal failure and verify that the failure counter and error detail both change. Skip one scheduled invocation entirely and verify that only the heartbeat system complains. That second test is the proof that the architecture covers silence instead of merely drawing a flat line. I would reject the design if both tests produced the same signal: an emitted failure and an absent run demand different operator actions, even when both eventually appear red.

Test the silence.

Do not call this full observability. Without notification delivery, distributed trace queries, replay, symbolication, and synthetic checks, it is a focused operations dashboard. For a solo founder, that can be the right boundary: enough information to decide whether to inspect errors, drain backlog, reduce agent steps, or investigate a scheduler, without paying an attention tax for every raw event.

Sources

Top comments (0)