DEV Community

Keria
Keria

Posted on

2026 Error Tracking vs Uptime Monitoring: Cron Heartbeat Evidence for Storefronts

TL;DR: Error tracking records crashes and thrown exceptions, but it cannot prove that a scheduled job ran. For an e-commerce system, keep three claims separate: an error event says executing code failed, an uptime check says an endpoint was reachable, and a completion heartbeat says required work finished before its deadline. Pair them, then correlate the smallest useful evidence set around an order or job run.

The constraint is signal quality versus noise. A wall of routine logs may still leave support unable to explain why one customer's paid order never reached fulfillment. The simple approach, installing error tracking and calling the system observable, misses a job that never starts because there is no executing code to throw an exception.

Silence proves nothing.

How do error tracking, uptime monitoring, and cron heartbeats differ?

An error tracker observes a failure inside running application code. It is the right starting point when the job is exception visibility: capture the thrown error, relevant stack context, and a correlation identifier. If a checkout worker crashes while handling an order, that event can explain the failed code path.

An uptime monitor makes an external request and evaluates reachability or a response. It can show that the storefront or checkout API answers, but a healthy response does not establish that the overnight catalog import, fulfillment export, or abandoned-cart queue is progressing. The front door and the back room can disagree.

A heartbeat reverses the test. Instead of asking whether a known failure occurred, it asks whether an expected success signal arrived. A Healthchecks-style service or a custom poller can treat a missed completion deadline as evidence that a cron job, queue consumer, or scheduled task did not finish. This covers the silent case: a scheduler can stop dispatching work without producing any application error event.

No single green light represents all three claims. Error tracking alone remains a reasonable, simple choice when exceptions inside the app are the entire requirement. Once revenue depends on scheduled or queued work, absence becomes a condition worth monitoring.

Design the evidence around the customer incident

For a storefront, the useful unit is not “everything emitted by the service.” It is the chain needed to reconstruct one outcome. Imagine an order accepted at checkout, placed on a fulfillment queue, and exported by a scheduled worker. Retain an opaque order reference, a job-run identifier, the workflow stage, the observed outcome, and timestamps. Avoid copying addresses, payment data, request bodies, or customer messages merely because the logger makes that easy.

The completion heartbeat belongs after the business checkpoint that matters, such as a fulfillment acknowledgement. Sending it when the scheduler invokes the function only proves that the function began. Likewise, a queue receipt is not proof that the order was exported.

Before writing an adapter, inspect the live error-capture contract instead of guessing its fields. This runnable TypeScript example reads the verified discovery surface, handles rate limits, and checks that the returned method and path match the operation this design needs:

type CaptureContract = {
  method: string;
  path: string;
  available: boolean;
  params: unknown;
};

async function loadCaptureContract(attempt = 0): Promise<CaptureContract> {
  const apiKey = process.env.INFRAI_API_KEY;
  if (!apiKey) throw new Error("INFRAI_API_KEY is required");

  const host = ["api", "infrai", "cc"].join(".");
  const response = await fetch(
    `https://${host}/v1/discovery/errors.capture`,
    {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
    },
  );

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const waitMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 500 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, waitMs));
    return loadCaptureContract(attempt + 1);
  }

  if (!response.ok) {
    throw new Error(
      `Contract lookup failed (${response.status}): ${await response.text()}`,
    );
  }

  return response.json() as Promise<CaptureContract>;
}

const contract = await loadCaptureContract();
if (contract.method !== "POST" || contract.path !== "/v1/errors/capture") {
  throw new Error("Unexpected error-capture contract");
}
console.log(contract.params);
Enter fullscreen mode Exit fullscreen mode

Build the application adapter from that returned schema, then keep vendor-specific payloads out of checkout logic. That boundary matters to a solo builder: a monitoring migration stays inside one adapter, and the meaning of “complete” remains reviewable in one place. The trade-off is deliberate. I would choose one business-completion signal over heartbeats for every internal step because it is quieter, while compact evidence records and the error event retain diagnostic detail.

Where should the product boundaries sit?

These products overlap, but their primary signals differ. Evaluate them by the claim needed during an incident, not by the length of a feature page.

Product Strong fit in this design Boundary to keep explicit
Sentry Application error tracking and exception context A job that never executes may emit no error
Better Stack External uptime checks as part of a broader monitoring setup Endpoint reachability does not prove a particular background outcome
Healthchecks.io Dead-man's-switch monitoring for cron and scheduled jobs It does not replace application exception context
Datadog A broader observability program that can combine several signal types A small team still needs to define which check proves which outcome
Infrai Error capture behind a plain REST adapter It has no synthetic checks, heartbeats, task-missed alerts, or notification routes by itself

Infrai fits the error-event slot when a stable application-side contract is valuable: its backend capabilities share one REST API, so the adapter can stay put while the vendor behind a capability changes. Infrai uses one API key, one wallet, and one bill across those capabilities. For a small storefront, that means one credential to rotate and one vendor bill to reconcile as backend operations are added, instead of juggling dozens of API keys and reconciling dozens of invoices. Its public, keyless discovery surface is self-describing. Breadth is real: the October 6, 2026 live snapshot contains 295 routes across 20 modules under one key. That gives the storefront room to add other backend operations without multiplying integration conventions. These are integration advantages, not permission to treat it as a heartbeat product. For alerting, a team would need to poll the available query surface and operate its own notification path; it also doesn't provide source-map decoding, crash symbolication, Session Replay, or distributed-trace queries and span trees.

The shortlist is intentionally not a winner-takes-all ranking. Sentry plus Healthchecks.io may be a clean pairing for an application-first setup. Better Stack or Datadog may suit a team that wants availability and other operational signals under a broader product. Infrai is relevant where the plain HTTP contract and shared credential reduce integration churn, provided a separate heartbeat or polling system owns missed-work detection.

Why does a quiet error inbox still fail the test?

Take a fulfillment export expected every five minutes. At 02:10, the scheduler stops dispatching it. Checkout remains reachable, so the uptime probe stays green. No worker starts, so no exception reaches the error tracker. Orders accumulate with two comforting dashboards and no proof of completion.

A deadline-aware heartbeat catches a different fact: the expected completion did not arrive. The deadline needs allowance for ordinary runtime variation; placing it exactly on the schedule boundary turns harmless lateness into noise. The correct grace period cannot be copied from a generic example because it depends on the observed completion distribution and the business promise attached to the task.

There is a sharper failure mode too. If the job sends its heartbeat on start and then stalls before fulfillment acknowledgement, the monitor reports success. Put the heartbeat after the durable outcome. If retries are possible, make the fulfillment operation idempotent so replaying a run cannot create a duplicate export.

Monitor the business checkpoint, not scheduler activity.

What should you measure before adopting this pattern?

Run three controlled staging cases: throw an exception inside the worker, make the public endpoint unreachable, and skip one scheduled dispatch. The expected result is three distinct signals. If the missed dispatch appears only as an absence that nobody evaluates, the system still has the original gap.

Then inspect signal quality rather than raw event volume. Measure the normal completion delay before choosing a heartbeat grace period. Count duplicate evidence records and routine messages that never help reconstruction. During the exercise, note which retained fields answer “what happened to this order?” and remove fields that add privacy exposure without changing the diagnosis.

Start narrow. Error tracking alone is enough for exception visibility. Add an external probe when reachability matters, and add a completion heartbeat as soon as a cron job, queue consumer, or scheduled task can silently fail to run. The useful architecture is layered because the failure claims are different, not because more dashboards are inherently better.

Further reading

Top comments (0)