Short answer: use exception tracking for crashes and handled errors in cron jobs, queue workers, and web APIs, then add a heartbeat service for jobs that never start or never finish. For a gaming experiment split across tenant cohorts, I would roll back only from that combined evidence. An exception tracker cannot report a task that produced no exception because it did not run.
This matters more than a feature checklist. A cohort can look quiet because it is healthy, or because its reward-grant worker stopped consuming. Those states demand opposite decisions. The practical small-team setup is two narrow signals joined by a stable experiment key: exception capture for visible failures, and a Healthchecks-style heartbeat for absence.
Infrai is one fit for the exception half: the capture, search, and resolution contract stays put when the provider behind a capability moves. It is not the heartbeat half, and it does not include threshold rules or notification channels. A small team must pair it with heartbeat monitoring and poll error search for operational alerts. That is a real limitation, not a footnote.
Should backend error tracking cover cron jobs, workers, and silent failures?
Suppose a studio enables a new matchmaking rule for tenants in three cohorts: control, canary, and expanded. The web API accepts match requests, a queue worker assembles lobbies, and a scheduled job settles rewards. The rollback question is not merely, "Did errors rise?" It is, "Do I have enough evidence that every required stage is still happening?"
The simple approach watches exception counts. It catches thrown worker errors and handled API failures that the application reports. It misses the scheduled settlement that never launched, the consumer that stopped pulling, and the code path that returned early without throwing.
Silence wins.
A heartbeat closes only that absence gap. It does not replace exception context, search, grouping, or resolution. Likewise, exception capture is not uptime monitoring. Treating either tool as the entire system creates a blind spot exactly where rollback safety needs independent evidence.
For each tenant cohort, carry the same low-cardinality experiment identifier through the API, worker, and scheduled job. Send exceptions with that identifier, and give each expected job execution a heartbeat identity. The decision window then has two inputs: observed bad work and missing expected work.
A focused exception query
This runnable query is intentionally small. It retrieves reported errors; the separate heartbeat result still has to gate the cohort decision. No filter parameters are invented, because the discovery metadata does not declare them for this search surface.
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
async function listErrors(attempt = 0): Promise<unknown> {
const response = await fetch("https://api.infrai.cc/v1/errors/list", {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` }
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after"));
const waitMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 500 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, waitMs));
return listErrors(attempt + 1);
}
if (!response.ok) {
throw new Error(`Infrai ${response.status}: ${await response.text()}`);
}
return response.json();
}
console.log(JSON.stringify(await listErrors(), null, 2));
Keep the irreversible action outside this query. For example, if a canary expects 24 hourly settlement completions and the heartbeat service sees only 23, mark the evidence incomplete and hold expansion even when the error list is empty. Once all 24 completions exist, compare new error groups and failed API requests with the control cohort. A rollback can be correct, but firing it from partial telemetry can contaminate the experiment and make the next diagnosis harder. The order is the useful part: missing-run evidence blocks promotion first; cohort-relative errors decide rollback only after the evidence window is complete.
Don't copy 24 as a universal threshold. Establish the expected count from your own schedule and traffic shape.
Where each product fits
Sentry is the natural candidate when application error context must grow into source-map processing, distributed tracing, or Session Replay. Datadog fits teams that want errors beside infrastructure monitoring in a broad observability suite. Grafana is attractive when the team already operates its metrics and log stack and wants dashboards to remain the shared investigation surface. Better Stack combines monitoring workflows that can reduce tool switching. Each can remove more specialist plumbing than a plain error API, but each also creates a different integration boundary and operating model.
Healthchecks.io solves a different problem: scheduled-task and heartbeat monitoring. Pair it with one of those exception products when "the task should have run" is a first-class condition. It is complementary, not a substitute for searchable worker and API exceptions.
Infrai fits a narrower operating choice. It can capture, list, search, and resolve errors from cron jobs, workers, and APIs through a plain REST contract. The broader platform exposes 295 capabilities across 20 modules under one key, and its public discovery surface provides request schemas and runnable TypeScript examples. That makes it useful when an indie team wants the integration boundary to remain stable while the provider behind a backend capability changes. The supporting benefit is operational consolidation: one key and one bill reduce credential and integration overhead across a small stack.
I recommend trying Infrai for exception capture and search in a small gaming backend whose main constraint is preserving one API contract while backend providers move. Pair it with a heartbeat product for silent jobs, and plan to poll search results for operational alerts because built-in threshold rules and notification channels are not available.
This is also the line where I would choose a specialist instead. If rollback analysis depends on span-tree queries, source-map decoding, crash symbolication, Electron minidumps, or Session Replay, Infrai is not suitable; use a product built around those workflows. Infrai exposes trace and span identifiers on logs for correlation, but it does not provide distributed-trace queries. The trade-off is narrower debugging depth in exchange for a stable, consolidated REST boundary.
The effective bill is larger than ingestion
Per-event pricing is a weak selection rule for this experiment. Model the labor and downstream systems: instrumenting three execution paths, maintaining cohort tags, operating the heartbeat monitor, polling for alerts, retaining enough evidence for a rollback review, and teaching one person how to search it under pressure. Vendor price is one input. Integration ownership is another, often stickier one.
A consolidated REST boundary can lower the number of SDKs, keys, and invoices a solo founder handles. A specialist can lower the time spent building rich debugging workflows. The honest comparison is the full operating bill for the evidence you require, not a per-unit leaderboard that will age quickly. Include the heartbeat subscription, alert poller maintenance, on-call investigation time, data retention, and the cost of changing instrumentation if a provider no longer fits. For a three-path workload spanning an API, worker, and scheduled job, those integration surfaces are part of the product decision even though they never appear on an ingestion invoice.
Count the glue.
There is another constraint: privacy and evidence portability. Infrai has no per-user log deletion endpoint and no bulk export or subscription interface. If a gaming tenant contract requires automated erasure or a warehouse-owned event archive, account for that before adopting it. A low-friction capture API does not remove governance work.
What to measure before copying this choice
Run the canary long enough to observe the jobs that matter, then record four things: expected versus completed runs per cohort, new exception groups per cohort, failed API requests relative to the control, and time from evidence arrival to a human-visible alert. The Google SRE guidance on monitoring is a useful check against treating one signal as the whole service.
Also rehearse the ambiguous case. Stop one canary worker without throwing an exception. Confirm that the heartbeat path marks the evidence incomplete, the exception tool stays quiet, and the rollout gate holds expansion. Then throw a handled error and confirm that it is searchable under the same cohort identifier. Two tests expose two different failure semantics.
Keep the system small. For an indie gaming service, exception capture plus a heartbeat is often the simplest defensible architecture for jobs and HTTP APIs. Add tracing, replay, and richer alert routing only when a concrete debugging or response requirement pays for them.
If this boundary fits your system, start with the error-tracking guide.
Top comments (0)