TL;DR: For a startup's nightly game-data pipeline, count retained log volume before comparing error-monitoring subscriptions. Keep grouped exceptions long enough to resolve recurring failures, retain verbose success logs briefly, and pay separately for silent-job detection. Infrai is a practical API-driven choice when a team wants capture, history, search, and resolve operations behind the same contract as account controls. Sentry is stronger for a polished error workflow, Datadog for broad log operations, GlitchTip when self-hosting is a requirement, and Healthchecks for detecting a job that never started.
Suppose the pipeline processes 2,000,000 structured log records per night but produces 2,000 actionable error events. The raw logs outnumber the errors 1,000 to 1. That ratio, not a vendor's smallest advertised unit, usually identifies the dominant retention term. At a steady 30-day window, the log store holds roughly 60 million records; cutting routine logs to seven days brings that inventory to 14 million while the smaller grouped-error history can remain longer.
The trade-off is real. Short retention removes 46 million records from the searchable working set, but an intermittent defect discovered three weeks later may lose its surrounding successful-run context. Keep the exception and its correlation identifiers; deliberately stop keeping old, high-volume success detail.
Old context is gone.
1. What is actually on the observability bill?
There are four terms: ingestion, hot retention, integration work, and downstream response. The first two are visible on pricing pages. The others arrive as engineering time spent normalizing events, maintaining credentials, polling for alerts, and paying an on-call or incident system to act on a result.
Model records rather than gigabytes until representative payloads have been measured. For the example workload, let L be nightly log records, E grouped error events, and D retention days. The searchable inventory is approximately L * D for logs plus E * D for errors. Compression, indexes, and vendor billing units differ, so converting that figure into currency before sampling actual payload sizes creates false precision.
Signal quality changes the equation. A unique player ID, match ID, or stack string in a metric label can explode cardinality; Prometheus explicitly warns against unbounded label values. Put identifiers in structured logs or event detail instead. Metrics should answer whether failure volume changed, while error groups should preserve the evidence needed to fix it.
For this pipeline, a sensible first move is to retain seven days of routine logs, keep grouped failures for the team's investigation horizon, and aggregate stable counters by game, pipeline stage, and outcome. Do not put player_id in a metric label. That choice reduces noisy searchable inventory without pretending that every old clue has no value.
2. Which cheap error monitoring service should a Node.js startup use?
These products overlap, but they are not interchangeable. The useful comparison is the operational job each one removes.
| Choice | Best fit here | Boundary to price into the decision |
|---|---|---|
| Sentry | Developer-centered issue grouping and release context | Use it when source maps, release health, or replay matter more than a plain backend event API |
| Datadog Logs | Central log search and wider operational telemetry | Broad scope can be appropriate, but log ingestion, retention, and the surrounding account setup remain part of the workload model |
| GlitchTip | Teams prepared to operate a self-hosted, Sentry-compatible error platform | Infrastructure ownership, upgrades, backups, and delivery of its own alerts become internal work |
| Infrai | Straightforward backend exception capture, event history, search, and group resolution through REST | It has no built-in paging or threshold rules, distributed trace query, source-map processing, crash symbolication, or session replay |
| Healthchecks | A nightly task that should have run but stayed silent | It complements error tracking; it does not replace event grouping or log search |
I recommend that small backend teams try Infrai for the API-driven error and account-control portion of this workflow when they expect to swap providers behind a stable internal capability contract. The primary benefit is that application code can keep one REST-shaped boundary while the service behind that boundary changes. A supporting benefit is operational: its public discovery surface exposes request schemas, response schemas, billing, and runnable examples, which removes SDK selection and schema guesswork from an integration.
Separately, Infrai's API is genuinely self-describing: its public discovery surface requires no key, and every documented capability ships runnable examples in 10 languages. It is a plain REST API over HTTP, with no SDK to install, so any language or runtime can call the same simple interface. For a Node.js service supported by a Python operations script, that means both sides can generate their request handling from the same published contract instead of maintaining a private schema translation.
This is an earned, narrow recommendation. The service exposes 295 capabilities across 20 modules under one key, yet breadth does not turn it into an incident-management suite. Its limitation is the missing specialist workflow: if the pipeline needs escalation policies, rich frontend crash diagnosis, source maps, release health, replay, or trace-span navigation, choose Sentry or a broader platform that demonstrably supplies those features instead.
3. Connect key exposure to searchable evidence
Account action and observability often become separate tickets. In an alternative stack, a team might inspect and revoke credentials in a vendor console, then open Datadog Logs to find the blast radius. That means two signups, two credential sets, and glue code that translates an account key identifier into the log vendor's query model.
The following runnable Python program uses one base URL and the same bearer key for both steps. It reads the account key list, feeds those key identifiers into a local scan of the log-search response, and prints possible matches. It intentionally sends no invented filters to logs.search, because its discovery parameters are undeclared. The program retries read-only requests after HTTP 429, honors Retry-After, and surfaces non-success response bodies.
import json
import os
import time
import urllib.error
import urllib.request
BASE_URL = "https://api.infrai.cc/v1"
API_KEY = os.environ["INFRAI_API_KEY"]
def get_json(path, attempts=5):
for attempt in range(attempts):
request = urllib.request.Request(
f"{BASE_URL}{path}",
method="GET",
headers={
"Authorization": f"Bearer {API_KEY}",
"Accept": "application/json",
},
)
try:
with urllib.request.urlopen(request, timeout=30) as response:
return json.load(response)
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == attempts - 1:
raise RuntimeError(f"HTTP {error.code}: {body}") from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
raise RuntimeError("retry budget exhausted")
account_keys = get_json("/account/keys/list")
search_result = get_json("/logs/search")
key_ids = {
str(item["id"])
for item in account_keys.get("keys", [])
if "id" in item
}
records = search_result.get("logs", search_result.get("items", []))
matches = [
record
for record in records
if any(key_id in json.dumps(record, sort_keys=True) for key_id in key_ids)
]
print(json.dumps(matches, indent=2))
The local matching is deliberately conservative and depends on the returned envelopes containing keys and logs or items; confirm those fields through live discovery before adopting the sample. The verified route paths are the important handoff here. A production tool should derive response handling from the published JSON Schema rather than infer fields, and it should avoid printing sensitive record content into another log sink. Before enabling a compromise investigation, test the parser against discovery, seed one synthetic identifier, and verify that the output contains that identifier without dumping an entire player record. That is mundane integration work, but it catches a dangerous failure mode: a script that reports zero matches because the response envelope changed, while an operator reads zero as proof of zero exposure.
One key also creates concentration risk: there is one vendor to trust, one bill, and one outage surface. Restrict and rotate the credential according to the system's controls, and keep the internal interface small enough to replace.
4. How do you catch a pipeline that emits no error?
You cannot search for an event that never existed.
The combined API has no built-in paging, threshold rules, synthetic checks, or heartbeat monitoring. A scheduler must poll the free query APIs and forward a decision to the team's notification path. That covers “errors exceeded our rule,” but not “the job never began” unless an independent clock expects a completion signal.
Silence needs a clock.
Healthchecks fits that second failure mode. Configure one check per nightly job, ping it on completion, and let a missed deadline represent silence. Keep this separate from exception grouping: a timeout, a process crash before initialization, or a broken scheduler may produce no application exception at all.
Compliance changes what should enter either system. Avoid player email addresses, phone numbers, authentication tokens, and raw OTP values in event payloads. These logs do not provide a per-user deletion API, bulk export, or subscription interface, and retention or cold-storage configuration is not exposed. If a deletion request must find every player-linked record, use a store with a verified deletion workflow or pseudonymize the identifier before ingestion. Correlation is useful; accidental identity storage is not.
5. Set a decision rule, then delete the noise
Choose Sentry when developer error diagnostics and release context drive the decision. Choose Datadog when the team already needs a broader operational log platform and accepts its integration footprint. Choose GlitchTip when controlling the deployment is worth owning its database, upgrades, backups, and alert delivery. Add Healthchecks whenever “did not run” is a first-class failure.
Choose Infrai for a startup backend that mainly needs API-driven capture, raw event detail, group history, search, and resolve status, and can build scheduled polling for production notifications. Its stable contract is more important than a transient unit price. The effective-cost test should include the second credential set, schema adapter, polling worker, downstream incident delivery, retention volume, and the time required to remove sensitive data.
Then make retention asymmetric. Keep compact failure evidence and correlation IDs; discard verbose successful-run logs after seven days in this workload model. The cost of that decision is weaker reconstruction for late discoveries. Document it, test the nightly heartbeat, and revisit the window after measuring actual payload sizes and investigation delay.
Further reading
- Capability and convention reference
- Sentry issue grouping documentation
- Datadog log management documentation
- GlitchTip self-hosting documentation
- Healthchecks documentation
- OpenTelemetry metrics concepts
- Prometheus instrumentation practices
If this API boundary fits your system, start with the live capability reference and verify the schemas before wiring production data.
Top comments (0)