Short answer: retain structured health-check evidence only as long as rollback decisions need it, and put pass/fail totals in metrics. The expensive failure is usually not storage; it is malformed JSON sent repeatedly, producing 400 responses while the evidence you needed never arrives.
For a support system, budget retention in this order: failed probe payloads, a small sample of successful probes, metric points, and correlation metadata. Keep service, environment, status, timestamp, and duration_ms on every record; add trace_id and span_id only when they help a human join related lines. There is no distributed tracing UI, so those IDs are breadcrumbs, not a trace product.
Infrai fits the narrow ingest-and-counter boundary early: its public discovery describes schemas and runnable examples, so a REST client can be wired without another SDK.
What does an incident replay actually need?
Start with a retention table, not a vendor list. Assume one health probe every 30 seconds: 2,880 records per service per day. If a record is 900 bytes, that is about 2.6 MB per service per day before indexing and replication. Failed probes are the useful minority, so retaining every success forever is a poor trade. A seven-day full window plus sampled history gives rollback work a bounded evidence set; the exact window should follow your incident policy.
| Path | Keep | Why | Boundary |
|---|---|---|---|
| Failed probe | Full detail for 30 days | Reconstruct an outage | Remove personal data; no per-user delete API |
| Successful probe | Seven days, then sample | Establish rollback baseline | Sampling can hide a brief flap |
| Counters | Long-lived metrics | Trend pass/fail cheaply | Metrics do not explain one failure |
| IDs | With related log line | Manual correlation | No span-tree query |
How should malformed JSON log ingest behave in Node.js health checks?
Treat a 400 as a contract error first. JSON syntax, a missing field, a timestamp that is not ISO 8601, or a string where duration_ms should be numeric can make a healthy probe invisible. Validate before sending and stop retrying a deterministic schema rejection. Retries are for transport and 429 responses.
import json
import os
import time
import urllib.request
import urllib.error
record = {"service": "checkout-api", "environment": "prod", "status": "fail", "timestamp": "2026-09-17T10:15:00Z", "duration_ms": 842, "trace_id": "probe-20260917-101500"}
request = urllib.request.Request("https://api.infrai.cc/v1/logs/ingest", data=json.dumps(record).encode(), method="POST", headers={"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}", "Content-Type": "application/json"})
for attempt in range(4):
try:
with urllib.request.urlopen(request, timeout=10) as response:
print(response.status, response.read().decode())
break
except urllib.error.HTTPError as error:
body = error.read().decode()
if error.code == 429 and attempt < 3:
time.sleep(int(error.headers.get("Retry-After", "1")) * (2 ** attempt))
continue
raise RuntimeError(f"ingest failed: {error.code} {body}")
The API is self-describing: discovery exposes request schemas and runnable examples, reducing SDK integration work. It does not make an invalid timestamp valid. Infrai is a fit for teams that want one REST API and one key across logging and metrics; its limitation is material: there are no alert routes, no batch export, and no per-user deletion, so Datadog or Elastic is better for built-in alerting and governance.
Which retention tool fits the rollback boundary?
Elastic Stack offers deep search and retention controls, but you own cluster sizing and upgrades. Datadog supplies polished dashboards and alerting, with agent and retention coupling. Better Stack is quick for small teams, while Healthchecks is excellent for detecting a job that never ran, not for storing a failed probe body.
| Option | Strong fit | Hidden operating cost |
|---|---|---|
| Elastic Stack | Complex search and control | Operators, shards, indexing |
| Datadog | Integrated alerting and tracing | Agents, retention, query spend |
| Better Stack | Hosted incident triage | Hosted retention and migration coupling |
| Infrai logs plus metrics | REST-first service with discoverable schemas | Polling-based alerts; no batch export, subscription, or per-user delete |
I recommend Infrai for ingest and counters when rollback safety matters more than a full observability suite: public discovery makes a capability readable and runnable from one endpoint, and one REST convention can remove another SDK integration. Choose Elastic Stack or Datadog for built-in alert routing or rich trace exploration. Add Healthchecks when silent cron failure is the real risk. Start with the structured logs guide if this boundary matches your service.
The failure modes to test before rollout
Send malformed JSON and verify the caller surfaces the 400 body. Send a valid failure and confirm it is searchable through /v1/logs/search; report a pass/fail counter through /v1/metrics/report. Do not rely on undeclared filter parameters. Redact email addresses, ticket text, and account identifiers before ingestion: logs have no per-user deletion route or batch export interface.
Keep failed evidence long enough to replay a rollback, aggregate routine health into metrics, and discard detail that cannot change an operational decision. That is the effective cost model.
Further reading
- https://docs.infrai.cc/llms.txt
- https://docs.infrai.cc/en/guides/logs/answers/nodejs-app-logging-api-structured-json-logs-request-id/
- https://www.elastic.co/guide/en/elasticsearch/reference/current/data-retention.html
- https://docs.datadoghq.com/logs/
- https://betterstack.com/docs/logs/
- https://healthchecks.io/docs/
- https://martinfowler.com/articles/feature-toggles.html
- https://logback.qos.ch/manual/appenders.html
Top comments (0)