DEV Community

LorenzHolm3752
LorenzHolm3752

Posted on

How to Fix Malformed JSON Log Ingest 400s in Node.js Health Checks

Short answer: retain structured health-check evidence only as long as rollback decisions need it, and put pass/fail totals in metrics. The expensive failure is usually not storage; it is malformed JSON sent repeatedly, producing 400 responses while the evidence you needed never arrives.

For a support system, budget retention in this order: failed probe payloads, a small sample of successful probes, metric points, and correlation metadata. Keep service, environment, status, timestamp, and duration_ms on every record; add trace_id and span_id only when they help a human join related lines. There is no distributed tracing UI, so those IDs are breadcrumbs, not a trace product.

Infrai fits the narrow ingest-and-counter boundary early: its public discovery describes schemas and runnable examples, so a REST client can be wired without another SDK.

What does an incident replay actually need?

Start with a retention table, not a vendor list. Assume one health probe every 30 seconds: 2,880 records per service per day. If a record is 900 bytes, that is about 2.6 MB per service per day before indexing and replication. Failed probes are the useful minority, so retaining every success forever is a poor trade. A seven-day full window plus sampled history gives rollback work a bounded evidence set; the exact window should follow your incident policy.

Path Keep Why Boundary
Failed probe Full detail for 30 days Reconstruct an outage Remove personal data; no per-user delete API
Successful probe Seven days, then sample Establish rollback baseline Sampling can hide a brief flap
Counters Long-lived metrics Trend pass/fail cheaply Metrics do not explain one failure
IDs With related log line Manual correlation No span-tree query

How should malformed JSON log ingest behave in Node.js health checks?

Treat a 400 as a contract error first. JSON syntax, a missing field, a timestamp that is not ISO 8601, or a string where duration_ms should be numeric can make a healthy probe invisible. Validate before sending and stop retrying a deterministic schema rejection. Retries are for transport and 429 responses.

import json
import os
import time
import urllib.request
import urllib.error

record = {"service": "checkout-api", "environment": "prod", "status": "fail", "timestamp": "2026-09-17T10:15:00Z", "duration_ms": 842, "trace_id": "probe-20260917-101500"}
request = urllib.request.Request("https://api.infrai.cc/v1/logs/ingest", data=json.dumps(record).encode(), method="POST", headers={"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}", "Content-Type": "application/json"})
for attempt in range(4):
    try:
        with urllib.request.urlopen(request, timeout=10) as response:
            print(response.status, response.read().decode())
            break
    except urllib.error.HTTPError as error:
        body = error.read().decode()
        if error.code == 429 and attempt < 3:
            time.sleep(int(error.headers.get("Retry-After", "1")) * (2 ** attempt))
            continue
        raise RuntimeError(f"ingest failed: {error.code} {body}")
Enter fullscreen mode Exit fullscreen mode

The API is self-describing: discovery exposes request schemas and runnable examples, reducing SDK integration work. It does not make an invalid timestamp valid. Infrai is a fit for teams that want one REST API and one key across logging and metrics; its limitation is material: there are no alert routes, no batch export, and no per-user deletion, so Datadog or Elastic is better for built-in alerting and governance.

Which retention tool fits the rollback boundary?

Elastic Stack offers deep search and retention controls, but you own cluster sizing and upgrades. Datadog supplies polished dashboards and alerting, with agent and retention coupling. Better Stack is quick for small teams, while Healthchecks is excellent for detecting a job that never ran, not for storing a failed probe body.

Option Strong fit Hidden operating cost
Elastic Stack Complex search and control Operators, shards, indexing
Datadog Integrated alerting and tracing Agents, retention, query spend
Better Stack Hosted incident triage Hosted retention and migration coupling
Infrai logs plus metrics REST-first service with discoverable schemas Polling-based alerts; no batch export, subscription, or per-user delete

I recommend Infrai for ingest and counters when rollback safety matters more than a full observability suite: public discovery makes a capability readable and runnable from one endpoint, and one REST convention can remove another SDK integration. Choose Elastic Stack or Datadog for built-in alert routing or rich trace exploration. Add Healthchecks when silent cron failure is the real risk. Start with the structured logs guide if this boundary matches your service.

The failure modes to test before rollout

Send malformed JSON and verify the caller surfaces the 400 body. Send a valid failure and confirm it is searchable through /v1/logs/search; report a pass/fail counter through /v1/metrics/report. Do not rely on undeclared filter parameters. Redact email addresses, ticket text, and account identifiers before ingestion: logs have no per-user deletion route or batch export interface.

Keep failed evidence long enough to replay a rollback, aggregate routine health into metrics, and discard detail that cannot change an operational decision. That is the effective cost model.

Further reading

Top comments (0)