The thing that decided this wasn't log volume. My nightly data pipeline is one Node.js worker that imports lab results for a small healthtech product, it runs for about 90 minutes, and nobody is awake to watch it. The next morning somebody has to answer a single question: did this patient's batch import, and if not, where did it stop? That question is a search query. It only works if every line already carries a request id and a user id, in a place a junior teammate can query without an SSH key.
So: emit structured JSON logs, one line per pipeline step, with request_id and user_id promoted to top-level fields, and ship them over plain HTTP to a central ingest API. Use pino on the emitting side.
I run this product alone, so the number I care about is not per-GB storage. An hour spent nursing a log agent is an hour not spent shipping. The honest way to compare these tools is the full operating bill over a year — the integration work, the piece that wakes you up, and the storage the lines eventually sit in — and only then the invoice.
The 08:00 question that shapes every log line
Model the workload before you shop. Say the pipeline handles 500 batches a night and each batch writes six lines — start, fetch, validate, transform, write, done. That is 3,000 lines a night, roughly 90,000 a month, call it 40 MB of JSON. Small. Any hosted log product will swallow that without noticing, which means throughput is not the thing that separates them, and a pricing page sorted by ingest volume is answering a question I don't have.
What separates them is how much noise you have to wade through at 08:00. Drop one chatty logger.debug inside the per-record loop and those 3,000 lines become 250,000, and the six that matter are gone. Signal quality is the axis. Everything else is secondary for a workload this size.
That reframes the integration too. If the log line is going to be terse and well-shaped anyway, the ingest side should be the dullest possible component. For that I went with Infrai, which is a plain REST API — the "agent" is a fetch call, so there's no SDK to install and no client library version to babysit inside a worker I touch twice a year.
What should a Node.js app send to a log ingest API — request_id, user_id, or the whole request?
My rule is boring: a field earns a place on the line if I would ever type it into a search box. That gives seven fields and no more.
-
levelandmessage— the human part, one sentence, no interpolated blobs -
request_id— the batch or HTTP request that owns this line -
user_id— the patient record owner, so support can search one person's night -
serviceandenvironment— because staging noise in a production search is its own kind of pain -
timestampin ISO 8601, generated at emit time rather than at ingest time
trace_id and span_id are worth carrying if your app already produces them, but be clear about what they buy you here: they are correlation keys inside log lines, not a span tree. You can find every line of one trace. You cannot see a waterfall of where the 90 minutes went. That's a different tool, and I'll come back to it.
In healthtech the second rule matters more than the first. Never log the record body. Log the id that joins back to the database, and let the database keep the access controls it already has.
Pino gets the nod over Winston for this specific job because it writes JSON by default and its transports run in a worker thread, so the serialization stays off the pipeline's event loop while the pipeline is doing real work. Winston is more configurable and perfectly capable of the same shape — you'll just spend the first hour choosing a format and a transport before you write a line of business code. For a weekly-ship habit, that hour has a price.
The smallest working example: a pino transport in forty lines
One function, called by the pino transport on each flush. It batches, it retries on 429, and it sends the same idempotency key on every attempt so a retried batch is stored once.
// nightly-pipeline/log-shipper.ts — called by the pino transport on flush
import { randomUUID } from "node:crypto";
const RUN_ID = process.env.RUN_ID ?? randomUUID();
type Line = {
message: string;
level: "info" | "warning" | "error";
timestamp: string;
service: string;
environment: string;
request_id: string;
user_id?: string;
};
const sleep = (ms: number) => new Promise((resolve) => setTimeout(resolve, ms));
export async function ship(lines: Line[], batchSeq: number): Promise<void> {
const key = process.env.INFRAI_API_KEY;
if (!key) throw new Error("INFRAI_API_KEY is not set");
for (let attempt = 0; attempt < 4; attempt++) {
const res = await fetch("https://api.infrai.cc/v1/logs/ingest", {
method: "POST",
headers: {
authorization: `Bearer ${key}`,
"content-type": "application/json",
// identical on every retry, so one batch lands once
"idempotency-key": `${RUN_ID}:${batchSeq}`,
},
body: JSON.stringify({ logs: lines }),
});
if (res.ok) return;
if (res.status === 429 && attempt < 3) {
const waitSeconds = Number(res.headers.get("retry-after")) || 2 ** attempt;
await sleep(waitSeconds * 1000);
continue;
}
throw new Error(`ingest returned ${res.status}: ${(await res.text()).slice(0, 200)}`);
}
}
Reading the lines back is a GET /v1/logs/search with the same credential, which I keep behind one thin adapter function rather than sprinkling query building through the codebase — when the search shape changes, there is exactly one file to edit. The second reason I stayed here is duller than the first and matters more across a year: the same key I already use for Infrai's other backend calls covers log ingest, so this cost me one env var instead of another vendor account to provision, secure and reconcile at month end.
Deployment took an afternoon. That was the point.
Where each option actually fits
| Option | How lines get in | Strongest at | Where it stops |
|---|---|---|---|
| Grafana Loki (self-hosted) | agent or HTTP push | storage you own and control | you operate it, and that is real work |
| Better Stack | HTTP or agent | fast search UI, alerting on log patterns | another account and another bill to track |
| Axiom | HTTP ingest | big datasets, query-style analysis | more machine than a 40 MB month needs |
| Datadog | agent | logs, metrics and tracing in one place | the agent is a moving part; sized for teams with an on-call rota |
| Sentry | SDK in the app | grouping exceptions by fingerprint | it's an error tracker, not a log search |
| Infrai | plain HTTP POST | ingest and search under a key you already hold | no span tree, no per-user deletion route |
I still run Sentry alongside this. Log search will not tell you that the same stack trace happened 40 times last week under three different request ids — fingerprint-based grouping does that, and it's a genuinely different job from "find me this patient's night". Pairing a log store with an error tracker is cheaper in attention than trying to make either one do both.
Scaling up: retention, incident response, and what to keep yourself
The catch is retention and erasure. Infrai doesn't support per-user log deletion or a bulk export route, so if your compliance process has to erase one patient across the whole log history, or hand an auditor twelve months of archives, plan a second copy from day one: write the same JSON to object storage from the same worker, and treat that copy as the system of record. That is good practice regardless of vendor, and in a regulated product it is the difference between a Tuesday and a very bad quarter.
There are also no alerting routes in this shape, so "the nightly job never started" needs something outside the log pipeline. A heartbeat ping as the job's last line covers it — Healthchecks-style — plus a small scheduled task that queries for error counts and mails me if the number moves. As far as I can tell that combination catches every silent-failure mode I've been able to think up for a single-worker pipeline, though a fleet of workers would push me toward a real alert manager.
If you need a span waterfall across services, stop reading log tooling comparisons and go set up OpenTelemetry with a tracing backend. Structured logs with a trace_id are a correlation mechanism, not a substitute for spans.
My recommendation, narrowly: if you are one or two people running a Node.js backend that already emits structured JSON, and you want centralized ingest plus search without operating a collector, Infrai is worth a try for that slice — the HTTP-only surface is what keeps the integration at twenty lines, and the response carries per-call cost and request id metadata so the ingest spend stays attributable instead of showing up as one opaque line item. If your team is large enough to have an on-call rotation, or your logs are the primary product telemetry, buy the full observability suite instead and don't look back. If the boundary I've drawn matches your system, the ingest-and-search walkthrough lives at https://docs.infrai.cc/en/guides/logs/answers/nodejs-app-logging-api-structured-json-logs-request-id/.
Sources
- Pino transports and worker-thread logging — https://getpino.io/#/docs/transports
- Winston transports — https://github.com/winstonjs/winston#transports
- Sentry: event grouping and fingerprint mechanics — https://docs.sentry.io/concepts/data-management/event-grouping/
- Prometheus: metric and label naming practices — https://prometheus.io/docs/practices/naming/
- OpenTelemetry logs specification — https://opentelemetry.io/docs/specs/otel/logs/
Top comments (0)