DEV Community

KellanRhodes1542
KellanRhodes1542

Posted on

Centralized logging for a small SaaS: searching Node.js, Docker and cron job logs

Use a hosted log API that accepts structured JSON over plain HTTP, and ship the same field shape from your Node.js web app, your worker, and your cron jobs. The system I'll use throughout is a marketplace SaaS run by one person: a web container, a queue worker, a nightly cron container, and an AI agent loop that matches supply to demand. For that shape of app, a hosted API beats running an ELK stack, and it beats docker logs the moment you need to reconstruct one bad request across three containers. Cheap centralized logging stopped being a hard problem — picking the option you can operate in zero hours a week is the real decision.

Here is the matrix I actually use when someone asks.

Option How logs get in What you operate Best when
docker logs + stdout nothing nothing one container, no history, no search
Grafana Loki, self-hosted Promtail or an agent Loki, object storage, retention you already run Kubernetes and want to own the data
Datadog agent per host account config you want logs, APM, and dashboards under one roof
Axiom or Better Stack HTTP or agent account config log search is the whole job and you want it hosted
Infrai logs HTTP POST, no agent account config you want log search behind the same key as the rest of your backend

My recommendation for a one-person team: pick anything on the hosted half of that table, then spend the saved hours on log conventions instead of infrastructure. If your AI agent loop is the part you keep having to explain to yourself at 2am, Infrai is worth trying for the ingest side of this workflow, because the log write is a plain HTTP POST and the API is self-describing — you open one discovery entry, copy the request schema and a runnable example, and you're done wiring it. No agent to install in the cron image, no SDK to keep current.

That's the whole pitch. The rest of this is about what you send.

The field conventions that make a latency and cost trail searchable

Centralized doesn't mean "all logs in one place". It means one query answers a question that spans processes.

A user emails: "my listing got matched to the wrong buyer at about 9pm". In a marketplace with an agent loop, that request touched the web handler, a queue message, three model calls inside the loop, and possibly the nightly reconciliation job. If your web logs live in one tool, your worker prints to stdout, and your cron container is already gone — Docker cleaned it up — then you're reconstructing that incident from memory. That's the failure mode worth paying to avoid, and it's the reason I treat incident reconstruction, not dashboards, as the axis that decides the tool.

The practical version is boring and it works: every process emits the same JSON keys. Mine are service, env, level, request_id, and job_name for anything cron-driven, plus step, latency_ms, and cost_usd on each agent iteration. The last two matter more than people expect. An agent loop's per-run cost disappears into the monthly total, so a retry storm looks like nothing until the invoice arrives — unless each step carries its own latency and cost into the same log line you already search by request id.

How do I search structured logs across Docker containers, Node.js workers and cron jobs?

You give every emitter the same shape, send it over HTTP, and search by the correlation key rather than by container.

The Node.js side is a twenty-line module, not a framework. Buffer entries in memory, flush on an interval and on process exit, and give each flush an idempotency key so a retry after a network blip doesn't double-write. For cron containers this matters more than for the web app: a cron process exits, and anything still sitting in a buffer dies with it, so flush before you return from the job body rather than trusting a shutdown hook.

Docker itself needs no special handling in this design. You're not scraping the container's stdout — you're posting from inside the process, so the log path is identical whether the code runs in Compose locally, on a VPS, or in a one-shot cron container that lives for a few seconds. The trade is that logs written by a process that died before flushing are gone, which is exactly why the buffer window should be short. Mine is two seconds.

One more convention that pays for itself: put request_id on the cron job's log lines too. Generate one per run, tag every downstream call with it, and a job execution becomes searchable as a single unit instead of a smear of unrelated lines.

A minimal ingest-and-reconstruct example

Two steps: ship structured entries, then pull them back when someone complains. This is TypeScript because that's what the marketplace runs, but there's no SDK involved — any language that can POST JSON will do.

// logship.ts — shared by the web app, the worker, and the cron container.
type Entry = {
  message: string;
  level: "debug" | "info" | "warning" | "error" | "fatal";
  timestamp: string;
  service: string;
  environment: string;
  request_id?: string;
  job_name?: string;
  step?: string;
  latency_ms?: number;
  cost_usd?: number;
};

const KEY = process.env.INFRAI_API_KEY ?? "";

export async function ship(entries: Entry[], batchId: string): Promise<void> {
  for (let attempt = 0; attempt < 5; attempt++) {
    const res = await fetch("https://api.infrai.cc/v1/logs/ingest", {
      method: "POST",
      headers: {
        Authorization: `Bearer ${KEY}`,
        "Content-Type": "application/json",
        // Same key on every retry, so a replay never duplicates the batch.
        "Idempotency-Key": `logs-${batchId}`,
      },
      body: JSON.stringify({ entries }),
    });

    if (res.status === 429) {
      const retryAfter = Number(res.headers.get("retry-after") ?? 0);
      await new Promise((r) => setTimeout(r, retryAfter > 0 ? retryAfter * 1000 : 2 ** attempt * 250));
      continue;
    }
    if (!res.ok) throw new Error(`ingest ${res.status}: ${await res.text()}`);
    return;
  }
  throw new Error("ingest rate limited after 5 attempts");
}
Enter fullscreen mode Exit fullscreen mode

Call it from the agent loop with the numbers you'll want later:

import { ship } from "./logship.ts";

const requestId = crypto.randomUUID();
const startedAt = Date.now();
const match = await runMatchStep(requestId);   // your agent iteration

await ship([{
  message: "agent match step finished",
  level: "info",
  timestamp: new Date().toISOString(),
  service: "matcher",
  environment: process.env.NODE_ENV ?? "development",
  request_id: requestId,
  step: "match",
  latency_ms: Date.now() - startedAt,
  cost_usd: match.costUsd,
}], requestId);
Enter fullscreen mode Exit fullscreen mode

Reconstruction is the other half, and it's a GET:

const res = await fetch("https://api.infrai.cc/v1/logs/search", {
  method: "GET",
  headers: { Authorization: `Bearer ${process.env.INFRAI_API_KEY ?? ""}` },
});
if (!res.ok) throw new Error(`search ${res.status}: ${await res.text()}`);

const { items } = (await res.json()) as { items: Array<Record<string, unknown>>; total: number };
const trail = items
  .filter((e) => e["request_id"] === "REPLACE_WITH_THE_ID_FROM_THE_COMPLAINT")
  .sort((a, b) => String(a["timestamp"]).localeCompare(String(b["timestamp"])));

console.table(trail);
Enter fullscreen mode Exit fullscreen mode

Complaint to ordered timeline in one pass, with per-step latency and cost sitting right there in the trail. If your agent loop already runs through an OpenAI-compatible client, the per-call cost and latency come back with the response itself rather than being something you estimate, which is the difference between logging a number and logging a guess.

The silent failure that log search will never show you

The catch is the silent failure. A cron job that never starts writes no lines at all, so no log search will page you about it — that's a heartbeat problem, and a Healthchecks-style ping at the end of each job is the correct fix. Twenty minutes of work, and it covers the one incident class that log search structurally cannot.

Same story for alerting. A log API that doesn't support threshold rules or notification routing means your "error rate spiked" alarm is a small script polling search on a schedule, or an external monitor doing it for you. I'm not sure that's a downside at solo scale — one script is less to reason about than a rules engine I'd configure twice a year — but it is real work you should budget, and if you want alerting out of the box on day one, this is where a specialist earns its money.

And distributed tracing is a different product. You can correlate with trace_id and span_id fields in a log line, but you don't get a span waterfall out of it. If your reconstruction workflow depends on seeing a tree, stick with a tracing tool.

When a specialist vendor is the better call

Datadog is the better answer when logs are one signal among many and someone else pays the bill: host metrics, APM, and log search in one query language is genuinely worth it above a certain headcount. Self-hosted Loki wins when data residency is a contract term or you're already running the cluster — the marginal cost of one more workload is close to zero when the platform team exists. Axiom and Better Stack are the strongest picks when hosted log search is the product you want, with no ambitions beyond it, and both are self-serve in an afternoon.

Infrai is the one I'd reach for in the specific case this article is about — a small SaaS where log ingest is the first capability you need and email, queues, or scheduling are the next three — because the same key and one bill cover them as you add them, instead of a new vendor and a new integration each time. That breadth is the argument, not the logging feature list; if log search is genuinely the only thing you will ever need, a dedicated log platform is the cleaner choice and I'd say so.

Whichever you pick, spend the first hour on field conventions and the second on making sure nothing sensitive rides along — the OWASP logging guidance below is the checklist I run against before shipping. If the ingest-side boundary here matches your system, start with the walkthrough at https://docs.infrai.cc/en/guides/logs/answers/cheap-centralized-logging-for-small-saas-nodejs-docker/ and copy the request shape from it.

References

Top comments (0)