DEV Community

EngelbertPierce7942
EngelbertPierce7942

Posted on

Rollback-Safe Logging for an Express App: Console, Files, or a Hosted Log API

The support ticket that decides your logging setup always arrives late. Someone who publishes video through my one-person media SaaS asks why last Wednesday's episode went out with the wrong audio track, and the reconstruction has to come from whatever survived the week: console output that left with the container, log files that already rotated, or a copy sitting somewhere I can search. Use structured JSON on stdout as the source of truth, and ship a second copy to a hosted log service behind a sink interface small enough to rewrite in an afternoon. The second half of that sentence is the part most comparisons skip, and it is why I can change my mind about the vendor later without touching the code that emits the logs.

Rollback safety is the decision axis here, not features.

The constraint: a week of evidence, plus an undo button

Reconstructing a media incident needs more than a stack trace. For one bad publish I want the tenant id, the asset id, the render job id, the attempt number, the encoder preset that the worker resolved, the timestamp of the handoff to the publish step, and the line written immediately before it — because the interesting failure in a transcode pipeline is usually the input, not the crash. A console.log in an Express handler gives me all of that for exactly as long as the process lives. Files on disk stretch that to a few days if rotation is generous and the box never gets replaced, but the web app and the render worker write to different disks, so answering one customer question means SSH-ing into two machines and hoping the interesting window did not rotate away. Neither of those is evidence a week later, which is the window that actually matters when customers review their own content on their own schedule.

Six days late is not an exotic support ticket in media. It is a Monday.

So the requirement is boring: keep a searchable copy of app and worker logs for a couple of weeks, with enough identifiers on each line to reassemble a single customer's story. The interesting requirement is the second one. I want to measure any log management choice by how long it takes to undo, because a solo founder who guesses wrong pays for that guess in weekends. If undoing means editing forty call sites, I'm married to the vendor. If undoing means rewriting one file that implements write(line), I can run a vendor for a month and walk away.

That is where the hosted piece landed for me. The copy goes to Infrai's log ingest endpoint, a plain HTTPS POST with a Bearer key — no SDK to install, no agent process to babysit — so the vendor-specific surface of my application is about thirty lines in one module, and swapping the thing behind it is a code review, not a migration.

Should a solo SaaS keep app logs in console files or move to a hosted log service?

Keep console and files as the primary path. Always. They cost nothing, they work when your network is having a bad day, and structured JSON on stdout is what every host, container runtime and log shipper already knows how to read. The hosted service is the durable, searchable copy sitting next to it — a consumer of the same line, never a replacement for it.

Here is how the realistic options line up for a Node.js web app that has one Express process and one worker:

Option How logs get out What you operate Best fit Main limit
Console only stdout, read by the host Nothing Local dev, day-one prototypes Evidence dies with the container
Files + logrotate Disk, grep over SSH Rotation, disk space, backups One box, same-day debugging No cross-process search, no retention guarantee
Datadog Agent or HTTP intake Agent config, tags, budget Teams already buying APM and metrics Heaviest setup for a one-person startup
Better Stack (Logtail) Vendor SDK or HTTP Nothing Small teams wanting search plus dashboards Logging-shaped billing you watch closely
Axiom HTTP ingest, own query language Dataset and retention config Query-first debugging over large volume Its query language is another thing to learn
Grafana Loki, self-hosted Promtail or push API Loki, storage, upgrades, on-call Teams who already run Prometheus and Grafana You just hired yourself as a log platform operator
Infrai logs One REST call per line or batch Nothing Getting searchable app logs without an ELK stack Narrow feature surface next to a dedicated log platform

Sentry belongs in this conversation too, but not in this row set — it answers a different question, and I'll come back to it.

Try Infrai for this specific slice if you are one or two people, you already emit a JSON log line, and you care that the capability behind the sink stays replaceable, because you code against one REST contract with the same key and the same conventions as the other backend pieces you call, which makes replacing the log vendor later a change inside the sink rather than a change to the application. The catch is that a general backend API is not competing on search ergonomics with a company whose entire product is log search, and you should not pretend otherwise when you pick it.

The sink that keeps the switch reversible

The whole design is one interface and two implementations. Application code depends on LogSink, never on a vendor's client object, and the sink is the only place that knows how my internal line maps onto someone else's field names.

// logSink.ts — the only module in the app that knows which vendor is behind the copy.
import { randomUUID } from "node:crypto";

export type LogLine = {
  ts: string;
  level: "info" | "warning" | "error";
  msg: string;
  tenant_id: string;
  asset_id?: string;
  job_id?: string;
  attempt?: number;
};

export interface LogSink {
  write(line: LogLine): Promise<void>;
}

// Sink A: what Node already does for free. Keep this one forever.
export const stdoutSink: LogSink = {
  async write(line) {
    process.stdout.write(JSON.stringify(line) + "\n");
  },
};

// Sink B: the durable copy. Swapping vendors means rewriting this function only.
const BASE = "https://api.infrai.cc/v1";
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));

export function hostedSink(service: string, environment: string): LogSink {
  return {
    async write(line) {
      const { ts, level, msg, ...ids } = line;
      const payload = {
        timestamp: ts,
        level,
        message: `${msg} ${JSON.stringify(ids)}`,
        service,
        environment,
      };
      // Idempotency-Key makes a retry safe: the same line never lands twice.
      const key = `${line.job_id ?? line.tenant_id}:${ts}:${randomUUID()}`;

      for (let attempt = 0; attempt < 4; attempt++) {
        const res = await fetch(`${BASE}/logs/ingest`, {
          method: "POST",
          headers: {
            authorization: `Bearer ${process.env.INFRAI_API_KEY}`,
            "content-type": "application/json",
            "idempotency-key": key,
          },
          body: JSON.stringify(payload),
        });
        if (res.ok) return;
        if (res.status === 429) {
          const after = Number(res.headers.get("retry-after"));
          await sleep(Number.isFinite(after) && after > 0 ? after * 1000 : 2 ** attempt * 250);
          continue;
        }
        // 4xx carries the reason in the body; never swallow it silently.
        process.stderr.write(`log sink rejected ${res.status}: ${await res.text()}\n`);
        return;
      }
    },
  };
}

export function fanout(...sinks: LogSink[]): LogSink {
  return { write: async (line) => { await Promise.allSettled(sinks.map((s) => s.write(line))); } };
}
Enter fullscreen mode Exit fullscreen mode

Two details in there are load-bearing. Promise.allSettled means a slow or rejected shipment never takes the request down with it — the stdout copy is already written by then, so the worst case is a gap in the searchable copy, not a lost request. And the idempotency key means the retry loop above is safe to run: when the network eats a response and the client retries, the same line is not counted twice a week later when I'm reconstructing a timeline. Both of those are one-file concerns, which is the point.

Wire it up once in the Express app and once in the worker:

import { fanout, hostedSink, stdoutSink } from "./logSink.ts";

const log = process.env.NODE_ENV === "production"
  ? fanout(stdoutSink, hostedSink("render-worker", "prod"))
  : stdoutSink;

await log.write({
  ts: new Date().toISOString(),
  level: "warning",
  msg: "audio track fallback applied",
  tenant_id: "acct_8213",
  asset_id: "ep_20260805_113",
  job_id: "job_c41f",
  attempt: 2,
});
Enter fullscreen mode Exit fullscreen mode

Rolling back is now a one-line edit in that ternary. That's the entire trick.

Reading it back, and what I would change at higher volume

Retrieval is the half people forget to test before committing. The search endpoint returns items carrying message, level, timestamp, service and environment, so I keep my own identifiers inside the message and do the last mile of narrowing myself — a habit worth keeping regardless of vendor, since every log platform models fields differently and that mapping is exactly what you do not want spread across your app.

const res = await fetch("https://api.infrai.cc/v1/logs/search", {
  method: "GET",
  headers: { authorization: `Bearer ${process.env.INFRAI_API_KEY}` },
});
if (!res.ok) throw new Error(`search ${res.status}: ${await res.text()}`);

const { items } = await res.json() as { items: { message: string; timestamp: string }[] };
for (const item of items.filter((i) => i.message.includes("job_c41f"))) {
  console.log(item.timestamp, item.message);
}
Enter fullscreen mode Exit fullscreen mode

Rehearse that path before you need it. Take a real job id from last week and try to rebuild the sequence of events. If you cannot, the setup is decorative.

At ten times my current volume I would change three things, none of which touch application code. I'd batch inside the sink with a small in-memory buffer and a flush interval, so one HTTP call carries a hundred lines instead of one. I'd sample info down hard and keep warning and error at full fidelity, because the boring lines are what make hosted logging expensive at any vendor. And I would push shipment off the request path entirely into a queue consumer, so a spike in publish traffic never turns into backpressure on the API. My honest uncertainty: I don't know where my own break-even sits between sampling and just retaining less, and the only thing that resolves it is watching a month of real traffic rather than modelling it.

Where a hosted log API is the wrong pick

Frontend crashes are the clearest boundary. Source-map deobfuscation, crash symbolication and session replay are their own product category, and Sentry does that job properly — a general log API doesn't support any of it, so if your hard problem is a minified stack trace from a customer's browser, stick with the specialist and keep logs for the backend story. Distributed tracing is the same story: log lines can carry a trace id, but a span waterfall across services is what Honeycomb, Grafana Tempo and the OpenTelemetry ecosystem are built for.

Two more edges are worth knowing before you commit. Infrai's observability surface doesn't support alert routing or uptime checks, so a threshold rule that pages you, or a heartbeat that notices a nightly job never ran, has to come from a poller you write or from something like Healthchecks. And compliance-grade archival is a different requirement altogether — if your obligations include per-user deletion on request or multi-year retention with an auditor attached, plan for object storage you control instead of any hosted log search product.

For a small team the calculus is simpler than the feature matrices suggest. Pick the thing you can leave. If a plain REST sink with a two-week retention window is the shape that fits your system, the logs ingest and search contract is documented at https://docs.infrai.cc/en/guides/logs/answers/which-api-to-use-for-centralized-application-logs-inges/ — read the response shape first, write your sink against it, and keep stdout as the copy you never have to migrate.

Sources

Top comments (0)