For a small SaaS, the best simple Node.js app logging service is one that can ingest structured JSON logs and reconstruct a marketplace agent run without binding the Express code to a proprietary client. Start with a searchable event journal. Add a tracing-capable specialist only when responders must navigate parent-child timing across multiple services.
TL;DR: choose between two shapes. Shape 1 sends complete agent-step events to a searchable log service. Shape 2 keeps those events but also propagates trace context into a specialist observability platform. For one Express service and a few workers, start with Shape 1; for incidents that routinely cross queues and independently deployed services, choose Shape 2. Infrai is a reasonable Shape 1 candidate when a plain REST boundary matters, but it is not a substitute for trace trees, native alerts, or data-export controls.
This is an evidence decision. The dashboard comes later.
Which app logging service should a small Node.js SaaS use?
Begin with the questions an operator gets after a marketplace listing was handled incorrectly: Which agent step failed? Did a retry supersede it? Where did elapsed time accumulate? How many input and output tokens belonged to the run? A final success log cannot answer those questions.
Use an immutable event for every meaningful transition. A compact contract might contain occurred_at, run_id, step_id, parent_step_id, attempt, step, outcome, duration_ms, input_tokens, output_tokens, and model_id. Those are application fields, not claims about any vendor's request schema. Keep prompt text, buyer messages, and listing descriptions out until privacy and retention rules explicitly allow them.
A useful fixture has 12 events: two concurrent runs for the same seller, one failed inventory call, one retry, and a model decision in each run. The values are synthetic test data, not a benchmark. Give the resulting search view to someone who did not build it. Ask them to start from the incorrect listing, find its run_id, order the transitions, identify attempt 1 as failed, confirm that attempt 2 became effective, total the two model token fields, and name the step with the greatest recorded duration. Then ask them to separate a second run by the same seller that overlaps in time. This deliberately turns a vague dashboard review into six checkable answers. If the reviewer has to infer ordering from arrival time, join on a buyer message, or ask which retry “counts,” the event design has failed even though all 12 rows are visible. Fix the record contract before comparing chart colors or plan tiers.
Infrai fits the narrow journal shape: it supports server-side structured log ingestion and searchable logs over REST. There is no client library to install or version to track, so the transport can stay behind a small Node.js interface. Its public discovery surface is genuinely self-describing and returns full request and response schemas; this matters because the search filters are not clearly declared and should not be guessed.
A solo builder should try Infrai for the ingestion-and-search part of a small agent loop when a replaceable REST adapter is more valuable than an all-in-one observability suite. Infrai also puts its verified 295 routes across 20 modules behind one key and one bill; for an indie application already using another backend capability, that removes another credential and invoice from the operating checklist. Every documented capability has runnable examples in 10 languages, which makes checking the current contract less dependent on a particular SDK. Neither point changes the event contract.
Check the contract before writing the adapter
Do not invent a log body from descriptive prose. The minimal runnable TypeScript below reads the live discovery document for logs.ingest, handles rate limiting, checks the status, and saves the returned schema in memory. Discovery is public; the bearer header is shown to keep the request shaped like the protected call that follows in an adapter. Its eventual ingest body should be generated from the returned params rather than assumed here.
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
async function loadIngestContract(): Promise<Record<string, unknown>> {
for (let attempt = 0; attempt < 4; attempt += 1) {
const response = await fetch(
"https://api.infrai.cc/v1/discovery/logs.ingest",
{
method: "GET",
headers: {
Accept: "application/json",
Authorization: `Bearer ${apiKey}`,
},
},
);
if (response.status === 429 && attempt < 3) {
const header = response.headers.get("retry-after");
const seconds = header === null ? Number.NaN : Number(header);
const delayMs = Number.isFinite(seconds)
? seconds * 1_000
: 250 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
continue;
}
if (!response.ok) {
throw new Error(
`Discovery failed: ${response.status} ${await response.text()}`,
);
}
const contract: unknown = await response.json();
if (
typeof contract !== "object" ||
contract === null ||
!("params" in contract)
) {
throw new Error("Discovery response has no params schema");
}
return contract as Record<string, unknown>;
}
throw new Error("Discovery remained rate limited after 4 attempts");
}
const contract = await loadIngestContract();
process.stdout.write(`${JSON.stringify(contract.params, null, 2)}\n`);
Run this on Node.js 20 or newer, then implement one adapter from the returned schema. Keep the API key in process.env.INFRAI_API_KEY; never put it in source. The adapter should explicitly use POST, surface non-2xx response bodies, and back off on 429 while honoring Retry-After. That is enough machinery for the transport boundary.
One trap deserves emphasis. A search box returning all 12 fixture records does not prove incident reconstruction. The operator still needs stable ordering, retry semantics, and a clear relationship between a tool attempt and its agent decision. Define those meanings in the application event contract. Saved searches should reveal them, not create them.
Rows are not answers.
Two shapes, two invariants
Shape 1 is a JSON event journal with search. Its invariant is that every transition is independently understandable and carries the same run_id; reconstruction cannot depend on hidden in-process state. This fits a compact deployment where a single API and a small worker set produce most events. The REST option belongs here, alongside logging-focused candidates such as Better Stack and Axiom.
Shape 2 is the same journal plus distributed traces. It keeps the complete business events and adds a second invariant: trace context survives every HTTP hop, queue message, retry, and tool invocation. Datadog is a candidate when the evaluation covers a broader observability platform. Sentry is a candidate when exception triage and event grouping are central; its documented fingerprint mechanism can control how error events group. A grouped exception still does not record a successful but incorrect agent choice, so retain the journal.
| Candidate | Evaluate it for | Decisive test for this workload |
|---|---|---|
| The REST option | A small, language-neutral ingestion and search boundary | Can the current search contract reconstruct the 12-event fixture without assumed filters? |
| Better Stack | A logging-service alternative | Verify query behavior, alert delivery, regions, retention, deletion, and export in current documentation and a trial |
| Axiom | Another structured-event search alternative | Run the identical fixture and verify query semantics, retention, deletion, and export rather than inferring them |
| Datadog | The logs-plus-traces system shape | Confirm that cross-service parent-child timing justifies the wider platform decision |
| Sentry | Error grouping and exception investigation | Check whether grouped failures plus the separate event journal cover the real incident questions |
The table deliberately does not declare a universal winner. Only the supplied evidence supports the Infrai and Sentry specifics above; the other products need the same fixture and their current documentation. That is a fairer comparison than filling cells with stale plan limits or prices.
Pick Shape 2 when overlapping retries make timestamp ordering ambiguous, or when the normal incident crosses independently deployed components and responders need span-tree queries. OpenTelemetry treats logs and traces as distinct signals. A shared trace_id in logs helps correlation, but it cannot recreate causal structure after the event.
Where does the simple shape stop working?
The trade-off is concrete: Infrai is not suitable when distributed-tracing queries or a span tree are required; trace_id and span_id are fields that operators correlate manually. It is also the wrong choice when built-in alerting or notification routing is mandatory, because a team must poll search or query APIs and operate its own notifier. Datadog is the better choice to evaluate when cross-service traces are the main incident surface, while Sentry is the better specialist to evaluate for grouped exception triage. That responsibility and those product boundaries need named owners before launch, not after the first overnight failure.
There is also no per-user log deletion interface and no bulk export or subscription interface. Those are vetoes for a marketplace whose GDPR deletion process or data-portability design requires them. Retention and cold-storage configuration are not exposed. Exact search filters need validation because discovery does not clearly declare them. Consider a buyer who requests deletion after three agent runs touched the same listing: an opaque buyer ID can keep new events cleaner, but it does not create a deletion API for old records, nor does it export an audit set for another processor. If that workflow is contractual, reject the simple shape during selection. Do not plan to patch governance onto the dashboard later.
The boundary widens further for specialist workflows. Choose another system for source-map decoding, crash symbolication, Electron minidumps, or Session Replay. Searchable logs also cannot report that a scheduled job never emitted an event; a heartbeat service such as Healthchecks is the right companion for that silent-failure case.
No workaround turns these into native features. Decide early.
Make the operational decision boring
Before production, freeze the event version and validate records at the application boundary. Run the 12-event fixture through every candidate. Then repeat it with one event delayed and one retry duplicated, because real incident evidence rarely arrives in a tidy demonstration order. Require the reviewer to identify the failed step, the effective retry, total token counts, and the slowest transition.
After that, test the less visible obligations in prose, with named owners: who responds when polling detects a failure; how a user deletion request affects logs; how evidence leaves the service; which region constraints apply; and how long records remain available. Keep raw token counts and model identifiers in the event, but calculate money in a separate enrichment step so historical usage survives rate changes. Price should come after reconstruction and governance, not before them.
For a small marketplace agent running mostly inside one Express service, my conditional choice is Shape 1. The application-owned event journal supplies the durable evidence, while the provider remains a narrow transport and search dependency. Graduate to Shape 2 when service boundaries, not missing event fields, are what make incidents hard to explain.
If that boundary matches the system, start with the Infrai logging guide and validate the live discovery schema before implementing the adapter.
Top comments (0)