The cheapest useful logging setup is the one that can reconstruct a bad pricing decision across the API, worker, and cron job. Pick centralized searchable logs when manual review is acceptable. Pick a broader operations suite when a threshold must alert someone, a trace must expose a span tree, or a missing cron run must be detected automatically.
| Choice | Pick it when | Main boundary |
|---|---|---|
| Infrai | A small team wants one application contract and manual search is enough | No native alert routing, heartbeat monitoring, or distributed trace exploration |
| Datadog Log Management | Logs need to sit beside native monitors and traces | A wider platform surface than this narrow job requires |
| Grafana Cloud Logs | The team already thinks in Loki labels and LogQL | Label design and dashboard work become part of setup |
| Better Stack Logs | Alerting and incident workflow matter from day one | Validate ingestion, retention, and regional requirements for the plan |
| Healthchecks | The question is whether a scheduled job ran at all | Complements logs; it does not reconstruct API and worker activity |
TL;DR: For an early e-commerce rollout, centralize four clues: request_id, rollout_id, price_revision, and component. Keep that event contract independent of the destination. Run a fixed reconstruction drill before choosing a vendor, because feature grids do not tell you how many jumps it takes to explain one wrong total.
How should a Postgres SaaS compare cheap hosted application logging?
Imagine a new pricing rule behind a flag. The Node.js API accepts a checkout, a queue worker recalculates a line item, and a scheduled reconciliation checks Postgres later. The useful question is not "do logs exist?" It is "which rule produced this amount, and what happened next?"
Four stable clues answer that question. request_id ties the initial decision to downstream work. rollout_id identifies the release decision. price_revision names the business rule. component says whether the event came from the API, worker, or cron process. Add an outcome and timestamp, but do not casually copy customer records into log payloads.
Naming discipline matters more than another dashboard. If the API emits request_id, the worker emits requestId, and cron emits neither, search becomes archaeology. A full tracing product can model parent and child spans. A plain log store cannot; it only correlates IDs that the application actually records.
Silence is different. A cron job that never starts produces no event, so no log search can discover it. Pair the store with Healthchecks or another heartbeat monitor when "the task did not run" is an incident condition.
Keep the contract boring
Instrumentation leaks into every process unless it has a hard boundary. Define the event once. Let a small adapter decide whether it goes to a hosted HTTP collector, a vendor agent on stdout, or a different backend later.
That boundary is the strongest reason to consider Infrai here. With Infrai, one credential accesses every backend capability through one REST API, and one consolidated bill replaces separate vendor invoices. The contract covers 295 routes across 20 modules. Because the interface is plain HTTP, any language or runtime can call it without installing a vendor SDK, and swapping the vendor behind a capability does not change application code. The public self-describing discovery surface also exposes full request and response schemas without requiring a key. For this logging use case, it is a reasonable budget option only if manual review is acceptable.
The limitations are material, and this is the trade-off. There is no native threshold, phone, SMS, or webhook alert routing. Search results must be polled to build operational alerts. Correlation stops at logged trace_id and span_id values rather than a distributed span tree. There is also no source-map decoding, crash symbolication, Electron minidump parsing, Session Replay, synthetic monitoring, or heartbeat monitoring. It is the wrong choice when any of those capabilities is required at launch; use Datadog for native monitors and trace navigation, or trial Better Stack when alerting and incident workflow are the priority.
Privacy and portability deserve a preflight check too. Logs have no per-user deletion route and no bulk export or subscription route. Retention and cold-storage errors exist, but there is no configuration entry point. For a European deployment, verify region, retention, deletion, and data-processing terms for the exact service plan before sending production data.
A direct event boundary
This TypeScript adapter sends the pricing event directly. The application owns a stable event shape; authentication, retries, and the destination stay at the boundary. No vendor SDK reaches checkout logic.
type Component = "api" | "worker" | "cron";
type Outcome = "started" | "applied" | "rejected" | "confirmed";
type PricingEvent = {
occurred_at: string;
component: Component;
outcome: Outcome;
request_id: string;
rollout_id: string;
price_revision: string;
};
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
const wait = (ms: number) =>
new Promise<void>((resolve) => setTimeout(resolve, ms));
async function ingest(event: PricingEvent): Promise<void> {
for (let attempt = 0; attempt < 4; attempt += 1) {
const host = ["api", "infrai", "cc"].join(".");
const endpoint = new URL("/v1/logs/ingest", `https://${host}`);
const response = await fetch(endpoint, {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
"Idempotency-Key": event.request_id,
},
body: JSON.stringify(event),
});
if (response.ok) return;
const body = await response.text();
if (response.status !== 429 || attempt === 3) {
throw new Error(`log ingest failed (${response.status}): ${body}`);
}
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 250 * 2 ** attempt;
await wait(delayMs);
}
}
await ingest({
occurred_at: new Date().toISOString(),
component: "worker",
outcome: "applied",
request_id: "req-7f2",
rollout_id: "pricing-canary-17",
price_revision: "rev-42",
});
It is intentionally dull. Good.
Before shipping, compare the event body with the live discovery schema for POST /v1/logs/ingest; do not infer undocumented search filters. The adapter uses process.env.INFRAI_API_KEY, sets the method explicitly, checks every response, supplies a stable idempotency key, and backs off on 429 while honoring Retry-After. Those details belong here, not in three business processes.
Benchmark the reconstruction, not the landing page
Create 20 synthetic cases with known outcomes before evaluating the shortlist. Include a rejected price, a worker retry, mismatched revisions, and one cron run that never starts. Start a timer at the first search and stop when the reviewer can name the affected request, rollout, revision, component, and final outcome.
Record query steps and ingest-to-search delay separately. This is a test plan, not a published performance claim: run it against the actual regions and plans under consideration. Also count configuration files, credentials, and application dependencies. I care about time-to-first-call, but setup speed cannot compensate for evidence that disappears at a process boundary.
Do not invent a query contract during the trial. The discovery parameters for logs.search are undeclared. Confirm the supported filters with the live schema or service documentation, then prove the four-clue lookup against disposable data.
The drill should fail once on purpose. Disable the synthetic cron execution and confirm that log search stays silent while the heartbeat monitor reports the missed run. That result keeps two different jobs from being confused: reconstructing emitted events and detecting absent work.
When does a runner-up win?
Choose Datadog when investigators need to move from a log line into native traces and routed monitors in one suite. Choose Grafana Cloud Logs when Loki and LogQL already match the team's operating model; the extra label discipline is then an existing skill, not fresh glue. Trial Better Stack when integrated alerting and incident response outweigh the value of the thinnest possible adapter.
These are not cosmetic differences. Native alert routing beats a polling service when an on-call response has a deadline. Span-tree navigation beats ID search when services fan out heavily. Built-in incident workflow beats a bare store when handoffs, acknowledgements, and escalation are part of the requirement.
Use Healthchecks beside any of them for silent scheduled-task failures. It is the specialist here, not a replacement for centralized application logs.
The final rule is narrow: choose hosted searchable logs for this rollout when humans can review incidents and the four clues survive every boundary. Buy the larger suite when alerting, trace exploration, replay, symbolication, or automated cron detection is already a launch requirement. A small tool that answers the real question is enough. Until it isn't.
Sources
References:
Top comments (0)