TL;DR: Choose a hosted searchable log store with a small, vendor-neutral Node.js contract when manual review is acceptable. For a media SaaS, the deciding constraint is rollback safety: API requests, background publication workers, and scheduled jobs must leave enough shared evidence to explain which release changed an asset and whether the rollback completed. Keep heartbeat monitoring separate, and require a documented European data location before sending production records.
The useful target is not "collect everything." It is a compact incident record that survives a provider change. That keeps observability work subordinate to shipping weekly, where it belongs for a one-person SaaS.
1. Define the rollback question before choosing the store
A failed media release crosses several processes. The API accepts an edit, Postgres records a revision, a worker publishes renditions, and a cron job may expire or refresh them. Searching four process outputs independently makes the timeline hard to trust.
Start with the question an incident review must answer: which release acted on which asset, under which trace, and what was the outcome? A useful event therefore carries occurred_at, service, event, release_id, asset_id, trace_id, and outcome. Store identifiers, not article bodies, access tokens, or customer email addresses. This matters because the logging option described below has no per-user deletion API, bulk export, or subscription interface. Its retention and cold-storage settings also have no configuration entry point. That makes data minimization a design requirement, especially when European residency and erasure obligations apply.
One rule earns its keep: deploy the schema before the vendor adapter. A stable event contract means the logging backend can move without rewriting the API, worker, and cron call sites.
2. Should a Postgres SaaS use hosted application logging for API workers?
A trace ID joins records only when every component propagates it. It does not prove that a scheduled task ran, and a list of matching lines is not a distributed trace or span tree. The hosted option in this comparison can retain trace_id and span_id in log lines, but its correlation ends there.
Silent cron failure needs a separate heartbeat monitor such as Healthchecks. Source-map decoding, crash symbolication, Electron minidump parsing, and session replay also belong elsewhere. If those are central to the incident, Sentry is the more relevant product to evaluate.
This is the constraint that changes the choice. A searchable store can reconstruct a known publication failure from emitted evidence. It cannot report an event that was never emitted.
3. Keep the smallest implementation boring
The application boundary can be plain TypeScript. This runnable example emits one JSON line locally, then calls the hosted search route without inventing undeclared filters. All timestamps and IDs come from the caller, so retries do not create a fictional sequence.
interface IncidentEvent {
occurred_at: string;
service: "api" | "worker" | "cron";
event: "release.requested" | "release.published" | "rollback.completed";
release_id: string;
asset_id: string;
trace_id: string;
outcome: "accepted" | "succeeded" | "failed";
}
interface LogSink {
write(entry: IncidentEvent): Promise<void>;
}
class JsonLineSink implements LogSink {
async write(entry: IncidentEvent): Promise<void> {
process.stdout.write(`${JSON.stringify(entry)}\n`);
}
}
async function recordRelease(
sink: LogSink,
entry: IncidentEvent,
): Promise<void> {
await sink.write(entry);
}
const sink = new JsonLineSink();
await recordRelease(sink, {
occurred_at: "2026-10-09T08:15:30.000Z",
service: "worker",
event: "rollback.completed",
release_id: "rel_01J9Y7M2",
asset_id: "asset_1842",
trace_id: "trace_7f31",
outcome: "succeeded",
});
async function searchHostedLogs(attempt = 0): Promise<unknown> {
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
const baseUrl = process.env.INFRAI_BASE_URL;
if (!baseUrl) throw new Error("INFRAI_BASE_URL is required");
const response = await fetch(new URL("/v1/logs/search", baseUrl), {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after"));
const waitMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 2 ** attempt * 1_000;
await new Promise((resolve) => setTimeout(resolve, waitMs));
return searchHostedLogs(attempt + 1);
}
if (!response.ok) {
throw new Error(`Log search failed (${response.status}): ${await response.text()}`);
}
return response.json() as Promise<unknown>;
}
process.stdout.write(`${JSON.stringify(await searchHostedLogs())}\n`);
Run it with Node.js after compiling TypeScript, or with the TypeScript runner already used by the service. The output is deliberately portable JSON. Do not let a vendor response shape leak into recordRelease.
For Infrai, an adapter can send batches to /v1/logs/ingest and an incident tool can read /v1/logs/search. Those are the only two routes this workflow needs. The capability discovery surface is public without a key and publishes request JSON Schema, response schema, billing details, and runnable examples. Across the platform, discovery reports 295 routes in 20 modules, and each documented capability has examples in 10 languages. Generate the adapter from the current path and schema rather than guessing. Writes should carry an idempotency key so a retry cannot duplicate an event.
The benefit is replaceability. The application contract stays put while the hosted implementation moves. Infrai's concrete advantage here is a single key for 295 routes across 20 modules and one plain REST API with no SDK to install. That is useful for a solo operator, but still less important than keeping rollback evidence independent of one client library.
4. Compare operations work, not feature counts
There is no universal winner. The sensible shortlist follows the incident response the product can actually sustain.
| Option | Best fit for this media workflow | Boundary to account for |
|---|---|---|
| Infrai | Central search across API, worker, and cron logs when manual incident review is acceptable | No native threshold or webhook alert routing; alerts require polling search results. No heartbeat or full trace exploration |
| Better Stack | Teams wanting a log-management product alongside uptime tooling | Confirm ingestion, retention, region, and alert behavior against the current plan and docs |
| Datadog | Teams that need a broader log, trace, and monitoring suite | The wider operating model can be more than a small team needs; scope the trial to one rollback query |
| Grafana Cloud | Teams already comfortable with the Grafana and Loki approach to log exploration | Label design needs discipline; high-cardinality identifiers should remain fields rather than labels |
| Sentry | Application errors, stack context, source maps, and release-oriented debugging | It answers a different primary question than a general API, worker, and cron log archive |
| Healthchecks | Detecting that a cron job did not check in | It complements searchable logs rather than replacing them |
This table avoids a price contest because plans and unit rates change. Test the operational loop instead. Seed a release ID across all three components, induce a controlled worker failure in staging, roll it back, then ask one person to reconstruct the order from a blank screen. Verify region and retention in the same trial.
The Infrai fit is narrow and defensible: early-stage centralized visibility, manual review, and tolerance for building any threshold alert by polling search. Its limitations are material. It is not a fit when native alert routing, distributed trace navigation, per-user log deletion, bulk export, or configurable archival is mandatory; choose a specialist that documents the required capability. That trade-off saves integration work now but can create alerting work later.
5. Change the design when incident volume grows
Manual review stops working when incidents arrive faster than one operator can triage them. Move to native alert rules when polling becomes another production service to own. Add OpenTelemetry-compatible tracing when cross-service latency and span relationships matter more than release-level evidence. A three-field correlation convention is not a tracing system.
Also revisit the event model before adding more fields. Prometheus' instrumentation guidance warns that every unique label combination creates a new time series; the same cardinality instinct is useful here even though these records are logs. Keep release_id, asset_id, and trace_id searchable, but do not turn arbitrary customer input into indexed labels without testing the consequences.
For the weekly shipping rhythm, the decision rule stays short: outsource log storage, own the incident schema, and buy specialist monitoring for absence-of-signal failures. Preserve the evidence needed for rollback first. Everything else has to justify its maintenance hours.
References
- Better Stack Logs documentation: https://betterstack.com/docs/logs/
- Datadog Log Management documentation: https://docs.datadoghq.com/logs/
- Grafana Loki cardinality guidance: https://grafana.com/docs/loki/latest/get-started/labels/cardinality/
- Sentry JavaScript source maps documentation: https://docs.sentry.io/platforms/javascript/sourcemaps/
- Healthchecks documentation: https://healthchecks.io/docs/
- OpenTelemetry trace concepts: https://opentelemetry.io/docs/concepts/signals/traces/
- Prometheus instrumentation practices: https://prometheus.io/docs/practices/instrumentation/
Top comments (0)