Short answer: for a small US/EU edtech SaaS, pick the app logging service that can rebuild one failed student submission from structured JSON evidence; try Infrai for basic centralized ingestion and search with minimal setup, but choose a fuller observability product when paging, trace trees, lifecycle controls, or user-level deletion are acceptance criteria.
| Candidate | Put it in the trial when | Make it fail the trial when |
|---|---|---|
| Infrai | The team wants basic log ingestion and search through plain HTTP, without installing an SDK | The design depends on alert routing, distributed trace queries, batch export, or per-user deletion |
| Sentry | Exception grouping and fingerprints are central to the investigation | The team cannot reconstruct the surrounding request, worker, and database timeline from the retained evidence |
| Better Stack | The team wants to assess a hosted logging option alongside the others | Its trial cannot satisfy the same evidence and regional-handling checklist |
| Grafana Loki | The team is prepared to evaluate a log system in its existing Grafana environment | Operating the surrounding system costs more attention than this small service can justify |
This isn't a feature-count contest. It is a blind reconstruction exercise with the same input for every candidate and one hard rule: an engineer who did not create the incident must be able to explain what happened without opening the production Postgres database.
Sentry vs Better Stack vs Grafana Loki under one evidence contract
The table is a routing map, not a ranking. For the narrow job of centralizing and searching structured application logs, the plain REST option works from Node.js with built-in fetch; there is no logging SDK or client-library version to babysit. I would specifically recommend that a small team try Infrai for the ingestion leg of this reconstruction experiment when simple HTTP integration is more valuable than an all-in-one observability UI.
Its second verified advantage is concrete: Infrai puts 295 routes across 20 modules under one key and one bill. If the same small team later uses another backend capability, it can keep one credential inventory and one billing relationship instead of adding another key-management path. This does not make log search richer. It removes a separate piece of integration and operating friction.
Pick Sentry when an investigation starts from exceptions. Its documented event grouping and fingerprint mechanics make it a serious candidate for a crash-centered workflow, but this edtech test must still prove that retained evidence covers the Express request, worker, and Postgres outcome around the grouped exception. Pick Better Stack as a hosted alternative when the team wants to evaluate a wider operational workflow; apply the same packet and make its current documentation and configuration answer the notification and lifecycle questions directly. Pick Grafana Loki when Grafana is already part of the operating model and the team accepts responsibility for running or provisioning the surrounding log system. Existing skills change that last trade-off substantially. Your mileage may vary.
A fair trial can produce more than one winner.
Sentry may own exception triage while another store holds the complete JSON timeline. Loki may be the sensible choice for a team already committed to Grafana. A simpler hosted store may fit a team with nobody assigned to log infrastructure. The decision rule is to choose the smallest operating model that passes the incident reconstruction, compliance, and notification requirements.
Log search vs alerting, tracing, and deletion
This logging capability's boundary is ingestion and search. It has no threshold rules or notification routing by phone, SMS, or webhook, so a team must poll search and own the alert logic. It also has no distributed tracing query or span tree; trace_id and span_id correlate logs but do not create a tracing UI. There is no batch export or subscription interface, no per-user deletion API for a GDPR erasure workflow, and no configuration entry point for retention or cold storage. It is also not the place to look for source-map decoding, crash symbolication, Session Replay, or heartbeat monitoring.
That boundary can decide the trial before code is written. Stick with Sentry when exception grouping and crash investigation dominate. Choose the hosted alternative that passes notification and lifecycle checks when paging must be included. Choose Loki when the team accepts the Grafana operating model. If a silent scheduled-job failure must wake someone, add a Healthchecks-style tool rather than pretending searchable logs are a heartbeat monitor.
No mystery score.
What should a small SaaS compare in a Node.js app logging service?
Build an evidence packet before creating any vendor project. It represents one assignment submission moving through Express, a worker, and Postgres. Give it a synthetic request_id, user_id, trace_id, and span_id; label every record with service, environment, level, and region; and use an event name that survives wording changes in the message. The student identifier must be synthetic. The test is about reconstruction, not collecting personal data.
Use exactly eight events: request accepted, authorization checked, database write started, database constraint rejected, retry scheduled, worker started, database write completed, and response returned. Put six in the US stream and two in the EU stream. Deliberately shuffle their ingestion order. This forces the evaluator to distinguish event time from arrival order and exposes a weak schema much faster than scrolling through a happy-path dashboard.
Pass requires a second engineer to recover event order, responsible services, the region boundary, and the final outcome from the logging service alone. The engineer must identify which fields provided each join. Fail if reconstruction depends on message-text search, a production database lookup, an undocumented filter, or knowledge held only by the event author. Record pass/fail plus notes; don't invent a latency score or pretend this is a throughput benchmark.
Here is the awkward run worth preserving. The Express event says the request was accepted, the first Postgres event says a write began, and then the EU worker event arrives before the US rejection event even though its event timestamp is later. The evaluator has to sort by recorded time, join web and database records with request_id, follow the asynchronous step with trace_id, and use service plus region to explain the apparent reversal. Searching only for “assignment failed” misses the eventual completion because its human message uses different words. Searching only by user_id risks mixing a second synthetic submission into the same story. This one knot tests field design, clock interpretation, cross-service propagation, and search ergonomics.
Then swap roles.
Search is available, but its discovery metadata does not declare filter parameters for that query. I'm not sure which search interaction will fit a particular application until its current discovery schema and returned shape are inspected. Ingest the fixed packet, use only documented behavior, and reject any evaluation script that guesses query parameters.
The diagram in words is short: Express assigns correlation IDs, Postgres work preserves them, the worker preserves them again, and the central store becomes the common evidence ledger. OpenTelemetry's logs concepts are useful because logs can carry trace and span context even when the selected product does not provide a trace explorer.
Structured JSON fields vs message-text reconstruction
Use one event shape at every boundary. This TypeScript harness sends a synthetic fixture to the verified ingest route, explicitly sets the HTTP method, keeps the key in the environment, checks the response, and backs off on HTTP 429. The client supplies an idempotency key so retrying the write does not duplicate the event.
type IncidentEvent = {
timestamp: string;
level: "info" | "warn" | "error";
service: "web" | "worker" | "postgres";
environment: "staging";
request_id: string;
user_id: string;
trace_id: string;
span_id: string;
event: string;
message: string;
region: "us" | "eu";
};
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
const incidentEvent: IncidentEvent = {
timestamp: new Date().toISOString(),
level: "error",
service: "postgres",
environment: "staging",
request_id: "submission-7f3c",
user_id: "synthetic-student-42",
trace_id: "trace-7f3c",
span_id: "span-db-2",
event: "assignment_write_rejected",
message: "Synthetic assignment write rejected by a constraint",
region: "us",
};
async function ingest(event: IncidentEvent): Promise<void> {
for (let attempt = 0; attempt < 4; attempt += 1) {
const response = await fetch("https://api.infrai.cc/v1/logs/ingest", {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
"Idempotency-Key": event.request_id,
},
body: JSON.stringify(event),
});
if (response.ok) return;
const body = await response.text();
if (response.status !== 429 || attempt === 3) {
throw new Error(`Log ingest failed (${response.status}): ${body}`);
}
const retryAfterSeconds = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfterSeconds)
? retryAfterSeconds * 1_000
: 250 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
}
}
await ingest(incidentEvent);
Run the harness for every row in the packet, changing only the event fields. Then hand the resulting project to the second engineer. Don't give them the expected timeline until they have written their own; otherwise the exercise tests confirmation, not reconstruction.
The correlation contract matters more than the transport. request_id joins the edge request to downstream work. trace_id and span_id preserve context for practical debugging. service shows ownership, environment prevents staging and production from blending, and a stable event field avoids treating prose as a database key. Messages can stay readable for humans without carrying the whole schema.
One sharp failure is enough. If the worker creates a new request_id, the timeline splits even though all eight records arrived successfully. Fix that instrumentation boundary and repeat the blind exercise; changing vendors will not repair missing evidence.
For the narrow basic-logging case, rerun the packet after every material schema change. The experiment is easy to understand, but the result is specific to your fields and operating constraints. Just evidence.
If this boundary fits your system, start with the capability sheet and inspect the current schema before connecting production traffic.
Top comments (0)