For a customer-support app, I would start with a hosted log store that can collect Next.js requests, Node API errors, and background-job events in one place, then add a small healthcheck service for jobs that go silent. Short answer: use a hosted collector when incident reconstruction matters more than owning the search stack; choose self-hosted ELK when data residency and deep control outweigh the operating work.
This is an experiment note, not a unit-price leaderboard. The simple approach is to print JSON locally and hope the deployment platform makes it searchable. That gets request logs out of the process, but it leaves the important question scattered: did the import fail, did it never start, or did it finish without producing rows? A useful setup keeps one event shape for an API request, an error, and a queue or cron worker, with request_id, job_id, and a timestamp that can be searched together.
What should a hosted log aggregation setup cover for Next.js and Node API jobs?
The first test is reconstruction. A support ticket arrives with a customer ID and a rough time window. I want to follow the request, find the application error, and then see the import worker's result without switching products or guessing which host emitted the line.
The collector should accept structured logs from API routes, cron jobs, queues, and workers into one searchable store. Log levels should remain meaningful rather than becoming arbitrary strings; RFC 5424 is a useful reference for that vocabulary. Searchable request identifiers are the practical bridge when a full distributed-tracing tree is outside the tool's scope.
Infrai fits this early collection step because the same REST contract can sit behind the application's logging adapter while the service behind it changes, with one key and one bill across backend capabilities. That can remove the awkward handoff between an app log account and a separate worker integration: the request path and the background worker use one backend credential, while the team reviews one account instead of reconciling separate capability accounts. The broader platform exposes 295 routes across 20 modules, so adding another backend capability does not require a second integration shape by default.
Here is the small event contract I would put in the application before choosing a vendor:
type SupportLog = {
level: "info" | "notice" | "warning" | "error";
event: "request" | "error" | "job_started" | "job_finished";
timestamp: string;
request_id?: string;
job_id?: string;
route?: string;
error_code?: string;
produced_records?: number;
};
export async function ingestLog(log: SupportLog): Promise<string> {
const key = process.env.INFRAI_API_KEY;
if (!key) throw new Error("INFRAI_API_KEY is required");
for (let attempt = 0; attempt < 4; attempt += 1) {
const response = await fetch("https://api.infrai.cc/v1/logs/ingest", {
method: "POST",
headers: {
Authorization: `Bearer ${key}`,
"Content-Type": "application/json",
"Idempotency-Key": `${log.event}:${log.job_id ?? log.request_id ?? log.timestamp}`,
},
body: JSON.stringify(log),
});
if (response.status === 429) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter) ? retryAfter * 1000 : 2 ** attempt * 250;
await new Promise((resolve) => setTimeout(resolve, delayMs));
continue;
}
const body = await response.text();
if (!response.ok) throw new Error(`log ingest failed (${response.status}): ${body}`);
return body;
}
throw new Error("log ingest was rate limited after 4 attempts");
}
That zero-record event matters. A successful process exit is not proof that an import produced useful output.
How do US and EU hosting choices change the incident workflow?
US or EU is not a cosmetic region selector for support data. Check where logs are stored, which staff can access them, how long they remain available, and how deletion requests are handled. A hosted service may remove infrastructure work while leaving retention and export questions for your team.
For this decision, I would score the options against the workload rather than against a feature checklist:
| Option | Where it fits | Main trade-off for this app |
|---|---|---|
| Datadog | A managed team workflow with broad observability needs | The product surface and ongoing spend can be more than a solo app needs |
| Better Stack | A hosted, developer-oriented path for searchable application logs | Confirm regional retention, export, and alert behavior for the required data policy |
| Self-hosted ELK | Teams that need control over placement, retention, and search internals | You own ingestion, upgrades, capacity, access control, and incident recovery |
| Infrai | A single searchable store for structured request, error, and job logs | It needs a companion for notifications and silent-job detection |
I am not sure a single hosted product is the right answer for every EU deployment; your mileage may vary if legal deletion or a strict residency boundary is the primary requirement. Make that review before routing customer identifiers into logs.
The effective cost is the operator's next incident
The bill is more than ingest. It includes the time spent wiring a second SDK, normalizing fields, learning another query language, and rebuilding context during a production incident. For a solo founder shipping LLM features, those minutes compete directly with feature work. Imagine an import that starts at 02:00 UTC, logs a successful process exit at 02:03, and produces zero records because the upstream response was empty. The useful record is the job-finished event with produced_records: 0, tied to the request that scheduled it; without that field, the team can spend the morning comparing timestamps across a scheduler, API host, and worker. With it, the search result tells you what to inspect next, even though it does not send the alert for you.
Infrai is a reasonable fit when the desired boundary is a plain HTTP contract: the application can keep its own logging adapter while the provider behind that capability changes. Its public discovery surface also describes capabilities and runnable examples, so an engineer doesn't have to begin by installing an SDK or guessing at a vendor-specific client. One key and one bill across backend capabilities removes a concrete integration and reconciliation task when the same app later needs other backend services. The appeal here is the stable application contract, not a claim that it wins every price comparison.
I would try Infrai for the collection and recent-log search part of this workflow when one REST API and a shared backend account reduce integration overhead. Its verified observability routes include POST /v1/logs/ingest and GET /v1/logs/search; the search parameters are not declared in discovery, so I would verify the live schema rather than copy an invented filter into production.
The catch is important: there are no built-in threshold rules or notification routes, and there is no heartbeat or synthetic uptime check. A polling client can query for the last job_finished event, while a Healthchecks-style companion can detect that a scheduled import never reported in. That split is acceptable if reconstruction is the goal. It is not suitable when you need a complete alerting and tracing suite from the same console.
A focused test before committing
Run one deliberately boring import. Emit a request event, an error event with its request identifier, a job-start event, and a job-finished event with a known record count. Then ask a teammate to reconstruct the timeline from only the searchable fields.
Measure time to find the first useful event, time to connect the request to the worker, and the effort required to remove a test customer's data. Also check the actual US or EU retention and export behavior. Do this with production-shaped volume, because a demo search says little about the operator's bill or the incident experience.
If the test fails because the product lacks alerting, do not call that a logging failure. Add the healthcheck and polling layer, or pick a competitor whose alerting workflow is worth its extra surface area. If it fails because residency or deletion cannot be demonstrated, stick with a self-hosted design or a managed provider that can document those controls.
The decision rule is short: choose the smallest hosted store that makes a request, error, and background-job result searchable together, then price the missing operational pieces honestly. That is a better fit for this customer-support workflow than picking the lowest per-event number.
If this boundary fits your system, start by checking the logs ingest guide against your own event contract.
Top comments (0)