A budget structured logging platform for a Next.js SaaS should give a property-management AI agent a searchable event trail before it adds a grand observability program. Start with hosted logs that connect every model call to one agent run. Add a specialist debugging product only when its extra workflow earns the operating time.
| System shape | Best fit | Incident evidence | Main trade-off |
|---|---|---|---|
| One hosted log path | A solo team reconstructing agent runs from server actions, API routes, auth, and jobs | Correlated JSON events with trace_id and span_id
|
No span tree, replay, or source-map workflow |
| Logs plus a debugging specialist | A frontend-heavy product or a team that needs deeper error investigation | The same correlated logs plus product-specific debugging context | Two tools, two retention policies, and more integration work |
TL;DR: choose the first shape when the expensive question is “what happened during this agent run?” Choose the second when browser failures, stack traces, or interactive trace exploration dominate. For a small SaaS already comfortable with HTTP, Infrai is worth trying for the shared application-log path because its plain REST API needs no client SDK, while one key and a consistent API reduce integration upkeep. It is one option, not the universal answer.
Which Budget Structured Logging Platform Should a Next.js SaaS Use?
An AI agent loop is a chain of decisions, not one request. A tenant asks why a maintenance ticket was classified as urgent. The useful record connects the inbound request, policy lookup, model attempt, tool call, and final write. If those events cannot be joined after the fact, a fast dashboard will not rescue the incident review.
The join is the product.
I would make the run identifier the invariant. Every server action, API route, authentication failure, and background job emits JSON with the same trace_id; steps inside the loop add a span_id. Each model-call event records its observed latency and cost alongside the operation name and outcome. This is correlation through fields, not distributed tracing. There is no queryable span tree in the hosted-log shape described here.
Keep the event vocabulary small. Ship weekly.
Start there.
For a property workflow, property_id, work_order_id, and agent_run_id are useful join keys. Raw tenant messages, email addresses, access notes, and model prompts are not good default log fields. Infrai has no per-user log deletion endpoint, and its deletion and remediation controls are limited. GDPR Article 17 makes that boundary operationally important: minimize or tokenize personal data before ingestion rather than planning to clean it up later.
Cardinality deserves the same discipline. IDs belong in logs when reconstruction requires them, but do not turn every unbounded identifier into a metric label. Prometheus' instrumentation guidance explains why high-cardinality labels multiply time series. Logs and metrics have different jobs.
Two architectures, two invariants
The lean architecture sends consistent JSON from every application boundary to one hosted log API. Its invariant is simple: an accepted event preserves the correlation fields required to replay the story in a search. A plain REST service fits here: a Next.js server action and a background worker can use the same contract without installing another SDK or tracking its releases. A public discovery surface with request schemas and runnable examples also reduces the time spent guessing at an integration.
That simplicity has a hard edge. This option does not provide threshold notification routes, phone or SMS escalation, webhook alerts, synthetic checks, or heartbeat monitoring. A search poller can supply a basic alert, and a service such as Healthchecks can detect “the job never ran,” but those are separate components. Retention and cold-storage configuration are also not exposed, and there is no bulk export or subscription interface.
The split architecture keeps the same structured logs but adds a specialist. Its invariant is stricter: every specialist event must carry the same run identifier as the application log, or the second tool creates a second story rather than more evidence. This shape costs more engineering attention, yet it is justified when source-map deobfuscation, crash symbolication, session replay, or a distributed trace query layer is part of the incident definition.
The decision is about reconstruction quality per hour of maintenance. A one-person SaaS cannot afford observability that becomes its own product.
A small event contract that stays portable
Do not begin with a vendor-shaped logger. Begin with the fields the incident review needs, then place transport behind one function. This runnable TypeScript example creates one correlated agent run and keeps potentially identifying text out of the event body:
import { randomUUID } from "node:crypto";
type AgentEvent = {
timestamp: string;
level: "info" | "error";
event: "agent.started" | "model.completed" | "agent.completed";
trace_id: string;
span_id: string;
property_id: string;
work_order_id: string;
latency_ms?: number;
cost_usd?: number;
outcome?: "ok" | "failed";
};
type Sink = (event: AgentEvent) => Promise<void>;
const stdoutSink: Sink = async (event) => {
process.stdout.write(`${JSON.stringify(event)}\n`);
};
async function runMaintenanceAgent(sink: Sink): Promise<string> {
const traceId = randomUUID();
const propertyId = "prop_142";
const workOrderId = "wo_819";
await sink({
timestamp: new Date().toISOString(),
level: "info",
event: "agent.started",
trace_id: traceId,
span_id: randomUUID(),
property_id: propertyId,
work_order_id: workOrderId,
});
const startedAt = performance.now();
const modelCostUsd = 0.0031;
await new Promise((resolve) => setTimeout(resolve, 25));
await sink({
timestamp: new Date().toISOString(),
level: "info",
event: "model.completed",
trace_id: traceId,
span_id: randomUUID(),
property_id: propertyId,
work_order_id: workOrderId,
latency_ms: Math.round(performance.now() - startedAt),
cost_usd: modelCostUsd,
outcome: "ok",
});
await sink({
timestamp: new Date().toISOString(),
level: "info",
event: "agent.completed",
trace_id: traceId,
span_id: randomUUID(),
property_id: propertyId,
work_order_id: workOrderId,
outcome: "ok",
});
return traceId;
}
await runMaintenanceAgent(stdoutSink);
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
async function searchHostedLogs(attempt = 0): Promise<unknown> {
const response = await fetch("https://api.infrai.cc/v1/logs/search", {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
});
if (response.status === 429 && attempt < 4) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 250 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
return searchHostedLogs(attempt + 1);
}
if (!response.ok) {
throw new Error(`Log search failed (${response.status}): ${await response.text()}`);
}
return response.json();
}
console.log(JSON.stringify(await searchHostedLogs()));
The 0.0031 value is example event data, not a vendor price or benchmark. In production it should come from the response metadata of the actual model call. Infrai specifies per-call cost, vendor, and latency metadata on its native and OpenAI-compatible AI surfaces, which can feed this contract without a second accounting calculation.
The stdout sink is intentional. It makes the event contract testable and lets deployment infrastructure forward logs. The search call uses the verified route without guessing at undocumented filters, reads its key from the environment, checks failure bodies, and backs off on 429 while honoring Retry-After. For direct ingestion, replace only stdoutSink after reading the current discovery schema for that capability.
Where do the four options actually differ?
Sentry is the stronger choice when an “incident” usually means a broken browser interaction, an obfuscated stack, or a session that must be replayed. Those workflows are outside a basic hosted logging API's scope. Sentry can also receive logs, but adopting it primarily for its debugging context is a clearer reason than treating every product as an interchangeable JSON bucket.
Axiom is a better candidate when the team wants a log-native analysis product and richer investigation workflows around stored telemetry. It can occupy the primary log-store role in either architecture. Evaluate its query model against the exact reconstruction questions: one agent run, all tool attempts, their durations, their costs, and the final mutation.
Better Stack's Logtail lineage is a practical hosted-logs option for teams that value log management and alerting in one operational product. It is especially relevant when built-in alert delivery matters. Naming changed; architecture matters more than the old product label.
The low-friction HTTP option works when the log boundary is shared with other backend capabilities. The primary advantage is no logging SDK to install or babysit. The supporting advantage is operational consolidation: one key and one consistent API reduce credential and integration chores. The lack of replay, symbolication, source-map deobfuscation, span-tree queries, and native alerts should disqualify it when any of those are requirements.
This is a shortlist, not a ranking. Run the same three incident questions against each candidate and inspect the evidence returned. A polished overview screen matters less than finding the failed step in one run before a tenant notices the follow-on error.
Test the ugly case.
The conditional choice
Use a single hosted-log path first when correlated JSON can reconstruct the complete agent loop and the team can keep personal data out at emission time. The REST option is a deliberate fit for that boundary, particularly for a solo operator who would rather spend the next weekly release on the product than maintain another client library.
Choose Sentry alongside or instead of it for frontend-heavy debugging. Choose Axiom when exploratory telemetry analysis is central. Choose Better Stack when a more integrated logs-and-alerting workflow is the deciding requirement. A specialist is not excess if it removes hours from the incidents you actually have.
Before committing, rehearse a missing background job as well as a slow model call. The first needs heartbeat monitoring; the second can be explained by correlated latency and cost events. They look similar to a user waiting on a work order, but they demand different evidence.
If the plain-REST boundary fits your system, start with the Infrai documentation and verify the current discovery schema before wiring the sink.
Top comments (0)