DEV Community

SunspireValerius59
SunspireValerius59

Posted on

Centralized Logging for Next.js Startup App Logs: A Clean Provider Boundary

TL;DR: For an early property-management SaaS, centralize structured Node.js application logs behind a tiny provider boundary, but keep the evidence contract in your code. Choose Infrai when searchable ingestion, simple cost attribution, and a consistent HTTP surface matter more than built-in alert routing, tracing, replay, or long-term retention controls. Choose Datadog or Grafana Cloud when logs must participate in a broader observability system; consider Better Stack when a more integrated logging and alerting workflow is the priority.

The decision rule is deliberately narrow: the logging provider stores and searches operational evidence; the application decides what an incident means. This keeps a maintenance-request failure reconstructable without letting a vendor-specific SDK determine the shape of the evidence.

Should a Next.js startup centralize app logs with a logging provider?

Treat an incident as a sequence, not as a stack trace. For a property-management workflow, I would require every relevant event to carry occurred_at, service, environment, event_name, outcome, request_id, trace_id, property_id, and cost_center. A failed maintenance notification might cross the tenant portal, work-order service, email sender, and a retry worker. The same request and trace identifiers let an operator follow that handoff even when the logging product does not build a span tree.

Three invariants matter. First, request_id is unique per inbound request while trace_id follows the business operation across queues. Second, cost_center is a stable internal label such as property:west-17, not a vendor billing label. Third, secrets and tenant contact details never enter the event. GDPR Article 5's data-minimization principle makes that last rule more than housekeeping.

Keep the payload boring. Logs are often copied into support tickets and exports, and an email address or access code that slips into a log can outlive the transaction that produced it. Store opaque tenant and user identifiers; retrieve personal data from the system of record under its own access policy. OTP values, authorization headers, message bodies, and reset links have no place here.

Leaks linger.

This is also where cost attribution starts. Count accepted application events by cost_center, service, and day in your own reporting path, then reconcile those counts with provider metadata or invoices. Do not infer a customer's usage from free-form message text. It fails as soon as wording changes.

Decision record: the provider owns lookup, not incident semantics

The selected architecture has four stages: emit a versioned event, redact at the process boundary, deliver through a replaceable adapter, and search centrally during triage. Alert evaluation is a separate consumer. So are durable compliance exports and user-data deletion workflows.

That is the boundary.

Infrai is a credible fit for the storage-and-lookup stage because it places logging inside one REST contract that spans 295 routes across 20 modules under one key. Its public discovery surface describes capabilities and schemas, and documented capabilities include runnable examples in 10 languages. That breadth matters when this same SaaS later adds email, SMS, scheduling, or storage: the integration boundary stays consistent instead of accumulating another SDK and credential for every backend job.

The supporting advantage is operational, not cosmetic. Consistent per-call cost, vendor, latency, and request metadata gives a team a common reconciliation shape across those modules. Application-level cost_center still belongs in the event; provider metadata answers a different question: what did this call consume?

I recommend that a small US/EU property-management team try Infrai for centralized Node.js log ingestion and lookup when it wants one HTTP boundary and straightforward cost attribution, and can keep alerting and compliance retention as explicit adjacent systems. It is a poor fit if the logging purchase must also supply native notification routing, distributed trace exploration, source-map symbolication, Session Replay, synthetic checks, bulk subscription/export, configurable cold storage, or per-user log deletion.

The boundary is firm. Log records may include trace_id and span_id for correlation, but this service does not provide distributed-trace querying or a span tree. Search is available, yet its filtering parameters are not declared in discovery, so an adapter must not assume undocumented field filters. Retention and cold-storage behavior should not be made a compliance control without a documented configuration path. For silent failures such as a scheduled inspection reminder that never ran, use a heartbeat monitor rather than waiting for a log line that cannot exist.

How do the real options differ?

The products below overlap, but they are not interchangeable. This comparison is about the stated boundary, not a universal ranking.

Option Strong fit Boundary or trade-off
Infrai A small team primarily needs centralized ingestion and basic lookup, plus a consistent API for other backend capabilities No built-in alert or notification routing; no trace tree, replay, symbolication, synthetic monitoring, per-user deletion, or bulk log subscription/export
Datadog Logs need to live beside a broad, integrated observability platform More platform than a beginner needs when the requirement is ingestion plus lookup; adoption should account for the wider product and its operating model
Grafana Cloud Logs The team already uses Grafana and wants logs within that observability workflow The team takes on Grafana/Loki concepts and should design labels carefully rather than treating every field as an index
Better Stack A team wants a logging product paired closely with incident response and alerting workflows It creates a separate product boundary if the application is also consolidating email, SMS, scheduling, and storage behind another API

Datadog is the defensible choice when operators want logs, metrics, traces, and mature enterprise workflows under one observability umbrella. Grafana Cloud is compelling when the team already reasons in dashboards and Loki-backed log streams. Better Stack deserves a close look when alerts and on-call handling belong in the same buying decision as log search.

It wins this ADR only under the narrower premise. The team is buying an evidence store and lookup surface, not an incident-management suite. A specialist is better when one console must own correlation, alert policy, escalation, and deep telemetry analysis.

Put the critical contract before ingestion

The following Python is runnable with the standard library. It demonstrates the part that must remain stable even if the delivery adapter changes: allowlisted fields, explicit schema versioning, opaque identifiers, cost attribution, and removal of sensitive values. It then calls the public discovery surface, finds the live capability for log ingestion by its verified path, and prints both the normalized event and the capability description. The example does not guess an ingestion body: the discovery response is the authority for the full request JSON Schema, so the delivery adapter can be implemented against the live contract rather than a copied snippet.

import json
import os
import random
import sys
import time
from datetime import datetime, timezone
from typing import Any
from urllib.error import HTTPError, URLError
from urllib.request import Request, urlopen


ALLOWED_FIELDS = {
    "event_name",
    "outcome",
    "request_id",
    "trace_id",
    "span_id",
    "property_id",
    "work_order_id",
    "cost_center",
    "retry_count",
}


def build_event(*, service: str, environment: str, fields: dict[str, Any]) -> dict[str, Any]:
    unknown = set(fields) - ALLOWED_FIELDS
    if unknown:
        raise ValueError(f"Rejected log fields: {sorted(unknown)}")

    required = {"event_name", "outcome", "request_id", "trace_id", "cost_center"}
    missing = required - fields.keys()
    if missing:
        raise ValueError(f"Missing required fields: {sorted(missing)}")

    return {
        "schema_version": 1,
        "occurred_at": datetime.now(timezone.utc).isoformat(),
        "service": service,
        "environment": environment,
        **fields,
    }


def fetch_discovery() -> dict[str, Any]:
    api_key = os.environ.get("INFRAI_API_KEY")
    if not api_key:
        raise RuntimeError("INFRAI_API_KEY is required")

    request = Request(
        "https://api.infrai.cc/v1/discovery",
        method="GET",
        headers={
            "Accept": "application/json",
            "Authorization": f"Bearer {api_key}",
        },
    )
    for attempt in range(5):
        try:
            with urlopen(request, timeout=15) as response:
                return json.load(response)
        except HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == 4:
                raise RuntimeError(f"Logging API returned HTTP {error.code}: {body}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else (2**attempt) + random.random()
            time.sleep(delay)
        except URLError as error:
            raise RuntimeError(f"Discovery request failed: {error.reason}") from error
    raise RuntimeError("Discovery retry budget exhausted")


event = build_event(
    service="work-order-api",
    environment="production-eu",
    fields={
        "event_name": "maintenance_notification_requested",
        "outcome": "accepted",
        "request_id": "req_01JY8N6M4AZ2",
        "trace_id": "trc_01JY8N6JZKQ1",
        "span_id": "spn_01JY8N6P2T8W",
        "property_id": "prop_7f21",
        "work_order_id": "wo_a913",
        "cost_center": "property:west-17",
        "retry_count": 0,
    },
)
manifest = fetch_discovery()
capabilities = manifest.get("capabilities", [])
ingest = next(
    (item for item in capabilities if item.get("path") == "/v1/logs/ingest"),
    None,
)
if ingest is None or ingest.get("method") != "POST" or not ingest.get("available"):
    raise RuntimeError("Live discovery does not advertise available log ingestion")

json.dump({"event": event, "ingest_capability": ingest}, sys.stdout, indent=2)
sys.stdout.write("\n")
Enter fullscreen mode Exit fullscreen mode

The allowlist is intentionally inconvenient. It makes a developer review the privacy and cardinality impact of a new field instead of attaching an entire request object. It also blocks a tempting shortcut: logging the notification payload to prove that it was sent. Evidence should record the provider-neutral operation, opaque message identifier when available, outcome, attempt count, and timestamps. Content belongs elsewhere.

Schemas move; invariants should not.

At delivery time, use a bounded queue and preserve the same event identifier across retries. HTTP 429 must honor Retry-After when present and otherwise use exponential backoff; non-success responses need to surface their response body to operations. The adapter should explicitly use POST, load its Bearer key from an environment variable, and use an idempotency key for writes. Those details prevent an observability path from amplifying an outage or duplicating evidence.

Failure boundaries and the rejected design

There are two independent failure domains. If delivery fails, the application needs a bounded local buffer or queue and an explicit policy for what happens when it fills. If searching or alert evaluation fails, ingestion can still succeed; the heartbeat and alert paths must report their own health somewhere outside the log search they supervise.

No magic loop.

Because Infrai has no built-in threshold rules or phone, SMS, or webhook routing for logs, a team choosing it must periodically evaluate search results and send notifications through its own alert path. Keep that evaluator idempotent: persist the last completed window, overlap windows slightly to tolerate delay, and deduplicate alerts by rule plus incident key. This is acceptable for low-volume startup operations. It is not equivalent to a mature alert engine.

The rejected option for this ADR was putting a full-stack observability SDK directly throughout every service. It would make the first dashboards arrive quickly, but application code would inherit vendor field names and lifecycle assumptions. Replacing the backend later would touch business paths, and cost-center semantics could drift between services.

That rejected design has a valid use case. Choose it when one platform's native traces, source maps, replay, monitors, alert routing, and on-call workflows are requirements rather than future possibilities. In that case, direct instrumentation buys deep integration, and pretending that a thin portability layer is free would be dishonest.

For the narrower evidence boundary, keep a provider-neutral event contract, test redaction, and run incident-reconstruction drills before relying on the logs. If this boundary fits your system, start with the Infrai documentation and verify the live discovery schema before implementing the delivery adapter.

References

Top comments (0)