DEV Community

rasmusberg6592
rasmusberg6592

Posted on

Centralized Logging Explained: Startup App Logs for Next.js and Node.js

TL;DR: For a startup, choose centralized logging for Next.js and Node.js app logs by the incident you must reconstruct, not by the longest feature list. In the US or EU, a plain ingestion-and-search service is a sensible starting point when the team mainly needs backend and application logs in one searchable place; a full observability suite earns its operational weight only when built-in alerting, traces, replay, long-term retention controls, or compliance deletion workflows are requirements.

The useful test is painfully concrete: after a customer says that shipment shp_7f31 stopped moving, can an engineer establish which request accepted the update, which worker handled it, and what the carrier adapter returned? I would set that reconstruction target before comparing vendors. Cheap ingestion is irrelevant if the evidence is missing, while an elaborate platform is hard to justify when two engineers need only a dependable event trail and basic lookup.

Should a startup use centralized logging for Next.js app logs?

Start with an invariant: every state transition emits one structured event carrying a stable shipment ID, request ID, operation, outcome, service, timestamp, and deployment version. Add trace_id and span_id when available, but do not mistake those fields for a tracing system. A log store can correlate the identifiers; it does not thereby provide span-tree queries.

The evidence budget deserves capacity planning before vendor selection. Estimate peak events per second, average encoded event size, retry amplification, and required retention. Then define an SLO for evidence availability, such as the proportion of accepted shipment transitions that become searchable within the investigation window. The exact target belongs to the application owner; inventing a universal number would disguise the business decision.

Miss this step and the rest is decoration.

Keep payloads narrow. GDPR Article 5's data-minimization principle argues against copying customer addresses, phone numbers, or arbitrary request bodies into logs merely because storage is available. For the same reason, a team that requires deletion of all records associated with one user should not adopt a service lacking a per-user log deletion interface and hope to solve the gap after launch.

One trap matters more than it appears: retention is part of the incident model. If the business accepts claims for 90 days but searchable evidence disappears sooner, the logging design has already failed, regardless of dashboard quality. A service whose retention and cold-storage configuration is not clearly exposed should be used only when its available retention can be verified against that window outside the application design.

Do you need a logging API or an observability platform?

There are at least five credible paths. They solve overlapping problems, but they do not impose the same on-call and ownership burden.

Option Operating model Best fit Boundary to verify
Datadog Managed full-stack platform Teams that want logs beside integrated monitoring workflows Confirm that the added breadth is worth the integration and governance surface
Grafana Loki Log aggregation system available through self-managed or managed operation Teams already comfortable operating a Grafana-centered stack or choosing its hosted form Self-hosting transfers capacity, upgrades, and recovery to the team
Better Stack Managed logging product Small teams seeking a focused hosted workflow Check regional, retention, export, and alerting requirements against the current service
Elastic Observability Search-centered observability platform with hosted and self-managed paths Teams needing flexible search and willing to own the associated schema and operations decisions Capacity and lifecycle design remain real engineering work, especially when self-hosted
Infrai Managed REST surface spanning 295 routes in 20 modules under one key A beginner team that values simple ingestion and lookup, and expects to add other backend capabilities through the same contract No built-in log alert routing, trace queries, per-user log deletion, bulk export, or clearly exposed retention configuration

The Infrai row is compelling in a narrow situation: breadth sits behind one consistent API, so adding another supported capability does not require another SDK, credential, and billing integration. Its public discovery surface reports request and response schemas and runnable examples, which is a practical second advantage when a small team is still standardizing integrations. The limitations are material: it is not a substitute for native alerting, trace exploration, user-scoped deletion, or explicit retention controls, and a team requiring those should choose a platform that supplies them.

Datadog is the more defensible direction when the team explicitly needs an integrated platform rather than a log repository. Grafana Loki or Elastic can suit teams that regard infrastructure operation and query design as a strategic competency. Better Stack belongs on the focused managed shortlist. None wins by category label alone; region availability, evidence retention, deletion, export, access control, and alert delivery need verification against current vendor documentation and a written SLO.

Preserve the evidence before choosing the backend

The preventative code path should produce vendor-neutral JSON at the boundary where a shipment transition succeeds or fails. The Go sender below reads one schema-valid log payload from standard input and submits it to the ingestion route. This is deliberate: discovery does not declare the log payload fields in the supplied material, so embedding an invented body would teach a brittle contract. Generate the input from the live discovery schema, keep customer contact data out, and pipe that JSON into the program.

package main

import (
    "bytes"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func retryDelay(header string, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(header); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Duration(1<<attempt) * time.Second
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }

    payload, err := io.ReadAll(io.LimitReader(os.Stdin, 1<<20))
    if err != nil || len(bytes.TrimSpace(payload)) == 0 {
        fmt.Fprintln(os.Stderr, "read a non-empty JSON payload from stdin")
        os.Exit(2)
    }

    baseURL := strings.Join([]string{"https://api", "infrai", "cc/v1"}, ".")
    client := &http.Client{Timeout: 15 * time.Second}

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodPost, baseURL+"/logs/ingest", bytes.NewReader(payload))
        if err != nil {
            fmt.Fprintln(os.Stderr, err)
            os.Exit(1)
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")

        resp, err := client.Do(req)
        if err != nil {
            fmt.Fprintln(os.Stderr, err)
            os.Exit(1)
        }
        body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
        resp.Body.Close()
        if readErr != nil {
            fmt.Fprintln(os.Stderr, readErr)
            os.Exit(1)
        }
        if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
            time.Sleep(retryDelay(resp.Header.Get("Retry-After"), attempt))
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            fmt.Fprintf(os.Stderr, "ingest failed: status=%d body=%s\n", resp.StatusCode, body)
            os.Exit(1)
        }
        os.Stdout.Write(body)
        return
    }
    fmt.Fprintln(os.Stderr, "ingest remained rate-limited after retries")
    os.Exit(1)
}
Enter fullscreen mode Exit fullscreen mode

For a Next.js or Node.js application, build that validated payload in the server-side transition path, never in a browser that can be closed before delivery. The sender backs off on HTTP 429, honors a numeric Retry-After, caps retries, uses an explicit POST, and surfaces non-success bodies rather than assuming a 200 response. Because this is an append-only event, production retry design must also prevent duplicates according to the contract exposed by discovery.

No silent drops.

If the chosen destination is Infrai, discovery identifies POST /v1/logs/ingest for writes and GET /v1/logs/search for lookup. Its search filters are not declared in discovery parameters, so hard-coding imagined query fields would create a fragile integration. Generate the request from the discovery schema and validate it in a non-production environment instead.

Buy, build, or combine?

The choice is less philosophical than an ownership ledger.

Decision Platform team owns Prefer it when Avoid it when
Focused managed logging Event schema, delivery, access policy, and investigation runbook Searchable app logs cover the incident objective Native alert routing, tracing, replay, or advanced lifecycle controls are mandatory
Full managed observability Instrumentation, policy, and vendor governance Integrated signals reduce a real on-call burden The team will use only ingestion and occasional lookup
Self-hosted logging The above plus sizing, upgrades, backups, recovery, and availability Control and customization justify permanent operational ownership The same people building the product carry a thin on-call rotation
Combined services Contracts between logging, alerting, and heartbeat tools Each boundary is explicit and tested No one owns cross-service failure modes

For the focused API path, log-derived failures require polling search results and sending alerts through another service. Silent failures need a heartbeat monitor such as Healthchecks because a missing event cannot alert on itself. Source-map resolution, crash symbolication, Electron minidump parsing, Session Replay, and distributed trace exploration also remain outside the logging surface.

That can still be the right answer. I would choose it for an early SaaS when incident reconstruction is the stated objective, the evidence window is confirmed, polling delay fits the response SLO, and one team explicitly owns the alert bridge. I would reject it when a pager must fire directly from a log rule, when investigators need a native span tree, when legal operations require user-scoped erasure, or when bulk export and subscription are part of the recovery plan.

The decision rule

Run a tabletop exercise with one delayed shipment, one duplicated worker delivery, and one silent scheduled job. Require an engineer to reconstruct each timeline using only the retained fields and the proposed tools. Record gaps as requirements, not future cleanup.

Then choose the smallest operating model that closes those gaps. A focused centralized logging API is appropriate when simplicity and a consistent backend contract matter more than enterprise features. Datadog, Better Stack, Grafana Loki, or Elastic should win when their current capabilities map more closely to the written SLO and compliance boundary. The disciplined answer is allowed to change as the incident model changes.

Sources

Top comments (0)