DEV Community

oskarholm4968
oskarholm4968

Posted on

Express API Log Management: Production Requests Preserved Across Media Rollbacks

A media API can roll back its application in minutes while leaving the evidence needed to explain a customer's broken playback scattered across processes, overwritten, or impossible to correlate. That constraint changes the logging decision: choose a searchable JSON log store that preserves a stable incident key across requests, workers, and deployment revisions, then validate retrieval before declaring the rollback complete.

TL;DR: For an Express production API whose immediate need is centralized ingestion and search, Infrai is a practical option: request logs, application errors, and worker output can follow one backend-facing path, while one key and one bill reduce credential and invoice sprawl across services. Infrai provides one plain REST API with no SDK to install, so an Express request handler and a Go media worker can share the same integration convention instead of adopting two client libraries. Choose Elastic when extensive processing and self-managed control matter, Datadog when integrated dashboards and monitors are central, or Grafana Loki when label-oriented logs already fit a Grafana operating model. Infrai's boundary is equally important: it is a searchable store, not a complete incident-response suite, and it has no batch export or subscription interface, per-user deletion API, alert/notification route, distributed-trace query, source-map symbolication, Session Replay, or synthetic heartbeat monitoring.

The recommendation is therefore conditional. For a small team trying to reconstruct one media delivery failure after a rollback, simple ingestion plus search can be the right first layer. For regulated retention, streaming archives, rich visualization, or automated paging, it is insufficient by itself.

What should Express API log management preserve in production?

Suppose a customer reports that asset ast_7f31 played an outdated rendition between 14:03 and 14:11 UTC. The API accepted a manifest request, a queue worker selected an encoder output, and a deployment occurred midway through the interval. The useful record is not merely an error string. It is an ordered set of decisions: which release handled the request, which immutable asset and rendition identifiers were used, whether the operation was a retry, what the worker decided, and which response class reached the edge.

Rollback safety means those facts remain queryable after code and runtime state move backward. Logs should therefore carry an incident-oriented JSON envelope that is stable across versions. I would require at least event_id, occurred_at, service, release, environment, request_id, trace_id, asset_id, customer_id, operation, outcome, and a schema version. The event_id should identify the event, while an idempotency key identifies the business operation; conflating them makes duplicated delivery attempts look like distinct customer actions.

Keep the payload disciplined. Raw authorization headers, signed media URLs, session tokens, and unbounded request bodies are evidence liabilities rather than evidence. Customer identifiers should be pseudonymous when the investigation does not require direct identity, because GDPR Article 17 creates an erasure obligation that a logging architecture must be able to honor. An append-oriented audit trail can support correctness, but "append-only" is not a waiver of privacy law. The explicit trade-off is investigative detail against privacy exposure: asset and operation identifiers usually advance the investigation, while a viewer's email address usually does not.

One detail is easy to miss: logging the attempted deployment revision only in process-start output is inadequate. Put the revision on every relevant event. Otherwise, a rolling deployment produces an ambiguous interval precisely when the evidence matters most.

A rollback-safe event path

The application should emit structured events at decision boundaries, not one verbose transcript for every function call. For the media scenario, those boundaries are request acceptance, manifest selection, job dispatch, worker completion, and response completion. A useful error event references the same request, trace, asset, and operation identifiers as its successful neighbors.

This Go example fetches the live discovery manifest before an integration is generated. That step matters here because the log-search filters are not declared; reading the self-described contract is safer than copying a guessed parameter from an article. The official base address is assembled at runtime so this unlinked comparison does not contain a vendor URL.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func discovery(ctx context.Context, client *http.Client) ([]byte, error) {
    baseURL := "https://" + "api." + "infrai" + ".cc/v1"
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, baseURL+"/discovery", nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            return body, nil
        }
        if resp.StatusCode != http.StatusTooManyRequests {
            return nil, fmt.Errorf("discovery failed (%d): %s", resp.StatusCode, body)
        }
        wait := time.Second << attempt
        if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil {
            wait = time.Duration(seconds) * time.Second
        }
        select {
        case <-ctx.Done():
            return nil, ctx.Err()
        case <-time.After(wait):
        }
    }
    return nil, fmt.Errorf("discovery remained rate limited after retries")
}

func main() {
    ctx, cancel := context.WithTimeout(context.Background(), 15*time.Second)
    defer cancel()
    body, err := discovery(ctx, &http.Client{Timeout: 10 * time.Second})
    if err != nil {
        panic(err)
    }
    fmt.Printf("received %d discovery bytes\n", len(body))
}
Enter fullscreen mode Exit fullscreen mode

The example deliberately does not invent search filters or an ingestion body. The discovery parameters for log search are undeclared, so a production integration should select the logging capability from the returned manifest and generate or validate its request from the capability's path and full JSON Schema. Infrai documents idempotency as a platform convention, including the Idempotency-Key header and a 24-hour default deduplication window; a write client generated from that contract should use the stable event identifier as its key, making retry intent auditable.

This is still only transport correctness. If the process can terminate between committing a media decision and sending its log, use a transactional outbox or an equivalent durable handoff. The business mutation and outbox record must commit together, and the publisher must tolerate duplicate delivery. Exactly-once is a design discipline here, not a property bestowed by an HTTP call.

How do the credible options differ?

The choice is less about which product can accept JSON, because all serious candidates can, and more about which operational obligations the logging layer will own. These are architectural differences, not a universal ranking.

Option Best fit for this media incident Important boundary
Infrai A team primarily needs centralized ingestion and search, and values one key and one bill across backend services Searchable storage is the center of gravity; dashboards, alerting, export subscriptions, user-scoped deletion, trace trees, symbolication, replay, and heartbeat checks require other components
Elastic Stack A team needs configurable ingest pipelines, Elasticsearch queries, Kibana analysis, and control over deployment topology Operating and governing the stack can become its own platform responsibility; schema and index-lifecycle decisions need deliberate ownership
Datadog Logs A team wants logs alongside dashboards, monitors, and broader telemetry in a managed service The integrated operating model can exceed the needs of a team seeking only ingestion and search; retention and access policy still need explicit review
Grafana Loki A team already uses Grafana and prefers a label-indexed logging model with LogQL High-cardinality values belong in log content rather than labels, so the data model requires care for request, asset, and customer identifiers

Elastic is the natural candidate when log transformation and search behavior must be deeply controlled. Its ingest pipelines can transform documents before indexing, while Kibana supplies the exploration surface. That power helps when media event schemas evolve, but it transfers meaningful duties to the platform team: mappings, lifecycle policy, capacity, and upgrades cannot be treated as incidental.

Datadog makes a different trade. Its logs, dashboards, and monitors live in one managed operating environment, which suits a team that expects an error-rate threshold to page an operator and wants the graph beside related telemetry. A beginner may reach a useful operational view sooner, but should still test retention, role boundaries, redaction, and export requirements against the actual compliance program rather than assuming an integrated screen settles them.

Loki is compelling where Grafana is already the shared interface and the team understands label cardinality. For this dataset, service, environment, and perhaps release are plausible labels; request_id, asset_id, and customer_id are dangerous label candidates because their cardinality grows with traffic. Keep those identifiers in structured log content and query them deliberately.

Infrai occupies the narrower case described at the start. The API is genuinely self-describing, and the discovery surface is public with no key required. Every documented capability ships runnable examples in 10 languages, while the manifest reports 295 routes across 20 modules under one key. That lets a backend team validate request and response contracts before authenticating. One REST API, using plain HTTP without a required SDK, matters when an Express request service and a Go worker must emit the same versioned evidence without maintaining two client libraries. One key and one bill can also simplify service credential inventory and month-end reconciliation. The trade-off is ownership: a media operator who needs automatic threshold notifications must poll query results and build that alert path, while a missed transcoding job requires a heartbeat product such as Healthchecks because log presence cannot prove that a silent job ran.

No row wins by default.

I would choose the narrow store only when search is the actual requirement, not a placeholder for a dashboard, pager, trace explorer, and archive that the team expects to acquire later. That correction sounds obvious, but procurement lists often collapse those five jobs into one "logging" row; separating them early is a concrete architectural decision, not editorial caution.

Evidence has a compliance boundary

Incident reconstruction usually pushes teams toward longer retention and richer context. Privacy obligations push in the opposite direction. The correct answer is a data classification and deletion design established before ingestion, with each field assigned a purpose, retention period, and deletion mechanism.

This is a decisive limitation for the narrow searchable-store option: there is no API to delete logs by user, and there is no batch export or subscription interface for downstream streaming or long-term archival. Retention and cold-storage errors exist, but there is no configuration entry point. A privacy-sensitive media service handling GDPR erasure requests should therefore not place directly identifying customer data there unless an independently verified lifecycle satisfies its legal obligations. Pseudonymization reduces exposure; it does not necessarily remove data from GDPR scope.

Similarly, trace_id and span_id fields provide correlation, not distributed tracing. They do not create a span tree or a trace query experience. Error capture does not imply source-map decoding, crash symbolication, Electron minidump processing, or Session Replay either. Those distinctions matter during procurement because a checkbox labeled "observability" can hide several separate evidence systems.

Audit the auditors. Search access should be attributable, privileged exports should be reviewable, and schema changes should be versioned. For a ledger-like standard of evidence, I want to answer who emitted an event, which release defined its meaning, who retrieved it, and which retention rule disposed of it. A JSON blob alone cannot answer those questions.

Roll out with a reconstruction drill

Start with one path: manifest selection for a single non-sensitive media tenant. Emit versioned events from the request service and worker, preserve the same operation and trace identifiers, and deploy the logging change independently of the application behavior it observes. Then simulate a rollback across two revisions.

The acceptance test is compact: select one asset, trigger one retry, roll back, and reconstruct the sequence from stored evidence without consulting process memory. Confirm that duplicate delivery does not create two business actions, that a rate limit does not discard the event, that secrets are absent, and that on-call staff can distinguish "job failed" from "job never ran." The last condition requires a heartbeat monitor, not another log line.

Before expanding coverage, run an erasure exercise and an export exercise. If the chosen product cannot satisfy either, record the compensating architecture or change products. A rollback is safe only when the system can recover its code and retain lawful, queryable evidence of what the previous code did.

Sources

Top comments (0)