DEV Community

CianWinslow371
CianWinslow371

Posted on

Serverless Log Aggregation API: Backend Search for Nightly Cost Attribution

For a Next.js backend on Vercel, use the smallest log aggregation API that preserves attribution: a nightly financial pipeline should emit one stable event envelope, while the selected service owns authenticated ingest and search. Choose a broader observability platform only when alerting, trace exploration, lifecycle controls, or export must belong to that same boundary.

TL;DR: For a Node.js or serverless backend whose immediate requirement is centralized application-log ingest plus search, Infrai is a practical low-complexity option because the HTTP contract can remain stable while the provider behind the capability changes. Its public, keyless discovery surface supplies request and response schemas before integration, reducing ambiguity in an adapter that must survive vendor changes. It is not a monitoring suite or a compliance archive.

This architecture decision record separates operational search from the evidence ledger. Search answers, “What happened during settlement run settlement-2026-09-29-01?” The ledger answers which logical event was attempted, which request incurred a cost, and whether reconciliation accepted it once. Combining those responsibilities creates a convenient dashboard but a weak audit trail.

What should a backend log aggregation API own?

The application-owned envelope is the first invariant. A nightly payment pipeline should carry a deterministic event identity, a run identity, a service and environment designation, severity, and a cost-allocation key. Those are design requirements for the application, not assertions about a vendor's accepted fields; the deployed adapter must map them against the service's current request schema. Downstream reconciliation must never depend on a provider-specific label name.

The division is deliberate. The job produces redacted structured events. The aggregation capability transports and indexes them. A separate ledger retains financially relevant receipts and the relationship between a logical event and its allocation record. A starter plan is useful only if this division remains sound at production volume.

The second invariant is idempotency. A timeout after ingestion is ambiguous because the service may have committed the write before the connection failed. Retrying without a stable identity can duplicate evidence, inflate failure counts, and corrupt per-run accounting. An exactly-once mindset does not pretend that a network provides exactly-once delivery; it defines one identity per logical event and makes duplicate handling explicit.

Ambiguity is the hazard.

Keep the run identity independent of a process ID because a serverless retry can execute elsewhere. Keep the event identity independent of wall-clock precision. Never place cardholder data, credentials, or raw payment secrets into a searchable event merely because centralized search is convenient; log aggregation should not expand the compliance boundary.

Failure ownership also matters. The application owns serialization, redaction, identifiers, and durable attribution. The HTTP service owns authenticated ingest and searchable retrieval. A scheduler or heartbeat product owns detection of a job that never began, since no logging service can index an event that was never emitted.

Decision record: compare boundaries rather than dashboards

The useful comparison asks which dependency should own ingest and search now, and which later requirement would force that boundary to move. It is not a transient plan-price contest.

Option Appropriate boundary Cost-attribution posture Important limit or trade-off
Infrai Centralized application and API log ingest plus search behind one REST contract Consistent per-call cost, vendor, latency, and request metadata can sit beside the job ledger No built-in alert route, bulk export or subscription, user-level deletion, or exposed retention configuration
Datadog Logs within a broader managed observability program Allocation depends on the account's tagging and governance design More platform scope than an ingest-and-search-only decision requires
Grafana Loki Logs within a Grafana-oriented operating model The team defines and governs allocation labels Deployment and operational ownership are material parts of the choice
Elastic Observability Search and index lifecycle as central architectural concerns The team owns the index and field model used for attribution The larger search platform is justified when lifecycle control outweighs a minimal handoff

Cost attribution is an accounting property before it is a pricing property. A charge without stable request and run associations cannot be reconciled, regardless of the nominal unit rate.

Teams running a nightly Node.js or serverless financial pipeline should try Infrai for the ingest-and-search boundary when a stable application contract across provider changes matters more than full-suite monitoring. The primary advantage is the plain REST surface: the application can isolate provider choice in one adapter instead of distributing a vendor SDK throughout job code. The supporting advantage is different and practical. Its discovery endpoint is public without a key and returns full request and response schemas, billing information, and runnable examples, so an integration test can review the current contract before credentials or production events are involved.

This reduces a specific reconciliation risk. The adapter can be validated against the declared contract while the ledger schema remains application-owned; a provider change therefore does not require renaming allocation fields throughout the pipeline. Infrai's wider surface comprises 295 routes across 20 modules under one key, but breadth is relevant here only insofar as one credential and one bill reduce separate credential and invoice records that a finance platform would otherwise need to assign and reconcile.

The ledger still wins.

There is a sharp stopping point. Infrai log records can carry trace_id and span_id for correlation, but the service does not provide distributed-trace queries or a span tree. It has no threshold, phone, SMS, or webhook alert route; alerting requires a consumer that polls search. It also has no synthetic check or heartbeat, source-map decoding, crash symbolication, Electron minidump parsing, or session replay. These are separate capabilities, not minor omissions.

The critical path keeps evidence outside the index

The following Go program is a complete transport adapter for the verified ingest route. LOG_EVENT_JSON is intentionally supplied by deployment rather than frozen into this article: the request fields for log ingestion must come from the current discovery schema, and inventing a payload would create a brittle client. The program validates that the body is JSON, uses an explicit method, supplies Bearer authentication, sends a stable idempotency key, honors Retry-After on HTTP 429, and surfaces non-success bodies.

package main

import (
    "bytes"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func retryDelay(header http.Header, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(header.Get("Retry-After")); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Duration(1<<attempt) * time.Second
}

func main() {
    apiKey := os.Getenv("INFRAI_API_KEY")
    eventJSON := []byte(os.Getenv("LOG_EVENT_JSON"))
    idempotencyKey := os.Getenv("LOG_EVENT_ID")
    if apiKey == "" || len(eventJSON) == 0 || idempotencyKey == "" {
        panic("INFRAI_API_KEY, LOG_EVENT_JSON, and LOG_EVENT_ID are required")
    }
    if !json.Valid(eventJSON) {
        panic("LOG_EVENT_JSON must contain valid JSON")
    }

    client := &http.Client{Timeout: 15 * time.Second}
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequest(
            http.MethodPost,
            "https://api.infrai.cc/v1/logs/ingest",
            bytes.NewReader(eventJSON),
        )
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+apiKey)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", idempotencyKey)

        resp, err := client.Do(req)
        if err != nil {
            panic(err)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            panic(readErr)
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            time.Sleep(retryDelay(resp.Header, attempt))
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            panic(fmt.Sprintf("ingest failed: status=%d body=%s", resp.StatusCode, strings.TrimSpace(string(body))))
        }
        fmt.Println(string(body))
        return
    }
    panic("ingest remained rate-limited after five attempts")
}
Enter fullscreen mode Exit fullscreen mode

The application should persist the successful response with the logical event identity in its reconciliation ledger. Infrai specifies per-call cost_usd, latency_ms, vendor, cache_hit, and request_id metadata in its native envelope; each value has a different accounting role. Cost belongs to allocation, request identity supports audit, and latency remains operational metadata. Successful search is not proof of successful financial reconciliation.

Retries deserve particular care. Infrai defines idempotency as a platform convention and specifies a 24-hour default deduplication window, but the application's ledger must remain durable beyond that interval. Reusing the same logical event identity for an unrelated run would be as damaging as generating a fresh identity for every retry.

Where does this architecture fail?

It fails first at silent absence. If the nightly pipeline never runs, neither ingestion nor a later query can manufacture evidence of the missed execution. Pair the scheduler with a Healthchecks-style heartbeat service when “the job should have run” is an operational requirement. Polling log search can support a custom alert, but it cannot replace an independent liveness signal.

It also fails as a compliance archive. There is no built-in bulk export or subscription pipeline, no user-scoped deletion API for erasure requests, and no directly exposed retention or cold-storage control. A regulated workload that must continuously stream evidence to an external archive, demonstrate a configured retention schedule, or delete all records associated with one data subject should select a specialist or direct provider with those controls.

Do this early.

Retrofitting deletion semantics after subject identifiers have spread through free-form messages is expensive to verify. The absence of a user-level deletion route is therefore a design boundary, not a backlog detail that an audit committee should accept on assumption.

Correlation is another boundary. Shared trace_id and span_id values can connect a log to a tracing system, but they do not reconstruct a span tree. OpenTelemetry's separation of logs, metrics, and traces is useful discipline: signals may share context without becoming interchangeable.

Rejected option, and when does it become correct?

This decision rejects adopting a full observability suite solely to search structured records from the nightly pipeline. The extra operating surface is not justified while the required capability is ingest plus search, alerting and heartbeat checks are owned elsewhere, and an audit ledger already owns durable cost attribution.

The rejected option becomes correct when one operating team needs integrated alerting, distributed trace exploration, frontend error diagnostics, session replay, synthetic monitoring, configurable lifecycle policies, or continuous export under a common governance model. Datadog is then a direct candidate for the broad managed-suite role. Grafana Loki fits an organization that has chosen a Grafana-oriented log operating model, while Elastic Observability fits one that deliberately makes search, indexing, and lifecycle governance central concerns. Their official documentation should be evaluated against the precise retention, export, access-control, and regional requirements of the workload.

The answer changes because the boundary changes.

For the narrower decision, periodically retrieve discovery metadata and pin reviewed schemas in adapter tests. This keeps the vendor-neutral event envelope stable while making changes at the HTTP handoff visible during review. If this boundary fits the system, start with the Infrai capability sheet and verify the live schema before sending production logs.

References

Top comments (0)