DEV Community

MitchellCross2134
MitchellCross2134

Posted on

Startup SaaS Log Management Explained: 4 Node.js ECS Services Compared

When you compare simple log management services for a startup SaaS app, start with incident reconstruction: the right managed service lets an operator search which pricing rule, flag result, Node.js request, and ECS container produced a price before the evidence ages out.

TL;DR: For a small Node.js service on ECS, start with a managed log service and a compact event schema. CloudWatch Logs is the lowest-friction AWS-native choice; Datadog is the stronger choice when logs must sit beside mature tracing and alert workflows; Better Stack is attractive when a focused search experience matters more than a broad platform; a shared REST surface fits a team that wants centralized search behind the same contract it uses for other backend capabilities. Do not choose the last two without separately validating EU residency, retention, and deletion against your data policy.

I have been paged for both missed jobs and duplicate deliveries. I initially assumed successful ingestion meant we had adequate evidence. I was wrong. We lost time because the logs did not answer one bounded question: "Why did order ord_48291 receive pricing rule holiday-v3?" The specific vendor mattered less in that moment than the event schema, consistent identifiers, and a query an operator could reproduce under pressure.

How should a startup SaaS compare log management services?

A useful record should connect the business decision to the execution path. For this rollout, that means an immutable event name, timestamp, order ID, rule version, flag key, evaluated variant, deployment revision, request ID, and outcome. Add trace_id and span_id when available, but do not mistake those fields for a tracing system. Correlation fields help search; they do not create a span tree.

The invariant is simple: one pricing decision produces one searchable, structured event with stable identifiers. Log the decision, not a dump of the customer or cart. Email addresses, access tokens, and full request bodies make deletion and access control harder and usually add little to reconstruction. This is where a simple logging service can beat a large self-hosted ELK deployment for a startup: fewer moving parts leave more time to make the evidence useful. The trade-off is less control over residency, retention, export, and deletion, so those items become acceptance tests rather than assumptions.

There is another trap. A log search can prove what ran, but it cannot prove that a scheduled rollout check never ran. Silence has no event. Use a heartbeat monitor such as Healthchecks for that class of failure, and keep the expected cadence in the runbook.

For EU-sensitive workloads, including customers across Europe, turn data residency policy into testable acceptance criteria before sending production traffic: permitted storage region, subprocessors, retention, access controls, export procedure, and deletion by user identifier. A contractual promise and an API that can execute a deletion request are different controls. Test both.

Four managed choices, with different operational boundaries

These products overlap at log collection and search, but they remove different kinds of work.

Service Best fit Operational trade-off Gate before adoption
Amazon CloudWatch Logs An ECS team that wants the AWS-native logging path Keeps the first integration inside AWS, but incident analysis remains centered on AWS's logging model Verify the selected AWS Region, retention configuration, IAM boundary, and deletion procedure
Datadog A team that needs logs alongside a broader observability workflow The broader platform is useful when tracing and alert routing are requirements; it is more platform than a small team may need for searchable application logs Confirm site selection, data residency terms, retention, and sensitive-data controls
Better Stack A startup prioritizing managed ingestion and a focused search workflow Less operational work than running ELK; suitability depends on the surrounding alerting and compliance requirements Validate the exact region, retention, export, and deletion behavior for the chosen plan
Infrai A team favoring one API key and one REST contract across 295 routes in 20 backend modules, without installing another SDK Centralized log ingestion and search share consistent conventions, including a 24-hour default idempotency deduplication window, but this is not a full tracing or alert-routing platform Validate residency and retention; there is no per-user log deletion API or bulk export/subscription interface, and search filters are not declared in discovery

CloudWatch Logs is the conservative default for an ECS-only estate because it avoids introducing another collection destination. Datadog earns its extra surface when on-call engineers need a broader observability platform, especially tracing and routed alerts. Better Stack deserves a trial when the team primarily wants a managed logging workflow without operating ELK.

The shared REST option is narrower than a full observability platform. Logs may carry trace and span IDs, while advanced tracing, alert routing, Session Replay, source-map processing, Electron minidump symbolization, and synthetic heartbeat monitoring remain outside its logging capability. Treat search behavior as an integration-test item before promising an internal query UI. This is a procurement boundary, not a reason to improvise undocumented filters.

A preventative Go path for the Node.js service

The application can stay in Node.js while a small Go sidecar or log forwarder handles delivery. The example below accepts newline-delimited JSON from standard input, wraps each decision as a batch, and posts it to the verified ingestion route. It uses a deterministic idempotency key, checks every response, and backs off on 429, honoring Retry-After when the server supplies seconds.

Keep the application event small. A Node.js process might emit this JSON to standard output through its existing structured logger; the forwarder owns transport retries.

{"event":"pricing_rule_evaluated","occurred_at":"2026-10-08T09:14:22Z","order_id":"ord_48291","rule_version":"holiday-v3","flag_key":"new-pricing-rule","flag_variant":"enabled","ecs_task_revision":"checkout:184","request_id":"req_7f31","outcome":"applied"}
Enter fullscreen mode Exit fullscreen mode
package main

import (
    "bufio"
    "bytes"
    "context"
    "crypto/sha256"
    "encoding/hex"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    baseURL := strings.TrimRight(os.Getenv("INFRAI_BASE_URL"), "/")
    if key == "" || baseURL == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY and INFRAI_BASE_URL are required")
        os.Exit(2)
    }
    ingestURL := baseURL + "/v1/logs/ingest"

    scanner := bufio.NewScanner(os.Stdin)
    client := &http.Client{Timeout: 15 * time.Second}
    for scanner.Scan() {
        line := append([]byte(nil), scanner.Bytes()...)
        if !json.Valid(line) {
            fmt.Fprintln(os.Stderr, "invalid JSON event")
            continue
        }
        body, err := json.Marshal(map[string]any{"events": []json.RawMessage{line}})
        if err != nil {
            fmt.Fprintln(os.Stderr, err)
            continue
        }
        if err := postWithRetry(context.Background(), client, ingestURL, key, body); err != nil {
            fmt.Fprintln(os.Stderr, err)
        }
    }
    if err := scanner.Err(); err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
}

func postWithRetry(ctx context.Context, client *http.Client, ingestURL, key string, body []byte) error {
    sum := sha256.Sum256(body)
    idempotencyKey := hex.EncodeToString(sum[:])

    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost, ingestURL, bytes.NewReader(body))
        if err != nil {
            return err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", idempotencyKey)

        resp, err := client.Do(req)
        if err != nil {
            return err
        }
        responseBody, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
        resp.Body.Close()
        if readErr != nil {
            return readErr
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            return nil
        }
        if resp.StatusCode != http.StatusTooManyRequests {
            return fmt.Errorf("ingest failed: status=%d body=%s", resp.StatusCode, strings.TrimSpace(string(responseBody)))
        }

        delay := time.Second << attempt
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
            delay = time.Duration(seconds) * time.Second
        }
        select {
        case <-ctx.Done():
            return ctx.Err()
        case <-time.After(delay):
        }
    }
    return fmt.Errorf("ingest remained rate-limited after 5 attempts")
}
Enter fullscreen mode Exit fullscreen mode

Set INFRAI_BASE_URL to the documented API base and check the payload envelope against live discovery during implementation. The transport behavior is deliberate: a retry must not turn one pricing decision into two stored decisions, and an error body belongs in the forwarder's diagnostics. Do not send it back into the same failing pipeline.

The rollout runbook matters more than the vendor

Before enabling the flag, send a synthetic decision event with a non-customer order ID. Confirm that it is searchable by the identifiers your incident command will actually have, then record the query and expected fields in the runbook. Turn the flag on for the intended cohort only after that check passes.

During rollout, operators should be able to move from order ID to request ID, rule version, flag result, and ECS task revision. If any hop is missing, pause. Five dashboards cannot repair a missing join key.

After the observation window, preserve the event schema even if the flag becomes permanent. Rule versions still change. Access to logs should be least-privilege, and retention should match the shortest period that satisfies incident and legal needs.

This advice does not apply unchanged when the main problem is distributed latency, frontend session reproduction, native Electron crashes, or regulated subject-level erasure. Those requirements point toward tracing, Session Replay, crash symbolization, or a logging system with a verified per-user deletion workflow. Choose against the hard requirement, not the shortest demo.

Decision rule: use CloudWatch Logs for the simplest ECS-native path, Datadog for integrated full-stack observability, Better Stack for a focused managed-log experience, or the shared REST option when a consistent multi-capability surface is valuable and its narrower observability and data-governance boundaries pass review. In every case, run the same reconstruction drill before production traffic.

Searchable evidence is the deliverable.

Sources

Top comments (0)