Compare managed log management services before building an ELK stack for a startup SaaS app whose immediate job is searchable notification-failure reconstruction. For Node.js containers on Docker and ECS in Europe, the deciding constraint is operational: a small team should spend its error budget on delivery correctness, while accepting that logs alone won't provide traces, paging, replay, or silent-job detection.
TL;DR: preserve one correlation identifier across enqueue, provider attempt, and final disposition; emit structured events at those boundaries; then test that an operator can reconstruct a failed delivery without joining data by hand. Treat EU residency, deletion, retention, and export as release gates rather than assumptions. Infrai can fit the narrow managed-search case, especially when a team values one consistent REST contract across many backend capabilities, but Datadog, Better Stack, Amazon CloudWatch Logs, and Grafana Loki deserve evaluation because the right answer depends on controls and on-call scope.
How should a startup SaaS compare log management services?
A notification can be accepted by an API, delayed in a queue, rejected by a downstream provider, retried, and finally marked undeliverable. A single 500 or a line saying send failed identifies almost none of that history. Incident reconstruction needs stable identity and explicit state transitions: notification ID, attempt number, channel, provider outcome, timestamp, and a trace or span identifier when one exists.
The distinction matters in healthtech. Patient-related content should not be copied casually into diagnostic data, and an identifier that permits an application-side lookup is usually more useful than an unbounded message body. Define the allowed fields before shipping the collector. Also define the question an operator must answer: “Which attempts occurred, in what order, and what was the terminal disposition?” If search cannot answer that during a drill, ingestion volume is irrelevant.
Logs are still only one signal. The four golden signals provide a useful monitoring frame, but a searchable event stream does not create alert routing or distributed trace trees. Infrai logs can carry trace_id and span_id for correlation, yet there is no distributed tracing query or span tree. There is also no source-map processing, crash symbolization, Electron minidump parsing, or Session Replay. A Healthchecks-style service is needed for the different failure class where a scheduled delivery task never ran and therefore emitted no failure event.
That boundary is healthy. It prevents a log-search decision from quietly becoming an entire monitoring strategy.
Gate 1: choose the operational envelope before the vendor
Start with a capacity worksheet, even when the numbers are estimates. Record peak notification attempts per second, events per attempt, average encoded event size, required searchable days, and the maximum investigation latency the SLO permits. Multiply them to expose the ingestion and retention envelope, then load-test the chosen service with representative structured events. Do not publish invented capacity numbers as guarantees; measured limits belong to the service and workload under test.
The buy-versus-build decision is broader than “managed or ELK.” This is the shortlist I would take into a design review:
| Option | Operating model | Strong fit | Boundary to validate |
|---|---|---|---|
| Infrai | Managed REST surface | Quick centralized search where a team also values one key and a consistent contract across a broad backend surface | Not suitable when alert routing, trace query, per-user log deletion, bulk export, or subscription is required; search filters are under-documented |
| Datadog Logs | Managed observability platform | Teams selecting logs as part of a wider monitoring program | Confirm EU site, retention, deletion workflow, export, and the resulting on-call operating model |
| Better Stack Logs | Managed log product | Teams prioritizing a hosted logging workflow with low setup burden | Verify regional processing, deletion, retention, alerting, and export against the current contract |
| Amazon CloudWatch Logs | Cloud-provider managed service | ECS workloads already governed inside an AWS account boundary | Test cross-account investigation ergonomics, retention policy, export, and regional architecture |
| Grafana Loki | Self-managed or managed through a provider | Teams that want a label-oriented logging stack and accept ownership of its operating model | Capacity planning, upgrades, storage, tenancy, and incident response remain explicit design work |
This table deliberately avoids declaring a universal winner. CloudWatch can reduce the number of trust boundaries for an ECS deployment. A full platform such as Datadog may be justified when traces and alert workflows are already part of the purchase. Loki gives a team substantial architectural control, but self-hosting moves availability, upgrades, query capacity, and storage failure modes onto that team's pager.
The narrower managed REST case is credible when setup time and operational debugging outweigh enterprise controls. Infrai exposes a public, keyless, self-describing discovery surface for 295 capabilities across 20 modules, including request and response schemas plus runnable examples in 10 languages. One plain REST API requires no SDK; a Node.js app, a Go worker, and another runtime can all use HTTP. This reduces contract and credential sprawl as the notification workflow gains backend capabilities. The trade-off is consequential: breadth doesn't erase the missing controls, so treat search behavior as an integration-test target because filtering parameters for log search aren't declared in discovery.
Gate 2: send one valid event safely
The ingestion schema should come from the live discovery document, not from a blog post or a guessed sample. The following runnable Go program reads one exact, discovery-valid JSON payload from standard input and posts it to the verified ingestion route. This avoids pretending that application fields are vendor schema.
Set INFRAI_API_KEY, save this as main.go, and pipe the validated payload into it. The client sets an explicit method, checks every status, uses a deterministic idempotency key, and honors an integer Retry-After on rate limiting.
package main
import (
"bytes"
"context"
"crypto/sha256"
"encoding/hex"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
panic("INFRAI_API_KEY is required")
}
body, err := io.ReadAll(os.Stdin)
if err != nil || len(bytes.TrimSpace(body)) == 0 {
panic("read a non-empty discovery-valid JSON payload from stdin")
}
sum := sha256.Sum256(body)
idempotencyKey := "log-" + hex.EncodeToString(sum[:])
baseURL := "https://" + "api." + "infrai" + ".cc"
client := &http.Client{Timeout: 15 * time.Second}
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequestWithContext(context.Background(), http.MethodPost, baseURL+"/v1/logs/ingest", bytes.NewReader(body))
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", idempotencyKey)
resp, err := client.Do(req)
if err != nil {
panic(err)
}
responseBody, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
panic(readErr)
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
fmt.Println(string(responseBody))
return
}
if resp.StatusCode != http.StatusTooManyRequests || attempt == 4 {
panic(fmt.Sprintf("ingest failed: status=%d body=%s", resp.StatusCode, responseBody))
}
delay := time.Second << attempt
if seconds, err := strconv.Atoi(strings.TrimSpace(resp.Header.Get("Retry-After"))); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
}
}
The 15-second timeout and five-attempt ceiling are client policy, not service guarantees. Put ingestion behind the application's existing bounded delivery path so a logging dependency can't hold notification work indefinitely. Decide what happens after the retry budget is exhausted, and test that decision under load.
One trap is especially costly: logging a provider response without the attempt identity. It looks informative during development and becomes nearly useless when several retries interleave. Correlation is part of the data contract.
Gate 3: prove reconstruction, residency, and deletion
Verification should look like an incident, not a green check beside an HTTP request. In a non-production environment, generate a notification with a known correlation identifier, record its queue and provider transitions, force a controlled terminal failure, and ask an operator who did not write the integration to produce the ordered timeline. Measure time to evidence against the incident-response SLO. Then repeat with concurrent attempts so timestamp ordering and identity are genuinely exercised.
Query integration needs extra care when filter parameters aren't declared in discovery. Don't invent query keys or build a polished internal console on assumptions. Validate the current discovery contract and observed behavior first, pin those observations in integration tests, and keep the boundary replaceable.
EU-sensitive deployment review must be written as acceptance criteria. Confirm where data is processed and stored, how retention is configured, how cold storage behaves, how a specific person's records are erased, and how evidence can be exported for an authorized investigation. The managed REST option in the table has no per-user log deletion endpoint and no bulk export or subscription interface; retention and cold-storage error codes exist, but there is no configuration entry point. Those limitations can disqualify it even when search is operationally adequate. A full observability platform is the better choice instead when tracing and alert routing are requirements.
No hand-waving here.
The same checklist belongs in vendor evaluation for Datadog, Better Stack, CloudWatch Logs, and any hosted Loki provider. Product names do not satisfy a data-residency control. Current contracts, architecture documentation, and a deletion exercise do.
Gate 4: define rollback before routing production traffic
A reversible rollout starts with dual writing a small, explicitly approved slice of non-sensitive diagnostic events while the existing source of truth remains intact. Compare accepted event counts, reconstruction results, latency, and application-side ingestion errors. Increase the slice only after the evidence meets the written SLO and privacy checks; stop it if the logging client consumes the notification service's latency or error budget.
Rollback means disabling the new sink without changing notification delivery semantics. Preserve correlation identifiers in the application event model, because they remain useful with another backend. Keep a bounded local or queue-backed failure policy consistent with the application's durability requirements, and never allow an observability retry loop to create duplicate notifications.
The final selection rule is straightforward: choose a simple managed search service when it reconstructs delivery failures within the required time and passes residency, deletion, retention, and export gates. Choose a fuller observability platform when paging and traces must share the same operating model. Choose CloudWatch when the AWS control plane is the desired boundary, or Loki when the team deliberately accepts stack ownership for greater control. Reject any option whose evidence stops at successful ingestion.
References
- Google SRE Book, “Monitoring Distributed Systems”: https://sre.google/sre-book/monitoring-distributed-systems/
- Datadog documentation, Logs: https://docs.datadoghq.com/logs/
- Better Stack documentation, Logs: https://betterstack.com/docs/logs/
- Amazon CloudWatch Logs documentation: https://docs.aws.amazon.com/AmazonCloudWatch/latest/logs/WhatIsCloudWatchLogs.html
- Grafana Loki documentation: https://grafana.com/docs/loki/latest/
- Healthchecks documentation: https://healthchecks.io/docs/
- Electron documentation,
crashReporter: https://www.electronjs.org/docs/latest/api/crash-reporter
Top comments (0)