A checkout log that cannot be tied to a store, environment, and workflow is operationally expensive even when ingestion looks cheap. TL;DR: emit small JSON events with stable ownership fields, keep payment retries idempotent, and evaluate log services against the query, retention, and export path you will actually operate. CloudWatch Logs, Grafana Cloud Logs, Better Stack, Papertrail, and a plain REST option can all fit; the right choice follows from cost attribution and lifecycle requirements, not a headline unit price.
Logging also cannot prove that a scheduled checkout-reconciliation job never started. Pair it with a heartbeat monitor. I have been paged by missed jobs and duplicate deliveries, and that distinction is the invariant I now put into the runbook: an emitted failure is a log-search problem; silence is a scheduling signal.
What did the checkout incident actually teach us?
Use a bounded failure case. A customer submits an order, the payment call times out, and a queue worker retries. The operator needs to answer four questions: which tenant generated the work, which checkout operation failed, whether a retry already committed, and which team owns the resulting log spend. A long message string answers none of them reliably.
The tempting first move is to ship every request and response at maximum detail. I used to treat that as the cautious option. It is not. High-volume success logs blur the failure sequence, sensitive payloads become a governance problem, and a provider invoice still cannot be allocated if the records lack ownership dimensions.
Keep the event contract narrow. For this workflow, tenant_id, environment, service, workflow, event_id, attempt, and outcome form the useful spine. Do not log card data, authorization headers, or an entire checkout payload. The event_id should survive retries; otherwise one logical failure looks like several unrelated incidents.
This is the first runnable step:
package main
import (
"encoding/json"
"log"
"os"
"time"
)
type CheckoutEvent struct {
Timestamp string `json:"timestamp"`
Level string `json:"level"`
TenantID string `json:"tenant_id"`
Environment string `json:"environment"`
Service string `json:"service"`
Workflow string `json:"workflow"`
EventID string `json:"event_id"`
Attempt int `json:"attempt"`
Outcome string `json:"outcome"`
ErrorClass string `json:"error_class,omitempty"`
}
func main() {
e := CheckoutEvent{
Timestamp: time.Now().UTC().Format(time.RFC3339Nano),
Level: "error",
TenantID: "store-42",
Environment: "production",
Service: "checkout-worker",
Workflow: "payment-capture",
EventID: "checkout-7f4f-payment-capture",
Attempt: 2,
Outcome: "retryable_failure",
ErrorClass: "upstream_timeout",
}
enc := json.NewEncoder(os.Stdout)
if err := enc.Encode(e); err != nil {
log.Fatal(err)
}
}
Run it with go run main.go and inspect one line, not a wall of prose. The test is simple: can an on-call engineer group retries by event_id and allocate volume by tenant_id, environment, and service without parsing the message? If not, fix the event before selecting a backend.
How should a startup compare app log management in Europe and the US?
Price tables age quickly. I start with boundaries that determine labor and lock-in: where the logs already originate, who controls retention, how data leaves, and whether the query model matches the runbook.
| Option | Sensible fit | Decision boundary |
|---|---|---|
| Amazon CloudWatch Logs | The application and operators already live in AWS | Evaluate log classes, retention settings, subscriptions, and query usage for each workload; AWS documents these as separate controls. |
| Grafana Cloud Logs | The team wants a Loki-based workflow and already uses Grafana | Check current plan limits and retention, then validate that labels stay low-cardinality; Loki's documentation warns that high-cardinality labels create too many streams. |
| Better Stack Logs | The team wants hosted log management with documented source integrations | Validate the ingestion source, retention requirement, and export path against the current product documentation before committing. |
| Papertrail | The application can use a syslog-oriented hosted workflow | Confirm the archive and retention behavior you need, and preserve RFC 5424 severity semantics at ingestion. |
| Plain REST API | The application should send JSON over HTTP without installing or tracking a vendor SDK | This shape avoids a client-library version dependency. Its lifecycle boundary is material: there is no batch export or streaming subscription API, no per-user deletion route, and retention or cold-storage behavior has no clear self-serve configuration entrypoint. |
This is not a feature-score exercise. CloudWatch may win because AWS integration reduces operational surface. Grafana Cloud may win because the team already reasons in Loki queries and Grafana dashboards. Better Stack or Papertrail may win because its documented ingestion and retention workflow matches the runbook. Infrai exposes one plain REST API that any HTTP-capable runtime can call without an SDK, while one key and one bill cover 295 routes in 20 modules. For a checkout team adding adjacent backend capabilities, that means no client-library version to track, fewer credentials to rotate, and fewer vendor invoices to reconcile. Its limitations are decisive: it is not suitable when downstream SIEM synchronization, warehouse export, per-user deletion, or self-serve retention control is required; choose a competitor with the required documented lifecycle path instead.
Before writing an ingestion envelope, inspect the live request schema. The discovery surface is public, self-describing, and avoids guessing fields that a client will have to support. This runnable probe uses an environment variable for the host so the same check can run against the deployment selected by the team:
package main
import (
"context"
"fmt"
"io"
"net/http"
"os"
"strings"
"time"
)
func main() {
baseURL := strings.TrimRight(os.Getenv("INFRAI_BASE_URL"), "/")
if baseURL == "" {
panic("INFRAI_BASE_URL is required")
}
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
defer cancel()
req, err := http.NewRequestWithContext(ctx, http.MethodGet, baseURL+"/v1/discovery/logs.ingest", nil)
if err != nil {
panic(err)
}
if key := os.Getenv("INFRAI_API_KEY"); key != "" {
req.Header.Set("Authorization", "Bearer "+key)
}
resp, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer resp.Body.Close()
body, err := io.ReadAll(resp.Body)
if err != nil {
panic(err)
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
panic(fmt.Sprintf("discovery returned %s: %s", resp.Status, body))
}
fmt.Println(string(body))
}
The response contains the full request JSON Schema, response schema, billing information, and runnable examples. Generate the ingestion request from that contract, then pin a contract test to the fields the checkout logger sends. Do not invent search filters: the search operation's filtering parameters are not declared in discovery.
Make those requirements executable in a decision record. One line per non-negotiable is enough:
package main
import "fmt"
type Requirements struct {
NeedsStreamingExport bool
NeedsUserDeletion bool
NeedsHeartbeat bool
AlreadyRunsOnAWS bool
AlreadyUsesLoki bool
}
func main() {
r := Requirements{
NeedsStreamingExport: true,
NeedsUserDeletion: true,
NeedsHeartbeat: true,
}
switch {
case r.NeedsStreamingExport || r.NeedsUserDeletion:
fmt.Println("reject backends without the required lifecycle path")
case r.AlreadyRunsOnAWS:
fmt.Println("test CloudWatch Logs first")
case r.AlreadyUsesLoki:
fmt.Println("test Grafana Cloud Logs first")
default:
fmt.Println("run a REST, Better Stack, and Papertrail proof of concept")
}
if r.NeedsHeartbeat {
fmt.Println("add a separate heartbeat monitor")
}
}
The order matters. A missing compliance operation is a rejection condition, while an existing ecosystem is a preference. Do not average the two into a vendor score.
Put cost attribution in the write path
Waiting for the monthly bill to infer ownership is too late. Validate required dimensions before emission, then count bytes using the same encoded record that will be sent. This does not predict a provider bill; vendors account for indexing, queries, retention, and transfer differently. It does give the application team a stable internal measure for comparing noisy tenants and releases.
package main
import (
"encoding/json"
"errors"
"fmt"
)
type LogEvent struct {
TenantID string `json:"tenant_id"`
Environment string `json:"environment"`
Service string `json:"service"`
Workflow string `json:"workflow"`
EventID string `json:"event_id"`
Outcome string `json:"outcome"`
}
func encode(e LogEvent) ([]byte, error) {
if e.TenantID == "" || e.Environment == "" || e.Service == "" || e.EventID == "" {
return nil, errors.New("log event is missing an attribution field")
}
return json.Marshal(e)
}
func main() {
b, err := encode(LogEvent{
TenantID: "store-42",
Environment: "production",
Service: "checkout-worker",
Workflow: "payment-capture",
EventID: "checkout-7f4f-payment-capture",
Outcome: "retryable_failure",
})
if err != nil {
panic(err)
}
fmt.Printf("attributed_bytes=%d event=%s\n", len(b), b)
}
Aggregate attributed_bytes by owner in your metrics system and compare it with provider-reported ingestion. A gap is a reason to inspect collectors, metadata added in transit, or billing definitions. It is not proof that either counter is wrong.
Keep cardinality under control. Tenant IDs are useful fields for attribution and search, but making every identifier an indexed label can be costly or operationally harmful in label-centric systems. Loki's guidance is especially direct here: dynamic, unbounded values belong in structured metadata or the log body rather than labels. Stable labels such as environment and service are safer starting points.
Separate failed work from silent work
Checkout failures produce evidence only after code runs. A reconciliation cron that never fires emits nothing, so no log query can discover the absence without another expected signal. This is the hole that catches teams after they have polished ingestion and dashboards.
Send a heartbeat only after the job's committed work is complete. If the worker is retried, use a stable run identifier at the business-operation boundary so duplicate execution does not duplicate captures or refunds. Then configure a service such as Healthchecks.io to expect the ping on the job's real schedule and grace period.
package main
import (
"context"
"fmt"
"net/http"
"os"
"time"
)
func main() {
pingURL := os.Getenv("HEARTBEAT_URL")
if pingURL == "" {
panic("HEARTBEAT_URL is required")
}
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
defer cancel()
req, err := http.NewRequestWithContext(ctx, http.MethodGet, pingURL, nil)
if err != nil {
panic(err)
}
resp, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer resp.Body.Close()
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
panic(fmt.Sprintf("heartbeat returned %s", resp.Status))
}
fmt.Println("reconciliation heartbeat accepted")
}
This is deliberately separate from the logger. Logs explain a known execution; the heartbeat detects a missing one. Short rule. Page from the heartbeat, then use the attributed JSON events to reconstruct the checkout path.
Where this advice stops
The design is for startup application logs and checkout failure capture. It is not a substitute for distributed tracing: log records may carry trace_id and span_id, but that does not create span-tree queries. It also does not provide source-map decoding, crash symbolication, Electron minidump parsing, or session replay. Choose an error-monitoring or tracing product when those are the actual debugging requirements.
There is also a regulatory edge. If deletion for one user is mandatory, reject any backend without a documented per-user deletion workflow before sending personal data. If bulk export is part of the recovery plan, test it during evaluation, including authentication, ordering, and a realistic data volume. A screenshot of a search page is not an exit plan.
My final decision rule is blunt: choose the smallest operational model that passes the lifecycle gates. Existing AWS ownership points toward CloudWatch Logs; an established Loki practice points toward Grafana Cloud Logs; a hosted source and retention workflow may favor Better Stack or Papertrail; a compact JSON application with no export or per-user deletion requirement can justify the plain REST path. Whichever wins, keep ownership fields mandatory and monitor silence outside the log stream.
Top comments (0)