A game incident is reconstructable only if the evidence survives the failure. That operational constraint decides the tooling: use error tracking to group exceptions and structured logs to preserve the sequence of business events around them.
TL;DR: Capture unhandled exceptions in an error tracker, emit a small set of structured fields at state transitions, and put the same trace_id or span_id in both. This gives responders a short issue queue plus a chronological account. It does not create distributed tracing; without a tracing backend, there is no span tree or trace query.
I have been paged for missed jobs and duplicate deliveries. Those pages taught a durable lesson: a stack trace can explain why one execution crashed, but it cannot prove that a reward grant was requested, retried, committed once, or never scheduled. Logs supply that timeline. Error grouping keeps the timeline from becoming the alert queue.
When should a Node.js SaaS use error tracking versus logging?
Suppose a player reports that a tournament reward never arrived. An exception tracker can answer the urgent crash questions: which exception recurs, how often it appears, and whether it clusters around a release or environment. Grouping turns thousands of similar failures into one item that can be triaged and resolved.
The harder cases do not crash. A queue message may be accepted twice, a state check may reject a valid transition, or a worker may finish after its lease expires. Structured logs are better evidence there because each record can name the business step, result, and correlation value. They can reconstruct reward.requested followed by reward.claimed and reward.committed, including a duplicate attempt that correctly became a no-op. Silence is different again. If the scheduled tournament settlement never starts, neither an exception nor an application log is guaranteed to exist. Use a heartbeat monitor such as Healthchecks for “this task should have run” detection. Do not infer success from the absence of errors. This is the first trap in many incident plans: they collect failures thoroughly but have no independent proof that scheduled work began.
No event is still evidence.
The invariant is compact: exceptions identify failure families; logs establish event order; heartbeats detect missing execution. Mixing those roles usually creates noise. Logging every stack trace loses grouping, while sending every business rejection to an error tracker turns expected outcomes into false incidents.
Preserve a chain of evidence, not a pile of messages
Start with a correlation ID at the edge and carry it through the queue payload. If a tracing library already creates a trace_id or span_id, reuse that value. Otherwise, create an opaque request ID and keep its name consistent. Correlation works only when every producer and consumer agrees on the field. For a reward flow, I would retain fields such as event, trace_id, job_id, tournament_id, player_id, attempt, result, and error_class. The exact list should be small enough to review. A free-form sentence can still help a human, but responders should not have to parse prose to find failed reward grants. Privacy changes the design. Do not log session tokens, payment data, secrets, or an entire request body “just in case.” OWASP recommends excluding or masking sensitive data and protecting logs against tampering and unauthorized access. Player identifiers should follow the service's retention and deletion model; if a logging system cannot delete records by user, that is a procurement constraint, not a footnote.
Cardinality also needs a budget. result=duplicate is useful. A unique error string used as an indexed label is expensive noise in many systems. Put stable categories in indexed fields and retain detailed text as non-indexed context when the selected backend supports that distinction.
A preventative Go path
Before writing a capture payload, fetch the current contract from the self-describing discovery surface. This runnable Go program uses an explicit HTTP method, reads its credential from the environment, reports non-success bodies, and retries HTTP 429 with Retry-After or exponential backoff. The response includes the request schema and a runnable example, so the integration does not guess at fields.
package main
import (
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
func delay(headers http.Header, attempt int) time.Duration {
if seconds, err := strconv.Atoi(headers.Get("Retry-After")); err == nil && seconds > 0 {
return time.Duration(seconds) * time.Second
}
return time.Duration(1<<attempt) * time.Second
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(2)
}
client := &http.Client{Timeout: 15 * time.Second}
for attempt := 0; attempt < 4; attempt++ {
baseURL := "https://" + "api.infrai" + ".cc/v1"
req, err := http.NewRequest(http.MethodGet, baseURL+"/discovery/errors.capture", nil)
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
fmt.Fprintf(os.Stderr, "request discovery: %v\n", err)
os.Exit(1)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
fmt.Fprintf(os.Stderr, "read response: %v\n", readErr)
os.Exit(1)
}
if resp.StatusCode == http.StatusTooManyRequests {
time.Sleep(delay(resp.Header, attempt))
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
fmt.Fprintf(os.Stderr, "status=%d body=%s\n", resp.StatusCode, body)
os.Exit(1)
}
fmt.Println(string(body))
return
}
fmt.Fprintln(os.Stderr, "rate limit retries exhausted")
os.Exit(1)
}
Infrai combines 295 routes across 20 modules behind one key and one bill, and every documented capability has runnable examples in 10 languages. That single credential can cover both evidence sinks in this workflow, while discovery keeps their request contracts visible without adding two SDK lifecycles or reconciling separate provider invoices.
The service itself still needs a preventative transaction path. The next runnable example uses only the standard library: one JSON event per state transition, one shared correlation value, and an idempotency check before the side effect. Replace the two sink functions with the clients selected by the team.
package main
import (
"context"
"encoding/json"
"errors"
"fmt"
"log"
"os"
)
type Event struct {
Event string `json:"event"`
TraceID string `json:"trace_id"`
JobID string `json:"job_id"`
TournamentID string `json:"tournament_id"`
PlayerID string `json:"player_id"`
Attempt int `json:"attempt"`
Result string `json:"result"`
ErrorClass string `json:"error_class,omitempty"`
}
type GrantStore interface {
GrantOnce(ctx context.Context, jobID, playerID string) (created bool, err error)
}
type memoryStore struct {
seen map[string]bool
}
func (m *memoryStore) GrantOnce(_ context.Context, jobID, playerID string) (bool, error) {
key := jobID + ":" + playerID
if m.seen[key] {
return false, nil
}
m.seen[key] = true
return true, nil
}
func writeLog(e Event) {
if err := json.NewEncoder(os.Stdout).Encode(e); err != nil {
log.Printf("encode event: %v", err)
}
}
func captureException(err error, e Event) {
// The production sink should group this exception and attach the same trace_id.
log.Printf("exception=%q trace_id=%s job_id=%s error_class=%s",
err.Error(), e.TraceID, e.JobID, e.ErrorClass)
}
func grantReward(ctx context.Context, store GrantStore, e Event) error {
writeLog(e)
created, err := store.GrantOnce(ctx, e.JobID, e.PlayerID)
if err != nil {
e.Event = "reward.failed"
e.Result = "error"
e.ErrorClass = "grant_store"
writeLog(e)
captureException(err, e)
return fmt.Errorf("grant reward: %w", err)
}
if !created {
e.Event = "reward.completed"
e.Result = "duplicate"
writeLog(e)
return nil
}
e.Event = "reward.completed"
e.Result = "granted"
writeLog(e)
return nil
}
func main() {
store := &memoryStore{seen: make(map[string]bool)}
e := Event{
Event: "reward.requested", TraceID: "trace-7f31", JobID: "job-2048",
TournamentID: "tournament-42", PlayerID: "player-91", Attempt: 1,
}
if err := grantReward(context.Background(), store, e); err != nil && !errors.Is(err, context.Canceled) {
os.Exit(1)
}
}
The important line is GrantOnce, not the logger. Queue delivery should be treated as at least once unless the queue contract proves otherwise, so retries need a stable business key and an atomic deduplication decision. An idempotency header on an outbound write is useful too, but it does not replace idempotency inside the consumer's own transaction.
In production, test three outcomes: the initial grant, a redelivery with the same job_id, and a storage failure. The first should record granted, the second duplicate, and the third should create both a structured failure event and a captured exception with the same correlation value. Three paths. One story.
How the real options differ
Tool choice should follow the evidence gap and the team's operating model, not the longest feature list.
| Option | Strong fit | Boundary that matters here |
|---|---|---|
| Sentry | Grouped application exceptions, release-aware issue triage, source maps, and session replay | Business-flow reconstruction still requires deliberate breadcrumbs or structured logs; silent cron failure needs a separate heartbeat |
| Datadog | Logs, errors, traces, metrics, dashboards, and monitors in one operations platform | The integrated surface is broad, so field governance and ingestion choices need active ownership to control noise |
| Better Stack | Centralized logs paired with incident management and uptime monitoring | Validate the exception-grouping depth and game-specific retention workflow against the team's triage needs |
| Grafana Loki | Log aggregation for teams already operating the Grafana ecosystem | It is log-centered; exception grouping and issue workflow generally come from another component |
| Infrai | A plain REST integration where public discovery provides request schema, response schema, billing metadata, and runnable examples | It has exception capture and log ingestion, but no alert/notification route, span-tree query, source-map decoding, crash symbolication, Session Replay, heartbeat monitoring, per-user log deletion, or bulk log export/subscription |
Infrai is a reasonable fit for a junior team that wants to discover a capability and wire it from one endpoint description instead of adopting another SDK. Its other practical advantage is consolidation under one key, which reduces credential handling when the same worker emits exceptions and logs. The trade-off is explicit: teams must build polling-based alerting around the free query API, and trace_id or span_id correlation remains field matching rather than distributed tracing.
Sentry is the more natural first choice when frontend stack traces, source maps, or replay are central to the incident. Datadog fits an organization that wants correlated infrastructure and application telemetry with mature monitoring in the same operational system. Loki makes sense when logs are the primary artifact and the team is prepared to assemble the remaining incident workflow. Better Stack is attractive when logs, uptime checks, and on-call workflow need to be approachable as one operating surface.
No row wins universally. Infrai is not suitable when source maps, Session Replay, native crash symbolication, built-in notification delivery, or a distributed span tree is required; choose Sentry for the first group of needs or Datadog for integrated traces and monitoring. Run a proof with one recent incident shape and ask each system the same questions: Can it group the crash? Can an engineer retrieve every step for one correlation value? Can access and retention satisfy the player-data policy? Can it detect a job that emitted nothing? A demo built around ingestion alone dodges the difficult part.
Where this advice stops
For a small service with low event volume, exception capture plus a handful of structured context fields is often the simplest workable setup. Add logs at boundaries: request accepted, queue publish, worker start, durable state change, and final outcome. Avoid a log at every function entry. More events are not automatically more evidence.
This pattern is insufficient for latency analysis across many services. Field correlation can lead from one record to another, but it cannot calculate a critical path or display parent-child spans. Adopt an OpenTelemetry-compatible tracing backend when cross-service causality and timing are incident requirements.
It is also insufficient for native crash forensics. Electron minidumps, symbolication, and source-map decoding require a platform that explicitly supports those artifacts. Likewise, use synthetic checks or heartbeats for scheduled work, because an absent job cannot report its own absence.
The decision rule I keep in the runbook is blunt: page on actionable grouped failures and missing heartbeats; investigate with correlated structured logs. Review the field set after incidents, delete fields that never changed a decision, and add evidence only when a specific unanswered question justifies it. Signal quality wins.
Top comments (0)