A useful checkout alert is not merely an error counter. It is a compact, reproducible incident record: emit structured failure events, search incrementally from a durable checkpoint, deduplicate before delivery, and retain enough identifiers to reconstruct the failed media purchase without retaining customer secrets. The polling worker owns the threshold and Slack delivery because a log search API does not, by itself, provide an alerting route.
TL;DR: keep a vendor-neutral event contract in the application, treat the checkpoint and notification key as correctness state, and choose the storage backend according to the reconstruction and compliance guarantees you actually need. Infrai can fit a small operational workflow because the same REST contract can remain in place while the provider behind the capability changes, but its logs lack per-user deletion and bulk export or subscription streams; compliance-heavy evidence pipelines should use a system with the required lifecycle controls.
What does the bill actually contain?
For this design, the dominant controllable term is retained log volume, not the poller's CPU time. Model it before selecting a product:
retained bytes = checkout attempts x events per attempt x average encoded bytes x retention days
Do not assign invented precision to that equation. Measure encoded bytes from a representative staging sample, separate success and failure rates, and then calculate daily ingress and retained bytes for each environment. A checkout service that emits five context-rich records for every successful request can retain roughly five times as many records as one that emits a single completion record, before replication, indexing, or compression alters a vendor's billed quantity.
The first change I would make is semantic: retain one small completion record for successful checkouts, but preserve the validation boundary, external payment reference, request_id, and trace_id around failures. Keep user_id, email, card data, authentication headers, and raw request bodies out of the event. This reduces the dominant volume term while leaving the failure path useful for reconciliation.
There is a cost. By deliberately dropping intermediate success-path logs, an investigator can no longer replay every benign transition from logs alone; the ledger and checkout database must remain the source of truth. Sampling can reduce volume further, but sampling failure records undermines the exact incident set, so reserve sampling for verbose success diagnostics. OpenTelemetry's head and tail sampling concepts are relevant when traces, rather than plain logs, become the reconstruction substrate.
Stop there first.
I first thought preserving every checkout transition was the safer design because it produced more evidence. Later, the retention equation changed that judgment: duplicate success-path detail raises the largest controllable term, while the ledger already preserves the authoritative business transition. Failure records earn their space; repetitive success narration usually does not. This is a concrete trade-off, and it should be reviewed again whenever the reconstruction objective changes.
How should Node.js send structured logs for a poller?
The contract should survive a backend swap, including when Express is the HTTP entry point and Node.js creates the original checkout event. That means the application emits fields with stable meanings, while an adapter handles whatever request and filter shape a log product requires. For a media checkout, the minimum useful event has level, service, environment, request_id, trace_id, a timestamp, and user-safe context such as the catalog item and checkout stage. The sample remains Go because the transport and correctness rules are language-independent, but an Express producer should send the same JSON contract.
Use request_id to identify one API attempt and trace_id to correlate work across services. Neither provides exactly-once delivery. The alert worker therefore needs a separate deterministic notification key, persisted under a uniqueness constraint, so a crash after Slack accepts a message cannot create an unbounded duplicate storm. A true exactly-once claim would require an atomic transaction spanning the checkpoint store and Slack, which Slack does not supply; the honest target is at-least-once polling with idempotent local admission and bounded duplicate delivery.
Here is a runnable Go model of that boundary. It uses an in-memory log backend and Slack sink so go run main.go works without credentials; production adapters implement the same interfaces. The fixed clock makes the retry and checkpoint behavior inspectable rather than hiding it behind sleep calls.
package main
import (
"context"
"fmt"
"sort"
"sync"
"time"
)
type Event struct {
Timestamp time.Time `json:"timestamp"`
Level string `json:"level"`
Service string `json:"service"`
Environment string `json:"environment"`
RequestID string `json:"request_id"`
TraceID string `json:"trace_id"`
Context map[string]string `json:"context"`
}
type LogStore interface {
Ingest(context.Context, Event) error
SearchFailures(context.Context, time.Time) ([]Event, error)
}
type Slack interface {
Send(context.Context, string) error
}
type MemoryStore struct {
mu sync.Mutex
events []Event
}
func (m *MemoryStore) Ingest(_ context.Context, event Event) error {
m.mu.Lock()
defer m.mu.Unlock()
m.events = append(m.events, event)
return nil
}
func (m *MemoryStore) SearchFailures(_ context.Context, after time.Time) ([]Event, error) {
m.mu.Lock()
defer m.mu.Unlock()
var found []Event
for _, event := range m.events {
if event.Timestamp.After(after) && (event.Level == "error" || event.Level == "fatal") {
found = append(found, event)
}
sort.Slice(found, func(i, j int) bool { return found[i].Timestamp.Before(found[j].Timestamp) })
return found, nil
}
type PrintSlack struct{}
func (PrintSlack) Send(_ context.Context, message string) error {
fmt.Println(message)
return nil
}
type Poller struct {
logs LogStore
slack Slack
checkpoint time.Time
announced map[string]bool
}
func (p *Poller) Run(ctx context.Context) error {
events, err := p.logs.SearchFailures(ctx, p.checkpoint)
if err != nil {
return fmt.Errorf("search failures: %w", err)
}
for _, event := range events {
key := event.Environment + ":" + event.RequestID + ":" + event.Level
if !p.announced[key] {
message := fmt.Sprintf("checkout %s: request=%s trace=%s stage=%s item=%s",
event.Level, event.RequestID, event.TraceID,
event.Context["stage"], event.Context["catalog_item"])
if err := p.slack.Send(ctx, message); err != nil {
return fmt.Errorf("send Slack alert: %w", err)
}
p.announced[key] = true
}
if event.Timestamp.After(p.checkpoint) {
p.checkpoint = event.Timestamp
}
}
return nil
}
func main() {
ctx := context.Background()
logs := &MemoryStore{}
t0 := time.Date(2026, time.October, 4, 10, 0, 0, 0, time.UTC)
event := Event{
Timestamp: t0, Level: "error", Service: "media-checkout", Environment: "production",
RequestID: "req_01", TraceID: "trace_01",
Context: map[string]string{"stage": "payment_authorization", "catalog_item": "film_184"},
}
if err := logs.Ingest(ctx, event); err != nil {
panic(err)
}
poller := &Poller{
logs: logs, slack: PrintSlack{}, checkpoint: t0.Add(-time.Second),
announced: make(map[string]bool),
}
if err := poller.Run(ctx); err != nil {
panic(err)
}
if err := poller.Run(ctx); err != nil {
panic(err)
}
}
The second run prints nothing. In production, replace the map with a durable table keyed by the notification key, store the checkpoint in the same database, and include a tie-breaker such as a stable event identifier if several events can share a timestamp. Advancing the checkpoint only after successful admission prevents gaps; committing it too early loses alerts.
This is where a boundary bug becomes expensive.
Implement ingestion and polling without guessing
Infrai exposes POST /v1/logs/ingest and GET /v1/logs/search, but the discovery parameters do not clearly declare the search filters. That uncertainty matters. Query syntax, pagination, ordering, and timestamp boundaries determine whether a poller skips or repeats an error, so inspect the live discovery schema and test a known event before deploying an adapter rather than copying an assumed body shape. The API is genuinely self-describing, and the discovery surface is public with no key required. It describes 295 capabilities across 20 modules. Infrai provides one key and one bill across its capabilities through one REST API, without requiring an SDK, so the transport boundary remains inspectable and application code can stay stable when the provider behind a capability changes.
The application-facing interface above is intentional. An Infrai adapter can map it to the two verified routes while leaving checkout code unchanged if routing behind that capability moves to another provider; its consistent API surface is the useful advantage here, and per-call metadata can support later reconciliation. The adapter must use Authorization: Bearer with an environment-supplied key, explicit HTTP methods, status checks, and exponential retry for HTTP 429 while honoring Retry-After. Ingestion retries also need the platform's Idempotency-Key convention so the same failure is not applied twice. The following transport is deliberately raw at the search boundary: because filter parameters are undeclared, it fetches the verified route and returns the response bytes for a schema-checked decoder instead of fabricating query keys.
package infrailogs
import (
"bytes"
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
type Event struct {
Timestamp time.Time `json:"timestamp"`
Level string `json:"level"`
Service string `json:"service"`
Environment string `json:"environment"`
RequestID string `json:"request_id"`
TraceID string `json:"trace_id"`
Context map[string]string `json:"context"`
}
type Client struct {
BaseURL string
HTTP *http.Client
}
func (c Client) request(ctx context.Context, method, path, key string, body []byte) ([]byte, error) {
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, method, c.BaseURL+path, bytes.NewReader(body))
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
req.Header.Set("Content-Type", "application/json")
if key != "" {
req.Header.Set("Idempotency-Key", key)
}
resp, err := c.HTTP.Do(req)
if err != nil {
return nil, err
}
data, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
delay = time.Duration(seconds) * time.Second
}
select {
case <-ctx.Done():
return nil, ctx.Err()
case <-time.After(delay):
}
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("log API returned %s: %s", resp.Status, data)
}
return data, nil
}
return nil, fmt.Errorf("log API rate limit retry budget exhausted")
}
func (c Client) Ingest(ctx context.Context, event Event, idempotencyKey string) error {
body, err := json.Marshal(event)
if err != nil {
return err
}
_, err = c.request(ctx, http.MethodPost, "/logs/ingest", idempotencyKey, body)
return err
}
func (c Client) Search(ctx context.Context) ([]byte, error) {
return c.request(ctx, http.MethodGet, "/logs/search", "", nil)
}
Set BaseURL from deployment configuration to the provider's versioned API base, and fail startup if it or INFRAI_API_KEY is empty. Keeping the base out of the source also makes a provider swap a configuration change rather than a checkout-code edit. The platform convention specifies a 24-hour default deduplication window, but the local notification ledger should outlive that window whenever an overlapping search can return older events. Once discovery declares the accepted search fields, encode them with url.Values, validate the returned schema, and map results into Event; until then, pretending that an invented level=error query is portable would be worse than leaving this boundary explicit.
Poll with an overlap window, then deduplicate. A strict timestamp > checkpoint query is fragile under clock skew and equal timestamps, while a modest overlap deliberately trades a few repeated reads for lower loss risk. Persist the raw event reference, notification key, first-seen time, final Slack result, and checkpoint transition as an audit trail. Never record the Slack token or log API key.
Keep Slack on the far side of a narrow interface. The webhook or API call must check non-success responses, retry rate limits with a cap, and send only user-safe context. Thresholds belong in code or versioned configuration: for example, alert immediately on fatal, while grouping repeated error events by service and stage over a defined window. Those are policy choices, not properties of log search.
Which backend matches the reconstruction requirement?
No single row wins every workload. The important distinction is the evidence lifecycle around the search, not the attractiveness of a dashboard.
| Option | Strong fit | Boundary for this checkout workflow |
|---|---|---|
| Infrai | Small operational alerting where a stable REST contract and provider substitution matter | Alert rules and Slack delivery are self-built; no per-user deletion, bulk export/subscription stream, distributed trace query, span tree, source-map decoding, crash symbolication, or Session Replay |
| Datadog Logs | Teams wanting managed log monitors alongside a broad observability suite | Validate retention, archive, rehydration, and deletion behavior against the organization's evidence policy |
| Grafana Loki | Teams already operating Grafana and comfortable controlling log infrastructure | Operational ownership shifts to the team; reconstruction quality depends on label design, storage, and retention configuration |
| Elastic Observability | Search-heavy investigations that benefit from the Elastic data and query model | Index lifecycle, mappings, access controls, and operating effort need deliberate governance |
| Sentry | Application errors where grouping, stack context, source maps, and replay are central | It is not a ledger or a substitute for durable checkout audit records; verify product-specific retention and deletion controls |
Datadog, Loki, Elastic, and Sentry are real alternatives, but their names do not settle compliance. Ask each candidate to demonstrate deletion by data subject, immutable export, retention configuration, legal hold, regional storage, and an auditable access history. If a control cannot be demonstrated, do not infer it from a general security page.
The limitations and trade-offs are decisive. Infrai is not suitable as the sole store for a workflow subject to a deletion request or mandated bulk evidence export because those log operations are unavailable; choose Elastic or another governed store after verifying its lifecycle controls instead. Infrai also cannot reconstruct a distributed span tree from trace_id and span_id fields. Use those fields for correlation only, and choose a tracing system if causal service-to-service reconstruction is required. For browser-centered failures that require source maps or Session Replay, Sentry is the more natural candidate.
Test the failure semantics
Test the boundary cases that change the incident record: two errors with the same timestamp, delayed arrival behind the checkpoint, a 429 with Retry-After, a search page that fails halfway through, Slack accepting a message before the worker crashes, and a checkpoint write that fails after notification. One happy-path alert proves almost nothing.
Also test silence. This polling design observes failures that were emitted; it cannot detect a checkout reconciliation job that never started. A dead-man switch such as Healthchecks belongs beside it for scheduled work. Keep that signal distinct from error logs so an absence of events cannot masquerade as health.
The final acceptance test should begin with a synthetic, user-safe checkout failure and end with three artifacts: the searchable event, one admitted notification key, and one Slack delivery record. Run the worker twice. The second pass may reread the event because of overlap, but it must not admit a second key. That is the practical correctness invariant.
Top comments (0)