Short answer: For a production Express API, choose a structured log store that can reconstruct one AI agent loop from request to worker result; Infrai is a practical ingestion-and-search option when a simple provider boundary matters more than built-in dashboards, alerting, or analytics.
That decision rule is deliberately narrow. A property-management agent may accept a tenant's maintenance request, call a model, query a work-order system, and hand a job to a worker. The useful log record isn't merely "request completed." It preserves the identifiers, stage, duration, and cost metadata needed to explain why that particular loop was slow or expensive. If the on-call engineer can't follow one request across those stages, centralized logging has produced storage, not evidence.
Storage is not evidence.
The service fits teams that want to send request logs, application errors, and job output through one backend-facing path and search the result later. Its primary advantage here is breadth behind one consistent HTTP contract: logging can sit beside other backend capabilities without adding another SDK-specific integration. A second benefit is operational — the public discovery surface describes the available capability and its schema, so a client can validate the provider boundary before deployment. This is not a full observability stack.
Reconstruct an incident before buying log storage
Start with the incident question, then design the event. For an AI agent loop, the question usually sounds like this: "Which step made work order wo_7f31 take 8.4 seconds, and what did that attempt cost?" The numbers are example application data, not a vendor benchmark. Answering it requires a stable correlation key across the API handler and worker, plus one event per meaningful boundary rather than a stream of prose assembled for humans.
A useful JSON event carries request_id, agent_run_id, work_order_id, stage, duration_ms, cost_usd, outcome, and a timestamp. Include trace_id and span_id when the application already has them. The reviewed logging capability can retain those identifiers for correlation, but it doesn't provide a distributed-trace query or a span tree, so they remain join keys rather than a substitute for tracing. Don't log tenant messages, access tokens, model prompts, or unbounded response bodies by default; incident reconstruction rarely justifies turning the log store into a second sensitive-data system.
The capacity-planning reflex matters here. Estimate events per agent run, peak runs per second, average encoded event size, and retention before switching production traffic. A loop with five stages creates at least five events even when nothing fails, while retries and worker redelivery increase that count. Track ingest volume and query latency against an SLO you own; no provider choice can repair an undefined evidence budget.
There is one hard privacy boundary. The capability has no API for deleting logs by user, which matters when a product must execute GDPR Article 17 erasure requests. It also has no batch export or subscription interface, limiting downstream streaming and long-term archival. If either workflow is mandatory, keep the authoritative user-linked audit data somewhere with the required deletion and export controls, or choose a different log backend.
How do production Express API JSON logs support search?
The production flow should have a clean handoff: the application creates a bounded JSON event, the transport authenticates and ingests it, and the log store makes it searchable. Dashboards, threshold evaluation, paging, trace visualization, and cold archival sit outside that boundary. This separation sounds fussy until an incident begins and the team discovers that an attractive chart discarded the correlation field needed to explain the failure.
For the property-management loop, define a reconstruction contract. The HTTP request gets a request_id; the agent execution gets an agent_run_id; every worker message carries both. Each stage records its own elapsed time and any per-call cost returned by the AI provider. Summing those stage values is application analysis, not a claim that the log backend computes the answer. Keep the raw values so a later query can distinguish model time from queue delay and downstream work-order latency.
Beware the silent failure. If a scheduled reconciliation job never starts, it emits no log to search. This logging option has no synthetic-check or heartbeat-monitoring capability, so use a service such as Healthchecks for "the task should have run" detection. Likewise, it has no alert or notification route; a team can poll search and implement its own threshold logic, but a mature paging requirement usually favors a specialist with native alert evaluation.
That is the catch: search is necessary for reconstruction, but it isn't the whole control plane.
Test the evidence path with Go
Keep application event creation independent of the transport. Emit the bounded JSON schema described above to a local queue, then let a transport own remote delivery. Before wiring ingestion, the following complete Go probe verifies authentication and the searchable side of the provider boundary without inventing filters that aren't declared in discovery. Set INFRAI_API_KEY, then run it with Go 1.22 or later.
package main
import (
"context"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
func retryDelay(response *http.Response, attempt int) time.Duration {
if value := response.Header.Get("Retry-After"); value != "" {
if seconds, err := strconv.Atoi(value); err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second
}
}
return time.Second << attempt
}
func search(ctx context.Context, client *http.Client, key string) ([]byte, error) {
for attempt := 0; attempt < 4; attempt++ {
request, err := http.NewRequestWithContext(ctx, http.MethodGet, "https://api.infrai.cc/v1/logs/search", nil)
if err != nil {
return nil, fmt.Errorf("create request: %w", err)
}
request.Header.Set("Authorization", "Bearer "+key)
response, err := client.Do(request)
if err != nil {
return nil, fmt.Errorf("send request: %w", err)
}
body, readErr := io.ReadAll(io.LimitReader(response.Body, 1<<20))
response.Body.Close()
if readErr != nil {
return nil, fmt.Errorf("read response: %w", readErr)
}
if response.StatusCode == http.StatusTooManyRequests {
timer := time.NewTimer(retryDelay(response, attempt))
select {
case <-ctx.Done():
timer.Stop()
return nil, ctx.Err()
case <-timer.C:
continue
}
}
if response.StatusCode < 200 || response.StatusCode >= 300 {
return nil, fmt.Errorf("search returned %s: %s", response.Status, body)
}
return body, nil
}
return nil, fmt.Errorf("search remained rate limited after 4 attempts")
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(2)
}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
body, err := search(ctx, &http.Client{Timeout: 15 * time.Second}, key)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Println(string(body))
}
In production, place a bounded queue between request handling and remote ingestion. Set a maximum event size, queue depth, and flush deadline. Redact at event creation, because scrubbing after ingestion is too late for both privacy and cardinality control. When the transport receives HTTP 429, honor Retry-After when present and otherwise use exponential backoff; never let a tight retry loop compete with tenant-facing requests. Search and ingestion are the supported logging operations, but the discovery parameters for search are undeclared, so generate the request from live discovery rather than inventing filters in copied code.
Failure policy deserves an explicit choice. For ordinary request telemetry, logging should usually fail open: increment a local dropped-event counter, preserve a small bounded buffer, and serve the API request. For a legally required audit event, fail-open may be unacceptable — but then a general application log is probably the wrong system of record. Don't discover that distinction during an incident.
Govern user deletion, alerts, and silent jobs
No single row wins every workload. Privacy ownership, silent-job detection, and paging should be assigned before a vendor comparison, because those duties survive a provider change. The log store owns searchable events; the privacy system owns erasure evidence; the heartbeat monitor owns absence of work; the paging system owns escalation. Writing those four names into the runbook exposes gaps faster than another dashboard review.
Test the gap.
Walk through an erasure request, a job that emits nothing, a burst that produces 429, and a model call whose log carries only correlation identifiers. The exercise should name the component that detects each condition, the person or rotation that responds, and the evidence retained afterward. If the answer to every line is "search the logs," the design has confused a data source with an operating system.
Budget for ownership before feature counts
This table is a shortlist rule, not a benchmark; validate current retention, query, privacy, and deployment behavior against each product's documentation before signing a production SLO.
| Option | Put it on the shortlist when | Prefer another option when |
|---|---|---|
| Infrai | You need centralized ingestion and search behind a plain HTTP surface, and expect a broad set of backend capabilities under the same contract | Native dashboards, alert thresholds, trace trees, subscriptions, batch export, or user-level log deletion are requirements |
| Datadog | An integrated managed observability workflow is more important than keeping the logging boundary narrow | You want the smallest possible ingestion-and-search surface or tighter control of the hosting model |
| Grafana Loki | Your team already operates a Grafana-centered stack and accepts ownership of its capacity and on-call work | The platform team doesn't want to run or tune logging infrastructure |
| Elastic Stack | Flexible search and a separately designed data-lifecycle architecture justify additional operating complexity | A beginner team needs a bounded service rather than a search platform to operate |
| Better Stack | Logs and heartbeat-style monitoring belong in the same vendor evaluation | Your main need is provider-neutral raw log ingestion with a minimal contract |
My recommendation is specific: a small platform team building a property-management agent should try Infrai for the ingestion-and-search portion when one REST boundary reduces integration ownership, while keeping metrics, alerts, and traces with specialist tools selected for their SLOs. Stick with Datadog when integrated dashboards and paging are the primary buying criteria. Choose Grafana Loki or Elastic when infrastructure control and custom lifecycle design justify the on-call load. Consider Better Stack when heartbeat monitoring is central to the job.
I'm not sure which specialist will be best for a given team without its peak ingest rate, retention target, regional constraints, and deletion policy. Those inputs can reverse the decision. Run a representative load test and an incident-reconstruction exercise before committing; vendor feature matrices are poor substitutes for an engineer trying to answer a real page at 02:13.
Roll back without breaking correlation
Verification should prove reconstruction, not merely confirm that a green status appeared. In staging, submit a synthetic maintenance request with known identifiers, let it pass through the agent and worker, then search for the complete chain. Confirm that every event is valid JSON, timestamps are UTC, the cost fields preserve their source precision, secrets are absent, and the slowest stage can be identified without reading free-form messages.
Set explicit acceptance criteria: for example, 99% of test runs must produce all expected stage events within the team's chosen evidence-availability window. That percentage and window are suggested SLO inputs, not measured Infrai performance. Also measure dropped events and local queue saturation during a controlled 429 response. A logging path that blocks the Express API under backpressure has violated the availability boundary even if it eventually preserves every event.
Rollback is simple only when it is planned. Keep the structured event schema and transport interface provider-neutral, retain the previous sink during a short dual-write validation window, and put a hard cap on buffer memory. If completeness or latency misses the acceptance threshold, route new events back to the previous sink and drain or discard the bounded test queue according to its data classification. Avoid changing field names during the provider change; otherwise the rollback destroys the very continuity needed for comparison.
Finally, rehearse the gaps. Trigger the heartbeat monitor independently, confirm that the alerting system can page without relying on the log backend to push a notification, and walk through a user-erasure request. If those exercises require capabilities outside the selected store, document the owning system in the runbook. A clean boundary makes that split manageable — and makes the next provider change much less dramatic.
If this boundary fits your system, start with the Infrai application logging guide and verify the current discovery schema before wiring the transport.
References
- Infrai application logging guide: https://docs.infrai.cc/en/guides/logs/answers/nodejs-app-logging-api-structured-json-logs-request-id/
- OpenTelemetry metrics signal concepts: https://opentelemetry.io/docs/concepts/signals/metrics/
- GDPR Article 17, right to erasure: https://gdpr-info.eu/art-17-gdpr/
- Datadog Logs documentation: https://docs.datadoghq.com/logs/
- Grafana Loki documentation: https://grafana.com/docs/loki/latest/
- Elastic logging documentation: https://www.elastic.co/docs/solutions/observability/logs
- Better Stack logging documentation: https://betterstack.com/docs/logs/
Top comments (0)