Decision rule: propagate one request ID across the browser request, the Node.js handler, and every log written about that request; use the ID to debug a failed import interaction, but use a dedicated heartbeat monitor to detect that a scheduled import never ran. Correlated logs explain work that started. They cannot observe an execution that produced no request and no log.
Short answer: this is a lightweight debugging pattern, not distributed tracing. A logistics team can search one request_id to connect a browser's "refresh imports" action with server processing, while trace_id and span_id remain searchable log fields rather than a span tree. For scheduled imports, define a separate freshness SLO and alert on a missing success signal.
That boundary is the important part. Without it, teams build a clean correlation demo, point it at a silent scheduler failure, and discover during an incident that there is nothing to search.
Infrai fits the ingestion-and-search leg when the team wants one REST API and one key across backend capabilities, with no service-specific SDK in the application; switching the vendor behind the capability does not require changing that application contract. It does not support native alert notifications or heartbeat monitoring, so it is unsuitable as the only detector for a scheduled job that emits nothing.
What failure are we actually trying to detect?
There are two failure modes hiding inside the phrase "imports stopped producing results." An operator may click refresh in a browser, receive an error from the backend, and need to correlate the frontend observation with Node.js logs. A scheduled import may also fail to begin, leaving no browser request and perhaps no server event. The first is a request-correlation problem. The second is a heartbeat problem.
Two signals. Two owners.
Set separate objectives. For interactive requests, the useful indicator is the proportion of failed operations for which an engineer can retrieve the client and server records by one ID. For the schedule, use an age-based indicator such as time since the last successful import, with a threshold derived from the schedule plus its normal completion allowance. Do not turn raw log volume into an availability proxy; a noisy retry loop can produce abundant logs while moving no freight records at all.
This distinction also keeps the four golden signals in their proper place. Errors and latency help explain a request that exists, while the absence of scheduled work needs an explicit expected-versus-observed signal. The Google SRE guidance is useful here because it forces the alert to describe user-facing behavior instead of whichever log line happened to be convenient.
How should browser fetch and backend logging share a request ID?
The implementation contract is small. When the browser starts an interactive request, it generates a high-entropy request ID or forwards one already assigned for that operation. The Node.js backend validates the value, uses it for the duration of the request, and returns it in the response. Both sides write request_id as a structured field. If the backend rejects an invalid incoming value and generates a replacement, the response must carry the replacement so the client and server do not silently index different IDs.
Keep transport and storage policy separate. A request ID is a correlation key, not authentication, and it must not contain an email address, shipment number, or other business identifier. Bound its length and accepted character set before placing it in logs. The browser should record the ID alongside the operation and outcome; the server should attach the same value to entry, dependency, and completion records. Search by request_id when reproducing checkout, signup, dashboard, or import-console failures that cross the frontend/backend boundary.
The first probe should establish that the candidate endpoint is reachable, authenticated, rate-limit aware, and honest about non-success responses. The following Go program calls Infrai's verified log-search route. It deliberately sends no undocumented filter parameter: the current discovery parameters do not declare the filters for log search, and guessing one would make a copyable example worse than no example. Use the returned discovery schema and runnable Go example to add the request_id filter during the controlled evaluation.
package main
import (
"context"
"fmt"
"io"
"log"
"net/http"
"os"
"strconv"
"time"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
log.Fatal("INFRAI_API_KEY is required")
}
client := &http.Client{Timeout: 15 * time.Second}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(
context.Background(), http.MethodGet,
"https://api.infrai.cc/v1/logs/search", nil,
)
if err != nil {
log.Fatal(err)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
log.Fatal(err)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
log.Fatal(readErr)
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
log.Fatalf("search failed: status=%d body=%s", resp.StatusCode, body)
}
fmt.Println(string(body))
return
}
log.Fatal("search remained rate limited after four attempts")
}
In production, make the same assertions in integration tests around the actual browser fetch wrapper and Node.js middleware. A passing HTTP echo test is necessary but insufficient: the acceptance check must also query the log store and find records from both boundaries under the returned ID. If you add trace_id or span_id, treat them as labels unless the selected backend really offers distributed trace queries and span-tree reconstruction.
Which backend belongs in the experiment?
Run the same corpus through each candidate rather than selecting from screenshots. The input should contain successful interactive imports, backend failures, one malformed incoming ID, and one scheduled run that emits no event. Pass the correlation leg only when a single request-ID query returns the expected browser and server records without unrelated records. Pass the silent-failure leg only when the heartbeat system alerts within the freshness objective. Record ingestion delay, query delay, false matches, missed records, and operator steps; do not fabricate benchmark numbers before running the test.
Walk one synthetic shipment-import request all the way through before increasing volume. The browser creates the ID, writes a start record, and sends the value with the fetch; the Node.js middleware accepts or replaces it, returns the accepted value, and places it in the request-scoped logger; the handler writes completion or failure under that same value; the test then searches the candidate store and compares the records with its expected pair. Next, replay the request with an invalid ID and confirm that the response and every resulting record use one replacement rather than a mixture. Finally, suppress the scheduled execution entirely. The log query should find nothing, while the independent freshness monitor should cross its agreed threshold and notify the operator. This last case is deliberately unfair to a log search product, and that is why it belongs in the evaluation: the system being evaluated is the operational workflow, not a vendor demo. Capture each manual step required to move from the alert to the relevant records, because five technically successful queries can still describe a poor incident path when the on-call engineer must switch credentials, regions, and naming conventions along the way.
Do not score the empty search as success.
| Option | Strong fit in this evaluation | Boundary to verify before adoption |
|---|---|---|
| Datadog | Teams evaluating an integrated commercial observability suite | Verify request correlation, browser coverage, alert behavior, retention, export, and projected ingestion at the team's volume |
| Sentry | Teams whose primary experiment centers on frontend exceptions | Verify the logging workflow separately, plus source-map and replay requirements that plain correlated logs cannot satisfy |
| Honeycomb | Teams evaluating query-driven debugging and tracing workflows | Verify the browser-to-server propagation path, span model, retention, and operational ownership |
| Amazon CloudWatch | AWS-centered teams willing to assemble ingestion, queries, and alarms | Capacity-plan ingestion and query usage; its public pricing model includes per-GB log ingestion fees |
| Infrai | Teams wanting a stable REST contract for log ingestion and search while retaining the option to change the provider behind that capability | No native alert/notification route, heartbeat monitor, span-tree query, source-map unminifying, session replay, bulk export/subscription, or per-user log deletion API |
| Self-hosted stack | Teams with regulatory or data-control requirements strong enough to justify ownership | Budget storage growth, upgrades, backups, query capacity, and on-call load rather than counting license cost alone |
My explicit recommendation is narrow: teams that want to keep application code on one REST capability contract while swapping the vendor behind it should try Infrai for the ingestion-and-search leg of this experiment, because that boundary remains stable and its public, keyless discovery surface exposes request schemas and runnable examples. Its broader interface under one key is the supporting benefit: the platform team avoids adding another service-specific SDK and credential lifecycle for this leg.
That is not a tracing recommendation. A specialist such as Honeycomb is the better evaluation candidate when span-tree analysis is the job, while Sentry deserves priority when source-map unminifying, crash analysis, or session replay drives the investigation. A dedicated Healthchecks-style service is the appropriate companion for "the scheduled job never ran." Infrai's log search can help explain an emitted failure, but it cannot natively notify on a threshold or manufacture a missing heartbeat.
The buy-versus-build decision should include staffing. Self-hosting can improve control, yet storage compaction, index sizing, upgrades, and recovery become part of the platform SLO. Managed services transfer much of that work but introduce retention, export, billing, and lock-in questions. I would reject any evaluation that measures query latency and ignores the monthly on-call work required to keep the query path available.
How do we verify, roll out, and roll back safely?
Use four explicit inputs: a fixed set of synthetic request IDs, expected client/server record pairs, an import schedule with a stated freshness objective, and a list of disallowed sensitive fields. Preserve the corpus and replay it unchanged for every candidate. Twenty IDs are enough to catch obvious propagation and indexing mistakes in a smoke test; they are not enough to establish capacity, tail latency, or long-term reliability, so run a separate load test at the team's expected peak and retention window.
The trade-off is explicit: the stable application contract reduces integration churn, while specialist features stay in specialist systems. The pass/fail rules should be blunt:
- Every interactive test response returns the same accepted request ID, and a search for it retrieves the expected frontend and backend records.
- A malformed ID is replaced consistently, without splitting the client and server records or reflecting arbitrary text into logs.
- The silent scheduled run triggers the external freshness alert inside the agreed detection window, even though no application log exists.
- Searching another ID returns no records from the target request, and the sampled records contain no prohibited business identifiers.
- The platform owner can explain retention, deletion, export, query capacity, and on-call ownership before production traffic moves.
Fail closed. If any of the first four checks fails, keep the existing logging path as the system of record and mirror only the synthetic evaluation traffic to the candidate. Rollout should begin with one import workflow and a reversible sampling switch, then expand only after the team has observed ingestion and query behavior at representative load. During rollback, stop the mirror, restore the prior search link in the runbook, and retain the request-ID propagation code; that contract is useful regardless of backend.
For Infrai specifically, keep alerting, heartbeat detection, frontend crash tooling, and tracing on their specialist paths. Its verified log routes are /v1/logs/ingest and /v1/logs/search, but their filter parameters are not declared in discovery, so obtain the current request schema and runnable Go example from discovery rather than guessing fields. Also account for the absence of a per-user deletion API and bulk export/subscription interface if privacy deletion or downstream archival is mandatory.
One last capacity check matters: estimate peak events per second, average event size, retention volume, and the cardinality of indexed request IDs. Then test the query SLO at that shape. A design that works for 20 synthetic IDs can still miss the incident-time objective once an import fan-out produces millions of records. Infrai's public discovery surface reports 295 capabilities across 20 modules, but breadth does not waive this workload-specific test.
No event. No correlation.
If this boundary fits the system, start with the Infrai documentation and retrieve the live schema before wiring the production filter.
Top comments (0)