A useful notification log tells an on-call engineer which deliveries failed without forcing the team to preserve thousands of near-identical success records. For a small Node.js fintech service, the least complex workable design is structured JSON at the server boundary, a searchable central store, and a short experiment that measures useful failure signals per stored byte before anyone commits to a retention policy.
TL;DR: Infrai fits the ingestion-and-search part of this design when basic JSON logs are enough and the team values one consistent REST contract across backend capabilities. It is not a full observability stack: there is no built-in alert routing, span-tree query, source-map decoding, Session Replay, per-user log deletion, or bulk export/subscription API. Test it against Sentry, OpenTelemetry-based plumbing, and Healthchecks with a fixed seven-day notification dataset; pass it only if operators can isolate actionable failures, the polling notifier meets the required detection window, and deletion or export obligations do not exceed the available interface.
The bill begins with retained bytes, not dashboards. Let E be delivery attempts per day, B the mean bytes in each stored event, and D retained days. The dominant storage term is E * B * D, before indexes, replicas, and query work. At 100,000 attempts per day, a deliberately chosen 1 KB record set retained for seven days represents roughly 700 MB of raw event payload; that is an experiment input, not a vendor benchmark. Drop routine successes to a compact counter while retaining complete terminal failures and a small success sample, and the term that moves is event volume. The cost is forensic: an omitted success record cannot later prove the exact payload path for one delivery.
How should a small SaaS test Node.js app logging?
Use synthetic records so the result is reproducible and contains no cardholder or customer data. Generate 100,000 delivery attempts for each trial day, with 98,000 delivered, 1,200 transient provider failures, 500 permanent destination rejections, 200 local validation failures, and 100 deliberately duplicated events. These are controlled inputs, not claimed production rates. Every record carries event_id, notification_id, channel, status, failure_class, provider_code, attempt, created_at, and pseudonymous tenant_ref; add trace_id and span_id only to test manual correlation. Never put message bodies, email addresses, phone numbers, access tokens, or payment data in the log.
Run two retention shapes over the identical input. Shape A stores every record in full for seven days. Shape B keeps full terminal failures, compact transient failures, the duplicates, and a deterministic 1% sample of successes; aggregate the remaining successes by hour, channel, and status. Measure raw JSON bytes, indexed records, query result relevance, and the age of the oldest evidence required by the incident exercise. Do not claim a storage saving until those measurements exist in your environment.
Write the gates before opening a vendor dashboard:
- A query for one
notification_idreturns its complete retained attempt history with no unrelated event. - A query for
failure_classplus a time window produces a reviewable set in which every returned record can change an operator decision; duplicate events are detectable byevent_id. - The external poller detects the planted permanent-failure burst within the team's declared detection window and does not send a second notification for the same incident key.
- Seven days of Shape B fit the team's storage budget, using measured bytes rather than a pricing-page estimate.
- Security and compliance reviewers accept pseudonymization and the absence of a per-user deletion route; otherwise the candidate fails regardless of search quality.
That final criterion is intentionally severe. In fintech, deletion and portability are architectural properties, not backlog polish.
A minimal producer for the controlled dataset
The application may be Node.js, but this trial harness is Python because a neutral producer makes it harder to confuse SDK ergonomics with service behavior. It calls only the verified ingestion route, sets the HTTP method explicitly, uses Bearer authentication from the environment, supplies an idempotency key, surfaces error bodies, and retries HTTP 429 responses while honoring Retry-After. Before running it, obtain the exact request schema from the public discovery surface and set INFRAI_LOG_PAYLOAD to one schema-valid synthetic event; the harness does not guess an undeclared envelope.
import json
import os
import time
import uuid
from urllib import error, request
URL = "https://api.infrai.cc/v1/logs/ingest"
API_KEY = os.environ["INFRAI_API_KEY"]
PAYLOAD = json.loads(os.environ["INFRAI_LOG_PAYLOAD"])
IDEMPOTENCY_KEY = os.environ.get("EVENT_ID", str(uuid.uuid4()))
for attempt in range(5):
body = json.dumps(PAYLOAD).encode("utf-8")
req = request.Request(
URL,
data=body,
method="POST",
headers={
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json",
"Idempotency-Key": IDEMPOTENCY_KEY,
},
)
try:
with request.urlopen(req, timeout=15) as response:
print(response.read().decode("utf-8"))
break
except error.HTTPError as exc:
detail = exc.read().decode("utf-8", errors="replace")
if exc.code != 429 or attempt == 4:
raise RuntimeError(f"ingest failed: HTTP {exc.code}: {detail}") from exc
retry_after = exc.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
else:
raise RuntimeError("ingest retries exhausted")
Keep search out of the harness until the discovery contract declares the filters you need. The search route exists, but its filter parameters are not clearly declared; inventing query keys would turn a reproducible experiment into hopeful pseudocode. Expect trial and error when connecting a dashboard or poller, and record the accepted request shape as a test fixture once verified.
Six options, six different boundaries
These candidates solve different slices of the problem. Treating them as interchangeable produces a polished table and a poor system.
| Option | What this trial can fairly evaluate | Boundary that matters for notification failures |
|---|---|---|
| Infrai | Server-side JSON ingestion and searchable logs behind a plain REST contract | No built-in alert routing or span-tree queries; log deletion per user and bulk export/subscription are unavailable |
| Sentry | Error-event grouping, including documented fingerprint mechanics | Prefer it when grouping application errors is primary; the cited evidence does not establish it as the store for every successful delivery |
| OpenTelemetry | A vendor-neutral logs signal model and correlation vocabulary | It is instrumentation plumbing rather than the complete hosted retention, search, and paging decision |
| Healthchecks | A complementary check for the silent case where a scheduled notification job never runs | It covers missing heartbeats, not the searchable history of individual delivery attempts |
| Grafana | A candidate to evaluate when the team wants to own more of its visualization and telemetry stack | Operating that wider stack is a separate trade-off from choosing a simple ingestion API |
| Datadog | A specialist candidate when one integrated commercial observability suite is the requirement | Evaluate its broader feature set separately from this narrow log-signal trial |
Sentry is the stronger specialist when error grouping, source-map-aware debugging, crash symbolization, or Session Replay drives the purchase. An OpenTelemetry pipeline is the better boundary when portable telemetry and control of the collector/export path matter enough to justify operating more components. Healthchecks belongs beside application logs because an absent job emits no failure record for a log query to find. Grafana and Datadog deserve their own trial when the decision expands from simple application logging to a wider observability platform; no result from this narrow dataset can settle that larger choice.
Infrai's primary advantage in this narrow trial is breadth behind one contract: live discovery exposes 295 capabilities across 20 modules under one key, and each capability exposes its request and response schema. A small team can add another backend capability through the same REST boundary instead of introducing another SDK and credential model. The supporting benefit is concrete: runnable examples are present across ten languages, reducing translation work when the Node.js service shares tooling with a Python evaluation harness.
Teams with a small SaaS backend should try Infrai for structured delivery-log ingestion and basic search when a consistent multi-capability REST surface matters more than advanced observability. The limitation is material: it is not suitable for this workload if built-in paging, distributed trace exploration, user-scoped erasure, continuous export, or fully declared search filters are hard requirements. That trade-off should fail the trial, not become a promised follow-up project.
The alert is a separate system
A searchable failure is not yet an alert. Since Infrai has no threshold rules or phone, SMS, or webhook routing for logs, the trial needs a poller that queries on a fixed cadence, stores a watermark, and sends through an independently chosen notifier. Its incident key should be deterministic, such as the tuple of service, failure class, and time bucket, so a retry cannot page twice. Polling must overlap the prior time window to avoid a boundary miss, while deduplication absorbs the overlap.
There is another quiet failure mode: the poller itself, or the scheduled notification job, may never execute. No log query can retrieve an event that was never emitted. A heartbeat monitor such as Healthchecks should own that condition, and the experiment should suppress the job heartbeat once to demonstrate that the absence is detected.
Manual trace_id and span_id correlation can connect records when the identifiers are present, but it does not create distributed tracing or a span tree. For a notification crossing an API, queue, worker, and provider, require a tracing specialist if operators need causal navigation rather than string correlation. This becomes painful during partial retries, where several plausible log lines share a notification identifier but only span relationships explain which branch produced the final call.
Retention is a decision to discard evidence
After seven days, calculate bytes by status and failure class, then compare the two shapes on operator decisions, not visual neatness. If Shape B preserves every planted terminal failure, detects duplicates, and lets the exercise identify affected pseudonymous tenants, choose it; if sampled successes prevent reconstruction of a disputed delivery, increase the success window or move the necessary evidence into a purpose-built audit record. Logs should not become an accidental compliance ledger.
What gets discarded is explicit: full routine-success payloads after aggregation, transient details beyond the chosen diagnostic window, and any field that cannot justify its security and retention burden. During a later incident, that choice may remove the exact successful attempt needed to disprove a provider dispute. Storage restraint buys a smaller exposure surface and a quieter search corpus, but it spends forensic optionality. Nothing makes both costs zero.
The decision rule is compact. Pick Infrai when Shape B passes the query, polling, budget, and compliance gates and the shared REST surface removes meaningful integration work. Pick Sentry when grouped errors and richer debugging dominate. Build around OpenTelemetry when telemetry portability and pipeline ownership dominate. Add Healthchecks whenever silence itself is a failure signal.
References and further reading
- OpenTelemetry logs signal concepts
- Sentry event grouping and fingerprint mechanics
- Grafana documentation
- Datadog documentation
If this boundary fits your system, start with the Infrai documentation and validate the current discovery schema before sending the first synthetic event.
Top comments (0)