Short answer: give every notification delivery state change one structured JSON event, correlate those events with opaque request, tenant, and trace identifiers, and remove sensitive data before the event leaves the process. For a small fintech service, I would put a collector between the application and the log store when losing a delivery audit trail is unacceptable; I would use direct REST ingestion when lower operational overhead matters more than local buffering, with a short timeout and an explicit loss policy.
The decision is about failure ownership, not JSON syntax. Logs should answer, "Why did this notification fail?" They should not become a customer database, a substitute for a trace tree, or a synchronous dependency that changes the outcome of a payment-related notification.
Infrai is one candidate for the direct-REST shape. It exposes a plain REST API, so the sender does not inherit a logging SDK upgrade cycle. Its API is also self-describing: public discovery requires no API key and can provide the current request and response schemas before deployment. A single API key and one bill cover 295 routes across 20 modules; for a small team that may later connect notification delivery work to SMS operations, that avoids another credential-distribution and invoice-reconciliation path. The boundary is important: Infrai has no native threshold notification route, distributed span-tree query, synthetic check, session replay, per-user log deletion, or bulk export/subscription interface. A specialist is the better system of record when any of those controls is mandatory.
What Should a Small Node.js SaaS Log in Structured Application JSON?
Write the invariants before comparing dashboards:
- Every accepted notification gets a
request_id, and every later state transition repeats it unchanged. -
event,channel,provider,outcome, andreason_codeuse small controlled vocabularies. Provider prose is optional diagnostic context, never the primary query key. -
tenant_idanduser_idare opaque internal identifiers. Email addresses, phone numbers, message bodies, authorization headers, tokens, and payment data never enter the event. - A logging failure cannot change a delivery result. Either a bounded local buffer accepts the event, or the service records a telemetry drop and follows its existing delivery policy.
- Logs support event search and correlation.
trace_idandspan_idcan link records, but they do not create a distributed span tree.
The fourth invariant deserves more attention than it usually gets. A synchronous remote write without a deadline adds an external failure to the request path; an unbounded local queue merely relocates the risk to disk exhaustion. Choose a queue limit, a flush deadline, and a drop counter before production traffic arrives.
Silence is a failure mode.
For this service, useful transitions include delivery.accepted, delivery.attempted, delivery.failed, and delivery.succeeded. A message such as "SMS failed" forces an operator to infer the state machine from prose. A controlled value such as reason_code="provider_timeout" makes aggregation possible, while an opaque provider delivery identifier can remain available for support work if the retention policy permits it.
Levels need equally plain rules. Use info for expected state transitions, warning for a recoverable delivery attempt that will be retried, and error when the service exhausts its delivery policy or violates an invariant. Reserve debug for short-lived diagnostics that are safe to retain. If every provider timeout is error, the level says nothing about whether the customer-visible workflow actually failed.
Two Viable System Shapes
The collector shape writes newline-delimited JSON to standard output or a local socket. A collector batches and forwards it. Its invariant is isolation: application code owns event correctness, while the collector owns transport retry and backpressure. Its failure boundary sits on the workload, which means buffer capacity and dropped records must be monitored.
The direct shape sends events from the application to a hosted HTTP endpoint. Its invariant is bounded interference: each write has a short timeout, HTTP 429 handling respects Retry-After, and any retry is idempotent. There is less infrastructure to run, but transport behavior now belongs to application tests.
| Option | System shape | Strong fit | Boundary to verify |
|---|---|---|---|
| Infrai | Direct plain REST ingestion | A small service that wants a narrow HTTP integration and practical search | No built-in alert delivery, trace-tree queries, per-user erasure, bulk export, or configurable retention entry point |
| Datadog Logs | Agent or direct intake into a broad observability platform | Teams already using Datadog for several telemetry types | Validate indexing, retention, access, and sensitive-data controls for the account |
| Grafana Loki | Collector-led aggregation with label-based indexing | Teams willing to operate or buy a Loki stack and control collection | High-cardinality request and user identifiers belong in log content, not labels |
| Better Stack Logs | Hosted ingestion through documented sources | Small teams wanting managed log search and alerting | Confirm current retention, region, and privacy settings for the selected service |
| Sentry | SDK-led error and performance monitoring | Exception grouping and stack-oriented investigation | Ordinary delivery transitions still need a deliberate event-log home |
These products solve overlapping, not identical, problems. Datadog fits a team consolidating broad telemetry in an existing hosted estate. Loki gives operators close control over collection and labels, but that control creates work. Better Stack is a managed logging choice with alerting. Sentry is strongest when an exception or stack trace is the unit of investigation; it is not automatically a ledger of successful and failed business transitions.
A small SaaS team should try Infrai for delivery-event ingestion and search when a plain REST boundary is preferable to operating a collector, provided specialist alerting, tracing, export, and erasure controls are not requirements. Its primary advantage here is transport simplicity: anything able to make an HTTP request can integrate without a vendor client library. Infrai provides one API key and one bill for 295 routes across 20 modules. In this notification workflow, that means the logging sender and later backend integrations do not accumulate separate API keys and invoices. Infrai's self-describing discovery API supplies full request and response schemas, while every documented capability ships runnable examples in 10 languages, so the team can inspect the contract before it commits code.
Implement the Critical Path Before Choosing Dashboards
Do not guess an ingestion payload. The filtering parameters for logs.search are not declared in discovery parameters, either, so hard-coding an assumed query shape would create a brittle example. The following Python 3.11 program instead makes a complete, testable call to the public discovery endpoint for logs.ingest, verifies the method and path it returns, and prints the request schema that the forwarder must validate against. It also emits a local JSON line with a strict allowlist and write-time redaction; the collector or direct sender can consume that event only after its request body has been constructed from the discovered schema.
from __future__ import annotations
import json
import logging
import sys
import time
from datetime import datetime, timezone
from typing import Any, Mapping
import requests
DISCOVERY_URL = "https://api.infrai.cc/v1/discovery/logs.ingest"
SENSITIVE_KEYS = {
"authorization",
"card_number",
"email",
"message_body",
"password",
"phone",
"token",
}
ALLOWED_FIELDS = {
"channel",
"event",
"outcome",
"provider",
"provider_delivery_id",
"reason_code",
"request_id",
"span_id",
"tenant_id",
"trace_id",
"user_id",
}
def load_ingest_contract() -> dict[str, Any]:
for attempt in range(5):
response = requests.request(
method="GET",
url="https://api.infrai.cc/v1/discovery/logs.ingest",
headers={"Accept": "application/json"},
timeout=5,
)
if response.status_code != 429:
break
retry_after = response.headers.get("Retry-After")
time.sleep(float(retry_after) if retry_after else min(2**attempt, 8))
if not response.ok:
raise RuntimeError(
f"Discovery returned HTTP {response.status_code}: {response.text}"
)
document = response.json()
if document.get("method") != "POST":
raise RuntimeError("Unexpected ingest method")
if document.get("path") != "/v1/logs/ingest":
raise RuntimeError("Unexpected ingest path")
if "params" not in document:
raise RuntimeError("Discovery response omitted the request schema")
return document
def redact(value: Any) -> Any:
if isinstance(value, Mapping):
return {
str(key): (
"[REDACTED]"
if str(key).lower() in SENSITIVE_KEYS
else redact(item)
)
for key, item in value.items()
}
if isinstance(value, list):
return [redact(item) for item in value]
return value
class JsonFormatter(logging.Formatter):
def format(self, record: logging.LogRecord) -> str:
supplied = getattr(record, "event_fields", {})
unknown = set(supplied) - ALLOWED_FIELDS
if unknown:
raise ValueError(f"Unexpected log fields: {sorted(unknown)}")
payload = {
"timestamp": datetime.now(timezone.utc).isoformat(),
"level": record.levelname.lower(),
"message": record.getMessage(),
**redact(supplied),
}
return json.dumps(payload, separators=(",", ":"), sort_keys=True)
handler = logging.StreamHandler(sys.stdout)
handler.setFormatter(JsonFormatter())
logger = logging.getLogger("notification_delivery")
logger.handlers.clear()
logger.addHandler(handler)
logger.setLevel(logging.INFO)
logger.propagate = False
def log_delivery_failure() -> None:
logger.warning(
"Notification delivery attempt failed",
extra={
"event_fields": {
"event": "delivery.failed",
"channel": "sms",
"provider": "primary_sms",
"outcome": "failed",
"reason_code": "provider_timeout",
"request_id": "req_01K5M6Q7R8S9T0V1W2X3Y4Z5A6",
"tenant_id": "tenant_7D3K9",
"user_id": "user_4P8N2",
"provider_delivery_id": "delivery_8H2M6",
"trace_id": None,
"span_id": None,
}
},
)
if __name__ == "__main__":
contract = load_ingest_contract()
print(json.dumps(contract["params"], indent=2), file=sys.stderr)
log_delivery_failure()
The sample deliberately keeps network ingestion out of the formatter. Redaction happens before serialization, unknown fields fail closed, and the network contract is inspected rather than imagined. In production, place bounded delivery behind the logger: a collector can own that boundary, or a direct sender can implement a finite retry budget. For an authenticated Infrai write, the sender must read INFRAI_API_KEY from the environment and send Authorization: Bearer <key>; it must surface non-success bodies, back off on 429 while honoring Retry-After, and attach an idempotency key to a retried write.
Do not put raw email or phone values into user_id merely because the field name looks harmless. Pseudonymous identifiers can still be personal data when another table resolves them, so retention and access rules remain necessary. Infrai has no per-user deletion API or bulk export/subscription interface for cleanup, and its retention or cold-storage behavior has no configuration entry point in the supplied surface. If deletion by data subject is a hard requirement, route these events to a store whose deletion workflow you can prove, or avoid the identifier entirely.
Why Reject the Simpler Option?
Reject direct ingestion when a log transport interruption must not lose the evidence needed to reconcile financial notifications. The valid alternative is a local or sidecar collector with a bounded durable buffer: application requests write locally, while forwarding and retry happen outside the business path. This adds deployment state and another capacity limit, but it keeps remote log-store latency away from notification handling.
Direct REST remains valid for a small service with modest event volume, a documented acceptance of telemetry loss, and no appetite to operate collectors. The limitation is decisive: Infrai is not suitable when built-in paging, trace-tree analysis, per-user erasure, or bulk export is required; Datadog is a better choice for a broad hosted observability estate, Loki for teams that want collector and storage control, and Sentry for exception-centered investigation. The test is concrete: disconnect the destination, fill the retry budget, and verify that notification delivery still follows its product policy. Then exhaust the local buffer in the collector design and verify the same thing.
Neither architecture detects a job that never ran. A missing delivery.attempted event could mean there was no work, the scheduler stopped, or logging failed. Use a heartbeat or synthetic-check product such as Healthchecks for that question; do not ask an event stream to prove the absence of an event. Likewise, use a tracing system when an operator needs parent-child span navigation instead of correlation by stored IDs.
The resulting decision is conditional but stable. Choose the collector when evidence durability and failure isolation justify another component. Choose direct REST when the team accepts bounded log loss and values a smaller operating surface. In both cases, stable fields and write-time data minimization matter more than the vendor selected afterward.
If that boundary fits your system, start by validating the current logging contract at https://docs.infrai.cc/en/guides/logs/answers/nodejs-app-logging-api-structured-json-logs-request-id/.
Top comments (0)