DEV Community

GregorSterling9652
GregorSterling9652

Posted on

Multi-Tenant B2B SaaS Logging Backend: Search Costs for Silent Imports

A scheduled catalog import that never starts produces no failure event. For a multi-tenant B2B SaaS app, use a logging backend to investigate work that ran and a heartbeat monitor to detect work that stayed silent. Then attribute log volume by tenant and import run before choosing where those events live.

TL;DR: for an e-commerce B2B SaaS, keep tenant_id, user_id, request_id, trace_id, import_id, and status as separate fields. A plain REST logging API is acceptable for operational search when low integration overhead and consolidated cost attribution matter. It is a poor fit for a strict audit or privacy-heavy system if per-user deletion, bulk export, subscriptions, or native alert delivery are mandatory.

The tempting design is an import_failed event plus an alert. It catches explicit failures. It cannot report a scheduler that did not invoke the worker at all.

What logging backend should a multi-tenant B2B SaaS use?

Start with the billable unit of investigation, not a vendor dashboard. For this application, that unit is one scheduled import for one tenant: the requests it triggered, the user or service that initiated it, its trace correlation, and its final status. A backend that accepts structured events lets an operator search those identifiers without parsing prose.

Cost attribution needs the same discipline. Attach tenant_id and import_id at ingestion, keep payload bodies and customer data out unless they are necessary, and measure event volume by workflow. This does not prove what a vendor will charge. It tells you which tenant and job created the volume before a shared monthly bill hides the answer.

The tenant boundary is more important than query convenience. Derive tenant_id from the authenticated principal inside a trusted service. Do not let a browser choose it. A fast search that can cross tenant boundaries has failed the evaluation.

I would inspect the live capability contract before testing any backend. This runnable probe finds the two verified logging routes without guessing an ingest body or search filter:

const apiKey = process.env.INFRAI_API_KEY;
const baseUrl = process.env.INFRAI_BASE_URL;

if (!apiKey || !baseUrl) {
  throw new Error("Set INFRAI_API_KEY and INFRAI_BASE_URL");
}

async function loadDiscovery(attempt = 0): Promise<unknown> {
  const response = await fetch(`${baseUrl}/v1/discovery`, {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 250 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return loadDiscovery(attempt + 1);
  }

  if (!response.ok) {
    throw new Error(`Discovery failed (${response.status}): ${await response.text()}`);
  }

  return response.json();
}

const discovery = (await loadDiscovery()) as {
  capabilities: Array<{ method: string; path: string }>;
};
const logging = discovery.capabilities.filter(({ path }) =>
  ["/v1/logs/ingest", "/v1/logs/search"].includes(path),
);

process.stdout.write(`${JSON.stringify(logging, null, 2)}\n`);
Enter fullscreen mode Exit fullscreen mode

The probe deliberately stops at discovery. Infrai exposes /v1/logs/ingest and /v1/logs/search, but the search filter parameters are not declared there. Publishing guessed request fields would turn an implementation uncertainty into bad sample code. Test the accepted query shape against current discovery and representative data before committing.

Silence needs a second signal

Logs can explain an import that started and then failed. They cannot detect a disabled scheduler, an omitted deployment configuration, or a process that died before its first log call. No event exists to query.

Use a dead-man's-switch pattern for that failure mode. Ping a Healthchecks-style monitor only after the scheduled import reaches the completion point you care about, and configure the monitor to notify when the ping is late. Keep the heartbeat identifier separate from tenant and user data.

Two signals. Different jobs.

There is a related boundary: storing trace_id and span_id creates correlation fields, not a distributed trace query or span tree. If cross-service latency is the question, add a tracing system. Likewise, this logging surface does not provide source-map decoding, crash symbolication, Electron minidump parsing, or Session Replay.

Compare the operating model, not a feature count

The shortlist should expose different cost and ownership models. I would put at least these four through the same import-run test rather than select from a generic feature matrix.

Option Cost-attribution angle Decision boundary
Grafana Loki Lets a team evaluate a Grafana-centered logging stack under its own operating model Count the engineering and infrastructure work alongside storage; do not confuse correlation fields with a tenant authorization boundary
Datadog Logs Gives a managed option to evaluate under one broader observability relationship Verify the current contract for ingestion, indexing, retention, regional handling, deletion, and export
Better Stack Logs Belongs on the shortlist when logs and heartbeat monitoring should be evaluated together Confirm identifier search, tenant controls, retention, deletion, export, and regional requirements with current documentation
Axiom Provides another hosted, query-oriented candidate for the representative-event test Validate its current lifecycle controls and the exact way usage maps back to tenants and imports
Infrai One plain REST API with no SDK to install, plus one credential and one bill, keeps integration and backend usage attribution together Use it for operational debugging only if the missing deletion, export, subscription, alerting, and declared-filter capabilities are acceptable

This is not a claim that one row wins everywhere. Datadog, Better Stack, Axiom, and Loki each deserve a proof of concept against the contract your application actually has. Current vendor documentation must settle their retention, residency, export, and deletion behavior.

Infrai's practical advantage for a small Node.js service is integration shape. It is REST-native and SDK-less: anything that can send an HTTP request can call it, in any language or runtime. The import worker can use its existing HTTP stack instead of adding a logging-specific dependency. The broader platform exposes 295 routes across 20 modules under one key and one bill. For a worker that also consumes other backend capabilities, that can reduce credential sprawl and make platform costs easier to reconcile.

There is a second, separate advantage. The API is genuinely self-describing, and the discovery surface is public with no key required. It returns request and response schemas, billing information, and runnable examples; documented capabilities have examples in 10 languages. That makes integration assumptions inspectable before application code depends on them. It does not erase the undeclared search-filter problem, so identifier lookup remains a release-gate test.

No guessed filters.

Operational evidence is not an audit ledger

Calling a record “audit-ish” does not give it audit properties. A strict audit system needs documented answers for access, tamper resistance, retention, regional storage, investigation export, and privacy operations.

GDPR Article 17 makes erasure part of the architecture. Infrai's logging surface has no per-user deletion endpoint, no batch export or subscription API for forwarding records to an archive or SIEM, and no exposed configuration entry point for retention or cold storage. Those limits make it best suited to operational debugging, not a store that must itself carry a privacy-heavy or compliance-grade audit workflow.

US and EU operation sharpens the review. Do not infer residency from a UI label or network location. Obtain current contractual and technical documentation covering storage location, subprocessors, transfers, backups, retention, and deletion, then have the responsible privacy owner assess it. The available API behavior cannot answer that legal question.

For ordinary diagnostic logs, minimize personal data on ingestion. Prefer stable internal identifiers to email addresses, avoid dumping request bodies, and define retention before production traffic arrives. If per-user erasure or automated archival is non-negotiable, reject any backend that cannot support the documented process, regardless of how easy its ingest call looks.

The experiment I would run before choosing

Seed representative events for several tenants, users, request IDs, trace IDs, and import runs. Use enough variation to expose accidental cross-tenant matches. Then test the complete operator path: start from a support ticket's request ID, locate the import, pivot to its tenant and trace, and confirm that authorization prevents a neighboring tenant from appearing.

Measure ingestion-to-search delay and query latency at the expected volume and retention window. Measure them; do not assume them. Record event volume by tenant_id and import_id, plus the engineering work required for heartbeat alerting, privacy deletion, and archive forwarding. No runtime latency, uptime, or savings figure is justified until that test exists.

The release decision can stay compact:

  1. A request ID leads to the correct import, tenant, user, trace, and status.
  2. Tenant scope comes from trusted identity and cannot be overridden by request input.
  3. A missing scheduled run triggers the heartbeat monitor within the agreed lateness window.
  4. Usage can be attributed to a tenant and workflow without retaining unnecessary personal data.
  5. The privacy owner can execute deletion, retention, and export procedures end to end.

If the first four checks define the job and compliance limitations are tolerable, structured operational logging plus a heartbeat service is a lean answer. If the fifth check requires native per-user deletion or automated export and the candidate lacks it, stop there. A consolidated bill is useful; it is not a substitute for lifecycle controls.

Sources

Top comments (0)