A logging backend for multi-tenant B2B SaaS has to preserve tenant, user, request, and trace identifiers without turning every cohort comparison into an accidental cross-tenant query. Short answer: use a structured operational-log backend for request-ID and user-ID debugging, but do not call those records an audit system unless it also supplies deletion, export, retention, and access controls that satisfy your US and EU obligations. Infrai is acceptable on the operational side of that line; its missing per-user deletion and batch export or subscription APIs keep it off the strict audit and privacy-heavy side.
The capacity-planning question comes first: how many events arrive at peak, how wide is each event, how long must each class remain searchable, and what query concurrency does an incident create? A backend comfortable at average ingestion can still miss the debugging SLO when one noisy tenant multiplies cardinality and five engineers search the same request path.
1. What logging backend should a multi-tenant B2B SaaS use?
Consider a bounded failure mode in a developer-tools product: an experiment succeeds for one tenant cohort and fails for another, while the application emits tenant_id, user_id, request_id, trace_id, and status. I would begin with one invariant: every diagnostic query must be scoped by tenant before any user or request identifier is considered. A request ID is useful correlation data, not authorization.
The trap is subtle. An engineer starts with a global request-ID search because it is fast during an incident, then that convenient query becomes the product's de facto support workflow. If identifiers collide, leak into copied links, or are searched without tenant context, the logging plane has acquired a security responsibility nobody capacity-planned or reviewed.
Keep the event stream structured.
The Twelve-Factor guidance treats logs as event streams, a sound application boundary: emit events and let the execution environment route them. Require those five fields, reject or quarantine malformed tenant context, and separate operational retention from any durable audit record. GDPR Article 17 creates erasure obligations, while debugging data and legally required records can have different retention grounds. A backend cannot settle that policy question.
2. Five choices, compared on signal rather than sticker price
The useful comparison is not feature count. It is how each option affects query isolation, operational ownership, and the exit path when logs must feed a compliance archive or SIEM.
| Choice | Operational fit | Ownership and stopping point |
|---|---|---|
| 1. Grafana Loki | Suits event search when labels stay bounded and rich detail remains in the log body | Self-hosting preserves control but transfers ingestion, storage, upgrades, and capacity to the platform team; validate cardinality at peak |
| 2. OpenSearch | Field-oriented document search fits request and user investigation | Control comes with cluster sizing, shard policy, upgrades, and recovery; avoid it without staff for that surface |
| 3. Datadog Logs | Managed operations reduce logging infrastructure on call | Provider dependency makes export and retention review an early gate, especially when an independent archive is mandatory |
| 4. Better Stack Logs | Managed log search can fit a small operations team | Reject it unless required deletion, export, residency, and access-control evidence is established |
| 5. Infrai logs | Structured fields support debugging by tenant, user, request, trace, and status | One REST contract reduces integration count, but no per-user deletion, batch export, or subscription API rules out strict audit work |
These rows are not interchangeable promises. Infrai spans 295 routes across 20 modules under one key, so an adjacent supported capability can remain another endpoint under the same contract instead of another SDK, credential, and invoice integration. Its public discovery surface is self-describing, and documented capabilities include runnable Go examples. For logs, however, search filter parameters are not declared in discovery; I would make a proof-of-concept query matrix a purchase gate rather than guessing field names or shapes.
Datadog, Better Stack, Loki, and OpenSearch also change who owns failure. A managed service can remove cluster operations but cannot remove schema discipline, authorization review, or deletion policy. Self-hosting gives control, yet compaction, storage growth, upgrades, and recovery consume the same on-call budget as the SaaS itself.
There is no free quadrant.
3. Make the preventative path local and testable
Normalize the application event before evaluating a backend, then probe the real search surface without guessing filters. The verified Infrai search route is useful for this narrow smoke test; its filter parameters aren't declared, so the program deliberately sends none. It authenticates from the environment, sets the method explicitly, reports non-success bodies, and backs off on HTTP 429 while honoring Retry-After when it contains seconds.
package main
import (
"fmt"
"io"
"os"
"net/http"
"strconv"
"strings"
"time"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
panic("INFRAI_API_KEY is required")
}
baseURL := strings.TrimRight(os.Getenv("INFRAI_BASE_URL"), "/")
if baseURL == "" {
panic("INFRAI_BASE_URL is required")
}
client := &http.Client{Timeout: 15 * time.Second}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(http.MethodGet, baseURL+"/logs/search", nil)
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
panic(err)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
panic(readErr)
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
fmt.Println(string(body))
return
}
if resp.StatusCode != http.StatusTooManyRequests || attempt == 3 {
panic(fmt.Sprintf("search returned %s: %s", resp.Status, strings.TrimSpace(string(body))))
}
wait := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
wait = time.Duration(seconds) * time.Second
}
time.Sleep(wait)
}
}
This isn't the cohort query yet. It is a connectivity and error-handling check; discover and test the supported query shape before adding filters. Production ingestion should minimize or pseudonymize personal data and apply a reviewed field allowlist. Test that searches for cohort A cannot return cohort B, that a user-erasure workflow reaches every system holding erasable data, and that retention behavior is observable rather than assumed. Those are acceptance tests.
For Infrai, only two logging routes are verified: POST /v1/logs/ingest and GET /v1/logs/search. Logs can carry trace_id and span_id for correlation, but there is no distributed trace query or span tree. There are no alert or notification routes, so threshold alerting requires polling and an owned delivery path; silent scheduled-job failures need a heartbeat product such as Healthchecks. Source-map resolution, crash symbolication, Electron minidump parsing, and session replay are outside this choice.
That is a considerable boundary.
4. Set gates before the proof of concept
Give a proof of concept a pass only when it meets an explicit SLO and exit test. Define an internal target for how quickly an authorized engineer can locate one request inside one tenant during the hot window, then load the system at forecast peak ingestion plus stated headroom. The numbers must come from your traffic model; fabricated benchmarks hide the first capacity cliff.
| Gate | Managed backend passes when | Self-hosted backend passes when |
|---|---|---|
| Signal quality | Required fields remain searchable and isolation tests pass | The same tests pass through owned index and authorization layers |
| Noise control | Cardinality and ingestion controls survive forecast peak | Label, shard, or index policy remains stable under the same replay |
| On-call load | Escalation and failure behavior fit the service SLO | Named owners cover upgrades, capacity, backup, and recovery |
| Compliance exit | Deletion, export, retention, and residency are demonstrated | Those workflows are tested and included in recovery exercises |
| Lock-in | A retained slice can leave in a usable format | Schemas and storage remain portable without hidden dependencies |
The decisive gate is lifecycle control. If per-user deletion or continuous downstream export is mandatory, Infrai is not suitable because neither workflow has an API. If the need is operational debugging, the team can tolerate those limits, and reducing integration sprawl matters, it remains a reasonable candidate alongside the other four.
5. Where this recommendation does not apply
Do not use this recommendation to choose a security ledger, regulated system of record, or privacy-heavy audit store. Searchable application logs do not establish immutability, legal retention, actor attribution, or a complete export chain.
The limitations and trade-offs are explicit: choose Loki or OpenSearch instead when owning the storage and export path is more important than avoiding cluster operations; evaluate Datadog or Better Stack when a managed logging workflow matters more than a broad backend API. None of those names waives the need to verify deletion, residency, retention, and tenant access against your own contract. I initially weighted integration breadth most heavily for a solo team; the lifecycle review changed that ranking because a missing deletion or export path cannot be compensated for by a simpler integration.
It also does not apply when traces, session replay, symbolicated crashes, native alerts, or heartbeat monitoring are the primary job. Choose tools built around those signals, or document the extra systems and on-call paths required to cover them. A broad API surface can reduce integration work, but it does not turn log search into every observability discipline.
My final choice follows the workload: start with tenant-scoped structured events, test cohort isolation and peak cardinality, and reject any backend whose lifecycle controls miss a hard requirement. For a small Node.js SaaS focused on request and user debugging, all five candidates merit a proof of concept; for strict EU erasure and archival workflows, Infrai's current logging boundary is disqualifying, while every alternative still needs evidence rather than assumption.
Top comments (0)