TL;DR: choose the backend you can replace without changing application code or weakening tenant boundaries. For a multi-tenant B2B service that searches nightly pipeline logs by request ID and user ID, emit structured events to stdout, keep the query contract behind a tiny adapter, route EU and US tenants deliberately, and prove rollback with a dual-write replay before committing. Search speed matters. A reversible cutover matters more.
The label "audit-ish" is a warning. Operational logs can help reconstruct an import, but they should not quietly become an authorization ledger or an immutable compliance archive. The useful design question is narrower: can an engineer find one tenant's failed nightly run quickly, while the team can still retreat from a bad backend migration?
What logging backend should a multi-tenant B2B SaaS use?
A backend is rollback-safe when the application owns its event shape, routing rules, and query semantics. Storage should receive events, index approved fields, enforce retention, and answer a small set of searches. It should not define the only representation of tenant, request, or user identity.
Start with the failure you need to reverse. Imagine a nightly pipeline release changes an index template. Fresh events arrive, but request-ID searches return incomplete results. If producers emit a vendor-specific payload and every internal tool calls a proprietary query language, rollback now means changing application code under pressure. If producers emit a stable envelope and a gateway translates the query, rollback is a routing change plus validation. That constraint changes my choice.
I would test four things before accepting any backend:
- Can the same canonical event be written to two destinations without mutating its meaning?
- Can a tenant-scoped request or user search be expressed through our own interface?
- Can retention and deletion be applied to a tenant or data subject without searching arbitrary message text?
- Can traffic return to the previous destination while the new one is repaired or reindexed?
No feature matrix answers those questions. A migration drill does.
This pattern has limits. It is not suitable when logs must serve as the system's formal, tamper-evident audit record; use a separately governed audit store for that requirement. Dual writing also consumes more ingestion capacity and creates a reconciliation job. Teams unable to operate that temporary overhead should prefer an offline export-and-replay test, accepting that it proves less about live delivery. The adapter itself is another component to own. I accept that trade-off because it keeps backend-specific query syntax out of every service and internal tool.
The smallest useful event contract
The Twelve-Factor App treats logs as event streams and says the application should not concern itself with routing or storage. That separation is a good producer boundary. For this workload, though, unstructured text is too weak: tenant scoping and erasure become dependent on parsing prose. Emit one JSON object per event, then let the execution environment route stdout.
type Region = "eu" | "us";
type PipelineLogEvent = {
schemaVersion: 1;
occurredAt: string;
level: "info" | "warn" | "error";
tenantId: string;
region: Region;
pipelineRunId: string;
requestId: string;
userId?: string;
eventName:
| "pipeline.started"
| "record.rejected"
| "pipeline.completed";
message: string;
attributes?: Record<string, string | number | boolean>;
};
function writePipelineEvent(event: PipelineLogEvent): void {
process.stdout.write(`${JSON.stringify(event)}\n`);
}
writePipelineEvent({
schemaVersion: 1,
occurredAt: new Date().toISOString(),
level: "warn",
tenantId: "tenant_7f2",
region: "eu",
pipelineRunId: "run_20261010_0042",
requestId: "req_01J9Y7Q4K2",
userId: "user_184",
eventName: "record.rejected",
message: "Input record failed schema validation",
attributes: { source: "nightly-crm-import", recordNumber: 381 }
});
Keep secrets, access tokens, raw payloads, and direct personal details out of this envelope. An opaque internal user identifier is searchable without copying a name or email into every event. More importantly, tenantId is mandatory. A request ID is correlation data, not an authorization boundary; the query layer must require both.
The interface can stay equally boring:
type LogQuery =
| { tenantId: string; requestId: string; from: Date; to: Date }
| { tenantId: string; userId: string; from: Date; to: Date };
type LogRecord = {
occurredAt: string;
pipelineRunId: string;
requestId: string;
eventName: string;
message: string;
};
interface LogSearchBackend {
search(query: LogQuery): Promise<LogRecord[]>;
}
async function searchTenantLogs(
backend: LogSearchBackend,
query: LogQuery
): Promise<LogRecord[]> {
if (!query.tenantId) throw new Error("tenantId is required");
if (query.to.getTime() <= query.from.getTime()) {
throw new Error("invalid time range");
}
return backend.search(query);
}
That union supports two access paths and no free-form pass-through. It prevents callers from smuggling a backend query into the application layer. It also gives contract tests a stable target.
Small is good.
Build the rollback path before the migration
Dual writing is useful only when you can detect disagreement. During a migration, send the same canonical event to the current and candidate destinations at the routing layer. Do not make each application process manage two client libraries. Record delivery outcomes separately, because a successful enqueue is not proof that an event became searchable.
Create a fixed verification corpus. Include one successful pipeline run, one rejected record, one event without userId, two tenants sharing a deliberately identical requestId, and events on both sides of a retention boundary. The identical request ID is the trap: a query that omits tenantId may look correct in ordinary fixtures and expose records across tenants.
Benchmark the workflow you operate. Measure ingestion-to-search delay for verification events, query latency for bounded time windows, missing and duplicate event counts, and time needed to switch reads and writes back. Report distributions instead of one lucky run. I would reject a candidate if rollback requires redeploying every Node.js service, even if its isolated query benchmark wins. That is an explicit trade: a thin adapter adds code, but confines migration risk to one boundary.
A staged change has three independent switches: write destination, read destination, and shadow verification. Move one at a time. First mirror writes. Next compare known queries. Then move internal reads. Keep the old path receiving data until the rollback window closes under your retention and recovery policy.
Do not call the exercise complete after a dashboard returns results. Force the reverse switch. Time it. Verify that on-call engineers can do it with documented permissions, and confirm that the old destination contains the events required for the tested window.
Region, deletion, and the limits of audit-ish logs
EU and US placement should be a routing decision made from authoritative tenant metadata, not guessed from an IP address or accepted from a client-supplied log field. The event's region value is useful for validation, but the collector should reject or quarantine a mismatch rather than silently ship it across the wrong route. Keep regional destinations and credentials separate.
GDPR Article 17 establishes a right to erasure and lists exceptions. That makes deletion semantics a design input, not a cleanup script promised for later. Maintain a lookup path from the stable user identifier to indexed events, define which fields can contain personal data, and test deletion against primary storage, replicas, and the documented backup lifecycle. Legal classification and retention policy need qualified review; calling a stream "audit-ish" does not decide either one.
Operational logs also have a different integrity goal from a true audit trail. They are optimized for diagnosis and bounded search. If the business needs evidence of who changed permissions or approved an action, define a separate append-oriented audit event model with explicit access controls, integrity requirements, and retention. Do not infer those records from mutable diagnostic messages.
What I would change at scale
I would resist adding fields until a real query needs them. Every indexed dimension increases schema and operational surface area, while arbitrary attributes tempt teams to log payload fragments. Promote a field into the contract only after naming its owner, allowed values, retention impact, and deletion behavior.
At higher volume, partitioning and index layout become workload-specific. Keep the public query contract unchanged while storage adapters evolve underneath it. Time-bounded searches remain mandatory. Very large tenants may need isolated capacity or partitions, but tenant placement should stay in routing metadata so moving one tenant does not alter producer code.
The final selection is a proof, not a brand decision. Use representative nightly-pipeline events and score candidates on tenant isolation, regional routing, deletion execution, bounded request and user searches, ingestion visibility, and measured rollback time. Choose only after the team has migrated forward and back in a non-production environment. The backend that survives that drill with the least application change is the better fit for this system.
Further reading
- The Twelve-Factor App, "Logs": https://12factor.net/logs
- GDPR Article 17, "Right to erasure": https://gdpr-info.eu/art-17-gdpr/
Top comments (0)