TL;DR
- We wanted multi-tenant exception ingest on a two-container stack: one Rust binary + PostgreSQL 16. No Redis. No Kafka.
- Dashboard reads go through Postgres RLS with
SET LOCAL app.current_org_id. App bugs must not cross orgs on a shared DB. - Ingest is DSN → project_id, not the app RLS role. RLS is defense-in-depth for dashboard
/api/v1, not magic on every path.
I build Epure: exception-only error monitoring you can self-host. This week I am writing down the stack bet, not a migrate recipe. We already shipped the cheap VPS shape and the DSN swap. The question people ask next is why Rust plus Postgres row-level security instead of a bigger fleet or app-only filters.
Why Rust + Postgres RLS for exception ingest?
We optimized for a small blast radius on a shared Postgres: org-scoped dashboard reads enforced in the database, ingest scoped by DSN, two containers total.
Tenancy on a side-project VPS is easy to hand-wave. One engineer, one DB, WHERE org_id = ? in every handler. That works until a handler forgets the filter, a join leaks a neighbor row, or a reporting query skips the org check. On a multi-tenant exception store that is the whole product surface: other people's stack traces.
So we put the tenant key in Postgres. The dashboard role cannot see another org's rows even when application code is wrong. Rust keeps the binary one process: ingest HTTP, migrations, dashboard API. Postgres 16 holds events, issues, and the RLS policies. That is the whole product runtime besides your TLS edge.
If you run GlitchTip or self-hosted Sentry, the portable lesson is the same: decide where the tenant boundary lives before you scale the number of tenants on one database.
Threat model: who can read what if X fails?
Cookie and membership checks gate the dashboard. RLS is the backstop if a query forgets the org filter. Ingest trusts the DSN, not a client-supplied org id.
| Failure | Without RLS (app filters only) | With epure_app + RLS |
|---|---|---|
Handler drops WHERE org_id = …
|
Neighbor org rows can return | Policy still filters on app.current_org_id
|
| Client sends another org's UUID | App must reject; easy to miss on one route | Session sets org from membership, not from the body |
| Stolen session cookie for Org A | Sees Org A (expected) | Still only Org A; no cross-org via forgotten filter |
| Leaked ingest DSN for Project P | Writes land on P | Same: ingest is DSN-scoped; not the app RLS role |
Dashboard uses migrate/epure role by mistake |
Full table access | Wrong. epure is for migrations. App must use epure_app
|
What we are not claiming: RLS stops a stolen DSN, a stolen cookie inside the right org, or a bug in the ingest writer. Different paths, different trust boundaries.
Why not a shared DB without RLS, app-only filters, or a heavier stack?
App-only filters fail open on one missed query. Heavier stacks buy isolation with ops cost we refused on a two-container box.
| Option | Why it lost for this job |
|---|---|
| Shared DB, no RLS | Every read path must remember the org filter. One miss is a tenancy bug. |
| App-only org filters | Fine for a prototype. Weak as the only line when exceptions are the payload. |
| Separate database per customer | Strong isolation. Painful backups, migrations, and $5-class ops. |
| Full self-hosted Sentry (~65 containers) | Real product. Wrong unit when you only need exceptions ($5 VPS notes). |
| Redis / Kafka / extra workers | Useful for big fleets. We did not want them for exception ingest on one VPS. |
GlitchTip is a fair lighter alternative on Django-shaped stacks. We still chose Rust + Postgres RLS because we are building the ingest path and want the tenant boundary in the database for dashboard reads. A GlitchTip reader can still steal the role split and SET LOCAL pattern below.
Working sketch: three roles, SET LOCAL, policy shape
Three Postgres roles on one epure database. Dashboard: cookie → membership → SET LOCAL app.current_org_id → queries. Ingest: DSN → project_id. Migrations: epure only.
| Role / URL | Job |
|---|---|
epure via DATABASE_URL
|
Migrations / bootstrap on startup |
epure_ingest via EPURE_INGEST_DATABASE_URL
|
DSN-scoped event writes |
epure_app via EPURE_APP_DATABASE_URL
|
Dashboard /api/v1/* with RLS |
Compose injects all three for the app container. Missing migrate URL exits on start. Missing ingest URL yields storage errors on write. Missing app URL 500s the dashboard.
Dashboard request shape:
- Valid session cookie.
- Confirm the user is a member of the org.
-
SET LOCAL app.current_org_id = '<uuid>'on that transaction. - Run queries as
epure_app. Never trust a client-supplied org id for the session setting.
Ingest request shape:
- Authenticate with the project DSN.
- Resolve
project_id. - Write as
epure_ingest. Do not setapp.current_org_idfrom the client.
Illustrative policy shape (paraphrased from the docs pattern; pin your own names to migrations):
-- illustrative: dashboard role reads events only for the session org
CREATE POLICY events_org_isolation ON events
FOR SELECT
TO epure_app
USING (org_id = NULLIF(current_setting('app.current_org_id', true), '')::uuid);
-- FORCE ROW LEVEL SECURITY so table owners cannot bypass by accident
ALTER TABLE events FORCE ROW LEVEL SECURITY;
Events land in monthly partitions events_YYYY_MM. Project retention TTL is 14 / 30 / 90 days; old partitions drop. Issue aggregates remain after event TTL. That retention story is separate from RLS, but it is part of the same cheap-ops bet.
You do not need our binary to try the pattern. Any multi-tenant read API on Postgres can use a restricted role, SET LOCAL for the tenant key, and FORCE ROW LEVEL SECURITY.
Honest limits and failure checklist
Exception-only. RLS is defense-in-depth for dashboard reads. Not a full security whitepaper. Never "100% Sentry compatible."
Limits we accept:
- No APM, session replay, profiling, logs product, or mobile suite.
- Partial Sentry wire for error tracking. Official SDKs via DSN swap (migrate post). Not a clone of every SaaS feature.
- Ingest is not protected by the same org RLS session as the dashboard. DSN leakage is a project-level problem.
- RLS does not replace authn, authz, TLS, or password hygiene.
- Idle stack ~82 MiB was measured 2026-09-12 on our box. Treat that as one data point, not a promise.
Failure checklist:
-
Dashboard connected with
DATABASE_URL/epure. Migrations role can see everything. UseEPURE_APP_DATABASE_URLfor/api/v1. -
Forgot
SET LOCAL app.current_org_id. Queries asepure_appreturn empty or error depending on policy. Fix the session setter, do not weaken the policy. -
Trusting org id from the JSON body. Membership must drive
SET LOCAL. Client UUID is not a credential. -
DSN pasted into a dashboard
/api/v1call. Wrong path. Ingest uses the envelope/store routes and the ingest role. -
Example passwords on a public
EPURE_PUBLIC_URL. Boot refuses documented laptop defaults and leftoverCHANGE_MEoff localhost. Use the production env template. - Expecting RLS to hide events after a leaked project DSN. Rotate the key. RLS does not rewrite ingest auth.
- Assuming event TTL deletes issue rows. Partitions drop; issue aggregates remain. Read retention docs before promising "everything gone."
Configuration map for the three URLs and roles: self-hosting configuration.
Discussion
Where do you put the tenant boundary for multi-tenant reads: app filters only, Postgres RLS, or separate databases per customer?
Top comments (0)