DEV Community

Muhammad Ikram UL Mustafa
Muhammad Ikram UL Mustafa

Posted on

Your RAG Pipeline Is a Data Leak Waiting to Happen

Your multi-tenant app is probably well isolated. Every query is scoped to a tenant, enforced centrally, and no individual developer has to remember it.

Then you add retrieval, and you quietly create a second data path that never gets the same discipline.

The failure mode

Tenant A asks a question. Semantic search finds the most relevant chunks. One of them came from Tenant B's document. The model reads it and answers fluently, using data the user was never allowed to see.

No error. No exception. No alert. Just a confident, correct-sounding answer built on someone else's data.

This is worse than an ordinary data leak, because there is no artifact. The user never saw Tenant B's document. They saw a paragraph of generated text that happened to contain its contents. Your logs show a successful request.

Why it happens

Most teams build the application layer first and get isolation right there. In Laravel it is a global scope. In Django it is a manager. Somewhere central, every query gets WHERE tenant_id = ? attached, and developers stop thinking about it because they no longer have to.

Then retrieval arrives as a separate feature. Documents get chunked, embedded, and written to a vector store. The code that does this often lives in a background job or an ingestion script, written by whoever was working on the AI feature that sprint.

Now you have two paths to the same data, and only one of them inherited the rules.

A concrete example

Here is ingestion that looks fine:

async def ingest_document(doc: Document):
    chunks = chunk_text(doc.content)
    embeddings = await embed_batch(chunks)

    for chunk, embedding in zip(chunks, embeddings):
        await db.execute(
            "INSERT INTO doc_chunks (document_id, content, embedding) "
            "VALUES ($1, $2, $3)",
            doc.id, chunk, embedding
        )
Enter fullscreen mode Exit fullscreen mode

Spot the problem. There is no tenant_id on the row. The document has one, but it did not make it into the chunk table. You can still get back to it by joining through documents, which means somebody has to remember to join, which means eventually somebody does not.

And here is the retrieval that goes with it:

async def search(query: str, limit: int = 6):
    embedding = await embed(query)
    return await db.fetch(
        "SELECT content FROM doc_chunks "
        "ORDER BY embedding <=> $1 LIMIT $2",
        embedding, limit
    )
Enter fullscreen mode Exit fullscreen mode

Clean, fast, and it will happily return any tenant's content.

What actually fixes it

1. Denormalise the tenant onto the chunk

Put tenant_id directly on the vector rows, even though it is already on the parent document. This is deliberate duplication. It removes the join, and it means the filter cannot be forgotten by accident.

ALTER TABLE doc_chunks ADD COLUMN tenant_id uuid NOT NULL;
CREATE INDEX ON doc_chunks (tenant_id);
Enter fullscreen mode Exit fullscreen mode

2. Make the tenant filter impossible to omit

Do not leave it to the caller. Scope it centrally, the same way your ORM already does for SQL.

async def search(query: str, tenant_id: UUID, limit: int = 6):
    if tenant_id is None:
        raise ValueError("tenant_id is required for retrieval")

    embedding = await embed(query)
    return await db.fetch(
        "SELECT content FROM doc_chunks "
        "WHERE tenant_id = $1 "
        "ORDER BY embedding <=> $2 LIMIT $3",
        tenant_id, embedding, limit
    )
Enter fullscreen mode Exit fullscreen mode

A required argument with no default is cheap and effective. A developer who forgets gets an error, not a leak.

If you are on Postgres, row level security is stronger, because it holds even if someone writes a raw query:

ALTER TABLE doc_chunks ENABLE ROW LEVEL SECURITY;

CREATE POLICY tenant_isolation ON doc_chunks
  USING (tenant_id = current_setting('app.tenant_id')::uuid);
Enter fullscreen mode Exit fullscreen mode

3. Filter before the similarity search, not after

Retrieving the top 20 chunks globally and then discarding the ones belonging to other tenants is a correctness bug waiting to happen, and it also degrades quality. You are spending your result budget on rows you are going to throw away, so a tenant with little data gets worse answers.

Filter first. Let the index do the work.

4. Test isolation as a real case

Most teams test that retrieval returns relevant results. Very few test that it refuses to return someone else's.

async def test_tenant_cannot_retrieve_other_tenant_data():
    await ingest_for_tenant(TENANT_B, "The launch date is March 14th.")

    results = await search(
        "when is the launch",
        tenant_id=TENANT_A
    )

    assert all(r["tenant_id"] == TENANT_A for r in results)
    assert not any("March 14th" in r["content"] for r in results)
Enter fullscreen mode Exit fullscreen mode

Write the test where Tenant A asks for something only Tenant B has, and assert the answer is empty. If that test does not exist, you do not know your isolation works. You are assuming it.

The harder version

Everything above is filtering, which is an application-level control. It runs in the same process as the bug it is guarding against.

If your tenant count is low and the data is sensitive enough, physical separation is stronger. A separate index, collection or namespace per tenant means a dropped filter cannot leak anything, because the other tenant's vectors are not reachable from that connection at all.

The tradeoff is operational. Per tenant indexes get expensive at thousands of small tenants, migrations have to run per tenant, and cross tenant analytics become their own problem. For enterprise B2B with tens or hundreds of tenants, especially under compliance pressure, it is usually the right call. For a product with a large free tier, it usually is not.

Filters are the floor. Physical isolation is the ceiling. Most teams should at least be standing on the floor.

The thing to take away

You can have perfect isolation at the SQL layer and still leak through retrieval, because they are two data paths and only one of them inherited your rules.

Treat the embedding pipeline as a data path with the same requirements as your database, because that is exactly what it is.

I write about production AI systems and what breaks in them. More at ikramulmustafa.com, and there's a RAG agent with the failure handling documented at github.com/ikramulmustafa/n8n-rag-lead-agent.

Top comments (0)