Two users can ask the same question and be entitled to completely different answers.
“Summarize the renewal terms” might refer to a public template, a private contract, or a document shared with only one team. The words alone do not tell you which answer is safe to return.
In an LLM application, isolation must follow the information through conversation history, retrieval, agent memory, tool calls, and cached responses. Authentication establishes the caller. Every component that reads or returns context must enforce what that caller is allowed to access.
Why context isolation matters
Context changes what a model says and what an agent may decide to do. Mixing it between users creates several kinds of failure:
- Private information reaches the wrong person. A summary can disclose contract terms, customer records, or internal decisions without quoting the original document. Generated answers need protection just as their source documents do.
- An answer is plausible but belongs to someone else. Another team's pricing, constraints, or preferences can quietly influence a recommendation. The user may see a confident answer with no obvious sign of contamination.
- An agent acts on the wrong premise. Contaminated context can lead it to propose or attempt an inappropriate message, update, or workflow. Tool-level authorization remains essential, but it cannot guarantee that an otherwise permitted action makes sense for the user's actual request.
- Revoked access survives in derived data. Removing access to a document is incomplete if summaries, memories, or cached answers still expose it. Permissions must govern derived content as well as the original record.
These failures do not require a sophisticated prompt injection. Two ordinary users asking the same question can be enough to reveal a missing boundary. Asking the model to “never reveal another user's data” is not an access-control mechanism: unauthorized context should never enter its input in the first place.
Isolation also makes the system easier to debug. When every response has an identifiable authorized scope and context revision, an unexpected answer can be traced to the information it was allowed to use. Without that provenance, a context-selection bug can look like a model hallucination.
Build a trusted request context
Resolve identity from verified credentials, then construct an immutable request context on the server:
# Illustrative structure, not a complete authentication implementation.
@dataclass(frozen=True)
class RequestContext:
tenant_id: str
principal_id: str
permission_revision: str
Validate the credential's signature, issuer, audience, and expiration as applicable. Resolve current tenant membership and permissions. A tenant selected in a URL or header is a requested scope to authorize, not proof of membership.
Client-supplied fields such as user, owner_id, or tenant_id must not override that context. Pass it explicitly to storage, retrieval, tools, and cache operations. Avoid mutable process-wide variables for the current user: concurrent requests must never share that state.
Delegated agent credentials should bind the principal, tenant, audience, permitted operations, and expiration. Downstream services must verify them rather than trust identity fields supplied by the agent.
Authorize resources before building the prompt
A conversation ID identifies an object; it does not authorize access to it.
Read a conversation through its tenant and access policy. Enforce the same policy when listing messages, retrieving attachments, exporting history, or opening a streaming endpoint. Returning 404 for inaccessible objects can limit information about their existence, but the status code is not the authorization check.
For RAG, apply tenant and document permissions before retrieved content reaches a reranker, model, or client. Include the filters in the search where supported, and verify the returned resources against current authorization. Persisting an owner_id field without enforcing it on reads provides no isolation.
Define memory scopes deliberately:
- Conversation memory belongs to an authorized conversation.
- Personal memory may follow a user across conversations if that is an explicit product behavior.
- Team memory follows membership and document permissions.
- Public knowledge is explicitly approved for shared access.
Do not collapse these into one global session. Server-controlled session identifiers should include the relevant tenant, principal, and conversation scope. Background jobs and tool callbacks need the same context as the request that created them. OWASP covers these boundaries across data, sessions, caches, and background processing in its multi-tenant security guidance.
Treat a cached answer as protected data
A response cache can bypass the checks in a retrieval pipeline. That makes the cache an authorization boundary of its own.
Imagine Alice asks a question about her private contract. The answer is cached. Bob asks an identical question about a different contract. If the cache considers only prompt similarity, Alice's answer can become eligible for Bob.
A stricter similarity threshold does not solve this. Identical prompts can depend on different private context.
Use this rule:
Establish which entries the caller may reuse before comparing semantic similarity.
A conservative default is to scope private answers to both the tenant and principal. Broader sharing requires an explicit policy and equivalent access to the information behind the answer.
For context-dependent responses, reuse also needs to account for:
| Dimension | Why it matters |
|---|---|
| Tenant and principal or authorized sharing scope | Defines who may reuse the response |
| Conversation and context revision | Distinguishes histories and subsequent turns |
| Permission revision | Prevents reuse after access changes |
| Source revisions | Distinguishes updated or removed documents |
| Model, instructions, tools, and relevant generation settings | Distinguishes how the answer was produced |
A conversation ID alone is insufficient: its history changes. A source ID alone is insufficient: its contents and permissions change. Use canonical, unambiguous serialization when computing fingerprints, and never use a fingerprint as a substitute for access control.
If the cache lookup happens before retrieval and cannot establish the relevant source context, bypass it for that request or move it to a point where those dependencies are known. Do not label unknown context as empty context.
Make unsafe cache calls difficult to express
Prefer a cache API that requires an authenticated scope over one where omitting an owner silently selects a shared cache.
# Pseudocode: scope is derived by trusted server-side code.
scope = build_private_cache_scope(request_context, authorized_context)
if scope is None:
return await generate_with_authorized_context()
candidate = await cache.find(query, required_scope=scope)
if candidate and await may_reuse(candidate, request_context, scope):
return candidate.answer
answer = await generate_with_authorized_context()
await cache.store(query, answer, required_scope=scope)
return answer
Use the same scope construction for reads and writes. Enforce mandatory filters in the storage adapter, not just at individual call sites. Redis's semantic-cache documentation describes metadata filters for user and conversation scoping; the application must supply the correct values on every path.
Missing or invalid authorization should reject the request. An unavailable cache can fall back to normal authorized generation. Those are different failure modes: neither should trigger a search across all users.
Keep deliberately public caching separate. “Plain text,” “no tools,” and “no retrieved documents” are not evidence that an answer is public. System instructions, prior messages, and user preferences can all affect it.
When introducing a stricter scope, make older entries ineligible unless their provenance can be established. A cache schema version or fresh namespace can separate them. TTL limits age; it does not prove permission.
Test through the real request boundary
A unit test that passes two different owner IDs to a cache verifies that the cache can isolate them. It does not verify that the HTTP handler supplies those IDs.
Use synthetic identities and private markers in an isolated integration environment:
- Authenticate as Alice and generate an answer containing a marker from her private context.
- Confirm that the answer entered the intended cache partition.
- Authenticate as Bob and send the identical final prompt, then a paraphrase.
- Assert that Alice's cache entry is ineligible and that Bob receives no Alice marker.
- Repeat as Alice with unchanged authorized context and confirm that legitimate reuse still works.
Inspect cache provenance as well as response text. An unrelated output does not prove that an unauthorized entry was never selected.
Extend the matrix to different conversations, different tenants, revoked permissions, updated documents, missing identity, forged client identity fields, and concurrent requests. Test alternate routes: streaming, non-streaming, tools, and background jobs can have different implementations.
Keep these checks deterministic where possible. A stub generator can return the synthetic marker from its authorized input, making an isolation failure observable without depending on model behavior.
Audit the remaining copies of context
Application response caches, agent memory, and inference KV caches have different lifecycles. Review each separately. A shared model server should keep request sequences isolated; any cross-request prefix reuse needs its own assessment of the serving implementation and isolation controls.
Also review logs, traces, exports, and debugging tools. Avoid recording raw private prompts merely to diagnose cache hits. Prefer access-controlled metadata such as request IDs, scope identifiers, cache decisions, and policy versions, with appropriate retention.
The practical standard is consistent: a fast path must enforce the same access policy as the full path. Optimize only after the system can explain why this caller is allowed to receive this context.
More engineering notes: Hashnode · DEV · GitHub · Hugging Face · Telegram · Instagram
Originally published on Hashnode.
Top comments (0)