Adapted from the Hebrew guide on AI NEWS IL.
A knowledge base for an AI assistant does not start by uploading every file into one folder. It starts with a source register: an owner, a validity window and an access rule for each source. When retrieval, permissions and the update path work together, you connect it to real users.
The goal is not a model that never errs. The goal is an error you can detect: which source was returned, when it was approved, who may see it, what was missing and who fixes it. A knowledge base that cannot retire an old document or honor a permission change will give fluent, wrong answers even if the model never changed.
1. Build a source register before you index
Create one row per source, not per folder. Each row answers: what is the document, who owns it, who is it for, when is it valid, what replaces it, and how is it removed from the index.
Do not assume the newest file wins. Upload date can be later than content date, and a copy moved between folders can look new while being stale. Define a precedence order per information type. A signed policy may beat a training deck, and a numbered correction notice may beat the policy until a new version ships. If there is no rule, the right state is conflict. Do not ask the model to pick the document that sounds more convincing.
2. Separate active, historical and draft content
A historical document can matter for audit, but it must not silently compete with the active one. Keep it in an archive behind access control and remove it from the normal answer path with a status field and an expiry date. A draft stays out of the active store until the owner approves it. A filename that says "final" is not approval.
Useful metadata at minimum:
source_id: expenses-policy-2026
owner: ops-team
audience: all-employees
status: active # draft | active | superseded | expired
effective_from: 2026-01-01
reviewed_at: 2026-09-01
version: 3
supersedes: expenses-policy-2025
Some search products can filter on file attributes. That capability does not set your policy. Define what each field means and who may change it first, then implement filters in the product you chose.
3. Keep permissions end to end
Access to the knowledge base is not permission to see everything in it. At ingest, store the audience of the source. At query time, filter results by the user's identity. In the answer, never show, quote or link a document the user may not open. A broad service account is not a reason to widen what the user sees.
Check the source permissions too. A product that respects user permissions can still surface content that was shared too widely in the past. Microsoft's documentation, for example, explains that SharePoint and OneDrive controls affect discovery without changing user permissions, and that lifecycle and sharing policies help reduce oversharing. That is a statement about Microsoft 365 Copilot, not a guarantee for every RAG tool.
Start read-only. The knowledge base admin can import, disable and rebuild. A normal user can ask. The model gets no write access to source documents. Deleting, changing an audience or replacing a version are logged admin actions, not the result of a chat.
4. Make ingestion verifiable
A typical pipeline is ingestion, text extraction, chunking, embeddings, indexing and retrieval. Google's RAG Engine documentation shows a similar chain. It describes the product. It does not prove your documents were chunked well or that the output is faithful to the source.
Log every run: run ID, source list, versions, successes and failures. A file is not searchable just because you sent it to an API. OpenAI's File Search documentation, for example, says to check that a file reached the completed state, and it describes file citations and metadata filtering. If ingestion is partial, keep the previous version active or block the topic. Do not mix half an update without a marker.
The tooling can be simple: a sheet or database for the register, an exporter or crawler, a parser, a search engine or vector store, a link checker, a set of reference questions, and a log that ties each question to the retrieved results and the answer.
5. Require an openable, versioned citation
The answer should show a clear source name and a link that opens with the user's permissions. For testing, also keep source_id, version and chunk ID. A filename alone can change or repeat across folders. If the source cannot be opened, the citation does not let the user verify the claim.
A citation shows the system linked a claim to a file. It does not prove the claim follows from it. When testing, compare each material claim to the source sentence, check that the context was not cut off, and check that the source was valid at answer time. Also sample answers where a similar paragraph was found but belongs to another audience, product or period.
6. Build an acceptance test for retrieval
Do not only check whether the wording sounds good. For each question, store the expected result at three levels: which sources should be retrieved, what may be concluded from them, and what the correct decision state is. That lets you tell a retrieval failure from a wording failure. If the right document never reached the context, a better prompt is not the fix.
Include:
- a normal question and a rephrased one
- a technical term in another language
- a superseded document
- two conflicting versions
- a user without permission
- a file that failed ingestion
- a question with no answer
- a deleted link
There is no universal number of questions that guarantees quality. Pick coverage by content variety and the harm a wrong answer can cause. Useful metrics: right-source retrieval rate, share of answers with a valid source, conflicts detected, permission leaks, unanswered questions that correctly stopped, and time from approval to active. An average cannot make up for a permission leak or use of a revoked document.
7. Update and remove without a gray zone
Define a change path: the owner approves a version, the admin ingests it, regression tests run, and only then does an alias switch to the new active index. Keep a way back to the previous index, but never restore a document that was revoked for safety or privacy reasons. A technical rollback is subject to the business status of the source.
When a permission is removed or a document is revoked, set a response time and check every copy: source, cache, index, saved results and the test environment. Do not promise instant deletion if the product describes another retention period. Record what was deleted, what remains under policy, and when you checked.
Example (hypothetical): expense policy
This is a proposed design that was not tested in production. An ops team wants the assistant to answer travel expense questions. There is an active policy, an old training deck and a draft for next year.
-
Input: the active policy and an approved correction notice. The deck is marked
superseded. The draft stays out of the active store. -
Output: a short answer, an amount or rule only if it appears in a valid source, a link to the section, the version, and a state of
answeredorneeds-clarification. - Permissions: all employees can read the policy. Only ops manages versions. The model cannot write or approve expenses.
- Tests: a regular employee, a unit with an exception, Hebrew and English phrasing, a question about next year, a broken link, and a conflict between the notice and the policy.
- Stop conditions: no active source found, two equal-rank sources conflict, the user is not authorized, a file did not finish ingestion, or the answer adds a condition the source does not contain.
"Contact ops" is an acceptable result when there is no valid source. It is expected behavior that keeps the model from inventing policy.
Failure cases
-
Two active versions: stop with
conflict, show both sources to the owner, do not pick automatically. - Partial ingestion: do not turn on the new index. Log the file and the error, retry without lowering the threshold.
- Source too widely shared: fix the sharing or remove the source until a decision. An app-level filter does not replace source access control.
- Valid document not retrieved: mark a retrieval failure. Check extraction, chunking, metadata and the query before you rewrite the answer.
- Citation does not support the claim: block the claim, fix the context check, add the case to regression.
-
No answer in the base: return
no-valid-sourcewith an escalation path. Do not fill in general knowledge as if it were company policy. - Broken link: show that the source cannot be verified and route it to the owner.
- Permission changed after indexing: run removal or rebuild, invalidate caches, and test with the user's identity that access is gone.
Limits
RAG or file search does not make a source correct, current or permitted. Chunking can separate a rule from its exception, metadata can be wrong, retrieval can miss a term, and a model can state a conclusion the source does not support even with the right context. Permissions in one product do not prove permissions in another.
This guide does not replace data classification, a privacy assessment, legal advice, records management or a vendor contract review. Retention, deletion, processing region and sub-processors depend on the product, the plan and the real settings. Before using personal, confidential or regulated material, get approval from the responsible people in your organization.
Sources
- NIST AI 600-1, Generative AI Profile (published July 26, 2024). A voluntary cross-cutting framework for AI risk management. The source register and worksheet here are a proposed implementation, not a NIST requirement.
- OpenAI API, File Search documentation. Describes ingestion status, file citations, search result inclusion and metadata filtering for that product.
- Microsoft Learn, Microsoft 365 Copilot data protection architecture. Explains how permissions, sensitivity labels, sharing and lifecycle policies affect Copilot. Not a deployment guide for every RAG system.
- Google Cloud, Vertex AI RAG Engine overview. Lists the ingestion, transformation, embeddings, indexing, retrieval and generation chain.
Originally published in Hebrew on AI NEWS IL (ainewsil.co.il), with a printable worksheet. This is an English adaptation.
The full Hebrew guide, with the printable worksheet, is on AI NEWS IL: read the original article. The Israeli Institute for AI runs AI training and implementation programs for organizations worldwide.
Top comments (0)