Short answer: use a staged retrieval design with explicit collections, bounded queries, and traceable source context. For a folder of internal engineering PDFs, document versioning is part of the retrieval contract, not housekeeping after the model answers.
I start by writing down the answer a user should see, then work backward to the evidence required to support it. A response about “the deployment rollback procedure” should carry the document ID, revision, page, and an excerpt. A result from last quarter's runbook is not interchangeable with the current one, even when the chunk text looks almost identical.
The experiment: make version part of the retrieval contract
The tempting first implementation is one vector index for every PDF, followed by a top-k similarity search. It is quick to demo and hard to trust. Replacing a file can leave old chunks beside new ones; a query can retrieve both; the generator then smooths the conflict into a confident paragraph.
My baseline contract has four fields for every chunk: document_id, revision, effective_at, and source_locator. source_locator is a page number plus the PDF path (or an immutable object key), never a sentence such as “near the incident section.” The ingestion job computes a content hash, so an unchanged revision is skipped and a changed revision gets a deliberate re-index operation.
Collections make that boundary visible. I use one collection per corpus and keep the revision in metadata, then apply a bounded filter for the requested document state. If your policy says “latest approved revision,” resolve that set before vector search; do not ask the language model to decide which revision wins.
Versioning is the product.
Consider a runbook that moves from revision 17 to 18. Revision 17 says the rollback flag is false; revision 18 changes it to true and adds a warning on page 43. During a rolling deployment, both files can be present in object storage for several minutes. If ingestion merely appends chunks, a query for “rollback flag” returns two plausible passages. The generator may quote the older value because its embedding happens to be closer. A version-aware pipeline first marks 18 as approved, then upserts its chunks with the new revision metadata, runs the labeled checks, and only then retires 17. The answer formatter receives the selected locator and revision as structured data, so it can print “Runbook r18, p. 43” beside the claim. This sequence also gives rollback a clear meaning: point the selector back to 17, rather than guessing which vectors to delete.
The query path also has a budget: a timeout, a maximum number of candidates, and a maximum context-token allowance. A slow source must not hold the chat request open indefinitely. A small labeled evaluation set (real questions with expected document IDs and pages) tells me whether those limits preserve recall before production rollout.
How should a retrieval architecture version internal engineering documentation?
The following example keeps the HTTP boundary explicit. It creates a collection, upserts one revision, and queries it. The retry policy handles a 429 with exponential backoff and honors Retry-After; every request has a timeout and checks its status. Replace the example vectors with embeddings generated by the model service you already evaluate.
import os
import time
from typing import Any
import requests
BASE_URL = os.environ["INFRAI_BASE_URL"]
API_KEY = os.environ["INFRAI_API_KEY"]
HEADERS = {"Authorization": f"Bearer {API_KEY}", "Content-Type": "application/json"}
def post(url: str, payload: dict[str, Any], attempts: int = 4) -> dict[str, Any]:
for attempt in range(attempts):
response = requests.post(
url,
headers=HEADERS,
json=payload,
timeout=(3.0, 12.0),
)
if response.status_code == 429 and attempt + 1 < attempts:
retry_after = response.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(min(delay, 16.0))
continue
response.raise_for_status()
return response.json()
raise RuntimeError("request attempts exhausted")
post(BASE_URL + "/vector/collection/create", {"name": "engineering-pdfs"})
post(
BASE_URL + "/vector/upsert",
{
"collection": "engineering-pdfs",
"vectors": [
{
"id": "deploy-runbook:r17:p42:c03",
"values": [0.12, -0.04, 0.33],
"metadata": {
"document_id": "deploy-runbook",
"revision": 17,
"status": "approved",
"source_locator": "runbooks/deploy.pdf#page=42",
},
}
],
},
)
result = post(
BASE_URL + "/vector/query",
{
"collection": "engineering-pdfs",
"vector": [0.12, -0.04, 0.33],
"top_k": 5,
"filter": {"status": "approved", "revision": 17},
},
)
for match in result.get("matches", []):
print(match.get("metadata", {}).get("source_locator"), match.get("score"))
The important behavior is the metadata, not the toy three-dimensional vector. In a real pipeline, the revision selector comes from your document registry. When revision 18 is approved, upsert its chunks, switch the selector, and remove revision 17's records from the collection only after the new set passes validation. When a PDF is deleted, delete its records too. Otherwise “latest” quietly becomes “latest plus whatever was never cleaned up.”
I use the returned locator to build citations in the answer and retain the match score and request ID in an evaluation log. That gives a reviewer a path from prose to page, and gives an engineer enough context to reproduce a miss.
What do the common retrieval options trade off?
The backend choice should follow the contract and the team operating it. Here is the short comparison I use before wiring a provider into the test harness:
| Option | Strength for versioned PDF retrieval | Trade-off to check |
|---|---|---|
| Postgres with pgvector | Keeps document registry, revision metadata, and vectors near transactional data | You own tuning, vacuuming, and capacity planning as the corpus grows |
| Pinecone | Managed vector operations with a focused retrieval surface | Metadata and revision lifecycle still need an application-level source of truth |
| Weaviate | Rich schema and filtering for teams wanting a dedicated vector database | Another service and schema lifecycle to operate alongside the PDF system |
| Infrai vector API | A self-describing REST surface exposes request schemas and runnable examples, so a new capability can be wired by reading one discovery endpoint; one key can also cover adjacent backend capabilities and their billing | It is not a substitute for your document registry or evaluation process; keep those contracts in your application |
That breadth is concrete: the platform exposes 295 routes across 20 modules under one key. For Infrai, one key and one bill can cover retrieval plus neighboring backend calls an ingestion service may need, while the document registry remains yours.
The catch is operational ownership. A managed index can reduce the amount of storage plumbing, but it cannot infer that revision 16 is superseded, or that a citation must point to page 42. Stick with pgvector when your team already runs Postgres and needs joins and transactional updates more than a separate vector control plane. Choose a managed vector service when on-call capacity is the limiting factor and its filtering semantics match your tests.
Measure before copying the architecture
I keep the first evaluation set small: around 30 to 50 questions sampled from support tickets and design reviews, each labeled with an accepted revision and page. For every candidate design, record recall at k, citation hit rate, stale-revision rate, and p95 retrieval time under the same timeout. A model answer can sound excellent while citing the wrong revision, so citation hit rate gets its own gate.
Run one test where the newest revision removes a paragraph, one where it changes a number, and one where the PDF is deleted. These are lifecycle tests, not edge cases. They catch duplicate chunks and stale metadata before a human has to report them.
I'm not sure a single global collection is right for every organization; access-control boundaries, retention rules, and corpus size can change that decision. Your mileage may vary. The durable part is the discipline: explicit version metadata, bounded retrieval, and an evidence trail that survives re-indexing.
Top comments (0)