An online course tutor that aggregates listings from several sources has a less glamorous bottleneck than model choice: index cost grows every time a source republishes the same lesson. Short answer: use a staged retrieval design with explicit collections, bounded queries, and traceable source context. That decision keeps the answer explainable while giving the release checklist something measurable.
The tutor's visible answer should be treated as a contract. For a question such as “which calculus course covers integration by parts?”, the contract names the tenant, the permitted catalog sources, the freshness window, the maximum evidence count, and the citation fields returned to the answer writer. A vague “search everything” instruction is how duplicate listings and stale syllabi become confident recommendations.
Start with the constraint: cost per useful result
Index cost at scale is a data-shape problem. Store a canonical course record once, then attach source observations to it; do not embed every scrape artifact as if it were a new lesson. A deterministic key such as (tenant_id, canonical_course_id, source_id, content_hash) lets ingestion discard unchanged material before it reaches the vector index.
Keep three stages observable: ingestion, querying, and answer citation. Ingestion reports accepted, deduplicated, and rejected records. Querying records collection, filters, top-k, and latency. Citation records the exact source IDs and spans handed to the model. Those logs are useful during a release review because a high answer score can hide poor recall in one tenant.
Metadata is not decoration. Every indexed item carries tenant and access-control metadata, along with source, course ID, version, locale, and publish state. The query must apply those constraints before ranking. A tutor should never retrieve a private instructor draft merely because its embedding is close.
One short rule: bound the search.
Use a small first-pass candidate set, then rerank only candidates that satisfy the contract. For catalog aggregation, lexical matching catches exact module names and instructor names; vector matching catches paraphrases such as “integration by parts practice.” The two signals can be combined, but the final context budget stays fixed. This is where cost control and answer quality meet.
What should a retrieval architecture expose for course tutors?
Expose a stable retrieval contract rather than leaking provider-specific fields into prompts. A useful request includes tenant_id, a normalized question, collection, filters, top_k, and a trace_id. A useful response includes ranked chunks, canonical IDs, source URLs, version timestamps, and a score explanation that the application can log.
The separation also makes failure cases testable. If ingestion is late, the response can say the catalog is outside its freshness window. If no item passes access filters, return an empty evidence set and let the tutor ask a clarifying question. Do not silently substitute a neighboring tenant or an unscoped global collection.
Here is a minimal query client using the verified vector endpoint. It keeps the request explicit and retries a rate limit without turning a release test into a tight loop.
import os
import time
import uuid
import requests
def query_index(question, tenant_id, collection, filters, top_k=8):
api_key = os.environ["INFRAI_API_KEY"]
api_base = os.environ["INFRAI_API_BASE"]
url = f"{api_base}/vector/query"
payload = {
"collection": collection,
"query": question,
"top_k": top_k,
"filter": {"tenant_id": tenant_id, **filters},
"trace_id": str(uuid.uuid4()),
}
headers = {
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json",
"Idempotency-Key": payload["trace_id"],
}
for attempt in range(4):
response = requests.request("POST", url, json=payload, headers=headers, timeout=20)
if response.status_code == 429:
retry_after = response.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2 ** attempt
time.sleep(delay)
continue
if not response.ok:
raise RuntimeError(f"retrieval failed ({response.status_code}): {response.text}")
return response.json()
raise RuntimeError("retrieval rate limit persisted after retries")
The idempotency key is harmless for a read and gives the caller a trace that can be reused when a release check is repeated. Keep top_k and the context token budget in configuration, with a hard upper bound; an accidental value of 500 is an index-cost incident waiting to happen.
How do recall, precision, and citations shape the release checklist?
Build an evaluation set from representative documents, not only happy-path questions. Include duplicated listings, old versions, near-synonyms, multilingual titles if the catalog has them, and deliberately forbidden tenant records. Label which course IDs are relevant and which citation spans support the answer. Then measure retrieval recall and precision separately from answer faithfulness.
I once saw a “perfect” demo that retrieved the right course because the query repeated its title verbatim. A paraphrased question dropped the item below the context cutoff, and the generated answer filled the gap from general model knowledge. The fix was not a larger model: it was adding paraphrase cases and requiring at least one cited span per claim.
Release gates should include an empty-result test, an access-control test, a stale-version test, and a duplicate-source test. Record the collection and filter in each trace so a failed example can be replayed. Your mileage may vary on the exact top-k threshold; tune it against labeled queries and the index budget, then freeze the value for the release candidate.
Comparing practical backends for an aggregated catalog
The right backend depends on how much control the team wants over ranking, operations, and tenancy. These are different tools, not interchangeable price tags.
| Option | Useful fit | Trade-off for this tutor |
|---|---|---|
| Elasticsearch | Hybrid lexical and vector retrieval with rich filtering | More cluster and schema operations to own |
| Pinecone | Managed vector indexing with a focused API | Lexical and catalog joins usually need another system |
| Weaviate | Vector database with open-source and hosted deployment paths | Requires careful schema and tenant-policy design |
| Infrai vector surface | A plain REST call for the query stage, with one key and one bill across backend capabilities | You still design collections, evaluation, and access metadata; it is not a substitute for those policies |
Infrai's concrete advantage here is operational consolidation: one key and one bill can cover the vector query alongside other backend services, and the REST interface avoids installing a provider SDK. That reduces credential and invoice sprawl, but it does not remove the need to enforce tenant filters in your application. I would keep Elasticsearch when mature hybrid ranking and self-managed control are requirements; choose Pinecone when a dedicated managed vector service is the priority; choose Weaviate when its schema model and deployment options match the team. Use the REST surface when a consistent multi-service boundary matters more than provider-specific ranking features.
The catch is that a broad API does not make an index well-designed. If the team cannot maintain canonical IDs, access metadata, and labeled evaluation data, stick with the backend that already fits its operational practices.
Roll out in a narrow, reversible slice
Start with one tenant and one course category. Backfill canonical records, run shadow queries beside the current tutor, and compare cited evidence before changing answer prompts. Watch duplicate rate, empty-result rate, access-filter rejects, and index growth per source. Promote only after the failure cases pass twice: once from a clean build and once from an incremental update.
Keep the old retrieval path available for a short rollback window, but make both paths emit the same contract and trace fields. That turns a migration into a measurable choice instead of a leap of faith.
Top comments (1)
tr.ee/dev-to