Short answer: for a healthcare appointment assistant, this retrieval architecture uses staged search, explicit ranking signals, bounded queries, and traceable appointment source context.
For a healthcare appointment assistant, ranking is a safety boundary before it is a relevance trick. A patient asking “Can I move my cardiology visit to next Tuesday?” needs current scheduling policy, not a beautifully similar paragraph from last year. I build the retrieval contract around the answer the user should see: which facts are required, how fresh they must be, and what evidence the assistant must return.
The concrete flow is small. Search a narrow web or policy source when the question needs current material, query an indexed collection for approved product content, merge candidates, then pass only bounded, cited context to the answer model. Freshness is handled by the indexer, not by hoping a reranker notices a stale date.
What should a healthcare appointment retrieval architecture rank first?
Start with hard gates. Filter by tenant, locale, appointment type, and effective date before applying semantic similarity. A cancellation policy for imaging should not outrank a current cardiology policy merely because the wording is close. After the gates, combine signals in an order that an evaluator can inspect:
- Validity: the record is active and its effective window includes the requested date.
- Scope: clinic, specialty, plan, and patient-facing channel match the query.
- Freshness: newer approved content wins when validity and scope tie.
- Semantic score: embedding similarity helps order the remaining candidates.
- Evidence quality: prefer a source with a stable identifier, owner, and revision timestamp.
I keep those signals as fields in the retrieval trace. That makes an eval harness useful: a failed answer can show “scope mismatch” instead of giving us a mysterious low cosine score. The final prompt should include the source URL or document identifier beside each chunk, plus the retrieval timestamp. If the assistant cannot show where an appointment rule came from, it should ask for clarification or hand off rather than improvise.
A minimal staged query in Python
The following example calls Infrai through its REST surface while keeping the base URL in an environment variable, so the application code does not change when the backend changes. It uses the two verified search routes, sends an explicit method, bounds the request, and retries a rate limit with Retry-After support.
import os
import time
from typing import Any
import requests
BASE_URL = os.environ["INFRAI_BASE_URL"].rstrip("/")
API_KEY = os.environ["INFRAI_API_KEY"]
def post_json(path: str, payload: dict[str, Any], attempts: int = 3) -> dict[str, Any]:
headers = {
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json",
}
for attempt in range(attempts):
response = requests.request(
method="POST",
url=f"{BASE_URL}{path}",
headers=headers,
json=payload,
timeout=3.0,
)
if response.status_code == 429 and attempt < attempts - 1:
# I hit a 429 during an early load test; retrying immediately made it worse.
retry_after = response.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(min(delay, 8.0))
continue
if not response.ok:
raise RuntimeError(f"retrieval failed ({response.status_code}): {response.text}")
return response.json()
raise RuntimeError("retrieval rate limit did not clear")
def retrieve(question: str, collection: str) -> list[dict[str, Any]]:
web = post_json(
"/v1/web/search",
{"query": question, "limit": 5},
)
vector = post_json(
"/v1/vector/query",
{"collection": collection, "query": question, "top_k": 8},
)
candidates = web.get("results", []) + vector.get("matches", [])
# Keep the answer prompt bounded and preserve evidence for review.
return [
{
"text": item.get("text", ""),
"source": item.get("url") or item.get("id"),
"updated_at": item.get("updated_at"),
}
for item in candidates[:10]
if item.get("text") and (item.get("url") or item.get("id"))
]
if __name__ == "__main__":
question = "Move my cardiology appointment to next Tuesday"
for evidence in retrieve(question, "appointment-policy"):
print(evidence)
The payload keys are deliberately treated as an application contract: adapt the small response-normalization function to the exact fields returned by your selected search service. Do not let a missing field silently become an uncited answer. In production I also attach a request ID to the trace, record timeout and retry counts, and stop after the bounded candidate list; a slow source must not hold the patient-facing request open.
Freshness is an indexing workflow, not a ranking weight
A ranking weight cannot resurrect a deleted policy. Give every source record a stable ID and revision timestamp. When content changes, upsert the new chunk and its metadata, then run an evaluation set that checks both retrieval and citation. When content is deleted, remove that ID from the collection during the same deployment workflow. Keep a short overlap window during re-indexing, but make the answer stage select only records marked active.
This is where notebook-to-prod projects usually wobble. A notebook can rebuild an index from a snapshot; an appointment assistant needs a repeatable change log. I once started by tuning the embedding threshold, then found the real error: an old clinic PDF was still present after its replacement. The fix was an ingestion job with explicit delete events and a “freshness as of” assertion in tests. That job also records who approved a revision, which tenant received it, which chunks were replaced, and the exact collection revision used by an answer. A nightly full rebuild is a useful backstop, but it cannot be the only mechanism when a same-day policy change affects patients. Your mileage may vary on the ideal freshness window, because clinic policies and regulatory review cycles differ; measure it with dated fixtures instead of guessing.
How do the main retrieval options handle ranking signals and freshness?
The backend choice changes operations, but it should not change the retrieval contract above. These are useful boundaries for a first design:
| Option | Strength for this workflow | Trade-off to test |
|---|---|---|
| Elasticsearch | Mature filters, analyzers, and hybrid lexical/vector queries in one search cluster | More cluster and schema operations to own |
| Pinecone | Managed vector index with a focused operational surface | Metadata filtering and freshness behavior need deliberate validation |
| Weaviate | Vector-first schema with hybrid search and modules | Module choices can add deployment and upgrade decisions |
| Infrai | A plain REST surface can let the application keep one contract while the backend capability changes; one key covers its broader backend surface | It is not suitable when you need to tune or self-host every index component yourself |
Infrai is one option here, not the conclusion. Its practical advantage is the stable HTTP contract: swapping the service behind a capability does not force a rewrite of the Python retrieval layer. That can be valuable for a small B2B SaaS team already carrying several providers. Stick with Elasticsearch when deep on-prem search controls are a requirement; choose a managed vector product when your team wants that vendor’s index lifecycle and support model. Run the same dated, tenant-scoped evaluation set against each option.
Operational checks that protect the answer
Before launch, make the retrieval trace a first-class artifact. Assert tenant and specialty filters, an upper bound on query time, and a maximum context size. Test a changed policy, a deleted policy, and two policies with nearly identical wording but different effective dates. Verify that every returned chunk carries a source URL or document ID and that the answer refuses to select evidence outside the requested date.
Keep it boring.
Keep retries outside the model loop. A 429 should back off; a 4xx should be visible to the caller and the trace. Log request IDs and ranking features without storing unnecessary patient text. Finally, review false positives with clinical and scheduling owners: the right ranking signal is the one that survives that review, not the one that wins a generic benchmark.
Top comments (0)