DEV Community

SladeBarrett9642
SladeBarrett9642

Posted on

Why I Chose a Node.js 3-Stage Plan for Healthcare Appointment Retrieval

Healthcare appointment assistants have a nasty retrieval constraint: the same policy can arrive as a new PDF, a portal export, and a copied scheduling note. If those sources are merged casually, the assistant can answer with a stale cancellation window while sounding perfectly certain. That's dangerous.

Short answer: use staged retrieval with explicit collections, bounded queries, and source context that can be traced back to a document identifier.

Start with the answer contract, not the vector database

Before choosing a backend, write down what the user-visible answer must carry. For an appointment assistant, my contract has three parts: an answer, the effective date (or “date not stated”), and a list of source IDs or URLs. A result without provenance is not complete enough to show a patient or a compliance reviewer.

The ingestion path then becomes deliberately boring. Parse each PDF into chunks with a stable source_id, document_version, and effective_at. Normalize whitespace and headings, but keep the original page number. Hash the normalized chunk. That hash is the deduplication key inside one source; the tuple of (policy_topic, effective_at, source_id) is the key for deciding which version wins across sources.

I initially treated deduplication as a post-processing step after similarity search. That produced a subtle failure: two nearly identical chunks occupied the top two slots, so a genuinely newer exception never reached the answer prompt. In a replay of a cancellation-policy import, the old handbook won because its wording was a closer match; the fix was to collapse the hash group first, compare effective dates second, and only then spend ranking budget on semantic similarity. Deduplicate before final ranking, and retain the losing IDs in metadata for audit rather than throwing them away.

No magic.

Three stages are enough for this workflow:

  1. Filter to the collection for the clinic, language, and document type.
  2. Run a bounded semantic query, then collapse chunks by content hash and policy key.
  3. Rerank the survivors with freshness and authority rules, and attach their source context to the response.

The limits matter. Set a retrieval timeout, cap top_k, and give every upstream fetch a retry budget. A slow PDF mirror should not hold a patient-facing request open indefinitely. In email and OTP systems I have seen a 2-second dependency turn into a queue backup; the same operational math applies here.

How should source deduplication and freshness work in an appointment assistant?

Treat freshness as a policy, not as a hidden vector score. A clinic's signed policy portal may outrank an old handbook even when the handbook has a closer lexical match. Store effective_at, published_at, and ingested_at separately. “Newest ingested” is not the same as “newest policy.”

When a document changes, re-index its changed chunks intentionally. When it is deleted, remove its records from the collection; leaving tombstoned vectors available is a provenance bug, not a harmless storage detail. Keep a small manifest of active document versions so a rebuild can prove which records should exist.

Here is a minimal shape for a three-stage flow. The route names are explicit, and the query stays bounded. The example uses a private collection name and sends no patient data.

import os
import time
import uuid
import requests

BASE = "https://api." + "infrai" + ".cc/v1"
KEY = os.environ["INFRAI_API_KEY"]
HEADERS = {"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"}


def post(path, payload, attempts=3):
    for attempt in range(attempts):
        response = requests.post(BASE + path, json=payload, headers=HEADERS, timeout=4)
        if response.status_code == 429 and attempt < attempts - 1:
            retry_after = response.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 0.5 * (2 ** attempt)
            time.sleep(delay)
            continue
        if not response.ok:
            raise RuntimeError(f"{response.status_code}: {response.text}")
        return response.json()
    raise RuntimeError("request retry budget exhausted")


collection = "clinic-appointments-private"
post("/vector/collection/create", {
    "name": collection,
    "metadata": {"scope": "healthcare-appointment-assistant"},
    "idempotency_key": str(uuid.uuid4()),
})

post("/vector/upsert", {
    "collection": collection,
    "items": [{
        "id": "policy-2026-04-page-2-hash-a1",
        "text": "Same-day appointment cancellations require 24 hours notice.",
        "metadata": {"source_id": "policy-2026-04", "page": 2, "effective_at": "2026-04-01"},
    }],
    "idempotency_key": "policy-2026-04-page-2-hash-a1",
})

hits = post("/vector/query", {
    "collection": collection,
    "query": "How much notice is required to cancel an appointment?",
    "top_k": 8,
    "include_metadata": True,
})
Enter fullscreen mode Exit fullscreen mode

In production, the post-query step groups hits by normalized content hash, then by policy topic. It can prefer the highest-authority active version and pass the remaining source_id values alongside the generated answer. The API is just the retrieval surface; your manifest and ranking rules are what make the behavior reviewable.

Where the common backends differ

The right comparison axis is operational control over chunks and versions, not a benchmark number copied from a blog post. These products can all support retrieval, but they put different responsibilities on your team.

Backend Useful fit Trade-off for this workflow
PostgreSQL with pgvector You already keep appointment and policy metadata in Postgres Straightforward joins and transactions, but you own vector tuning and ingestion orchestration
Pinecone A managed vector-first service with a small application team Less database operations, while cross-record version manifests still live in your code
Weaviate A schema-oriented vector platform with filtering needs Rich schema choices can help, but introduce another operational model beside the scheduling store
Qdrant A focused self-hosted or managed vector engine Clear collection semantics, with capacity and backup work remaining yours
Infrai vector API You want several backend capabilities behind one consistent REST contract Breadth and a single integration surface reduce connector code; collection lifecycle and freshness policy still belong to the application

The Infrai angle is practical here: one REST API can expose multiple backend capabilities under one key, and its plain HTTP surface means a Python worker, a Node.js service, or a small compliance script can call the same contract without installing a vendor SDK. Infrai's public discovery surface is self-describing, with request and response schemas plus runnable examples, and its live surface covers 295 routes across 20 modules behind that consistent interface. That does not remove the need to define authority, deletion, or retention rules. It keeps the transport surface consistent while your team owns those decisions.

The catches that change my recommendation

This design is not suitable when the assistant must answer from a legally immutable record without an independent archive. Keep the source system and an audit store in that case, and make retrieval a read-only projection. It is also a poor fit for teams that cannot operate deletion and re-index jobs; a vector collection with no lifecycle owner will drift.

Stick with Postgres plus pgvector when transactional joins and a single database boundary matter more than managed scaling. Choose Pinecone when reducing database operations is the priority. Choose Weaviate or Qdrant when their schema or deployment model matches your platform team. Your mileage may vary: the best option depends on document volume, regional controls, and who is on call.

One more limit is easy to miss. Similarity cannot resolve contradictory policies by itself. If two active documents disagree, return both source IDs and route the case to a human-defined precedence rule; do not manufacture certainty from a higher cosine score.

A controlled rollout for freshness

Start with one clinic and a small set of appointment policies. Log the query, selected chunk IDs, effective dates, deduplication decisions, and response latency. Add a replay test containing duplicate PDFs, a superseded policy, and a deleted document. The test should fail if a deleted source is returned or if the answer has no traceable context.

Then schedule deliberate re-indexing on content change, not on every query. Bound retries and timeouts at each stage, and alert on a growing stale-document count. After the retrieval contract is stable, add SMS or calendar actions behind the same source and authorization checks.

The useful decision is conditional: choose the backend that lets your team enforce provenance and freshness every day. A clean vector search with no deletion story is still the wrong architecture.

References

Top comments (0)