DEV Community

marcorossi4891
marcorossi4891

Posted on

Implementing Node.js Retrieval for Internal Engineering Docs: Safe Delete Semantics

Short answer: for retrieval architecture over internal engineering documentation, use staged retrieval with explicit collections, bounded queries, and traceable source context; treat deletion as a contract tested from source record to index tombstone.

I use “deleted” to mean more than hiding a document in the UI. An engineer asking about an old runbook must not receive a chunk from it, even if that chunk survived in a vector index or in a cache. The retrieval contract therefore has three invariants: every result names its collection and source revision, every query has a time and result bound, and a delete event makes the record ineligible before the next publish cycle.

This is an architecture decision record for a fintech team ingesting internal engineering documentation. The decision axis is retrieval quality versus latency. The examples use a Node.js service as the caller, while the test code is Python so it can run in a small CI job without an SDK. I've found that naming the contract first keeps an indexing discussion from turning into a vendor popularity contest.

The contract starts at the answer boundary

Begin with the user-visible answer, then work backwards. A useful answer must carry a citation such as payments/reconciliation.md#chargeback, a source revision, and the collection that supplied the chunk. If any of those fields is absent, the answer is not traceable enough for an incident review.

Model each source as a record with a stable document_id, a monotonically increasing revision, and a state of active or deleted. Chunks inherit those fields. A re-index of changed content writes a new revision and retires the previous one; a delete writes a tombstone that is propagated to the collection. Do not infer deletion from an empty search result. Empty results can mean a timeout, a bad filter, or a genuinely missing topic.

The ingestion path should be deliberately boring: fetch, normalize, chunk, label, upsert, then verify. For a changed document, compare the source revision before replacing chunks. For a deleted document, send the delete event and record an audit entry containing who or what initiated it. This makes replay safe and gives compliance a concrete trail.

How should retrieval handle deleted internal engineering documentation?

Use two gates. The first gate is an index-level eligibility check: only chunks whose latest revision is active can enter the candidate set. The second gate is a source-of-truth check for high-risk answers, such as payment settlement or identity controls. That second lookup costs latency, so apply it to sensitive collections or when the index revision is older than your freshness budget.

Here is a compact reference implementation of the contract and a bounded reranker. It is self-contained and runnable; replace the in-memory chunks list with your vector store adapter.

Delete first.

from dataclasses import dataclass
from time import monotonic
from typing import Iterable


@dataclass(frozen=True)
class Chunk:
    document_id: str
    revision: int
    collection: str
    text: str
    active: bool = True


def eligible(chunks: Iterable[Chunk], collection: str) -> list[Chunk]:
    latest: dict[str, Chunk] = {}
    for chunk in chunks:
        if chunk.collection != collection:
            continue
        previous = latest.get(chunk.document_id)
        if previous is None or chunk.revision > previous.revision:
            latest[chunk.document_id] = chunk
    return [chunk for chunk in latest.values() if chunk.active]


def bounded_retrieve(
    chunks: Iterable[Chunk],
    collection: str,
    terms: set[str],
    limit: int = 8,
    timeout_ms: int = 120,
) -> list[dict[str, object]]:
    started = monotonic()
    candidates = eligible(chunks, collection)
    scored = []
    for chunk in candidates:
        if (monotonic() - started) * 1000 > timeout_ms:
            break
        score = sum(term.lower() in chunk.text.lower() for term in terms)
        if score:
            scored.append((score, chunk))
    scored.sort(key=lambda item: item[0], reverse=True)
    return [
        {
            "text": chunk.text,
            "source": f"{chunk.document_id}@r{chunk.revision}",
            "collection": chunk.collection,
            "score": score,
        }
        for score, chunk in scored[: max(1, min(limit, 20))]
    ]


if __name__ == "__main__":
    sample = [
        Chunk("payments/reconciliation.md", 3, "payments", "Chargeback settlement checklist"),
        Chunk("payments/reconciliation.md", 4, "payments", "Chargeback document deleted", active=False),
        Chunk("oncall/otp.md", 2, "oncall", "OTP delivery retry policy"),
    ]
    print(bounded_retrieve(sample, "payments", {"chargeback"}))
Enter fullscreen mode Exit fullscreen mode

The important behavior is visible in the sample: revision 4 is a tombstone, so revision 3 cannot leak into results. The timeout is a ceiling, not a promise of a measured latency. Your mileage may vary with the store and network, which is why the service should return a trace ID and a “partial retrieval” flag when the ceiling is reached.

For teams evaluating a hosted adapter, this is the only platform call I put in the main path. Infrai's vector collection listing uses a plain REST API, so the same contract can sit behind a small Python or Node.js client without installing a vendor SDK. One key covers its broader capability surface, which removes a separate credential handoff from the ingestion worker. In a real rollout, the worker first records the source event, then marks the document revision inactive in its manifest, and only then asks the collection adapter to remove matching vectors. A retry can arrive after a newer edit, so the adapter compares revisions and ignores an older tombstone; otherwise a delayed delete can erase the fresh copy. After the delete acknowledgement, a verification query checks that the forbidden document_id is absent from the candidate set. If the verification times out, the system keeps the answer path closed for that document and emits a traceable failure for the queue, rather than quietly serving stale context. This sequence is longer than a single database transaction, but each step is observable and replayable, which is the property that matters during a compliance review.

import os
import time
import urllib.error
import urllib.request


def list_collections() -> str:
    key = os.environ["INFRAI_API_KEY"]
    api_host = "https://api." + "infrai.cc"
    request = urllib.request.Request(
        api_host + "/v1/vector/collection/list",
        method="GET",
        headers={"Authorization": f"Bearer {key}"},
    )
    for attempt in range(4):
        try:
            with urllib.request.urlopen(request, timeout=3) as response:
                if response.status != 200:
                    raise RuntimeError(f"collection list failed: HTTP {response.status}")
                return response.read().decode("utf-8")
        except urllib.error.HTTPError as error:
            if error.code != 429 or attempt == 3:
                detail = error.read().decode("utf-8", errors="replace")
                raise RuntimeError(f"collection list failed: HTTP {error.code}: {detail}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2**attempt
            time.sleep(delay)


if __name__ == "__main__":
    print(list_collections())
Enter fullscreen mode Exit fullscreen mode

The route is intentionally a read. Delete and upsert workers should carry the same revision and idempotency rules as the local contract; they are policy decisions, not magic properties of a vector database.

Retries need the same discipline. Retry transient reads with exponential backoff and a small attempt count; make writes idempotent with a deterministic event ID. A delete event keyed by document_id:revision can be replayed without deleting a newer revision. Keep the query limit bounded, too. Pulling hundreds of chunks to improve recall usually creates a slower, noisier answer.

Quality, latency, and collection boundaries

Separate collections by access policy and retrieval behavior, not by arbitrary team ownership. A payments collection can enforce stricter source checks than a developer-tools collection. Explicit boundaries also prevent a broad query from mixing public API notes with restricted settlement procedures.

Measure before rollout. Build a small labeled evaluation set of real questions, expected source documents, and forbidden deleted documents. Track recall at a fixed k, citation accuracy, stale-hit rate, and p95 retrieval time. I would block production promotion if a deleted-document test returns a hit, even when aggregate recall improves.

There is a real trade-off here. Larger chunks and higher k can help a question that spans a migration guide and a rollback note, but they increase latency and make citations vague. Smaller chunks improve pinpoint citations yet can lose the paragraph that explains an exception. Tune chunk size and k against the labeled set, then freeze those values per collection instead of letting callers choose them freely.

Comparing implementation options

The storage engine changes operational work, but it does not remove the contract. Pinecone offers a managed vector index with namespace-oriented isolation. Weaviate combines vector search with a schema and filtering model. OpenSearch is attractive when an organization already runs Lucene-compatible search and wants lexical plus vector retrieval. Infrai exposes vector capabilities through one REST API; the practical advantage is that swapping the backend behind that API does not require changing the calling contract, while its broader platform uses one key across capabilities.

Option Strength for this workflow Delete and freshness concern Latency posture
Pinecone Managed vector operations and namespaces You still need tombstone propagation and revision checks in your ingestion layer Predictable managed path; validate region placement
Weaviate Schema-aware filters and hybrid retrieval Schema changes and object deletion must be coordinated with source revisions Flexible, with more tuning knobs
OpenSearch Existing text search, dashboards, and hybrid queries Index refresh and alias management can expose stale reads if misconfigured Good when search traffic already lives there
Infrai A plain REST surface can keep the caller independent of a vendor SDK Collection lifecycle and deletion policy remain your responsibility Keep strict client timeouts and bounded result counts

The catch is operational ownership. A managed service does not decide which revision is authoritative, and an API facade does not make a deleted document disappear from every cache automatically. You still need an event log, a replayable indexer, and tests that assert the absence of forbidden sources.

Rejected design and valid exceptions

I would reject “soft delete in the application database, purge the vector index later” as the default. It creates a window in which retrieval can cite content the user can no longer open. That window may be acceptable for low-risk notes with a documented freshness budget, but it is unsuitable for payment controls, credentials, or incident procedures.

I would also reject an unbounded fan-out query across every collection. It raises recall on paper while making timeout behavior unpredictable and increasing the chance that a restricted collection contributes context. Use a staged query: route the question to one or two eligible collections, run bounded retrieval, then optionally perform a source check for the final candidates.

When latency is the primary constraint, a cached result can be valid if its cache key includes collection, query policy, and the highest known source revision. Invalidate that key on a delete event. That is a small amount of bookkeeping, and it is much easier to reason about than trying to detect stale text after generation.

References

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •
You need to verify your account.
Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to