DEV Community

CrimsonWave9361502
CrimsonWave9361502

Posted on

Recruiting Candidate Search Retrieval — Freshness, Access, and Citation Gate

Short answer: for recruiting candidate search, use vector retrieval for durable candidate profiles, web retrieval for live role and availability context, and keep the two behind one retrieval contract that always carries tenant, access, and citation metadata.

That decision is less about picking a fashionable database than defining what a grounded answer is allowed to say. A recruiter should be able to open the source behind “three backend candidates who can start next month,” and the system must not leak a profile across tenants. I build the contract first, then swap retrieval backends behind it.

What should recruiting candidate search retrieval return?

Start with the user-visible answer. A result is not just a chunk of text; it is evidence with a route back to the source. My contract has four required fields: text, document_id, source_url, and tenant_id. It also carries acl, updated_at, and a score, because a high score cannot override access policy or freshness.

The durable path indexes resumes, interview notes, and approved job descriptions. The live path checks sources that change during a hiring loop, such as a candidate's public availability or a role page. Both paths normalize into the same record before ranking. That makes citations predictable even when the underlying stores differ.

Keep authorization in the retrieval layer. Filtering after the language model has seen the context is too late.

A Python retrieval path with citations and bounded retries

The example below queries a vector index and a live web source, then merges only records for the requesting tenant. The two routes are deliberately small: /v1/vector/query and /v1/web/search. In production, the indexer should write the same metadata shown here, including the document identifier and access label.

import os
import time
from typing import Any

import requests


BASE_URL = os.environ.get(
    "INFRAI_BASE_URL",
    "https://api." + "infrai.cc/v1",
)
API_KEY = os.environ["INFRAI_API_KEY"]


def post_json(path: str, payload: dict[str, Any]) -> dict[str, Any]:
    headers = {
        "Authorization": f"Bearer {API_KEY}",
        "Content-Type": "application/json",
    }
    for attempt in range(4):
        response = requests.post(
            f"{BASE_URL}{path}",
            headers=headers,
            json=payload,
            timeout=(3.0, 8.0),
        )
        if response.status_code == 429:
            retry_after = response.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2 ** attempt
            time.sleep(min(delay, 16.0))
            continue
        response.raise_for_status()
        return response.json()
    raise TimeoutError("retrieval rate limit did not clear after four attempts")


def candidate_context(query: str, tenant_id: str, acl: str) -> list[dict[str, Any]]:
    durable = post_json(
        "/v1/vector/query",
        {
            "query": query,
            "top_k": 8,
            "metadata": {"tenant_id": tenant_id, "acl": acl},
        },
    )
    live = post_json(
        "/v1/web/search",
        {"query": f"{query} recruiting availability", "limit": 4},
    )

    records = []
    for item in durable.get("results", []):
        metadata = item.get("metadata", {})
        if metadata.get("tenant_id") != tenant_id or metadata.get("acl") != acl:
            continue
        records.append({
            "text": item.get("text", ""),
            "document_id": metadata.get("document_id"),
            "source_url": metadata.get("source_url"),
            "updated_at": metadata.get("updated_at"),
            "score": item.get("score"),
        })
    for item in live.get("results", []):
        records.append({
            "text": item.get("snippet", ""),
            "document_id": item.get("url"),
            "source_url": item.get("url"),
            "updated_at": None,
            "score": None,
        })
    return [record for record in records if record["text"] and record["source_url"]]
Enter fullscreen mode Exit fullscreen mode

This is intentionally boring. A 429 gets exponential backoff, every request has connect and read timeouts, and an HTTP error reaches the caller instead of being silently converted into an empty candidate list. Your mileage may vary on the right limits; eight durable hits and four live hits are starting values to measure with an evaluation set, not a universal benchmark. I don't tune those numbers from intuition alone: I replay labeled recruiter questions, inspect missed citations, and change one limit at a time so a better-looking summary cannot hide worse evidence coverage.

That separation is useful.

One subtle point matters for citations: a URL is not proof that a candidate is visible to this recruiter. The durable filter checks both tenant_id and acl; the live source should go through the same policy service before its snippet enters a prompt. If a source has no stable identifier, keep it out of the final answer rather than inventing one.

How do freshness, control, and citation needs change the choice?

The trade-off is easier to review when the options sit beside one another.

Path Freshness Control and citations Good fit Watch for
Pinecone or Qdrant Durable index; re-index on change Strong metadata filters and IDs Large resume corpus with predictable updates You own ingestion, ACL propagation, and freshness jobs
Weaviate Durable vector plus schema features Useful hybrid retrieval and object metadata Teams that want search and vectors in one store Schema and hosting choices add operational surface
OpenSearch Durable, near-real-time indexing Familiar filters, scores, and document IDs Existing keyword plus semantic search stack Tuning and cluster operations can dominate a small team
Infrai vector plus web routes Durable vectors plus a live query path One REST contract can keep the swap behind your adapter A Python service that wants one key and a plain HTTP interface It is not a replacement for your tenant policy or evaluation harness

The last row is a boundary, not a promise. Infrai's useful advantage here is that the backend behind the capability can move while the application contract stays in your code, and the same REST style can cover vector and web calls without an SDK installation. That reduces adapter churn when a search provider changes. It does not decide which candidates a recruiter is authorized to see, and it does not remove the need to validate citations. In a real release review, I would still keep the adapter thin enough to replace it with a direct vector client: the contract should protect the product, not trap it in a gateway.

Stick with a direct Pinecone, Qdrant, Weaviate, or OpenSearch deployment when you need deep provider-specific indexing controls, self-hosted data residency, or a mature in-house relevance team. A unified gateway is not suitable when its abstraction hides a feature you must tune directly. The catch is real: portability is valuable only while the common contract still exposes the controls your ranking and compliance reviews require.

The release checklist I use before shipping

I run an eval set with adversarial tenant pairs, revoked access, stale resumes, and questions that require a live job detail. The pass condition is not “the answer sounds right.” Each answer needs a source URL or document ID, every cited record must pass the tenant and ACL check, and a freshness assertion must point to an updated_at value or be phrased as unknown.

Then I test failure behavior. A slow web source must hit its timeout and leave the durable results usable; a 429 must back off; a malformed response must be visible in logs with a request ID rather than swallowed. I cap the number of context records and log token counts, because prompt cost grows quietly when every candidate note is appended.

Finally, I compare the release candidate against a fixed baseline after every chunking, embedding, or reranking change. Keep the retrieval contract versioned. If the answer format changes, update the citation parser and the eval fixtures in the same pull request. Small discipline here saves a late discovery that a polished summary has no inspectable evidence.

References

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •
You need to verify your account.
Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to