DEV Community

Cover image for Retrieval is a routing problem. Your RAG stack just hides it.
TokenLat
TokenLat

Posted on

Retrieval is a routing problem. Your RAG stack just hides it.

Most RAG failures aren't model failures. They're routing failures wearing a retrieval costume.

We spend enormous effort choosing the right model for a RAG pipeline — the right context window, the right reasoning strength, the right price tier — and then we feed every query through the exact same retrieval path. Same vector index. Same top_k. Same chunking. Same reranker. A yes/no factual lookup and a ten-document synthesis get treated as the same shape of problem.

That isn't a retrieval step. That's a routing decision we forgot to make.

The retrieval step decides your model for you

Retrieval is upstream of model selection. The chunks you pull determine what context the model sees, and the shape of that context determines which model is even appropriate. If retrieval drowns a simple lookup in noise, no model choice saves it. If retrieval starves a synthesis of breadth, the strongest model just hallucinates confidently.

Yet we treat retrieval as fixed preprocessing — a hardcoded preamble the model has to swallow whole. The leverage is right there, and we've bolted it shut.

Three intents, three retrieval strategies

The same top_k=5 vector search serves three very different jobs badly:

  • Lookup — "What's our refund policy for cross-border orders?" Needs precision. Narrow, high-recall-on-target, maybe keyword + metadata filter, low tolerance for irrelevant context. Top-1 to top-3.
  • Synthesis — "Summarize the Q3 incident reports." Needs breadth. Multi-source, high-recall, dedup, section-aware assembly. Starved by a small k.
  • Comparison — "How does our p99 latency compare to the regional average?" Often isn't a vector problem at all. Structured extraction, maybe a SQL or semi-structured path, not free-text chunks the model has to reverse-engineer into numbers.

Feed all three through one undifferentiated retrieval call and you get the worst of each: lookups drowned, syntheses starved, comparisons given prose it can't reason over.

Concretely, a retrieval plan keyed by intent — not hardcoded — looks like this:

# retrieval plan, selected by classified intent
strategies:
  lookup:
    index: policy_dense
    k: 3
    filter: "doc_type:policy AND region:{user_region}"
    reranker: none
  synthesis:
    index: [incidents_dense, incidents_faq]
    k: 12
    filter: "date >= now() - 90d"
    reranker: cross_encoder
    fusion: rrf
  comparison:
    backend: sql_extract        # not a vector path at all
    source: metrics_db
    output: structured_rows
Enter fullscreen mode Exit fullscreen mode

The defensible lever: make retrieval a routable step

The fix is to pull retrieval up into the routing layer. Before the query hits the index, classify intent and select a retrieval strategy — which index, what k, what filter expression, which reranker, what fusion method. This is the same shape as model routing: you're deciding which backend serves this step.

And here's the constraint that shapes the whole architecture: context is a budget, not a free resource. More retrieved tokens are not better — irrelevant context dilutes attention (the same dilution problem that makes an unmeasured step expensive). So retrieval routing is simultaneously a quality router and a cost router. With 25+ models behind an OpenAI-compatible interface, the router can pair "cheap retrieval + cheap model for lookup" against "broad retrieval + stronger model for synthesis" — but only if retrieval is itself a routable step, not a hardcoded preamble.

Compliance adds a second axis. For PDPA-aligned workloads, the source index and the model that reads it may need to live in the same region — SG-hosted, not shipped offshore. Retrieval routing then has a compliance gate before a quality gate: you can't route a strategy that would move regulated data out of jurisdiction, no matter how good its recall looks.

The pattern, concretely

A retrieval router takes (query, metadata) and emits a retrieval plan: { index, k, filter, reranker, fusion }. The model router then consumes the projected context shape as an input — not just the query text. Routing becomes two-stage: retrieve-shape → model-shape.

And the failover lesson carries straight over from model routing: if the primary retrieval backend degrades, the router should have a fallback retrieval path — keyword, SQL, a secondary index — not just a blind retry of the same call. A retried vector search that's already degraded returns the same degraded result, just billed twice.

The two-stage flow, side by side with the naive path:

flowchart LR
  Q[Query] --> C{Intent classifier}
  C -->|lookup| L[policy index · k=3 · filtered]
  C -->|synthesis| S[incidents · k=12 · RRF + rerank]
  C -->|comparison| P[SQL extract · structured rows]
  L --> M[Model router]
  S --> M
  P --> M
  M --> A[Answer]
  subgraph Naive[Naive RAG]
    NQ[Query] --> NV[one vector index · top_k=5]
    NV --> NM[Model]
  end

And the function that produces the plan — note the compliance gate sits before the quality gate, and the failover is a different path, not a retry:

def route_retrieval(query, metadata):
    intent = classify_intent(query)      # lookup | synthesis | comparison
    plan = STRATEGY_TABLE[intent]        # {index, k, filter, reranker, ...}

    # compliance gate BEFORE quality gate
    if not in_jurisdiction(plan.index, metadata.region):
        plan = fall_back_to(region_local_index)

    context = retrieve(plan, query)
    if context.degraded:                 # primary backend soft-failed
        context = retrieve(FAILOVER[plan], query)  # different path, not a retry
    return context   # model router consumes context.shape, not just the query
Enter fullscreen mode Exit fullscreen mode

The question I keep landing on

Do you route retrieval by query intent, or does the same top_k serve every question in your stack? If you split retrieval strategies, where does that decision live — before the query reaches the index, or after the model has already committed to an answer?

Top comments (1)

Collapse
 
max_quimby profile image
Max Quimby •

"Comparison often isn't a vector problem at all" is the sentence I'd tattoo on a few RAG pipelines I've seen. Forcing a p99-vs-average question through free-text chunks and hoping the model reverse-engineers numbers out of prose is where so much silent inaccuracy comes from.

The one tension I'd surface: the moment you make retrieval routable, the intent classifier becomes a new single point of failure sitting upstream of everything. If it misroutes a synthesis query as a lookup, you've starved it before the model sees a token — and that failure is invisible because retrieval "succeeded." Two things that made this safe for us: keep the classifier cheap and mostly deterministic (keyword + metadata rules first, small model only for the ambiguous tail), and always fail open to a reasonable default strategy rather than erroring. Also worth deciding early how you handle genuinely mixed-intent queries ("summarize the incidents and compare our latency to baseline") — do you run multiple strategies and merge, or force a single classification? Curious where you landed on that.