DEV Community

Cover image for I Measured My RAG Pipeline Honestly. It Was 40x Slower Than I Thought.
Victor
Victor

Posted on

I Measured My RAG Pipeline Honestly. It Was 40x Slower Than I Thought.

A few days ago I published the architecture behind Vicquant’s RAG Vault; a 9-stage retrieval pipeline built to ground financial AI answers in actual source documents, with strict citations, so a user asking “what’s my 401(k) contribution limit” gets an answer tied to a real document, not a model’s confident guess. The quality evaluation backed it up: +5.2% answer relevance, +16% context precision, +4% context recall over the simpler pipeline it replaced, with zero hallucinations measured on either version.

What I didn’t have yet was an honest measurement of how fast it actually was. This is the story of getting that number, not liking it, and fixing it properly.

The number that was true and meaningless

My first latency measurement showed the pipeline completing in about 4.56 milliseconds. That number was real, and completely useless, because the test suite that produced it had the network-dependent calls (the hosted embedding API, the LLM calls for HyDE and generation) running against mocked test doubles, not the real OpenRouter API. It measured the pipeline’s own logic, correctly, while saying nothing about what a user would actually experience.

So I built a live benchmark: the same 25-question golden evaluation set, run through the real pipeline against the real API, three times, with full per-stage timing.

Mean latency: 14.81 seconds. p95: 21.28 seconds.

That’s not a number you publish as a feature. That’s a number you fix.

Finding the actual levers

The full pipeline runs several sequential steps that each cost real time: HyDE generates a hypothetical document via an LLM call, multi-query expansion generates alternate phrasings via another LLM call, then retrieval, reranking, and a final generation call. Three separate cloud LLM calls, done carefully but done on every single query, including ones that had no reason to need any of it.

_The first and biggest fix was routing: _Most of what a financial advisor app gets asked; budgeting questions, “how am I doing this month,” general conversation, doesn’t reference an uploaded document at all. Only questions actually grounded in RAG Vault content need the full apparatus. I added a lightweight, sub-millisecond heuristic classifier to route document-grounded questions into the full pipeline and everything else into the existing fast direct-chat path. Tested against a realistic 75-query mix (25 document-grounded, 50 general), it classified with 100% accuracy at roughly 0.09 milliseconds of overhead — a rounding error compared to what it saves.

The second fix was concurrency: HyDE generation and multi-query expansion don’t depend on each other’s output, but they were running one after another. Running them with asyncio.gather instead cut roughly 20% off the document-grounded path on its own.

The third fix was streaming: Total response time matters, but felt response time matters more for whether something feels broken. Adding SSE token streaming meant the fast path now renders its first token in under a second, instead of a user staring at nothing until the full response assembles.

Combined, these three changes brought mean latency from 14.81s down to 6.24s, and median latency from 14.71s down to 3.61s. A real improvement. Still not where I wanted to land.

The variance that wasn’t random

Even after those fixes, something was off: identical questions, identical model, wildly different response times from one call to the next. It looked like noise. It wasn’t.

OpenRouter doesn’t always serve a given model from the same place. A model string like meta-llama/llama-3.1-8b-instruct can be fulfilled by several different upstream inference providers, and OpenRouter rotates between them per request. I added logging to capture which provider actually served every call, then ran twelve identical back-to-back requests with the exact same prompt and model string to see what would happen.

Three different providers answered across those twelve calls; Groq, Novita, and DeepInfra and the latency gap between them was stark:

DeepInfra was roughly 11x slower than Groq for identical work. That single provider, showing up unpredictably in the rotation, was the entire source of the long tail in my earlier benchmarks, the 2-to-14-second HyDE call spread I’d measured wasn’t jitter, it was DeepInfra showing up some of the time.

OpenRouter supports explicit provider ordering in the request body, so I pinned it: Groq first, Novita second, DeepInfra kept as a last-resort fallback (never removed entirely because availability matters more than shaving off a few hundred more milliseconds when both fast providers happen to be down).

Every one of the 75 test queries now completes in under one second. That’s a 97.5% reduction in mean latency from where this started.

The cost of pinning to faster providers was functionally zero about $0.000003 extra per query, something like three-thousandths of a cent. And at every single stage of this entire process, I re-ran the quality evaluation: faithfulness, answer relevance, context precision, and context recall moved by exactly 0.0000 throughout. None of these fixes touched what the pipeline retrieves or generates, only when the calls happen and which infrastructure answers them.

What this actually proves

Three things I’d rather have learned on purpose than by accident, written down for whoever reads this next:

A benchmark is only honest if it measures the thing a user actually experiences. Mocked-out network calls will always look fast, that’s not a benchmark, it’s a logic test wearing a benchmark’s clothes.

“Random” variance usually isn’t. The twelve-identical-calls test took a few minutes to write and turned an unexplained 6x latency spread into a fully understood, fully fixable root cause.

And the last one: speed and correctness aren’t actually in tension here, even though it’s tempting to assume they are. Every optimization in this piece made the pipeline faster without moving a single quality metric. The tradeoff I was bracing for, give up some grounding rigor for responsiveness and was never had to be made.

RAG Vault is done. Next up: what Vicquant needed a mobile-native rebuild to get right, and why a web wrapper wasn’t going to cut it.

Top comments (0)