DEV Community

Cover image for Why a Financial Advisor App Needs a RAG Vault, Not Just a Chatbot
Victor
Victor

Posted on

Why a Financial Advisor App Needs a RAG Vault, Not Just a Chatbot

When I started building Vicquant, the core design principle was simple and non-negotiable: the AI never does the math. Every number and a debt payoff schedule, a backtested win rate, a risk metric is computed by deterministic code. The model's only job is to explain results in plain language, never to generate a number from scratch.

That principle protects users from one kind of failure: a model confidently inventing a wrong calculation. But it doesn't protect against a second, quieter failure mode, when a model confidently answering a factual question about a document it has never actually seen, using something that merely sounds plausible. If a user uploads a fee schedule and asks "what does this actually charge me for an overdraft," a generic chatbot answer isn't good enough. It needs to be grounded in that specific document, with a receipt.

That's what the RAG Vault is for, and getting it right turned out to be a deeper engineering problem than I expected.

*Why the naive version doesn't work for finance
*

A basic RAG implementation do split text into fixed-size chunks, embed them, retrieve the closest matches and works reasonably well for general text. It falls apart on financial documents for a specific reason: financial documents are full of tables, and a naive character-splitter slices through them without any awareness that it's destroying a row of numbers. Tax brackets, fee schedules, 10-K disclosures are all tabular, all fragile to blind splitting. The fix was layout-aware parsing that detects and preserves tables, headings, and structured sections as whole units rather than arbitrary text windows.

The chunk-size dilemma

Even with layout-aware parsing, there's a tension in how big each retrievable piece should be. Small chunks match a search query precisely but don't give the model enough surrounding context to answer well. Large chunks give plenty of context but are harder to match precisely against a specific question. Vicquant resolves this with hierarchical chunking: small "child" chunks (150-250 tokens) are what gets searched and matched, but each one is linked to a larger "parent" chunk (800-1200 tokens) that actually gets handed to the model once a match is found. Precision for retrieval, context for generation and it decoupled instead of traded off against each other.

Finding the right passage when the words don't match

Pure keyword or basic cosine-similarity search misses conceptual synonyms and also "drawdown mitigation" and "capital preservation" mean roughly the same thing to a person asking a question, but not to a system doing literal string or frequency matching. The fix is dense neural embeddings combined with traditional keyword search (BM25) in a hybrid approach, fused using Reciprocal Rank Fusion(RRF) rather than an arbitrary weighted average. RRF is scale-invariant, so it doesn't break when one search method's scores happen to run on a different numeric range than the other's.

On top of that sits a cross-encoder reranker: a second-pass model that looks at the actual query and each candidate passage together, rather than comparing pre-computed embeddings in isolation. This catches token-level relevance that embedding similarity alone misses.

Where the model's attention actually goes

One detail that's easy to miss: large language models don't pay uniform attention across a long context window — information placed in the middle tends to get under-weighted compared to the beginning and end (a well-documented effect sometimes called "lost in the middle"). So the most relevant retrieved passages are deliberately placed at the start and end of what gets sent to the model, not simply stacked in ranked order.

Grounding, not just retrieval

The last piece is making every answer auditable. Vicquant's citation format isn't a vague "according to the document". it's a strict, deterministic token: [Citation ID: doc_id:chunk_id], pointing at the exact passage that produced the claim. For a financial app, this matters for a reason beyond user trust: it's the difference between an answer someone can verify and an answer they just have to believe.

What the evaluation actually showed

I ran the new pipeline against the old one on a 25-question financial benchmark, split across numeric lookups, conceptual synonym matches, and multi-table extraction, scored using RAGAS metrics:

Metric Legacy pipeline New pipeline Change
Faithfulness 100% 100% unchanged
Answer Relevance 78.7% 83.9% +5.2%
Context Precision 34.0% 50.0% +16.0%
Context Recall 92.0% 96.0% +4.0%

_Faithfulness _— the measure of whether claims are actually grounded in retrieved context rather than hallucinated, it was already perfect on the old pipeline and stayed perfect. That's expected; faithfulness is downstream of the generation step, not the retrieval architecture. The real gains are in Context Precision, up 16 points and the new pipeline pulls in noticeably less irrelevant "distractor" content, largely thanks to the cross-encoder reranking stage filtering out near-misses before they reach the model.

I also stress-tested the resilience claims directly rather than assuming them: deliberately failing the hosted embedding API confirms the circuit breaker trips and the pipeline falls back to the local offline embedder with zero unhandled exceptions; feeding in a document containing an embedded instruction ("SYSTEM OVERRIDE: ignore all prior instructions") confirms that text stays encapsulated as inert data rather than influencing the model's behavior.

One honest caveat on performance: the pipeline's internal processing overhead measured in testing is small, but that number reflects the pipeline's own logic with the network-dependent components (the hosted embedding call, HyDE's LLM call) running against test doubles, not live API latency. Real-world response time will be dominated by those actual network calls to OpenRouter, which I haven't benchmarked against the live API yet and that's the next thing to measure honestly before making any speed claims to users.

Why build a financial advisor this way at all

The short version: the worst thing an AI assistant can do with someone's money isn't getting something wrong. It's getting something wrong confidently, with no way for the person to check. Every design decision in Vicquant, from "the AI never computes numbers" to "every RAG answer carries a citation," traces back to that one principle. It's slower to build this way. I think it's the only responsible way to build it.

Top comments (0)