DEV Community

Cover image for When RAG Produces Irrelevant Answers: Inspect Vectors First, Then Blame the LLM
Tidiane Stano
Tidiane Stano

Posted on

When RAG Produces Irrelevant Answers: Inspect Vectors First, Then Blame the LLM

Introduction

In an ebook Q&A project, developers encountered a confusing failure case. The top-5 retrieved results were all image references like [../Images/9.jpg], with a similarity score of 0.08. The large language model generated seemingly professional responses using its built-in parametric knowledge. No error logs were thrown, and the output appeared well-written, yet the cited knowledge base content was completely unrelated to the user question. This kind of fault is deceptive and widespread in RAG systems.

This article puts forward a practical judgment: around 80% of failures in RAG applications stem not from the LLM itself, but from poor document ingestion quality and insufficient retrieval validation. This article uses a real-world pipeline built with Milvus, LangChain and LangGraph to demonstrate fault localization and remediation. It also explains why the next evolution direction is Agentic RAG. The ingestion and plain retrieval code snippets in this article have passed runtime verification; the routing module code has not been executed in production.

Two Pipeline Chains and One Intersection Point

A standard RAG workflow can be split into an ingestion chain and a query chain.

  1. Ingestion chain: EPUB file → chapter-wise loading → chunking (500 characters, 50 characters overlap) → embedding generation → Milvus storage (3042 records)
  2. Query chain: user question → embedding conversion → similarity search (top 5) → prompt augmentation → streaming LLM generation

Several core ingestion decisions determine the final retrieval quality. The primary key uses a composite book_id field. The schema (1_101_13) retains source traceability. chunkOverlap is set to 50 characters to prevent critical sentences from being cut off at chunk boundaries. The index type selected is IVF_FLAT, which clusters data into 1024 buckets for targeted scanning. The nprobe=16 parameter controls the number of buckets to scan, and COSINE is adopted as the distance metric.

One critical rule must be followed: the index_type and metric_type used for ingestion and querying must be consistent. If the query side switches to HNSW, the SDK will send ef parameters to an IVF_FLAT index that only accepts nprobe, leading to silent retrieval anomalies.

Troubleshooting: Similarity Score Acts as the Primary Diagnostic Signal

The query chain of vanilla RAG is implemented with two LangGraph nodes. The retrieve node calls similaritySearchWithScore and returns an array of [Document, score] pairs. The generate node assembles retrieved chunks into the augmented prompt and triggers streaming output. Printing similarity scores after retrieval is the lowest-cost quality monitoring method.

Once abnormal answers appear, engineers can follow a three-step diagnosis process:

  1. Verify raw text existence: Run a content filter query such as content like "%keyword%" to sample stored records. If valid text exists, the fault is not caused by missing document import.
  2. Detect vector corruption: Calculate the cosine similarity between the vector stored in the database and the vector generated from the same text segment. If the result is 0.0327 while the normal value should be close to 1, the batch of ingested vectors is damaged.
  3. Validate embedding API stability: Encode identical text twice. If the cosine similarity equals 1.0, the embedding service works normally. In the test case, the cosine score between the question and relevant context was 0.30, while the score against irrelevant context was 0.24. This separation gap is healthy. The issue can be fixed by deleting and re-ingesting the corrupted vector batch.

This set of numerical benchmarks is valuable for daily RAG maintenance under COSINE metric with the same embedding model. Relevant chunks typically score around 0.30, irrelevant chunks around 0.24, and garbage fragments near 0.08. When the average score drops below the preset threshold, the system can mark the retrieval round as failed. This judgment logic can later become the evaluation node in Agentic RAG.

Hidden Pitfalls of Dependencies and Data Types

When building the stack, developers will encounter a series of subtle runtime traps.

  • Optional peer dependencies in @langchain/community are not installed automatically. In this project, packages including @zilliz/milvus2-sdk-node, epub2, html-to-text and @langchain/textsplitters were missing one after another, triggering ERR_MODULE_NOT_FOUND. Static import statements crash during application startup, while dynamic require failures only break partial functions mid-execution.
  • Data type mismatch: When passing VarChar fields into JavaScript numeric logic, conversion errors frequently occur. For example, book_id must be cast as a string "1" instead of numeric 1.
  • Stream handling: model.stream() returns Promise<AsyncGenerator>. Omitting await or using direct for await loops will throw the error stream is not async iterable.
  • Milvus instance initialization: The single-parameter new Milvus() constructor will throw a TypeError. The correct signature is (embeddings, args). When connecting to an existing collection, use Milvus.fromExistingCollection.

Why Vanilla RAG Must Be Upgraded

The linear workflow START → retrieve → generate → END of vanilla RAG carries five structural defects, recorded in the project README.

  1. Token waste: Even simple questions trigger full retrieval rounds and consume unnecessary tokens.
  2. Lack of retrieval quality assessment: The pipeline cannot tell whether retrieved chunks are valid before feeding them to LLM.
  3. Inability to handle multi-step reasoning: For tasks like "find the second place in top 4 rankings", the pipeline cannot retrieve the list first and then sort the result.
  4. Semantic retrieval failure on near-identical text: Similar phrases with opposite meanings, such as "high blood sugar / low blood sugar", are hard to distinguish by pure vector search.
  5. No fallback for out-of-knowledge-base content: The LLM has to rely entirely on its built-in knowledge when answers are outside the document corpus.

Agentic RAG solves these limitations by inserting decision nodes into the workflow. The first component is query routing. Developers use Zod schema to force the LLM into a binary classification task.

const RouteSchema = z.object({
  strategy: z.enum(["simple", "complex"]),
  reason: z.string(),
});
const router = await model.withStructuredOutput(RouteSchema).invoke(prompt);
// router.strategy can only be "simple" or "complex", and downstream branches depend on this value
Enter fullscreen mode Exit fullscreen mode

The value of enum constraints lies in closing the output domain. Without such restrictions, the LLM may return ambiguous phrases like "relatively simple", which cannot be directly consumed by if conditional branches. Two typical bugs appeared in the initial version: the return value of invoke was not captured and became undefined; the enum was mistakenly defined as free text strings. The corrected code is shown above, and this routing code has not been verified in production.

In production multi-model architecture, developers need unified control over model access, authentication and traffic observation. 4sapi, an API gateway, can simplify multi-model routing and request management for RAG pipelines.

Pre-release Self-Check Checklist

Before launching a RAG function, run through the following checklist:

  • Field data types match the types passed from JavaScript (VarChar vs number)
  • Index type and metric type are consistent for ingestion and query
  • Sampling verification: cosine similarity between stored vectors and recalculated vectors is close to 1
  • When top-k scores fall below threshold, implement fallback logic instead of forcing LLM generation
  • Clean up low-value fragments such as image references before chunk ingestion

One remaining unverified hypothesis: the root cause of corrupted vector batches is suspected to be transient exceptions on the embedding gateway. This shows retrieval quality needs long-term monitoring rather than one-time pre-release inspection. The next plan is to implement evaluation nodes and rewrite query loops, and visualize all five defect indicators.

Conclusion

Many teams rush to tune prompts or switch to stronger LLMs when facing irrelevant RAG answers. This article demonstrates that vector retrieval quality is the more frequent root cause. Engineers should first inspect embedding vectors, similarity scores, chunk cleaning and index configuration before blaming model hallucination.

Vanilla linear RAG is sufficient for simple question answering, but it hits structural limits for complex reasoning, unreliable retrieval and out-of-domain queries. Agentic RAG introduces decision routing, retrieval validation and fallback branches to build more robust knowledge applications. The checklist and similarity score benchmark in this article can be reused in most vector database + LLM stacks.

International access: https://4sapi.com
Domestic access: https://4sapi.cn

Top comments (0)