Previous articles covered the individual pieces of RAG. You now know what embeddings are, how to chunk documents, and how Pinecone stores and searches vectors. This article connects those pieces into a working pipeline in Spring Boot — and shows exactly what happens at each step when a real question comes in.
Two flows, one shared store
Every RAG system has two separate flows:
Ingest — runs once per document. Takes raw text, splits it into chunks, converts each chunk into a vector, stores it in Pinecone.
Query — runs on every user message. Takes the question, converts it to a vector, finds the most relevant chunks, injects them into the prompt, sends the enriched prompt to the model.
These two flows are completely independent. They share exactly one thing: the vector store. The ingest flow writes to it. The query flow reads from it.
The ingest flow
Splitting and storing
@Service
public class IngestService {
private final VectorStore vectorStore;
private final TokenTextSplitter splitter;
public IngestService(VectorStore vectorStore) {
this.vectorStore = vectorStore;
this.splitter = new TokenTextSplitter(350, 50, 10, 10000, true);
}
public void ingest(String sourceId, String title, String content) {
vectorStore.delete(List.of(sourceId));
Document doc = new Document(content, Map.of(
"source", sourceId,
"title", title
));
List<Document> chunks = splitter.split(doc);
vectorStore.add(chunks);
}
}
vectorStore.delete() runs first — it removes any previously stored chunks from this document before adding new ones. Skip this step and re-uploading a document creates two sets of chunks in Pinecone. Retrieval will silently mix answers from both versions.
vectorStore.add(chunks) does three things in one call: embeds each chunk using text-embedding-004, then stores the text, the vector, and the metadata together in Pinecone. You don't call the embedding API separately — Spring AI handles it.
Ingest endpoint
@RestController
@RequestMapping("/api/documents")
public class IngestController {
private final IngestService ingestService;
@PostMapping
public ResponseEntity<String> ingest(@RequestBody IngestRequest request) {
ingestService.ingest(request.sourceId(), request.title(), request.content());
return ResponseEntity.ok("Ingested: " + request.sourceId());
}
}
Try it:
curl -X POST http://localhost:8080/api/documents \
-H "Content-Type: application/json" \
-d '{
"sourceId": "return-policy-v2",
"title": "Electronics Return Policy",
"content": "Electronics purchased at full price may be returned within 30 days with original packaging and receipt..."
}'
The response is just "Ingested: return-policy-v2". Nothing tells you how many chunks were created or what they contain — the work happened inside vectorStore.add().
The query flow
Wiring retrieval into the chat client
QuestionAnswerAdvisor is the piece that connects the vector store to every chat request. Configure it once on the ChatClient bean and it runs automatically on every message — you never call Pinecone directly from your controller.
@Configuration
public class ChatConfig {
@Bean
public ChatClient chatClient(ChatModel chatModel, VectorStore vectorStore) {
return ChatClient.builder(chatModel)
.defaultAdvisors(
new MessageChatMemoryAdvisor(new InMemoryChatMemory()),
new QuestionAnswerAdvisor(
vectorStore,
SearchRequest.defaults().withTopK(5)
)
)
.defaultSystemPrompt("""
You are a helpful assistant.
Answer ONLY based on the provided context.
If the context does not contain enough information, say so clearly.
Do not use knowledge from outside the provided context.
""")
.build();
}
}
Two things here worth understanding deeply:
withTopK(5) tells the advisor to retrieve the 5 most similar chunks for every query. That number directly affects what the model receives — too few and the answer might be in the 6th chunk, too many and the model gets overwhelmed with context. Five is a reasonable starting point.
The system prompt instruction is the enforcement mechanism for grounding. Answer ONLY based on the provided context tells the model not to use its training knowledge to fill gaps. Without this instruction, the model will — confidently, and invisibly — supplement retrieved context with things it learned during training. That's the exact problem RAG is supposed to solve.
Chat endpoint
@RestController
@RequestMapping("/api/chat")
public class ChatController {
private final ChatClient chatClient;
@PostMapping
public String chat(@RequestBody ChatRequest request) {
return chatClient
.prompt()
.user(request.message())
.advisors(a -> a.param(
AbstractChatMemoryAdvisor.CHAT_MEMORY_CONVERSATION_ID_KEY,
request.conversationId()
))
.call()
.content();
}
}
What the model actually receives
This is the detail that makes the whole system make sense.
When a user asks "What is the return policy for electronics?", the model doesn't receive that question alone. By the time QuestionAnswerAdvisor runs, the prompt looks like this:
System:
You are a helpful assistant.
Answer ONLY based on the provided context.
If the context does not contain enough information, say so clearly.
Context information is below.
---------------------
[Chunk 1: "Electronics purchased at full price may be returned within 30 days
with original packaging and receipt. Items must be in original condition."]
[Chunk 2: "Exceptions: laptops and tablets cannot be returned after the seal
is broken unless the item is defective."]
[Chunk 3: "To initiate a return, contact customer support with your order
number. Refunds are processed within 5–7 business days."]
---------------------
Given the context information and not prior knowledge, answer the question.
User: What is the return policy for electronics?
The model reads those chunks and answers from them. It did not search your documents — you retrieved the relevant sections and placed them directly in front of it. The model's job is to read and summarise. Your job is to retrieve well.
When there is no answer in the documents
Ask the system something that isn't in any of your uploaded documents. Pinecone still returns 5 chunks — the 5 least-dissimilar ones regardless of actual relevance. The model receives those chunks, finds no useful information in them, and because of the system prompt responds: "I don't have enough information to answer that based on the available documents."
That is the correct behaviour. If the system gives a detailed, confident answer to a question that isn't in your documents, check the system prompt — the grounding instruction is either missing or being ignored.
You can also add a similarity threshold to reject low-quality retrievals outright:
SearchRequest.defaults()
.withTopK(5)
.withSimilarityThreshold(0.70)
Chunks with similarity below 0.70 are discarded before they reach the prompt. If nothing passes the threshold, the model gets an empty context section and says so.
Tuning retrieval quality
Once the pipeline is running, these are the variables to adjust — in order of how much impact they have:
Chunk size — the biggest lever. Vague answers that feel like summaries usually mean chunks are too large. Drop from 350 to 200–250 tokens and re-ingest. Fragmented answers that miss context usually mean chunks are too small — try going larger.
topK — if the correct answer is in your documents but the model keeps missing it, the relevant chunk might be ranked 6th or 7th. Try increasing topK to 7 or 8.
Similarity threshold — add it only after you observe irrelevant chunks appearing in the context and causing wrong answers. Starting without a threshold is fine.
Test changes by writing 10–15 representative questions, running them before and after, and comparing the answers. Intuition alone won't tell you which direction to go.
A common failure mode
There's one failure that appears after the pipeline is running and the initial tests look fine: a question comes in that the documents don't answer, and the system gives a confident, detailed response anyway.
This happens when the system prompt grounding instruction is missing or soft. During initial testing, questions tend to be things the documents do answer — so the failure is invisible. It surfaces when real users ask things outside the scope of what was uploaded.
Always test the "not in the documents" case before considering the pipeline done. Ask something the system could not possibly know. If the answer sounds specific and confident, the system prompt is not enforcing grounding.
What's next
RAG gives the model knowledge from your documents. But knowledge isn't enough for every task. Some things require the model to act — query a database, call an API, run a multi-step process.
That's what AI agents do. The next phase of this series covers how agents work, what the ReAct loop is, and how to build one in Spring Boot.
Sham Prakash K — Backend Engineer, 4+ years in Java, Spring Boot, and distributed systems. Building AI backend infrastructure. Writing about what I actually learned, mistakes included.
Top comments (1)
Official Platform Update
Security protocols have been updated for all developer accounts.