Pinecone with Java in AI: Building Semantic Search and RAG Pipelines
Introduction: why keyword search is not enough
Most enterprise applications still rely on keyword search. A user types "wire transfer limits" and the system looks for those exact words. A customer who asks "how much can I send abroad per day?" gets nothing, even though the answer is in the documentation. Semantic search fixes this by comparing meaning instead of words, and large language models (LLMs) make that comparison practical.
Retrieval-Augmented Generation (RAG) builds on semantic search. Instead of relying only on what a model learned during training, the application retrieves relevant passages from your own data and passes them to the model as context. The result is answers that are grounded in current, company-specific information.
A vector database is the storage layer that makes this work. Pinecone is a managed vector database that stores embeddings (numeric representations of text, images, or other content) and returns the nearest matches to a query vector with low latency. In this article we will build a Java service that indexes documents, stores them in Pinecone, and retrieves the most relevant chunks for a question.
Core concepts
Embeddings. An embedding model converts text into a vector, typically a few hundred to a few thousand floating-point numbers. Texts with similar meaning produce vectors that are close together under a distance metric such as cosine similarity.
Index. A Pinecone index stores vectors with a fixed dimension and a similarity metric. The dimension must match the embedding model you use. If you switch models later, you usually need a new index.
Namespace. A namespace partitions the vectors inside one index. It is a convenient way to separate tenants, environments, or document collections without creating many indexes.
Metadata. Each vector can carry a small JSON object such as source, department, or updatedAt. Pinecone can filter queries by metadata, which is essential in multi-tenant and access-controlled systems.
Chunking. Documents are split into passages before embedding. Chunks that are too large dilute meaning; chunks that are too small lose context. A common starting point is 300 to 800 tokens with a small overlap between neighboring chunks.
Architecture of the solution
The pipeline has two flows.
- Ingestion: load documents, split them into chunks, generate embeddings, and upsert vectors with metadata into Pinecone.
- Query: embed the user question, query Pinecone for the top-k nearest chunks (optionally filtered by metadata), build a prompt that includes those chunks, and call the LLM.
In a Spring Boot application, these responsibilities map cleanly to separate components: a DocumentChunker, an EmbeddingClient, a PineconeVectorStore adapter, and a RagAnswerService that orchestrates the flow. Keeping the Pinecone code behind an interface also makes it easy to test with a fake store and to swap providers later.
Setting up the Java client
Add the Pinecone Java SDK to your build. Check the official documentation for the current artifact coordinates and version before adding it, since the SDK evolves.
<dependency>
<groupId>io.pinecone</groupId>
<artifactId>pinecone-client</artifactId>
<version>VERIFY_LATEST_VERSION</version>
</dependency>
Store the API key outside source code, for example in an environment variable or a secrets manager:
pinecone:
api-key: ${PINECONE_API_KEY}
index-name: docs-index
namespace: product-docs
Create a configuration class that builds a single client instance per application:
@Configuration
public class PineconeConfig {
@Bean
public Pinecone pinecone(@Value("${pinecone.api-key}") String apiKey) {
return new Pinecone.Builder(apiKey).build();
}
}
Ingesting documents
Each chunk gets a stable ID, its embedding, and metadata that lets you trace the answer back to its source. Stable IDs matter: if you re-ingest a document, the same chunk should overwrite the old vector instead of creating a duplicate.
@Service
public class DocumentIngestionService {
private final EmbeddingClient embeddingClient;
private final Index index;
private final String namespace;
public DocumentIngestionService(EmbeddingClient embeddingClient,
Pinecone pinecone,
@Value("${pinecone.index-name}") String indexName,
@Value("${pinecone.namespace}") String namespace) {
this.embeddingClient = embeddingClient;
this.index = pinecone.getIndexConnection(indexName);
this.namespace = namespace;
}
public void ingest(String documentId, List<String> chunks, String source) {
List<float[]> embeddings = embeddingClient.embedAll(chunks);
for (int i = 0; i < chunks.size(); i++) {
String vectorId = documentId + "#" + i;
Map<String, Object> metadata = Map.of(
"source", source,
"chunkIndex", i,
"text", chunks.get(i)
);
index.upsert(vectorId, embeddings.get(i), namespace, metadata);
}
}
}
Two practical notes. First, batch your upserts when ingesting large corpora to reduce round trips and stay within rate limits. Second, storing the chunk text in metadata is convenient for small and medium datasets, but for large documents you may prefer to store only an identifier and fetch the text from your own database.
Querying for relevant context
The query path embeds the question with the same model used during ingestion, then asks Pinecone for the closest matches. Using a different embedding model at query time is one of the most common and hardest-to-debug mistakes in RAG systems.
@Service
public class SemanticSearchService {
private static final int TOP_K = 5;
private final EmbeddingClient embeddingClient;
private final Index index;
private final String namespace;
public SemanticSearchService(EmbeddingClient embeddingClient,
Pinecone pinecone,
@Value("${pinecone.index-name}") String indexName,
@Value("${pinecone.namespace}") String namespace) {
this.embeddingClient = embeddingClient;
this.index = pinecone.getIndexConnection(indexName);
this.namespace = namespace;
}
public List<RetrievedChunk> search(String question, String department) {
float[] queryVector = embeddingClient.embed(question);
Map<String, Object> filter = Map.of("department", Map.of("$eq", department));
QueryResponseWithUnsignedIndices response =
index.query(TOP_K, queryVector, null, null, namespace, filter, true, true);
return response.getMatchesList().stream()
.map(match -> new RetrievedChunk(
match.getId(),
match.getScore(),
match.getMetadata().getFieldsOrThrow("text").getStringValue()))
.toList();
}
public record RetrievedChunk(String id, double score, String text) {}
}
The metadata filter above is the important part for enterprise use. A user from one department should never receive chunks from another department's restricted documents, and filtering inside the vector query is far safer than filtering the results afterwards in application code. Filtering after retrieval can silently return fewer than k results and can also leak information through timing or logs.
Note that the exact method signatures and return types of the SDK can differ between versions. Verify them against the current Pinecone Java documentation before you copy the code into production.
Building the RAG answer
With the retrieved chunks in hand, the final step is to construct a prompt that asks the model to answer only from the provided context and to say so when the context is insufficient.
@Service
public class RagAnswerService {
private final SemanticSearchService searchService;
private final ChatClient chatClient;
public RagAnswerService(SemanticSearchService searchService, ChatClient chatClient) {
this.searchService = searchService;
this.chatClient = chatClient;
}
public String answer(String question, String department) {
List<SemanticSearchService.RetrievedChunk> chunks = searchService.search(question, department);
String context = chunks.stream()
.map(c -> "[" + c.id() + "] " + c.text())
.collect(Collectors.joining("\n\n"));
String prompt = """
Answer the question using only the context below.
If the context does not contain the answer, say that you do not know.
Cite the bracketed identifiers you used.
Context:
%s
Question: %s
""".formatted(context, question);
return chatClient.prompt(prompt).call().content();
}
}
Spring AI provides a ChatClient abstraction and vector store integrations that can handle much of this plumbing. Whether you use the Spring AI abstraction or the Pinecone client directly, the design principles are the same: keep retrieval, prompt construction, and generation separate so each can be tested and tuned on its own.
Operational considerations
Dimension and model alignment. Create the index with the same dimension as your embedding model. Record the model name in configuration and, ideally, in metadata, so you can detect mismatches during a migration.
Re-ingestion strategy. Use deterministic IDs and version your chunking logic. When the chunking strategy changes, re-index the whole corpus into a new namespace, validate it, and then switch the application over.
Cost and latency. Embedding calls and LLM calls usually dominate cost, not the vector store. Cache embeddings for repeated questions and set reasonable limits on top-k and context length.
Evaluation. Build a small evaluation set of real questions with known source documents. Measure retrieval quality (is the right chunk in the top k?) separately from answer quality. Most RAG problems that look like "the model is wrong" are actually retrieval problems.
Security and privacy. Apply metadata filters based on the authenticated user's permissions, never on user-supplied filter values. Avoid putting sensitive personal data into metadata unless you have a clear retention and deletion policy.
Observability. Log query latency, the number of matches returned, the similarity scores, and the IDs of the chunks used in each answer. Scores that drop sharply are often an early warning that the corpus or the embedding model has drifted.
Best practices checklist
- Use one embedding model per index and document it.
- Keep chunks semantically coherent; split on headings and paragraphs before falling back to fixed sizes.
- Store stable, deterministic vector IDs so re-ingestion is idempotent.
- Enforce access control with metadata filters inside the query.
- Keep API keys in a secrets manager, never in source control.
- Set timeouts and retries around every external call.
- Build an evaluation set before tuning anything.
- Treat retrieved text as untrusted input when building prompts, and guard against prompt injection hidden in documents.
Conclusion and key takeaways
Pinecone gives Java teams a production-grade vector store that fits naturally into Spring Boot services. Combined with an embedding model and an LLM, it turns static documents into a searchable knowledge base that answers questions in context.
The key ideas to take away:
- Semantic search compares meaning, which is why it succeeds where keyword search fails.
- The embedding model used at ingestion and query time must be the same.
- Metadata filtering inside the query is the right place to enforce access control.
- Most RAG quality problems are retrieval problems, so measure retrieval first.
- Keep the Pinecone integration behind an interface so the rest of your application stays clean and testable.
Start small: index one document collection, build one question-answering endpoint, and evaluate it honestly. The architecture described here scales from a prototype to a multi-tenant platform without changing its core shape.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.