DEV Community

Cover image for Vector Databases — What They Are and How to Use Pinecone with Spring AI
Sham Prakash K
Sham Prakash K

Posted on AI-assisted

Vector Databases — What They Are and How to Use Pinecone with Spring AI

After chunking and embedding, you have a collection of vectors — one for each chunk of your documents. Now you need somewhere to store them, and more importantly, somewhere that can answer the question: "which of these stored vectors is closest in meaning to this new query vector?"

That's not a question a regular database is designed to answer.

What a dimension actually is

Before talking about vector databases, it's worth making "dimensions" concrete — because this is where most explanations lose people.

A dimension is just one number in the vector. A 3-dimensional vector has 3 numbers. A 5-dimensional vector has 5 numbers.

Imagine you described every piece of text along just two dimensions: how technical it is and how positive its tone is. You could plot every chunk on a flat 2D plane:

positive tone
      ↑
      │  "Great product, fast shipping!"        ● 
      │                          ● "Easy setup guide"
      │
──────┼──────────────────────────────────────────→ technical
      │           "JWT token validation fails" ●
      │  "Database connection pool exhausted" ●
      ↓
negative tone
Enter fullscreen mode Exit fullscreen mode

Chunks that are similar in both dimensions end up close together. That's the entire idea.

Real embedding models don't use 2 dimensions — they use hundreds or thousands, because no two dimensions can capture all the nuance of language. But the intuition is the same: each dimension captures some aspect of meaning, and vectors with similar meanings end up close together in that space — regardless of how many dimensions there are.

The number of dimensions is fixed per embedding model. Google's text-embedding-004 uses 768. OpenAI's text-embedding-3-small uses 1536. You don't choose it — the model decides it. What matters is that you use the same model (and therefore the same number of dimensions) consistently.


Why a regular database can't search by similarity

A standard database index — a B-tree — works by sorting values so it can binary-search through them. Looking for rows where price < 50? The database jumps to the right part of the sorted index and reads from there.

Vector similarity search asks a different question: "which stored vectors point in the most similar direction to this query vector?" You can't sort vectors in a way that puts similar ones next to each other in a linear index. Similarity in high-dimensional space doesn't work like sorting numbers on a line.

The naive solution — compare the query vector against every single stored vector and rank them — works technically, but it scans the entire table on every query. With 10,000 chunks it's slow. With a million it's unusable.

A vector database solves this with specialised index structures that group similar vectors together, so a similarity search only has to examine a fraction of your data rather than all of it.


Pinecone — a managed vector database

Pinecone is a purpose-built vector database, fully managed. You don't install anything. You don't maintain a server. You create an index, store vectors, and query them through an API. There's a free tier that covers everything you need for learning and small projects.

This is why it's a good first choice for RAG: it removes all the infrastructure overhead so you can focus on understanding how vector storage and retrieval actually work.

The core concepts in Pinecone:

Index — a named collection of vectors. You create one per use case. Each index is configured with a dimension count that must match your embedding model.

Namespace — a partition inside an index. Useful for separating vectors by user, by document type, or by environment (dev vs production).

Record — one stored vector, with an ID, the vector values, and optional metadata.


Storing and searching with Pinecone

After creating an account and a free index at pinecone.io, you get an API key and an index host URL.

In Spring AI, add the Pinecone dependency:

<dependency>
    <groupId>org.springframework.ai</groupId>
    <artifactId>spring-ai-pinecone-store-spring-boot-autoconfigure</artifactId>
</dependency>
Enter fullscreen mode Exit fullscreen mode

Configure in application.properties:

spring.ai.vectorstore.pinecone.api-key=${PINECONE_API_KEY}
spring.ai.vectorstore.pinecone.index-name=your-index-name
spring.ai.vectorstore.pinecone.namespace=documents
Enter fullscreen mode Exit fullscreen mode

That's the entire setup. Spring AI handles the rest — embedding, storing, and searching all go through the same VectorStore interface you already know:

@Autowired
private VectorStore vectorStore;

// Ingest — stores chunks + their vectors in Pinecone
vectorStore.add(chunks);

// Search — embeds the query and finds the closest chunks
List<Document> results = vectorStore.similaritySearch(
    SearchRequest.query("What is the return policy?").withTopK(5)
);
Enter fullscreen mode Exit fullscreen mode

similaritySearch embeds your query using the same embedding model configured for the app, sends the query vector to Pinecone, and returns the top K chunks ranked by similarity. The chunks come back as Document objects with content and metadata — ready to inject into a prompt.


Spring AI Advisors — automatic retrieval in every chat call

Manually calling similaritySearch and building the prompt yourself works, but Spring AI provides a higher-level abstraction that does it automatically: Advisors.

QuestionAnswerAdvisor intercepts every user message before it reaches the model, retrieves relevant chunks from your vector store, and injects them into the prompt as context:

@Bean
public ChatClient chatClient(ChatModel chatModel, VectorStore vectorStore) {
    return ChatClient.builder(chatModel)
        .defaultAdvisors(
            new MessageChatMemoryAdvisor(new InMemoryChatMemory()),
            new QuestionAnswerAdvisor(
                vectorStore,
                SearchRequest.defaults().withTopK(5)
            )
        )
        .build();
}
Enter fullscreen mode Exit fullscreen mode

With this wired up, every call to the chat client automatically:

  1. Takes the user's message and embeds it
  2. Searches the vector store for the top 5 similar chunks
  3. Injects those chunks into the system prompt before sending to the model

The injected context looks like this in the actual prompt:

Context information is below.
---------------------
[Chunk 1: "Electronics purchased at full price may be returned within 30 days..."]
[Chunk 2: "Exceptions: laptops and tablets cannot be returned after the seal is broken..."]
[Chunk 3: "To initiate a return, contact customer support with your order number..."]
---------------------
Given the context information and not prior knowledge, answer the question.
Enter fullscreen mode Exit fullscreen mode

The model reads the injected chunks and answers from them — not from its training data. That's RAG running end to end through a single advisor.


The dimension mismatch bug

When you create a Pinecone index, you set a dimension count. That count must match the number of dimensions your embedding model produces.

If you create an index with dimension 768 and embed with Google's text-embedding-004 (which outputs 768 dimensions), everything works. If you later switch to a different embedding model that outputs a different number of dimensions, Pinecone rejects the vectors with a dimension mismatch error — and that error is visible, so you catch it.

The silent version is harder: switching to a different model that happens to output the same number of dimensions. Pinecone accepts the vectors. No error. But vectors from two different models are not comparable — they live in completely different spaces, like two maps drawn at different scales and orientations. Similarity scores between them are meaningless. Retrieval returns confidently wrong results with no warning.

The rule is the same as the one from the embeddings article: same model at ingest, same model at query, always. If you change the embedding model, delete your index and re-embed all documents from scratch.

A practical habit: store the embedding model name in your chunk metadata so you can always tell which model produced which vectors.

Document doc = new Document(chunkText, Map.of(
    "source",          "return-policy-v2",
    "embeddingModel",  "text-embedding-004"
));
Enter fullscreen mode Exit fullscreen mode

What's next

You now have all the pieces: documents chunked, embedded, stored in Pinecone, and searchable via Spring AI. The next article puts it all together — a full working RAG pipeline in Spring Boot from the ingest endpoint to the retrieval-augmented chat response, with the complete architecture wired up and running.


Setting up Pinecone for the first time? Drop any issues in the comments — dimension mismatches and API key config are the two that catch people most often.

Sham Prakash K — Backend Engineer, 4+ years in Java, Spring Boot, and distributed systems. Building AI backend infrastructure. Writing about what I actually learned, mistakes included.

Top comments (0)