DEV Community

Said Olano
Said Olano

Posted on

Vector Search & Embeddings in Java: Building Semantic Search Engines

Vector Search & Embeddings in Java: Building Semantic Search Engines

Introduction

Traditional keyword-based search is dead. When users search for "best restaurants near me," they don't expect results matching those exact words—they expect restaurants that mean the right thing.

This is where vector search and embeddings enter the picture.

Vector search is the technology powering modern AI applications: ChatGPT's retrieval-augmented generation (RAG), Netflix's recommendation engine, Spotify's "Discover Weekly," and enterprise semantic search platforms. Yet many Java developers still think of search as Elasticsearch queries with exact terms.

This gap is costing you:

  • Missed user intent: Returning exact matches instead of semantic matches
  • Irrelevant results: Users abandon apps when search doesn't understand context
  • Lost competitive advantage: Your competitors' AI-powered features work better

In this guide, you'll learn:

  1. What embeddings and vector search actually are (the math isn't complex, the concepts are elegant)
  2. How to implement semantic search in Java using production libraries
  3. Integration patterns with PostgreSQL (pgvector), Elasticsearch, and Pinecone
  4. Real-world use cases from e-commerce to support chatbots
  5. Performance tuning for millions of vectors

By the end, you'll understand why vector search is essential for modern applications and how to build it with Java.


Part 1: Understanding Embeddings and Vector Space

What Is an Embedding?

An embedding is a numerical representation of text, images, or other data. Instead of storing "restaurant recommendations," you store a vector of numbers: [0.25, -0.15, 0.82, ..., 0.41].

These aren't random numbers. They're learned through neural networks trained on massive datasets. Similar concepts produce similar vectors. This property is the entire foundation of vector search.

Example:

  • Query: "best pizza place"
  • Vector: [0.12, 0.88, -0.31, 0.45, ...]
  • Document: "authentic Italian pizzeria"
  • Vector: [0.14, 0.87, -0.29, 0.46, ...]

These vectors are close in vector space. Measuring that distance (using cosine similarity, Euclidean distance, etc.) gives you a relevance score.

How Modern Embedding Models Work

In production, you don't hand-craft embeddings. You use pre-trained embedding models:

  • OpenAI's text-embedding-ada-002: 1,536 dimensions, state-of-the-art for general text
  • Sentence Transformers (open-source): 384-768 dimensions, fast inference, great for local deployments
  • Cohere's embedding API: 4,096 dimensions, excellent domain-specific options
  • Google's Gecko Embedding: Lightweight, optimized for cost

These models map text → vector in a way that preserves semantic meaning.

Vector Space Fundamentals

All vectors live in an N-dimensional space. When you have 1,536-dimensional vectors (from OpenAI), you're working in 1,536-dimensional space.

Similarity metrics:

  1. Cosine Similarity (most common)

    • Measures angle between vectors
    • Range: -1 to 1 (1 = identical direction)
    • Formula: (A · B) / (||A|| × ||B||)
  2. Euclidean Distance (L2)

    • Measures straight-line distance
    • Smaller = more similar
    • Formula: √(Σ(ai - bi)²)
  3. Manhattan Distance (L1)

    • Sum of absolute differences
    • Faster but less intuitive than Euclidean

For text search, cosine similarity is almost always the right choice.


Part 2: Building Semantic Search in Java

Architecture Overview

A semantic search system has these components:

┌─────────────────────────────────────────────┐
│  User Query                                  │
└─────────────┬───────────────────────────────┘
              │
┌─────────────▼───────────────────────────────┐
│  1. Embedding Model (Convert text → vector) │
│     (OpenAI API / Local Sentence Transformer)│
└─────────────┬───────────────────────────────┘
              │
┌─────────────▼───────────────────────────────┐
│  2. Vector Search Engine                     │
│     (Pinecone / PostgreSQL pgvector)        │
└─────────────┬───────────────────────────────┘
              │
┌─────────────▼───────────────────────────────┐
│  3. Similarity Ranking                       │
│     Return top-K most relevant results      │
└─────────────┬───────────────────────────────┘
              │
┌─────────────▼───────────────────────────────┐
│  Results with scores (0.0 - 1.0)            │
└─────────────────────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode

Example 1: Using OpenAI Embeddings + Pinecone

Step 1: Add Dependencies

<!-- pom.xml -->
<dependency>
    <groupId>com.theokanning.openai-gpt3-java</groupId>
    <artifactId>api</artifactId>
    <version>0.18.1</version>
</dependency>

<dependency>
    <groupId>io.pinecone</groupId>
    <artifactId>pinecone-client</artifactId>
    <version>0.1.0</version>
</dependency>
Enter fullscreen mode Exit fullscreen mode

Step 2: Create Embedding Service

import com.theokanning.openai.embedding.Embedding;
import com.theokanning.openai.embedding.EmbeddingRequest;
import com.theokanning.openai.embedding.EmbeddingResult;
import com.theokanning.openai.service.OpenAiService;

public class EmbeddingService {
    private final OpenAiService openAiService;
    private final String modelId = "text-embedding-ada-002";

    public EmbeddingService(String apiKey) {
        this.openAiService = new OpenAiService(apiKey);
    }

    public List<Double> embedText(String text) {
        EmbeddingRequest request = EmbeddingRequest.builder()
            .model(modelId)
            .input(Collections.singletonList(text))
            .build();

        EmbeddingResult result = openAiService.createEmbeddings(request);
        return result.getData().get(0).getEmbedding();
    }

    public List<List<Double>> embedTexts(List<String> texts) {
        EmbeddingRequest request = EmbeddingRequest.builder()
            .model(modelId)
            .input(texts)
            .build();

        EmbeddingResult result = openAiService.createEmbeddings(request);
        return result.getData().stream()
            .sorted(Comparator.comparingInt(Embedding::getIndex))
            .map(Embedding::getEmbedding)
            .collect(Collectors.toList());
    }
}
Enter fullscreen mode Exit fullscreen mode

Step 3: Pinecone Vector Search

import io.pinecone.clients.Pinecone;
import io.pinecone.clients.Index;

public class PineconeVectorStore {
    private final Index index;
    private final EmbeddingService embeddingService;

    public PineconeVectorStore(String apiKey, String projectName, 
                               String indexName, String embeddingApiKey) {
        Pinecone client = new Pinecone.Builder()
            .withApiKey(apiKey)
            .build();

        this.index = client.getIndex(projectName, indexName);
        this.embeddingService = new EmbeddingService(embeddingApiKey);
    }

    // Index documents with embeddings
    public void indexDocument(String docId, String content, 
                             Map<String, String> metadata) {
        List<Double> embedding = embeddingService.embedText(content);

        index.upsert(
            docId,
            embedding,
            metadata
        );
    }

    // Search: returns top K similar documents
    public List<SearchResult> search(String query, int topK) {
        List<Double> queryEmbedding = embeddingService.embedText(query);

        var results = index.query(queryEmbedding)
            .withTopK(topK)
            .withIncludeMetadata(true)
            .execute();

        return results.getMatches().stream()
            .map(match -> new SearchResult(
                match.getId(),
                match.getScore(),
                match.getMetadata()
            ))
            .collect(Collectors.toList());
    }
}

record SearchResult(String id, Double score, Map<String, String> metadata) {}
Enter fullscreen mode Exit fullscreen mode

Step 4: End-to-End Usage

public class SemanticSearchApp {
    public static void main(String[] args) {
        PineconeVectorStore store = new PineconeVectorStore(
            System.getenv("PINECONE_API_KEY"),
            "my-project",
            "restaurant-index",
            System.getenv("OPENAI_API_KEY")
        );

        // Index documents
        store.indexDocument("rest-1", "Best pizza in town, authentic Italian, family-owned", 
            Map.of("name", "Pasta Paradise", "type", "pizza"));

        store.indexDocument("rest-2", "Fast casual ramen, Tokyo-style noodles, great broth",
            Map.of("name", "Noodle House", "type", "ramen"));

        // Search with semantic understanding
        var results = store.search("where to find good Italian food", 3);

        results.forEach(r -> 
            System.out.printf("ID: %s, Score: %.4f, Name: %s%n",
                r.id(), r.score(), r.metadata().get("name"))
        );
    }
}
Enter fullscreen mode Exit fullscreen mode

Example 2: Local Vector Search with PostgreSQL + pgvector

For privacy-sensitive applications, you might want embeddings to stay within your infrastructure.

Step 1: Setup PostgreSQL pgvector

-- Enable pgvector extension
CREATE EXTENSION IF NOT EXISTS vector;

-- Create table
CREATE TABLE documents (
    id SERIAL PRIMARY KEY,
    content TEXT NOT NULL,
    embedding vector(1536),
    metadata JSONB,
    created_at TIMESTAMP DEFAULT NOW()
);

-- Create HNSW index for fast similarity search
CREATE INDEX ON documents USING hnsw (embedding vector_cosine_ops);
Enter fullscreen mode Exit fullscreen mode

Step 2: Java Implementation with Sentence Transformers

import org.springframework.ai.document.Document;
import org.springframework.ai.embedding.Embedding;
import org.springframework.ai.embedding.EmbeddingModel;
import org.springframework.jdbc.core.JdbcTemplate;

@Component
public class PostgresVectorStore {
    private final JdbcTemplate jdbcTemplate;
    private final EmbeddingModel embeddingModel;

    @Autowired
    public PostgresVectorStore(JdbcTemplate jdbcTemplate, 
                               EmbeddingModel embeddingModel) {
        this.jdbcTemplate = jdbcTemplate;
        this.embeddingModel = embeddingModel;
    }

    public void indexDocument(String content, String metadata) {
        Embedding embedding = embeddingModel.embed(content);

        String sql = "INSERT INTO documents (content, embedding, metadata) " +
                     "VALUES (?, ?::vector, ?::jsonb)";

        jdbcTemplate.update(sql, content, vectorToString(embedding), metadata);
    }

    public List<DocumentResult> similaritySearch(String query, int limit) {
        Embedding queryEmbedding = embeddingModel.embed(query);

        String sql = "SELECT id, content, metadata, " +
                     "1 - (embedding <=> ?::vector) as similarity " +
                     "FROM documents " +
                     "ORDER BY embedding <=> ?::vector " +
                     "LIMIT ?";

        return jdbcTemplate.query(sql, new Object[]{
            vectorToString(queryEmbedding),
            vectorToString(queryEmbedding),
            limit
        }, (rs, rowNum) -> new DocumentResult(
            rs.getInt("id"),
            rs.getString("content"),
            rs.getDouble("similarity"),
            rs.getString("metadata")
        ));
    }

    private String vectorToString(Embedding embedding) {
        return "[" + embedding.getOutput().stream()
            .map(String::valueOf)
            .collect(Collectors.joining(",")) + "]";
    }
}

record DocumentResult(int id, String content, double score, String metadata) {}
Enter fullscreen mode Exit fullscreen mode

Part 3: Real-World Use Cases

1. E-Commerce Product Search

Traditional: Search for "comfortable shoes" → matches only products with those exact words

Semantic: Matches "ergonomic footwear," "supportive sneakers," "cushioned athletic shoes"

public class ProductSearchService {
    private final VectorStore vectorStore;

    public List<Product> findSimilarProducts(String query) {
        return vectorStore.search(query, 10).stream()
            .map(this::toProduct)
            .collect(Collectors.toList());
    }
}
Enter fullscreen mode Exit fullscreen mode

2. Support Chatbot with RAG

Retrieval-Augmented Generation: Embed your knowledge base, find relevant docs, feed them to LLM

@RestController
public class SupportChatbot {
    private final VectorStore knowledgeBase;
    private final OpenAiService openAiService;

    @PostMapping("/ask")
    public ResponseEntity<String> ask(@RequestBody String question) {
        // Step 1: Find relevant docs using vector search
        var relevantDocs = knowledgeBase.search(question, 3);

        // Step 2: Build context from retrieved docs
        String context = relevantDocs.stream()
            .map(SearchResult::content)
            .collect(Collectors.joining("\n\n"));

        // Step 3: Ask LLM with context
        String prompt = String.format(
            "Based on this knowledge base:\n%s\n\nAnswer: %s",
            context, question
        );

        ChatCompletionRequest request = ChatCompletionRequest.builder()
            .model("gpt-4")
            .messages(List.of(new ChatMessage(ChatMessageRole.USER.value(), prompt)))
            .build();

        String answer = openAiService.createChatCompletion(request)
            .getChoices().get(0).getMessage().getContent();

        return ResponseEntity.ok(answer);
    }
}
Enter fullscreen mode Exit fullscreen mode

3. Duplicate Detection

Find near-duplicate documents or suspicious fraud patterns

public class DuplicateDetector {
    private final VectorStore vectorStore;

    public boolean isProbablyDuplicate(String document, double threshold) {
        var similarDocs = vectorStore.search(document, 1);

        return !similarDocs.isEmpty() && 
               similarDocs.get(0).score() > threshold;
    }
}
Enter fullscreen mode Exit fullscreen mode

Part 4: Performance Optimization

Batching Embeddings

Don't embed one document at a time. Batch them.

public void indexManyDocuments(List<Document> documents) {
    // ❌ Slow: N API calls
    // documents.forEach(doc -> index(doc));

    // ✅ Fast: 1 API call per batch
    Iterables.partition(documents, 100).forEach(batch -> {
        List<String> texts = batch.stream()
            .map(Document::getContent)
            .collect(Collectors.toList());

        List<List<Double>> embeddings = embeddingService.embedTexts(texts);

        for (int i = 0; i < batch.size(); i++) {
            vectorStore.index(batch.get(i).getId(), embeddings.get(i));
        }
    });
}
Enter fullscreen mode Exit fullscreen mode

Caching Embeddings

Store computed embeddings to avoid redundant API calls

@Component
public class CachedEmbeddingService {
    private final EmbeddingService service;
    private final Map<String, List<Double>> cache = new ConcurrentHashMap<>();

    public List<Double> embed(String text) {
        return cache.computeIfAbsent(text, key -> service.embedText(key));
    }
}
Enter fullscreen mode Exit fullscreen mode

Dimension Reduction

Not all 1,536 dimensions are necessary for your use case. PCA can reduce them:

public List<Double> reduceDimensions(List<Double> embedding, int targetDim) {
    // Use Apache Commons Math or similar
    // This trades accuracy for speed/storage
    return PCA.reduce(embedding, targetDim);
}
Enter fullscreen mode Exit fullscreen mode

Part 5: Production Considerations

1. Latency Budget

  • Embedding API call: 50-200ms
  • Vector search: 10-50ms (with proper indexing)
  • Total for user query: <500ms

If this is too slow, cache frequent queries.

2. Cost Management

  • OpenAI embeddings: ~$0.02 per million tokens
  • Pinecone: ~$0.70 per 1M vectors/month
  • Self-hosted (pgvector): Your infrastructure cost

For high-volume applications, self-hosting saves money.

3. Handling Updates

When document content changes, re-embed and update:

public void updateDocument(String docId, String newContent) {
    List<Double> newEmbedding = embeddingService.embedText(newContent);
    vectorStore.update(docId, newEmbedding, newContent);
}
Enter fullscreen mode Exit fullscreen mode

4. Monitoring

Track embedding quality:

@Component
public class EmbeddingQualityMonitor {
    private final MeterRegistry meterRegistry;

    public void recordSimilarityScore(double score) {
        Timer.builder("search.similarity.score")
            .publishPercentiles(0.5, 0.95, 0.99)
            .register(meterRegistry)
            .record(Duration.ofMillis((long) (score * 1000)));
    }
}
Enter fullscreen mode Exit fullscreen mode

Conclusion

Vector search and embeddings are no longer bleeding-edge. They're essential for:

  • Better search relevance (understand user intent, not just keywords)
  • Semantic recommendations (similarity-based, not just collaborative filtering)
  • RAG-powered chatbots (ground LLMs in your data)
  • Anomaly detection (spot unusual patterns in vector space)

Key Takeaways:

  1. Embeddings are numerical representations where similar concepts = close vectors
  2. Use cosine similarity to measure relevance
  3. Choose between managed (Pinecone, Weaviate) or self-hosted (PostgreSQL + pgvector)
  4. Batch API calls, cache embeddings, monitor quality
  5. Start with OpenAI embeddings; optimize to self-hosted later if cost-prohibitive

Your next semantic search implementation is just these components away. Build it today.


Further Reading

Top comments (0)