DEV Community

Cover image for How AI Understands Meaning — Embeddings Explained for Backend Engineers
Sham Prakash K
Sham Prakash K

Posted on AI-assisted

How AI Understands Meaning — Embeddings Explained for Backend Engineers

The previous article explained that RAG retrieves "semantically similar" chunks before generating an answer. But that phrase hides an entire question: how does a computer measure semantic similarity?

A computer doesn't understand meaning. It processes numbers. You can't tell a computer that "return policy" and "refund rules" are about the same thing — at least not directly. What you can do is convert both into numbers in such a way that similar meanings produce similar numbers. That conversion is what an embedding is.

What an embedding is

An embedding is a fixed-length list of numbers that represents a piece of text.

Take the sentence "What is the return policy for electronics?" — an embedding model converts that into something like:

[0.023, -0.417, 0.882, 0.001, -0.334, ... ]  ← 768 numbers total
Enter fullscreen mode Exit fullscreen mode

Every piece of text — every word, sentence, paragraph — maps to a list of numbers like this. The length of the list depends on the model (768, 1024, 1536 are common). This list is called a vector.

The numbers themselves mean nothing individually. What matters is their relationship to other vectors. Similar texts produce vectors that are numerically close to each other.


Why numbers can capture meaning

This sounds like it shouldn't work. How can a list of numbers "know" that "return policy" and "refund rules" are related?

The answer is in how embedding models are trained.

An embedding model is trained on a massive amount of text — books, articles, documentation, code. During training, it learns to predict which words and phrases appear near each other. The model learns that "dog" and "cat" appear in similar contexts: they're both followed by words like "food", "vet", "sleep", "play". They're both preceded by words like "my", "the", "a cute". Over billions of examples, the model builds an internal map where things that appear in similar contexts end up in similar positions.

That internal map is the vector space. When you ask the model to embed a word, it returns the coordinates of that word on the map.

"Dog" and "cat" end up close together on the map — not because anyone programmed that, but because the training data made them end up near each other. "Cricket" and "sports" end up close for the same reason. "Dog" and "democracy" end up far apart — almost no overlap in the contexts where they appear.


Vector space intuition

Imagine your embedding model produces 2-dimensional vectors. Every piece of text maps to a point on a flat plane.

                    Y
                    |
   "cricket" •  •  "football"
                    |
"sports"       •   |
               "basketball"
─────────────────────────────── X
                    |
"democracy"  •      |
   "parliament" •   |
                    |
Enter fullscreen mode Exit fullscreen mode

Words and phrases with similar meanings cluster together. Words with very different meanings are far apart.

Real embeddings have hundreds or thousands of dimensions instead of 2. You can't visualize them, but the math works the same way: similar meanings land in nearby positions in that high-dimensional space.


How similarity is measured — cosine similarity

Once you have two vectors, you measure how close they are using cosine similarity.

Cosine similarity measures the angle between two vectors. If two vectors point in almost the same direction, the angle between them is small and the similarity score is close to 1. If they point in opposite directions, the score is close to -1. Perpendicular vectors — totally unrelated — score around 0.

High similarity:   vector A and vector B point in the same direction → score ≈ 1.0
Low similarity:    vector A and vector C point in different directions → score ≈ 0.1
Enter fullscreen mode Exit fullscreen mode

Why angle instead of distance? Because distance is sensitive to the magnitude of the vector — a longer document would produce a vector further from the origin even if it's about the same topic. Cosine similarity ignores magnitude and only looks at direction, which makes it much more reliable for comparing meaning regardless of text length.

In practice, your vector database handles all of this for you. You ask "give me the 5 vectors closest to this query vector" and it returns them, ranked by cosine similarity. You never implement the math yourself.


Cosine similarity vs dot product

Vector databases give you a choice of similarity metric. The two most common are cosine similarity and dot product. They look similar — and under one condition they're identical — but they behave differently and the choice matters.

Dot product multiplies the two vectors element by element and sums the results:

dot(A, B) = A[0]×B[0] + A[1]×B[1] + ... + A[n]×B[n]
Enter fullscreen mode Exit fullscreen mode

The result depends on both the direction and the magnitude of the vectors. A longer vector (one with larger numbers overall) will produce a higher dot product score even if it isn't pointing in a more similar direction.

Cosine similarity normalises both vectors to unit length first, then computes the dot product:

cosine(A, B) = dot(A, B) / (magnitude(A) × magnitude(B))
Enter fullscreen mode Exit fullscreen mode

The result depends only on direction — magnitude is cancelled out.

When they're identical: If your embedding model produces unit-normalised vectors (vectors with magnitude = 1), dot product and cosine similarity give exactly the same score. Many modern embedding models, including Google's text-embedding-004, output unit-normalised vectors specifically so you can use the faster dot product operation without needing to normalise first.

When the difference matters:

Metric Sensitive to magnitude? Use when
Dot product Yes Vectors are unit-normalised (most modern models)
Cosine similarity No Vectors have varying magnitude; you want direction only

For RAG with a modern embedding API, dot product is the practical choice — it's faster and the model is already normalising for you. Cosine similarity is the safer default if you're unsure whether the model normalises its output.

In pgvector (the PostgreSQL vector extension we'll use in the next article), the operators are:

-- Cosine distance (1 - cosine similarity; lower = more similar)
ORDER BY embedding <=> query_vector

-- Dot product (negative; lower = more similar, because pgvector uses distance)
ORDER BY embedding <#> query_vector

-- Euclidean distance (less common for semantic search)
ORDER BY embedding <-> query_vector
Enter fullscreen mode Exit fullscreen mode

Spring AI's VectorStore handles the metric selection for you when you configure the store, so you rarely write this SQL directly — but knowing what's underneath matters when retrieval results look unexpected.


How this connects to RAG

When you embed your documents during the ingest phase, each chunk gets converted into a vector and stored in the vector database:

"Electronics purchased at full price may be returned within 30 days..."
     ↓ embedding model
[0.023, -0.417, 0.882, ...] → stored in vector DB
Enter fullscreen mode Exit fullscreen mode

When a user asks a question, the same process runs on the question:

"What is the return policy for electronics?"
     ↓ same embedding model
[0.021, -0.408, 0.879, ...]  ← numerically close to the chunk above
Enter fullscreen mode Exit fullscreen mode

The vector database compares the question vector against all stored chunk vectors and returns the ones with the highest cosine similarity. Those chunks are the semantically relevant ones — and they get injected into the prompt.

The reason semantic retrieval works is that the question vector and the relevant chunk vector end up close in vector space, even if they share no words. You never searched for "return" or "electronics" as keywords. You searched by meaning.


The same-model constraint

In previous article mentioned this briefly: the embedding model used during ingest must be the same one used during query.

Here's why that matters.

Every embedding model has its own internal map — its own way of laying out text in vector space. Google's text-embedding-004 and OpenAI's text-embedding-3-small both produce vectors, but those vectors live in completely different spaces. The numbers are incomparable across models.

If you embed your documents with model A and embed your query with model B, the vectors don't exist in the same space. There's no meaningful angle between them. Cosine similarity returns a number, but that number is noise — it has no relationship to actual semantic similarity.

The failure is silent. The vector database still runs the search, still returns 5 results, still looks like it's working. But the results will be wrong or random. This is one of the hardest bugs to catch because there's no error — just quietly bad retrieval.

Same model, both ways, always.


Embedding dimensions

When you call an embedding API, you get back a vector of a specific length — this is called the number of dimensions.

Google's text-embedding-004 produces 768-dimensional vectors by default (it supports up to 3072). OpenAI's text-embedding-3-large produces 3072-dimensional vectors. More dimensions generally means more nuance — the model can encode more information about meaning. But it also means larger storage and slower similarity search.

The dimensions are fixed per model and model version. If you switch from one model to another during a project, you have to re-embed all your documents from scratch because the vector length changes. Your old vectors become unusable — a 768-dimensional vector and a 1024-dimensional vector can't be compared.

This is another reason to choose your embedding model carefully at the start and not change it.


Calling the embedding API in Java

Using Spring AI, embedding a text string looks like this:

@Autowired
private EmbeddingModel embeddingModel;

public float[] embed(String text) {
    EmbeddingResponse response = embeddingModel.embedForResponse(List.of(text));
    return response.getResults().get(0).getOutput();
}
Enter fullscreen mode Exit fullscreen mode

Spring AI's EmbeddingModel interface abstracts the provider — whether you're using Google's text-embedding-004, OpenAI's embedding API, or another provider, the call looks the same. The dependency injection handles the HTTP call and API key.

To embed multiple chunks at once (more efficient during ingest):

public List<float[]> embedBatch(List<String> texts) {
    EmbeddingResponse response = embeddingModel.embedForResponse(texts);
    return response.getResults().stream()
        .map(EmbeddingResultData::getOutput)
        .toList();
}
Enter fullscreen mode Exit fullscreen mode

Batching is significantly faster and cheaper than embedding one chunk at a time — the API processes the batch in a single call instead of making one HTTP request per chunk.


A mistake worth knowing

Early in the project, I was testing the ingest pipeline with Google's text-embedding-004 and the retrieval looked fine in isolation. Then I swapped to OpenAI's API to test something unrelated, and in the process I accidentally left the embedding configuration pointing to OpenAI's embedding model while the stored vectors had been produced by Google's model.

Retrieval completely broke. Every question returned the same two chunks regardless of what the question was. But nothing threw an error. The search ran fine, just silently against incompatible vectors.

It took two hours to realise the issue was in the embedding configuration, not in the retrieval logic or the vector database setup. The fix was one line — pointing the query embedding back to text-embedding-004 — and then re-testing with a fresh batch of stored chunks to confirm the vectors were now in the same space.

The lesson: whenever retrieval looks wrong, check your embedding model configuration first, both ingest-side and query-side. It's the most common silent failure in RAG systems.


What's next

You now know what embeddings are and how semantic search works at the level that matters for building. The next piece is: how do you split your documents into chunks before embedding them?

The chunk size you choose has a bigger impact on retrieval quality than almost any other variable. Too large and you get vague results. Too small and you get fragments that miss context. There's a specific pattern that works, and a specific mistake I made that cost me two days of debugging — that's what the next article covers.


Built an embedding pipeline and hit unexpected results? Drop what you observed in the comments — it's almost always one of the issues above.

Sham Prakash K — Backend Engineer, 4+ years in Java, Spring Boot, and distributed systems. Building AI backend infrastructure. Writing about what I actually learned, mistakes included.

Top comments (0)