To make an AI work with your own company data, you first have to turn that data into something a machine can measure. Words are not measurable. Numbers are. The bridge between the two is the embedding.
In plain terms: Imagine a giant map where every sentence in your company has a pin. Sentences about the same topic are pinned close together, and unrelated ones are far apart. An embedding is the pin's coordinates.
Diagram: An embedding model maps a chunk of text to a list of numbers. Texts about similar things, like network latency and packet drop, get similar numbers. Real embeddings have hundreds or thousands of values, and the ones shown are illustrative. See the animated version.
1. The mathematical map
An embedding model reads a chunk of text and outputs a dense vector: a list of numbers, usually 768 or 1,536 of them. You never read these numbers yourself. What matters is how they relate to each other. The numbers are produced so that text with a similar meaning gets a similar list.
2. Semantic closeness
Because similar meanings get similar numbers, they cluster together when you plot them. The vectors for "network latency" and "packet drop" sit close together. The vector for "financial forecast" sits far away from both. A question you type also becomes a vector, and it lands next to the text that answers it.
Diagram: Embeddings place texts on a map where meaning is distance. Network terms cluster together, finance terms cluster elsewhere, and a question lands near the text that answers it. The map is a flattened, illustrative view of a space with hundreds of dimensions. See the animated version.
3. The power of distance
Once all your documents are vectors, finding the right one becomes arithmetic. The common measure is cosine similarity, which compares the direction of two vectors. A score near 1 means they point the same way, so the texts mean much the same thing. A score near 0 means they are unrelated.
Diagram: Cosine similarity measures the angle between two vectors. Vectors pointing the same way score close to 1, meaning similar meaning, and unrelated ones score near 0. The scores shown are computed from the four-number example vectors, which are illustrative. See the animated version.
This is why embedding search beats keyword search. Ask "Why is the site slow?" and a keyword search finds nothing if no document contains those words. An embedding search finds the paragraph about high latency on the web tier, because the meaning is close.
Diagram: Keyword search needs the same words. Embedding search compares meaning, so a question about a slow site can find a paragraph about high latency on the web tier even though the words differ. See the animated version.
Here is the arithmetic in Python, using the same four-number example vectors from the diagrams. Real embedding models return far longer vectors, and you would call one from a library or an API.
import numpy as np
def cosine(a, b):
a, b = np.array(a), np.array(b)
return float(a @ b / (np.linalg.norm(a) * np.linalg.norm(b)))
latency = [0.82, -0.11, 0.47, 0.05] # "network latency"
drop = [0.79, -0.08, 0.51, 0.02] # "packet drop"
forecast = [-0.12, 0.55, 0.20, 0.78] # "financial forecast"
print(round(cosine(latency, drop), 2)) # 1.0 -> almost the same meaning
print(round(cosine(latency, forecast), 2)) # -0.03 -> unrelated
The numbers in this example are made up to show the idea. In a real system you embed every document once, store the vectors, and compare each new question against them.
Coming Up Next
Day 12: Vector databases: storing and searching enterprise knowledge.
VectorEmbeddings #MachineLearning #NLP #DataEngineering #AIArchitecture
Originally published at https://sureshpallapothu.in/blog/day-11-embeddings, where this post includes animated diagrams.
Top comments (0)