DEV Community

anne sonia
anne sonia

Posted on

Word Embeddings: How Machines Learn the Meaning Behind Words

When I started learning Natural Language Processing (NLP), one of the things that confused me was a simple question: How can a computer understand the meaning of words?

Computers work with numbers, but humans communicate using words, sentences, and different ways of expressing the same idea. This became especially interesting to me when I started building a simple FAQ chatbot for a clothing shop.

My first approach used TF-IDF to represent questions. It worked by looking at the words in a question and comparing them with the words in stored FAQ questions. But I wanted my chatbot to understand more than just matching words.

That is where word embeddings became important.

What are word embeddings?

Word embeddings are numerical representations of words or pieces of text. Instead of representing a word only as a label such as "dress", an embedding represents it as a vector of numbers.

For example, we can imagine:

dress → [0.21, -0.14, 0.73, 0.42, ...]
shirt → [0.19, -0.10, 0.68, 0.39, ...]
car → [-0.51, 0.83, -0.22, 0.11, ...]

The actual vectors contain many more dimensions, and the numbers are learned from data.

The important idea is that words or sentences with related meanings can have vectors that are closer together in the embedding space.

This is useful because language is not only about exact words. Two people can ask the same question using completely different wording.

For example:

"What is the price of this dress?"

and:

"How much does this dress cost?"

They are different sentences, but they have almost the same meaning.

A system based only on exact word matching can have difficulty with this. Embeddings help us represent semantic relationships.

Word2Vec

One of the important embedding methods I learned about is Word2Vec.

Word2Vec was introduced by Mikolov and colleagues in 2013. It learns vector representations of words from their surrounding context. The basic idea is that words appearing in similar contexts tend to develop similar representations.

There are two well-known Word2Vec approaches:

CBOW (Continuous Bag of Words)
Skip-gram

CBOW uses surrounding words to predict a target word.

For example:

"The customer bought a new ___ yesterday."

The surrounding context can help the model predict a word such as:

dress

Skip-gram works in the opposite direction. It uses a target word to predict words that are likely to appear around it.

During training, the model learns vectors that capture relationships between words.

One famous type of example associated with word embeddings is:

king - man + woman ≈ queen

The interesting part is that the vectors can capture relationships rather than simply storing dictionary definitions.

Why do embeddings matter?

Embeddings are useful because many NLP problems require understanding relationships between words and text.

They can be used for:

Semantic search
Question answering
Chatbots
Text classification
Recommendation systems
Similarity detection
Information retrieval
Clustering

For my FAQ chatbot, the most important use is semantic similarity.

Imagine a clothing customer asks:

"Do you have something I can wear to a party?"

The FAQ database might contain:

"Do you sell party clothes?"

The two questions don't contain exactly the same words, but their meanings are related.

This is the kind of problem where embeddings can be useful.

From TF-IDF to embeddings in my chatbot

My first chatbot version used TF-IDF.

TF-IDF converts text into numerical vectors based mainly on the importance of words in the documents. I then used cosine similarity to find the FAQ question most similar to the user's question.

The improved version uses a pre-trained Sentence Transformer model:

from sentence_transformers import SentenceTransformer

model = SentenceTransformer(
"sentence-transformers/all-MiniLM-L6-v2"
)

questions = [
"Do you sell party clothes?",
"What sizes do you have?",
"How can I order clothes?"
]

embeddings = model.encode(questions)

print(embeddings.shape)

The Sentence Transformers documentation shows the same basic workflow: load a pretrained model, encode text into embeddings, and compare the resulting vectors for similarity. The all-MiniLM-L6-v2 model produces 384-dimensional embeddings.

For example, my chatbot can encode both the customer's question and the stored FAQ questions, then compare their vectors.

from sklearn.metrics.pairwise import cosine_similarity

user_question = model.encode(
["How much does a dress cost?"]
)

similarities = cosine_similarity(
user_question,
embeddings
)[0]

best_index = similarities.argmax()

print(questions[best_index])
print(similarities[best_index])

The chatbot can then return the FAQ answer associated with the question having the highest similarity.

What surprised me

The most interesting thing I learned is that similar meaning does not always require the same words.

Before learning about embeddings, I mainly thought of text as words that needed to be matched. Embeddings changed how I think about NLP.

For example, these questions are different:

"How much is a dress?"
"What is the price of a dress?"
"How expensive is that dress?"

A human immediately understands that they are asking about the same thing.

The goal of semantic embeddings is to represent that relationship numerically.

What I struggled with

One challenge was understanding what the numbers in an embedding actually mean.

At first, seeing a vector such as:

[0.034, -0.218, 0.731, ...]

does not tell us anything directly. The meaning is not stored in one specific number. Instead, useful information is distributed across the vector, and relationships can be studied by comparing vectors.

I also learned that choosing an embedding model matters. Different models can behave differently depending on the task and the data. A model that works well for general semantic similarity may not automatically be perfect for a specific FAQ dataset.

What I learned from the project

Building the clothing-shop FAQ chatbot helped me understand why embeddings are important in practical NLP systems.

TF-IDF is simple, fast, and useful when important keywords are enough to identify the correct answer.

Embeddings provide another way to represent text by capturing semantic relationships. This can be especially useful when users express the same idea using different words.

The next step in my project is to evaluate both approaches using the same test questions and compare their results.

For me, word embeddings are not just another machine learning technique. They are a way of moving from thinking about text as a collection of individual words toward thinking about relationships and meaning in language.

Sources
Mikolov et al., Efficient Estimation of Word Representations in Vector Space (2013).
Mikolov et al., Distributed Representations of Words and Phrases and their Compositionality (2013).
Sentence Transformers documentation, Quickstart and Computing Embeddings.
REFERENCES

nlp #machinelearning #python

Top comments (1)

Collapse
 
code250 profile image
Richard Munyemana •

I loved the idea of explaining word embeddings using your actual NLP project. However, you could use code snippets and headings so that we can be able to read your blog much more easily.

Overall, this was great and gives an idea of whether someone can start to learn word embeddings without going through the hard mathematics