DEV Community

Rose Umutesi
Rose Umutesi

Posted on

Understanding Word Embeddings: How Machines Learn the Meaning of Words

When I started learning Natural Language Processing (NLP), one question kept coming back to me: How can a computer understand words when words are just text?

A computer does not naturally understand that king and queen are related, or that cat and dog are more similar than cat and car. To a computer, these are simply sequences of characters.

This is where word embeddings become important.

What are word embeddings?

Word embeddings are a way of representing words as numerical vectors.

Instead of giving a machine the word:

"coffee"

we can represent it as something like:

[0.21, -0.43, 0.87, 0.12, ...]

The actual vectors usually have many more dimensions than this example. The important idea is that words with similar meanings or similar uses can have similar positions in a mathematical space.

This is based on a useful idea in linguistics: words that appear in similar contexts tend to have related meanings.

For example, consider:

I drank coffee in the morning.

and:

I drank tea in the morning.

Because coffee and tea appear in similar contexts, an embedding model can learn that they are related.

This is one reason embeddings are useful for NLP tasks such as text classification, recommendation systems, search, sentiment analysis, and information retrieval.

Word2Vec

One embedding method I studied is Word2Vec.

Word2Vec was introduced by researchers at Google and is designed to learn vector representations of words from text. Two important Word2Vec approaches are Continuous Bag of Words (CBOW) and Skip-gram.

The basic idea is surprisingly simple.

CBOW

CBOW tries to predict a word from the words surrounding it.

For example:

The farmer planted potatoes in the ___.

The surrounding words can help the model predict a word such as field.

Skip-gram

Skip-gram works in the opposite direction. It starts with a target word and tries to predict the words that appear around it.

For example:

The farmer planted potatoes in the field.

If the target word is potatoes, the model learns from words appearing around it, such as planted, farmer, and field.

After seeing many examples, the model adjusts its numerical representations so that words appearing in similar contexts develop related vectors.

A small Python example

We can experiment with Word2Vec using the gensim library.

from gensim.models import Word2Vec

sentences = [
    ["the", "farmer", "plants", "potatoes"],
    ["the", "farmer", "grows", "crops"],
    ["potatoes", "are", "important", "crops"],
    ["farmers", "grow", "maize", "and", "potatoes"],
    ["the", "farmer", "harvests", "crops"]
]

model = Word2Vec(
    sentences,
    vector_size=50,
    window=3,
    min_count=1,
    workers=4,
    seed=42
)

similar_words = model.wv.most_similar("potatoes", topn=3)

print(similar_words)
Enter fullscreen mode Exit fullscreen mode

The output from my experiment was:

[('plants', 0.3308681547641754),
 ('the', 0.20481298863887787),
 ('are', 0.10614515095949173)]
Enter fullscreen mode Exit fullscreen mode

The most_similar() method compares the vector for potatoes with the other word vectors and returns the words that are closest to it in the learned vector space.

In this experiment, plants had the highest similarity score among the three returned words. The scores indicate how close the word vectors are, with higher values representing greater similarity.

However, there is an important limitation. This example uses only five sentences, so the results should not be interpreted as meaningful knowledge about the English language. The model has very little data to learn from. A real Word2Vec model would normally be trained on a much larger and more representative corpus.

This experiment helped me understand an important point about word embeddings: the model is not learning dictionary definitions. Instead, it is learning patterns from how words appear in relation to other words.

What surprised me?

The most interesting thing I learned is that we can represent something as complicated as language using numbers.

At first, seeing a word represented by hundreds of numbers felt strange. I wondered how those numbers could possibly represent meaning.

The key realization was that one number does not represent one specific meaning. Instead, the entire vector represents a position in a learned mathematical space. Relationships between vectors can then capture useful patterns between words.

This also explains why word embeddings can perform interesting similarity tasks. If two words frequently occur in similar contexts, their vectors can end up relatively close together.

However, embeddings are not perfect. They can learn biases and patterns from the data used to train them. They also depend heavily on the quality and size of the training corpus.

What I struggled with

My biggest challenge was understanding the difference between text processing and meaning representation.

Techniques such as tokenization and TF-IDF are useful for converting text into numbers, but embeddings go a step further by learning relationships between words based on context.

I initially thought that an embedding was simply another way of encoding words. Now I understand that the important part is the relationship between the vectors, not just the individual numbers.

Final thoughts

Word embeddings helped me understand an important idea in NLP: computers do not need to understand language exactly like humans do in order to find useful patterns in text.

By representing words as vectors and learning from their contexts, models such as Word2Vec can capture relationships between words mathematically.

There are also other embedding approaches, including GloVe and FastText, each with different ways of learning word representations. Modern NLP has gone further with contextual representations from models such as BERT and transformer-based systems.

For me, learning word embeddings was a good reminder that machine learning often works by turning things that seem difficult to quantify—such as language—into representations that algorithms can process.

And that is what makes embeddings so interesting: words become numbers, but the relationships between those numbers can carry information about language.

References

Top comments (1)

Collapse
 
code250 profile image
Richard Munyemana •

This is a great blog you clear explained word embeddings and how they differ from TF-IDF and tokenization