DEV Community

immaculee byukusenge
immaculee byukusenge

Posted on

How I Started Understanding Word Embeddings

Introduction

As I continue learning Artificial Intelligence and Natural Language Processing (NLP), I have started learning about word embeddings.

Before learning about embeddings, I understood text mainly as words and sentences. However, computers cannot directly understand the meaning of words in the same way humans do. NLP needs a way to represent words as numbers so that machine-learning models can process them.

This is where word embeddings become useful.

What are Word Embeddings?

Word embeddings are numerical representations of words in the form of vectors.

Instead of representing a word simply as a label such as:
** "king"**

an embedding represents it using a vector of numbers:

king → [0.21, 0.45, -0.13, 0.72, ...]
These numbers allow words to be represented in a mathematical space.

The interesting part is that words with similar meanings or that appear in similar contexts can have similar vector representations.

For example, words such as:

king
queen
prince
princess
may have relationships in the embedding space.

This is important because NLP systems need more than just knowing that two words are different. They need representations that can capture relationships between words.
Why are Word Embeddings Important?

Before learning about embeddings, I came across traditional approaches such as Bag-of-Words and TF-IDF.

These approaches represent text using numerical values, but they do not represent the meaning of words in the same way that embeddings can.

For example, consider these two sentences:

The cat is sleeping.
The kitten is sleeping.

A traditional representation may treat cat and kitten as completely different words.

Word embeddings can learn that these words are related because they often appear in similar contexts.

This allows NLP models to work with relationships between words rather than treating every word as completely independent.

Word2Vec

One word-embedding method I learned about is Word2Vec.

Word2Vec learns word representations by looking at the contexts in which words appear.

There are two common approaches:

  1. CBOW (Continuous Bag of Words)

  2. Skip-gram

CBOW tries to predict a target word from surrounding words.

For example:

The cat is ___ on the sofa.

The surrounding words can help the model predict the missing word.

Skip-gram works in the opposite direction. It uses a target word to predict surrounding words.
A Small Word2Vec Experiment

I wanted to see what happens when a computer learns words from their context, so I created a very small dataset:

sentences = [
["cat", "eats", "fish"],
["cat", "likes", "milk"],
["dog", "eats", "meat"],
["dog", "likes", "food"]
]

Then I trained a Word2Vec model:

from gensim.models import Word2Vec

model = Word2Vec(
sentences,
vector_size=10,
window=2,
min_count=1,
workers=1
)

After training, I asked the model for the vector representation of cat:

print(model.wv["cat"])

And I asked it to find words similar to cat:

print(model.wv.most_similar("cat"))

The exact results depend on the data used to train the model. With a very small dataset like this one, the results should not be treated as meaningful real-world language knowledge. A larger dataset would allow the model to learn better representations.

Something I Found Challenging

One thing I confused when I learned about word embeddings was the idea of representing a word using numbers.

At first, I wondered:

How can numbers represent the meaning of a word?

I learned that the individual numbers themselves do not have a simple human-readable meaning. Instead, the position of a word's vector in the learned space and its relationship to other vectors are what make the representation useful.

This helped me understand that machine learning does not need to understand words exactly as humans do. Instead, it can learn numerical patterns from data that represent useful relationships between words.

What I Learned

The main thing I learned is that word embeddings provide a way of representing words as dense numerical vectors.

I also learned that the context in which a word appears is important. Words used in similar contexts can have similar representations.

Word2Vec showed me how a model can learn these representations from text instead of manually assigning meanings to every word.

This has made me more interested in Natural Language Processing because it shows how mathematical representations can be used to work with human language.

Conclusion

Word embeddings are an important concept in NLP because they transform words into numerical vectors that machine-learning models can process.

Word2Vec is one approach for learning these representations from the context of words.

As a beginner in Artificial Intelligence, I am still learning how embeddings work and how they can be applied to real NLP projects. However, understanding the basic idea of representing words as vectors is an important step in my AI learning journey.

REFERENCES

[1] T. Mikolov, K. Chen, G. Corrado, and J. Dean,
“Efficient Estimation of Word Representations in Vector Space,”
International Conference on Learning Representations (ICLR), 2013.
[Online]. Available: https://arxiv.org/abs/1301.3781

[2] J. Pennington, R. Socher, and C. D. Manning,
“GloVe: Global Vectors for Word Representation,”
in Proceedings of the 2014 Conference on Empirical Methods in
Natural Language Processing (EMNLP), 2014, pp. 1532–1543.
[Online]. Available: https://nlp.stanford.edu/projects/glove/

Top comments (0)