While learning Natural Language Processing (NLP), I came across a concept that helped me understand better how computers work with human language: word embeddings.
I read the Dive into Deep Learning material on natural language processing, especially the section about word embeddings and Word2Vec. Before reading it, I knew that text usually needs to be converted into numbers before a machine learning model can process it, but I did not fully understand how the representation of a word could contain information about its relationship with other words.
After going through the material, I understood word embeddings as a way of representing words using numerical vectors. Instead of treating a word as just a piece of text, we give it a position in a mathematical space. Words that are used in similar situations can have vectors that are closer to each other.
Why Are Word Embeddings Important?
One of the first things I learned is that there are different ways of converting words into numbers.
A simple method is one-hot encoding. For example, if our vocabulary contains four words:
cat
dog
car
house
we could represent them like this:
cat → [1, 0, 0, 0]
dog → [0, 1, 0, 0]
car → [0, 0, 1, 0]
house → [0, 0, 0, 1]
This tells the computer which word we are talking about, but it does not really tell the computer how the words are related.
For example, "cat" and "dog" are both animals, but their one-hot vectors do not show that relationship.
This is where word embeddings become useful. Instead of giving every word an isolated position, embeddings can represent words in a way that allows relationships and similarities to be captured.
A simplified example could look like this:
cat → [0.21, 0.73, 0.45]
dog → [0.24, 0.70, 0.48]
car → [0.91, 0.12, 0.08]
These numbers are only for illustration. Real word embeddings normally have many more dimensions.
What Is Word2Vec?
The embedding method I focused on was Word2Vec.
The main idea I took from the D2L material is that a word can be understood partly by looking at the words that appear around it. If two words regularly appear in similar contexts, their representations can become similar.
For example:
The child plays with a ball.
and
The boy plays with a ball.
The words "child" and "boy" occur in similar contexts. A model can learn from these patterns instead of someone manually telling it that the two words are related.
Word2Vec has two main approaches: Skip-gram and Continuous Bag of Words (CBOW).
With Skip-gram, we start with a word and use it to predict words that appear around it.
With CBOW, we do the opposite: we use surrounding words to predict the word in the middle.
A simple way I remember the difference is:
Skip-gram:
center word → surrounding words
CBOW:
surrounding words → center word
The model learns during training, and the resulting numerical representations become the word embeddings.
Seeing Similarity in Practice
One thing I found interesting is that once words have been represented as vectors, we can compare those vectors.
For example, with a pretrained Word2Vec model, we could use Python to find words that are similar to another word:
similar_words = model.most_similar("computer", topn=5)
for word, score in similar_words:
print(word, score)
The output might look something like:
software 0.78
technology 0.76
device 0.73
machine 0.71
internet 0.69
The actual results depend on the particular pretrained model and the data it learned from.
The important part for me is understanding what the similarity score represents. It gives us an indication of how close two words are in the learned vector space.
Word embeddings can also be used for word analogies. A well-known example is:
king - man + woman ≈ queen
What interested me here is that the model is not being given a rule saying "king relates to queen in this way." Instead, relationships can emerge from the patterns learned from large amounts of text.
What I Found Difficult
The most difficult part for me was understanding what each individual number in a word vector actually means.
At first, I thought that each dimension might have an obvious meaning. For example, I imagined that one number could represent whether a word was an animal, another could represent whether it was a person, and so on.
But that is not really how it works.
The information is distributed across the dimensions of the vector, which makes the representation useful but not always easy for humans to interpret.
This was one of the things I found most interesting because it showed me that a machine learning model can learn useful patterns even when those patterns are not represented in a way that is easy for us to explain dimension by dimension.
What I Took Away
My biggest takeaway from studying word embeddings is that the way we represent language matters.
A computer does not understand words in the same way humans do. We need to convert language into a form that a model can work with. Word embeddings provide a way to represent words as numbers while still preserving useful information about relationships between words.
Word2Vec made this idea clearer to me because it learns from context. The model looks at how words are used and learns representations from those patterns.
Reading the Dive into Deep Learning material also helped me see that Word2Vec is only one part of the bigger picture of NLP. There are other approaches, including GloVe and subword embeddings, and modern NLP systems have developed much further.
For me, the most important lesson is that NLP is not simply about giving text to a model. How we represent the text can have a major impact on what the model is able to learn.
Reference
Zhang, A., Lipton, Z. C., Li, M., & Smola, A. J. Dive into Deep Learning. Natural Language Processing: Pretraining — Word Embedding (word2vec). https://d2l.ai/chapter_natural-language-processing-pretraining/word2vec.html
Top comments (0)