Before I studied word embeddings, I thought of text as something you count. Take a document, tally the words, and turn the tallies into numbers a model can use. That works up to a point, but I learned it has a real weakness: counting tells a model which words appear, not how the words relate to each other.
Word embeddings try to fix that. In this post I explain what they are, why they matter, how Word2Vec works at a high level, and what I found difficult while learning them.
What is a word embedding?
A word embedding represents a word as a list of numbers called a vector. The word "king" might look like this:[0.21, -0.34, 0.57, 0.12, ..]
Real vectors are longer, usually 100 to 300 values. What matters is that words used in similar ways end up with vectors that sit close together. Models can only work with numbers, so this gives us a way to feed language into them without throwing away the relationships between words.
How Word2Vec works
Word2Vec, introduced by Mikolov and colleagues in 2013, learns from context. The idea is that the neighbours of a word tell you a lot about it. It comes in two versions:
CBOW (Continuous Bag of Words) predicts a target word from the words around it.
Skipgram does the reverse and uses a target word to predict the words around it.
Take the sentence "The cat is sleeping on the trees" so if "cat" is the target and the window size is two, the model looks at "the", "is" and "sleeping" as its context. The model is trained on this prediction task over millions of sentences. The vectors are a by product of that training: the weights the model adjusts to make better predictions become the word vectors. Words that appear in similar contexts end up with similar weights, and so similar vectors.
he famous analogy
The best known example of vector arithmetic is:
text:
king - man + woman is approximately queen
python:
model.most_similar(positive=["king", "woman"], negative=["man"], topn=3)
I found this fascinating because it shows the vectors hold structure and not just random numbers. But it has limits. The result depends on the model, and the search leaves out the input words from the answer. Without that, the nearest word is often "king" itself. So I treat this as a neat demonstration and not as proof that the model understands language.
What I learned
The main lesson is that turning text into numbers does not have to mean counting. How you represent language changes what a model can learn. I also came to see how much context matters, since the words around a word tell you a lot about how it is used.
One more point that surprised me is that similar vectors do not always mean similar meaning. They mean similar usage. Words like "hot" and "cold" are opposites, but they appear in the same kinds of sentences, so their vectors are often close.
What I found hard
The first thing that confused me was the individual numbers. It is tempting to look at a vector and ask what the 0.57 means. On its own it means nothing. The information is spread across the whole vector, not stored in one value.
The second was data quality. Embeddings are only as good as the text they were trained on. Small or unrepresentative data gives weak embeddings, and biased data gives biased embeddings. Models trained on news text, for example, have been shown to link some jobs with one gender more than another.
The third is that Word2Vec gives every word exactly one vector. The word "bank" gets the same vector whether it means a river bank or a financial bank. Newer contextual models such as BERT solve this by producing a different vector depending on the sentence.
What is next
I want to try GloVe and FastText and compare all three on the same dataset, so I can see the differences for myself instead of only reading about them.
References
.Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient Estimation of Word Representations in Vector Space.
.Mikolov, T., Sutskever, I., Chen, K., Corrado, G., & Dean, J. (2013). .Distributed Representations of Words and Phrases and their Compositionality.
.Pennington, J., Socher, R., & Manning, C. (2014). GloVe: Global Vectors for Word Representation.
.Bojanowski, P., et al. (2017). Enriching Word Vectors with Subword Information.
.gensim documentation and course materials on Natural Language Processing.
Top comments (0)