Building natural language processing models always brings you face-to-face with a core challenge: computers have no clue what words actually mean. Feed a machine a sentence like "the programmer writes code," and it sees nothing. Computers don't process text natively; they need numbers.
For a long time, traditional strategies like bag-of-words or TF-IDF handled this conversion. But when you look closely at how they represent text, you realize they hit a hard ceiling when it comes to capturing actual meaning. That's where word embeddings and the word2vec framework change the game.
Let's break down why traditional approaches fall short and how self-supervised vector representations solve the problem.
Why One-Hot Vectors Fall Short
A straightforward way to numericize text is using one-hot encoding. If your vocabulary contains a specific number of unique words, say a dictionary size of $\vert{}V\vert{}$, you assign each word an integer index from $0$ to $\vert{}V\vert{}-1$. To represent any specific word, you build a vector of length $\vert{}V\vert{}$ filled entirely with zeros, except for a single 1 placed at that word's index.
While these vectors are simple to set up, they are a terrible choice for machine learning models. The biggest issue is that they cannot capture semantic similarity. If you measure the cosine similarity between two different one-hot vectors, the result is always zero.
Because every orthogonal vector is completely independent, a one-hot encoder treats words like "cat" and "dog" as just as distant from each other as "cat" and "refrigerator." They completely fail to encode relationships or shared meanings.
Self-Supervised Word2Vec
To fix this, researchers introduced word2vec, mapping each word to a dense, fixed-length vector that captures semantic similarity and analogies.
The clever part about word2vec is that it relies on self-supervised learning. Instead of manual data labeling, the model extracts supervision straight from raw text by predicting words based on their surrounding context. It splits into two primary model architectures: Skip-Gram and Continuous Bag of Words (CBOW).
The Skip-Gram Model
The skip-gram architecture operates on the assumption that a central target word can generate its surrounding context words in a sequence.
Take the text sequence: "the", "man", "loves", "his", "son". If we select "loves" as our center word and set a context window size of 2, skip-gram looks at the conditional probability of generating the words that appear no more than two steps away: "the", "man", "his", and "son".
In this model, every word gets two separate vector representations: one used when it acts as a center word, and another used when it acts as a context word. During training, we maximize the likelihood function using stochastic gradient descent. Once trained, the center word vectors are typically pulled out and used as the final word embeddings for downstream tasks.
The Continuous Bag of Words (CBOW) Model
The continuous bag of words (CBOW) model flips the script. Instead of predicting context from a center word, CBOW assumes that a center word is generated based on its surrounding context words.
Using that same sequence ("the", "man", "loves", "his", "son"), with "loves" as the center word and a window size of 2, CBOW takes the surrounding context words ("the", "man", "his", "son") and averages their vectors to predict the target center word.
While training follows a similar gradient optimization process to skip-gram, CBOW typically uses the context word vectors as the final word representations.
Wrapping Up
Moving away from sparse frequency matrices to dense vector representations completely transforms how text data is handled in machine learning. By mapping words into a continuous vector space where geometric distance reflects semantic meaning, models can finally understand relationships, synonyms, and context.
References & Further Reading
- Dive into Deep Learning (D2L.ai): Chapter 15.1 – Word Embedding (word2vec). Available at d2l.ai.
- Mikolov, T., et al. (2013): Efficient Estimation of Word Representations in Vector Space. arXiv preprint arXiv:1301.3781.
- Mikolov, T., et al. (2013): Distributed Representations of Words and Phrases and their Compositionality. Advances in Neural Information Processing Systems (NeurIPS).
Top comments (1)
Great work in explaining why One-Hot Encoding fails and how Word Embeddings actually work. However, it is really hard to connect your paragraphs and capture your target.
Also, you should have told us what you learned from your research. What surprised you, and what did you struggle with while learning about word embeddings
You could also provide us with code snippets so that we can understand where you used word embeddings
Overall, you have done well figuring out that word2Vec uses self-supervised learning.