DEV Community

Cover image for The Ultimate Guide to Word Embeddings & Word2Vec
Marie Claire NIYOMUGENGA
Marie Claire NIYOMUGENGA

Posted on

The Ultimate Guide to Word Embeddings & Word2Vec

In modern Natural Language Processing (NLP), translating human language into numerical formats that machine learning
models can process is foundational. Historically, text processing relied on One-Hot Encoding representing each word as
a sparse vector of size |V| (vocabulary size) with a single 1 and zeros elsewhere

One-Hot Encoding has two major drawbacks:

  1. High Dimensionality & Sparsity: A vocabulary of 100,000 words requires 100,000-dimensional vectors.
  2. Lack of Semantic Connection: The dot product between any two one-hot vectors is always 0, meaning the model
    cannot infer that 'cat' and 'feline' share similar contexts.
    Word Embeddings solve this by mapping words into dense, continuous vector spaces Rd (d ˛ [50, 300]). Words appearing
    in similar contexts cluster closely together in space

  3. Core Problem & Visual Intuition
    In dense vector spaces, semantic similarity corresponds to geometric proximity, evaluated using Cosine Similarity

Vector Arithmetic & Semantic Analogies: Word embeddings preserve linear relational structures. As visualised in Jay
Alammar's The Illustrated Word2vec and StatQuest, semantic operations correspond to consistent spatial directions:

Real-World Applications: 'Item2Vec'
As Jay Alammar notes, Word2vec principles extend far beyond text. Companies like Spotify, Airbnb, and Alibaba treat
user interaction sequences as 'sentences' and items as 'words' to build high-performance recommendation engines

  1. Under the Hood: Neural Architecture & Math
    Sliding Context Windows: Word2vec generates training data automatically from unlabelled text corpora using a sliding
    context window of size c. Accounting for bi-directional context (looking left and right) provides richer semantic signals.
    Word2vec introduces two primary architectures:

  2. Continuous Bag-of-Words (CBOW): Predicts the target center word given its context words.

  3. Skip-gram: Predicts surrounding context words given a single input center word.


Figure 2: Architecture comparison between CBOW and Skip-gram

Softmax Bottleneck & Negative Sampling (SGNS): Standard language models compute context probabilities using
Softmax over the full vocabulary |V|, creating an O(|V|) computational bottleneck.
Skip-gram with Negative Sampling (SGNS) converts the task into a fast binary logistic regression: determining whether
two words are real context neighbors (label 1) or random noise words (label 0)

Subsampling Frequent Words: High-frequency stopwords (e.g. 'the', 'is') carry minimal semantic signal. As detailed in D2L
and Lena Voita's course, Word2vec discards frequent words probabilistically:

Dual Weight Matrices: Word2vec maintains a Target Embedding Matrix W for input words and a Context Matrix W' for
context words. After training, W' is discarded and W is retained for downstream tasks

  1. Practical Python Implementation with Gensim Page 2 The following script demonstrates Skip-gram model training using Gensim, matching the pipeline described in D2L and Lena Voita

  1. Architecture Comparison & Recommended Resources

Recommended Learning Resources:
• Jay Alammar The Illustrated Word2vec (Visual intuitive guide to embeddings and Skip-gram).
• Lena Voita :NLP Course Word Embeddings (Detailed mathematical derivations of SGNS and distribution
hypothesis).
• StatQuest with Josh Starmer Word Embedding and Word2Vec, Clearly Explained!!! (Step-by-step neural network
mechanics).
• Aston Zhang, Zachary C. Lipton, Mu Li, & Alexander J. Smola Dive into Deep Learning (D2L.ai) Chapter 15:
Natural Language Processing: Pretraining.

Top comments (0)