In modern Natural Language Processing (NLP), translating human language into numerical formats that machine learning
models can process is foundational. Historically, text processing relied on One-Hot Encoding representing each word as
a sparse vector of size |V| (vocabulary size) with a single 1 and zeros elsewhere
One-Hot Encoding has two major drawbacks:
- High Dimensionality & Sparsity: A vocabulary of 100,000 words requires 100,000-dimensional vectors.
Lack of Semantic Connection: The dot product between any two one-hot vectors is always 0, meaning the model
cannot infer that 'cat' and 'feline' share similar contexts.
Word Embeddings solve this by mapping words into dense, continuous vector spaces Rd (d ˛ [50, 300]). Words appearing
in similar contexts cluster closely together in spaceCore Problem & Visual Intuition
In dense vector spaces, semantic similarity corresponds to geometric proximity, evaluated using Cosine Similarity
Vector Arithmetic & Semantic Analogies: Word embeddings preserve linear relational structures. As visualised in Jay
Alammar's The Illustrated Word2vec and StatQuest, semantic operations correspond to consistent spatial directions:
Real-World Applications: 'Item2Vec'
As Jay Alammar notes, Word2vec principles extend far beyond text. Companies like Spotify, Airbnb, and Alibaba treat
user interaction sequences as 'sentences' and items as 'words' to build high-performance recommendation engines
Under the Hood: Neural Architecture & Math
Sliding Context Windows: Word2vec generates training data automatically from unlabelled text corpora using a sliding
context window of size c. Accounting for bi-directional context (looking left and right) provides richer semantic signals.
Word2vec introduces two primary architectures:Continuous Bag-of-Words (CBOW): Predicts the target center word given its context words.
Skip-gram: Predicts surrounding context words given a single input center word.

Figure 2: Architecture comparison between CBOW and Skip-gram
Softmax Bottleneck & Negative Sampling (SGNS): Standard language models compute context probabilities using
Softmax over the full vocabulary |V|, creating an O(|V|) computational bottleneck.
Skip-gram with Negative Sampling (SGNS) converts the task into a fast binary logistic regression: determining whether
two words are real context neighbors (label 1) or random noise words (label 0)
Subsampling Frequent Words: High-frequency stopwords (e.g. 'the', 'is') carry minimal semantic signal. As detailed in D2L
and Lena Voita's course, Word2vec discards frequent words probabilistically:
Dual Weight Matrices: Word2vec maintains a Target Embedding Matrix W for input words and a Context Matrix W' for
context words. After training, W' is discarded and W is retained for downstream tasks
- Practical Python Implementation with Gensim Page 2 The following script demonstrates Skip-gram model training using Gensim, matching the pipeline described in D2L and Lena Voita
- Architecture Comparison & Recommended Resources
Recommended Learning Resources:
• Jay Alammar The Illustrated Word2vec (Visual intuitive guide to embeddings and Skip-gram).
• Lena Voita :NLP Course Word Embeddings (Detailed mathematical derivations of SGNS and distribution
hypothesis).
• StatQuest with Josh Starmer Word Embedding and Word2Vec, Clearly Explained!!! (Step-by-step neural network
mechanics).
• Aston Zhang, Zachary C. Lipton, Mu Li, & Alexander J. Smola Dive into Deep Learning (D2L.ai) Chapter 15:
Natural Language Processing: Pretraining.






Top comments (0)