<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Marie Claire NIYOMUGENGA</title>
    <description>The latest articles on DEV Community by Marie Claire NIYOMUGENGA (@claire_123).</description>
    <link>https://dev.to/claire_123</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4151834%2Fdc831286-bf6b-43e9-b6b5-734b5a0e3c3c.jpg</url>
      <title>DEV Community: Marie Claire NIYOMUGENGA</title>
      <link>https://dev.to/claire_123</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/claire_123"/>
    <language>en</language>
    <item>
      <title>The Ultimate Guide to Word Embeddings &amp; Word2Vec</title>
      <dc:creator>Marie Claire NIYOMUGENGA</dc:creator>
      <pubDate>Wed, 30 Sep 2026 14:32:14 +0000</pubDate>
      <link>https://dev.to/claire_123/the-ultimate-guide-to-word-embeddings-word2vec-1lpp</link>
      <guid>https://dev.to/claire_123/the-ultimate-guide-to-word-embeddings-word2vec-1lpp</guid>
      <description>&lt;p&gt;In modern Natural Language Processing (NLP), translating human language into numerical formats that machine learning&lt;br&gt;
models can process is foundational. Historically, text processing relied on One-Hot Encoding representing each word as&lt;br&gt;
a sparse vector of size |V| (vocabulary size) with a single 1 and zeros elsewhere&lt;/p&gt;

&lt;p&gt;One-Hot Encoding has two major drawbacks:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;High Dimensionality &amp;amp; Sparsity: A vocabulary of 100,000 words requires 100,000-dimensional vectors.&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Lack of Semantic Connection: The dot product between any two one-hot vectors is always 0, meaning the model&lt;br&gt;
cannot infer that 'cat' and 'feline' share similar contexts.&lt;br&gt;
Word Embeddings solve this by mapping words into dense, continuous vector spaces Rd (d ˛ [50, 300]). Words appearing&lt;br&gt;
in similar contexts cluster closely together in space&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Core Problem &amp;amp; Visual Intuition&lt;br&gt;
In dense vector spaces, semantic similarity corresponds to geometric proximity, evaluated using Cosine Similarity&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fde0j12d5dcaao1glity3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fde0j12d5dcaao1glity3.png" alt=" " width="745" height="138"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Vector Arithmetic &amp;amp; Semantic Analogies: Word embeddings preserve linear relational structures. As visualised in Jay&lt;br&gt;
Alammar's The Illustrated Word2vec and StatQuest, semantic operations correspond to consistent spatial directions:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu9a0erafs4ztmyk4o24f.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu9a0erafs4ztmyk4o24f.png" alt=" " width="721" height="260"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Real-World Applications: 'Item2Vec'&lt;br&gt;
As Jay Alammar notes, Word2vec principles extend far beyond text. Companies like Spotify, Airbnb, and Alibaba treat&lt;br&gt;
user interaction sequences as 'sentences' and items as 'words' to build high-performance recommendation engines&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Under the Hood: Neural Architecture &amp;amp; Math&lt;br&gt;
Sliding Context Windows: Word2vec generates training data automatically from unlabelled text corpora using a sliding&lt;br&gt;
context window of size c. Accounting for bi-directional context (looking left and right) provides richer semantic signals.&lt;br&gt;
Word2vec introduces two primary architectures:&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Continuous Bag-of-Words (CBOW): Predicts the target center word given its context words.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Skip-gram: Predicts surrounding context words given a single input center word.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6acacu9j3rla1hsafck3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6acacu9j3rla1hsafck3.png" alt=" " width="710" height="231"&gt;&lt;/a&gt;&lt;br&gt;
Figure 2: Architecture comparison between CBOW and Skip-gram&lt;/p&gt;

&lt;p&gt;Softmax Bottleneck &amp;amp; Negative Sampling (SGNS): Standard language models compute context probabilities using&lt;br&gt;
Softmax over the full vocabulary |V|, creating an O(|V|) computational bottleneck.&lt;br&gt;
Skip-gram with Negative Sampling (SGNS) converts the task into a fast binary logistic regression: determining whether&lt;br&gt;
two words are real context neighbors (label 1) or random noise words (label 0)&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn92qela6lxrhms3ysiok.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn92qela6lxrhms3ysiok.png" alt=" " width="721" height="145"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Subsampling Frequent Words: High-frequency stopwords (e.g. 'the', 'is') carry minimal semantic signal. As detailed in D2L&lt;br&gt;
and Lena Voita's course, Word2vec discards frequent words probabilistically:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3cp6awzm6pjtq9m1tqmm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3cp6awzm6pjtq9m1tqmm.png" alt=" " width="673" height="123"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Dual Weight Matrices: Word2vec maintains a Target Embedding Matrix W for input words and a Context Matrix W' for&lt;br&gt;
context words. After training, W' is discarded and W is retained for downstream tasks&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Practical Python Implementation with Gensim
Page 2
The following script demonstrates Skip-gram model training using Gensim, matching the pipeline described in D2L and
Lena Voita&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh8xs8ppldrb8qgc795cf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh8xs8ppldrb8qgc795cf.png" alt=" " width="668" height="435"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Architecture Comparison &amp;amp; Recommended Resources&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feemtkmm9jzzrb7ljmz3k.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feemtkmm9jzzrb7ljmz3k.png" alt=" " width="688" height="190"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Recommended Learning Resources:&lt;br&gt;
• Jay Alammar  The Illustrated Word2vec (Visual intuitive guide to embeddings and Skip-gram).&lt;br&gt;
• Lena Voita :NLP Course  Word Embeddings (Detailed mathematical derivations of SGNS and distribution&lt;br&gt;
hypothesis).&lt;br&gt;
• StatQuest with Josh Starmer Word Embedding and Word2Vec, Clearly Explained!!! (Step-by-step neural network&lt;br&gt;
mechanics).&lt;br&gt;
• Aston Zhang, Zachary C. Lipton, Mu Li, &amp;amp; Alexander J. Smola  Dive into Deep Learning (D2L.ai) Chapter 15:&lt;br&gt;
Natural Language Processing: Pretraining.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>nlp</category>
      <category>python</category>
    </item>
  </channel>
</rss>
