<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Inezayimana Henriette</title>
    <description>The latest articles on DEV Community by Inezayimana Henriette (@ineza250).</description>
    <link>https://dev.to/ineza250</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4151837%2Fe1b4a9bc-8d51-4dcf-8565-9fb5bf7653fc.png</url>
      <title>DEV Community: Inezayimana Henriette</title>
      <link>https://dev.to/ineza250</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ineza250"/>
    <language>en</language>
    <item>
      <title>words embeddings</title>
      <dc:creator>Inezayimana Henriette</dc:creator>
      <pubDate>Wed, 30 Sep 2026 14:57:00 +0000</pubDate>
      <link>https://dev.to/ineza250/words-embeddings-4pf4</link>
      <guid>https://dev.to/ineza250/words-embeddings-4pf4</guid>
      <description>&lt;p&gt;words embeddings are a way of representing a words as numbers(a vectors) so that the computer can understand the relationship and similalities  between words, &lt;br&gt;
the vectors usually 50 to 300 numbers long. These numbers are learned from large amounts of text, and they place words with similar meanings close to each other.&lt;br&gt;
why  words embedding is matters to help computer to understands human languege becouse computes hear only numbers. ex on map, cities that are near each other are similar in location. In an embedding space, words that are near each other are similar in meaning. "Doctor" ends up close to "nurse" and "hospital", and far from "bride".&lt;br&gt;
the simplest way is bag-of- word  is a way to converting text into a numbers so a computer can process it where each words gets its own position in a long list ,and we count  how is appears. example   "i love my school" on this  computer convert it like ,i:1, love:1,my:1 school:1 [1,1,1,1} this what computer receive , related method is tf-idf,which gives more weight to rare, informative words. both work, but they have a big weakness:they don&lt;code&gt;t understand meaning becouse they count words they don&lt;/code&gt;t care the meaning of words.&lt;br&gt;
in bag-of-words, "good" and "great" are as different as "good" and "banana". It treats every word as a separate item with no connection to any other.&lt;br&gt;
how  word2vec works  on high levels&lt;br&gt;
 first  Word2Vec was created at Google in 2013.it trains a small neural network  on task like predicting words and  find their neighbours&lt;br&gt;
 and also the words embedded  its capture meaning of the word not spelling  and also  a model trained on one tsk can reuse embedding that were learned from high amount of text even if your dataset is small it help you for that &lt;br&gt;
how it work&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;they look at the neghbours: it read  a lot of text and slides asmall window across eachStep 1: Look at the neighbours.
Word2Vec reads a lot of text and slides a small window across each sentence. Take the sentence:"I drink hot coffee every morning"For the word coffee with a window of 2, the neighbours are: drink, hot, every, morning. These pairs (coffee → drink, coffee → hot, and so on) become the training examples.
2: Give every word random numbers.
At the start, each word gets a vector of random numbers. These numbers mean nothing yet.
3: Practise a prediction game.
A small neural network tries to guess words from their neighbours. There are two versions of the game:
v1:continous bag of words  they try to predict the middle word  colles the target words ex "i love cuckies " the target will be the "love"  so you have to find love inorder to fill a sentences "i -cockies" they have  need love to fill the sentence  this it predit it
v2.skip-gram they trie to predict the sarrounding words  you get a middle input  ex :i love  this it find cockies depend on "i love cockies" CBOW is faster compare  skip-gram works better for rare words
4: Learn from mistakes.
Each time the network guesses wrong, it adjusts the word vectors slightly. After millions of sentences, words used in similar contexts end up with similar vectors. Coffee and tea both appear near drink, hot and cup, so their vectors move close together.
5: Keep the vectors, throw away the game.
We don't care about the predictions. The learned vectors are the useful result, and they are the word embeddings.
on training model i got code  from chart gbt
from gensim.models import Word2Vec&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;sentences = [&lt;br&gt;
    ["i", "drink", "hot", "coffee", "every", "morning"],&lt;br&gt;
    ["she", "drinks", "hot", "tea", "every", "morning"],&lt;br&gt;
    ["he", "drives", "a", "car", "to", "work"],&lt;br&gt;
]&lt;/p&gt;

&lt;p&gt;model = Word2Vec(&lt;br&gt;
    sentences,&lt;br&gt;
    vector_size=50,   # length of each word vector&lt;br&gt;
    window=2,         # how many neighbours to look at&lt;br&gt;
    min_count=1,      # keep all words&lt;br&gt;
    sg=1              # 1 = Skip-gram, 0 = CBOW&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;print(model.wv["coffee"])                    # the vector for "coffee"&lt;br&gt;
print(model.wv.most_similar("coffee"))       # closest words&lt;br&gt;
what i learned  and what suprise me&lt;br&gt;
 first what it`s suprise me how computer understand our  sentences or words in mathematics lic to convert it  with encording  as numbers and decording  as a word or answer   that make a sence i also learn  the method of embedding  called word2vec how it work depend on   what they look at the neighbour of the word lik i drink hot coffee and give   word a random number  and even how it train it&lt;br&gt;&lt;br&gt;
what supprise me  is how you can do a math with a words&lt;br&gt;
i run this analogy&lt;br&gt;
king -man+woman=Qween &lt;br&gt;
the model was never taught what a king or a queen is. It only read a lot of text, and it still found the right answer. I didn't expect a computer to pick up relationships like this on its own&lt;/p&gt;

</description>
      <category>webdev</category>
    </item>
  </channel>
</rss>
