words embeddings are a way of representing a words as numbers(a vectors) so that the computer can understand the relationship and similalities between words,
the vectors usually 50 to 300 numbers long. These numbers are learned from large amounts of text, and they place words with similar meanings close to each other.
why words embedding is matters to help computer to understands human languege becouse computes hear only numbers. ex on map, cities that are near each other are similar in location. In an embedding space, words that are near each other are similar in meaning. "Doctor" ends up close to "nurse" and "hospital", and far from "bride".
the simplest way is bag-of- word is a way to converting text into a numbers so a computer can process it where each words gets its own position in a long list ,and we count how is appears. example "i love my school" on this computer convert it like ,i:1, love:1,my:1 school:1 [1,1,1,1} this what computer receive , related method is tf-idf,which gives more weight to rare, informative words. both work, but they have a big weakness:they dont understand meaning becouse they count words they dont care the meaning of words.
in bag-of-words, "good" and "great" are as different as "good" and "banana". It treats every word as a separate item with no connection to any other.
how word2vec works on high levels
first Word2Vec was created at Google in 2013.it trains a small neural network on task like predicting words and find their neighbours
and also the words embedded its capture meaning of the word not spelling and also a model trained on one tsk can reuse embedding that were learned from high amount of text even if your dataset is small it help you for that
how it work
- they look at the neghbours: it read a lot of text and slides asmall window across eachStep 1: Look at the neighbours. Word2Vec reads a lot of text and slides a small window across each sentence. Take the sentence:"I drink hot coffee every morning"For the word coffee with a window of 2, the neighbours are: drink, hot, every, morning. These pairs (coffee → drink, coffee → hot, and so on) become the training examples. 2: Give every word random numbers. At the start, each word gets a vector of random numbers. These numbers mean nothing yet. 3: Practise a prediction game. A small neural network tries to guess words from their neighbours. There are two versions of the game: v1:continous bag of words they try to predict the middle word colles the target words ex "i love cuckies " the target will be the "love" so you have to find love inorder to fill a sentences "i -cockies" they have need love to fill the sentence this it predit it v2.skip-gram they trie to predict the sarrounding words you get a middle input ex :i love this it find cockies depend on "i love cockies" CBOW is faster compare skip-gram works better for rare words 4: Learn from mistakes. Each time the network guesses wrong, it adjusts the word vectors slightly. After millions of sentences, words used in similar contexts end up with similar vectors. Coffee and tea both appear near drink, hot and cup, so their vectors move close together. 5: Keep the vectors, throw away the game. We don't care about the predictions. The learned vectors are the useful result, and they are the word embeddings. on training model i got code from chart gbt from gensim.models import Word2Vec
sentences = [
["i", "drink", "hot", "coffee", "every", "morning"],
["she", "drinks", "hot", "tea", "every", "morning"],
["he", "drives", "a", "car", "to", "work"],
]
model = Word2Vec(
sentences,
vector_size=50, # length of each word vector
window=2, # how many neighbours to look at
min_count=1, # keep all words
sg=1 # 1 = Skip-gram, 0 = CBOW
)
print(model.wv["coffee"]) # the vector for "coffee"
print(model.wv.most_similar("coffee")) # closest words
what i learned and what suprise me
first what it`s suprise me how computer understand our sentences or words in mathematics lic to convert it with encording as numbers and decording as a word or answer that make a sence i also learn the method of embedding called word2vec how it work depend on what they look at the neighbour of the word lik i drink hot coffee and give word a random number and even how it train it
what supprise me is how you can do a math with a words
i run this analogy
king -man+woman=Qween
the model was never taught what a king or a queen is. It only read a lot of text, and it still found the right answer. I didn't expect a computer to pick up relationships like this on its own
Top comments (0)