DEV Community

abijuru diana
abijuru diana

Posted on

word embeddings in NLP

what is word embeddings?

word embedding is techniques of NLP that mapps the text into lists of vectors to represent word's meaning.

how does word embeddings works?

numerical conversion

vector space

semantic clustering

Why does this word embeddings matters?

word embedding matters because they were the step to make machine work with meaning instead of just spelling.

Ex: before word embeddings machine could take like "good" and "great" as unrelated data but after word embeddings model can treat them as related because of the "vector space"

they enable transfer learning.
--> word embeddings can be reuse in many small project,the model starts from the general langauge knowledge instead of starting from scracth. you dont need to relearn what langauge mean everytime you run a new project

they are the sourced applications

Word embedding Methods:

Normally in word embedddings we have the 3-be-called-approaches :

  1. count/frequency-based :This methods describe how often the words are in the sentence.
    ...This includes:

    • BoW (Bag-Of-Word)
    • TF-IDF(Term Frequency-Inverse Document Frequency)
    • Co-occurrence Matrix
  2. prediction-based :Methods use neural networks to learn word representations based on surrounding context.
    ...This includes:

    • Word2Vec
    • GloVe
    • FastText
  3. Contextual embeddings: This method generates dynamic vectors that change depending on the surrounding sentence, effectively handling polysemy
    ...This includes:
    -ELMo (Embeddings from Language Models)
    -Transformers models(BERT, GPT

FOCUS: Prediction-Based Methods:

  1. Word2Vec Here we train neural network to guess which words appear near each other.Word2Vec looks at the context of words in a large text dataset. Words that appear in similar settings get vector coordinates close to each other in a multi-dimensional space. It uses two main architectures:

• CBOW (Continuous Bag of Words): Predicts a target center word using its surrounding context words. It works faster for frequent words.
• Skip-gram: Predicts the surrounding context words by taking a single target center word as input. It performs better with smaller datasets and rare words.

example of codes:
`from gensim.models import Word2Vec

sentences = [
["Paul", "is", "driving", "the", "car"],
["mary", "is", "driving", "a", "biycle"],
["Cat", "drank", "the", "whole", "milk",]

]
model = Word2Vec(
sentences,
vector_size=100,
window=5,
min_count=1,
sg=1,

negative=5,
epochs=50,
)

print(model.wv["driving"])

print(model.wv.most_similar("driving"))

print(model.wv.similarity("biycle", "car"))
`

from gensim.models import Word2Vec
Enter fullscreen mode Exit fullscreen mode

1.-->Load the Word2Vec tool

sentences = [[...], [...], [...]]
Enter fullscreen mode Exit fullscreen mode

2.-->A list of sentences, each a list of words

Word2Vec(sentences, ...)
Enter fullscreen mode Exit fullscreen mode

3.-->Learn a vector for every word

model.wv[...], .most_similar(), .similarity()
Enter fullscreen mode Exit fullscreen mode

4.-->Look at the results

What i learned:

1.What is embeddings
2.How is word embedding giving high volume to today's tech
3.How is it applied through different ways
4.Word2Vec approach

Sources

-Chatgpt : some informations were prompted from chatgpt

  • Google

Top comments (0)