as today's we learnt about word Embendding ,first we started with learning Natural Language Processing(NLP) we learn how computers turn words into numbers and the thing i have got today's leason is that word embending is used to see the simillarity of the words like kigali and nairobi have the simillarity of both being cities and If I see the words “king” and “queen,” I understand that they are related, now i can understand that computers dont see words the way we do , a computer utlimately works with numbers .
1. what ar5e word embenddings?
==> a word embendding is a way of representing a word as alist of numbers ,also called vector.
For example, a very simplified representation could look something like this: "got this example from chatgpt"
king → [0.25, 0.81, -0.14, 0.63]
queen → [0.27, 0.79, -0.12, 0.61]
apple → [-0.72, 0.10, 0.54, -0.31]
The real vectors are usually much larger than this example. They might contain dozens, hundreds, or more dimensions.
the important idea is that words with similar meanings or similar5 uses tend to have vectors that are close to each other in the emmbedding space .
this is useful because traditional approaches such as representing every word simply as an ID don't naturally tell a machine that two words are related.
Word embeddings instead allow machine learning models to work with relationships between words.
they are usefull in applications such as we saw today
- in sentiment analysis
- in text classification
- in search engines 4.in recommendation systems
- in chatbots
- in machine tran
he basic idea is often summarized by the phrase “you shall know a word by the company it keeps.” In other words, the words appearing around a word can tell us a lot about its meaning.
""2 Word2Vec""
One embedding method I studied is Word2Vec.
Word2Vec was introduced by Tomas Mikolov and his colleagues in 2013. The approach learns continuous vector representations of words from large amounts of text and was designed to capture useful syntactic and semantic relationships between words.
there are two things we looked about today which are:
- cbow(continuous bag of words)
- skip-gram
CBOW
CBOW tries to predict a target word from the words around it.
Imagine this sentence:
The cat is sleeping on the sofa.
If the target word is sleeping, the surrounding words provide context:
The cat is [sleeping] on the sofa
The model learns from many examples like this and gradually adjusts the word vectors.
Skip-gram
Skip-gram does almost the opposite.
Instead of using surrounding words to predict the target word, it uses a target word to predict words that appear around it.
For example:
The cat is sleeping on the sofa.
Given the word:
cat
the model tries to predict nearby words such as:
The
is
By repeating this process over a large corpus, the model learns useful representations of words.
The original Word2Vec work described both CBOW and Skip-gram architectures and showed that the learned vectors could capture syntactic and semantic relationships. Later work also introduced techniques such as negative sampling to make training more efficient.this is from google
Trying Word2Vec with Python
One of the easiest ways to experiment with Word2Vec in Python is with the gensim library.
Here is a small example:
from gensim.models import Word2Vec
sentences = [
["the", "cat", "is", "sleeping"],
["the", "dog", "is", "sleeping"],
["the", "cat", "likes", "milk"],
["the", "dog", "likes", "food"],
["a", "cat", "is", "an", "animal"],
["a", "dog", "is", "an", "animal"]
]
model = Word2Vec(
sentences,
vector_size=50,
window=2,
min_count=1,
workers=4
)
print(model.wv.most_similar("cat", topn=3))
** this are the similarity result**
The vector_size determines the number of dimensions used for each word vector.
The window controls how many nearby words are considered as context.
min_count=1 tells the model to keep words that appear at least once.
Finally, most_similar() lets us find words whose vectors are close to the vector for cat.
Because this is a tiny artificial dataset, the results won't be as meaningful as those from a model trained on a large real-world corpus. That is actually an important lesson: the quality of an embedding depends heavily on the data used to train it.
Gensim also provides functionality for retrieving vectors and finding similar words from trained Word2Vec models.
*the anology that we talked about to day *
the one the we saw today the things about word embedding is that they can sometime capture relationshios between words
example:
king - man + woman = queen😊
Conceptually, the model can learn directions in its vector space that correspond to certain relationships.
This doesn't mean the model understands gender or royalty in the same way a human does. Instead, those relationships have been encoded statistically through patterns in the training data.
That's one of the things I found most surprising about embeddings: simple numerical vectors can contain surprisingly rich information about language.
Something I struggled with
The hardest part for me was initially understanding what the individual numbers in an embedding actually mean.
If I see:
[0.21, -0.48, 0.73, 0.11, ...]
I cannot simply look at the first number and say, “This represents whether the word is an animal.”
The dimensions don't normally have such simple human-readable meanings. Meaning is distributed across the vector.
That was an important shift in how I thought about machine learning.
Instead of asking:
“What does each numberr mean?"
I started thinking:
“What relationships does the entire vector represent?”
Similarity is often measured using cosine similarity, which compares the direction of two vectors rather than simply comparing their individual values.
For example, if two words have vectors pointing in similar directions, their cosine similarity will be relatively high.
Word embeddings are powerful, but not perfect
There is another important lesson: embeddings learn from their training data.
If the training data contains biases, the resulting embeddings can also contain those biases.
They can also struggle with things such as:
-Words with multiple meanings
-Rare words
-Misspellings
-New words
-Context-dependent meanings
Another limitation of traditional Word2Vec-style embeddings is that a word generally has one learned vector, even when the word has different meanings depending on context.
For example, consider:
I deposited money at the bank.
and:
We sat beside the river bank.
The word bank has different meanings, but a traditional Word2Vec model gives bank a single representation.
This is one reason later NLP approaches moved toward contextual representations, where the representation of a word can change depending on the sentence.
What I learned
The biggest thing I learned from studying word embeddings is that NLP isn't simply about turning words into numbers.
It is about finding a numerical representation that preserves useful information about how language is used.
Word embeddings create a bridge between human language and mathematical models.
Once words become vectors, we can calculate similarities, compare concepts, perform classifications, and give other machine learning algorithms a numerical representation of language.
And honestly, seeing something like:
king - man + woman ≈ queen
come out of a mathematical representation of language is pretty wild.
It made me realize that machine learning doesn't necessarily need to “understand” language in the human sense to discover useful patterns in it.
That is what makes word embeddings such an interesting foundation for NLP.
Sources
- Google — Used to research and verify information about word embeddings, Word2Vec, NLP, and related concepts.
- ChatGPT (OpenAI) — Used as a learning and writing assistant to help explain and organize the concepts discussed in this article.
---------------------------end------------------------------
Top comments (0)