For humans, a word like "paris" has meaning.We know that paris is a city and we might associate it with words like "tokyo","kigali","berlin".But a computer doesn't automatically understand those relationships.
So that's what i found hard to hard to understand: "How a computer could actually work with words".
So, what are word embeddings?
Word embeddings are a way of presenting words as numbers.
Instead of giving a machine a word like:
paris
we represent it as a vector:
Paris → [0.18, -0.42, 0.73, 0.11, -0.56, ...]
What interested me most was that these numbers aren't random in the final result. Words that are used in similar contexts can end up with similar representations.
For example, we might expect words like: nairobi,kigali,bujumubura to have representations that are closer to each other than completely unrelated words.
This made me understand that a computer doesn't need to understand a word exactly the way a human does in order to learn useful relationships between words. It can learn those relationships from patterns in text.
Before learning about embeddings, I thought turning words into numbers would simply mean assigning a number to each word but it doesn't have a meaningful relationship between those words.Embedding represent words as vectors where relationships between words can be represented mathematically
Things I learned about Word2Vec
One of the embedding methods i studied was Word2Vec. It has a basic idea behind it:Look at the words surrounding a word and learn from the context.
For example, with sentences like:
I drink coffee every morning.
I drink fanta every morning.
The words "coffee" and "fanta" appear in a very similar context and by looking at many examples like this, Word2Vec can learn that these words are related.
Word2Vec includes approaches such as Skip-gram, which learns word representations by using a target word to predict surrounding context words. The original Word2Vec work showed that these learned vectors could capture semantic and syntactic relationships between words.
So instead of someone manually telling the computer:
coffee is related to tea, the model can discover relationships from the text it is trained on.
Knowing about this was interesting.
A small Python example
I also wanted to see what this idea looks like in Python. Using gensim, I can train a very small Word2Vec model on a few example sentences:
from gensim.models import Word2Vec
sentences = [
["i", "drink", "coffee", "every", "morning"],
["i", "drink", "tea", "every", "morning"],
["i", "drink", "fanta", "every", "morning"]
]
model = Word2Vec(
sentences,
vector_size=50,
window=2,
min_count=1,
sg=1
)
print(model.wv.most_similar("coffee"))
The important part for me was most_similar("coffee"). Instead of simply asking the computer to define coffee, I can ask the model which words have vectors that are close to the vector it learned for "coffee". With a real, much larger training dataset, this type of similarity can reveal useful relationships between words.
This small example helped me connect the theory to something I could actually run in Python.
A simple example of what embedding can capture
One famous example of word-vector relationships is :
King-man+woman = queen
The idea is that relationships between words can sometimes be represented through vector arithmetic. It is important to understand that this doesn't mean the computer "understands" the words exactly like a person does. Rather, patterns in the training data can resukt in meaningful mathematical relationships between the vectors.
The biggest surprise for me was realizing that meaning can emerge from patterns.
At first, vectors just looked like a bunch of numbers:
[0.21, -0.43, 0.67, 0.12, ...]
It was difficult to see how those numbers could represent anything related to language.
But after learning that similar words can have similar vector representations, and that vector differences can sometimes capture relationships between words, the idea became much clearer.
I also realized that the quality of the embeddings depends heavily on the text used to train them. If the training data contains certain biases or doesn't represent particular ways of using language, those patterns can end up being reflected in the learned representations.
What I struggled with
One thing I struggled with was understanding how a model actually learns relationships between words. At first, I thought words had to be manually labeled as similar or related for the model to understand their connection. It was confusing to realize that the model can learn these relationships simply by analyzing which words appear together and how often they occur in similar contexts. This helped me understand that embeddings are built from patterns in language rather than from manually assigned meanings.
My main lesson learned is that computers can learn useful representations of language from patterns in data.
Word embeddings create a bridge between human language and machine learning by turning words into numerical vectors that models can work with.
I started this topic thinking that representinf words with numbers would be straightforward but I ended up learning that the interesting part isn't simply turning words into numbers, It is how numbers can capture relationships between words.
This is what makes word embeddings such an important idea in NLP
Sources
- Mikolov et al., Efficient Estimation of Word Representations in Vector Space (2013).
- Mikolov et al., Distributed Representations of Words and Phrases and their Compositionality (2013).
- Stanford NLP, GloVe: Global Vectors for Word Representation.
Top comments (0)