When I started learning about Natural Language Processing (NLP), I came across the term word embedding. At first, I was confused about how a computer could understand words.
As humans, we understand that words have meanings and relationships. For example, we know that king and queen are related, and that cat and dog are both animals. But computers work with numbers. So, how can we make a computer understand words?
This is where word embeddings come in.
What is a word embedding?
A word embedding is a way of representing words as numbers (vectors) so that a computer can understand the relationships and meanings between words.
The key idea is that words used in similar contexts get similar numbers. So king and queen end up close together, and cat and dog are close, while cat and table are far apart.
This matters because older methods, like bag-of-words, treat every word as completely different from every other word. With those methods, "rain" is no closer to "storm" than it is to "table", so the computer misses the meaning. Word embeddings fix this by placing related words near each other, which helps in tasks like search, translation, and sentiment analysis.
How does Word2Vec work?
Word2Vec is a technique used to convert words into numerical vectors. It learns these vectors by looking at the context in which words appear.
The main idea behind Word2Vec is that words used in similar contexts are likely to have similar meanings or relationships. For example, look at these sentences:
I drink coffee every morning.
I drink tea every morning.
The words coffee and tea are different, but they appear in a very similar context. Both come after the word "drink" and before "every morning."
Because coffee and tea appear in similar contexts, Word2Vec learns that they are related and gives them similar numerical representations (vectors). It learns this by reading a large amount of text, so nobody has to write the meanings by hand.
GloVe is another embedding method. It also gives each word a vector, but it learns from how often words appear together across a whole text collection. I used GloVe in my example because a pre-trained version was small, quick to download, and easy to run. Glove allowed me to work with word embeddings in Google Colab and explore how words are represented as vectors and how their relationships can be measured.
My example
To see word embeddings in action, I used the gensim library with GloVe, a pre-trained model where every word is a list of 50 numbers. I did not train it myself: I downloaded a version already trained on Wikipedia and news text. I asked it for the words closest to "rain":
import gensim.downloader as api
model = api.load("glove-wiki-gigaword-50")
for word, score in model.most_similar("rain", topn=5):
print(f"{word}: {score:.3f}")
Result:
rains: 0.878
torrential: 0.843
winds: 0.833
downpour: 0.801
snow: 0.795
The scores are cosine similarity: the closer to 1, the more similar the two words are.
Then I tried the famous analogy: king - man + woman. In the code, positive words are added and negative words are subtracted.
print(model.most_similar(positive=["king", "woman"], negative=["man"], topn=1))
Result:
[('queen', 0.852)]
What I noticed
The five words closest to "rain" were rains, torrential, winds, downpour, and snow. They all make sense, because they are about weather. Even "snow" appears, which shows the model groups words by how they are used, not by exact meaning: rain and snow are different, but they show up in very similar contexts. The model was never given a dictionary, so it learned this only from how words appear together in text.
The analogy surprised me most. When I did king - man + woman, the model returned "queen," with a similarity score of about 0.85. Seeing a computer solve a word puzzle with math made the idea of word embeddings finally click for me. Not every analogy works this well, though, and embeddings can also pick up biases from the text they learn from.
Testing it on my own field
My research is on rainwater harvesting, so I tested three words from my field:
for w in ["catchment", "cistern", "runoff"]:
print(w)
for word, score in model.most_similar(w, topn=5):
print(f" {word}: {score:.3f}")
Result:
catchment
floodplain: 0.808
reservoir: 0.774
drainage: 0.768
catchments: 0.752
streams: 0.747
cistern
cisterns: 0.804
outfall: 0.695
chimney: 0.670
earthen: 0.655
cavern: 0.652
runoff
run-off: 0.830
landslide: 0.675
elections: 0.667
reelection: 0.660
election: 0.659
"Catchment" worked well: all five neighbors are about water, and the scores stay high (0.75 to 0.81). "Cistern" was mixed. Its top neighbor is just the plural "cisterns" (0.804), and the scores then drop to around 0.65 to 0.70, where words like "chimney" and "cavern" appear. "Runoff" was the surprise: apart from its spelling variant "run-off", most of its neighbors are about elections, because a runoff is also a type of vote, and that meaning is common in news text. GloVe gives each word only one vector, so it cannot separate the two meanings. I only tested the small 50-dimension version, but this shows that a general-purpose model can misunderstand words from a specialized field, and that lower scores can be a warning sign.
"Catchment" worked well, and most neighbors of "cistern" made sense, although "chimney" and "cavern" did not. "Runoff" was the surprise: most of its neighbors are about elections, because a runoff is also a type of vote, and that meaning is common in news text. GloVe gives each word only one vector, so it cannot separate the two meanings. I only tested the small 50-dimension version, but this shows that a general-purpose model can misunderstand words from a specialized field.
My Google Colab Notebook
You can find the complete code and implementation here:
[view my google Colab Notebook]
https://colab.research.google.com/drive/1O_FaOlHn9sduGY8tQ8HmScnF1vIF3_UT?usp=sharing
What I learned and what's next
At first, "words as numbers" sounded confusing. What helped me was running a real example and seeing that the numbers carry meaning: similar words end up close together. I also learned that the model only knows what it has seen in its training text. GloVe was trained on general text like Wikipedia and news, and my test showed the limit: for "runoff", it mostly returned election words instead of water words.
Next, I want to use word embeddings in my NLP project to see if they improve my results compared with a simple bag-of-words approach. I also want to try training a small Word2Vec model on my own rainwater harvesting text, to see if it handles words like "runoff" better.
Sources
- Pennington, J., Socher, R., & Manning, C. (2014). GloVe: Global Vectors for Word Representation. Stanford NLP.
- Mikolov, T., et al. (2013). Efficient Estimation of Word Representations in Vector Space.
- gensim library documentation.
- Notes from my NLP class.
Top comments (0)