DEV Community

Tuyishimire claudine
Tuyishimire claudine

Posted on

I Was Lost With Word Embeddings, So I Built My First Example

When I started learning about Natural Language Processing (NLP), I came across the term word embedding. At first, I was confused about how a computer could understand words.

As humans, we understand that words have meanings and relationships. For example, we know that king and queen are related, and that cat and dog are both animals. But computers work with numbers. So, how can we make a computer understand words?

This is where word embeddings come in.

What is a word embedding?

A word embedding is a way of representing words as numbers (vectors) so that a computer can understand the relationships and meanings between words.

The key idea is that words used in similar contexts get similar numbers. So king and queen end up close together, and cat and dog are close, while cat and table are far apart.

This matters because older methods, like bag-of-words, treat every word as completely different from every other word. With those methods, "rain" is no closer to "storm" than it is to "table", so the computer misses the meaning. Word embeddings fix this by placing related words near each other, which helps in tasks like search, translation, and sentiment analysis.

How does Word2Vec work?

Word2Vec is a technique used to convert words into numerical vectors. It learns these vectors by looking at the context in which words appear.

The main idea behind Word2Vec is that words used in similar contexts are likely to have similar meanings or relationships. For example, look at these sentences:

I drink coffee every morning.

I drink tea every morning.

The words coffee and tea are different, but they appear in a very similar context. Both come after the word "drink" and before "every morning."

Because coffee and tea appear in similar contexts, Word2Vec learns that they are related and gives them similar numerical representations (vectors). It learns this by reading a large amount of text, so nobody has to write the meanings by hand.

GloVe is another embedding method. It also gives each word a vector, but it learns from how often words appear together across a whole text collection. I used GloVe in my example because a pre-trained version was small, quick to download, and easy to run. Glove allowed me to work with word embeddings in Google Colab and explore how words are represented as vectors and how their relationships can be measured.

My example

To see word embeddings in action, I used the gensim library with GloVe, a pre-trained model where every word is a list of 50 numbers. I did not train it myself: I downloaded a version already trained on Wikipedia and news text. I asked it for the words closest to "rain":

import gensim.downloader as api
model = api.load("glove-wiki-gigaword-50")

for word, score in model.most_similar("rain", topn=5):
    print(f"{word}: {score:.3f}")
Enter fullscreen mode Exit fullscreen mode

Result:

rains: 0.878
torrential: 0.843
winds: 0.833
downpour: 0.801
snow: 0.795
Enter fullscreen mode Exit fullscreen mode

The scores are cosine similarity: the closer to 1, the more similar the two words are.

Then I tried the famous analogy: king - man + woman. In the code, positive words are added and negative words are subtracted.

print(model.most_similar(positive=["king", "woman"], negative=["man"], topn=1))
Enter fullscreen mode Exit fullscreen mode

Result:

[('queen', 0.852)]
Enter fullscreen mode Exit fullscreen mode

What I noticed

The five words closest to "rain" were rains, torrential, winds, downpour, and snow. They all make sense, because they are about weather. Even "snow" appears, which shows the model groups words by how they are used, not by exact meaning: rain and snow are different, but they show up in very similar contexts. The model was never given a dictionary, so it learned this only from how words appear together in text.

The analogy surprised me most. When I did king - man + woman, the model returned "queen," with a similarity score of about 0.85. Seeing a computer solve a word puzzle with math made the idea of word embeddings finally click for me. Not every analogy works this well, though, and embeddings can also pick up biases from the text they learn from.

Testing it on my own field

My research is on rainwater harvesting, so I tested three words from my field:

for w in ["catchment", "cistern", "runoff"]:
    print(w)
    for word, score in model.most_similar(w, topn=5):
        print(f"  {word}: {score:.3f}")
Enter fullscreen mode Exit fullscreen mode

Result:

catchment
  floodplain: 0.808
  reservoir: 0.774
  drainage: 0.768
  catchments: 0.752
  streams: 0.747
cistern
  cisterns: 0.804
  outfall: 0.695
  chimney: 0.670
  earthen: 0.655
  cavern: 0.652
runoff
  run-off: 0.830
  landslide: 0.675
  elections: 0.667
  reelection: 0.660
  election: 0.659
Enter fullscreen mode Exit fullscreen mode

"Catchment" worked well: all five neighbors are about water, and the scores stay high (0.75 to 0.81). "Cistern" was mixed. Its top neighbor is just the plural "cisterns" (0.804), and the scores then drop to around 0.65 to 0.70, where words like "chimney" and "cavern" appear. "Runoff" was the surprise: apart from its spelling variant "run-off", most of its neighbors are about elections, because a runoff is also a type of vote, and that meaning is common in news text. GloVe gives each word only one vector, so it cannot separate the two meanings. I only tested the small 50-dimension version, but this shows that a general-purpose model can misunderstand words from a specialized field, and that lower scores can be a warning sign.

"Catchment" worked well, and most neighbors of "cistern" made sense, although "chimney" and "cavern" did not. "Runoff" was the surprise: most of its neighbors are about elections, because a runoff is also a type of vote, and that meaning is common in news text. GloVe gives each word only one vector, so it cannot separate the two meanings. I only tested the small 50-dimension version, but this shows that a general-purpose model can misunderstand words from a specialized field.

My Google Colab Notebook

You can find the complete code and implementation here:
[view my google Colab Notebook]
https://colab.research.google.com/drive/1O_FaOlHn9sduGY8tQ8HmScnF1vIF3_UT?usp=sharing

What I learned and what's next

At first, "words as numbers" sounded confusing. What helped me was running a real example and seeing that the numbers carry meaning: similar words end up close together. I also learned that the model only knows what it has seen in its training text. GloVe was trained on general text like Wikipedia and news, and my test showed the limit: for "runoff", it mostly returned election words instead of water words.

Next, I want to use word embeddings in my NLP project to see if they improve my results compared with a simple bag-of-words approach. I also want to try training a small Word2Vec model on my own rainwater harvesting text, to see if it handles words like "runoff" better.

Sources

Top comments (0)