DEV Community

Marie Claire Uwiringiyimana
Marie Claire Uwiringiyimana

Posted on

From Words to Vectors: Understanding Word Embeddings with FastText

When I started learning Natural Language Processing (NLP), one question I had was, "How can a computer work with words when computers understand numbers better than human language?"

This is where word embeddings come in. Word embeddings are a way of representing words as numerical vectors. Instead of giving a machine only the word "cat," we represent it using a list of numbers. These numbers allow machine-learning models to work with words mathematically and learn relationships between them [1].

For example, a word could have a representation like

cat → [0.21, 0.45, -0.13, 0.72, ...]

A real embedding can contain hundreds of numbers. The important thing is not what one individual number means, but the position of the whole vector in the vector space. Words that have similar meanings or are used in similar contexts can have similar vector representations [1].

Why Do Word Embeddings Matter?

Before learning about embeddings, I looked at one-hot encoding. With one-hot encoding, every word is represented using a vector containing mostly zeros and one value of 1.

For example:

cat → [1, 0, 0]
dog → [0, 1, 0]
banana → [0, 0, 1]

This tells the computer that the words are different, but it does not show that "cat" and "dog" are more related than "cat" and "banana."

Word embeddings provide a better representation because they can learn relationships between words from how those words are used in text. They also use dense vectors with fewer dimensions than a vocabulary-sized one-hot representation, making them more useful for machine-learning applications [1].

This is important for NLP tasks such as text classification, sentiment analysis, search, machine translation, and other applications that need numerical representations of language [1].

A famous example of relationships captured by word embeddings is:

king - man + woman ≈ queen

This does not mean that every embedding model will produce this exact result, but it shows how vector representations can capture useful relationships between words [2].

The Embedding Method I Chose: FastText

FastText is based on the skip-gram approach and extends the idea of word embeddings by using subword information. Instead of looking at a word only as a complete unit, FastText also represents a word using character n-grams [3].

A character n-gram is simply a group of consecutive characters.

For example, parts of:

playing

can include:

pla
lay
ayi
yin
ing

The exact n-grams depend on the model settings.

This is useful because related words often share parts of their spelling:

play
played
playing
player
plays

FastText can learn information from these shared character patterns. This is especially useful for rare words and words that were not directly seen during training [3].

My Python Experiment

I used the gensim Python library to train a small FastText model on a few example sentences. The model was configured with 100-dimensional word vectors, a context window of 3, and the Skip-gram approach.

After training, I used most_similar() to find the words closest to “playing”. In my experiment, “plays” had the highest similarity score, followed by “played”, “love”, “scored”, and “player”. The similarity scores show how close their learned vector representations are.

I also inspected the vector generated for “playing”. The model represents the word as a list of numerical values rather than as text. Since I used vector_size=100, the complete representation contains 100 values.

What I Found Interesting

I then tested “playingly”, a word that was not included in my original training sentences. FastText was still able to generate a vector for it. This helped me understand one of the main advantages of FastText: it uses character n-grams (subword information) to build word representations. This means the model can make use of parts of a word, even when the complete word was not seen during training [3][4].

The screenshots below show the actual code and outputs from my Jupyter Notebook.


            **Figure 1**
Enter fullscreen mode Exit fullscreen mode

What I Learned

Before working with embeddings, I thought a vector was simply a list of numbers and did not understand how those numbers could represent language.

One thing I learned is that we should not try to give every individual number a specific human meaning. The whole vector is what matters.

I also initially found character n-grams confusing. I was thinking about words as complete units, but FastText showed me that words can also be understood through smaller pieces. For example, words such as "play," "played," "playing," and "player" share many character patterns.

This made me understand why FastText can be useful for languages and datasets containing rare words, different word forms, or words that may not appear frequently in the training data.

Conclusion

Word embeddings provide a way to connect human language with machine learning by representing words as numerical vectors. Instead of treating every word as completely independent, embeddings can capture relationships based on how words are used.

FastText extends this idea by using character n-grams, allowing it to learn from smaller parts of words. My small Python experiment helped me see this concept in practice rather than only reading about it.

My main takeaway is simple:

Words
↓
Context + subword information
↓
Numerical vectors
↓
Compare relationships.
↓
NLP applications

For me, the biggest lesson was that word embeddings are not just about converting words into numbers. They are about creating numerical representations that allow machines to work with relationships in language.

References

[1] IBM. What Are Word Embeddings?

[2] Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient Estimation of Word Representations in Vector Space.

[3] Bojanowski, P., Grave, E., Joulin, A., & Mikolov, T. (2017). Enriching Word Vectors with Subword Information.

[4] Gensim Documentation. FastText Model and Similarity Operations.

Top comments (0)