DEV Community

Cover image for Word Embeddings Explained: From Words to Vectors
Yvette Kazeneza
Yvette Kazeneza

Posted on

Word Embeddings Explained: From Words to Vectors

Understanding Word Embedding
When I first started working on Natural Language Processing (NLP), I thought the main challenge was simply getting a computer to understand text. During my movie-review sentiment analysis project, I learned that there is a much deeper problem: computers do not naturally understand words the way humans do.
For a computer, a word such as “amazing” is initially just text. To process it, we need to convert it into numbers. But simply assigning a number to each word does not tell a machine that “amazing” and “excellent” have related meanings, or that “terrible” expresses a very different sentiment.
This is where word embeddings become useful.
Word embeddings are a way of representing words as numerical vectors in a continuous mathematical space. Instead of treating every word as an unrelated symbol, embeddings allow words with similar meanings or usage patterns to have similar representations

What Are Word Embeddings?
A word embedding is a numerical representation of a word as a vector of real numbers. In simple terms, it is a way of turning words into numbers while trying to preserve information about their meaning and how they are used.
For example, imagine that we have these words:
king
queen
man
woman
apple
orange
If we represented each word using a simple token IDs, we might have:
king → 1
queen → 2
man → 3
woman → 4
apple → 5
orange → 6
The problem is that these numbers do not tell us anything about the relationships between the words. The fact that “king” has ID 1 and “apple” has ID 5 does not mean anything useful to the model.
Word embeddings work differently. Each word is represented by a vector containing many numerical values, for example:
king → [0.25, -0.13, 0.72, 0.41, ...]
queen → [0.27, -0.11, 0.70, 0.45, ...]
apple → [-0.62, 0.31, 0.08, -0.54, ...]
These numbers are learned from text. Words that appear in similar contexts tend to develop similar vector representations.
This is based on an important idea in NLP: a word's meaning can be understood partly from the words that appear around it.
For example, words such as “doctor”, “hospital”, “patient”, and “medicine” may frequently occur in related contexts. An embedding model can learn representations that reflect some of these relationships.
This is one of the major advantages of embeddings: instead of treating every word as completely independent, they give a model a way to capture relationships between words.
Why Do Word Embeddings Matter?
Before embeddings, one common way to represent text was using approaches such as Bag-of-Words or TF-IDF.
These methods can be useful, especially for tasks such as text classification. However, they generally represent words based on their occurrence in documents rather than learning a rich representation of their meaning.
Consider these two sentences:
The movie was excellent.
The movie was amazing.
A traditional representation can tell us that “excellent” and “amazing” are different words. However, it does not naturally understand that they express similar ideas.
Embeddings can capture this kind of relationship because they learn from how words are used in context.
This is particularly useful in NLP tasks such as:
• Sentiment analysis
• Text classification
• Machine translation
• Recommendation systems
• Question answering
• Named Entity Recognition
• Text similarity
For example, in my movie-review sentiment analysis project, the goal is to determine whether a review expresses a positive or negative opinion. Understanding relationships between words can be useful because different words and phrases can express similar sentiments.
The important point is that embeddings do not simply give a word a number. They create a representation in which the relationships between vectors can contain useful information about language.
How Does Word2Vec Work?
One of the most well-known methods for creating word embeddings is Word2Vec. Word2Vec was introduced by researchers at Google and became popular because it can learn useful relationships between words from large amounts of text.
• Converts words into numerical vectors for machine learning models
• Captures semantic relationships between words
• Words with similar meanings have similar vector representations
• Developed by Google researchers
The main idea behind Word2Vec is simple: words that appear in similar contexts tend to have similar meanings or relationships.
For example, consider these sentences:
The cat drinks milk.
The cat eats food.
The dog drinks water.
The dog eats food.
The words “cat” and “dog” appear in similar contexts. They are often surrounded by words such as “eats”, “drinks”, “food”, and “water”. Because of these patterns, Word2Vec can learn vector representations for “cat” and “dog” that are relatively close to each other in the embedding space.
Word2Vec mainly uses two approaches:
1. CBOW — Continuous Bag of Words
CBOW tries to predict a target word from the words around it,the model predicts the current word given context words within a specific window. The input layer contains the context words and the output layer contains the current word. The hidden layer contains the dimensions we want to represent the current word present at the output layer.
For example, consider:
The cat drinks milk.
If the target word is “drinks”, the surrounding words can provide context:
The cat ___ milk
The model learns from many examples like this and gradually adjusts the word vectors so that they become useful for predicting words from their context.
2. Skip-gram
Skip-gram works in the opposite direction. Instead of using surrounding words to predict the target word, it uses a target word to predict surrounding words.
The Skip gram predicts the surrounding context words within specific window given current word. The input layer contains the current word and the output layer contains the context words. The hidden layer contains the number of dimensions in which we want to represent current word present at the input layer.
For example:
The cat drinks milk.
If “cat” is the target word, the model tries to predict nearby words such as:
the
drinks
By repeating this process across a large collection of text, the model learns vector representations for words.
The difference can be summarized as:
CBOW:
Context words → Target word

Skip-gram:
Target word → Context words
One of the interesting things about Word2Vec is that the resulting vectors can capture relationships between words. For example, with a suitable trained embedding, vector operations can produce relationships such as:
king - man + woman ≈ queen
This does not mean the model is performing human-like reasoning. Rather, the mathematical relationships between the learned vectors can reflect patterns that were present in the training data.
This is what makes word embeddings different from simply assigning an ID to every word. The numbers in an embedding are learned from language patterns, and the distance and direction between vectors can contain useful information.

A Practical Word2Vec Experiment
To understand word embeddings beyond the theory, I decided to run a small experiment using the IMDB movie review dataset. Instead of using a pre-trained embedding, I trained my own Word2Vec model on the reviews.
I first cleaned and tokenized the reviews so that each review became a list of words. I then used Gensim to train a Word2Vec model using the Skip-gram approach.

import pandas as pd
from gensim.models import Word2Vec
import re

Load the dataset

df = pd.read_csv(
"/kaggle/input/datasets/rehanliaqat17/imbd-dataset/IMDB_Dataset.csv"
)

Clean and tokenize reviews

def tokenize(text):
text = re.sub(r"<.*?>", "", text)
text = text.lower()
text = re.sub(r"[^a-z\s]", "", text)
return text.split()

sentences = df["review"].apply(tokenize).tolist()

Train Word2Vec

model = Word2Vec(
sentences=sentences,
vector_size=100,
window=5,
min_count=5,
workers=4,
sg=1
)
chose a vector size of 100, meaning that every word in the vocabulary is represented by a vector containing 100 numerical values.

The window=5 parameter tells the model to look at nearby words within a context window of five words. I also used min_count=5, so words that appeared fewer than five times were ignored. Finally, sg=1 specifies that the model should use the Skip-gram architecture.

Looking at a Word's Embedding

After training the model, I could inspect the numerical representation learned for a word:

model.wv["movie"]

The result is a vector containing 100 numbers. The individual numbers do not have an obvious meaning by themselves. What matters is how this vector relates to the vectors of other words.

Finding Similar Words

I also used Word2Vec to find words that were close to "movie" in the learned embedding space:

similar_words = model.wv.most_similar("movie", topn=10)

similar_words

The model returned a list of words together with similarity scores.

My actual result was:

[('film', 0.919495165348053),
('movieand', 0.8412395715713501),
('filmit', 0.8398792147636414),
('filmbut', 0.835565447807312),
('movieit', 0.8224337697029114),
('moviebut', 0.8216363787651062),
('filmand', 0.793904721736908),
('filmso', 0.7863085269927979),
('flick', 0.7821688652038574),
('filmwhat', 0.7720096111297607)]

This was interesting because the model was not given a dictionary telling it which words were related. It learned these relationships from the contexts in which words appeared throughout the reviews.

** What I Learned**

One thing I found surprising while working with Word2Vec was that the numbers in a word embedding do not have an obvious meaning by themselves. When I looked at the 100 numerical values representing a word such as "movie", I could not look at one number and say that it represented something specific like sentiment or meaning.

Instead, the important information comes from the relationships between the vectors. Words that are used in similar contexts can end up close to each other in the embedding space.

I also found it interesting that Word2Vec could discover relationships between words simply by learning from the text. I did not manually tell the model that certain words were related. The model learned these patterns from the surrounding words in the movie reviews.

One thing I struggled with at first was understanding how converting words into vectors could actually preserve information about language. I initially thought that turning a word into numbers would just give the model numerical data without much meaning. After working with most_similar() and seeing the words that the model learned to associate with "movie", I understood better how the relationships between vectors can represent useful information.

This experiment also helped me understand why word embeddings are important in NLP. They provide a way for machine learning models to work with text while preserving some of the relationships between words, rather than treating every word as completely unrelated.

** References**

  1. Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient Estimation of Word Representations in Vector Space. This is the original Word2Vec paper describing the CBOW and Skip-gram approaches.

  2. Gensim Documentation. Word2Vec Model. Used as a reference for the Python implementation and methods such as most_similar().

  3. Pennington, J., Socher, R., & Manning, C. D. (2014). GloVe: Global Vectors for Word Representation. Stanford NLP. Used as a reference for the broader concept of word embeddings and vector representations.

Top comments (0)