In my first NLP project, I built an SMS spam classifier. I cleaned the text, turned each message into numbers with TF-IDF, compared three models, and the best one, a Linear SVM, caught 93% of spam on messages it had never seen.
Then I noticed a blind spot. To my model, the words "free" and "complimentary" were as unrelated as "free" and "banana." A spammer who switched to different words with the same meaning could slip right past it. That's the problem word embeddings solve, and learning about them changed how I think about text.
The problem with counting words
TF-IDF gives every word in the vocabulary its own column. With single words and word pairs, my TF-IDF setup gave every message 5,790 columns, and almost all of them were zeros for any given message. That's called a sparse representation.
The bigger issue is that each column is independent. "Cheap" and "affordable" are separate columns with no connection between them. The model has no idea they mean nearly the same thing unless it sees enough examples of both.
What are word embeddings?
A word embedding represents each word as a short, dense list of numbers, usually 50 to 300 of them, where most values aren't zero. Those numbers are learned so that words used in similar contexts end up with similar vectors.
I like to think of it as a map. Every word gets coordinates, and words with related meanings sit close together: "cheap" near "affordable," "king" near "queen," and "banana" far away from both. To measure how close two words are, we usually use cosine similarity, which checks whether two vectors point in the same direction.
The idea behind this comes from the linguist J.R. Firth, who wrote in 1957: "You shall know a word by the company it keeps." If two words keep showing up next to the same neighbours, they probably mean something similar.
Why they matter
Similarity: a model can tell that "prize" and "reward" are related, even if it saw only one of them during training.
Better generalization: knowledge learned about one word transfers to its neighbours on the map.
Compact input: 100 meaningful numbers per word instead of thousands of mostly-zero columns.
The foundation of modern NLP: today's large language models still start by turning words (or pieces of words) into embedding vectors.
How Word2Vec works (at a high level)
Word2Vec was introduced by Mikolov and colleagues at Google in 2013. The clever part is that it learns embeddings through a fake task.
It slides a small window over huge amounts of text. Take the sentence "text FREE to claim your prize." If the centre word is "free," its context words are "text," "to," and "claim." Word2Vec has two versions:
Skip-gram: given the centre word, predict the words around it.
CBOW (Continuous Bag of Words): given the surrounding words, predict the centre word.
A small neural network gets better and better at this prediction game. But here's the twist: we don't actually care about the predictions. What we keep are the weights the network learned for each word. Those weights are the embeddings. Words that help predict similar neighbours, like "prize" and "reward," end up with similar vectors.
Two other methods are worth knowing:
GloVe (Stanford, 2014) builds vectors from global statistics: how often every pair of words appears together across the whole corpus.
FastText (Facebook, 2016) breaks words into smaller character pieces, so it can build a vector even for a word it has never seen, like an unusual spelling or slang.
Trying it myself
I used Gensim to load pre-trained GloVe vectors, trained on Wikipedia and news text, and explored them:
import gensim.downloader as api
# Load small pre-trained GloVe vectors (50 numbers per word)
model = api.load("glove-wiki-gigaword-50")
# Which words are closest to "free"?
print(model.most_similar("free", topn=5))
# How similar are these pairs? (closer to 1 = more similar)
print(model.similarity("cheap", "affordable"))
print(model.similarity("cheap", "banana"))
# The famous analogy: king - man + woman = ?
print(model.most_similar(positive=["king", "woman"], negative=["man"], topn=1))
When I ran it, the closest words to "free" were "allowing," "allowed," "giving," "for," and "without." "Cheap" and "affordable" scored 0.71, while "cheap" and "banana" scored only 0.40. And the analogy returned "queen."
I didn't expect that: GloVe learned "free" in the sense of freedom or permission, because that's how the word is used in Wikipedia and news. In my spam data, "free" almost always meant a free prize. Same word, different meaning, depending on the training text.
What surprised me, and what I struggled with
The analogy surprised me the most. Nobody told the model that a king is a male queen. That relationship emerged purely from which words appear near each other in text. It felt like the model had discovered meaning on its own.
The numbers themselves mean nothing. At first I expected each of the 50 numbers to stand for something, like one for "royalty" and one for "gender." They don't. Meaning only exists in the relationships between vectors, which took me a while to accept.
Slang is trickier than it looks. SMS messages are full of words like "wif," "lar," and "lor." Interestingly, my TF-IDF model (in spam_classifier.ipynb) learned that slang like "lor" and "da" was a strong sign of a normal message, because friends text casually and spammers don't. But when I checked whether the pre-trained GloVe vectors knew these words, all four I tested ("wif," "lar," "lor," and "tkts") were in its vocabulary, which I didn't expect. Being in the vocabulary only means GloVe has a vector for that string, though, not that it learned the SMS meaning: the vector comes from however the word happens to appear in Wikipedia and news text. For slang that isn't in the vocabulary at all, FastText's subword approach is designed to help, because it can build a vector from character pieces.
Embeddings learn our biases. Since they learn from human writing, they absorb the stereotypes in it. Bolukbasi and colleagues showed in 2016 that analogies can reproduce gender stereotypes about jobs. That's something to keep in mind before using embeddings in real systems.
One word, one vector. "Bank" gets the same vector whether it means a river bank or a money bank. Newer contextual models like BERT fix this by giving a word a different vector depending on the sentence around it.
What's next for me
My capstone project involves understanding short messages that residents send about water supply in Rwanda, many of them in Kinyarwanda. Kinyarwanda builds words by adding prefixes and suffixes, so a single root can appear in many forms. FastText's subword idea seems like a promising fit, and it's what I want to explore next.
Key takeaways
TF-IDF treats every word as unrelated; embeddings capture that some words mean similar things.
Word2Vec learns embeddings as a side effect of predicting a word's neighbours.
Pre-trained embeddings are powerful but reflect their training text, including its gaps and biases.
Sources
Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient Estimation of Word Representations in Vector Space. https://arxiv.org/abs/1301.3781
Pennington, J., Socher, R., & Manning, C. (2014). GloVe: Global Vectors for Word Representation. https://nlp.stanford.edu/projects/glove/
Bojanowski, P., Grave, E., Joulin, A., & Mikolov, T. (2017). Enriching Word Vectors with Subword Information. https://arxiv.org/abs/1607.04606
Bolukbasi, T., et al. (2016). Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings. https://arxiv.org/abs/1607.06520
Firth, J.R. (1957). A Synopsis of Linguistic Theory, 1930–1955.
Gensim documentation: https://radimrehurek.com/gensim/
Top comments (0)