My classifier could recognize “full bin.” Then I wondered: what if a user writes “the waste container is overflowing”? Same problem. Very different words. Awkward.
When I first built my Waste Complaint Classifier, I used TF-IDF with Logistic Regression. The goal was simple: a resident submits a waste complaint, and the model automatically places it into a service category.
My first version used 1,000 labeled complaints, five balanced categories, and an 80/20 train-test split. The categories included Missed Pickup, Delayed Pickup, Full Bin, Illegal Dumping and Recycling Question.
The admin dashboard: complaints come in, the model classifies them, and the admin gets something useful instead of a mountain of text.
TF-IDF did its job. Very proudly.
TF-IDF helped me turn complaint text into numbers by giving more weight to useful words. So if someone wrote:
“The bin is completely full.”
Words like “bin” and “full” can become strong signals for the Full Bin category. Nice. Efficient. No drama.
Then I looked at these two complaints:
“The bin is full.”
vs.“The waste container is overflowing.”
As humans, we immediately see the connection. TF-IDF mainly sees different words. At that point I realized my model was good at spotting useful vocabulary, but I wanted to explore whether it could also capture meaning.
Me: “These two sentences mean almost the same thing.”
TF-IDF: “I have never seen these people together in my life.”
Enter word embeddings
Word embeddings represent words as dense numerical vectors. The interesting part is that words used in similar contexts can end up closer together in the vector space.
So words such as waste, garbage and rubbish - or full and overflowing - can potentially be represented as related rather than completely separate features.
TF-IDF asks: “Which words are important?”
Embeddings add: “Which words are related?”
Word2Vec got mentioned first... but nobody should get jealous
I am starting with Word2Vec. At a high level, Word2Vec learns word representations from the words that tend to appear around each other. If “collector,” “pickup,” “waste,” and “truck” repeatedly appear in related contexts, Word2Vec can learn that connection (Mikolov et al., 2013).
A quick concrete example
One simple way I can inspect what Word2Vec has learned is to ask it for the words it considers most similar to a word such as “waste”:
similar_words = model.wv.most_similar("waste", topn=5)
for word, score in similar_words:
print(word, score)
This lets me check which words the model sees as most similar to “waste.” If the results include words such as “garbage,” “rubbish,” or “collection,” that is a useful sign that the embedding space is learning meaningful relationships from the complaint text.
For this next version, I am starting with Word2Vec, but I am not stopping there. I also want to test GloVe and BERT embeddings to compare how they represent complaints and which one gives the best results for my classifier. Word2Vec should not get too comfortable just because I mentioned it first - GloVe and BERT are also getting their chance. In the end, I will compare their performance using the same evaluation metrics and let the results decide which embedding approach works best for Isuku.
GloVe: looking at the bigger picture
GloVe stands for Global Vectors for Word Representation. It learns from global word-word co-occurrence statistics - basically, patterns in how often words appear together across a large collection of text (Pennington et al., 2014). That gives me another way to capture semantic relationships in complaints.
BERT: context matters
BERT Word Embeddings TutorialBERT goes a step further because its embeddings are contextual. A word does not have to keep exactly the same representation everywhere; its meaning can depend on the sentence around it. That is useful because language is messy, and one word can mean different things in different situations. I found the practical BERT Word Embeddings Tutorial especially helpful for understanding how token embeddings can be extracted and interpreted in practice (McCormick & Ryan, 2019).
Word2Vec: local context
GloVe: global co-occurrence
BERT: context-aware representations
Three candidates. One complaint dataset. May the best embedding win.
Plot twist: newer does not automatically mean better
I am not assuming embeddings will automatically beat TF-IDF. TF-IDF can be very strong for text classification, especially on smaller labeled datasets like mine.
So I will compare the approaches using the same evaluation metrics: accuracy, precision, recall and F1 score. I also want to check examples where the wording changes but the underlying complaint stays the same.
Machine-learning survival rule: never announce the winner before the test set has spoken.
What this is teaching me
The biggest lesson for me is that text representation matters just as much as the classifier. My Logistic Regression model never sees a complaint the way I do - it sees numbers.
With TF-IDF, those numbers mainly represent word importance. With embeddings, those numbers can also represent relationships between words and, with BERT, the context in which those words appear.
So no, I am not breaking up with TF-IDF. I am simply letting Word2Vec, GloVe and BERT interview for the same job.
The next step is to run the experiments, compare the results fairly, and see which representation helps Isuku understand complaints best.
Because in machine learning, the technique with the fanciest name does not get the job. The metrics do.
References:
McCormick, C., & Ryan, N. (2019, May 14). BERT word embeddings tutorial. McCormick ML. Bert word embeddings
Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv. vector space
Pennington, J., Socher, R., & Manning, C. D. (2014). GloVe: Global vectors for word representation. Stanford NLP Group. GloVe

Top comments (0)