DEV Community

Cover image for What I Learned About Word Embeddings While Building a FAQ Chatbot
Mbonimpa Kelia
Mbonimpa Kelia

Posted on

What I Learned About Word Embeddings While Building a FAQ Chatbot

When I set out to build a simple FAQ chatbot for Python programming questions, I thought I'd be done in a weekend. I was wrong — but in the best way. Along the way, I ran into a concept that changed how I think about text: word embeddings.

This post is a lessons-learned write-up of what I discovered. If you're a developer getting into NLP, this might save you a few hours of confusion.

What Are Word Embeddings, Anyway?
Computers don't understand words. They understand numbers. So any NLP system has to convert text into numbers before it can do anything useful.

The naive way to do this is one-hot encoding or bag-of-words: assign each word a number, and represent a sentence as a list of word counts. This works, but it has a big problem — it treats every word as completely independent. To a bag-of-words model, "car" and "automobile" are as unrelated as "car" and "banana."

Word embeddings solve this. An embedding represents each word as a dense vector of real numbers — usually 100 to 300 dimensions — where the position of the vector in space reflects the word's meaning. Words that appear in similar contexts end up close together in this vector space. "Dog" and "puppy" end up near each other. "King" and "queen" end up near each other. "Python" and "programming" end up near each other.

Why does this matter? Because once words are vectors, you can do math on them. You can measure similarity, find clusters, and even solve analogies.

The Method That Clicked for Me: Word2Vec
The assignment notes I was working from mentioned Word2Vec, and it's the one that finally made embeddings click.

Word2Vec comes in two flavors:

CBOW (Continuous Bag of Words): predicts a word from its surrounding context

Skip-gram: predicts the surrounding context from a word

The key insight is the distributional hypothesis: words that appear in similar contexts have similar meanings. If "cat" and "dog" both frequently appear near "pet," "feed," and "vet," Word2Vec learns that they're semantically related.

Training happens by sliding a window over a large corpus. For each word, the model tries to predict its neighbors. Over millions of examples, the weights of a small neural network adjust until each word has a useful vector. What surprised me is that nobody labels this data — the text itself is the supervision. You just throw a lot of text at it and it figures out meaning on its own.

The Analogy That Blew My Mind
The classic Word2Vec demo is this:

text
king - man + woman ≈ queen
Translated: take the vector for "king," subtract "man," add "woman" — and the closest vector to the result is "queen." The model wasn't told that kings and queens are related, or that man and woman are a gender pair. It learned both facts from raw text alone.

I tried to reproduce this with a Python example using gensim:

python
from gensim.models import Word2Vec

Assume sentences is a list of tokenized sentences

model = Word2Vec(sentences, vector_size=100, window=5, min_count=5)

Find words most similar to "python"

similar = model.wv.most_similar("python", topn=5)
print(similar)

[('programming', 0.71), ('java', 0.68), ('code', 0.65), ...]

The fact that a model trained on raw Wikipedia text knows "python" is close to "programming" still amazes me.

What I Struggled With
Here's the honest part: I did not use word embeddings in my final FAQ chatbot.

My project used TF-IDF, which is a sparse, frequency-based vectorization method — not an embedding. TF-IDF is fast, interpretable, and works great for small FAQ datasets. But it has a real limitation: it can't handle paraphrases.

For example, if a user asks "How do I set up Python?" and my FAQ has "How do I install Python?", TF-IDF sees "set up" and "install" as completely different words. It might miss the match. A word embedding model would recognise that "set up" and "install" mean the same thing and would match them correctly.

So I learned embeddings as theory but implemented TF-IDF in practice. If I were to expand this project to hundreds or thousands of FAQs, the very first upgrade I'd make is swapping TF-IDF for sentence embeddings like Sentence-BERT, which can encode entire questions into a single dense vector and handle paraphrasing gracefully.

Other Embedding Methods Worth Knowing
While learning, I also came across:

GloVe (Global Vectors): similar to Word2Vec but trained on global word co-occurrence statistics rather than a sliding window.

FastText: extends Word2Vec by treating words as bags of character n-grams. This means it can generate vectors for words it has never seen — very handy for typos or rare words.

Contextual embeddings (BERT, ELMo): the same word gets different vectors in different sentences. "Bank" in "river bank" vs. "money bank" gets two different vectors. This is the foundation of modern LLMs.

What I'd Tell Another Developer
If you're just starting with NLP:

Learn TF-IDF first. It's simpler, faster, and good enough for many tasks.

Learn Word2Vec next. It builds the intuition for why embeddings work.

Then try sentence embeddings (SBERT). This is the practical modern approach for semantic search.

Don't skip the theory. Understanding why embeddings work — the distributional hypothesis — makes it much easier to know when to use them.

Embeddings turned text into geometry. That's the mental shift that changed how I look at NLP problems.

Sources:

Mikolov et al., "Efficient Estimation of Word Representations in Vector Space" (2013)

Pennington et al., "GloVe: Global Vectors for Word Representation" (2014)

Bojanowski et al., "Enriching Word Vectors with Subword Information" (FastText, 2017)

Reimers & Gurevych, "Sentence-BERT" (2019)

Dev.to community posts on TF-IDF vs. embeddings

Top comments (0)