This note follows the concepts and terminology presented in the
supplied lecture material and adds practical Python implementations so
that the ideas can be tested directly.
1. Why Do We Need Text Representation?
Machine Learning models work with numerical features. Human language,
however, is naturally represented as text.
Human Language
│
▼
"People watch movies"
│
▼
Text Representation
│
▼
Numerical Vector
│
▼
Machine Learning Model
│
▼
Prediction / Classification
Main idea
Text Representation = converting text into numbers that a
machine-learning algorithm can process.
This process is also commonly described as feature extraction from
text.
2. Basic NLP Terminology
Before learning the algorithms, understand four terms.
Term Meaning Example
Corpus Collection of All reviews in a dataset
documents/texts
Document One individual text "people watch movie"
Vocabulary Unique words/features people, watch, movie
found in the corpus
Token An individual unit/word movie
Visual representation
CORPUS
│
┌───────────┼───────────┐
▼ ▼ ▼
Document 1 Document 2 Document 3
│ │ │
▼ ▼ ▼
Tokens Tokens Tokens
└───────────┼───────────┘
▼
Vocabulary
(unique words)
3. The Overall Evolution
A useful way to remember the methods:
Raw Text
│
▼
One-Hot Encoding
"Which word is this?"
│
▼
Bag of Words
"How many times does each word occur?"
│
├──────────────► N-Gram
│ "Which words occur together?"
│
▼
TF-IDF
"How important is this word in this document?"
One-line memory trick
Method Main idea
One-Hot Identity
BoW Frequency
TF-IDF Importance
N-Gram Local sequence/context
4. One-Hot Encoding
4.1 Concept
Suppose our vocabulary is:
Vocabulary = [apple, book, movie, people]
Each word receives a vector whose length equals the vocabulary size.
apple → [1, 0, 0, 0]
book → [0, 1, 0, 0]
movie → [0, 0, 1, 0]
people → [0, 0, 0, 1]
Only one position contains 1; all other positions contain 0.
Diagram
Vocabulary
┌────────┬────────┬────────┬────────┐
│ apple │ book │ movie │ people │
└────────┴────────┴────────┴────────┘
│ │ │ │
▼ ▼ ▼ ▼
apple = [1, 0, 0, 0]
book = [0, 1, 0, 0]
movie = [0, 0, 1, 0]
people = [0, 0, 0, 1]
4.2 Main Problems
Sparsity
Most values are zero.
[0, 0, 0, 0, 0, 1, 0, 0, 0, 0, ...]
If the vocabulary becomes very large, the vectors become very
high-dimensional.
No semantic relationship
For example:
king → [1, 0, 0]
queen → [0, 1, 0]
car → [0, 0, 1]
The representation does not naturally express that king and queen
are semantically more related than king and car.
Out-of-vocabulary problem
A word not represented in the vocabulary cannot simply receive a normal
existing one-hot position.
5. Bag of Words (BoW)
5.1 Core Idea
Bag of Words represents a document using the frequency/count of
vocabulary words.
Consider:
D1 = "people watch campus"
D2 = "people watch movie"
D3 = "people like movie"
Vocabulary:
[people, watch, campus, movie, like]
Now count every vocabulary word in every document.
people watch campus movie like
D1 1 1 1 0 0
D2 1 1 0 1 0
D3 1 0 0 1 1
This is called a Document-Term Matrix.
6. BoW Workflow
Documents
│
▼
Tokenization
│
▼
Create Vocabulary
│
▼
Count Word Occurrences
│
▼
Document-Term Matrix
│
▼
Numerical Features
│
▼
Machine Learning
Important property
Every document gets a vector of the same length:
Vocabulary size = 5
D1 → [1, 1, 1, 0, 0]
D2 → [1, 1, 0, 1, 0]
D3 → [1, 0, 0, 1, 1]
7. BoW in Python
CountVectorizer from scikit-learn can construct a Bag-of-Words
representation.
from sklearn.feature_extraction.text import CountVectorizer
documents = [
"people watch campus",
"people watch movie",
"people like movie"
]
vectorizer = CountVectorizer()
X = vectorizer.fit_transform(documents)
print("Vocabulary:")
print(vectorizer.get_feature_names_out())
print("\nDocument-Term Matrix:")
print(X.toarray())
What happens internally?
fit_transform()
│
├── fit → learn vocabulary
│
└── transform → convert documents to vectors
Useful inspection:
print(vectorizer.vocabulary_)
The exact integer positions are implementation details, so
get_feature_names_out() is easier for beginners to read.
8. Binary BoW
Sometimes we only care whether a word is present, not how many times it
occurs.
Use:
vectorizer = CountVectorizer(binary=True)
X = vectorizer.fit_transform(documents)
print(X.toarray())
Count vs Binary
Suppose:
D1 = "movie movie good"
Count representation:
movie = 2
good = 1
Binary representation:
movie = 1
good = 1
So:
Count → "How many?"
Binary → "Is it present?"
9. Limiting Vocabulary with max_features
A real corpus may contain thousands or millions of unique terms.
You can restrict the vocabulary:
vectorizer = CountVectorizer(max_features=1000)
X = vectorizer.fit_transform(documents)
Conceptually:
Huge Vocabulary
│
▼
Select limited number of features
│
▼
Smaller Feature Matrix
│
▼
Lower memory / computation
10. Advantages and Limitations of BoW
Advantage Limitation
Simple Ignores word order
Easy to understand Sparse representation
Fixed-size vectors Weak semantic representation
Easy to implement Unseen words are ignored
Useful baseline "not good" can be problematic
Important example
"This movie is good"
"This movie is not good"
A simple unigram BoW can contain:
movie
good
not
But the relationship between not and good is not explicitly
represented.
This motivates N-Grams.
11. N-Gram
An N-Gram is a sequence of N consecutive tokens.
Suppose:
"I love this movie"
Unigram
N = 1
I
love
this
movie
Bigram
N = 2
I love
love this
this movie
Trigram
N = 3
I love this
love this movie
Diagram
Sentence:
I | love | this | movie
Unigram:
[I] [love] [this] [movie]
Bigram:
[I love] [love this] [this movie]
Trigram:
[I love this] [love this movie]
12. N-Gram as a Sliding Window
Think of an N-Gram as a window moving across the sentence.
Sentence:
I | love | this | movie
Bigram window:
[I | love] this movie
↓
I [love | this] movie
↓
I love [this | movie]
For a sequence of m tokens, the number of contiguous N-Grams is:
m - n + 1
Example:
m = 4 tokens
n = 2
4 - 2 + 1 = 3 bigrams
13. N-Gram with CountVectorizer
Unigram
from sklearn.feature_extraction.text import CountVectorizer
documents = [
"this movie is good",
"this movie is not good"
]
vectorizer = CountVectorizer(ngram_range=(1, 1))
X = vectorizer.fit_transform(documents)
print(vectorizer.get_feature_names_out())
print(X.toarray())
Bigram
vectorizer = CountVectorizer(ngram_range=(2, 2))
X = vectorizer.fit_transform(documents)
print(vectorizer.get_feature_names_out())
You will see features similar to:
this movie
movie is
is good
is not
not good
Now the phrase:
not good
is explicitly represented.
14. Unigram + Bigram
You can combine multiple N-Gram sizes:
vectorizer = CountVectorizer(ngram_range=(1, 2))
X = vectorizer.fit_transform(documents)
print(vectorizer.get_feature_names_out())
This creates:
Unigrams
+
Bigrams
For example:
this
movie
is
good
not
this movie
movie is
is good
is not
not good
15. Bigram vs Unigram
Feature Unigram Bigram
N 1 2
Captures individual words Yes Yes
Captures adjacent word pairs No Yes
Feature count Lower Higher
Context Low Better local context
Computation Lower Higher
Memory trick
Unigram → WORD
Bigram → WORD + WORD
Trigram → WORD + WORD + WORD
16. Why Increasing N Can Become Expensive
As N increases, possible combinations increase.
Unigram
↓
More features
Bigram
↓
More features
Trigram
↓
Even more features
4-gram
↓
Potentially much larger vocabulary
Therefore:
Higher N
│
├──► Larger vocabulary
├──► More memory
├──► More computation
└──► Longer training/prediction
This trade-off is important when designing NLP features.
17. TF-IDF
Bag of Words treats word counts as the main signal.
TF-IDF asks a different question:
How important is this word to this particular document compared with
the whole corpus?
The lecture's core intuition is:
High importance
=
frequent in this document
+
rare across the corpus
18. TF --- Term Frequency
TF measures how frequently a term occurs in a document.
The lecture presents:
TF(term, document)
=
count of term in document
/
total number of terms in document
Example
Document:
"movie movie good"
Total terms:
3
For movie:
count = 2
TF(movie)
= 2 / 3
= 0.667
For good:
TF(good)
= 1 / 3
= 0.333
19. IDF --- Inverse Document Frequency
IDF reduces the importance of terms appearing in many documents.
The lecture gives the basic form:
IDF(term)
=
log(
total number of documents
/
number of documents containing the term
)
So:
Appears in many documents
↓
Lower IDF
Appears in fewer documents
↓
Higher IDF
Intuition
Suppose:
Document 1 → people watch movie
Document 2 → people like movie
Document 3 → people watch campus
people appears everywhere.
Therefore:
people → low IDF
campus appears in only one document.
Therefore:
campus → higher IDF
20. TF-IDF
The basic relationship is:
TF-IDF = TF × IDF
Workflow
DOCUMENT
│
┌─────────┴─────────┐
▼ ▼
Term Frequency Document Frequency
│ │
▼ ▼
TF IDF
│ │
└─────────┬─────────┘
▼
TF × IDF
│
▼
TF-IDF Vector
21. TF-IDF in Python
from sklearn.feature_extraction.text import TfidfVectorizer
documents = [
"people watch campus",
"people watch movie",
"people like movie"
]
vectorizer = TfidfVectorizer()
X = vectorizer.fit_transform(documents)
print("Features:")
print(vectorizer.get_feature_names_out())
print("\nTF-IDF Matrix:")
print(X.toarray())
Inspect the feature names
features = vectorizer.get_feature_names_out()
for i, feature in enumerate(features):
print(i, feature)
This makes it easier to connect each matrix column to the corresponding
word.
22. BoW vs TF-IDF
Suppose a word appears in almost every document:
people
BoW:
people = 1
TF-IDF:
people → relatively lower importance
Now suppose a word is specific to one document:
campus
BoW:
campus = 1
TF-IDF:
campus → relatively higher importance
So:
BoW
↓
Counts occurrence
TF-IDF
↓
Weights importance
23. Cosine Similarity with BoW
The lecture also connects similar documents with similar vector
patterns.
We can measure similarity using cosine similarity.
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.metrics.pairwise import cosine_similarity
documents = [
"people watch movie",
"people watch film",
"people like campus"
]
vectorizer = CountVectorizer()
X = vectorizer.fit_transform(documents)
similarity = cosine_similarity(X)
print(similarity)
Conceptually:
Document A → Vector A
Document B → Vector B
│
▼
Cosine Similarity
│
▼
Similarity Score
Higher similarity means the vectors point in more similar directions.
24. Complete Practical Workflow
RAW TEXT
│
▼
Clean / Prepare
│
▼
Tokenize
│
▼
Choose Method
│
┌───────────┼────────────┐
▼ ▼ ▼
One-Hot BoW TF-IDF
│
└──────┐
▼
N-Gram
│
▼
Numerical Matrix
│
▼
ML / NLP Model
│
▼
Prediction
25. Side-by-Side Comparison
Property One-Hot BoW TF-IDF N-Gram
Main purpose Word identity Word frequency Word Word sequence
importance
Uses counts No Yes Weighted Usually
counts count/weighted
Word order No No No Local order
captured
Semantic Very limited Limited Limited Better local
meaning context
Feature Vocabulary Vocabulary Vocabulary Can become much
dimension size size size larger
Sparse Yes Usually Usually Usually
Simple Very Very Moderate Moderate
Context None None None Local context
Important: N-Gram is not necessarily a completely separate
weighting family. It is a way of defining features. For example, you
can have Bag of Bigrams or TF-IDF over unigrams + bigrams.
26. One Example Through All Methods
Input:
"this movie is not good"
One-Hot
Each word gets its own identity vector.
this → [1,0,0,0,0]
movie → [0,1,0,0,0]
is → [0,0,1,0,0]
not → [0,0,0,1,0]
good → [0,0,0,0,1]
BoW
this movie is not good
1 1 1 1 1
TF-IDF
Each word gets a weight based on:
How frequent is it here?
×
How rare is it across the corpus?
Bigram
this movie
movie is
is not
not good
The important phrase:
not good
is now preserved as a feature.
27. A Practical Python Comparison
from sklearn.feature_extraction.text import (
CountVectorizer,
TfidfVectorizer
)
documents = [
"this movie is good",
"this movie is not good",
"people watch movie"
]
# -------------------------
# 1. Bag of Words
# -------------------------
bow = CountVectorizer()
X_bow = bow.fit_transform(documents)
print("BoW features:")
print(bow.get_feature_names_out())
print("BoW matrix:")
print(X_bow.toarray())
# -------------------------
# 2. TF-IDF
# -------------------------
tfidf = TfidfVectorizer()
X_tfidf = tfidf.fit_transform(documents)
print("\nTF-IDF features:")
print(tfidf.get_feature_names_out())
print("TF-IDF matrix:")
print(X_tfidf.toarray())
# -------------------------
# 3. Bigram
# -------------------------
bigram = CountVectorizer(ngram_range=(2, 2))
X_bigram = bigram.fit_transform(documents)
print("\nBigram features:")
print(bigram.get_feature_names_out())
print("Bigram matrix:")
print(X_bigram.toarray())
28. What the Python Pipeline Is Doing
documents
│
▼
CountVectorizer / TfidfVectorizer
│
├── tokenize text
│
├── build vocabulary
│
├── construct features
│
└── create sparse matrix
│
▼
X
│
▼
Machine Learning Algorithm
The returned X is usually a sparse matrix, which is useful because
text representations often contain many zeros.
29. Important scikit-learn Parameters
CountVectorizer
CountVectorizer(
binary=False,
max_features=None,
ngram_range=(1, 1)
)
binary
CountVectorizer(binary=True)
Changes counts into presence/absence values.
max_features
CountVectorizer(max_features=1000)
Limits the number of selected vocabulary features.
ngram_range
ngram_range=(1, 1) # unigram
ngram_range=(2, 2) # bigram
ngram_range=(3, 3) # trigram
ngram_range=(1, 2) # unigram + bigram
30. Beginner Decision Guide
What do you need?
│
┌─────────────┼─────────────┐
▼ ▼ ▼
Word identity Frequency Importance
│ │ │
▼ ▼ ▼
One-Hot BoW TF-IDF
│
▼
Need word order?
│
YES
│
▼
N-Gram
Simple rule
- Learning the basic idea of numerical word representation → One-Hot
- Need a simple frequency-based baseline → BoW
- Want to down-weight common words and emphasize document-specific words → TF-IDF
- Need local word combinations such as
not good→ N-Gram - Need both single words and phrases → N-Gram range
(1, 2)
31. Common Beginner Confusions
Confusion 1: Is BoW the same as N-Gram?
Not exactly.
BoW with unigrams
=
one-word features
BoW with bigrams
=
two-word features
N-Gram defines the feature unit; BoW describes the "bag/count"
representation.
Confusion 2: Is TF-IDF a neural network?
No.
TF-IDF is a feature-weighting method.
Raw Text
↓
TF-IDF
↓
Numerical Features
↓
ML Model
The ML model can then be something such as Logistic Regression, SVM,
etc.
Confusion 3: Does TF-IDF understand meaning?
Not in the same semantic sense as modern contextual language models.
It mainly uses:
term frequency
+
document frequency
Therefore, semantic relationships are still limited.
Confusion 4: Why does N-Gram increase feature size?
Because combinations are added.
Unigram:
movie
Bigram:
good movie
watch movie
...
Trigram:
I watch movie
people watch movie
...
More combinations → more possible features.
32. Final Mental Model
Imagine a document as a bag of words.
One-Hot
WORD
↓
Identity
BoW
DOCUMENT
↓
Count words
TF-IDF
DOCUMENT
↓
Count + corpus-level rarity
↓
Importance
N-Gram
DOCUMENT
↓
Group consecutive words
↓
Local sequence/context
33. Quick Revision Table
Question Answer
What is a corpus? Collection of documents
What is vocabulary? Unique features/words
What is a token? Individual text unit
What does One-Hot represent? Word identity
What does BoW represent? Word frequency
What does TF measure? Term frequency in a document
What does IDF measure? Inverse document frequency
TF-IDF formula? TF × IDF
What is a unigram? One token
What is a bigram? Two consecutive tokens
What is a trigram? Three consecutive tokens
What does binary=True do? Converts count presence to binary
What does max_features do? Limits vocabulary/features
What does ngram_range=(1,2) mean? Unigrams + bigrams
Main BoW weakness? Ignores word order
Why use bigrams? Capture local word combinations
34. Final Summary
TEXT REPRESENTATION
│
┌─────────────────┼─────────────────┐
│ │ │
▼ ▼ ▼
ONE-HOT BoW TF-IDF
Identity Frequency Importance
│ │ │
└─────────────────┼─────────────────┘
│
▼
N-GRAM
Local sequence
│
▼
Numerical Features
│
▼
ML / NLP Model
Remember
One-Hot = "Which word?"\
BoW = "How many?"\
TF-IDF = "How important?"\
N-Gram = "Which words occur together?"
Source Note
This document is based primarily on the supplied lecture material on
text representation, feature extraction, One-Hot Encoding, Bag of Words,
TF-IDF and N-Grams. The Python sections provide practical
implementations of those concepts using scikit-learn; they are
included to make the lecture easier to reproduce and experiment with.
Top comments (0)