DEV Community

Ragib Hasan
Ragib Hasan Subscriber

Posted on

One-Hot Encoding, Bag of Words > (BoW), TF-IDF, Unigram, Bigram, Trigram and N-Gram

This note follows the concepts and terminology presented in the
supplied lecture material and adds practical Python implementations so
that the ideas can be tested directly.


1. Why Do We Need Text Representation?

Machine Learning models work with numerical features. Human language,
however, is naturally represented as text.

Human Language
      │
      ▼
"People watch movies"
      │
      ▼
Text Representation
      │
      ▼
Numerical Vector
      │
      ▼
Machine Learning Model
      │
      ▼
Prediction / Classification
Enter fullscreen mode Exit fullscreen mode

Main idea

Text Representation = converting text into numbers that a
machine-learning algorithm can process.

This process is also commonly described as feature extraction from
text
.


2. Basic NLP Terminology

Before learning the algorithms, understand four terms.


Term Meaning Example


Corpus Collection of All reviews in a dataset
documents/texts

Document One individual text "people watch movie"

Vocabulary Unique words/features people, watch, movie
found in the corpus

Token An individual unit/word movie


Visual representation

                    CORPUS
                      │
          ┌───────────┼───────────┐
          ▼           ▼           ▼
      Document 1  Document 2  Document 3
          │           │           │
          ▼           ▼           ▼
       Tokens       Tokens       Tokens
          └───────────┼───────────┘
                      ▼
                 Vocabulary
                (unique words)
Enter fullscreen mode Exit fullscreen mode

3. The Overall Evolution

A useful way to remember the methods:

Raw Text
   │
   ▼
One-Hot Encoding
"Which word is this?"
   │
   ▼
Bag of Words
"How many times does each word occur?"
   │
   ├──────────────► N-Gram
   │                "Which words occur together?"
   │
   ▼
TF-IDF
"How important is this word in this document?"
Enter fullscreen mode Exit fullscreen mode

One-line memory trick

Method Main idea


One-Hot Identity
BoW Frequency
TF-IDF Importance
N-Gram Local sequence/context


4. One-Hot Encoding

4.1 Concept

Suppose our vocabulary is:

Vocabulary = [apple, book, movie, people]
Enter fullscreen mode Exit fullscreen mode

Each word receives a vector whose length equals the vocabulary size.

apple  → [1, 0, 0, 0]
book   → [0, 1, 0, 0]
movie  → [0, 0, 1, 0]
people → [0, 0, 0, 1]
Enter fullscreen mode Exit fullscreen mode

Only one position contains 1; all other positions contain 0.

Diagram

Vocabulary
┌────────┬────────┬────────┬────────┐
│ apple  │ book   │ movie  │ people │
└────────┴────────┴────────┴────────┘
    │        │        │        │
    ▼        ▼        ▼        ▼

 apple   = [1, 0, 0, 0]
 book    = [0, 1, 0, 0]
 movie   = [0, 0, 1, 0]
 people  = [0, 0, 0, 1]
Enter fullscreen mode Exit fullscreen mode

4.2 Main Problems

Sparsity

Most values are zero.

[0, 0, 0, 0, 0, 1, 0, 0, 0, 0, ...]
Enter fullscreen mode Exit fullscreen mode

If the vocabulary becomes very large, the vectors become very
high-dimensional.

No semantic relationship

For example:

king  → [1, 0, 0]
queen → [0, 1, 0]
car   → [0, 0, 1]
Enter fullscreen mode Exit fullscreen mode

The representation does not naturally express that king and queen
are semantically more related than king and car.

Out-of-vocabulary problem

A word not represented in the vocabulary cannot simply receive a normal
existing one-hot position.


5. Bag of Words (BoW)

5.1 Core Idea

Bag of Words represents a document using the frequency/count of
vocabulary words
.

Consider:

D1 = "people watch campus"
D2 = "people watch movie"
D3 = "people like movie"
Enter fullscreen mode Exit fullscreen mode

Vocabulary:

[people, watch, campus, movie, like]
Enter fullscreen mode Exit fullscreen mode

Now count every vocabulary word in every document.

             people  watch  campus  movie  like
D1              1      1      1       0      0
D2              1      1      0       1      0
D3              1      0      0       1      1
Enter fullscreen mode Exit fullscreen mode

This is called a Document-Term Matrix.


6. BoW Workflow

Documents
   │
   ▼
Tokenization
   │
   ▼
Create Vocabulary
   │
   ▼
Count Word Occurrences
   │
   ▼
Document-Term Matrix
   │
   ▼
Numerical Features
   │
   ▼
Machine Learning
Enter fullscreen mode Exit fullscreen mode

Important property

Every document gets a vector of the same length:

Vocabulary size = 5

D1 → [1, 1, 1, 0, 0]
D2 → [1, 1, 0, 1, 0]
D3 → [1, 0, 0, 1, 1]
Enter fullscreen mode Exit fullscreen mode

7. BoW in Python

CountVectorizer from scikit-learn can construct a Bag-of-Words
representation.

from sklearn.feature_extraction.text import CountVectorizer

documents = [
    "people watch campus",
    "people watch movie",
    "people like movie"
]

vectorizer = CountVectorizer()

X = vectorizer.fit_transform(documents)

print("Vocabulary:")
print(vectorizer.get_feature_names_out())

print("\nDocument-Term Matrix:")
print(X.toarray())
Enter fullscreen mode Exit fullscreen mode

What happens internally?

fit_transform()
     │
     ├── fit → learn vocabulary
     │
     └── transform → convert documents to vectors
Enter fullscreen mode Exit fullscreen mode

Useful inspection:

print(vectorizer.vocabulary_)
Enter fullscreen mode Exit fullscreen mode

The exact integer positions are implementation details, so
get_feature_names_out() is easier for beginners to read.


8. Binary BoW

Sometimes we only care whether a word is present, not how many times it
occurs.

Use:

vectorizer = CountVectorizer(binary=True)

X = vectorizer.fit_transform(documents)

print(X.toarray())
Enter fullscreen mode Exit fullscreen mode

Count vs Binary

Suppose:

D1 = "movie movie good"
Enter fullscreen mode Exit fullscreen mode

Count representation:

movie = 2
good  = 1
Enter fullscreen mode Exit fullscreen mode

Binary representation:

movie = 1
good  = 1
Enter fullscreen mode Exit fullscreen mode

So:

Count → "How many?"
Binary → "Is it present?"
Enter fullscreen mode Exit fullscreen mode

9. Limiting Vocabulary with max_features

A real corpus may contain thousands or millions of unique terms.

You can restrict the vocabulary:

vectorizer = CountVectorizer(max_features=1000)

X = vectorizer.fit_transform(documents)
Enter fullscreen mode Exit fullscreen mode

Conceptually:

Huge Vocabulary
       │
       ▼
Select limited number of features
       │
       ▼
Smaller Feature Matrix
       │
       ▼
Lower memory / computation
Enter fullscreen mode Exit fullscreen mode

10. Advantages and Limitations of BoW

Advantage Limitation


Simple Ignores word order
Easy to understand Sparse representation
Fixed-size vectors Weak semantic representation
Easy to implement Unseen words are ignored
Useful baseline "not good" can be problematic

Important example

"This movie is good"
"This movie is not good"
Enter fullscreen mode Exit fullscreen mode

A simple unigram BoW can contain:

movie
good
not
Enter fullscreen mode Exit fullscreen mode

But the relationship between not and good is not explicitly
represented.

This motivates N-Grams.


11. N-Gram

An N-Gram is a sequence of N consecutive tokens.

Suppose:

"I love this movie"
Enter fullscreen mode Exit fullscreen mode

Unigram

N = 1

I
love
this
movie
Enter fullscreen mode Exit fullscreen mode

Bigram

N = 2

I love
love this
this movie
Enter fullscreen mode Exit fullscreen mode

Trigram

N = 3

I love this
love this movie
Enter fullscreen mode Exit fullscreen mode

Diagram

Sentence:
I | love | this | movie

Unigram:
[I] [love] [this] [movie]

Bigram:
[I love] [love this] [this movie]

Trigram:
[I love this] [love this movie]
Enter fullscreen mode Exit fullscreen mode

12. N-Gram as a Sliding Window

Think of an N-Gram as a window moving across the sentence.

Sentence:
I | love | this | movie

Bigram window:

[I | love] this movie
      ↓
I [love | this] movie
         ↓
I love [this | movie]
Enter fullscreen mode Exit fullscreen mode

For a sequence of m tokens, the number of contiguous N-Grams is:

m - n + 1
Enter fullscreen mode Exit fullscreen mode

Example:

m = 4 tokens
n = 2

4 - 2 + 1 = 3 bigrams
Enter fullscreen mode Exit fullscreen mode

13. N-Gram with CountVectorizer

Unigram

from sklearn.feature_extraction.text import CountVectorizer

documents = [
    "this movie is good",
    "this movie is not good"
]

vectorizer = CountVectorizer(ngram_range=(1, 1))

X = vectorizer.fit_transform(documents)

print(vectorizer.get_feature_names_out())
print(X.toarray())
Enter fullscreen mode Exit fullscreen mode

Bigram

vectorizer = CountVectorizer(ngram_range=(2, 2))

X = vectorizer.fit_transform(documents)

print(vectorizer.get_feature_names_out())
Enter fullscreen mode Exit fullscreen mode

You will see features similar to:

this movie
movie is
is good
is not
not good
Enter fullscreen mode Exit fullscreen mode

Now the phrase:

not good
Enter fullscreen mode Exit fullscreen mode

is explicitly represented.


14. Unigram + Bigram

You can combine multiple N-Gram sizes:

vectorizer = CountVectorizer(ngram_range=(1, 2))

X = vectorizer.fit_transform(documents)

print(vectorizer.get_feature_names_out())
Enter fullscreen mode Exit fullscreen mode

This creates:

Unigrams
   +
Bigrams
Enter fullscreen mode Exit fullscreen mode

For example:

this
movie
is
good
not

this movie
movie is
is good
is not
not good
Enter fullscreen mode Exit fullscreen mode

15. Bigram vs Unigram

Feature Unigram Bigram


N 1 2
Captures individual words Yes Yes
Captures adjacent word pairs No Yes
Feature count Lower Higher
Context Low Better local context
Computation Lower Higher

Memory trick

Unigram → WORD
Bigram  → WORD + WORD
Trigram → WORD + WORD + WORD
Enter fullscreen mode Exit fullscreen mode

16. Why Increasing N Can Become Expensive

As N increases, possible combinations increase.

Unigram
   ↓
More features

Bigram
   ↓
More features

Trigram
   ↓
Even more features

4-gram
   ↓
Potentially much larger vocabulary
Enter fullscreen mode Exit fullscreen mode

Therefore:

Higher N
   │
   ├──► Larger vocabulary
   ├──► More memory
   ├──► More computation
   └──► Longer training/prediction
Enter fullscreen mode Exit fullscreen mode

This trade-off is important when designing NLP features.


17. TF-IDF

Bag of Words treats word counts as the main signal.

TF-IDF asks a different question:

How important is this word to this particular document compared with
the whole corpus?

The lecture's core intuition is:

High importance
=
frequent in this document
+
rare across the corpus
Enter fullscreen mode Exit fullscreen mode

18. TF --- Term Frequency

TF measures how frequently a term occurs in a document.

The lecture presents:

TF(term, document)
=
count of term in document
/
total number of terms in document
Enter fullscreen mode Exit fullscreen mode

Example

Document:

"movie movie good"
Enter fullscreen mode Exit fullscreen mode

Total terms:

3
Enter fullscreen mode Exit fullscreen mode

For movie:

count = 2

TF(movie)
= 2 / 3
= 0.667
Enter fullscreen mode Exit fullscreen mode

For good:

TF(good)
= 1 / 3
= 0.333
Enter fullscreen mode Exit fullscreen mode

19. IDF --- Inverse Document Frequency

IDF reduces the importance of terms appearing in many documents.

The lecture gives the basic form:

IDF(term)
=
log(
    total number of documents
    /
    number of documents containing the term
)
Enter fullscreen mode Exit fullscreen mode

So:

Appears in many documents
        ↓
Lower IDF

Appears in fewer documents
        ↓
Higher IDF
Enter fullscreen mode Exit fullscreen mode

Intuition

Suppose:

Document 1 → people watch movie
Document 2 → people like movie
Document 3 → people watch campus
Enter fullscreen mode Exit fullscreen mode

people appears everywhere.

Therefore:

people → low IDF
Enter fullscreen mode Exit fullscreen mode

campus appears in only one document.

Therefore:

campus → higher IDF
Enter fullscreen mode Exit fullscreen mode

20. TF-IDF

The basic relationship is:

TF-IDF = TF × IDF
Enter fullscreen mode Exit fullscreen mode

Workflow

                    DOCUMENT
                       │
             ┌─────────┴─────────┐
             ▼                   ▼
        Term Frequency       Document Frequency
             │                   │
             ▼                   ▼
             TF                  IDF
             │                   │
             └─────────┬─────────┘
                       ▼
                  TF × IDF
                       │
                       ▼
                 TF-IDF Vector
Enter fullscreen mode Exit fullscreen mode

21. TF-IDF in Python

from sklearn.feature_extraction.text import TfidfVectorizer

documents = [
    "people watch campus",
    "people watch movie",
    "people like movie"
]

vectorizer = TfidfVectorizer()

X = vectorizer.fit_transform(documents)

print("Features:")
print(vectorizer.get_feature_names_out())

print("\nTF-IDF Matrix:")
print(X.toarray())
Enter fullscreen mode Exit fullscreen mode

Inspect the feature names

features = vectorizer.get_feature_names_out()

for i, feature in enumerate(features):
    print(i, feature)
Enter fullscreen mode Exit fullscreen mode

This makes it easier to connect each matrix column to the corresponding
word.


22. BoW vs TF-IDF

Suppose a word appears in almost every document:

people
Enter fullscreen mode Exit fullscreen mode

BoW:

people = 1
Enter fullscreen mode Exit fullscreen mode

TF-IDF:

people → relatively lower importance
Enter fullscreen mode Exit fullscreen mode

Now suppose a word is specific to one document:

campus
Enter fullscreen mode Exit fullscreen mode

BoW:

campus = 1
Enter fullscreen mode Exit fullscreen mode

TF-IDF:

campus → relatively higher importance
Enter fullscreen mode Exit fullscreen mode

So:

BoW
↓
Counts occurrence

TF-IDF
↓
Weights importance
Enter fullscreen mode Exit fullscreen mode

23. Cosine Similarity with BoW

The lecture also connects similar documents with similar vector
patterns.

We can measure similarity using cosine similarity.

from sklearn.feature_extraction.text import CountVectorizer
from sklearn.metrics.pairwise import cosine_similarity

documents = [
    "people watch movie",
    "people watch film",
    "people like campus"
]

vectorizer = CountVectorizer()
X = vectorizer.fit_transform(documents)

similarity = cosine_similarity(X)

print(similarity)
Enter fullscreen mode Exit fullscreen mode

Conceptually:

Document A → Vector A
Document B → Vector B
                    │
                    ▼
             Cosine Similarity
                    │
                    ▼
             Similarity Score
Enter fullscreen mode Exit fullscreen mode

Higher similarity means the vectors point in more similar directions.


24. Complete Practical Workflow

                 RAW TEXT
                    │
                    ▼
              Clean / Prepare
                    │
                    ▼
                Tokenize
                    │
                    ▼
              Choose Method
                    │
        ┌───────────┼────────────┐
        ▼           ▼            ▼
    One-Hot        BoW        TF-IDF
                    │
                    └──────┐
                           ▼
                         N-Gram
                           │
                           ▼
                  Numerical Matrix
                           │
                           ▼
                   ML / NLP Model
                           │
                           ▼
                    Prediction
Enter fullscreen mode Exit fullscreen mode

25. Side-by-Side Comparison


Property One-Hot BoW TF-IDF N-Gram


Main purpose Word identity Word frequency Word Word sequence
importance

Uses counts No Yes Weighted Usually
counts count/weighted

Word order No No No Local order
captured

Semantic Very limited Limited Limited Better local
meaning context

Feature Vocabulary Vocabulary Vocabulary Can become much
dimension size size size larger

Sparse Yes Usually Usually Usually

Simple Very Very Moderate Moderate

Context None None None Local context


Important: N-Gram is not necessarily a completely separate
weighting family. It is a way of defining features. For example, you
can have Bag of Bigrams or TF-IDF over unigrams + bigrams.


26. One Example Through All Methods

Input:

"this movie is not good"
Enter fullscreen mode Exit fullscreen mode

One-Hot

Each word gets its own identity vector.

this  → [1,0,0,0,0]
movie → [0,1,0,0,0]
is    → [0,0,1,0,0]
not   → [0,0,0,1,0]
good  → [0,0,0,0,1]
Enter fullscreen mode Exit fullscreen mode

BoW

this movie is not good
 1     1    1  1   1
Enter fullscreen mode Exit fullscreen mode

TF-IDF

Each word gets a weight based on:

How frequent is it here?
            ×
How rare is it across the corpus?
Enter fullscreen mode Exit fullscreen mode

Bigram

this movie
movie is
is not
not good
Enter fullscreen mode Exit fullscreen mode

The important phrase:

not good
Enter fullscreen mode Exit fullscreen mode

is now preserved as a feature.


27. A Practical Python Comparison

from sklearn.feature_extraction.text import (
    CountVectorizer,
    TfidfVectorizer
)

documents = [
    "this movie is good",
    "this movie is not good",
    "people watch movie"
]

# -------------------------
# 1. Bag of Words
# -------------------------
bow = CountVectorizer()
X_bow = bow.fit_transform(documents)

print("BoW features:")
print(bow.get_feature_names_out())

print("BoW matrix:")
print(X_bow.toarray())


# -------------------------
# 2. TF-IDF
# -------------------------
tfidf = TfidfVectorizer()
X_tfidf = tfidf.fit_transform(documents)

print("\nTF-IDF features:")
print(tfidf.get_feature_names_out())

print("TF-IDF matrix:")
print(X_tfidf.toarray())


# -------------------------
# 3. Bigram
# -------------------------
bigram = CountVectorizer(ngram_range=(2, 2))
X_bigram = bigram.fit_transform(documents)

print("\nBigram features:")
print(bigram.get_feature_names_out())

print("Bigram matrix:")
print(X_bigram.toarray())
Enter fullscreen mode Exit fullscreen mode

28. What the Python Pipeline Is Doing

documents
    │
    ▼
CountVectorizer / TfidfVectorizer
    │
    ├── tokenize text
    │
    ├── build vocabulary
    │
    ├── construct features
    │
    └── create sparse matrix
    │
    ▼
X
    │
    ▼
Machine Learning Algorithm
Enter fullscreen mode Exit fullscreen mode

The returned X is usually a sparse matrix, which is useful because
text representations often contain many zeros.


29. Important scikit-learn Parameters

CountVectorizer

CountVectorizer(
    binary=False,
    max_features=None,
    ngram_range=(1, 1)
)
Enter fullscreen mode Exit fullscreen mode

binary

CountVectorizer(binary=True)
Enter fullscreen mode Exit fullscreen mode

Changes counts into presence/absence values.

max_features

CountVectorizer(max_features=1000)
Enter fullscreen mode Exit fullscreen mode

Limits the number of selected vocabulary features.

ngram_range

ngram_range=(1, 1)  # unigram
ngram_range=(2, 2)  # bigram
ngram_range=(3, 3)  # trigram
ngram_range=(1, 2)  # unigram + bigram
Enter fullscreen mode Exit fullscreen mode

30. Beginner Decision Guide

                 What do you need?
                        │
          ┌─────────────┼─────────────┐
          ▼             ▼             ▼
     Word identity   Frequency     Importance
          │             │             │
          ▼             ▼             ▼
      One-Hot          BoW         TF-IDF
                        │
                        ▼
                 Need word order?
                        │
                       YES
                        │
                        ▼
                     N-Gram
Enter fullscreen mode Exit fullscreen mode

Simple rule

  • Learning the basic idea of numerical word representation → One-Hot
  • Need a simple frequency-based baseline → BoW
  • Want to down-weight common words and emphasize document-specific words → TF-IDF
  • Need local word combinations such as not good → N-Gram
  • Need both single words and phrases → N-Gram range (1, 2)

31. Common Beginner Confusions

Confusion 1: Is BoW the same as N-Gram?

Not exactly.

BoW with unigrams
=
one-word features

BoW with bigrams
=
two-word features
Enter fullscreen mode Exit fullscreen mode

N-Gram defines the feature unit; BoW describes the "bag/count"
representation.


Confusion 2: Is TF-IDF a neural network?

No.

TF-IDF is a feature-weighting method.

Raw Text
   ↓
TF-IDF
   ↓
Numerical Features
   ↓
ML Model
Enter fullscreen mode Exit fullscreen mode

The ML model can then be something such as Logistic Regression, SVM,
etc.


Confusion 3: Does TF-IDF understand meaning?

Not in the same semantic sense as modern contextual language models.

It mainly uses:

term frequency
+
document frequency
Enter fullscreen mode Exit fullscreen mode

Therefore, semantic relationships are still limited.


Confusion 4: Why does N-Gram increase feature size?

Because combinations are added.

Unigram:
movie

Bigram:
good movie
watch movie
...

Trigram:
I watch movie
people watch movie
...
Enter fullscreen mode Exit fullscreen mode

More combinations → more possible features.


32. Final Mental Model

Imagine a document as a bag of words.

One-Hot

WORD
 ↓
Identity
Enter fullscreen mode Exit fullscreen mode

BoW

DOCUMENT
 ↓
Count words
Enter fullscreen mode Exit fullscreen mode

TF-IDF

DOCUMENT
 ↓
Count + corpus-level rarity
 ↓
Importance
Enter fullscreen mode Exit fullscreen mode

N-Gram

DOCUMENT
 ↓
Group consecutive words
 ↓
Local sequence/context
Enter fullscreen mode Exit fullscreen mode

33. Quick Revision Table

Question Answer


What is a corpus? Collection of documents
What is vocabulary? Unique features/words
What is a token? Individual text unit
What does One-Hot represent? Word identity
What does BoW represent? Word frequency
What does TF measure? Term frequency in a document
What does IDF measure? Inverse document frequency
TF-IDF formula? TF × IDF
What is a unigram? One token
What is a bigram? Two consecutive tokens
What is a trigram? Three consecutive tokens
What does binary=True do? Converts count presence to binary
What does max_features do? Limits vocabulary/features
What does ngram_range=(1,2) mean? Unigrams + bigrams
Main BoW weakness? Ignores word order
Why use bigrams? Capture local word combinations


34. Final Summary

                 TEXT REPRESENTATION
                         │
       ┌─────────────────┼─────────────────┐
       │                 │                 │
       ▼                 ▼                 ▼
   ONE-HOT              BoW             TF-IDF
   Identity           Frequency        Importance
       │                 │                 │
       └─────────────────┼─────────────────┘
                         │
                         ▼
                       N-GRAM
                    Local sequence
                         │
                         ▼
                 Numerical Features
                         │
                         ▼
                   ML / NLP Model
Enter fullscreen mode Exit fullscreen mode

Remember

One-Hot = "Which word?"\
BoW = "How many?"\
TF-IDF = "How important?"\
N-Gram = "Which words occur together?"


Source Note

This document is based primarily on the supplied lecture material on
text representation, feature extraction, One-Hot Encoding, Bag of Words,
TF-IDF and N-Grams. The Python sections provide practical
implementations of those concepts using scikit-learn; they are
included to make the lecture easier to reproduce and experiment with.

Top comments (0)