One example in my sentiment project still bothers me. The dataset contains the text "I have a feeling i will fail french #fuckfrench", and its label is joy. A human would read that as anxious or frustrated.
I will come back to that. First, the project.
What I built
This is not a positive/negative classifier. It predicts one of eight emotions: joy, sadness, fear, anger, surprise, neutral, disgust or shame. The dataset has 34,792 short, mostly Twitter-style texts, and the classes are very uneven:
| Emotion | Count | Share |
|---|---|---|
| joy | 11,045 | 31.8% |
| sadness | 6,722 | 19.3% |
| fear | 5,410 | 15.6% |
| anger | 4,297 | 12.3% |
| surprise | 4,062 | 11.7% |
| neutral | 2,254 | 6.5% |
| disgust | 856 | 2.5% |
| shame | 146 | 0.4% |
The pipeline:
- Clean the text with NeatText (strip user handles and stopwords).
- Vectorize with
CountVectorizer. - Train a
LogisticRegressionmodel inside a scikit-learn pipeline. - Compare with a
MultinomialNBbaseline.
I used a 70/30 split, stratified by emotion.
Results
| Model | Test accuracy |
|---|---|
| Logistic Regression | 63.5% |
| Naive Bayes | 57.2% |
| Always predict "joy" | 31.8% |
Logistic Regression beats the majority baseline by a wide margin. Macro F1 is 0.60 and weighted F1 is 0.63.
A single accuracy number hides the spread between classes:
| Emotion | Precision | Recall | F1 | Test support |
|---|---|---|---|---|
| fear | 0.74 | 0.68 | 0.71 | 1,623 |
| joy | 0.63 | 0.77 | 0.69 | 3,313 |
| anger | 0.64 | 0.55 | 0.59 | 1,289 |
| surprise | 0.60 | 0.42 | 0.49 | 1,219 |
| disgust | 0.54 | 0.18 | 0.27 | 257 |
Disgust is the weakest, with a recall of 0.18. Shame scores the highest F1 (0.75), but it has only 44 test examples, so I would not lean on that number.
Back to the "joy" label
If the model predicts joy for text like that, it is not necessarily wrong. It may simply have learned the pattern that is in the training labels, flaws included.
The dataset looks like it was built from hashtag-based labeling, which is a common way to build emotion datasets. I cannot verify how every row was labeled, but a hashtag does not guarantee the emotion in the sentence, so some noise in the labels would not be surprising.
What I took from it
This was my first time dealing with class imbalance and reading per-class metrics instead of one accuracy figure. It was also the first time a "wrong" prediction sent me to the raw labels instead of the model. Look at your labels before you blame your model.
Code: github.com/bluntjudg/Text-Sentiment-Analysis-
Series: Part 4 of 7 in my ML fundamentals revisit. Next in the series: Spam Classifier.
Live projects I built after these basics:
- ATS Resume Analyzer, a Streamlit app that scores resumes against job descriptions
- AI-Based Loan Verification System, a Streamlit app for automated loan eligibility checks
- Subreach, a two-agent Reddit tool
More of my work is on GitHub.
Top comments (0)