DEV Community

Ayush S Pangaonkar
Ayush S Pangaonkar

Posted on

96.5% Accurate Spam Filter, and the 58 Spam Messages It Let Through

My spam classifier scores 96.5% accuracy on the test set. That sounds great. The confusion matrix tells a more useful story.

The setup

The dataset is mail_data.csv, with 5,572 messages: 4,825 ham (86.6%) and 747 spam (13.4%), no missing values. I mapped labels to 0 (spam) and 1 (ham), split 70/30 with random_state=3, fit a TfidfVectorizer (English stopwords removed, vocabulary of 6,896 terms) on the training messages, and trained a LogisticRegression model.

Metric Value
Train accuracy 96.6%
Test accuracy 96.5%
Always predict ham 86.6%

Test accuracy beats the always-ham baseline by about ten points, so the model is learning something real.

What the accuracy hides

On the test set (232 actual spam, 1,440 actual ham):

Predicted spam Predicted ham
Actual spam 174 58
Actual ham 1 1,439
  • Spam precision is 99.4%. Only one real message was flagged as spam.
  • Spam recall is 75.0%. 58 of 232 spam messages got through.
  • Ham recall is 99.9%.

So the filter almost never blocks a real message, but it misses about one spam message in four. For a spam filter, letting spam through is usually the safer failure compared with blocking real mail, but that is a trade-off worth stating plainly instead of only quoting 96.5%.

Two corrections I had to make

When I went back through this project, my own write-up did not match my code.

  1. I had described the project as using Naive Bayes and SVM. The script only trains and evaluates Logistic Regression. There is no Naive Bayes or SVM anywhere in it.
  2. The notebook builds a clean_message column with NeatText, but the model is trained on the raw Message text. The cleaning step had no effect on the results.

I fixed the README to describe what the code does.

Takeaway

This was the first project where I looked past one accuracy number into a confusion matrix. The model trades recall for precision on the minority class, and that is a measurable choice, not just a score to quote. It also taught me to re-read my own project descriptions against the code before I publish them.


Code: github.com/bluntjudg/Spam-Classifier-

Series: Part 5 of 7 in my ML fundamentals revisit. Next in the series: Speed Distance Model.

Live projects I built after these basics:

More of my work is on GitHub.

Top comments (0)