DEV Community

Ayush S Pangaonkar
Ayush S Pangaonkar

Posted on

Five Classifiers, One Dataset: What I Learned About Model Choice

If you hand five different classifiers the exact same data, how different do the results really look? That was the question behind my Admissions Predictions project, and the answer was more interesting than I expected.

The setup

The task is binary: given a student's entrance exam score and percentage, predict whether they get admitted.

The data is synthetic, 1,000 students, with both features drawn uniformly between 0 and 100. I generated the labels myself from a logistic function, then sampled each outcome from a Bernoulli distribution:

admission_prob = 1 / (1 + exp(-(0.1*X1 + 0.2*X2 - 10)))
Enter fullscreen mode Exit fullscreen mode

That gave 755 admitted and 245 not admitted. I used an 80/20 split (800 train, 200 test) with random_state=42, and trained all five models on identical splits so the comparison was fair.

The results

Model Train acc Test acc F1 (admitted)
Logistic Regression 94.13% 92.50% 95.05%
SVM (RBF) 93.75% 91.50% 94.46%
Decision Tree 100.00% 90.50% 93.77%
Random Forest (100 trees, entropy) 100.00% 93.00% 95.45%
KNN 95.50% 92.00% 94.81%

Three things I noticed

The decision tree memorized the training set. A perfect 100% on train and the lowest score on test, 90.5%. That roughly 9.5 point gap is what overfitting looks like in numbers.

Logistic regression held its own. It is the simplest model here and still scored 92.5%. Since I built the labels from a logistic function, a logistic model matching the data-generating process makes sense.

The gap between models was small. All five finished between 90.5% and 93.0%. The test set has 200 rows, so half a point is a single student. I would not call Random Forest a clear winner over Logistic Regression from this experiment alone.

What it changed for me

This was my first side-by-side model comparison, and it changed how I evaluate everything after it. I no longer trust train accuracy by itself, and I look at precision and recall next to accuracy, especially when the classes are not balanced (here, about 75% of students are admitted).


Code: github.com/bluntjudg/Admissions-Predictions

Series: Part 1 of 7 in my ML fundamentals revisit. Next in the series: Book Recommendation System.

Live projects I built after these basics:

More of my work is on GitHub.

Top comments (0)