DEV Community

Ayush S Pangaonkar
Ayush S Pangaonkar

Posted on

When My House Price Model Hit 100% Accuracy, I Should Have Been Worried

My first script on the Kaggle House Prices data reported 100% training accuracy. I was pleased for about as long as it took to look at my feature list.

The bug: the answer was in the inputs

The script was meant to predict SalePrice, but I had included SalePrice itself as one of the input columns. The model could simply read the answer off its own inputs. Once the leaked column was gone, the same setup gave a meaningless 52% on held-out data, because I was treating a price as if it were a class label.

A suspiciously perfect number is usually a sign that something is broken, not that the model is good.

Fixing it

The data is the standard Kaggle split: train.csv with 1,460 rows and 81 columns. I used these features: LotArea, BsmtFinSF1, TotalBsmtSF, GrLivArea, FullBath, HalfBath, BedroomAbvGr, GarageArea and SaleCondition.

Predicting SalePrice properly, as regression and without the leak:

Model R2 RMSE MAE
Decision Tree Regressor 0.75 $43,830 $29,724
Random Forest Regressor 0.84 $34,656 $21,784

The mean sale price is $180,921, so the Random Forest's average error of about $21.8k is roughly 12% of the average price. That is a believable result.

The second lesson: beat the baseline first

My second script classifies SaleCondition. About 82% of houses sell under "Normal" conditions, so the target is heavily imbalanced.

Model Train acc Test acc
Decision Tree 99.9% 70.2%
Random Forest 99.9% 80.8%
Always guess "Normal" n/a 82.1%

Neither model beats a rule that always says "Normal". An 80% accurate model sounds good until you learn that guessing the majority class gets you 82%.

What I do now

Two habits came out of this project. I check my feature list for the target (or anything derived from it) before I train anything. And I compute a naive baseline before I trust any classification score.


Code: github.com/bluntjudg/House-Price-Predictions-

Series: Part 3 of 7 in my ML fundamentals revisit. Next in the series: Text Sentiment Analysis.

Live projects I built after these basics:

More of my work is on GitHub.

Top comments (0)