My first script on the Kaggle House Prices data reported 100% training accuracy. I was pleased for about as long as it took to look at my feature list.
The bug: the answer was in the inputs
The script was meant to predict SalePrice, but I had included SalePrice itself as one of the input columns. The model could simply read the answer off its own inputs. Once the leaked column was gone, the same setup gave a meaningless 52% on held-out data, because I was treating a price as if it were a class label.
A suspiciously perfect number is usually a sign that something is broken, not that the model is good.
Fixing it
The data is the standard Kaggle split: train.csv with 1,460 rows and 81 columns. I used these features: LotArea, BsmtFinSF1, TotalBsmtSF, GrLivArea, FullBath, HalfBath, BedroomAbvGr, GarageArea and SaleCondition.
Predicting SalePrice properly, as regression and without the leak:
| Model | R2 | RMSE | MAE |
|---|---|---|---|
| Decision Tree Regressor | 0.75 | $43,830 | $29,724 |
| Random Forest Regressor | 0.84 | $34,656 | $21,784 |
The mean sale price is $180,921, so the Random Forest's average error of about $21.8k is roughly 12% of the average price. That is a believable result.
The second lesson: beat the baseline first
My second script classifies SaleCondition. About 82% of houses sell under "Normal" conditions, so the target is heavily imbalanced.
| Model | Train acc | Test acc |
|---|---|---|
| Decision Tree | 99.9% | 70.2% |
| Random Forest | 99.9% | 80.8% |
| Always guess "Normal" | n/a | 82.1% |
Neither model beats a rule that always says "Normal". An 80% accurate model sounds good until you learn that guessing the majority class gets you 82%.
What I do now
Two habits came out of this project. I check my feature list for the target (or anything derived from it) before I train anything. And I compute a naive baseline before I trust any classification score.
Code: github.com/bluntjudg/House-Price-Predictions-
Series: Part 3 of 7 in my ML fundamentals revisit. Next in the series: Text Sentiment Analysis.
Live projects I built after these basics:
- ATS Resume Analyzer, a Streamlit app that scores resumes against job descriptions
- AI-Based Loan Verification System, a Streamlit app for automated loan eligibility checks
- Subreach, a two-agent Reddit tool
More of my work is on GitHub.
Top comments (0)