An earlier churn-prediction project of mine hit a data-quality snag partway through: 33% of rows had a missing value somewhere across the dataset's 20 columns. My fix at the time was to cut the feature set down to just 6 columns - the ones with the highest individual correlation to churn which brought missingness down to a more manageable 10%, and then drop those remaining incomplete rows. It worked, in the sense that the models ran and produced reasonable-looking accuracy numbers. Going back to check whether that was actually the right call: it wasn't, and the gap between the two approaches is bigger than I expected.
One more thing worth flagging upfront: the project's own README describes this as a "telecom" dataset. It's actually an e-commerce dataset (customer churn for an online retailer) a second documentation mismatch, on top of the modeling issue below.
The original approach
5,630 customer records, 20 columns (tenure, order history, complaints, satisfaction score, and more), predicting whether a customer churns. Missingness was scattered thin across many columns individually (roughly 4.5–5.5% each), but because different rows were missing different columns, the union came out to 33% of all rows having something missing somewhere.
The fix: keep only the 6 features with the highest individual (univariate) correlation to the churn label Tenure, Complain, DaySinceLastOrder, CashbackAmount, MaritalStatus, Gender then drop the now-smaller set of incomplete rows (10.1% of the data, down from 33%). Train Logistic Regression, SVM, Decision Tree, Random Forest (9 trees), and XGBoost on that.
What I checked
Two design choices, both untested at the time:
- Does dropping 13 features because they didn't individually correlate well with churn actually make sense? A feature can be a poor univariate predictor on its own and still contribute real information once combined with others — correlation with the target in isolation isn't the same thing as usefulness in a model.
- Is dropping rows better or worse than imputing the missing values and keeping all 18 usable features (everything except customer ID)?
I rebuilt the pipeline using median imputation for missing numeric values (rather than dropping), kept all 18 features, one-hot encoded the categoricals, and trained a properly-sized Random Forest (300 trees instead of 9, with class_weight='balanced' since churn here is a 16.8%-minority class) evaluated the same way on a held-out test split, and additionally checked with 5-fold cross-validation since a single split can be misleading on its own (a lesson from an earlier project in this series).
Results
| Approach | Data used | Test F1 | Test Recall | Test Precision | 5-fold CV F1 (mean) |
|---|---|---|---|---|---|
| Original (6 features, dropna, 9-tree RF) | 89.9% of rows | 0.796 | 0.767 | 0.828 | 0.774 |
| Full features (18), imputed, 300-tree RF, balanced | 100% of rows | 0.869 | 0.805 | 0.944 | 0.888 |
Using the full feature set with imputation instead of dropping rows and features improved cross-validated F1 from 0.774 to 0.888 a meaningful jump, not a marginal one, and it held up consistently across folds (not just a lucky single split).
Checking which of the previously-excluded 13 features actually mattered once given the chance:
WarehouseToHome, OrderAmountHikeFromlastYear, NumberOfAddress, and SatisfactionScore all ranked in the model's top 10 most important features none of these made the original cut, because none of them had strong individual correlation with churn. In combination with other features, they clearly carried real signal the original approach never had access to.
The actual lesson
Filtering features by univariate correlation with the target, before any modeling, silently assumes that a feature's value can only show up as a direct linear relationship with the outcome on its own. Interaction effects and non-linear contributions exactly the kind of pattern a Random Forest is supposed to be good at finding get filtered out before the model ever gets a chance to use them. The "clean the data by cutting to a small, high-correlation feature set" instinct, applied here, wasn't really solving the missing-data problem efficiently it was solving it by discarding most of the dataset's actual information content.
Where this still has limits
- I used median imputation, the simplest reasonable option a more sophisticated imputation method (e.g., model-based imputation) might close the gap further or reveal it's already near the ceiling for this dataset.
-
class_weight='balanced'and imputation were changed together in this comparison, not in isolation, so this result shows the combined effect of the improved pipeline rather than isolating exactly how much came from imputation alone versus the other changes (more trees, class weighting). A cleaner ablation would test each change separately. - This is still one dataset the "don't pre-filter by univariate correlation" lesson is a reasonable general instinct, but not a universal law; on a different dataset with truly independent-effect features, both approaches might land closer together.
Code
Full comparison, including the exact rows-dropped calculation and both pipelines: github.com/DostonUr/Customer_churn
If you've run into the same "filter by univariate correlation before modeling" habit and found similar or different results, I'd like to hear about it reply here or find me on LinkedIn.
Top comments (0)