What happened when I let a genetic algorithm pick my model's settings instead of guessing them myself.
The problem I was trying to solve
I built a model that predicts whether a phone company customer is going to cancel their plan ("churn") or stick around. I used real data from IBM — about 7,043 customers, and only about 1 in 4 of them actually churned.
Here's the thing about this kind of problem: the model can mess up in two different ways, and they are NOT equally bad.
Example 1: Imagine the model says "this customer is totally fine, they're staying" — but they actually cancel next month. That's a real customer, and real money, walking out the door. Ouch.
Example 2: Now imagine the model says "uh oh, this customer might leave!" — but they were actually going to stay the whole time. What happens? The company probably just sends them a coupon or a "we miss you" email they didn't need. A little wasteful, but not a big deal.
So missing a real churner (Example 1) is way more expensive than falsely worrying about a loyal customer (Example 2). That means I want my model to lean toward catching as many real churners as possible — even if it means a few false alarms along the way. In machine learning terms, that means I care more about recall than precision. (Quick refresher: recall = "out of everyone who really did churn, how many did I catch?" Precision = "out of everyone I flagged as a churn risk, how many actually churned?")
My first version of the model, using pretty normal, hand-picked settings, looked like this:
| Metric | My original model |
|---|---|
| Recall (Churn) | 0.481 (48%) |
| Precision (Churn) | 0.623 (62%) |
| F1 (Churn) | 0.543 |
48% recall means my model was basically a coin flip on catching real churners — it missed more than half of them! Not great. I later nudged this up to 68% by manually adjusting a setting called the "decision threshold" (basically, lowering how confident the model needs to be before it raises its hand and says "churn risk"). That helped, but I had to fiddle with it by trial and error to find that number.
Enter GASearchCV: let evolution do the guessing
Here's an analogy. Imagine you're trying to bake the perfect chocolate chip cookie, and you have 6 things you can change: how much sugar, how much flour, oven temperature, baking time, how much butter, and whether you chill the dough first. Trying every single combination would take forever.
A genetic algorithm does something smarter, kind of like how evolution works in nature:
- Bake 20 random batches of cookies (the "population").
- Taste-test all of them and rank which ones are best.
- Take the best ones, mix and combine their recipes a little (like "breeding" them), and throw in a few random tweaks (mutations).
- Bake a new batch of 20 with these improved recipes.
- Repeat this a bunch of times (called "generations").
Each generation, the recipes tend to get a little better, because you're always building on your best results so far instead of starting from scratch. That's exactly what GASearchCV, from the Python library sklearn-genetic-opt, does — except instead of cookie recipes, it's tweaking your model's settings (like how many trees to use, how deep each tree can grow, etc.), and instead of a taste test, it scores each "recipe" using cross-validation.
Here's the actual code — don't worry, I'll explain each part with an example:
from sklearn_genetic import GASearchCV
from sklearn_genetic.space import Integer, Categorical, Continuous
param_grid = {
'n_estimators': Integer(50, 300), # e.g. "try between 50 and 300 trees"
'max_depth': Integer(3, 20), # e.g. "try tree depths from 3 to 20"
'min_samples_split': Integer(2, 20),
'min_samples_leaf': Integer(1, 10),
'max_features': Continuous(0.1, 1.0),
'class_weight': Categorical([None, 'balanced']), # e.g. "try both options"
}
evolved_rf = GASearchCV(
estimator=RandomForestClassifier(random_state=42),
param_grid=param_grid,
cv=5, # taste-test each recipe 5 different ways
scoring='recall', # judge recipes by recall, since that's what I care about
population_size=20, # 20 recipes per generation
generations=15, # repeat the process 15 times
n_jobs=-1,
)
evolved_rf.fit(X_train, y_train)
Think of Integer(50, 300) like telling the algorithm "you're allowed to try anywhere from 50 to 300 trees in the forest — go find the sweet spot yourself," instead of me just guessing "let's do 100 trees" and hoping for the best.
What actually happened
After the algorithm ran through its 15 generations of "baking and taste-testing," here's what it landed on:
| Metric | My original model | GA-tuned model |
|---|---|---|
| Recall (Churn) | 0.481 (48%) | 0.826 (83%) |
| Precision (Churn) | 0.623 (62%) | 0.475 (48%) |
| F1 (Churn) | 0.543 | 0.604 |
Recall basically doubled — from 48% to 83%! That's a huge jump, and it's even better than the 68% I got earlier by manually fiddling with the threshold. And here's the part that surprised me most: even though precision dropped, the F1 score (which balances both) still went up. So this isn't just "trading one thing for another" — it's an actual, real improvement.
The winning "recipe" the algorithm found was:
{'n_estimators': 299, 'max_depth': 3, 'min_samples_split': 14,
'min_samples_leaf': 3, 'max_features': 0.112, 'class_weight': 'balanced'}
The part I found genuinely surprising: max_depth=3. That means each tree in the forest is really shallow — like, barely a tree at all, more like a bush. I never would have guessed that on my own; my instinct would've been "deeper trees = smarter model." But apparently, for this dataset, a bunch of simple, shallow trees working together (plus class_weight='balanced', which tells the model "pay extra attention to the rare churn cases") worked way better than one big complicated tree.
What I took away from this
If you already know which mistake is worse for your specific problem (like how missing a churner is worse than a false alarm), you can literally tell GASearchCV "go optimize for that" using the scoring parameter, and it'll search way more of the possibility space than you'd ever try by hand — kind of like having a tireless assistant testing hundreds of cookie recipes overnight while you sleep. It doesn't replace understanding why the trade-off matters in the first place — you still need to know your problem. But it definitely saves you from a lot of manual, trial-and-error guessing.
Full code: tune_churn_rf_with_gasearchcv.py
Top comments (0)