TL;DR: Before anything clever, fit a model that guesses the majority class and a plain logistic regression. Together they take seconds, and they're the only way to tell what the expensive model actually bought you. On six public datasets, guessing alone was 95.3% accurate on hypothyroid. Default boosted trees beat logistic regression by 2 to 15 AUC points. A 200-fit hyperparameter search added more than half a point on only one dataset, and across four random splits its gain on the other five averaged between −0.3 and +0.2 points. Skip the cheap models and the complex one takes credit for all of it.
Ask an AI coding assistant to "build a model to predict churn" from a CSV and you'll often get a notebook that looks finished: preprocessing, a gradient-boosted model, a hyperparameter search, a classification report, maybe a feature-importance chart. Before you read any of the numbers, search the notebook for one thing: a model that ignores the features entirely.
If there isn't one, none of the numbers mean anything yet. "95% accuracy" is excellent if guessing gets you 60%, and embarrassing if guessing gets you 95%. A 0.93 AUC from a tuned ensemble is a big win if a plain logistic regression gets 0.80, and the ensemble is mostly overhead if the logistic regression gets 0.92. The assistant didn't do anything wrong. It answered the question you asked, and "build a model" doesn't ask for a floor to measure it against.
So the first model you fit should be one you'd be embarrassed to ship. It isn't there to be deployed. It's there to set the score everything else has to beat.
Four models, six datasets
To see how much the floor changes the story, here are four models of increasing effort on six public binary-classification datasets from the Penn Machine Learning Benchmark (PMLB) collection:
-
Guess the majority class. scikit-learn ships this as
DummyClassifier, and its documentation is blunt about the job: it "makes predictions that ignore the input features" and "serves as a simple baseline to compare against other more complex classifiers." - Logistic regression, with one-hot encoded categories and scaled numeric columns. Nothing tuned.
-
Gradient-boosted trees on default settings (scikit-learn's
HistGradientBoostingClassifier, a histogram-based implementation inspired by LightGBM). On the two datasets with more than 10,000 rows, the defaults also switch on early stopping, which is a small amount of built-in tuning. - The same boosted trees after a hyperparameter search: 40 random configurations over learning rate, tree size, leaf size, regularisation and number of rounds, each scored with 5-fold cross-validation. That's 200 model fits plus a final refit of the winner, and it's the step that makes a notebook look thorough.
Each dataset gets one stratified 80/20 train/test split. The search only sees the training data. Scores are on the held-out 20%. Categorical columns are one-hot encoded for the logistic regression and passed to the trees as native categories. The churn data's phone-number column, which is a unique ID per customer, is dropped.
| Dataset (rows) | Guess the majority | Logistic regression | Boosted trees, defaults | Boosted trees, tuned |
|---|---|---|---|---|
| Census income (48,842) | 76.1% · 0.500 | 85.6% · 0.909 | 87.7% · 0.931 | 87.6% · 0.931 |
| Telecom churn (5,000) | 85.9% · 0.500 | 86.6% · 0.797 | 95.5% · 0.929 | 95.3% · 0.923 |
| Spam email (4,601) | 60.6% · 0.500 | 93.2% · 0.970 | 96.0% · 0.990 | 96.0% · 0.990 |
| Gamma telescope (19,020) | 64.8% · 0.500 | 79.3% · 0.847 | 89.0% · 0.942 | 89.3% · 0.946 |
| Hypothyroid (3,163) | 95.3% · 0.500 | 96.2% · 0.917 | 98.3% · 0.989 | 98.0% · 0.983 |
| Phoneme (5,404) | 70.7% · 0.500 | 74.0% · 0.803 | 90.1% · 0.957 | 92.0% · 0.967 |
Each cell shows test accuracy, then ROC AUC. The majority guess scores 0.500 AUC by definition, because it gives every row the same score. Bold is the best AUC in the row, after rounding. The default trees fit in under a second on every dataset. The search took between 43 and 152 seconds per dataset on a 4-core machine. Run with scikit-learn 1.9.1.
Each rung of that ladder answers a different question, and the rung below it supplies the number you need to answer it.
What the bottom rung tells you
Look at the accuracy column for hypothyroid. A model that has never looked at a patient, and says "no" every time, is 95.3% accurate. That's what 4.8% prevalence does to accuracy. The logistic regression's 96.2% sounds strong on its own, and it's less than one point above doing nothing. Telecom churn is the same: 85.9% for predicting that nobody leaves, 86.6% for the logistic regression. Report "our churn model is 87% accurate" in a meeting and it sounds like a result. Next to the majority-class guess, it's close to nothing.
This is the cheapest check in applied ML, and it catches the most embarrassing failures before anyone else does. It also shows that accuracy is the wrong headline metric for imbalanced problems like these two. That's a larger topic, but you only notice the problem once you have the floor.
What the middle rung tells you
The jump from logistic regression to default boosted trees is where the real decisions are, and the table shows two very different situations.
On telecom churn, phoneme and the gamma telescope data, the trees add 9 to 15 points of AUC (a point here is 0.01). Hypothyroid is less settled: 7 points on this split, but anywhere from 2 to 8 across the four splits described below. The gap comes from interactions and non-linear thresholds that a plain linear model can't represent. Some of it could be won back with feature engineering, such as binning the number of customer-service calls in the churn data, so treat this logistic regression as a floor, not the best simple model possible. Even so, on those datasets the trees are easily worth the loss of interpretability. On census income and spam, the gain is about 2 points. Whether 2 points is worth giving up coefficients you can read out to a risk committee, a regulator or a customer is a business decision. It might well be. But you can only make that decision if you have the logistic regression number. Start with the ensemble and the question never comes up.
What the top rung tells you
The tuned search is the most expensive step by two orders of magnitude, and on this run it was the least informative. It improved AUC by more than half a point on one dataset of six (phoneme, +1.0). On census income and spam it changed nothing. On churn and hypothyroid the tuned model scored slightly worse on the test set than the defaults. That's most likely the search fitting noise in the cross-validation folds, which is easy to do when you pick the best of 40 configurations scored on a few thousand rows.
This is not a claim that tuning never matters. Phoneme gained a real point, and on other data it can matter a lot. McElfresh et al., comparing 19 algorithms across 176 datasets, found that "for a surprisingly high number of datasets, either the performance difference between GBDTs and NNs is negligible, or light hyperparameter tuning on a GBDT is more important than choosing between NNs and GBDTs." Note what that compares: light tuning against switching model families, not tuning against good defaults. On these six datasets, the defaults were already close to what a 200-fit search found. The only way to know which situation your dataset is in is to have the untuned number to compare against. Without it, the search's best score looks like something the search earned.
A caveat that works in the argument's favour: the test sets here range from 633 rows to about 9,800. A half-point AUC difference on a thousand held-out rows is within the variation you'd get from a different random split. Re-running the whole comparison on three more random splits didn't change the picture. Phoneme gained about a point from tuning every time. On the other five datasets the tuning gain averaged between −0.3 and +0.2 points across the four splits. On four of them it landed on both sides of zero from one split to the next. Gamma telescope gained every time, but only 0.1 to 0.4 points. The honest reading of the tuned column is "indistinguishable from defaults" on most rows, which is exactly the thing you can't see without a default to compare.
The research literature keeps getting embarrassed the same way
None of this is new. Whole subfields have found out, sometimes years late, that their sophisticated methods hadn't been compared against a properly built simple one.
In recommender systems, Ferrari Dacrema, Cremonesi and Jannach tried to reproduce 18 neural recommendation algorithms from top conferences in 2019. They could reproduce 7 with reasonable effort, and 6 of those 7 could often be beaten by simple nearest-neighbour or graph-based heuristics. The paper won best paper at RecSys that year. The problem hasn't gone away. A May 2026 study (Han et al.) tested an "intentionally simple graph heuristic" against modern generative recommenders. It starts from only the last one or two items a user interacted with and uses no training at all. It "matches or outperforms many modern baselines, with relative NDCG@10 improvements of 38.10% and 44.18% over the best competing baseline" on two widely used Amazon review datasets, and stays competitive on 10 of the 14 datasets tested.
In time-series forecasting, Zeng et al. introduced what they called "a set of embarrassingly simple one-layer linear models" as a sanity check against the Transformer architectures that had dominated long-horizon forecasting research. Their result: "LTSF-Linear surprisingly outperforms existing sophisticated Transformer-based LTSF models in all cases, and often by a large margin." Later Transformer designs such as PatchTST did beat those linear models, which is the point: once a simple baseline existed, the bar went up.
In clinical prediction, a systematic review by Christodoulou et al. (Journal of Clinical Epidemiology, 2019) pooled 282 comparisons between logistic regression and machine-learning models from 71 studies. Among the 145 comparisons at low risk of bias, the difference in discrimination between the two was 0.00 on the logit(AUC) scale (95% CI −0.18 to 0.18). A machine-learning advantage only appeared in the 137 comparisons at high risk of bias. Their conclusion: "We found no evidence of superior performance of ML over LR."
And on ordinary tabular data, Grinsztajn, Oyallon and Varoquaux benchmarked deep learning against tree ensembles on 45 datasets and found that "tree-based models remain state-of-the-art on medium-sized data (∼10K samples) even without accounting for their superior speed", which is why trees, not neural networks, are rung three above.
The common thread isn't "simple models are better". Sometimes they're not, as the churn row shows. It's that the complicated method was never properly compared against the simple one. In a research paper that costs credibility. In a company it costs a production system that's harder to explain, slower to retrain and more fragile than it needed to be, delivering an improvement nobody measured.
Reproducing the bottom three rungs
Here are the first three rungs for the census-income dataset. It needs pip install pmlb scikit-learn and runs in a few seconds. It reproduces the table's numbers on scikit-learn 1.9.1; other versions can differ in the third decimal.
from pmlb import fetch_data
from sklearn.compose import make_column_transformer
from sklearn.dummy import DummyClassifier
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
df = fetch_data("adult")
y = (df.pop("target") == 0).astype(int) # 1 = income above $50K
cats = ["workclass", "education", "marital-status", "occupation",
"relationship", "race", "sex", "native-country"]
nums = [c for c in df.columns if c not in cats]
X_tr, X_te, y_tr, y_te = train_test_split(
df, y, test_size=0.2, stratify=y, random_state=0)
ladder = {
"guess the majority": DummyClassifier(strategy="most_frequent"),
"logistic regression": make_pipeline(
make_column_transformer(
(OneHotEncoder(handle_unknown="ignore"), cats),
(StandardScaler(), nums)),
LogisticRegression(max_iter=2000)),
"boosted trees, defaults": HistGradientBoostingClassifier(
categorical_features=[c in cats for c in df.columns], random_state=0),
}
for name, model in ladder.items():
model.fit(X_tr, y_tr)
auc = roc_auc_score(y_te, model.predict_proba(X_te)[:, 1])
print(f"{name:24} accuracy {model.score(X_te, y_te):.3f} AUC {auc:.3f}")
Swap in your own data and you have the three numbers every later model has to be reported against. If you're working with an assistant, ask for the floor explicitly, because "build a model" won't reliably produce it. Something like: "Before any other model, fit a majority-class DummyClassifier and a plain logistic regression on the same split, and report every later model as an improvement over both." That's a one-sentence addition to the prompt, and it changes what the rest of the notebook can prove.
The list of impressive first moves keeps growing. After gradient boosting came tabular deep learning, and now pretrained tabular foundation models, which TabArena, a continuously maintained benchmark, reports "excel on smaller datasets." Each one is worth trying. Each one also makes it easier to skip the floor, because the first result already looks like a finished product.
Two lines in every model report
The fix is a reporting habit, not a technique. Every model summary, whether it's a notebook, a slide or a pull request description, should include two numbers next to the headline score: the improvement over the majority-class guess, and the improvement over the simplest reasonable model. If the first number is small, you don't have a model yet. If the second is small, you have a decision to make about whether the complexity is worth keeping, and now you have what you need to make it.
It's the same judgement the tuning lesson in SophiArch's Introduction to Machine Learning course is built around: it starts from untouched defaults and gives a whole section to deciding when to stop. The table above is that decision with real numbers attached.
The embarrassing model isn't the one you ship. It's the one that tells you whether the model you ship was worth building.
References
- Olson, R. S., La Cava, W., Orzechowski, P., Urbanowicz, R. J., & Moore, J. H. (2017). PMLB: a large benchmark suite for machine learning evaluation and comparison. BioData Mining, 10, 36. Source of the six datasets; the data is at github.com/EpistasisLab/pmlb.
- scikit-learn.
DummyClassifier. The majority-class baseline used as rung one; quoted wording checked against the docstring in scikit-learn 1.9.1. - McElfresh, D., Khandagale, S., Valverde, J., Prasad C, V., Feuer, B., Hegde, C., Ramakrishnan, G., Goldblum, M., & White, C. (2023). When Do Neural Nets Outperform Boosted Trees on Tabular Data? NeurIPS 2023 Datasets and Benchmarks Track. 19 algorithms, 176 datasets; source of the light-tuning finding.
- Ferrari Dacrema, M., Cremonesi, P., & Jannach, D. (2019). Are We Really Making Much Progress? A Worrying Analysis of Recent Neural Recommendation Approaches. RecSys '19. 7 of 18 algorithms reproducible, 6 of those often beaten by simple heuristics.
- Han, H., Ma, L., Wang, H., Li, B., Zha, D., et al. (2026). An Embarrassingly Simple Graph Heuristic Reveals Shortcut-Solvable Benchmarks for Sequential Recommendation. arXiv:2605.07125. Simple graph heuristic competitive on 10 of 14 datasets.
- Zeng, A., Chen, M., Zhang, L., & Xu, Q. (2023). Are Transformers Effective for Time Series Forecasting? AAAI 2023. Source of the one-layer linear (LTSF-Linear) result.
- Nie, Y., Nguyen, N. H., Sinthong, P., & Kalagnanam, J. (2023). A Time Series is Worth 64 Words: Long-term Forecasting with Transformers. ICLR 2023. PatchTST, a Transformer design that beat the LTSF-Linear baselines.
- Christodoulou, E., Ma, J., Collins, G. S., Steyerberg, E. W., Verbakel, J. Y., & Van Calster, B. (2019). A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models. Journal of Clinical Epidemiology, 110, 12–22.
- Grinsztajn, L., Oyallon, E., & Varoquaux, G. (2022). Why do tree-based models still outperform deep learning on tabular data? NeurIPS 2022 Datasets and Benchmarks Track. 45-dataset benchmark of trees against deep learning.
- Erickson, N., Purucker, L., Tschalzev, A., Holzmüller, D., et al. (2025). TabArena: A Living Benchmark for Machine Learning on Tabular Data. arXiv:2506.16791. Continuously maintained tabular benchmark; source of the foundation-model finding.
- SophiArch. Introduction to Machine Learning. Course covering supervised learning fundamentals, model evaluation metrics, and a hyperparameter tuning workflow that starts from defaults and includes explicit stopping criteria.
Top comments (0)