DEV Community

Cover image for Credit Risk Scoring: The Data Problems That Haven't Changed
Emmanuel R for CobuildX AI

Posted on Originally published at cobuildx.ai

Credit Risk Scoring: The Data Problems That Haven't Changed

Better models have not solved the thin-file problem, the data quality issues in alternative data, or the validation requirements that govern what can be deployed in regulated lending. The data problems are still the hard part.

Credit risk modelling has changed significantly over the past decade. Gradient boosting has replaced logistic regression as the standard approach at most sophisticated lenders. Alternative data — cash flow data, rental payment history, utility payments — has broadened the credit-visible population. Explainability tooling has made complex models more deployable in regulated environments. What has not changed: the data quality problems that determine whether any model, regardless of sophistication, produces reliable predictions.

The limiting factor in credit risk modelling is not model architecture. Given reasonable data, a well-tuned gradient boosting model and a well-tuned logistic regression produce different but both useful predictions. The limiting factors are: the thin-file problem (a significant fraction of credit-worthy borrowers have insufficient traditional credit history to model reliably), the alternative data quality problem (much of the data that could address thin-file has its own reliability and coverage gaps), and the model validation requirements that govern what can be deployed.

Key insight: Lenders who have improved credit risk modelling in recent years have done so primarily through better data — specifically cash flow data and expanded credit bureau data — not better model architectures. The model is less important than the signal.

"We spent six months building a more sophisticated model. When we went back and cleaned the training data, the simple model performed just as well."

The Thin-File Problem

Approximately 45 million Americans have thin or no credit files — too little traditional credit history for the major bureaus to generate a reliable score. This population is not uniformly high risk. It includes recent immigrants with strong financial management histories in their home countries, young adults who have avoided debt but have stable income, and individuals who have historically used cash and community lending. A lender that cannot score them declines credit-worthy borrowers and cedes that business to competitors who can.

Traditional bureau data does not solve the thin-file problem because the thin-file is definitionally not in the bureau data. Expanding the modelling approach requires either incorporating non-traditional data sources (rental payments, utility payments, cash flow data) or accepting that a segment of the credit-worthy population will not be scoreable through traditional approaches.

45 million Americans are not scoreable through traditional bureau data — addressing the thin-file problem requires non-traditional data sources, not better models on the same data

Alternative Data: What Works and What Does Not

Cash flow data — transaction-level bank account data showing income deposits, recurring payments, and spending patterns — is the most consistently useful alternative data source for credit risk in consumer lending. It is highly predictive of ability to pay, has reasonable coverage through open banking integrations, and is relatively clean compared to other alternative data sources. Cash flow features (income consistency, average balance, payment-to-income ratios) add meaningful lift for thin-file borrowers on top of bureau data.

Rent payment history, when available and reliably reported, is also useful — a demonstrated track record of on-time rent payments is predictive of future credit performance. The coverage problem is significant: most rent payments are not reported to any bureau, and the coverage of rent payment data from alternative providers is incomplete and variable by geography.

Data sources that have shown less consistent value: social media signals, education and employment data from non-traditional sources, and 'character' signals derived from digital behaviour. These data types often introduce fair lending risk (proxies for protected characteristics appear in unexpected places) without delivering equivalent predictive lift. Several lenders have retreated from these approaches after fair lending testing revealed disparate impact.

Cash flow data is the most consistently predictive alternative data source. Character and digital behaviour signals introduce fair lending risk without reliable predictive lift

The Recency Problem

Credit risk models trained on pre-2020 data reflect credit behaviour in a low-interest-rate, low-unemployment economic environment. Models trained primarily on 2020–2022 data reflect an unprecedented economic disruption period with stimulus payments, forbearance programs, and unusual payment behaviour. Neither training period is a reliable guide to current credit risk in a normalising interest rate environment.

The practical implication: models trained on data more than two to three years old should be retested against recent performance data before being treated as reliable. Performance metrics that looked stable during model development may have degraded significantly as the economic environment shifted.

This is not a new problem — economic cycle sensitivity in credit models has been understood for decades. It became more acute after 2020 because the economic disruption was severe enough to make pre-crisis training data significantly less predictive of post-crisis behaviour. Lenders whose models date from 2018 or earlier and have not been recalibrated against post-2022 performance data are running on models whose accuracy may be materially worse than their validation metrics suggest.

Models trained on pre-2020 data should be retested against recent performance data — the economic environment shift has degraded accuracy in ways validation metrics from that period do not reflect

Model Validation in a Regulated Environment

For banks and supervised lenders, credit risk models are subject to model risk management requirements under SR 11-7. The validation requirement is not just that the model performs well by internal metrics — it is that the model has been independently validated, the validation covers conceptual soundness and outcome analysis, limitations are documented, and ongoing monitoring is in place.

For the models with the most regulatory scrutiny — models used in underwriting decisions for consumer credit at scale — the validation needs to cover fair lending analysis as a standard component. A model that validates well on predictive performance metrics but has not been tested for disparate impact is not ready for production at a regulated lender.

The practical implication for AI/ML models: more complex models require more documentation and a more sophisticated validation function. The validation team needs to understand the feature engineering, the training data, the hyperparameter selection rationale, and the limitations of the model well enough to write a challenge that addresses each. This is achievable but it requires planning — presenting a finished ML model to model risk management without having engaged them in the development process is a common reason for long validation timelines.

Engaging model risk management during development, not at the end of it, is the single most effective way to avoid long validation delays


Originally published on the CobuildX blog.

Top comments (0)