DEV Community

Cover image for AI in Underwriting: What Regulators Actually Care About
Emmanuel R for CobuildX AI

Posted on Originally published at cobuildx.ai

AI in Underwriting: What Regulators Actually Care About

Lenders deploying AI in underwriting face a compliance landscape that shapes what models can be used, what documentation is required, and how adverse action notices must work. Here is what the regulatory framework actually requires and what you can deploy today.

AI underwriting models can assess creditworthiness faster, more consistently, and — when built carefully — more accurately than traditional scorecards. Lenders know this. The compliance and legal teams know it too, and they have a set of questions that need answers before any AI model touches a credit decision. Those questions are not obstacles to deploying AI in underwriting. They are a framework that, once understood, makes it straightforward to separate what is deployable today from what requires more work.

Most AI underwriting projects stall not because the model does not work but because the compliance path was not scoped upfront. Model development, validation, and business user testing proceed for six to nine months — and then the model sits in legal review because the explainability requirements, the adverse action notice process, and the fair lending testing were not designed in from the start.

Key insight: Regulatory requirements for AI in underwriting are not ambiguous. ECOA, the Fair Housing Act, SR 11-7 model risk management guidance, and the relevant CFPB guidance give a reasonably clear picture of what is required. The compliance path is knowable upfront — the projects that get stuck are the ones that did not engage it early.

"The model was excellent. We spent four months in legal trying to figure out how to write the adverse action notice."

The Adverse Action Notice Requirement

Under ECOA and Regulation B, lenders must provide applicants who are denied credit — or offered less favourable terms — with a written statement of the specific reasons for the adverse action. This requirement applies regardless of what type of model made the decision. A black-box neural network that cannot produce interpretable reason codes does not satisfy the adverse action notice requirement.

For traditional scorecard models, adverse action reasons are generated directly from the score factors — the variables that contributed most negatively to the applicant's score. For complex ML models, producing equivalent reason codes requires either post-hoc explanation methods (SHAP values, LIME) or designing the model architecture to produce interpretable outputs.

SHAP-based reason codes are increasingly accepted by compliance teams as a defensible approach, but they require validation. The reason codes need to be consistent — the same input should produce the same reason codes across runs — and they need to reflect the actual model mechanics rather than being post-hoc rationalisations that do not match the model's actual decision logic. Both of these properties require testing.

Adverse action notices require specific reasons for the decision — post-hoc explanation methods like SHAP are increasingly accepted but need to be validated for consistency and accuracy

Fair Lending: What Testing Is Required

ECOA and the Fair Housing Act prohibit discrimination in credit decisions on the basis of race, colour, national origin, sex, religion, marital status, age, and other protected characteristics. This applies to AI models.

The fair lending testing requirement for AI models is not that the model cannot use protected characteristics as inputs — that is a floor, not the full requirement. The actual requirement is disparate impact analysis: does the model produce materially different approval rates or pricing outcomes across protected groups that cannot be justified by legitimate credit risk differences?

Fair lending testing should be run during model development, not after. If disparate impact is detected during development, the model can be adjusted — through feature selection, recalibration, or architectural changes — before deployment. If it is detected during regulatory examination after deployment, the remediation is more expensive and the reputational risk is significant.

The specific statistical tests used for disparate impact analysis in lending — the 80% rule, regression-based control methods, the Bayesian Improved Surname Geocoding (BISG) proxy methodology for race and ethnicity — are well-documented. Fair lending review should be designed into the model validation process, not added as a final gate.

Fair lending testing belongs in the development process, not at the end of it — detecting disparate impact early is cheaper and safer than detecting it in an exam

SR 11-7 Model Risk Management

SR 11-7 is the Federal Reserve and OCC's guidance on model risk management, and it applies to AI underwriting models at banks and other supervised institutions. The core requirements: models must be validated by a function independent of the model development team, documentation must cover model purpose, methodology, data, limitations, and performance metrics, and models must be subject to ongoing monitoring and periodic revalidation.

For AI models specifically, SR 11-7 requires that the model validation team understand the model well enough to challenge it. This is harder for complex ML models than for traditional scorecards. Validators need to understand the training data, the feature engineering, the hyperparameter selection process, and the performance metrics used to evaluate the model — and document that they have assessed each for soundness.

The practical implication: complex ML models require more documentation and a more sophisticated validation team than traditional models. This is not a reason to avoid them, but it is a reason to involve model risk management early in the development process rather than presenting them with a finished model for validation.

SR 11-7 validation requires the ability to challenge the model — complex ML models require more documentation and more sophisticated validators than traditional scorecards

What You Can Deploy Today

Given this framework, the question is which AI approaches fit within the compliance requirements without requiring regulatory engagement beyond standard model validation.

Gradient boosting models (XGBoost, LightGBM) with SHAP-based reason codes are the current sweet spot for lenders who want meaningfully better predictive performance than traditional logistic regression scorecards without taking on interpretability risk they cannot manage. These models are significantly more predictive than scorecards on most credit datasets, produce consistent SHAP explanations, and have a validation pathway that most model risk management teams can work with.

Logistic regression models with well-engineered features remain a defensible choice for regulated lenders who want maximum interpretability and a simpler validation process. The performance gap versus gradient boosting depends on the dataset, but for standard consumer lending use cases the gap is often smaller than the implementation complexity difference.

Deep learning models and LLM-based approaches require more care. They can produce excellent predictive performance, but the interpretability pathway, the validation approach, and the adverse action reason code methodology all require more work to make defensible. They are deployable, but they require more upfront investment in the compliance infrastructure.

Gradient boosting with SHAP reason codes is the current practical standard for regulated lenders wanting ML performance without interpretability complexity


Originally published on the CobuildX blog.

Top comments (0)