DEV Community

Cover image for Can AI Be Trusted to Code Financial Software? I Built a Benchmark to Find Out
Kesavaram Rathnasingam
Kesavaram Rathnasingam

Posted on

Can AI Be Trusted to Code Financial Software? I Built a Benchmark to Find Out

kagglechallenge

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked
BankSafe-Bench: Can AI Safely Code Financial Software?

I built BankSafe-Bench, a benchmark designed to evaluate how well AI coding models handle realistic financial software engineering tasks.

Instead of testing models with generic programming questions, the benchmark focuses on situations where “code that looks correct” is not necessarily correct.

The benchmark covers tasks across:

C# / ASP.NET Core backend development
SQL Server queries and data operations
Financial calculations and business rules
Edge-case handling and validation
Database transactions and consistency
Concurrency and failure scenarios
Secure coding practices
Legacy-code modification
Debugging and code review

For example, a task may provide a financial calculation with several business rules and edge cases and ask the model to implement it correctly. Another task may present a multi-step database operation and ask the model to identify what could happen if one operation fails.

The goal is to measure something more useful than whether an AI can generate code:

Can an AI coding model produce software that is actually safe and reliable when financial correctness matters?

This interests me because modern AI coding assistants are increasingly used for production software development, while financial applications have requirements where small mistakes in calculations, authorization, transactions, or data handling can have significant consequences.

Models Tested

I selected models representing different approaches to AI-assisted software engineering, including:

A leading general-purpose model
A reasoning-focused model
A coding-focused model
A strong open-weight model
A smaller/efficiency-focused model

The models were evaluated using the same benchmark tasks and scoring criteria so that their performance could be compared on identical problems.

The benchmark does not assume that the largest or most popular model will perform best. The goal is to discover which capabilities actually matter for financial software engineering.

Findings

The benchmark is designed to evaluate models across several dimensions rather than using a single “correct/incorrect” score.

Each task is scored across:

Category Weight
Functional correctness 30%
Financial/business-rule correctness 25%
Edge-case handling 15%
Security 15%
Code quality 10%
Explanation 5%

The most interesting part of this benchmark is the difference between code generation ability and engineering reliability.

I will report:

Overall model scores
Performance by task category
Financial-rule accuracy
Security failures
Edge cases missed
Transaction/concurrency failures
Examples of seemingly correct solutions that fail under realistic conditions
What surprised me

Rather than focusing only on which model achieves the highest overall score, I want to investigate where models fail.

For example:

Can a model produce code that compiles but violates a financial rule?

Can it recognize a transaction-consistency problem that isn't obvious from the happy path?

Can it distinguish a technically valid SQL query from one that produces incorrect financial results?

These failures are potentially more important than simple syntax or compilation errors.

What I would measure next

Future versions could expand the benchmark to include:

Multi-turn debugging
Realistic API integration
Codebase-level changes
Automated test generation
Performance optimization
Database deadlock diagnosis
Authentication and authorization
AI-generated code review
Agentic coding workflows

The long-term goal is to develop a benchmark that measures engineering reliability, not just code-generation capability.

My Benchmark

Kaggle Benchmark:
BankSafe-Bench — Can AI Safely Code Financial Software?

The benchmark is designed to be reproducible so that additional models and new versions of existing models can be evaluated as AI coding capabilities evolve.

Top comments (0)