<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Kesavaram Rathnasingam</title>
    <description>The latest articles on DEV Community by Kesavaram Rathnasingam (@kesavaram_rathnasingam_5c).</description>
    <link>https://dev.to/kesavaram_rathnasingam_5c</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3587355%2Fafaff36c-5f1e-4ab2-83d0-b0535a9be271.png</url>
      <title>DEV Community: Kesavaram Rathnasingam</title>
      <link>https://dev.to/kesavaram_rathnasingam_5c</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kesavaram_rathnasingam_5c"/>
    <language>en</language>
    <item>
      <title>Can AI Be Trusted to Code Financial Software? I Built a Benchmark to Find Out</title>
      <dc:creator>Kesavaram Rathnasingam</dc:creator>
      <pubDate>Thu, 08 Oct 2026 04:10:23 +0000</pubDate>
      <link>https://dev.to/kesavaram_rathnasingam_5c/can-ai-be-trusted-to-code-financial-software-i-built-a-benchmark-to-find-out-2g60</link>
      <guid>https://dev.to/kesavaram_rathnasingam_5c/can-ai-be-trusted-to-code-financial-software-i-built-a-benchmark-to-find-out-2g60</guid>
      <description>&lt;p&gt;This is a submission for the Kaggle Benchmarking Challenge&lt;/p&gt;

&lt;p&gt;What I Benchmarked&lt;br&gt;
BankSafe-Bench: Can AI Safely Code Financial Software?&lt;/p&gt;

&lt;p&gt;I built BankSafe-Bench, a benchmark designed to evaluate how well AI coding models handle realistic financial software engineering tasks.&lt;/p&gt;

&lt;p&gt;Instead of testing models with generic programming questions, the benchmark focuses on situations where “code that looks correct” is not necessarily correct.&lt;/p&gt;

&lt;p&gt;The benchmark covers tasks across:&lt;/p&gt;

&lt;p&gt;C# / ASP.NET Core backend development&lt;br&gt;
SQL Server queries and data operations&lt;br&gt;
Financial calculations and business rules&lt;br&gt;
Edge-case handling and validation&lt;br&gt;
Database transactions and consistency&lt;br&gt;
Concurrency and failure scenarios&lt;br&gt;
Secure coding practices&lt;br&gt;
Legacy-code modification&lt;br&gt;
Debugging and code review&lt;/p&gt;

&lt;p&gt;For example, a task may provide a financial calculation with several business rules and edge cases and ask the model to implement it correctly. Another task may present a multi-step database operation and ask the model to identify what could happen if one operation fails.&lt;/p&gt;

&lt;p&gt;The goal is to measure something more useful than whether an AI can generate code:&lt;/p&gt;

&lt;p&gt;Can an AI coding model produce software that is actually safe and reliable when financial correctness matters?&lt;/p&gt;

&lt;p&gt;This interests me because modern AI coding assistants are increasingly used for production software development, while financial applications have requirements where small mistakes in calculations, authorization, transactions, or data handling can have significant consequences.&lt;/p&gt;

&lt;p&gt;Models Tested&lt;/p&gt;

&lt;p&gt;I selected models representing different approaches to AI-assisted software engineering, including:&lt;/p&gt;

&lt;p&gt;A leading general-purpose model&lt;br&gt;
A reasoning-focused model&lt;br&gt;
A coding-focused model&lt;br&gt;
A strong open-weight model&lt;br&gt;
A smaller/efficiency-focused model&lt;/p&gt;

&lt;p&gt;The models were evaluated using the same benchmark tasks and scoring criteria so that their performance could be compared on identical problems.&lt;/p&gt;

&lt;p&gt;The benchmark does not assume that the largest or most popular model will perform best. The goal is to discover which capabilities actually matter for financial software engineering.&lt;/p&gt;

&lt;p&gt;Findings&lt;/p&gt;

&lt;p&gt;The benchmark is designed to evaluate models across several dimensions rather than using a single “correct/incorrect” score.&lt;/p&gt;

&lt;p&gt;Each task is scored across:&lt;/p&gt;

&lt;p&gt;Category    Weight&lt;br&gt;
Functional correctness  30%&lt;br&gt;
Financial/business-rule correctness 25%&lt;br&gt;
Edge-case handling  15%&lt;br&gt;
Security    15%&lt;br&gt;
Code quality    10%&lt;br&gt;
Explanation 5%&lt;/p&gt;

&lt;p&gt;The most interesting part of this benchmark is the difference between code generation ability and engineering reliability.&lt;/p&gt;

&lt;p&gt;I will report:&lt;/p&gt;

&lt;p&gt;Overall model scores&lt;br&gt;
Performance by task category&lt;br&gt;
Financial-rule accuracy&lt;br&gt;
Security failures&lt;br&gt;
Edge cases missed&lt;br&gt;
Transaction/concurrency failures&lt;br&gt;
Examples of seemingly correct solutions that fail under realistic conditions&lt;br&gt;
What surprised me&lt;/p&gt;

&lt;p&gt;Rather than focusing only on which model achieves the highest overall score, I want to investigate where models fail.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;Can a model produce code that compiles but violates a financial rule?&lt;/p&gt;

&lt;p&gt;Can it recognize a transaction-consistency problem that isn't obvious from the happy path?&lt;/p&gt;

&lt;p&gt;Can it distinguish a technically valid SQL query from one that produces incorrect financial results?&lt;/p&gt;

&lt;p&gt;These failures are potentially more important than simple syntax or compilation errors.&lt;/p&gt;

&lt;p&gt;What I would measure next&lt;/p&gt;

&lt;p&gt;Future versions could expand the benchmark to include:&lt;/p&gt;

&lt;p&gt;Multi-turn debugging&lt;br&gt;
Realistic API integration&lt;br&gt;
Codebase-level changes&lt;br&gt;
Automated test generation&lt;br&gt;
Performance optimization&lt;br&gt;
Database deadlock diagnosis&lt;br&gt;
Authentication and authorization&lt;br&gt;
AI-generated code review&lt;br&gt;
Agentic coding workflows&lt;/p&gt;

&lt;p&gt;The long-term goal is to develop a benchmark that measures engineering reliability, not just code-generation capability.&lt;/p&gt;

&lt;p&gt;My Benchmark&lt;/p&gt;

&lt;p&gt;Kaggle Benchmark:&lt;br&gt;
BankSafe-Bench — Can AI Safely Code Financial Software?&lt;/p&gt;

&lt;p&gt;The benchmark is designed to be reproducible so that additional models and new versions of existing models can be evaluated as AI coding capabilities evolve.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>LankaLens — AI Nature Explorer</title>
      <dc:creator>Kesavaram Rathnasingam</dc:creator>
      <pubDate>Thu, 08 Oct 2026 04:06:30 +0000</pubDate>
      <link>https://dev.to/kesavaram_rathnasingam_5c/lankalens-ai-nature-explorer-mjc</link>
      <guid>https://dev.to/kesavaram_rathnasingam_5c/lankalens-ai-nature-explorer-mjc</guid>
      <description>&lt;p&gt;This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass&lt;/p&gt;

&lt;p&gt;What I Built&lt;br&gt;
🌿 LankaLens — AI Nature Explorer&lt;/p&gt;

&lt;p&gt;I built LankaLens, a mobile application that encourages people to put their phones down and explore the nature around them.&lt;/p&gt;

&lt;p&gt;The idea is simple: when you're walking, hiking, gardening, or exploring outdoors, you can take a photo of a plant, tree, flower, insect, or other part of nature. LankaLens uses open-source AI to help identify and explain what you found.&lt;/p&gt;

&lt;p&gt;But the goal isn't to keep people looking at the app.&lt;/p&gt;

&lt;p&gt;The app gives you a short explanation and an observation challenge, such as:&lt;/p&gt;

&lt;p&gt;"Look around you and see if you can find another tree with similar leaves."&lt;/p&gt;

&lt;p&gt;Then the phone goes back into your pocket.&lt;/p&gt;

&lt;p&gt;I designed it for people who are curious about nature but don't necessarily know the names of the things they see around them.&lt;/p&gt;

&lt;p&gt;The longer-term goal is to make it particularly useful for exploring Sri Lankan plants, trees, birds, insects, and local biodiversity, with support for English, Sinhala, and Tamil.&lt;/p&gt;

&lt;p&gt;Demo&lt;/p&gt;

&lt;p&gt;🎥 Demo video: [Coming soon]&lt;/p&gt;

&lt;p&gt;The demo will show the application being used outdoors with the device disconnected from the internet, demonstrating the local-first AI workflow.&lt;/p&gt;

&lt;p&gt;Code&lt;/p&gt;

&lt;p&gt;💻 GitHub: [Coming soon]&lt;/p&gt;

&lt;p&gt;The project will be open source so others can experiment with the models, improve the identification pipeline, add local species knowledge, and adapt it to their own regions.&lt;/p&gt;

&lt;p&gt;How I Built It&lt;/p&gt;

&lt;p&gt;The application is built around the idea that AI should work for the outdoor experience rather than become the experience.&lt;/p&gt;

&lt;p&gt;The mobile application handles:&lt;/p&gt;

&lt;p&gt;📷 Camera-based nature observations&lt;br&gt;
📍 Location-aware exploration&lt;br&gt;
💾 Offline storage&lt;br&gt;
🌿 Nature observations and history&lt;br&gt;
🧭 Small outdoor exploration challenges&lt;br&gt;
🌐 English, Sinhala, and Tamil support&lt;/p&gt;

&lt;p&gt;At the core is an open-weight AI model running locally rather than sending every photograph to a proprietary cloud AI API.&lt;/p&gt;

&lt;p&gt;The architecture is designed so that the AI model can be replaced or upgraded without rebuilding the entire application.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;         Mobile App
             │
    ┌────────┼────────┐
    │        │        │
 Camera     GPS    Offline DB
    │
    ▼
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Local AI Inference&lt;br&gt;
        │&lt;br&gt;
        ▼&lt;br&gt;
  Open-Weight Model&lt;br&gt;
        │&lt;br&gt;
        ▼&lt;br&gt;
 Nature Identification&lt;br&gt;
        │&lt;br&gt;
        ▼&lt;br&gt;
 Short Explanation&lt;br&gt;
        │&lt;br&gt;
        ▼&lt;br&gt;
 Exploration Challenge&lt;br&gt;
        │&lt;br&gt;
        ▼&lt;br&gt;
      📱 → Pocket&lt;br&gt;
      🌳 → Explore&lt;br&gt;
Why Does Open Innovation Matter?&lt;/p&gt;

&lt;p&gt;For this project, open AI isn't just a technology choice. It changes what the application can do.&lt;/p&gt;

&lt;p&gt;Nature exploration often happens in places where mobile connectivity is unreliable or unavailable. A cloud-only AI application would make the experience dependent on an internet connection and an external service.&lt;/p&gt;

&lt;p&gt;With an open-weight model and local inference, the goal is to make the core experience possible without an internet connection.&lt;/p&gt;

&lt;p&gt;It also means users don't have to upload every photograph they take to a third-party AI service.&lt;/p&gt;

&lt;p&gt;More importantly, an open approach gives the project room to evolve.&lt;/p&gt;

&lt;p&gt;I can experiment with different models, improve the prompts and inference pipeline, add local knowledge about Sri Lankan species, and eventually fine-tune or replace components without being locked into a single proprietary API.&lt;/p&gt;

&lt;p&gt;For this project, open innovation makes AI more accessible, private, customizable, and useful in the places where people actually go outside.&lt;/p&gt;

&lt;p&gt;My Agent Session&lt;/p&gt;

&lt;p&gt;Coming soon.&lt;/p&gt;

&lt;p&gt;I will include the DevRelay session showing how the project was designed and built.&lt;/p&gt;

&lt;p&gt;Prize Categories&lt;br&gt;
Open-Source AI&lt;br&gt;
Touch Grass&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>hf26challenge</category>
    </item>
  </channel>
</rss>
