DEV Community

Cover image for Ask an AI for your home loan EMI and it will be off by ₹10. Every month.
Reet Singh
Reet Singh

Posted on

Ask an AI for your home loan EMI and it will be off by ₹10. Every month.

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

Last week I built a little app about Ananya. She's made up, but you know someone like her. She's 31, she's been saving for eight years, and she's about to sign for a ₹72 lakh 2 BHK whose brochure says 1,050 sq ft. The registration says the carpet area is 720.

While I was building it I kept asking chatbots for help with the numbers: what's ₹72 lakh in crore, what's the price per sq ft of carpet, what's the EMI on the loan. The answers sounded confident. I had no idea whether they were right.

So I benchmarked it. Then my benchmark turned out to be too easy, so I made it harder. That second part is where the real findings are.

What I Benchmarked

Lakh, Crore and Carpet Area is 100 deterministic questions about the money and property maths Indian buyers actually do. Every answer was computed in Python, not by hand.

Tier Group Cases Example
Core Conversion 15 "₹1.2 crore minus ₹45 lakh, in lakh?"
Core Area 15 "Super built-up 1,300 sq ft, carpet 910. Loading as a % of carpet?"
Core EMI and tax 15 "₹72,00,000 at 8.4% for 25 years. Monthly EMI?"
Core Format 15 "Write 98765432 with lakh/crore commas" → 9,87,65,432
Hard Multi-step loans 15 prepay ₹5 lakh after 36 EMIs, keep the EMI: how many EMIs are left?
Hard Tax rules 7 80C and 24(b) caps, LTCG with cess, circle rate vs agreement price
Hard Regional land units 8 guntha, cent, Chennai ground, bigha-biswa, Punjab vs Haryana marla
Hard Hinglish 10 "72 lakh ka flat hai, 20% down payment, baaki 8.5% pe 20 saal. EMI kitni?"

Each prompt asks for at most five lines of working, then ANSWER: <value>. Numbers are checked against the true value (a rupee or two of slack where rounding is involved, an exact string match for the format group). Tax questions state the rule inside the prompt, so they test arithmetic, not whether a model knows this year's tax law. The scorer also flags any response that slips into Western grouping (7,200,000 instead of 72,00,000), even when the final answer is right. I call that western comma drift.

Why this matters to me: hundreds of millions of people think in lakh and crore, and property listings in India are famous for blurring carpet and super built-up area. If a model is going to sit inside a home-buying assistant, these are its first questions, not edge cases.

Models Tested

A spread you'd actually reach for in a product, across three providers and two price tiers: Gemini 3.7 Flash, Gemini 3.1 Pro Preview, Claude Haiku 5.5, Claude Sonnet 4.6, Gemma 4 31B (open-weight) and GPT-5.4 mini.

Findings

1. My first benchmark was too easy

Version 1 was just the 60 core questions. Every model scored between 0.92 and 1.00:

Model Core score (60)
Gemini 3.7 Flash 1.00
Gemini 3.1 Pro Preview 0.98
Claude Haiku 5.5 0.98
Claude Sonnet 4.6 0.95
Gemma 4 31B 0.95
GPT-5.4 mini 0.92

A benchmark where everyone gets an A tells you nothing about which model to ship. So I added 40 harder cases, the kind of questions a buyer asks after the first EMI: prepayments, rate hikes, balance transfers, tax caps, regional land units, and the same maths in Hinglish.

2. The hard tier opened a 35-point gap

Hard tier results

Model Score (100) Hard tier (40) Multi-step loans Tax rules Regional units Hinglish Comma drift
Claude Haiku 5.5 0.97 37/40 12/15 7/7 8/8 10/10 16%
Gemini 3.7 Flash 0.94 34/40 10/15 6/7 8/8 10/10 8%
Claude Sonnet 4.6 0.85 29/40 7/15 7/7 7/8 8/10 13%
GPT-5.4 mini 0.77 23/40 2/15 7/7 7/8 7/10 6%

On the core 60, these four were 8 points apart. On the hard 40, they're 35 points apart, and the ranking flipped: Flash, the v1 winner, is now second.

Gemini 3.1 Pro Preview and Gemma 4 31B runs on the v3 task were still finishing on Kaggle when I updated this post. Their scores will be on the leaderboard linked below.

3. Multi-step loans break everyone

Models that nailed single EMIs still stumbled once the loan had a history. Even the best model missed 3 of 15. GPT-5.4 mini got 2 of 15.

The scariest miss was l13: "You can afford an EMI of ₹50,000 at 8.5% for 20 years. What's the largest loan you can take?" The answer is about ₹57.6 lakh. GPT-5.4 mini said 537,000. That's roughly ₹5.4 lakh, a tenth of the real figure, delivered with the same confidence as everything else.

The near-misses are worse in a quieter way. On e04 (₹72,00,000 at 8.4% for 25 years) the correct EMI is ₹57,492. In v1, Claude Sonnet 4.6 said ₹57,487 and Gemini 3.1 Pro Preview said ₹57,482. None of them said "approximately". (1 + r)^300 is exactly where mental arithmetic slips.

What this changed for me: in anything I ship now, the model explains the formula and a tool computes the EMI.

4. The bigger model lost to the cheaper one

Claude Haiku 5.5, the small tier, beat Claude Sonnet 4.6 by 12 points overall, and by 8 cases on the hard tier. Price tier was a worse predictor of accuracy than I expected, so test the model you actually plan to pay for.

5. What I thought would be hard, wasn't

I predicted regional land units would be the bloodbath, since bigha and marla change meaning from state to state. But when the conversion is given, models handle it: 30 of 32 right across the four models. Tax rules stated in the prompt were nearly perfect too (27 of 28). The only tax miss was Gemini 3.7 Flash on t02, which forgot the buyer had already used ₹90,000 of their 80C limit and answered ₹3,30,000 instead of ₹2,40,000.

The difficulty is in chaining steps, not in unfamiliar units.

6. Hinglish costs accuracy for some models

Haiku and Flash got all 10 Hinglish questions right. Sonnet dropped 2 and GPT-5.4 mini dropped 3, on sums like total interest and carpet-from-loading that are routine in English. If your users type "EMI kitni banegi?", test in that language.

7. Models drift into Western commas even in an Indian question

Every model aced the format group when asked directly. Yet 6–16% of responses still wrote something like 4,000,000 somewhere in their working. It doesn't make the answer wrong, but it's exactly the slip that makes "₹40 lakh" read as "₹4 crore" to a tired buyer at 11 PM.

8. My benchmark lied to me first

My first run said Claude Opus 5 scored 0.08, and GPT-5.5 errored completely. I almost wrote a post about it.

Then I opened the run logs. Every failure was a 403 PermissionDeniedError: max estimated cost exceeds your available quota. I'd fired cases in parallel, the expensive models blew through the quota, and my scorer counted each errored call as a wrong answer. Opus 5's printed table actually showed 100% on every group it got to answer.

I fixed it: sequential calls, per-case progress logging, and every run prints answered: X of 100 before any score. I dropped the two most expensive models rather than publish a number I can't stand behind.

The lesson: a benchmark score is only as honest as its error handling. If a leaderboard can't tell you how many questions were actually answered, don't trust the leaderboard.

What I'd measure next

  • Give the models a calculator tool. Do the multi-step loan errors vanish, or do models skip the tool when they feel confident?
  • Leave the unit definition out. Ask "2 bigha in sq ft?" with no state, and score whether the model asks which state instead of guessing.

My Benchmark

Fork it, add your city's stamp duty, and tell me which model gets your home loan right.

Top comments (0)