This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
I build software for Indian small businesses (POS, billing, payments), and this month I've been contributing Kestra workflow blueprints for Indian finance: GSTIN validation, UPI end-of-day close, GSTR-2B reconciliation. Every one of them exists because one small mistake costs real money:
- one mistyped character in a supplier's GSTIN and the input tax credit is at risk,
- IGST booked as CGST + SGST and the buyer can't claim it,
- "5 crore" written as 50 crore,
- a UPI payment that "went through" on the customer's phone but never reached the shop.
So I asked: can AI models handle everyday Indian business data? I built a Kaggle benchmark with 6 tasks, 16 cases each:
| Task | What the model has to do |
|---|---|
india_gstin_check_digit |
Validate a GSTIN and compute its Luhn mod 36 check character (algorithm given) |
india_gstin_from_memory |
Same, without being told the algorithm |
india_gst_tax_split |
Split GST into CGST + SGST, CGST + UTGST or IGST, with real traps: supplies to an SEZ unit in your own state, and Union Territories with and without a legislature |
india_invoice_number_rules |
Apply GST invoice-number rules (16 characters, only letters/digits/-//, unique per financial year) with look-alike traps: en dash, division slash, trailing space, 17 characters, extra leading zero |
india_lakh_crore_formats |
Indian digit grouping (12,34,56,789), lakh/crore conversions, cheque amounts in words |
india_upi_reconciliation |
Match 18–30 UPI bills against a settlement report and list the bills whose money never arrived, arrived short, or share a UPI reference |
How it's graded:
- Every case is generated from a fixed seed with an exact answer key. No LLM judge.
- Answers come back as structured output, so grading is exact.
- The GSTIN answer key is cross-checked against
python-stdnum(0 disagreements). The other answer keys use the same logic as my Kestra blueprints. - Every model gets the same answer budget. A case that errors (provider down, credit limit) is not counted as wrong. It's left blank.
Models Tested
I picked a mix to see what matters most: size, "thinking", or being open-weight.
- Frontier: GPT-5.5, Claude Sonnet 5, Gemini 3.7 Flash, Gemini 3.5 Flash
- Fast / small: GPT-5.4 nano, Claude Haiku 4.5, Gemini 3.1 Flash-Lite
- Open-weight: gpt-oss-20b, Gemma 4 31B, Qwen3 235B Instruct, DeepSeek R1
(Grok 4.6 and gpt-oss-120b were unavailable through Kaggle's model proxy while I ran this, so they're not on the board.)
Findings
Overall (average of task scores): GPT-5.5 1.00 · Gemini 3.7 Flash 1.00 · Gemini 3.5 Flash 0.99 · Claude Sonnet 5 0.98 · Gemma 4 0.97 · DeepSeek R1 0.96 · gpt-oss-20b 0.92 · Gemini 3.1 Flash-Lite 0.75 · Qwen3 235B 0.51 · Claude Haiku 4.5 0.45 · GPT-5.4 nano 0.34
1. Puducherry is a State (for GST), and three famous models disagree
This was my favourite result. Under the CGST Act, section 2(103), a Union Territory with its own legislature (Puducherry, Delhi) counts as a State. So a sale inside Puducherry is CGST + SGST. UTGST is only for UTs without a legislature, like Chandigarh or Ladakh.
- Claude Sonnet 5, DeepSeek R1 and Gemma 4 all charged UTGST on a Puducherry sale.
- Claude Haiku charged UTGST and IGST on the same sale, so it taxed it twice.
- GPT-5.5, every Gemini model, Qwen and even tiny GPT-5.4 nano got it right.
Bigger isn't uniformly better. Models fail on different local rules.
2. Checksums are a cliff, not a slope
On the GSTIN check character, every model that reasons step by step scored 100%, and every fast model scored 6–12%. That's no better than guessing. Small open models that reason (gpt-oss-20b, Gemma 4) beat bigger fast ones. In an early test, one model needed about 15,000 thinking tokens and 84 seconds for a single GSTIN. Correct isn't always cheap.
3. The SEZ trap
A supply to a Special Economic Zone unit in your own state is still an inter-state supply, so it's IGST (IGST Act, section 7(5)(b)). Claude Haiku, gpt-oss-20b and GPT-5.4 nano split it into CGST + SGST anyway.
4. 10× errors with lakh and crore
The small models made the most dangerous kind of mistake, being off by exactly one digit group:
- "5 crore" →
500000000(it's50000000) - "Rs 98,64,000 in lakh" →
986.4(it's98.64) - Indian grouping written the Western way:
8,263,425,718instead of8,26,34,25,718
5. Invisible characters
On GST invoice numbers, the look-alikes caught the small models: a trailing space, a 17-character number, and a number with an extra leading zero (a different number, so it's allowed) that they called a duplicate. All the big models were perfect here.
6. UPI: the misses are the dangerous part
On days with 18–30 UPI bills, weaker models missed bills whose money never arrived (Claude Haiku 44%, Qwen 25%). A false alarm costs a minute. A missed bill is silent lost money.
7. The best value is open
On Kaggle's score-vs-cost chart, the efficient frontier runs through the open-weight gpt-oss-20b and Gemma 4 31B. They get most of the way to the top score at a small fraction of the cost.
What I'd measure next
- GST rate schedules and HSN codes (which rate applies, not just how to split it)
- E-invoice JSON validation
- Invoices in Hindi and regional languages
My Benchmark
👉 Indian Business Data: GST, UPI and Lakh-Crore on Kaggle
All 6 tasks and their notebooks are public, so you can run them on any model.
I used an AI assistant to draft this post. The idea, benchmarking, the Indian accounting rules and the final review are mine.



Top comments (0)