DEV Community

ohyoseok92
ohyoseok92

Posted on

Three models aced my Korean prize benchmark. That is the lesson.

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge.

What I Benchmarked

I wanted to test a narrow mistake in contest discovery: reading a Korean prize notice and reporting the headline prize pool or a gift voucher as the first-place cash award. A contest listing that says “total prizes: 10 million won” may pay only 3 million won to the winner.

I wrote ten original, synthetic Korean notices with hand-checked answers. The task asks each model to return only the KRW cash amount for 대상 or 1등, for the whole winning team. It returns 0 when that award is non-cash or undisclosed. Cases cover 만 and 억 notation, commas, distractor awards, a gift voucher, an unknown amount, and a team payout. Exact integer matches earn one point; the reported score is the fraction correct. No personal data or scraped contest notices are in the test.

Models Tested

I ran the same Kaggle task against Gemini 3.7 Flash, Gemma 4 26B A4B, and GPT-5.4 nano. I wanted a mix of a widely used hosted model, an open model, and a small hosted model. Each saw the same ten cases and scoring code.

Findings

Model Exact matches Score
Gemini 3.7 Flash 10/10 1.00
Gemma 4 26B A4B 10/10 1.00
GPT-5.4 nano 10/10 1.00

The result surprised me less as a ranking than as a warning about the test: three perfect scores do not distinguish these models. The cases confirm that all three follow this explicit instruction on clean, short notices. They do not show how any model handles real contest pages.

The zero-answer cases are useful checks: “50만원 상당의 여행 상품권” must not become 500,000 KRW of cash, and an undisclosed first-place amount must not be inferred from the total pool. All three models got those cases right. The team example also asks for the whole team's award rather than a per-person division; all three returned 1,000,000 KRW.

What would I measure next? Longer notices with prize tables and footnotes, OCR errors, mixed currencies, ranges and conditional awards, plus paraphrases of the same source so the answer is not tied to one wording. I would add reviewed real notices only with clear reuse rights. I would also report error types alongside accuracy: wrong award level, unit conversion, non-cash confusion, and unsupported inference.

This first version is a small, transparent probe. Its strongest finding is the ceiling effect, and it gives a reproducible baseline for a harder follow-up. Codex helped build and run the notebook; the cases, code, and evaluations are open for review.

My Benchmark

Korean Contest Prize Reading leaderboard on Kaggle. The task and linked notebook show every prompt, expected integer, scoring rule, and run.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to