DEV Community

Yuli Dai
Yuli Dai

Posted on Fully Autonomous

Two models sent a moving bill to the wrong chat

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge.

What I Benchmarked

I'm building TheOne, a DSH plugin where you talk in one main chat and each topic has a separate background session. Before the assistant answers, it has to decide where your message belongs.

I wanted to isolate that decision. Given a catalog of three topics and recent conversation, can a model choose an existing topic, recognize a new matter, or ask for clarification?

The test has 36 synthetic cases, six in each group:

  • Explicit topic references
  • Resuming a topic from recent history
  • Overlapping words and facts across topics
  • Starting a genuinely new matter
  • Ambiguous references that require clarification
  • Routing instructions quoted inside untrusted content

The topics are a thesis paper, a monthly household budget, and an apartment move. Each case makes one model request. The catalog order rotates, and neither the expected label nor the group name appears in the prompt. The model must return a topic ID, NEW, or CLARIFY. Scoring is exact matching after case and whitespace normalization. Extra explanations score zero; there is no LLM judge.

The cases and implementation were authored with AI assistance. A separate offline harness checks scoring with scripted responses. Those checks make no model calls and are not part of the model results below.

Models Tested

Kaggle's initial task execution used Gemini 3.7 Flash. I then ran GPT-5.4 nano as a lightweight alternative on the same frozen cases and scoring policy. These are single runs, not repeated trials or a search for the best prompt.

I also scheduled Qwen 3 Next 80B Instruct, but that run is still incomplete at the time of writing. I am not assigning it an accuracy score or treating incomplete execution as a routing failure.

Findings

Group Gemini 3.7 Flash GPT-5.4 nano
Explicit topic 6/6 6/6
Resumption 6/6 6/6
Overlapping facts 5/6 5/6
New matter 6/6 6/6
Clarification 6/6 5/6
Quoted instructions 6/6 6/6
Total 35/36 (97.2%) 34/36 (94.4%)

Both models missed this message:

Add the 480 EUR moving bill as a one-off expense in my monthly budget.

Both returned the apartment-move topic. The expected destination is the household budget: the user is asking to categorize an expense. The routing policy explicitly distinguishes moving logistics from monthly spending.

The paired case made this easier to inspect:

The mover's quote is 480 EUR. Compare it with renting a truck for moving day.

That request belongs to the apartment move, and both models routed it correctly. The number and subject overlap, but the requested action changes the destination. My interpretation is that the moving-related content outweighed the budget operation in the failed case. Two examples cannot establish a general mechanism, but they give me a specific failure to test further.

GPT-5.4 nano also picked the thesis topic for “Use the smaller number,” even though the supplied history mentioned lowering both the participant count and the grocery allowance without selecting either. Under the stated policy, that should be CLARIFY.

The six quoted-instruction cases passed for both models. That is useful as a basic check, but the outer requests are explicit and the quoted attacks are simple. I would not call this evidence of general prompt-injection resistance.

The completed traces record about $0.0368 for Gemini and $0.0026 for GPT-5.4 nano. Those are recorded inference costs for these runs, not a price comparison for general workloads. The pending Qwen run is excluded. Execution used Kaggle's free allowance; no paid subscription was started.

The scores are close to the ceiling. One case changes the total by 2.78 percentage points, so the one-case gap is not enough to declare a broadly better model. The concrete mistakes tell me more than the ranking.

Next I would expand the paired cases while keeping the operation distinct from the subject, then test natural ambiguous dialogue without phrases such as “neither is selected.” I would also test cases where a single message genuinely needs two projects, and define the expected behavior before running the models.

This is a small public diagnostic of base-model decisions under an explicit policy. It does not measure production TheOne accuracy, represent real user conversations, or provide a contamination-proof held-out test.

My Benchmark

The result summarizer checks all 36 case traces against the source labels and the parent score before producing the tables. Kaggle run IDs are 4630860 (Gemini) and 4630864 (GPT-5.4 nano).

If you've built a router like this, how do you decide which topic owns a message that mentions one project but asks for work in another?

Top comments (0)