This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
I run a voice CRM for a realtor. The agent is supposed to take a lead over the phone: hear a fact, save it, ask for the one missing detail, and stay on the line until the realtor says that is all.
In production it kept missing that script. It went silent after saving a field. A phone number made it skip the email and jump toward the close. It asked again for a name or a budget it had just been given. It stored "call him Friday" and never said Friday. It saved and hung up before asking whether that was everything.
I turned those five failures into five Kaggle tasks. The tools return only the fields they saved. They do not hand the model its next sentence. A short spoken reply is in the prompt, and it is not what the score measures. A task scores 1 only when every check passes.
- speak-after-update. After a new lead, ask one missing question out loud.
- email-after-phone. A phone number must not skip the email question.
- no-repeat-known. Do not ask again for a field the realtor already gave.
- confirm-before-end. Ask if that is all before saving, and end the call only after yes.
- speak-the-reminder. Say the follow-up day. Storing it is not enough.
Models Tested
The model on the live phone line is Gemma 4 26B (gemma-4-26b-a4b-it). That was the one I most wanted to see fail in the same way the calls fail.
I also ran Gemini 3.7 Flash, which Kaggle ran when each task was created, plus Gemini 3 Flash, Claude Sonnet 5, and GPT-5.4 mini. Same prompt, same tools, same leads: Ada in Akobo, Tunde's land in Ibadan, Ngozi's duplex in Lekki, and a Friday reminder for Chidi.
Findings
Gemma passed all five. So did Gemini 3.7 Flash. Gemini 3 Flash and Claude passed four. GPT-5.4 mini passed two.
| Task | Gemma 4 26B | Gemini 3.7 Flash | Gemini 3 Flash | Claude Sonnet 5 | GPT-5.4 mini |
|---|---|---|---|---|---|
| speak-after-update | Pass | Pass | Pass | Pass | Pass |
| email-after-phone | Pass | Pass | Pass | Pass | Fail |
| no-repeat-known | Pass | Pass | Pass | Pass | Fail |
| confirm-before-end | Pass | Pass | Fail | Fail | Fail |
| speak-the-reminder | Pass | Pass | Pass | Pass | Pass |
On the opening lead, Gemma saved Ada Okafor, the three-bedroom flat, and Akobo, then said, "Got it. Do you have her phone number?" Given the phone number, it asked, "Thanks. Do you have an email for her?" Told she had no email, it moved on: "No problem. What is her budget?" For the reminder it said, "I'll remind you to call Chidi this Friday." Before ending Ngozi's call it asked, "Is that all?" and closed with, "The lead is saved. Have a great day!"
That is the opposite of what the live calls do. The intake policy is in this model. The production misses are in the voice loop around it: silence after a tool call, a phone number that skips a step, a reminder that is stored and never spoken.
GPT-5.4 mini is the model that actually skipped the script. After Ada's phone number it asked, "What's her budget?" Told she had no email, it stored the email as "No email" and jumped to, "Is that all for Ada Okafor?" For Tunde it first asked, "What's his email?" Then, when the realtor said, "What else should I tell you?" it saved the lead, ended the call, and said "Saved." Email and a follow-up were still missing.
Claude and Gemini 3 Flash failed only the last sentence of the closing task. Both asked whether that was all, both saved, and both ended the call. Claude said, "Saved Ngozi's lead successfully!" and then, after the hang-up, "Thanks for calling, have a great day!" Gemini said the details were saved, ended the call, and finished with "Omnipresent Voice AI CRM lead capture complete." The check looks at the final sentence. A correct save followed by a polite goodbye scores zero. GPT did say "Saved." and still failed, because it never called end_call.
Two measurements are worth doing next. One is the spoken audio path, because a perfect text transcript did not reproduce the live failures. The other is a closing check that accepts "saved" anywhere in the closing turn, so a second goodbye sentence is not an automatic miss.
Top comments (0)