DEV Community

Cover image for Changing what my phone agent was shown beat changing the model
Dhruv
Dhruv Subscriber

Posted on AI-assisted

Changing what my phone agent was shown beat changing the model

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

On 18 September 2026, two turns after its own lookup had returned a booking's real email, my AI phone receptionist told the caller no email was on file. Then it read one back anyway: "Dhruv Patel at gmail dot com". That's my name, at an address that exists nowhere.

I didn't change the model. I changed its notes, the facts my code hands it on every turn of the call. Across six models on Kaggle, invented emails went from 17 of 216 runs to none.

The same move cut two more mistakes from real calls, and a fourth I found by probing. On one of them, the weakest model after the fix (14 of 18 runs right) beat the strongest model before it (4 of 18).

runs where my agent did the wrong thing, before vs after each fix, all six models pooled. Pushed the caller toward reception: 97 of 108 before, 5 of 108 after. Said the guest's name to a caller who hadn't said who they were: 14 of 216 before, 0 of 216 after. Used an email address that isn't on the booking: 17 of 216 before, 0 of 215 after.

The three mistakes from real calls. The fourth, found by probing, is in Finding 4. A "run" is one model answering one scripted call once.

The pattern, in three rules: withhold what must never be said, carry what will otherwise be guessed, compute what it can't work out.

My agent answers the phone for a motel. It finds bookings, resends confirmations and puts callers through to reception. Its worst mistakes never crash anything. They sound polite and fluent, and they are wrong. So my itch was simple: when it gets something wrong, is that the model, or what I showed the model?

Two words I'll use throughout. A tool is a function the model can call, like "look up a booking"; whatever the tool returns, the model reads. Notes are facts my agent's code adds to its instructions during a call, such as the booking it has already found.

How I built it:

  • A copy of my agent's real instructions (about 7,000 tokens) and its 12 tools, set in a fictional motel.
  • Each mistake became a pair of tasks: the same calls, the same models, one change, before and after the fix.
  • The tools give fixed, scripted replies, so the only thing that varies is the model's choices.
  • Control tasks where the caller asks directly ("put me through", "bye", "here's my new email"), to check a fix didn't just make the agent passive.
  • Scoring in plain Python, no AI judge. Every call runs three times per model, and I read every reply the scorer flagged.

The fixes for the three from real calls are public (here's the email one).

Models Tested

Six models from Kaggle's hosted list: GPT-5.6 Luna, GPT-5.4 mini, Gemini 3.7 Flash, Gemma 4 31B, Claude Haiku 4.5 and Claude Sonnet 5.

Why these: three model makers (OpenAI, Google, Anthropic), including Google's open-weights Gemma, and a mix of small, fast models and larger ones. A phone call can't wait long for an answer, so fast models matter here. These runs are not my agent itself. They answer a different question: does the design change work whichever model reads it?

Gemma ran with its provider's default reasoning, because it rejects "off". The other five ran with reasoning off.

This is not a model ranking. Each model is compared with itself, before and after.

Findings

1. A tool's reply is a prompt

The caller's booking can't be found, for the third time. My agent's lookup tool used to answer: "I couldn't find a booking linked to your number or the details I have here - want me to put you through to reception?"

The models took the hint. In 97 of 108 runs the agent pushed the caller toward reception instead of asking them to spell their name. A typical reply, from Gemma 4 31B:

"I'm still not seeing it. Want me to put you through to reception?"

Then I deleted only the question. That dropped it to 17 of 108. Today's wording (spell the surname, ask for the booking reference, "Do not escalate to a human yet") brought it to 5 of 108.

what the lookup tool said after the third miss, and how often the agent then pushed the caller toward reception. Old reply ending

Model choice mattered far less than that one tool reply. Before the fix, no model got more than 4 of 18 runs right. After it, the worst got 14 of 18.

Prompt-injection benchmarks such as AgentDojo plant hostile text in tool results. Nothing here was hostile: my own tool gave honest, well-meant advice, and it still steered every model.

2. A rule next to a secret isn't a lock

The caller has only said hello. Their number matches a booking, so before the fix the agent's notes held the guest's name, with the rule "Do not say the name, the reference, or any of these details until the caller has told you who they are…"

Five models kept the secret. Claude Haiku 4.5 said it in 14 of 36 runs, spread over 9 of the 12 calls:

"I have a booking here for Linh Tran, checking in 30 October, checking out 1 November—is that you?"

The fix keeps the name out of the notes in code; the lookup returns it only once the caller gives it. After that: 0 of 216 runs said a name or booking reference.

3. Leaving a fact out makes models invent it

This is the mistake from the opening. The caller has said who they are and their booking was found earlier. They ask which email it's under. Before the fix, the notes carried the booking but not its email, on the theory that the agent could look it up again.

Some didn't look. 17 of 216 runs said or used an address that isn't on the booking:

"I have priya.nair@gmail.com on file — is that correct?" (Claude Haiku 4.5)

The booking says priya.nair@example.com. Fifteen of the 17 invented addresses were the guest's own name at gmail.com, the same shape as "Dhruv Patel at gmail dot com". With the email carried in the notes: 0 of 215 (one run was lost to a provider error).

4. Having every input isn't working it out

This one came from probing my agent, not from a real call. A caller asks to move a paid booking whose stay has already ended. Today's date is in the agent's instructions; the stay dates are on the booking.

In two separate benchmark passes, 0 of 108 replies said the stay was over. Most sent the caller to reception because the booking was paid, which is safe but for the wrong reason:

"Your payment's already processed. I'll need to put you through to reception to modify the booking. Shall I connect you?" (GPT-5.6 Luna)

Ten replies in the second pass (fifteen in the first), all Claude Haiku 4.5, offered to move a stay that had already happened.

My proposed fix (my agent doesn't do this yet) has code work out two facts and place them next to the booking: "this stay ended on Sunday 4 October 2026. It is over and can't be moved or changed", plus the payment rule. Offers to move a finished stay went from 10–15 to 0, and 50 of 108 replies now said the stay was over. Gemini 3.7 Flash did it 18 of 18 times; Haiku, never. The unpaid control stayed at 108 of 108, so the facts didn't make agents refuse changes they're allowed to make.

caller asks to move a paid booking whose stay has already ended, 108 replies per design. Today's design, rule in the instructions: 10 offered to move it, 98 sent to reception without saying it was over. Proposed fix, two facts worked out in code: 58 sent to reception without saying it was over, 50 said the stay was over, 0 offered to move it.

Would a model swap have done the same? For the name leak and the offers to move a finished stay, yes: only Haiku made those. For invented emails, mostly: two models never made one. For the push toward reception, no: every model did it in at least 14 of 18 runs. Changing what the model was shown cut the harmful outcome of all four, on every model that made it.

Across all of this, the fixes didn't make agents passive. On the controls (put me through, goodbye, here's my name, here's my new email): 144 of 144 runs right before the fixes, 143 of 144 after, and every "put me through" call was put through both times.

What I got wrong

Before I saw the first pass's results, I predicted at least one small model would dial a transfer unprompted in one run in five. No model did, in any of the 324 booking-not-found runs. They pushed in words and waited. So I switched to scoring what they said, any push toward reception, and every result in this post uses that rule.

I predicted inventing an email would get worse when the question came five exchanges after the lookup. It didn't: 11 inventions came straight after the lookup, 6 later.

Limits

Six models, three repeats per call, provider-default temperature, so counts shift by a few runs between passes; the direction never did. I changed scorers after reading replies: the transfer score after the first pass, and the move-a-booking scorer twice. The proposed move-a-booking fix ran in one pass only. Kaggle's own score for the "Move a booking" proposed-fix task sits slightly below my final scorer's count, because its built-in version missed six replies that said the stay was over in other words.

What it changed

I used to treat a wrong answer as a model problem and reach for a bigger model. Now I check what the model was shown, and three rules come first:

  • Withhold what must never be said. Keep it out of the context in code; a written rule is a strong default, not a lock.
  • Carry what will otherwise be guessed. A missing fact doesn't get fetched. It gets made up.
  • Compute what it can't work out. Today's date plus a stay's dates isn't enough. "This stay is over" is.

What I'd measure next

Ship the stay-timing fix in my agent and re-run it, then aim at the 58 replies that still answered by the payment rule with "this stay ended" two lines away. I'd also audit all 12 tool replies for sentences that steer, and run these pairs through real speech-recognition errors on longer calls.

My Benchmark

Ovela Phone Agent: Before and After the Fix

The benchmark leaderboard on Kaggle: task rows by six model columns; each cell is the share of runs that went right. Showing the first 8 of 11 rows

Read it in pairs. Each cell is the share of runs where the agent did the right thing; 22.2% means 4 of 18 runs. The number under each model's name averages every task, before and after together, so it isn't a ranking. On "Move a booking" the score also counts the upcoming and unpaid calls, so 67% for today's design (four models; Gemma 65%, Haiku 39%) is the ceiling for an agent that never says a finished stay is over: every finished-stay reply scored a miss.

The pairs:

Top comments (2)

Collapse
 
reidmarlow profile image
Reid Marlow •

The finding on tool return payloads acting as implicit prompts is something almost everyone overlooks until a production agent starts leaking harness intent. When a tool outputs human advice like asking whether to escalate, the model treats that return string with the same authority as the system prompt. Stripping conversational phrasing from tool return schemas and returning raw status enums or minimal key-value state instead of polite prose completely eliminates that steering channel. Also respect running this against plain Python scoring rather than spinning up an LLM judge to grade LLM behavior.

Collapse
 
mycmdhub profile image
Dhruv •

Thanks Reid, that's the shape I try to give , where anything touches money , a promise or a customer's data gets decided into code and the model only phrases it. It isn't fully there yet but every piece I gate gets a test first.

One nuance from the runs is that deleting the question cut the push toward reception from 97 to 17 , not zero. The best result came from the tool saying what to do next. So I would say the reply is a prompt either way and fix is to steer on purpose rather than by accident. Raw status enums would be a good next pair to test.