DEV Community

Cover image for Faster model, slower chatbot. What we learned from testing Jev
Michal Nowikowski
Michal Nowikowski

Posted on Originally published at webspeaker.pro AI-assisted

Faster model, slower chatbot. What we learned from testing Jev

Originally published on the WebSpeaker blog.

In short: We tested Jev, a new model for rapid decisions, as a way to speed up WebSpeaker. Jev made decisions quickly and slightly more accurately than the model we use today. Yet in every variant we tested, visitors waited longer for an answer. We did not deploy it. The main lesson: in an AI product, the time and quality of the entire conversation matter more than one model's benchmark result.

Why we considered it

WebSpeaker is a chatbot companies put on their websites to answer visitors' questions using the site's content and documentation. Every answer begins with a few decisions: should the chatbot search the company's materials, what should it look for, and does the visitor want to speak with a person? Only then does the chatbot write an answer and check it before sending it.

Today, Gemini makes those initial decisions while also preparing search queries and notes about the conversation. TypeSafe AI released Jev, a model designed to make fast, almost reflexive choices from a given list rather than write text. In our first test, a single Jev decision took a median of 232 ms, while Gemini took 1,824 ms to prepare the full plan. These were not the same tasks (Gemini did more), but the difference was large enough to make us test a division of labor: Jev decides, Gemini writes.

What we tested

We ran four trials with synthetic questions in Polish and English. No customer data was sent to the model under test.

One decision. Should the chatbot retrieve information, answer immediately, or hand the conversation over to a person? Jev and Gemini were equally accurate: 96.7% across 60 questions. Jev was much faster, but did only part of the work.

Jev decisions instead of Gemini decisions. We replaced some decisions in Gemini's plan with Jev's and ran the conversation through the whole chatbot. Jev more often chose the answer we expected (91% versus 87%), but the plan slightly more often needed an additional check by another model. A decision that is correct in isolation may not fit the rest of the plan: for example, "search the documentation" when no one has prepared a search query.

Jev first, then Gemini. Gemini had less work and finished faster, but it had to wait for Jev first. With a larger set of decisions, Jev's median call time rose from 232 to 872 ms. The analysis stage took longer, and the full answer arrived about 27% later on average.

Jev and Gemini in parallel. This time the analysis stage was finally faster: by 9% for a typical request and by 24% for slower ones. Decision accuracy was also higher (92% versus 87%). Even so, the full answer arrived about 22% later, and in one case the chatbot gave an internally contradictory answer.

Why a faster stage didn't produce a faster answer

Analyzing the question is only one of several steps. It is followed by retrieval, sometimes an additional check of the plan, writing the answer, and checking that answer. The 151 ms saved on analysis were lost in what happened afterward.

And what happened afterward was not always the same. Changing the division of labor made the chatbot choose a different route in some conversations, search more or less broadly, or run an additional check. When two models work in parallel, neither knows what the other has decided: Gemini writes a search query without knowing whether Jev decided retrieval was needed at all. With such a small sample, we cannot say which of these factors mattered most, but the overall direction is clear.

When you split work between models, you change what each of them knows. The speed of one component tells you nothing about the speed of the whole system.

What this means for WebSpeaker customers

We only add a new model to WebSpeaker when it improves the quality, speed, or cost of the entire conversation without making the other aspects worse. A fast result in an isolated test is a reason to investigate, not a reason to deploy.

What happens to the data matters just as much. Before a new model provider sees visitors' questions, we check where it processes them, how long it retains them, whether it uses them to train models, and how it performs under real traffic. Jev has not reached that stage: it did not enter customer conversations, and these experiments stayed outside production code.

Will we revisit Jev?

Possibly. This was a pilot with eight test scenarios, each repeated three times, and we did not systematically assess the quality of every final answer. That is not enough to settle the question either way. We also have not calculated the full cost of the two variants.

We will revisit it if, on a larger and more varied set of questions, the division of labor consistently delivers a faster full response at a lower cost without reducing quality, and the provider passes our data-handling review.


Have you seen a faster component make your whole AI pipeline slower? I'd love to hear how you measured it in the comments.

Top comments (0)