DEV Community

Cover image for How we prove a cheaper model can do the job
Pranav Dhoolia
Pranav Dhoolia

Posted on Originally published at operantlabs.com

How we prove a cheaper model can do the job

Disclosure: I'm building Operant, so I'm biased. This post first appeared on our blog. I'm sharing it here because I'd like honest critique of the method, and my questions are at the end.

You built an agent, and most of its day is the same few tasks, over and over. Every one of those calls still goes to the most expensive model you have.

So your ops assistant pays top price to answer the kind of question it answered yesterday. A cheaper model could probably do it. But one bad answer in front of a customer costs more than the savings.

Operant moves a repeated task to a cheaper model. It does that only after the model passes a test on your own past work. We call that test the gate. This post walks through it step by step, with screenshots from the live console. The values in them come from a sample workspace.

It starts with a rule that does nothing

When Operant spots a task your agent keeps repeating, it writes a skill for it. A skill is short notes learned from your past chats that went well. Then it proposes a rule. The rule says to send this task to a cheaper model, with these notes. A new rule starts in shadow, which means it runs but moves no live traffic. The gate is the only way out of shadow.

A rule for the ops assistant

The test uses chats the notes never saw

Before the notes are written, about a quarter of the chats for that task are set aside. The notes learn only from the rest. The gate then tests on the chats that were set aside, never on the ones that taught the notes. A test on the teaching chats would only show that the model can repeat what it was shown.

Where the notes came from

The gate picks the most expensive chats first, because that's where a wrong switch would hurt you most. By default it runs up to five of them. Some tasks don't have enough saved text. For those, the gate makes short practice tasks from goal summaries. It labels those scores, so you know the evidence is weaker.

Both models answer the same chats

For each test chat, the gate sends the same user messages twice. One run uses the model that the chat actually used. The other run uses the cheaper model with the notes. So the cheaper model is compared with what your agent really did, not with some fixed reference model.

A judge grades the two answers

Next, a judge reads both answers side by side. The judge is a strong model that compares two answers and gives a score:

Score What it means
1.0 The cheaper model did as well as the old one, or better.
0.5 It did part of the task.
0.0 It didn't get the task done.

The judge checks whether the task got done, and whether the answer was correct and complete. It ignores style and length. It also writes one sentence that explains each score. The default judge is claude-opus-4-8, and the console shows it on every rule.

The gate passes when the average score is 0.8 or higher.

A person still has to say yes

Passing the gate moves no traffic on its own. Someone on your team has to approve the rule. Only then does the cheaper model take over the task, with the notes added to its instructions.

The next steps for a task

Change anything, and it tests again

The gate approves one exact setup: the cheaper model, the version of the notes and the compression settings. Change any of them and the old result is cleared. The gate has to run again before anyone can approve. A live rule can't move to a different model at all. You put it back in shadow first, edit it, and test again.

Undo takes one click

Unsure after a switch? The Demote to shadow button sends the task straight back to the model it used before. Undo needs no test, so you can stop a change the moment you're unsure.

What the score does and doesn't tell you

In the sample workspace above, Claude Haiku scored 95% against the old model on the ops assistant's task. That's well over the 0.8 bar.

That score has limits. Readers on Reddit pointed out three of them, and they were right.

It measures agreement, not a correct outcome. The judge compares the cheaper model with your agent's past answers. If a past answer was wrong, a cheaper model that gives the same answer still scores well. Nobody grades the real outcome of each case.

It only sees old inputs. The test chats come from the past. So the gate can't see new kinds of requests that arrive after the switch. A next step is to keep a small share of live traffic on the old model for a while. Then the two sets of outcomes can be compared. Until then, the undo button is the fast way back.

A good average can hide a bad case. The bar is on the average, so one costly failure can hide behind good scores. The judge also reads each answer as plain text, with no separate check that required fields are there. Format checks and a second bar on the worst cases are next steps.

So a cheaper model only gets a task after it proves itself on your past chats and you approve.

Sign up at app.operantlabs.com and see which of your agent's tasks could move.

Questions for you

I'd really like your critique:

  1. Is a mean parity of 0.8 the right bar, or should every conversation have to pass on its own?
  2. The baseline is the model each conversation used most. Is that a fair reference, or does it hide its own mistakes?
  3. A default run uses 5 conversations. How many would you want before you trusted the result?

I'll read and reply to every comment.

Top comments (0)