DEV Community

Cover image for We asked a decision model to check our AI's work. It got good the day we stopped asking it to think.
Tom Jones
Tom Jones

Posted on

We asked a decision model to check our AI's work. It got good the day we stopped asking it to think.

Microsoft put a small, strange model on OpenRouter this month. You hand it a situation and a question with named answers, and it hands back a probability for each answer. Microsoft-Decision-1 does that one thing.

We run a gateway that sends most requests to cheap models and checks the work before it goes out. A model that only answers yes or no, with a number attached, sounded like the checker we had been trying to build for months. We spent a Saturday finding out.

This is what happened, including the parts that did not work.

1. At a glance: four versions took wrong approvals from 9 to 0

The first thing it taught us

We started by asking it the questions we would ask a careful colleague. Is this note relevant to what the agent is about to do? Does this tool call match what the user asked for?

It wobbled. We sent identical requests twice and got different probabilities back, sometimes by a lot. On a relevance question the same input moved by as much as 0.17 between calls. A check that changes its mind about the same input is useless.

Then we noticed where it held still. When the question was about a fact, it was steady. "Does the user's message contain this exact city name?" came back the same every time, and came back right.

That turned into the rule the rest of the day was built on. Let plain code find the facts: is the value in the user's words, is it the right type, does the list have the right number of items. Then ask the model one clear-cut question about a fact that code has already laid in front of it. When a question needs real judgement, do not force an answer. Say UNSURE and send it up to something stronger.

Give it a fact, not a judgement.

2. Asked for a judgement, it wobbled. Asked about a fact, it held still.

The checker

What we built is boring on purpose. Code checks the call. Decision-1 answers a few one-fact questions. If anything is wrong and fixable, a cheap model gets one try to repair the call, and the repaired call goes through the same checks. The output is APPROVE, REJECT or UNSURE, with a dossier of every check that ran.

3. Code finds the facts; the model only answers about them

We measured three things, because a checker can fail in three ways:

  • How often it approves a call that is wrong. We call this "less wrong", and wanted it at 5% or under.
  • How often it takes a wrong call, fixes it, and approves the fix. "More right." We wanted at least 20%.
  • How often its repair breaks a call that was already right. "Nothing broken." 2% or under.

Every version was tested the same way. Before any data existed, a separate model that had not seen our results wrote the pass and fail rules, and we locked them. Each version ran once on 300 tool calls it had never seen. Then fresh labellers, separate AI sessions given only a written rubric, who could not tell an original call from a repaired one, marked each call right or wrong, with a tenth of the calls labelled twice as a check. We scored only after the labels were in.

4. A checker can fail three ways, so we measured all three

Four versions in one evening

version wrong calls it approved wrong calls fixed and approved right calls it broke
2.1 9 of 48 25 of 100 2 of 200
2.2 4 of 33 18 of 92 2 of 203
2.3 2 of 34 22 of 88 1 of 211
2.4 0 of 72 20 of 96 0 of 201

5. Leaks fell every round; fixes dipped once and came back

Each step came from reading the calls the last version got wrong.

Version 2.2 raised the bar for a yes, and the leaks fell. The fixes fell with them. It was trading one good thing for another, so we set a floor on fixes and kept going.

Version 2.3 looked at the thirteen wrong calls that 2.1 and 2.2 had approved, and found they came in four shapes, all of them things code can see. A list sent as text inside a list. An optional value the user never mentioned. "Washington" when the user said "Washington D.C." A limit of 5 when the user said "less than 5". Four new facts for the checker, and the leaks fell to 2 while the fixes came back.

One of those two leaks was a list of dietary needs sent as ['["vegan"]']. A list, inside a list, written as text. The exact shape 2.3 was built to catch. It slipped through because one part of the code wrapped the text in a list before the next part could notice it was a list already. The other surprise was a repair that turned a growth rate of 0.015 into 1.5, the same number written as a percent, which is a very different forecast.

6. A list, inside a list, written as text, slipped past the check made for it

7. A repair turned a 1.5% growth rate into 150%

Version 2.4 fixed the order of those two checks, stopped the repair model from rescaling numbers, and added one quiet rule: a repair is kept only if the checker approves it. Otherwise the original call goes out, untouched. That rule is why nothing broke. A repair nobody could vouch for was never allowed to replace anything.

8. A repair nobody could vouch for never replaces anything

What zero means

When we told the reviewer who writes our rules that 2.4 had approved no wrong calls, it pushed back hard. Part of its review held. Every version ran on a different draw of calls, with different labellers, so the trend across versions carries noise. And the fix rate of 20 out of 96 clears our floor of 20% by a hair.

Part of it did not hold. It suggested the zero was flattered by a larger denominator. We counted the plain number underneath. Across the four versions the checker said yes 164, 150, 120 and 110 times, and was wrong on 9, 4, 2 and 0 of those. When 2.4 says yes, it has been right every time so far. That count stands on its own.

9. When 2.4 says yes it has been right every time; it says yes less often

The real price is visible in the same numbers. It says yes less often.

The bill

Decision-1 is cheap enough to disappear. The whole 300-call test cost under a cent, and on 1,417 more recorded calls it ran at about two thousandths of a cent per call, in under a second at the median.

The interesting cost is somewhere else. A checker changes what you are buying. Without one, you pay per token and hope the answer was right. With one, the unit becomes a checked answer, and its price includes every call the checker would not approve and had to send somewhere more expensive.

That pushed us to think in price per task instead of price per token. On our own recorded runs, a cheap model plus the checker plus escalation lands at a fraction of what the same tasks cost on a premium model called directly. How small a fraction depends almost entirely on one number: how often the checker can say yes.

10. A checker turns price per token into price per checked task

The problem we have not solved

Four in ten calls come back UNSURE. Most of them hinge on an argument the checker cannot pin to a fact in the user's words. Code cannot see what such a value means, and the decision model, asked to judge it, wobbles.

11. Four in ten calls come back UNSURE

The question we are carrying into next week is this. How do you turn a judgement about meaning into a fact that plain code can check, without teaching the checker to say yes to things it should not?

If you have built something that does this, we would like to hear how.

Tom Jones, Spanda Works. The checker and the test harness are ours; Microsoft-Decision-1 is Microsoft's model, used through OpenRouter.

Top comments (0)