DEV Community

Cover image for I Moved Decision Logic Out of an LLM and Measured What Broke
sunnydachs
sunnydachs

Posted on

I Moved Decision Logic Out of an LLM and Measured What Broke

Two earlier posts in this series compared two "decision models" — text models that return a probability instead of prose — and found they fail in opposite ways (part 1, part 2).

But that was still a model-vs-model question. The operational question I actually cared about:

if your pipeline asks a general-purpose LLM to make a judgment call,
what actually changes when you carve that decision out
into a decision model?
Enter fullscreen mode Exit fullscreen mode
pitch:      "faster, cheaper, safer"
my policy:  believe it exactly as far as I can measure it
Enter fullscreen mode Exit fullscreen mode

The measurement harness and all raw data are public:

https://github.com/sunnydachs/decision-flow

What I measured

Task: classify customer-inquiry emails into 5 categories, one of which is an urgent claim with a high cost of being missed.

The three decision models are Jev, Mercury, and pplx-decider — text models that return a probability instead of prose.

Seven approaches, all given the same data, the same category definitions,
and the same concurrency/timeout budget:

 1. keyword rules
 2. embeddings + logistic regression (local)
 3. general LLM (ask for JSON in the prompt)
 4. general LLM (JSON-schema-enforced)
 5. mercury-decide (decision model)
 6. pplx-decider (decision model)
 7. Jev (decision model)
Enter fullscreen mode Exit fullscreen mode

Thresholds were tuned on a separate 300-item set only, then applied unchanged to the 500-item evaluation set.

The evaluation data itself is LLM-generated synthetic data, so treat absolute numbers as indicative and the differences between methods as the signal.

Result 1: accuracy is a tie

 keyword rules        0.732
 embeddings + LR      0.878
 mercury-decide       0.886
 general LLM (prompt) 0.908
 general LLM (schema) 0.904
 pplx-decider         0.902
 Jev                  0.900
Enter fullscreen mode Exit fullscreen mode

The top five sit between 0.886 and 0.908 — within noise of each other. Neither "decision models are smarter" nor "the general LLM is smarter" survived measurement. Accuracy does not decide this comparison, so the rest of this post is about everything else.

Result 2: how much you can automate splits from 0% to 47%

This is the operational question that matters: holding recall of urgent claims at ≥95%, what share of the inbox can be auto-processed?

                      recall (95% CI low)   automation rate
 embeddings + LR       0.922 (0.885)       47.2%
 mercury-decide         0.969 (0.943)       37.8%
 pplx-decider           0.979 (0.958)       28.6%
 general LLM (schema)   0.984 (0.963)       20.2%
 general LLM (prompt)   0.990 (0.974)       14.2%
 Jev                    1.000 (1.000)        0.0%
 keyword rules          1.000                0.0%
Enter fullscreen mode Exit fullscreen mode

Same safety floor, and the automatable share ranges from nothing to nearly half.

risk-coverage curves: x = share auto-processed, y = urgent-claim recall; dotted line is the 95% floor; circles mark each method's chosen threshold

Look at Jev: zero missed urgent claims — the most careful judge in the set — yet 0% automation. The reason is the shape of its probabilities.

Result 3: "probability" means different things

Decision models return probabilities, but the distributions are nothing alike:

                    P == 0.0   P == 1.0
 Jev                  53%        24%      <- nearly binary
 pplx-decider          0%         0%      <- continuous
 embeddings + LR       0%         0%      <- continuous
 mercury-decide        0%         0%      <- extreme-leaning but never exactly 0/1
Enter fullscreen mode Exit fullscreen mode

Jev returned a probability of exactly 0.0 for over half the items. Its judgments are right, but with probabilities pinned to the extremes there is no "route the uncertain ones to a human" band to operate in. Raise the threshold at all and misses appear — so the only safe operating point routes everything to review. That is what 0% automation is made of.

probability mass by band: Jev and mercury pile up at the extremes; pplx-decider and embeddings+LR stay continuous across the whole band

mercury-decide leaned extreme too (90% of its mass below 0.05 or above 0.95) — enough to cap automation at 37.8% — but it never returns exactly 0 or 1, which is why it still leaves a thin band to threshold on. This extends part 2's observation (one model binary, one spread) to 500 items — and shows it landing directly on an operational metric.

Result 4: latency and output size are a different league

                      p50 latency   output tokens / item
 keyword rules             0 ms      -
 embeddings + LR           0 ms      -
 Jev                     201 ms      53
 mercury-decide          308 ms      3
 pplx-decider            325 ms      1
 general LLM (prompt)   4133 ms      234
 general LLM (schema)   3372 ms      154
Enter fullscreen mode Exit fullscreen mode

The decision models answered in 0.2–0.3 s with 1–53 output tokens per item; the general LLM took 3–4 s and wrote 154–234. "Faster" held up this time — a generation model writes prose, a decision model returns a value. One nuance: enforcing the JSON schema cut the general LLM's p95 latency by 1.5 s — output-format enforcement helps speed too.

Result 5: stability differs in kind

Same input, 20 runs:

                     20 identical runs   paraphrased instructions
 general LLM (prompt)        8%                    8%
 general LLM (schema)        4%                   10%
 pplx-decider                0%                    2%
Enter fullscreen mode Exit fullscreen mode

The decision model was perfectly deterministic on identical input — but paraphrasing the instructions still flipped 2% of its labels (the LLM flipped 8–10%). Decision models look less prompt-sensitive, not insensitive.

What I take from this

Stated only as far as the numbers go:

all seven methods: horizontal bar = automation rate at the safety floor, right-side columns = recall / misses / latency, sorted by automation

  • Accuracy didn't separate the methods. 0.886–0.908 for the top five.
  • Decision models were genuinely fast. 0.2–0.3 s vs 3–4 s.
  • "Safe" was not automatic. A model whose probabilities collapse to 0 and 1 achieves zero misses and zero automation. A continuous-probability model automated up to 47%, but in one case let the recall floor's lower CI slip to 0.885.

So carving decisions out of the LLM is worth it on speed alone in this task — but which model's probabilities you operate on determines the shape of the whole pipeline. Before migrating, plot that model's probability distribution on your own data. That is the one check that would have predicted every operational number in this post.

The decision models also expose other question types (yes/no, score) — combining those is the next measurement.

Honest limitations

  • The evaluation data is LLM-generated synthetic data with self-labels; no human double-annotation was performed. This is a relative comparison between methods, not a quality guarantee. For reference, trained human annotators agree with each other around 79% on hard classification tasks (Gilardi et al. 2023, PNAS) — all methods landing at 0.88–0.91 is at or above that band.
  • One general-purpose LLM was tested. Conclusions are about that model, not about larger ones.
  • Jev's stability (20 identical runs) was not measured for budget reasons.
  • Costs were measured as output-token counts per item, not dollar amounts; see the repo for the per-1k estimates.
  • No real customer data was used; validate with human review before production.

This is a personal OSS project — no warranty. Issues and improvement ideas are welcome.

Repo: sunnydachs/decision-flow

Top comments (0)