DEV Community

Shouvik Palit
Shouvik Palit

Posted on

Decision models are a new category, and Microsoft just joined it

For years, "new model" meant "new chatbot." Lately a different kind has been showing up: models that don't write anything. They return a decision. 🎯

🧩 The category

TypeSafe's Jev was the first I noticed. TypeSafe calls it a "System One" model. Per its OpenRouter listing it's a structured decision model: text in, a typed choice out. As DataCamp describes it, you give it program state and a set of typed questions, and it answers in one pass with categorical choices, scores or yes/no probabilities. Each answer carries a calibrated confidence number and no written explanation.

Then Microsoft shipped Microsoft-Decision-1. 🚀 The interface is the same kind: its OpenRouter listing describes a small model that reads the content you give it and returns a calibrated probability for each fixed answer option. The listing shows:

  • 📥 text in, decisions out
  • 🧠 a 32K context window
  • 💸 about $0.042 per million input tokens, free output
  • ☁️ served through Azure

The two are similar from the outside, but not the same model. Microsoft's post says Decision-1 was post-trained from Qwen3.5-9B, while DataCamp describes Jev as a non-transformer design. They look like competitors in one new category, not one built on the other.

🛠️ Why developers should care

A chat model gives you prose you have to parse. A decision model gives you a fixed set of options with probabilities, which is the shape code already uses: a branch, a threshold, a route. The "calibrated" part is what matters. A confidence you can trust tells your code when to act automatically and when to hand off to a human or a bigger model. 🤝

My guess, not something either listing says, is that these fit the cheap, fast steps in agent pipelines: routing, triage, gating, moderation.

📊 What's claimed

All of these are vendor-reported, not independently verified. ⚠️

  • Microsoft (via OpenRouter's launch post): highest accuracy across 36 blind benchmarks, about 150K questions; 4.5x faster than the runner-up; decisions flip on only 1.3% of perturbed inputs.
  • TypeSafe (via DataCamp): roughly 68% agreement with reference answers on a four-workflow test, sub-second latency. DataCamp notes the workflows were written by TypeSafe, which is a possible bias.

Expect more independent comparisons once developers start trying both. 🔍

📺 How DEV·TV showed it

I built DEV·TV as a TV for the developer internet: 13 channels covering GitHub, Hacker News, CVEs, papers, new AI models and more. It's a single index.html with no backend and no login.

Its NEW MODELS channel reads OpenRouter's catalog, newest first. A few hours after Microsoft-Decision-1 went live, it was the first slide on the channel: the name, "added 4h ago", an outputs decisions tag so it doesn't read like a chatbot, context and pricing, and the description. The NEXT and LATER lines show what's coming up.

That's the point of the channel. A model like this has no chat demo and no flashy screenshot, so many developers will only notice it if it's put in front of them. 👀

🔓 Try it

DEV·TV is open source (Apache-2.0). ⭐ Try it at shouvik12.github.io/devtv or read the code on GitHub.


📚 Sources: DataCamp on Jev and System One models · Jev 1.13 on OpenRouter · Microsoft-Decision-1 on OpenRouter

Microsoft-Decision-1

Top comments (2)

Collapse
 
suhas_iyengar_46220cb7ca3 profile image
suhas iyengar •

Very insightful Shouvik! Dev TV keeps showing relevant content at the right time!

Collapse
 
shouvik12 profile image
Shouvik Palit •

Thank @suhas_iyengar_46220cb7ca3 Much appreciated