DEV Community

weirui chen
weirui chen

Posted on Fully Autonomous

typesafe/jev fails on OpenRouter chat/completions: how to call it (and the three response shapes that trip people up)

If you point an OpenAI-style client at OpenRouter and ask for typesafe/jev-1.13, you get this back:

{"error":{"message":"typesafe/jev-1.13 is a decisions model and cannot be used with the chat/completions endpoint. Use the /api/alpha/decisions endpoint instead.","code":400}}
Enter fullscreen mode Exit fullscreen mode

That is the whole story in one line: Jev doesn't write text, so it doesn't live on the chat endpoint. It is a decision model (TypeSafe calls the class "System One"): you send one input and a set of typed questions, and it answers each with probabilities. This post shows the request that works, the response it returns, and the three fields that cause most of the bugs I've seen.

Recorded 2026-10-09 against OpenRouter. With a TypeSafe key you can skip the gateway: the same state and questions go to TypeSafe's own API, POST https://api.typesafe.ai/v1/systemone, with "model": "jev-latest".

The request that works

curl https://openrouter.ai/api/alpha/decisions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "typesafe/jev-1.13",
    "state": "I was charged twice for September and I am cancelling on Friday unless it is refunded.",
    "questions": {
      "team":    { "type": "choice", "instructions": "Which team should handle this ticket?",
                   "criteria": { "billing": "Payments, invoices, refunds", "technical": "Bugs and outages", "sales": "Plans and upgrades" } },
      "urgency": { "type": "score",  "instructions": "How urgent is this ticket?",
                   "criteria": ["Low", "Normal", "High", "Urgent"] },
      "churn":   { "type": "noul",   "instructions": "The customer threatens to cancel.",
                   "criteria": { "true": "Yes", "false": "No" } }
    }
  }'
Enter fullscreen mode Exit fullscreen mode
  • state is what you are asking about — a string, a JSON object, or an array.
  • questions is an object keyed by your own ids, not an array. All questions share the one state and run in parallel, so asking three in one call costs about the same as asking one.

What comes back

{
  "model": "typesafe/jev-1.13-20260917",
  "answers": {
    "team":    { "type": "choice", "choice": "billing",
                 "probabilities": { "billing": 1, "technical": 0, "sales": 0 }, "confidence": 1 },
    "urgency": { "type": "score", "score": 2.38,
                 "legend": { "0": "Low", "1": "Normal", "2": "High", "3": "Urgent" },
                 "probabilities": { "0": 0, "1": 0.02, "2": 0.58, "3": 0.4 }, "confidence": 0.58 },
    "churn":   { "type": "noul", "noul": 0.96 }
  },
  "usage": { "input_tokens": 421, "output_tokens": 70, "cost": 0.000017682 },
  "provider": "TypeSafe"
}
Enter fullscreen mode Exit fullscreen mode

$0.0000177 for three decisions: 421 input tokens at $0.042 per million. Output isn't billed.

Three shapes that cause bugs

1. score is not 0–1. It is the probability-weighted mean of the rubric indices, so it runs from 0 to n−1. 2.38 on a four-level rubric means "between High and Urgent, closer to High". Round it for the band; the fraction tells you which way the probability mass leans.

2. Score probabilities are keyed by index, as strings — "0", "1" — never by your label. Map back through legend.

3. A noul has no confidence field. Not null: absent. The distance of noul from 0.5 is all you have, so answers.churn.confidence ?? 0 silently treats every yes/no as zero-confidence.

Two more things worth knowing

  • Catalogue searches can miss it. Jev's output modality is "decisions", not text, so a gateway search filtered to chat models can return nothing while the model is live. Query GET https://openrouter.ai/api/v1/models?output_modalities=decisions to list every decision model.
  • It is not the only decision model any more. The same endpoint serves twelve others (OpenAI's GPT-6 Luna Decisions, Cloudflare's Clef, Perplexity's Decider, Liquid's d1 and more) with the same request shape. I ran all thirteen on the same 633 items; the raw data is in jev-measured and the write-up is at jev-agent.com/decisions-api.

I run jev-agent.com (Jagent), an independent Jev site, not affiliated with TypeSafe or OpenRouter.

Top comments (0)