TypeSafe's Jev made a simple idea popular: a model that doesn't write text, only makes decisions. You send some state and typed questions, and you get back choices, scores and probabilities your code can act on. Jev itself is closed, and you pay per token.
When it launched, I saw everyone hyping it up and thought: I can build an open one. So I did. WaterSheep is an open-source alternative to Jev. It answers the same kinds of questions about any text:
- noul: yes/no, with the probability of yes
- choice: one of your options
- score: a rating scale, with the expected level
- multi: every option that applies (Jev doesn't have this one)
Every answer includes a probability for every option.
Already using Jev? Keep your code
Run WaterSheep as a local server:
pip install git+https://github.com/SamratDuttaOfficial/WaterSheep
watersheep --model samratduttaofficial/WaterSheep --serve
It answers POST /v1/systemone with the same request and response shape as Jev, so TypeSafe's own Python SDK works against it. Point the SDK at it and run your code as usual:
export TYPESAFE_BASE_URL=http://127.0.0.1:8766
python triage.py
The script is plain Jev code:
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
client = TypeSafeClient() # or TypeSafeClient(base_url="http://127.0.0.1:8766")
r = client.system_one(
state={"ticket": "I was charged twice for order #4411. Refund the duplicate today."},
questions={
"refund": Noul(instructions="Does the customer ask for a refund?"),
"team": Choice(
instructions="Which team should handle this?",
criteria={"billing": None, "shipping": None, "technical": None},
),
"urgency": Score(
instructions="How urgent is this?",
criteria=["low", "medium", "high"],
),
},
)
a = r.answers
print("refund:", a["refund"].noul) # 0.9451
print("team:", a["team"].choice, a["team"].probabilities) # billing {'billing': 0.8817, ...}
print("urgency:", a["urgency"].score, a["urgency"].legend) # 1.4805 {0: 'low', 1: 'medium', 2: 'high'}
Any API key value works locally. I tested this with typesafe-sdk 0.7.2; Jev's JavaScript SDK is untested so far.
No SDK? It's plain HTTP
curl http://127.0.0.1:8766/v1/systemone -d '{
"state": "My card was charged twice and the box arrived crushed.",
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "payments, refunds", "shipping": "delivery, damage"}},
"issues": {"type": "multi", "instructions": "Which problems are reported?",
"criteria": ["double charge", "damaged item", "late delivery"]}
}
}'
The multi question returns both problems, each with its own probability. A single choice has to pick one.
Python without a server
from transformers import pipeline
ws = pipeline(model="samratduttaofficial/WaterSheep", trust_remote_code=True)
ws("I was charged twice.", question="Which team should handle this?", options=["billing", "shipping", "support"])
It also runs as ONNX, and the live demo lets you try it without installing anything.
How good is it?
On the in-distribution test split it gets 77.8% accuracy; on held-out datasets it never saw, 61.2%. The probabilities are calibrated (expected calibration error 0.026 and 0.043), so you can set thresholds on them: act automatically when the model is confident, and send the rest to a person. Per-benchmark results, including the weak ones, are in the README, and the details are in the paper.
Limits: English only, long inputs are truncated, rating questions are the weakest type, and it isn't meant for high-stakes decisions on its own.
Open all the way down
Code, weights and the training pipeline are Apache 2.0, and you can retrain it on your own data with one script. WaterSheep is independent and not affiliated with TypeSafe.
- GitHub: https://github.com/SamratDuttaOfficial/WaterSheep
- Hugging Face: https://huggingface.co/samratduttaofficial/WaterSheep
I'd love feedback, especially examples where it gets things wrong.

Top comments (1)
tr.ee/dev-to